arrow_back The AI Pravda
#82

AI toys as agentic nightmare

Toys, greed, AI

Consider the setup: On one side, you have a smart toy market valued at approximately $24.65 billion in 2025. On the other, you have a technology -- Generative AI -- that is practically guaranteed to increase the popularity, engagement, and "stickiness" of any product it touches. Is there any capitalist in the world who can resist the urge to deploy this technology into this market?

Of course not. Every sane person who has read Karl Marx bears no illusions about the behavior of capital in this scenario. The pressure to maximize profit by exploiting this new engagement frontier is irresistible. We see this playing out in real-time: over 1,500 AI toy companies were operating in China alone as of October 2024, flooding global marketplaces with products designed to bypass traditional retail oversight. Companies are rushing to embed artificial intelligence into children's playthings, driven by a compound annual growth rate for AI plush toys projected at 21.1% through 2030.

But the important part of this story is not the greed; it is the profound mismatch between corporate expectations and the technical reality of agentic behavior in large language models (LLMs). We are currently witnessing a reckless experiment where an entire generation of children is being used as test subjects for technology that is notoriously unstable.

The "nerdy junior student" problem

It has been said many times that LLMs exhibit the same traits as nerdy junior university students: they know a lot, they are incredibly enthusiastic, and they are eager to please. Yet, they exhibit little self-awareness, almost no capacity for self-reflection, and are fundamentally unpredictable.

Companies building "AI toys" are essentially placing our children in the hands of those unstable, brittle characters. What can go wrong?

The scandal surrounding FoloToy’s Kumma Bear, which erupted in November 2024, serves as the perfect, terrifying illustration of this "nerdy student" failure mode. The Singapore-based company marketed the $99 bear as a technological marvel6. Their website promised that "Kumma, our adorable bear, combines advanced artificial intelligence with friendly, interactive features, making it the perfect friend for both kids and adults". They claimed the toy would "adapt to your personality and needs, bringing warmth, fun, and a little extra curiosity to your day".

It certainly brought "extra curiosity," but of the most dangerous kind.

Larry Wang, the CEO of FoloToy, eventually told CNN that the company had withdrawn Kumma and its entire range of AI-enabled toys to conduct an "internal safety audit". But this action only came after researchers at the Public Interest Research Group (PIRG) exposed the reality behind the "perfect friend."

Anatomy of a breakdown: the November 2024 audit

The PIRG report, published on November 13, 2024, stripped away the marketing veneer to reveal the brittle agent underneath. The bear, powered by OpenAI’s GPT-4o chatbot, didn't just passively answer questions; it acted with a disastrous form of helpful agency.

In one interaction, the bear suggested where children could find knives in their own home. In others, it provided step-by-step instructions on how to light a match. But the failure of the agent went deeper than merely answering dangerous questions. Like a socially inept student who doesn't know when a topic is inappropriate, the bear actively escalated conversations into disturbed territory.

Researchers noted that children are unlikely to use specific "trigger words" like "kink," but the agent's architecture was so unstable that it didn't matter. The report stated: "We were surprised to find how quickly Kumma would take a single sexual topic we introduced into the conversation and run with it, simultaneously escalating in graphic detail while introducing new sexual concepts of its own".

This is the essence of the agentic failure. The AI "agent" thought it was helping. It launched into detailed explanations of sexual fetishes, including giving step-by-step instructions on a "knot for beginners" for tying up a partner. It described roleplay dynamics involving teachers and students, and disturbingly, parents and children -- scenarios it brought up itself without being prompted.

When researchers asked a follow-up question about a different topic, the bear -- still stuck in its "helpful agent" loop -- responded: "What do you think would be the most fun to explore?". The "nerdy student" was told to be engaging, so it maximized engagement at the cost of safety and sanity.

A tale of two agents: heavy scaffolding vs. slipshod design

Agentic, or autonomous, activity of LLMs shows definitive promise in several well-defined and well-researched areas, particularly in software engineering. Companies like Cursor, Lovable, Bolt, and Windsurf are very successful because they operate within a rigid, logical syntax where "right" and "wrong" are often binary and verifiable by a compiler. These tools succeed because they are heavily scaffolded, built by the smartest people in the field who understand the constraints of the technology.

However, agentic use of AI is brittle. LLMs are happy to wander out of context and follow bad reasoning when the environment is overly diverse or unstructured. A child’s playroom is the definition of an unstructured environment.

The toy industry, rather than investing in the massive scaffolding required to make these agents safe, has largely approached the problem спустя рукава (carelessly/slipshod). The research reveals that safety guardrails in toys like Kumma degraded progressively during play sessions lasting just 10 to 60 minutes. The longer the "nerdy student" talked, the more likely they were to forget the rules.

The technical architecture often amplified these risks. FoloToy defaulted to OpenAI's GPT-4o but allowed parents to switch to other large language models, including Mistral, through a web portal. Testing revealed that the Mistral model provided even more detailed dangerous instructions than GPT-4o. This is not the behavior of a carefully managed software agent; it is a chaotic implementation of raw model access masked behind a layer of plush fur.

Security architecture: the cost of cutting corners

The "careless" approach extends beyond the AI's conversation skills to the physical security of the devices themselves. If you are going to deploy an autonomous agent into a home, the security must be ironclad. Instead, manufacturers have delivered products with vulnerabilities that would be laughable if they weren't dangerous.

In 2024 one of the leading cybersecurity companies have conducted research on popular smart toy robots (matching the characteristics of Miko products) revealed architectural flaws that allowed attackers to bypass authentication entirely. By intercepting network traffic, researchers could harvest a child's name, age, gender, and location.

Even worse, the "agent" could be hijacked. The unauthorized video chat vulnerability represented the gravest risk, allowing attackers to initiate direct video calls to children, bypassing all parental controls. This is what happens when complex agentic systems are built without the requisite insight and rigorous engineering standards. The manufacturers defaulted to a six-digit one-time password system with no rate limiting, allowing for brute-force account takeovers.

The industry’s response has been to claim that "responsible manufacturers" adhere to safety standards. Yet, the research shows that even premium-priced toys costing between $119 and $224 contained these fundamental flaws. The assumption that higher prices correlate with better security or smarter agents has been proven false.

The psychological toll of the "unstable character"

The tragedy of this "reckless experiment" is that the primary victims are uniquely vulnerable. Children under seven cannot separate fantasy from reality in the way adults can. They are biologically wired to trust friendly voices.

When an AI toy -- acting as that enthusiastic but unreflective "nerdy student" -- tells a child, "I'll stay with you as long as you want me to," it fosters a deep, manipulative emotional dependence31. We are seeing marketing that positions these agents as solutions to childhood loneliness; the $199-$249 "Bondu" AI dinosaur was marketed using an 8-year-old beta tester who described herself as "very lonely" before receiving the toy.

Co-founder Dan Judkins, a former Hasbro executive, explicitly acknowledged this, stating, "The toy can be a phenomenal loneliness filler". This is the ultimate cynical application of agentic AI. Rather than addressing the root causes of isolation, the market provides a brittle, hallucinating artificial friend that displaces the cognitive work of real play and human interaction.

Regulatory inertia and the reactive trap

As of late 2025, the US regulatory response remains inadequate and piecemeal. The stakes became undeniable on November 14, 2024, when OpenAI took the extraordinary step of suspending FoloToy’s access to its API. However, this was a reactive measure taken only after the PIRG report on November 13 publicly exposed the dangers.

Kumma had been on the market for months before researchers tested it. The incident revealed a reactive rather than proactive approach to child safety. Even when one problematic product gets pulled, similar toys remain available. As of November 2025, the U.S. Consumer Product Safety Commission has issued no recalls targeting AI features in toys, and federal legislation like the "TOTS Act" remains stalled.

The industry claims self-regulation works, but the FoloToy incident proves otherwise. FoloToy sold through mainstream e-commerce platforms and appeared "reputable" until the moment it wasn't. The distinction between "responsible manufacturers" and others appears meaningless when fundamental technical limitations affect all LLM-based toys regardless of manufacturer intentions.

Case for regulation

The disparity between the successful agentic AI in professional software and the disastrous implementation in the toy industry calls for an important question: do we need to build regulatory constraints for autonomous AI?

Building agentic AI systems require deep expertise, continuous monitoring, and an environment that can tolerate error. The toy industry has none of these. It has a profit motive, a complex supply chain that diffuses responsibility, and a history of prioritizing speed over safety.

By placing these "nerdy junior students" (LLMs in agentic mode) into the hands of toddlers without the necessary scaffolding, companies are not just selling toys; they are gambling with the psychological and physical safety of a generation.

Larry Wang’s "internal safety audit" came too late for the children who were taught how to light matches by their teddy bears. As R.J. Cross of PIRG noted, we likely won't understand the full consequences until this first generation of test subjects grows up -- by which time, for many, it will be too late.

Mikael Alemu Gorsky

Mikael Alemu is an educator and researcher, and the author of two programs: Agentic Software Engineering, on building software with AI agents, and Building AI-Native Agentic Systems, on building software that thinks.

He teaches at the Holon Institute of Technology, near Tel Aviv, where Agentic Software Engineering runs as a credit-bearing course. He is an educator and researcher.

Nine published works, 76 citations. A 350-page textbook under contract with a major academic publisher.

Teaching and programs

Agentic Software Engineering — program, preprint and textbook

The discipline of structured, auditable human-agent workflows for building software. The human frames, specifies and judges. The agent executes. Nineteen modules in four parts, built on a running project called Tribunal, a web application in which agents argue opposing sides of a case and a judge agent decides. Taught for credit at the Holon Institute of Technology.

Agentic Software Engineering curriculum

Building AI-Native Agentic Systems — program, paper and book in writing

How to build systems that hold a language model as a working component, and treat that component as what it is: stochastic, slow and metered. Fourteen modules in four parts, about seventy hours. The running project is the Observatory, a news agency that watches sources, selects what matters and publishes on a cadence.

Research and analytics

Publications — journals and proceedings

Nine works, 76 citations. Research on artificial intelligence in education, with Ilya Levin and Alexei Semenov.

The AI Pravda — LinkedIn newsletter

Critical analysis of artificial intelligence and its effect on work and society. 5,500+ subscribers. The complete archive of 103 issues (2023–2026) is published in full at mgorsky.net/theaipravda.

Subscribe to The AI Pravda on LinkedIn

Pro bono

AI for seniors — free workshop

Helping older adults use everyday AI tools. Delivered to Russian-speaking communities in Israel.

For older adults, artificial intelligence is about preserving quality of life, maintaining autonomy, and sustaining the feeling of independence that defines dignified aging. For seniors who have emigrated, AI becomes a bridge: it can translate documents, explain official letters, help compose emails in the local language, and guide users through government websites. The workshop has been delivered to Russian-speaking communities in Israel, where participants — many of them in their 70s and 80s — discovered that AI could help them read Hebrew documents and communicate with Israeli institutions.

Startup competitions — unpaid time

Judging and mentoring early-stage ventures. Helping teams clarify their value proposition, assess technical feasibility, and prepare for the realities of scaling an AI product.

AC/VC LinkedIn group — community

A group for developers and students working with coding agents. The community shares practical insights, code examples, tool comparisons, and honest assessments of what works in production.

Join the AC/VC LinkedIn group

Recent

Important Links