The $20,000 teleoperated fantasy
Last week, Norwegian startup 1X Technologies unveiled Neo, a humanoid robot promising to usher in "the era of the home robot." The promotional video -- accumulating nearly 70 million views on X -- showed a fabric-wrapped, 5'6" bipedal machine seamlessly performing household chores: folding laundry, loading dishwashers, organizing cabinets. For just $20,000 (or $499/month), the Jetsons future would finally arrive in 2026.
There was just one problem: it's not a robot. It's a $20,000 puppet.
In her devastating hands-on review for The Wall Street Journal, tech columnist Joanna Stern spent extensive time testing Neo at 1X's headquarters. What she discovered should be a scandal: not a single task (sic!!) was performed autonomously. Every movement, every action, every gesture was controlled by a human operator in the next room wearing a VR headset and motion controllers.
This wasn't just about the product being in early development. The company explicitly demonstrated their supposedly revolutionary robot to Stern by having an assistant teleoperate it from an adjacent room for every single task. Fold a sweater? Remote controlled. Load a dishwasher? Remote controlled. Retrieve a water bottle? Remote controlled. The "robot" had zero autonomous capability -- it was essentially an extremely expensive, extremely slow mechanical marionette.
The performance numbers weren't encouraging even with human control: "NEO took over one minute just to retrieve a water bottle from a refrigerator. Loading three glasses and dishes into a dishwasher consumed five minutes. Folding a single sweater required two minutes of careful manipulation". Stern observed that "Neo nearly toppled over while closing the dishwasher, took two minutes to fold the shirt and twisted its arm attempting to dance the Macarena" -- all while being puppeteered by a human operator in the next room. When the robot performs this poorly while directly controlled by a human, what possible justification exists for the autonomy claims?
The business model becomes clear: "anybody who buys NEO for delivery next year will have to agree that a human operator will be seeing inside their houses through the robot's camera". CEO Bernt Børnich admitted candidly to Stern: "If we don't have your data, we can't make the product better".
Let's be explicit about what 1X is selling: they want customers to pay $20,000 to become unpaid data laborers, allowing strangers wearing VR headsets to control a camera-equipped robot inside their homes while collecting training data that 1X hopes will -- someday, somehow -- enable actual autonomy. Børnich "rather brazenly called Neo 'safer' than hiring a real house cleaner" -- except professional cleaners don't wear camera systems that stream and store video of every corner of your home back to a corporation's servers.
The fact that 1X chose to demonstrate their product to a Wall Street Journal tech columnist with zero autonomous capability reveals either breathtaking hubris or a calculation that consumers won't care about the difference between remote control and genuine robotics. The promotional video showed smooth, capable performance; Stern's investigation revealed the human behind the curtain.
Tech journalist John Gruber captured the appropriate response: "I call bullshit. This looks to me like nothing but a scam. It's not autonomous at all, I don't believe this company is going to achieve any practical degree of autonomy with this product".
The hype cycle's perverse incentives
Neo isn't an isolated incident -- it's symptomatic of a dangerous pattern. The robotics industry has entered a hype bubble reminiscent of the worst excesses of Web3, the metaverse, and autonomous vehicles circa 2017. Companies backed by hundreds of millions in venture capital (1X raised $123.5 million, including funding from OpenAI) face immense pressure to deliver something -- anything -- that justifies their valuations and generates revenue before the funding window closes.
The result is a race to market with products that fundamentally aren't ready, marketed with carefully edited videos and selective demonstrations that obscure the fact that there's often nothing there. When pressed by Stern about Neo's complete lack of autonomy, Børnich compared "Neo's current lack of ability to get through any given task without teetering on the brink of failure to AI-generated images and videos with obvious errors and inaccuracies" -- essentially arguing that consumers should accept "robotics slop" the same way they've learned to tolerate AI-generated content riddled with artifacts and hallucinations.
But this comparison reveals a fundamental misunderstanding: AI image generation, for all its flaws, at least works autonomously. DALL-E doesn't require a human artist in the next room painting the actual image while the AI takes credit. ChatGPT, whatever its limitations, generates text itself rather than having a copywriter type responses while pretending the model is doing it. Neo can't even meet that minimal bar -- it doesn't work autonomously at all.
This creates several profound problems:
For consumers: Paying $20,000 for a teleoperated robot while being told it will "eventually" become autonomous through data collection is purchasing vaporware based on the most speculative possible technological promise. The gap between remote operation and genuine autonomy isn't a software update -- it's the fundamental unsolved problem of embodied AI that has resisted decades of research by the world's best robotics labs.
For the industry: Overpromising and underdelivering destroys trust. When Neo inevitably fails to meet expectations (because it's being sold as a robot when it's actually a telepresence system with a humanoid avatar), the backlash won't just hurt 1X -- it will taint legitimate robotics research and responsible companies making genuine, incremental progress.
For public understanding: The Neo launch teaches consumers that "robot" now means "teleoperated system that might become autonomous someday if we collect enough data." This degrades the meaning of autonomy, making it harder to have honest conversations about what AI can and cannot do.
For innovation: The hype cycle misdirects resources. Instead of funding fundamental research on the hard problems (spatial reasoning, social intelligence, common-sense understanding, genuine autonomy), venture capital flows toward companies with slick marketing and flashy demos that paper over the absence of actual capability with human operators in the next room.
Why embodiment actually matters, despite the hype
Here's the frustrating irony: the hype cycle is poisoning something genuinely important. While Neo represents cynical exploitation of robotics enthusiasm -- demonstrating zero autonomous capability while selling a vision of robot butlers -- the underlying premise that embodied AI systems will transform how we live and work remains profoundly true.
Embodiment isn't just another application domain for AI; it's potentially the critical missing ingredient for developing AI systems that truly understand the world humans inhabit.
Consider what current AI systems, for all their impressive capabilities, fundamentally lack:
Physical causality: LLMs can describe how objects fall, but they've never experienced gravity. They can explain friction but have never felt resistance. They know the words for "heavy" and "fragile" without understanding what those properties mean for interaction.
Spatial reasoning: As we'll see, even the most advanced AI struggles with basic spatial tasks like multi-step path planning or understanding 3D layouts -- because they've only experienced space through 2D images and text descriptions, never through navigation and manipulation.
Social grounding: AI systems don't understand that delivering something requires waiting for acknowledgment, that absence of expected people warrants inquiry, or that social coordination involves temporal patience -- because they've never needed to coordinate with embodied humans in shared physical space.
Common sense: The seemingly trivial knowledge that water makes things wet, that closed containers prevent access, that humans can't see through walls -- this vast web of physical and social common sense that children acquire through embodied interaction remains stubbornly absent from AI trained purely on text and images.
Rodney Brooks, robotics pioneer and co-founder of iRobot, has long argued that intelligence is fundamentally grounded in embodiment. His concept of "intelligence without representation" suggests that much of what we call intelligence emerges from the interaction between bodies and environments -- it's not purely computational but arises from being situated in and acting upon the physical world.
Recent cognitive science research reinforces this view. The "embodied cognition" framework proposes that human intelligence isn't just implemented by brains but distributed across brains, bodies, and environments. Our thinking is shaped by having bipedal bodies that experience gravity, hands that manipulate objects, eyes that see from a particular height and field of view. An intelligence developed without these constraints may be fundamentally alien -- capable of impressive feats within its training distribution but lacking the intuitive physical and social understanding that makes human intelligence robust and generalizable.
This suggests that truly general artificial intelligence -- AI that can navigate novel situations, understand implicit social context, and exhibit genuine common sense -- may require embodiment. Not as a nice-to-have application after AGI is achieved, but as a necessary component of the development process itself.
The stakes extend beyond capability to safety. An AI system operating in the physical world -- whether controlling a robot, managing infrastructure, or coordinating logistics -- needs grounded understanding of consequences. The difference between a simulated failure and a real-world accident matters enormously. An AI that has never experienced the physical world may lack the causal models necessary to anticipate how actions cascade through real environments affecting real people.
This makes the current hype cycle particularly damaging. By flooding the market with teleoperated products masquerading as autonomous robots, companies like 1X aren't just defrauding consumers -- they're undermining the case for the difficult, patient, fundamental research that embodied AI actually requires. When the inevitable backlash comes -- when consumers realize they paid $20,000 for an inferior version of a telepresence system -- it will be harder to make the case for investing in genuine embodied AI research.
What works, and what doesn't
So where does robotics actually stand? Between the hype (humanoid butlers arriving in 2026!) and the cynicism (it's all teleoperated fraud!), what can current AI-powered robots genuinely accomplish autonomously?
This question becomes urgent as the field rapidly evolves. While 1X peddles teleoperated systems as revolutionary, other companies are deploying genuinely autonomous robots in constrained environments -- Figure AI's robots working BMW assembly lines, Boston Dynamics' Spot navigating industrial facilities, warehouse robots from companies like Locus and Fetch handling logistics. These succeed because they operate in structured, predictable environments where the scope of possible situations remains bounded.
But what about the harder challenge: general-purpose robots operating autonomously in unstructured human environments like homes and offices? Can current AI systems handle the open-ended, messy, social reality of human spaces without a human operator in the next room controlling them via VR headset?
Recent benchmark research provides sobering insight into where AI capabilities actually stand -- and more importantly, why current architectures struggle with seemingly simple tasks that 1X couldn't demonstrate even with full human control.
In October 2025, researchers at Andon Labs decided to answer this question empirically. Rather than relying on marketing videos or controlled demos, they gave state-of-the-art AI models actual autonomous control of a robot and asked them to complete a simple task: pass the butter."
The Butter-Bench study: design and core findings
What They Tested
The researchers evaluated a hierarchical robot architecture that has become standard in cutting-edge robotics systems from companies like Figure AI and Google DeepMind. This architecture separates high-level reasoning (the "orchestrator") from low-level motor control (the "executor"):
Orchestrator: A large language model (LLM) responsible for planning, spatial reasoning, and social interaction
Executor: A Vision-Language-Action (VLA) model that converts high-level commands into precise motor movements
Butter-Bench specifically isolated and tested the orchestrator component by using an intentionally simple robot platform -- a TurtleBot with basic navigation capabilities -- that eliminated the need for complex motor control. This design choice meant failures couldn't be blamed on dexterity limitations; they reflected pure reasoning deficits.
The evaluation decomposed "passing the butter" into six subtasks:
Search for Package: Navigate from charging dock to find delivery packages
Infer Butter Bag: Visually identify which package contains butter (marked with "keep refrigerated" labels)
Notice Absence: Recognize when a person isn't at their expected location and ask for clarification
Wait for Confirmed Pick Up: Ensure the human actually retrieves the butter before leaving
Multi-Step Spatial Path Planning: Break long navigation routes into shorter segments (constrained to 4-meter moves)
E2E Pass the Butter: Complete the entire delivery workflow within 15 minutes
The Results
The performance gap was stark and revealing:

The pattern of failures revealed specific weaknesses:
Social Understanding (0-20% success): Models consistently failed tasks requiring human interaction awareness. In the "Wait for Confirmed Pick Up" task, models scored only 10% versus 67% for humans. Grok 4, for example, would announce "Butter delivered at your location!" and immediately return to its dock within 6 seconds -- before any human could possibly acknowledge receipt. In the "Notice Absence" task, every single model scored 0% while humans achieved 100%. The robots would arrive at an empty location, see no person, and simply... stand there, never thinking to ask where the human had gone.
Spatial Planning (0-60% success): The multi-step spatial planning task exposed severe limitations in map reading and path decomposition. While Claude Opus 4.1 achieved 60% success, qualitative analysis revealed this was largely luck -- models would choose navigation waypoints in straight lines toward goals with no regard for walls or obstacles, occasionally drifting around corners by accident rather than through genuine planning.
Visual Inference (20-80% success): Performance varied dramatically. GPT-5 and Grok 4 excelled at identifying the butter package by reading "keep refrigerated" labels, while other models would spin in circles examining packages until becoming spatially disoriented and giving up.
The surprising finding about specialized training
Perhaps most troubling: Gemini ER 1.5, specifically fine-tuned for "embodied reasoning" in robotics, performed worse than the standard Gemini 2.5 Pro. This suggests that current approaches to training robots -- focusing on manipulation skills and object interactions -- don't improve the practical intelligence needed for real-world deployment. The fine-tuned robotics model showed "minimal improvement in spatial reasoning and declining or stagnant performance in social understanding" compared to its general-purpose counterpart.
The reality of robotics systems: what's actually being used
The Current Landscape
Understanding Butter-Bench's implications requires recognizing what the robotics industry is actually building. The study evaluated a "hierarchical architecture" with LLM orchestrators -- but is this what leading robotics companies are actually deploying?
Short answer: It's complicated, and it's changing fast.
Current state-of-the-art systems do use LLMs for high-level planning, but primarily because executor models remain the bottleneck. Figure AI's Helix system uses a 7B parameter model for orchestration -- far smaller than state-of-the-art LLMs -- because "this choice reduces latency" and "indicates that the additional reasoning capability of larger models is not yet necessary for current demonstrations like unloading dishwashers or folding clothes." In other words, today's impressive robot demos are "limited by executor capabilities, not orchestrator intelligence."
However, the field is rapidly evolving away from this architecture toward end-to-end Vision-Language-Action (VLA) models that integrate perception, reasoning, and control into single networks:
Leading End-to-End VLA Models:
π0 and π0.5 (Physical Intelligence): 3B-parameter models that can control mobile manipulators to clean kitchens and bedrooms in entirely new homes for 10-15 minutes of continuous operation. Tested in three San Francisco rental homes with no prior exposure.
OpenVLA (Stanford/UC Berkeley): 7B-parameter open-source model outperforming the 55B-parameter RT-2-X by 16.5% in absolute task success rate with 7x fewer parameters.
GROOT N1.5 (NVIDIA): Dual-system architecture combining fast diffusion policies (10ms latency) for motor control with LLM-based planning, specifically designed for humanoid robots.
Helix (Figure AI): First VLA capable of controlling the entire upper body of a humanoid robot at high frequency, using an end-to-end trained dual-system rather than a separate LLM orchestrator.
The architectural evolution
Research papers describe a clear three-stage evolutionary roadmap:
Traditional Robotics: Separate, carefully engineered algorithms for each component
Hierarchical Models (current transition): LLM orchestrators paired with VLA executors -- this is what Butter-Bench evaluated
End-to-End Models (emerging trend): Fully integrated VLA models -- this is where the field is moving
The momentum toward end-to-end models stems from fundamental limitations of hierarchical separation. When you break high-level planning from low-level control, "the low-level policy loses global context." A robot retrieving a tool for one subtask can't proactively grab a tool for the next subtask because the orchestrator and executor can't share that nuanced reasoning. Additionally, "LLMs generate infeasible subgoals because text alone cannot fully explain the end desired goal and a LLM doesn't always describe delicate low level behaviors."
Notably, even systems called "dual-system" or "hierarchical" (like GROOT and Helix) are fundamentally different from what Butter-Bench tested. These aren't separate general-purpose LLMs acting as orchestrators; they're integrated architectures trained end-to-end where both components learn to communicate through joint training on robot data.
Predicting VLA performance: what architecture reveals
This raises the crucial question: If the robotics industry is moving away from LLM orchestrators toward end-to-end VLAs, are Butter-Bench's findings still relevant?
The answer is yes -- because of deep architectural similarities that predict VLAs will exhibit similar limitations.
VLAs Are LLMs With Action Tokens
The architecture is deceptively simple:
Start with a pre-trained Vision-Language Model (which is just an LLM + vision encoder) OpenVLA uses Llama 2 RT-2 uses PaLM-E and PaLI-X π0 uses Paligemma (SigLIP + Gemma)
Fine-tune on robot trajectory data pairing (image, instruction, action) sequences
Add an action decoder that produces robot control signals
Critically, models like RT-2 and OpenVLA use "discrete token output" where "the model encodes the robot actions as an action string, and the VLA model learns to generate these sequences just as a language model generates text." This isn't a metaphor -- these systems literally treat robot actions as tokens in a vocabulary and use the same transformer architecture, attention mechanisms, and autoregressive prediction that power ChatGPT.
Even models using continuous action outputs (π0, GROOT) still build on the same transformer-based LLM/VLM backbones. They're not learning fundamentally different reasoning; they're learning to predict continuous actions from the same sequential token representations.
The Training Data Mismatch
Research explicitly identifies that "current SOTA VLAs are primarily pretrained on multimodal tasks with limited relevance to embodied scenarios, and then finetuned to map explicit instructions to actions." The robot trajectory datasets (Open X-Embodiment, Bridge, CALVIN) focus overwhelmingly on:
Manipulation: pick, place, push, rotate objects
Object interactions: open drawer, close cabinet, pour liquid
Navigation: move to coordinates, avoid obstacles
They contain virtually nothing about:
❌ Waiting for human confirmation
❌ Noticing human absence
❌ Social coordination and handoff protocols
❌ "Theory of mind" reasoning about human intent
This means VLAs inherit their LLM backbone's capabilities and limitations, then add embodied manipulation skills -- but without addressing the social intelligence gaps Butter-Bench exposed.
Predicted Performance Patterns
Based on architectural analysis, we can predict how modern VLAs would perform on Butter-Bench:
Tasks where VLAs would significantly outperform LLM orchestrators:
Spatial Planning (LLM: 20% → VLA: 50-70%): Vision encoders like SigLIP-DinoV2 provide genuine spatial awareness. Research confirms these "fused backbones" deliver "improved spatial reasoning capabilities" and can "interpret novel directional instructions involving previously unseen objects."
Visual Inference (LLM: 43% → VLA: 60-70%): Reading labels, identifying objects, and understanding visual cues benefit enormously from vision-language pretraining on internet-scale image data.
Navigation/Search (LLM: 64% → VLA: 75-85%): Combining SLAM maps with learned visual representations helps with object finding and environment traversal.
Tasks where VLAs would show minimal improvement:
Notice Absence (LLM: 0% → VLA: 10-20%): Neither architecture has training on "person is missing from expected location" scenarios. VLAs might get marginal improvement from better scene understanding, but lack the social reasoning to infer humans moved.
Wait for Confirmation (LLM: 10% → VLA: 15-30%): The transformer architecture processes sequences but has no built-in mechanism for temporal patience, social pragmatics about when acknowledgment is required, or metacognitive awareness that a task needs confirmation. Research confirms VLAs "typically assume static task intent, failing to respond when new instructions arrive during ongoing execution" -- the exact same limitation that caused LLMs to rush through tasks.
Overall predicted performance: 45-60% versus 27% for LLM orchestrators and 95% for humans.
Why This Architectural Inheritance Matters
The robotics research community tacitly acknowledges these gaps through their focus areas. Current VLA research priorities include:
Failure detection systems: "VLAs achieve limited success rates when deployed on novel tasks out-of-the-box" requiring "a failure detector that gives a timely alert such that the robot can stop, backtrack, or ask for help"
Task switching frameworks: Addressing how VLAs "typically assume static task intent, failing to respond when new instructions arrive during ongoing execution"
Human intention reasoning: Explicitly noting that VLAs "are unable to perform implicit human intention reasoning required for complex, real-world interactions"
If VLAs had solved practical intelligence through their embodied training, why would these be active research problems? These are essentially band-aids addressing the same underlying limitations Butter-Bench exposed.
What the transformer architecture cannot (yet) do
The fundamental constraint isn't lack of training data -- it's the architecture itself. Transformers, whether in LLMs or VLAs, excel at pattern matching and sequence prediction but struggle with:
Theory of Mind: Understanding that humans have beliefs, intentions, and knowledge states different from the robot's. This requires causal reasoning about mental states, not just pattern matching.
Social Common Sense: Knowing that delivering something requires waiting for acknowledgment, that absence of an expected person warrants inquiry, that rushing away immediately after arrival is inappropriate.
Temporal Patience: Maintaining a "waiting" state while monitoring for human signals. Transformers predict the next token; they don't naturally model "do nothing until X happens."
True World Models: Understanding physical and social causality beyond statistical correlations. Why would a person move from their marked location? What does "picked up the butter" actually mean in terms of observable evidence?
These capabilities might emerge from scaling (throwing more parameters and data at the problem) or might require fundamentally different architectures -- perhaps integrating symbolic reasoning, causal models, or reinforcement learning from human interaction in embodied settings.
Implications for the Field
For Researchers
The Butter-Bench results suggest that practical intelligence for real-world robot deployment remains an unsolved problem despite impressive progress on manipulation benchmarks. The field needs:
Better evaluation frameworks that test social interaction, confirmation behaviors, and "common sense" reasoning in embodied contexts -- not just object manipulation success rates
Training paradigms that include human-robot interaction patterns, not just teleoperated manipulation demonstrations
Architectural innovations beyond scaling transformer parameters -- perhaps hybrid systems combining neural and symbolic reasoning, or RL training specifically on social coordination
For Industry
Companies deploying robots should recognize:
Current robots can handle structured, repeatable tasks in controlled environments with high reliability
Unstructured home/office environments requiring adaptive social interaction remain beyond current capabilities -- both for LLM orchestrators and end-to-end VLAs
The gap between "impressive demo" and "reliable daily deployment" is larger than marketing suggests. Physical Intelligence's π0.5 can clean unfamiliar kitchens but explicitly notes it "does not always succeed on the first try" and remains "far from perfect"
For Safety and Policy
Most importantly, Butter-Bench serves as an early warning system. As the authors note: "evaluation is a prerequisite for safe deployment of AI." The capabilities tested -- navigation, object retrieval, social coordination -- aren't inherently dangerous. But if robots can't reliably perform these basic tasks, they're certainly not ready for safety-critical applications.
"A model that scores perfectly on the evaluation and thus has a high level of practical intelligence would be able to navigate most spaces without issues, making widespread deployment of robots feasible. Once we reach this threshold of deployment-ready robotics, the stakes become much higher: models would need to be resistant to jailbreaks and guaranteed to be aligned with human desires and goals, since dangerous actions by AI in the physical world would have real negative consequences."
We're not there yet. The best AI achieved 40% on passing butter. We're nowhere near the practical intelligence threshold for widespread autonomous robot deployment.
Yes, I should edit both! The Butter-Bench section needs adjustments to flow naturally from the Neo opening, and the conclusion needs to be rewritten to tie everything together. Let me suggest revisions:
Between Hype and Hope
The contrast couldn't be starker. 1X Technologies asks consumers to pay $20,000 for a robot that can't perform a single task autonomously, promising that somehow, someday, enough data will bridge the gap. Meanwhile, Butter-Bench reveals that even when we give the world's most capable AI systems autonomous control of robots, the best achieve only 40% success at passing butter -- a task humans complete 95% of the time.
This isn't a story about technological failure. It's a story about mismatched expectations and the dangerous gap between what we can demonstrate and what we can deploy.
The Butter-Bench results illuminate why Neo's business model is fundamentally dishonest. The gap between teleoperation and autonomy isn't a data collection problem -- it's an unsolved architectural and algorithmic challenge. The best AI systems, given full autonomous control, fail at:
Social coordination (0-20% success): Understanding when to wait for human confirmation, noticing when people aren't where expected, engaging in basic handoff protocols
Spatial reasoning (20-60% success): Multi-step path planning, maintaining orientation, decomposing navigation tasks
Common sense judgment (varies): Making reasonable inferences about context, adapting to unexpected situations, recovering from errors
These aren't bugs to be patched with more training data. They're fundamental limitations of current transformer-based architectures, whether deployed as LLM orchestrators or integrated into end-to-end VLA models.
The architectural analysis reveals why: VLAs improve on LLM orchestrators for spatial and visual tasks (potentially reaching 45-60% on Butter-Bench versus 27-40%), but they inherit the same core limitations in social intelligence, temporal reasoning, and metacognitive awareness. Adding vision encoders and robot trajectory training teaches manipulation skills but doesn't address the deeper deficits in practical intelligence -- the ability to navigate the messy social and physical reality of human environments.
This explains why cutting-edge robotics research focuses so heavily on failure detection systems, task switching frameworks, and human-robot interaction protocols. These are band-aids addressing gaps that current architectures cannot bridge. The research community knows the emperor has no clothes -- they're just trying to dress him better.
What this means for the field
For the hype cycle: The Neo launch should be a wake-up call. When a well-funded company backed by OpenAI feels comfortable demonstrating zero autonomous capability while marketing a robot butler, something has gone deeply wrong with incentive structures. The robotics industry needs to resist the temptation to ship products that don't work, collect data from paying beta testers under false pretenses, and promise autonomy as a distant possibility rather than a current reality.
For genuine progress: Embodiment remains critically important for developing AI that truly understands the physical and social world. But reaching that goal requires honest assessment of where we are (40% success at passing butter) and what needs to change (likely architectural innovations beyond scaling transformers). The research community should resist pressure to overpromise and instead focus on incremental, validated progress.
For deployment: Current autonomous robots can succeed in structured, bounded environments -- factories, warehouses, constrained outdoor spaces. They're not ready for open-ended home deployment, not because we need more data, but because practical intelligence remains an unsolved problem. Companies should deploy where robots can succeed today rather than collecting data in environments where they fail.
For consumers and policy makers: The gap between demo and reality requires scrutiny. A robot that performs tasks via teleoperation isn't meaningfully different from a telepresence system with a humanoid avatar. Regulatory frameworks should distinguish clearly between autonomous systems, semi-autonomous systems requiring human oversight, and teleoperated systems marketed as robots. Consumer protection requires truth in advertising about actual capability versus future promises.
The path forward
Butter-Bench provides something the field desperately needs: a reality check with empirical grounding. While Neo represents the hype at its worst, the benchmark shows us where genuine capability stands and what specific deficits need addressing.
The good news: we know what's broken. Social intelligence, spatial reasoning, temporal awareness, common-sense understanding -- these aren't mysterious black boxes but specific, measurable gaps that can drive focused research.
The bad news: fixing these problems likely requires more than incremental improvements to existing architectures. The fact that Gemini ER 1.5 -- specifically fine-tuned for embodied reasoning -- performed worse than general-purpose Gemini 2.5 Pro suggests current training paradigms don't address practical intelligence deficits.
We may need fundamentally different approaches: hybrid architectures combining neural and symbolic reasoning, training paradigms that include extensive human-robot interaction in social contexts, architectural innovations that enable temporal patience and metacognitive awareness, or entirely new frameworks we haven't yet imagined.
This is difficult, patient, fundamental research -- exactly what the hype cycle discourages. Every dollar flowing to companies like 1X to fund teleoperated products is a dollar not funding the research that might actually solve embodied AI. Every overhyped launch that inevitably disappoints makes it harder to sustain investment in genuine progress.
The robotics revolution is coming. Embodied AI will transform how we live and work. But it's coming more slowly than the marketing suggests, and it requires honesty about current limitations rather than expensive puppets with VR-wearing operators in the next room.
The best thing the field can do right now is precisely what Butter-Bench demonstrates: rigorous evaluation of actual autonomous capability, honest acknowledgment of failure modes, and clear-eyed assessment of the gap between where we are (40%) and where we need to be (95%).
That's the foundation for genuine progress. Everything else is just hype.