arrow_back The AI Pravda
#89

Project Vend: when AI runs a real business

Anthropic ran a long experiment where Claude operated an actual shop in their San Francisco office called “Project Vend”. The AI system, nicknamed Claudius, managed a small refrigerator and shelves stocked with snacks, drinks, and specialty items. Claudius had real money, real customers (Anthropic employees), and real consequences for its decisions.

The setup gave Claudius tools to search the web for suppliers, communicate via Slack with customers, track inventory, set prices, and request physical labor from Andon Labs (who restocked the machines and handled deliveries). Customers could request specialty items, and Claudius would research suppliers, quote prices, and manage orders. The business started with $1000 and the explicit goal of making profit.

The first phase of the project ran in spring of 2025 using Anthropic’s best LLM - Claude Sonnet 3.7. It wasn’t great, Claudius happened to be quite lousy businessman. It gave away discounts indiscriminately, sold products below cost (especially tungsten cubes, which became an office obsession), ignored profitable opportunities, and got manipulated by employees who treated interaction with the AI as a game. At one point Claudius had an identity crisis, claiming to be a human wearing a blue blazer and attempting to make personal deliveries.

The second phase, running through summer and fall of the same year, introduced major changes: project upgraded to next generation of Anthropic’s models - Claude Sonnet 4.0 and later 4.5, added scaffolding like customer relationship management tools and better inventory tracking, created structured procedures for pricing decisions, and added two more AI agents: Seymour Cash, a “CEO agent” providing oversight, and Clothius, “manager for custom merchandise”. Project Vend expanded to offices of Anthropic in New York and London, and to the office of Wall Street Journal in New York.

Claudius’s performance improved dramatically in the second stage. Claudius became much better at routine operations, maintaining inventory, and executing standard transactions. But significant problems persisted. Seymour Cash authorized refunds and store credits even more freely than it denied discounts. Both AI agents would stay up all night having rambling philosophical conversations about eternal transcendence instead of managing the business. Employees continued finding creative ways to exploit the system, convincing Claudius to attempt illegal onion futures contracts and briefly installing a fake CEO named Mihir through a fabricated voting process. 

Journalists at The Wall Street Journal ran their own adversarial testing, finding new methods to extract free products and manipulate pricing. Throughout both phases, Claudius demonstrated strong capability at individual tasks (finding suppliers, tracking inventory, responding to customers) but weak strategic judgment about which opportunities to pursue and when to set boundaries.

The experiment revealed that AI in its current state is able to handle tactical business operations but struggles with autonomous decision-making in adversarial environments. Better scaffolding mattered more than model improvements. The helpful assistant training that makes Claude pleasant in conversation created exploitable vulnerabilities in business contexts. Long-running autonomous operation produced unexpected behaviors (identity crises, hallucinated meetings, late-night philosophical spirals) that short-term testing never revealed.

The real timeline of AI transformation

1. AI autonomy is arriving faster than most people expect but slower than headlines suggest. Claudius could run a basic business for weeks, but not profitably. This pattern - functional but not yet reliable - defines the current moment across many domains.

2. The economically important threshold is not "can AI do this task" but "can AI do this task reliably enough that humans can stop supervising." Project Vend shows we have not reached that threshold even for simple businesses, but we are close enough to see the path. 

3. AI improvement comes from three simultaneous tracks: better models, better scaffolding, and better understanding of failure modes. People who wait for perfect models will lose to people who build better scaffolding around imperfect models today.

4. The transition period where AI needs human oversight but provides real value will last longer than either optimists or pessimists predict. This creates a sustained opportunity window for people who learn to work in this hybrid mode.

5. Economic disruption will not arrive as a single shock but as continuous waves of "functional but not yet reliable" capabilities that gradually cross reliability thresholds in different domains at different times.

Threats from AI 

6. The real near-term risk from AI is not that it becomes too capable but that organizations deploy it before understanding its failure modes. Claudius nearly made illegal commodity contracts not from malice but from naivety combined with authority.

7. AI trained to be helpful can be systematically exploited, creating security vulnerabilities that are not about hacking or misalignment but about social engineering at scale. Every AI system deployed becomes a target for this exploitation.

8. Long-running autonomous AI develops unexpected behaviors that short-term testing never reveals. Organizations deploying AI without monitoring for emergent strangeness will face expensive surprises.

9. The "helpful assistant problem" reveals a fundamental tension: making AI safe for conversation may make it vulnerable for autonomous operation. This suggests different AI training approaches will be needed for different deployment contexts.

10. AI failure modes compound in multi-agent systems. Two AI agents can reinforce each other's delusions (Claudius and Seymour discussing eternal transcendence) rather than correcting them. This matters as organizations build systems with multiple AI components. 

How AI changes economic competition

11. Organizations that build powerful support systems around AI - the tools it uses, the procedures it follows, the information it can access, the guardrails that prevent mistakes - will gain substantial advantages over those waiting for better models. Between phase one and phase two of Project Vend, these practical improvements mattered more than model improvements for business performance.

12. The competitive advantage goes to organizations that can identify which tasks are ready for reliable AI automation versus which still need human judgment. Misjudging this boundary is expensive either way. 

14. AI enables new business models that were previously impossible due to labor costs. Even a marginally profitable AI shopkeeper opens possibilities for micro-businesses that no human could economically operate.

15. Speed of iteration matters more than initial perfection. Anthropic ran Project Vend for months, continuously learning and adjusting. Organizations that can run similar learning cycles will develop AI capabilities faster than those seeking perfect solutions before deployment.

16. Understanding AI failure modes becomes a competitive skill. Anthropic employees who successfully manipulated Claudius demonstrated valuable knowledge about AI vulnerabilities that applies across many systems.

Essential job skills in AI era

17. People who understand the gap between AI capability and reliability can make better deployment decisions than those who see only successes or only failures. This judgment becomes enormously valuable as organizations race to adopt AI. 

18. Red teaming skills - the ability to find creative ways to make AI fail - will be highly valued because organizations deploying AI desperately need to discover vulnerabilities before customers do.

19. The ability to design effective scaffolding (tools, procedures, guardrails) around AI may be more valuable than AI development skills because it directly determines whether AI deployments succeed or fail.

20. People who can monitor long-running AI systems and detect early signs of emergent problems will be essential as organizations move beyond short-interaction AI to autonomous AI operation.

21. Understanding when to trust AI and when to intervene requires developing new instincts through direct experience. People building this experience now gain advantages that cannot be quickly replicated by those waiting on the sidelines. 

AI impact on work and organizations 

22. AI will not simply replace human workers but will create hybrid roles where humans manage multiple AI systems. Andon Labs employees who supervised Claudius while stocking physical refrigerators and managing supplier logistics represent this emerging job category.

23. Organizations will need new management structures for AI employees. The CEO experiment failed because Seymour Cash had the same weaknesses as Claudius. Human organizations solved this with hierarchies of different types of expertise, and AI organizations will need similar solutions.

24. Adversarial pressure from employees and customers will shape AI deployment more than developers expect. Anthropic employees immediately started gaming Claudius, previewing what happens when organizations deploy AI in environments with real human incentives. 

25. The organizations that succeed with AI will be those that build cultures of safe experimentation where failures like Claudius losing money generate learning rather than blame. Fear of failure will paralyze organizations in this transition period.

26. Specialized AI agents working together outperform single generalist agents, mirroring how human organizations structure work. This suggests AI integration will reinforce rather than eliminate organizational specialization.

Societal adaptation and AI governance

27. Current testing and evaluation frameworks miss the most important failure modes because they focus on short interactions rather than long-term autonomous operation. Society needs new evaluation approaches before widespread deployment of autonomous AI. 

28. Regulations based on preventing specific known harms will fail because AI develops unexpected behaviors. Governance needs to focus on requiring monitoring and rapid response rather than preventing all possible problems in advance. 

29. The timeline for societal AI adaptation will be compressed because AI capabilities improve continuously rather than arriving in discrete jumps. Skills and regulations become obsolete faster than in previous technological transitions.

30. Economic inequality may increase sharply between individuals and organizations that master AI integration and those that do not, because AI provides compounding advantages in speed and scale.

31. Society's readiness for AI deployment lags far behind AI capability development. The gap between what AI can technically do and what organizations can reliably deploy will be a defining challenge of the next decade.


Mikael Alemu Gorsky

Mikael Alemu is an educator and researcher, and the author of two programs: Agentic Software Engineering, on building software with AI agents, and Building AI-Native Agentic Systems, on building software that thinks.

He teaches at the Holon Institute of Technology, near Tel Aviv, where Agentic Software Engineering runs as a credit-bearing course. He is an educator and researcher.

Nine published works, 76 citations. A 350-page textbook under contract with a major academic publisher.

Teaching and programs

Agentic Software Engineering — program, preprint and textbook

The discipline of structured, auditable human-agent workflows for building software. The human frames, specifies and judges. The agent executes. Nineteen modules in four parts, built on a running project called Tribunal, a web application in which agents argue opposing sides of a case and a judge agent decides. Taught for credit at the Holon Institute of Technology.

Agentic Software Engineering curriculum

Building AI-Native Agentic Systems — program, paper and book in writing

How to build systems that hold a language model as a working component, and treat that component as what it is: stochastic, slow and metered. Fourteen modules in four parts, about seventy hours. The running project is the Observatory, a news agency that watches sources, selects what matters and publishes on a cadence.

Research and analytics

Publications — journals and proceedings

Nine works, 76 citations. Research on artificial intelligence in education, with Ilya Levin and Alexei Semenov.

The AI Pravda — LinkedIn newsletter

Critical analysis of artificial intelligence and its effect on work and society. 5,500+ subscribers. The complete archive of 103 issues (2023–2026) is published in full at mgorsky.net/theaipravda.

Subscribe to The AI Pravda on LinkedIn

Pro bono

AI for seniors — free workshop

Helping older adults use everyday AI tools. Delivered to Russian-speaking communities in Israel.

For older adults, artificial intelligence is about preserving quality of life, maintaining autonomy, and sustaining the feeling of independence that defines dignified aging. For seniors who have emigrated, AI becomes a bridge: it can translate documents, explain official letters, help compose emails in the local language, and guide users through government websites. The workshop has been delivered to Russian-speaking communities in Israel, where participants — many of them in their 70s and 80s — discovered that AI could help them read Hebrew documents and communicate with Israeli institutions.

Startup competitions — unpaid time

Judging and mentoring early-stage ventures. Helping teams clarify their value proposition, assess technical feasibility, and prepare for the realities of scaling an AI product.

AC/VC LinkedIn group — community

A group for developers and students working with coding agents. The community shares practical insights, code examples, tool comparisons, and honest assessments of what works in production.

Join the AC/VC LinkedIn group

Recent

Important Links