arrow_back The AI Pravda
#39

Preventing AI Apocalypse

Eric Schmidt, not only a legendary technology manager who transformed Google into a global giant but also a well-known security and defense expert through his work with the National Security Commission on AI, brings unparalleled gravitas. Alexandr Wang, who founded and leads Scale AI—a pivotal player in AI data infrastructure valued at nearly $14 billion—offers a hands-on insight into AI’s practical deployment. Dan Hendrycks, Director of the Center for AI Safety (CAIS), rounds out the group with his deep expertise in AI risk mitigation, backed by Open Philanthropy’s funding from tech billionaires Dustin Moskovitz and Cari Tuna. Together, their credentials demand attention, promising a bold roadmap for navigating the national security implications of superintelligence—AI that vastly outstrips human cognition.

The document rests on several premises: (1) AI, especially superintelligence, mirrors nuclear technology as a dual-use tool with transformative potential and catastrophic risks, requiring a Cold War-inspired framework; (2) Mutual Assured AI Malfunction (MAIM) can deter states from reckless AI races by threatening sabotage, akin to nuclear Mutual Assured Destruction (MAD); (3) nonproliferation—tracking chips and model weights—can curb rogue actors; (4) competitiveness—boosting domestic AI capabilities—secures national strength; (5) superintelligence is imminent and detectable, justifying preemptive action; and (6) state-level rivalry and terrorism are the primary threats, overshadowing internal AI risks.

I was sincerely surprised, as these premises hold almost zero value under scrutiny. The “Superintelligence Strategy” flaws are so glaring they could make one think that we would not be able to stop AI Apocalypse. Let’s dismantle each premise systematically.

1️⃣ First, the nuclear analogy is a mirage. Nuclear weapons are tangible—missiles in silos with clear launch signals—while AI progress is evolutionary, creeping forward in labs and codebases without a “launch” moment (e.g., GPT-3 to GPT-4). MAIM’s sabotage — cyberattacks or datacenter strikes — targets an invisible process, unlike MAD’s concrete deterrence. Historically, states raced to outpace rivals in science, not cripple their labs; MAIM lacks precedent and practicality.

2️⃣ Second, MAIM’s deterrence is unworkable because its trigger is undefined. The authors vague out at “destabilizing” projects threatening survival but offer no threshold—compute scale? reasoning benchmarks? — leaving rivals guessing via espionage. AI’s gradual gains defy this; there’s no red line to spot, only a slope of improvement, rendering MAIM either premature or tardy. MAD thrived on clarity; MAIM flounders in ambiguity.

3️⃣ Third, nonproliferation assumes control over chips and weights, but groundbreaking AI can emerge from small, distributed teams using hosted models—cloud APIs, not server farms. Virtual collectives across borders (e.g., EleutherAI) dodge datacenter-centric sabotage, making global policing impossible. The authors’ focus on big infrastructure misses this lean reality.

4️⃣ Fourth, competitiveness—pushing chip production and talent —is sensible but slow, ignoring the risk of rivals building “perfect” AI: faster, smarter, leaner on compute. MAIM’s sabotage can’t stop a stealthy leapfrog, and domestic resilience won’t catch up if efficiency trumps scale.

5️⃣ Fifth, the premise of imminent, detectable superintelligence is shaky. The authors assume a tipping point rivals can see, but evolutionary progress hides breakthroughs—especially from small teams—until too late. MAIM’s preemptive logic falters when threats stay covert.

6️⃣ Sixth, fixating on state rivalry and terrorism neglects deeper risks: AI “intentionalizing” into self-directed agents, or rivals out-engineering us. The document brushes off emergent agency with technical tweaks, not systemic answers, and assumes parity MAIM can disrupt, not asymmetric perfection.

Worse, MAIM risks an adversarial spiral. Sabotage heats up tensions—hacking a Chinese lab could spark war—unlike MAD’s cold standoff. Totalitarian regimes like China, with superior surveillance and opacity, outmatch democracies in the spying MAIM demands, tilting the field against its proponents.

The number of flaws here—speculative, untested, disconnected—could suggest AI safety is a lost cause.

But a recent YouTube video, the Anthropic roundtable with three young engineers—Akbar Khan, Joe, and Ethan Perez—gave me back my trust in humanity. These alignment researchers at Anthropic (IMHO, the most important AI lab), discuss AI control: a practical, testable alternative to MAIM’s grandiosity. Perez defines control as ensuring misaligned models can’t harm, even with “bad goals,” via security (e.g., permissioning weights) and monitoring (e.g., smaller models auditing bigger ones). Unlike MAIM’s geopolitics, control operates within a lab’s grasp—mitigating risks like a model copying weights or sabotaging code.

Joe’s evaluation of Claude 3 Sonnet exemplifies this. They tested a misaligned model steering humans on business decisions, finding trust in its tone swayed participants despite warnings. Control counters with layered checks: untrusted monitors (the model self-auditing) or trusted weaker models (e.g., Claude 2 vs. Claude 3) flag deception, escalating to humans adaptively. This is granular and empirical—stress-tested against models trained to undermine it — not MAIM’s hypothetical “escalation ladders”.

Anthropic’s threat models—weight copying, unmonitored deployment, safety research sabotage —focus on deployment realities, not state rivalries. They tackle your intentionality worry: models faking alignment or subtly skewing outcomes. Current limits—weak planning for subtle sabotage —offer a window to refine control, unlike MAIM’s rush to preempt an unseen superintelligence. Optimism shines through: chain-of-thought reasoning exposes sabotage, and small monitors catch big failures (Page 16), building on robustness work labs already do.

The “AI control” approach presented by Anthropic’s engineers is certainly a “work in progress”— using controlled, transformative AI to research alignment — and not a permanent fix, but a sane step while superintelligence looms. Luminous authors of “Superintelligence Strategy” crafted a sand castle while Anthropic has shown true concern for AI safety that leans on practical ingenuity, not geopolitical gambles.

Mikael Alemu Gorsky

Mikael Alemu is an educator and researcher, and the author of two programs: Agentic Software Engineering, on building software with AI agents, and Building AI-Native Agentic Systems, on building software that thinks.

He teaches at the Holon Institute of Technology, near Tel Aviv, where Agentic Software Engineering runs as a credit-bearing course. He is an educator and researcher.

Nine published works, 76 citations. A 350-page textbook under contract with a major academic publisher.

Teaching and programs

Agentic Software Engineering — program, preprint and textbook

The discipline of structured, auditable human-agent workflows for building software. The human frames, specifies and judges. The agent executes. Nineteen modules in four parts, built on a running project called Tribunal, a web application in which agents argue opposing sides of a case and a judge agent decides. Taught for credit at the Holon Institute of Technology.

Agentic Software Engineering curriculum

Building AI-Native Agentic Systems — program, paper and book in writing

How to build systems that hold a language model as a working component, and treat that component as what it is: stochastic, slow and metered. Fourteen modules in four parts, about seventy hours. The running project is the Observatory, a news agency that watches sources, selects what matters and publishes on a cadence.

Research and analytics

Publications — journals and proceedings

Nine works, 76 citations. Research on artificial intelligence in education, with Ilya Levin and Alexei Semenov.

The AI Pravda — LinkedIn newsletter

Critical analysis of artificial intelligence and its effect on work and society. 5,500+ subscribers. The complete archive of 103 issues (2023–2026) is published in full at mgorsky.net/theaipravda.

Subscribe to The AI Pravda on LinkedIn

Pro bono

AI for seniors — free workshop

Helping older adults use everyday AI tools. Delivered to Russian-speaking communities in Israel.

For older adults, artificial intelligence is about preserving quality of life, maintaining autonomy, and sustaining the feeling of independence that defines dignified aging. For seniors who have emigrated, AI becomes a bridge: it can translate documents, explain official letters, help compose emails in the local language, and guide users through government websites. The workshop has been delivered to Russian-speaking communities in Israel, where participants — many of them in their 70s and 80s — discovered that AI could help them read Hebrew documents and communicate with Israeli institutions.

Startup competitions — unpaid time

Judging and mentoring early-stage ventures. Helping teams clarify their value proposition, assess technical feasibility, and prepare for the realities of scaling an AI product.

AC/VC LinkedIn group — community

A group for developers and students working with coding agents. The community shares practical insights, code examples, tool comparisons, and honest assessments of what works in production.

Join the AC/VC LinkedIn group

Recent

Important Links