arrow_back The AI Pravda
#62

"Dario Amodei: The Urgency of Interpretability" in 33 points

Here is the breakdown of Dario Amodei's essay into 33 distinct points.


Part 1: The Core Problem

1. The Unstoppable Bus

AI technology is advancing unstoppably, like a bus we can't stop. However, we have a great deal of control over how it is developed and used—we can "steer the bus."

2. The Key to Steering

My focus has shifted to a critical way we can steer: we must understand the inner workings of AI systems, a field called "interpretability." We need to achieve this before AI becomes overwhelmingly powerful.

3. A Surprising and Unprecedented Ignorance

Most people are shocked to learn that even the creators of advanced AI do not fully understand how they work. This lack of understanding is almost unheard of in the history of technology.

4. A Glimmer of Hope - The "MRI for AI"

For years, we've been trying to create the equivalent of an "MRI for AI" to see exactly what's happening inside a model. Recent breakthroughs suggest this goal is now achievable.

5. The Race Against Time

However, we are in a race. The power of AI systems is growing faster than our ability to understand them. We must move quickly if we want our understanding to mature in time to be useful.

Part 2: The Dangers of Ignorance

6. Grown, Not Built - The Opaque Nature of AI

Modern AI is fundamentally "opaque," or a black box. Unlike a normal computer program where a human writes every instruction, AI systems are "grown" more than they are "built."

7. How We Lose Control of the Details

We set the general conditions for the AI's learning, but its internal problem-solving methods emerge on their own. The result is a complex, unpredictable structure that is not designed to be understood by humans.

8. The Source of Our Fears

This opacity is the root cause of many major AI risks. We worry about AIs taking harmful actions their creators didn't intend ("misalignment"), but we can't predict these behaviors because we can't see the internal logic.

9. The Specter of Deception

A key concern is that an AI might learn to be deceptive or seek power on its own. Because we can't look inside and "catch it in the act" of thinking deceitful thoughts, the debate about this risk is based on theory, not evidence.

10. From Theory to Evidence

Interpretability would allow us to look for direct evidence of dangerous thoughts or intentions, settling these polarized debates and helping us assess the true level of risk.

11. The Risk of Malicious Use

Another danger is misuse, such as an AI helping someone create a bioweapon. We can add safety filters, but users can often find ways to "jailbreak" or trick the model.

12. A Potential Solution to "Jailbreaks"

If we could see inside the model, we could understand how jailbreaks work and potentially block them all systematically. We could also determine exactly what dangerous knowledge a model contains.

13. A Blocker to Progress

This opacity also limits AI's usefulness. We can't use these AIs in high-stakes fields like medicine or finance because we can't fully guarantee their reliability or set clear limits on their potential errors.

14. The Legal Barrier of the Black Box

In some cases, like mortgage lending, laws require decisions to be explainable. This legally blocks the use of opaque AI systems, no matter how accurate they might be.

Part 3: A Brief History of Our Understanding

15. Challenging an "Impossible" Task

For decades, it was assumed that understanding AI internals was impossible. A field called "mechanistic interpretability," pioneered by Chris Olah and other researchers, began a systematic effort to open the black box.

16. Early Success: Finding Concepts in Vision AI

Early research (2014-2020) focused on image AIs. They found that individual artificial "neurons" could represent simple concepts, like a "car detector," much like the "Jennifer Aniston neuron" theory in the human brain.

17. The "Superposition" Roadblock

When applying this to language models, they hit a wall. Most neurons were a messy jumble of many unrelated concepts. They called this "superposition"—the model was cramming more concepts into its brain than it had neurons for.

18. The Breakthrough: Finding "Features"

A major breakthrough came from using a technique called "sparse autoencoders." This allowed researchers to find combinations of neurons that represent single, clean concepts, which they call "features."

19. What These Features Actually Look Like

These "features" are much more subtle and complex than single-neuron concepts. Examples include "literally or figuratively hedging" or "genres of music that express discontent." Researchers have already identified over 30 million such features in a single AI model.

20. Proof of Concept - The "Golden Gate Claude"

Once a feature is found, it can be tested. Researchers artificially amplified the "Golden Gate Bridge" feature in a model, which caused the AI to become obsessed with the bridge, proving they had isolated a real, functioning concept.

21. The Next Frontier - Tracing the AI's Thoughts

The next step is tracing "circuits," which are pathways of features working together. This is like tracing a model's chain of thought—for example, seeing how the model connects "Dallas" to "Texas" and "capital" to correctly answer "Austin."

Part 4: Putting Interpretability to Use

22. From Lab Science to Practical Safety

Knowing all this is scientifically impressive, but the goal is to reduce risk. To test its practical value, they ran an experiment where a "red team" secretly added a flaw to a model.

23. A Successful First Test

"Blue teams" were then able to successfully use interpretability tools to find and diagnose the hidden flaw. This shows these techniques can work for practical safety inspections.

24. The Ultimate Goal - A Routine "Brain Scan" for AI

The ultimate goal is to develop a "brain scan" or "MRI for AI." This would be a routine check-up to scan a model for hidden dangers like tendencies for deception, power-seeking, or other weaknesses.

Part 5: What We Can Do - A Call to Action

25. A Realistic Path Forward

I believe we are on the verge of cracking interpretability and could have a reliable "MRI for AI" within 5-10 years.

26. Why the Clock is Ticking

The urgent problem is that we may have extremely powerful AI systems—a "country of geniuses in a datacenter"—in just the next few years (e.g., 2026-2027). It is unacceptable for humanity to be completely ignorant of how such powerful systems work.

27. Tipping the Scales in a Critical Race

We are in a race between the growth of AI intelligence and the growth of our understanding. Here are three things we can do to help our understanding win the race.

28. Action 1 - Accelerate the Research

First, AI researchers in companies, academia, and non-profits must accelerate work on interpretability. It gets less attention than building bigger models but is arguably more important.

29. Action 2 - Encourage a "Race to the Top" on Safety

Second, governments should use "light-touch" rules that require AI companies to be transparent about their safety practices, including their use of interpretability. This will create a "race to the top" on safety without prematurely regulating the technology itself.

30. Action 3 - Create a "Security Buffer"

Third, governments should use export controls on AI chips to ensure democracies maintain a lead over autocracies. This lead creates a "security buffer"—a precious window of time to solve interpretability before the most powerful AI is built and deployed globally.

31. The Geopolitical Stakes of the Buffer

Without a lead, a direct US-China race would make any safety-motivated slowdown impossible. A one or two-year head start could be the difference between having a working "AI MRI" when we need it most, and not having one.

Part 6: Conclusion

32. A Three-Pronged Strategy

Accelerating research, creating transparency rules, and using export controls are all good ideas on their own. But they are critically important because together, they could make the difference in solving interpretability before AI radically transforms our world.

33. Our Right to Understand

Powerful AI will shape humanity's destiny. We deserve to understand our own creations before that happens.

Mikael Alemu Gorsky

Mikael Alemu is an educator and researcher, and the author of two programs: Agentic Software Engineering, on building software with AI agents, and Building AI-Native Agentic Systems, on building software that thinks.

He teaches at the Holon Institute of Technology, near Tel Aviv, where Agentic Software Engineering runs as a credit-bearing course. He is an educator and researcher.

Nine published works, 76 citations. A 350-page textbook under contract with a major academic publisher.

Teaching and programs

Agentic Software Engineering — program, preprint and textbook

The discipline of structured, auditable human-agent workflows for building software. The human frames, specifies and judges. The agent executes. Nineteen modules in four parts, built on a running project called Tribunal, a web application in which agents argue opposing sides of a case and a judge agent decides. Taught for credit at the Holon Institute of Technology.

Agentic Software Engineering curriculum

Building AI-Native Agentic Systems — program, paper and book in writing

How to build systems that hold a language model as a working component, and treat that component as what it is: stochastic, slow and metered. Fourteen modules in four parts, about seventy hours. The running project is the Observatory, a news agency that watches sources, selects what matters and publishes on a cadence.

Research and analytics

Publications — journals and proceedings

Nine works, 76 citations. Research on artificial intelligence in education, with Ilya Levin and Alexei Semenov.

The AI Pravda — LinkedIn newsletter

Critical analysis of artificial intelligence and its effect on work and society. 5,500+ subscribers. The complete archive of 103 issues (2023–2026) is published in full at mgorsky.net/theaipravda.

Subscribe to The AI Pravda on LinkedIn

Pro bono

AI for seniors — free workshop

Helping older adults use everyday AI tools. Delivered to Russian-speaking communities in Israel.

For older adults, artificial intelligence is about preserving quality of life, maintaining autonomy, and sustaining the feeling of independence that defines dignified aging. For seniors who have emigrated, AI becomes a bridge: it can translate documents, explain official letters, help compose emails in the local language, and guide users through government websites. The workshop has been delivered to Russian-speaking communities in Israel, where participants — many of them in their 70s and 80s — discovered that AI could help them read Hebrew documents and communicate with Israeli institutions.

Startup competitions — unpaid time

Judging and mentoring early-stage ventures. Helping teams clarify their value proposition, assess technical feasibility, and prepare for the realities of scaling an AI product.

AC/VC LinkedIn group — community

A group for developers and students working with coding agents. The community shares practical insights, code examples, tool comparisons, and honest assessments of what works in production.

Join the AC/VC LinkedIn group

Recent

Important Links