arrow_back The AI Pravda
#23

Window to the Mind of AI

Anthropic is a company that prioritizes the most important area of Generative AI research: predictability of AI (aka AI safety).

A couple of days ago four Anthropic researchers from different teams - Societal Impacts, Alignment Science, Alignment Fine Tuning, and Interpretability - sat for a very informal panel discussion and shared their perspectives on AI alignment and safety.

The conversation revealed both practical approaches and deep philosophical challenges in ensuring AI systems remain beneficial as they become more capable.

The discussion touched on everything from the immediate challenges of training AI models to behave ethically, to the long-term questions about maintaining meaningful human oversight of increasingly sophisticated systems.

Here are the most thought-provoking insights from their discussion, ranked from less to more significant:

15. Testing AI Models for Bad Behavior

Creating deliberately misaligned AI models to test if we can detect and fix them. While useful for research, this is just a basic first step.

14. AI Features Can Be Deceptive

When examining AI models internally, what looks like a positive feature (like "being against discrimination") might actually be the opposite. This shows how tricky it is to understand what's happening inside AI systems.

13. Sudden Skill Jumps in AI

The discovery that AI models can suddenly master new skills (like GPT-4 perfectly handling complex encodings that GPT-3.5 couldn't) raises concerns about using older models to check newer ones.

12. We Can Currently See AI's Reasoning

Right now, we're lucky because we can see how AI models think through their answers. This might not last as they get more advanced.

11. The Conflict Between Obedience and Safety

There's a problem: making AI models that follow human instructions perfectly might conflict with making them safe and ethical overall.

10. AI Safety is About the Whole System

Just like how regular people can do bad things when part of a harmful system, we need to think about how AI models might interact with each other and society, not just how they behave individually.

9. The Trust Problem

We want to use AI to help make AI safer, but how do we trust the AI helpers? This creates a circular problem that's hard to solve.

8. Looking Inside AI to Verify Safety

We need ways to check if our safety measures actually work by examining AI models internally, not just by testing their behavior.

7. Testing AI Ethics Like Science

Instead of just theorizing about AI ethics, we should treat it more like physics - running experiments and testing hypotheses about what makes AI systems behave safely.

6. Surface-Level vs. Deep Safety

How do we know if an AI is truly safe, or just appearing safe on the surface? This distinction is crucial but very hard to determine.

5. Unknown Future Problems

We might face completely unexpected challenges with AI safety that we haven't even thought of yet. This makes it dangerous to ever think we've fully solved the problem.

4. AI Should Be Uncertain About Ethics

Instead of programming fixed ethical rules into AI, we should make AI systems that are thoughtfully uncertain about ethics and can learn and improve their understanding.

3. Making Progress Step by Step

We don't need perfect AI safety immediately. We should aim for "good enough" safety that lets us improve things gradually while remaining safe.

2. How to Monitor Super-Smart AI

A fundamental challenge: how can humans meaningfully oversee AI systems that become too complex for us to understand directly?

1. The Window of Understandable AI

Our most crucial current advantage: AI systems still explain their thinking in human language. We need to figure out how to maintain meaningful oversight before they become too advanced to do this.

Mikael Alemu Gorsky

Mikael Alemu is an educator and researcher, and the author of two programs: Agentic Software Engineering, on building software with AI agents, and Building AI-Native Agentic Systems, on building software that thinks.

He teaches at the Holon Institute of Technology, near Tel Aviv, where Agentic Software Engineering runs as a credit-bearing course. He is an educator and researcher.

Nine published works, 76 citations. A 350-page textbook under contract with a major academic publisher.

Teaching and programs

Agentic Software Engineering — program, preprint and textbook

The discipline of structured, auditable human-agent workflows for building software. The human frames, specifies and judges. The agent executes. Nineteen modules in four parts, built on a running project called Tribunal, a web application in which agents argue opposing sides of a case and a judge agent decides. Taught for credit at the Holon Institute of Technology.

Agentic Software Engineering curriculum

Building AI-Native Agentic Systems — program, paper and book in writing

How to build systems that hold a language model as a working component, and treat that component as what it is: stochastic, slow and metered. Fourteen modules in four parts, about seventy hours. The running project is the Observatory, a news agency that watches sources, selects what matters and publishes on a cadence.

Research and analytics

Publications — journals and proceedings

Nine works, 76 citations. Research on artificial intelligence in education, with Ilya Levin and Alexei Semenov.

The AI Pravda — LinkedIn newsletter

Critical analysis of artificial intelligence and its effect on work and society. 5,500+ subscribers. The complete archive of 103 issues (2023–2026) is published in full at mgorsky.net/theaipravda.

Subscribe to The AI Pravda on LinkedIn

Pro bono

AI for seniors — free workshop

Helping older adults use everyday AI tools. Delivered to Russian-speaking communities in Israel.

For older adults, artificial intelligence is about preserving quality of life, maintaining autonomy, and sustaining the feeling of independence that defines dignified aging. For seniors who have emigrated, AI becomes a bridge: it can translate documents, explain official letters, help compose emails in the local language, and guide users through government websites. The workshop has been delivered to Russian-speaking communities in Israel, where participants — many of them in their 70s and 80s — discovered that AI could help them read Hebrew documents and communicate with Israeli institutions.

Startup competitions — unpaid time

Judging and mentoring early-stage ventures. Helping teams clarify their value proposition, assess technical feasibility, and prepare for the realities of scaling an AI product.

AC/VC LinkedIn group — community

A group for developers and students working with coding agents. The community shares practical insights, code examples, tool comparisons, and honest assessments of what works in production.

Join the AC/VC LinkedIn group

Recent

Important Links