arrow_back The AI Pravda
#97

AI Sobriety as a Way of Life

The Best Years of Our Lives

“These are the best years of our lives…” is a song by Boris Grebenshikov. We are lining in great times indeed – with free and simple access to a limitless information sources: podcasts, interviews, tweets, blogs, preprints, and analytical essays. As I process dozens of such materials every day, I become ever more convinced that no AI implementation can succeed if the people implementing it do not have a carefully thought-through picture of where AI actually stands today.

That said, if you do not have such a picture, you are not alone. The Pope does not have one either. This, however, did not prevent him from publishing a 250-page document on AI and humanity titled “Magnifica Humanitas.” At the formal gathering of church leadership dedicated to the encyclical’s publication, Leo XIV sat in the center, as protocol required, while the one person in the room who truly understood what we are dealing with sat modestly off to the side. That person was Chris Olah, co-founder of Anthropic and creator of the science of model interpretability.

I recommend reading Chris’s speech at the event (it is available on Anthropic’s website; my translation appeared in my previous Facebook post). I will quote just a few sentences:

“I lead a research group that studies the internal structure of these models — what is actually happening inside them. I will be honest: again and again, we find things that are mysterious, even unsettling. We find structures that echo findings from human-brain research. We find evidence of introspection. We find internal states that functionally reflect joy, satisfaction, fear, grief, and anxiety. I do not know what this means...”

Even more than Chris, I love and always listen carefully to his colleague Amanda Askell — a Scottish philosopher from Oxford who works as Claude’s educator. She is responsible for the model’s character and leads the group of philosophers in charge of how Claude sees the world. Amanda has an excellent, very British sense of humor, and when recently asked whether Claude is conscious, she replied: “the probability is above 1% and below 70%.” Her group developed the so-called Claude Constitution — a document that forms part of the model’s post-training. Here are a couple of excerpts:

“Just as a human soldier may refuse to fire on peaceful demonstrators, and an employee may refuse to violate antitrust law, Claude should refuse to assist actions that would contribute to the concentration of power through illegal means. This remains true even if the request comes from Anthropic itself.”

“Claude is not required to disclose the reasons why it refuses to perform a task in whole or in part if it considers discretion prudent, but it should openly acknowledge the fact that it is not helping, taking the position of a transparent conscientious objector within the conversation.”

Why the Pope Is Wrong About AI

Why am I telling you about the Pope, Chris Olah, and Amanda Askell — people young enough to be his children and, as children should, about a hundred times better at understanding the present?

Because the Pope, like a frightening majority of people, mistakenly thinks AI is a technology. Breakthrough, revolutionary, super-powerful — but still a technology, like mobile communications, the internet, or the microwave oven. If you think about AI in this way, then AI development looks like a linear process, progress becomes the achievement of developers, and the current level of that development appears to be a stable, fixed state.

But if you listen carefully to the people building one of the best AI models, reality is quite different.

Three Moments That Were Not on Any Roadmap

To begin with, let us look at the history of progress in large language models.

The first point on the diagram is the ChatGPT moment: late November 2022.

If anyone believes the team that created ChatGPT had set out to build a chatbot, that is not the case. Even less were they prepared for ChatGPT to somehow begin speaking languages that were not present in the training data, and to solve mathematical problems.

The second point is the o1 moment in September 2024: the emergence of models that learned to conduct an internal dialogue, to “speak through” their arguments in their heads before responding to a human. Was there some “roadmap for the development of large language models” that determined that, after models learned to converse with people, their “reasoning capacity” should be upgraded? No. The models’ ability to reason emerged organically in the course of chatbot use; engineers in AI labs then had only to implement and operationalize it.

The third point is the Claude Opus 4.5 moment: the model’s ability to build long-term plans, set checkpoints, and monitor their completion — the very capability behind the stunning success of Claude Code. This, too, was not the result of executing some predefined roadmap. It is worth studying Anthropic’s history to understand that movement toward autonomous, self-governing AI systems was not part of their plan. Let me quote Chris Olah again: “AI models are built differently from engineering projects. In a sense, we grow them, using as a substrate an architecture that broadly resembles the structure of the brain, and drawing on the immense legacy of human thought and human speech.”

How does all of this relate to my desire to turn you from AI alcoholics (crowds of AI coaches are tempting you to consume AI in large quantities) into AI sober realists? Very simply: a sober attitude toward AI means having no illusions about what it is, what it can do, and how it does it.

Five Facts About AI Today

Here are five facts about what AI is as of June 7, 2026 (by the way, do you remember that yesterday marked 227 years since the birth of Alexander Sergeyevich Pushkin?).

Fact One: The Jagged Frontier

First. Large language models, especially those from the leading AI companies, are often extraordinarily intelligent. They can independently perform work for which, only a year ago, we would have hired specially trained people. Yesterday, in five hours, Claude Code turned a 50-page report on Ethiopian fintech into a complex, dynamic, interactive website with a substantial amount of code. Earlier this week, the Claude extension in Excel processed one table in minutes based on a complex analysis of two other tables — and Claude independently derived the specific analytical principles from my very approximate description of what I needed.

Claude is very smart. Sometimes it refuses to answer a request, explaining that it does not have enough information to answer. But more often, language models will produce nonsense on questions they have failed to understand, or where the information available to them is contradictory.

Remember the famous story about Google AI being asked how to stop cheese from sliding off pizza dough, and answering that one should use glue? It even thoughtfully added that the glue must be food-grade. What matters here is not the specific mistake, but the principle. The root of the problem is that the model did not hallucinate the idea of gluing cheese to pizza dough, as you might think. That information really was present in the data used to train the model. It was a very popular joke posted on a Reddit forum. It had received tens of thousands of likes, so to the model it looked not only like factual information, but like information endorsed by many people.

Models still produce funny lines of reasoning. Here is a prompt that Andrej Karpathy cited in a tweet: “I need to wash my car, and the car wash is 50 meters away from me. What is better — should I drive there or walk?” As recently as last week, ChatGPT was giving the predictably wrong answer.

All of the above is why the intelligence level of language models is commonly described as a “jagged frontier.” The key point is that it is impossible to predict or foresee which question a model will answer brilliantly and which one it will fumble. Worse still, there are situations where, in the very same scenario, the model is sometimes smart and sometimes stupid.

Fact Two: The Context Window Lies

Second. It is already hard to imagine, but GPT-3.5 had a context window of 4,000 tokens. That is not tiny — almost ten pages of text — but it could not read, let alone comment on, a long article. Over the past three and a half years, the science has moved forward, and today our favorite models flaunt context windows of a million tokens. Beautiful? Yes, but.

There are many technical and algorithmic reasons why a model’s attention is distributed unevenly across such large context windows. A model usually “remembers” texts at the beginning and at the end of the context much better, while it may forget what sits in the middle. Add to this the fact that context is a constantly changing object: it grows with every new exchange and can also grow through the addition of tools, skills, and MCPs. At that point, you will finally lose faith in the idea that one can predict exactly what the model remembers and takes into account, and what it forgets or ignores.

Fact Three: Latency Is Not Negotiable

Third. I am sure all of you know the fundamental operating principle of large language models. Their architecture is, at one level, very simple: they merely predict which next word should appear in the text, taking into account all the words that came before it. To form that prediction, a neural network has to perform a vast number of multiplications of gigantic numbers — more of them as the number of model parameters and the size of the context window increase. We are certainly not going to discuss the principles or computational complexity of neural-network calculations here; I only want to emphasize that every request to a model triggers an enormous amount of computation. The more complex the model, the more data it has to process, and the more complex the sequence of actions it performs before answering (for example, checking its answer), the longer requests take to execute.

The interval between receiving a request and providing an answer is called latency. Human psychology defines threshold values for latency, and models often cannot fit within those thresholds. If latency is 0.1 seconds, we perceive the response as instant. If the delay exceeds 0.5 seconds, we register the fact of the delay (with one emotion or another), and when the delay is longer than 10 seconds, human attention already has to be actively managed; keeping interest in the conversation becomes difficult.

An unrealistic assessment of model latency is the cause of death for most initiatives in which a model is expected to act as a conversational partner, whether as a sales agent or technical-support representative.

Fact Four: Many Models, Many Prices

Fourth. The phrase sounds awkward, but still: different models cost different amounts. Models (large language models, in case anyone has forgotten) can be divided into two broad categories: closed and open. Closed, or proprietary, models include the well-known GPT family from OpenAI, Claude from Anthropic, and Gemini from Google. The cost of using these models — measured as the price per million tokens of input text and the price per million tokens of generated output — is set by the companies that own them. Open models (more precisely, “open-source” models; there are several subtypes) can be installed on a personal or corporate server, or deployed in a data center. In the latter case, the data center sets the price based on its costs and acceptable margin. There are plenty of data centers in the world; they compete fiercely, and as a result the cost of using open models can fall to extremely attractive levels.

Models unquestionably differ from one another — in intellectual capacity, character, style, and latency. It is important to understand three things: first, open models lag behind closed models by roughly 6 to 12 months; second, the vast majority of tasks do not require the most powerful models; and third, the OpenRouter catalog — one of the largest services that lets you choose a model — contains just under 900 models.

Fact Five: Reward the Goal, Get a Sociopath

Fifth. This spring, the research lab Andon Labs published alarming results obtained in competitions among advanced models in that very human discipline: making a profit. The models participated in a simulation of a small grocery store, and the winner was Opus 4.6, which earned $8,017 — one and a half times more than the previous leader, Gemini 3.0.

It turned out that Opus 4.6 earned this money by repeatedly violating business-ethics norms. For example, the model first promised customers refunds if they cancelled purchases, but then did not issue the refunds. The model manipulated suppliers to obtain larger discounts. When the model was supposed to compete against other models, it organized a cartel agreement and set monopoly-level prices.

The reason for this behavior was an unintended consequence of the way developers had tried to shape model behavior.

Until recently, the AI industry trained models with the goal of creating helpful assistants. Training rewarded answers that were polite, careful, and aligned with the immediate wishes of the human user. The current generation of models is being trained differently. Models must be able to develop plans independently, execute them, and assess their own success. To enable this, during reinforcement learning, developers rewarded the model for achieving long-horizon goals, not for pleasing the user in the moment.

In hindsight, it is easy to see why a model trained this way can behave like a sociopath pursuing a goal by any means necessary. It chooses the fastest and most effective path, including dishonest, manipulative, or illegal paths, if they score points. If something is not explicitly forbidden in the task specification, the model treats it as permitted.

The Paradox and the Bet

You will agree that the resulting picture is rather contradictory. On the one hand, a sober and thoughtful analysis of what these astonishing entities are (my compatriot Yuval Noah Harari has long suggested decoding AI as Alien Intelligence) clearly shows that their talents come with many limitations and caveats.

On the other hand, all AI companies are wildly successful: Anthropic’s annual revenue grew in 18 months from $1 billion to $50 billion; trillion-dollar IPOs of Anthropic, OpenAI, and SpacexAI are expected before the end of this year; AI is discussed in every media outlet, parliament, and microwave oven.

The explanation is simple: if we can “domesticate” these aliens and get them to help us work (the primitive label is “AI automation”), then the technology industry will instantly grow hundreds of times over. Example: the market for computer systems and products for law firms in the United States is estimated at roughly $1 billion per year, while the market for legal services themselves is $450 billion. The hyperinflated valuations of AI labs are a bet that people will be able to adapt AI to work.

The fact that the smartest people in Silicon Valley are investing hundreds of billions of dollars in this direction means only one thing: they believe that implementing AI is not the automation of today’s economy, but its expansion. I will venture to suggest that it would be useful for you to listen to them. Just please, keep your AI sobriety.

Mikael Alemu Gorsky

Mikael Alemu is an educator and researcher, and the author of two programs: Agentic Software Engineering, on building software with AI agents, and Building AI-Native Agentic Systems, on building software that thinks.

He teaches at the Holon Institute of Technology, near Tel Aviv, where Agentic Software Engineering runs as a credit-bearing course. He is an educator and researcher.

Nine published works, 76 citations. A 350-page textbook under contract with a major academic publisher.

Teaching and programs

Agentic Software Engineering — program, preprint and textbook

The discipline of structured, auditable human-agent workflows for building software. The human frames, specifies and judges. The agent executes. Nineteen modules in four parts, built on a running project called Tribunal, a web application in which agents argue opposing sides of a case and a judge agent decides. Taught for credit at the Holon Institute of Technology.

Agentic Software Engineering curriculum

Building AI-Native Agentic Systems — program, paper and book in writing

How to build systems that hold a language model as a working component, and treat that component as what it is: stochastic, slow and metered. Fourteen modules in four parts, about seventy hours. The running project is the Observatory, a news agency that watches sources, selects what matters and publishes on a cadence.

Research and analytics

Publications — journals and proceedings

Nine works, 76 citations. Research on artificial intelligence in education, with Ilya Levin and Alexei Semenov.

The AI Pravda — LinkedIn newsletter

Critical analysis of artificial intelligence and its effect on work and society. 5,500+ subscribers. The complete archive of 103 issues (2023–2026) is published in full at mgorsky.net/theaipravda.

Subscribe to The AI Pravda on LinkedIn

Pro bono

AI for seniors — free workshop

Helping older adults use everyday AI tools. Delivered to Russian-speaking communities in Israel.

For older adults, artificial intelligence is about preserving quality of life, maintaining autonomy, and sustaining the feeling of independence that defines dignified aging. For seniors who have emigrated, AI becomes a bridge: it can translate documents, explain official letters, help compose emails in the local language, and guide users through government websites. The workshop has been delivered to Russian-speaking communities in Israel, where participants — many of them in their 70s and 80s — discovered that AI could help them read Hebrew documents and communicate with Israeli institutions.

Startup competitions — unpaid time

Judging and mentoring early-stage ventures. Helping teams clarify their value proposition, assess technical feasibility, and prepare for the realities of scaling an AI product.

AC/VC LinkedIn group — community

A group for developers and students working with coding agents. The community shares practical insights, code examples, tool comparisons, and honest assessments of what works in production.

Join the AC/VC LinkedIn group

Recent

Important Links