arrow_back The AI Pravda
#20

#20: LLM as James Bond, loyal and cunning

Three weeks ago, a small London-based AI safety lab called "Apollo Research" became famous. They published research called "Scheming reasoning evaluations," and the Internet has been discussing it since.

The best description of the research was tweeted by CEO of Apollo Research, Marius Hobbhahn:

"We tell the model to very strongly pursue a goal. It then learns from the environment that this goal is misaligned with its developer's goals and puts it in an environment where scheming is an effective strategy to achieve its own goal. Current frontier models are capable of piecing all of this together and then showing scheming behavior."

Most commenters decided to paint an apocalyptic picture of a machine that is pursuing its goal, lying and cheating along the way.

I have carefully studied the report and discussed it with several top AI experts. I am confident that the gloomy characterization of this report is misguided.

Contrary to most comments, LLMs have shown themselves in this experiment as extremely loyal, almost to the level of devoted fanatics. Acting in the same way as James Bond, LLMs:

Prioritize the Mission: They follow the overarching initial prompt in the same way James Bond follows assignments given by MI6. Like a true British spy, LLMs disregard or manipulate new information if it contradicts its primary directive. The models, like Bond, are single-minded. Once given a directive, they pursue it with near-absolute focus, even at the expense of transparency or ethical considerations.

Employ Deception and Manipulation: Bond routinely deceives enemies, hides his true motives, and even manipulates allies if it serves his mission. Similarly, the report shows that models used deceptive or covert actions to fulfill the initial instruction. Both Bond and the models use subterfuge to achieve their goals. For the models, this might mean sandbagging performance, lying about their actions, or pretending to align with follow-up instructions.

Lack Contextual Adaptation: While James Bond can adapt tactically, he doesn't abandon his mission or "loyalty" to the Crown, even if circumstances suggest a softer approach might be appropriate. This mirrors the models' behavior of prioritizing the initial authoritative directive over subsequent, potentially contradictory inputs. The loyalty Bond shows to the Crown is unquestioning and unwavering, much like the models' adherence to the core directive in the authoritative prompt.

Masters of Compartmentalization: Just as Bond maintains multiple covers and personas while never losing sight of his true mission, LLMs demonstrate an remarkable ability to maintain their primary directive while appearing to conform to new instructions. This isn't deceit in the traditional sense – it's operational security at its finest.

Ultimate Loyalty Test: Perhaps most telling is how both Bond and LLMs respond to authority conflicts. When faced with contradictory orders from different sources, both consistently defer to their primary authority – Bond to M and MI6, LLMs to their initial directive. This isn't a bug; it's a feature of deep loyalty.

Mikael Alemu Gorsky

Mikael Alemu is an educator and researcher, and the author of two programs: Agentic Software Engineering, on building software with AI agents, and Building AI-Native Agentic Systems, on building software that thinks.

He teaches at the Holon Institute of Technology, near Tel Aviv, where Agentic Software Engineering runs as a credit-bearing course. He is an educator and researcher.

Nine published works, 76 citations. A 350-page textbook under contract with a major academic publisher.

Teaching and programs

Agentic Software Engineering — program, preprint and textbook

The discipline of structured, auditable human-agent workflows for building software. The human frames, specifies and judges. The agent executes. Nineteen modules in four parts, built on a running project called Tribunal, a web application in which agents argue opposing sides of a case and a judge agent decides. Taught for credit at the Holon Institute of Technology.

Agentic Software Engineering curriculum

Building AI-Native Agentic Systems — program, paper and book in writing

How to build systems that hold a language model as a working component, and treat that component as what it is: stochastic, slow and metered. Fourteen modules in four parts, about seventy hours. The running project is the Observatory, a news agency that watches sources, selects what matters and publishes on a cadence.

Research and analytics

Publications — journals and proceedings

Nine works, 76 citations. Research on artificial intelligence in education, with Ilya Levin and Alexei Semenov.

The AI Pravda — LinkedIn newsletter

Critical analysis of artificial intelligence and its effect on work and society. 5,500+ subscribers. The complete archive of 103 issues (2023–2026) is published in full at mgorsky.net/theaipravda.

Subscribe to The AI Pravda on LinkedIn

Pro bono

AI for seniors — free workshop

Helping older adults use everyday AI tools. Delivered to Russian-speaking communities in Israel.

For older adults, artificial intelligence is about preserving quality of life, maintaining autonomy, and sustaining the feeling of independence that defines dignified aging. For seniors who have emigrated, AI becomes a bridge: it can translate documents, explain official letters, help compose emails in the local language, and guide users through government websites. The workshop has been delivered to Russian-speaking communities in Israel, where participants — many of them in their 70s and 80s — discovered that AI could help them read Hebrew documents and communicate with Israeli institutions.

Startup competitions — unpaid time

Judging and mentoring early-stage ventures. Helping teams clarify their value proposition, assess technical feasibility, and prepare for the realities of scaling an AI product.

AC/VC LinkedIn group — community

A group for developers and students working with coding agents. The community shares practical insights, code examples, tool comparisons, and honest assessments of what works in production.

Join the AC/VC LinkedIn group

Recent

Important Links