1. The distinction that media, influencers, and pundits ignore
The phrase "AI in warfare" covers two technologies that share nothing except a label. The first has existed in military systems for decades. The second appeared in combat for the first time in 2024. Their capabilities differ, their failure modes differ, and the risks they pose to civilians differ. Mixing them up is not a harmless simplification. It leads to the wrong fears, the wrong debates, and the wrong policies.
Machine learning (pattern recognition, optimization, classification) covers computer vision on drone cameras and satellite imagery; sensor fusion across radar, infrared, acoustic, and electromagnetic spectrum; scoring algorithms like Israel's Lavender (supervised learning that classifies individuals by statistical correlation with known militants); combinatorial optimization for strike planning (matching weapons to targets, optimizing sortie schedules); autonomous navigation for drones in GPS-denied environments; and missile defense interception decisions at superhuman speed.
All of these produce outputs verifiable against physical reality. A classifier either correctly identifies a tank or it does not. A navigation algorithm either reaches the target or it does not. The error modes are measurable. Systems can be tested, calibrated, their false-positive rates calculated. Lavender's 10% error rate was a known, measured quantity, even if the decision to accept it was reckless. The lineage of these systems runs through signals processing, operations research, and statistical pattern recognition. The Israeli Iron Dome's interception algorithms, the US Navy's Aegis combat system, and WWII-era fire control computers all sit on the same continuum. Calling them "AI" is technically accurate (we are talking about technology based on neural networks) but obscures their continuity with pre-LLM military technology.
Large language models (generative text systems) cover natural-language querying of intelligence databases; synthesis of multi-source intelligence into briefings; document summarization (orders, reports, intercepted communications); generation of textual components for target packages (descriptions, legal rationale); battlefield simulation through scenario generation; and advisory functions ("how should I prosecute this target given these constraints").
The outputs of LLMs cannot be verified against physical reality by the user at the point of consumption. A summary either captures the essential content of the source documents or it does not, and the commander reading it has no way to know which without reading the sources themselves. A legal justification sounds authoritative whether or not it correctly applies the law of armed conflict. A recommendation sounds confident whether or not it accounts for all relevant factors.
2. Machine learning in the current wars
ML-based military systems are already mature, deployed at scale, and consequential. They do not need LLMs to be dangerous.
Israel: Matzpen and the C2 architecture. The IDF's Matzpen unit builds the command-and-control software layer. Maestro provides multi-arm situational awareness. Lohem plans strikes by matching weapons to targets. Gannt-IT synchronizes operational schedules. Sensor-fed platforms push data to field commanders in real time. These are integration and optimization systems, not language models. The architecture resembles a corporate ERP system: sensors at the edge, planning in the middle, command at the top, single source of truth. The difference is that errors kill people.
Ukraine: autonomous drones at scale. Ukraine produced 4.5 million drones in 2025 (up from 2.2 million in 2024). In January 2026 alone, 7,495 robotic operations were conducted. Adding ML-enabled autonomous navigation to FPV drones raised the strike success rate from 10-20% to 70-80%. Training to operate ML-enabled drones takes 30 minutes to one day. Russian electronic warfare jamming has forced Ukrainian drones toward greater autonomy: when remote operation becomes impossible, the drone must decide on its own. This is the practical path to autonomous weapons, not a policy decision but an operational necessity in the electromagnetic environment. Russia fields V2U attack drones with Nvidia Jetson Orin processors for autonomous navigation in GPS-denied areas. Both sides collect massive datasets: one Ukrainian nonprofit has gathered 2 million hours (228 years) of battlefield footage since 2022. In March 2026, Ukraine launched a program to share its combat data with allied nations to train ML models.
US missile defense interoperability. US-Israel data integration in the current Iran war operates at machine speed. When an American radar in the Gulf detects a launch, the target appears on Israeli Arrow or Patriot battery screens, sometimes before the Israeli radar has detected it. The Pax Silica agreement (January 2026) extends this data-sharing architecture to a coalition including Japan, South Korea, Singapore, UK, Australia, UAE, and Qatar. These are sensor fusion and track-assignment systems, not LLMs.
Humanoid robots: early stage. Foundation's Phantom MK-1 was deployed to Ukraine in February 2026 for battlefield testing, the first known humanoid robot in an active warzone. The robot stands 5'9", weighs 175-180 lbs, uses computer vision for navigation. Plans call for 50,000 units by end of 2027 at a target cost below $20,000 at scale. The technology is early-stage (the robot fell repeatedly during demonstrations). But the trajectory from testing to deployment is expected to take 1-3 years. This is ML in a new form factor, not LLM territory.
3. LLMs enter the battlefield: what actually happened
The Maven Smart System architecture. Maven is Palantir's data integration platform, born in 2017 as a computer vision project for drone video analysis. Over the years it grew into a multi-source intelligence fusion system operating across every US combatant command, serving over 25,000 users. The core is ML analytics: object detection, pattern matching, geospatial correlation, change detection. This layer uses classical machine learning and has been the workhorse for years.
In late 2024, Anthropic's Claude 3 and 3.5 family models were integrated as a language interface layer via Palantir's AI Platform (AIP) on AWS. They received Impact Level 6 accreditation (Secret level). In June 2025, Anthropic launched "Claude Gov" for classified environments. The LLM layer allows analysts to query the system in natural language, receive synthesized briefings, summarize intelligence reports, and generate structured outputs. Palantir describes the system's "Ontology objects" (Satellite Image, Detection, etc.) that provide the LLM with structured information to reason about.
The model version matters. The models that received IL6 accreditation were the Claude 3 and 3.5 family. Security clearance and accreditation processes take many months. The model operating during the Iran strikes in February 2026 was very likely Opus 3.5 at best, possibly older. The current frontier models (Claude 4.x series) would not have passed through the accreditation pipeline. The capabilities that dominate public discussion (advanced reasoning, extended context) are far ahead of what was actually deployed.
The "1,000 targets in 24 hours" headline needs disaggregation. The ML analytics pipeline (computer vision, sensor fusion, database lookup) did the heavy lifting of identifying physical locations and matching them to intelligence. Claude's contribution was in the synthesis layer: turning raw analytical outputs into structured target packages with textual components. The headline "Claude identified 1,000 targets" conflates the two layers. More accurate: Maven's ML analytics identified the targets; Claude wrote the paperwork.
The replacement problem confirms this. When the Pentagon ordered Claude removed, Reuters reported that the architecture "relies on numerous prompts and workflows built with Anthropic's Claude Code developer framework." The replacement requires "extensive redesign of internal components" and "several months." This tells us Claude is embedded in the text generation and workflow automation pipeline, not in sensing or detection. If Claude were doing the target identification, replacing it would require new detection algorithms. Instead, it requires rewriting prompts. That is an LLM-shaped problem.
Israel's Genie. The IDF's classified LLM (described in the Ynet article on the Matzpen unit) initially served clerical purposes: summarizing classified documents, transcribing recordings. Then it was "called up for combat" and began "advising commanders on how to manage the battle." This migration from clerical tool to battlefield advisor mirrors the trajectory described below: from high-confidence to low-confidence applications.
4. How LLMs could be used in military operations: a reliability gradient
Frontier LLMs (2025-2026 capability) can plausibly serve military functions across a spectrum from safe to dangerous, determined by whether the task is primarily linguistic or primarily judgmental.
High-confidence applications (language tasks, error is non-lethal). Translation and transcription of intercepted communications, including low-resource languages where human translators are scarce. The PLA explicitly identified this as a priority, noting insufficient foreign-language personnel. Document summarization for intelligence analysts drowning in report volume; a single Predator mission generates hundreds of hours of video and corresponding text reports. Drafting routine operational documents: orders, status updates, logistics requests. Training and simulation: generating scenarios for wargaming, creating dialogue for simulated adversaries, populating computer-generated forces.
Medium-confidence applications (judgment-adjacent, error matters). Natural-language querying of intelligence databases, where the risk is that the LLM presents an incomplete synthesis as complete. Battle damage assessment, where plausible-sounding textual analysis may overstate or understate damage. Legal review generation, where the Maven system reportedly produces initial drafts of IHL justifications that sound correct regardless of whether they actually apply the law accurately. Anomaly detection in reporting, where the LLM might flag false inconsistencies or miss genuine ones.
Low-confidence applications (genuinely dangerous). Advisory functions: "recommend a course of action given these constraints." This is where Genie and Claude-in-Maven are reportedly heading. The LLM generates a recommendation that sounds like it came from an experienced staff officer. The commander has no way to evaluate whether it accounts for all relevant factors. Real-time decision support under time pressure, where at 86 seconds per target the LLM's output becomes the de facto decision. Cross-domain intelligence synthesis, where the LLM produces text that reads like a unified assessment across signals, human, imagery, and open-source intelligence. Whether this synthesis is accurate remains unverifiable at the point of consumption.
The pattern. LLMs add the most value where the task is primarily linguistic and the stakes of error are recoverable. They add the most risk where the task is primarily judgmental and the stakes are lethal. The military's current trajectory pushes LLMs toward the high-risk end, because that is where the time savings are greatest.
5. The "human in the loop" fiction
When ML or LLM systems generate targets, the human's task changes from "what should I strike" (initiative) to "is there a reason not to strike what the system recommends" (skepticism under time pressure). These are cognitively opposite tasks. Humans are bad at the second one.
David Leslie (Queen Mary University of London) described the situation as "a potential scaled hazard of rubber stamping, where because of the speed involved, you don't have active human, critical human engagement." Craig Jones (Newcastle University): "Humans are technically in the loop. That doesn't mean they are in the loop enough to have effective decision-making power."
A Defense One investigation (March 2026) found that the Pentagon's rapid adoption of ML and LLM tools may be eroding the military's ability to tell fact from fiction. Research shows that reliance on automated systems to perform tasks degrades the human's native ability to perform them independently. French Admiral Pierre Vandier (NATO): "The more you use AI, the more you will use your brain in a different way."
An Oxford International Affairs paper (January 2026) found that humans naturally anthropomorphize automated systems, attributing reasoning and morality to machines that possess neither. This misconception is amplified when the system produces authoritative-sounding natural language, which is specifically what LLMs do. An ML system that outputs a numerical score (Lavender's 0-100 risk rating) invites at least some scrutiny. An LLM that outputs a paragraph of fluent analysis invites trust. The paper proposes confidence scores, uncertainty indicators, and justification fields as countermeasures, but notes that no such mechanisms were in place during the Iran operations.
The automation bias problem is well documented in aviation and medicine. When an automated system recommends an action, the human operator agrees in the overwhelming majority of cases, not because of laziness but because the operator has no independent information source to justify disagreement. In the military context, the "default" for each generated target becomes the strike. Not striking requires effort and justification. Approval requires nothing.
6. The Pentagon's doctrine
On January 9, 2026, Defense Secretary Hegseth issued the "Artificial Intelligence Strategy for the Department of War." Core premise: "We must accept that the risks of not moving fast enough outweigh the risks of imperfect alignment." The strategy mandates frontier models deployed to soldiers within 30 days of public release.
Budget: $13.4 billion requested for autonomous and ML/LLM-driven systems in 2026. Total Pentagon allocation to such programs since 2016 exceeds $75 billion (Brennan Center estimate, likely understated due to classified programs). Additional $9 billion in data center infrastructure.
In March 2026, Maven was designated a "program of record," making its ML and LLM capabilities permanent military infrastructure. The Maven contract has grown to $1.3 billion. The Army is shifting acquisition to a "commercial-first model." Under Secretary of the Army Obadal: "This month, our army is engaged in active combat." Acquisition is now treated as "a warfighting function."
Pentagon contracts of up to $200 million each were awarded to Anthropic, OpenAI, Google, and xAI. The European defense industry moves in parallel: UK's ASGARD targeting system (ML-based) claims a kill chain under one minute; Germany's Uranos KI is planned for 2026 deployment; Helsing has delivered thousands of ML-guided loitering munitions to Ukraine.
7. China and military LLMs
China is pursuing military LLM applications with specific, documented efforts.
ChatBIT was published in June 2024 by researchers from three institutions, including two under the PLA's Academy of Military Science. They adapted Meta's Llama 2 (13B parameters) with fine-tuning on 100,000 military dialogue records, creating a model "optimized for dialogue and question-answering tasks in the military field." The Jamestown Foundation notes this was the first substantive indication of PLA experts actively investigating open-source LLMs for military applications.
PLA LLM applications under development: intelligence analysis via natural-language queries over fused data; computer-generated forces in training simulations with LLM-enhanced autonomous behavior; electronic warfare strategy training (a PLA-connected aviation firm used Llama 2 for this); cognitive and psychological warfare content generation; and predictive policing already deployed domestically.
DeepSeek's military potential. Chinese reports suggest DeepSeek can integrate data from drones, satellites, and radars to aid military decisions. The Pentagon's 2025 report to Congress notes China has "narrowed the performance gap" in LLMs and is pursuing "algorithmic" and "network-centric" warfare at various levels of autonomy by 2030.
Scale. By mid-2025, China accounted for 1,509 of the world's approximately 3,755 publicly released LLMs. Civil-military fusion doctrine means commercial capabilities flow into military applications by design. Open-source models from both the US (Llama) and Chinese ecosystems provide the PLA with building blocks requiring modest adaptation.
Limitation. ChatBIT's training data (100,000 records) is tiny by Western standards. Compute access is constrained by semiconductor export controls (though enforcement is imperfect; Nvidia chips reach Russia through India and similar channels likely supply China). China's military LLM effort is real but probably 1-2 years behind the frontier. For many military applications, a 90% capable model is sufficient.
8. International governance
In 2024, the UN General Assembly voted to begin formal negotiations on a treaty governing autonomous weapons. The Convention on Certain Conventional Weapons adopted eleven non-binding guidelines but cannot agree on definitions: China considers only unstoppable systems autonomous; France includes anything that can choose its own targets. In November 2025, 156 nations supported a UN resolution calling for a legally binding treaty. Major military powers (US, Russia, Israel) resist binding commitments.
The pace of governance lags the pace of deployment by years. The US-led Political Declaration on Responsible Military Use of AI, endorsed by over 30 countries, remains voluntary. Meanwhile, the Pax Silica technology alliance (US, Israel, Japan, South Korea, Singapore, UK, Australia, UAE, Qatar) focuses on data-sharing and joint development. Technology alliances may prove more consequential than treaty efforts, as they historically have.
Sources and references
Architecture and technical analysis: Palantir blog, "Maven Smart System: Innovating for the Alliance" (March 2026) — blog.palantir.com/maven-smart-system-innovating-for-the-alliance-5ebc31709eea Striving Space, "What Is The Maven Smart System?" (March 2026) — strivingspace.com/what-is-maven-smart-system-pentagon-ai/ Escudo Digital, "Beyond Claude: the reality of Palantir's Maven" (March 2026) — escudodigital.com/en/technology/artificial-intelligence/beyond-claude-the-reality-of-palantirs-maven-in-modern-warfare.html Wikipedia, "Project Maven" (updated April 2026) — en.wikipedia.org/wiki/Project_Maven Cybershafarat, "Analysis of Maven Smart System" (March 2026) — cybershafarat.com/2026/03/16/analysis-of-ai-driven-command-and-control-maven-smart-system/
Human oversight and automation bias: Defense One, "The real danger of military AI isn't killer robots; it's worse human judgement" (March 2026) — defenseone.com/technology/2026/03/military-ai-troops-judgement/412390/ Oxford International Affairs, "Can AI behave ethically during military crises?" (January 2026) — academic.oup.com/ia/article/102/1/63/8355995 Psychology Today, "The Clash Over the Use of AI in Military Decision-Making" (March 2026) — psychologytoday.com/us/blog/psychology-through-technology/202603/the-clash-over-the-use-of-ai-in-military-decision-making ResultSense, "Military Decision-Making and Oversight Risks" (March 2026) — resultsense.com/news/2026-03-04-ai-in-warfare-us-military-speeds-decision-making-but-oversight-concerns-grow
Israel's ML targeting systems: +972 Magazine, "'Lavender': The AI machine directing Israel's bombing spree in Gaza" (April 2024) — 972mag.com/lavender-ai-israeli-army-gaza/ HRW, "Israeli Military's Use of Digital Tools in Gaza" (September 2024) — hrw.org/news/2024/09/10/questions-and-answers-israeli-militarys-use-digital-tools-gaza Lieber Institute (West Point), "The Gospel, Lavender, and the Law of Armed Conflict" (September 2024) — lieber.westpoint.edu/gospel-lavender-law-armed-conflict/ Foreign Policy, "Israel's Algorithmic Killing Sets Dangerous Precedent" (May 2024) — foreignpolicy.com/2024/05/02/israel-military-artificial-intelligence-targeting-hamas-gaza-deaths-lavender/ RUSI, "The IDF's Use of AI in Gaza: A Case of Misplaced Purpose" (July 2024) — rusi.org/explore-our-research/publications/commentary/israel-defense-forces-use-ai-gaza-case-misplaced-purpose Ynet, "How the IDF's digital brain works" (March 2026) — ynet.co.il/digital/technology/article/yokra14713296
Pentagon doctrine and budget: Hegseth memo, "Artificial Intelligence Strategy for the Department of War" (January 9, 2026) — media.defense.gov/2026/Jan/12/2003855671/-1/-1/0/ARTIFICIAL-INTELLIGENCE-STRATEGY-FOR-THE-DEPARTMENT-OF-WAR.PDF Brennan Center, "The Military's Use of AI, Explained" (March 2026) — brennancenter.org/our-work/research-reports/militarys-use-ai-explained Military.com, "Army Speeds AI Warfighting Push" (March 2026) — military.com/daily-news/headlines/2026/03/26/us-troops-active-combat-army-speeds-ai-warfighting-push.html
Ukraine and autonomous drones: IEEE Spectrum, "Rise of the Autonomous Attack Drones" (April 2026 print) — spectrum.ieee.org/autonomous-drone-warfare CSIS, "Ukraine's Vision for Autonomous Warfare" (March 2025) — csis.org/analysis/ukraines-future-vision-and-current-capabilities-waging-ai-enabled-autonomous-warfare Military Times, "Ukraine opens battlefield data to allies" (March 2026) — militarytimes.com/flashpoints/ukraine/2026/03/13/ukraine-opens-battlefield-ai-data-to-allies-in-world-first-move/ MIT Technology Review, "Autonomous warfare is unfolding in Europe" (January 2026) — technologyreview.com/2026/01/06/1129737/autonomous-warfare-europe-drones-defense-automated-kill-chains/ Lieber Institute (West Point), "The Continuing Autonomous Arms Race" (February 2025) — lieber.westpoint.edu/continuing-autonomous-arms-race/
Humanoid combat robots: TIME, "The Race to Build AI Humanoid Soldiers for War" (March 2026) — time.com/article/2026/03/09/ai-robots-soldiers-war/ Interesting Engineering, "Phantom MK-1 robots reach Ukraine" (March 2026) — interestingengineering.com/military/humanoid-soldier-robots-arrive-in-ukraine
China and military LLMs: Jamestown Foundation, "PRC Adapts Meta's Llama for Military and Security Applications" (November 2025) — jamestown.org/prcs-adaptation-of-open-source-llm-for-military-and-security-purposes/ DefenseScoop, "Pentagon report on China's progress on LLMs" (December 2025) — defensescoop.com/2025/12/26/dod-report-china-military-and-security-developments-prc-ai-llm/ Observer Research Foundation, "China's LLM Bet: Military Dominance" (July 2025) — orfonline.org/expert-speak/china-s-llm-bet-the-push-for-ai-driven-military-dominance National Defense Magazine, "China Seeking AI to Counter US Military Strengths" (March 2026) — nationaldefensemagazine.org/articles/2026/3/23/algorithmic-warfare-china-seeking-ai-to-counter-us-military-strengths CIGI, "Chinese AI Models and the Fight for AI Neutrality" (January 2026) — cigionline.org/articles/chinese-ai-models-and-the-high-stakes-fight-for-ai-neutrality/ ACM, "LLMs and Their Applications in the Military Field" (December 2025) — dl.acm.org/doi/10.1145/3773365.3773415
Broader analysis: RAND, "How AI Could Reshape Four Essential Competitions in Future Warfare" (January 2026) — rand.org/pubs/research_reports/RRA4316-1.html Jacobin, "Artificial Intelligence Is Already Making War More Horrific" (March 2026) — jacobin.com/2026/03/artificial-intelligence-us-war-military Foreign Affairs Forum, "Project Maven Explained" (March 2026) — faf.ae/home/2026/3/31/project-maven-explained Global Challenges Foundation, "Artificial Intelligence" in Global Catastrophic Risks 2026 — globalchallenges.org/gcr-2026/artificial-intelligence/