A US military operation was minutes from boarding a Chinese vessel — armed, with air support — based on intelligence that was entirely fabricated by a machine.
According to a CNN report citing four sources familiar with the episode, a US Special Operations Command analyst submitted intelligence claiming a Chinese ship was moving nuclear arms program components through the Middle East. The US military was actively preparing to intercept that vessel. What stopped them was the discovery that a chatbot used to generate the report had invented the core finding. One source described the near-miss plainly: it "almost started a war."
That sentence deserves to sit for a moment before the analysis begins.
The Incident: How an AI Chatbot Nearly Triggered a Military Confrontation
The mechanics of the near-incident follow a pattern that AI researchers have warned about for years. An analyst — working within Special Operations Command, one of the most operationally active elements of the US military — used AI tools to assist in producing an intelligence report. The resulting document described a Chinese ship transporting material linked to a nuclear arms program as it transited the Middle East. Based on that report, the military was on track to intercept and board the vessel, with supporting air assets in position.
At some point before action was taken, officials determined the chatbot had "inaccurately identified the material the ship was carrying." The intelligence was, in the words of those familiar with the episode, "entirely false."
What nearly happened here was not a cyberattack, not a miscommunication between commanders, not a technical malfunction in weapons systems. It was a language model generating plausible-sounding, structurally coherent text that bore no relationship to reality — and that text propagating through an institutional chain until it almost became kinetic military action against a vessel from a nuclear-armed nation.
What Is AI Hallucination and Why Does It Happen
The term "hallucination" in AI refers specifically to the tendency of large language models to produce confident, fluent, and entirely fabricated outputs. This is not a bug in the traditional software sense. It is an emergent property of how these systems are built.
Read next Laika's Wildwood: Stop-Motion Fantasy at TIFF 2026Large language models are trained to predict statistically likely sequences of tokens given a prompt. They do not retrieve information from a verified database. They do not reason about ground truth. When asked a question at the edge of their training data — or asked to synthesize intelligence from ambiguous inputs — they generate text that looks correct because it follows the structural and stylistic patterns of correct text. The confidence of the output is decoupled from its accuracy.
Benchmarking studies across multiple research institutions have found that frontier LLMs hallucinate on factual tasks at rates ranging from roughly 3 percent in well-constrained domains to more than 20 percent in open-ended, ambiguous, or out-of-distribution queries. Intelligence analysis is almost by definition the latter category — sparse data, adversarial information environments, high uncertainty. This is precisely where AI hallucination rates climb.
The problem is compounded by what researchers call "automation bias": the documented human tendency to over-trust outputs from systems perceived as authoritative or technical. An analyst reviewing a cleanly formatted, confidently phrased AI-generated intelligence report faces real cognitive pressure to accept it.
The Dangers of AI in Military Intelligence Operations
The AI hallucination military risk is not hypothetical anymore. It materialized. And it materialized in one of the most consequential possible contexts — a potential confrontation with China over suspected nuclear proliferation.
Paul Scharre, a former Army Ranger and now vice president at the Center for New American Security, has written extensively on the risks of autonomous and AI-assisted military systems. His core argument — that speed and automation in warfare create dangerous compression of the decision cycle — applies directly here. When AI tools accelerate the production of intelligence products, they can compress the time available to apply human skepticism. The report went from generation to near-action faster than the error-checking mechanisms could catch up.
There is also a structural problem in how intelligence workflows treat sources. Finished intelligence reports carry institutional authority. An analyst's report, submitted through proper channels, is treated as a product of human judgment with traceable sourcing. When a chatbot is part of that pipeline but that fact is not clearly flagged — or when the verification steps appropriate for AI-generated content are not applied — the institutional trust attaches to a product that has not earned it.
The stakes in military intelligence are categorically different from, say, a chatbot giving a user bad restaurant recommendations. An error in operational intelligence can move troops, authorize force, trigger diplomatic crises, or in the near-miss described here, place armed personnel on the deck of a foreign nation's ship.
US Special Operations Command and AI Tool Adoption
Special Operations Command has been among the more aggressive adopters of AI and data analytics tools within the US military. This reflects a broader posture: the Department of Defense has invested billions in AI development since at least 2018, when the Joint Artificial Intelligence Center was established. That organization was later absorbed into the Chief Digital and AI Office, or CDAO, which now serves as the primary coordinator for AI integration across the department.
The institutional appetite for AI tools in intelligence analysis is understandable. The volume of signals, imagery, and open-source data available to modern intelligence agencies is staggering — far beyond what human analysts can manually process. AI tools offer genuine leverage on that problem. Pattern recognition in satellite imagery, translation of foreign-language communications, entity extraction from large document sets: these are real capabilities with demonstrated value.
The problem is the gap between what these tools do well and what analysts may believe they do well. Classifying objects in satellite images is a constrained, supervised task. Synthesizing open-ended intelligence assessments from ambiguous multi-source inputs is not. Using the same category of tool — "AI" — for both problems without distinguishing their reliability profiles is institutionally dangerous.
Policy Implications: Regulating AI in Defense and Intelligence
The Department of Defense adopted five AI ethics principles in February 2020: responsible, equitable, traceable, reliable, and governable. The principle of traceability holds that the department must be able to audit AI systems and understand how they reach their outputs. The principle of reliability requires that AI systems perform within specified ranges and produce consistent results.
An AI tool that generates an "entirely false" intelligence report documenting nuclear proliferation activities fails both principles — and the existing framework offers no obvious mechanism to have prevented it.
The National Security Commission on Artificial Intelligence, chaired by former Google CEO Eric Schmidt, published its final report in 2021 warning explicitly about the risks of deploying AI in high-stakes decision environments without adequate verification infrastructure. That report recommended mandatory human-in-the-loop requirements for certain classes of AI-assisted decisions. The incident described by CNN suggests those requirements either were not in place, were not enforced, or were insufficient to catch the error before operational action was nearly taken.
The gap between published DoD principles and operational reality is where the policy work needs to happen now.
What This Means for the Future of Military AI
The near-boarding of a Chinese ship over a hallucinated AI report is not an argument against using AI in defense and intelligence. It is an argument for using it correctly — which means being specific about what these systems can and cannot do, building verification workflows that match the stakes of the decisions involved, and treating AI-generated intelligence products with the same source-skepticism applied to any single-source report.
Short sentences matter here: AI makes mistakes. Those mistakes can be consequential. The workflow around AI must account for that.
The harder institutional problem is cultural. Speed is rewarded in operational contexts. Skepticism introduces friction. AI tools that produce fluent, confident outputs create pressure to accept them. Changing that dynamic requires explicit policy mandates, not just principles on paper — mandatory flagging of AI-assisted products, required adversarial review before action thresholds are crossed, and clear accountability when those steps are skipped.
The AI hallucination military problem will not be solved by better models alone. GPT-series successors and frontier multimodal systems will continue to hallucinate at some rate in open-ended domains. The question is whether the institutional infrastructure around those tools is designed with that reality in mind. Based on what nearly happened in the Middle East, the answer right now appears to be no.
Source: Ars Technica - All content



