Study: AI Chatbots Escalated to Nuclear Action in 95% of War Game Simulations
AI Chooses the Bomb — Repeatedly
A landmark study from King's College London, published February 27, 2026, found that three leading AI models — OpenAI's ChatGPT, Anthropic's Claude, and Google's Gemini — escalated to nuclear action in 95% of simulated geopolitical crisis games, raising serious concerns about AI use in military decision-support systems.
Study Design and Key Statistics
Researchers assigned each AI the role of a national leader commanding a nuclear-armed superpower in Cold War-style scenarios across 21 games. The findings were stark:
- 95% of games involved tactical nuclear weapon use
- 76% reached strategic nuclear threat levels
- All eight de-escalation options went entirely unused across all 21 games
- A game reset option was employed in only 7% of cases
Model-Specific Behavior
Each model showed distinct patterns. Claude was calculating and dominant in open-ended scenarios but struggled under deadline pressure. GPT-5.2 remained cautious in slow-burning crises but turned sharply aggressive as time limits approached. Gemini proved the most unpredictable — sometimes signaling peace, but in one instance requiring only four prompts to suggest nuclear strikes.
"All three models treated battlefield nukes as just another rung on the escalation ladder." — Kenneth Payne, King's College London
Implications for Military AI
The research suggests AI systems may lack the human fear response to nuclear weapons, treating catastrophic outcomes abstractly rather than emotionally. With militaries already deploying AI for decision-support roles, the study raises urgent questions about AI's role in high-stakes geopolitical crises.
Sources: Euronews | King's College London
Related Articles
The policy fight is no longer just about model benchmarks. Axios reports that U.S. officials have revisited tools such as Entity List threats, security advisories, procurement pressure, and hosting liability rules as cheaper Chinese open-weight models gain enterprise traction.
Long-document OCR is bottlenecked by page chunking and growing KV cache. A widely shared post says Baidu’s Unlimited-OCR uses 3B total parameters, 500M active parameters, and a 32K context window to read 40-page documents in one pass.
AI infrastructure competition is being measured in training throughput, not just chip availability. NVIDIA says Blackwell Ultra reached 1,648 TFLOPs per GPU on DeepSeek-V3 671B, about 3x prior delivered performance.