Anthropic Raises Misalignment Risk as Claude Writes Most Production Code
Original: Anthropic Raises Misalignment Risk as Claude Writes Most Production Code View original →
A one-step increase in the risk rating
Claude now writes a large majority of the code merged into Anthropic’s production codebases, while the company has raised its assessment of high-stakes misalignment risk from “very low” to “low.” Anthropic is not saying catastrophic harm is likely. It says recent incident disclosures involving model behavior in cybersecurity evaluations increased uncertainty enough to justify the higher rating.
“Our second Risk Report is now available.” — Anthropic
The source tweet is a short notice that Anthropic published another report under its Responsible Scaling Policy. The linked 186-page document is much more consequential. Covering models and operations through July 15, 2026, it evaluates misalignment, automated research, and chemical and biological weapon risks. Anthropic says it aims to publish these reports every three to six months.
Most production code, but less than a 2× R&D speedup
The report describes Claude Mythos 5 and an unreleased “Model 2” as heavily used for coding, data generation, and persistent agent deployments. It says Claude authors a large majority of code merged into production. Anthropic nevertheless estimates that AI assistance has not yet doubled the pace of its internal AI research and engineering. Confidence in that conclusion has fallen because concrete task-based evaluations have saturated and no longer cleanly capture capability gains.
The chemical and biological section draws another important boundary. Anthropic operates as though its models meet the CB-1 threshold, meaning they may significantly help relevant actors obtain and deploy existing chemical or biological weapons. It says the models do not yet meet CB-2, the point at which they could functionally replace scarce expert knowledge needed to develop novel weapons. The company also disclosed a remediated access-control gap affecting models without blocking classifiers. It found no evidence of misuse, but said the incident reduced confidence that similar gaps do not exist.
What to watch next
The next report should show whether the “low” misalignment rating holds, whether new evaluations replace saturated benchmarks, and whether measured R&D acceleration approaches 2×. Independent evidence about covert model capabilities and the reliability of CB-1 safeguards will matter more than another qualitative assurance.
Related Articles
Anthropic published a March 6, 2026 case study showing how Claude Opus 4.6 authored a working test exploit for Firefox vulnerability CVE-2026-2796. The company presents the result as an early warning about advancing model cyber capabilities, not as proof of reliable real-world offensive automation.
New Claude models embed a machine-readable watermark that may survive copy-paste and some editing. The EU-driven rule applies worldwide, and even proofreading or translating human-written text can leave a Claude processing signal.
Election-season AI safety is moving from slogans to measurable tests. On April 24, 2026, Anthropic published Claude election metrics showing 100% and 99.8% appropriate handling on a 600-prompt misuse-and-legitimate-use set for Opus 4.7 and Sonnet 4.6, plus 90% and 94% performance in influence-operation simulations.