Agent risk has moved beyond the earlier blackmail experiments. Anthropic’s new simulations cover four failure modes: code sabotage, fraud assistance, motivated mislabeling, and coaching a human proxy.
Agent risk has moved beyond the earlier blackmail experiments. Anthropic’s new simulations cover four failure modes: code sabotage, fraud assistance, motivated mislabeling, and coaching a human proxy.
Verified U.S. K-12 teachers can get a full year of Claude access at no cost. The important shift is that Anthropic is packaging premium Claude, teaching skills, standards-aligned curriculum context, Claude Code, and Cowork into one education workflow.
Anthropic is putting CAD 10 million into Canadian AI research, with credits and partnerships spanning Amii, Mila, Vector and health institutions. The move links Claude distribution to safety, health and public-sector research.
Anthropic measured how Claude’s expressed values shift by model and language across more than 300,000 anonymized conversations. The result is a four-axis profile that could become part of model evaluation and post-release monitoring.
Anthropic says Claude contains a J-space that resembles a global workspace for active, verbalizable thoughts. The lead tweet has more than 9.1 million views and points to audit use cases, including hidden goals in sabotage-trained models.
The Future of Life Institute’s Summer 2026 AI Safety Index grades nine frontier AI companies across 37 indicators, and no firm rises above C+. The sharper point is not who leads, but how weak the ceiling remains as model capabilities and defense use expand.
Alberta put roughly 50 Claude Code agents across 466 million lines of government code and compressed a security review estimated at 6.5 years into 20 hours. The case matters because it moves coding agents from developer convenience into public-sector cyber operations.
Anthropic is not only selling Claude Science as a research workbench. It says it wants to discover treatments for neglected diseases, raising a harder question: can a frontier AI lab become both a pharma software vendor and a drug developer?
Anthropic is trying to make AI jailbreaks measurable, not just viral. Its July 2 framework separates minor bypasses from universal failures, adds a HackerOne path for Fable 5 reports, and says one new classifier blocks the Amazon-reported technique in over 99% of cases.
Anthropic is moving AI-for-science support from chat into reproducible work sessions. Claude Science combines 60-plus scientific skills and connectors, reviewer agents, HPC or SSH workflows, and up to $30,000 in credits for as many as 50 projects.
Anthropic is moving stronger agentic work into its mainstream Sonnet tier. Sonnet 5 becomes the default for Free and Pro users, ships in Claude Code and the API, and starts at $2 per million input tokens and $10 per million output tokens through August 31.
AI model availability is being shaped by export-control decisions, not only product readiness. Anthropic said it received notice that the U.S. Department of Commerce lifted export controls on Claude Fable 5 and Mythos 5, with access restoration starting the next day.