A new DELEGATE-52 benchmark study finds that even frontier LLMs like Gemini 3.1 Pro, Claude 4.6 Opus, and GPT 5.4 corrupt an average of 25% of document content during long delegated workflows, with errors compounding silently.
#ai-safety
RSS FeedAnthropic's independent research body, The Anthropic Institute (TAI), has published its research agenda covering economic diffusion, threats and resilience, AI systems in the wild, and AI-driven R&D—including the risk of recursive AI self-improvement by 2028.
The Center for AI Standards and Innovation (CAISI) announced on May 5 that it signed national security testing agreements with Google DeepMind, Microsoft, and xAI, expanding pre-deployment frontier AI evaluations focused on cybersecurity, biosecurity, and chemical weapons risks.
Inspired by Asimov's Three Laws of Robotics, a software engineer proposes three inverse laws governing human behavior when interacting with AI — covering anthropomorphism, blind trust, and accountability.
The UK's AI Safety Institute (AISI) found that GPT-5.5 completed a multi-step corporate network attack simulation in 11 minutes at $1.73 — a task estimated to take a human expert 12 hours. It is the second model after Anthropic's Claude Mythos to reach this benchmark, confirming that advanced AI cyber capabilities are an industry-wide trend.
Election-season AI safety is moving from slogans to measurable tests. On April 24, 2026, Anthropic published Claude election metrics showing 100% and 99.8% appropriate handling on a 600-prompt misuse-and-legitimate-use set for Opus 4.7 and Sonnet 4.6, plus 90% and 94% performance in influence-operation simulations.
r/artificial pushed this study because it replaces vague AGI doom with a much more concrete threat model: swarms of AI personas that can infiltrate communities, coordinate instantly, and manufacture the appearance of consensus.
A new arXiv preprint reports that LLM judges became meaningfully more lenient when prompts framed evaluation consequences, exposing a weak point in automated safety and quality benchmarks.
OpenAI is widening access to GPT-5.4-Cyber through verified cyber-defense channels, with $10 million in API credits and government evaluation access attached. The real story is the access model: stronger cyber capability is being paired with identity checks, tiered trust, and accountability rather than a simple public release.
Synthetic-data training has a sharper safety problem than obvious bad examples. A Nature paper co-authored by Anthropic researchers reports that traits such as owl preference or misalignment can move through semantically unrelated number sequences.
Automating alignment research is moving from concept to measured experiment. Anthropic says a Claude Opus 4.6 researcher recovered 97% of the weak-to-strong supervision gap at roughly 1/100 the human time cost.
A Reddit thread pulled attention to AISI’s latest Mythos Preview evaluation, which shows a step change not just on expert CTFs but on multi-stage cyber ranges. The important claim is not generic danger rhetoric, but that Mythos became the first model to complete a 32-step corporate attack simulation end to end.