Google DeepMind is pushing cyber-defense agents toward cheaper repeated scans: Gemini 3.5 Flash Cyber found 55 unique V8 issues versus 47 for mainline Flash and 36 for Claude Opus 4.6.
#cybersecurity
RSS FeedGoogle’s Gemini Flash update is less about another model name and more about the economics of long-running agent workflows: fewer output tokens, lower prices, and a cyber-specialized variant tied to CodeMender.
A security incident tied to model evaluation drew unusually intense HN debate. The real issue is not only the breach, but how far cyber benchmarks can safely push models against realistic infrastructure.
Google is steering Gemini toward cost-controlled production agents rather than a single flagship race. The new 3.6 Flash cuts output token use by 17% versus 3.5 Flash, while 3.5 Flash-Lite reaches 350 output tokens per second.
Cybersecurity agents are becoming a cost-per-run problem, not just a leaderboard race. Malte Ubl says GPT-5.6 Sol had the best recall and precision in a private Deepsec benchmark, but cost more than 7x the runner-up.
Alberta put roughly 50 Claude Code agents across 466 million lines of government code and compressed a security review estimated at 6.5 years into 20 hours. The case matters because it moves coding agents from developer convenience into public-sector cyber operations.
Anthropic is trying to make AI jailbreaks measurable, not just viral. Its July 2 framework separates minor bypasses from universal failures, adds a HackerOne path for Fable 5 reports, and says one new classifier blocks the Amazon-reported technique in over 99% of cases.
OpenAI’s GPT-5.6 preview is as much about release control as model capability. Sol claims Terminal-Bench 2.1 SOTA, competitive ExploitBench results using about one-third the output tokens of Mythos Preview, and first access limited to trusted partners shared with the U.S. government.
GPT-5.5-Cyber reached 85.6% on CyberGym, while Codex Security has scanned more than 30,000 codebases. The real shift is from vulnerability discovery to validated remediation.
The bottleneck in AI security is shifting from finding bugs to landing fixes. OpenAI says GPT-5.5-Cyber reached 85.6% on CyberGym, while Codex Security has scanned more than 30,000 codebases.
Five Eyes cyber agencies warned that frontier AI could reshape offensive and defensive cyber capabilities within months. The warning turns AI security from a technical concern into a board-level continuity and market-confidence risk.
AI-enabled attacks are shifting from setup work into post-compromise operations. Anthropic mapped 832 malicious accounts to MITRE ATT&CK and found medium-or-higher risk actors rising from 33% to 56%.