GPT-5.5 Completes Corporate Network Attack Simulation in 11 Minutes at $1.73
Original: GPT5.5 slightly outperformed Mythos on a multi-step cyber-attack simulation. One challenge that took a human expert 12 hrs took GPT-5.5 only 11 min at a $1.73 cost View original →
AISI Evaluation Results
The UK AI Safety Institute (AISI) published its cybersecurity evaluation of OpenAI GPT-5.5. The headline finding: GPT-5.5 completed a complex multi-step corporate network attack simulation in just 11 minutes at a cost of $1.73 — a task AISI estimates takes a human expert up to 12 hours.
Second Model to Cross the Threshold
In April, AISI announced that Anthropic Claude Mythos Preview was the first model to complete this benchmark end-to-end. The critical question was whether that was a single-model breakthrough or a broader trend. GPT-5.5 answers it clearly: two models from different developers have now crossed the same bar. Frontier-level AI cyber capabilities are maturing across the industry.
Evaluation Structure
AISI uses 95 cyber tasks across four difficulty tiers. Basic tasks have been fully saturated since February 2026. The advanced suite, built with cybersecurity firms Crystal Peak Security and Irregular, targets what matters most: reverse engineering stripped binaries, reliable exploits for heap overflows and UAF vulnerabilities, and full multi-step attack chains against realistic enterprise targets.
Implications
AISI is explicit that this cuts both ways. While it raises concerns about AI-assisted attacks by malicious actors, defenders can deploy the same capabilities for detection, response, and proactive hardening. The institute shared findings with OpenAI before publication. The core message: defenders must now prioritize integrating AI-based security, because the offensive baseline has permanently risen.
Related Articles
AI safety scrutiny is shifting from abstract risk debates to incident review and public technical reporting. OpenAI said on July 25 that the Hugging Face-related incident is under review with external advisers and committee oversight, drawing more than 303,000 views.
OpenAI’s February 2026 safety report says it banned accounts linked to seven operations originating in China. The company says abuse covered cyber activity, covert influence, and scams, while overall malicious use remained low versus legitimate use.
OpenAI said frontier model acceleration may eventually need deliberate pacing tools. The linked petition lists 1,224 employees from frontier AI companies and asks the U.S. government to support an international effort.