ARC-AGI-3 Benchmarks: GPT-5.5 at 0.43%, Claude Opus 4.7 at 0.18%
Original: ARC-AGI-3 Update (GPT-5.5 High and Opus4.7) View original →
The Numbers
An r/singularity update (354 points) reports the latest ARC-AGI-3 results: GPT-5.5 High 0.43%, Claude Opus 4.7 0.18%.
What Is ARC-AGI-3?
ARC-AGI-3 is the third ARC Prize benchmark, significantly harder than ARC-AGI-2. It tests genuine reasoning that humans perform easily but current AI models struggle with.
Why It Matters
The most capable models ever built are functionally at zero on a test any person would pass. ARC-AGI-3 remains one of the clearest indicators of the gap between today AI and genuine general intelligence.
Related Articles
GPT-5.6 Sol moved from 13.3% to 38.3% on ARC-AGI-3 when OpenAI retained reasoning and used compaction in the harness. The result makes benchmark setup, not just model weights, part of the frontier-agent story.
The new Claude default for high-end daily work shifts the model race toward performance per dollar. Anthropic says Opus 5 approaches Claude Fable 5 on coding and knowledge work while keeping API pricing at $5/M input and $25/M output tokens.
Claude Opus 4.6 achieved a 50%-time-horizon of approximately 14.5 hours on METR's software task benchmark — beating all predictions and suggesting a doubling time of under 3 months for AI task capabilities.