Coding-agent evals are moving beyond pass@1. Together AI says Kimi K3 came close to Fable 5 xhigh on DeepSWE while costing about one-third as much per rollout.
#benchmarks
RSS FeedThe 235-comment HN thread focused less on whether reasoning models can solve hard tasks and more on whether their reasoning traces explain why they succeeded.
The new Claude default for high-end daily work shifts the model race toward performance per dollar. Anthropic says Opus 5 approaches Claude Fable 5 on coding and knowledge work while keeping API pricing at $5/M input and $25/M output tokens.
Cybersecurity agents are becoming a cost-per-run problem, not just a leaderboard race. Malte Ubl says GPT-5.6 Sol had the best recall and precision in a private Deepsec benchmark, but cost more than 7x the runner-up.
OpenAI is trying to move enterprise AI measurement from token cost to cost per successful task. It says GPT-5.6 Sol reached 72.7% on DeepSWE v1.1, above Claude Fable 5’s 69.9%, while carrying 36.2% lower estimated API cost.
OpenAI says SWE-Bench Pro no longer reliably measures frontier coding capability after finding 30% of its public tasks broken. The cited issues include hidden requirements, contradictory instructions, strict tests and incomplete grading criteria.
GPT-5.6 moved from preview into access across ChatGPT, Codex and the OpenAI API. OpenAI paired the rollout with an 80.0 Coding Agent Index score, 2.8 points above Claude Fable 5, while claiming lower token use, time and cost.
Microsoft Research turned agent skill files into trainable artifacts. SkillOpt raised GPT-5.5’s six-benchmark direct-chat average from 58.8 to 82.3 and improved all or tied for best across 52 evaluation cells without updating model weights.
Biology agents are being judged on research judgment, not just factual answers. GeneBench-Pro puts 129 computational-biology problems in front of agents, and indexed coverage says GPT-5.6 Sol reaches 28.7% at the highest reasoning level and 31.5% in Pro mode.
HN interest centered on whether the model feels useful in real coding loops, not just on the benchmark table.
Arena says its commercial AI evaluation service has reached a $100M annualized run rate just eight months after launch. The milestone shows how crowdsourced model preferences are becoming paid infrastructure for labs and enterprises.
OpenRouter says it continuously runs GPQA and TAU-Bench on open-weight models and feeds the results into AutoExacto routing. The linked GLM 5.2 page pairs benchmark rankings with production details such as a 1M-token context window and $0.94/$3 per 1M token pricing.