Grok 4.6 Scores 61, Matching GPT-5.6 Sol at $0.84 per Task
Original: Grok 4.6 Scores 61, Matching GPT-5.6 Sol at $0.84 per Task View original →
A 61 score at $0.84 per task
Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol at its maximum setting. That puts it only one point behind Claude Fable 5 with fallback and two points behind Claude Opus 5 at maximum effort. The more consequential number is cost: Artificial Analysis measured $0.84 per task, while SpaceXAI prices input at $2 and output at $6 per million tokens. Frontier model selection is increasingly about the cost of completing useful work, not a single leaderboard position.
“Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol.”
That is the central finding in the Artificial Analysis post, published at 15:39 UTC on August 12, 2026. The evaluator says the model gained five index points over Grok 4.5 in just over a month and 23 points over Grok 4.3. Minutes earlier, the official SpaceXAI launch post said the new version improves on Grok 4.5 without raising its headline price.
Agent results carry more weight than the aggregate
The breakdown is more useful for buyers than the index alone. Grok 4.6 reached 1577 Elo on AA-Briefcase, Artificial Analysis’s private test of long-horizon knowledge work, placing it in the Claude Fable 5 tier. It completed those tasks in roughly 53 turns and 0.5 billion input tokens on average. Claude Opus 5 at maximum effort used about 103 turns and 2.0 billion input tokens, according to the same post.
Grok 4.6 also scored 50.7% on τ³-Banking, close to Qwen3.8 Max at 51.3%, and 88.4% on Terminal-Bench v2.1. Its 500,000-token context window is unchanged from Grok 4.5. Cache-hit input costs, however, rose from $0.30 to $0.50 per million tokens, a detail that could matter for repetitive agent workloads.
Artificial Analysis regularly publishes independent comparisons of model quality, speed, price and agent performance. Its post adds comparable execution data to SpaceXAI’s brief release claim. Still, one composite score cannot reveal every failure mode, and the private AA-Briefcase set cannot be reproduced directly by outside researchers.
What to watch next
The next test is whether the low turn count holds in production across coding, browser use and long-running tool calls. Independent evaluations should also examine stability near the 500K context limit and calculate full workflow cost after cache usage. If Grok 4.6 preserves its measured efficiency outside this suite, the frontier could be judged less by peak score and more by reliable tasks completed per dollar.
Source: Artificial Analysis on X
Related Articles
HN’s 718-point discussion focused less on the leaderboard slot and more on what cheap, fast reasoning changes for daily agent work.
The practical question drew the attention: can a 30B agentic model live on one consumer GPU? Meta combines Apache 2.0 weights, roughly 4-bit quantization, and DFlash speculative decoding to make that case.
Open-weight LLMs are moving from cost comparisons into production agent design. OpenRouter singled out four June 2026 models, including DeepSeek V4 Flash at 79.0% on SWE-bench Verified and GLM 5.2 as the top open model on Artificial Analysis v4.1.