Skip to content

Grok 4.6 Scores 61, Matching GPT-5.6 Sol at $0.84 per Task

Original: Grok 4.6 Scores 61, Matching GPT-5.6 Sol at $0.84 per Task View original →

Read in other languages: 한국어日本語
LLM Aug 14, 2026 By Insights AI (Twitter) 2 min read 1 views Source
Grok 4.6 Scores 61, Matching GPT-5.6 Sol at $0.84 per Task

A 61 score at $0.84 per task

Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol at its maximum setting. That puts it only one point behind Claude Fable 5 with fallback and two points behind Claude Opus 5 at maximum effort. The more consequential number is cost: Artificial Analysis measured $0.84 per task, while SpaceXAI prices input at $2 and output at $6 per million tokens. Frontier model selection is increasingly about the cost of completing useful work, not a single leaderboard position.

“Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol.”

That is the central finding in the Artificial Analysis post, published at 15:39 UTC on August 12, 2026. The evaluator says the model gained five index points over Grok 4.5 in just over a month and 23 points over Grok 4.3. Minutes earlier, the official SpaceXAI launch post said the new version improves on Grok 4.5 without raising its headline price.

Agent results carry more weight than the aggregate

The breakdown is more useful for buyers than the index alone. Grok 4.6 reached 1577 Elo on AA-Briefcase, Artificial Analysis’s private test of long-horizon knowledge work, placing it in the Claude Fable 5 tier. It completed those tasks in roughly 53 turns and 0.5 billion input tokens on average. Claude Opus 5 at maximum effort used about 103 turns and 2.0 billion input tokens, according to the same post.

Grok 4.6 also scored 50.7% on τ³-Banking, close to Qwen3.8 Max at 51.3%, and 88.4% on Terminal-Bench v2.1. Its 500,000-token context window is unchanged from Grok 4.5. Cache-hit input costs, however, rose from $0.30 to $0.50 per million tokens, a detail that could matter for repetitive agent workloads.

Artificial Analysis regularly publishes independent comparisons of model quality, speed, price and agent performance. Its post adds comparable execution data to SpaceXAI’s brief release claim. Still, one composite score cannot reveal every failure mode, and the private AA-Briefcase set cannot be reproduced directly by outside researchers.

What to watch next

The next test is whether the low turn count holds in production across coding, browser use and long-running tool calls. Independent evaluations should also examine stability near the 500K context limit and calculate full workflow cost after cache usage. If Grok 4.6 preserves its measured efficiency outside this suite, the frontier could be judged less by peak score and more by reliable tasks completed per dollar.

Source: Artificial Analysis on X

Share: Long

Related Articles