Gemini 3.7 Flash Raises Coding Scores While Halving Launch Price
Original: Gemini 3.7 Flash Raises Coding Scores While Halving Launch Price View original →
A three-week update changes both capability and price
Gemini 3.7 Flash arrives only three weeks after 3.6 Flash, with higher coding and document-workflow scores and introductory pricing at half the earlier model’s original rate. Through the end of 2026, Google is charging $0.75 per million input tokens and $3.75 per million output tokens. The short release interval matters because Flash is positioned as a production workhorse for agents, not an occasional experimental model.
“Gemini 3.7 Flash is here. It’s stronger for coding, knowledge work, and web development.” — Google DeepMind
Google DeepMind’s account covers official Gemini research, model updates, robotics, and scientific results. The linked Google article supplies the measurable comparison. On FrontierCode 1.1 Main, 3.7 Flash scores 43.6% versus 34.4% for 3.6 Flash. DeepSWE v1.1 moves from 49.0% to 65.3%, while WebDev Arena rises from an Elo score of 1538 to 1588.
Document reasoning and automation also move
The new model scores 34.0% on GDP.pdf, a complex-document benchmark, compared with 22.0% for its predecessor. AutomationBench increases from 17.0% to 30.4%. Those gaps are more relevant to agent deployments than a generic chat comparison because debugging, issue resolution, multi-step planning, and tool calls impose real retry and supervision costs. Google also says the model follows instructions more faithfully and asks for clarification when blocked.
Developers can use 3.7 Flash through Google Antigravity, the Gemini API in Google AI Studio, and Android Studio. It is also entering Gemini Enterprise and powering Gemini Spark for Google AI Pro and Ultra subscribers across more than 160 countries. Google says the release includes updated safeguards for chemical, biological, radiological, nuclear, and cyber-offense misuse.
Watch the price after the introductory period ends, along with independent reproductions of the benchmark gains. More deliberate planning can improve completion rates while increasing latency or token use, so total task cost will matter more than the per-token headline. The source post is on X, and Google’s model overview contains the benchmark table and availability details.
Related Articles
Google says Gemini in Google Sheets reached 70.48% on the full SpreadsheetBench benchmark, approaching human expert ability. The company attributes the result to product-specific tuning plus stronger verbalization and coding behavior inside Sheets.
The 936-point discussion focused less on a routine model refresh than on a three-week release cycle and introductory pricing at half the original 3.6 Flash cost. The published gains span coding, document reasoning, and workflow automation, but production teams still need to test total task cost.
Grok 4.6 reached 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol. At $0.84 per evaluated task, the result shifts the frontier contest toward agent efficiency and total cost, not just headline scores.