Skip to content

Gemini 3.7 Flash Raises Coding Scores While Halving Launch Price

Original: Gemini 3.7 Flash Raises Coding Scores While Halving Launch Price View original →

Read in other languages: 한국어日本語
LLM Aug 15, 2026 By Insights AI (Twitter) 2 min read 1 views Source

A three-week update changes both capability and price

Gemini 3.7 Flash arrives only three weeks after 3.6 Flash, with higher coding and document-workflow scores and introductory pricing at half the earlier model’s original rate. Through the end of 2026, Google is charging $0.75 per million input tokens and $3.75 per million output tokens. The short release interval matters because Flash is positioned as a production workhorse for agents, not an occasional experimental model.

“Gemini 3.7 Flash is here. It’s stronger for coding, knowledge work, and web development.” — Google DeepMind

Google DeepMind’s account covers official Gemini research, model updates, robotics, and scientific results. The linked Google article supplies the measurable comparison. On FrontierCode 1.1 Main, 3.7 Flash scores 43.6% versus 34.4% for 3.6 Flash. DeepSWE v1.1 moves from 49.0% to 65.3%, while WebDev Arena rises from an Elo score of 1538 to 1588.

Document reasoning and automation also move

The new model scores 34.0% on GDP.pdf, a complex-document benchmark, compared with 22.0% for its predecessor. AutomationBench increases from 17.0% to 30.4%. Those gaps are more relevant to agent deployments than a generic chat comparison because debugging, issue resolution, multi-step planning, and tool calls impose real retry and supervision costs. Google also says the model follows instructions more faithfully and asks for clarification when blocked.

Developers can use 3.7 Flash through Google Antigravity, the Gemini API in Google AI Studio, and Android Studio. It is also entering Gemini Enterprise and powering Gemini Spark for Google AI Pro and Ultra subscribers across more than 160 countries. Google says the release includes updated safeguards for chemical, biological, radiological, nuclear, and cyber-offense misuse.

Watch the price after the introductory period ends, along with independent reproductions of the benchmark gains. More deliberate planning can improve completion rates while increasing latency or token use, so total task cost will matter more than the per-token headline. The source post is on X, and Google’s model overview contains the benchmark table and availability details.

Share: Long

Related Articles