Reddit Highlights Gemini 3 Deep Think Upgrade for Science and Engineering
Original: Google upgraded Gemini-3 DeepThink: Advancing science, research and engineering View original →
How the update spread in community channels
A popular r/singularity thread amplified Google’s new Gemini 3 Deep Think announcement. At collection time, the post had a score of 675 and 51 comments. The linked primary source is Google’s official article, Gemini 3 Deep Think: Advancing science, research and engineering, which frames Deep Think as a specialized reasoning mode for difficult research and engineering workflows.
Numbers Google published
Google reports several headline metrics for the updated mode: 48.4% on Humanity’s Last Exam without tools, 84.6% on ARC-AGI-2 (noted as verified by the ARC Prize Foundation), and a Codeforces Elo of 3455. The post also states gold-medal-level performance on the 2025 International Math Olympiad. Beyond math and coding, Google adds claims for science-heavy evaluation settings, including gold-medal-level written performance for the 2025 Physics and Chemistry Olympiads and 50.5% on CMT-Benchmark.
Availability and deployment signal
According to the announcement, the upgraded Deep Think is available in the Gemini app for Google AI Ultra subscribers. Google also opened an early-access path through the Gemini API for researchers, engineers, and enterprises. The write-up includes applied examples: identifying a subtle logical flaw in a technical math paper at Rutgers and helping optimize a crystal-growth recipe at Duke University’s Wang Lab for thin films larger than 100 μm.
Why this Reddit post mattered
The thread gained traction because it combined benchmark-oriented claims with a concrete API access path. That pairing matters for teams evaluating whether frontier reasoning models can move from demo-level tasks into reproducible science and engineering pipelines. In practice, production value will depend less on headline scores alone and more on domain-specific validation, failure analysis, and tool integration quality. Still, this release is a clear signal that specialized reasoning modes are being positioned as operational infrastructure rather than one-off showcase features.
Related Articles
Cybersecurity agents are becoming a cost-per-run problem, not just a leaderboard race. Malte Ubl says GPT-5.6 Sol had the best recall and precision in a private Deepsec benchmark, but cost more than 7x the runner-up.
Google’s Gemini Flash update is less about another model name and more about the economics of long-running agent workflows: fewer output tokens, lower prices, and a cyber-specialized variant tied to CodeMender.
Google is steering Gemini toward cost-controlled production agents rather than a single flagship race. The new 3.6 Flash cuts output token use by 17% versus 3.5 Flash, while 3.5 Flash-Lite reaches 350 output tokens per second.