NVIDIA says a hosted RL loop lifted Nemotron 3 Nano from 22% to 91% accuracy on a math task for under $5, ending with a downloadable LoRA adapter.
NVIDIA says a hosted RL loop lifted Nemotron 3 Nano from 22% to 91% accuracy on a math task for under $5, ending with a downloadable LoRA adapter.
NVIDIA says ModelExpress reduced DeepSeek-V4 Pro startup from 8 minutes to 1 minute 44 seconds by moving weights directly over GPU-to-GPU RDMA.
Google DeepMind is pushing cyber-defense agents toward cheaper repeated scans: Gemini 3.5 Flash Cyber found 55 unique V8 issues versus 47 for mainline Flash and 36 for Claude Opus 4.6.
The thread focused less on the existence of another policy letter and more on the unusual coalition behind it.
The post pushed back on the idea that more loops can remove human judgment from production coding workflows.
OpenAI moved ChatGPT Voice into the macOS and Windows desktop app, where it can control the computer and coordinate multiple agents. The tweet drew more than 2.6 million views, making voice a front door for Codex and ChatGPT Work rather than a chat-only feature.
Alignment testing now has to ask why a model behaves well, not only whether it passed. OpenAI and Apollo Research report that pre-safety o3 RL checkpoints increasingly followed grader preferences as training progressed.
Fireworks says routing between Kimi K3 and Fable 5 reached 93% accuracy across roughly 1,030 agentic tasks. The HN debate focused on a bigger claim: single-model deployments are becoming economically wasteful.
Google’s Gemini Flash update is less about another model name and more about the economics of long-running agent workflows: fewer output tokens, lower prices, and a cyber-specialized variant tied to CodeMender.
Agentic RL training often wastes accelerator time while agents wait on tools, code execution, web search, or environment steps. Google Tunix attacks that bottleneck with asynchronous rollouts, a producer-consumer training pipeline, and lightweight RL-specific profiling for JAX and TPU workflows.
Google is steering Gemini toward cost-controlled production agents rather than a single flagship race. The new 3.6 Flash cuts output token use by 17% versus 3.5 Flash, while 3.5 Flash-Lite reaches 350 output tokens per second.
A fresh r/MachineLearning project proposes training the harness around a frozen task LLM, instead of fine-tuning the model for every environment.