A fresh r/MachineLearning project proposes training the harness around a frozen task LLM, instead of fine-tuning the model for every environment.
LLM
RSS FeedThe HN discussion around Claude Code’s bundled Bun focused less on speed and more on what Anthropic’s ownership means for a developer runtime.
Open coding agents are maturing around harnesses, protocols, and SDK compatibility. OpenInterpreter says its Kimi K3 native harness is written in Rust, Apache licensed, and compatible with ACP and the Codex SDK.
Cybersecurity agents are becoming a cost-per-run problem, not just a leaderboard race. Malte Ubl says GPT-5.6 Sol had the best recall and precision in a private Deepsec benchmark, but cost more than 7x the runner-up.
A 27B model running on phones would shift the boundary for private, offline AI. RunAnywhere says Bonsai uses 1-bit weights, fits in 3.9GB, and keeps about 90% of full-precision quality in its own evals.
HN’s interest landed on the tradeoff Bionic represents: local models, cloud fallback, coding workflows, and a closed-source desktop app all in one package.
HN treated Grok Build less as a feature drop and more as a test of control: an open Rust coding agent is useful, but telemetry, forks, and provider lock-in shaped the discussion.
HN focused less on the launch framing and more on the pressure Kimi K3 puts on model economics: a 2.8T open model with a 1M-token context is expensive, capable, and hard to ignore.
OpenAI is trying to move enterprise AI measurement from token cost to cost per successful task. It says GPT-5.6 Sol reached 72.7% on DeepSWE v1.1, above Claude Fable 5’s 69.9%, while carrying 36.2% lower estimated API cost.
NVIDIA says Nemotron 3 Embed now leads LMEB, with the 8B model ranked first and the 1B model second. The linked Hugging Face discussion cites LMEB scores of 64.4 for 8B and 61.5 for 1B BF16, extending the release beyond its earlier RTEB win.
OpenAI tied GPT-5.6 Sol’s new “The Last Ones” cyber-range result to Codex Security, a plugin meant to find, validate, and fix vulnerabilities in real repositories. The important comparison is controlled benchmark success versus code review work that security teams can actually run.
Kimi K3 raises the open-weight scale race to 2.8T parameters with a 1M-token context window and native vision. Full weights are scheduled by July 27, 2026, while the model is already available through Kimi.com, Kimi Code, and the API.