A LocalLLaMA build with five RTX PRO 6000 cards and a 5090 made the practical cost of serious local inference hard to ignore.
LLM
RSS FeedWebKit’s new Safari Technology Preview tool gives coding agents access to DOM state, network requests, console output, and screenshots.
HN pushed the campaign because the real question is who gets to decide whether people can run capable models on their own machines.
Starting July 2, organizations that have not run inference on a fine-tuned model in the past 60 days can no longer create new fine-tuning jobs. Active existing customers lose new job creation on January 6, 2027, while inference on existing fine-tuned models continues until the base model is deprecated.
xAI’s Grok Build is now available inside Railway sandboxes, giving developers a direct terminal path through `ssh [email protected]`. The xAI post drew more than 175,000 views, while Railway’s quoted demo passed 145,000 views.
Microsoft Research turned agent skill files into trainable artifacts. SkillOpt raised GPT-5.5’s six-benchmark direct-chat average from 58.8 to 82.3 and improved all or tied for best across 52 evaluation cells without updating model weights.
Copilot now has its first selectable open-weight model. GitHub says Kimi K2.7 Code starts in VS Code for Pro tiers, with Business and Enterprise admins required to enable it by policy.
The 1.6T-total, 48B-active MoE numbers are only half the story. HN discussion treated domestic AI ASIC training as the bigger signal.
The interesting part is not just the score table. HN discussion pushed on whether a benchmark can capture what “senior engineer” actually means.
NVIDIA is testing a different route to faster LLM decoding. Nemotron-Labs-TwoTower adapts a 30B backbone into a two-tower diffusion model that keeps 98.7% of baseline quality while reaching 2.42x throughput.
Anthropic is moving stronger agentic work into its mainstream Sonnet tier. Sonnet 5 becomes the default for Free and Pro users, ships in Claude Code and the API, and starts at $2 per million input tokens and $10 per million output tokens through August 31.
HN interest centered on whether the model feels useful in real coding loops, not just on the benchmark table.