Sakana AI released KAME, a tandem speech-to-speech architecture that pairs a low-latency front-end S2S model with a back-end LLM via an oracle stream, achieving MT-Bench 6.43 with near-zero response latency and eliminating the typical 2.1-second pipeline delay.
LLM
RSS FeedPoolside AI released Laguna XS.2 on April 28, 2026 under Apache 2.0 — a 33B total/3B active MoE model purpose-built for agentic coding, scoring 68.2% on SWE-bench Verified and deployable on a single consumer GPU.
Released April 29, 2026 under Modified MIT license, Mistral Medium 3.5 consolidates the company's chat, reasoning, and coding models into one 128B dense open-weight model with 256K context, scoring 77.6% on SWE-bench Verified.
Anthropic unveiled Claude Opus 4.7 and ten pre-built financial services AI agents at an invite-only Wall Street briefing on May 5, alongside full Microsoft 365 integration and a Moody's data partnership covering 600 million businesses.
llama.cpp's Multi-Token Prediction (MTP) support has entered beta, currently covering Qwen3.5 MTP. Combined with maturing tensor-parallel support, most token generation speed gaps between llama.cpp and vLLM are expected to close.
DeepClaude keeps Claude Code's complete agent loop — file editing, bash, subagent spawning — while routing API calls to DeepSeek V4 Pro or other backends, cutting output token costs from $15/M to $0.87/M.
Andrej Karpathy shared highlights from his Sequoia Ascent 2026 fireside chat, arguing that LLMs open genuinely new categories of functionality, not just faster versions of what already existed.
A local LLM researcher achieved 95.7% on SimpleQA using Qwen3.6-27B with agentic search on a single consumer GPU.
The latest ARC-AGI-3 scores show GPT-5.5 High at 0.43% and Claude Opus 4.7 at 0.18% — the most powerful models today remain effectively at zero on this AGI benchmark.
The technique GPT-5.4 Pro used to solve Erdos Problem 1196 has been applied to other problems, including another conjecture unsolved for 60 years.
AWS customers can now access OpenAI's GPT models and Codex coding agent through Amazon Bedrock, marking OpenAI's first major deployment outside Microsoft Azure. General availability is expected within weeks.
DeepSeek released DeepSeek-V4-Pro (1.6T total parameters, 49B active) and V4-Flash (284B total, 13B active), both Mixture-of-Experts models with MIT license and 1M token context. V4-Pro is the largest open-weights model released so far, and its pricing at $1.74/M input undercuts GPT-5.4 and Claude Sonnet 4.6 by more than half.