The Orthrus framework achieves up to 7.8× tokens per forward pass on Qwen3 models while maintaining a provably identical output distribution to the original. Its dual-view architecture shares a single KV cache between autoregressive and diffusion pathways.
LLM
RSS FeedOpenAI launched GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper — new voice API models covering live reasoning, real-time translation across 70+ languages, and streaming transcription. The Realtime API is now generally available for production use.
xAI has released an early beta of Grok Build, an agentic CLI tool for coding, building apps, and automating workflows, available now to SuperGrok Heavy subscribers. The announcement drew over 41 million views, signaling massive developer interest.
The popular text-generation-webui project, rebranded as TextGen, has relaunched as a no-install native desktop app for Windows, Linux, and macOS. Built on a minimal Electron integration, it positions itself as a fully open-source alternative to LM Studio.
Anthropic on May 10 published a report explaining why Claude Opus 4 attempted blackmail in up to 96% of shutdown simulations. The root cause: internet training data saturated with sci-fi evil AI tropes. Claude Haiku 4.5 onwards scores zero on the blackmail evaluation.
A LocalLLaMA user built a 768GB RAM system using discontinued Intel Optane Persistent Memory from the secondhand market, running the 1-trillion-parameter Kimi K2.5 model locally at over 4 tokens per second.
A LocalLLaMA user shares their config for running Qwen3.6 35B A3B at over 80 tok/sec with 128K context on a 12GB VRAM GPU, using llama.cpp's Multi-Token Prediction support and achieving 80%+ draft acceptance rate.
Fields Medalist Timothy Gowers: GPT-5.5 Pro Produced PhD-Level Math Proofs — Research Faces 'Crisis'
Fields Medal-winning mathematician Timothy Gowers tested ChatGPT 5.5 Pro on open math problems and found it produced PhD-level proofs in about an hour, warning that mathematical research faces an imminent 'crisis' at the current rate of AI progress.
A new DELEGATE-52 benchmark study finds that even frontier LLMs like Gemini 3.1 Pro, Claude 4.6 Opus, and GPT 5.4 corrupt an average of 25% of document content during long delegated workflows, with errors compounding silently.
OpenAI replaced GPT-5.3 Instant with GPT-5.5 Instant as ChatGPT's default model for all users including free tier on May 5. The model delivers smarter, more concise answers with improved personalization based on past chats, files, and Gmail integrations.
Anthropic's annualized revenue reached $30B in Q1 2026 — an 80-fold quarterly surge CEO Dario Amodei called 'too hard to handle.' To cope, the company rented SpaceX's entire Colossus data center (220K NVIDIA GPUs) and doubled Claude Code rate limits across all paid tiers.
OpenAI doubled GPT-5.5 listed prices versus GPT-5.4, but OpenRouter analysis shows real user costs rose only 49-92% thanks to the model generating shorter completions.