Skip to content

Kimi K3 Opens 2.8T Weights and Resets Open-Model Stakes

Original: Kimi K3 Opens 2.8T Weights With 1M-Token Context View original →

Read in other languages: 한국어日本語
LLM Jul 28, 2026 By Insights AI (Twitter) 1 min read 1 views Source
Kimi K3 Opens 2.8T Weights and Resets Open-Model Stakes

Frontier Scale Moves Into Open Weights

Kimi K3 turns the open-model debate from a policy argument into a deployment question. Moonshot AI’s Kimi account posted on July 27, 2026 that the model weights and technical report were available, and the post drew roughly 9.2 million views, 40,000 likes, and more than 6,000 reposts within the freshness window used for this crawl.

"Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window."

The Hugging Face model card describes Kimi K3 as a 2.8T-parameter mixture-of-experts model with 104B activated parameters per token. It uses 896 experts, selects 16 per token, and supports a 1,048,576-token context length. The card also frames it as a native multimodal model for text and image tasks, rather than a text-only checkpoint with peripheral tooling.

The architectural claim is the sharper signal. Moonshot says Kimi Delta Attention, Attention Residuals, and Stable LatentMoE produce about 2.5x better scaling efficiency over Kimi K2. Alongside the model, the company points to high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale. That makes the release useful not only for benchmark watchers but also for teams assessing whether open frontier models can be served, audited, and adapted in production.

The next things to watch are license constraints, verified serving stacks, and the first independent benchmark replications. Open weights do not make a 2.8T model cheap to run, but they do change who can inspect and adapt the system. Source tweet: Kimi_Moonshot on X.

Share: Long

Related Articles