Skip to content

Blackwell Ultra reaches 1,648 TFLOPs per GPU on DeepSeek-V3

Original: NVIDIA Blackwell Ultra hits 1,648 TFLOPs per GPU on DeepSeek-V3 View original →

Read in other languages: 한국어日本語
AI Jul 22, 2026 By Insights AI (Twitter) 1 min read 1 views Source
Blackwell Ultra reaches 1,648 TFLOPs per GPU on DeepSeek-V3

NVIDIA AI is putting a hard number on Blackwell Ultra training performance: 1,648 TFLOPs per GPU on DeepSeek-V3 671B pre-training. The company says that is a world record and about three times the delivered performance of the previous generation.

The tweet highlighted “1,648 TFLOPs per GPU” and performance “about 3x” the prior generation.

The NVIDIA AI account typically posts GPU, framework, robotics, and AI infrastructure updates. This July 21 post had about 90,500 views when fetched. Its attached chart compares GB200 NVL72 and GB300 NVL72 throughput, framing the gain as a combination of Blackwell Ultra hardware and software optimization across Megatron-Core, TorchTitan, and JAX.

The workload matters. DeepSeek-V3 671B is a large mixture-of-experts model, so pre-training stresses routing, communication, memory movement, parallelism, and fault tolerance. High TFLOPs per GPU only matters if the cluster can keep thousands of accelerators fed and synchronized. That is why the benchmark is also a story about networking, orchestration, and framework maturity.

CoreWeave’s related MLPerf Training v6.0 write-up gives broader context. It says an 8,192-GPU NVIDIA Blackwell Ultra cluster completed the DeepSeek-V3 671B benchmark in 2.02 minutes, with 4,096 GPUs at 3.09 minutes and 2,048 GPUs at 5.54 minutes. CoreWeave also reported a 9.77-minute Llama-3.1-405B result on 4,096 Blackwell Ultra GPUs, 2.8x faster than its MLPerf v5.0 comparison.

The next thing to watch is how much of this benchmark performance appears in long-running customer jobs. Enterprises buying training capacity care about wall-clock time, queueing, recoverability, data loading, and cost predictability, not just peak chart numbers. If GB300 NVL72 clusters can keep this efficiency across messy production training runs, Blackwell Ultra becomes a stronger argument for renting large clusters rather than assembling smaller in-house pools. The source tweet is here.

Share: Long

Related Articles