CoreWeave Smashes AI Training Records with DeepSeek-V3 in Two Minutes

  • CoreWeave trained DeepSeek-V3 671B in 2.02 minutes using 8,192 NVIDIA GB300 NVL72 GPUs, the largest cluster in MLPerf Training v6.0.
  • The company demonstrated near-linear scaling efficiency, completing the same training in 3.09 minutes with 4,096 GPUs and 5.54 minutes with 2,048 GPUs.
  • CoreWeave was the only submitter to scale a GB300 platform beyond 2,048 GPUs on DeepSeek-V3, showcasing full-stack optimization advantages.
  • Results were achieved on production infrastructure, not a benchmark-only setup, validating real-world performance for customers.

CoreWeave's record-breaking MLPerf results highlight the growing importance of full-stack optimization in AI training infrastructure. As frontier models reach trillion-parameter scale, the ability to efficiently scale training performance becomes a critical differentiator. CoreWeave's achievements position it as a leader in AI-native cloud services, particularly for teams under pressure to iterate quickly and bring models to production.

Scaling Efficiency
How CoreWeave's near-linear scaling efficiency will impact AI teams operating under compute budgets, potentially accelerating development cycles.
Production Readiness
Whether CoreWeave can maintain this performance edge as new hardware generations arrive and AI workloads become more complex.
Competitive Positioning
The pace at which competitors adopt similar full-stack optimization strategies to challenge CoreWeave's leadership in AI training infrastructure.