CoreWeave Smashes AI Training Records with DeepSeek-V3 in Two Minutes
Event summary
- CoreWeave trained DeepSeek-V3 671B in 2.02 minutes using 8,192 NVIDIA GB300 NVL72 GPUs, the largest cluster in MLPerf Training v6.0.
- The company demonstrated near-linear scaling efficiency, completing the same training in 3.09 minutes with 4,096 GPUs and 5.54 minutes with 2,048 GPUs.
- CoreWeave was the only submitter to scale a GB300 platform beyond 2,048 GPUs on DeepSeek-V3, showcasing full-stack optimization advantages.
- Results were achieved on production infrastructure, not a benchmark-only setup, validating real-world performance for customers.
The big picture
CoreWeave's record-breaking MLPerf results highlight the growing importance of full-stack optimization in AI training infrastructure. As frontier models reach trillion-parameter scale, the ability to efficiently scale training performance becomes a critical differentiator. CoreWeave's achievements position it as a leader in AI-native cloud services, particularly for teams under pressure to iterate quickly and bring models to production.
What we're watching
- Scaling Efficiency
- How CoreWeave's near-linear scaling efficiency will impact AI teams operating under compute budgets, potentially accelerating development cycles.
- Production Readiness
- Whether CoreWeave can maintain this performance edge as new hardware generations arrive and AI workloads become more complex.
- Competitive Positioning
- The pace at which competitors adopt similar full-stack optimization strategies to challenge CoreWeave's leadership in AI training infrastructure.
