Multiverse Computing's AI Model Compression Now Runs on Intel Xeon 6 Processors
Event summary
- Multiverse Computing's CompactifAI-compressed Llama* 3.3 70B model now runs on Intel® Xeon® 6 processors with Performance-cores, utilizing vLLM* CPU and Intel® Advanced Matrix Extensions (AMX).
- The compressed model demonstrated significant performance improvements, including a 93.6% increase in output throughput and a 48.6% reduction in latency compared to the uncompressed baseline.
- The compressed model retained over 97% of the baseline accuracy across standard benchmarks, with some benchmarks showing improved accuracy due to the re-training process.
- The model size on disk is reduced by approximately 50%, from ~130 GiB to ~65 GiB, significantly reducing storage requirements and provisioning times for large-scale deployments.
The big picture
Multiverse Computing's breakthrough in running its CompactifAI-compressed models on Intel Xeon 6 processors highlights the growing trend of optimizing AI workloads for mainstream hardware. This development is strategically significant as it enables more energy-efficient and scalable AI deployments, addressing the increasing demand for efficient AI solutions across various industries.
What we're watching
- Performance Scaling
- How the performance improvements of CompactifAI-compressed models will scale across different AI workloads and concurrency levels.
- Market Adoption
- Whether enterprises in industries such as finance, healthcare, and manufacturing will widely adopt this compressed model for their AI applications.
- Competitive Dynamics
- The pace at which competitors in the AI model compression space will respond with their own optimizations for Intel Xeon 6 processors.
Related topics
