- 9X speedup in time-to-first-token (TTFT) demonstrated by VAST-AMD solution
- 9.7X increase in token throughput for high-concurrency workloads
- 20-50GB per user potential KV cache size, a critical bottleneck in AI inference
Experts would likely conclude that the VAST and AMD alliance presents a significant advancement in AI inference infrastructure, offering scalable, cost-effective solutions to overcome key technical bottlenecks while promoting an open ecosystem approach.
VAST and AMD Forge Alliance to Tackle AI’s Next Big Hurdle: Inference
NEW YORK CITY & SAN FRANCISCO, CALIF. – July 24, 2026 – As the artificial intelligence gold rush moves beyond model training and into the complex world of real-world deployment, a critical new bottleneck is emerging: inference. In a significant move to address this, VAST Data and AMD today announced an expanded collaboration aimed at redefining the infrastructure for AI factories, promising a more open, efficient, and cost-effective alternative to the market's dominant players.
The partnership integrates VAST’s AI Operating System with AMD’s upcoming 6th Gen EPYC™ CPUs and powerful AMD Instinct™ GPUs. The goal is to create a unified, high-performance platform that streamlines the deployment of large-scale inference services and the next wave of "agentic AI" applications. This alliance directly challenges the notion that AI infrastructure must be a closed, single-vendor ecosystem, offering enterprises and cloud providers a flexible foundation built for the operational realities of AI.
"AI is entering an operational phase where infrastructure efficiency matters as much as model performance," said John Mao, Vice President, Global Technology Alliances at VAST Data. "The industry is discovering that inference is fundamentally a data problem."
The Data Problem in the Inference Era
While the industry has spent years optimizing the massive, compute-heavy task of training AI models, the operational phase of AI—inference—presents a different and arguably more complex set of challenges. Inference, where a trained model makes predictions on new data, is becoming the dominant workload, especially with the rise of conversational AI, Retrieval-Augmented Generation (RAG), and multi-turn AI agents.
These applications generate enormous context windows and session states, leading to a critical issue with what is known as the KV cache. This cache stores intermediate calculations, allowing the model to "remember" the context of a conversation without recomputing everything from scratch. As context windows grow, the KV cache can swell to 20-50GB per user, quickly overwhelming the expensive, high-bandwidth memory (HBM) built directly onto GPUs. When HBM is full, the system must either discard the cache and suffer the performance penalty of re-computation or find a way to offload it.
This is where the VAST and AMD collaboration introduces its most compelling innovation. By offloading the KV cache from the GPU's limited memory to VAST’s high-performance, NVMe-based storage platform, the system can manage virtually unlimited context sizes. Early tests conducted by VAST are striking: using an AMD Instinct MI355X GPU, the integrated solution demonstrated a 9X speedup in time-to-first-token (TTFT) and a 9.7X increase in token throughput for high-concurrency workloads. This isn't just a marginal improvement; it's a potential game-changer for deploying responsive, long-running AI agents at a manageable cost.
“Agentic AI moves the inference bottleneck beyond raw compute to context," noted Pin Siang Tan, CTO of Embedded LLM, a software partner in the ecosystem. “The innovations AMD and VAST are bringing to inference infrastructure help address a critical bottleneck for organizations deploying agentic AI at scale.”
Building the Open AI Factory
The partnership's strategic core is its "open ecosystem" approach, a direct counterpoint to the vertically integrated, proprietary stacks that have characterized the first wave of the AI boom. By combining best-in-class components from multiple vendors, the alliance aims to offer customers greater flexibility and avoid vendor lock-in.
“The future of AI will be built on an open ecosystem that gives organizations the flexibility to choose the technologies that best meet their performance, operational and business requirements,” stated Derek Dicker, Corporate Vice President, Enterprise Business Group at AMD.
At the heart of this new stack is a trifecta of powerful technologies:
* AMD’s 6th Gen EPYC CPUs: Codenamed "Venice," these next-generation processors will power VAST’s new server platforms. Crucially, they are among the first server CPUs to support PCIe® Gen-6, an interconnect standard that doubles I/O bandwidth. This leap is vital for rapidly moving massive datasets between storage and compute, slashing latency for the data-intensive services that underpin modern AI.
* VAST’s DASE Architecture: VAST’s Disaggregated Shared Everything (DASE) architecture is the software-defined foundation that ties everything together. It separates compute logic from storage media, allowing all compute nodes to access a global data pool in parallel. This enables independent, linear scaling of performance and capacity, a key factor in managing the unpredictable and explosive growth of AI data while controlling costs.
* A Validated Reference Architecture: To simplify deployment, VAST, AMD, and networking specialist DriveNets have co-developed a reference architecture for AI factories. This blueprint combines AMD Helios rack-scale infrastructure with the VAST AI OS and DriveNets AI Fabric networking, providing enterprises with a validated, predictable path to building high-performance AI infrastructure.
This collaborative model is already gaining traction with AI cloud providers. "Our infrastructure strategy is built on a silicon agnostic approach to give customers the best optionality," said Raghu Chakravarthi, EVP of Engineering and General Manager – Americas at Core42. Similarly, Kevin Cochrane, CMO at Vultr, highlighted the customer demand for flexibility: "The collaboration between AMD and VAST gives customers more flexibility in how and where they deploy AI."
Lowering Costs and Simplifying Compliance
Beyond raw performance, the partnership delivers tangible business benefits aimed at lowering the total cost of ownership (TCO) and reducing operational friction. For many organizations, the journey from AI pilot to production is fraught with complexity, hidden costs, and compliance hurdles.
By unifying storage, database, and compute services on a single platform, the VAST AI OS eliminates the need for IT teams to stitch together and manage a patchwork of disparate systems. This simplification extends to critical governance functions. The platform features Automated KV Cache Lifecycle Management, which uses native data policies to automatically expire and delete sensitive cached information. This feature is crucial for enterprises in regulated industries, helping them maintain security, privacy, and compliance without manual intervention.
The efficiency gains from KV cache offloading also translate directly to the bottom line. By ensuring GPUs are constantly fed with data rather than sitting idle or re-computing information, organizations can maximize the return on their most expensive hardware assets. This improved utilization, combined with the DASE architecture's efficient scaling, promises a more sustainable economic model for large-scale inference.
"Every frontier lab, every AI-native company building the future of AI needs infrastructure that scales as fast as their ambitions," said Erwan Menard, SVP Product Management at Crusoe, another cloud partner. "Our collaboration with AMD and VAST gives Crusoe Cloud customers a validated foundation purpose-built for AI."
This focus on operational reality and economic viability signals a maturation of the AI market. As the initial hype cycle gives way to pragmatic deployment, solutions that balance groundbreaking performance with practical, enterprise-grade features are poised to define the next phase of AI adoption. The alliance between VAST Data and AMD represents a formidable entry in this new race, betting that an open, efficient, and data-centric approach is the key to unlocking the true business value of artificial intelligence.
Topics & Related
Agentic AI
Semiconductors
GPUs
📝 This article is still being updated
Are you a relevant expert who could contribute your opinion or insights to this article? We'd love to hear from you. We will give you full credit for your contribution.
Contribute Your Expertise →