QumulusAI announced a signed two-year agreement valued at more than $32 million to supply NVIDIA Blackwell B300 capacity to an AI inference platform provider focused on generative AI applications. The agreement includes renewal options, with capacity expected to come online in the fall of 2026.
The customer’s platform serves developers and enterprises running image, video, and other generative AI models, where QumulusAI will provide dedicated GPU clusters to give the platform committed, high-performance capacity as demand scales.
Capacity will be served from QumulusAI’s U.S. data center footprint, using what the company describes as a demand-led deployment model that places capacity into available pockets of power across a distributed network of colocation and owned facilities, allowing it to bring GPU capacity online within months rather than years. The agreement adds a two-year commitment of more than $32 million to QumulusAI’s book of business.
KEY QUOTE:
“Generative media workloads put real pressure on inference infrastructure, images and video are compute-intensive to serve, and the user experience depends on speed. This agreement reflects a pattern we’re seeing in our own business: inference customers want dedicated, committed capacity they can count on, and our model is built to put that capacity to work quickly.”
Mike Maniscalco, CEO, QumulusAI