DeepInfra has surpassed $100 million in annual recurring revenue, reflecting rapidly growing demand for the infrastructure that runs artificial intelligence applications at production scale. The company said revenue has increased by more than 900% since the beginning of 2026, while its infrastructure is now processing 22 trillion tokens per week, more than triple its volume in May.
DeepInfra reported an annualized revenue run rate of more than $105 million, representing nearly 14x year-over-year growth. The company is positioning its growth as evidence of a broader shift in AI spending from experimentation toward production workloads.
As AI workloads scale, companies increasingly have to optimize factors including model selection, infrastructure utilization, latency, caching, reliability, and price-performance rather than simply acquiring raw computing capacity. DeepInfra has built its platform around high-throughput AI inference, where those variables can significantly affect the economics of deploying AI at scale.
The milestone follows DeepInfra’s $107 million Series B financing announced in May. The company is using the funding to expand its global infrastructure footprint, increase computing capacity, and accelerate product development. Since the round, DeepInfra has continued adding infrastructure, including its first data center capacity in Canada.
DeepInfra operates a vertically integrated inference platform combining dedicated AI compute, optimized software, and a managed developer experience. The platform supports more than 200 open-source AI models, including multiple versions of leading models, enabling customers to select models for different workloads while maintaining production stability.
Customers using DeepInfra for production AI workloads include LiveKit, Humans&, and OpenCode. LiveKit uses the platform for real-time voice and video AI applications, which require particularly demanding inference performance.
LiveKit’s infrastructure requirements include maintaining time-to-first-token below 500 milliseconds and throughput above 100 tokens per second, while avoiding downtime. The company said DeepInfra has helped it scale to billions of tokens per day, with workloads expected to continue increasing.
DeepInfra is continuing to expand internationally while scaling its enterprise and custom DeepCluster offerings. It is also growing its commercial organization and recently appointed Behzad Nouri to lead go-to-market efforts. Nouri previously spent more than a decade leading sales initiatives at Twilio.
KEY QUOTES:
“AI spending is moving decisively from the lab into production, and that shift is creating an entirely new set of demands. The companies building with AI need inference infrastructure that can deliver the performance, reliability and economics to operate at enormous scale, and that is exactly what we’ve built DeepInfra to provide.”
Nikola Borisov, Co-Founder and CEO of DeepInfra
“Conversational AI demands inference infrastructure that can keep pace with our real-time architecture. LiveKit inference has strict requirements to keep time-to-first-token below 500ms and tokens-per-second above 100, all with no downtime. DeepInfra has been instrumental in helping LiveKit scale to many billions of tokens per day, and we are looking forward to continuing our work together into the trillions.”
Neil Dwyer, Inference Lead at LiveKit

