Infinity.inc has raised $15 million in seed funding at a $100 million post-money valuation to expand its AI infrastructure platform. The round included significant participation from Touring Capital, Principal Venture Partners, chip industry executives, researchers from OpenAI and Anthropic, and other angel investors.
The San Francisco-based company is developing software intended to make new AI chips ready for inference workloads in days rather than months or years. Inference is the process through which a trained AI model generates responses, predictions or other outputs when it receives new data.
Infinity plans to use the funding to scale its automated research platform and expand its engineering team. The company will also accelerate partnerships with semiconductor developers, including its existing work with AI chip company d-Matrix.
Infinity said it is already generating millions of dollars in annual recurring revenue from chip design partnerships. Its customers are seeking to shorten the time required to build production-quality software for new AI processors.
The company’s core product is Ignition, an autonomous AI research agent that generates, tests and optimizes the low-level software kernels used to run AI models on specialized chips. These kernels control how efficiently computational operations are performed across a processor’s architecture.
New AI chips can offer significant theoretical computing performance but may struggle to achieve that performance without optimized software. Infinity is targeting the gap between what a chip can deliver under ideal conditions and how it performs when running production AI models.
NVIDIA has maintained a strong position in AI computing partly through CUDA, a software ecosystem developed over decades to help programmers use NVIDIA processors efficiently. Rival chipmakers must create similarly optimized software stacks before customers can deploy their hardware at scale.
Infinity’s platform is designed to automate much of this development process. Human engineers provide architectural direction while Ignition repeatedly writes code, measures performance and uses the results to improve subsequent versions.
This recursive improvement process allows the platform to continuously tune software around the characteristics of each target processor. Infinity said the approach enables hardware engineers to focus on higher-level chip architecture rather than manually optimizing thousands of individual software operations.
In one test, Infinity said Ignition increased inference throughput for the Qwen3-8B model from approximately 1,400 tokens per second to more than 20,000 tokens per second in one day. The company described this as a 14-fold increase and said the resulting framework outperformed the widely used vLLM software by more than 34%.
Infinity has also worked with d-Matrix to optimize software for the company’s Corsair inference chip. Its AI agents reportedly reached as much as 92% of the processor’s theoretical peak performance approximately 10 hours after gaining access to the hardware.
Within 10 days, Infinity had the Qwen3, Qwen3.5 and Gemma4 models running from end to end on the processor. The company said every model layer was written from scratch for the new architecture.
Infinity intends for Ignition to work across proprietary instruction sets and memory architectures. This could help emerging chip companies support major AI models even when their hardware was not represented in the data used to train general-purpose coding systems.
The company’s business model is tied to the performance improvements and cost savings delivered to its semiconductor partners. Rather than relying solely on upfront software licensing fees, Infinity plans to share in the economic value created by faster and more efficient inference.
The opportunity is growing as more AI computing spending shifts from training models to operating them in production. Infinity said inference is expected to represent approximately two-thirds of AI compute spending during 2026.
The company was founded in August 2025 by Jeremy Nixon, a former Google Brain researcher and co-founder of the AGI House community. Infinity is also hiring research engineers focused on hardware enablement and developers responsible for maintaining its inference libraries.
Infinity believes automated software development could allow more semiconductor companies to compete in the AI accelerator market. Its broader goal is to make specialized processors commercially useful more quickly by removing the software bottleneck between chip design and customer deployment.
KEY QUOTES:
“The AI industry has operated under an artificial constraint that only a handful of chips could run AI well, because only NVIDIA spent decades building the software to make them work optimally. Ignition eliminates that constraint. We believe the next era of AI will be defined not just by who makes the best chip, but by who can make any chip run state-of-the-art models at blazing speeds. For hardware providers, the difference between having that optimized software stack and not is the difference between mere potential and true performance. Our role is to ensure every partner reaches that potential.”
Jeremy Nixon, Founder and CEO of Infinity.inc
“The approach to AI-driven model enablement has the potential to significantly shorten the time required to bring new AI compute architectures into production, helping accelerate deployment and time-to-first-revenue. We’re excited about Infinity’s mission to help unlock the full potential of the next generation of AI compute.”
Sid Sheth, Founder and CEO of d-Matrix
“Every major technology platform ultimately creates the most value when it becomes broadly accessible. Infinity is tackling one of the defining challenges of the AI era: making powerful AI affordable and available at scale. We are proud to back a team building the infrastructure needed to extend the benefits of this revolution to many more people.”
Songyee Yoon, Founder and Managing Partner of Principal Venture Partners
“As neural network architectures evolve at a rapid pace, AI hardware companies are caught in a constant cycle of keeping their software stacks current with state-of-the-art models and approaches. Infinity solves this fundamental software bottleneck. Its AI agents automate kernel development and build optimized inference libraries tailored to each target hardware platform, dramatically shortening the time it takes new silicon to achieve its intended peak performance. We’re excited to support Jeremy and his team as they change the game in how chip companies maximize the real-world performance of their silicon.”
Samir Kumar, General Partner at Touring Capital