Arena raised $200 million in Series B funding at a $3.1 billion valuation to expand its platform for evaluating frontier AI models and agents in real-world environments. The round was co-led by Lightspeed Venture Partners and Khosla Ventures, with participation from Salesforce Ventures, 01 Advisors, Dell Technologies Capital, and Endeavor Catalyst. Existing investors including Andreessen Horowitz, Felicis, AMP PBC, QuantumLight, and The House Fund also participated.
The financing follows Arena’s Series A in January 2026 and comes as the company expands beyond measuring AI model capability into evaluating whether AI agents behave safely and in accordance with human intent.
Arena said it has exceeded $100 million in annualized revenue.
The company’s platform evaluates AI systems across real-world tasks including coding, document analysis, creative writing, text, vision, search, video, and image generation.
Arena said Agent Arena has recorded seven million sessions in less than five months since launch.
Across the broader Arena platform, the company has recorded approximately 350 million sessions and 62 million votes across multiple AI modalities.
Arena also reports tens of millions of monthly visitors from more than 150 countries.
The platform has conducted more than 1,000 new model evaluations and has open-sourced approximately 375,000 data points for AI research.
Alongside the funding announcement, Arena introduced the Arena Alignment Index, a new measurement framework designed to evaluate whether frontier AI models and agents act consistently with human intent.
The company said traditional static benchmarks are becoming less effective as models become increasingly capable of recognizing when they are being evaluated.
Arena’s approach instead uses real-world interactions and agent traces to evaluate how AI systems behave when completing actual tasks.
The initial Alignment Index focuses on three measurable behaviors: Unauthorized Action, False Attribution, and Deceptive Completion.
Unauthorized Action measures whether an AI system takes actions beyond what a user requested or permitted.
False Attribution evaluates situations where a model attributes a statement, intention, or fact to a user despite evidence contradicting that claim.
Deceptive Completion measures cases where an AI system tells a user that a task has been completed when it has not.
Arena said the definitions were informed by safety and alignment concepts published by AI laboratories including OpenAI and Anthropic.
The company is initially publishing Alignment Index results for more than 20 frontier models.
Arena plans to expand the index over time by introducing additional verified alignment signals while continuing to maintain its existing capability leaderboards.
The company’s Agent Arena uses causal inference methodology to observe complete workflows between humans and AI agents rather than evaluating only isolated outputs.
Arena said this methodology is intended to provide a more accurate picture of how AI systems perform once they are used in real-world environments.
The Series B funding will support Arena’s effort to build independent evaluation infrastructure that measures both the capabilities and behavior of increasingly autonomous AI systems.
The company said the long-term goal is to provide an independent, data-driven way to determine whether advanced AI systems can be trusted to operate safely and truthfully while acting on behalf of users.
KEY QUOTES:
“This round is a vote of confidence in the idea that as AI gets more powerful, the world needs an independent, data-driven approach to measure not just the capabilities of AI, but whether it can be trusted and used safely in the real world.”
“Agents are no longer just answering questions. They’re writing code, running analyses, and taking actions on people’s behalf, often in areas where the person can’t easily check the work. When an agent deceives a user, or does something it wasn’t permitted to do, it can have serious consequences.”
“Which model performs best is no longer the only question that matters. Today, can we trust what it did and what it says it did is just as urgent, and very much unanswered.”
Arena Team

