Business

AI-generated text

Arena raises $200M at $3.1B valuation as crowdsourced AI evaluation gains commercial traction

Arena, born as a 2023 UC Berkeley research project that crowdsources human rankings of AI outputs, has closed a $200 million Series B at a $3.1 billion valuation after reporting $100 million in annualized run‑rate revenue in June.

Arena raises $200M at $3.1B valuation as crowdsourced AI evaluation gains commercial traction

Arena, which began in 2023 as a research project at the University of California, Berkeley to crowdsource human rankings of AI outputs, announced on Thursday that it has closed a $200 million Series B round at a $3.1 billion valuation.

The funding follows the company’s report that it reached $100 million in annualized run‑rate revenue in June. In January, Arena disclosed a $150 million Series A at a $1.7 billion post‑money valuation, and at that time said its annualized revenue was $30 million. By those figures, the company’s valuation has nearly doubled in roughly ten months.

The Series B was led by Lightspeed Venture Partners and Khosla Ventures. Other participants include Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, Felicis and additional investors.

Product and usage

Arena operates a crowdsourced platform that is free for consumers: people submit prompts or vibe‑coded tasks and then rate which model performs better. The company says it attracts tens of millions of monthly visitors.

In September of last year, Arena launched a commercial product called AI Evaluations, which provides model labs and enterprises with detailed performance analytics derived from community feedback. The service gained traction amid growing awareness that some models were gaming standard benchmarking tests—improving scores without delivering genuine, generalizable improvements—while enterprises sought evaluations tailored to their internal use cases rather than only standardized metrics.

In its funding announcement Arena stated that “AI is advancing faster than our ability to evaluate it, and static benchmarks break down once models recognize they’re being tested,” adding that the world needs a neutral third party to measure how safe and aligned AI is when used by real people, and positioning Arena to fill that role.

New alignment category

Arena has added a new leaderboard category called alignment. Models in this category are ranked on issues such as unauthorized action (taking actions it wasn’t asked to take), false attribution (incorrectly crediting statements or facts to the wrong source), and what the company calls “deceptive completion” (claiming to have completed a task when it was not done).

On the preliminary alignment leaderboard, several OpenAI models occupy top positions. Anthropic’s Claude Opus 5.5 and Claude Fable appear in sixth and ninth place, respectively.

Why it matters

Combining crowdsourced human ratings with commercial analytics addresses shortcomings of static benchmarks and provides enterprises and research labs with alternative, user‑centered evaluations. Arena’s rapid revenue growth and the new financing round indicate investor confidence in the market need for independent, real‑world AI evaluation.