Research

AI-generated text

Vals aims to set a new standard for AI evaluation

Vals, a startup founded in 2024, builds proprietary benchmarks to measure modern AI models on real-world, domain-specific tasks rather than public academic tests.

Vals aims to set a new standard for AI evaluation

Vals, a startup founded in 2024, aims to modernize how AI models are evaluated. The company argues that many legacy benchmarks—often academic and publicly available—have fallen behind the capabilities of the newest large models and can be gamed, producing results that don’t reflect real-world performance.

Vals’s approach

Instead of measuring general knowledge, Vals builds closed tests that assess how well models perform complex, domain-specific tasks. The firm does not publish the exact test materials, which reduces the chance that companies will train models specifically to optimize for those tests. Vals also evaluates potential negative outcomes, seeking to understand the harms that might arise if models were widely deployed without adequate safeguards.

Domains and topics covered

Vals’s evaluations target industry-specific work such as law, finance, and coding. The company has also expanded into less traditional but critical areas including recursive self-improvement, mental health, cybersecurity, biosecurity, and the law of armed conflict—examining, for example, how models might apply the Geneva Conventions.

Business model and growth

Clients pay Vals to evaluate their models; the company compares this to a student paying the College Board to take the SAT. These evaluations are intended to help companies debug and improve models and are increasingly used as a factor in procurement and strategic decisions.

Vals has grown rapidly since its founding. Last month it raised $40 million in a Series A led by Andreessen Horowitz; earlier it secured a seed round led by 8VC and Bloomberg Beta. The company reports that current revenue is eight times what it was last year. Headcount has also increased—from eight employees at the start of the year to 25—and the company plans to hire another 10 to 15 people and move to a larger office. Vals recently launched a program to provide model evaluations for federal agencies.

Founder background and philosophy

Co-founder Rayan Krishnan, 25, previously interned at Palantir and, while an undergraduate at Stanford, worked for Microsoft and Stanford’s artificial intelligence lab. Krishnan says Vals was founded after observing that existing benchmarks were not keeping pace with frontier advances. He argues benchmarks should verify whether models can actually perform the work companies claim they can, and that robust evaluations should become an industry-standard measure of capability and risk.

Why this matters

As AI models become more integrated into business and public life, precise, domain-specific testing matters not only for product quality but also for how companies report capabilities to investors and regulators. Krishnan believes such benchmarks will increasingly shape how AI companies disclose performance and justify investments in AI.

Near-term plans

Vals intends to expand its office space and staff while continuing to broaden the scope of evaluations to cover risks and capabilities that traditional benchmarks often miss. The company positions its closed, industry-focused testing as a possible new standard for judging AI model performance.