Vals, Backed by Andreessen Horowitz, Aims to Set New Standard for AI Benchmarking
Vals, a 2024-founded AI benchmarking startup, has raised $40 million in a Series A led by Andreessen Horowitz. It says its private, industry-specific evaluations can better verify model capabilities than legacy public benchmarks.
Benchmarking has become a standard way for AI companies to validate models and market them; strong results often serve as evidence of superiority. But companies have learned to game legacy benchmarking systems, many of which predate current models, TechCrunch reported. Vals says it was founded to address that problem.
Rayan Krishnan, Vals’ 25-year-old co-founder, previously interned at Palantir and, as a Stanford undergraduate, worked for Microsoft and Stanford’s artificial intelligence lab, according to the report. He said Vals grew out of his observation that capable new models were reaching the market faster than academic benchmarks could keep up. With AI spreading across society, he said, benchmarks should verify that models can do what companies claim.
Vals does not publicly disclose its specific test materials, unlike many benchmarking systems whose tests are publicly available. That public availability can allow companies to train models against tests, which Krishnan described as a form of cheating. Vals instead evaluates models on complex tasks tied to industries such as law, finance, and coding, and examines real-world impacts rather than only abstract knowledge. Krishnan said the goal is to see whether models can produce work of the same quality as humans in each domain, and also to understand negative outcomes if models were allowed to operate freely.
The firm is expanding into more areas. It has a benchmark for recursive self-improvement and is working on evaluations in mental health, cybersecurity, biosecurity, and the law of armed conflict, including how models apply the Geneva Convention, Krishnan said. Companies pay Vals to test their models, a business model Krishnan compared to a student paying the College Board to take the SAT. He said effective measurement helps companies troubleshoot and improve, and that evaluations are becoming key factors for companies choosing new AI models.
Vals said its revenue is currently eight times what it was last year. The startup began the year with eight people and has tripled to 25, Krishnan said, with plans to move to a larger office and hire another 10 to 15 people. It recently launched a program to provide model evaluations to federal agencies.
Krishnan said he expects benchmarking to shape how AI companies grow and build public trust. He pointed to AI companies beginning to go public, saying SpaceX went public, Anthropic is slated for later this year, and OpenAI will likely be public soon. As AI models become a core part of the economy, he said, the benchmarks and evaluations Vals performs will drive usage and become central to public filings and discussions of prospective AI investments.