Vals raises $40 million to make AI benchmarking more resistant to gaming

Vals, backed by Andreessen Horowitz, raised $40 million to build confidential, task-based AI benchmarking for companies and federal agencies as older tests s…

AI News

AI benchmarking startup Vals is trying to turn model evaluation into a more practical and trusted layer of the AI market after raising a $40 million Series A led by Andreessen Horowitz. The company argues that many established tests are no longer sufficient for judging increasingly capable models, particularly when companies can see the test material in advance.

Vals evaluates AI systems on complex, domain-specific tasks rather than relying only on public examinations of general knowledge. Its approach is designed to help model developers identify weaknesses, help buyers compare systems for actual work, and give regulators or investors more meaningful evidence about performance and risk.

A different model for AI evaluation

Founded in 2024, Vals was created after co-founder and CEO Rayan Krishnan observed that academic benchmarks were struggling to keep pace with new model capabilities. In comments reported by TechCrunch, Krishnan said the industry was releasing capable models faster than traditional evaluation systems could measure them.

The company’s testing is intended to resemble professional work in areas such as law, finance, and coding. Rather than asking only whether a model can answer difficult questions, Vals says it examines whether the system can produce work comparable in quality to a human’s output within a particular field.

A notable feature is that Vals does not publicly disclose the specific materials used in its tests. The company’s reasoning is that fully public benchmarks can become targets for optimization or training. A model may appear to improve on a test without demonstrating a comparable improvement in broader, unseen work.

Vals also says its evaluations consider harmful or undesirable outcomes, not just successful task completion. Its reported areas of interest include cybersecurity, biosecurity, mental health, recursive self-improvement, and the law of armed conflict. The evidence provided does not detail the methodology, scoring systems, or independent validation for each of those programs.

What the funding and growth claims show

The Series A follows an earlier seed round led by 8VC and Bloomberg Beta. TechCrunch reported that Vals’ latest financing was completed in the month before its September 2026 article and that the company now has a team of 25, up from eight at the beginning of the year.

Those growth figures, along with the company’s claim that revenue is eight times higher than last year, were reported through Krishnan and should be treated as company-reported metrics rather than independently audited evidence. The available source material does not identify Vals’ customers, total revenue, testing volume, or the number of models assessed.

The company makes money by charging AI developers to evaluate their models. That creates a potentially useful but delicate relationship: the party paying for the test is also the party whose system is being judged. Vals presents the service as a development and quality-control tool, comparing it to paying for a standardized examination that reveals where improvement is needed.

The company has also launched a program focused on providing model evaluations to federal agencies. The source material does not specify which agencies are participating or whether the program has produced public results.

Why benchmark neutrality matters to buyers

For AI builders, the value of an evaluation depends on whether it predicts performance outside the test environment. A public benchmark can be easy to communicate, but it may be less useful if models have been tuned against its questions or if the test measures abstract knowledge instead of a real workflow.

Vals’ task-based model could be more relevant to enterprise AI buyers choosing systems for legal review, financial analysis, software development, or other operational tasks. A buyer may care less about a model’s position on a general leaderboard than about whether it completes a defined workflow accurately, consistently, and with acceptable failure modes.

That makes evaluation quality a deployment issue, not simply a marketing issue. Product teams need to know where a model breaks, how often it requires human intervention, and whether a seemingly strong result hides safety or reliability problems. Confidential testing may reduce benchmark gaming, but it also makes outside scrutiny harder. Buyers will need enough methodological transparency to understand what a score means without receiving the test answers themselves.

The commercial stakes may grow as AI systems become embedded in business processes. If model evaluations begin influencing procurement, public filings, investment decisions, or regulatory reviews, the organizations producing those evaluations will face pressure to demonstrate independence, repeatability, and resistance to conflicts of interest.

Evidence, limits, and competitive pressure

The case for Vals rests mainly on the problem it identifies and on the company’s stated approach. TechCrunch reported Krishnan’s description of a market where older benchmarks are falling behind modern models, but the source does not provide comparative results showing that Vals’ tests are more predictive or harder to game than competing systems.

There is also no independent evidence in the available reporting that Vals has become the industry’s preferred evaluator. Its funding from Andreessen Horowitz and earlier backers is a signal of investor confidence, not proof of benchmark accuracy or market adoption. Likewise, the reported revenue growth and expansion to federal-agency evaluations indicate momentum claimed by the company, but do not establish broad customer acceptance.

Vals will also face competition from model developers, research labs, open benchmark communities, specialist testing firms, and internal evaluation teams. Some organizations may prefer transparent public tests, while others may pay for private assessments that reflect their own data and workflows. The company’s challenge is to show that its controlled process produces insights customers cannot obtain through internal testing or existing evaluation platforms.

What to watch next

The clearest signals will be public evidence of how Vals scores models and whether those scores correlate with performance in production. Buyers should watch for disclosed methodology, repeatability studies, model comparisons conducted under consistent conditions, and examples of evaluations changing deployment or procurement decisions.

Its federal-agency program is another important test. Named participants, published findings, or formal procurement relationships would provide stronger evidence than an announcement alone. The company’s planned hiring and office expansion may show continued commercial activity, but customer retention and recurring evaluation volume will matter more than headcount.

The market should also watch whether Vals can preserve confidentiality without becoming a black box. If its assessments influence claims about safety, capability, or investment value, independent oversight and clear explanations of scoring will become central to its credibility.

Creati.ai perspective

Vals is entering the market at a useful moment: AI companies increasingly need evaluations that reflect work people actually perform, while buyers are becoming more skeptical of leaderboard gains that do not translate into dependable deployments. Its focus on hidden tests and domain-specific tasks addresses real weaknesses in public benchmarking.

But becoming a trusted standard requires more than venture backing and rapid growth. Vals will need to demonstrate that its private evaluations are rigorous, reproducible, and sufficiently transparent for customers and outsiders to assess their meaning. The company’s next phase will be defined less by the ambition of its benchmark catalog than by whether its results reliably improve decisions about which AI models to build, buy, and deploy.

Ads