
Vals AI Raises $40M to Expand Independent AI Benchmarking
Vals AI Raises $40 Million After Finding Frontier Models Fail Over Half of Real Finance Tasks
An AI benchmarking startup just raised serious capital on the back of a finding that should give any business relying on AI for professional work real pause. Vals AI closed a $40 million Series A at a $400 million valuation, led by Andreessen Horowitz, according to Tech Funding News's reporting on the round, announced August 13. Existing investors 8VC, Pear VC, and Bloomberg Beta returned, joined by new backers HRT Ventures and Next Ladder Ventures.
What Vals AI Actually Does
Vals AI builds independent benchmarks that test AI models on real professional tasks rather than academic or general-purpose question answering, evaluating work performed by lawyers, bankers, engineers, and doctors, according to Citybiz's reporting on the company. Using task taxonomies developed with reference institutions, Vals then builds automated grading systems capable of evaluating a model's actual work product against an expert standard.
The Statistic That Explains the Investor Interest
The specific finding driving attention to this round is genuinely striking. Vals AI's independent testing found frontier AI models correctly complete fewer than 52% of real financial analysis tasks, according to TechTimes' reporting on the results. That's a meaningful gap between vendor-marketed capability and actual professional-grade reliability, exactly the kind of discrepancy that makes independent evaluation genuinely valuable to enterprises deciding which models to actually trust with real work.
Vals AI at a Glance
Metric | Figure |
|---|---|
Series A raised | $40 million |
Valuation | $400 million |
Revenue growth vs. 2025 | 8x |
Customer base growth | Doubled in 6 months |
Team growth | Tripled in 6 months |
Frontier model accuracy on finance tasks | Under 52% |
Why Traditional Benchmarks Are Breaking Down
The underlying market problem Vals AI is solving connects to a genuine structural issue across AI evaluation broadly. Academic benchmarks worked for years, giving the industry a common way to compare models, but that system has broken down as frontier models increasingly ace the same tests, since public datasets get saturated, leak into training data, or simply become the exact target a model gets optimized against, according to Tech Funding News's analysis. A related study cited by Pebblous found 29 of 60 widely used public benchmarks are now considered saturated, meaning they no longer meaningfully distinguish strong models from weak ones.
Vals AI's response is deliberately fast-moving: the company says it can turn around benchmark results within hours of gaining access to a new model, and it actively retires tests once they stop separating capable models from weaker ones. In May, for instance, the company swapped out a corporate-finance benchmark called CorpFin for a new Excel-modeling test specifically because CorpFin had stopped producing meaningful differentiation between models.
Real Industry Adoption, Not Just Investor Enthusiasm
Vals AI's evaluations have already been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI, according to Tech Funding News's reporting, giving the company genuine credibility across the entire frontier AI lab landscape rather than favoring any single provider. The company's growth metrics support that traction: revenue has grown eightfold compared to all of 2025, while its customer base doubled and its team tripled over the past six months.
Vals isn't operating in an empty category. Competitor LMArena, which rebranded to Arena in early 2026, raised a $150 million Series A at a $1.7 billion valuation in January, though it takes a different approach, collecting millions of human preference votes to measure which model people prefer in open-ended conversation, a fundamentally different question than which model completes professional work correctly, according to TechTimes' analysis of the competitive landscape. Cambridge-based Trismik has raised roughly $3 million applying psychometric testing methods to the same broad problem.
Alongside the funding, Vals launched Vals Smith, a new tool letting users create custom coding benchmarks directly from their own GitHub repositories, extending the company's evaluation approach beyond its existing standardized tests.
Why This Matters for Business
This funding round is worth taking seriously for any business currently deploying AI models for finance, legal, healthcare, or engineering work based purely on vendor-published benchmark scores. A sub-52% accuracy rate on real financial analysis tasks is a genuinely important data point for any company assuming frontier AI models are production-ready for high-stakes professional work without meaningful human review.
For businesses evaluating AI vendors, Vals AI's rapid growth signals that independent, third-party evaluation is becoming standard due diligence practice, not an optional extra, particularly as public benchmarks lose credibility and vendor-reported capability claims become harder to independently verify.
Frequently Asked Questions
What does Vals AI do?
Vals AI builds independent benchmarks that test AI models on real professional tasks across law, finance, healthcare, and software, using automated grading systems that evaluate a model's actual work product against expert standards.
How much funding did Vals AI raise?
Vals AI raised $40 million in a Series A round led by Andreessen Horowitz, valuing the company at $400 million, announced August 13, 2026.
How accurate are AI models at real finance tasks, according to Vals AI?
Vals AI's testing found frontier AI models correctly complete fewer than 52% of real-world financial analysis tasks, revealing a significant gap between vendor-marketed capability and actual professional-grade performance.
The Fast Version
Vals AI raised $40 million in Series A funding at a $400 million valuation, led by Andreessen Horowitz, to expand its independent AI benchmarking platform for professional tasks. The company's testing found frontier AI models correctly complete fewer than 52% of real financial analysis tasks, highlighting a meaningful gap between marketed capability and actual reliability. Vals AI's evaluations have been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI, with revenue growing eightfold and its customer base doubling over the past six months.



