Vals AI says legacy benchmarks are no longer keeping up with the frontier. After raising a $40M Series A led by Andreessen Horowitz, the startup is pushing a new model: evaluate AI systems not on general knowledge, but on whether they can actually do real work in law, finance, coding and other domains — without leaking the test set to model vendors.
Benchmarking has become the industry norm for validating AI capabilities. It is also, increasingly, a PR game. Companies have learned how to game legacy benchmarks that were never designed for today's models.
Vals, founded in 2024, is trying to change that. The company raised a seed round led by 8VC and Bloomberg Beta, then closed a $40M Series A last month led by Andreessen Horowitz. Revenue is already 8x what it was a year ago, and headcount has tripled from 8 to 25.
The core pitch is simple: stop measuring abstract intelligence, and start measuring whether a model can produce work of human quality inside a specific profession.
Instead of publishing its test materials publicly — which allows model vendors to train directly against the exam — Vals keeps its evaluation sets private. The company tests not only for positive outcomes, but for potential harms, including work in mental health, cybersecurity, biosecurity and even applications of the Geneva Convention.
Companies pay Vals to test their models, similar to how students pay the College Board to take the SAT. The evaluations are also becoming procurement tools for organizations deciding which models to deploy.
Krishnan expects the benchmarking market to grow sharply as AI companies go public and AI models become embedded in core economic infrastructure.




