Vals AI exposes AI benchmark flaws, spurs enterprise shift

Drip

The gist

Vals AI is shaking up the AI industry by exposing how traditional benchmarks wildly overstate model performance—and fueling a stampede toward specialized, real-world testing.

What to know

  • Vals AI just landed $40M from Andreessen Horowitz after eightfold revenue growth and a $400M valuation.
  • New research shows top AI models fail nearly half of real financial analyst tasks, revealing a glaring gap between benchmark scores and real-world capability.
  • Enterprises are ditching generic models for domain-specific AI—powered by custom, perishable benchmarks that better reflect true business workflows.

Vals AI’s Funding Surge

Backed by top investors, Vals AI is scaling globally and launching advanced tools like Vals Smith to transform AI audit standards across regulated industries.

Vals AI’s recent $40 million Series A funding round, led by Andreessen Horowitz with participation from 8VC, Bloomberg Beta, HRT Ventures, and Next Ladder Ventures, underscores robust investor confidence in the company’s innovative AI evaluation platform. This infusion comes as Vals AI’s revenue has surged eightfold, propelling its valuation to an impressive $400 million and positioning it as a formidable player in the AI audit space.

With this capital injection, Vals AI is aggressively scaling its testing infrastructure and broadening its enterprise customer acquisition on a global scale, while simultaneously expanding its engineering and machine learning teams. This strategic growth aims to meet rising demand across regulated industries, supported by recent platform innovations like Vals Smith, which enables custom coding benchmarks from any GitHub repository, and the Frontier Risk Benchmarks developed in collaboration with CoreWeave and academic researchers.

Sources

Benchmarks Broken by Contamination

AI models are gaming saturated, compromised benchmarks—forcing researchers to retire static tests in favor of dynamic, perishable ones that better reflect real-world complexity.

By mid-2026, frontier AI models demonstrated a striking inability to reliably perform real-world financial analyst tasks, with the top model achieving only about 52% accuracy—essentially failing half of the tasks expected of professional analysts. This glaring performance gap exposes a critical disconnect between high benchmark scores and actual work capability, underscoring that traditional evaluation metrics do not translate into practical utility in complex, domain-specific environments.

Traditional AI benchmarks have become increasingly saturated and ineffective due to contamination and over-optimization, as revealed by a February 2026 preprint from ETH Zurich and Stanford. Nearly half of the 60 widely used benchmarks showed performance differences shrinking to within measurement noise, partly because test questions like those from MMLU appeared verbatim in widely used training datasets such as Common Crawl. This contamination leads models to memorize rather than reason, severely undermining the benchmarks’ ability to distinguish true model capabilities.

Efforts to preserve benchmark integrity through private test sets have proven insufficient, as saturation persists once the underlying distributional characteristics become widely known, regardless of question privacy. The ETH Zurich and Stanford study highlights that only the continuous retirement and replacement of expert-curated benchmarks can maintain meaningful evaluation validity, emphasizing the need for dynamic, perishable benchmarks to prevent score compression and overfitting.

Sources

Real-World Tasks, Real Rigor

Vals AI’s evolving, domain-specific benchmarks and independent audits expose where frontier models fall short, holding them to the standards of actual professional workflows.

Vals AI revolutionizes AI evaluation by partnering with domain experts across fields like finance, law, and healthcare to craft benchmarks that mirror authentic professional tasks rather than relying on traditional multiple-choice or static test sets. These expert-curated benchmarks challenge models to execute complex, multi-step analytical workflows—retrieving precise data, synthesizing it without hallucination, and maintaining chained reasoning—thereby providing a far more valid and nuanced measure of real-world AI capabilities.

Recognizing that benchmarks can become obsolete as AI models evolve, Vals AI treats these evaluation tools as perishable assets, systematically retiring and replacing them to prevent saturation and preserve meaningful differentiation among frontier models. For instance, in May 2026, Vals retired its CorpFin benchmark in favor of a more challenging Excel test once the former ceased to effectively distinguish top-tier models, ensuring continuous rigor and relevance in AI assessment.

Vals AI’s independent evaluation platform extends beyond static tests by subjecting AI systems to complex, unscripted tasks across regulated industries such as corporate law, banking, and healthcare, effectively serving as an impartial auditor for enterprises, research labs, and governments. Complementing this, tools like Vals Smith enable users to create custom benchmarks from any GitHub repository, while collaborations with partners like CoreWeave have produced specialized Frontier Risk Benchmarks, including a novel RSI Index and cybersecurity assessments, further enhancing the precision and applicability of AI evaluations.

Sources

Vertical AI Powers Enterprise Edge

Domain-specialized AI models are outpacing generalists by embedding industry expertise, workflow context, and defensible outputs—creating competitive moats and driving rapid enterprise adoption.

Enterprise adoption of specialized AI models is rapidly accelerating as companies recognize the critical advantage of deeply embedding domain expertise and workflow context into their AI solutions. Firms like Harvey in legal services and MagicSchool in education demonstrate how vertical GPTs outperform horizontal models by tailoring responses to the nuanced demands of their fields, encoding how work is actually produced, and integrating citation and review mechanisms. This specialization creates a robust competitive moat, especially in high-stakes industries like semiconductor design, where ChipAgents’ $21M funding underscores the premium placed on domain-specific vocabulary and error minimization, areas where generalist models often produce 'confident nonsense.'

The key to enterprise success with specialized AI lies in seamless integration with existing workflows and delivering defensible, context-aware answers that generalist models cannot replicate. OpenEvidence’s $150M ARR and widespread daily use among US physicians highlight how mapping frequent medical queries and citing peer-reviewed literature—while embedding relevant pharma advertising—builds trust and utility. Similarly, TrunkTools’ ability to return exact specifications and RFI responses in commercial construction exemplifies how vertical models navigate complex, domain-specific documents to answer thousands of project-specific questions, a feat horizontal models struggle to achieve.

Beyond precision and domain fluency, enterprises are pursuing AI sovereignty by owning and controlling specialized AI capabilities that leverage proprietary knowledge and workflows, rather than relying solely on external frontier models. Thomson Reuters exemplifies this strategy by optimizing AI for professional environments demanding verifiability and precision, reinforcing that the future lies in orchestrating and routing tasks to the most suitable AI—whether frontier or specialized—to maximize outcomes. This hybrid approach not only enhances operational effectiveness but also creates a competitive advantage that is difficult for rivals to replicate using generic AI models.

Sources

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.