Technology7 min read

AI Benchmarks Are Broken. Vals AI Wants to Fix That

Vals AI, backed by Andreessen Horowitz, is building a neutral AI benchmarking platform to restore trust in AI model evaluation. Here's why it matters.

AI Benchmarks Are Broken. Vals AI Wants to Fix That

Key takeaways

  1. 1Why AI Benchmarks Are Losing Credibility MMLU—the Massive Multitask Language Understanding benchmark—was once considered a rigorous test of model knowledge across 57 academic subjects.
  2. 2By 2024, frontier models had pushed past 90 percent accuracy on it, a threshold that once seemed impossibly distant.
  3. 3HumanEval, the widely used coding benchmark from OpenAI, faces similar saturation—models now routinely exceed 90 percent, yet developers still encounter basic failure modes when deploying those same models in production.
  4. 4Stanford's Center for Research on Foundation Models, through its annual AI Index reports, has repeatedly flagged the inconsistency between lab-reported benchmark performance and real-world utility.
Sections · 6

The number on the leaderboard is not what you think it means. When a major AI lab announces that its latest model scored 90 percent on a flagship reasoning benchmark, that figure travels through press releases, analyst briefings, and procurement decisions—often without anyone questioning whether the benchmark itself still measures anything meaningful. That credibility gap is exactly the problem a startup called Vals AI, backed by Andreessen Horowitz, is positioning itself to solve.

Why AI Benchmarks Are Losing Credibility

MMLU—the Massive Multitask Language Understanding benchmark—was once considered a rigorous test of model knowledge across 57 academic subjects. By 2024, frontier models had pushed past 90 percent accuracy on it, a threshold that once seemed impossibly distant. The problem: researchers at Stanford and elsewhere began documenting signs of benchmark contamination, where training data overlaps with test sets, inflating scores without reflecting genuine capability gains. HumanEval, the widely used coding benchmark from OpenAI, faces similar saturation—models now routinely exceed 90 percent, yet developers still encounter basic failure modes when deploying those same models in production.

This is not a fringe concern. Stanford's Center for Research on Foundation Models, through its annual AI Index reports, has repeatedly flagged the inconsistency between lab-reported benchmark performance and real-world utility. MLCommons, the industry consortium behind the MLPerf benchmark suite, has noted that evaluation standards have failed to keep pace with the speed of model releases. The result is an ecosystem where benchmark scores function more like marketing assets than technical specifications.

The structural incentive problem compounds the measurement problem. AI labs benchmark their own models on tests they often design or select, then self-report results without independent verification. A 2023 analysis from researchers at the University of Washington found that models fine-tuned on benchmark-adjacent data could achieve dramatically higher scores without corresponding improvements in downstream tasks. Self-reported numbers, in other words, tell you how well a model was trained to take a specific test—not how well it will perform on your actual workload.

Meet Vals AI: The a16z-Backed Startup Redefining AI Evaluation

Into this environment steps Vals AI, a startup that has secured backing from Andreessen Horowitz with the explicit goal of becoming a neutral, trustworthy resource for AI benchmarking. The pitch is straightforward: as enterprises and developers face an ever-growing menu of models from OpenAI, Anthropic, Google, Mistral, Meta, and dozens of smaller players, they need an evaluation layer that doesn't have a financial stake in any particular model winning.

Read next Laika's Wildwood: Stop-Motion Fantasy at TIFF 2026

The specific mechanics of Vals AI's approach remain sparse in early reporting, but the mission is clear—to make AI evaluation more credible and independent at a moment when the industry needs exactly that. The startup is positioning itself as infrastructure for trust, which is a more durable business than any single model or application.

Andreessen Horowitz's involvement is not incidental. The firm has been explicit in its thesis that the AI stack has three layers—foundation models, infrastructure and tooling, and applications—and that the middle layer, the infrastructure and tooling category, represents some of the most defensible value creation in the current cycle. Neutral evaluation infrastructure sits squarely in that thesis. A firm that has backed both AI model developers and enterprise AI consumers has an obvious commercial interest in a credible referee.

What a Neutral AI Benchmarking Standard Would Actually Look Like

What a Neutral AI Benchmarking Standard Would Actually Look Like — robot and human hands reaching toward ai text
What a Neutral AI Benchmarking Standard Would Actually Look Like — robot and human hands reaching toward ai text

The credibility of any benchmarking organization rests on a few specific properties: independence from the entities being tested, transparency of methodology, resistance to gaming, and relevance to actual use cases. Each of these is harder to achieve than it sounds.

Independence is the easiest to declare and the hardest to maintain. Vals AI's backing from a venture firm that also invests in AI companies creates a structural tension worth naming. The comparison to financial ratings agencies—which theoretically evaluate bonds independently but are paid by the issuers—has already circulated in AI research circles. The analogy is imperfect but instructive. Sustained credibility requires governance structures that ring-fence evaluation decisions from investor relationships.

Transparency of methodology matters because it allows external auditors to catch flaws before they propagate. MLCommons has done this well by publishing its benchmark construction process; the AI community can debate, fork, and improve the methodology. A proprietary black-box evaluation service, by contrast, asks enterprises to trust a score without being able to inspect what produced it.

Resistance to gaming is perhaps the deepest technical challenge. Any benchmark that becomes widely adopted immediately becomes a target for optimization. This is Goodhart's Law applied to AI: once a measure becomes a target, it ceases to be a good measure. Effective benchmarking organizations address this by continuously refreshing test sets, using held-out evaluation data that model developers never see, and designing tasks that are difficult to game without genuinely improving the underlying capability being measured.

Relevance to actual use cases is where existing benchmarks most visibly fail enterprises. A business deploying an AI model for contract review, customer support, or financial analysis doesn't particularly care about MMLU. They care about how the model performs on their specific documents, in their specific context, with their specific failure tolerance. Benchmarks that approximate real enterprise workloads are scarcer and harder to build than academic-style knowledge tests.

The Stakes: Why Reliable AI Benchmarks Matter for Businesses and Developers

The practical consequences of unreliable benchmarks are not abstract. A procurement team at a financial institution choosing between models based on published scores is making a capital allocation decision—often at six or seven figures annually—on data that may not reflect production performance. A developer integrating a model into a healthcare application based on reported reasoning scores may discover critical failure modes only after deployment.

Gartner estimated in 2024 that more than 30 percent of enterprise AI projects that reached deployment were subsequently scaled back or discontinued, with performance falling short of expectations cited as a leading cause. Better evaluation infrastructure would not eliminate all such failures, but it would shift more of the discovery earlier in the process, where it costs far less to course-correct.

For developers, the problem manifests as a constant calibration exercise. Experienced practitioners have long known to distrust published benchmarks and run their own internal evaluations—a practice that is both time-consuming and accessible primarily to organizations with dedicated ML engineering resources. A credible third-party benchmarking service reduces that overhead and democratizes access to reliable model evaluation.

Can One Startup Become the Gold Standard for AI Testing?

The aspiration to become the gold standard for any infrastructure category is ambitious. The comparison benchmark—no pun intended—is something like Underwriters Laboratories for electronics safety, or the credit rating agencies before their credibility was undermined. Building that kind of institutional trust takes time, consistent accuracy, and the willingness to publish unflattering results about well-funded models.

Vals AI enters a field that is not empty. MLCommons, the Holistic Evaluation of Language Models (HELM) project from Stanford HAI, EleutherAI's Language Model Evaluation Harness, and Scale AI's evaluation products all compete for some version of this territory. Each has different strengths: HELM offers breadth and academic rigor; Scale AI brings commercial scale and enterprise relationships; MLCommons brings industry consortium legitimacy.

The a16z backing gives Vals AI real advantages: capital, network access to both model developers and enterprise buyers, and the credibility that comes with a prominent lead investor. It also creates the perception challenge described earlier. Navigating that tension publicly and proactively—rather than treating it as a PR problem to be managed—will likely determine whether the startup achieves genuine neutral authority or merely becomes another vendor in a crowded evaluation market.

What This Means for the Future of AI Model Comparison

The emergence of Vals AI reflects a broader maturation signal in the AI industry. Early-stage technology markets tolerate self-reported performance claims because buyers lack the sophistication to demand more. As AI procurement becomes routine at large enterprises, that tolerance erodes. Buyers want auditable, reproducible, third-party validation—the same standards applied to cybersecurity tools, financial software, and clinical diagnostics.

The pressure to establish rigorous evaluation infrastructure is also coming from regulators. The EU AI Act's conformity assessment requirements, and emerging AI procurement guidance from U.S. federal agencies, both push toward documented, verifiable performance claims rather than marketing materials dressed as technical benchmarks.

What the industry ultimately needs is not a single dominant benchmarking authority but an ecosystem of credible evaluation approaches—task-specific, domain-specific, and general—that together give buyers enough signal to make informed decisions. Whether Vals AI becomes one node in that ecosystem or the central hub depends on execution, governance choices, and whether the startup can resist the commercial temptations that have corrupted evaluation credibility in other industries.

The benchmark problem is real, the business case is solid, and the timing is right. The rest is proof of work.


Source: TechCrunch

Published

29 September 2026

Author

Editorial

Comments

No comments yet. Be the first.

Leave a comment