Forty million dollars in fresh capital. That’s what Vals closed in August 2026, with Andreessen Horowitz leading the Series A, according to TechCrunch on September 19, 2026. The startup, founded in 2024 by Stanford alum Rayan Krishnan, wants its Vals AI benchmarking system to be the bar everyone else clears. The bet: replace stale test sets and shiny leaderboards with evaluations that act more like a real job audition for models.
TechCrunch reports that Vals first raised a seed round in 2025 from 8VC and Bloomberg Beta, then quickly followed with the $40 million round as demand grew. Krishnan is 25, spent time at Palantir and Microsoft, and says academic benchmarks haven’t kept pace with what frontier systems now claim to do. The team is building in San Francisco, out of a brick Folsom Street building that once housed a brewery. A fitting image for a company trying to clean up a frothy market for AI bragging rights.
Why Vals AI benchmarking wants to replace static tests
Benchmarks have long shaped model marketing. Score higher than a rival and you get the headline, the demo slot, and often the deal. The trouble is that fixed test sets age fast, and models can overfit or even memorize parts of them through training data bleed. Researchers have spent years trying to broaden evaluations beyond trivia and test-set recall. Stanford’s HELM project mapped dozens of scenarios to judge models on accuracy, efficiency, and fairness, not just one number.
According to TechCrunch, Krishnan argues the field needs evaluations that mirror the jobs models are hired to do. That means tasks that change, inputs that aren’t predictable, and scoring that reflects outcomes a business cares about. It also implies ongoing monitoring rather than a one-and-done test. If Vals can make that repeatable at scale, the industry gets fewer victory laps and more hard data about whether a model can, say, draft a compliant email sequence, triage support tickets, or reconcile invoices without drifting.
That shift would also blunt the allure of leaderboard chasing. Crowdsourced arenas, like LMSYS Chatbot Arena, already reward overall usefulness in blind pairwise battles. Vals aims to bring a similar spirit of outcome-first testing into enterprise workflows, with enough control and auditability that buyers can trust the results in a contract discussion.
From leaderboards to job auditions for models
Procurement teams don’t buy abstract intelligence. They buy task performance under constraints: latency caps, cost ceilings, privacy rules, and service-level targets. If Vals turns evaluations into dynamic job auditions, model selection could look less like comparing SAT scores and more like a paid pilot. That would change incentives for vendors and integrators.
Here’s the practical consequence for buyers: proof moves closer to production. Instead of citing a model’s MMLU or MT-Bench score, a vendor would run a standardized “bake-off” on the buyer’s tasks, with metrics mapped to business outcomes and guardrails. The testing rig could plug into CI/CD, rerun after model or prompt changes, and flag regressions. Hardware has had this kind of shared yardstick for years through MLPerf. Vals wants something analogous for foundation models and the apps that depend on them.
That would also expose trade-offs more plainly. A cheaper model might deliver near-equal outcomes once you add retrieval, caching, or domain prompts. A flashier model might stumble on edge cases that matter in a regulated workflow. If Vals AI benchmarking becomes routine in RFPs, the conversation shifts from “who’s number one on a leaderboard” to “who meets our service, safety, and cost bars over time.”
Regulatory tailwinds: rules are starting to favor better evaluations
Europe’s AI law is marching toward enforcement, and it tilts the table toward testing that looks like operations. The EU AI Act expects risk management, documented testing, and post-market monitoring for higher-risk systems. That mindset aligns with an evaluation stack that’s continuous, auditable, and tied to a system’s stated purpose.
Even outside formal “high-risk” use, governments are signaling that marketing claims must match behavior. Buyers will need evidence, not just demo reels. If Vals can provide independent, repeatable measures for specific tasks, that evidence becomes easier to gather and defend. It won’t eliminate the need for domain-specific validation or human oversight. It would, however, make the first pass of vendor screening faster and less prone to hype.
There’s also a research upside. Academic groups could spend less time stitching together ad hoc tests and more time probing failure modes uncovered by shared evaluations. If Vals publishes schemas, error taxonomies, or red-team findings that are reusable, the feedback loop between labs and industry tightens. The company’s pitch, as relayed by TechCrunch, is that the field needs to move beyond abstract IQ tests. Regulators and researchers are inching that way too.
What could make or break the standard
Three questions will decide whether Vals becomes the gold standard or another vendor tool:
- Independence. If model makers can tailor to hidden tests, trust erodes. Clear governance and rotating task pools matter.
- Openness where it counts. Enterprises will want private data kept private, yet they’ll also want public, inspectable methods.
- Cost and speed. Evaluations that take weeks or blow through budgets won’t become a procurement staple.
TechCrunch’s reporting underscores the momentum: fast growth, a serious round, and a founder who watched benchmarks fall behind. The harder path is making results both business-relevant and hard to game. That likely means mixing public tasks with client-specific ones, decay schedules for prompts and datasets, and scoring that penalizes unsafe shortcuts. If Vals gets that balance right, it earns a seat at the table when buyers set requirements.
What to watch next for Vals AI benchmarking
Watch the first public proof points. Do major enterprises start naming Vals in RFP language? Do model vendors cite wins on its tests without cherry-picking? Also track whether researchers engage, the way they did around HELM and other shared efforts. Early collaborations or shared artifacts would hint that the company is building a community, not only a product.
On September 19, 2026, TechCrunch painted a picture of a young company with outsized ambition and fresh backing. The next phase will be measured in boring, valuable details: stable APIs, clear scoring, and audits that hold up under legal scrutiny. If Vals AI benchmarking can deliver that, the leaderboard era starts to fade, and buyers get something far more useful—a repeatable audition that predicts how a model will perform on the job. For more on this, see bloomberg.com and nytimes.com.
