Midterms 2026See who we think should earn your vote, based on our standardsThe guide →
WRITTEN IN PLAIN AMERICAN ENGLISH.
CLAY TRIBUNE.
Advertisement

Andreessen Horowitz-backed Vals wants its AI tests to replace the benchmarks everyone else uses

Vals, backed by Andreessen Horowitz, seeks to become the gold standard for AI benchmarking, measuring model performance across law, finance, and coding.

By mitch·5 min read
A modern office with glowing monitors displays code and graphs, with a gold trophy symbolizing a startup's pursuit of AI excellence.

Vals, a startup founded in 2024, is trying to become the gold standard for testing how well AI models actually work. The company just raised $40 million in a series A led by Andreessen Horowitz, after securing a seed round led by 8VC and Bloomberg Beta last year. Now it wants to measure whether models can pass tests that humans can’t — including the Geneva Convention — and charge companies to find out their models aren’t working.

The idea sounds simple: run a model through a task and see what it produces. The execution is anything but. Vals isn’t just asking whether a model knows enough to pass a bar exam. It’s asking whether a model can produce code, legal work, or cybersecurity analysis that a human would consider acceptable. And it’s not releasing the details of those tests to anyone who might want to cheat.

The Founder’s Story

The co-founder is Rayan Krishnan, 25. Before founding Vals, he spent time at Palantir as an intern and held a job at Microsoft while earning his undergraduate degree at Stanford University, where he took part in the campus’s artificial intelligence lab. According to Krishnan, the company came about after he noticed that benchmarking had fallen behind the models it was meant to gauge.

Advertisement

“We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance,” Krishnan shares. He argues that benchmarks should verify what companies actually claim their models can do, rather than measure abstract intelligence.

The company is based on San Francisco’s Folsom Street, in an old brick building that once housed a large brewery. Today it’s shared space for several startups pushing the tech industry forward.

What Vals Actually Tests

Companies can tailor their models to succeed on published benchmark tests, since most benchmarking systems share their tests openly. This arrangement undercuts the point of the tests entirely. Vals avoids that trap by taking a distinct path.

Vals judges models based on how well they handle demanding assignments tied to particular fields rather than general knowledge. The list of those domains includes law, finance, coding, mental health, cybersecurity, biosecurity, and the law of armed conflict. According to Krishnan, the firm has set a standard for recursive self-improvement and is also engaged in applying the Geneva Convention to models.

“What we’re doing is actually looking at what are the real impacts of the models,” Krishnan said. “Can they do work that produces a product of the same quality as a human within every domain?”

He also highlighted the need to look for harmful results. The firm wants to study how models could act on their own in the world, and what the drawbacks would be.

The Pricing Problem

When firms ask Vals to check how their models work, the question becomes whether they’s unusual. Vals is charging for its service, which raises the question: why would a company pay to learn its model isn’t doing what they’re supposed to do.

The revenue model operates like a student paying the College Board to take the SAT, according to Krishnan. The point is that having an effective measurement enables companies to diagnose and improve their performance over time.

Companies searching for new AI models are increasingly treating these assessments as central elements in their acquisition decisions.

Growth So Far

The company began the year with a staff of eight people. By last week, its team had grown to a size three times that original number, reaching 25. Its revenue now stands at eight times what it was in the previous year.

The firm intends to move into a much larger office space. It also aims to hire 10 to 15 more individuals. A recent initiative from Vals focuses on offering model reviews to federal government offices.

Milestone Date
Seed round led by 8VC and Bloomberg Beta Last year
Series A led by Andreessen Horowitz Last month
Revenue eight times prior year Current
Staff grew from 8 to 25 Current

The Bigger Picture

The growth and public trust of AI companies is headed toward a model that Krishnan believes represents the future, with his own firm’s system serving as an example. To support that view, he pointed to recent public offerings in the sector as evidence of where the industry is going.

SpaceX has gone public, and Anthropic is set for later this year. OpenAI may follow soon, and the speaker expects AI models to play a central role in the economy. He believes the benchmarks and evaluations used to judge those models will shape how companies disclose information and discuss future AI investments in public filings.

What We Make Of It

Vals depends on private testing to judge how its models perform. The argument made is that closed tests are more reliable than open ones, since open tests can be rigged.

Another concern is size. Vals remains a modest operation still adding people, and there’s doubt about whether it can match the speed of model growth in fields like law, finance, coding, and biosecurity.

Investors with real money have committed to the company, which shows the market backs the idea. Vals keeps its test materials under wraps, so no one outside can see how it runs its tests.

Right now, Vals is acting like any other ambitious startup by attempting to build a business around a problem it says it solves better than anyone else. The question of whether it succeeds rests entirely on whether its tests truly measure what they say they measure.

Source material: “Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking,” TechCrunch.

The Notebook

Get the Notebook.

The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

We send one note to confirm. Every issue has a one-click way out.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

As an Amazon Associate, Clay Tribune earns from qualifying purchases.