top of page

Vals AI’s $40M Raise Tests the Market for Independent AI Evaluation

Vals AI raised $40 million at a $400 million valuation, turning a funding headline aggregated by Techmeme into a larger test of independent AI measurement.

Andreessen Horowitz led the Series A, with participation from 8VC, Bloomberg Beta, HRT Ventures, and Next Ladder Ventures. Vals announced the financing on August 13, 2026.

The company says its revenue has grown eightfold compared with all of 2025. Its customer base has doubled, while its team has tripled within six months.

Those figures remain company-reported and lack the context available from public financial statements. However, they point toward a market forming around a difficult question: who can credibly measure whether AI models perform useful work?

Model developers already publish scores across coding, mathematics, reasoning, and safety tests. Enterprises also run internal tests before deploying models.

Vals is betting that neither approach offers enough independence, freshness, or coverage. Its primary opponent is not another evaluation startup. It is the practice of allowing AI vendors to define, run, and publicize the tests used to support their own product claims.

That conflict explains why the round matters. The financing does not establish Vals as an industry referee, but it gives the company more resources to compete for that role.

The Funding Round Backs an Independent Referee

Vals is raising money to build measurement infrastructure, not another general-purpose AI model.

According to the company’s Series A announcement, a16z led the $40 million investment at a $400 million valuation. Existing investors 8VC and Bloomberg Beta joined the round.

HRT Ventures and Next Ladder Ventures participated as new investors. Vals did not disclose how much each firm contributed or provide audited revenue figures.

The San Francisco company builds evaluations and benchmarks for tasks in finance, law, software, healthcare, mathematics, and AI research. A benchmark is a structured test used to compare model performance under defined conditions.

Vals says its results have appeared in model cards from OpenAI, Anthropic, Google, Meta, and xAI. Model cards document a system’s capabilities, limitations, and evaluation results.

That claimed adoption matters more than a public leaderboard alone. A benchmark becomes commercially relevant when labs cite it and buyers use it to select models.

Vals also says large enterprise deployments use its evaluations when choosing models and measuring products. The company has supported the Department of Commerce and members of Congress working on AI policy.

These statements describe several distinct markets. Model labs want evidence that a new release improved. Enterprises need tests tied to their own workflows, while governments need measurements related to risk and national capability.

A single evaluation provider serving all three groups faces potential conflicts. However, it also gains visibility across the model development, procurement, and policy cycles.

Vals paired the funding announcement with three product releases. Vals Smith became generally available for creating custom coding benchmarks from GitHub repositories.

The company also announced frontier-risk work covering autonomous AI research, cybersecurity, and mental health. Vals 2.0 added a rebuilt website and an expanded Vals Index.

The Vals Index combines performance across several economically relevant domains. It differs from a single exam-style benchmark because it includes work involving tools, files, code, and longer task sequences.

This financing therefore supports two related businesses. One produces public comparisons that establish credibility and distribution. The other sells customized evaluation infrastructure to organizations making private decisions.

That combination creates the central tension. Public rankings attract attention, but private evaluations are more likely to reflect the conditions surrounding a real deployment.

The next test is whether Vals can connect those businesses without weakening confidence in either one.

Why Techmeme Based Interest Reached AI Benchmarks

The funding drew attention because model capability claims have become harder to interpret just as enterprise spending decisions have grown more consequential.

Many established benchmarks are now too familiar to separate leading systems. Models can approach the top score while still failing on tasks that involve incomplete information or several connected tools.

Public test material can also appear in training data. This contamination can make a model look better without showing that it can generalize to unseen work.

Research on private benchmarking describes the same problem. Public benchmark leakage can weaken comparisons, while private holdouts can reduce exposure to training data.

The issue is not merely academic. An enterprise selecting a model for financial analysis does not need another score on general trivia.

It needs to know whether the system can read relevant documents, use required tools, maintain numerical accuracy, and flag uncertainty. It also needs evidence gathered under deployment-like conditions.

That pressure has increased as AI products have shifted from chat interfaces toward agents. An agent is a system that uses a model, tools, memory, and repeated actions to pursue a task.

Agent performance depends on more than the underlying model. Search quality, tool design, prompts, permission boundaries, and error recovery can change the outcome.

A model that performs well in a controlled question-answering test might fail after receiving a spreadsheet, a browser, and several conflicting instructions. Another model might compensate for weaker reasoning through better tool use.

Vals focuses on these combined systems as well as standalone models. Its methodology includes tool use, multiple input formats, long contexts, and tasks that require sustained work.

This shift pressures model laboratories in a specific way. A lab-controlled benchmark can highlight a system’s strongest configuration while leaving unfavorable settings or failures less visible.

An independent evaluator can standardize conditions across providers. It can also preserve private test material that models have not encountered during development.

That promise helps explain investor interest. Evaluation sits at a control point between technical progress and commercial adoption.

Every model lab needs evidence. Every serious buyer needs acceptance criteria, and regulators increasingly need methods for assessing capability and risk.

Still, the size of those markets does not guarantee that one company will dominate them. Large customers often build internal evaluations because their data, policies, and failure costs are specific.

Labs also possess deeper access to model internals and unreleased systems. Independent evaluators usually test through public interfaces or specially arranged access.

Vals must therefore show that independence produces enough additional trust to outweigh the evaluator’s narrower technical access. It must also update tests faster than models can learn their patterns.

This is the real meaning behind the techmeme based interest. The funding story signals that evaluation is becoming its own commercial layer, not a supporting research function.

Independent Testing Challenges Lab-Controlled Scores

Vals is selling separation between the company making a model and the organization judging its performance.

That separation resembles familiar institutions in finance, medicine, security, and accounting. A claim becomes more useful when an outside party can test it under consistent conditions.

AI lacks a universally accepted equivalent. Model companies publish extensive technical reports, but each provider chooses different tasks, prompts, settings, and comparison points.

Those choices can be reasonable and still make direct comparison difficult. Small differences in prompts, tool access, or reasoning settings can materially affect performance.

Vals argues that model development is moving faster than the community can create demanding new tests. Older tests become saturated, while public examples can enter future training sets.

Its proposed answer uses privately held test sets and domain-specific tasks. Vals develops those tasks with subject-matter experts and publishes selected validation examples for context.

The company’s evaluation methodology separates public validation material, private validation sets, and test sets that remain confidential. Only the private test set determines published benchmark scores.

That structure reduces direct contamination risk. It does not eliminate every route through which a model or developer can adapt to a benchmark.

Labs can still optimize against public descriptions, previous results, or related tasks. Repeated submissions can also reveal information about a hidden test through score changes.

Vals reports accuracy alongside cost, latency, uncertainty, and qualitative failure patterns. This broader view matters because the highest-scoring model is not automatically the best deployment choice.

A slightly less accurate model might run faster or handle failures more predictably. Another might produce stronger results only when given expensive reasoning budgets.

Real-world tasks also produce less tidy outputs than multiple-choice exams. A legal memorandum can contain a correct conclusion but omit a controlling case.

A financial model can use the proper structure while carrying a numerical mistake through several tabs. An agent can complete a coding task but introduce an unrelated regression.

Scoring those outputs often requires detailed rubrics, deterministic checks, or another model acting as a judge. Each method introduces its own assumptions.

Exact checks provide clarity but cover only outputs with objectively testable answers. Human review offers nuance but costs more and can vary between reviewers.

Model judges scale more easily, yet their preferences and errors can shape the final ranking. A benchmark provider must show how it validates those judges.

This creates an important difference between independent and infallible. Independence can reduce a vendor conflict, but it does not guarantee that the test measures the correct thing.

Competitors approach the problem from different positions. Patronus AI offers automated evaluators, datasets, experiments, and production monitoring for enterprise systems.

Its evaluator documentation describes continuous testing of evaluator accuracy across proprietary and public datasets. This emphasizes evaluation within product development and operations.

METR focuses heavily on frontier capabilities and risks, including whether AI systems can accelerate AI research. Epoch AI has developed difficult benchmarks with unreleased questions to limit contamination.

Scale AI combines data services with evaluations for frontier labs and enterprises. Academic groups continue to create public benchmarks that offer transparency and broad participation.

Vals sits across several of these categories. It publishes public model comparisons, creates proprietary tests, supports custom benchmarks, and works on frontier-risk evaluations.

That breadth can become an advantage if its methods transfer across domains. It can become a liability if customers question whether one organization can maintain expertise in every field.

The company’s funding gives it time to prove that model-independent scoring is a durable business. It does not settle which evaluation model will become standard.

Private Tests Solve One Problem and Create Another

Keeping test sets private limits contamination, but it also asks users to trust claims they cannot fully inspect.

Transparency and test integrity pull in opposite directions. Publishing every question lets researchers audit a benchmark, reproduce results, and identify mistakes.

It also gives model developers direct access to the material. That access makes deliberate training, accidental contamination, and repeated optimization easier.

Private tests protect the questions but restrict outside scrutiny. Users must trust the benchmark creator’s sampling, labels, rubrics, and scoring implementation.

Vals tries to balance those demands through public validation examples and private scoring data. It also open-sources parts of its evaluation infrastructure.

Open infrastructure can improve reproducibility. It lets researchers examine how models are called, how runs are distributed, and how common interfaces handle different providers.

The underlying private questions remain unavailable. That means an outsider cannot fully reproduce a published score without access from Vals.

This tradeoff is not unique to Vals. FrontierMath and other newer benchmarks have withheld most problems because capable models could encounter public questions during training.

The broader AI Index report also highlights differences between developer-reported performance and independent testing. It identifies contamination as a source of inflated results.

Private data is therefore a rational response to a documented weakness. Yet secrecy can hide weak questions as easily as it can protect strong ones.

Benchmark design also determines what a score means. A finance evaluation based on standardized documents might not represent a company’s unusual reporting systems or approval rules.

A coding benchmark built from repositories might capture issue resolution but miss security review, architecture decisions, or long-term maintenance.

The company addresses part of this gap through Vals Smith. The product turns a selected GitHub repository into a custom coding benchmark.

That approach can test a model against code closer to the buyer’s actual environment. It also raises questions about repository permissions, sensitive code, and the quality of generated tasks.

A custom benchmark becomes useful only when its tasks reflect work people genuinely need completed. Easy or artificial issues can produce comforting scores without predicting deployment success.

The same problem applies outside coding. Evaluators need representative cases, clear success criteria, and a record of costly failures.

Organizations already maintain some of this material in bug reports, rejected analyses, customer complaints, and review comments. Turning those records into tests requires careful curation.

For teams building their own evidence base, a searchable engineering knowledge base can help organize specifications, incident notes, and past decisions. It does not replace formal model evaluation.

The skeptical view is straightforward. Vals might produce more credible comparisons than vendor-selected tests while still falling short of predicting production outcomes.

A benchmark measures performance inside its chosen frame. Production systems encounter shifting data, adversarial inputs, permission errors, and human behavior outside that frame.

No leaderboard can compress those conditions into one definitive number. Buyers should treat rankings as evidence for investigation, not automatic procurement decisions.

Vals will need to publish enough methodological detail for experts to challenge its conclusions. It must do so without exposing the material that protects each test.

That balancing act will determine whether independent AI model evaluations become trusted infrastructure or another layer of competing claims.

The Business Depends on Turning Scores Into Decisions

The strongest case for Vals is not that it can rank models, but that it can help organizations decide what to deploy.

Public leaderboards create visibility because model releases attract immediate comparison. Enterprise contracts depend on narrower questions with direct operational consequences.

A legal team cares whether an agent finds controlling authority and cites it accurately. A bank needs consistent calculations, traceable assumptions, and controlled data access.

A software organization needs tests based on its repositories, tools, and review standards. Government evaluators need evidence concerning cyber capability, autonomy, and misuse.

These customers do not share one definition of success. Vals must create reusable infrastructure while allowing each buyer to define local requirements.

That is a difficult product boundary. Too much standardization weakens relevance, while too much customization turns the business into labor-intensive consulting.

Vals Smith offers one possible mechanism. Repository-specific tasks can be generated and run through a common evaluation platform.

If the system automates most task creation and validation, Vals can deliver custom evidence without rebuilding its process for every customer. If expert review remains extensive, scaling becomes harder.

The company’s eightfold revenue claim suggests early demand, but it reveals little about durability. Vals did not disclose annual recurring revenue, customer concentration, retention, or contract length.

It also compared current revenue with all of 2025 rather than providing a conventional year-over-year growth rate. That makes the statement difficult to interpret precisely.

The doubled customer count presents a similar limitation. A larger customer base is encouraging, but contract value and expansion matter more than account count alone.

The tripled team shows that Vals is investing ahead of anticipated demand. It also raises the amount of revenue needed to support ongoing benchmark development.

Fresh evaluations require domain specialists, engineers, reviewers, infrastructure, and access to many model APIs. Long-horizon agent tests can consume substantial compute and reviewer time.

Vals must update benchmarks whenever models saturate them, exploit scoring weaknesses, or gain new interfaces. Static tests lose value in a market where capabilities change within months.

The company’s new RSI Index illustrates the cost and ambition involved. It evaluates whether models can conduct AI research under fixed time and compute budgets.

Its RSI benchmark reports that tested systems completed real experiments but did not reach the leading human frontier. The models also performed worse during hidden confirmation than development measurements suggested.

That gap demonstrates why holdout testing matters. A promising result during experimentation can weaken when repeated against unseen data or controlled seeds.

It also shows the limits of a headline score. The tested agents differed in research strategy, cost, experimentation rate, and follow-through.

Enterprises need the same depth when reviewing a model for production. A single percentage cannot show why failures happen or whether safeguards can contain them.

Vals can create value by connecting scores to deployment choices. It must identify which model, configuration, and tool setup fits a specific risk tolerance.

That position places pressure on internal evaluation teams and competing platforms. Buyers will compare the cost of an external evaluator with building tests themselves.

Large AI labs can also expand their customer-facing evaluation products. Cloud providers already control model hosting, monitoring, security, and procurement relationships.

An independent provider has less distribution but greater separation from model sales. Whether customers value that separation enough to pay remains the commercial question.

The $400 million valuation assumes Vals can capture a meaningful share of this decision layer. Revenue growth alone will not prove that thesis.

Repeat usage will. Customers must return for new releases, update their private benchmarks, and use the results in actual purchasing or deployment reviews.

Three Signals Will Test the Techmeme Based Thesis

The next phase should be judged through adoption, methodological scrutiny, and competitive response, not the financing headline alone.

The first signal is repeat enterprise use. Vals should show that customers run evaluations across several model releases and connect results to deployment decisions.

One-time benchmark projects would suggest a services-heavy market. Recurring use would support the idea that evaluations are becoming permanent infrastructure.

Evidence could include renewal patterns, expanding benchmark coverage, or customer accounts describing how scores changed a model selection. Specific case studies would be more useful than aggregate growth claims.

The second signal is outside validation of Vals AI benchmarks. Independent researchers should be able to examine task design, scoring reliability, error bars, and judge behavior.

They do not need access to every hidden question. They do need enough evidence to determine whether published rankings remain stable under reasonable changes.

Challenges will strengthen the company if Vals responds with corrections, versioning, and transparent methodology updates. Unresolved inconsistencies would weaken its claim to neutral authority.

The third signal is the response from model labs and competing evaluators. More citations in model cards would increase Vals’ influence, especially when results are unfavorable to the citing lab.

Competing private benchmarks could fragment the market instead. Buyers might face several incompatible scores, each produced under different rules.

The market can support multiple evaluators, but decision-makers will still need common reporting standards. Scores should identify model versions, settings, tools, costs, sample sizes, and uncertainty.

These signals matter over the next several months because Vals now has the capital to expand. The funding removes one constraint while increasing expectations around execution.

The company must refresh public leaderboards, grow custom evaluation products, and preserve confidence among labs, enterprises, and policymakers. Serving those groups simultaneously will test its neutrality.

Readers should also resist treating any benchmark as a substitute for their own acceptance tests. Independent evidence is most valuable when combined with deployment-specific evaluation.

That means testing real documents, tools, edge cases, and failure costs. It also means rerunning evaluations whenever a model, prompt, retrieval system, or agent workflow changes.

Vals has identified a genuine measurement gap. Research on contamination and the rapid saturation of older tests supports its diagnosis.

What remains uncertain is who will own the solution. Vals could become an important independent layer, or evaluation could remain distributed across labs, customers, nonprofits, and specialized vendors.

The techmeme based funding story is therefore only the opening result. Watch whether organizations repeatedly trust Vals when money, safety, and model selection depend on the score.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

​Add Search Bar in Your Brain

Just Ask remio

Remember Everything

Organize Nothing

bottom of page