Hook
At 10:00 AM ET today, a16z announced a $40 million Series A investment in Vals AI, a company few outside the AI infrastructure circles have heard of. But the signal is louder than the check size. Vals AI is not building a foundation model—it’s building an evaluation tool. In a market where AI agents are starting to manage on-chain treasuries, execute trades, and vote in DAOs, the ability to independently verify these agents’ decisions is no longer a luxury. It’s a survival necessity. The question is: does this tool actually work, or is it just another piece of infrastructure that will be eaten by the platforms it sits on?
Context
We are in a bear market. Survival matters more than gains. Over the past six months, I’ve watched three DeFi protocols lose a combined $12 million in total value locked (TVL) after their AI-powered trading bots went rogue—one misread a liquidity pool, another panic-sold into a flash loan. The common thread? No team had a robust evaluation framework before deployment. They trusted the model’s output without a third-party verification layer.
Vals AI enters this void. The company positions itself as the “testing ground” for AI applications—offering automated benchmarks, scenario-based stress tests, and risk scoring. The $40 million round, led by a16z, places it squarely in the AI infrastructure race. But the crypto-angle is critical: as autonomous agents begin to handle real economic value, the evaluation layer becomes the new audit layer.
From my own work in mid-2025, I deployed a custom AI agent to monitor a lending protocol’s reentrancy vulnerabilities. The agent found a hidden bug in 48 hours—but only because I had a makeshift evaluation script. Without a standardized tool, I was essentially flying blind. That experience taught me that the evaluation infrastructure is the weakest link in the crypto-AI stack.
Core
Let’s break down the numbers and the tech. a16z is injecting $40 million into a company that, according to the press release, is launching a new product “for reliable AI evaluation.” The key phrase is “reliable.” In the AI evaluation landscape, reliability is a moving target. The industry has shifted from static benchmarks like MMLU to dynamic, agentic evaluation that tests multi-step reasoning, tool use, and adversarial robustness. Vals AI’s new product likely follows this trend, but the article provides zero technical details—no architecture, no supported dimensions, no open-source transparency.
Based on my audit experience, I can make a few educated guesses. The tool probably uses a “LLM-as-Judge” approach, where a frontier model (like GPT-4o or Claude) scores the target model’s outputs. This is the standard in the industry, adopted by players like LangSmith, Galileo, and Patronus AI. The differentiation comes from the quality of test cases, the coverage of edge cases, and the ability to scale. Vals AI’s funding implies they have something unique—perhaps a proprietary dataset or a novel evaluation methodology—but it’s not disclosed.
Speed is the asset, but silence is the warning. The silence around Vals AI’s technical specifics is a red flag. In the crypto world, we’ve learned that opacity often hides undercooked products. If Vals AI is serious about being the evaluation layer for crypto AI agents, it needs to publish its methodology. Otherwise, it’s just another black box.
Contrarian
Here’s the angle the mainstream coverage misses: evaluation tools, by their very nature, suffer from a fundamental paradox—Quis custodiet ipsos custodes? Who evaluates the evaluator? If Vals AI’s tool relies on another AI to judge outputs, the entire system’s trust hinges on the impartiality of that judge. In a crypto context, where incentives are often misaligned, this creates a potential attack vector. What if a malicious actor compromises the evaluation model? The entire evaluation pipeline collapses.
Moreover, the competitive landscape is crowded but not yet frozen. Vals AI faces competition from both vertical startups (Galileo, Arthur AI, Patronus AI) and platform-native tools (OpenAI Evals, Anthropic’s evaluation suite). The platform giants have a built-in advantage: they can integrate evaluation directly into their model APIs, making third-party tools redundant. For Vals AI to survive, it must prove that its “independent” evaluation is more trustworthy than the platform’s own. That’s a hard sell—especially when a16z is also a major investor in OpenAI and Anthropic. The house didn’t collapse; it was sold. The conflict of interest is subtle but real.

Another unseen risk is “benchmark overfitting.” If Vals AI’s evaluation becomes the de facto standard, model developers will optimize their models to score high on Vals AI’s tests, potentially ignoring real-world performance. This is the same problem that plagued the SAT and MMLU—the number becomes the goal, not the underlying capability. In crypto, this could lead to AI agents that pass evaluation but fail in production, causing catastrophic losses.
Finally, the article ignores the ethical dimension of evaluation tools being used to “game” safety. A malicious actor could use Vals AI’s tool to identify weaknesses in a model and then exploit them in a live attack. This is the dark side of red-teaming infrastructure. Without proper access controls, Vals AI could become a weapon, not a shield.
Takeaway
The next 12 months will tell us if Vals AI is building infrastructure or a mirage. I’ll be watching two things: first, whether they open-source their evaluation dataset. Transparency is the only cure for the audit theater problem. Second, whether they land a crypto-native enterprise customer—a DeFi protocol, a DAO, or a crypto exchange that uses their tool for agent evaluation. That will separate the signal from the noise.
FOMO drove the bus; reality will hit the brakes. The $40 million is a bet on a thesis, not a product. We’ll know soon enough if the thesis holds.