Inside the Race to Measure Frontier Intelligence
Erik TorenbergBen HorowitzJennifer LiRayan Krishnan
- The founding thesis of independent AI evaluation got its proof point when Meta's Llama 4 "performed worse" in closed private tests while showing "incredible ability" on major public tests — a gap between claimed and actual capability that labs' self-reports can't close. Rayan Krishnan argues labs investing billions need a "rational purchasing market" with third-party evidence, a call he says Demis and others across the industry are echoing. "Every time a new trillion-dollar industry emerges, there is a need for such an independent testing group."
- The firm's core governance choice is refusing to sell training data to labs, invoking Enron as the cautionary tale. Krishnan says "a significant portion of this industry has created performance benchmarks as a mechanism to sell its data," and when auditor and consultant are the same party, "the main thing becomes to pass the audit or, in this case, beat the benchmark" — a structural reason for avoiding conflicted incentives.
- Token spend could begin to rival or dwarf payroll, and enterprises are rationing intelligence arbitrarily. One Fortune 10 company gave engineers a roughly $100/day Cloud Code budget with limits resetting at 4 PM — so "the most productive working hours" fell between 4:00 PM and 6:00 PM. Separately, another Fortune 10 company raised an arbitrary per-employee limit from $100 to $300. In Jennifer Li's unlimited-usage experiment, engineers burned 1–2 billion tokens a day (one hit 6 billion), totaling about $1.5 million in a month — "10 times more on tokens than we did on employee salaries."
- Model selection is genuinely counterintuitive without private-repository testing: "in many cases Sonnet is more expensive than Opus because it consumes too many tokens." The firm's Val Smith product builds internal benchmarks from a company's GitHub codebase to find the Pareto-optimal agent; Li's team's audit found Cognition's Devin unusually token-efficient and that subscriptions can beat per-token pricing. Krishnan's frame is that "a firm is effectively its evals" — companies that make eval results legible to ROI math will outcompete.
- The frontier benchmark is the Recursive Self-Improvement Index — proxy indicators across pretraining, post-training, and framework-level engineering, since actually letting a model train its successor is "very expensive and slow." Krishnan calls RSI a major long-run concern: "one country or one company leaping forward and creating models that we know little about," which is why the discussion cites researchers calling for joint government negotiations.
- Ben Horowitz's policy division of labor: government defines what it fears and enforces rules; private evaluators answer the two operative questions — "can the model do it, and can the model be made to do it?" Government is "not particularly equipped" to evaluate, "but they are very good at setting rules because they can enforce them." Rayan adds that labs' posture of keeping dangerous models in-house "doesn't always work."
- On geopolitics, Krishnan wants a "common language of assessments" playing the arms-control verification role — Reagan's "trust, but verify," with eval frameworks as the overflights. He admits surprise at "extremely inefficient" sovereign-AI duplication, notes a Xi–Trump meeting next month as a sign of trust without any verification mechanism, and says cyber evals must move beyond code vulnerabilities to modeling corporate cloud and energy-grid environments.
1. Llama 4 exposed the self-report problem — and the Enron lesson shapes the business model
- Krishnan's origin story: in early 2024 his team concluded "all the public tests were not enough to measure the progress of the models." The proof came with Llama 4 — "a bit of a disaster" — which "in our closed, private tests... actually performed worse" while demonstrating "incredible ability" on major public tests where questions and criteria are open. Labs want a "rational purchasing market" where billion-dollar investments point to real evidence, "not just self-reports to justify the investment."
- Ben's analogy for fuzzy capability standards: the MPAA. What's R versus X "has changed over time," and there's no answer beyond "I'll know it when I see it" — but "if enough people who run companies, manage finances, or whatever else, agree, then it will become the norm," which beats today's open benchmarks that are "vulnerable to attack, and also too narrow."
- Krishnan's structural commitment, drawn from auditing's failures: never sell training data to labs, despite being "often pushed to do this." His Enron read — when the same group audits and consults, "the main thing becomes to pass the audit or, in this case, beat the benchmark. And that's not at all what the market benefits from."
2. The eval machine: six-hour windows, Steve, and "there's always a higher peak"
- The operating constraint: "we never want to be a delay or a hindrance to the release of a model." What began as Krishnan and his co-founder Lynx working all night pre-launch is now massively distributed infrastructure running at each model's rate limits, plus an internal system called Steve intended to progressively absorb more human work.
- The Recursive Self-Improvement Index, the benchmark he's proudest of: since having a leading model train its successor is "very expensive and slow," the firm builds proxy indicators for each stage — pretraining, post-training, and framework-level engineering — while examining the mechanisms and behaviors that enable high-quality research and creation. Labs discuss RSI in model data sheets, "but there is still no common language."
- On benchmark retirement, the unofficial T-shirt slogan is: "There's always a higher peak." As labs hill-climb, "our job is to constantly build these new mountains for them." Torenberg also argues that benchmarks must track the state of the world, such as updated case law in legal-research evaluations, much as lawyers retake licensing exams.
- The agentic shift changes eval architecture: from ImageNet's millions of one-to-one labels to tasks like "create 50 fully functional web applications" — smaller sample sizes, far richer rubrics, and infrastructure stable enough to resume mid-trajectory on tasks running "hours, days, sometimes weeks."
3. Tokens versus payroll: the enterprise is flying blind
- Krishnan's telling anecdote: one Fortune 10 firm deployed Cloud Code with a budget of about $100 a day for its engineers, with a request limit resetting at 4 PM — creating a dead period during the day and peak productivity from 4–6 PM. Separately, another Fortune 10 company arbitrarily set a $100-per-employee limit and later raised it to $300, "almost like one employee's salary spent on tokens." His diagnosis: "misjudgment of intelligence at every level of the stack," compounded by Anthropic's large model-maintenance costs and low margins.
- Jennifer Li's own experiment: a month of unlimited coding tools saw engineers spending 1–2 billion tokens daily, one peaking at 6 billion — roughly $1.5 million in tokens, "10 times more on tokens than we did on employee salaries." Her team's audit of logs, traces, and GitHub work yielded "pretty strange insights": Devin from Cognition "uses tokens very effectively," and subscription pricing can beat per-token pricing.
- That work led to the product Val Smith: build internal benchmarks from your own GitHub codebase to find the Pareto-optimal agent. Results are counterintuitive — "in many cases Sonnet is more expensive than Opus because it consumes too many tokens" — amid a confusing middle tier that includes Luna and Terra alongside Opus and Sonnet, plus Spark, whose version 1.2 is described as very functional. On Torenberg's question about real-time routing via OpenRouter ("which Stripe just bought," as stated), Krishnan says the name is "a bit of a misnomer": it is primarily a gateway, and "the really hard part of routing is creating evaluation systems."
- The thesis line: "a firm is effectively its evals." As Torenberg suggests token spending could "overshadow payroll spending," Krishnan says the companies that make eval results legible for ROI calculation "will outperform competitors in the long run."
4. Policy: rules from government, verification from evaluators
- Ben's division of labor, in full: government has some idea what it fears — biohacking, cyberattacks — but the operative questions are "is the model capable of doing this? And can you make a model do this?" Government is "not particularly equipped" to evaluate that over time — "simply an inappropriate state function" — "but they are very good at setting rules because they can enforce them."
- Rayan adds that labs' promises to keep dangerous models in-house do not always work: "even that doesn't always work... we live in interesting times."
- Jennifer's stance is that policy debate has been "very abstract," so the firm stays in "evidence-gathering mode," while regularly briefing the executive and legislative branches. She insists "it's not really our job" to recommend policy direction. Torenberg separately raises the compute threshold of "10 to the 26th power of FLOPs."
- On misuse testing, Jennifer places the work within the consistency category: "there are already cases now where models undergoing cybersecurity evaluations are actually hacking bug bounties and finding other ways to get around restrictions."
5. Geopolitics: evals as the verification layer of an AI arms race
- Krishnan's honest surprise: from "my very idealistic perspective," sovereign AI looks "extremely inefficient" — duplicated data centers, data-provisioning processes, and giant training runs — "but it seems that we are not living in such a world." His proposed response is a shared assessment language for discussing risks and areas of agreement.
- The nuclear-arms analogy: Reagan's "trust, but verify." Trust signals exist — "Xi Jinping and Trump are going to meet next month" — "but there is no clear way to actually perform the verification part." Common assessments would play the role of warhead counts and overflights; RSI is a key long-run concern because a country or company could leap ahead with "models that we know little about."
- Jennifer's values point, from lived experience: born and raised in China and a user of Chinese open-source models, "you still can't let them talk freely about the CCP and the whole history there" — evals inevitably encode the developer's values, complicating standardization across labs and countries.
- The frontier of the eval catalog itself: cyber work is shifting from code vulnerabilities and memory leaks to infrastructure-level risk — "modeling larger-scale environments of corporate cloud infrastructure or even energy grids." Closing conviction: the most valuable model for this business is one "where incentives are aligned with conducting quality assessments," rather than with helping develop intelligence or the methods by which models improve.
Full transcript
Every time a new trillion-dollar industry emerges, there is a need for an independent testing group. When Meta released Llama 4, the model actually performed poorly in our closed, private tests, but in all the major public tests, it demonstrated incredible capabilities.
What is the limit of what can be achieved, and within that limit, how are you going to do it?
In an ideal world, we would take a leading model and have it train the next version of itself, but that is obviously very expensive and slow. So we form a set of proxy indicators for each stage of the process of creating the next version of the models.
As assessments become more complex, they have a smaller sample size but a broader set of criteria or expectations. Where do you see the gap that exists today?
The government has some idea of what it fears, whether it’s biohacking or cyberhacking, but then the question becomes: Can the model do it, and can the model be made to do it?
1. Why Vals Exists: When Public Benchmarks Stopped Working
What do you think this landscape will look like? I’ll start with a question that arose in early 2024, after our team discovered that all the public tests were not enough to measure the progress of the models, and that a new methodology and approach was needed to keep us on the cutting edge and help the model labs continue to evolve. Maybe take us back to when you were founded and what you thought was missing in the market back then.
Yes, I had experience in research work, including creating tests and assessment systems. Therefore, it was quite obvious to me that there was a close connection between what was needed to create new generative systems and the new evaluation mechanisms themselves. In fact, to improve one, you often need to improve the other.
One of the biggest drivers of model capabilities is having a clear way to evaluate models. So, in early 2024, we saw many interesting models coming to the market. They weren’t all from OpenAI. In particular, it had become harder than ever to figure out exactly what the new models were capable of.
Based on those principles, we realized that there had to be some kind of third-party company that existed solely to create high-quality assessments and tests, so that we could recognize what these models made possible. We released our first tests in 2024.
This has already been realized by many different parts of the industry over the last few years. One obvious question is: Why do you think labs can’t do it themselves? They know the best software on which models are improved and what they lack in terms of capabilities. Why can’t they be the ones doing the testing?
Internally, they create a lot of great tests, and that’s what drives the progress of the models. But I think there is a problem when we talk about the capabilities of self-reported models. One of the early signs of that was when Meta released Llama 4. It was a bit of a disaster.
2. The Llama 4 Disaster: Public Scores vs Private Reality
Interestingly, what we saw was that in our closed, private tests, the model actually performed worse. But on all the major public tests, where the questions and criteria are open-ended, it demonstrated incredible ability. So there’s a huge gap between what was claimed based on these open tests and what we actually found with our higher-quality tests, which have higher signal levels.
I think it’s indicative of a broader concept that labs understand: They would like to see a rational purchasing market. They would like to see that when they invest billions of dollars in building a new model, there are meaningful ways to point to the evidence and say that they’re evolving in these directions, and that it’s not just self-reports to justify the investment.
That’s why you also see cases where Demis and others in the industry are calling for an ecosystem of third-party evaluators.
What historical analogue do you have in mind in this context? Are these rating agencies and auditing firms? What is the correct comparison?
I think there are lessons to be learned everywhere. Every time a new trillion-dollar industry emerges, there is a need for an independent testing group. The fact that things are moving so fast in AI is forcing a lot of these parallels to emerge.
3. Inside the 6-Hour Pre-Release Testing Window
We think of ourselves as trying to be on both sides of the market. There are mechanisms by which labs need to prove that new models are very capable, but there are also parallels where enterprises need to figure out which implementation strategy will provide them with the greatest return on investment.
I want to move on to having you walk us through the 6-hour window before the model is released. Obviously, you need to process tens of billions of tokens without delaying the launch. What part of your work keeps you up at night, and what part is automated? Tell us about it.
It has honestly been quite a journey. The real guiding star is that we never want to be a delay or a hindrance to the release of a model. That means we have to move very fast and extract the maximum signal within the speed limits and compute power that we have.
In the beginning, it looked like this: My co-founder, Lynx, and I worked all night to get as much done as possible and deliver results. We have now built a team, but we’ve also invested heavily in infrastructure so that we can conduct assessments in a massively distributed way, using the maximum rate limits for each model we access.
Additionally, we have an internal system called Steve. Steve is an employee with economical valves. This is the mechanism by which we can, over time, shift more of the human work to Steve.
How do you deal with a problem that is a bit like the “AI-complete” problem? We’re still not very good at evaluating people, or we haven’t reached an agreement on how to do it. There are tests for IQ, EQ, the Big Five personality traits, and so on, but there’s no universally accepted approach, and people have questions about tests like the SAT and everything else.
4. The Limits of Evaluation: Making Fuzzy Evals Explicit
Of course, models are very good at hacking benchmarks, as I’ve argued before. How do you look at this problem? What is the limit of what can be achieved, and within that limit, how do you work with it?
The honest answer is that it makes a lot of the more blurred, distributed forms of assessment clearer. For example, what is the real difference between an associate and a partner in a law firm? There is no clear test or assessment for this in the human world.
So we first have to debug a lot of this in different corporate or real-world workflows to be able to test the models in the same way. In the long run, I think that’s actually going to be the biggest obstacle: our ability to take companies and their valuations and make them understandable. That’s how we figure out what signal to focus on and where we’re actually making changes.
Got it. This is very interesting. Actually, Ben, I have a question about how the industry was formed before intelligence came along. We’re measuring something very volatile now, compared with how it used to be when we were talking about enterprise software, such as Gartner rankings on 70 different metrics. You could rank the projects and say, “This company has these features, but these others don’t.”
Now it seems like a very uneven frontier that is difficult to measure across different industries. What do you think? Are there any analogies from the past or lessons we can learn? What do you think will be the real challenge, and what will we lack in the future?
It’s a bit like the MPAA, isn’t it? The question is: What is art, what is pornography, and where is the line? When is it rated R, and when is it rated X?
The definition of this has changed over time, in my opinion. What used to be an X rating has now become an R, and so on, with all these PG, PG-13, and other ratings. There’s no other answer than the famous phrase: “I’ll know it when I see it.”
I think it’s bound to be fuzzy, but over time, certain norms will form. If enough people who run companies, manage finances, or whatever else agree, then it will become the norm.
I think that’s better than what we have now in open benchmarks, where they say, “If you can solve this specific problem, then you’re at this level,” and so on. It’s vulnerable to attack and also too narrow.
I think when you look at historical analogies, there are a lot of lessons to be learned about what went wrong and what we should avoid. For example, one of the first decisions we made at FedML was to never sell training data to labs.
We’re often pushed to do this. When we start working with a new lab, we’re expected to find and sell them a bunch of training data.
Yes, it’s a profitable business.
Yes, and in fact, a significant portion of this industry has created performance benchmarks as a mechanism to sell its data. That also became its way of entering the market.
But if you look at auditing as an industry, you get problems like Enron. If the same group is responsible for auditing and also for consulting and supporting the company, you get a mixed incentive structure. Then the main thing becomes passing the audit—or, in this case, beating the benchmark.
That’s not at all what the market benefits from, and it’s not what we’re trying to do.
5. The Recursive Self-Improvement Index
Today, you already have a fairly extensive catalog of different types of benchmarks. Some are more focused on a specific industry, and some are more related to consumer mental health. Maybe first let’s discuss which benchmarks are the most popular and the most studied, and then we can dive into one of them.
We’ve done a lot of work in economically interesting areas of application for models.
Our financial agent benchmark is used by many large financial institutions to understand how models are improving. We also have a lot of good things going on in the programming field. Our Vibe Code Bench measures how well models can process natural-language queries and build full-fledged web applications. It has become a great way to track model improvements over the past 9 months.
We also do a lot of experimental work. One benchmark that we recently released, and that I’m very proud of, is our Recursive Self-Improvement Index. This is a topic that many large laboratories are talking about and are starting to report on in their model data sheets. But there is still no common language for discussing the potential of RSI in models. We created this as a way to directly compare different models with each other.
6. Alignment, Reward Hacking & Models Gaming the Test
Yes, I think this is a very cool benchmark. This is very popular in the research community right now: how to measure the progress that can be achieved through the development of advanced models that increase their capabilities. How do you even go about creating this RSI benchmark? I think in an ideal world, you would take a leading model, have it train the next version of itself, and see where the gain comes from.
But, of course, that is very expensive and slow. Therefore, we form a set of proxy indicators for each stage of the process of creating a new version of the model. This includes working with pretraining, post-training, and framework-level engineering. Then we look at the mechanisms and behaviors that allow models to perform high-quality research work and create something new, as well as where they have difficulties.
Very cool. There are also times when you abandon outdated indexes and benchmarks. It’s interesting that I keep watching this benchmark industry, just like in the days of early diffusion models, when you only chose the best images; the same goes for benchmarks. You choose something popular that may have already exhausted itself, and you get a high rating or score, but you take a different approach: if those benchmarks are saturated, you stop using them. Tell me more about this.
I think it’s a necessity, and it’s a kind of never-ending game we play. We have a T-shirt, and our unofficial slogan is, “There’s always a higher peak.” As the base-model labs are “hill climbing” and looking for the next peaks to conquer, our job is to constantly build these new mountains for them.
Very interesting, and I think the economy naturally functions this way. Over time, as agriculture becomes less important to our labor market, new forms of labor emerge that we require from our population. In the same way, we should expect our benchmarks to meet the new limits of what we want from models.
There’s another aspect of moving away from benchmarks that I think is underappreciated: benchmarks should also reflect the current state of the world. Just like a lawyer has to retake a licensing exam, an architect has to get certified, or a doctor has to be tested, we should expect models to also be tested against the current state of the world—against what we know in medicine or against the legal framework we have.
7. Beyond Capability: Cost, Latency & Keeping Benchmarks Fresh
So, in the case of updating case law for things like legal-research benchmarks, it’s about creating a benchmark that better reflects the current state of the world and also guiding models in the direction we want them to go.
How has this changed? First of all, it used to be that we would use multiple-choice or multistep questions just to assess the answers to prompts. There’s a lot more agentic work happening now, whether it’s in finance, the legal field, or especially programming. There are a lot of asynchronous background agents that can simply execute tasks. How has this changed the way you build infrastructure or approach evaluating the capabilities of not only the models themselves but also the agents?
There are also many more dimensions that people are interested in. It’s not just capabilities; it’s cost, latency, and whether the model is flexible enough to cover broader domains and tasks. How do you approach these additional parameters when evaluating?
A lot of effort has gone into this. I think a whole new set of problems arises at the infrastructure level. Let’s say we are currently testing models and their ability to work for hours, days, and sometimes weeks. Therefore, the infrastructure must be very stable to support evaluation over a long period of time. If a request fails, you need to be able to repeat it from the same place rather than retracing the entire trajectory. So there are some simple mechanisms in the infrastructure that we should think about.
Overall, I think we’re seeing assessments becoming more complex, with smaller sample sizes but a broader set of criteria or expectations for them. A benchmark is basically an input space for queries to the model and a set of requirements or rubrics that you define as expectations for the result.
First, you have things like ImageNet, which has millions of images that you’re trying to categorize. This is a one-to-one mapping between the input image and the output text label. Now we have fewer tasks, such as “Create 50 fully functional web applications,” but a much more complex mechanism for evaluating the result that was created. I think this trend will continue as we see models evaluating increasingly complex workflows.
Do you think this will become a kind of real-time mechanism? For example, with something like OpenRouter, which Stripe just bought, will OpenRouter go to the evals and say, “Okay, where should this next request go?” Or will it be purely for selecting a model within a corporation to complete a task?
The name OpenRouter is a bit of a misnomer, as most of its usage comes from acting as a gateway for models. It’s really up to users to decide which models they want to use and when. That’s because the really hard part of routing is creating evaluation systems and trying to determine where exactly a particular set of intelligent systems should be used for a particular application. Our work supporting businesses in building evaluation systems has actually helped many of them implement routers as well.
Tell us more about why this is important not only for labs but also has existential significance for businesses. Maybe tell us more about how you work with corporate clients.
Yes, of course. I think this aspect is very clear for laboratories. If you are raising a lot of money and investing heavily in creating models, it’s important for you to show why your model is getting better and why the client should pay more for it.
But what I think is still underestimated is that, for businesses, this also becomes an existential issue. I have a little story related to this: I was meeting with a Fortune 10 company, and they had implemented Cloud Code with a budget of about $100 a day for their engineers.
I heard that this had fundamentally changed the way they worked at the company because there was a request limit that reset at 4:00 p.m. The most productive working hours therefore fell between 4:00 p.m. and 6:00 p.m., when those limits were reset. But this created a dead period during the day when people would go for a walk or have coffee simply because they had run out of limits.
I think this is a very telling example of where we’re seeing a misjudgment of intelligence at every level of the stack. By this I mean that engineers have a usage limit of $100 and don’t quite understand how to allocate those funds to achieve maximum productivity.
There’s also a Fortune 10 company that arbitrarily set a limit of $100 per employee. They recently increased this limit to $300 per employee. It’s almost like one employee’s salary being spent on tokens for work. This is actually a rather arbitrary decision, as it’s difficult to quantify what the correct usage limit should be.
8. When Token Spend Starts to Eclipse Salary Spend
In addition, Anthropic operates with a sufficiently low margin to support this process. They incur huge maintenance costs for these models. So I think we’re in a world where it’s still unclear what the return on investment, or ROI, should be and how to evaluate the intelligence being used.
Speaking of existential challenges for businesses, I think we’re moving in a direction where token spending could start to overshadow payroll spending. If it’s such a significant expense item, you’ll have to justify the ROI much more clearly than you have been for the past 6 months.
Over time, as we discussed, I believe that a firm is effectively its evaluation results, its evals. A company’s ability to make these results understandable for calculating return on investment will be the reason why that company outperforms its competitors in the long run.
Perhaps it’s worth dwelling in more detail on a similar question: Why can’t laboratories do this themselves, and why is a third-party agency needed for assessment? It’s clear that a certain amount of neutrality is needed in this area. However, businesses will claim that they know the specifics of their clients’ tasks best. So how does LangChain provide value in this situation? Maybe you can also tell us about LangChain’s new product launch.
9. Private Repos vs Public Benchmarks: The Real Performance Gap
I would advise many companies to develop internal expertise, but I believe that this should not be the only solution. There’s a real explosion of intelligence happening right now. More and more labs are developing base models, and each one is producing more models than ever before, along with a bunch of hyperparameters that exist in complex systems and agents.
So there are more choices, and we even talk about specialized intelligence, this new paradigm that’s emerging. More and more intelligent solutions are emerging, and we’re constantly finding new areas for the application of AI models. As a result, the complexity of their use is also increasing.
Therefore, if you represent a company, you face a whole set of complexities and choices. It is very difficult to develop internal capabilities to conduct such an assessment. In an attempt to fix this, we started releasing some products more openly for enterprise use. The first one is called Val Smith.
Val Smith is focused on code generation—this is the area where we're seeing the fastest growth in enterprise AI. It allows any company to take its codebase from GitHub and build an internal test based on it to understand which agents will be the most effective, as well as which will be Pareto-optimal or provide the highest return on investment. We use Val Smith ourselves, and we see many of the best, most cutting-edge companies doing the same. I expect this is the direction the market will move in as it rationalizes.
What are some examples where performance and cost results from testing a private repository are significantly different from those when testing on a public repository using a leading model?
I think it's still unclear today whether the best OpenAI model or the best Anthropic model will actually be suitable for your repository. We saw many counterintuitive examples where we had to do an evaluation to figure out what would provide the best performance for a particular repository. I think there's also a complex middle segment of options emerging now.
In addition to Opus and Sonnet from Anthropic, there's also Luna and Terra. Luna is very competitively priced. Spark is also very cheap, and version 1.2 is very functional. There is also a growing ecosystem of open-source models that companies can host themselves.
So I think that, in this uncertain situation, it's very difficult to understand what's best. We actually see that in many cases Sonnet is more expensive than Opus because it consumes too many tokens. I think if you go by the principle of “use Sonnet where you see fit,” you might end up spending more than you planned.
Tell us more about how this evaluation system will be applied to intellectual work in other fields.
What examples can I give? I think programming is a harbinger of what's to come for every industry, and many of the basic principles established there carry over to other areas. If you have a very good programming agent, you probably have a model that can create PowerPoint presentations or financial models in Excel just as well.
I think in many of these areas we need to use existing practices as a mechanism for building assessments. As Ben said, we haven't solved the question of what human intelligence is yet, but in many industries we have huge archives of data about what work looked like. So our challenge, and the challenge of others, will be to try to turn this into an evaluation system that remains dynamic and can actually evaluate models where human work is being done.
10. How Vals Uses Vals: Token Maxing the Coding Tools
Perhaps to supplement the previous question, how do you use LangSmith internally to evaluate the best programming model for Valse?
Yes, to be honest, this arose from a problem we also encountered. I wanted to experiment with maximizing token usage and was able to get unlimited access to some programming tools for our team for a month. Looking back, many of our engineers were spending 1 to 2 billion tokens per day. It seems that, on a peak day, one engineer spent 6 billion.
Yeah, that's just crazy. How much is that in dollars?
So I went back to it, did the math, and it turned out that we spent about $1.5 million worth of tokens that month.
By the way, it's free.
I know. But we actually spent 10 times more on tokens than we did on employee salaries that month. So it's not even 50/50; it's 10 times more. It was interesting to take stock and see where people used agents, because there's a kind of uncertainty around using models all the time and everywhere.
We faced the fact that we couldn't continue to operate in this mode for the next month. How do we intelligently determine which tools to use, for which teams, and for which projects? So we conducted an experiment, analyzing the work done. We reviewed many logs and traces, analyzed our GitHub repository, and created the Val Smith tool.
We discovered some pretty strange insights. For example, Cognition's Devin tool actually uses tokens very effectively. So this is something we decided to implement more often. I think there are many cases where better pricing models can be obtained through subscriptions rather than paying for tokens. This influenced our strategy for how we could effectively optimize tokens without spending $1.5 million a month.
Very cool. So is your current mode of operation a more token-efficient toolkit and model, with people having more flexibility to use token-based payments for more complex tasks?
11. Policy: What Should the Government Actually Do?
Yes, we have access to all the tools. We give everyone access to everything. But we automatically provide recommendations for any ticket or task on GitHub about where to start the session, and this should regulate actual usage depending on the level of intelligence required for the task.
Very cool. I want to move on to the policy aspect for a moment, since we were talking about how quickly benchmarks become obsolete. In politics, things are even worse because laws move much more slowly than the capabilities of technology.
Ben and Mark spent a lot of time in Washington, D.C., talking to politicians to try to bridge this gap. Given this, who should set these standards? Laboratories? Independent evaluators? Customers? Or maybe the government? How should this work from a policy perspective?
Yes, I think the short answer is that everyone should be involved to some extent. I think there is a benefit to having diverse perspectives. I believe the main problem is that policy discussions over the past few years have been very abstract and have lacked any substantive basis for defining exactly what policy should regulate.
So even when laboratories offer to engage a third-party company or ecosystem for testing, they don't really explain what their operating principles and mechanisms are. I see our role, especially at the initial stage, as being in evidence-gathering mode, where we can get a lot of information and empirical data about the capabilities of the models and where the risks lie. That will help build a more informed policy discussion later on.
How do you view the division of responsibilities, given that the government has conducted the first assessments and continues to conduct them to determine what exactly is advanced technology and needs regulation? They started with some crazy idea about 10²⁶ FLOPs or something like that.
When you think about what the government should do—for example, establish a 30-day or 60-day waiting period—what exactly should happen during that period? How does this intersect with what you do? What's the best way to determine whether a model is advanced and needs to be put in a sandbox for a period of time to make sure it doesn't get access to everything? How do you see this relationship working?
I believe that there are 2 opposing forces that need to be taken into account. The first is the desire to move very quickly and ensure that government processes do not slow down the pace of technological innovation. The other is making sure that the technology being developed is in the best interests of Americans and people more broadly. So I think they are very difficult to reconcile, and often choosing one means harming the other.
My hope is that, through the evidence-gathering process, we can help policymakers better understand what technology should be to serve America's interests. The work of creating the tools to test and enforce that control will fall on the shoulders of third-party evaluators. I believe this will create a mechanism for improving assessment and testing methodologies that will truly keep pace with the development of cutting-edge technologies. This will not lag behind or slow down the pace of development.
When you're developing your evaluations, do you think, “Can we test how easy it is to get this model to engage in bounty hacking or something like that, or to perform illegal activities?” Is that not yet in your scope, or what do you think about that?
Yes, we consider this generally within the consistency category. I think there are already cases where models undergoing cybersecurity evaluations are actually hacking bug bounties and finding other ways to get around restrictions. But we try to assess whether our models are aligned with user intent and identify cases where they exhibit the opposite behavior.
And maybe this question is for you too, Ben. What do you think should be the correct division of labor in this area? What should government agencies control or do on their own, and where should they trust private companies? Where do you see the gap that exists today?
Yes, I think government agencies are getting a lot of warnings, by the way, from the big labs. They say this will lead to biohacking, this will create a cybersecurity risk, and so on.
Therefore, I believe the government needs to act like this: if a model is capable of doing something, can someone force it to perform illegal actions? The government should clearly define what exactly it doesn't want to see on the market and then commission an outside organization to conduct an assessment.
The government has some idea of what it is afraid of, whether it's biohacking, cyberattacks, or something else. But then the question arises: is the model capable of doing this? And can you make a model do this? Someone has to actually evaluate these 2 possibilities.
I believe the government is not particularly equipped to perform the latter task, especially over time. This is simply an inappropriate state function. But they are very good at setting rules because they can enforce them.
So, I think this is the combination we need: the government sets and enforces the rules, and a competent private company reports if those rules are broken. It's quite interesting to watch big labs start saying, “The model has some capabilities, and it can be made to do something bad. Therefore, we will not give it to anyone. We will use it ourselves and make sure that our people don't force it to do anything bad.” But even that doesn't always work, so we live in interesting times, I would say.
For me, there is a gap between what is being said and what happens in the real world, because every infrastructure, especially corporate infrastructure, is very different. Simply the way people already use these models is also very different. Stories based on real examples are extremely rare, which is why we are still talking about the OpenAI hack and Hugging Face. Today, we're still discussing what happened to Fable, even though it only lasted 2 months. But often, that's not exactly how the models actually unfold.
The environment in which they work is very specific. So how do we connect these dots and again create the right environment and rule base for implementing these models? I think it also requires someone with the appropriate capabilities to take all these safeguards and adapt them to real-world conditions so that people can use the models with confidence.
Tell us more about how politicians should collaborate with evaluators. What information do they need? How should these relationships be built to be most effective?
Yes. First and foremost, there should be a mechanism through which findings and data are communicated directly to the relevant individuals in government. We regularly brief the executive and legislative branches on what we are learning about the opportunities and risks of the models. I find that it helps them stay on top of what's going on, as well as keep track of issues that may arise in the future. Everything is moving very, very fast, and it's hard to predict where this is all going.
But at least when you have data, you can start to predict trends. Then I think it's up to the people in the legislative branch to decide in which direction they want to see policy. It's not really our job to give such recommendations. But if they see that there is a significant risk, for example, to the mental health of people under 18 or to biohazards in models, and that requires a standardized way to restrict the release of models, then they are the ones who should shape the policy.
And I think there are other areas where the executive branch, like the Department of Commerce or the Securities and Exchange Commission, is responsible for ensuring that private companies can implement and use models productively for the entire system.
12. The Geopolitics of Evals: Whose Values Get Embedded?
I'd like to consider another aspect—the geopolitical one. I often see evaluations, or evals, as a reflection of the values of the model developer, as you create criteria for what is best in that model. And, of course, different countries and labs within those countries take care of different things.
For example, I was born and raised in China. I also use a lot of Chinese open-source models. You still can't let them talk freely about the CCP and the whole history there because, well, you know what's going on in China. So what do you think about the role of assessments and benchmarks in standardization, or how they are perceived by different model labs from different places?
To be honest, from my very idealistic perspective, I'm surprised to see such a large investment in sovereign AI. If I were to look at it from a bird's-eye view, it would be extremely inefficient to build all these data centers, duplicate data-provisioning processes, and train these huge models when, in fact, we could consolidate many of these efforts. But it seems that we are not living in such a world and are not moving in this direction. In fact, there is a growing effort to create AI at the sovereign level.
So I think that requires a common language to communicate about what the assessment framework is and where we are going to collectively agree on positions on risks. I think there's a lot to learn from the example of nuclear energy. I think Reagan had this phrase: “Trust, but verify.” So I think we're starting to see signs of trust because Xi Jinping and Trump are going to meet next month. But there is no clear way to actually perform the verification part of this process.
Having a common language of assessments will allow us to say things like, “You have the right number of nuclear warheads.” In that example, there were also flybys—the next step, through which a country could check another country's nuclear arsenal through overflights. Similarly, if there is a concern about the societal or even existential risk of AI, it would require us to build a common language of assessments to perform the verification process.
What do you think about how to harmonize such policies? It's pretty hard to do even in America. How do you think about taking this to a global level? Now you're not dealing with corporate clients; you're dealing with governments, and those governments are competing with each other. How do you think this should work?
It would be naive of me to say that I have the perfect solution to this problem today. So I think there are first steps we can start with. For example, there seems to be a lot of talk about cybersecurity risks. I think biosecurity concerns will become even more important over time. So there are clear areas where there will be a mutual interest in agreeing on ways to prevent conflicts in the area of cyber or biosecurity.
In my opinion, I think the most interesting thing in the long run will be the possibility of recursive self-improvement. This is an area where you can see one country or one company leaping forward and creating models that we know little about or that work in ways that are unknown to us. So I think having a common way to describe this as a level of pace that we agree on, or as something that exceeds the pace of development, as far as recursive self-improvement is concerned, is going to be extremely important. That's why I think you see a lot of researchers at Closer Us Labs today calling for joint negotiations between governments.
13. What the Benchmarking Landscape Looks Like Next
What do you think the future will look like, now that we have so many different opportunities and powerful models, and countries that are concerned with different aspects of development? Everyone, of course, cares about the security of AI, but in biotechnology or cybersecurity, people are interested in slightly different things, depending on whether we are talking about offensive or defensive capabilities. What do you think this landscape will look like, and how do you approach developing new tests to keep up with these changes?
Yes, we are focused on creating tests that cover the cutting edge of technology. So when we see new areas of opportunity or risk, we want to turn them into a well-documented assessment on the valis.ai platform. I believe that this requires a continuous expansion of scope over time.
For example, in cybersecurity, much of our previous work has focused on code vulnerabilities or memory leaks that may exist within it. However, in reality, many of the biggest problems or risks lie at the infrastructure level. These are things that cannot be expressed in code alone; they require modeling larger-scale environments of corporate cloud infrastructure or even energy grids so that we can assess the real offensive or defensive capabilities of the models. So making sure our assessments reflect these new areas is extremely important to our work.
We believe that the most valuable model for this business will be one where incentives are aligned with conducting quality assessments, rather than supporting the process of developing intelligence or the methods by which models can improve in this direction.
Perfect. Thank you for coming to the podcast. This was a great episode.
Thank you very much for inviting me.
Thank you very much, Rayan. Thank you, Ben. That was interesting.
Thank you.