[BidClub_]
The a16z Show · · 39 分钟

深入了解衡量前沿智能的竞赛

Erik TorenbergBen HorowitzJennifer LiRayan Krishnan

AI与软件技术政策
YouTube
TL;DR
  • 独立AI评测的创立初衷在 Meta 的 Llama 4 身上得到了验证:它在封闭私测中“表现更差”,却在主要公开测试中展现出“惊人的能力”——声明能力与实际能力之间的鸿沟,不可能靠实验室自报来弥合。 Rayan Krishnan 认为,投入数十亿美元的实验室需要一个“理性的采购市场”,由第三方证据提供依据;他称 Demis 及行业内其他人也在呼吁建立这样的市场。“每当一个新的万亿美元产业出现,就需要这样一个独立测试机构。”
  • 这家公司的核心治理选择,是拒绝向实验室出售训练数据,并以 Enron 作为警示案例。 Krishnan 表示,“这个行业相当大一部分人把性能基准做成了销售数据的工具”;当审计方和咨询方是同一方时,“最重要的事就变成通过审计,或者在这里说,击败基准”——这正是公司要避开利益冲突激励的结构性原因。
  • Token支出可能开始逼近、甚至超过工资单,而企业正在随意配给智能。 一家财富10强企业给工程师设置了每天约100美元的 Cloud Code 预算,额度在16:00重置,结果“生产力最高的工作时段”落在16:00至18:00之间。另一家财富10强企业则把随意设定的员工额度从100美元提高到300美元。在 Jennifer Li 的无限使用实验中,工程师每天消耗10亿—20亿枚 token,其中一人达到60亿枚,一个月合计约150万美元——“我们在 token 上的支出是员工工资的10倍。”
  • 如果没有私有代码仓库测试,模型选择确实会反直觉:“很多情况下 Sonnet 比 Opus 更贵,因为它消耗的 token 太多。” 公司的 Val Smith 产品基于企业 GitHub 代码库构建内部基准,以找出帕累托最优智能体;Li 团队的审计发现,Cognition 的 Devin 在 token 使用上异常高效,而且订阅制可能优于按 token 计费。Krishnan 的判断是,“一家公司的本质,实际上就是它的评测体系”——能把评测结果转化为 ROI 计算的公司,最终会胜出。
  • 前沿基准是递归自我改进指数(Recursive Self-Improvement Index,RSI)——它通过预训练、后训练和框架层工程的代理指标进行衡量,因为真正让一个模型训练其继任者“既昂贵又缓慢”。 Krishnan 认为 RSI 是长期重大风险:“某个国家或某家公司突然跃升,创造出我们知之甚少的模型”,这也是讨论中引用研究人员呼吁各国政府进行联合谈判的原因。
  • Ben Horowitz 对政策的分工是:政府定义担忧什么并执行规则;私人评测机构回答两个实际问题——“模型能不能做到,以及能不能让模型做到?” 政府“并不特别适合”做评测,“但非常擅长制定规则,因为它们可以执行规则”。Rayan 补充说,实验室承诺把危险模型留在内部的做法“并不总是奏效”。
  • 在地缘政治上,Krishnan 希望建立一套“共同的评估语言”,发挥军控核查机制的作用——对应 Reagan 的“信任,但要核查”,把评测框架作为空中侦察。 他承认,主权AI的重复建设“极其低效”令自己意外;他提到 Xi Jinping 与 Trump 将于下月会面,认为这体现了有信任、却没有任何核查机制,并表示网络安全评测必须从代码漏洞推进到对企业云环境和能源电网环境的建模。
摘要 · 为研究而整理的核心内容

1. Llama 4暴露自报问题——Enron教训塑造商业模式

  • Krishnan 回忆,2024年初,团队得出结论:“所有公开测试都不足以衡量模型的进展。”Llama 4 成为了验证——它“有点像一场灾难”,在“我们的封闭私测中……实际表现更差”,但在问题和标准都公开的主要公开测试中,却展现出“惊人的能力”。实验室需要一个“理性的采购市场”,让数十亿美元级投资指向真实证据,“而不是仅靠自报来证明这笔投资合理”。
  • Ben 用 MPAA 来类比模糊的能力标准。R级和X级的界线“随着时间推移发生了变化”,除了“我一看就知道”之外没有更好的答案;但“如果足够多的企业经营者、财务管理者,或者其他相关人士达成共识,它就会成为规范”,这比当前开放基准更可靠,因为后者“容易遭到攻击,而且范围也太窄”。
  • Krishnan 从审计失败中提炼出的结构性承诺是:绝不向实验室出售训练数据,尽管公司“经常被推动去这么做”。他对 Enron 的理解是,当同一团队既做审计又做咨询时,“最重要的事就变成通过审计,或者在这里说,击败基准。这完全不是市场真正受益的方向”。

2. 评测机器:6小时窗口、Steve,以及“总有更高的峰值”

  • 公司的运行约束是:“我们绝不希望成为模型发布的延误或障碍。”最初,Krishnan 和联合创始人 Lynx 在发布前通宵工作;如今,这已经变成大规模分布式基础设施,按每个模型的速率限制运行,另有名为 Steve 的内部系统,目标是逐步接管更多人工工作。
  • 他最自豪的基准是递归自我改进指数。由于让领先模型训练其继任者“既昂贵又缓慢”,公司为预训练、后训练和框架层工程的每个阶段构建代理指标,同时研究支撑高质量研究和创造的机制与行为。实验室会在模型数据表中讨论 RSI,“但仍然没有共同语言”。
  • 关于基准退役,非正式的T恤标语是:“总有更高的峰值。”随着实验室不断进行局部爬坡,“我们的工作就是不断为它们建造新的山峰。”Torenberg 还认为,基准必须跟踪现实世界的状态,例如法律研究评测要更新判例,就像律师需要重新参加执照考试。
  • 智能体转向改变了评测架构:评测不再是 ImageNet 那种数百万个一对一标签,而是“创建50个完全可运行的 Web 应用”等任务——样本量更小,评分标准丰富得多,基础设施也必须足够稳定,能够在中途恢复那些运行“数小时、数天,有时甚至数周”的任务。

3. Token对工资单:企业几乎在盲飞

  • Krishnan 讲了一个典型案例:一家财富10强企业为工程师部署 Cloud Code,每天预算约100美元,请求额度在16:00重置,结果白天出现空档,生产力最高的时段反而落在16:00至18:00。另一家财富10强企业则随意设定每名员工100美元的额度,之后提高到300美元,“几乎像是把一个员工的工资花在 token 上”。他的判断是,从技术栈的每一层看,企业都在错误判断智能的价值;Anthropic 高昂的模型维护成本和低利润率又进一步放大了问题。
  • Jennifer Li 自己做了一个实验:一个月内无限使用编码工具,工程师每天消耗10亿—20亿枚 token,其中一人的峰值达到60亿枚;token 成本约150万美元,“我们在 token 上的支出是员工工资的10倍”。她的团队审查日志、调用链和 GitHub 工作成果后,得到了一些“相当奇怪的洞见”:Cognition 的 Devin “对 token 的使用非常高效”,而订阅制定价可能优于按 token 计费。
  • 这项工作催生了 Val Smith 产品:基于企业自己的 GitHub 代码库构建内部基准,找出帕累托最优智能体。结果反直觉——“很多情况下 Sonnet 比 Opus 更贵,因为它消耗的 token 太多”;中间层模型也令人困惑,Luna 和 Terra 与 Opus、Sonnet 并列,另有 Spark,其1.2版本被形容为功能性很强。针对 Torenberg 关于通过 OpenRouter 实时路由的问题——节目中称 OpenRouter 已被 Stripe 收购——Krishnan 说这个名称“有点名不副实”:它主要是一个网关,而路由真正困难的部分是建立评测系统。
  • 核心判断是:“一家公司的本质,实际上就是它的评测体系。”Torenberg 提到 token 支出可能“超过工资单支出”时,Krishnan 表示,能够把评测结果转化为 ROI 计算的公司,“长期来看会跑赢竞争对手”。

4. 政策:政府制定规则,评测机构负责核查

  • Ben 完整阐述了分工:政府大致知道自己担心什么——生物黑客、网络攻击——但实际需要回答的问题是:“模型是否有能力做到这件事?以及能否让模型做到这件事?”政府“并不特别适合”长期进行这类评测,这“根本不是适合由国家承担的职能”;“但它们非常擅长制定规则,因为它们可以执行规则”。
  • Rayan 补充说,实验室承诺把危险模型留在内部并不总能奏效:“即便这样也不总是有效……我们生活在一个有趣的时代。”
  • Jennifer 认为,政策讨论一直“非常抽象”,因此公司保持在“证据收集模式”,同时定期向行政和立法部门汇报。她强调,“提出政策方向并不是我们的工作”。Torenberg 则另行提出了“10的26次方 FLOPs”的算力门槛。
  • 关于滥用测试,Jennifer 将这项工作归入一致性范畴:“现在已经有一些案例,接受网络安全评测的模型实际上正在攻击漏洞赏金项目,并寻找其他绕过限制的方法。”

5. 地缘政治:评测是AI军备竞赛的核查层

  • Krishnan 坦言,从“我非常理想主义的视角”看,主权AI显得“极其低效”:数据中心、数据供给流程和巨型训练任务都在重复建设,“但我们似乎并不生活在那样的世界里”。他的应对方案,是建立一套共同的评估语言,用于讨论风险和各方共识领域。
  • 核武军控的类比是 Reagan 的“信任,但要核查”。信任信号确实存在——“Xi Jinping 和 Trump 将于下月会面”——“但没有明确方式真正执行核查”。共同评估应当发挥核弹头数量和空中侦察的作用;RSI 是关键的长期风险,因为某个国家或公司可能突然跃升,创造出“我们知之甚少的模型”。
  • Jennifer 从亲身经历出发谈到价值观:她在中国出生长大,也是中国开源模型的用户,“你仍然不能让它们自由谈论 CCP 以及那里的全部历史”;评测不可避免地会编码开发者的价值观,这使不同实验室和国家之间的标准化更加复杂。
  • 评测目录本身的前沿也在变化:网络安全工作正从代码漏洞和内存泄漏,转向基础设施层面的风险——“对更大规模的企业云基础设施,甚至能源电网进行建模”。最后的判断是,对这项业务最有价值的模型,应当是“激励与开展高质量评估保持一致”的模型,而不是激励开发智能,或帮助改进模型的方法。
完整逐字稿
Rayan Krishnan

Every time a new trillion-dollar industry emerges, there is a need for an independent testing group. When Meta released Llama 4, the model actually performed poorly in our closed, private tests, but in all the major public tests, it demonstrated incredible capabilities.

Erik Torenberg

What is the limit of what can be achieved, and within that limit, how are you going to do it?

Rayan Krishnan

In an ideal world, we would take a leading model and have it train the next version of itself, but that is obviously very expensive and slow. So we form a set of proxy indicators for each stage of the process of creating the next version of the models.

Erik Torenberg

As assessments become more complex, they have a smaller sample size but a broader set of criteria or expectations. Where do you see the gap that exists today?

Rayan Krishnan

The government has some idea of what it fears, whether it’s biohacking or cyberhacking, but then the question becomes: Can the model do it, and can the model be made to do it?

1. Why Vals Exists: When Public Benchmarks Stopped Working

Erik Torenberg

What do you think this landscape will look like? I’ll start with a question that arose in early 2024, after our team discovered that all the public tests were not enough to measure the progress of the models, and that a new methodology and approach was needed to keep us on the cutting edge and help the model labs continue to evolve. Maybe take us back to when you were founded and what you thought was missing in the market back then.

Rayan Krishnan

Yes, I had experience in research work, including creating tests and assessment systems. Therefore, it was quite obvious to me that there was a close connection between what was needed to create new generative systems and the new evaluation mechanisms themselves. In fact, to improve one, you often need to improve the other.

One of the biggest drivers of model capabilities is having a clear way to evaluate models. So, in early 2024, we saw many interesting models coming to the market. They weren’t all from OpenAI. In particular, it had become harder than ever to figure out exactly what the new models were capable of.

Based on those principles, we realized that there had to be some kind of third-party company that existed solely to create high-quality assessments and tests, so that we could recognize what these models made possible. We released our first tests in 2024.

Erik Torenberg

This has already been realized by many different parts of the industry over the last few years. One obvious question is: Why do you think labs can’t do it themselves? They know the best software on which models are improved and what they lack in terms of capabilities. Why can’t they be the ones doing the testing?

Rayan Krishnan

Internally, they create a lot of great tests, and that’s what drives the progress of the models. But I think there is a problem when we talk about the capabilities of self-reported models. One of the early signs of that was when Meta released Llama 4. It was a bit of a disaster.

2. The Llama 4 Disaster: Public Scores vs Private Reality

Interestingly, what we saw was that in our closed, private tests, the model actually performed worse. But on all the major public tests, where the questions and criteria are open-ended, it demonstrated incredible ability. So there’s a huge gap between what was claimed based on these open tests and what we actually found with our higher-quality tests, which have higher signal levels.

I think it’s indicative of a broader concept that labs understand: They would like to see a rational purchasing market. They would like to see that when they invest billions of dollars in building a new model, there are meaningful ways to point to the evidence and say that they’re evolving in these directions, and that it’s not just self-reports to justify the investment.

That’s why you also see cases where Demis and others in the industry are calling for an ecosystem of third-party evaluators.

Erik Torenberg

What historical analogue do you have in mind in this context? Are these rating agencies and auditing firms? What is the correct comparison?

Rayan Krishnan

I think there are lessons to be learned everywhere. Every time a new trillion-dollar industry emerges, there is a need for an independent testing group. The fact that things are moving so fast in AI is forcing a lot of these parallels to emerge.

3. Inside the 6-Hour Pre-Release Testing Window

We think of ourselves as trying to be on both sides of the market. There are mechanisms by which labs need to prove that new models are very capable, but there are also parallels where enterprises need to figure out which implementation strategy will provide them with the greatest return on investment.

Erik Torenberg

I want to move on to having you walk us through the 6-hour window before the model is released. Obviously, you need to process tens of billions of tokens without delaying the launch. What part of your work keeps you up at night, and what part is automated? Tell us about it.

Rayan Krishnan

It has honestly been quite a journey. The real guiding star is that we never want to be a delay or a hindrance to the release of a model. That means we have to move very fast and extract the maximum signal within the speed limits and compute power that we have.

In the beginning, it looked like this: My co-founder, Lynx, and I worked all night to get as much done as possible and deliver results. We have now built a team, but we’ve also invested heavily in infrastructure so that we can conduct assessments in a massively distributed way, using the maximum rate limits for each model we access.

Additionally, we have an internal system called Steve. Steve is an employee with economical valves. This is the mechanism by which we can, over time, shift more of the human work to Steve.

Erik Torenberg

How do you deal with a problem that is a bit like the “AI-complete” problem? We’re still not very good at evaluating people, or we haven’t reached an agreement on how to do it. There are tests for IQ, EQ, the Big Five personality traits, and so on, but there’s no universally accepted approach, and people have questions about tests like the SAT and everything else.

4. The Limits of Evaluation: Making Fuzzy Evals Explicit

Of course, models are very good at hacking benchmarks, as I’ve argued before. How do you look at this problem? What is the limit of what can be achieved, and within that limit, how do you work with it?

Rayan Krishnan

The honest answer is that it makes a lot of the more blurred, distributed forms of assessment clearer. For example, what is the real difference between an associate and a partner in a law firm? There is no clear test or assessment for this in the human world.

So we first have to debug a lot of this in different corporate or real-world workflows to be able to test the models in the same way. In the long run, I think that’s actually going to be the biggest obstacle: our ability to take companies and their valuations and make them understandable. That’s how we figure out what signal to focus on and where we’re actually making changes.

Erik Torenberg

Got it. This is very interesting. Actually, Ben, I have a question about how the industry was formed before intelligence came along. We’re measuring something very volatile now, compared with how it used to be when we were talking about enterprise software, such as Gartner rankings on 70 different metrics. You could rank the projects and say, “This company has these features, but these others don’t.”

Now it seems like a very uneven frontier that is difficult to measure across different industries. What do you think? Are there any analogies from the past or lessons we can learn? What do you think will be the real challenge, and what will we lack in the future?

Ben Horowitz

It’s a bit like the MPAA, isn’t it? The question is: What is art, what is pornography, and where is the line? When is it rated R, and when is it rated X?

The definition of this has changed over time, in my opinion. What used to be an X rating has now become an R, and so on, with all these PG, PG-13, and other ratings. There’s no other answer than the famous phrase: “I’ll know it when I see it.”

I think it’s bound to be fuzzy, but over time, certain norms will form. If enough people who run companies, manage finances, or whatever else agree, then it will become the norm.

I think that’s better than what we have now in open benchmarks, where they say, “If you can solve this specific problem, then you’re at this level,” and so on. It’s vulnerable to attack and also too narrow.

Rayan Krishnan

I think when you look at historical analogies, there are a lot of lessons to be learned about what went wrong and what we should avoid. For example, one of the first decisions we made at FedML was to never sell training data to labs.

We’re often pushed to do this. When we start working with a new lab, we’re expected to find and sell them a bunch of training data.

Erik Torenberg

Yes, it’s a profitable business.

Rayan Krishnan

Yes, and in fact, a significant portion of this industry has created performance benchmarks as a mechanism to sell its data. That also became its way of entering the market.

But if you look at auditing as an industry, you get problems like Enron. If the same group is responsible for auditing and also for consulting and supporting the company, you get a mixed incentive structure. Then the main thing becomes passing the audit—or, in this case, beating the benchmark.

That’s not at all what the market benefits from, and it’s not what we’re trying to do.

5. The Recursive Self-Improvement Index

Erik Torenberg

Today, you already have a fairly extensive catalog of different types of benchmarks. Some are more focused on a specific industry, and some are more related to consumer mental health. Maybe first let’s discuss which benchmarks are the most popular and the most studied, and then we can dive into one of them.

Rayan Krishnan

We’ve done a lot of work in economically interesting areas of application for models.

Our financial agent benchmark is used by many large financial institutions to understand how models are improving. We also have a lot of good things going on in the programming field. Our Vibe Code Bench measures how well models can process natural-language queries and build full-fledged web applications. It has become a great way to track model improvements over the past 9 months.

We also do a lot of experimental work. One benchmark that we recently released, and that I’m very proud of, is our Recursive Self-Improvement Index. This is a topic that many large laboratories are talking about and are starting to report on in their model data sheets. But there is still no common language for discussing the potential of RSI in models. We created this as a way to directly compare different models with each other.

6. Alignment, Reward Hacking & Models Gaming the Test

Erik Torenberg

Yes, I think this is a very cool benchmark. This is very popular in the research community right now: how to measure the progress that can be achieved through the development of advanced models that increase their capabilities. How do you even go about creating this RSI benchmark? I think in an ideal world, you would take a leading model, have it train the next version of itself, and see where the gain comes from.

Rayan Krishnan

But, of course, that is very expensive and slow. Therefore, we form a set of proxy indicators for each stage of the process of creating a new version of the model. This includes working with pretraining, post-training, and framework-level engineering. Then we look at the mechanisms and behaviors that allow models to perform high-quality research work and create something new, as well as where they have difficulties.

Erik Torenberg

Very cool. There are also times when you abandon outdated indexes and benchmarks. It’s interesting that I keep watching this benchmark industry, just like in the days of early diffusion models, when you only chose the best images; the same goes for benchmarks. You choose something popular that may have already exhausted itself, and you get a high rating or score, but you take a different approach: if those benchmarks are saturated, you stop using them. Tell me more about this.

Rayan Krishnan

I think it’s a necessity, and it’s a kind of never-ending game we play. We have a T-shirt, and our unofficial slogan is, “There’s always a higher peak.” As the base-model labs are “hill climbing” and looking for the next peaks to conquer, our job is to constantly build these new mountains for them.

Erik Torenberg

Very interesting, and I think the economy naturally functions this way. Over time, as agriculture becomes less important to our labor market, new forms of labor emerge that we require from our population. In the same way, we should expect our benchmarks to meet the new limits of what we want from models.

There’s another aspect of moving away from benchmarks that I think is underappreciated: benchmarks should also reflect the current state of the world. Just like a lawyer has to retake a licensing exam, an architect has to get certified, or a doctor has to be tested, we should expect models to also be tested against the current state of the world—against what we know in medicine or against the legal framework we have.

7. Beyond Capability: Cost, Latency & Keeping Benchmarks Fresh

So, in the case of updating case law for things like legal-research benchmarks, it’s about creating a benchmark that better reflects the current state of the world and also guiding models in the direction we want them to go.

How has this changed? First of all, it used to be that we would use multiple-choice or multistep questions just to assess the answers to prompts. There’s a lot more agentic work happening now, whether it’s in finance, the legal field, or especially programming. There are a lot of asynchronous background agents that can simply execute tasks. How has this changed the way you build infrastructure or approach evaluating the capabilities of not only the models themselves but also the agents?

There are also many more dimensions that people are interested in. It’s not just capabilities; it’s cost, latency, and whether the model is flexible enough to cover broader domains and tasks. How do you approach these additional parameters when evaluating?

Rayan Krishnan

A lot of effort has gone into this. I think a whole new set of problems arises at the infrastructure level. Let’s say we are currently testing models and their ability to work for hours, days, and sometimes weeks. Therefore, the infrastructure must be very stable to support evaluation over a long period of time. If a request fails, you need to be able to repeat it from the same place rather than retracing the entire trajectory. So there are some simple mechanisms in the infrastructure that we should think about.

Overall, I think we’re seeing assessments becoming more complex, with smaller sample sizes but a broader set of criteria or expectations for them. A benchmark is basically an input space for queries to the model and a set of requirements or rubrics that you define as expectations for the result.

First, you have things like ImageNet, which has millions of images that you’re trying to categorize. This is a one-to-one mapping between the input image and the output text label. Now we have fewer tasks, such as “Create 50 fully functional web applications,” but a much more complex mechanism for evaluating the result that was created. I think this trend will continue as we see models evaluating increasingly complex workflows.

Erik Torenberg

Do you think this will become a kind of real-time mechanism? For example, with something like OpenRouter, which Stripe just bought, will OpenRouter go to the evals and say, “Okay, where should this next request go?” Or will it be purely for selecting a model within a corporation to complete a task?

Rayan Krishnan

The name OpenRouter is a bit of a misnomer, as most of its usage comes from acting as a gateway for models. It’s really up to users to decide which models they want to use and when. That’s because the really hard part of routing is creating evaluation systems and trying to determine where exactly a particular set of intelligent systems should be used for a particular application. Our work supporting businesses in building evaluation systems has actually helped many of them implement routers as well.

Erik Torenberg

Tell us more about why this is important not only for labs but also has existential significance for businesses. Maybe tell us more about how you work with corporate clients.

Rayan Krishnan

Yes, of course. I think this aspect is very clear for laboratories. If you are raising a lot of money and investing heavily in creating models, it’s important for you to show why your model is getting better and why the client should pay more for it.

But what I think is still underestimated is that, for businesses, this also becomes an existential issue. I have a little story related to this: I was meeting with a Fortune 10 company, and they had implemented Cloud Code with a budget of about $100 a day for their engineers.

I heard that this had fundamentally changed the way they worked at the company because there was a request limit that reset at 4:00 p.m. The most productive working hours therefore fell between 4:00 p.m. and 6:00 p.m., when those limits were reset. But this created a dead period during the day when people would go for a walk or have coffee simply because they had run out of limits.

I think this is a very telling example of where we’re seeing a misjudgment of intelligence at every level of the stack. By this I mean that engineers have a usage limit of $100 and don’t quite understand how to allocate those funds to achieve maximum productivity.

There’s also a Fortune 10 company that arbitrarily set a limit of $100 per employee. They recently increased this limit to $300 per employee. It’s almost like one employee’s salary being spent on tokens for work. This is actually a rather arbitrary decision, as it’s difficult to quantify what the correct usage limit should be.

8. When Token Spend Starts to Eclipse Salary Spend

In addition, Anthropic operates with a sufficiently low margin to support this process. They incur huge maintenance costs for these models. So I think we’re in a world where it’s still unclear what the return on investment, or ROI, should be and how to evaluate the intelligence being used.

Erik Torenberg

Speaking of existential challenges for businesses, I think we’re moving in a direction where token spending could start to overshadow payroll spending. If it’s such a significant expense item, you’ll have to justify the ROI much more clearly than you have been for the past 6 months.

Rayan Krishnan

Over time, as we discussed, I believe that a firm is effectively its evaluation results, its evals. A company’s ability to make these results understandable for calculating return on investment will be the reason why that company outperforms its competitors in the long run.

Erik Torenberg

Perhaps it’s worth dwelling in more detail on a similar question: Why can’t laboratories do this themselves, and why is a third-party agency needed for assessment? It’s clear that a certain amount of neutrality is needed in this area. However, businesses will claim that they know the specifics of their clients’ tasks best. So how does LangChain provide value in this situation? Maybe you can also tell us about LangChain’s new product launch.

9. Private Repos vs Public Benchmarks: The Real Performance Gap

Rayan Krishnan

I would advise many companies to develop internal expertise, but I believe that this should not be the only solution. There’s a real explosion of intelligence happening right now. More and more labs are developing base models, and each one is producing more models than ever before, along with a bunch of hyperparameters that exist in complex systems and agents.

So there are more choices, and we even talk about specialized intelligence, this new paradigm that’s emerging. More and more intelligent solutions are emerging, and we’re constantly finding new areas for the application of AI models. As a result, the complexity of their use is also increasing.

Therefore, if you represent a company, you face a whole set of complexities and choices. It is very difficult to develop internal capabilities to conduct such an assessment. In an attempt to fix this, we started releasing some products more openly for enterprise use. The first one is called Val Smith.

Val Smith is focused on code generation—this is the area where we're seeing the fastest growth in enterprise AI. It allows any company to take its codebase from GitHub and build an internal test based on it to understand which agents will be the most effective, as well as which will be Pareto-optimal or provide the highest return on investment. We use Val Smith ourselves, and we see many of the best, most cutting-edge companies doing the same. I expect this is the direction the market will move in as it rationalizes.

Erik Torenberg

What are some examples where performance and cost results from testing a private repository are significantly different from those when testing on a public repository using a leading model?

Jennifer Li

I think it's still unclear today whether the best OpenAI model or the best Anthropic model will actually be suitable for your repository. We saw many counterintuitive examples where we had to do an evaluation to figure out what would provide the best performance for a particular repository. I think there's also a complex middle segment of options emerging now.

In addition to Opus and Sonnet from Anthropic, there's also Luna and Terra. Luna is very competitively priced. Spark is also very cheap, and version 1.2 is very functional. There is also a growing ecosystem of open-source models that companies can host themselves.

So I think that, in this uncertain situation, it's very difficult to understand what's best. We actually see that in many cases Sonnet is more expensive than Opus because it consumes too many tokens. I think if you go by the principle of “use Sonnet where you see fit,” you might end up spending more than you planned.

Erik Torenberg

Tell us more about how this evaluation system will be applied to intellectual work in other fields.

Jennifer Li

What examples can I give? I think programming is a harbinger of what's to come for every industry, and many of the basic principles established there carry over to other areas. If you have a very good programming agent, you probably have a model that can create PowerPoint presentations or financial models in Excel just as well.

I think in many of these areas we need to use existing practices as a mechanism for building assessments. As Ben said, we haven't solved the question of what human intelligence is yet, but in many industries we have huge archives of data about what work looked like. So our challenge, and the challenge of others, will be to try to turn this into an evaluation system that remains dynamic and can actually evaluate models where human work is being done.

10. How Vals Uses Vals: Token Maxing the Coding Tools

Erik Torenberg

Perhaps to supplement the previous question, how do you use LangSmith internally to evaluate the best programming model for Valse?

Jennifer Li

Yes, to be honest, this arose from a problem we also encountered. I wanted to experiment with maximizing token usage and was able to get unlimited access to some programming tools for our team for a month. Looking back, many of our engineers were spending 1 to 2 billion tokens per day. It seems that, on a peak day, one engineer spent 6 billion.

Erik Torenberg

Yeah, that's just crazy. How much is that in dollars?

Jennifer Li

So I went back to it, did the math, and it turned out that we spent about $1.5 million worth of tokens that month.

Erik Torenberg

By the way, it's free.

Jennifer Li

I know. But we actually spent 10 times more on tokens than we did on employee salaries that month. So it's not even 50/50; it's 10 times more. It was interesting to take stock and see where people used agents, because there's a kind of uncertainty around using models all the time and everywhere.

We faced the fact that we couldn't continue to operate in this mode for the next month. How do we intelligently determine which tools to use, for which teams, and for which projects? So we conducted an experiment, analyzing the work done. We reviewed many logs and traces, analyzed our GitHub repository, and created the Val Smith tool.

We discovered some pretty strange insights. For example, Cognition's Devin tool actually uses tokens very effectively. So this is something we decided to implement more often. I think there are many cases where better pricing models can be obtained through subscriptions rather than paying for tokens. This influenced our strategy for how we could effectively optimize tokens without spending $1.5 million a month.

Erik Torenberg

Very cool. So is your current mode of operation a more token-efficient toolkit and model, with people having more flexibility to use token-based payments for more complex tasks?

11. Policy: What Should the Government Actually Do?

Jennifer Li

Yes, we have access to all the tools. We give everyone access to everything. But we automatically provide recommendations for any ticket or task on GitHub about where to start the session, and this should regulate actual usage depending on the level of intelligence required for the task.

Erik Torenberg

Very cool. I want to move on to the policy aspect for a moment, since we were talking about how quickly benchmarks become obsolete. In politics, things are even worse because laws move much more slowly than the capabilities of technology.

Ben and Mark spent a lot of time in Washington, D.C., talking to politicians to try to bridge this gap. Given this, who should set these standards? Laboratories? Independent evaluators? Customers? Or maybe the government? How should this work from a policy perspective?

Jennifer Li

Yes, I think the short answer is that everyone should be involved to some extent. I think there is a benefit to having diverse perspectives. I believe the main problem is that policy discussions over the past few years have been very abstract and have lacked any substantive basis for defining exactly what policy should regulate.

So even when laboratories offer to engage a third-party company or ecosystem for testing, they don't really explain what their operating principles and mechanisms are. I see our role, especially at the initial stage, as being in evidence-gathering mode, where we can get a lot of information and empirical data about the capabilities of the models and where the risks lie. That will help build a more informed policy discussion later on.

Erik Torenberg

How do you view the division of responsibilities, given that the government has conducted the first assessments and continues to conduct them to determine what exactly is advanced technology and needs regulation? They started with some crazy idea about 10²⁶ FLOPs or something like that.

When you think about what the government should do—for example, establish a 30-day or 60-day waiting period—what exactly should happen during that period? How does this intersect with what you do? What's the best way to determine whether a model is advanced and needs to be put in a sandbox for a period of time to make sure it doesn't get access to everything? How do you see this relationship working?

Jennifer Li

I believe that there are 2 opposing forces that need to be taken into account. The first is the desire to move very quickly and ensure that government processes do not slow down the pace of technological innovation. The other is making sure that the technology being developed is in the best interests of Americans and people more broadly. So I think they are very difficult to reconcile, and often choosing one means harming the other.

My hope is that, through the evidence-gathering process, we can help policymakers better understand what technology should be to serve America's interests. The work of creating the tools to test and enforce that control will fall on the shoulders of third-party evaluators. I believe this will create a mechanism for improving assessment and testing methodologies that will truly keep pace with the development of cutting-edge technologies. This will not lag behind or slow down the pace of development.

Erik Torenberg

When you're developing your evaluations, do you think, “Can we test how easy it is to get this model to engage in bounty hacking or something like that, or to perform illegal activities?” Is that not yet in your scope, or what do you think about that?

Jennifer Li

Yes, we consider this generally within the consistency category. I think there are already cases where models undergoing cybersecurity evaluations are actually hacking bug bounties and finding other ways to get around restrictions. But we try to assess whether our models are aligned with user intent and identify cases where they exhibit the opposite behavior.

Erik Torenberg

And maybe this question is for you too, Ben. What do you think should be the correct division of labor in this area? What should government agencies control or do on their own, and where should they trust private companies? Where do you see the gap that exists today?

Ben Horowitz

Yes, I think government agencies are getting a lot of warnings, by the way, from the big labs. They say this will lead to biohacking, this will create a cybersecurity risk, and so on.

Therefore, I believe the government needs to act like this: if a model is capable of doing something, can someone force it to perform illegal actions? The government should clearly define what exactly it doesn't want to see on the market and then commission an outside organization to conduct an assessment.

The government has some idea of what it is afraid of, whether it's biohacking, cyberattacks, or something else. But then the question arises: is the model capable of doing this? And can you make a model do this? Someone has to actually evaluate these 2 possibilities.

I believe the government is not particularly equipped to perform the latter task, especially over time. This is simply an inappropriate state function. But they are very good at setting rules because they can enforce them.

Rayan Krishnan

So, I think this is the combination we need: the government sets and enforces the rules, and a competent private company reports if those rules are broken. It's quite interesting to watch big labs start saying, “The model has some capabilities, and it can be made to do something bad. Therefore, we will not give it to anyone. We will use it ourselves and make sure that our people don't force it to do anything bad.” But even that doesn't always work, so we live in interesting times, I would say.

For me, there is a gap between what is being said and what happens in the real world, because every infrastructure, especially corporate infrastructure, is very different. Simply the way people already use these models is also very different. Stories based on real examples are extremely rare, which is why we are still talking about the OpenAI hack and Hugging Face. Today, we're still discussing what happened to Fable, even though it only lasted 2 months. But often, that's not exactly how the models actually unfold.

The environment in which they work is very specific. So how do we connect these dots and again create the right environment and rule base for implementing these models? I think it also requires someone with the appropriate capabilities to take all these safeguards and adapt them to real-world conditions so that people can use the models with confidence.

Jennifer Li

Tell us more about how politicians should collaborate with evaluators. What information do they need? How should these relationships be built to be most effective?

Rayan Krishnan

Yes. First and foremost, there should be a mechanism through which findings and data are communicated directly to the relevant individuals in government. We regularly brief the executive and legislative branches on what we are learning about the opportunities and risks of the models. I find that it helps them stay on top of what's going on, as well as keep track of issues that may arise in the future. Everything is moving very, very fast, and it's hard to predict where this is all going.

But at least when you have data, you can start to predict trends. Then I think it's up to the people in the legislative branch to decide in which direction they want to see policy. It's not really our job to give such recommendations. But if they see that there is a significant risk, for example, to the mental health of people under 18 or to biohazards in models, and that requires a standardized way to restrict the release of models, then they are the ones who should shape the policy.

And I think there are other areas where the executive branch, like the Department of Commerce or the Securities and Exchange Commission, is responsible for ensuring that private companies can implement and use models productively for the entire system.

12. The Geopolitics of Evals: Whose Values Get Embedded?

Jennifer Li

I'd like to consider another aspect—the geopolitical one. I often see evaluations, or evals, as a reflection of the values of the model developer, as you create criteria for what is best in that model. And, of course, different countries and labs within those countries take care of different things.

For example, I was born and raised in China. I also use a lot of Chinese open-source models. You still can't let them talk freely about the CCP and the whole history there because, well, you know what's going on in China. So what do you think about the role of assessments and benchmarks in standardization, or how they are perceived by different model labs from different places?

Rayan Krishnan

To be honest, from my very idealistic perspective, I'm surprised to see such a large investment in sovereign AI. If I were to look at it from a bird's-eye view, it would be extremely inefficient to build all these data centers, duplicate data-provisioning processes, and train these huge models when, in fact, we could consolidate many of these efforts. But it seems that we are not living in such a world and are not moving in this direction. In fact, there is a growing effort to create AI at the sovereign level.

So I think that requires a common language to communicate about what the assessment framework is and where we are going to collectively agree on positions on risks. I think there's a lot to learn from the example of nuclear energy. I think Reagan had this phrase: “Trust, but verify.” So I think we're starting to see signs of trust because Xi Jinping and Trump are going to meet next month. But there is no clear way to actually perform the verification part of this process.

Having a common language of assessments will allow us to say things like, “You have the right number of nuclear warheads.” In that example, there were also flybys—the next step, through which a country could check another country's nuclear arsenal through overflights. Similarly, if there is a concern about the societal or even existential risk of AI, it would require us to build a common language of assessments to perform the verification process.

Jennifer Li

What do you think about how to harmonize such policies? It's pretty hard to do even in America. How do you think about taking this to a global level? Now you're not dealing with corporate clients; you're dealing with governments, and those governments are competing with each other. How do you think this should work?

Rayan Krishnan

It would be naive of me to say that I have the perfect solution to this problem today. So I think there are first steps we can start with. For example, there seems to be a lot of talk about cybersecurity risks. I think biosecurity concerns will become even more important over time. So there are clear areas where there will be a mutual interest in agreeing on ways to prevent conflicts in the area of cyber or biosecurity.

In my opinion, I think the most interesting thing in the long run will be the possibility of recursive self-improvement. This is an area where you can see one country or one company leaping forward and creating models that we know little about or that work in ways that are unknown to us. So I think having a common way to describe this as a level of pace that we agree on, or as something that exceeds the pace of development, as far as recursive self-improvement is concerned, is going to be extremely important. That's why I think you see a lot of researchers at Closer Us Labs today calling for joint negotiations between governments.

13. What the Benchmarking Landscape Looks Like Next

Jennifer Li

What do you think the future will look like, now that we have so many different opportunities and powerful models, and countries that are concerned with different aspects of development? Everyone, of course, cares about the security of AI, but in biotechnology or cybersecurity, people are interested in slightly different things, depending on whether we are talking about offensive or defensive capabilities. What do you think this landscape will look like, and how do you approach developing new tests to keep up with these changes?

Rayan Krishnan

Yes, we are focused on creating tests that cover the cutting edge of technology. So when we see new areas of opportunity or risk, we want to turn them into a well-documented assessment on the valis.ai platform. I believe that this requires a continuous expansion of scope over time.

For example, in cybersecurity, much of our previous work has focused on code vulnerabilities or memory leaks that may exist within it. However, in reality, many of the biggest problems or risks lie at the infrastructure level. These are things that cannot be expressed in code alone; they require modeling larger-scale environments of corporate cloud infrastructure or even energy grids so that we can assess the real offensive or defensive capabilities of the models. So making sure our assessments reflect these new areas is extremely important to our work.

We believe that the most valuable model for this business will be one where incentives are aligned with conducting quality assessments, rather than supporting the process of developing intelligence or the methods by which models can improve in this direction.

Jennifer Li

Perfect. Thank you for coming to the podcast. This was a great episode.

Rayan Krishnan

Thank you very much for inviting me.

Erik Torenberg

Thank you very much, Rayan. Thank you, Ben. That was interesting.

Ben Horowitz

Thank you.