[评测现状] LMArena 的 17亿美元愿景——Anastasios Angelopoulos,LMArena
Arena 融资 1亿美元买来的是战略重试权,而不是烧钱授权。 Anastasios Angelopoulos 称,如果第一笔下注失败,资本就是可以翻开的“牌”;眼下的支出包括为免费推理买单、招聘,以及从 Gradio 迁移到 React,但他强调,Arena 不必把这笔融资全部花掉。
Arena 的护城河来自原生使用的规模与真实度,而不是一套静态基准测试目录。 主持人将其社区规模描述为 500万 MAU;Angelopoulos 提到,平台上已有约 2.5亿次对话,每月新增数千万次。他还表示,约 25%的用户以软件为业。不同于围绕预生成输出搭建的竞技场,Arena 记录用户提出的真实问题,持续刷新评测分布。
公开排行榜是建立可信度的亏损引流产品,并明确承诺不 pay-to-play。 Angelopoulos 称其既是“慈善项目”,也是“亏损引流产品”:无论是否付费、得分多少,已发布模型都会出现,模型提供商也不能付费要求下架。他对“排行榜幻觉”批评的回应是,相关分析存在事实错误,包括声称开源模型抽样率为 9%,而 Arena 的实际构成约为 60/40;同时,该分析也误读了长期运行的预览模型测试。
Nano Banana 展现出图像生成的经济吸引力后,嘉宾与主持人都扭转了此前的怀疑。 Angelopoulos 如今预计,多模态系统将成为 AI 在消费端和企业端最具经济价值的能力之一,营销与设计是增长最快的采用领域。主持人的案例是:把 DeepSeek V3.2 对强化学习环境的解释输入 Nano Banana Pro,得到一张论文级图示;这项工作过去可能需要一个博士生花上一个月。
Arena 正从单一综合排名扩展到职业、模态和智能体专属评测类别。 其庞大用户群中,已有个位数比例分别构成医疗、法律、金融、会计、创意和营销等群体;视频评测计划于今年晚些时候或明年初推出。Code Arena 也可能从评估模型,进一步比较 Devin 这类完整智能体 harness。
用户留存与创业公司的聚焦能力仍是执行约束。 持久化历史记录显著推动了用户登录,但 Angelopoulos 表示,“每个用户都要靠自己赢得”,用户在一次“昙花一现”的爆发后仍可能随时离开。API 仍有可能推出,但 Arena 当前应对战略发散的答案很简单:“我们确实应该把一件事做好”——竞技场。
1. 公司成立:把学术基准变成基础设施
Angelopoulos 最终判断,Arena 无论作为学术项目还是非营利组织,都无法获得所需的分发能力、平台质量和运营规模。其使命早已明确:通过真实用户与原生反馈,“衡量、理解并推动前沿 AI 能力”;成立公司,成为将这项使命规模化落地的实际组织形式。
Anjney 在创始团队决定创业前孵化了 Arena,提供早期资助与资源;Sequoia 也提供了一笔资助。Anjney 设立实体时,还做出了一个不同寻常的承诺:团队随时可以退出。Angelopoulos 最终同意,成立公司是“我们实现规模化的唯一方式”。
1亿美元融资买来的是选择权:“一家公司拿到钱的目的,是给自己留下可以翻开的牌。”Arena 为平台上的免费使用承担全部推理成本,并持续招聘;此前支撑其达到约 100万 MAU 的 Gradio,如今已迁移至 React,以获得更丰富的组件,以及更容易吸引熟悉该技术栈的开发者人才。
2. 原生提示词是 Arena 的核心数据优势
主持人将 Arena 社区规模描述为 500万 MAU。Angelopoulos 则单独提到“超过 500万”,但没有说明对应指标;平台累计约有 2.5亿次对话,每月新增数千万次。他表示,约 25%的用户以软件为业,如今约一半用户会登录,使 Arena 能够结合问卷研究用户行为——但他明确保留了样本存在回答偏差的限制。
主持人转述了 Artificial Analysis 打造“AI 界 Gartner”的目标。该机构将公开基准测试汇总并独立重跑,产出分析与报告;Arena 则让用户输入自己的使用场景和问题,而不是仅仅评判预先生成的输出。
主持人提出了另一面:策划好的示例可以让不擅长提示词的用户了解模型能做到什么。Angelopoulos 同意,查看他人的提示词具有教育意义,但他强调,Arena 让用户输入自身使用场景,才是其真实度的来源。
3. 榜单公正是 Arena 拒绝出售的资产
《Leaderboard Illusion》论文称,预览模型测试制造了未披露的不平等。Angelopoulos 认为这一批评“不科学”,并指出其中一些事实错误已经得到修正;具体而言,论文声称开源模型抽样率约为 9%,而他表示实际分布更接近 60/40。
他为预览模型辩护,依据的不只是统计学,也包括社区文化:Arena 长期以来都会以秘密代号向用户开放预发布模型,社区乐于自行发现它们。Nano Banana 就是在这一机制中出现,随后成为“全球现象”。
Angelopoulos 表示,Nano Banana 的那一刻本身就改变了 Google 的市场份额。主持人认为,它在当时肉眼可见地领先于其他产品,并将其与 Google 股票中数十亿美元的资金流动联系起来。
主持人更尖锐的质疑是,并非每个预览模型都会登上排行榜。Angelopoulos 的回答是:尚未发布的模型不必出现,但每个已发布模型都会基于数百万张选票获得统计上可靠的分数。模型提供商不能购买上榜资格,也不能付费要求下架;排行榜必须继续成为“透明、公平反映模型表现”的工具。
4. 多模态模型改写了双方对经济价值的判断
主持人承认,自己曾认为图像生成是 AGI 的边缘能力,还可能带来声誉风险:AI 的上行空间为何不集中在语言、编程和推理上?Nano Banana 改变了他的看法;Angelopoulos 也表示:“我对这件事的判断也有些错了。”
他们修正后的判断是务实的,而非哲学性的。营销与设计是 AI 采用速度最快的领域之一;创作者则获得了“无限供应”的图表、解释图和信息图,使多模态系统有潜力成为 AI 在消费端和企业端最具经济价值的能力之一。
最具代表性的案例是 DeepSeek V3.2:主持人把其对强化学习环境的解释输入 Nano Banana Pro,生成了一张帮助自己理解论文的图示。他认为,如果手工完成同等工作,过去可能要一个博士生“花上一个月”。
5. Arena 扩大评测范围,但不放弃聚焦
Angelopoulos 希望 Arena 继续成为行业的“北极星”:一个持续更新的基准,通过不断进入的新数据点抵抗过拟合。Arena 跟踪新模型与新使用场景,并已发布数百万条真实对话,供研究人员研究和改善真实世界表现。
职业视图如今可以分别展示医疗、法律、商业、金融、会计、创意和营销用户的结果。在 Arena 的规模下,即使个位数百分比也足以形成有意义的用户群体;视频评测预计在今年晚些时候或明年初推出。
API 仍有可能推出,但 Angelopoulos 从创业公司聚焦的角度对其提出疑问:公司应该把“一件事做好”。在留存方面,持久化历史记录提高了登录率,但任何功能都无法免除一项基本义务:每天都要“赢得每一个用户”。
他还在寻找消费产品、机器学习、B2B 市场进入、营销及相关领域的专家加入 Arena。
Code Arena 可能将评测单位从模型扩展到完整的智能体 harness。在一项潜在的 Cognition 合作中,Angelopoulos 提议把 Devin 放进 Arena,测试它是否“在所做的事情上成为全球最强或最强之一”,以回应外界关于 Devin 已经消亡的说法。
All right, we're here with Anastasios from Arena. I actually don't know your last name.
Angelopoulos.
Yeah, there you go. Congrats on all the success. You got the Arena handle.
Yeah, we did. We got the Arena handle. Thank you. Big branding moment.
I think X is becoming more commercial. Obviously, you bought it, but at least you have a place to go where you can be like, “Hey, we really like this.” But I do think dropping the LM has changed the feel of it.
I don't know. The reason we kept the LM at the beginning is because we started as LMSYS, right, out of the LMSYS organization at Berkeley. So we decided—
Language models.
Exactly. So we wanted to maybe broaden a little bit. We were the first Arena, so we felt like, “Let's kind of try to own that.”
Last time we had you guys on, you hadn't really spun out yet. I did a call with Alessio and I was like, “These guys are going to start a company.” [laughter]
I didn't know. I think you actually were already started at the time.
I don't remember. Maybe because I had chatted with Anjney, and he said he was your founding CEO.
He was indeed, which people don't know.
Anjney is a very interesting character. We have a podcast scheduled with him.
Yeah. Yeah.
He does a lot more than normal VCs.
He does. He's been incredible to us.
Do you want to shout out some of the stuff that he did?
Yeah, absolutely. The way the company started was as an incubation by Anjney.
Yeah. What he did was find us at Berkeley and pick us out of the basement. He was like, “Hey, these guys seem like they're onto something.”
He started working with us really early and gave us some grants. a16z was not the only one to do this; we also had a great grant from Sequoia. But Anjney was particularly supportive of us and gave us some resources to continue building out Arena before we were even committed to starting a business.
In that capacity, he formed an entity for us and said, “Hey, you guys can walk away at any time if you don't want to start a business.” It was really incredible—a very aggressive investment move by him. [laughter]
Right, because any money that he spent, at the end of the day, we could walk away and leave him with nothing.
But I think he wisely knew that the right thing for Arena was to start a company out of it. It was the only way that we could scale, and he knew that Wayne Yan and I would ultimately see that and be excited about doing it ourselves, which ended up being the case.
Was there a moment for you where you decided? I'm sure you were debating it yourself, and you had other opportunities. What was the deciding factor for you?
It became clear that the only way to scale what we were building was to build a company out of it. The world really needed something like Arena: a place to measure, understand, and advance frontier AI capabilities through real-world usage with real-world users, based on organic feedback.
In order to achieve the scale and distribution necessary—and, of course, the quality of the platform necessary—to do this effectively, we would need to start a company out of it. We considered other options: Were we going to keep doing this as an academic project? Were we going to do it as a nonprofit? But ultimately, under those constructs, we didn't feel like we'd have the resources necessary to accomplish our mission.
So you raised $100 million, or $80 million?
$100 million. That's a lot of resources. It's great.
What's it for?
Well, obviously—
On behalf of everyone—
Yeah, everybody—
Dude, it's an arena. How are you going to spend your money? [laughter]
First of all, we don't necessarily need to spend all that money, right? The purpose of money at a company is to give you cards to flip. It's to say, “Hey, you have enough resources so that if your first bet fails, you can make another bet and another bet.”
Of course, that's not to say that we're going to spend all of it. You want to spend things responsibly. Having said that, the platform is actually quite expensive to run. We fund all of the inference on the platform. The way it works is the platform—
You pay market rates. They don't give you discounts.
No, no, we get discounts, but they are standard enterprise discounts, the same that would be given to any other customer.
Have you disclosed any numbers? I see numbers of votes, but what's that in monthly tokens?
I don't know about tokens. I'd have to back that out. We have, let's say—again, this is off the cuff, but I can safely say—we have more than 5 million now, 5 or 6 million. We have probably 250 million conversations that happen on the platform. We're on the order of mid-tens of millions of conversations every month happening on the platform.
It's actually quite a large consumer platform for LLMs. Of course, nothing really compares to ChatGPT, but—
But no, still, I think the largest-scale ones—
The benefit of this is that it's actually quite a diverse population. For example, 25% of the people on our platform do software for a living.
Still, at this scale, how do you know?
We do all sorts of things. We either survey them or analyze the prompt distribution that's coming into the platform. I'm happy to share more. We've done something called Expert Arena, which is trying to understand the distribution of experts that are coming to the platform.
A lot of that can be unauthenticated, or whatever the usage is. But about half of our users now are logged in.
Yeah.
So we have some ability to understand them. We also have surveys that we run on the platform that tell us a bit more about who the actual users are. Of course, there's always response bias in surveys, so you have to take it with a grain of salt. Nonetheless—
If you don't know Anastasios's background, he's the guy to correct for response bias.
Okay. Yeah, there are a lot of guys like that—and girls.
Guys and girls. You're not the only player. There's Artificial Analysis. They started an arena. It's like some crypto people who started this. I don't know if you've had a conversation with them, like, “Hey, this is our thing,” or, “Let's work together on something.” I don't know if you've—
No, I've talked to Artificial Analysis. I've actually talked to both groups. Both seem—
Am I missing any major players? Is it just those two?
No, I think those are some of the larger ones. Depending on how you define the term, Artificial Analysis obviously has huge market mindshare around the analysis—
Of different AI systems.
Yeah, they told me they were going to be the Gartner of AI.
Yeah. That's kind of their goal, and I think they're going after that consulting market and so on. Artificial Analysis, from what I understand, is a team of consultants doing this.
Yeah.
They seem like really, really nice guys going after that particular market.
Their analysis is based on aggregating public benchmarks and turning those into analytics—
Independently rerunning.
And independently rerunning, which matters.
Yes. Using those to compile reports and so on that educate the field on the performance of all these different models.
But they also have arenas.
They have arenas, but the arenas are not based on organic usage. The thing that distinguishes our platform from theirs is that users are actually inputting their own use case. They're asking their own question. That gives a level of realism that their platform doesn't have.
Of course, they specialize in a slightly different thing, but I see those platforms diverging in that sense.
Yeah, and sometimes it's the only way to do this. For example, for Artificial Analysis, their video arena is pre-generated videos. You can't enter your own video.
That's correct. But we are doing it organically.
Yeah. Exactly.
As a voter, it does help in terms of not having to wait.
It does, but also, why would you go?
Do you actually care about other people's videos? Form your own intuition.
Maybe. Yeah. Maybe you're interested in comparing.
I'm a shitty prompter, right?
Are you? I don't believe that.
I'm a terrible prompter. I learn by example.
Don't denigrate yourself. Don't denigrate yourself.
There are many people whose prompts are much better than mine. Let's say that. That's a fact.
People have all sorts of cool ways of prompting. Oh, yeah. It's educational to see.
The only way to learn is by looking at other people’s prompts, seeing the results, and thinking, “Oh, I didn’t know you could do that.”
Totally. Yeah. Okay, so let’s go back to Arena. One thing I do want to say is that the number-one use of funds is getting off Gradio.
Oh, yeah. Yes, we—well, listen: Gradio is an incredible platform. Gradio scaled us to a million MAU.
Yeah, that’s incredible. And, of course, you tell the Hugging Face fans that—
Of course. Yeah, we’re really grateful to Gradio for taking us this far. Eventually, it became time for us to move off of it and go to React.
I’m sure Hugging Face would have loved you to stay on.
They would have. I’m sure.
Was there a technical reason? You just couldn’t get the performance?
Yeah, it just became hard to develop, and there were all these tools that we wanted in React. To do all the fancy things that you can do in React became kind of difficult. One example—I don’t know, what’s a feature that you really wanted?
Let’s say we wanted to create our own custom loading icons for video with notifications.
Okay.
How are we going to do that in React? It’s hard.
Yeah, I mean, you can make a custom component. I’m sure the Hugging Face guys are going to come in and say, “You can do that in Gradio,” which maybe you can. But fewer developers know how to do that. How are we going to hire for that? We’d have to reskill them, and people are less familiar with that stack.
So it’s all React, all that?
Yeah, all that.
Okay, cool. Other uses of funds that might be interesting?
Basically, that’s compute resources. It’s primarily inference; that funds the free usage of the platform, and then also hiring, of course—headcount. Yeah, we have an office. You know, it’s in San Francisco.
I’ll tackle one of the major things this year, which I’m sure you’re tired of thinking about, but for people who are not in the loop, this is going to be news to them: the Leaderboard Illusion, the whole thing with Cohere. Let’s summarize what they said and then your response.
Mhm. The Leaderboard Illusion is a paper that critiques Chatbot Arena, and the main—
Pretty brutally.
Well, I would say unscientifically. [Laughter] Let’s be clear: Cohere wasn’t doing that well on—
Cohere was around 74.
It’s all good. It’s actually a respectable place to have on the leaderboard. I don’t even think it was really Cohere—the model developers—doing this. It was more their research side. But in any case, what does the Leaderboard Illusion say? It says that Chatbot Arena was doing undisclosed, quote-unquote, private testing on our platform: model providers would send us prerelease models, and we would expose them, and so on and so forth. The claim is that this creates inequities in the leaderboard due to that prerelease testing. For example, they cited that Meta had, at some point, tested a number of models with us.
Of course, we can’t disclose all the details of how all that was done.
But that is the main claim of the paper. Our response to that paper is online; you can find it by searching for “Response to Leaderboard Illusion.” Our response is essentially pointing out a series of factual mistakes in the paper that question the validity of its claims. You can look at the first version of the paper on arXiv yourself and see the claims. I think most scientists would view that as—
Oh, they’ve corrected it.
They corrected it, of course, but they didn’t correct everything. They just corrected some aspects that were blatantly unscientific and false. For example, they said that we only sampled around 9% open-source models and around 60% closed-source models, and that this created a gap between open and closed source. In reality, we’re very supportive of open-source models, and it was more like 60/40. That was one example of an error in the claims.
Another example is that they claimed there was some sort of bias introduced by this prerelease testing and that it was undisclosed. In reality, as you probably know, we’ve been doing this prerelease testing for a long time. Our community loves it. They love getting the secret code names.
Yeah, the secret code names—Nano Banana and all that. So Nano Banana, by the way, started on you.
It started on us, right? And people loved it. It became a global sensation. A nonzero fraction of the global population was using it. About naming it Banana, was that their decision?
It was their decision, I believe. It was sort of randomly generated, and it just went—
No, no. So apparently, Nana, who’s a PM—
Yeah, yeah.
—is named after her because her nickname is Nana.
Oh, that’s sweet. I didn’t know that.
Nana put Banana on it.
Oh, that’s sweet. Yeah, I didn’t know that. I didn’t know the origin story.
To us, it just looked like a random thing.
But it was clearly head and shoulders above everything. Before that, there was Recraft Image, remember?
Yeah, I do. Of course. And also BFL and all those. All those models are great, and I think those teams are also improving quite quickly. But Nano Banana was a sensation. That moment alone changed Google’s market share.
Yeah, market share.
Yeah.
Seriously. Google’s stock—billions of dollars are moving because of Nano Banana.
And now there’s an OpenAI code red and everything.
I don’t know about that, but The Information reported this.
I would say image generation has been this weird part of AI overall because it’s not strictly AGI-critical. It’s not reasoning, and it’s not feeding more context into the model. It’s the model generating a visual representation. I always think, well, Gemini used to get a lot of complaints for generating racist images or whatever.
That was a hilarious moment, and ChatGPT also had it in the past. I’m like, can we just get rid of this? Do we have to do image generation? Let’s focus the positive reputation of AI in general on language models and coding and the other stuff.
Yeah.
But I’m wrong. [Laughter] I’m such a huge Nano Banana Pro shill.
Yeah, I totally agree. I was also kind of wrong about this. I didn’t see the positive benefits, but actually, I think these multimodal models are going to become some of the most economically valuable aspects of AI, both in the consumer space and in enterprise. One of the fastest-growing market segments in AI adoption is marketing and design.
Yeah, ads. I’m a content creator, right?
Yeah, of course. I’m sure you’re using it all the time.
Infinite supplies of diagrams and explainers and infographics.
Yeah. Soon, we’re not even going to be making them for our papers. Our paper figures are just going to be made by—
Yes. I do think they actually one-shot. DeepSeek came out with V3.2 recently. I took their explanations of the RL environment stuff and fed them into Nano Banana Pro, and it generated an image that helped me understand the paper better. The fact that I can casually generate a paper-quality diagram that would usually take a PhD student a month in Photoshop or something is incredible.
Yeah, it’s incredible. It is amazing.
I want to ask about your principles for running Arena. I think you manage a giant community—5 million MAU. What have you decided are the core principles, both before becoming a company and now that you’re a company? I don’t know if anything has changed for you.
I don’t think anything has really changed. We want to provide the north star of the industry and center the use cases of real users, foregrounding those so that people know what to target. The goal is to create a benchmark that is constantly fresh and does not suffer from overfitting because we constantly have new data points coming in. It tracks all the different new models and all the different new use cases of AI, and gives the whole world a sort of ground truth for how real users are using these models and how good they are on those use cases.
We continue to do quite a few open-source data releases. We’ve probably released more data than basically anybody on the real-world use cases of AI: millions and millions of conversations from real users that the community is using to study and improve on.
In terms of what you will build versus what you will not build, I’m not necessarily caught up on everything that you’ve launched. I know recently you’ve done the Dev or Code Arena.
Yeah, Code Arena.
Expert Arena.
Yeah, Expert Arena.
What’s in the critical path for you, let’s say, for next year? What have you decided you’ll never do?
Let me first talk about things that I’ll never do on the platform.
Integrity comes first to the platform. Basically, the public leaderboard that we show on LMArena, I think of as a charity. It’s a loss leader for us. We don’t really make money on the public leaderboard. You can’t pay to get on the public leaderboard.
It’s not like a Gartner in that sense. It’s not like any of these pay systems. It’s never going to be like that.
Models are going to be listed on the leaderboard—
Whether or not the providers pay and whether or not they’re getting a good score. They can’t pay to take it off either. And so, what that means is very important.
And so, what that means is that the leaderboard has a certain integrity that will never be compromised, of course.
Yeah. But not all preview models will make it onto the leaderboard.
No, but that’s okay. Those preview models have never been released.
Yeah, right. Who cares about putting unreleased models on a leaderboard? The point is that for every released model, the score that you see on the leaderboard is statistically sound. It reflects the real-world capabilities of the model.
Yeah. Why? Because millions of people from around the world have voted for it, and that’s where that number comes from. All we do to compute that number is let millions of people vote. We take those votes and turn them into a number that’s always going to remain a transparent and fair reflection of model performance.
Where are we going? Lots of different new categories. I don’t know if you recently saw that we exposed occupational and expert categories. Now, a single-digit percentage of our user base—we’re at millions to tens of millions of users, right? So a single-digit percentage means a lot.
A single-digit percentage of our user base is in medicine, legal, business, finance, accounting, creative, marketing, and things like this. We’re able to show the performance of these models in all these different verticals because we have all these users in our user base. We’re working more toward multimodal. Video is soon to launch on the site, at some point later this year or early next year. So, there are lots of things in the pipeline.
Amazing. Will you expose an API?
We’ve thought about it. Yeah, I think it’s a possibility.
Yeah. What are the counterarguments? Why not?
Well, there’s obviously a need for an API. The question is more one of focus for our company, just because we’re a startup, and so we really should be doing one thing well—
Arenas.
Yeah, arenas. So, I’m not sure how far we want to splay out, or on what timeline we’d want to do that.
Yeah. Any other community-management tips, more broadly? Every AI company really wants to grow its community. You’re obviously one of the strongest in the world. What’s really worked?
Well, first of all, I want to give a shout-out to our community manager, Greg, who is doing an awesome job managing our community, whether that’s on Discord or on LMArena. He’s really incredible. So, I would say, hire Greg—but don’t hire Greg.
Don’t hire Greg.
Don’t hire Greg. He’s ours. Find a Greg. I agree.
But in general, the question is: How do you get to so many users?
That is a tough question.
And keep and retain them.
That is a tough question because consumer is one of the hardest markets in the world. There are a lot of websites in the world that people can go to. Why should they go to yours? The reality is, if you want to create a really dominant product, you have to provide people value.
To be frank, I don’t think we’re all the way there yet. It’s not like I have the solution and answer for how to build a great consumer product. If I did, we wouldn’t be at tens of millions of users. We’d be at hundreds of millions, or we’d be at a billion users. We’d be like—
Is there a world where you’re bigger than ChatGPT?
I don’t know. I don’t know that we need to be.
Yeah.
And I don’t know that we ever will be, because that’s an extraordinary generational product that they built, right? It took a lot of time, and to some extent it also involved luck. There are a lot of lightning-in-a-bottle moments, like Nano Banana was for us, where our user base just goes up by a lot.
But when those users come, they can just as easily leave. The way I think about it is that every user is earned. You have to earn them every single day. They can leave at any moment. They’re fickle.
All the time, you have to be thinking about: How do I provide this person value? Learning how they’re using my website, what more could I give them, and how do I build in all the retention mechanisms so that they stay and then also bring their friends?
Is there one thing that’s working in terms of retention? You said a lot of people are signing in now.
Yeah, sign-in was a big driver of retention.
No, but what did you give them in order to encourage that?
Persistent history.
That’s it? That’s enough?
Yeah, that’s one thing that has had a big impact.
Okay. Yeah. What do you want from people? What are you looking for help on? Any call to action?
Yeah, we are always looking for people to come and join us. If you are one of the best people in the world in your area, whether that’s consumer product, machine learning, B2B go-to-market, marketing, or any of these things, we need you at Arena.
We’re building a high-performance team of real experts in everything that they do, and I’m always looking for excellent people to work with.
What about partnerships? Let’s say I’m at Cognition. I want to partner with LMArena, or just Arena. What works for you? What existing partnerships do you already have that are really fruitful?
Yeah. So, we of course partner with all the major model labs.
Yeah.
And that’s straightforward—
And that’s straightforward: “Hey, we have a new model here. Here you go.”
Exactly. So, I think the most straightforward thing for someone like Cognition would be: Let’s evaluate that—
Agent.
Yeah. But we should be continuing to shape our—well, Code Arena is an agent evaluation.
That’s true. That’s true.
And it’s more focused on— all these arenas tend to focus on the model rather than the harness. But maybe that should change.
Maybe we should be evolving in that direction.
And I think Code Arena is a good example of an arena that would support—
A full-featured harness like Devin.
Yeah. And so, in my view, if I’m talking to Cognition, I’m saying, “Hey, let’s get Devin on the arena and figure out how to loop the harness together so that we can—” I’m sure there’s something that could be really valuable there, especially given Devin.
Last week, people were talking about Devin being dead.
Did you see that?
Yeah. People were saying Devin’s gone. Devin’s not gone.
Devin’s everywhere.
Well, people were saying Devin was dead. So, can we highlight that for people and show them, “Hey, Devin is actually the best, or one of the best in the world, at doing what it does”?
LMArena can actually do that. Our place as a central evaluation platform allows that to happen.
Yeah, love it. All right. Thank you for owning the state of the art.
Thanks so much, and congrats on a wonderful year.
Appreciate it. Congrats to you, too. Congrats on all the growing momentum in your podcast and in your career. Thank you.
It’s really impressive to see.