[BidClub_]
Latent Space · · 24 min

[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena

Anastasios Angelopoulos

YouTube
TL;DR
  • Arena’s $100 million raise buys strategic retries, not a mandate to burn. Anastasios Angelopoulos calls capital “cards to flip” if the first bet fails; immediate costs include funding free inference, hiring, and replacing Gradio with React, but he stresses that Arena need not spend the entire raise.

  • Arena’s moat is the scale and realism of organic usage, not a static benchmark catalog. The host described the community as 5 million MAU; Angelopoulos cited roughly 250 million conversations on the platform and mid-tens of millions monthly. He also said roughly 25% of users do software for a living. Unlike arenas built around pre-generated outputs, Arena captures users asking their own questions, continuously refreshing the evaluation distribution.

  • The public leaderboard is a credibility-building loss leader with an explicit no-pay-to-play covenant. Angelopoulos calls it both a “charity” and a “loss leader”: released models appear regardless of payment or score, and providers cannot pay for removal. His answer to the “Leaderboard Illusion” critique is that its analysis contained factual errors—including a claimed 9% open-source sampling rate versus Arena’s roughly 60/40 mix—and mischaracterized long-running preview testing.

  • Both guest and host reversed their skepticism on image generation after Nano Banana demonstrated its economic pull. Angelopoulos now expects multimodal systems to become among AI’s most economically valuable consumer and enterprise capabilities, with marketing and design among the fastest-growing adoption segments. The host’s specimen was feeding DeepSeek V3.2 explanations of RL environments into Nano Banana Pro and receiving a paper-quality diagram that might once have taken a PhD student a month.

  • Arena is expanding from one aggregate ranking into occupational, multimodal, and agent-specific evaluation categories. Single-digit shares of its large audience already represent medicine, legal, finance, accounting, creative, and marketing cohorts; video is planned for later in the year or early the next. Code Arena could also evolve from evaluating models toward comparing full harnesses such as Devin.

  • Consumer retention and startup focus remain the execution constraints. Persistent history made sign-in a meaningful retention driver, but Angelopoulos says “every user is earned” and can leave at any moment after a “lightning-in-a-bottle” spike. An API remains possible, though Arena’s current answer to strategic sprawl is simple: “We really should be doing one thing well”—arenas.

Digest · the substance, structured for research

1. Company formation turned an academic benchmark into infrastructure

  • Angelopoulos’s deciding premise was that Arena could not reach the necessary distribution, platform quality, or operating scale as an academic project or nonprofit. The mission was already clear: “measure, understand and advance frontier AI capabilities” through real users and organic feedback; the company became the practical scaling structure.

  • Anjney incubated Arena, providing early grants and resources before the founders had committed to a business; Sequoia also provided a grant. Anjney formed an entity with the unusual assurance that the team could walk away. Angelopoulos ultimately agreed that forming a company was “the only way that we could scale.”

  • The $100 million raise is optionality: “The purpose of money at a company is to give you cards to flip.” Arena funds all inference for free platform usage, hires headcount, and migrated from Gradio—which carried it to roughly 1 million MAU—to React for richer components and a developer pool more familiar with that stack.

2. Organic prompts are Arena’s core data advantage

  • The host framed Arena’s community as 5 million MAU. Angelopoulos separately cited “more than 5 million” without specifying the metric, about 250 million conversations on the platform, and mid-tens of millions of conversations each month. He said roughly 25% of users do software for a living, while approximately half now log in, giving Arena more ability to study behavior alongside surveys—though he explicitly preserves the caveat of response bias.

  • The host relayed Artificial Analysis’s “Gartner of AI” ambition. That group aggregates and independently reruns public benchmarks into analytics and reports, whereas Arena lets people enter their own use cases and questions rather than merely judge pre-generated outputs.

  • The host supplied the counterpoint that curated examples can teach weak prompters what is possible. Angelopoulos agreed that viewing other people’s prompts is educational, while distinguishing Arena’s own-use-case input as the source of its realism.

3. Leaderboard integrity is the asset Arena refuses to sell

  • The “Leaderboard Illusion” paper alleged that preview-model testing created undisclosed inequities. Angelopoulos called the critique “unscientific,” pointing to some corrected factual errors and, specifically, its claim of roughly 9% open-source sampling against what he says was closer to a 60/40 distribution.

  • His defense of previews is cultural as well as statistical: Arena has long exposed prerelease models under secret code names, and its community enjoys discovering them. Nano Banana began there and became a “global sensation.”

  • Angelopoulos said Nano Banana’s moment alone changed Google’s market share. The host argued that it was visibly ahead of everything else and linked it to billions of dollars moving in Google’s stock.

  • The host’s sharper objection was that not every preview reaches the leaderboard. Angelopoulos’s answer: unreleased models need not appear, but every released model receives a statistically sound score derived from millions of votes. Providers cannot buy inclusion or pay for removal—the ranking must remain a “transparent and fair reflection of model performance.”

4. Multimodal models changed both speakers’ economic forecasts

  • The host admitted he once viewed image generation as peripheral to AGI and reputationally troublesome: why not concentrate AI’s upside on language, coding, and reasoning? Nano Banana changed his view; Angelopoulos likewise said, “I was also kind of wrong about this.”

  • Their revised thesis is practical rather than philosophical. Marketing and design are among the fastest-growing AI adoption segments, while creators gain an “infinite supply” of diagrams, explainers, and infographics—making multimodal systems potentially among AI’s most economically valuable consumer and enterprise capabilities.

  • The best specimen was DeepSeek V3.2: the host fed its RL-environment explanations into Nano Banana Pro and generated a diagram that helped him understand the paper. Producing comparable work manually, he argued, might once have taken a PhD student “like a month.”

5. Arena’s roadmap broadens evaluation without abandoning focus

  • Angelopoulos wants Arena to remain the industry’s “north star”: a continuously fresh benchmark that resists overfitting because new data points constantly enter. Arena tracks new models and use cases and has released millions of real conversations so researchers can study and improve real-world performance.

  • Occupational views now expose results for medicine, legal, business, finance, accounting, creative, and marketing users. Even single-digit percentages become meaningful cohorts at Arena’s scale; video evaluation is expected later in the year or early the next.

  • An API is possible, but Angelopoulos questions it on startup-focus grounds: a company should do “one thing well.” On retention, persistent history increased sign-ins, but no feature removes the basic obligation that “every user is earned” each day.

  • He is also looking for experts in consumer product, machine learning, B2B go-to-market, marketing, and related areas to join Arena.

  • Code Arena may extend the unit of evaluation beyond a model to the full agent harness. In a potential Cognition partnership, Angelopoulos proposed putting Devin into Arena and testing whether it is “the best, or one of the best in the world, at doing what it does,” in response to people saying Devin was dead.

Speaker 1

All right, we're here with Anastasios from Arena. I actually don't know your last name.

Anastasios Angelopoulos

Angelopoulos.

Speaker 1

Yeah, there you go. Congrats on all the success. You got the Arena handle.

Anastasios Angelopoulos

Yeah, we did. We got the Arena handle. Thank you. Big branding moment.

Speaker 1

I think X is becoming more commercial. Obviously, you bought it, but at least you have a place to go where you can be like, “Hey, we really like this.” But I do think dropping the LM has changed the feel of it.

Anastasios Angelopoulos

I don't know. The reason we kept the LM at the beginning is because we started as LMSYS, right, out of the LMSYS organization at Berkeley. So we decided—

Speaker 1

Language models.

Anastasios Angelopoulos

Exactly. So we wanted to maybe broaden a little bit. We were the first Arena, so we felt like, “Let's kind of try to own that.”

Speaker 1

Last time we had you guys on, you hadn't really spun out yet. I did a call with Alessio and I was like, “These guys are going to start a company.” [laughter]

Anastasios Angelopoulos

I didn't know. I think you actually were already started at the time.

Speaker 1

I don't remember. Maybe because I had chatted with Anjney, and he said he was your founding CEO.

Anastasios Angelopoulos

He was indeed, which people don't know.

Speaker 1

Anjney is a very interesting character. We have a podcast scheduled with him.

Anastasios Angelopoulos

Yeah. Yeah.

Speaker 1

He does a lot more than normal VCs.

Anastasios Angelopoulos

He does. He's been incredible to us.

Speaker 1

Do you want to shout out some of the stuff that he did?

Anastasios Angelopoulos

Yeah, absolutely. The way the company started was as an incubation by Anjney.

Speaker 1

Yeah. What he did was find us at Berkeley and pick us out of the basement. He was like, “Hey, these guys seem like they're onto something.”

Anastasios Angelopoulos

He started working with us really early and gave us some grants. a16z was not the only one to do this; we also had a great grant from Sequoia. But Anjney was particularly supportive of us and gave us some resources to continue building out Arena before we were even committed to starting a business.

In that capacity, he formed an entity for us and said, “Hey, you guys can walk away at any time if you don't want to start a business.” It was really incredible—a very aggressive investment move by him. [laughter]

Speaker 1

Right, because any money that he spent, at the end of the day, we could walk away and leave him with nothing.

Anastasios Angelopoulos

But I think he wisely knew that the right thing for Arena was to start a company out of it. It was the only way that we could scale, and he knew that Wayne Yan and I would ultimately see that and be excited about doing it ourselves, which ended up being the case.

Speaker 1

Was there a moment for you where you decided? I'm sure you were debating it yourself, and you had other opportunities. What was the deciding factor for you?

Anastasios Angelopoulos

It became clear that the only way to scale what we were building was to build a company out of it. The world really needed something like Arena: a place to measure, understand, and advance frontier AI capabilities through real-world usage with real-world users, based on organic feedback.

In order to achieve the scale and distribution necessary—and, of course, the quality of the platform necessary—to do this effectively, we would need to start a company out of it. We considered other options: Were we going to keep doing this as an academic project? Were we going to do it as a nonprofit? But ultimately, under those constructs, we didn't feel like we'd have the resources necessary to accomplish our mission.

Speaker 1

So you raised $100 million, or $80 million?

Anastasios Angelopoulos

$100 million. That's a lot of resources. It's great.

Speaker 1

What's it for?

Anastasios Angelopoulos

Well, obviously—

Speaker 1

On behalf of everyone—

Anastasios Angelopoulos

Yeah, everybody—

Speaker 1

Dude, it's an arena. How are you going to spend your money? [laughter]

Anastasios Angelopoulos

First of all, we don't necessarily need to spend all that money, right? The purpose of money at a company is to give you cards to flip. It's to say, “Hey, you have enough resources so that if your first bet fails, you can make another bet and another bet.”

Of course, that's not to say that we're going to spend all of it. You want to spend things responsibly. Having said that, the platform is actually quite expensive to run. We fund all of the inference on the platform. The way it works is the platform—

Speaker 1

You pay market rates. They don't give you discounts.

Anastasios Angelopoulos

No, no, we get discounts, but they are standard enterprise discounts, the same that would be given to any other customer.

Speaker 1

Have you disclosed any numbers? I see numbers of votes, but what's that in monthly tokens?

Anastasios Angelopoulos

I don't know about tokens. I'd have to back that out. We have, let's say—again, this is off the cuff, but I can safely say—we have more than 5 million now, 5 or 6 million. We have probably 250 million conversations that happen on the platform. We're on the order of mid-tens of millions of conversations every month happening on the platform.

It's actually quite a large consumer platform for LLMs. Of course, nothing really compares to ChatGPT, but—

Speaker 1

But no, still, I think the largest-scale ones—

Anastasios Angelopoulos

The benefit of this is that it's actually quite a diverse population. For example, 25% of the people on our platform do software for a living.

Speaker 1

Still, at this scale, how do you know?

Anastasios Angelopoulos

We do all sorts of things. We either survey them or analyze the prompt distribution that's coming into the platform. I'm happy to share more. We've done something called Expert Arena, which is trying to understand the distribution of experts that are coming to the platform.

A lot of that can be unauthenticated, or whatever the usage is. But about half of our users now are logged in.

Speaker 1

Yeah.

Anastasios Angelopoulos

So we have some ability to understand them. We also have surveys that we run on the platform that tell us a bit more about who the actual users are. Of course, there's always response bias in surveys, so you have to take it with a grain of salt. Nonetheless—

Speaker 1

If you don't know Anastasios's background, he's the guy to correct for response bias.

Anastasios Angelopoulos

Okay. Yeah, there are a lot of guys like that—and girls.

Speaker 1

Guys and girls. You're not the only player. There's Artificial Analysis. They started an arena. It's like some crypto people who started this. I don't know if you've had a conversation with them, like, “Hey, this is our thing,” or, “Let's work together on something.” I don't know if you've—

Anastasios Angelopoulos

No, I've talked to Artificial Analysis. I've actually talked to both groups. Both seem—

Speaker 1

Am I missing any major players? Is it just those two?

Anastasios Angelopoulos

No, I think those are some of the larger ones. Depending on how you define the term, Artificial Analysis obviously has huge market mindshare around the analysis—

Anastasios Angelopoulos

Of different AI systems.

Speaker 1

Yeah, they told me they were going to be the Gartner of AI.

Anastasios Angelopoulos

Yeah. That's kind of their goal, and I think they're going after that consulting market and so on. Artificial Analysis, from what I understand, is a team of consultants doing this.

Speaker 1

Yeah.

Anastasios Angelopoulos

They seem like really, really nice guys going after that particular market.

Speaker 1

Their analysis is based on aggregating public benchmarks and turning those into analytics—

Anastasios Angelopoulos

Independently rerunning.

Speaker 1

And independently rerunning, which matters.

Anastasios Angelopoulos

Yes. Using those to compile reports and so on that educate the field on the performance of all these different models.

Speaker 1

But they also have arenas.

Anastasios Angelopoulos

They have arenas, but the arenas are not based on organic usage. The thing that distinguishes our platform from theirs is that users are actually inputting their own use case. They're asking their own question. That gives a level of realism that their platform doesn't have.

Of course, they specialize in a slightly different thing, but I see those platforms diverging in that sense.

Speaker 1

Yeah, and sometimes it's the only way to do this. For example, for Artificial Analysis, their video arena is pre-generated videos. You can't enter your own video.

Anastasios Angelopoulos

That's correct. But we are doing it organically.

Speaker 1

Yeah. Exactly.

Anastasios Angelopoulos

As a voter, it does help in terms of not having to wait.

Speaker 1

It does, but also, why would you go?

Anastasios Angelopoulos

Do you actually care about other people's videos? Form your own intuition.

Speaker 1

Maybe. Yeah. Maybe you're interested in comparing.

Anastasios Angelopoulos

I'm a shitty prompter, right?

Speaker 1

Are you? I don't believe that.

Anastasios Angelopoulos

I'm a terrible prompter. I learn by example.

Speaker 1

Don't denigrate yourself. Don't denigrate yourself.

Anastasios Angelopoulos

There are many people whose prompts are much better than mine. Let's say that. That's a fact.

Speaker 1

People have all sorts of cool ways of prompting. Oh, yeah. It's educational to see.

Anastasios Angelopoulos

The only way to learn is by looking at other people’s prompts, seeing the results, and thinking, “Oh, I didn’t know you could do that.”

Speaker 1

Totally. Yeah. Okay, so let’s go back to Arena. One thing I do want to say is that the number-one use of funds is getting off Gradio.

Anastasios Angelopoulos

Oh, yeah. Yes, we—well, listen: Gradio is an incredible platform. Gradio scaled us to a million MAU.

Speaker 1

Yeah, that’s incredible. And, of course, you tell the Hugging Face fans that—

Anastasios Angelopoulos

Of course. Yeah, we’re really grateful to Gradio for taking us this far. Eventually, it became time for us to move off of it and go to React.

Speaker 1

I’m sure Hugging Face would have loved you to stay on.

Anastasios Angelopoulos

They would have. I’m sure.

Speaker 1

Was there a technical reason? You just couldn’t get the performance?

Anastasios Angelopoulos

Yeah, it just became hard to develop, and there were all these tools that we wanted in React. To do all the fancy things that you can do in React became kind of difficult. One example—I don’t know, what’s a feature that you really wanted?

Speaker 1

Let’s say we wanted to create our own custom loading icons for video with notifications.

Anastasios Angelopoulos

Okay.

Speaker 1

How are we going to do that in React? It’s hard.

Anastasios Angelopoulos

Yeah, I mean, you can make a custom component. I’m sure the Hugging Face guys are going to come in and say, “You can do that in Gradio,” which maybe you can. But fewer developers know how to do that. How are we going to hire for that? We’d have to reskill them, and people are less familiar with that stack.

Speaker 1

So it’s all React, all that?

Anastasios Angelopoulos

Yeah, all that.

Speaker 1

Okay, cool. Other uses of funds that might be interesting?

Anastasios Angelopoulos

Basically, that’s compute resources. It’s primarily inference; that funds the free usage of the platform, and then also hiring, of course—headcount. Yeah, we have an office. You know, it’s in San Francisco.

Speaker 1

I’ll tackle one of the major things this year, which I’m sure you’re tired of thinking about, but for people who are not in the loop, this is going to be news to them: the Leaderboard Illusion, the whole thing with Cohere. Let’s summarize what they said and then your response.

Anastasios Angelopoulos

Mhm. The Leaderboard Illusion is a paper that critiques Chatbot Arena, and the main—

Speaker 1

Pretty brutally.

Anastasios Angelopoulos

Well, I would say unscientifically. [Laughter] Let’s be clear: Cohere wasn’t doing that well on—

Speaker 1

Cohere was around 74.

Anastasios Angelopoulos

It’s all good. It’s actually a respectable place to have on the leaderboard. I don’t even think it was really Cohere—the model developers—doing this. It was more their research side. But in any case, what does the Leaderboard Illusion say? It says that Chatbot Arena was doing undisclosed, quote-unquote, private testing on our platform: model providers would send us prerelease models, and we would expose them, and so on and so forth. The claim is that this creates inequities in the leaderboard due to that prerelease testing. For example, they cited that Meta had, at some point, tested a number of models with us.

Speaker 1

Of course, we can’t disclose all the details of how all that was done.

Anastasios Angelopoulos

But that is the main claim of the paper. Our response to that paper is online; you can find it by searching for “Response to Leaderboard Illusion.” Our response is essentially pointing out a series of factual mistakes in the paper that question the validity of its claims. You can look at the first version of the paper on arXiv yourself and see the claims. I think most scientists would view that as—

Speaker 1

Oh, they’ve corrected it.

Anastasios Angelopoulos

They corrected it, of course, but they didn’t correct everything. They just corrected some aspects that were blatantly unscientific and false. For example, they said that we only sampled around 9% open-source models and around 60% closed-source models, and that this created a gap between open and closed source. In reality, we’re very supportive of open-source models, and it was more like 60/40. That was one example of an error in the claims.

Another example is that they claimed there was some sort of bias introduced by this prerelease testing and that it was undisclosed. In reality, as you probably know, we’ve been doing this prerelease testing for a long time. Our community loves it. They love getting the secret code names.

Speaker 1

Yeah, the secret code names—Nano Banana and all that. So Nano Banana, by the way, started on you.

Anastasios Angelopoulos

It started on us, right? And people loved it. It became a global sensation. A nonzero fraction of the global population was using it. About naming it Banana, was that their decision?

Speaker 1

It was their decision, I believe. It was sort of randomly generated, and it just went—

Anastasios Angelopoulos

No, no. So apparently, Nana, who’s a PM—

Speaker 1

Yeah, yeah.

Anastasios Angelopoulos

—is named after her because her nickname is Nana.

Speaker 1

Oh, that’s sweet. I didn’t know that.

Anastasios Angelopoulos

Nana put Banana on it.

Speaker 1

Oh, that’s sweet. Yeah, I didn’t know that. I didn’t know the origin story.

Anastasios Angelopoulos

To us, it just looked like a random thing.

Speaker 1

But it was clearly head and shoulders above everything. Before that, there was Recraft Image, remember?

Anastasios Angelopoulos

Yeah, I do. Of course. And also BFL and all those. All those models are great, and I think those teams are also improving quite quickly. But Nano Banana was a sensation. That moment alone changed Google’s market share.

Speaker 1

Yeah, market share.

Anastasios Angelopoulos

Yeah.

Speaker 1

Seriously. Google’s stock—billions of dollars are moving because of Nano Banana.

Anastasios Angelopoulos

And now there’s an OpenAI code red and everything.

Speaker 1

I don’t know about that, but The Information reported this.

I would say image generation has been this weird part of AI overall because it’s not strictly AGI-critical. It’s not reasoning, and it’s not feeding more context into the model. It’s the model generating a visual representation. I always think, well, Gemini used to get a lot of complaints for generating racist images or whatever.

That was a hilarious moment, and ChatGPT also had it in the past. I’m like, can we just get rid of this? Do we have to do image generation? Let’s focus the positive reputation of AI in general on language models and coding and the other stuff.

Anastasios Angelopoulos

Yeah.

Speaker 1

But I’m wrong. [Laughter] I’m such a huge Nano Banana Pro shill.

Anastasios Angelopoulos

Yeah, I totally agree. I was also kind of wrong about this. I didn’t see the positive benefits, but actually, I think these multimodal models are going to become some of the most economically valuable aspects of AI, both in the consumer space and in enterprise. One of the fastest-growing market segments in AI adoption is marketing and design.

Speaker 1

Yeah, ads. I’m a content creator, right?

Anastasios Angelopoulos

Yeah, of course. I’m sure you’re using it all the time.

Speaker 1

Infinite supplies of diagrams and explainers and infographics.

Anastasios Angelopoulos

Yeah. Soon, we’re not even going to be making them for our papers. Our paper figures are just going to be made by—

Speaker 1

Yes. I do think they actually one-shot. DeepSeek came out with V3.2 recently. I took their explanations of the RL environment stuff and fed them into Nano Banana Pro, and it generated an image that helped me understand the paper better. The fact that I can casually generate a paper-quality diagram that would usually take a PhD student a month in Photoshop or something is incredible.

Anastasios Angelopoulos

Yeah, it’s incredible. It is amazing.

Speaker 1

I want to ask about your principles for running Arena. I think you manage a giant community—5 million MAU. What have you decided are the core principles, both before becoming a company and now that you’re a company? I don’t know if anything has changed for you.

Anastasios Angelopoulos

I don’t think anything has really changed. We want to provide the north star of the industry and center the use cases of real users, foregrounding those so that people know what to target. The goal is to create a benchmark that is constantly fresh and does not suffer from overfitting because we constantly have new data points coming in. It tracks all the different new models and all the different new use cases of AI, and gives the whole world a sort of ground truth for how real users are using these models and how good they are on those use cases.

We continue to do quite a few open-source data releases. We’ve probably released more data than basically anybody on the real-world use cases of AI: millions and millions of conversations from real users that the community is using to study and improve on.

Speaker 1

In terms of what you will build versus what you will not build, I’m not necessarily caught up on everything that you’ve launched. I know recently you’ve done the Dev or Code Arena.

Anastasios Angelopoulos

Yeah, Code Arena.

Speaker 1

Expert Arena.

Anastasios Angelopoulos

Yeah, Expert Arena.

Speaker 1

What’s in the critical path for you, let’s say, for next year? What have you decided you’ll never do?

Anastasios Angelopoulos

Let me first talk about things that I’ll never do on the platform.

Integrity comes first to the platform. Basically, the public leaderboard that we show on LMArena, I think of as a charity. It’s a loss leader for us. We don’t really make money on the public leaderboard. You can’t pay to get on the public leaderboard.

It’s not like a Gartner in that sense. It’s not like any of these pay systems. It’s never going to be like that.

Speaker 1

Models are going to be listed on the leaderboard—

Anastasios Angelopoulos

Whether or not the providers pay and whether or not they’re getting a good score. They can’t pay to take it off either. And so, what that means is very important.

Speaker 1

And so, what that means is that the leaderboard has a certain integrity that will never be compromised, of course.

Speaker 1

Yeah. But not all preview models will make it onto the leaderboard.

Anastasios Angelopoulos

No, but that’s okay. Those preview models have never been released.

Speaker 1

Yeah, right. Who cares about putting unreleased models on a leaderboard? The point is that for every released model, the score that you see on the leaderboard is statistically sound. It reflects the real-world capabilities of the model.

Anastasios Angelopoulos

Yeah. Why? Because millions of people from around the world have voted for it, and that’s where that number comes from. All we do to compute that number is let millions of people vote. We take those votes and turn them into a number that’s always going to remain a transparent and fair reflection of model performance.

Where are we going? Lots of different new categories. I don’t know if you recently saw that we exposed occupational and expert categories. Now, a single-digit percentage of our user base—we’re at millions to tens of millions of users, right? So a single-digit percentage means a lot.

A single-digit percentage of our user base is in medicine, legal, business, finance, accounting, creative, marketing, and things like this. We’re able to show the performance of these models in all these different verticals because we have all these users in our user base. We’re working more toward multimodal. Video is soon to launch on the site, at some point later this year or early next year. So, there are lots of things in the pipeline.

Speaker 1

Amazing. Will you expose an API?

Anastasios Angelopoulos

We’ve thought about it. Yeah, I think it’s a possibility.

Speaker 1

Yeah. What are the counterarguments? Why not?

Anastasios Angelopoulos

Well, there’s obviously a need for an API. The question is more one of focus for our company, just because we’re a startup, and so we really should be doing one thing well—

Speaker 1

Arenas.

Anastasios Angelopoulos

Yeah, arenas. So, I’m not sure how far we want to splay out, or on what timeline we’d want to do that.

Speaker 1

Yeah. Any other community-management tips, more broadly? Every AI company really wants to grow its community. You’re obviously one of the strongest in the world. What’s really worked?

Anastasios Angelopoulos

Well, first of all, I want to give a shout-out to our community manager, Greg, who is doing an awesome job managing our community, whether that’s on Discord or on LMArena. He’s really incredible. So, I would say, hire Greg—but don’t hire Greg.

Speaker 1

Don’t hire Greg.

Anastasios Angelopoulos

Don’t hire Greg. He’s ours. Find a Greg. I agree.

Speaker 1

But in general, the question is: How do you get to so many users?

Anastasios Angelopoulos

That is a tough question.

Speaker 1

And keep and retain them.

Anastasios Angelopoulos

That is a tough question because consumer is one of the hardest markets in the world. There are a lot of websites in the world that people can go to. Why should they go to yours? The reality is, if you want to create a really dominant product, you have to provide people value.

To be frank, I don’t think we’re all the way there yet. It’s not like I have the solution and answer for how to build a great consumer product. If I did, we wouldn’t be at tens of millions of users. We’d be at hundreds of millions, or we’d be at a billion users. We’d be like—

Speaker 1

Is there a world where you’re bigger than ChatGPT?

Anastasios Angelopoulos

I don’t know. I don’t know that we need to be.

Speaker 1

Yeah.

Anastasios Angelopoulos

And I don’t know that we ever will be, because that’s an extraordinary generational product that they built, right? It took a lot of time, and to some extent it also involved luck. There are a lot of lightning-in-a-bottle moments, like Nano Banana was for us, where our user base just goes up by a lot.

But when those users come, they can just as easily leave. The way I think about it is that every user is earned. You have to earn them every single day. They can leave at any moment. They’re fickle.

All the time, you have to be thinking about: How do I provide this person value? Learning how they’re using my website, what more could I give them, and how do I build in all the retention mechanisms so that they stay and then also bring their friends?

Speaker 1

Is there one thing that’s working in terms of retention? You said a lot of people are signing in now.

Anastasios Angelopoulos

Yeah, sign-in was a big driver of retention.

Speaker 1

No, but what did you give them in order to encourage that?

Anastasios Angelopoulos

Persistent history.

Speaker 1

That’s it? That’s enough?

Anastasios Angelopoulos

Yeah, that’s one thing that has had a big impact.

Speaker 1

Okay. Yeah. What do you want from people? What are you looking for help on? Any call to action?

Anastasios Angelopoulos

Yeah, we are always looking for people to come and join us. If you are one of the best people in the world in your area, whether that’s consumer product, machine learning, B2B go-to-market, marketing, or any of these things, we need you at Arena.

We’re building a high-performance team of real experts in everything that they do, and I’m always looking for excellent people to work with.

Speaker 1

What about partnerships? Let’s say I’m at Cognition. I want to partner with LMArena, or just Arena. What works for you? What existing partnerships do you already have that are really fruitful?

Anastasios Angelopoulos

Yeah. So, we of course partner with all the major model labs.

Speaker 1

Yeah.

Anastasios Angelopoulos

And that’s straightforward—

Speaker 1

And that’s straightforward: “Hey, we have a new model here. Here you go.”

Anastasios Angelopoulos

Exactly. So, I think the most straightforward thing for someone like Cognition would be: Let’s evaluate that—

Speaker 1

Agent.

Anastasios Angelopoulos

Yeah. But we should be continuing to shape our—well, Code Arena is an agent evaluation.

Speaker 1

That’s true. That’s true.

Anastasios Angelopoulos

And it’s more focused on— all these arenas tend to focus on the model rather than the harness. But maybe that should change.

Speaker 1

Maybe we should be evolving in that direction.

Anastasios Angelopoulos

And I think Code Arena is a good example of an arena that would support—

Speaker 1

A full-featured harness like Devin.

Anastasios Angelopoulos

Yeah. And so, in my view, if I’m talking to Cognition, I’m saying, “Hey, let’s get Devin on the arena and figure out how to loop the harness together so that we can—” I’m sure there’s something that could be really valuable there, especially given Devin.

Last week, people were talking about Devin being dead.

Speaker 1

Did you see that?

Anastasios Angelopoulos

Yeah. People were saying Devin’s gone. Devin’s not gone.

Speaker 1

Devin’s everywhere.

Anastasios Angelopoulos

Well, people were saying Devin was dead. So, can we highlight that for people and show them, “Hey, Devin is actually the best, or one of the best in the world, at doing what it does”?

LMArena can actually do that. Our place as a central evaluation platform allows that to happen.

Speaker 1

Yeah, love it. All right. Thank you for owning the state of the art.

Anastasios Angelopoulos

Thanks so much, and congrats on a wonderful year.

Speaker 1

Appreciate it. Congrats to you, too. Congrats on all the growing momentum in your podcast and in your career. Thank you.

Anastasios Angelopoulos

It’s really impressive to see.

[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena | BidClub