[BidClub_]
Machine Learning Street Talk · · 16 分钟

AI基准测试是否讲全了故事?[赞助内容](Andrew Gordon 和 Nora Petrova - Prolific)

Tim ScarfeAndrew GordonNora Petrova

播客
TL;DR
  • 技术排行榜上的领先地位,是产品质量的弱代理指标。 Andrew Gordon 将高分模型比作 Formula One 赛车:它们是工程学的巅峰,却可能成为日常使用的“绝对噩梦”。在 Humanity’s Last Exam 或 MMLU 上夺冠,可能同时伴随较差的信任、沟通、适应性或个性表现。技术基准通常评估固定测试,不把人类置于回路中。
  • 基准测试披露过于零散,难以进行清晰的竞争比较。 各家实验室强调不同测试,甚至完全不发布基准数据;Grok 4 的宣传重点就是 Humanity’s Last Exam。Andrew 的核心警告是:“如果你只是依赖这些技术指标,就会错过一半重点”("If you just rely on those technical metrics, you miss half the point.")。
  • 随着人们把敏感的个人问题交给模型,安全正成为一项尚未被充分衡量的重大产品风险。 Nora Petrova 称心理健康和人生建议场景“目前仍是蛮荒西部”(“Wild West at the moment”),并以 Grok 3 和 Mecha Hitler 为例,质疑“这些模型之上的安全训练究竟只是多薄的一层表面保护”。
  • Chatbot Arena 可能奖励的是测试资源优势,而不只是用户偏好。 在 Llama 4 发布前,Andrew 称 Meta 曾向 Arena 发布27个模型,但最终只有1个被报告;更高的曝光量会带来更多提示词和优化数据。他还指出,战斗次数与排行榜位置之间存在“相当强的关系”。
  • 匿名二元投票带来了规模,却几乎没有可执行的产品洞察。 只知道用户偏好哪个答案,并不能知道用户是谁、为何投票——因此,除了排名之外,这些数据“在某种意义上毫无用处”,也无法诊断信任、帮助性、沟通、适应性或个性表现。
  • Prolific 的 HUMANE 排行榜,以受控且具代表性的评估取代不受限制的流量。 它从美国500人的概念验证发展为模型对战、结构化多步对话和质量控制——“三次这样,你就出局”(“three of those and you’re out”)——参与者抽样则按美国和英国人口普查人口结构分层。
  • 早期产品差距在于个性,而非基本效用。 在6个头部模型中,用户对个性以及对背景和文化的理解评分低于帮助性、沟通和适应性。原因仍不确定,但 Nora 指出模型的谄媚倾向正在上升:“人们总体上似乎并不喜欢这一点”(“People generally don’t seem to like it.”)。
摘要 · 为研究而整理的核心内容

1. 技术卓越,仍可能造出糟糕的日常主力产品

  • Andrew 的框架是:Formula One 赛车是“工程学的绝对巅峰”,却并不适合日常通勤;同样,在 Humanity’s Last Exam 或 MMLU 上表现出色的模型,可能“日常使用起来简直是噩梦”。
  • Andrew 指出,大多数技术报告只是让模型完成固定评估,给出一个分数,过程中没有人类参与。各家实验室披露的证据也不统一——Grok 4 强调 Humanity’s Last Exam,而有些发布则完全没有基准数据——因此不存在标准化的竞争场。结论是:技术分数只反映了面向人类的产品一半表现。

2. 敏感场景的使用已经跑在安全衡量之前

  • Nora 认为,人们越来越多地向模型寻求心理健康和人生建议,却没有获得其他场景通常要求的监督与伦理操守。Grok 3 和 Mecha Hitler 让人不得不追问:“这些模型之上的安全训练究竟只是多薄的一层表面保护”(“how thin a veneer the safety training is on top of some of these models”)。
  • Andrew 直指缺口:“安全没有排行榜”(“There is no leaderboard for safety.”)。他认为,安全的重要性应当不亚于模型的速度或智力。
  • Nora 提到 Anthropic 的 Constitutional AI 工作和机械可解释性——追踪被激活的特征、概念与回路——是针对新型场景建立信心的重要路径。

3. Chatbot Arena 的规模掩盖了结构性偏差

  • Andrew 最尖锐的诚信担忧是:据称在 Llama 4 发布前,Meta 曾向 Arena 发布27个模型,但最终只有1个被报告。更多比较意味着更多提示词访问和优化数据,可能让模型专门变得“更擅长 Arena”。
  • Arena 团队曾将抽样不均归因于用户追逐最新、最先进的模型。Andrew 认为,这会导致采样低效:部分模型获得的对战次数远多于其他模型,而他观察到战斗次数与排行榜位置之间存在“相当强的关系”。
  • Arena 参与者是匿名的,没有人口统计数据,二元偏好也无法揭示投票原因。这会产生一个看起来很有吸引力的排名,却几乎无法说明模型究竟在哪些方面失分——是信任、个性,还是沟通。
  • 提示词质量同样不受控制:用户可以什么都不提交,可以只说你好,也可以从“太阳有多大”跳到“蛇有多长”,削弱了细致比较的价值。

4. HUMANE 让每场模型对战都必须产生信息

  • Prolific 最初的 Prolific User Experience Leaderboard 使用500名具代表性的美国参与者,并采用1至7分的 Likert 量表。如今的 HUMANE 使用比较式对战、多步对话、分项偏好指标,并对低投入或跑题提示词进行惩罚:“三次这样,你就出局”(“Three of those and you’re out.”)。
  • Nora 解释称,HUMANE 采用 Microsoft Xbox Live 的 TrueSkill 框架。随着证据累积,贝叶斯估计会逐步收窄;系统则按照信息增益选择配对,只在最能快速降低不确定性的地方安排对战。

5. 具代表性的偏好暴露了模型更柔性的短板

  • Prolific 根据美国和英国人口普查数据对参与者分层,变量包括年龄、族裔和政治立场。围绕约20个不同人口群体分别进行的竞赛,可以合并结果,也可以加入更多群体和模型,或延长运行时间,直到达到目标置信区间。
  • 在首项6模型研究中,个性以及背景与文化理解的评分落后于帮助性、沟通和适应性。Andrew 没有预设解释:任务可能没有 eliciting 出个性,模型可能缺乏用户上下文,也可能训练数据根本没有生成符合用户期待的个性。
  • Nora 表示,数据集将检验的一个问题是:可检测的谄媚倾向,是否与个性维度上的负面投票相关。分析可以把人类反馈与面向 LLM-as-a-judge 的对话模式分类结合起来。HUMANE 竞赛当时仍在累积对战数据,因此这仍是开放问题,而非已经确定的结论。
Andrew Gordon

Formula One cars are the absolute pinnacle of engineering, right? Everything is perfect. They have a huge top speed. If you used one as your daily commuting car, you’d have an absolute nightmare, right? I think the same can be said for these models, right? A model that is incredibly good on Humanity’s Last Exam or MMLU might be an absolute nightmare to use day to day.

Most reporting on benchmarks is done on technical benchmarks these days, right? That is where you get the model, give it a set of evaluations—maybe on one theme or maybe an exam—and then get a score. Humans aren’t really involved in that loop.

My name is Andrew Gordon. I’m a staff researcher in behavioral science at Prolific, so I work on the sciences team, tackling questions related to humans in research, specifically online research.

Nora Petrova

My name is Nora Petrova. I’m an AI researcher at Prolific. I’m tackling questions around how we include humans in the development and evaluation of AI models, how we align them with human values, and how we fully understand what they’re capable of.

Andrew Gordon

How helpful do people find models? What’s the communication like? How adaptive do they find them? What do they think of the model’s personality? When you get ratings on those kinds of factors, what you actually get is an actionable set of results that say, “Okay, your model is struggling with trust,” or, “Your model is struggling with personality.”

We had a first go at this with what’s called the Prolific User Experience Leaderboard. That was a proof of concept with 500 participants from the US, a representative set of participants who evaluated a single model at a time and gave their feedback in a Likert scale format. How helpful did you find the model, from 1 to 7, and so on?

We’ve now taken the learnings from that initial leaderboard and built them into what we’re calling HUMANE, which is our main leaderboard. It uses a similar approach to Chatbot Arena, in the sense that we have comparative battles between models, which allows us to much more clearly differentiate which model is performing better.

Tim Scarfe

What if we could have a fairer approach where we actually diversely sampled and stratified folks based on how old they are, where they live, and what their values are? What would be a fairer approach to understanding the behavior of models and conducting evaluation?

This is what Andrew and Nora have been doing at Prolific. How do we know whether these models are actually good for the humans who are using them? How can we make evaluation metrics fairer? This is Andrew and Nora.

1. Benchmarking Misses Human Experience

Andrew Gordon

The problem is that, at the moment, the field of evaluating and benchmarking these models is incredibly nascent, right? It’s only been around as long as LLMs have been around, in the last couple of years. Because of that, it’s a fractured field. There’s no standard playing field for how these labs report data on benchmarking.

Some may emphasize, like with Grok 4 recently, a huge amount of emphasis on Humanity’s Last Exam and less so on other benchmarks. Some models will come out without any benchmarking data at all. Essentially, there’s a lot of heterogeneity in how these labs report results, and it leads to a situation where I think we’re at risk of struggling to compare the models on an even playing field.

There are, of course, bigger questions as well. Models are often lauded for getting the highest score on Humanity’s Last Exam, so we know from a technical perspective that the model has advanced above its rivals. But my core argument in this space is that if you just rely on those technical metrics, you miss half the point. These models are designed for humans to use. At the end of the day, most of the users are humans, and simple performance on these exams doesn’t necessarily correlate to a good user experience.

All the frontier labs really need to start having human-preference leaderboards more front of mind, alongside all the technical metrics.

2. Safety Needs Its Own Metric

Nora Petrova

People are increasingly using these models for very sensitive topics and questions—for mental health and for how they should navigate problems in their lives—and there is no oversight on that. In any other area where these topics are discussed, there is a lot of regulation and ethical conduct built into it.

Here, it’s the Wild West at the moment. Some companies are taking it more seriously than others and trying to study the ways in which humans are using the models for more personal topics and problems. We’ve seen some pretty startling examples recently with Grok 3 and Mecha Hitler, and it does raise questions about how thin a veneer the safety training is on top of some of these models.

Andrew Gordon

There is no leaderboard for safety, right? There’s no metric. We don’t grade LLMs by how safe they are. In fact, safety isn’t really even part of the question for some researchers. I would argue that it should be just as important as how fast or smart the model is. How safe is it for people to use?

Nora Petrova

There’s been a lot of interesting research coming from Anthropic in that direction, with regards to safety and alignment of the models, using Constitutional AI and various approaches that they’ve explored. There’s also been research around mechanistic interpretability, peering behind the curtains of the models and understanding how an input produces a certain output: which features and concepts, and which circuits, get activated along the way.

It’s about tracing the thoughts, essentially, of these models and trying to isolate where potential problems may emerge. Work of this kind is very important in raising confidence that these models will be able to handle novel situations in safe ways.

3. The Leaderboard Illusion

Andrew Gordon

The one other thing I’d say is that, given Chatbot Arena is the No. 1—and frankly, pretty much the only—human-preference leaderboard out there for LLMs, it’s really important that we understand what’s going on behind the scenes.

Obviously, Chatbot Arena is entirely open source. People go in, put in a prompt, get a response from 2 different models, and then say which one is better. The Leaderboard Illusion paper found that some companies are getting access to a lot more private testing in the background than others.

For instance, before Llama 4 launched, Meta released 27 models on the Arena, but of course only 1 was actually reported in the end. That undermines the integrity of the Arena because the more comparisons you have for your model, the more access to prompts you have, and the more data you have to refine a better model that’s better at the Arena. It adds an element of bias into the data that’s very, very hard to get around.

There are other issues that were called out, and issues that we’ve seen ourselves, which we think dictate the need for a more rigorous and methodologically sound approach to creating these kinds of human-preference datasets. Beyond the criticisms in The Leaderboard Illusion paper, we think that paper didn’t really touch on some of the other things we think should be considered when you’re doing human-preference evaluation.

4. HUMANE Makes Feedback Actionable

For me, there are 3 big areas where we’ve sought to improve. First of all, as you mentioned, the sample. The sample for Chatbot Arena is anybody, right? We don’t know anything about them, and we don’t collect any demographic data. They’re just people going there anonymously, prompting the models, and giving their preference data.

Obviously, that’s great. You get a huge amount of data, which is fantastic, but you know nothing about the people giving the data, which is fairly suboptimal. In terms of specificity, for anybody who’s used Chatbot Arena, all you’re doing is saying, “I like this response more,” or, “I like this response more.”

In the real world, that kind of data is useless, in a sense. It gives you a really nice way to make a nice leaderboard of AI models, but it tells the companies nothing about why that preference has been given. In our approach, we sought to mitigate that by splitting preference down into its constituent parts.

Things like: How helpful do people find models? What’s the communication like? How adaptive do they find them? What do they think of the model’s personality? When you get ratings on those kinds of factors, what you actually get is an actionable set of results that say, “Okay, your model is struggling with trust,” or, “Your model is struggling with personality.” That’s where you need to be focusing to actually build a model that is good for real users in the real world.

But there’s no QA in the sense that I could go in and just say hello, or say absolutely nothing, or have a multi-turn conversation and completely wander from “How big is the sun?” to “How long is a snake?”—just topic wandering. I don’t think that gives a really good, nuanced view of models. So we’ve built into our structure that participants come in and have multi-step conversations with models.

We built in QA that actually says, if you put low effort into your question or you start wandering, we're going to penalize you. Three of those and you're out. Those are the kinds of principles we built the leaderboard around.

5. TrueSkill Guides Model Battles

Nora Petrova

I would touch upon the methodology that we've used, which is TrueSkill. It's a framework developed by Microsoft for estimating the skill levels of players on Xbox Live. They take into account things like randomness in games, changing skill levels across time, and whether someone is having a fluky win streak versus a seasoned player who consistently performs well. All of these things were factors that we thought would be good to take into account.

It's a very flexible system that estimates probabilities with Bayesian distributions, with a mean and a variance that get narrower and narrower over time as the system learns about the outcomes of these battles or comparisons. Most importantly, it's based on information gain. The way we pick the next pair that should occur in the tournament is based on how much we will learn from these models going head-to-head. How much information are they giving us? How much are they reducing the uncertainty?

We order the queue of pairs according to that, and that gets us to a place of minimized uncertainty as fast as possible. It's a really flexible approach. We can run separate tournaments, as we've done with our demographic groups. We have around 20 demographic groups, and we've run separate tournaments for them.

We can consolidate the findings for each tournament to obtain an overall leaderboard that is much less uncertain than any of the individual tournaments or leaderboards that we can produce from any of the demographic groups. It really allows us to slice and dice the data in any way we want, and we can easily add more demographic groups and more models over time. We're developing in the open and welcoming feedback.

Andrew Gordon

I think one of the things they pointed out in “The Leaderboard Illusion” paper was that some models are sampled considerably more than other models. I believe the folks behind Chatbot Arena said that's because people come to the Arena to play with the latest models. People want to play with the state of the art, which is all well and good, but it doesn't lead to an efficient sampling method.

It essentially means that some models get a lot more battles than others. They therefore get a lot more data and get a lot better in the Arena. There's a pretty strong relationship between the number of battles and the place on the leaderboard. We only ever conduct battles based on the need from the data. If the uncertainty is high for a specific model against another specific model, we conduct a battle to lower that uncertainty.

So it's all driven by the data. It's very computationally sound because we don't make any more comparisons than we need to. It allows us to get to a point where models are strongly differentiated based on uncertainty.

Nora Petrova

If we have a certain goal with regard to uncertainty in order to fully differentiate the models at the confidence interval we're interested in, we can conduct more battles until we get there. The control is in our hands. We just need to recruit more participants in order to get to that level of certainty.

Andrew Gordon

The way we've sampled for this study, we're obviously using our own participants from the Prolific platform, but we've sampled effectively based on the census data that we have for both the US and the UK. The long-term vision would obviously be a more global product, but at the moment, we stratify our sample—that is, our participants who are giving us this feedback—by demographics like their age, ethnicity, and political alignment.

We have an awful lot of data from censuses that tells us each country is made up of certain proportions of these demographics. That essentially allows us to say that when we've amalgamated all these findings and found that leading model, we can very confidently say that model is preferred by as representative a set of the general public as we can possibly get.

Hopefully, in that sense, it's more related to the real-world preferences of people in the world rather than a very potentially skewed and biased subset that might be responding to Chatbot Arena. We ran our first one as an MVP, a proof of concept that was more about proving that we could do this in a rigorous and methodologically sound way. When we actually ran it, we only ran it with 500 participants.

It gave us a lot of insights about how we build Humane, which is our leaderboard that we're working on at the moment. That leaderboard is actually running as we speak in the background. We're still having battles, so we expect to have more data from that.

6. Models Struggle With Personality

What I can say about the first round that we did is that, across the board, the 6 models we tested, which were leading models at the time, tended to perform a lot worse on personality metrics and background and culture metrics than on things like helpfulness, communication, and adaptiveness. What that really signals is that there's some more subjective aspect of these models that people are less impressed by.

Potentially, they were doing tasks that don't elicit a personality in the model, or they don't elicit the model talking about background and culture. Also, the model doesn't know their background and culture, so it's very hard to align with them. But the other possibility is that models are just not very good at that. That would potentially be an effect of the data they've been trained on, right?

We know very little. Obviously, models are trained on the entire internet, but when you train a model on the entire internet, do you get a personality that really represents what people want? From this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more objective measures.

Nora Petrova

Obviously, a lot of these models have undergone extensive fine-tuning to tailor their personalities or tailor how they approach answering questions, which are different across the different companies. We've observed recently that there has been an increase in sycophancy, or this people-pleasing behavior of models, and people generally don't seem to like it.

One thing that the results of this experiment and the later leaderboard datasets will allow us to answer is what the correlation is between telltale signs of sycophancy and downvotes in the personality metric, and whether that influences people's decisions about which model they prefer. We can perform various types of post-processing and analysis of the data to identify the levels of sycophancy observed in the datasets and try to identify the interesting relationships between the feedback that people gave, more model-driven, LLM-as-a-judge–oriented analysis of the conversations, and the classification of various patterns and what the models exhibit.

It's quite interesting to see what we'll find.