模拟人类:从生成式智能体到80亿个数字孪生体——Joon Sung Park,Simile AI
Vibhu × swyx × Joon Sung Park
- Joon Sung Park 的核心判断是,模拟也有自己的规模定律:在模拟和预测人类时,摄入的关于人类的数据越多、使用的算力越大,模型表现就会获得可预测的提升。 Simile 正在训练自己的模型,并看到了这条曲线的“早期一瞥”。
- 护城河的核心在于,Park 将生成式模型追求“超级理性、客观的机器”,与 Simile 所需的“和我一样笨”的模型作对比——模型必须犯下人类会犯的同样错误。 通过网络训练的模型,本质上包含的是人类主动暴露出来的态度数据,只掺杂了部分行为数据,而不是“人类的暗知识”(dark knowledge of humanity)——即人们实际上会做什么。在利基人群上,Park 称前沿模型的行为预测准确率会降至20–30%,一般人群也只有50–60%;相比之下,Simile 已验证的基准是以85%的准确率复现人们的态度和行为,“大致与人们复现自身的准确度相当”。
- 验证方面的核心资产是《Generative Agent Simulations of 1,000 People》论文:对具有代表性的1,000名美国人进行抽样,采集2小时数据,2周后用问卷、Big Five、行为经济学博弈、General Social Survey 以及已发表的RCT测试数字孪生。 后续研究显示,在 Open Science Foundation 平台数万项预注册实验上进行后训练,还能带来显著的进一步提升;因果数据和随机对照试验数据最稀缺、也最有价值,因为“世界是我们的 ground truth,但它只发生一次”(the world is our ground truth, but it happens once)。
- 产品定位是用模拟来塑造结果,而不是预测结果。 Park 说:“得知两季度后销售额会崩掉……并不能真正帮到你。”客户真正想知道的是,现在需要做什么,才能避开那个未来。他对 Foundation 及心理史学的讨论——把科学家流放至Terminus这一反直觉的第一步——说明了逐步推进的因果模拟为何胜过点预测;例如,一套电动车营销方案可能提升 EV 销量,却让整体汽车销量下降。
- 在TAM问题上,Park明确拒绝将市场研究定义为1000亿美元级市场,并直言:“模拟不是市场研究工具。” 模拟是人类决策工具。当前落地场景包括概念测试、焦点小组、上市公司的模拟财报电话会,以及与 Gallup 的战略合作;Simile 每周从数万人处采集数据,并拥有覆盖全球数千万人的样本库合作。由于风险偏好等特征“不会随时间真正改变”,采集到的受访者可以跨研究复用。
- 给投资人的成熟度标尺是,Park认为模拟行业所处阶段,类似 AGI 叙事中的 GPT-3.5 和 GPT-4:在当前垂直领域已足以造成真实影响,但激进扩展仍在前方。 他猜测,模拟最终的成本会“与训练一个基础模型相当”;谈到一次社会规模的气候变化模拟时,他说:“我现在就会筹钱,只为跑一遍。”在主持人对谈中,swyx 称一项非洲 UBI 研究得出的结果是“否”,并提出5年约1400万美元的成本;Park 没有确认这一数字,并质疑问题是否出在落地执行上。
- Park 将 Simile 描述为研究实验室与产品公司的结合体:公司约60人,在 Mission Rock 设有旧金山总部,并新开纽约办公室;联合创始人包括 Michael Bernstein、Percy Liang(“foundation model”一词的创造者)和 Lainie Ellen,员工约15–20%来自 Park 的实验室。 他对市场的收束判断是:“你看任何一部科幻作品里的先进文明,都会有两项双支柱技术:一项是某种形式的 AGI,另一项就是模拟。”
1. 从油画家到 Smallville 论文——通过一场“时间机器游戏”
- Park 的人生路径始于韩国,11岁时移居波士顿,之后认真接受写实油画训练——“那不是爱好,而是:我们要靠这个谋生”——直到他意识到,“伟大的艺术家往往会创造自己的媒介,而今天我们能用的最佳媒介,其实是计算。”
- 他2020年开始在 Stanford 攻读博士,加入 Percy Liang 牵头、撰写《Opportunities and Risks of Foundation Models》的团队,并抓住了其中真正新颖的部分:一种“并没有被训练去完成任何特定事情”的模型,但由于吸收了来自互联网的广泛人类数据,“只要从正确的角度戳一下,人类行为就会以相当逼真的方式冒出来”。
- 他与未来的联合创始人 Michael Bernstein、Percy Liang 做过一次创始讨论:把时间快进10年,再回头看,什么才是最重要的单一应用?答案是:“如果我们能直接重建自己生活的世界呢?很难再提出比这更有野心的目标。”这催生了 Social Simulacra,随后是2023年的 Generative Agents(Smallville)论文。Park 认为这篇论文被记住了错误的重点:“记忆组件其实被严重低估了……那是一个非常好的早期记忆系统。”
2. 为什么模拟必须先于个人智能体梦想
- “时间机器游戏”中紧随其后的答案是个性化智能体,但 Park 押注的是推进顺序:如果要做出真正出色的个人助理,“首先需要的其实是一个出色的用户模型”。他的刻意简化版例子是:让智能体替你点晚餐,它选了夏威夷披萨,而你讨厌菠萝——“它彻底失败了”。
- 他的强判断带着明确保留:“我不认为我们已经看到过真正有用、且达到这一工作应有野心的个人助理……我不认为所有正确要素已经就位。”谈到 OpenClaw 一类智能体,他认为用 Markdown 文件存储记忆“相当聪明”,这与 Generative Agents 在2022年的直觉相同——“把所有东西放进 Markdown 文件或文本文件里,就结束了”;但在规模极大的记忆中做检索,“需要大量工作”。
- 他给出了仅靠提示词不够的界线:如果模型必须学习其运行世界的底层物理规律,也就是“新的社会物理学”,就需要触及参数层。公开模型“还没有学会人类社会物理学的完整映射”,这正是 Simile 的核心论点。
3. 行为基础模型:3类数据,只有1类稀缺
- 第一类是丰富的定性访谈,问题可以直截了当地问:“讲讲你的人生故事。”其中包括“童年记忆、创伤或初恋……这些信息以很难预测的方式提供了大量洞见”。第二类是观察性行为数据,包括交易记录和网络抓取活动,用来建立基础统计。
- Park 称第三类数据“可能是最重要的”:来自随机对照试验的因果机制数据,也就是那些“为什么”。这类数据天然稀缺:“世界是我们的 ground truth,但它只发生一次。”因此 Simile 会在虚拟实验室中,邀请经过同意并获得激励的参与者自行开展 RCT,借鉴社会科学方法:一项决定是否真正具有行为意义,取决于“你在这个决定上的利益是否真实”;例如,实验性线上商店中的购买商品必须真的送到参与者手中。
- Park 将生成式 AI 模型追求成为“超级理性、客观的机器”,与 Simile 的目标对照:这里讨论的模型要“和我一样笨”(as dumb as I am),我犯什么错,模型就得犯同类的错。swyx 接话说:“你是在解决 Moravec 悖论。”
4. 模拟是为了塑造未来,而不是预测未来
- Vibhu 以 West Wing 中一场虚构的多发性硬化症披露民调为例:“我们知道这很糟,只是不知道会糟到什么程度。”这引出了主持人的现实质疑:如果我凭直觉就能判断影响方向,为什么还要为模拟付费?Park 的两个回答是,影响幅度和冲击强度确实很难凭直觉判断——“每次有人在网上发言引发巨大反弹,你都会想:‘真是个蠢货。’但这事很难预判”——以及模拟真正交付的是路径,而不是一个点估计。
- 在 Vibhu 提到 Foundation 后,Park 给出了他的标志性类比:心理史学要把3万年的动荡压缩至1000年,反直觉的第一步是把发出警报的科学家流放到Terminus——“太反直觉了……但结果证明,在这次特定模拟里,那确实是正确的一步。”你给系统的是一个目标,而不是一道调查问卷;系统返回的是实现目标所需的步骤。
- 放到市场中,一家只优化 EV 销量的车企可能发现,获胜的电动车营销活动“会改变人们对非 EV 汽车的看法,最终反而让整体销量下降”。swyx 提到 Shopify 在 SimGym 中的类似思路:干预的是多轮购物轨迹,而不是态度本身。
5. 真实世界校准:85%的自我复现准确率,以及前沿模型的失分区
- 主持人的问题是:同一个目标能不能直接交给 Opus 或 GPT-4.5?答案会有多大不同,又该如何确认模拟建立在真实世界之上?Park 指向2024年底发表的《Generative Agent Simulations of 1,000 People》论文:研究对1,000名具有代表性的美国人进行抽样,开展2小时的广泛数据采集,包括由 American Voices Project 编写脚本的访谈和行为数据;随后构建数字孪生,并在2周后把真人请回来,接受问卷、实验、行为经济学博弈、Big Five、General Social Survey 以及发表在 PNAS 上的 RCT测试。
- 结果显示,数字孪生“能够以85%的准确率复现人们的行为和态度,大致与人们复现自身的准确度相当”——这是首次经过验证地证明,个体可以被准确建模。前沿模型提供了“正确的基础”,却缺少态度和行为的细节质感:在客户真正关心的利基人群上,表现会“一路降到20%或30%”,一般人群也只有约50–60%,“你不会想基于这类发现做决策”。
- 后续论文利用了重复危机带来的一线希望:Open Science Foundation 平台上的预注册研究,包含数万项由专业人士设计、假设预先锁定的实验;Simile 用这些数据对模型进行后训练,行为预测能力获得显著提升。Simile 训练了两个不同模型,分别面向人群层面和个体层面;其中个体任务“在很多方面都更难”。
6. 模型遗漏的人性,以及 Park 愿意买下的数据集
- 当被问及人类会做什么而模型做不到时,Park 将差距重新表述为:“人类会犯哪些偏差或错误,而模型遗漏了它们。”他的例子是自己住在 Palo Alto 时,宁愿花40分钟步行回家,也不叫 Uber——“不是为了效率。那段路真的很帮助我思考……这是非常人类的活动。”目标是建模那些“本质上属于人类”的东西:它“可能不是最高效的做法……但正是这些事让我们成为我们自己”。
- 如果必须在 LinkedIn、Twitter 和 Facebook 中选择收购标的,他会选 Facebook:LinkedIn 上的人“戒备心很强”,Twitter 充斥着“疯狂的人设”,而 Facebook “是更私密的空间之一”,更能展现一个人的默认状态。Amazon 交易数据确实是行为数据,但“也是最常见、最容易获得的数据”。
- 谈到 Tencent 那篇将职业与背景交叉组合成提示词的“十亿人格”论文,他认可其规模,但认为如果底层统计已经足够,“那我们其实就已经解决了模拟”——只需调取模型参数中已经嵌入的知识。“但遗憾的是,我们看到的并不是这样”:要获得关于特定人群的细节化利基知识,仍需要定制化采集,并“关注并尊重人们每天真实过的生活”。
7. 规模定律、Schelling 的点,以及诺贝尔级别的野心
- Park 说,Simile 已经看到“模拟规模定律的早期一瞥”:随着人类数据和算力增加,模型表现会获得可预测的提升。swyx 回应:“我们需要一个 scaling walker。”Park 说:“你什么时候找到一个,那会是一件很美的事。”
- 他再次讲起10年愿景:“我们能不能模拟地球上生活的80亿人?”这将打开气候变化、“民主正在崩塌的信号”以及“货币体系的起源”等棘手问题。Park 认为“那里有一座诺贝尔奖等着被拿下”,并确认主持人的限定:是经济学诺贝尔奖。
- 他引用的先例是 Thomas Schelling 在20世纪70—80年代构建的种族隔离模型:红点和蓝点在网格上移动,展示了即便极轻微的同色偏好,也会“随着时间推移导致社会彻底隔离”。这挑战了“只有明确、公开的种族主义才能解释隔离”的看法,并影响了混合收入住房政策。Schelling 后来因奠定早期模拟的基础获得诺贝尔奖。基于智能体的建模后来式微,因为“红点和蓝点并不是对人的丰富描述”;生成式智能体或许能恢复这种逼真度。swyx 补充说,新加坡公共住房实行强制种族配额,正体现了这套逻辑的现实相关性。
8. 成本经济学:今天是数万人,未来是基础模型级别的运行
- 当前的规模已经能从数千到数十万人的建模中提炼丰富洞见;每周从数万人处采集数据,样本库合作覆盖全球数千万人。更关键的是,同一批样本可以在后续研究中重复使用:智能体与具体领域无关,而且部分特征相对稳定。Park 说,风险偏好“不会随时间真正改变”。更大的样本规模,与其说是为了提高统计功效,不如说是为了筛选出“真正关心的目标子人群”。
- swyx 担心,如果数千个智能体与另外数千个智能体彼此交互,成本会出现组合式增长。但他随后指出,现实世界研究往往更昂贵,很多研究根本无法实施,而这些研究所服务的决策可能牵涉数亿美元。Park 将经济账拆成两层:部署时应替代已有预算;但“捕获长期价值的方式……在于上行空间”——更好的决策可能价值数亿美元乃至数十亿美元。
- Park 明确把这说成一个猜测:“未来几年,我们会开始创建成本与训练一个基础模型相当的模拟。”谈到社会规模的气候模拟时,他说:“我现在就会筹钱,只为跑一遍。”在主持人讨论 UBI 时,swyx 称一项非洲研究的结果是“否”,并询问成本是否达到5年约1400万美元;Park 质疑问题是否出在落地执行。Vibhu 的保留意见仍然成立:有时人们付费是为了验证自己已经相信的结论,而不是为了真正模拟它。
9. 超越市场研究的TAM、GPT-3.5时刻,以及背后的公司
- 当前用例包括概念测试、焦点小组、行为实验和 A/B 测试,以及模拟财报电话会。Walmart 是早期客户之一,希望在询问人们想法之外,用模拟测试产品,因此需要处理图片、Figma 原型和网站等多模态输入。Worldfront 是早期客户之一,对让智能体使用网站 URL 这类领域的能力很感兴趣。Gallup 是战略合作伙伴,但 Park 有意暂缓进入政治领域,直到公司拥有足够的“护栏和视角”。
- 令 Park 意外的是部署规模:许多组织声称“我们会倾听人们……但实际上,这种情况极少发生”。模拟可以确保“人们的声音始终出现在那些替他们做决定的会议室里”。
- 在市场空间估算上,市场研究是一个1000亿美元行业,但“模拟不是市场研究工具。模拟是人类决策工具……这种工具的TAM是多少?真的说不清。”Park 坦言,自己从未带着计算TAM的目的入场,只能“假设……它一定很大”。在成熟度上,他认为“这很像 AGI 叙事中的 GPT-3.5 和 GPT-4 阶段”。
- 公司约60人,在 Mission Rock 设有旧金山总部,并新开纽约办公室。联合创始人包括 Park、Michael Bernstein(据 Park 介绍,他是 ImageNet 论文共同作者)、Percy Liang 和负责业务的 Lainie Ellen。员工约15–20%是 Park 过去实验室的同事,其中许多人后来去了 OpenAI、Google Gemini 等机构。公司也在继续招聘产品和基础设施工程师、研究人员,以及“各个部门”的其他人才。
- 最后的落点回到了画家的比喻:“模拟很像绘画……没有一幅画是完美的……但它会试图突出主体最重要的部分”,也就是“本质”。至于我们是否身处某种模拟之中,他说:“无论我们是不是身处模拟,我们的体验都不会因此变得不真实……我会在死后担心这件事。”
Today, we have Joon Sung Park on the podcast. I'm excited to kick this one off. Very exciting company. I want to kick off by asking you: Talk us through the story of your life. How have you gotten here?
Yeah, for sure. Really excited to be here. A story of my life: I was born in Korea, and I lived there for a good 11 years of my life. Then my family moved to Boston. We moved when I was 11, and my parents were doctors, so they were going through their postdoctoral studies. My dad was a surgeon, so he was doing his sabbatical at Boston Children's Hospital.
I grew up there, not too close to tech, actually. I was very much an artsy, painting kind of guy.
Painting.
I actually got into painting a little bit later, in high school, but that's what I used to do. I grew up mostly on the East Coast after Korea, so I lived a good number of years in New Hampshire, and then I went to college in Pennsylvania, where I got into more of this tech scene.
1. From Painting To Computation
I was originally trained to be an artist. I actually thought that would be my professional career, so it wasn't a hobby. I actually thought, "Hey, let's make a living out of this." Then gradually, I got really interested in this idea that the greatest artists often create their own medium, and the best medium that we had available today was actually computation.
I decided to go deeper into that, and one thing led to another. Obviously, we can go deeper into this, but I decided that research was something that I gradually got interested in, and here I am.
So there's obviously a lot that you packed into the research component. You had one of the best papers of 2023, which was Generative Agents paper, commonly known as the Smallville paper.
Yeah.
Feel free to call back to anything else that you mentioned, but most people would have heard of you from this, obviously. Do you have any statistics on how many people have read it? arXiv gives you something, right? Some stats.
Yeah, it's a good question. How many people have read it? I'm actually not sure. I know that we do keep track of the number of citations, which I know is going up quite fast.
Google Scholar. Yeah, we've got Google Scholar—
Google Scholar.
Google Scholar has 7,200—
But I feel like it made a bigger hit, and it was actually a pretty instrumental paper. It was one that got cited so many times—
People frequently ask—
Yes.
“What is the best paper of the year?” Very recently, it's this one.
I thought the memory component was pretty underrated. It was a very good early memory system. But yeah, one of the biggest papers.
Yeah, yeah.
2. The Smallville Origin Story
Yeah. So maybe I can talk a little bit about how this particular paper came together. When I got into research, it was back in 2020, when I started my PhD program at Stanford. That was the year when we were about to get GPT-3.5, GPT-3 to be available.
We already had GPT-2, and you could sense that there was this new class of models that was just becoming available in the market, and the team got very intrigued. The general consensus was, “Is this model actually going to be useful for anything?”
Mm.
It was really strange that these models were not trained to do any particular task, but we decided to take a bet. A large group of scholars at Stanford, led by one of my co-founders, Percy Liang, came together and—
Who coined “foundation models”—
—who coined the term “foundation model.” We wrote this paper, where that term came from, called “Opportunities and Risks of Foundation Model.”
During that process, the thing that I started to think deeply about was this: Here is a model that is fundamentally new in our ecosystem. The reason why this was new was that it wasn't, again, trained to do anything in particular, but its premise was that it could do anything and everything. It was like a stem cell, if you were to use a biology analogy.
I got really interested in this idea that, if we were to really think about what the killer applications were that this particular technology would enable, what would that be? Many of my colleagues were using this for simple classification and simple generation. It's interesting that these models can do that, but from an interaction perspective, it's not that interesting. We've known how to do that for many decades.
What we came down to was that these models were actually trained on this very broad data from the web. These are human behavioral data: social media, Wikipedia, all this kind of data. If you poke at the right angle, then you could see human behavior that would just pop out and be quite realistic, and we'd never seen that before.
Mm.
So that got us really interested. The exercise that we decided to do—and this is something that this particular group of colleagues, myself, Micah Burnstein, and Percy Liang, who ended up becoming my co-founder at Simili, did—we sat down and played this game that we called the time machine game.
Imagine we were to get on a time machine, fast-forward 10 years, and look back. What would have been the single application that would have mattered, that would have been the most interesting and inspiring?
We thought, “What if we could just recreate the world that we live in?” It's really hard to get more ambitious than that. Let's just create a world. That's where we started. Initially, we had this paper that was a precursor to the Generative Agents paper called “Social Simulacra.”
Before you go further—
Yeah.
Were there other candidates for the most ambitious thing in the time machine exercise?
Exercise?
Yeah.
What else could it have been?
I'm just—what could have been—
What were the next—
What was number 2 and number 3?
Okay.
If you remember.
3. The Personal Assistant Bet
There is a close second that we were considering, which basically ended up becoming more of these automation tools, but especially the vision around really personalized agents that would actually—
Mm.
—do things for you.
That's also happening.
It's also happening. But it was interesting for us, right? The reason why we decided to go with the idea of simulation was, first, I was a huge science fiction nerd. This idea of creating a simulation fascinated me. I loved the idea. It's really cool to see a game town like this and just see these agents live in it.
But at the same time, my bet was that if you were to create a really amazing personal assistant out of this technology, what you actually need first is an amazing model of your users. For instance, I told the model, “Hey, can you make dinner for me?” It orders Hawaiian pizza, but I do not like pineapples on my pizza, so it totally failed.
The way for it not to make that mistake is only by having a deep understanding of who I am. I gave a very simple and dumb example here, but you can imagine how this core understanding of people is instrumental. This is how, for instance, our family and closest friends have a good mental model of who we are. That's the basis of our social connection.
So our bet also was that this technology around simulation—
Mm.
—creating accurate representations of people ought to precede the more complex agents that would automate the world that we live in. So that was the bet.
That was a very close second, and I'm still very much fascinated by it. I think there's a lot of interesting work going around. My hot take, actually, is that I don't think we've seen a true personal assistant that's actually useful in ways that meet the ambition of that particular line of work.
I think there are early applications that are obviously interesting, and if you talk to even ChatGPT nowadays, or Claude, they obviously know a lot about us. So a lot of the generation it's doing, I do think it's much more tailored, but I think the ambition is quite large in that field, and I don't think we quite have all the right ingredients just yet.
OpenClaw, all these kinds of—
These agents—
—personal agents, what do you want to see from them that they don't currently have?
I do think it's slowly getting there, but I do generally want them to have a much deeper understanding of the person. Right now, you look at the models—I mean, OpenClaw—it's basically leveraging a Markdown file, and I think it's quite clever, right?
If you look at the Generative Agents paper, this was actually the same intuition that we had. Initially, when we were creating the memory architecture for the Generative Agents, back in 2022, we didn't really quite have the idea of even agentic architecture or the term “agent.”
But the intuition that we shared with some of the work that's coming out today was that we initially thought: Do we want to make the memory into, let's say, a knowledge graph? Do we want to train a bespoke model? All of these kinds of things.
And what we decided to do was forget about all this. These language models are actually quite good at modeling text and understanding and reasoning about text. So just put everything in a markdown file or a text file. You're done.
I thought that was quite interesting—that we could do that—and there's a lot of strength in doing that, but also there is a limitation. It's the way you retrieve and make sense of extremely large data that's difficult and takes a lot of work. So I think that technology is getting better.
I also do, however, think there are certain things you just cannot shape just by prompting the model. To some degree, you do need to touch the parameters of the model itself. So there's this kind of work that I do think needs to happen, and obviously it is happening. The question is, how far can we take it? How do we source data? And how do you also create an ecosystem where people are continuously feeding data to this model, so it's learning about you?
What's the intuition behind why you need to do it in the model?
My intuition behind when you train or even post-train the model versus just prompt the model is: does the model have to learn the underlying physics of the world that it's operating in? So it has to learn new social physics.
The places where it doesn't have to train are where it already has the physics. We trust the physics. It already has the base statistics, but it's just trying to react to an environment. Then I think you can just prompt your way into getting the actions out of it.
I don't think the models that are out in the open have yet learned the complete mapping of the social physics of humanity. This is actually one of the core theses of Simile, right? And one of the core reasons why that is the case is, if you look at the data that the model was trained on, these models were trained on web data—whatever was available on the web.
These are really interesting datasets, but they are fundamentally self-exposed attitudinal data, with some behavioral data sprinkled around here and there. They have yet to learn the really deep behavioral nature of people—not just what people say they do online, but what they actually do in real life. This is one of what I would consider to be the dark knowledge of humanity that we haven't quite captured. It's this kind of data that would also need to get factored into the model creation.
You call it a behavior foundation model.
Yeah.
There's a good one-liner here, but outside of that, what type of data do you need? What are you changing at the model level? How do you go about actually modeling—doing a behavior foundation model?
4. The Behavior Data Stack
We think about data in 3 buckets. One bucket is interview data, for instance. Qualitative, rich qualitative data is interesting. It's not behavioral, but we would literally ask people, “Hey, tell me the story of your life.”
Yeah.
It's what we're doing here.
Exactly. The question that you all asked at the beginning of this interview literally is the question we also ask. Obviously, we ask our participants to go a little bit deeper than how far I went. Maybe I can actually give more of my life story in lieu of this.
The reason why that data is interesting is, by learning about this very long-tail information about people, you actually get a lot of texture around this model—this person as a model. Even understanding their childhood memories, their trauma, or their first love is quite informative in ways that are really hard to predict. So that's one.
Then there are 2 tranches of what I would consider to be behavioral data. One type of behavioral data is observational, so this might be transaction data, or it might be data that you can get by scraping the web. You can imagine why these datasets would be interesting, right? They give you the base statistics of people's behavior.
But then there's the last category of data that I personally think is perhaps the most important, which is the data that basically describes the causal mechanism—the whys of people. Some of this is covered by the interview data, the qualitative data, because people talk about why they made certain decisions. But really, where you get to see the most behavioral aspect of this is in randomized controlled trials, like RCTs.
Imagine you basically have the same setup, but you have a few different variables that you are trying to tweak. Can you actually get realistic human behavior out of it? Imagine you had this particular option. Imagine you're even trying to choose whether you're going to drink coffee or not. The day you drank coffee versus the day you didn't drink coffee, does your behavior change? That's a dataset that describes a causal mechanism.
This is quite important in modeling people. The reason why this is important is that oftentimes, when people come to us—or not just to us, but the reason why people are interested in simulation actually isn't because they want to predict the future. If you're trying to win against the stock market, predicting the future is interesting. But most people, most decision-makers, what they want to know is: How can we shape the future?
It doesn't really help you to hear that your sales are going to tank in 2 quarters. They're just going to say, “Wow, that sucks.” What they want to know is, “Well, what do we need to do now to avoid that future?” That's causal mechanism, and this is also very hard data to come by, right? Because the world is our ground truth, but it happens once.
So in a very controlled setup where everything is equal except for 1 variable, this kind of dataset rarely happens. This is the reason why this dataset is both hard to come by, but also quite important if you're trying to model human behavior.
So behavioral data, I think, is the hardest dataset to acquire. What is out there? What is possible, even? Because you're not going to know a lot of details about my life. I don't even have data for myself on—I want to analyze my own health or habits, and I just don't log everything. So how can you have that data?
So we actually run a lot of randomized controlled trials.
Yeah.
But you put people in a lab? Do they watch them sleep, or what?
We do actually care a lot about the consent process, so people know that we invite them to be a member of this community—
Huh.
—to both share data and also have themselves represented in different forms. But we bring a lot of people to the lab, or virtual lab, where we design experiments that would actually pose them real behavioral decisions.
Often, in these kinds of experimental setups, what makes the difference between what is attitudinal and what is behavioral is whether the stake in your decision is real. That's ultimately what makes it behavioral.
In these kinds of setups, we are inspired by our colleagues in the social sciences, psychology, and so forth. When they run studies, the kinds of techniques they utilize are—for example, imagine there's an online store that you're inviting people to come by. Whatever they purchase in this experiment, they actually get that item delivered, for instance. These are the kinds of things that make the stakes real.
So we run a lot of these experiments, and we also partner with firms. Right now, we also have customers who are quite excited to at least give us a glimpse of the kinds of behaviors that their users exhibit, so that we can get a deeper understanding of how people behave on these different platforms.
I think on the customer side, they have a lot of data about their users and who has bought. They have the action data.
Mm-hmm.
Can you walk us through an example of what someone comes to you for? What questions would they want solved in the process? Do you customize a model for them? Do you have something off the shelf? What does that look like?
Yeah.
5. Enterprise Simulation Use Cases
Today, when people leverage our models, it's often to better understand the population of interest. Usually, at the start of the relationship, we basically come together and hear about what population they want us to model.
It might be that if you're a CPG company selling to all of the U.S., then it might be fairly straightforward: you want to model the general population of the U.S. But at the same time, if there's a vertical or a market that they're trying to go into—imagine they want to better understand, let's say, people in their 20s and 30s living in California—that's a much more specific population.
So we hear about these populations, and we go recruit these people, with consent and incentives, and we basically collect some of their data and create a model of these people. Then what our product allows you to do is query them.
It can take as input a filter that is a description of the population that you want to talk to, just like the one I just mentioned, and an environment. An environment can literally be survey questions, behavioral experiments, or A/B testing. Oftentimes, the core use cases are things like concept testing, to start with.
But also, people sometimes want to do focus groups. One of the fun use cases that we also serve is modeling things like the earnings call for public companies. These are the use cases that we often start with.
Concept testing—is that an established term? I've never heard of concept testing.
Yeah. It basically has to do with when they have, let's say, different messaging, different products, or different ideas.
It's like a marketing exercise?
Yeah.
Okay. Got it. Politics?
We do have a strategic partnership with Gallup. Of course, Gallup is deep into the policy space and so forth. Right now, we have not worked deeply with politics, that area, just yet.
I'm curious if there is demand, or if they would really have different needs that somehow fundamentally don't mix with your existing users or people.
I think there's certainly demand.
Yeah.
But we are very much mindful of how this technology gets adopted and the societal impact that we'll end up having with this technology. I do see politics as an area where a company has to be particularly thoughtful about the way they operate and the impact they make. This is where we also want to make sure that we form enough guardrails and perspective on how to leverage this technology before we go on to serve markets like politics.
I'll give people an example. One of my favorite shows is The West Wing. I don't know if people have watched it.
Mm-hmm.
One of the key storylines is that the president has multiple sclerosis, but they haven't disclosed it; they need to figure out how to disclose it. So they run a poll with a fake governor and ask people to respond to it, and they try to make decisions based on the results of that poll—how well they'll be received and how they should play it.
Right.
They try to figure out: How should we play this?
Mm-hmm.
I'm like, well, I think those kinds of counterfactual things, I would actually use a simulation for this if I could trust it.
Oh, sure.
Yeah.
In that show, how'd it go?
In that show, it basically was a foregone conclusion. They were like, "We know it's bad; we just don't know how bad." Then the poll came back and was like, "It's really bad," and then they just did it anyway.
Part of it is that it's a show, right? So you're—
They're maximizing drama.
—looking at the idea of how bad it could be. Oh, it's horrible.
And, to some extent, I think that is part of the trick—or the challenge—of being a customer of yours, which is that if I roughly know and can intuit—
Yeah.
—what the effect is going to be, do I need you? What sensitivity of the effect do I need in order to make a decision, right? So, for example, if my approval rating is 50% and this negative news item comes out and it drops to 30%—if it drops to 20%, if it drops to 40%, do I care? No. I know it drops. It's negative. So when do I care about simulations?
You do something that's clearly bad, that's not popular, and people are not going to like you. Yeah, I mean, it's like—
You don't need a simulation.
Of course. Yeah. Well, so there are a couple of things. One, obviously, is that there are use cases where, every day, for instance, developers, designers, policymakers, and marketers create assets and new products. It turns out that many of the decisions, in hindsight, are sort of obvious. Yes, of course, this is bad, but we still run those studies because understanding the magnitude and how acute something is is actually quite difficult.
Even if we feel like, of course, this makes sense, I mean, this is the reason why we make so many mistakes. Every time somebody goes online and says something that has huge backlash, you look at that and think, "What an idiot." However, it's tough. That's one.
There's also another aspect here, which is, again, the reason why simulation is actually different from prediction. In simulation, in the ideal-case scenario, what simulation is trying to show is each step of the way, or each step that we need to take, to get to a certain outcome, right? In the most advanced simulations, sometimes the next step that we're suggesting might actually be quite counterintuitive.
The analogy that I sometimes give, and I ground it in a more realistic example, but, as I mentioned, I'm a huge fan of science fiction. I don't know how many audience members have read things like the Foundation series by Asimov.
We've mentioned psychohistory a number of times.
Okay, fantastic. So I might actually be talking to the right crew. If you read the Foundation series, literally the first act is that there's a group of scientists who have found out that—oh, our Galactic Empire is going to collapse, and we're going to have 30,000 years of unrest.
They basically run psychohistory, the simulator that tries to teach them: How can we keep this unrest to 1,000 years? They plan this out, and the first step of that plan is to get the scientists who say, "Okay, this is coming," exiled into this random place in this galaxy.
Terminus.
Exactly. And that's so counterintuitive. What a strange move: You literally sent the group of scientists who was raising its voice about this potential collapse of the Galactic Empire into nowhere. How is that the right first move? Well, it turns out that in this particular simulation, that actually was the move. It's these kinds of things, right?
The reason why this kind of reasoning is possible is because you're showing the step function, or each step that results in a particular outcome. So really, what simulation allows you to do in its highest form is give it not a problem or a question, like, "What would people answer to the survey?" That's not what we do.
What we tell it is: Here is a goal that we have. In the context of Foundation, we want to keep the unrest to 1,000 years. What is the path that we need to take now to get to that particular future? And that's what simulation allows you to do.
Now, translating that into a real market, imagine you're an automobile company and you're about to release an EV, and you're trying to understand, well, how do we market an EV to make sure that our stock price goes up? But what if the answer comes down to this: You can market your EV in an XYZ way, but that might change people's perception of cars that aren't EVs and actually make your overall sales go down?
Not very intuitive, especially if all you're trying to optimize is EV sales, and that's the only thing that you're tracking. That might actually result in a completely wrong solution, or at least a different solution than what you would have expected, whether it's right or wrong. That's the power of simulation.
For listeners, we covered a similar topic with Mikhail Parakhin from Shopify, where they are working on SimGym. I don't know if he ever talked to you about it. It's very similar.
The journey is unusual.
The goal is increased conversion, but then the journey is very unusual.
The journey is unusual.
Yeah. He's actually trying to look for interventions on a shopping trajectory, which is similar to what you're saying. It's not about the attitudinal—is that your word for it?
Yeah.
It's about behavior.
It's about behavior.
And that's exactly the difference, right? It's not about the near-term direction, but it's more about how you affect multiple turns of interactions.
Right. You had a good quote at the start about this as well. It's not about people wanting to know the outcome; it's about how they can change it, change the way to get there. Something like that. But I want to take it back to how do we know this is grounded?
Yes.
How do you run evals? How do you test that simulations come through? Basically, if I were to do the same thing that you described—
Mm-hmm.
—with, say, your favorite LLM, Opus or GPT-4.5, and have some agent map out these things—
Yeah.
—how different are the answers we would get? If I give it the same goal, the same objective, and make a decent system, you're saying that you need to change the model weights. You have your own solution to this. But how far off are we, and how do you check if it's grounded?
You have some interesting stuff on your site that actually points to how you run real evals, but if you could take us through that side. I think that's one of the big concerns that people have. They're like, "LLMs hallucinate."
6. Testing Digital Twins
The way we do this—and this is actually the paper that we worked on after the Generative Agents paper that really became the foundation, at least for Simile and also for the field of simulation and synthetic panels—is this paper. The paper is called Generative Agent Simulations of 1,000 People.
Here's what we've done. For this paper, we actually brought 1,000 people who were representatively sampled from the U.S. to a virtual lab, and what we basically did was spend 2 hours collecting fairly wide-ranging data.
In this particular study, we focused a lot on the interview data, whose script was taken from a project called the American Voices Project. We paired that with a lot of behavior data and whatever else we could collect within 2 hours. We then sent these participants away for a couple of weeks, and during that time, I used this data to create their digital twins. We brought the human participants back after 2 weeks and had them complete a battery of surveys, experiments, and behavioral studies.
We actually have the list here, which included things like behavioral economic games. We ran Big Five personality tests and the General Social Survey. We also ran the randomized controlled trials that were published in PNAS. We had their digital twins predict how the source individuals would have acted in these studies and surveys. This is where we found that we could replicate people’s behaviors and attitudes with 85% accuracy, about as accurately as people could replicate their own.
That was the first paper that gave us really validated results showing that we could model individuals accurately. Of course, in the AI space, this paper came out at the end of 2024, and 1.5 or 2 years is a lifetime.
Yeah, for listeners who aren’t watching YouTube, I just want to say that the headline figure is 85% accuracy, which is a big improvement over all the other methods that you showed.
Yeah.
You’re solving Moravec’s paradox.
That’s very hard.
Do we want to keep going on the paper route?
Yeah, for sure. But what was particularly striking to us, especially as we improved this technology even further, was that generative AI models like ChatGPT and Claude do give you the right foundation. However, what they don’t consider is the true attitudinal and behavioral aspect of people, especially in the population that you care about.
What these models are really, really good at today is trying to become these super-rational, objective machines, right? They get their data from places like Markov scale. You talk to professional programmers and scientists to create models that are amazing at reasoning. That’s what they do. Simile actually doesn’t care about any of this. The models we’re talking about here—what we’re trying to create—are models that are as dumb as I am, right? If I make a mistake, the model has to make the same kind of mistake.
Oh, that’s very hard.
That’s exactly it. This is actually a completely different kind of data and training objective. This is also where we see quite a bit of discrepancy in performance in human behavior prediction between frontier models and Simili’s models—the models being created in this space.
In some cases, the performance of frontier models goes all the way down to 20% or 30%, especially if you go into a more niche population and topics that our customers would actually care about. For the general population, it might be around 50% to 60%, so it’s not very robust. You wouldn’t want to make your decision based on these kinds of findings. If you can bring that up to 85%, that’s ultimately what people get very excited about.
Yeah. Do we want to keep going on the paper route?
Yeah, for sure. The last one was an interesting one. This was the follow-up to the Generative Agent Simulations of 1,000 People paper. The idea was: can we augment the models even further and post-train them based on a lot of randomized controlled trials? This was an interesting one. The data is always the most interesting part of modeling, in many ways. The data we got here was from a platform called the Open Science Foundation.
Some of the audience might be familiar with this. Especially in the social sciences over the past 5 years or so, there has been concern around the replicability of studies. It’s a bit of a crisis that scientists have acknowledged, where we rerun the study and don’t actually see the same finding.
Oof.
It’s tough. The reason that was often the case was basically survival bias: the papers that get published often need to maintain what we call a p-value of less than 0.05 in the experiments that we ran. That basically suggests that there’s only a 5% chance that the results we saw are a false positive.
But the tricky part was all the papers that were not published, and there’s still a 5% chance that whatever we publish is actually totally randomly generated. There’s a 5% chance that this effect is not real, but it happened to appear real because of sampling bias.
Because of that, scientists started to preregister their studies. Before running an experiment, they would go to this platform and say, “Here is the data, here is the population we’re collecting, and here’s the hypothesis—this is what we believe.” You cannot retroactively change those hypotheses. This is what actually gives us more scientific and statistical confidence that whatever effect we ended up seeing is actually true.
That created a really interesting platform containing tens of thousands of real-world experiments and hypotheses. A lot of these are actually high-quality, professionally designed behavioral studies and randomized controlled trials. We got the data and the studies from this platform and used them to make a point. Obviously, this particular model is not something that we’re serving commercially, because it was part of open science. But this particular data set helped us make the point that by collecting a lot of these well-designed randomized controlled trials, we can significantly improve the model’s ability to predict human behavior. That’s what this paper was about.
Is this stuff done at an individual level? Do I need to tune the model per individual, per company? Are there changes to the foundation model followed by some slight post-training? Anything you can share there?
This particular model was trained on data at the level of individuals, but we experimented with both, and this is what we end up doing at Simili, too. We always train 2 distinct models. One is what we call the population-level model, and the other is what we call the individual-level model. Both take very similar input, which is a description of a subpopulation or individual and a stimulus.
In this particular work, we did the same. The results we’re reporting are much more geared toward individuals because we do think that’s a harder task in many ways. But that’s what we’ve done.
Have you seen anything involving questions that humans can solve but models can’t? For example, I live 5 minutes’ walk away from a car wash. It’s a 10-minute drive. Should I walk or drive?
Uh-huh.
The model will say, “Walk to the car wash,” and you don’t have your car.
Yeah.
Is anything like this a problem in simulation? You’d assume it’s very simple for a human to think about, but if the model says you should walk to the car wash, you know. Anything here?
It’s less about what we can solve, but I think it’s more about what biases or mistakes people make that models miss. For instance, imagine that when I was still at Stanford, I lived in Palo Alto, so it was about a 40-minute walk from campus. You ask the model, “Okay, let’s go home. What can I do?” It would likely call an Uber or give me the bus times.
But for the longest time, I really liked walking back. The reason I wanted to do that wasn’t efficiency. It actually really helped me think. I like to walk for half an hour or 40 minutes a day, where I get to think about ideas and research and get lost in my thoughts. That’s a very human activity.
Unless the model has seen that and actually understands the importance of that activity, it would miss these kinds of features. That, I think, is fundamentally what we’re trying to model: what is fundamentally human. It might not be the most efficient thing to do, and it might not be the right thing to do, but these are the things that make us who we are.
I’m curious whether there are some data sets that you really want that would materially help you. One version of this might be interesting: which would be more valuable for you to acquire as a data set—all of LinkedIn, all of Twitter, or all of Facebook?
To be honest, it’s a little bit hard to rank, partly because there’s this product saying that no feedback is bad feedback, because it teaches you something about your users, no matter what kind of feedback it is.
Mm.
I think it’s a little bit like that.
So just whatever is bigger.
What about a different domain? Say it was all of Amazon’s data?
Oh, yeah.
Shopping data, right?
Shopping data. Amazon data is interesting in that it’s very much behavioral. Although what people do on social media, you could squint and say that’s also behavioral. But transaction data is always interesting. It is also most commonly available, however.
Yeah.
If we were to look at purely social media, if I really had to pick, Facebook would likely be interesting because I do think it is more of a default version of people. You go to LinkedIn, and it’s very much a professional environment, so people put up their guards, right? That’s still interesting because it reflects true human attitude and behavior, but it is not your base state.
You go to Twitter—in Twitter, people have their own crazy personas, depending on who you are. My Twitter profile and persona were initially very much academic: “Hey, I’m here to share my studies.” Now I share things related to Simile. But Facebook is one of those more private spaces where people just connect with their friends. In that way, I do think it shows you a little bit more about who that person is. So if I had to pick, I’d likely pick Facebook.
Yeah. You’re interested in the whole person and their background and philosophy. Is it too clinical or too machine-learning-oriented to say that this is just a way to inject variance and biases? The broad question, I guess, is: Is this any better than a randomized, combinatorial-explosion version?
We have a link to Tencent’s billion-persona paper, where they basically did not do any of the groundwork that you are doing.
Mm-hmm.
They just did a cross-product: Here are all the professions in the world, here are all the possible backgrounds in the world, take dot products across all of them, and that’s it. That’s your prompt for a billion people.
Yeah.
This will do something. I don’t know if it’ll do what you do, but it gets you some percentage of the way there.
This was actually an interesting paper. What I admired about it when it came out was the scale. Obviously, you gradually do want to be able to simulate really large societies and interactions.
Yeah.
The scale is definitely admirable. It is relying heavily on the known statistics that went into training the model. To the extent that you believe those statistics are correct, this is actually not a bad way to go about it. But the thesis here—and this is something that we have also seen in the market—is that if this works, then we have actually solved simulation.
Right.
Because I survey, say, 5% of the U.S. population is in construction. Another 5% is in medicine, whatever, right? Then you just keep going down the list, and then you do the other side: 5% has the Big Five personality traits, neuroticism or whatever. That’s it.
That’s it. So if you believe that the underlying data set and the platform we’re leveraging have all the right statistics, then this will actually have solved it. At that point, you’re merely retrieving the knowledge that is already embedded in the model parameters.
That’s not, unfortunately, what we see. There is such detailed and niche knowledge about people that, if you just take one example, it might feel very mundane, but it’s actually quite rich when you put it together. You do need to do a lot of bespoke data collection to better understand people.
This is also what makes this particular job fun. You want to deeply understand people, and the process of deeply understanding them requires a lot of attention to detail. You do need to pay attention to and respect the daily lives that people lead.
I want to talk about scaling simulation. What can’t we simulate? What can we simulate, and how does scaling affect all of this? How big are the models? What if we go from 8B, like a couple hundred million, like hundred billion parameters, trillion? Do we get any interesting emergent scaling at a certain scale or a certain amount of training? Do you uncover anything unusual, any learnings from that?
7. Scaling Society Simulations
What we’re seeing at Simile is that we do post-train our own model. We’re seeing an early glimpse of scaling laws in simulation. The more data about humans and more compute you ingest, you start to get predictable gains in model performance when simulating and predicting people.
Ooh. We need a scaling walker.
It’s a scaling law. Whenever you find one, it’s a beautiful thing. We’re starting to see a glimpse of it, which is quite exciting.
But if you talk about the ambition of simulation as a whole, it’s not merely about building a model. It’s about building a model, then creating the agents that become the individuals in a much larger ecosystem. So you’re basically creating this multi-agent simulation. Down the line, you want this multi-agent simulation to also live in a very rich environment, right?
What we’re really trying to get to at that point is: Let’s do a time-machine game again. Five or 10 years into the future, can we create a simulation of 8 billion people living on Earth? I think that’s quite interesting, and that really is the vision.
Once you get to that kind of state, the kinds of questions that you can help answer for society also start to change, from my perspective. The answers are fundamentally about the emergent behavior of society and large groups of people.
Yeah.
For instance, the kinds of questions that I get excited by—and maybe this is still a bit of my academic side—are questions like: Can we help solve climate change? If you look at climate change as a problem space, this is what social scientists would often call a wicked problem: one where you have many actors with competing incentives who are trying to make a very complex decision, and coordinating that decision is very difficult to do in real life, which is also the reason we couldn’t solve it. Can simulation help us solve that?
Another question is: Can we actually understand the signals for a collapsing democracy? Or can we understand or uncover the origin story of the monetary system? These are societal questions that we never really had a good way of answering. If we can create simulations of our society, you have to believe that these are the kinds of problems we can solve.
That’s really the ambition of this field. I also think there’s a Nobel Prize to be won there, which wouldn’t be surprising. There’s some amazing societal impact we can have to help people make better decisions.
Nobel Prize in economics?
In economics.
I see. I see. We’re rooting for you to write that paper.
One of these days. One of the scholars I was deeply inspired by when I was coming into the space of simulation was a scholar named Thomas Schelling.
Schelling point.
The canonical example of the work he did was that he was one of the creators of agent-based modeling. This was in the 1970s and ’80s. It was very early days, but this was truly one of the first exemplars of simulations.
One of the canonical models from that time—and, of course, many of these simulations are trying to tackle the societal problems most relevant to their era—was called the model of segregation. Racial segregation was a big topic that we cared about. They created this grid world with red dots and blue dots. These dots were, back in the day, the agents, and they had a simple rule that governed their behavior: If a certain percentage of your neighbors are of a different color, and that goes above a certain threshold, then you move to a new location at random.
One of the striking findings of this paper, or this agent-based model, was that, for the longest time, people thought segregation within society was caused by explicit and overt racism. But if you look at this model, people’s preference for living with people of the same color can be very minute, and that very small difference actually causes society to segregate completely over time. This was very counterintuitive to a lot of people, and this particular work ended up informing housing policies. Mixed-income housing, for instance, was really inspired by this kind of work.
Thomas Schelling ended up winning the Nobel Prize for having laid the groundwork for very early versions of simulations. The opportunity I see here in more scientific terms is that agent-based models had an impact for the longest time—in the 1980s and ’90s, and to some extent the early 2000s—but they’ve now gotten a little forgotten by the community because, as you can imagine, red dots and blue dots are not really a rich description of people.
But with the emergence of things like generative AI and, in particular, generative agents, we do have an opportunity to create these kinds of agent-based models that are high-fidelity enough to help us make really complex decisions, and that's the opportunity that I see. If that truly works, then yes, that is the kind of work that will result in a Nobel Prize.
Yeah. For what it's worth, I grew up in Singapore. 80% of Singapore is in public housing—
Yeah.
—and public housing has enforced racial quotas for exactly that reason, which is very interesting. Okay, so we talk about scaling. We talk about all these sort of agent-based applications.
Mm-hmm.
I'm scared about the cost. Let's just keep it to the US—not a billion people. How much does it cost to model so many hundreds of millions of people?
Oftentimes today, obviously, we don't start at that scale. At this stage of the industry and simulation as a technology, we can actually give our users extremely rich and meaningful insights even by modeling thousands or tens of thousands of people.
Today, what we do is, every week, we collect data from tens of thousands of people, and we actually have panel partnerships that get us access to tens of millions of people globally. That's what we do today.
And just as a side note—
Mm-hmm.
—once you've collected one person for one study—
Yeah.
—can you reuse that same person for all the subsequent studies?
That's exactly right.
Okay.
The beauty of this model and these agents is the fact that they are domain-agnostic. What you're really trying to understand is the fundamental nature of these people—what's their social physics?
Obviously, there's a lot about people that does change over time. Even things like, how many times have you been to CVS in the past week? Obviously, that will change. But there are so many traits about people that are also known to never change. Your risk tolerance doesn't really change over time; it's very consistent. So it's these kinds of things that we're trying to learn.
The scale we are operating at right now is tens of thousands to hundreds of thousands. And in many of the core use cases in which we are deployed, this is more than enough of a population to cover those. Really, at that point, what you care about is less the number of people and more whether you have the right subpopulation of interest covered.
This is also the reason why people want a larger sample. It's not because they actually want stronger statistical guarantees; it's more about whether they can actually filter down to any population of interest.
However, you can also imagine in 10 years, if we truly believe that compute is going to scale, that we'll have much more availability for compute, and our ambition for simulation is also going to scale accordingly. I mean, there's definitely a reason for us to create an entire data center's worth of simulations.
My hunch here is that, in the next several years, we will start creating simulations that will actually cost as much as training a foundation model. But perhaps it's going to be so valuable to society that it would be a no-brainer. Right now, even today, we're training a bunch of new foundation models just so we can say we trained one, and we're spending tens of millions.
But if we can create a simulation at the level of society that would actually solve climate change, I would run that today. I would raise the money right now just to run that.
Amazing. I guess the follow-up question is, does it also compound if you let the simulations talk to each other? Or do they already do that today? They don't, right, as far as I understand.
It depends on what kind of simulation you're trying to run.
Yeah.
In the multi-agent simulation setup, the agents do talk to each other.
Right. Which is exactly Smallville, right?
That's right.
But a lot of times, for example, in e-commerce, you're just by yourself, so there's no point talking, which is way cheaper—
But people are always social, right? You decide what you buy based on what other people around you buy and talk about, right?
It depends.
It depends.
I'm coming at this from a cost point of view. I'm like, oh my God—if there is some combinatorial thing of thousands of people talking to thousands of people, then that 1,000,000× may cost—
I think—
—I have a very different view when it comes to cost. Running these studies in reality is actually a lot more expensive, right? Running any study like this, you gotta have people do it, you gotta sign people up. It's very expensive and sometimes not feasible to actually run the study.
But the outcomes or decisions you make have very expensive consequences, right? So spending X million on something where the overall process costs $100 million might as well, right? There's a lot of value to be had there. It's a small cost, but I'm excited about the cost side, actually.
To some extent, yeah. Obviously, when you deploy technology, you often want to deploy it in a way where you can replace existing budget or you can basically make things more efficient, and that is the best way to deploy.
However, the way you capture the long-term value of the technology is actually by making the argument that now it's the upside: by making this better decision using simulation, you have saved yourself or made yourself hundreds of millions or even billions of dollars. That's a case to be made.
Random tangent question. So if you're doing a lot of inference, a lot of multi-agent model stuff, are you at the point where it makes sense to train a model that's very sparse, expecting to do multimillion-dollar runs? Are you thinking about this from a model architecture standpoint or inference efficiency? Or are you still at the research phase of, "If it works, it works; we're not super there yet"?
Efficiency, we actually do think quite a bit about. I mean, this is technology that is deployed now in some of the largest enterprise companies in the world, and we do process a significant number of queries that are trying to simulate the populations in the world. So efficiency is a consistent consideration.
Obviously, we don't want to over-optimize too early, so I wouldn't say this is the highest priority right now, but this is definitely something that we think pretty carefully about.
Yeah. Are there other case studies? You talked about CVS, Gallup, and Deloitte. Worldfront?
Walmart is an interesting one because one of the things they were trying to do—they were one of the first customers that wanted to actually do product testing that goes beyond just asking people what they think about, let's say, behavioral experiments and so forth.
So there, really, what we had to do was reason about multimodal inputs: images, but you can also imagine these agents traversing through Figma mockups or websites. Some of the things our agents can also do are be given a domain, like a website URL, and actually use it for a while. Worldfront was one of the first customers that was very excited about this possibility.
Have people been asking, is there any demand that we have not covered? UI testing, right? I want to try a new feature, ship a new feature, test the UI, simulate how people will use it. Any interesting things that you're seeing demand for?
Today, a lot of the demand does come from basically the places where people have historically used human panels. We can now basically replace those with agents and these synthetic populations. This is obviously not replacing human panels.
In many ways, the simulation that Simili is building is grounded. The way that I think about this is that we are trying to represent humanity at scale. In that way, the use cases are what we would expect, but it's the scale of deployment that surprises me.
It turns out there are so many decisions that people make every day in these organizations and groups, and we want to be able to say, "We listen to people. We have consulted our users." But in reality, that is rarely the case.
Because getting to people and actually asking them many questions is difficult. It's both costly and time-consuming, but most importantly, people are just not available. If I had to answer 1,000 survey questions for this one particular vendor, even if I wanted to do that, I would never do it. And that's very much the case.
What simulation can do is ensure that the voices of people are always represented in rooms where the decisions for them are made, right? So all the stakeholders of this particular product launch, ideally, they are consulted. That's what this technology really is trying to enable.
In my mind, that means it skews toward more consumer focus, right? Anything with a wide enough customer base where you benefit from the diversity that you represent. What are some rough statistics, just for people who are not familiar with this market in general? What's the market size that—
I'm sure you have some rough numbers. Obviously, market size is a vague question.
Yeah.
But how much do people spend?
So market research is a $100 billion industry. But the thing about simulation is that simulation is not a tool for market research. Simulation is a tool for human decision-making. So the question around what the TAM here is actually quite tricky, right? Because it's easy to say, “Well, the market research TAM is roughly $100 billion or $100 billion. Is that a TAM?” And not really, right?
In many ways, you're trying to inform all human decision-making. You're trying to basically inform every decision that is made about humans, for humans. What is a TAM for that? Really unclear. I'll be honest: I have a scientific background, I have a research background, so I didn't come into the field calculating, “What is a TAM for human decision-making?” But I just had to assume, well, if we can inform every decision that is made about humans, for humans, that has to be big.
Something valuable.
Exactly.
I mean, to some extent, you are a unicorn founder now, and you have to care as a CEO. But I do think that, when you go into these boardrooms with people that you're quoting millions of dollars in contracts for, you have to say, “Well, here's what you spend on humans—and here's what we save you,” and it's 85% similar.
Certainly, the value case is something that we care deeply about: What is the value that we actually provide to the users and the decision-makers? But this is also where, as a founder, I think valuation only tells one very superficial aspect of the story, and I try not to think too much about valuation in general, because that's not what motivates the team.
The interesting thing about researchers is that we're happy living in academia. We get paid okay. I mean, we don't get paid that much as a researcher if you're in academia, but it's the impact and the value that we can provide to individuals and society that really drives us. In that way, ultimately, what drives us is the impact. Does the simulation we provide have a real impact on people's decision-making in ways that progress our society? If the answer is yes, then yes. I mean, that has to be great business, and we see that in numbers, and we do care deeply about that upside story, but that's the harder bit.
Do you have any timeline predictions? We talked about scaling laws of simulations.
Yeah.
You brought up, okay, maybe one day we can simulate how to solve climate change. Where are we now? If that's not the end state, what is an end state, and what does progress look like?
So what I sometimes tell people is that the simulation industry feels a lot like where GPT-3.5 and GPT-4 were for the AGI saga. We now have technology that is powerful enough to do real damage on the verticals that we are tackling. At the same time, there's a lot of progress that is yet to come. And that's, I think, where this is.
The way I see it, I do think there will continue to be breakthroughs, both in data and, obviously, in algorithms, and there will be much more aggressive scaling over the next few years. But I think that's roughly where we are.
I think that was about the rough set of topics. Is there anything else that we should have asked you, or that you wish people asked you more about Simili?
8. Simulation As Human Meaning
I think what's actually fascinating about simulation is that it is very impactful technology, but it is also very interesting technology, both in terms of what it means for human society and our philosophy.
The way I sometimes interpret simulation is—going back to my background, I actually started my career as a painter. It was a professional pursuit, and I did oil painting for figures. So I got my training originally in realism studios, and that's what I spent a lot of my years doing. Simulation is a lot like painting, right? The best paintings teach you something deep about the subject that you're trying to represent. It is never a perfect representation. No painting is perfect. There's always some small difference or discrepancy. But what it does is try to highlight the thing that matters most about the subject.
The essential essence.
The essential essence.
You brought up some of your work.
It's nice to put it up.
Yeah. So these are some of the works. This is actually from my personal website that I maintained when I was still a researcher.
I think a lot of people will say that a Picasso, like anything postmodern, is very much focused on the essence.
Right. But I don't know if any one of these evokes something that you'd like to tell the story of.
No, it's one of those things where each of these paintings, drawings, whatever it may be, is trying to surface something about the subject that you feel deeply about.
When I was a painter and artist, the topic that I cared really deeply about was the more mundane aspect of human lives. This shows up in some of the work that I've done, where I did an entire study of a rural town. I basically went around and took photos of people who weren't really doing anything special, but just living their everyday lives. I thought that was the most interesting thing.
I'm somebody who has this perspective where the world is oriented around this fractal shape, and you have 2 choices for understanding the fractal shape. You either go outward and try to explore as much as you can to understand the broader shape of the fractal, or you go inward because the outward resembles the inward shapes. Understanding the mundane aspect of it was very much that.
Simulation has a lot of this, right? You're trying to understand even the most mundane aspects of people. When put together, they teach you something really deep about that individual and society. I think that's what's interesting about simulation. In the same way that AGI helped us better understand, or really think critically about, humanity and human intelligence, simulation is really an exercise in understanding more about human society and our collective lives. I find that particularly interesting.
Yeah. Now you're reminding me that some of the best biographers, documentarians, and even photographers take a photo of you. But before I take a photo of you, I must follow you for a week just to understand you, which some artists do.
That reminds me of a very famous book called Working. I don't know if anyone has referred you to it before. It's very, very famous, to the point of having a Wikipedia page, about this kind of really in-depth understanding and interviewing of people about their lives, which seems mundane but is told in a very compelling way. It's from the 1970s, as well.
It was an amazing decade.
Actually, before the closing question.
Oh, go ahead. Go ahead.
You said that you started Simile with your 10-year question, right? If we do that now, 10 years down, what can we simulate? What would you simulate if you've made significant progress? Are there any questions outside of the ones that we brought up? Anything that you think is most impactful? Anything that you would envision 10 years out?
In many ways, as I mentioned, I am somebody who is very much impact-driven, so what would actually inspire me is this: I would want to ask, 10 years later, what would actually be the most important societal question that we as a society have to ask? I would love to tackle that. For instance, do we need UBI? That could be an interesting one.
Ooh, has anyone done that?
Well, we're thinking about it.
Can we get access?
I think Sam Altman actually funded a study on this—
Yes.
—in Africa, and the answer was no.
The answer was no. But was it something about the implementation?
Yeah, I know. It was a scale issue.
But this is the thing. See, when Sam—
Funny news article.
Altman funded this particular study—
He spent $14 million? Oh my God.
That's a lot.
Quite a bit. But this is the thing. This is the reason why you want to run a simulation. You spend 5 years and $14 million on this 1 study and have 1 finding. But if you can run a simulation many, many times instantly, then that's the value.
Mm-hmm. I feel like that one could have been done in a simulation.
If you can do the housing study, you can do the UBI one. I mean, come on.
I think sometimes people will spend the money because they want to verify what they think, right? Sometimes you just want to know: is it actually right? You have to test it.
Okay, closing question: What are the chances we are in a simulation right now?
It's a fun question, and I started—at some point, I just answer, “Yeah, we're definitely in a simulation.” But what I do feel, however, is whether we are in a simulation or not, I don't think that makes our experience any less real. And I think that's fundamentally what I believe in. Maybe we live in a simulation, maybe not, but—
It's real to us. Yeah.
For me, yeah, I don't really care.
Yeah. Unless you die and you wake up at a higher level or below.
That'll be interesting.
I feel like you wouldn't care, you know? I mean, once you die, then you find out you're in a higher level, like—
I worry about it when I die.
Yeah.
I think the other thing is that I like the mathematical answer to this, which is, like, the sheer number of possibilities that you are in a simulation far outweighs the sheer number of possibilities that you're not.
Yes.
Except for the simplest answer, which is that it is computationally very expensive to have you be a simulation.
Okay, great. You've been very generous with your time. Congrats on all your success. You know, I met you just after your paper, Smallville paper and had no idea that you could build such an enormous company.
And now you're like, “Well, it's a $100 billion market, but that's just where we're starting.” So this is very exciting.
A $100 billion market was not the TAM. That was only part.
Yeah, yeah. Exactly. Exactly. It's that you're thinking too small.
Well, I do believe that. Maybe my final note here might be that, again, I love science fiction. You look at any advanced civilization in science fiction, and there are two twin-pillar technologies. One's AGI in some form, and the other is simulation. So I think the market's pretty big here.
Yeah.
Tell us about the company. You guys just raised a lot. You're half a research lab, half a company. I guess you're hiring. Where are you based?
9. Building Simile
Yeah. We're based in Mission Rock, so not too far away from where we are right now. We're in San Francisco, but we are also bicoastal. We have our headquarters in San Francisco, and a lot of our technical talent is in San Francisco. We also have a smaller office that just opened up in New York.
As a company, we're an interesting one. Today, obviously, there are AI neo labs, and then there are AI product companies. Simili truly is both.
This is a company that was founded by 4 co-founders: myself, Michael Bernstein, Percy Liang, and Lainie Ellen. Michael, Percy, and I are all researchers. Of course, Michael was one of the co-authors of ImageNet, which kickstarted the AI revolution back in 2013, and has been instrumental in human-centered AI. Percy coined the term “foundation model” and obviously is one of the great AI researchers today.
Lainie is my business counterpart. She led some of the fastest-growing AI-native companies from their C to A and B.
But we have this DNA at the company where the vision for the technology that we're creating is continuously developing. We are getting people who were basically my lab mates. Right now, about 60 or so people—15%, almost 20%, of the company—are actually just my lab mates from Mechanical Process Lab.
It's actually quite fun because many of them had gone on to OpenAI, Google Gemini, and these places. It's been a few years since we really got together and had a chance to work together, but now they're coming back and really building out this vision that I find to be quite exciting, and that excitement is shared.
There is that motion at Simile where we are a group of researchers trying to do something that no one is working on, that we find to be potentially the most impactful. But at the same time, this is technology that can make an impact today.
We have an amazing group of engineers, product people, and designers who are sitting here with us, basically trying to imagine what it looks like to help people understand what simulation can do and make real-world decisions with it. Having both, and then deploying it to some of the largest customers in the world today, feels quite unique.
Yeah. It's very compelling. One part of it is the call to action: Who are you hiring? You've done part of it, which is that you've got a very talented group. Who are you hiring? What roles?
Honestly, at this point, we are hiring across all sections.
Everything.
We're always excited to bring on amazing research talent.
Yeah.
So if you're interested in working with our lab mates, we're always welcoming amazing researchers. But we also hire amazing engineers, some of whom I like and respect the most.
Many of them actually come from places where we have personal connections, so many of our members are from Figma, Notion, Harvey, and so forth, but also more broadly from the companies that we as a team have really admired. Engineers both on the product side and the infrastructure side—we're looking for those hires.
Well, lots of people. I think you made a really good case, so thanks, and we'll see you in the simulation.
Amazing. See you all there.