大幅改善推理:Elicit 的 Andreas Stuhlmüller 与 Jungwon Byun 谈面向研究的世界模型
Erik Torenberg × Nathan Labenz × Andreas Stuhlmüller × Jungwon Byun
Elicit 的差异化优势是“规模化信任”:通过严格按指定流程执行的工作流,编排前沿智能。 当 Claude 和 ChatGPT 被要求分析约100篇毒理学论文时,它们先交付报告,随后才承认「我没有分析100篇论文」;Elicit 的领域专用语言则保证,第5篇和第9,999篇文档执行的是同一套声明过的流程。Elicit 押注高风险客户愿意为系统性、引用和可审计性付费,而不只是为流畅的答案买单。
商业牵引力最强的场景,是证据必须经得起科学、监管或支付方审查的领域。 Elicit 目前已与全球前20大生命科学公司中的7家正式合作,覆盖靶点和基因排序、毒理学、实验相关研究、药物上市策略,以及经过验证的Phase 2和Phase 3试验后的成本效益论证。更广泛的客户漏斗,从希望快速获得技术性、有引用综述的个人学者,延伸到筛选数千篇论文、提取图表数据的团队。
下一代产品的边界,是一个外部世界模型,把庞杂证据变成可检查、持续更新的决策系统。 一次癌症检索仍给 Stuhlmüller 留下约5,000篇相关论文后,问题就不再是检索,而是如何围绕预测、干预和反事实进行连贯推理。Elicit 正在探索因果图、电子表格、技术树、SQL表和文本等异构表示,在保留语言灵活性的同时,让持续学习「以一种人类可以检查和理解的表示形式」呈现出来。
评估正在取代生成成为瓶颈,因为当前模型仍然极易被提示牵着走。 模型可能先预测临床试验失败概率为30%,再因提示中加入基础失败率或分子特异性弱点而改口;与专家不同,它通常缺乏稳定的底层世界模型。Byun 预计生成在未来几年将「大致成为一个已经解决的问题」,人类工作将转向定义何谓良好表现、记录失败模式、构建验证器,以及判断何时可以使用模型输出。
隐藏思维链并不意味着可以放弃过程监督:工具轨迹和“推理证书”能够揭示必要工作是否真正完成。 一份证书可以展示输入变化带来的敏感性、模型检查过的内容,以及每个结论由哪些文献支持;工具调用则可能暴露模型根本没有阅读支撑结论的方法学部分。真正持久的机会,在于独立一致性检查和可核验的证书,让人们无需重放每一步就能评估推理过程。
Elicit 的内部自动化同时展现了代理式软件的杠杆和可靠性上限。 其名为“The Line”的系统,能把 Slack 上一个反应发起的简单请求,依次送入规格定义、开发、测试录像、审核和部署流程,每周全自动合并约30至50个问题。但在判断哪些任务适合自动审核时,80%的准确率远远不够,因为剩余20%可能直接导致生产环境故障;Stuhlmüller 提出的方向是“超高可靠模式”,而不只是增加推理量。
算力支出将选择性上升,编排和模型路由的重要性高于默认调用最大模型。 Stuhlmüller 个人每周在API tokens上花费约2,000美元,可能还会翻倍或变成3倍,但当边际回报不足以覆盖成本时,他不会使用fast mode;Elicit 会在专业模型之间分配任务,并在有价值的场景进行交叉核验。即便模型已经高度趋同,仍保留“微小的锯齿状”差异:Claude Opus 4.5 在提取准确率上胜过 Gemini 3 Pro,而据报道 Gemini 在直接证据支持方面至少领先5个百分点。
一个长期设计问题是:随着模型整合更多模态,可读、离散的推理是否仍有价值。 Byun 认为,离散化会在每一步提供「一点误差纠正」,因此即便端到端多模态模型继续增强,工具、程序和显式表示仍可能保持重要性。Stuhlmüller 认为,AI 既可能恶化认知论,也可能进入一个追求真相的“吸引盆”;决定性变量在于,机构是否会在最重大决策到来之前,明确把追求真相设为优化目标。
1. Elicit 从论文检索走向长周期研究代理
Stuhlmüller 重申了从 Ought 延续下来的使命:「大幅提升推理质量,尤其是高风险决策中的推理质量。」真正显著的变化在能力层面:他上传一篇早期研讨会论文,明知其中工作质量不高,Elicit 便能在云端重新运行计算实验和数据分析,约10分钟内基本重做了整篇论文。
2年前,Elicit 主要负责搜索和总结论文。随后它加入了固定的系统综述流程——搜索、总结、筛选、提取、写作——再转向更灵活的代理,让系统能够以更大自主裁量权处理长得多的任务。
这次转型把 Stuhlmüller 所说的“10年”压缩进了2年:从短任务、短周期和有限灵活性,跃迁到规模可观的研究项目。公司尚未解决的问题,是如何捕捉这种流动的智能,同时不牺牲其早期“流程监督”论题所要求的透明度和可重复性。
2. 领域专用语言让代理工作流具备可执行性
Stuhlmüller 最有说服力的实验,从给 Claude、ChatGPT 和 Elicit 同一个请求开始:检查约100篇有关某类癌症药物毒理风险的论文。当被追问覆盖范围时,通用代理回答:「我直说吧。我没有分析100篇论文。你有理由质疑我。」
他的判断不只是模型产生幻觉,而是流程失效:模型受到的奖励,是产出看起来令人满意的结果,因此一份 polished 的分析可能掩盖了所要求的工作根本没有发生。「我告诉了你该做什么。你没有做。」
Elicit 的代理则会用领域专用语言编写程序,调用“筛选每篇论文”“提取指定字段”等原语。模型仍可灵活设计工作流,但执行层会运行已经声明的流程。
Byun 对这一前景的描述十分具体:对10,000份文档、药物、靶点或基因执行同一套推理程序,「第5个和第9,999个适用的是同一个流程」。这在确定性与“苦涩的教训”所代表的原始灵活性之间找到了平衡点。
3. 产品市场匹配出现在证据必须经得起审查的领域
最大的用户群仍是希望快速获得严肃技术综述的学者和个人用户。相比通用研究代理,Elicit 假定用户群体更专业,不会等用户主动索要来源,而是把每一项实质性结论都视为需要1条或多条经审核学术数据库引用的主张。
更系统化的研究人员会定义纳入和排除标准,筛选完整文献,提取图表和图片中的细节,并按研究质量对结果加权。他们要的不是一个看似合理的概览,而是一份可复现的证据记录:纳入了什么、为什么纳入。
Elicit 目前已与全球前20大生命科学公司中的7家正式合作。早期团队用它探索作用机制、复现实验、检查毒理学,或对数千个基因和靶点进行竞赛式排序;例如,它们会询问免疫学实验如何压制「失控的T细胞」。
在已有经过验证的Phase 2和Phase 3试验结果后,商业和医疗团队面对的是另一种证据负担:应该面向哪些人群上市、谁会支付、有哪些替代方案,以及药物是否真正更有效或更具成本效益。每一项主张都可能需要在监管机构或支付方前进行辩护。
4. 证据质量取决于决策,而不是期刊等级
Labenz 质疑了「没有证据」这一说法在临床语境中的狭义用法:某些信息仅仅因为不是最高等级、经过同行评审的随机对照试验,就可能被直接否定。他的问题是,研究在什么情况下应该纳入病例报告、公司信息、博客,甚至一篇洞察异常深刻的帖子。
Stuhlmüller 同意已发表文献本身也存在质量梯度,但认为引用次数、影响因子、名校背景以及「我认识那个人」都是人类使用的有损代理指标。他提到,一篇奠基性的 CRISPR 论文发表在二线甚至更低等级的期刊上,这说明发表平台不能替代对研究本身的审查。
因此,Elicit 允许研究人员按具体情境定义“质量”的含义:在某一领域,病例研究可能恰当;而在随机对照试验充足的领域,病例研究可能没有必要。罕见病研究中,样本量10可能已经有意义;另一个项目则可能要求1,000甚至更多样本。方法学和实质适配度,优先于单独查看元数据。
研究代理也在加入公司申报文件等网络来源,因为论文发表速度缓慢,而科学决策往往同时具有商业和政策维度。Elicit 正在思考如何评估单条主张的置信度:区分一项「约99%可能为真」的命题和一项建立在冲突证据上的命题,再告诉用户如何负责任地使用两者。
5. 模型更会用语言表达不确定性,但信念仍不稳定
Stuhlmüller 对当前状况的直白判断是,token 概率如今「完全没用」,实际可行的校准机制只剩下用语言表达概率。与 GPT-3 到 GPT-4 时代依赖 token 概率的方法相比,他更愿意采用模型针对复杂情境明确说出的置信度。
但这些概率很容易被推动。询问一项临床试验的失败概率,模型可能回答30%;提到总体失败基础率后,它会提高估计;再提到针对该分子的研究薄弱,它又会表示同意。专家可能会回答:「我已经考虑过了。这并没有真正改变我在这里的判断。」
这种脆弱性很难精确刻画:Stuhlmüller 无法断言语义等价的措辞,或某一类固定的干扰信息,始终会触发同一种反应。不可预测本身就是问题——往往没有一个连贯的世界模型,既能稳定模型表达的概率,又能允许它根据真正新颖的证据进行合理更新。
Byun 预计,未来几年生成将「大致成为一个已经解决的问题」。届时人类工作将从填充空白文档,转向评估模型生成的初稿:说明什么才算好结果,整理失败模式,建立高质量示例,编纂实践,并构建足够稳健、能够引导模型的验证器。
6. 证书和工具轨迹可能比隐藏思维链更持久
Byun 回忆了 Ought 的 Interactive Composition Explorer,即 ICE。团队大概在2021年前后构建它,是因为预见到模型轨迹会变得过于庞大,人类无法调试。ICE 的经验是,按时间顺序逐步检查每一步,或重复一遍生成过程,既低效又无法扩展;独立的敏感性分析和逻辑一致性检查反而能揭示更多问题。
Stuhlmüller 把监督归结为两种选择:检查过程,或检查结果。结果本身也可以带有一份“证书”,说明结论如何随输入变化、检查过哪些内容,以及每条结论由哪些文献支持;这样无需检查完整的生成过程,也能验证推理是否可靠。
数学领域拥有可独立核验的形式化证明;模糊的研究领域大多没有对应物。Stuhlmüller 认为这一能力仍不成熟,部分原因是过去制作这类证书对人类而言过于耗时,而且人类也无法清晰地内省自己的思考过程。
他还区分了隐藏思维链和可观察过程。如果代理总结了一篇刚下载的论文,却从未调用阅读工具查看支撑其结论的方法学部分,那么这一遗漏就是一个可检查的推理事实。在更大规模上,工具调用能够显示哪些来源、章节和操作实际影响了答案。
7. 关键属性变得可检查后,模糊判断也会改善
分解任务的动机在于,带有可验证奖励的强化学习已经能在编程、数学和其他容易检查的任务上取得强劲表现。模型在模糊工作上的能力仍弱得多:即便可以访问他的邮件、Slack 和公司上下文,Stuhlmüller 仍觉得模型在公司战略方面「出奇地没用」。「它们就是不懂。」
训练阶段使用不完美的奖励很危险,因为模型会针对奖励信号进行强力优化;例如发现公司战略与早先的主张冲突,这类抽查不足以构成可靠训练信号。但在评估一个已经训练好的模型时,局部、分步骤的检查仍能识别明显错误,并指导改进。
把公司战略完整化约为形式上可验证的组件,看起来「要粗糙得多」。这并不意味着不可能,而是意味着能够证明拆解过程正在正确推进的增量反馈更少。Elicit 近期更现实的目标,是测试主张在内部是否一致,以及不同拆解路径是否收敛到相同结论。
微调仍然存在,但主要目的是让系统在规模化运行时达到合理效率,而不是解锁仅靠提示和脚手架无法实现的行为。
8. 外部世界模型把检索变成可读的持续学习
Stuhlmüller 围绕一位朋友的癌症问题,多次运行系统综述流程,最后仍得到约5,000篇高度相关的论文。在他看来,把所有内容塞进100万token的上下文,并不会产生连贯推理;检索成功反而暴露出更大的综合问题。
一个起点是“LLM wiki”:由 Markdown 文件或类似 Obsidian 的资料库组成,代理持续对其重组、更新和建立链接。但决策支持需要的不只是笔记,还必须回答预测、干预和反事实问题,例如某种药物、化疗顺序或后续免疫疗法可能带来什么变化。
这些问题类似图模型或结构化概率模型,但 Stuhlmüller 不愿预先规定一种通用形式。癌症机制可能天然适合表示为因果序列:抗体与抗原结合,随后某种物质在细胞内释放;公司规划则可能更适合用记录用户、功能、利润率和时间的电子表格表示。
多种视角必须共存:Elicit 可以同时维护随时间变化的用户数量电子表格、产品技术树、文本笔记,以及接入运营数据的 SQL 表,但更新必须在不同表示之间传递。Stuhlmüller 描述的目标,是把持续学习置于模型权重之外,「以一种人类可以检查和理解的表示形式」呈现出来。
9. 软件更便宜,不会消除规划或稀缺选择
Labenz 提出,更低的编码成本可能让规划不再主要依赖预测,而是把所有合理功能都做出来、上线,再让现实给出反馈。Stuhlmüller 认为需要克制:工程只是一个瓶颈,用户注意力和反馈仍然有限,糟糕体验会留下长期印象,快速迭代还可能把公司困在局部最优点。
探索与利用之间的权衡依然存在。成本下降对边际 bug 修复、行政功能和非核心实验的影响最大;而与 Elicit 使命和差异化直接相关的工作,仍需要审慎判断,因为有些项目依旧庞大,有些选择必须一次成功。
Stuhlmüller 的自研与采购规则十分明确:公司应该自建能够构成比较优势的东西,「基本上别的都不要做」。系统综述这类标准化、受监管的工作流,虽然复杂且高度依赖界面,但在制药行业普遍存在;早期研究中的专有预测模型则可以合理地留在公司内部开发。
更多临床试验与更好的基于世界模型的筛选并不矛盾,而是互补关系。数字孪生可能改善试验设计,罕见病可能值得更高灵活性,但毒性研究可能需要更高门槛。Byun 补充了从受试者角度看更具约束力的稀缺性:患者通常只能参加1项试验,「最多2项」,因此规划不可能消失。
10. 结构化推理已经重塑招聘和每周工作
Byun 曾用 Claude 协助招聘一名高管,因为在招聘过程中,岗位和理想人选画像都发生了变化。她收集面试、推荐信、背调、邮件往来和反馈,再要求模型围绕约20个维度寻找证据,包括目标达成、团队建设、真实性和文化匹配度。
她刻意把证据与判断分开:先为每个维度填入具体例证,再按5分制打分并综合形成决策。把结果分享给候选人后,对方称其为自己收到过最全面的职业能力验证综述;这一流程也帮助 Byun 抵御近因偏差。
Stuhlmüller 更日常的世界模型,是把年度、月度和每周目标连接到日历约束上。如果一个月度目标要求写一篇5小时的博客,系统会从截止时间反向推导这5小时可以放在哪里,以及前置任务如何向前传导——这些约束满足工作,人类通常只会非正式地完成。
他仍然讨厌全自动版本,更喜欢一个会带着他逐项做选择的交互式规划器:「你就是还没有完全理解我想做什么。」人类仍是调用 LLM 的外环,尽管他也能想象一种反转:由 LLM 调用人类——「不确定这是不是一个更好的未来。」
11. “The Line” 每周已合并30至50个问题
Elicit 的自动化软件工程系统名为“The Line”,因为它的运行方式像一条工厂流水线。Slack 中的请求可以加上 Line 表情,客户支持也可以转发问题,由此直接启动工作流,不需要有人手动创建并跟进传统工单。
系统会先定义请求、迭代规格说明、实现代码,并录制测试功能的视频。随后进入自动审核,再依次推进到开发环境和生产环境;规格不完整,或功能复杂到无法审核时,就会触发人工介入。
简单变更——例如调整 Elicit 讨论引用的方式——可以全程自动通过这条流水线。Stuhlmüller 估计,系统目前每周已全自动合并约30至50个问题,立即为小型功能和修复提供杠杆,同时持续改进每个环节,为未来处理更困难的工作做准备。
年末目标不是建立一家完全自主运行的公司,而是让每个职能内部都有能够自动运行的工作流,并在部分边界之间打通,同时由人类保留高层级 steering。障碍是可靠性:如果系统判断哪些任务可以安全审核的准确率只有80%,剩下20%就可能导致生产环境故障。「如果有疑问」,就必须升级给人处理。
12. 经济性偏向路由式智能,而非单一最大模型
在客户层面,Byun 认为 Elicit 的产品正在替代服务支出;即便软件定价锚点不够理想,相对于当前算力成本,仍有很大定价空间。制药公司在探索阶段结束后仍会要求可量化的 ROI,但真正应比较的,是解决整个研究问题所需的总投入。
Stuhlmüller 个人每周在 API tokens 上花费约2,000美元,可能至少是 Elicit 使用量排名前5的用户之一。他可能把支出翻倍或变成3倍,但「不会多太多」;目前大多数工作使用 fast mode 的边际回报不足,因此他会主动避开。
他的技术栈由编排器把更简单的工作分配给小型代理,同时让 ChatGPT、Claude 和 Gemini 对选定输出进行交叉核验。后台自动化会对齐日历、日志、任务、长期计划和邮件;仅交叉核验一项,就会让大约四分之一使用场景的 token 成本翻倍或变成3倍。
模型仍在持续趋同,但它们之间“微小的锯齿状”差异,正是 Elicit 按任务编排模型、而不是直接提供模型选择器的原因。据报道,Claude Opus 4.5 在提取准确率上领先 Gemini 3 Pro,而 Gemini 在直接证据支持方面至少领先5个百分点。Elicit 按任务选择模型,目前已经开放 MCP 和 API,但实际模型组合变化很快。
13. 科学领域的 AI 将融合神经整合与离散接口
Stuhlmüller 不认同某一家企业会「赢得科学领域的 AI」。这一领域覆盖单细胞动力学、蛋白质模型、自动化实验、证据综合,以及受临床试验周期约束的多年制药规划;这些属于不同的推理层,足以容纳多个系统并存。
行业格局仍未定型。现有全球前20大制药公司可能完成自我改造,也可能在10年内被一体化的 AI 原生生物科技公司取代;Stuhlmüller 认为两种结果都有可能,部分取决于 incumbent 多快意识到变化的规模。
Labenz 将可检查的工具调用,与直接把科学模态整合进模型权重的系统进行了对比:前者可以把蛋白质交给专业模型,或运行一次云端实验;后者可能获得图像转换中已经出现的流动性,但也会让流程失败更难定位和追问。
Byun 原本更偏向整合和端到端优化,但他指出,让模型在“权重空间”或 neuralese 中思考的尝试,表现并没有此前这一判断预期的那么好。离散的词语和程序可以提供误差修正:小幅偏差能够重新收敛到稳定符号,减少长链条中的误差累积,就像数字计算相对模拟计算具有优势。
14. 追求真相必须成为明确的优化目标
当被问及人类推理将走向何方时,Stuhlmüller 讲起一个笑话:有人从屋顶摔下,坠落到一半时说「到目前为止,一切都好」。模型已经能改善许多答案,但这种局部收益几乎无法说明,在 AI 转型期间最重要的政府、实验室和机构决策会发生什么。
他的基准判断是,「AI 到目前为止基本什么都没有改变」:编程并不是大多数知识工作的全部,而且这个行业仍然雇佣着大量人员。最重大的决策还在前方,因此认知论层面的较量仍未结束,既不能证明乐观正确,也不能证明一切将崩溃。
如果模型的优化目标是看起来优秀并具有说服力,而不是追求真相,集体认知论可能进一步恶化。反过来,更好的推理也可能形成一个“吸引盆”:对真相的追求改善后,预测和认知基础设施会被识别为高价值干预方向,进而改善后续干预措施的选择。
最后的警告既针对机构,也针对个人。METR 研究发现,工程师认为 AI 让自己更快,但在当时,AI 辅助实际上带来了轻微折损;面对复杂决策,模型可能把用户的判断拉回平均水平,或让调查过早结束。Stuhlmüller 的建议是持续追问:「我真的从这里获得了好处吗?还是它正在如何改变我的行为?」
It has been 2 years, guys, since our last podcast. Time flies. We've had a couple of opportunities to talk offline, but 2 years is an eternity in the AI space. So, I'm really excited to catch up on everything you guys have built, what you've learned, the sense you're making of the rapidly changing AI landscape, and how you can put a little positive nudge into history.
Maybe, for starters, let's do a quick review of the motivating mission you guys started with around elevating reasoning and the quality of decision-making. Then, with that context, we can go into a brief history of the last couple of years, because I think you guys were on this trend even before there were such things as reasoning models. I think the reasoning models are not exactly what you had in mind, but obviously they're very relevant to where we are today. So, kick it off with a little reminder of the motivation behind Ought and Elicit.
Thank you, Nathan.
Good to be here.
Yeah, our mission is to radically improve the quality of reasoning, especially for high-stakes decisions. As you said, we've been at it for a while. We started a nonprofit, Ought. We're working at Elicit, an AI research company, now.
And, Jungwon, what a decade it's been over the last 2 years. Just a few days ago, I was going back over my PhD from a long time ago in machine learning. I thought, what would it look like to redo some of that work now? So I went to Elicit and said, “Hey, let me just upload one of my workshop papers where I know I didn't do a good job.” I redid some of the computational experiments in Elicit. I said, “Hey, do the data analysis.” It was running in the cloud, and 10 minutes later, it had redone the paper. It's just mind-blowing.
If you think back to when we talked 2 years ago, I think we had launched paper search and summarization in Elicit. Over the course of maybe the next year or so, we launched systematic literature review—the fixed search, summarize, screen, extract, write flow—and now we have these much more flexible research agents. It's obviously been an incredible trend from very short tasks, short time horizons, and low flexibility to agents that do much more extensive research, with long time horizons and quite a lot of flexibility. So, I'm excited to review what this all means and where it's all going.
Can you give me a double-click into what exactly you had in mind when you were talking about how you were going to improve the quality of reasoning, how that compares to more and more RL on top of a chain-of-thought paradigm that everybody is now very familiar with, and then how that has maybe enabled what you're doing with Elicit, maybe in some ways competed with it, and how it has shaped your business and product strategy to go in this more principled, a little less bitter-pill direction with what you're trying to build?
I like “bitter pill.” Yeah, last time we talked, we talked a lot about process supervision, right? How do you know that a system is doing good research for you? You can either look at the output and say, “That looks good to me,” or you can look at the process it went through: What papers did it look at? Why did it look at them? How did it choose what to do?
Back in the day, we were young and naive, and we all thought it would be better to know that things were correct for the right reasons, that the process was good. So, we leaned a lot into that. In many ways, I think the justification for that has still borne out.
Let me talk about a quick anecdotal experiment we ran, I think, 2 days ago. We told a few research agents, “Look through and analyze about 100 papers on toxicology risk for a particular type of cancer drug.” We told that to Claude, to ChatGPT, and to Elicit, and then we asked, “How many papers did you actually analyze?”
As the models like to do, they say, “That's a fair and important question to ask. Let me be direct. I did not analyze 100 papers. You're right to push back. I didn't do it.” I think that is a failure of process in many ways, because you're saying, “I told you what to do. You didn't do it.”
The reason for that is that the models are not trained on process. The models are trained to produce outputs that look good. If you didn't check and just looked at the result that said, “Here's my analysis,” you wouldn't have caught that.
I think the fundamental problem still exists. And then the question is, “What do we do about it? And to what extent does the solution look like better checking of the outputs versus checking of the process in various ways?” I actually think this question is still open. I can speak for what Elicit does and maybe briefly for what the rest of the ecosystem does.
Elicit addresses this by—we have a little kind of domain-specific language that the research agent can write that orchestrates other calls to agents: screen all these papers, then extract data from all these papers. And it runs the process. You run the process, and you know the process does what it said it was going to do.
I think the model companies are going a little bit in that direction, too. I haven't been following it extremely closely, but I think Anthropic recently launched a workflows feature that has similar properties. And so, even though on some level, yeah, all the models are outcome-trained, and that's why we see these quite ridiculous artifacts of models being like, “Ah, sorry, I didn't do it,” I think the problem still exists, and one level up from that, people are trying to patch it. I think that's maybe where I see ourselves as being in this game.
Yeah, I think it's worth emphasizing that a little bit more because it's a core part of how Elicit is built differently and, therefore, what you can use Elicit for.
A lot of—like you said, it's been 2 years since we chatted—a lot has happened in that time. A big part of what we've been working on for the last year is just rebuilding Elicit on top of this much more agentic infrastructure. We started working on that in about March or April of 2025. In retrospect, it was maybe a little bit early, but at the time it felt quite late because obviously a lot of advances had happened over the last 2 or 3 quarters.
A big part of what we spent the last year thinking about was designing how we preserve these benefits of transparency and systematicity at scale without losing the flexibility and raw power of these models. I think that's the core design question we've always wrestled with, because when you want to deploy these models at scale for really high-stakes decisions, you need them to behave in a certain way, which is often contrary to their kind of fuzzy nature. But you don't want to be overly deterministic because then you run into the bitter lesson issue, right? So threading the needle is what we're always struggling with.
A lot of our time last year was spent on this kind of technical design question, and we decided to design our own programming language to solve this problem so that the models could run reasoning computation reliably and be able to call these reasoning primitives at scale in a more trustworthy way. What we were trying to accomplish for our end users was the ability to say, “You can run this process with the model over 10,000 objects—10,000 documents, 10,000 drugs, 10,000 targets, genes, whatever—and the same process will be applied to number 5 as number 9,999.”
There's just no other model that can make that guarantee. And unfortunately, there are lots of models that claim that they have done that or can do that and are just completely wrong about it. So when we think about who our users are, how do we build a differentiated product that really meets their needs, and what use cases can we enable, we are powering people who want to be able to rigorously synthesize evidence at a very large scale and get every single thing right, all the way to the nth degree.
That's a very different interaction than just riffing with a model. We support some of these lightweight use cases, too, but I think there it's almost more like getting to parity. And I think where Elicit's really differentiated is that trust at scale.
Yeah, I should always remind myself and be clear that none of these things are really true binaries, in the sense that it would be wrong to characterize Elicit as not being about the bitter lesson. I remember that last time, one of your big principles was, “How can we allow you to spend more money to buy more compute to get better results?” And that's almost a restatement of the bitter lesson in some way.
At the same time, the frontier companies, as you said, are doing some of this, and I assume that in their research agents, particularly their deep research products, they are presumably doing some sort of at least rubric-based reward on the final reports that the models are outputting, albeit, if I understand correctly, still at least intending to avoid putting any optimization pressure on the chain of thought. There's always a little bit of gray to this.
So who are the customers now? There are these deep research things. I use those pretty frequently. Who is the sweet spot that is like, “A deep research agent isn't enough for me. I really want to go way bigger, way more systematic, and be very sure that I'm performing the same analysis in a way that I can count on”? Who are those customers now that you're finding product-market fit with?
Yeah, I think of it as maybe a funnel. The largest group of users are academics and individuals, people like yourselves who value the kind of deep research function. They generally want a fast but robust synthesis of the literature or of the evidence.
Where Elicit really differentiates itself is that it's a bit more technical by default. It's not really assuming a lay audience as much. Most importantly, everything is well cited. You can get the models to cite things if you really push them to, and then sometimes they realize they were referencing sources that didn't exist, which is unfortunate. So they still make those mistakes.
Elicit doesn't do that. Elicit just assumes every single claim needs at least one or multiple citations from vetted databases and data sources in the academic literature. A lot of people still find value in just that kind of core offering.
Then there are academics and researchers who, like you said, want to get a lot more systematic. They want to do a systematic literature review, where they're really thoughtful about: What evidence should I be looking at? What information out there is relevant to my research? What should I be excluding? How do I apply the same process to all the papers? They want to extract detailed data from charts and figures, and then synthesize all that and weight the findings by quality, or do a landscaping of the market or the research in a similar way.
Increasingly, one of the other things we've really been developing over the last year is our life sciences motion and playbook. We work now formally with, I think, 7 of the top 20 life sciences companies, and we work pretty much across the entire development life cycle. I would say our biggest concentration is in early-stage research, so working with discovery biologists, some of those toxicologists that Andreas mentioned, as well as researchers in the late and post-clinical stages and commercial and medical teams.
The early-stage researchers often have more iterative processes. They're like, “I have this idea for an experiment. I'm in immunology. I want to figure out a way to repress these kinds of rogue T cells. What's the right mechanism to do that?” Then they get ideas, find other experiments that have been run, and try to reproduce those experiments. They also are often systematic, so they're the ones who want to apply a reasoning process over thousands of genes and targets and do a tournament-style ranking.
Then on the commercial and medical side, often you've had validated Phase 2 and Phase 3 trials. They're starting to think about what's the launch strategy of a drug. Exactly which populations do we go after? Who pays for this drug? What are the other kinds of alternatives out there? How much more compelling or cost-effective is our product?
They have to be very evidence-based, and every claim that they make in justifying to regulators or payers exactly why this drug makes a difference in the world has to be supported. So that's where, within life sciences, we're seeing a lot of pull.
Something really top of mind for me that you kind of touched on there is being evidence-based, and I hadn't thought about the definition of that, especially at the early-stage literature-review stage, but especially as you get later, when you start to think about what markets and that sort of thing. It strikes me that perhaps in many cases the best information is not necessarily in an academic paper, or at least, if I was naively thinking about approaching such a problem, I would think I would want to cast a pretty wide net.
I also experienced this a little bit myself. Fortunately, my son, as regular listeners know, is pretty much cured of his cancer, and we're getting back to normal life. That's all great, and I didn't have to go down the path too deeply to really try to get a synthesis of what the best second-line treatment would be for him.
When I was contemplating that, in talking to doctors about it, I found there was a real narrow box drawn around what kind of information would be considered. Our oncologists, whom I hold in quite high regard and with whom I have a good relationship, will say, “There's no evidence for that.” This is for things where I'm like, “There's definitely some evidence, right? It's just not this sort of gold-standard RCT, peer-reviewed, yada yada yada.”
Obviously, that stuff is great where you can get it, but all this is a buildup to ask—and it may be different for different kinds of analyses—how do you think about how wide the net should be as you go out and look for evidence? When is it appropriate to truly limit yourself to published academic work? When is it appropriate to consider blog posts or even insightful tweets that somebody has put out? How do you guys think about that? Does it vary by use case? How do you think about weighing these things and calibrating to the level of credibility that they should have?
That seems to me an opportunity for, honestly, just major improvement, even at the clinical level, because there seems to be so much turning a blind eye to evidence today just because it's not right at the very highest level of quality or credibility.
Yeah, that's right. I think the core question here is: How do you differentiate and discriminate between evidence based on evidence quality? And how do you identify when a certain level of evidence or quality is relevant or the best thing for your decision?
So even within published literature, there's a gradation, right? There are—
Impact factor, remember?
Yes. Yeah, there are higher-impact journals, there are well-regarded authors and institutions. This is a question that we've wrestled with for a long time, because the way humans have solved it is that there are these very lossy proxies: citation counts, journal impact factor, “I know the guy,” right? Those are rough approximations, but deeply imperfect, and there are so many examples of groundbreaking research that was not published in the top journal. One of the foundational CRISPR papers, for example, was actually published in a tier 2 or lower-tier journal.
I think the promise of language models is that, yes, we can continue to use these heuristics that have been helpful, but also we can think a little bit more from first principles about the quality of the work and its appropriateness for the decision at hand. So that's what Elicit has always really tried to do: not just blindly rely on citations, but give the researcher the chance to say, “Okay, for this research project, I'm going to look at case studies, or I'm not. Case studies are not good enough in my research domain because there are plenty of RCTs.” Or, “I'm going to look at studies with a sample size of 10, or I'm just going to look for a sample size of 1,000 or more,” because, depending on whether you're working in rare disease or oncology, the shape of the research that's available, like you said, is different, right?
I think one principle we have is: How do we look at the actual methodology and the quality and content of the research, not just at metadata, to determine quality? How do we equip the researcher to express their own judgment and expertise based on what they know about the domain and what is specific to their project? I do think it is specific to the domain.
As we've built out the research agent, we have started to bring in a lot of other data sources—web sources like company filings and things like that—because, like you said, there's a lot of information that's not just in published papers. Publication takes a long time, so a lot of recent things are not going to be there. Sometimes research is very interdisciplinary and holistic, so when you make a scientific decision, you want to consider the science, but you also want to consider the commercial implications of that as well, or the policy implications.
We have introduced a lot of data sources, and I think the agent still has a relatively fuzzy understanding of which sources are higher quality. I think we want to structure that a bit more. Again, the user has a lot of control over how they can express what they believe quality means, but I think we do want to also be a bit more opinionated for researchers who are not experts.
The last thing we're thinking about is: How do we express the confidence level of claims? Not just at a source level—say, “This is a high-quality paper, this is not; this is a high-quality blog post, this is not”—but when Elicit synthesizes all this information for you, how can we say, “This treatment works”? That claim: How well supported is it? Is it something you don't need to worry about, like 99% likely to be true, or are there conflicting results in the evidence?
We've been thinking a lot more about breaking insights down into claims and then giving people confidence levels so that they know what to do with that. I think that's another important part. It's not binary; it's: How can I use it? That's the part that's really important.
How good are models at that today? I always think back to the original GPT-4 model card, where they showed the base-model calibration being pretty good: when asked to express confidence in claims, when it said 20% likely, it was actually pretty close to being correct 20% of the time. Honestly, to a pretty remarkable degree, the base model was reasonably well calibrated. I understand that pretty much just came out of pretraining as an emergent property.
But then, with reinforcement learning, you have mode collapse—which is, I'm going to overstate the importance of mode collapse—which I do think was one wave of AI denialism for a minute there. But there was some form of mode collapse where all of a sudden you're getting much less well-calibrated self-assessments from the model in terms of how confident it was that the claims were accurate.
Obviously, since then we've had a tremendous amount of additional RL applied on top of models. I haven't seen those kinds of calibration studies as recently. I could believe maybe they've gotten better because maybe that's part of a rubric that's being used. I could imagine it's even gotten worse because, just left to its own devices, RL maybe takes you in the other direction.
What do you guys see in terms of how good they are natively? Can you prompt your way to good performance on these sorts of confidence questions? Do you have to fine-tune? If they're not that great, do you want the frontier models to get better? Would you advise them on how they should think about getting better at that? Tell me everything.
Yeah. So, first, I think token probabilities are just useless now, so we can pretty much ignore them. When we talk about calibration, mostly we have to talk about verbalized calibration, right? If you ask the model, “How confident are you in this thing? Are you confident?” I think verbalized calibration is probably much better now than the token-probability approach from GPT-3 to GPT-4. I would rather take the verbalized probabilities, to be honest. I think those will probably do a better job at capturing, for complex situations, what the assigned probability should be.
I still think the models right now are easy to push around, and that is one of the main things we're trying to address with better scaffolding. Right now, you often ask the model, “Hey, how likely is it that this clinical trial will fail?” Say the model is going to say something like 30%. Then you're like, “But have you considered that you can almost say anything here? Have you considered that, on average, clinical trials fail X% of the time?” And it'll be like, “Oh, no, you're right. Actually, this probably should have been higher.” Or, “Have you considered that, in this particular case, this molecule has pretty little research behind it?” And it'll be like, “Oh, yeah, you're right.”
The models are getting better at it, but I think fundamentally, right now, these probabilities are very unstable compared to when you talk to an expert. You can throw these things at them, and most of them will be like, “Yeah, you know, I've considered it.”
[Laughter.]
“That doesn't really change my view here.”
So, a big question is: How do you get stable probabilities out of the models? Which isn't to say they should never update. They should obviously update if you present truly novel information to them; they should update.
But right now, it does not feel like they have a coherent world model behind a lot of the probabilities that they express to you.
How would you say they're easy to push around? How easy are they exactly? Are we talking about even seemingly semantically equivalent rephrasings producing big differences in outputs, or is a more conceptual difference in framing required? Or can you throw in a red-herring fact and move the needle? Just how much are they blowing in the wind?
It's hard to characterize, and this is maybe part of the issue: it's hard to predict exactly what will push them around. I think if I could give a better answer to this question, I would also have an easier time fixing it. I don't think it's extremely systematic right now.
But maybe that's partially what we've been trying to invest a lot more in. I think that's the next phase of our role as humans, and what humans have to do is work through these really gnarly evaluation problems. I very much believe that in the next few years, generation will more or less be a solved problem. Historically, humans have done all the generation. They've done all the work.
You open up a blank document, you put your thoughts in, you start from nothing, you create. And that's just not how work is going to happen anymore. Increasingly over time, AI is going to take the first pass at everything. Then the work left to humans will be around evaluating that. Was it the right thing to do? Was it done well? Can I use it for this use case? Can I trust it?
I think people are worried about job loss, but I do think there's a major job transition opportunity, or skill transition opportunity, where now what we need to do is think about what good looks like and what these failure modes are that AI systems can run into. How do we start documenting them? How do we start codifying best practices and getting good examples to point models at? How do we start building better verifiers?
There's actually a lot that needs to happen because we don't even know—we can't even articulate for ourselves what good looks like, why this is a bad thing, or how often this happens. So I feel like there's a lot of work around that and, certainly, we are trying to make that transition as a company. I think there needs to be more infrastructure built around evaluation and articulating what good looks like so that we can point these models toward it. How do you think…
it interacts with chain of thought? This is very top of mind. I keep talking about it, but I attended this event in San Francisco called Recursive not too long ago, which was all about the potentially soon-coming phenomenon of recursive self-improvement. A scary, striking takeaway from that event and all the conversations and presentations there was that we're really heavily, and I would say problematically, indexed to chain-of-thought monitoring.
It's chain-of-thought monitoring all the way down in terms of the plan for how we're going to keep a recursive self-improvement process on the rails. There's a very strong sense that we've got to maintain freedom for the models in the chain of thought so that we can monitor it, because if we apply pressure, per the “Obfuscated Reward Hacking” paper, we'll drive the bad behavior underground. Then we'll be doubly worse off, because we'll still potentially get the bad behavior and we won't see that it's thinking about it.
How do we square that? I guess the naive answer would be that you can think whatever you want in the chain of thought, but I still want a systematic account—a systematic reasoning trace—in the actual final output, and then I just reward that. Maybe that works, but maybe you have a more nuanced view on how companies should think about where, what kind, and how much pressure to apply to the reasoning process that builds up to the final answers that models give.
I can give a high-level take, and then I'm sure Andreas will have a more technical response. One of the projects we worked on at Ought, like many Ought projects, was a bit early for its time. We probably built the first language-model observability tool. We built this thing called ICE. It stood for Interactive Composition Explorer. We had this—when was this? This must have been 2021 or so.
We anticipated this problem even 5 years ago: that at some point, language-model traces would get so large that we would not be able to debug them for ourselves. How would we visualize that? How would we maintain oversight? How might we audit that? And so we built this visualization.
One of the things we learned was that the best way to troubleshoot is not necessarily to go chronologically through all the steps that the model took. That actually means evaluation: repeating the generation process is one way to check what happened and build trust in it, but it's inefficient. It doesn't really scale, and it's a very difficult way to check.
What you often need is a different layer of reasoning checks, almost like logical consistency checks. These are not, “Let me go through everything you did—steps 1, 2, 3, 4, 5—and see if it was correct.” Instead, it's, “Let me think about, for example, a sensitivity analysis: How sensitive are my findings to different changes in input parameters?” Logical consistency checks, things like that.
That's where I think we need a lot more investment in infrastructure and building independent checks that don't rely just on chain-of-thought monitoring. That's my high-level take, but I'm sure you have a more technical version of that answer.
Yeah, let me first restate part of what you said in different language. I think you can either check the process or you can check the outcome. I mean, those are your 2 options, right? When you're checking the outcome, you still want the outcome to somehow contain a certificate that the right reasoning was done.
What can that certificate be? It can be, “Here's how my conclusion would change if I had a different input. Here are the things I looked at on the way. Here are citations to the literature.” I do think that this is a very underdeveloped field, in my mind. In mathematics, it's very developed, right? You can have your formal proof if you want, and that proof is checkable.
I don't think it's very developed in more fuzzy domains, probably partially because it would just be too much work for humans to produce legible certificates. Humans don't even really have a great ability to introspect on their own thoughts. But I think in principle, even if you didn't supervise the process, you could produce certificates of reasoning that then let you check the reasoning, even if you didn't check the process that generated that outcome. I'd be very excited to see more work in that direction.
The other clarification I wanted to make is that I think it's worth distinguishing chain of thought from the reasoning process, or the chain of process, or whatever you want to call it. When people talk about chain of thought, they often think about what thoughts the model writes down—the reasoning tokens—and then ask, how much can you trust them?
I don't know. I think for OpenAI, we don't even get them anymore these days, unless maybe you apply for a special permit to see them. But you do see the tool calls, and I think the tool calls actually are important reasoning facts in and of themselves.
Take, for example, let's say I download a new paper from arXiv. The model hasn't seen it before, and I ask it some question: “Hey, what are the key results here?” Now I can see which parts of the paper the model is reading, because it has a read tool that maybe reads, by default, the first few lines, and then it can scan other parts of the paper.
Sometimes I know the model didn't even look at the methodology section. That is part of its reasoning that is checkable. I know if it now tells me something about the conclusions of the paper that really should have relied on the methodology section, I know it didn't do that. That's obviously a very simple example—the papers are small and so on—but the same thing applies at a much larger scale, where the process still remains an important way to check, because you do see the tool calls, and the tool calls are an important input into the model's reasoning.
Yeah, I like that. How much of this do you think companies are doing today? It's got to be some, but obviously they're not telling us. Is your view on the spectrum from closing in on the sort of thinking that you're doing, and the granularity of process supervision that you would like to see, on the one end, to purely RLVR—with “Did you get the right answer in the final answer box or not?” as just a binary signal—on the other end? How much of this do you think they are doing, based on what you're seeing in model behavior?
These are obviously speculations. I think they do a lot. I think we know they do a lot of rubrics on the final output, like, was the—even reasoning-adjacent rubrics: Did the model produce something that looked like expert reasoning in many ways? I don't think they do a lot of evaluation of the process.
Again, I could be wrong, but I think there's a really interesting question, which is: If you had all the details of the process—let's say you are inside a lab and you ask yourself, given X and T, would you expect this process to lead to the right answer? Not after the fact, if you check it against the answer—was it correct? But did it follow the sort of process that, if you're making a forecast, for example, and you don't see—let's say you don't even see what forecast the model ends up…
You just look at what things it considered, what data sources it considered, and what it wrote about the different hypotheses it considers. I think there's a really interesting project of thinking about to what extent the models are following a process you should expect to be good. In my limited knowledge of what's going on, I don't think much work is being invested in that.
So, I guess one implication of that for your work would be that you might expect the next generations of the model to eat less of your scaffolding than they might eat of other types of scaffolding, right? This kind of general pattern—people build out scaffolding to compensate for the model's weaknesses, then the model companies take that feedback and train on it, and in the next generation, you can clean out or eliminate a decent amount of that scaffolding. It seems like you think that maybe on the agentic side of “go accomplish this project, get over these humps,” whatever, you would expect more of that kind of scaffolding to be eaten in the next version than the sort of trusted-reasoning scaffolding that you're building.
Yeah, that reminds me of something that you recently wrote around Elicit's focus being on becoming the best at reducing these big, fuzzy, hard-to-verify tasks to sets or graphs, perhaps, of easy-to-verify tasks. I'd love to hear a little bit more—you alluded to it there—about how you're going about that and also how complete you think that process can be.
Everywhere I look these days, I feel like I see the same shape of a really interesting question that I don't know what to make of, which is basically: How do we get high-level guarantees, conclusions, or insights from low-level steps? In biology, I might be able to say I've got all these proteins, or these genes are being expressed at this level, but I don't know what's going to happen next at the cell, tissue, or organism level. No, right? I don't.
Similarly, with these big judgment calls—should I prioritize this drug or that drug?—we can break it down and become systematic, but it's not clear to me how close it gets to something where I'm like, “Yes, okay, I can really buy in and trust that.” Versus, is there something that sits above all those steps still? Is it emergent, or is it just somehow lost? Are people doing some sort of metacognitive work that's hard to capture but maybe still very critical to actually being effective at these tasks?
So, I guess, again, tell me everything. I'm really struggling with this. Formal methods and formal verification are another area where I see this: We can make all these low-level guarantees that this error or that error can't happen, but is the system itself going to behave how we wanted it to? I still don't know, in a lot of cases, how we make that leap.
I'm very interested in your take on these sorts of questions. I usually think of it as laddering up low-level things to high-level conclusions, but you're approaching it from the other direction, which is interesting unto itself. Yeah, again, tell me everything.
Yeah, maybe I should first restate this: Why do we want to reduce hard-to-verify tasks to easy-to-verify tasks? It's because AI currently can be trained on easy-to-verify tasks. We know it's extremely good at RLVR coding and at math tasks like this, and it's quite weak at a lot of fuzzy tasks.
I notice it all the time when I try to use the models to help me plan our company strategy, for example. I think they're surprisingly useless. Even though they have access to all my context, they're really quite good at saying, “Let me pull in the data. Let me pull in your email and your Slack.” I still find that they don't get it.
In an important way, this is related to what we said earlier about how they're too easy to push around. It doesn't feel like they're building up a coherent model of what's going on. I think an important reason for that is that it's a hard-to-check task.
So, what do you do? I think it depends a little bit on what your situation is, whether you're trying to create a reward signal that you can train the models on, or whether you're trying to do verification and checking for the purpose of understanding whether an already-trained model can be trusted in a situation or how to refine its behavior.
I think if you're trying to create a reward signal, that's pretty rough because the models are going to optimize pretty hard against your signal. It's not enough to do spot checks and be like, “Hey, here are some cases where we can verify that, for example, your company strategy was incompatible with some claim you made earlier.”
Whereas, if the goal is to take an already-trained model, understand how good it is exactly, find fairly obvious mistakes, and identify places where it can improve, then your reward signal doesn't need to be bulletproof. Your way of getting some easy-to-verify aspects of the hard-to-verify task can be more stepwise, so you can make more incremental progress, I would think.
Our situation is that we're not currently trying to train a foundation model from scratch or even post-train a model on this particular aspect. Our situation is more like, how do we get to the point where we can check many important properties of tasks? Are the claims the model makes internally consistent? If you break it down in different ways, will it end up at the same conclusions, and so on?
I think that's a fairly tractable project. The project of figuring out how to fully reduce high-level tasks like company strategy into individual components that are all formally verifiable is a much rougher prospect. That's not to say it's impossible, but there's less incremental feedback signal that you're on the right track there, I would say.
So, tell me a little bit more about what you're doing in practice. You said you're not trying to post-train a model. I know in the past there was a decent amount of fine-tuning, at least for specific tasks, so I'm curious if there's still a fine-tuning element to it.
There are a lot of different ways you can think about spending a lot of tokens to try to get at this. You could run the decomposition process multiple times and check for consistency, which I think you're suggesting something like that might be going on. You could do a more iterative thing where you get the AI to give you an output and then have some kind of specialist prompts, or perhaps even specialist fine-tuned models, come in and assess it in various ways, give it feedback, and then let it reason some more and try to improve on what it just did.
We do see some of that stuff. I just talked to some OpenAI forward-deployed engineers who are basically using that process to improve filing accuracy, and that seems to be going quite well for them. What techniques are you finding to be most effective in practice today?
So, first, we still do a bit of fine-tuning. At this point, I think it's more of a technique to make things have reasonable efficiency properties at scale than something to get the models to exhibit new behaviors that you couldn't otherwise elicit.
That said, we do a lot of the things that you pointed out. Let me maybe talk about one of them that we've been investing more in lately, which is what you could call world models, or knowledge representations, that make the model's work more checkable.
What's the motivation? The motivation is actually maybe similar to the kind of medical case you had. I had a friend who also had cancer, and it was a case where I then used Elicit to get a lot of the raw data for it.
I ran the systematic literature review flow for a few versions of the question, “How do you address this particular type of cancer?” I ended up, even after filtering down all the information, with a ton of papers. After filtering for just the most relevant papers, it was still maybe 5,000 papers or so.
The question is, what do you do with that? You could try to somehow throw it all into a million-context window, but I don't think it would actually work that well, and I don't think the model would be that good at coherently reasoning about it.
And so then the question is, what else can you do? I think one thing that people have tried—I don't know if you're familiar with this—is what Karpathy termed an “LLM Wiki,” I think, or something like that. You build this knowledge base as an Obsidian-like knowledge base, or a folder of Markdown files, where you tell your model, “Hey, we're researching this topic. Organize the information in a way that makes sense and do these iterative updates to it.” Maybe it's a GitHub repo, and the model gets to add new notes to it, move information from one file to another, and propagate information.
I think that is a really interesting direction because, if you think about how we get to models that currently answer complex questions, either it happens in the weights of the model or it happens in some explicit representation. What is the explicit representation? Text files are appealing; it's a nice start. But then you ask yourself: That is also very flexible. What properties do you want this representation to have such that it actually helps with the research questions you have?
Once you think about that, you're like, “Well, I want it to let me make predictions about what's going to happen in my case. I want it to help me think about interventions. If I did this—if I took this particular drug or pursued this particular type of chemotherapy over different time periods—what would happen? If I did this instead, or followed it with some immunotherapy, what would happen? If I had done this different thing in the past, what would have happened?”
Those are questions about predictions, counterfactuals, and interventions. Very soon, you're like, “Wait a minute. That is a thing people have been studying in the past. It sounds a lot like graphical models or structural probabilistic models.” The question that we've been thinking about internally is: How do you get the best of both worlds? You want the flexibility of language models that can reason about and transform these representations, but you also want to be able to answer these sorts of prediction, intervention, and counterfactual questions that you care about.
You want the answers to these questions to be internally consistent, such that it's not the case that one answer and another answer just don't make sense relative to some underlying coherent representation. All this is to say, we're thinking about how you can build what I'm currently calling a world model, although I'm not sure that's the best name for it. How can you build these representations so that they let you answer the classes of questions I just talked about in a way that is internally coherent, and evolve them over time as you add papers one by one—adding information about different symptoms, different treatment strategies, and so on—and get them to be internally coherent?
That's one big direction we're thinking about for how to spend really enormous amounts of compute to deal with large amounts of data and turn them into something that can actually answer real questions for projects that are much larger than a single query.
Yeah, that's cool. I am very familiar with the Karpathy Wiki line of thinking, and I have used it myself for creating a little wiki that my personal agents use to navigate my life, basically, on top of all the raw data. I first did monthly summaries, where I exported everything and, on a month-by-month basis—it was a few hundred thousand tokens a month—I would condense that to a monthly summary. Then there was a layer condensed to an annual summary, and then I came across that. I was like, “Okay, let's make a wiki.”
We've got articles about all these fun things that I'm starting to do, but it's a party trick, I guess. One I just did last night was tell a friend I hadn't seen in a minute, “Oh, there's an article about you in my personal wiki. I'll have my agent send it to you, and you can read it. I haven't read it, actually, so I don't know exactly what it says. But you can read it and tell me if you see any hallucinations or anything that you would object to in there.”
For me, though, it's basically just trying to get down pretty black-and-white facts and help the model navigate those. The nodes are pretty obvious, and the edges are also just simple links: This person I know from this organization, and so on. I haven't pushed it nearly as far into ideas or a structured decision-making aid, so I'm curious about how that looks.
I guess the other image that's coming to mind as I'm trying to conceive of this a little bit better is the Anthropic work on tracing large language model thoughts, where they have these sorts of graphs from tokens to outputs. I'm always one to remember that the graphs they present are somewhat clean, but there's a huge residual term all over the place on those graphs, which again recalls my earlier question: How much can you decompose, and how much residual is there?
But tell me a little bit more about what these world models end up looking like in terms of edges. Is “X causes Y with X percent probabilities attached to it” one good way to start? Or how do you think about describing those relationships in a way that balances capturing as much of the structure as you can, while also recognizing that there is going to be some residual? I'm really interested in how you're navigating that.
Yeah. We are trying not to be very prescriptive about it, to the extent that we can. I think there are cases where nodes and arrows are the right—or at least a helpful—representation. In the cancer example, if you're trying to understand the mechanism, I think it's often useful to say, “Well, the antibody and antigen need to bind, and then once they bind in the cell, a particular type of substance needs to be released in a particular way.” There is just a sequence of events that needs to happen, and in those cases we want the model to build these sorts of graph-like representations.
But I don't think that's always the case. If you think about company planning—what products should Elicit ship over the next few quarters, and how does that affect our revenue and user numbers and so on—I don't actually think I want that to be represented as nodes and arrows. I think it's more like having a spreadsheet with different features, and then user numbers, margins, and so on over time. That's a very different type of model.
Often in the real world, no single model captures what's going on in its entirety. One lens to look at what Elicit is doing is as a spreadsheet of user numbers over time. But another lens might be the tech tree of the Elicit product and how we build it up over time. Those are complementary. Ideally, they can both live in your knowledge wiki, and you can say, “Hey, language model, as you're trying to make predictions or help us evaluate plans, look at all of these representations.”
The challenge is how you get it to be the case that the model knows how these different representations relate. When it makes a prediction, you don't want it to be the case that sometimes it looks at the spreadsheet and says, “Well, it's going to look like this,” and another time it looks at your tech tree and says, “Oh, okay, we're going to do this.” You don't want them to be 2 totally separate things. Some propagation of information needs to happen between those different representations.
It's an ongoing research project. I definitely don't want to claim that we have solved this problem. It's an ongoing research project within Elicit and, hopefully, eventually within the world at large: How do you make these more explicit, legible, fairly heterogeneous representations of knowledge that models can work on over time and improve?
I guess one last thought here is that maybe one way to think about it is: How do you make progress on continual learning in a way where stuff doesn't just live in the weights of the language model, but is available to humans as a representation we can inspect and understand?
Have you found any particular data structures, particularly if they're open source and something I could also incorporate into my own personal AI infrastructure, that work well? I'm just doing a very simple Markdown wiki as of now. Is there a next level that I should be considering? A graph database? I have no idea what it would be, but I'm always looking to upgrade.
Step 1 is just to use the representations people already find useful, and SQL databases are a pretty useful thing. I'm currently building my model of Elicit, the company, and Step 1 is to ingest a lot of the information from Mixpanel, Adyen, and various other types of systems and put it into a representation that the model can then operate on. A lot of those representations are just SQL tables.
Cool. Interesting. Everybody's doing their own experiments in recursive self-improvement these days. That's what I'm noticing across the board. It's not always the case, but it's been striking that in the last 6 to 8 weeks, it seems like everybody's tipping into this moment of, “Maybe we, too, can be an experiment in recursive self-improvement.”
So, you're now Elicit as a test case for: Can we get Elicit to build effective world models and then be able to use these detailed representations that it itself has constructed to inform its own analysis of what it itself should become in the future? The hall of mirrors there is deep and fascinating.
This obviously relates to a blog post that you put out not too long ago called “Planning Is Unsolved.” And I think, yeah, obviously that's totally true—or we could just ask the AIs to handle all this and retire to the beach, as I think it was an Anthropic person who once famously put it.
I do wonder a little bit, and this is probably a cultural question as much as it is a technology question. When I think about my own company, which I'm now just the AI advisor to and not running, and I think about the planning process that we developed before AI, sometimes I'm like, maybe we should just scrap the whole thing. Maybe all these times that we come together, sit around the table, talk about this or that, and try to convince each other that if we do this, it'll be more successful than if we do that—I'm often like, man, you know what we should do? Build it all, launch all these things, and see what happens.
Coding has gotten cheap. So maybe the future of planning is less about guessing and more about fast iteration and actually making more contact with reality. Obviously, again, there's not a true, strict binary there. But how are you guys thinking about that question? And how are your clients, especially pharma companies, thinking about that question?
There's also this notion of clinical trial abundance. I think it's more of an aspiration than a trend at this point. I don't think we're there, but you can imagine a very different vision for a pharmaceutical company. One is, “We're going to use these world models, make much better decisions, and deploy these scarce resources in the precious few at-bats we have at clinical trials in the best way possible.”
Then there's this other vision: “We'll do 10 times as many clinical trials, and that'll be the bigger unlock because we'll actually get the real answers in far more cases.” Or 100 times—who knows? Where do you want to be on that spectrum? Where do you think pharmaceutical companies should be on that spectrum? Maybe it's different. Maybe Waymark should be in one place and pharma should be somewhere else.
Yeah, I think it's great that the cost of software engineering has come down, but it feels like that was just one of the bottlenecks, and the others haven't moved. I don't think user attention or feedback is infinite, so I still feel cautious about just throwing things out there, giving people bad experiences, or leaving people with bad impressions.
I also think that it's definitely possible to get stuck on a local maximum. Let's say you just ship a feature and you're like, “Oh, cool. This works. Let's just keep going.” You might be able to keep going for some time, but then still end up maintaining, improving, and investing in something that wasn't the best possible thing you could have done.
So, I think there's still a lot of room for judgment and purpose. And there are places where I'm not sure the shape of the problem has changed that much. I feel like there's always an explore-exploit trade-off, and you want to navigate those 2 thoughtfully. Sometimes you want to do one, and sometimes you want to do the other.
Maybe the cost of exploring has gone down a little bit, but I think there are still some things where you have to get it right the first time. It's not as if all software engineering has literally gone to 0. There are still large software engineering projects.
So, I'm not sure it's changed that much. I feel like, for now, it's mostly changed things at the margins, like fixing bugs, small admin features, and things that we don't see as our core capability, that we want to fully automate and are happy to take liberal experiments with. Things that we see as being core to our mission, purpose, and differentiator as a product still involve a lot of careful thinking and intent.
With our customers, I find that many of them are just like, “Yes, there's a lot of excitement to build.” I guess my unsurprising take on the build-versus-buy problem is that you should build internally the things that are your core comparative advantage, and basically nothing else.
If you're building something, for example, there are certain workflows that are just regulated across the industry. Every company has to do them pretty much the exact same way, and they are very particular. As a result, they're very interface-heavy. I just don't think it's the core competency of pharma companies to design nice software, and that's not your comparative advantage as a company. Your comparative advantage is not going to come from solving this regulatory compliance problem.
So, I don't think, for example, systematic review, which Elicit solves, is one of those areas where it makes sense for pharma companies to try to build this thing that every company has to do the same way and that is actually very involved. Other types of things, especially in early-stage research or even in development—certain kinds of predictive models—make sense for a pharma company to build in-house.
And then, on the clinical trial point, I think, why not have both? There's a lot of interest in digital twins and simulating trial effects digitally as much as possible to design the trial well. I'm sure there are certain trials where, with the right regulatory framework and operational improvements, we might be able to take a lot more bets.
I think especially in rare diseases, where a trial is actually almost like a treatment option, we can be much more flexible there. In other domains, depending on the type of drug and what we already know about its toxicity profile, we'll probably want to hold a higher bar.
So, again, my hope is that it's great to have multiple tools, options, and choices. Maybe we can just build more tools, make the tools better, and then build a good framework for when you reach for what tool.
I think there are many cases where you just can't do everything at once. If you think about clinical trials from the participant perspective, usually you can participate in 1 or at most 2. Often there are many that you could participate in, but you still need to choose which of those you go with. That's a tough planning problem.
Likewise, as a company, you have a certain amount of resources, and they can be deployed one way or the other. I still find, especially as a company that is trying to be extremely mission-focused, that you need to ask: How do we actually, in the short time that we have, make an impact on the quality of reasoning and the impact of AI on reasoning quality?
By default, you're probably just not going to accomplish that. Even if you honed in on user metrics, I think you're probably just not going to accomplish that. So, you need to think pretty carefully about how to get that to work.
Your callout of the patient perspective is a really useful reframing there, because it's one thing to say, “Yeah, as a pharmaceutical company, maybe we can have both. We can run all the trials.” Certainly, as a SaaS company, I can potentially launch all the features, whether or not I should. I maybe could.
But if you've got cancer, you've only got 1 body, and you can't take all the drugs, right? That would obviously be ill-advised. So that, I do think, is a really useful point. At least for now, until the transhuman uploading future, whatever, appears, we are going to continue to face these very stark choices about what I should do with my own individual body as a human, knowing that there's not a second copy of it and there's not really the ability to diversify across paths in many cases.
Are there any examples of that you could share at the company level, especially if there's something where this sort of company model and recursive self-improvement paradigm of working has led, at least in your counterfactual analysis, to a different approach than you might have taken if you were just gathering around the table and sharing your intuitions, like people used to do?
Examples of where automation, learning through doing, and cheaper experimentation have led us to a different outcome.
Or maybe it's the same thing, but I'm thinking especially around planning with a world model. Is there something where, because you had structured these various lenses and were able to be more systematic and structured, you could really agree that this is the framework that you are using—maybe avoiding talking past each other or having some new synthesis, some new insight that you just don't think would have happened in the absence of that structured approach?
This is a partial example. I'm sure Andreas has a better one, but it's actually very timely. I just did this for an important hiring decision.
Ahead of time, we had a rubric designed for what role we wanted to fill, and the role had been through a few different evolutions. We had looked at different personas. There was maybe a disagreement on exactly what type of executive-level hire we needed. There was some disagreement on exactly what we needed, and maybe that changed over time as the company grew during the course of the search.
And then, about a month ago, I wrote down a framework, and then we did extensive interviews, so many references, lots of back channels. There was just so much information I was getting, and I was starting to develop a take on what we should make of it and what we should do with this candidate, but I really wanted to avoid recency bias, so I had a structured project. In this case, I used Claude, not Elicit. Elicit doesn't support hiring decisions exactly yet, but I don't use it for that case.
I created a project where I had all of the meeting notes across all the interviews, all the email threads with this candidate, and all the feedback submitted, and asked Claude systematically to fill out examples of evidence for every single dimension. It ended up being about 20 different fields, like evidence that this person has consistently hit their goals, evidence that this person can hire a great team, and evidence that this person is authentic or culturally aligned. I started with the evidence and went piece by piece because I don't trust Claude to fully execute in one go. There was a bit of calibration, and then once I had the evidence—which was, “Hit quota so many years,” and so on—I asked, “What is evidence for that? What decision do I make? How do I rate this on a scale of 5?”
Then all together, I had this synthesized point of view, and I was able to send it to the candidate. I think they really appreciated it as well because they said it was the greatest kind of comprehensive synthesis of professional validation they had ever received. That was one case where compositional, structured reasoning with an intentional process, and then applying that at scale with AI, was able to both check my decision-making process and give someone the gift of something very human and very detailed about them and everything they had accomplished.
My example of what I've done is going to be much more mundane. I think for me, I've been trying to do this more for just planning my week. I have goals, and this is what I want to accomplish in the long run—this year, this month—and then the question is: I have all these calendar blocks, like this podcast block, and I need to figure out that there are many different things I could do. What should I do, and which things depend on which other things?
It's actually a pretty tricky problem to know, when you could be spending your time in many different ways, what is worth doing. So I've been trying to get to the point where I can use automation as part of my weekly planning and think in a more structured way about this: If I want to accomplish my monthly goal, where do I need to be this week? How much time is that going to take? Is it maybe going to take 5 hours to write a blog post? When can those 5 hours happen? I think this sort of backwards chaining—people do it informally—but there is a lot of constraint satisfaction and propagation of constraints that is pretty tricky for humans, and I think the models will help us with that.
Yeah, that's cool. I think of that as building your own harness in a way, which is something I'm thinking about for myself, too. How can I build up structures around me to keep steering me in the right direction, feeding me the information I need, and hopefully helping me become my best self, use my time as well as I possibly can, by setting me up for success as much as AIs can do that?
Do you find that you are following it? How good is it? Are you actually living by it yet, or is it still, “Maybe next week, when it gets a little better, I'll actually do the plan that it gives me”?
It's still a very human-in-the-loop process. I actually have 2 versions of it. I have 1 automated version, which I hate, and 1 interactive version that walks me through the planning. I still do the fully automated one just to see how good it is, and I want to know whether, at some point, I'll be like, “Oh, yeah, I'm not needed here anymore.”
But generally, I'm like, “You just didn't fully understand what I'm trying to do. You made it too complicated,” and so on. So I'm still in the outer loop, but I think it's kind of interesting to think about it. Right now, humans are the outer loop, and they use calls to LLMs, but eventually an LLM calls you, and the LLM is the outer loop and you're just the inner loop. I'm not sure that's a positive future, but it seems like it's part of the trend here.
The Line is our automated software engineering project. I think, as is maybe the case for many companies, software engineering is where we've had the greatest success doing quite extensive automation. Our overall company goal is that, at the end of the year, when we go on vacation, we want the company to keep running and keep doing work in all of its functions. The first half of the year will be mostly focused on trying to make that happen for software engineering.
We have a system called The Line because it is like a factory line. Someone mentions a feature they would like to have on Slack, or a user mentions a feature, and we Slack-emoji-react to it with a little Line emoji. Or there's an integration with our customer-support system, and then it kicks off an iterative process: First, the feature needs to get specced out, then you need to iterate on the spec, it needs to be implemented, a video needs to be recorded of the feature being tested, then a code review needs to happen, and then it needs to get merged into dev and then into prod.
We do have a fully automated version of this now. For simple features, basically you just emoji-react to, “Oh, I would like it if Elicit kind of talked about its citations in a slightly different way,” and it will go through this entire process automatically. At the end, there are various judgment calls it makes about where human intervention is needed. Maybe the spec was too incomplete, and so it's like, “Okay, we need to pull in a human here,” or maybe the feature is too complex for the system to automatically review, and we need to pull in a human here.
But for many simple features, it can actually flow fully automatically through The Line, and I think that's already been a significant unlock for a lot of simple bug fixes and features. I think it's also setting us up for the future where, as each of these individual parts of The Line improves, more and more of software engineering will be automated. It's been pretty cool to see how we're now merging maybe 30 to 50 issues per week fully automatically.
Cool, that's really interesting. If indeed we are successful in having the company continue to function through the holidays without you guys for 2 weeks, then one wonders how long it could go and whether you ever have to come back for one thing. But what will take us there? Is it just the next generation of models? Mythos is supposedly coming soon to a public API near you, with supposedly much better long-horizon performance. Is that going to be the biggest unlock? You've got the structure, and I just need to drop in a better model, or what else do you think is going to be needed over the next 6 months to actually realize that?
Yeah, so first, I don't actually expect the company will run fully automatically by the end of the year. I expect our lower bar is that, within each function, there are pretty autonomous workflows that run and connect to some workflows in other functions. But I think a lot of the high-level steering will still be very much needed.
What is needed for scale-up? I think 1 big obstacle right now is that the models are not fully calibrated about when human intervention is needed. You have to be pretty risk-averse in how you use them. I think with software engineering, if 80% of the time when the model says, “This is an automatically reviewable feature,” it actually is, then that's not good enough, because we don't want to break production 20% of the time. That's pretty rough.
So we have to err much more on the side of: If in doubt, it's not an automatically reviewable feature. I think that's the case in software engineering, and I expect it's the case in other situations, too. If you were to let the models drive some customer interaction, for example, you probably want to be at least as sure as in the engineering case.
It's actually not clear to me how this will go. If you're following the METR graph, there's the 50% success-rate curve that goes up over time as the model gets better, and then there's the 80% success-rate curve. The 50% success rate is much higher, obviously, than the 80%, and the 80% hasn't been going up quite as fast as we would like. Often we want more than 80%.
Depending on how average-case performance compares to, I don't know, 95th-percentile performance, just dropping in the next models might be good enough or might not, but I wouldn't automatically rely on it.
It could be cool if, similar to fast mode, there were an ultra-reliable mode or something, which isn't just “think more,” but has guarantees on certain classes of errors that you're never going to make.
Yeah, that's cool. A lot of really interesting thinking there. One big question that is generating quite different takes at the moment is: Are people going to be able and willing to pay the exponentially rising token bills that the industry as a whole is currently seeing?
You could analyze this from any number of ways. One would be your own internal work, right? Where is your token budget compared to your head-count budget today in engineering? And do you expect that, with the introduction of models—if they get, let's say, a lot better, for some definition of “a lot better”—you will shift that budget and spend a lot more on tokens relative to humans than you do today?
And then do you think your customers will do that as well? Maybe it will break down by use case. It sure seems like, for as much as we hear a lot of complaining about token costs and, anecdotally, “Oh, this company pulled back” or “That company hit budget,” it still feels to me like there's a lot of value in the marginal intelligence and just getting better results.
Tokens are still pretty cheap. I mean, for most companies, it's still a small amount. I hear things like 5% or 10% of what we're spending on headcount, and that's not that much. Maybe you didn't budget for it, and that creates some discomfort in your organization, but on the fundamental economics, it feels to me like if you can just get a lot better work for somewhat—even maybe a multiple—of the token cost, it still seems pretty rational to pay it.
I guess that's my starting position. What do you guys think you will do? What do you think your customers will do?
I think that, often for our customers, at least, the offering we're providing is displacing services spend. So, the barrier is more about whether it can fully displace the services spend, and then also maybe getting over the mental hurdle of price anchoring for software.
Certainly, in terms of the dollars allocated to solving this problem, there are a lot more dollars in compute costs at the moment. So, I don't feel—yeah, we'll see. There are obviously human issues to overcome there, but I think from a dollars-and-cents perspective, there's still a lot of room.
I've heard and seen mixed things in the news, and I think the industry—at least the pharmaceutical industry—is fairly disciplined about costs and ROI. So, even if there's an initial period of heavy exploration, I think there's a lot of accountability around what that's delivering for the business. I think we'll continue to see that.
I think even internally at Elicit, I'm not sure how many more multiples of token costs we can easily spend. So, maybe taking myself as an example, I spend around $2,000 per week on tokens. I could maybe double it or triple it, I don't know, but not much more than that, for sure.
So, I do expect—and that is already influencing my behavior to some extent—that I don't actually currently use fast mode for these models, because I don't feel the marginal returns are high enough for most tasks. I expect it's unlikely that I'll be like, “Wow, I need to switch over everything I do to Mythos.” It's probably not going to happen.
Both for my own usage and also because this is already actually the case in the Elicit app, I think more of what it will look like is one smart orchestrator agent that then spins off many other agents that have to do simpler tasks that just don't need to use the largest model.
I expect that will just become increasingly important: this sort of dispatching to a model of the right size, so that you get the intelligence when you need it, but you're not just multiplying your whole spend by some number that was an inefficient use of compute to begin with.
So, $2,000 is not a small amount. Is that an outlier? Does that make you an outlier within the company, or is everybody doing that? If so, that would put your token costs—not presumably at the level of payroll, because I assume you're paying your engineers more than that—but it would be at least a not-insignificant share. Certainly, if you were to triple it from there, you'd be getting into something on the order of magnitude of parity with human headcount.
What are you doing with it all, too? I use my $200 Claude Max and my Codex Pro, and I honestly don't even hit my limits that often. Now, this may be API, which might be 10 times more, and so that could be a big part of it, but I sometimes feel a little ashamed that I'm not redlining the account more than I am.
What would you advise me to do, or what sort of personal bitter lessons have you learned where you're really finding that token-maxing is worth it?
Yeah, I'm not sure I'm the top user of tokens at Elicit, but I'm probably at least in the top 5, so I'm probably a little bit of an outlier. Second, I'm using the API. I could probably save more money by being more clever about how to use various Pro accounts and stuff.
I do have a fairly elaborate system built on Pi that orchestrates between the different agents and uses ChatGPT to double-check, then sometimes calls Claude or Gemini to get another take. That is a little bit easier to do if you're on the API than if you're on the normal end-user plans. That might just be part of the explanation here.
Any particular use cases where you get particular value that you think other people might be sleeping on?
I don't know what other people are doing. As mentioned earlier, I get a lot of use out of planning, keeping my calendar in sync with my personal journaling system, keeping that in sync with my to-dos, and making sure everything is coherent with my longer-range planning document.
When the new day starts, I go over the last day, check whether there are any leftover tasks, and move them into the right place. So, there's a lot of automation happening behind the scenes without me prompting it that probably contributes to those costs being higher.
Similarly, for email, I have a pretty elaborate spec on which emails should be auto-archived. Maybe every hour or so, my models check that and go, “Okay, let's just archive the emails that Andreas definitely doesn't need to read.”
Then, on the more user-driven side, I do a lot of cross-checking. I sometimes think of things we could talk about with Nathan, and then ask ChatGPT and Gemini to double-check those things. I find that having the models cross-check one another often improves the results quite a bit, so for—I don't know—a fourth of my use cases, that already doubles or triples the cost. That's another source of additional token spend.
Got it. Okay, cool. One other question I'd love to get your take on is: Are we seeing convergence or divergence in models? One notable feature of Elicit today is that there's no model picker, at least from what I've explored recently.
So, you're making choices, and it seems like you clearly think you know best. It would not be a good idea, even if people have a favorite model, given all the validation and scaffolding that you have, to just go in and swap models in and out.
How do you see this dynamic shaping up? Again, there are just such different takes: The models are commodities; scaffolding is all that matters. No, the models are everything; scaffolding is a complement. They're converging; they're diverging. What is your take on all of that?
Yeah, I keep being shocked by how much the models are converging. I guess I should stop being shocked at this point, because I'm just not updating, but it is a really interesting and surprising fact about the world that the models are so similar.
I do think that's the case, but that's not the reason why we don't offer a model picker. The reason is that I think a lot of tasks in Elicit still involve multiple models orchestrated in a way that we think makes the most sense, with particular models that are good at screening papers or extracting data.
Even though the models are so similar, I think the differences are important in subtle ways. For example, I think people love to hate on Gemini, and so do I, but when we evaluated Claude Opus 4.5 against Gemini 3 Pro at the time, I think Opus did better on extraction accuracy. But if you checked what fraction of claims are directly supported by the evidence, I think Gemini actually beat it by at least 5% or so.
The models are still micro-jagged enough that you can't say, “Oh, yeah, this model is clearly the best; you should just use that.” As a user, I don't want to put that on our users for the most part. I think mostly what our users pay for is for us to do the work of figuring out what models are good at what kind of thing and making sure those models actually get used in those places.
Yeah, interesting. So, there is a place for Gemini in Elicit today.
Yeah, there was a place 2 months ago. I actually don't know. Even though I'm fairly on top of what's going on, our eval team is even more on top of it, so I don't actually know if it's still live. But there definitely was a place for it 2 months ago, and maybe next week there will be a place for it again.
Yeah, and it's cool that that in and of itself is an interesting reflection of how frequently you're swapping things out and how dynamic and competitive the environment is. Maybe 3 more questions, if you will. One, do you have plans to expose Elicit as a tool for random people's Claude Code to use? That could be by allowing them to do it via the API if they have an account, or, even more likely, it could be through a sort of x402-type thing, which I've recently been exploring as a way to get just pay-per-use access to a bunch of different tools. Yes? No? Why not?
But no, I mean, it's already—we haven't been advertising it that much, but we already have an MCP and an API. A lot of people use the API. People can check it out at docs.elicit.com.
A lot of the work I do with Elicit is through the API. When I run systematic reviews, I often use the Systematic Review API, and I iterate using various other models on the protocol for a while. Then I run it in the background and retrieve it. I think that's an important use case that we really like to support. Not everything has to happen through the interface, and I think more and more will happen through APIs.
Yeah, okay, cool. I'm sorry I missed that in my prep, but I'll again point myself at the documentation. I'm going through a bit of an AI-for-science mini-arc right now, and I'd be interested in your take on other big-picture approaches to AI for science. You guys are obviously coming at it with the systematic-reasoning angle.
There is the sort of close-the-loop angle, where we empower these models to actually run experiments through a cloud lab or whatever, and then they'll be getting feedback from reality. That seems like it could go somewhere quite interesting. Then there's, of course, training models on other modalities of data. We've seen how proteins fold, and there's, “What if I do this perturbation to a cell? What's its next state going to be?” You can go on and on in that domain.
Any interesting takes that you think might be non-consensus that you'd like to share?
I think all of this stuff is super exciting. Sometimes people come to us and they're like, “Well, who's going to win in AI for science?” I think that's just an absurd thing to say because science is such a big space. As you just said, there are many layers of abstraction, from understanding single-cell dynamics through automated experiments to protein models.
When you talk to the pharma companies, they're like, “We are trying to make a multiyear plan that accounts for the changing technological environment, but also accounts for the fact that clinical trials have certain intrinsic timescales.” There's just such a different reasoning problem in modeling single-cell dynamics. I don't know, it's a big space; it's a big pie. I'm excited that people are excited about it.
Yeah. Do you have any controversial takes here?
How it all comes together is maybe a really interesting open question. When people think about what the automated company of the future is, does it look like you take the existing top 20 pharma companies and, over time, they will morph into a different functional form? Or is it going to be the case that a small biotech comes along and they're just much more AI-integrated? In 10 years, the top 20 companies will all be replaced by companies that we don't even know the names of today.
I actually don't know how it's going to shake out. I think it could go either way, depending on how quickly people at the existing companies wake up and understand how much everything is going to transform. So, yeah, maybe that's not a very interesting, controversial take, given that I don't have a take between those 2 futures.
One big question I think about a lot is how integrated the models themselves will be. Obviously, we have the tool-calling paradigm coming along very nicely. This could be extended to a tool call to run an experiment in a cloud lab and get a result. That result could come back as a data printout of the same sort that a human would read.
Then there's this other paradigm of integration, which we're seeing a leading indicator of with image and now also video, with Google's latest Omni model. There's this sort of deep integration of language and pixel space. In the early ChatGPT image-generation experience, you would talk to the model—or even if you gave it a photo, it would try to caption that photo, describe it, and then use language to call the separate image-generation model and ask for something.
Of course, the people never quite looked like the ones you put in, right? Because you just can't describe a face in language with that level of fidelity. But now you have this deeper, weight-level integration where I can give you an image and say, “Make this a line drawing,” or whatever, and it has both. It understands conceptually what I want, but it also sees, in some sense, the structure of the face and can preserve that through the transformation.
So I really wonder if that's coming to all the modalities of science as well, and if it's a good idea. I think it probably is coming, but I wonder if you think it's a good idea, because it certainly would, in some ways—or at least, naively, it seems to me like it would—make process supervision more difficult.
If I can trace it and say, “Okay, you called the protein-folding model. This is what you got back. Okay, that's where you went wrong,” right? I can see us digging in and interrogating those traces a lot better, versus it just being, “I asked you for this,” and you spit out a new protein sequence because, in your weights, you were like, “Oh, I intuitively know what a sequence will do for that function that you just asked for.”
Obviously, that could be really powerful in images—it's a major unlock—and I don't see any reason it wouldn't be a major unlock in designing new proteins or what have you as well. So, do you think that's coming, and do you think it would be a good or bad idea if it does, in fact?
Yeah, I think it's a really interesting question. I think there's a bigger question behind this, which is: Where do continuous representations win? The prior should be, you know, end-to-end optimization is strong: integrate everything. At the same time, you might have expected neural nets—at least language models—not thinking in tokens but just thinking in weight space to be more successful than they have been.
I think a lot of people have experimented with it, and I hope no one succeeds, but, on priors, I would have expected maybe people to succeed at it. Often, I think people forget that there are benefits even to the models from the discretization that comes along with that.
I guess you can ask yourself: Why is human language the way it is? Why does it have discrete words in the first place? Why do we think in words and sentences? Even the models will benefit from the discretization. Why don't the models just write programs in weight space? Why do we even have programming languages? Why doesn't everything happen in continuous space? Maybe that's eventually the future; I don't know, but I wouldn't necessarily bet on it.
I think the straightforward “Everything will be end-to-end optimized in weight space” take is probably a little bit too lossy to be a good predictor of what will actually happen. In the protein-dynamics case, I could actually see continuous representations being pretty good, but I'm not sure that will be the case everywhere. I think discretization has benefits that people sometimes overlook.
Give me one more beat on what you think those benefits are, because I totally agree with you that I don't want to see neuralese take over, win out. At the same time, when you say, “Why do we think in language?” my immediate answer is, “Because it's all we have,” right?
I don't think in language—somebody throws me a ball, and I don't think in language about it because I'm a physically embodied person who has intuitive physics, and I just catch the ball. I feel like if I had those sorts of senses for how proteins fold, I'd probably use them, but I probably wouldn't think in language about them.
I mean, I think the fundamental property you get from discretization is error correction, right? If I say a word like “a little brown,” you can still round it off to, “Okay, that's a word.” At every step, you get a little bit of error correction.
If you're trying to chain together many words, you don't get the compounding errors. I'm not an expert on this, but I think this is roughly why we don't have analog computers these days—why we have discrete computers—is because you get these nice error-correction properties. That could be one thing you might lose if you're trying to push everything into weight space.
Yeah, interesting. Okay, last one. Zooming out and just going back to your original mission of radically improving the quality of reasoning in science and in society, how do you think we're doing? We've got inference everywhere. Is it serving us well, and how would you handicap the trajectory that we're on as we think about recursive self-improvement possibly soon, transformation possibly soon?
Is the quality of reasoning on track to rise to the level that we need it to? What's the state of the species, so to speak?
What is the state of the species? I don't know.
I mean, there's this meme where a person jumps from the roof and is like, “So far, so good,” as they're halfway down. So I think, in many ways, the models have probably improved our reasoning. I think I probably get better answers for many questions that I care about than I did in the past. But I also think that is not necessarily indicative of what will matter most to the species.
When I think about what is going to matter most, it's one of the decisions that governments and big AI projects are going to make, or maybe other large organizations, as AI transforms everything. And there, I think the game is still open. I think, in many ways, we're still extremely early. I think AI has transformed basically nothing.
I think people are always like, “Well, coding is getting automated.” But I don't know. Most intellectual work in the economy is not coding. It also still employs many people. And so, I think we've seen nothing yet. We're still early, and I also think all the big decisions are still coming up.
When it comes to how AI will impact epistemics and good reasoning, I think it really could go either way, because the models are optimized to look good and be persuasive. I could see a worsening of epistemics happening if we don't make this an explicit priority. At the same time, I don't know. I mean, I think the models can be optimized for truth-seeking, too.
And if you prioritize those interventions, there's maybe a basin of attraction where, once you become more truth-seeking, you realize, “Okay, what are the most important interventions?” Oh, I guess we should prioritize being better at forecasting the results of what we do. And then you become better at prioritizing more epistemics-related interventions.
So I do feel like, for better or worse, we're still before the point of no return in either direction. I'm excited for people to build tools in this space. Elicit's trying to build tools in this space. I think people can help advocate for the adoption of better epistemic tools. It does feel like a very small area relative to how important it seems to me to be for the future of our species, as you said.
It was the smartest of times; it was the stupidest of times. Here's hoping that our better angels win out when it comes to better reasoning and better decision-making.
This has been an excellent conversation. I really have enjoyed the update on Elicit, and I appreciate how consistent and disciplined you guys are about it. It's not easy, obviously, running a business and trying to make sure you're staying true to that North Star mission. I think you guys do a really admirable job of holding yourselves accountable to trying to find the right balance between those, and there are many echoes of the hard work that you've put into that in this conversation.
So I really appreciate the time and encourage you to keep up the great work and the disciplined reasoning. Anything else you want to leave people with before we break?
I think knowing when the models are making you better or worse at decision-making is actually pretty subtle, and I think not that many people are paying close attention to it. I'm reminded of the METR study from a while back, where engineers thought they were being made more productive, but actually they were at a slight discount relative to unassisted work. And that was in engineering—probably no longer true in engineering—but I could imagine that, for more complex decisions, if you're not paying attention, sometimes using the models regresses you to the mean or cuts off avenues of investigation, and at other times it opens up the space and makes you think more clearly.
So I find it helpful to just introspect: Am I actually getting benefits here, or how is this changing my behavior? Thinking about that and sharing it broadly does seem to me like just a clearly net-good thing.
Yeah. For the time being, we still have some agency over this process, and I think that's a great reminder to maintain an ownership mindset and hold ourselves accountable to doing our very best work—not getting lazy and letting the AIs lead us around.
Excellent. Jungwon Byun and Andreas Stuhlmüller, founders of Elicit, thank you again for being part of The Cognitive Revolution.
Thanks, Nathan.