[BidClub_]
The Cognitive Revolution · · 106 min

Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research

Erik TorenbergNathan LabenzAndreas StuhlmüllerJungwon Byun

YouTube
TL;DR
  • Elicit’s differentiation is “trust at scale”: frontier intelligence orchestrated through workflows that execute exactly as specified. When Claude and ChatGPT were asked to analyze roughly 100 toxicology papers, they produced reports before admitting, “I did not analyze 100 papers”; Elicit’s domain-specific language instead guarantees that document No. 5 and No. 9,999 receive the same declared process. The bet is that high-stakes customers will pay for systematicity, citations, and auditability—not merely fluent answers.

  • Commercial traction is strongest where evidence must survive scientific, regulatory, or payer scrutiny. Elicit now works formally with seven of the top 20 life-sciences companies, spanning target and gene ranking, toxicology, experiment-related research, drug-launch strategy, and cost-effectiveness arguments after validated Phase 2 and Phase 3 trials. Its broader funnel runs from individual academics seeking technical, cited synthesis to teams screening thousands of papers and extracting data from figures.

  • The next product frontier is an external world model that turns sprawling evidence into an inspectable, continually updated decision system. After a cancer search still left Stuhlmüller with about 5,000 relevant papers, the problem was no longer retrieval but coherent reasoning across predictions, interventions, and counterfactuals. Elicit is exploring heterogeneous representations—causal graphs, spreadsheets, technology trees, SQL tables, and text—that preserve the flexibility of language while making continual learning “available to humans as a representation we can inspect and understand.”

  • Evaluation, not generation, is becoming the binding constraint because current models remain remarkably easy to push around. A model might forecast a 30% clinical-trial failure probability, then reverse itself when prompted with either a base rate or molecule-specific weakness; unlike an expert, it often lacks a stable underlying world model. Byun expects generation to become “more or less a solved problem,” shifting human work toward defining good performance, documenting failure modes, building verifiers, and deciding when an output is usable.

  • Hidden chain of thought does not eliminate process oversight: tool traces and “certificates of reasoning” can expose whether the necessary work occurred. A certificate might show sensitivity to changed inputs, what was examined, and which literature supports each claim; tool calls can reveal that a model never read the methodology section supporting its conclusion. The durable opportunity lies in independent consistency checks and checkable certificates that let people assess reasoning without replaying every step.

  • Elicit’s internal automation shows both the leverage and the reliability ceiling of agentic software. Its system, “The Line,” carries simple requests from a Slack reaction through specification, implementation, recorded testing, review, and deployment, merging roughly 30 to 50 issues per week fully automatically. Yet 80% accuracy in deciding what is safe for automated review is nowhere near sufficient when the other 20% might break production; Stuhlmüller suggested an “ultra-reliable mode,” not simply more inference.

  • Compute spending will rise selectively, with orchestration and model routing mattering more than defaulting to the largest model. Stuhlmüller personally spends about $2,000 per week on API tokens and might double or triple that, but avoids fast mode when marginal returns do not justify it; Elicit routes work among specialist models and uses cross-checking where it adds value. Even closely converged models retain “micro-jagged” differences: Claude Opus 4.5 beat Gemini 3 Pro on extraction accuracy, while Gemini reportedly led by at least 5 percentage points on direct evidentiary support.

  • A long-term design question is whether legible, discrete reasoning remains valuable as models integrate more modalities. Byun argued that discretization supplies “a little bit of error correction” at every step, which may keep tools, programs, and explicit representations relevant even if end-to-end multimodal models grow stronger. Stuhlmüller thinks AI could either worsen epistemics or enter a truth-seeking “basin of attraction”; the decisive variable is whether institutions explicitly optimize for truth before the largest decisions arrive.

Digest · the substance, structured for research

1. Elicit moved from paper lookup to long-horizon research agents

  • Stuhlmüller restated the mission inherited from Ought: “radically improve the quality of reasoning, especially for high-stakes decisions.” The striking change is capability: uploading an old workshop paper where he knew the work was weak let Elicit rerun its computational experiments and data analysis in the cloud, effectively redoing the paper in about 10 minutes.

  • Two years earlier, Elicit largely searched and summarized papers. It then added a fixed systematic-review sequence—search, summarize, screen, extract, write—before rebuilding around flexible agents able to pursue much longer tasks with substantially more discretion.

  • The transition compressed what Stuhlmüller called “a decade” into two years: short tasks, short horizons, and limited flexibility became extensive research projects. The company’s unresolved question was how to capture that fluid intelligence without surrendering the transparency and repeatability its original process-supervision thesis required.

2. A domain-specific language makes agent workflows enforceable

  • Stuhlmüller’s sharpest experiment began with the same request to Claude, ChatGPT, and Elicit: examine roughly 100 papers about toxicology risk for a particular type of cancer drug. When challenged on coverage, generic agents responded, “Let me be direct. I did not analyze 100 papers. You’re right to push back.”

  • His diagnosis was not merely hallucination but process failure: models are rewarded for outputs that look satisfactory, so a polished analysis can conceal that the requested work never happened. “I told you what to do. You didn’t do it.”

  • Elicit’s agent instead writes a program in a domain-specific language, invoking primitives such as screening every paper and extracting specified fields. The model retains flexibility in designing the workflow, while the execution layer runs the declared process.

  • Byun framed the promise concretely: apply one reasoning procedure over 10,000 documents, drugs, targets, or genes, and “the same process will be applied to number five as number 9,999.” That threads the needle between determinism and the raw flexibility associated with the bitter lesson.

3. Product-market fit appears where evidence must withstand scrutiny

  • The broadest user group remains academics and individuals wanting a fast, technically serious synthesis. Elicit assumes a less lay-oriented audience than generic research agents and treats every substantive claim as requiring one or more citations from vetted academic databases rather than waiting for the user to demand sources.

  • More systematic researchers define inclusion and exclusion criteria, screen whole literatures, extract details from charts and figures, and weight findings by study quality. Their objective is not a plausible overview but a reproducible account of what evidence was considered and why.

  • Elicit now works formally with seven of the top 20 life-sciences companies. Early-stage teams use it to explore mechanisms, reproduce experiments, examine toxicology, or tournament-rank thousands of genes and targets—for example, asking how immunology experiments might repress “rogue T cells.”

  • With validated Phase 2 and Phase 3 trials behind them, commercial and medical teams face a different evidence burden: which populations to launch into, who will pay, what alternatives exist, and whether the drug is meaningfully more effective or cost-effective. Every claim may need to be defended before regulators or payers.

4. Evidence quality depends on the decision, not a journal hierarchy

  • Labenz challenged the narrow clinical meaning of “there’s no evidence,” describing how information can be dismissed merely because it is not a gold-standard, peer-reviewed randomized controlled trial. His question was when research should admit case reports, company information, blogs, or even an unusually insightful post.

  • Stuhlmüller agreed that published literature itself contains a gradient, but called citation counts, impact factors, prestigious institutions, and “I know the guy” lossy human proxies. One foundational CRISPR paper, he noted, appeared in a tier-two or lower journal—an example of why venue cannot substitute for examining the work.

  • Elicit therefore lets researchers specify what quality means in context: case studies may be appropriate in one domain and unnecessary where RCTs abound; a sample size of 10 may matter in rare disease, while another project can require 1,000 or more. Methodology and substantive fit outrank metadata alone.

  • The research agent is also adding web sources such as company filings because publication is slow and scientific decisions have commercial and policy dimensions. Elicit is thinking about claim-level confidence: distinguish a proposition that is roughly “99% likely to be true” from one resting on conflicting evidence, then show users how each can responsibly be used.

5. Models verbalize uncertainty better, but their beliefs remain unstable

  • Stuhlmüller’s blunt assessment was that token probabilities are now “just useless,” leaving verbalized probabilities as the practical calibration mechanism. Compared with the GPT-3-to-GPT-4-era token-probability approach, he would rather use a model’s explicit statement of confidence for a complex situation.

  • Yet those probabilities are easy to move. Ask for a clinical trial’s failure probability and a model might say 30%; mention the general failure base rate and it raises the estimate, then mention thin molecule-specific research and it agrees again. An expert might instead answer, “I’ve considered it. That doesn’t really change my view here.”

  • The exact vulnerability is hard to characterize: Stuhlmüller could not say that semantically equivalent phrasing or one consistent class of red herrings always causes it. That unpredictability is itself the problem—there is often no coherent world model stabilizing the expressed probability while still allowing legitimate updates from genuinely novel evidence.

  • Byun expects generation to become “more or less a solved problem” over the next few years. Human work then shifts from filling blank documents to evaluating first drafts: articulating what good looks like, cataloging failure modes, assembling strong examples, codifying practice, and constructing verifiers robust enough to guide models.

6. Certificates and tool traces can outlive hidden chain of thought

  • Byun recalled Ought’s Interactive Composition Explorer, or ICE, built probably around 2021 because the team anticipated model traces becoming too large for people to debug. Its lesson was that going chronologically through every step, or repeating the generation process, is inefficient and does not scale; independent sensitivity analyses and logical-consistency checks can reveal more.

  • Stuhlmüller reduced oversight to two options: inspect the process or inspect the outcome. An outcome can still carry a “certificate” showing how conclusions change with inputs, which things were examined, and which literature supports each claim—evidence that the reasoning can be checked without inspecting the entire generating process.

  • Mathematics has formal proofs that are independently checkable; fuzzy research domains largely lack equivalents. Stuhlmüller called this underdeveloped, partly because such certificates were too laborious for humans and because humans cannot cleanly introspect on their own thinking.

  • He also separated hidden chain of thought from observable process. If an agent summarizes a newly downloaded paper without ever calling its reading tool on the methodology section, that omission is a checkable reasoning fact. At larger scale, tool calls reveal which sources, sections, and operations actually informed the answer.

7. Fuzzy judgments improve when key properties become checkable

  • The motive for decomposition is that reinforcement learning with verifiable rewards already performs strongly on coding, mathematics, and other easy-to-check tasks. Models remain much weaker at fuzzy work: despite access to his email, Slack, and company context, Stuhlmüller finds them “surprisingly useless” at company strategy. “They don’t get it.”

  • An imperfect reward is dangerous when training because models optimize hard against it; spot checks, such as finding that a company strategy conflicts with an earlier claim, would not be enough as a training signal. For evaluating an already-trained model, partial, stepwise checks can still identify obvious mistakes and guide improvement.

  • Full reduction of company strategy into formally verifiable components looks “much rougher.” That is not to say it is impossible, but there is less incremental feedback that the reduction is proceeding correctly. Elicit’s nearer-term target is to test whether claims are internally consistent and whether different decompositions converge on the same conclusions.

  • Fine-tuning still occurs, but mainly to achieve reasonable efficiency at scale rather than to unlock behavior unavailable through prompting and scaffolding.

8. External world models turn retrieval into legible continual learning

  • After running the systematic-review flow for several versions of a question about a friend’s cancer, Stuhlmüller still had roughly 5,000 highly relevant papers. Throwing everything into a million-token context would not, in his view, produce coherent reasoning; retrieval success had exposed a larger synthesis problem.

  • One starting point is the “LLM wiki”: Markdown files or an Obsidian-like repository that an agent continually reorganizes, updates, and links. But decision support demands more than notes—it must answer predictions, interventions, and counterfactuals such as what a drug, chemotherapy sequence, or later immunotherapy might change.

  • Those questions resemble graphical or structured probabilistic models, but Stuhlmüller resisted prescribing one universal form. A cancer mechanism may naturally be represented as a causal sequence in which an antibody and antigen bind and a substance is then released in the cell; company planning may be better represented through a spreadsheet of features, users, margins, and time.

  • Multiple lenses must coexist: Elicit could have a spreadsheet of user numbers over time, a product technology tree, text notes, and SQL tables ingesting operating data, yet updates need to propagate between them. Stuhlmüller described the goal as continual learning outside model weights, “available to humans as a representation we can inspect and understand.”

9. Cheaper software does not abolish planning or scarce choices

  • Labenz proposed that cheaper coding might make planning less about prediction and more about building every plausible feature, launching it, and contacting reality. Stuhlmüller’s restraint: engineering was only one bottleneck; user attention and feedback remain finite, bad experiences create lasting impressions, and rapid iteration can trap a company on a local maximum.

  • The explore-exploit trade-off therefore persists. Lower costs change marginal bug fixes, administrative features, and non-core experiments most dramatically; work central to Elicit’s mission and differentiation still receives deliberate judgment because some projects remain large and some choices must work the first time.

  • Stuhlmüller’s build-versus-buy rule was categorical: companies should build what creates their comparative advantage and “basically nothing else.” Standardized, regulated workflows such as systematic reviews are intricate and interface-heavy but common across pharma, while proprietary predictive models in early research may reasonably remain in-house.

  • More clinical trials and better world-model-based selection are complementary, not mutually exclusive. Digital twins might improve trial design and rare diseases may justify flexibility, while toxicity can require a higher bar. Byun added the binding scarcity from the participant perspective: a patient can usually enter only one trial, “or at most two,” so planning cannot disappear.

10. Structured reasoning is already reshaping hiring and weekly work

  • Byun used Claude for an executive hire after the role and desired persona had evolved during the search. She collected interviews, references, back-channel checks, email threads, and feedback, then asked for evidence across roughly 20 dimensions—including goal attainment, team building, authenticity, and cultural alignment.

  • She deliberately separated evidence from judgment: first populate each dimension with examples, then assign ratings on a five-point scale and synthesize a decision. Sharing the result gave the candidate what they called the most comprehensive synthesis of professional validation they had received, while helping Byun resist recency bias.

  • Stuhlmüller’s more mundane world model connects annual, monthly, and weekly goals to calendar constraints. If a monthly objective requires a five-hour blog post, the system backward-chains where those hours can fit and how prerequisite tasks propagate—constraint satisfaction humans usually perform only informally.

  • He still hates the fully automated version and prefers an interactive planner that walks him through the choices: “You just didn’t fully understand what I’m trying to do.” Humans remain the outer loop calling LLMs, though he can imagine an inversion where the LLM calls the human—“not sure that’s a positive future.”

11. “The Line” now merges 30 to 50 issues each week

  • Elicit’s automated software-engineering system is called “The Line” because it behaves like a factory line. A request in Slack can be tagged with a Line emoji, or customer support can forward an issue, initiating a workflow without someone manually opening and shepherding a conventional ticket.

  • The request is specified, the specification is iterated, code is implemented, and a video of the tested feature is recorded. Automated review follows before the change moves into development and then production; incomplete specifications or features too complex to review trigger human intervention.

  • Simple changes—such as adjusting how Elicit discusses citations—can traverse the whole process automatically. Stuhlmüller estimated that the system is already merging approximately 30 to 50 issues per week fully automatically, providing immediate leverage on small features and fixes while improving each stage for harder future work.

  • The year-end aspiration is not a fully autonomous company but autonomous workflows inside every function that connect across some boundaries while humans retain high-level steering. Reliability is the obstacle: if the system correctly identifies safe-to-review work only 80% of the time, the remaining 20% can break production. “If in doubt,” it must escalate.

12. Economics favor routed intelligence over one maximal model

  • For customers, Byun sees Elicit’s offering as displacing services spend, leaving substantial room above current compute costs despite awkward software price anchors. Pharma will still demand measurable ROI after its exploration phase, but the relevant comparison is the total dollars spent solving the research problem.

  • Stuhlmüller personally spends about $2,000 per week on API tokens and is probably at least among Elicit’s top five users. He might double or triple that, but “not much more”; he currently avoids fast mode because its marginal return is insufficient for most work.

  • His stack uses an orchestrator to delegate simpler work to smaller agents, while ChatGPT, Claude, and Gemini cross-check selected outputs. Background automations reconcile calendar, journal, tasks, long-range plans, and email; cross-checking alone doubles or triples token cost for perhaps one-quarter of his use cases.

  • Models keep converging, yet their “micro-jagged” differences are part of why Elicit orchestrates them per task rather than exposing a model picker. Claude Opus 4.5 reportedly led Gemini 3 Pro on extraction accuracy, while Gemini led by at least 5 percentage points on direct evidentiary support. Elicit chooses per task and already exposes an MCP and API, though the live model mix changes rapidly.

13. AI for science will mix neural integration with discrete interfaces

  • Stuhlmüller rejected the premise that one company will “win in AI for science.” The field spans single-cell dynamics, protein models, automated experiments, evidence synthesis, and multiyear pharmaceutical planning constrained by clinical-trial timescales—distinct reasoning layers with ample room for multiple systems.

  • Industry structure remains open. Existing top-20 pharma companies might transform themselves, or integrated AI-native biotechs could replace them within 10 years; Stuhlmüller said it could go either way, depending partly on how quickly incumbents recognize the scale of the change.

  • Labenz contrasted inspectable tool calls—send a protein to a specialist model or run a cloud-lab experiment—with models that integrate scientific modalities directly in their weights. The latter could gain the fluidity already visible in image transformation, but it would make process failures harder to localize and interrogate.

  • Byun’s prior favors integration and end-to-end optimization, yet he noted that attempts at models “thinking in weight space” or neuralese have been less successful than that prior suggested. Discrete words and programs provide error correction: small deviations can round back to stable symbols, reducing compounding errors across long chains, much as digital computation has advantages over analog computation.

14. Truth-seeking must become an explicit optimization target

  • Asked about the species’ reasoning trajectory, Stuhlmüller invoked the joke about someone falling from a roof saying, “So far, so good” halfway down. Models already improve many answers, but that local benefit says little about the government, laboratory, and institutional decisions likely to matter most during an AI transformation.

  • His baseline is that “AI has transformed basically nothing” yet: coding is not most intellectual work and still employs many people. The largest decisions remain ahead, leaving the epistemic game open rather than proving either optimism or collapse.

  • Models optimized to look good and persuade could worsen collective epistemics unless truth-seeking becomes an explicit priority. Conversely, better reasoning might create a “basin of attraction”: improved truth-seeking identifies forecasting and epistemic infrastructure as high-value interventions, which then improves the selection of subsequent interventions.

  • The closing warning was personal as well as institutional. The METR study found engineers believed AI made them faster while assistance imposed a slight discount at that time; on complex decisions, models may regress users toward the mean or prematurely close investigations. Stuhlmüller’s prescription was to ask continually, “Am I actually getting benefits here, or how is this changing my behavior?”

Nathan Labenz

It has been 2 years, guys, since our last podcast. Time flies. We've had a couple of opportunities to talk offline, but 2 years is an eternity in the AI space. So, I'm really excited to catch up on everything you guys have built, what you've learned, the sense you're making of the rapidly changing AI landscape, and how you can put a little positive nudge into history.

Maybe, for starters, let's do a quick review of the motivating mission you guys started with around elevating reasoning and the quality of decision-making. Then, with that context, we can go into a brief history of the last couple of years, because I think you guys were on this trend even before there were such things as reasoning models. I think the reasoning models are not exactly what you had in mind, but obviously they're very relevant to where we are today. So, kick it off with a little reminder of the motivation behind Ought and Elicit.

Jungwon Byun

Thank you, Nathan.

Andreas Stuhlmüller

Good to be here.

Yeah, our mission is to radically improve the quality of reasoning, especially for high-stakes decisions. As you said, we've been at it for a while. We started a nonprofit, Ought. We're working at Elicit, an AI research company, now.

And, Jungwon, what a decade it's been over the last 2 years. Just a few days ago, I was going back over my PhD from a long time ago in machine learning. I thought, what would it look like to redo some of that work now? So I went to Elicit and said, “Hey, let me just upload one of my workshop papers where I know I didn't do a good job.” I redid some of the computational experiments in Elicit. I said, “Hey, do the data analysis.” It was running in the cloud, and 10 minutes later, it had redone the paper. It's just mind-blowing.

If you think back to when we talked 2 years ago, I think we had launched paper search and summarization in Elicit. Over the course of maybe the next year or so, we launched systematic literature review—the fixed search, summarize, screen, extract, write flow—and now we have these much more flexible research agents. It's obviously been an incredible trend from very short tasks, short time horizons, and low flexibility to agents that do much more extensive research, with long time horizons and quite a lot of flexibility. So, I'm excited to review what this all means and where it's all going.

Nathan Labenz

Can you give me a double-click into what exactly you had in mind when you were talking about how you were going to improve the quality of reasoning, how that compares to more and more RL on top of a chain-of-thought paradigm that everybody is now very familiar with, and then how that has maybe enabled what you're doing with Elicit, maybe in some ways competed with it, and how it has shaped your business and product strategy to go in this more principled, a little less bitter-pill direction with what you're trying to build?

Andreas Stuhlmüller

I like “bitter pill.” Yeah, last time we talked, we talked a lot about process supervision, right? How do you know that a system is doing good research for you? You can either look at the output and say, “That looks good to me,” or you can look at the process it went through: What papers did it look at? Why did it look at them? How did it choose what to do?

Back in the day, we were young and naive, and we all thought it would be better to know that things were correct for the right reasons, that the process was good. So, we leaned a lot into that. In many ways, I think the justification for that has still borne out.

Let me talk about a quick anecdotal experiment we ran, I think, 2 days ago. We told a few research agents, “Look through and analyze about 100 papers on toxicology risk for a particular type of cancer drug.” We told that to Claude, to ChatGPT, and to Elicit, and then we asked, “How many papers did you actually analyze?”

As the models like to do, they say, “That's a fair and important question to ask. Let me be direct. I did not analyze 100 papers. You're right to push back. I didn't do it.” I think that is a failure of process in many ways, because you're saying, “I told you what to do. You didn't do it.”

The reason for that is that the models are not trained on process. The models are trained to produce outputs that look good. If you didn't check and just looked at the result that said, “Here's my analysis,” you wouldn't have caught that.

I think the fundamental problem still exists. And then the question is, “What do we do about it? And to what extent does the solution look like better checking of the outputs versus checking of the process in various ways?” I actually think this question is still open. I can speak for what Elicit does and maybe briefly for what the rest of the ecosystem does.

Elicit addresses this by—we have a little kind of domain-specific language that the research agent can write that orchestrates other calls to agents: screen all these papers, then extract data from all these papers. And it runs the process. You run the process, and you know the process does what it said it was going to do.

I think the model companies are going a little bit in that direction, too. I haven't been following it extremely closely, but I think Anthropic recently launched a workflows feature that has similar properties. And so, even though on some level, yeah, all the models are outcome-trained, and that's why we see these quite ridiculous artifacts of models being like, “Ah, sorry, I didn't do it,” I think the problem still exists, and one level up from that, people are trying to patch it. I think that's maybe where I see ourselves as being in this game.

Nathan Labenz

Yeah, I think it's worth emphasizing that a little bit more because it's a core part of how Elicit is built differently and, therefore, what you can use Elicit for.

Jungwon Byun

A lot of—like you said, it's been 2 years since we chatted—a lot has happened in that time. A big part of what we've been working on for the last year is just rebuilding Elicit on top of this much more agentic infrastructure. We started working on that in about March or April of 2025. In retrospect, it was maybe a little bit early, but at the time it felt quite late because obviously a lot of advances had happened over the last 2 or 3 quarters.

A big part of what we spent the last year thinking about was designing how we preserve these benefits of transparency and systematicity at scale without losing the flexibility and raw power of these models. I think that's the core design question we've always wrestled with, because when you want to deploy these models at scale for really high-stakes decisions, you need them to behave in a certain way, which is often contrary to their kind of fuzzy nature. But you don't want to be overly deterministic because then you run into the bitter lesson issue, right? So threading the needle is what we're always struggling with.

A lot of our time last year was spent on this kind of technical design question, and we decided to design our own programming language to solve this problem so that the models could run reasoning computation reliably and be able to call these reasoning primitives at scale in a more trustworthy way. What we were trying to accomplish for our end users was the ability to say, “You can run this process with the model over 10,000 objects—10,000 documents, 10,000 drugs, 10,000 targets, genes, whatever—and the same process will be applied to number 5 as number 9,999.”

There's just no other model that can make that guarantee. And unfortunately, there are lots of models that claim that they have done that or can do that and are just completely wrong about it. So when we think about who our users are, how do we build a differentiated product that really meets their needs, and what use cases can we enable, we are powering people who want to be able to rigorously synthesize evidence at a very large scale and get every single thing right, all the way to the nth degree.

That's a very different interaction than just riffing with a model. We support some of these lightweight use cases, too, but I think there it's almost more like getting to parity. And I think where Elicit's really differentiated is that trust at scale.

Nathan Labenz

Yeah, I should always remind myself and be clear that none of these things are really true binaries, in the sense that it would be wrong to characterize Elicit as not being about the bitter lesson. I remember that last time, one of your big principles was, “How can we allow you to spend more money to buy more compute to get better results?” And that's almost a restatement of the bitter lesson in some way.

At the same time, the frontier companies, as you said, are doing some of this, and I assume that in their research agents, particularly their deep research products, they are presumably doing some sort of at least rubric-based reward on the final reports that the models are outputting, albeit, if I understand correctly, still at least intending to avoid putting any optimization pressure on the chain of thought. There's always a little bit of gray to this.

So who are the customers now? There are these deep research things. I use those pretty frequently. Who is the sweet spot that is like, “A deep research agent isn't enough for me. I really want to go way bigger, way more systematic, and be very sure that I'm performing the same analysis in a way that I can count on”? Who are those customers now that you're finding product-market fit with?

Jungwon Byun

Yeah, I think of it as maybe a funnel. The largest group of users are academics and individuals, people like yourselves who value the kind of deep research function. They generally want a fast but robust synthesis of the literature or of the evidence.

Where Elicit really differentiates itself is that it's a bit more technical by default. It's not really assuming a lay audience as much. Most importantly, everything is well cited. You can get the models to cite things if you really push them to, and then sometimes they realize they were referencing sources that didn't exist, which is unfortunate. So they still make those mistakes.

Elicit doesn't do that. Elicit just assumes every single claim needs at least one or multiple citations from vetted databases and data sources in the academic literature. A lot of people still find value in just that kind of core offering.

Then there are academics and researchers who, like you said, want to get a lot more systematic. They want to do a systematic literature review, where they're really thoughtful about: What evidence should I be looking at? What information out there is relevant to my research? What should I be excluding? How do I apply the same process to all the papers? They want to extract detailed data from charts and figures, and then synthesize all that and weight the findings by quality, or do a landscaping of the market or the research in a similar way.

Increasingly, one of the other things we've really been developing over the last year is our life sciences motion and playbook. We work now formally with, I think, 7 of the top 20 life sciences companies, and we work pretty much across the entire development life cycle. I would say our biggest concentration is in early-stage research, so working with discovery biologists, some of those toxicologists that Andreas mentioned, as well as researchers in the late and post-clinical stages and commercial and medical teams.

The early-stage researchers often have more iterative processes. They're like, “I have this idea for an experiment. I'm in immunology. I want to figure out a way to repress these kinds of rogue T cells. What's the right mechanism to do that?” Then they get ideas, find other experiments that have been run, and try to reproduce those experiments. They also are often systematic, so they're the ones who want to apply a reasoning process over thousands of genes and targets and do a tournament-style ranking.

Then on the commercial and medical side, often you've had validated Phase 2 and Phase 3 trials. They're starting to think about what's the launch strategy of a drug. Exactly which populations do we go after? Who pays for this drug? What are the other kinds of alternatives out there? How much more compelling or cost-effective is our product?

They have to be very evidence-based, and every claim that they make in justifying to regulators or payers exactly why this drug makes a difference in the world has to be supported. So that's where, within life sciences, we're seeing a lot of pull.

Nathan Labenz

Something really top of mind for me that you kind of touched on there is being evidence-based, and I hadn't thought about the definition of that, especially at the early-stage literature-review stage, but especially as you get later, when you start to think about what markets and that sort of thing. It strikes me that perhaps in many cases the best information is not necessarily in an academic paper, or at least, if I was naively thinking about approaching such a problem, I would think I would want to cast a pretty wide net.

I also experienced this a little bit myself. Fortunately, my son, as regular listeners know, is pretty much cured of his cancer, and we're getting back to normal life. That's all great, and I didn't have to go down the path too deeply to really try to get a synthesis of what the best second-line treatment would be for him.

When I was contemplating that, in talking to doctors about it, I found there was a real narrow box drawn around what kind of information would be considered. Our oncologists, whom I hold in quite high regard and with whom I have a good relationship, will say, “There's no evidence for that.” This is for things where I'm like, “There's definitely some evidence, right? It's just not this sort of gold-standard RCT, peer-reviewed, yada yada yada.”

Obviously, that stuff is great where you can get it, but all this is a buildup to ask—and it may be different for different kinds of analyses—how do you think about how wide the net should be as you go out and look for evidence? When is it appropriate to truly limit yourself to published academic work? When is it appropriate to consider blog posts or even insightful tweets that somebody has put out? How do you guys think about that? Does it vary by use case? How do you think about weighing these things and calibrating to the level of credibility that they should have?

That seems to me an opportunity for, honestly, just major improvement, even at the clinical level, because there seems to be so much turning a blind eye to evidence today just because it's not right at the very highest level of quality or credibility.

Andreas Stuhlmüller

Yeah, that's right. I think the core question here is: How do you differentiate and discriminate between evidence based on evidence quality? And how do you identify when a certain level of evidence or quality is relevant or the best thing for your decision?

So even within published literature, there's a gradation, right? There are—

Nathan Labenz

Impact factor, remember?

Andreas Stuhlmüller

Yes. Yeah, there are higher-impact journals, there are well-regarded authors and institutions. This is a question that we've wrestled with for a long time, because the way humans have solved it is that there are these very lossy proxies: citation counts, journal impact factor, “I know the guy,” right? Those are rough approximations, but deeply imperfect, and there are so many examples of groundbreaking research that was not published in the top journal. One of the foundational CRISPR papers, for example, was actually published in a tier 2 or lower-tier journal.

I think the promise of language models is that, yes, we can continue to use these heuristics that have been helpful, but also we can think a little bit more from first principles about the quality of the work and its appropriateness for the decision at hand. So that's what Elicit has always really tried to do: not just blindly rely on citations, but give the researcher the chance to say, “Okay, for this research project, I'm going to look at case studies, or I'm not. Case studies are not good enough in my research domain because there are plenty of RCTs.” Or, “I'm going to look at studies with a sample size of 10, or I'm just going to look for a sample size of 1,000 or more,” because, depending on whether you're working in rare disease or oncology, the shape of the research that's available, like you said, is different, right?

I think one principle we have is: How do we look at the actual methodology and the quality and content of the research, not just at metadata, to determine quality? How do we equip the researcher to express their own judgment and expertise based on what they know about the domain and what is specific to their project? I do think it is specific to the domain.

As we've built out the research agent, we have started to bring in a lot of other data sources—web sources like company filings and things like that—because, like you said, there's a lot of information that's not just in published papers. Publication takes a long time, so a lot of recent things are not going to be there. Sometimes research is very interdisciplinary and holistic, so when you make a scientific decision, you want to consider the science, but you also want to consider the commercial implications of that as well, or the policy implications.

We have introduced a lot of data sources, and I think the agent still has a relatively fuzzy understanding of which sources are higher quality. I think we want to structure that a bit more. Again, the user has a lot of control over how they can express what they believe quality means, but I think we do want to also be a bit more opinionated for researchers who are not experts.

The last thing we're thinking about is: How do we express the confidence level of claims? Not just at a source level—say, “This is a high-quality paper, this is not; this is a high-quality blog post, this is not”—but when Elicit synthesizes all this information for you, how can we say, “This treatment works”? That claim: How well supported is it? Is it something you don't need to worry about, like 99% likely to be true, or are there conflicting results in the evidence?

We've been thinking a lot more about breaking insights down into claims and then giving people confidence levels so that they know what to do with that. I think that's another important part. It's not binary; it's: How can I use it? That's the part that's really important.

Nathan Labenz

How good are models at that today? I always think back to the original GPT-4 model card, where they showed the base-model calibration being pretty good: when asked to express confidence in claims, when it said 20% likely, it was actually pretty close to being correct 20% of the time. Honestly, to a pretty remarkable degree, the base model was reasonably well calibrated. I understand that pretty much just came out of pretraining as an emergent property.

But then, with reinforcement learning, you have mode collapse—which is, I'm going to overstate the importance of mode collapse—which I do think was one wave of AI denialism for a minute there. But there was some form of mode collapse where all of a sudden you're getting much less well-calibrated self-assessments from the model in terms of how confident it was that the claims were accurate.

Obviously, since then we've had a tremendous amount of additional RL applied on top of models. I haven't seen those kinds of calibration studies as recently. I could believe maybe they've gotten better because maybe that's part of a rubric that's being used. I could imagine it's even gotten worse because, just left to its own devices, RL maybe takes you in the other direction.

What do you guys see in terms of how good they are natively? Can you prompt your way to good performance on these sorts of confidence questions? Do you have to fine-tune? If they're not that great, do you want the frontier models to get better? Would you advise them on how they should think about getting better at that? Tell me everything.

Andreas Stuhlmüller

Yeah. So, first, I think token probabilities are just useless now, so we can pretty much ignore them. When we talk about calibration, mostly we have to talk about verbalized calibration, right? If you ask the model, “How confident are you in this thing? Are you confident?” I think verbalized calibration is probably much better now than the token-probability approach from GPT-3 to GPT-4. I would rather take the verbalized probabilities, to be honest. I think those will probably do a better job at capturing, for complex situations, what the assigned probability should be.

I still think the models right now are easy to push around, and that is one of the main things we're trying to address with better scaffolding. Right now, you often ask the model, “Hey, how likely is it that this clinical trial will fail?” Say the model is going to say something like 30%. Then you're like, “But have you considered that you can almost say anything here? Have you considered that, on average, clinical trials fail X% of the time?” And it'll be like, “Oh, no, you're right. Actually, this probably should have been higher.” Or, “Have you considered that, in this particular case, this molecule has pretty little research behind it?” And it'll be like, “Oh, yeah, you're right.”

The models are getting better at it, but I think fundamentally, right now, these probabilities are very unstable compared to when you talk to an expert. You can throw these things at them, and most of them will be like, “Yeah, you know, I've considered it.”

[Laughter.]

“That doesn't really change my view here.”

So, a big question is: How do you get stable probabilities out of the models? Which isn't to say they should never update. They should obviously update if you present truly novel information to them; they should update.

But right now, it does not feel like they have a coherent world model behind a lot of the probabilities that they express to you.

Nathan Labenz

How would you say they're easy to push around? How easy are they exactly? Are we talking about even seemingly semantically equivalent rephrasings producing big differences in outputs, or is a more conceptual difference in framing required? Or can you throw in a red-herring fact and move the needle? Just how much are they blowing in the wind?

Andreas Stuhlmüller

It's hard to characterize, and this is maybe part of the issue: it's hard to predict exactly what will push them around. I think if I could give a better answer to this question, I would also have an easier time fixing it. I don't think it's extremely systematic right now.

Jungwon Byun

But maybe that's partially what we've been trying to invest a lot more in. I think that's the next phase of our role as humans, and what humans have to do is work through these really gnarly evaluation problems. I very much believe that in the next few years, generation will more or less be a solved problem. Historically, humans have done all the generation. They've done all the work.

You open up a blank document, you put your thoughts in, you start from nothing, you create. And that's just not how work is going to happen anymore. Increasingly over time, AI is going to take the first pass at everything. Then the work left to humans will be around evaluating that. Was it the right thing to do? Was it done well? Can I use it for this use case? Can I trust it?

I think people are worried about job loss, but I do think there's a major job transition opportunity, or skill transition opportunity, where now what we need to do is think about what good looks like and what these failure modes are that AI systems can run into. How do we start documenting them? How do we start codifying best practices and getting good examples to point models at? How do we start building better verifiers?

There's actually a lot that needs to happen because we don't even know—we can't even articulate for ourselves what good looks like, why this is a bad thing, or how often this happens. So I feel like there's a lot of work around that and, certainly, we are trying to make that transition as a company. I think there needs to be more infrastructure built around evaluation and articulating what good looks like so that we can point these models toward it. How do you think…

Nathan Labenz

it interacts with chain of thought? This is very top of mind. I keep talking about it, but I attended this event in San Francisco called Recursive not too long ago, which was all about the potentially soon-coming phenomenon of recursive self-improvement. A scary, striking takeaway from that event and all the conversations and presentations there was that we're really heavily, and I would say problematically, indexed to chain-of-thought monitoring.

It's chain-of-thought monitoring all the way down in terms of the plan for how we're going to keep a recursive self-improvement process on the rails. There's a very strong sense that we've got to maintain freedom for the models in the chain of thought so that we can monitor it, because if we apply pressure, per the “Obfuscated Reward Hacking” paper, we'll drive the bad behavior underground. Then we'll be doubly worse off, because we'll still potentially get the bad behavior and we won't see that it's thinking about it.

How do we square that? I guess the naive answer would be that you can think whatever you want in the chain of thought, but I still want a systematic account—a systematic reasoning trace—in the actual final output, and then I just reward that. Maybe that works, but maybe you have a more nuanced view on how companies should think about where, what kind, and how much pressure to apply to the reasoning process that builds up to the final answers that models give.

Jungwon Byun

I can give a high-level take, and then I'm sure Andreas will have a more technical response. One of the projects we worked on at Ought, like many Ought projects, was a bit early for its time. We probably built the first language-model observability tool. We built this thing called ICE. It stood for Interactive Composition Explorer. We had this—when was this? This must have been 2021 or so.

We anticipated this problem even 5 years ago: that at some point, language-model traces would get so large that we would not be able to debug them for ourselves. How would we visualize that? How would we maintain oversight? How might we audit that? And so we built this visualization.

One of the things we learned was that the best way to troubleshoot is not necessarily to go chronologically through all the steps that the model took. That actually means evaluation: repeating the generation process is one way to check what happened and build trust in it, but it's inefficient. It doesn't really scale, and it's a very difficult way to check.

What you often need is a different layer of reasoning checks, almost like logical consistency checks. These are not, “Let me go through everything you did—steps 1, 2, 3, 4, 5—and see if it was correct.” Instead, it's, “Let me think about, for example, a sensitivity analysis: How sensitive are my findings to different changes in input parameters?” Logical consistency checks, things like that.

That's where I think we need a lot more investment in infrastructure and building independent checks that don't rely just on chain-of-thought monitoring. That's my high-level take, but I'm sure you have a more technical version of that answer.

Andreas Stuhlmüller

Yeah, let me first restate part of what you said in different language. I think you can either check the process or you can check the outcome. I mean, those are your 2 options, right? When you're checking the outcome, you still want the outcome to somehow contain a certificate that the right reasoning was done.

What can that certificate be? It can be, “Here's how my conclusion would change if I had a different input. Here are the things I looked at on the way. Here are citations to the literature.” I do think that this is a very underdeveloped field, in my mind. In mathematics, it's very developed, right? You can have your formal proof if you want, and that proof is checkable.

I don't think it's very developed in more fuzzy domains, probably partially because it would just be too much work for humans to produce legible certificates. Humans don't even really have a great ability to introspect on their own thoughts. But I think in principle, even if you didn't supervise the process, you could produce certificates of reasoning that then let you check the reasoning, even if you didn't check the process that generated that outcome. I'd be very excited to see more work in that direction.

The other clarification I wanted to make is that I think it's worth distinguishing chain of thought from the reasoning process, or the chain of process, or whatever you want to call it. When people talk about chain of thought, they often think about what thoughts the model writes down—the reasoning tokens—and then ask, how much can you trust them?

I don't know. I think for OpenAI, we don't even get them anymore these days, unless maybe you apply for a special permit to see them. But you do see the tool calls, and I think the tool calls actually are important reasoning facts in and of themselves.

Take, for example, let's say I download a new paper from arXiv. The model hasn't seen it before, and I ask it some question: “Hey, what are the key results here?” Now I can see which parts of the paper the model is reading, because it has a read tool that maybe reads, by default, the first few lines, and then it can scan other parts of the paper.

Sometimes I know the model didn't even look at the methodology section. That is part of its reasoning that is checkable. I know if it now tells me something about the conclusions of the paper that really should have relied on the methodology section, I know it didn't do that. That's obviously a very simple example—the papers are small and so on—but the same thing applies at a much larger scale, where the process still remains an important way to check, because you do see the tool calls, and the tool calls are an important input into the model's reasoning.

Nathan Labenz

Yeah, I like that. How much of this do you think companies are doing today? It's got to be some, but obviously they're not telling us. Is your view on the spectrum from closing in on the sort of thinking that you're doing, and the granularity of process supervision that you would like to see, on the one end, to purely RLVR—with “Did you get the right answer in the final answer box or not?” as just a binary signal—on the other end? How much of this do you think they are doing, based on what you're seeing in model behavior?

Andreas Stuhlmüller

These are obviously speculations. I think they do a lot. I think we know they do a lot of rubrics on the final output, like, was the—even reasoning-adjacent rubrics: Did the model produce something that looked like expert reasoning in many ways? I don't think they do a lot of evaluation of the process.

Again, I could be wrong, but I think there's a really interesting question, which is: If you had all the details of the process—let's say you are inside a lab and you ask yourself, given X and T, would you expect this process to lead to the right answer? Not after the fact, if you check it against the answer—was it correct? But did it follow the sort of process that, if you're making a forecast, for example, and you don't see—let's say you don't even see what forecast the model ends up…

You just look at what things it considered, what data sources it considered, and what it wrote about the different hypotheses it considers. I think there's a really interesting project of thinking about to what extent the models are following a process you should expect to be good. In my limited knowledge of what's going on, I don't think much work is being invested in that.

Nathan Labenz

So, I guess one implication of that for your work would be that you might expect the next generations of the model to eat less of your scaffolding than they might eat of other types of scaffolding, right? This kind of general pattern—people build out scaffolding to compensate for the model's weaknesses, then the model companies take that feedback and train on it, and in the next generation, you can clean out or eliminate a decent amount of that scaffolding. It seems like you think that maybe on the agentic side of “go accomplish this project, get over these humps,” whatever, you would expect more of that kind of scaffolding to be eaten in the next version than the sort of trusted-reasoning scaffolding that you're building.

Yeah, that reminds me of something that you recently wrote around Elicit's focus being on becoming the best at reducing these big, fuzzy, hard-to-verify tasks to sets or graphs, perhaps, of easy-to-verify tasks. I'd love to hear a little bit more—you alluded to it there—about how you're going about that and also how complete you think that process can be.

Everywhere I look these days, I feel like I see the same shape of a really interesting question that I don't know what to make of, which is basically: How do we get high-level guarantees, conclusions, or insights from low-level steps? In biology, I might be able to say I've got all these proteins, or these genes are being expressed at this level, but I don't know what's going to happen next at the cell, tissue, or organism level. No, right? I don't.

Similarly, with these big judgment calls—should I prioritize this drug or that drug?—we can break it down and become systematic, but it's not clear to me how close it gets to something where I'm like, “Yes, okay, I can really buy in and trust that.” Versus, is there something that sits above all those steps still? Is it emergent, or is it just somehow lost? Are people doing some sort of metacognitive work that's hard to capture but maybe still very critical to actually being effective at these tasks?

So, I guess, again, tell me everything. I'm really struggling with this. Formal methods and formal verification are another area where I see this: We can make all these low-level guarantees that this error or that error can't happen, but is the system itself going to behave how we wanted it to? I still don't know, in a lot of cases, how we make that leap.

I'm very interested in your take on these sorts of questions. I usually think of it as laddering up low-level things to high-level conclusions, but you're approaching it from the other direction, which is interesting unto itself. Yeah, again, tell me everything.

Andreas Stuhlmüller

Yeah, maybe I should first restate this: Why do we want to reduce hard-to-verify tasks to easy-to-verify tasks? It's because AI currently can be trained on easy-to-verify tasks. We know it's extremely good at RLVR coding and at math tasks like this, and it's quite weak at a lot of fuzzy tasks.

I notice it all the time when I try to use the models to help me plan our company strategy, for example. I think they're surprisingly useless. Even though they have access to all my context, they're really quite good at saying, “Let me pull in the data. Let me pull in your email and your Slack.” I still find that they don't get it.

In an important way, this is related to what we said earlier about how they're too easy to push around. It doesn't feel like they're building up a coherent model of what's going on. I think an important reason for that is that it's a hard-to-check task.

So, what do you do? I think it depends a little bit on what your situation is, whether you're trying to create a reward signal that you can train the models on, or whether you're trying to do verification and checking for the purpose of understanding whether an already-trained model can be trusted in a situation or how to refine its behavior.

I think if you're trying to create a reward signal, that's pretty rough because the models are going to optimize pretty hard against your signal. It's not enough to do spot checks and be like, “Hey, here are some cases where we can verify that, for example, your company strategy was incompatible with some claim you made earlier.”

Whereas, if the goal is to take an already-trained model, understand how good it is exactly, find fairly obvious mistakes, and identify places where it can improve, then your reward signal doesn't need to be bulletproof. Your way of getting some easy-to-verify aspects of the hard-to-verify task can be more stepwise, so you can make more incremental progress, I would think.

Our situation is that we're not currently trying to train a foundation model from scratch or even post-train a model on this particular aspect. Our situation is more like, how do we get to the point where we can check many important properties of tasks? Are the claims the model makes internally consistent? If you break it down in different ways, will it end up at the same conclusions, and so on?

I think that's a fairly tractable project. The project of figuring out how to fully reduce high-level tasks like company strategy into individual components that are all formally verifiable is a much rougher prospect. That's not to say it's impossible, but there's less incremental feedback signal that you're on the right track there, I would say.

Nathan Labenz

So, tell me a little bit more about what you're doing in practice. You said you're not trying to post-train a model. I know in the past there was a decent amount of fine-tuning, at least for specific tasks, so I'm curious if there's still a fine-tuning element to it.

There are a lot of different ways you can think about spending a lot of tokens to try to get at this. You could run the decomposition process multiple times and check for consistency, which I think you're suggesting something like that might be going on. You could do a more iterative thing where you get the AI to give you an output and then have some kind of specialist prompts, or perhaps even specialist fine-tuned models, come in and assess it in various ways, give it feedback, and then let it reason some more and try to improve on what it just did.

We do see some of that stuff. I just talked to some OpenAI forward-deployed engineers who are basically using that process to improve filing accuracy, and that seems to be going quite well for them. What techniques are you finding to be most effective in practice today?

Andreas Stuhlmüller

So, first, we still do a bit of fine-tuning. At this point, I think it's more of a technique to make things have reasonable efficiency properties at scale than something to get the models to exhibit new behaviors that you couldn't otherwise elicit.

That said, we do a lot of the things that you pointed out. Let me maybe talk about one of them that we've been investing more in lately, which is what you could call world models, or knowledge representations, that make the model's work more checkable.

What's the motivation? The motivation is actually maybe similar to the kind of medical case you had. I had a friend who also had cancer, and it was a case where I then used Elicit to get a lot of the raw data for it.

I ran the systematic literature review flow for a few versions of the question, “How do you address this particular type of cancer?” I ended up, even after filtering down all the information, with a ton of papers. After filtering for just the most relevant papers, it was still maybe 5,000 papers or so.

The question is, what do you do with that? You could try to somehow throw it all into a million-context window, but I don't think it would actually work that well, and I don't think the model would be that good at coherently reasoning about it.

And so then the question is, what else can you do? I think one thing that people have tried—I don't know if you're familiar with this—is what Karpathy termed an “LLM Wiki,” I think, or something like that. You build this knowledge base as an Obsidian-like knowledge base, or a folder of Markdown files, where you tell your model, “Hey, we're researching this topic. Organize the information in a way that makes sense and do these iterative updates to it.” Maybe it's a GitHub repo, and the model gets to add new notes to it, move information from one file to another, and propagate information.

I think that is a really interesting direction because, if you think about how we get to models that currently answer complex questions, either it happens in the weights of the model or it happens in some explicit representation. What is the explicit representation? Text files are appealing; it's a nice start. But then you ask yourself: That is also very flexible. What properties do you want this representation to have such that it actually helps with the research questions you have?

Once you think about that, you're like, “Well, I want it to let me make predictions about what's going to happen in my case. I want it to help me think about interventions. If I did this—if I took this particular drug or pursued this particular type of chemotherapy over different time periods—what would happen? If I did this instead, or followed it with some immunotherapy, what would happen? If I had done this different thing in the past, what would have happened?”

Those are questions about predictions, counterfactuals, and interventions. Very soon, you're like, “Wait a minute. That is a thing people have been studying in the past. It sounds a lot like graphical models or structural probabilistic models.” The question that we've been thinking about internally is: How do you get the best of both worlds? You want the flexibility of language models that can reason about and transform these representations, but you also want to be able to answer these sorts of prediction, intervention, and counterfactual questions that you care about.

You want the answers to these questions to be internally consistent, such that it's not the case that one answer and another answer just don't make sense relative to some underlying coherent representation. All this is to say, we're thinking about how you can build what I'm currently calling a world model, although I'm not sure that's the best name for it. How can you build these representations so that they let you answer the classes of questions I just talked about in a way that is internally coherent, and evolve them over time as you add papers one by one—adding information about different symptoms, different treatment strategies, and so on—and get them to be internally coherent?

That's one big direction we're thinking about for how to spend really enormous amounts of compute to deal with large amounts of data and turn them into something that can actually answer real questions for projects that are much larger than a single query.

Nathan Labenz

Yeah, that's cool. I am very familiar with the Karpathy Wiki line of thinking, and I have used it myself for creating a little wiki that my personal agents use to navigate my life, basically, on top of all the raw data. I first did monthly summaries, where I exported everything and, on a month-by-month basis—it was a few hundred thousand tokens a month—I would condense that to a monthly summary. Then there was a layer condensed to an annual summary, and then I came across that. I was like, “Okay, let's make a wiki.”

We've got articles about all these fun things that I'm starting to do, but it's a party trick, I guess. One I just did last night was tell a friend I hadn't seen in a minute, “Oh, there's an article about you in my personal wiki. I'll have my agent send it to you, and you can read it. I haven't read it, actually, so I don't know exactly what it says. But you can read it and tell me if you see any hallucinations or anything that you would object to in there.”

For me, though, it's basically just trying to get down pretty black-and-white facts and help the model navigate those. The nodes are pretty obvious, and the edges are also just simple links: This person I know from this organization, and so on. I haven't pushed it nearly as far into ideas or a structured decision-making aid, so I'm curious about how that looks.

I guess the other image that's coming to mind as I'm trying to conceive of this a little bit better is the Anthropic work on tracing large language model thoughts, where they have these sorts of graphs from tokens to outputs. I'm always one to remember that the graphs they present are somewhat clean, but there's a huge residual term all over the place on those graphs, which again recalls my earlier question: How much can you decompose, and how much residual is there?

But tell me a little bit more about what these world models end up looking like in terms of edges. Is “X causes Y with X percent probabilities attached to it” one good way to start? Or how do you think about describing those relationships in a way that balances capturing as much of the structure as you can, while also recognizing that there is going to be some residual? I'm really interested in how you're navigating that.

Andreas Stuhlmüller

Yeah. We are trying not to be very prescriptive about it, to the extent that we can. I think there are cases where nodes and arrows are the right—or at least a helpful—representation. In the cancer example, if you're trying to understand the mechanism, I think it's often useful to say, “Well, the antibody and antigen need to bind, and then once they bind in the cell, a particular type of substance needs to be released in a particular way.” There is just a sequence of events that needs to happen, and in those cases we want the model to build these sorts of graph-like representations.

But I don't think that's always the case. If you think about company planning—what products should Elicit ship over the next few quarters, and how does that affect our revenue and user numbers and so on—I don't actually think I want that to be represented as nodes and arrows. I think it's more like having a spreadsheet with different features, and then user numbers, margins, and so on over time. That's a very different type of model.

Often in the real world, no single model captures what's going on in its entirety. One lens to look at what Elicit is doing is as a spreadsheet of user numbers over time. But another lens might be the tech tree of the Elicit product and how we build it up over time. Those are complementary. Ideally, they can both live in your knowledge wiki, and you can say, “Hey, language model, as you're trying to make predictions or help us evaluate plans, look at all of these representations.”

The challenge is how you get it to be the case that the model knows how these different representations relate. When it makes a prediction, you don't want it to be the case that sometimes it looks at the spreadsheet and says, “Well, it's going to look like this,” and another time it looks at your tech tree and says, “Oh, okay, we're going to do this.” You don't want them to be 2 totally separate things. Some propagation of information needs to happen between those different representations.

It's an ongoing research project. I definitely don't want to claim that we have solved this problem. It's an ongoing research project within Elicit and, hopefully, eventually within the world at large: How do you make these more explicit, legible, fairly heterogeneous representations of knowledge that models can work on over time and improve?

I guess one last thought here is that maybe one way to think about it is: How do you make progress on continual learning in a way where stuff doesn't just live in the weights of the language model, but is available to humans as a representation we can inspect and understand?

Nathan Labenz

Have you found any particular data structures, particularly if they're open source and something I could also incorporate into my own personal AI infrastructure, that work well? I'm just doing a very simple Markdown wiki as of now. Is there a next level that I should be considering? A graph database? I have no idea what it would be, but I'm always looking to upgrade.

Andreas Stuhlmüller

Step 1 is just to use the representations people already find useful, and SQL databases are a pretty useful thing. I'm currently building my model of Elicit, the company, and Step 1 is to ingest a lot of the information from Mixpanel, Adyen, and various other types of systems and put it into a representation that the model can then operate on. A lot of those representations are just SQL tables.

Nathan Labenz

Cool. Interesting. Everybody's doing their own experiments in recursive self-improvement these days. That's what I'm noticing across the board. It's not always the case, but it's been striking that in the last 6 to 8 weeks, it seems like everybody's tipping into this moment of, “Maybe we, too, can be an experiment in recursive self-improvement.”

So, you're now Elicit as a test case for: Can we get Elicit to build effective world models and then be able to use these detailed representations that it itself has constructed to inform its own analysis of what it itself should become in the future? The hall of mirrors there is deep and fascinating.

This obviously relates to a blog post that you put out not too long ago called “Planning Is Unsolved.” And I think, yeah, obviously that's totally true—or we could just ask the AIs to handle all this and retire to the beach, as I think it was an Anthropic person who once famously put it.

I do wonder a little bit, and this is probably a cultural question as much as it is a technology question. When I think about my own company, which I'm now just the AI advisor to and not running, and I think about the planning process that we developed before AI, sometimes I'm like, maybe we should just scrap the whole thing. Maybe all these times that we come together, sit around the table, talk about this or that, and try to convince each other that if we do this, it'll be more successful than if we do that—I'm often like, man, you know what we should do? Build it all, launch all these things, and see what happens.

Coding has gotten cheap. So maybe the future of planning is less about guessing and more about fast iteration and actually making more contact with reality. Obviously, again, there's not a true, strict binary there. But how are you guys thinking about that question? And how are your clients, especially pharma companies, thinking about that question?

There's also this notion of clinical trial abundance. I think it's more of an aspiration than a trend at this point. I don't think we're there, but you can imagine a very different vision for a pharmaceutical company. One is, “We're going to use these world models, make much better decisions, and deploy these scarce resources in the precious few at-bats we have at clinical trials in the best way possible.”

Then there's this other vision: “We'll do 10 times as many clinical trials, and that'll be the bigger unlock because we'll actually get the real answers in far more cases.” Or 100 times—who knows? Where do you want to be on that spectrum? Where do you think pharmaceutical companies should be on that spectrum? Maybe it's different. Maybe Waymark should be in one place and pharma should be somewhere else.

Andreas Stuhlmüller

Yeah, I think it's great that the cost of software engineering has come down, but it feels like that was just one of the bottlenecks, and the others haven't moved. I don't think user attention or feedback is infinite, so I still feel cautious about just throwing things out there, giving people bad experiences, or leaving people with bad impressions.

I also think that it's definitely possible to get stuck on a local maximum. Let's say you just ship a feature and you're like, “Oh, cool. This works. Let's just keep going.” You might be able to keep going for some time, but then still end up maintaining, improving, and investing in something that wasn't the best possible thing you could have done.

So, I think there's still a lot of room for judgment and purpose. And there are places where I'm not sure the shape of the problem has changed that much. I feel like there's always an explore-exploit trade-off, and you want to navigate those 2 thoughtfully. Sometimes you want to do one, and sometimes you want to do the other.

Maybe the cost of exploring has gone down a little bit, but I think there are still some things where you have to get it right the first time. It's not as if all software engineering has literally gone to 0. There are still large software engineering projects.

So, I'm not sure it's changed that much. I feel like, for now, it's mostly changed things at the margins, like fixing bugs, small admin features, and things that we don't see as our core capability, that we want to fully automate and are happy to take liberal experiments with. Things that we see as being core to our mission, purpose, and differentiator as a product still involve a lot of careful thinking and intent.

With our customers, I find that many of them are just like, “Yes, there's a lot of excitement to build.” I guess my unsurprising take on the build-versus-buy problem is that you should build internally the things that are your core comparative advantage, and basically nothing else.

If you're building something, for example, there are certain workflows that are just regulated across the industry. Every company has to do them pretty much the exact same way, and they are very particular. As a result, they're very interface-heavy. I just don't think it's the core competency of pharma companies to design nice software, and that's not your comparative advantage as a company. Your comparative advantage is not going to come from solving this regulatory compliance problem.

So, I don't think, for example, systematic review, which Elicit solves, is one of those areas where it makes sense for pharma companies to try to build this thing that every company has to do the same way and that is actually very involved. Other types of things, especially in early-stage research or even in development—certain kinds of predictive models—make sense for a pharma company to build in-house.

And then, on the clinical trial point, I think, why not have both? There's a lot of interest in digital twins and simulating trial effects digitally as much as possible to design the trial well. I'm sure there are certain trials where, with the right regulatory framework and operational improvements, we might be able to take a lot more bets.

I think especially in rare diseases, where a trial is actually almost like a treatment option, we can be much more flexible there. In other domains, depending on the type of drug and what we already know about its toxicity profile, we'll probably want to hold a higher bar.

So, again, my hope is that it's great to have multiple tools, options, and choices. Maybe we can just build more tools, make the tools better, and then build a good framework for when you reach for what tool.

Jungwon Byun

I think there are many cases where you just can't do everything at once. If you think about clinical trials from the participant perspective, usually you can participate in 1 or at most 2. Often there are many that you could participate in, but you still need to choose which of those you go with. That's a tough planning problem.

Likewise, as a company, you have a certain amount of resources, and they can be deployed one way or the other. I still find, especially as a company that is trying to be extremely mission-focused, that you need to ask: How do we actually, in the short time that we have, make an impact on the quality of reasoning and the impact of AI on reasoning quality?

By default, you're probably just not going to accomplish that. Even if you honed in on user metrics, I think you're probably just not going to accomplish that. So, you need to think pretty carefully about how to get that to work.

Nathan Labenz

Your callout of the patient perspective is a really useful reframing there, because it's one thing to say, “Yeah, as a pharmaceutical company, maybe we can have both. We can run all the trials.” Certainly, as a SaaS company, I can potentially launch all the features, whether or not I should. I maybe could.

But if you've got cancer, you've only got 1 body, and you can't take all the drugs, right? That would obviously be ill-advised. So that, I do think, is a really useful point. At least for now, until the transhuman uploading future, whatever, appears, we are going to continue to face these very stark choices about what I should do with my own individual body as a human, knowing that there's not a second copy of it and there's not really the ability to diversify across paths in many cases.

Are there any examples of that you could share at the company level, especially if there's something where this sort of company model and recursive self-improvement paradigm of working has led, at least in your counterfactual analysis, to a different approach than you might have taken if you were just gathering around the table and sharing your intuitions, like people used to do?

Jungwon Byun

Examples of where automation, learning through doing, and cheaper experimentation have led us to a different outcome.

Nathan Labenz

Or maybe it's the same thing, but I'm thinking especially around planning with a world model. Is there something where, because you had structured these various lenses and were able to be more systematic and structured, you could really agree that this is the framework that you are using—maybe avoiding talking past each other or having some new synthesis, some new insight that you just don't think would have happened in the absence of that structured approach?

Jungwon Byun

This is a partial example. I'm sure Andreas has a better one, but it's actually very timely. I just did this for an important hiring decision.

Ahead of time, we had a rubric designed for what role we wanted to fill, and the role had been through a few different evolutions. We had looked at different personas. There was maybe a disagreement on exactly what type of executive-level hire we needed. There was some disagreement on exactly what we needed, and maybe that changed over time as the company grew during the course of the search.

And then, about a month ago, I wrote down a framework, and then we did extensive interviews, so many references, lots of back channels. There was just so much information I was getting, and I was starting to develop a take on what we should make of it and what we should do with this candidate, but I really wanted to avoid recency bias, so I had a structured project. In this case, I used Claude, not Elicit. Elicit doesn't support hiring decisions exactly yet, but I don't use it for that case.

I created a project where I had all of the meeting notes across all the interviews, all the email threads with this candidate, and all the feedback submitted, and asked Claude systematically to fill out examples of evidence for every single dimension. It ended up being about 20 different fields, like evidence that this person has consistently hit their goals, evidence that this person can hire a great team, and evidence that this person is authentic or culturally aligned. I started with the evidence and went piece by piece because I don't trust Claude to fully execute in one go. There was a bit of calibration, and then once I had the evidence—which was, “Hit quota so many years,” and so on—I asked, “What is evidence for that? What decision do I make? How do I rate this on a scale of 5?”

Then all together, I had this synthesized point of view, and I was able to send it to the candidate. I think they really appreciated it as well because they said it was the greatest kind of comprehensive synthesis of professional validation they had ever received. That was one case where compositional, structured reasoning with an intentional process, and then applying that at scale with AI, was able to both check my decision-making process and give someone the gift of something very human and very detailed about them and everything they had accomplished.

Andreas Stuhlmüller

My example of what I've done is going to be much more mundane. I think for me, I've been trying to do this more for just planning my week. I have goals, and this is what I want to accomplish in the long run—this year, this month—and then the question is: I have all these calendar blocks, like this podcast block, and I need to figure out that there are many different things I could do. What should I do, and which things depend on which other things?

It's actually a pretty tricky problem to know, when you could be spending your time in many different ways, what is worth doing. So I've been trying to get to the point where I can use automation as part of my weekly planning and think in a more structured way about this: If I want to accomplish my monthly goal, where do I need to be this week? How much time is that going to take? Is it maybe going to take 5 hours to write a blog post? When can those 5 hours happen? I think this sort of backwards chaining—people do it informally—but there is a lot of constraint satisfaction and propagation of constraints that is pretty tricky for humans, and I think the models will help us with that.

Nathan Labenz

Yeah, that's cool. I think of that as building your own harness in a way, which is something I'm thinking about for myself, too. How can I build up structures around me to keep steering me in the right direction, feeding me the information I need, and hopefully helping me become my best self, use my time as well as I possibly can, by setting me up for success as much as AIs can do that?

Do you find that you are following it? How good is it? Are you actually living by it yet, or is it still, “Maybe next week, when it gets a little better, I'll actually do the plan that it gives me”?

Andreas Stuhlmüller

It's still a very human-in-the-loop process. I actually have 2 versions of it. I have 1 automated version, which I hate, and 1 interactive version that walks me through the planning. I still do the fully automated one just to see how good it is, and I want to know whether, at some point, I'll be like, “Oh, yeah, I'm not needed here anymore.”

But generally, I'm like, “You just didn't fully understand what I'm trying to do. You made it too complicated,” and so on. So I'm still in the outer loop, but I think it's kind of interesting to think about it. Right now, humans are the outer loop, and they use calls to LLMs, but eventually an LLM calls you, and the LLM is the outer loop and you're just the inner loop. I'm not sure that's a positive future, but it seems like it's part of the trend here.

The Line is our automated software engineering project. I think, as is maybe the case for many companies, software engineering is where we've had the greatest success doing quite extensive automation. Our overall company goal is that, at the end of the year, when we go on vacation, we want the company to keep running and keep doing work in all of its functions. The first half of the year will be mostly focused on trying to make that happen for software engineering.

We have a system called The Line because it is like a factory line. Someone mentions a feature they would like to have on Slack, or a user mentions a feature, and we Slack-emoji-react to it with a little Line emoji. Or there's an integration with our customer-support system, and then it kicks off an iterative process: First, the feature needs to get specced out, then you need to iterate on the spec, it needs to be implemented, a video needs to be recorded of the feature being tested, then a code review needs to happen, and then it needs to get merged into dev and then into prod.

We do have a fully automated version of this now. For simple features, basically you just emoji-react to, “Oh, I would like it if Elicit kind of talked about its citations in a slightly different way,” and it will go through this entire process automatically. At the end, there are various judgment calls it makes about where human intervention is needed. Maybe the spec was too incomplete, and so it's like, “Okay, we need to pull in a human here,” or maybe the feature is too complex for the system to automatically review, and we need to pull in a human here.

But for many simple features, it can actually flow fully automatically through The Line, and I think that's already been a significant unlock for a lot of simple bug fixes and features. I think it's also setting us up for the future where, as each of these individual parts of The Line improves, more and more of software engineering will be automated. It's been pretty cool to see how we're now merging maybe 30 to 50 issues per week fully automatically.

Nathan Labenz

Cool, that's really interesting. If indeed we are successful in having the company continue to function through the holidays without you guys for 2 weeks, then one wonders how long it could go and whether you ever have to come back for one thing. But what will take us there? Is it just the next generation of models? Mythos is supposedly coming soon to a public API near you, with supposedly much better long-horizon performance. Is that going to be the biggest unlock? You've got the structure, and I just need to drop in a better model, or what else do you think is going to be needed over the next 6 months to actually realize that?

Andreas Stuhlmüller

Yeah, so first, I don't actually expect the company will run fully automatically by the end of the year. I expect our lower bar is that, within each function, there are pretty autonomous workflows that run and connect to some workflows in other functions. But I think a lot of the high-level steering will still be very much needed.

What is needed for scale-up? I think 1 big obstacle right now is that the models are not fully calibrated about when human intervention is needed. You have to be pretty risk-averse in how you use them. I think with software engineering, if 80% of the time when the model says, “This is an automatically reviewable feature,” it actually is, then that's not good enough, because we don't want to break production 20% of the time. That's pretty rough.

So we have to err much more on the side of: If in doubt, it's not an automatically reviewable feature. I think that's the case in software engineering, and I expect it's the case in other situations, too. If you were to let the models drive some customer interaction, for example, you probably want to be at least as sure as in the engineering case.

It's actually not clear to me how this will go. If you're following the METR graph, there's the 50% success-rate curve that goes up over time as the model gets better, and then there's the 80% success-rate curve. The 50% success rate is much higher, obviously, than the 80%, and the 80% hasn't been going up quite as fast as we would like. Often we want more than 80%.

Depending on how average-case performance compares to, I don't know, 95th-percentile performance, just dropping in the next models might be good enough or might not, but I wouldn't automatically rely on it.

It could be cool if, similar to fast mode, there were an ultra-reliable mode or something, which isn't just “think more,” but has guarantees on certain classes of errors that you're never going to make.

Nathan Labenz

Yeah, that's cool. A lot of really interesting thinking there. One big question that is generating quite different takes at the moment is: Are people going to be able and willing to pay the exponentially rising token bills that the industry as a whole is currently seeing?

You could analyze this from any number of ways. One would be your own internal work, right? Where is your token budget compared to your head-count budget today in engineering? And do you expect that, with the introduction of models—if they get, let's say, a lot better, for some definition of “a lot better”—you will shift that budget and spend a lot more on tokens relative to humans than you do today?

And then do you think your customers will do that as well? Maybe it will break down by use case. It sure seems like, for as much as we hear a lot of complaining about token costs and, anecdotally, “Oh, this company pulled back” or “That company hit budget,” it still feels to me like there's a lot of value in the marginal intelligence and just getting better results.

Tokens are still pretty cheap. I mean, for most companies, it's still a small amount. I hear things like 5% or 10% of what we're spending on headcount, and that's not that much. Maybe you didn't budget for it, and that creates some discomfort in your organization, but on the fundamental economics, it feels to me like if you can just get a lot better work for somewhat—even maybe a multiple—of the token cost, it still seems pretty rational to pay it.

I guess that's my starting position. What do you guys think you will do? What do you think your customers will do?

Jungwon Byun

I think that, often for our customers, at least, the offering we're providing is displacing services spend. So, the barrier is more about whether it can fully displace the services spend, and then also maybe getting over the mental hurdle of price anchoring for software.

Certainly, in terms of the dollars allocated to solving this problem, there are a lot more dollars in compute costs at the moment. So, I don't feel—yeah, we'll see. There are obviously human issues to overcome there, but I think from a dollars-and-cents perspective, there's still a lot of room.

I've heard and seen mixed things in the news, and I think the industry—at least the pharmaceutical industry—is fairly disciplined about costs and ROI. So, even if there's an initial period of heavy exploration, I think there's a lot of accountability around what that's delivering for the business. I think we'll continue to see that.

Andreas Stuhlmüller

I think even internally at Elicit, I'm not sure how many more multiples of token costs we can easily spend. So, maybe taking myself as an example, I spend around $2,000 per week on tokens. I could maybe double it or triple it, I don't know, but not much more than that, for sure.

So, I do expect—and that is already influencing my behavior to some extent—that I don't actually currently use fast mode for these models, because I don't feel the marginal returns are high enough for most tasks. I expect it's unlikely that I'll be like, “Wow, I need to switch over everything I do to Mythos.” It's probably not going to happen.

Both for my own usage and also because this is already actually the case in the Elicit app, I think more of what it will look like is one smart orchestrator agent that then spins off many other agents that have to do simpler tasks that just don't need to use the largest model.

I expect that will just become increasingly important: this sort of dispatching to a model of the right size, so that you get the intelligence when you need it, but you're not just multiplying your whole spend by some number that was an inefficient use of compute to begin with.

Nathan Labenz

So, $2,000 is not a small amount. Is that an outlier? Does that make you an outlier within the company, or is everybody doing that? If so, that would put your token costs—not presumably at the level of payroll, because I assume you're paying your engineers more than that—but it would be at least a not-insignificant share. Certainly, if you were to triple it from there, you'd be getting into something on the order of magnitude of parity with human headcount.

What are you doing with it all, too? I use my $200 Claude Max and my Codex Pro, and I honestly don't even hit my limits that often. Now, this may be API, which might be 10 times more, and so that could be a big part of it, but I sometimes feel a little ashamed that I'm not redlining the account more than I am.

What would you advise me to do, or what sort of personal bitter lessons have you learned where you're really finding that token-maxing is worth it?

Andreas Stuhlmüller

Yeah, I'm not sure I'm the top user of tokens at Elicit, but I'm probably at least in the top 5, so I'm probably a little bit of an outlier. Second, I'm using the API. I could probably save more money by being more clever about how to use various Pro accounts and stuff.

I do have a fairly elaborate system built on Pi that orchestrates between the different agents and uses ChatGPT to double-check, then sometimes calls Claude or Gemini to get another take. That is a little bit easier to do if you're on the API than if you're on the normal end-user plans. That might just be part of the explanation here.

Nathan Labenz

Any particular use cases where you get particular value that you think other people might be sleeping on?

Andreas Stuhlmüller

I don't know what other people are doing. As mentioned earlier, I get a lot of use out of planning, keeping my calendar in sync with my personal journaling system, keeping that in sync with my to-dos, and making sure everything is coherent with my longer-range planning document.

When the new day starts, I go over the last day, check whether there are any leftover tasks, and move them into the right place. So, there's a lot of automation happening behind the scenes without me prompting it that probably contributes to those costs being higher.

Similarly, for email, I have a pretty elaborate spec on which emails should be auto-archived. Maybe every hour or so, my models check that and go, “Okay, let's just archive the emails that Andreas definitely doesn't need to read.”

Then, on the more user-driven side, I do a lot of cross-checking. I sometimes think of things we could talk about with Nathan, and then ask ChatGPT and Gemini to double-check those things. I find that having the models cross-check one another often improves the results quite a bit, so for—I don't know—a fourth of my use cases, that already doubles or triples the cost. That's another source of additional token spend.

Nathan Labenz

Got it. Okay, cool. One other question I'd love to get your take on is: Are we seeing convergence or divergence in models? One notable feature of Elicit today is that there's no model picker, at least from what I've explored recently.

So, you're making choices, and it seems like you clearly think you know best. It would not be a good idea, even if people have a favorite model, given all the validation and scaffolding that you have, to just go in and swap models in and out.

How do you see this dynamic shaping up? Again, there are just such different takes: The models are commodities; scaffolding is all that matters. No, the models are everything; scaffolding is a complement. They're converging; they're diverging. What is your take on all of that?

Andreas Stuhlmüller

Yeah, I keep being shocked by how much the models are converging. I guess I should stop being shocked at this point, because I'm just not updating, but it is a really interesting and surprising fact about the world that the models are so similar.

I do think that's the case, but that's not the reason why we don't offer a model picker. The reason is that I think a lot of tasks in Elicit still involve multiple models orchestrated in a way that we think makes the most sense, with particular models that are good at screening papers or extracting data.

Even though the models are so similar, I think the differences are important in subtle ways. For example, I think people love to hate on Gemini, and so do I, but when we evaluated Claude Opus 4.5 against Gemini 3 Pro at the time, I think Opus did better on extraction accuracy. But if you checked what fraction of claims are directly supported by the evidence, I think Gemini actually beat it by at least 5% or so.

The models are still micro-jagged enough that you can't say, “Oh, yeah, this model is clearly the best; you should just use that.” As a user, I don't want to put that on our users for the most part. I think mostly what our users pay for is for us to do the work of figuring out what models are good at what kind of thing and making sure those models actually get used in those places.

Nathan Labenz

Yeah, interesting. So, there is a place for Gemini in Elicit today.

Andreas Stuhlmüller

Yeah, there was a place 2 months ago. I actually don't know. Even though I'm fairly on top of what's going on, our eval team is even more on top of it, so I don't actually know if it's still live. But there definitely was a place for it 2 months ago, and maybe next week there will be a place for it again.

Nathan Labenz

Yeah, and it's cool that that in and of itself is an interesting reflection of how frequently you're swapping things out and how dynamic and competitive the environment is. Maybe 3 more questions, if you will. One, do you have plans to expose Elicit as a tool for random people's Claude Code to use? That could be by allowing them to do it via the API if they have an account, or, even more likely, it could be through a sort of x402-type thing, which I've recently been exploring as a way to get just pay-per-use access to a bunch of different tools. Yes? No? Why not?

Andreas Stuhlmüller

But no, I mean, it's already—we haven't been advertising it that much, but we already have an MCP and an API. A lot of people use the API. People can check it out at docs.elicit.com.

A lot of the work I do with Elicit is through the API. When I run systematic reviews, I often use the Systematic Review API, and I iterate using various other models on the protocol for a while. Then I run it in the background and retrieve it. I think that's an important use case that we really like to support. Not everything has to happen through the interface, and I think more and more will happen through APIs.

Nathan Labenz

Yeah, okay, cool. I'm sorry I missed that in my prep, but I'll again point myself at the documentation. I'm going through a bit of an AI-for-science mini-arc right now, and I'd be interested in your take on other big-picture approaches to AI for science. You guys are obviously coming at it with the systematic-reasoning angle.

There is the sort of close-the-loop angle, where we empower these models to actually run experiments through a cloud lab or whatever, and then they'll be getting feedback from reality. That seems like it could go somewhere quite interesting. Then there's, of course, training models on other modalities of data. We've seen how proteins fold, and there's, “What if I do this perturbation to a cell? What's its next state going to be?” You can go on and on in that domain.

Any interesting takes that you think might be non-consensus that you'd like to share?

Andreas Stuhlmüller

I think all of this stuff is super exciting. Sometimes people come to us and they're like, “Well, who's going to win in AI for science?” I think that's just an absurd thing to say because science is such a big space. As you just said, there are many layers of abstraction, from understanding single-cell dynamics through automated experiments to protein models.

When you talk to the pharma companies, they're like, “We are trying to make a multiyear plan that accounts for the changing technological environment, but also accounts for the fact that clinical trials have certain intrinsic timescales.” There's just such a different reasoning problem in modeling single-cell dynamics. I don't know, it's a big space; it's a big pie. I'm excited that people are excited about it.

Nathan Labenz

Yeah. Do you have any controversial takes here?

Andreas Stuhlmüller

How it all comes together is maybe a really interesting open question. When people think about what the automated company of the future is, does it look like you take the existing top 20 pharma companies and, over time, they will morph into a different functional form? Or is it going to be the case that a small biotech comes along and they're just much more AI-integrated? In 10 years, the top 20 companies will all be replaced by companies that we don't even know the names of today.

I actually don't know how it's going to shake out. I think it could go either way, depending on how quickly people at the existing companies wake up and understand how much everything is going to transform. So, yeah, maybe that's not a very interesting, controversial take, given that I don't have a take between those 2 futures.

Nathan Labenz

One big question I think about a lot is how integrated the models themselves will be. Obviously, we have the tool-calling paradigm coming along very nicely. This could be extended to a tool call to run an experiment in a cloud lab and get a result. That result could come back as a data printout of the same sort that a human would read.

Then there's this other paradigm of integration, which we're seeing a leading indicator of with image and now also video, with Google's latest Omni model. There's this sort of deep integration of language and pixel space. In the early ChatGPT image-generation experience, you would talk to the model—or even if you gave it a photo, it would try to caption that photo, describe it, and then use language to call the separate image-generation model and ask for something.

Of course, the people never quite looked like the ones you put in, right? Because you just can't describe a face in language with that level of fidelity. But now you have this deeper, weight-level integration where I can give you an image and say, “Make this a line drawing,” or whatever, and it has both. It understands conceptually what I want, but it also sees, in some sense, the structure of the face and can preserve that through the transformation.

So I really wonder if that's coming to all the modalities of science as well, and if it's a good idea. I think it probably is coming, but I wonder if you think it's a good idea, because it certainly would, in some ways—or at least, naively, it seems to me like it would—make process supervision more difficult.

If I can trace it and say, “Okay, you called the protein-folding model. This is what you got back. Okay, that's where you went wrong,” right? I can see us digging in and interrogating those traces a lot better, versus it just being, “I asked you for this,” and you spit out a new protein sequence because, in your weights, you were like, “Oh, I intuitively know what a sequence will do for that function that you just asked for.”

Obviously, that could be really powerful in images—it's a major unlock—and I don't see any reason it wouldn't be a major unlock in designing new proteins or what have you as well. So, do you think that's coming, and do you think it would be a good or bad idea if it does, in fact?

Jungwon Byun

Yeah, I think it's a really interesting question. I think there's a bigger question behind this, which is: Where do continuous representations win? The prior should be, you know, end-to-end optimization is strong: integrate everything. At the same time, you might have expected neural nets—at least language models—not thinking in tokens but just thinking in weight space to be more successful than they have been.

I think a lot of people have experimented with it, and I hope no one succeeds, but, on priors, I would have expected maybe people to succeed at it. Often, I think people forget that there are benefits even to the models from the discretization that comes along with that.

I guess you can ask yourself: Why is human language the way it is? Why does it have discrete words in the first place? Why do we think in words and sentences? Even the models will benefit from the discretization. Why don't the models just write programs in weight space? Why do we even have programming languages? Why doesn't everything happen in continuous space? Maybe that's eventually the future; I don't know, but I wouldn't necessarily bet on it.

I think the straightforward “Everything will be end-to-end optimized in weight space” take is probably a little bit too lossy to be a good predictor of what will actually happen. In the protein-dynamics case, I could actually see continuous representations being pretty good, but I'm not sure that will be the case everywhere. I think discretization has benefits that people sometimes overlook.

Nathan Labenz

Give me one more beat on what you think those benefits are, because I totally agree with you that I don't want to see neuralese take over, win out. At the same time, when you say, “Why do we think in language?” my immediate answer is, “Because it's all we have,” right?

I don't think in language—somebody throws me a ball, and I don't think in language about it because I'm a physically embodied person who has intuitive physics, and I just catch the ball. I feel like if I had those sorts of senses for how proteins fold, I'd probably use them, but I probably wouldn't think in language about them.

Jungwon Byun

I mean, I think the fundamental property you get from discretization is error correction, right? If I say a word like “a little brown,” you can still round it off to, “Okay, that's a word.” At every step, you get a little bit of error correction.

If you're trying to chain together many words, you don't get the compounding errors. I'm not an expert on this, but I think this is roughly why we don't have analog computers these days—why we have discrete computers—is because you get these nice error-correction properties. That could be one thing you might lose if you're trying to push everything into weight space.

Nathan Labenz

Yeah, interesting. Okay, last one. Zooming out and just going back to your original mission of radically improving the quality of reasoning in science and in society, how do you think we're doing? We've got inference everywhere. Is it serving us well, and how would you handicap the trajectory that we're on as we think about recursive self-improvement possibly soon, transformation possibly soon?

Is the quality of reasoning on track to rise to the level that we need it to? What's the state of the species, so to speak?

Jungwon Byun

What is the state of the species? I don't know.

Andreas Stuhlmüller

I mean, there's this meme where a person jumps from the roof and is like, “So far, so good,” as they're halfway down. So I think, in many ways, the models have probably improved our reasoning. I think I probably get better answers for many questions that I care about than I did in the past. But I also think that is not necessarily indicative of what will matter most to the species.

When I think about what is going to matter most, it's one of the decisions that governments and big AI projects are going to make, or maybe other large organizations, as AI transforms everything. And there, I think the game is still open. I think, in many ways, we're still extremely early. I think AI has transformed basically nothing.

I think people are always like, “Well, coding is getting automated.” But I don't know. Most intellectual work in the economy is not coding. It also still employs many people. And so, I think we've seen nothing yet. We're still early, and I also think all the big decisions are still coming up.

When it comes to how AI will impact epistemics and good reasoning, I think it really could go either way, because the models are optimized to look good and be persuasive. I could see a worsening of epistemics happening if we don't make this an explicit priority. At the same time, I don't know. I mean, I think the models can be optimized for truth-seeking, too.

And if you prioritize those interventions, there's maybe a basin of attraction where, once you become more truth-seeking, you realize, “Okay, what are the most important interventions?” Oh, I guess we should prioritize being better at forecasting the results of what we do. And then you become better at prioritizing more epistemics-related interventions.

So I do feel like, for better or worse, we're still before the point of no return in either direction. I'm excited for people to build tools in this space. Elicit's trying to build tools in this space. I think people can help advocate for the adoption of better epistemic tools. It does feel like a very small area relative to how important it seems to me to be for the future of our species, as you said.

Nathan Labenz

It was the smartest of times; it was the stupidest of times. Here's hoping that our better angels win out when it comes to better reasoning and better decision-making.

This has been an excellent conversation. I really have enjoyed the update on Elicit, and I appreciate how consistent and disciplined you guys are about it. It's not easy, obviously, running a business and trying to make sure you're staying true to that North Star mission. I think you guys do a really admirable job of holding yourselves accountable to trying to find the right balance between those, and there are many echoes of the hard work that you've put into that in this conversation.

So I really appreciate the time and encourage you to keep up the great work and the disciplined reasoning. Anything else you want to leave people with before we break?

Andreas Stuhlmüller

I think knowing when the models are making you better or worse at decision-making is actually pretty subtle, and I think not that many people are paying close attention to it. I'm reminded of the METR study from a while back, where engineers thought they were being made more productive, but actually they were at a slight discount relative to unassisted work. And that was in engineering—probably no longer true in engineering—but I could imagine that, for more complex decisions, if you're not paying attention, sometimes using the models regresses you to the mean or cuts off avenues of investigation, and at other times it opens up the space and makes you think more clearly.

So I find it helpful to just introspect: Am I actually getting benefits here, or how is this changing my behavior? Thinking about that and sharing it broadly does seem to me like just a clearly net-good thing.

Nathan Labenz

Yeah. For the time being, we still have some agency over this process, and I think that's a great reminder to maintain an ownership mindset and hold ourselves accountable to doing our very best work—not getting lazy and letting the AIs lead us around.

Excellent. Jungwon Byun and Andreas Stuhlmüller, founders of Elicit, thank you again for being part of The Cognitive Revolution.

Andreas Stuhlmüller

Thanks, Nathan.

Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research | BidClub