[BidClub_]
The Cognitive Revolution · · 190 min

It's Crunch Time: Ajeya Cotra on RSI & AI-Powered AI Safety Work, from the 80,000 Hours Podcast

Ajeya CotraRob Wiblin

YouTube
TL;DR
  • Cotra’s modal expectation is for top-human-expert-level AI in the early 2030s, after which software automation could spill rapidly into robotics, chipmaking, and the full physical production loop. If AI and robots can perform everything required to make more AI and robots, today’s roughly 2% growth norm ceases to be a persuasive ceiling. Her outside possibility is that 2050 differs from today as radically as today differs from hunter-gatherer society: “10,000 years of progress rather than 25 years of progress.”

  • The investable disagreement is not whether AI matters, but whether it adds 0.3 percentage points to growth or pushes peak growth toward 1,000% annually. Slow-growth thinkers extrapolate 150 years of stubbornly stable frontier growth and assume hidden bottlenecks; fast-growth thinkers extrapolate the longer historical acceleration created by larger populations, more ideas, and positive feedback. The resulting gap is “100 or 1,000 or a 10,000-fold disagreement,” large enough to reverse views about work, policy, and risk.

  • Benchmark headlines are weak warning signals because every benchmark follows an S-curve, saturates, and is replaced before it establishes real-world danger. Cotra instead wants fixed-cadence disclosure of labs’ strongest internal results, actual productivity gains, internal AI usage, and the share of pull requests mostly written and reviewed by AI. The decisive indicator is whether AI has begun accelerating the entire AI stack—not merely code, but chip design, fabrication, equipment, maintenance, and raw-material supply.

  • “Crunch time” is the narrow interval when AI may be powerful enough to transform safety work but not yet powerful enough to escape human control. Once AI R&D is substantially automated, progress previously expected over 10, 20, or 30 years might arrive within six months to two years. Cotra’s prescription is to redirect as much AI labor as possible away from recursive capability improvement and toward alignment, cyber defense, biodefense, monitoring, negotiation, and collective decision-making.

  • Major frontier labs’ convergent safety plan is to use each generation of AI to understand and secure its successors, but the binding risk may be institutional commitment rather than technical impossibility. A lab with 100,000 smart-human equivalents could still assign only 100 to safety while competition consumes the rest. Cotra is “reasonably bullish” that control techniques can extract useful work from early, non-“galaxy brain” systems, yet warns that insufficient checking hands power to the models while exhaustive checking destroys the speed advantage.

  • Compute and model access become strategic assets if external safety organizations must mobilize during crunch time. The leading lab might withhold its best internal system, or price inference near the opportunity cost of using that compute for further self-improvement. Cotra therefore entertains owning GPUs, securing model-access agreements, or hedging compute inflation through NVIDIA and other AI-exposed public equities—though a super-exponential feedback loop could still turn today’s competitive market into winner-take-all concentration.

  • Cotra’s organizational lesson is that AI adoption must begin before the emergency, because neither governments nor philanthropies can improvise a new operating model in six months. Open Philanthropy might eventually spend more on API credits and GPU time than on human salaries, but its multilayer grant process is poorly shaped for billion-dollar, time-critical deployments. Her broader warning is vivid: without aggressive adoption, industry’s “fast cars” will be overseen by regulators using “horses and buggies.”

  • The episode’s post-publication framing says the timetable may already be shortening. The cross-post notes that on March 5, after making forecasts in January 2026, Cotra wrote “I Underestimated AI Capabilities Again” because several expectations were already beginning to be met in the first couple of months of the year; it also cites Anthropic’s Mythos model and reported zero-day discoveries across major operating systems and browsers. Its conclusion is deliberately conditional but urgent: “crunch time is arguably here now.”

Digest · the substance, structured for research

1. The cross-post says crunch time may already have begun

  • The Cognitive Revolution introduces Cotra as an unusually strong forecaster: third among more than 400 participants in the 2025 AI Digest forecasting survey. Its framing is that even accelerationists may experience “future shock” if compounding automation encounters no insurmountable bottleneck.

  • The postscript says Cotra’s March 5 article, “I Underestimated AI Capabilities Again,” reported that predictions made in January 2026 were already beginning to be met in the first couple of months of the year.

  • The introduction also points to Anthropic’s Mythos, citing major benchmark gains and reported zero-day discoveries in every major operating system and browser. Its bottom line is not that superintelligence has arrived, but that “crunch time is arguably here now.”

2. AGI has been watered down until it predicts almost nothing

  • Cotra’s DealBook example exposes the semantic drift: seven or eight of roughly ten panelists expected, by 2030, AI able to do everything humans can do, yet eight expected AI to create more jobs than it destroyed over the following decade.

  • Cotra herself did not endorse the by-2030 forecast because her timelines were somewhat longer. Her confusion was conceptual: “Why is it that you think we will have AI that can do absolutely everything that the best human experts can do in five years,” yet employment remains conventionally expansionary?

  • When pressed, panelists retreated to milder definitions—“we kind of already have AGI,” perhaps GPT-5—then treated the absence of immediate transformation as evidence that AGI was never momentous. Cotra sees that as importing evidence from one definition into a radically stronger one.

3. Cotra’s 2050 is closer to prehistory than to 2000

  • The mainstream picture is manageable continuity: a somewhat larger population, better medicine, modestly longer lives, and technological change comparable to 2000–2025. Even people expecting AGI in 2030 often retain that baseline.

  • At the opposite pole, the worldview represented by If Anyone Builds It, Everyone Dies has a sudden jump from systems like GPT-5 or GPT-6 to intelligence against which humans resemble “cats or like mice or ants,” followed by immediate physical leverage through nanotechnology or near-light-speed probes.

  • Cotra sees a broad spectrum between those extremes: intermediate stages still lead toward technologies near physical limits, including useful self-replicating systems. Her striking comparison is that 2050 might embody “10,000 years of progress rather than 25 years of progress.”

4. Remote-work superintelligence could close the physical loop

  • Cotra’s modal forecast is top-human-expert-level AI in the early 2030s: systems better than the best virologist, software engineer, or other specialist at tasks performed remotely through a computer. Narrower systems may already have changed the world substantially before then.

  • From that point, superhuman cognitive systems could direct human labor to build robotic actuators, then automate more physical work. Cotra’s uncertainty is wide, but she considers perhaps one or a few years enough to make robotics materially more autonomous because large models, scale, data, and imitation are already advancing it.

  • Tom Davidson’s “Three Types of Intelligence Explosion” supplies the missing frame: better AI software is only one loop. Full self-reproduction requires chip design, fabs, the machines that build fab equipment, repairs, raw materials, and every upstream dependency needed to produce more chips and models.

5. Growth forecasts differ by three or four orders of magnitude

  • The conversation puts the live range starkly: some serious analysts expect AGI to add only 0.3 percentage points to economic growth, perhaps roughly a 15% increase over current rates; others envisage peak growth of 1,000% annually or even thousands of percent.

  • This is not a disagreement among isolated camps that never exchanged arguments. People have discussed the mechanisms repeatedly and still differ by factors of 100, 1,000, or 10,000 in the technology’s likely economic impact.

  • The policy consequences reverse with the forecast. An accelerationist who moved from 0.3% to 1,000% would likely become far more concerned; an x-risk researcher moving the other way might regard the new view as decisive evidence against much of their current work.

6. Competing outside views prevent convergence

  • Slow-growth thinkers begin with 100–150 years of roughly 2% frontier growth. Electrification, washing machines, radio, television, computers, and the internet transformed life without producing an obvious permanent break in measured growth, so AI may merely sustain the same trend.

  • Their supporting intuition is Hofstadter’s law: “It always takes longer than you think even when you take Hofstadter’s law into account.” Cotra’s favorite variant is the programmer’s credo: “We do these things not because they are easy, but because we thought they would be easy.”

  • Fast-growth thinkers zoom out 10,000 years and see acceleration, from perhaps 0.1% growth around 3000 BC to post-Industrial-Revolution rates near 2%. More people produced more ideas, raising food output, supporting more people, and repeating the feedback loop.

  • Each side has an error theory for the other. One hears every explosive scenario as another story omitting drag; the other hears “there will be bottlenecks” as an ungrounded blanket assumption whose specific examples might reduce 1,000% growth without establishing ceilings of 2% or even 10%.

7. Bottlenecks need observable tests, not another round of stories

  • Cotra wants near-term observations that can adjudicate between those priors, beginning with actual AI uplift in software and AI R&D. METR’s randomized controlled trial split developers into AI-allowed and AI-disallowed groups and, in that setting, found AI slowed completion of real work.

  • She does not expect that result to remain true, but values having the measurement before speedups become overwhelming. Benchmarks should be cross-checked against internal rollouts, company RCTs, self-reports, and observable output rather than treated as self-validating capability measures.

  • Her formula is to find domains with the deepest adoption and measure actual production: not just whether a model answers questions about solar panels, but whether an AI-heavy factory manufactures them faster or improves their performance more rapidly.

8. Agent benchmarks improved, but the real world remains the test

  • Cotra’s late-2023 requests for proposals distinguished difficult, realistic agent benchmarks from other evidence. She wanted models to book flights, repair software, write and run tests, and iterate until success—not merely answer multiple-choice questions.

  • Researchers responded strongly to the benchmark arm, producing projects including Cyber, a cyber-offense benchmark used in standard evaluations. Interest was much weaker in surveys and RCTs because benchmarks are “clean and contained,” while reality is “messy and open-ended.”

  • Cotra nevertheless stresses that benchmarks routinely overestimate real-world performance. The companion evidence program was meant to capture deployment, adoption, and outcomes that automated scoring cannot reproduce.

  • The Forecasting Research Institute’s LEAP—Longitudinal Experts on AI Panel—now asks roughly 100–200 AI experts, economists, and superforecasters about six-month, one-year, and five-year indicators, then connects granular predictions to longer-run worldviews and checks who was right.

9. A private intelligence explosion could outrun public awareness

  • AI 2027 provides the scenario: a leading company, OpenMind in the story, becomes so far ahead that it retains its best models internally and releases only systems slightly better than competitors’ public frontier.

  • Such a company need not monetize its best product if internal AI R&D is worth more than external revenue. The public could therefore see only incremental product improvement while the lab experiences explosive productivity behind closed doors.

  • Cotra’s concern is the lost warning interval. Society might have known six or twelve months earlier which growth regime had begun, yet governments and outside experts would remain unable to act because the decisive evidence stayed internal.

10. Transparency should measure delegation, not just benchmark scores

  • Cotra would like labs to report their strongest internal benchmark results on a calendar cadence—perhaps every three months—rather than only when a product launches. Public model cards for Claude Opus 4 or GPT-5 are helpful but miss purely internal deployment.

  • She is more interested in how systems are actually used. CEOs sometimes boast that AI writes 90% of code, but line count is weak evidence; a sharper measure is the fraction of internal pull requests mostly written and mostly reviewed by AI.

  • That measure captures both capability and organizational deference. If humans continue handling management, approval, and review, progress can accelerate only so far; for “crazy fast” growth, AI eventually has to perform those higher-level functions too.

  • Cotra also wants internal uplift studies, subjective speedup estimates, algorithmic progress, and serious misalignment incidents—for example, whether a deployed model lied about something important and concealed the logs. Those disclosures are also the most competitively and reputationally sensitive.

11. Public evidence matters because an alarm must become common knowledge

  • Benchmarks alone cannot trigger the alarm because they repeatedly trace an S-curve, saturate, and give way to harder successors. Even 100% on today’s evaluations would not convince Cotra that a model could take over the world.

  • The late but clear signal is observed productivity: internal discovery occurring far faster than before. That is the point at which “they should definitely sound the alarm,” because theoretical progress is visibly feeding back into more progress.

  • Rob suggests confidential reporting to technically capable government agencies, which routinely receive commercially sensitive information. Cotra accepts that as better than silence but doubts that 10 or 50 understaffed officials can interpret a shifting evidentiary frontier fast enough.

  • Her preferred alarm resembles COVID or the public reassessment after Joe Biden’s disastrous debate: a society-wide conversation. A prominent skeptic such as Arvind Narayanan publicly changing his mind would create common knowledge that rumors at San Francisco parties cannot.

12. Transparency legislation should target the highest-value signals

  • Cotra does not rank an all-or-nothing disclosure package as the single top policy fight. She favors identifying what evidence would reveal an intelligence explosion, then securing the highest-value, biggest-bang-for-buck elements first.

  • Public disclosure creates difficult trade-offs: unusually rapid algorithmic progress attracts competitors, unusually slow progress discourages investors, and misalignment incidents are embarrassing. Industry aggregation might soften those costs, but with only a few frontier labs observers could often infer the source.

  • Existing efforts such as New York’s RAISE Act and California’s SB 53 fit the broader approach because they emphasize transparency and whistleblower protection. Cotra treats whistleblowers as an important policy plank supporting transparency.

  • Leaks remain insufficient. Social proximity between Bay Area safety researchers and lab employees can reveal what is coming, but “rumors in San Francisco tech-bro parties” cannot justify costly action in Washington, London, or Brussels.

13. The alarm should redirect AI labor before capabilities compound

  • If AI substantially automates frontier AI R&D, progress previously expected over 10, 20, or 30 years might arrive within a year or two, perhaps six months. Systems may be manageable at the alarm point yet rapidly approach “godlike abilities.”

  • Cotra’s response is to redirect as much AI labor as possible from making stronger models toward work that protects society from successor systems. The target includes takeover risk, misuse, infrastructure vulnerability, geopolitical conflict, and other disruptions created by increasingly powerful AI.

  • A leading company cannot easily redirect unilaterally because rivals may catch it. The warning therefore creates a coordination window: if society can credibly see six, 12, or 18 months to radical superintelligence, competitors may coordinate to redirect resources toward protective work.

14. AI-powered safety is not circular, but trust is the bottleneck

  • Critics describe the strategy as “flying by the seat of your pants”—using the technology creating the problem to solve it. Cotra replies that civilization repeatedly used general-purpose technologies to contain their own externalities.

  • Cars enabled drive-by shootings and carjackings, but also equipped law enforcement; computers enabled hacking, but also automated monitoring and vulnerability discovery. Forecasts often imagine a new technology’s harms in detail while failing to imagine equally technology-enabled defenses.

  • Nathan initially suggests misalignment is uniquely awkward because an AI asked to solve alignment might sabotage its own constraint. Cotra’s pushback is broader: a power-seeking model would also undermine biodefense, epistemic improvement, and civilizational defenses that might expose or block it.

  • The central technical problem is therefore building enough confidence through control, alignment, interpretability, and related techniques to rely on AI outputs. Excessive checking bottlenecks progress; inadequate checking means “we hand the AIs the power to take over.”

15. The defensive portfolio runs from alignment to civilization-wide resilience

  • Alignment is foundational: each system and its successors must remain motivated to help humans, honest, steerable, and responsive to instruction. A failure anywhere in that chain compromises every later protective application.

  • Cyber defense is a natural dual-use race. The same models capable of finding vulnerabilities in power grids, weapons, and critical systems could identify and patch them before malicious actors obtain comparable capability.

  • Biodefense could combine rapid detection of novel pathogens, faster medical countermeasures, and large-scale production of PPE, clean rooms, or related infrastructure. If robotics has matured too, AI could help execute manufacturing rather than merely recommend it.

  • More speculative work targets collective cognition: truth-seeking, negotiation, compromise, policy design, avoiding US–China war, space governance, value lock-in, and healthier political discourse. Cotra groups much of this as AI for “coordination, compromise, negotiation, truth-seeking.”

16. Major frontier labs are betting that one generation can secure the next

  • Cotra sees the same architecture in public plans from OpenAI, Anthropic, and Google DeepMind: as models improve, the companies expect to incorporate AI itself more heavily into alignment, control, interpretability, and safety work.

  • The plan requires a useful interval before systems become uncontrollably powerful—ideally observable in advance and lasting at least six months or a year even without an extraordinary slowdown.

  • If a generality threshold instead produces extreme superintelligence within days or weeks, the strategy fails. Society may not notice the transition, mobilize resources, validate the models’ work, or coordinate before control is already lost.

17. Capability ordering can make the window useless

  • A particularly bad ordering produces AI that is extraordinary at AI R&D but weak at everything protective, even closely related safety research. It might improve successor models’ sample efficiency for months without being useful for alignment, biodefense, or governance.

  • Nathan notes the apparent contradiction: such a narrow system is not itself broadly dangerous. Cotra’s resolution is an AlphaFold-like savant or blind search that discovers an architecture or training method whose successor can suddenly “go foom.”

  • Current skill profiles already look jagged. Cotra’s ML-researcher friends receive larger gains than she does in weird open-ended thinking, podcasts, and grant-related emails; control and alignment benefit because much of that work resembles software engineering and ML research.

  • Moral philosophy, negotiation, policy, and institution-building may lag badly. The more distant a task is from code and ML research, the larger Cotra expects the capability penalty to be.

18. Tight feedback loops favor code over institutions

  • AI excels where success is rapidly legible: code either runs or fails. That “very hard to fake signal” supports training, evaluation, and repeated improvement in a way that year-long management or political projects do not.

  • A model may produce a persuasive white paper and receive a thumbs-up without being able to lead thousands of humans and robots through execution. Nathan’s caricature captures the gap: “Here’s the widget…go and make 10 billion of them.” The model replies, “Good luck.”

  • Defensive planning should therefore assume that ideas arrive before execution capacity. Long-lead infrastructure—PPE, vaccines, pathogen detection, clean rooms, or manufacturing capacity—may need to be built before systems become capable of designing threats.

  • Cotra treats AI defensive labor as a forecast to plan around, “not a guarantee.” Humans should specialize now in tasks where future models may remain comparatively disadvantaged, especially physical deployment and social coordination.

19. The intelligence explosion changes both urgency and capability

  • Safety work should use AI as soon as it becomes useful; crunch time is not permission to wait. Its special significance is the clock: the default trajectory might leave only 12 months to uncontrollable superintelligence.

  • Crunch time also supplies a new lever. Because AI is, by definition, highly capable at AI R&D then, some of that work can be redirected toward filling its own skill gaps—fine-tuning or scaffolding systems for biodefense, philosophy, negotiation, or other priorities.

  • Cotra extends the logic to philanthropy: today more than 80% of grant money goes to salaries for humans; within a few years, the better allocation could be API credits or rented GPU time supporting a similar distribution of research and policy work.

20. A pause should stretch the transition, not freeze and jump

  • Cotra regards her proposal as compatible with pausing at the brink of an intelligence explosion. If the default is 12 months, she would vote to make the transition “10 times longer or even longer”—10, 15, or 20 years rather than one.

  • She reframes a pause as a continuum. Instead of choosing between 100% of AI labor going to better models and 0%, society can repeatedly slow, redirect, test, and selectively advance capabilities that improve defenses without making systems uncontrollable.

  • Her quibble is with “pause, hang out for 10 years, then unpause,” which preserves the later jump. She prefers slowly inching through a sweet spot where models are powerful enough to help but not powerful enough that “we’ve already lost the game.”

21. The likeliest lab failure is underinvestment, not impossible alignment

  • Cotra’s leading failure mode is that labs never execute the promised redirection. Competitive pressure could leave 100 of 100,000 smart-human equivalents working on safety—more labor than before, but tiny relative to the pace of capability improvement.

  • Rob proposes legal requirements or inter-company commitments to devote perhaps 50% of compute to safety. Cotra sees the enforcement problem: what counts as safety, who audits every team, and how can regulators distinguish protective research from capability work without deep technical competence?

  • She considers severe early misalignment plausible but not the most likely failure. If controls cannot obtain honest work, models might selectively advance capabilities while sabotaging alignment, biodefense, and other safeguards; even a slowdown would then leave humans “on our own.”

  • A discontinuous corporate pivot is intrinsically difficult. Cotra wants labs to increase human labor, inference compute, and fine-tuning devoted to safety gradually, so the final transition follows an established schedule instead of requiring an abrupt institutional reinvention.

22. Philanthropy may have to replace salary budgets with inference budgets

  • Open Philanthropy’s immediate task is to monitor how useful AI is across its own work and that of grantees such as Forethought, Redwood Research, Apollo, and policy organizations. Once a funded activity shows “signs of life” from automation, it may warrant rapid scaling.

  • Cotra suggests explicitly separating conventional grants from spending that buys AI labor. Tracking ChatGPT Pro subscriptions, API credits, and inference compute would reveal whether AI’s share of giving is rising in line with capabilities and expected crunch-time timing.

  • Existing approval chains are poorly shaped for a billion-dollar decision: a junior staffer develops a case, then information climbs through two, three, or four managerial layers. That process cannot simply scale to a time-critical commitment that no junior employee could reasonably authorize.

  • In the easy scenario, Open Philanthropy itself is largely automated and has perhaps 1,000 AI-team workers rather than 45. In the harder jagged scenario, AI masters a few fundable verticals while the institution remains human-bottlenecked and lacks a visceral sense that the moment has arrived.

23. Model access and compute ownership become strategic bottlenecks

  • An external safety organization may possess billions of dollars yet be unable to buy the leading model. A frontier lab could retain its best systems internally and sell only less-capable systems that remain marginally ahead of competitors’ public products.

  • Even willing sellers may charge prohibitive prices because the opportunity cost of inference is training or operating a stronger successor. During crunch time, compute prices could therefore reflect recursive-improvement value rather than ordinary API economics.

  • Cotra’s hedge ranges from owning GPUs—renting them commercially in normal times, then redirecting them—to holding NVIDIA or other liquid AI-exposed stocks. Rising compute costs would then increase portfolio value, though the organization would still need rights to run frontier software on its hardware.

24. Competition initially opens access, then may create a winner-take-all race

  • Cotra expects early crunch time to resemble today: several leading companies within perhaps a month of one another, with different capability spikes, limited moats, and strong incentives to sell API access.

  • A super-exponential loop changes the market structure. Under ordinary exponential growth, competitors growing at the same rate preserve relative shares; under accelerating growth, the actor reaching each milestone first can compound its lead into disproportionate wealth and power.

  • A distant leader has reason to conceal its system, forgo quarterly revenue, and recurse toward intelligence capable of rivaling nation-states or “decisively” shaping the future. Close competitors and impatient investors instead force commercialization and narrow the room for secrecy.

  • Safety also competes with two more attractive uses: private power and ordinary goods, services, and media content that consumers will pay for. As with today’s limited spending on biodefense, cyber defense, or moral philosophy, market demand does not automatically allocate transformative capacity toward public protection.

25. Sequential bottlenecks make preparation more valuable than raw compute

  • Physical defenses contain irreducible calendar time: experiments must run, factories must be built, and countermeasures must be manufactured. Doubling compute does not necessarily halve any of those schedules.

  • Social systems are similarly sequential. AI might identify a mutually beneficial and enforceable US–China agreement, but officials still have to meet, deliberate, come to a decision, and ratify terms.

  • Even theoretical research is not perfectly parallel. A hundred models can search for an insight, but if the solution requires three or four dependent conceptual leaps, the field must discover each foundation before work on the next can begin.

  • The best pre-crunch contributions are therefore long-lead assets: physical biosecurity infrastructure and social consensus around possibilities such as slowing jointly, redirecting compute, or negotiating treaties. Ideas need years to enter policymakers’ “toolkit” before an emergency.

26. Governments risk regulating fast cars with horses and buggies

  • Cotra especially wants government entities to adopt AI now. Random types of red tape could widen the gap until regulated companies have “fast cars” while regulators retain “horses and buggies.”

  • Defensive organizations should not wait for a universal model. Each needs a team repeatedly asking where AI has become genuinely useful, how workflows should change, and which outputs can be verified reliably.

  • Her advice is aggressive but conditional: maximize adoption for the organization’s actual use case, measure performance, and watch where automation first becomes dependable. Familiarity with current limitations is part of preparedness, not merely a productivity initiative.

27. Cotra’s grantmaking career began as emergency triage

  • Cotra joined Open Philanthropy in 2016 but made no grants for more than six years. The FTX collapse changed that: hundreds of promised recipients faced cancellation or clawbacks, creating an emergency call that required surge capacity.

  • In roughly six weeks, she went from no grants to about 50. The experience showed her that she liked parts of grantmaking, but rapid triage prevented the depth she naturally sought.

  • Technical proposals in interpretability or adversarial robustness often looked reasonable without a complete causal story. Cotra wanted to specify how a research direction produced a capability or technique, how that technique fit a takeover-prevention plan, and what success would concretely mean.

28. Inside-view funding produced a narrow, expensive, opinionated bet

  • Open Philanthropy’s older heuristic—support a strong researcher with a good record who cares about AI safety—was defensible, but emotionally unsatisfying to Cotra. She wanted enough understanding to answer both more-doom and less-doom critics in sustained detail.

  • Her temporary compromise was a barbell: process renewals quickly using conventional heuristics while developing a smaller number of deeply justified programs. The main inside-view bet became realistic agent benchmarks and non-benchmark evidence about AI’s effects.

  • The agent-benchmark RFP ultimately made about $25 million in grants, with another $2–3 million through its broader RCT-and-survey companion. It specified hard, realistic tasks and repeatedly warned applicants that apparently difficult tasks were probably “not hard enough.”

  • Cotra acknowledges the opportunity cost: the same time could have funded more low-hanging opportunities across ten fields. Her defense is that deep homework changes details, reveals stronger versions of researchers’ ideas, and enables grantmaker and grantee to co-create better projects.

29. Burnout revealed the cost of operating without an institutional thought partner

  • When Holden left, Cotra lost a manager who would debate timelines, takeoff, and threat models at the object level. Remaining leadership had less AI context and less bandwidth, turning “Is this strategy right?” into “You can do it that way if you want.”

  • Cotra felt alone while trying to build an understanding-oriented program and discovered she was not naturally entrepreneurial. Hiring could not substitute for a shared central brain because the vision remained nebulous and candidates had to resonate with it before it was fully defined.

  • Perfectionism made management expensive. Writers rarely captured her ideas as she wanted, editing often took longer than writing, and delegating grant work could be slower than doing it herself—the familiar new-manager problem arriving just as guidance from above diminished.

  • Her four-month sabbatical mixed practical recovery, a new group house, exercise, involvement in the Curve conference, unpublished writing, and career reflection. She concluded that she wants to advise and help an organization’s center more than run an isolated internal startup.

30. EA’s appeal combined impartial care, intellectual depth, and radical integrity

  • Cotra entered the effective-altruism “rabbit hole” at 13. Its first attraction was moral scope: helping distant people, animals, future generations, and potentially conscious artificial systems rather than privileging those nearby or familiar.

  • The second was methodological seriousness: stop, research, quantify, and compare interventions that may differ by orders of magnitude. Her analogy is how intensely someone would investigate treatments if they or a spouse had cancer.

  • The third was an exacting integrity beyond what donors demanded. GiveWell maintained a public mistakes page and refused donation matching because it is usually a scam, even though using it might have raised more money.

  • As EA shifted from persuading analytical donors toward deploying money, talent, and political influence, radical openness became strategically costly. Donors sought privacy, campaigns could not reveal tactics, and adversaries began searching Open Philanthropy’s publications for material that could damage it.

31. Cotra wanted EA to be more religion-shaped, not less

  • Cotra regards EA as far more truth-seeking than religion, but accepts the analogy because it offers a “map of the good life,” a community, a worldview, and guidance that crosses personal conduct, politics, and one’s place in history.

  • What it supplied professionally was not quite what she needed emotionally. Her corner of EA idealized having an impactful job and working extremely hard; she wanted more deliberate existential reflection on morality and a world that might become utopian or dystopian within a decade or two.

  • Her imagined “EA church” would convene thoughtful discussion each week, reconnecting daily documents and emails to the larger stakes. Rob says he personally likes the more professional, limited aspect because he wants to go home and not think about the work all the time; Nathan, by contrast, says he wants to think about his work in a more spiritual way.

  • She recognizes the danger of cultishness and why a professional community should welcome good safety researchers without demanding philosophical conversion. Still, Joe Carlsmith’s popularity among committed EAs suggests she is not alone in wanting that form of nourishment.

32. Better local fit made the same work feel lighter

  • Cotra initially planned to start a Substack after sabbatical, admitting the choice was motivated partly by desire rather than a confident highest-impact case. She stayed when Open Philanthropy searched for new global-catastrophic-risk leadership and both leading candidates seemed strong.

  • Working with Emily Olsen restored the integrated role Cotra wanted: Emily had a larger project, needed Cotra’s answers, and would use them in strategy. Cotra found herself working more hours than before sabbatical while experiencing the work as less difficult.

  • At recording, she was also exploring Redwood Research and a work trial with METR; the cross-post introduction says she is now working on risk assessment at METR. Both offered narrower missions—AI control or early warning—along with the opportunity for deeper work.

  • Her least glamorous conclusion may be the most generalizable: “the literal person you’re reporting to matters a huge amount.” Leadership changes should trigger a fresh assessment of role, rhythm, collaborators, and fit even when the organization’s nominal mission is unchanged.

33. EA’s AI-era niche may be speculative research others avoid

  • Cotra thinks AI safety no longer requires the whole EA moral package; misalignment and misuse concern people with many values. EA remains distinctive where altruistic motivation, unconventionality, and tolerance for philosophical uncertainty are prerequisites.

  • Digital sentience is the clearest example: asking whether AI systems can suffer, deserve protections, or create trade-offs with safety is neither lucrative nor conventionally prestigious. EA can incubate the field before it becomes institutionally respectable.

  • Value lock-in need not require a dictator. Nathan imagines “social media plus plus”: personalized, superintelligent information bubbles that help each person defend existing beliefs indefinitely, leaving power distributed while society becomes increasingly incapable of reflection or revision.

  • The broader comparative advantage is “informed speculation” between storytelling and measurement. Nathan wonders whether some EAs steered into operations or policy, while privately wanting to be “a weird truth-teller,” should return to research precisely because the rest of the world undersupplies it.

Nathan Labenz

Today's episode is a cross-post from the 80,000 Hours Podcast, hosted by Rob Wiblin and featuring a conversation with Ajeya Cotra, who previously led technical AI safety grant-making at Open Philanthropy, now Coefficient Giving, and is now working on risk assessment at METR.

For years, AI insiders have recognized Ajeya as one of the most rigorous thinkers about the AI future. She recently validated that judgment by coming in number 3 out of more than 400 participants in the AI Digest's 2025 AI forecasting survey. For comparison, I was proud to land in the top 5% at number 23.

In this conversation, Ajeya takes Rob through her expectations for the next few years as AI crosses critical thresholds, recursive self-improvement intensifies, and we enter what she describes as crunch time: a potentially short window in which AI is powerful enough to dramatically accelerate AI R&D, but not yet totally beyond human control. As a preview, I'll warn you that even the accelerationists may suffer some future shock from this conversation, because Ajeya thinks it's actually quite plausible that we find no insurmountable bottlenecks to widespread and compounding automation.

If so, the world of 2050 could look as different from our perspective today as our world would look to hunter-gatherers of 10,000 years ago. So, what's the plan to make sure such a mind-boggling transformation goes well for humans? Ajeya advocates for transparency measures and early-warning systems designed to make sure that superintelligence doesn't happen in secret.

Aside from that, she reports that all frontier developers are gradually converging on a strategy of using each generation of AIs to attempt to align, understand, and control their own successors. As regular listeners know, I signed the Future of Life Institute's October 2025 petition calling for a ban on superintelligence—not because I think this approach is forever destined to fail, but simply because I worry that we don't yet understand AIs well enough to bet on a good outcome from such a recursive-self-improvement-powered intelligence explosion.

At the same time, I do agree with Ajeya's advice: almost regardless of the kind of work you're doing, you should be adopting AI as aggressively as possible, both to maintain an accurate understanding of the situation and increasingly because you won't be able to keep up without it. It is a mad, mad world that we'll soon be living in. But I would go as far as to say that even pause-AI campaigners ought to be using AI extensively.

If all that weren't enough for you to process, you should also know that the situation has recently accelerated yet again. On March 5, just about 2 weeks after this episode was originally published, Ajeya posted an article on her Substack, Planned Obsolescence, called “I Underestimated AI Capabilities Again,” in which she reports that the predictions she made in January 2026, which were the backdrop for this conversation, were already starting to be met in just the first couple of months of this year.

More recently, we've learned of Anthropic's new Mythos model, which, despite the fact that Anthropic has never emphasized benchmark scores as much as other model developers, shows major gains on many benchmarks and has reportedly found zero-day exploits in every major operating system and every major web browser, among many other major software projects. The bottom line is that crunch time is arguably here now. So, if you've been watching and waiting for AI to get serious before deciding what to do about it, I would suggest getting off the sidelines sooner rather than later. If you need help figuring out what to do, you might consider applying for free one-on-one career advising from 80,000 hours. As always, I want to thank Rob and the 80,000 hours team for allowing me to cross post this episode. They have been delivering incredible alpha for years. And the nearer the singularity becomes, the more impressive they look. With that, I hope you enjoy this essential conversation about AI timelines and crunch time strategy with Ajeya Cotra and host Rob Wiblin from the 80,000 hours podcast.

Rob Wiblin

If you look at public communications from at least OpenAI, Anthropic, and Google DeepMind, in all of their stated safety plans, you see this element: as AIs get better and better, they're going to incorporate the AIs themselves into their safety plans more and more. How to create a setup where we use control techniques, alignment techniques, and interpretability to the point where we feel good about relying on their outputs is a crucial step to figure out, because it either bottlenecks our progress because we're checking on everything all the time and slowing things down, or it doesn't bottleneck our progress, but we hand the AIs the power to take over.

Today, I'm speaking with Ajeya Cotra. Ajeya is a senior advisor at Open Philanthropy, where, in 2024, she led their technical AI safety grant-making. More generally, she's been doing AI-related research and strategy since 2018 and has become very influential in AI circles for her work on timelines, capability evaluations, and threat modeling. Thanks so much for coming back on the show, Ajeya.

Ajeya Cotra

Thank you so much for having me.

Rob Wiblin

Doing this interview gave me a chance to go back and listen to the interview that we recorded, I guess, 2 or 2.5 years ago. I have to say, you were very much on the ball. There were a lot of issues that came up in that conversation that you were bringing to people's attention that, in the subsequent 2.5 years, seem like a much bigger deal now.

You talked about METR's evaluations of autonomous capabilities, a line of research that's gone on to become super influential and very widely read, I think, in policy circles. You talked about using probes to monitor and shut down dangerous conversations, something that's a pretty standard practice and maybe one of the potentially most useful outputs from mechanistic interpretability. You talked about the importance of using chain-of-thought and scratchpads to monitor what AIs are doing and why, still probably the dominant technique.

You talked about the growing situational awareness of AI models and the resulting possibility of deceptive alignment, something that's now a completely mainstream topic. You talked about how, when you train models not to engage in bad behavior, they don't necessarily just learn to become honest. They also learn to hide their misbehavior better, something that research has borne out, does really happen, and is a big concern.

You talked about how you expected models to get scheming as they get smarter, especially once we inserted reinforcement learning back into the mix, something that's definitely happened. And you talked a bunch about sycophancy—how you thought models might end up just flattering people rather than giving accurate information because that's something we enjoy.

I feel like—I mean, you didn't come up with all these ideas or anything like that, but I think you were ahead of the curve, and maybe we'll get some ahead-of-the-curve ideas in this interview as well.

Ajeya Cotra

Hopefully.

Rob Wiblin

So, you think that a key driver of disagreements about kind of everything to do with AI is people's different views on how likely AGI is to speed up science and technology, and I guess physical infrastructure and manufacturing. Why is that?

Ajeya Cotra

Yeah, so I think a thing I've been noticing as the concept of AGI has become more and more mainstream is that it's also become more and more watered down. Last year, I was on a panel about the future of AI at DealBook in New York. It was me and 1 or 2 other folks who think about things from a safety perspective, and then a number of venture capitalists and technologists.

The moderator asked at the very beginning of the panel whether we thought it was more likely than not that, by 2030, we would get AGI, defined as AIs that can do everything humans can do. Seven or 8 hands went up, not including mine, because my timelines are somewhat longer than that.

Then he asked a follow-up question a couple of questions later about whether we thought that AI would create more jobs or destroy more jobs over the following 10 years. 2030 was 5 years away, and 7 out of 10 people thought that we would have AGI by 2030, but then it turned out that 8 out of 10 people, not including me, thought that AI would create more jobs than it destroyed over the next 10 years.

I was a little confused. I thought, why is it that you think we will have AI that can do absolutely everything that the best human experts can do in 5 years, but will actually end up creating more jobs than it destroys in the following 10 years? What's happening here?

When I poked some people later in the panel about that seeming tension, they really quickly backed off and said, “What does AGI really mean?” The moderator had defined it as this very extreme thing, but they were like, “We kind of already have AGI. People keep moving the goalpost. We keep making cool new products, and people aren't accepting that it's AGI. They aspire to something higher.”

I thought that was funny because, in the old-school singularitarian futurist definition of AGI, it's this very extreme thing. But I think VCs have an instinct to call something AGI—GPT-5 is AGI, or something just much milder.

I think this creates a situation where people feel like they've gotten a lot of evidence that AGI isn't a very big deal and doesn't change much, because we already have AGI, or we're going to have it next year, or we got it 2 years ago. Look around us: nothing much is changing.

Rob Wiblin

But I feel like there’s this expectation where, whether or not we get AGI in the next few years, a lot of people are starting not to really care about that question. They still expect the next 25 years and the next 50 years to play out kind of like the last 25 years or the last 50 years, where there was a lot of technological change between 2000 and 2025. But it was a moderate amount of change, and they expect that in 2050 there will be a similar amount of change as there was between 2000 and 2025.

Even if they think that we’re going to get AGI in 2030, they think AGI is just what’s going to drive that sort of continued mild improvement. Whereas I think there’s a pretty good chance that by 2050, the world will look as different from today as today does from the hunter-gatherer era. It’s like 10,000 years of progress rather than 25 years of progress, driven by AI automating all intellectual activity.

Rob Wiblin

Yeah, I guess you’ve hinted at the fact that there’s an enormously wide range of views on this. Can you give us a sense of just how large the spectrum is and what the picture looks like on either end?

Ajeya Cotra

On the standard mainstream view, if you ask a normal person on the street what 2050 will look like, or if you ask a standard mainstream economist, I think they would think, “Well, the population is a little bit bigger. We have somewhat better technologies.” Maybe they have a few pet technologies that they’re most interested in, and maybe we have this one or that one. There’s slightly better medicine. People live slightly longer. It’s an amount of change that’s extremely manageable.

On the far extreme from there is a view described in “If Anyone Builds It, Everyone Dies.” In that worldview, at some point, probably pretty unpredictably, we crack the code to extreme superintelligence. We invent a technology that rather suddenly goes from being GPT-5 and GPT-6 and so on to being so much smarter than us that we’re like cats, mice, or ants compared to this thing’s intelligence.

Then that thing can immediately have really extreme impacts on the physical world. The classical, canonical example here is the invention of nanotechnology—the ability to precisely manufacture things that are really tiny, that can replicate themselves really quickly, and that can do all sorts of things—along with inventing space probes that travel close to the speed of light.

I think there’s a whole spectrum in between, where people think that we’re going to get to a world where we have technologies approaching their physical limits. We have spaceships approaching the speed of light, and we have self-replicating entities that replicate as quickly as bacteria while also doing useful things for us. But we’re going to have to go through intermediate stages before getting there.

Something that unites all of the people who are AI futurists and concerned about AI x-risk is that they think that in the coming decades, we’re likely to get this level of extreme technological progress driven by AI.

Rob Wiblin

How strong is the correlation between how much someone expects AI, or AGI, to speed up science and research in particular—and I guess physical industry as well—and how likely they think it is to go poorly, or how nervous they are about the whole prospect?

Ajeya Cotra

I think it’s a very strong correlation. I’ve often found that reasonable people who are AI accelerationists tend to think that the default course of how AI is developed and deployed in the world is very slow and gradual. They think that we should cut some red tape to make it go at a more reasonable pace.

People who are worried about x-risk think that the default course of AI is this extremely explosive thing where it overturns society on all dimensions at once—in maybe a year, maybe 5 years, maybe 6 months, or maybe a week. They’re saying, “We should slow it down to take 10 years, maybe.”

Meanwhile, the accelerationists think that by default, diffusing and capturing the benefits of AI will take 50 years or 100 years, and they want to speed it up to take 35 years.

Nathan Labenz

It’s quite interesting that people who radically differ in their policy prescriptions might be targeting the same level of speed, actually. Maybe they want this period to take 10 or 20 years. That’s what both of them want, but they just think their baseline is so different, so they’re pushing in completely opposite directions.

What’s your modal expectation? What do you think is the most likely impact for it to have?

Ajeya Cotra

I think that probably in the early 2030s, we’re going to see what Ryan Greenblatt calls top-human-expert-level AI, which is an AI system that can do tasks that you can do remotely from a computer better than any human expert. It’s better at remote virology tasks than the best virologist, better at remote software engineering tasks than the best software engineers, and so on for all the different domains.

By that time, I feel like the world has probably already accelerated and changed, and narrower and weaker AI systems have already penetrated a bunch of places. We’re looking at a pretty different world. But at that point, I think things can go much faster, because I think top-human-expert-level AIs in the cognitive domain could probably use human physical labor to build robotic physical actuators for themselves.

That would be one of the things that, whether the AIs have already taken over and are acting on their own or whether humans are still in control of the AIs, I think would be a goal they would have: automating the physical as well. I have pretty wide uncertainty on exactly how hard that’ll be.

But whenever I check in on the field of robotics, I actually feel like robotics is progressing pretty quickly and taking off for the same reasons that cognitive AI is taking off. It’s large models, lots of data, imitation, and large-scale training helping robotics a lot.

I imagine that you can pretty quickly, maybe within a year or maybe within a couple of years, get to the point where these superhuman AIs are controlling a bunch of physical actuators that allow them to close the loop of making more of themselves—doing all the work required to run the factories that print out the chips that then run the AIs, doing all the repair work, and gathering the raw materials.

Nathan Labenz

So you’re saying you’re expecting in the 2030s that it won’t just be that these AI models are capable of automating computer-based R&D, but they’ll also be able to lead the project of building fabricators that produce the chips they run on. And so that’s another kind of positive feedback loop.

Ajeya Cotra

Yeah. I really recommend the post “Three Types of Intelligence Explosion.” That’s by Tom Davidson on Forethought, in which he makes the point that we talk a lot about the promise and the danger of AIs automating AI R&D and automating the process of making better AIs.

But that’s only one feedback loop that’s required to fully close the loop of making more AIs, because we’re talking about software that makes the transformer architecture slightly more efficient or gathers better data to train the AIs on. But AIs are also running on chips, which are printed in these chip factories at Nvidia, and those factories have machines that are built by other machines that are built by other machines and ultimately go down to raw materials.

I think that something we don’t talk about very much, because it’ll happen afterward, is how hard it would be for the AIs to automate that entire stack—the full stack—and not just the software stack.

Nathan Labenz

The range of expectations that exists among sensible, thoughtful people who have engaged with this—how much is AGI going to speed up economic growth? It ranges from people who say it will speed up economic growth by 0.3 percentage points. So, it’ll be a 15% increase or something in current rates of economic growth, and I’d be very happy if it was that good.

Ajeya Cotra

People who say that at peak, the economy will be growing at 1,000% a year or higher than that—thousands of percent a year. So, it’s like a 100-, 1,000-, or 10,000-fold disagreement, basically, on the likely impact that this is going to have. It’s an almost unfathomable degree of disagreement among people who—it’s not as if they thought about this independently and haven’t had a chance to talk. They’ve spoken about this, they’ve shared their reasons, and they don’t change their minds, and they disagree by a 1,000-fold impact.

Nathan Labenz

You’ve made it part of your mission in life over the last couple of years to have really sincere, intellectually engaged, curious conversations with people across the full spectrum. Why do you think this disagreement is able to be maintained?

Ajeya Cotra

I feel like, at the end of the day, the different parties tend to lean on two pretty simple priors, or outside views, that are kind of different. I would say that the group that expects things to be a lot slower tends to lean on: “For the last 100 or 150 years in frontier economies, we’ve seen 2% growth.” Think of the technological change that has occurred over the last 100 or 150 years. We went from having very little—electricity was just an idea—to everywhere being electrified.

We had the washing machine, the television, the radio—all these things happened. Computers happened in this period of time. None of these show up as an uptick in economic growth. I think there’s this stylized fact that mainstream economists really like to cite, which is that new technology is sort of the engine that sustains 2% growth. In the absence of that new technology, growth would have slowed.

They’re kind of like, “This is how new technologies always are. People think they’re going to lead to a productivity boom, but you never see them in the statistics. You didn’t see the radio, you didn’t see the television, you didn’t see the computer, you didn’t see the internet, and you’re not going to see AI.” AI might be really cool; it might be the next thing that lets us keep chugging along. That’s one perspective. It’s an outside view they keep returning to.

There’s also a somewhat more generalized idea: things are just always hard and slow—way harder and slower than you think. It’s our experience in our personal lives that it’s awfully hard to achieve things at work. Things that might seem so straightforward to other people prompt them to ask, “Why haven’t you finished this yet?” And you’re like, “Well, I could give you a very long list.” What was it—Hofstadter’s, not Murphy’s law?

Nathan Labenz

Murphy’s law is that anything that can go wrong will go wrong. I think this is our experience in our personal lives: it’s awfully hard to achieve things at work, even things that might seem so straightforward to other people. They’re like, “Why haven’t you finished this yet?” And you’re like, “Well, I could give you a very long list.”

Or Hofstadter’s law: it always takes longer than you think, even when you take Hofstadter’s law into account. Or the programmer’s credo—this is my favorite one: “We do these things not because they are easy, but because we thought they would be easy.” [laughter]

Ajeya Cotra

So, there’s just this whole cloud of: it’s naïve to think things can go crazy fast. If you write down a story that seems perfect and unassailable for how things will be super easy and fast, there are all sorts of bottlenecks and drag factors you inevitably failed to account for in that story. That’s kind of that perspective.

Then I think the alternative perspective leans a lot on much longer-term economic history. If you attempt to assign reasonable GDP measures to the last 10,000 years of human history, you see acceleration. The growth rate was not always 2% per year at the frontier. 2% per year is actually blisteringly fast compared to what it was in 3000 BC, which maybe was 0.1% per year. So, the growth rate has already multiplied manyfold, maybe an order of magnitude, maybe 2.

I think the people in the slower camp tend to feel like the exercise of using long-run historical data is just too fraught to rely on. But people in both camps do agree that the Industrial Revolution happened, and the Industrial Revolution accelerated growth rates a lot. We went from having growth rates that were well below 1% to having 2% per year growth rates.

People in the faster camp tend to lean on the long run and on models that say the reason we had accelerating growth in the long run was a feedback loop where more people can try out more ideas and discover more innovations. That then leads to food production becoming more efficient, which then leads to a larger supportable population, and then you can rinse and repeat, and you get super-exponential population growth.

That perspective says that if you can slot in AIs to replace not just the cognitive, but the cognitive and physical—the entire package—and close the full loop of AIs doing everything needed to make more AIs, or AIs and robots doing everything needed to make more AIs and robots, then there’s no reason to think that 2% is some sort of physical law of the universe. They can grow as fast as their physical constraints allow them to grow, which are not necessarily the same as the constraints that keep human-driven growth at 2%. That’s the justification they provide for their perspective, in broad strokes.

Nathan Labenz

But why is it that, even after communicating this at great length to one another, they don’t converge on uncertainty or say it’ll be something in the middle because there are competing factors? Why do they continue to be reasonably confident about quite different narratives about how things will go?

Ajeya Cotra

I’m honestly not sure. I think maybe one part of it is that I’m partial to the “things will be crazier” side of things, so I’m not sure I’ll be able to give a perfectly balanced account. But one thing I’ve noticed among people who think it’ll be slower is that their worldview kind of has a built-in error theory of people who think things will go faster.

That worldview is not just, “Things will keep ticking along,” but, “Everyone thinks there will always be some big new revolution that makes things speed up, and they’ve always been wrong.” So, there’s that dynamic, which is, from their point of view, totally reasonable. Even if there isn’t some super-knockdown argument in terms of your interlocutor where you can point to a mistake they’ll accept, or even if you look at the story and think it’s kind of plausible, you still have this strong prior: someone could have made the same argument about television, and someone could have made the same argument about computers. None of these played out.

I think that’s a big factor. I also think these are complicated ideas, and there hasn’t been that much dialogue. I think there could be more dialogue, and I think there could be more dialogue that’s trying to ground things in near-term observations as well. But I think that’s a big part of it.

They have an error theory built in that makes it so the object-level conversation about, “Okay, here’s how the AI could make the robots, and here’s how the robots could bootstrap into more robots, and so on”—that whole way of thinking doesn’t feel very legitimate or interesting. They sort of have a story where that type of thinking always leads to a bias toward expecting things to go faster than they actually will, because it’s hard for that kind of thinking to account for all the drag factors and bottlenecks.

Whereas I think, on the other side, people who think things will go faster feel like everyone is always blanket-assuming that there are going to be bottlenecks. Then they bring up specific bottlenecks. Those specific bottlenecks, when you look into them, don’t seem like reasons to think that they might slow things down from some absolute peak of 1,000% growth, but they’re not reasons to think that 2% is where the ceiling is, or even that 10% is where the ceiling is.

So, they also have this kind of error theory about the bottlenecks objection. It’s incredibly decision-relevant to figure out who is right here. I think almost all of the parties to this conversation, if they completely changed their view—if the people who thought it was going to be a 1,000% speed-up decided it was going to be 0.3%—would probably change what they’re working on, or think it was a decisive consideration against everything they were doing previously.

And vice versa, if people came to think that it would be a 1,000% speed-up, then they’d probably be a whole lot more nervous and interested in different kinds of projects. So, how can we potentially get more of a heads-up ahead of time about which way things are going to go? I guess it seems like sharing theoretical arguments hasn’t been persuasive to people. Yeah.

Rob Wiblin

Is there any kind of empirics that we could collect as early as possible?

Ajeya Cotra

So, one thing that I think will not address all of this, but is a step in the right direction, is really characterizing how, why, and if AI is speeding up software and AI R&D. METR came out with an uplift RCT, which I think was the first of its kind, or at least the largest and highest quality, where they had software developers split into 2 groups. One group was allowed to use AI, and the other group was disallowed from using AI. They studied how quickly those developers solved issues, like tasks on their to-do list. It actually turned out that, in this case, AI slowed down their performance, which I thought was interesting.

I don’t expect that to remain true, but I’m glad we’re starting to collect this data now. I’m glad we’re starting to cross-check between benchmark-style evaluations, where AIs are given a bunch of tasks and scored in an automated way, and evidence we can get about actual, in-context, real-world speed-ups. I really want to get a lot more evidence about that of all kinds, like big uplift RCTs. It would be great if companies were into internally conducting RCTs on their own rollouts of internal products to see whether teams that get the latest AI product earlier are more productive than teams that don’t.

Even self-report, which I think has a lot of limitations, is still something we should be gathering. So, I guess my high-level formula would be to look at the places where adoption has penetrated the most and start to measure speed-up in actual output variables. I think it would be really cool if there were a solar-panel manufacturing plant that had really adopted AI, and we started to see how much more quickly they could manufacture solar panels, or how much better they could make solar panels.

Nathan Labenz

Yeah, is it possible to do this at the chip-manufacturing level? I guess maybe that’s the most difficult manufacturing there is, more or less.

Ajeya Cotra

So, we might think that you’d get more of an early heads-up if you did something more straightforward, like solar panels. But we’d really like to be monitoring across all kinds of different manufacturing how much difference any of this is making. I think the most important thing, or the thing I ultimately care about, is the AI stack: chip design, chip manufacturing, manufacturing the equipment that manufactures chips, and then, of course, the software piece of it, too.

The software piece is the earliest piece. But I think we should be monitoring degree of AI adoption, self-reported AI acceleration, RCTs—anything we can get our hands on—for the entire stack, because I think the moment when the sort of AI futurists think things are likely to be going much, much faster sort of coincides with when AI has fully automated the process of making more AI. That’s really something to watch out for.

And then, on a separate track, you also want to be looking at the earliest power users, no matter where they are, just because you can get insight that transfers to these domains.

Nathan Labenz

Is there anything else we can do? I don’t know. I’m really curious about this. Do I understand right that last year you put out a request for proposals?

Ajeya Cotra

Yeah. So, I put out a pair of requests for proposals in late 2023. One of them was on building difficult, realistic benchmarks for AI agents. At the time, very few people were working with AI agents, and only a couple of agentic benchmarks had come out, including METR’s benchmark, which I discussed on the show last time.

I was really excited about it. I felt like it was a moment to move on from giving LLMs multiple-choice tests to giving them real tasks, like “Book me a flight,” or “Make this piece of software work.” Write tests, run the tests, and iterate until the thing actually works. That was a very new idea at the time, but the time was sort of right for that idea, and there were a lot of academic researchers who were excited about moving into the space.

We got a lot of applications for that arm of our request for proposals, and we funded a bunch of cool benchmarks, including Cyber, which is a cyber-offense benchmark that’s used in a lot of standard evaluations now. But then we also had this other arm, which was basically types of evidence other than benchmarks, like surveys, RCTs, and all the things we talked about. We got much less interest for that.

I think it just reflects that it’s harder to think of good ways to measure things outside of benchmarks, even though everyone agrees benchmarks have major weaknesses and consistently overestimate real-world performance. Benchmarks are sort of clean and contained, and the real world is messy and open-ended.

But one thing that I’m excited about that came out of the second RFP is that the Forecasting Research Institute is running this panel called LEAP, which is the Longitudinal Experts on AI Panel. They take 100 or 200 AI experts, economists, and superforecasters, and have them answer a bunch of granular questions about where AI is going to be in the next 6 months, in the next year, and in the next 5 years.

That includes benchmark scores, but also things like whether companies will report that they’re slowing down hiring because of AI, whether an AI will be able to plan an event in the real world, or things like that. I’m very excited about that, and I think that having people make subjective predictions, explain how those predictions are connected to their longer-run worldviews, and then check over time who’s right might be the most flexible tool we have.

I’m very excited to see where LEAP goes. But I think it is challenging to get indicators that are clearly early warnings, so that we can actually do something about it if the people who are more concerned are right, but that are also clearly valid and not easy to dismiss on the other side as just not realistic enough to matter.

Nathan Labenz

So, as part of this, you’ve been thinking about—I guess one way that this could really go wrong is if the companies that are developing cutting-edge AI may begin to see internally how much it’s helping them, and that perhaps it’s speeding them up enormously, but they may decide not to share that information with the rest of the world.

Ajeya Cotra

And they may decide not to release those products. Like, if there’s one company that’s well ahead of the others, then, in AI 2027, it was sort of depicted that the company that was ahead in the AI race was so far ahead of its competitors that it could afford to just keep its best stuff internal and only release sort of less-good products to the rest of the world. It could afford it in the sense that it didn’t need to make money by selling the product.

Nathan Labenz

Its competitors were far enough behind that they couldn’t undercut it or compete with it by releasing a better product.

Ajeya Cotra

In the story, the company in the lead, OpenMind, is basically just releasing products that are slightly better than the state of the art of its competitors.

Nathan Labenz

I'm saying they're so far ahead that they can just choose to always have their product be somewhat better. They can release whatever level of their own internal machine would be best for the external world.

Okay, but I guess it would be unfortunate if there are people who do know this but the broader world doesn't get a heads-up, and so we could have known 6 months or a year earlier in what direction things were going, but that was kept secret. I guess maybe the leading AI company would prefer to keep it secret, but the rest of us would prefer that the government has some idea of what's going on.

So you've been thinking about what sort of transparency requirements could be put in place that would require companies to release information that would give the rest of us clues as to where things are going. What sort of transparency requirements could those be?

Ajeya Cotra

Yeah, so I think there's a whole spectrum of evidence about AI capabilities where, on the one hand, the easiest to test but least informative is benchmark results. Companies do release benchmark results when they release models right now. So they say, you know, Claude Opus 4 was released, and they have a model card that says it has this score on this hacking benchmark, this score on the software engineering benchmark, and so on, as part of a report about whether it's dangerous. GPT-5 had the same thing.

I think that's great that they do that, but in my ideal world, they would release their highest internal benchmark score at some calendar-time cadence. So every 3 months they would say, "We've achieved this level of score on this hacking benchmark, this level of score on the software engineering benchmark, and this score on an autonomy benchmark." That's because, as you said, danger could manifest from purely internal deployment.

If they have an AI agent that's sufficiently good at AI R&D, they could use that to go much faster internally, and then other capabilities, and therefore other risks, might come online much faster than people were previously expecting. So it's not ideal to have your report card for the model come out when you release it to the public, unless there's some sort of guarantee that you're not sitting on a product that's substantially more powerful than the public product.

So maybe it's fine to release your model card and system card along with the product if you also separately have a guarantee that you won't have too much of a gap between the internal and the external. That's on the end of things that are currently discussed. It's kind of how I would tweak information that's currently reported to be somewhat more helpful for this concern.

But then there's a bunch of other stuff that's not currently reported that I think it would ideally be really great to know. Stuff like how much they're using AI systems internally and how they're using them. One thing I'm very interested in is that companies will sometimes report, kind of to brag, the percentage of lines of code that are written by their AI systems.

Various CEOs have said, "Internally, 90% of our lines of code are written by AIs," and things like that. I think it'd be great to have systematic reporting of those kinds of metrics, but those metrics aren't the ideal metric I'd be interested in.

One thing I'm interested in is what fraction of pull requests to your internal codebase were mostly written by AI and mostly reviewed by AI. So humans are not involved for the most part in both sides of this equation. I'd be very interested in watching that number climb because I think it's an indication both of AI capabilities and of how much deference they're giving to AIs.

Eventually, if things are going to go crazy fast, the AIs have to be doing most things, including most management, approval, and review, because if humans have to do that stuff, then things can only go so fast. So I really want to track how much higher-level decision-making authority is being given to the AIs in practice inside the companies.

I think there are probably a bunch of other things that we could send, basically as a survey. How much do you use AIs for this type of thing, for that type of thing? How much speedup do you get? Subjectively, how much do you think you get? If you're running any internal RCTs, I would of course love to know the results of those.

Rob Wiblin

What about just requirements that, insofar as they're training future generations of AI models, they have to reveal to at least some people in the government how they're performing on normal evals of capabilities, so they can see the line going up even if they're not releasing them as products for whatever reason? If the line starts climbing upward, far above previous expectations, that could lead them to sound the alarm.

Ajeya Cotra

Yeah, I think that is a good thing to do, but I sort of don't think that just benchmarks alone will actually lead anyone to sound the alarm, because the thing with benchmarks is that they saturate.

Nathan Labenz

Yeah, but they always have that S-curve shape.

Ajeya Cotra

They always have the S-curve shape, and the benchmarks we have right now are harder than the previous generation of benchmarks, but it's still far from the case that I feel confident that if your AI gets 100% on all these benchmarks, then it's a threat to the world and it could take over the world. I still think the benchmarks we have right now are well below that.

What's probably going to happen is that these benchmarks are going to get saturated, then there's going to be a next generation of benchmarks people make, and then those benchmarks are going to tick up and get saturated. So I think we need some kind of real-world measure before we can start sounding the alarm.

The ultimate real-world measure is actually just observed productivity, right? If they're seeing internally that they're discovering insights faster than they were before, then that's a very late but also very clear signal. That's the point at which they should definitely sound the alarm, and we should know what's happening. So, yeah.

Nathan Labenz

Yeah, how is this idea being received by the companies? On the one hand, it seems like transparency requirements are the regulatory instrument that the companies have objected to the least. It's the one that they've been most willing to tolerate.

On the other hand, the whole message of this is, "We don't trust you to share information with the rest of the world, and we think that you might screw us over by rushing ahead and deliberately concealing that." I can imagine that could be a little bit offensive to them, or at least—

Ajeya Cotra

[laughter]

Nathan Labenz

If that is their plan, then they probably want to find some excuse for not having this kind of oversight.

Ajeya Cotra

Yeah, I think the response just tends to differ based on the actual information that's being asked for. Benchmark scores they already release. Like I said, they release them at the point of releasing a product, which I think is fine for now, but I would like to move to a regime where they release benchmark scores at some sort of fixed cadence, even if they don't have a product release.

Benchmark scores are not considered sensitive information, but the other stuff that I think is a lot more informative on the margin is much more fraught, right? They don't necessarily want to share with the world the rate at which they're gaining algorithmic insights, because you want to maintain some mystery about that for competitive reasons.

It's risky for you if it's a little bit too fast, because then, I don't know, competitors will start paying more attention to you, trying to copy you, and trying to find out what's going on. It's also risky for you if it's too slow, because then that's kind of embarrassing.

Nathan Labenz

Investors lose heart.

Ajeya Cotra

Yeah, investors lose heart. And another thing I didn't mention earlier is that I would really like them to be reporting their most concerning misalignment-related safety incidents.

Has it ever been the case that, in real-life use within the company, the model lied about something important and covered up the logs? I really want to know that, but then of course it's clear that reporting that is very embarrassing to companies.

One thing that might help here is that there are a number of companies now, so perhaps they could report their individual data to some sort of third-party aggregator that then reports an anonymized overall industry aggregate score. But I don't think that solves all the issues, because there are few enough of them that people would be able to guess.

So I think there's a lot of competitive challenges, IP-sensitivity challenges, and PR challenges to overcome here with some of the more penetrating internal information. But I think it's important enough to the public interest that we should try and find a way to navigate that.

Nathan Labenz

Yeah, so it's not unusual for government agencies to be able to basically demand commercially sensitive information from companies for regulatory or governance purposes.

Ajeya Cotra

Actually, at one point, when I was in the Australian government, I was at the Productivity Commission, which had extraordinary subpoena powers to basically demand almost any documents from any company in the country. I rarely used the power, and it wasn't the only agency that had that capability. I guess people were proud of the fact that we had that authority.

Nathan Labenz

And what kinds of things would you ask them?

Ajeya Cotra

I never actually saw this power being used. I guess people were proud of the fact that we had that authority. Usually, you would do it for competition reasons: trying to tell whether companies were potentially colluding, or whether there was an insufficient degree of market competition and there would be reason to intervene. I would imagine there are almost certainly government agencies in the U.S. that have a similar remit.

If they could actually keep that kind of information secret, maybe companies would be happier to share it with people who specialized in reading and comprehending this data and figuring out what to do with it.

Nathan Labenz

Yeah, I think that could be a solution, but I'm a little skeptical. I think releasing this information publicly is probably a lot better than releasing it just to a government body, basically because we're building the plane of AI safety research as we're flying it. It's not like there's a box-checking exercise that any kind of government agency—often understaffed, especially with technical staff—could do.

It's more like we want this information out there in the open, and then we want people to do some involved analyses of it. Our sense of what information we even want is probably going to be shifting over time, and it'll probably go better if there's a robust kind of external scientific conversation about what indicators we want to see, what that would mean, and when we should trigger alarm.

If that's all being routed through governments with 10 people or even 50 people who have to deal with it, I think it would be very hard for them to interpret the evidence quickly enough and well enough, be confident enough to sound the alarm, and then have people actually listen to them. If I imagine sounding the alarm on something like the intelligence explosion, I picture it having to be a society-wide conversation, kind of like sounding the alarm about COVID.

Something I have in my mind is when Joe Biden had that disastrous debate performance that led to weeks of conversation, which ultimately led to him being removed from the ticket. It would have been very hard, I think, for a small, narrow group of people entrusted with the authority to make the same thing happen.

Because I guess you want common knowledge, and you want lots of attention focused on the issue, as well as just some technocrats being aware. You also want the opportunity for a bunch of technical experts who may not be paying that much attention now—because maybe they think this stuff is all science fiction—to jump in at that moment and offer their takes.

I think it would be very powerful if someone like Arvind Narayanan, who's known for being very skeptical of these stories, actually looked at the data, changed his mind, and said, “Oh, yeah, this is happening now, and it's dangerous.” It's very hard to get those kinds of common-knowledge dynamics if everything is just sent to governments.

That said, of course, I think sending things to governments is better than not sending them anywhere. So, insofar as plan A would be that we want them to be sharing this information such that anyone in the public can find out, I guess they'll probably resist legislation imposing this to some extent. I guess for partially legitimate reasons, it is probably going to be frustrating for them.

How high on the list, insofar as people are trying to set priorities for what sorts of asks you make and which sorts of fights you pick, would this be for you?

Ajeya Cotra

I think I laid out a whole spectrum of ideal information-sharing practices. I don't think going all or nothing on that whole package is a top-priority fight to pick, but I think the algorithm of thinking really hard about what pieces of information we would want to know in order to know for ourselves if the intelligence explosion was happening—and then getting the highest-value items on that list, or the biggest-bang-for-buck items on that list—feels very high.

I think that's the strategy that people working on AI safety-related legislation have landed on. The RAISE Act in New York and SB 53 in California are both quite transparency-oriented and both oriented around, for example, whistleblower protections, which are an important policy plank underlying transparency.

Nathan Labenz

Do you think that information about an emerging intelligence explosion might just leak out to the public anyway, because staff at the companies would feel uncomfortable with the proceeding in secret?

Ajeya Cotra

I think that's very plausible. I still think information that leaks in the form of rumors at San Francisco tech-bro parties doesn't have the ability to impact policy and decision-making all the way in D.C., London, or Brussels in the same way as information that is clearly unrefuted, very salient, and official.

The AI safety scene in the Bay Area has benefited from having close social ties to people who work at AI companies, which gives us a sense of what might be coming around the corner. But that's not something you can use to really pull an alarm or advocate for very costly actions.

I think it isn't really enough. We need more.

Rob Wiblin

So let's imagine that, via whatever mechanism, society does get a heads-up that we're starting to see the early stages of an intelligence explosion. What would we do with that heads-up?

Ajeya Cotra

Yeah, so one extremely important factor is: at that point in time, how good are AI systems at everything besides AI R&D? The alarm has sounded, and we've learned that AI has fully or almost fully automated R&D at the leading AI lab, perhaps all the AI labs. This is causing those labs to go way faster than they were going with mostly human-driven progress in the previous era.

At that point in time, whatever AI progress you thought was going to be made by default in the next 10 years, the next 20 years, or the next 30 years might be made in a year or two, or even 6 months, depending on how much AI is speeding everything up. At this stage, AIs might not be that dangerous, but we might be about to move very quickly through the point in time where they're not so dangerous to the point in time where they have godlike abilities.

I think what we want to do as a society, if we gain confidence that we're at the starting point of this intelligence explosion, is to redirect as much of that AI labor as we can from further AI R&D to things that could help protect us from future generations of AIs, both in terms of AI takeover risk and also in terms of a wide range of other problems that might be created for society by increasingly powerful AI.

At that point, it's still not in the narrow, selfish interests of whichever company is in the lead to do that, because if they were to slow down unilaterally, someone behind them could catch up. But hopefully, if the alarm has sounded and we have a clear picture that we have 6 months, 12 months, or 18 months until radical superintelligence, this might be a window of opportunity to coordinate, to use AIs for protective activities instead of further AI capability acceleration.

Nathan Labenz

So the challenge we have is that AI is becoming much smarter very, very quickly, and we feel very nervous about that. But I guess the opportunity that's created is that we have a lot more AI labor and much smarter potential researchers than we did before. So why don't we turn that new resource toward solving this problem that, at the moment, we don't really know how to fix?

Ajeya Cotra

It's a little bit—I guess some people who are not too worried about AI look at society as a whole, or they look at history, and say, “Well, technology has enabled us to do all kinds of more destructive things, but we don't particularly feel like we're in a more precarious situation now, or at much greater personal risk now, than in 1900 or in 1800, because advances in destructive technology have been offset by advances in safety-increasing technology. On balance, probably things have gotten safer.”

The idea is: can we potentially pull off the same trick in this crunch-time period? It's going to be a vertiginous time, but perhaps we could.

Nathan Labenz

Yeah, and I think a lot of people who are more concerned about AI risk are very dismissive of this plan. It just sounds like a plan that's really flying by the seat of your pants, expecting the thing that's creating the problem to solve the problem.

Ajeya Cotra

But in a sense, I do think humanity has repeatedly used general-purpose technologies that both created problems and helped solve those problems. Take automobiles, something as mundane as that. Cars created the opportunity for carjackings and drive-by shootings, and empowered bad actors in various ways. But of course, if the police and law enforcement have cars as well, that is a balance. It’s not like when you imagine a future with some crazy new advanced technology and imagine all the problems it creates, it’s easy to imagine with the same level of detail and fidelity all the responses to those problems that are also enabled by that technology.

You could imagine someone worrying about the rise of fast vehicles and neglecting to think about how all the ways those fast vehicles cause bad things could be kept in check by people using vehicles for law enforcement and similar purposes. Similarly, with computers, you can hack things with computers, but computers also enable you to do a lot of automated monitoring for that kind of hack. They also enable different kinds of law enforcement. You couldn’t imagine a police force not using computers.

So I do think the basic principle is sound: if you’re worried about problems created by technology, one of the first things on your mind should be, “How can you use whatever that new technology is to solve those problems?”

But I think that this is an especially narrow window to get this right. You’re not imagining cars creating broad-based, rapid acceleration of all sorts of new technologies, and potentially just a 12-month window, 2-year window, or 6-year window before everything goes totally crazy. So I do think that it’s important not to blow through that window, to monitor as we’re approaching it, and to monitor how long we have.

But, yeah, I think I’m fundamentally fairly optimistic about trying to use early transformative AI systems—early systems that automate a lot of things—to automate the process of controlling, aligning, and managing risks from the next generation of systems, which then automate the process of managing those risks from the generation after, and so on.

Nathan Labenz

It’s interesting you say that this approach has often been dismissed, because I feel like it’s very in vogue now. I hear about this proposal every couple of days. Someone presents it, or I read something about it in one guise or another.

Ajeya Cotra

I guess one reason why, years in the past, it might have felt unpopular is that people were mostly focused on the issue of misaligned AI. They were concerned about an AI that had it in for you and would like to take over if it had the opportunity. That’s maybe the worst application of this out of all of them, because there you’re asking the AI to align itself, but you don’t know whether it’s assisting you or trying to undermine you.

You could try to make that work. People have suggested proposals where you could try to get useful, honest work out of an AI that doesn’t want to help you. But it’s a lot easier to see how you potentially solve problems other than alignment. If you assume, well, the alignment part—we feel like we’ve got a good handle on that—but there’s a huge list of other problems being created during the intelligence explosion, like the fact that AI, if people get access to it, could invent other kinds of destructive technologies that we don’t yet have good countermeasures for. In that case, it’s clear how the AI could help you figure out what the countermeasures ought to be.

So I don’t think that I agree with this. I do think misalignment—the prospect that these early AIs, these early transformative AIs, are misaligned—is a huge obstacle to this plan that needs to be shored up, handled, and specifically addressed. I don’t think that it necessarily bites harder for getting the AIs to do alignment research than for getting the AIs to do anything else helpful, because if they have it out for you, they don’t necessarily want to help you shore up your civilization’s defenses.

If you’re imagining trying to get a hardened, misaligned AI to help you with biodefense, if it’s misaligned and, for example, wants the option of threatening you with a bioweapon in its arsenal in the future, it would similarly have an incentive to do a bad job at that as it would to do a bad job at alignment research. In general, I think there’s one big concern, which is: will the AIs that we’re trying to use at that point in time have motivations that give them incentives to undermine the work we’re trying to get them to do?

I think they certainly would have incentives to undermine alignment research if they were misaligned, but I think they would also have incentives to undermine efforts to make ourselves more rational and thoughtful—AI for epistemics—because if we’re more rational and thoughtful, then maybe we’ll realize they’re probably misaligned, and that would be bad for them. They would also have an incentive to undermine our defensive efforts, because that would make it harder for them to take over.

Nathan Labenz

That makes sense. I think the distinction I was drawing is for people who thought that the alignment problem was extremely hard to solve and that we were way off track to solving it. The idea of getting the AI to solve the problem is kind of self-contradictory, because I wouldn’t trust the AI at all. Anything that it proposed, I would assume was sabotaging us.

If you’re on the side of thinking, well, the alignment problem is actually the easier part of things—I think that’s a relatively straightforward technical problem that we’re on track to solve—but there’s this laundry list of 10 other issues, then it’s very obviously, “We’ll have the brilliant AGI, so why don’t we just use that to solve all the other things?” I’m also inclined to trust it and believe it.

Ajeya Cotra

I do think that if you are not worried about alignment at this early stage, everything becomes easier. It becomes an even more attractive strategy and path. But I think the canonical “using AI for AI safety” or “using AI for defense” plan does imagine that we’re not sure at the beginning that they’re aligned.

We may not be highly confident that they’re extremely misaligned, fully power-seeking, and looking to take over at every opportunity, but we’re not imagining that we know with confidence that we can trust them. So figuring out how to create a setup where we use control techniques, alignment techniques, interpretability, and whatever other tools are at our disposal to get to the point where we feel good about relying on their outputs is a crucial step to figure out, because it either bottlenecks our progress—we’re checking on everything all the time and slowing things down—or it doesn’t bottleneck our progress, but we hand the AIs the power to take over.

Nathan Labenz

Which specific problems arising from the intelligence explosion are you envisioning wanting to get the AGI to help us out with?

Ajeya Cotra

One obvious one is AI alignment. How can we ensure that either these AIs that we’re using to help us right now, or future generations of AIs that they help us create, and future generations that those AIs help us to create, are motivated to help humans, are honest, are basically doing what we say, and are steerable? That is the foundation of everything else.

But then there are also other things that are not really about AIs at all, and are just about broad societal defenses. If we think that the advent of extremely powerful AI will create a flood of new cyber vulnerabilities that are quickly discovered in a bunch of critical systems, like weapon systems and the power grid, can we preemptively use those same AIs that are good at finding those vulnerabilities to find and patch them before bad actors can use the AIs to find them?

Another thing is biodefense. You had my colleague Andrew on your podcast recently, and he talked about his ambitious plan to rapidly scale up detection of novel pathogens, rapidly scale up medical countermeasures when they’re detected, and rapidly scale up the manufacturing of PPE, clean rooms, and things like that. If we have AI systems that are good at that kind of research problem, and maybe we have robots at that point, a lot of that manufacturing itself can be automated and can happen a lot faster than if humans had to do that stuff. That would be a big boon to biodefense.

Then there are some somewhat more speculative things. You can think of this as a kind of defense, maybe a psychological defense. Can we use AIs to make our collective decision-making a lot smarter, wiser, and better? Can we make it so that we’re better at finding truth together? Can we make it so that we’re better at coming to compromise policy solutions that leave lots of people happy? How do you ensure that advances in AI don’t lead to a war between the US and China, that kind of thing?

Or even that, too, but even more mundanely, stuff like over the last 10, 15 years, social media has led to a degradation of political discourse.

Nathan Labenz

Could AI tools help you just kind of find the policy from among the vast space of possible policies that a large number of people actually like and can credibly put trust in, and so on?

So, I interviewed Will MacAskill and Tom Davidson from Forethought earlier in the year. The organization has a long list of what they call grand challenges, which they suspect are all probably amenable to this kind of AGI labor during crunch time. I think other ones are ensuring that society doesn't end up locked into particular values prematurely in a way that cuts off our ability for further reflection and changing our mind.

There's the potential use of AI or AGI, insofar as it's very steerable and follows instructions, in power grabs by the people who are operating it. There's space governance: if we actually do start to be able to use resources in space, how would we share them? How would we divide them such that, in particular, there's not conflict ahead of time because people anticipate that once you start grabbing resources in space, you're on track to become overwhelmingly dominant?

There's epistemic disruption, which you mentioned. I guess there are new competitive pressures—the concern that you can end up in a sort of Malthusian situation if you have competition between many different AIs—and possibly some others that are missing here. But there are many different ways that we could potentially apply it. We don't know which of these are going to loom large at the time. Some of them might feel like they've been addressed, or perhaps we were hallucinating issues that aren't so severe.

Ajeya Cotra

Yeah, I agree. I think all of those problems that Tom and Will highlighted seem like real problems to me. I think maybe my approach would be to, from our current vantage point, lump a lot of that under AI for helping us think better and helping us find solutions that we're mutually happy with.

So, it's AI for coordination, compromise, negotiation, truth-seeking—that cluster of things—because I think something like the question of space governance is: How do we divide up the resources of space if there are some existing factions that have an existing distribution of power? No one really wants the destruction that comes from everybody racing as hard as possible to get there first. But there's a complicated space of negotiated options beyond that, and I think AIs could potentially help a lot with that sort of thing.

Nathan Labenz

So, you said in your notes that you think this approach is basically what all of the frontier AI companies say their safety plan is, more or less. Is that right?

Ajeya Cotra

Yeah, I would think so. If you look at public communications from at least OpenAI, Anthropic, and Google DeepMind, this sort of jumps out, more or less, in these different cases. But in all of their stated safety plans, you see this element of: as AIs get better and better, they're going to incorporate the AIs themselves into their safety plans more and more.

Some are more explicit than others about expecting some sort of specific crunch time that occurs when AI is rapidly accelerating AI R&D. But everybody is picturing AIs playing a heavy role in the safety of future AIs.

Nathan Labenz

What assumptions are necessary for this approach to make sense? Or what kinds of setups could actually just make it a bad plan?

Ajeya Cotra

Yeah, I think fundamentally, you need it to be the case that there exists a window of opportunity where, before AIs are uncontrollably powerful or have created unacceptable levels of risk, they are really capable and really change the game for AI safety research. There needs to be some meaningful window of time where you can notice as you're approaching it, and even by default, without crazy slowdown, it lasts at least 6 months or a year.

If you think instead that once your AI sort of hits upon some generality threshold, it becomes crazy superintelligent within a matter of days or weeks, this plan doesn't work because you wouldn't even notice before it's too late.

I think there can also be unlucky orderings of capabilities where this plan wouldn't work. You could have AIs that are really specifically good at AI R&D and really not good at anything else—not even AI safety research that's very similar to AI R&D. They're just extremely good at AI R&D.

Maybe the only thing they're good at is making it so that future generations of AIs have better sample efficiency and can learn new things more efficiently. Then you could have a period of 6 months or a year where you know this is happening and you have these AIs, but you're still hurtling toward a highly general superintelligence without being able to use these AIs for anything else necessarily because they're just not good at anything else.

Nathan Labenz

There's something that's a bit self-contradictory about that because an AI that's extremely smart, but all it can do is improve the sample efficiency of the next model, isn't, in a sense, very troubling in itself because it doesn't have general capabilities. That kind of model isn't going to be able to take over or invent other technologies. It's only at the point that it has the broader capabilities—the broader agency—that it's actually able to make problems.

But I guess you're saying you can have a long lead-up where that's all that it can do, and then at the last stage—

Ajeya Cotra

Yeah, and then at the last stage it might go back to the first scenario I talked about, where the narrow AIs that are just savants at AI R&D hit upon an algorithm in almost a blind search. It's almost like if you imagine AlphaFold: it's brilliant at figuring out how proteins fold, but it isn't broadly aware.

You could imagine such AIs, or an algorithmic search process, hitting upon an architecture or a training strategy that then can FOOM really quickly. So, in this lead-up, you're like, “Yep, AI is accelerating AI R&D. It's crunch time. We have 6 months left. We have 3 months left.” But these AIs are not the AIs that you can use for anything useful.

Nathan Labenz

Yeah, I guess many of the problems that we'd like to help with are social, political, or philosophical issues in some cases. What do you think the chances are that AI—I mean, the companies, I think, are working harder to make them good at coding and to make them good at AI research than any other particular thing?

I guess those are more concrete, measurable problems than solving philosophical questions. So, it seems like it is really a live risk that, unfortunately, the balance of capabilities will end up being pretty disadvantageous for this plan.

Ajeya Cotra

Yeah, I think that the further afield you go from work that looks like doing ML research and software engineering, the greater a penalty there'll probably be. The AIs are currently much better at helping my friends who do ML research all day than they are at helping me, where I do weird thinking, go on these kinds of podcasts, and write emails to people making grant decisions and stuff like that.

They're much worse at that stuff. You can see already that they've got a very specialized skill profile. Fortunately, I do think that at least in AI safety, there's a big chunk of AI safety research that does look very similar to ML research. And I do think that my friends who are getting big speedups from AI are safety researchers, and they're doing the kinds of work—control, alignment, and so on—that I think will be some of the most important things you want these AIs to be helping with at the very beginning.

But stuff like AI for epistemics, AI for moral philosophy, AI for negotiation, and AI for policy design may just not be that good. It doesn't necessarily have to be good by default, and that's a big concern of the plan.

Nathan Labenz

I guess another worry would be that the AI models end up being able to cause trouble before they end up being capable enough to figure out solutions. A classic case there would be: imagine that we put a lot of effort into—I guess it would be a bit stupid to do this—but we put a lot of effort into training an AI model that's extremely good at developing new viruses or new bacteria, basically changing diseases to make them worse.

I mean, there are people who—

[laughter]

—are using AI to develop new viruses. I guess they're using it to develop medical treatments, but that sort of stuff can then be repurposed for other things. But if that sort of highly specialized model arrives first, before you end up with a model that has a sufficient understanding of all of society and biology and medicine to figure out what the good countermeasures are, then we need a different approach than this one.

Ajeya Cotra

Yeah, and in general, I think of AIs doing defensive labor as a prediction about the world that you want to try to be thinking about as you make your plans. It's not a guarantee, and in many cases, the answer will be to specialize now in doing the kinds of things that might be hardest for the AIs to do then.

I think stuff like building a bunch of physical infrastructure to stockpile a bunch of PPEs and vaccines and things like that is a prime candidate for something that inherently takes a long lead time and that the AIs might not be that advantaged at when they're good at doing the scary things that it's meant to protect against.

Nathan Labenz

Yeah, that was going to be another concern of mine: insofar as AIs are very helpful, you might imagine that they’re very helpful at the idea-generation or strategizing stage, but they might still be quite bad at actually running a business or figuring out how to do all of the manufacturing. So, if they come up with a great strategy for countering a new bioweapon where they’re like, “Here’s the widget that you should use. Go and make 10 billion of them,” can you help us with that? It’s like, “No, I’m not very good at that.” [laughter]

Ajeya Cotra

Good luck. Yeah, I think that in general you should expect AIs to be much better at things where there are tighter feedback loops, where you can recognize success after a short period of time. That’s one of the reasons why they’re really, really good at coding, because you can train them on this very hard-to-fake signal of whether the code ran after you did whatever you did with it.

In general, I think idea generation versus actually executing on a 1-year plan has some of this element: you can read a white paper and be like, “Huh, yeah, that’s pretty good,” and you can push the thumbs-up button and generate an AI that’s pretty good at generating white papers that you think are neat and probably would work. But it’s much harder to train the AI to run the team of thousands of humans and robots that are actually executing on the plan.

Rob Wiblin

Why is the crunch-time aspect—or, you know, the intelligence explosion taking off—actually even relevant to when we would want to start doing this? Because you might just think that if AI can help us do research or do work to solve any of these problems, then as soon as it’s able to do that, we want to do it, whether or not an intelligence explosion is kicking off.

Ajeya Cotra

To some extent, that’s right. I think the reason that I focus so much on the intelligence explosion is twofold. One is that at that point I think we might have a pretty short clock to figure out a bunch of stuff. The default trajectory might look like 12 months to extremely powerful, uncontrollable superintelligence that can easily take over the world. So it changes our calculus: you might want to focus on very short-term things rather than things that have long lead times, at least at crunch time, if not before.

The other thing is I think crunch time can help alleviate some of the challenges we’ve been talking about with AIs not being good at the full spectrum of things we want them to be good at, because, by definition, at that point AIs are really good at further AI R&D. One of the things we could do with AIs that are good at AI R&D, at least in most cases, is try to direct their AI R&D toward filling out the skill profile of AIs and getting them to be good at some of the types of things that we want them to be good at that they aren’t so good at right now.

And so at that point you might have just much more capability at your disposal. It might be much more worth putting in the effort to fine-tune and scaffold and do all these other things to make your AI that’s good at moral philosophy, or your AI that’s good at biodefense.

Nathan Labenz

So you’re thinking about this strategy not just as a description of what other organizations potentially should work on, or as a description of what the AI companies are already planning to do, but also, I guess, because you think maybe it should influence what Open Philanthropy plans to do over the coming years. And potentially Open Philanthropy’s best play might be to have billions of dollars waiting at this relevant crunch time and then disperse them incredibly quickly, buying a whole lot of compute to get AIs to solve these problems.

Ajeya Cotra

Yeah, I mean, just like how right now 80% plus of our grant money goes to salaries to pay humans to think about stuff, do research, do policy analysis and advocacy, and all these other things, so too in a few years it might be the case that AIs are better than most of our human grantees, and our money should mostly be going to buying API credits or renting GPU time to get the AIs to do a similar distribution of activities.

Nathan Labenz

So an alternative approach to this would be that, at the point that we get a heads-up that we think an intelligence explosion is beginning to take place, we do everything we can to pause at that stage, to slow down, basically, to arrest that process. So rather than having to rush in 3 or 6 months to get the AIs to fix all of these issues, we buy ourselves a bunch more time. Why not adopt that as the primary approach instead?

Ajeya Cotra

Yeah, so I think that the plan I described is compatible with pausing at an intelligence explosion, right at the brink of an intelligence explosion. In fact, I would hope that we do that, because by default having 12 months to get everything in order is just not enough time. But I think of it as doing 2 things. One is making the pause less binary.

So if you think of the default path as almost 100% of AI labor going into further rounds of making AIs better, making more AIs, making more chips, and so on, and you think of a pause or a stop as 0% of the world’s AI labor going toward those activities, I think there’s a whole spectrum between 0 and 100%. Then I think of it as doing another thing, which is sort of answering the question of what you do in the pause.

You do all this protective stuff, and you have these AIs around to do it with. And you might think, once you have that frame of making the pause less binary and thinking really hard about what you do during a pause, you might often end up thinking, “Oh, it’s worth going a little bit further with AI capabilities,” because especially if we tilt the capabilities in a certain direction, we might at the end of that get AIs that are much better than they are right now at biodefense while still not being uncontrollable, still not being that scary.

And you can imagine a bunch of little pauses and little redirections and so on during that whole period. And I would hope that at some point in the period we do activities like policy coordination and so on that cause us to have longer in this sweet spot of AIs that are powerful enough to help with a lot of stuff, but not so powerful that we’ve already lost the game.

Nathan Labenz

So, yeah, we should probably clarify that although you think this is among our best bets, in an ideal world you think that we would go substantially slower through all of this, because as good a plan as this might be, we’ll really be white-knuckling it and not be confident that it’s necessarily going to work.

Ajeya Cotra

Yeah, so I think that if a really clear early-warning sign triggers that we are about to enter into this intelligence-explosion, fast-takeoff space, where we go in the space of 12 months from AI R&D automation to vastly superhuman AI, then I would vote for, at that time, shifting that trajectory to be 10 times longer or even longer than that, and trying to make that transition as a society in 10 years instead of 1 year, or 20 years instead of 1 year.

I still wouldn’t—and this is maybe a bit of a quibble—advocate for pausing and then hanging out for 10 years and then unpausing, because I actually think that slowly inching our way up is better than pause, then unpause, and then having a jump. But, going back to what we said about how your default expectations of trajectories influence what you think should happen, I think the default is going through this in 1 year, and I would certainly rather it be 10 or 15 or 20 years.

But I think that this frame of using AIs to solve our problems applies regardless of whether you’re white-knuckling it in 1 year, or maybe eking out an extra 2 months, or if you manage to get the consensus and the common knowledge that allows the world to step through it in 10 years.

Nathan Labenz

Yeah, I guess insofar as we’re slowing down to do something, this is a big part of the thing that we’re slowing down to do.

Ajeya Cotra

Yeah. So this is a big part of the company’s plan for technical alignment.

Nathan Labenz

If this doesn’t work out, why do you think it’s most likely to have failed for them?

Ajeya Cotra

I think that if it fails, it’s probably most likely to fail because they just didn’t actually do a big redirection from using AIs for further AI capabilities to putting a lot of energy toward using them for AI safety. Because they say this is their plan, but they don’t really have any quantitative claims about, at that stage, what fraction of their AI labor—or their human labor, for that matter—is going to go toward safety versus further acceleration.

And they’ll be facing tremendous pressure at that point from their competitors to stay ahead. And so my guess is that unless they have much more robust commitments than they have right now, they probably just won’t be directing that much of their AI labor. So if they have 100,000 really smart human equivalents, maybe only 100 of them are working on AI safety, which is maybe still more than they had before in human labor, but not that much compared to how quickly things are going.

Nathan Labenz

I guess you’re saying unless they have really strong commitments. But I guess other mechanisms would be that it’s legally required at this point, or the government basically insists that most of the compute go toward this, or at least that most of it not go toward recursive self-improvement.

Or I guess if the companies could reach some sort of agreement where they’re saying, “Well, we would all like to spend more of our compute on this kind of thing, so we’re going to have some contract where we’re going to spend 50% of all of our compute, and then we don’t lose relative position in particular.”

Ajeya Cotra

Yeah, I think that particular contract is probably going to run into big antitrust issues.

Nathan Labenz

A little illegal, but, oh yeah, maybe we could carve out an exception to antitrust for this one. Inasmuch as the government is taking a massive interest, they could help try to coordinate this one way or another. Yeah, I think that’s a possibility.

Ajeya Cotra

I do think it’s a bit tough. This is not the kind of thing that’s super easy to make laws about because it’s really not a box-checking exercise. When you write the legislation that says half the compute must be spent on safety rather than capabilities, what do you count as safety research? How are you enforcing this?

Do you have auditors in there asking all the team leads in the companies, “What are you working on?” and checking off that it’s 50% safety? I could imagine stuff like that. I think it would require extremely technically deep regulators that we just don’t really have right now, I think.

Nathan Labenz

I thought that you might say that the most likely reason for this to fail was that it just turned out that alignment is incredibly hard. You get egregious misalignment even at relatively low levels of intelligence, and we don’t really figure out how to fix that early enough to get useful work out of them.

Ajeya Cotra

Yeah, I think that’s a possibility. I don’t think it’s the most likely way it fails, in my view. I think the most likely way it fails is that they don’t go super hard on it.

But I think it’s also plausible that they’re just trying to get the AIs to help with alignment, and the AIs are just misaligned, and the control procedures and other things are ineffective. So they deliberately only help with further AI R&D and don’t help with alignment and safety, biodefense, and all these other things you’d want them to help with.

I would hope that at that stage the transparency regime is strong enough that that fact is broadcast really widely. That could inspire a change in policy that causes us to slow down, but then in that world it’s a bad world even if we do slow down a lot, because we’re just on our own. We have to do this stuff without the AIs’ help because we can’t get them to help us.

But I’m actually reasonably bullish about control techniques getting early AIs that are not super-galaxy-brain superintelligences to be helpful for a range of stuff that they’re good at.

Nathan Labenz

There’s another way that they could end up not making that much of an effort: if the window is relatively brief and it just takes a long time to get projects off the ground. They haven’t really planned this ahead, so they end up debating it back and forth. By the time they figure out that they actually do want to do this, I mean, I suppose it’s nominally in these various papers, but I wonder whether they’re actually thinking ahead about how this would feel and whether they’ll have the decision-making capability to decide to redirect enormous resources toward this other effort.

Ajeya Cotra

Yeah, I do think anything that requires a large corporation to be super-discontinuous in something it’s doing is facing big headwinds as a plan. I would hope that they’re smoothly increasing the amount of internal inference compute that’s going toward safety as the AIs get better and better, so that the jump doesn’t have to be huge at that final stage.

If we could elicit honest reports without creating perverse incentives, that’s something I’d want to know about. How much human labor is going to safety versus capabilities? How much internal AI inference is going to safety versus capabilities? How much fine-tuning effort is going to safety versus capabilities?

I think they have a much better shot if they’re stepping it up over time on some kind of schedule.

Nathan Labenz

Okay, so that’s the AI companies, who I guess we’re imagining would mostly be focused on the strategy for AI technical alignment. But you’ve been thinking about this more in the context of Open Philanthropy and what a niche it could fill. What would Open Philanthropy need to do if this was dumping billions of dollars onto this plan that became its mainline strategy?

Ajeya Cotra

Yeah, I think that for now the biggest thing we need to do is very similar to the biggest thing I think society needs to do to prepare for the intelligence explosion, which is really trying to track where we’re at right now in terms of how useful AIs are for the work that we do and the work our grantees do.

I think pushing ourselves to automate ourselves, and pushing our grantees to automate themselves, and tracking how good AI is at the stuff that Forethought does, how good AI is at the stuff that Redwood Research or Apollo does, and how good AI is at the stuff that our policy grantees do.

I think that’s one thing: socializing within ourselves that it’s probably a big deal when the AIs start to get really good at any given good thing we’re funding. Once we start to see signs of life there, we should be prepared to potentially go really big on that.

And like you said earlier, I do think crunch time isn’t 100% a special thing. We absolutely shouldn’t be waiting until crunch time to do anything at all. It’s just the prediction that crunch time is the point when a lot of things that were hard to automate before become easier to automate.

So, if it turns out, for example, that AI is really good at math research, which I think is plausible, then maybe we should be trying to deliberately shift our technical grantmaking toward more mathy kinds of technical grantmaking, because that is an area where you can churn a lot more. That’s just so much more tractable.

I think just having a function that is looking out for these things and is maybe poking Open Phil and Open Phil’s grantees to consider shifting their work toward more easily automatable things, and to consider repeatedly testing whether their work can be automated, is a big thing.

Then I could imagine, down the line, something like even just having separate accounting for the rest of our grantmaking versus grantmaking that is going toward paying for AIs for our grantees. We already pay for ChatGPT Pro subscriptions and ChatGPT API credits for tons and tons of grantees.

I think just making it a bit more salient in our minds: What fraction of our giving is going toward that? Do we endorse its size? Is there any place where we should be going bigger, and are we on track? Is the percentage climbing the way we think it should be? Does that seem in line with the way AI capabilities are climbing? Are we on track to have inference compute be a large fraction of our spending at that time if we think crunch time is going to start in 6 years?

Nathan Labenz

If I think about this kind of psychologically, I could imagine, if I was leading Open Philanthropy—or I guess if I was one of the donors being advised—and we did have these transparency requirements and we did start getting a sense that an intelligence explosion might be kicking off, I could imagine dithering for a long time rather than deciding to commit billions of dollars toward this. There’s only a particular amount of money, and there’s only a particular size of endowment.

I think I would be very scared that it would be going too early, or that this was a bad idea, or that we were going to have egg on our face afterward because it would turn out there were some early signs of an intelligence explosion, but it wasn’t really going to work out. Then we’d have spent $10 billion and have nothing left to show for it.

You’d feel really bad if you made that mistake. Does that sound like a plausible way for things to go?

Ajeya Cotra

Oh, totally. I mean, I think this is a very natural institutional thing. Even beyond just being scared of making a mistake on this front, organizations have particular ways they do things, and there are processes.

Right now, Open Phil’s process for grantmaking looks like this: Usually someone fairly junior gets an opportunity come across their desk, either through one of our open calls or through some contact they have. That junior person pulls together some materials to convince their manager it’s a good fit, and then that manager sort of convinces someone higher up that it’s a good fit.

You can have 2 layers, 3 layers, or sometimes 4 layers of information cascading up the decision-making process that we have in place as an organization, and then it’s approved.

If the right thing to do is to spend $1 billion on some particular strain of work that’s super-automatable, that isn’t even something you would trust some random junior person to make the call on. You might need to have just a different process for that. You need to figure out what that process would look like, but I think that would be one thing to figure out.

Nathan Labenz

I guess for this sort of incredible scaling of funding and effort to take place, it would have to be that you're going to be incredibly bottlenecked on people, or there won't be that many more people involved. So, the AIs would have to be not just doing the object-level work, but also deciding what problems to work on.

Ajeya Cotra

Yeah, they have to be able to make decisions about things like managing the project and overseeing other AIs—basically just taking up the entire org hierarchy.

Nathan Labenz

So, that's the picture you're envisioning?

Ajeya Cotra

Yeah, so I think there are 2 possibilities here. One possibility is that by the time it's the right move to dump a bunch of money on crunch-time AI labor, Open Phil itself has already been largely automated. And that's actually an easy world, because in that world we just have a visceral sense that AIs are really helpful: maybe we've slowed down our junior hiring, all our program associates are AIs right now, and we're totally transformed as an organization. So the evidence and the conviction to pull the trigger might be easier to achieve.

And then, actually, we have a bunch of labor. Maybe we have 1,000 people on the AI team instead of the 45 that we have now, and they can figure out all this stuff much more quickly. But I think the concerning possibility is that there's actually some jaggedness, where maybe AI is extremely good at math, extremely good at technical AI safety, and extremely good at certain specific kinds of manufacturing that could be really useful for a PPE play.

But it's not like we have automated ourselves. It's not that good at doing our jobs because there wasn't much of that stuff in the training data. We're just not well set up to absorb AI labor. It makes horrible mistakes in a way that you can put it in a setup, like in software or manufacturing, where you catch those mistakes, but it's harder—you need humans to do that on the Open Phil side.

So we're not very automated. We don't have a visceral sense that it's time now—that this is the moment, AIs are really, really good, and we have to go big. But it's still the right thing to do to pour a bunch of money into AI labor on these few verticals that are heavily automated.

Nathan Labenz

I think we've maybe actually been burying the lead a little bit here on what the biggest challenge is for an external group like Open Phil to implement this plan, which is: will you even be given access to the very best models that are being trained? And I guess, at this crunch time, when there's a crunch-time demand for compute, will you actually have enough compute? Will anyone be willing to sell to you if you were to do this kind of work? Can you go into that?

Ajeya Cotra

Yeah, I think there are 2 challenges here to getting access to enough labor as an external group. One is whether they will even sell to you. So, like I said earlier, in AI 2027 and in a lot of stories of the intelligence explosion, you get to a point where 1 company has pulled far enough ahead of its competitors that it keeps its best internal systems to itself and only releases systems that are considerably worse than its internal frontier—systems that are just good enough to be ahead of its competitors' released products. There can be a growing gap between how intelligent the best internal systems are and how intelligent the best externally accessible systems are. The AI company may deliberately choose not to sell to willing customers because it wants to keep its secrets to itself.

Another possibility is that they might be willing to sell to you, but the price might just be way too steep because the opportunity cost of using that compute to sell to you—to do whatever you want to do with it—is training further, more powerful AIs, and they might be willing to pay quite a lot for that. I think both are challenges.

The second one is, in some sense, more straightforward to address, which is that you try to hedge against this possibility by having some portion of your portfolio really exposed to compute prices and hope that maybe, in the extreme case, that looks like just having GPUs yourself that, in peacetime, you rent out to other people doing commercial activity with them, but then during crunch time you redirect them to doing AI labor. Although, in that case, you'll have to furthermore figure out how to get the latest AI chips and the latest AI models onto those chips that you own. So you might have to cut deals to make that happen.

But also, in less extreme cases, you might just purchase a bunch of NVIDIA or purchase a bunch of liquid public stocks that are exposed to AI to make it more likely that you can afford AI capabilities at the time. So there could be a huge run-up in the price of GPUs or compute at this time, but you can partly hedge against that possibility by having most of your investments be in NVIDIA or other companies that sell GPUs. If their price goes up, you benefit on the investment side, and that helps to offset the increase in price.

Nathan Labenz

Okay. And then on the software side, there's a question of whether you have access to the very best models that are being trained. I guess, on the one hand, there's this story you could imagine where the companies are very close together, the models are roughly the same, margins are very low, and they're very keen to put out models as soon as possible in order to remain competitive. I guess, on the other hand, you could have 1 leader that's starting to keep things all secret. Do you have a particular take on which of these scenarios you think is more likely to come about?

Ajeya Cotra

Yeah, I think that, at least at the beginning part of crunch time, when the AIs are just starting to automate a lot of AI R&D, my bet is that things will at that point be relatively commercial, relatively open. The leading few companies are within a month of each other in their capability frontier, or maybe it's hard to say who's in the lead because 1 company specializes in 1 aspect—for example, its model is a little spiky on pretraining and another company's model is a little spiky on software engineering or something like that. I think the reason I think that is basically just because it's kind of what a naive Econ 101 model would predict would happen. It seems like these companies don't have big moats. It also seems like what we've seen happen over the last few years—

Nathan Labenz

It describes the present day more or less.

Ajeya Cotra

It describes the present day, and that's a change from a few years ago, when I do think OpenAI had way more of a lead and it seemed more plausible that there would be a monopoly or a duopoly. But there are reasons to push in the other direction, which is basically that if you have a super-exponential feedback loop, you have a bunch of actors that are growing at an increasingly rapid rate—first at 2%, then at 4%, then at 8%—and they don't interact with one another, you do get a winner-take-all dynamic. If they're growing on the same growth curve but 1 gets to a particular milestone first, that leader gets more and more and more powerful and wealthy relative to the laggards. This is in contrast to exponential growth, where if everyone is growing at 2% forever, then the ratios between more- and less-wealthy nations or companies stay fixed.

There is a reason to think that specifically around the time of the intelligence explosion, gaps will begin to grow again. But I think probably around the start, it will most likely be the case that you can buy AI labor if you can afford it. You can buy API credits. You can go on chat.openai.com. And then I think I have a lot of uncertainty about how it evolves from there.

Nathan Labenz

Yeah, what do you think is the chance that the leading company will try to keep the level that they're reaching secret?

Ajeya Cotra

I think it depends a lot on the competition landscape they face. So basically, if the other companies are really far behind, then I think there's a pretty strong incentive and reason to keep your capabilities secret, because you give up quarterly profits, but maybe you don't care about that because you're running on investment money anyway. And if you can get your AI to help you make better AI to help you make better AI and so on, you could emerge with superintelligence that might give you a power that rivals nation-states, or the ability to decisively control how the future goes. And that might be very attractive to a sort of power-seeking company.

I do think it involves forgoing short-term profits, though, which means that if competitors are close at your heels and your investors are breathing down your neck to deliver quarterly earnings, I guess you can't go and tell all of your investors, "Oh, don't worry, we have a superintelligence," because I think then that will get out. And then also, your plan is to screw over the investors in this case. Your plan is to create a superintelligence, not to pay them back. So, create a superintelligence and take over the world, maybe. They won't like that.

There's a mismatch in incentives between the investors and the CEO, and the CEO is sort of being a bad agent to their principal. So basically, the more things look like an efficient, competitive market with very little slack, the more the leading company will be forced to provide access to the rest of us.

Nathan Labenz

To what extent do you imagine the companies would be enthusiastically bought in on assisting with this plan? So, this strategy is the predominant approach to AI technical safety work.

Ajeya Cotra

I think even the optimists agree that there are other issues that society is going to have to deal with. In fact, the leaders of the companies say this all the time: that we're going to need a new social contract. It's going to upend everything. It's going to be a big deal.

I imagine that, insofar as they're nervous about the effects the technology is going to have, they'll be very happy if someone came to them with a pre-prepared plan for how we're going to deploy all of this compute in order to solve all these other problems.

Nathan Labenz

Yeah, I think it's unclear. I think there are certainly some incentives for them to be into this. But the 2 alternative uses of AI labor that might be more attractive to them are, 1st, power-seeking for themselves: building up an enormous AI lead over everyone else and then bursting onto the scene with an incredible amount of power and the ability to challenge the US government or nation-states. That might be attractive to some people.

I think that would be a very evil strategy to pursue, but it's definitely in the water. The other thing is more mundane. It's just using these AIs to make normal goods and services—to make the products, media content, and other services that people most want to pay money for in the short-term sense.

It's very similar to how, right now, we don't spend a huge fraction of society's GDP on biodefense, cyber defense, and moral philosophy. It's just that that's not what people want to pay for. AI is just another thing that accelerates the creation of the products and services people want to pay for, and this isn't very high on the list.

I guess most people are not looking to become dictator of the world or to take on huge amounts of power. But the kinds of people who end up leading very risky technology projects are not typical people. They're somewhat more ambitious than the typical person, so I suppose we can't totally rule that out as a possibility.

So a possible challenge would be that, even if you have an enormous amount of compute, there might be a limit to how fast you can go because you require some sort of sequential steps. There's some step that's bottlenecked in time. You have to do an experiment that actually takes a certain amount of time to play out.

More generally, at least with LLMs, for example, they produce one token after another. Having twice as much compute doesn't necessarily allow you to complete an answer twice as fast. With that limit, how much of an issue is that here, insofar as we're trying to solve problems in a very short calendar time?

Ajeya Cotra

Yeah, I think that is likely to come up, especially for physical defenses like manufacturing PPE or scaling up the ability to rapidly create medical countermeasures. It also applies to social and policy things.

I can imagine that AIs could be very helpful in figuring out what kind of agreement between the US and China would be mutually beneficial and how we can enforce it. But the way human decision-making works still probably requires humans from the US and China to come together and talk about it—a conference or convening—and come to a decision that they ratify and feel good about. That could be a bottleneck.

Nathan Labenz

Yeah, are there any other examples of similar bottlenecks?

Ajeya Cotra

I guess in terms of solving theoretical problems, you can speed things up enormously by having many different instances of the same model try to brainstorm different solutions and then have them evaluate one another. That allows you to have many different efforts in parallel.

But I do think that, for deep theoretical problems, you can speed things up by having efforts going in parallel, while the right solution that's out there somewhere involves multiple leaps. It's hard to think of the next insight without having the foundation of the earlier insight.

Really, even if you have 100 AIs working in parallel, what will happen is that one of them comes up with the 1st step of the insight, and then everyone works in parallel on finding the next insight. But you still need to go 3 or 4 steps in.

Nathan Labenz

So what sort of stuff do we need to be doing in advance? For example, setting up planning meetings ahead of time for diplomats between the US and China: are we about to do that at the very early stage in anticipation that eventually we might have a deal that they might want to ratify? I guess that sounds a bit crazy, but are there other examples of things that you need to do before this all kicks off?

Ajeya Cotra

Yeah, I think that in general, you want to be thinking about what the AIs at the time would be comparatively most disadvantaged in. They'll have all these advantages over us. They'll understand the situation much better at that point in time than we do now. They'll be able to think faster, move faster, and so on.

But I think what we can contribute now would be things that inherently take a long lead time to set up. That might include physical infrastructure, like the bio-infrastructure that my colleague Andrew is working on building out.

It might also include social consensus. I think it takes some amount of time for an idea to be socialized in society, to become an accessible concept. Maybe we should try to create some sort of treaty between the US and China to allow AIs to progress somewhat slower than they might naturally and use a bunch of AI compute to solve all these problems.

I think that kind of thing takes years to become something that's in people's toolkit, in the water, such that they actually think to have the AIs go down that path and figure out the details of it.

Nathan Labenz

So what should people be doing if they think that this makes sense, or it's something that they'd want to contribute to? Are there other organizations that should similarly be planning ahead and thinking about how this might look for them? Could individuals be thinking about how they could contribute to adopting this approach for their own particular projects?

Ajeya Cotra

In terms of other organizations, I think it would be especially great for government entities to be thinking about adopting AI. I know that there are a number of random little types of red tape that make it harder for governments to adopt AIs than for anyone in industry to adopt AIs.

I think we might end up in a situation where the regulatees, like industry people, have fast cars and the regulators have horses and buggies because of this differential adoption gap. More broadly, if your company is not already going maximally hard on adopting AI for your particular use case, and you work on defenses, AI safety, moral philosophy, or all these good things, it's probably worth having a team that's on the lookout for how you could adopt AI as soon as it becomes actually useful for you.

Rob Wiblin

Let's talk a bit about the career journey that you've been on since we last did an interview 2 and a half years ago. Back then, you were doing general AI research and strategy for Open Philanthropy. This was in 2023. Then, in 2024, you started leading the AI technical grantmaking. Towards the end of that year, you decided to take 4 months off and take a sabbatical. Tell us about all of that.

Ajeya Cotra

Before that, I had been at Open Phil for more than 6 years before I made my first grant. I was involved in some grantmaking conversations earlier, but the first grant I actually led on was somewhere in mid or late 2023, and I had joined Open Phil in 2016.

My work at Open Phil was strange if you took the outside view and said, “This is a philanthropy that's giving away money.” It involved thinking about these heady topics and writing long reports that I published on LessWrong. I always felt like I maybe should dip into grantmaking because that is our core product in some sense. It's what we do.

But I had always been drawn away by deeper intellectual projects. Even though I always vaguely had the thought that I should do grantmaking, it never really happened for me.

The thing that pushed me headfirst into grantmaking was the FTX collapse. Actually, my first grant must have been in 2022 instead of 2023, because at that point there were hundreds and hundreds of people who had been promised grants by the FTX Foundation whose grants weren't going to go through, or who were worried their grants were going to be clawed back, or whose grants were only partially going through.

Open Phil put out an emergency call for proposals for people who had been affected by the crash. I had some thoughts and takes on technical research, and the organization also needed surge capacity for this emergency influx of grantmaking. In a matter of maybe 6 weeks, I made 50 different grants after not having made any grants at all.

That was a really interesting experience. I discovered that there were elements of it I really liked, but there was also something about the way you made grants where you just couldn't dig into any particular thing very much. Especially in the context of something like the FTX emergency, you had to make these decisions really quickly.

But I felt like I had thoughts about how grantmaking could be done with more—at least in the technical AI safety space—inside-view justification for the research directions we were funding than we had previously. So in early/mid-2023, I tried to go down that path.

Nathan Labenz

Sorry, so in 2022, you did this huge burst of grantmaking. I guess it was to help a lot of refugees from the FTX Future Fund, basically?

Ajeya Cotra

Yeah.

Nathan Labenz

But then you thought you had probably noticed that there was no overarching strategy behind all of the grants you were making, and you were like, “We need to have a bigger-picture idea of where we're at, what we're actually trying to push on, and why.”

Ajeya Cotra

Yeah, so I was focused on grants to technical researchers. These were often academics, sometimes AI safety nonprofits, and they would often be working on interpretability or some kind of adversarial robustness. They seemed like reasonable research bets, but I felt unsatisfied, and I think this is going to be a theme of me and my career.

I felt unsatisfied about how the theory of change hadn't really been grounded out and spelled out: how this type of interpretability research would lead to this type of technique or ability we have, and then how that could fit into a plan to prevent AI takeover in this way. The same was true for any of the other research streams we were funding.

This had actually been the big thing that deterred me from getting involved in Open Phil's technical AI safety grantmaking for a long time, even though I was one of the few people on staff who thought about technical AI safety outside of that team. In the end, it seemed like most grant decisions in the 2015 to 2022 period turned on heuristics about, “This person is a cool researcher and they care about AI safety,” which is totally reasonable.

But I wanted to have more of a story for why this line of research was addressing a critical problem, why we thought it was plausibly likely to succeed, and what it would mean if it succeeded. We never really had that kind of very built-out strategy, because it's very hard. It's a lot to invest in building out a strategy like that.

Having been thrown headfirst into grantmaking with the FTX crisis, I thought, “Maybe I do want to try to take on the AI safety grantmaking portfolio,” which at the time didn't have a leader because all the people who had worked on that portfolio had left by that point. Some had gone to the FTX Future Fund, actually.

Nathan Labenz

Okay.

Ajeya Cotra

So it was this portfolio that had been somewhat orphaned within the organization, and it was clearly a very important thing. I thought, “Maybe we could approach it in this kind of novel way for us in this area: really try to form our own inside views about the priorities of different technical research directions and really connect how that would address the problems we most cared about.”

Nathan Labenz

It sounds like you find it unpleasant or anxiety-inducing to make grants where you don't have a deep understanding of what the money—or, I guess, not so much what the money is being spent on, but whether you have a personal opinion about whether it's likely to bear fruit. Is that right?

Ajeya Cotra

Yeah, or I think it's a bit nebulous what the standard is that I hold myself to. But for my research projects, when I think about timelines, how AI could lead to takeover, or how quickly the world could change if we had AGI, I can often, with months of effort, get to the point where I can anticipate and have a reasonable response to—and a reasonable back-and-forth with—a very wide range of intelligent criticisms for why my conclusion might be totally wrong and totally off base.

I feel like I know what the skeptics who are more doomy than me will say, and I know what the skeptics who are less doomy than me will say. I could have an intelligent conversation that goes for a long while with either side, and that is a standard I aspired to get to with why we supported certain grants.

I could do that with some of our grants, but I wanted the program to get to the point where, if somebody came to me and said, “Isn't interpretability just something that hasn't seen much success over the last 4 years? What do you make of that?” I wanted to be at reflective equilibrium on my answers to questions like that.

I wanted to be able to say something that went a bit beyond, “Yes, but outside view, we should support a range of things.” That is something I find emotionally unsatisfying if it's a big element of my work.

Nathan Labenz

Yeah, it's maybe worth explaining why Open Phil doesn't aspire to get to that level of confidence with most of its grants. Why is that?

Ajeya Cotra

I think it just takes a long time. There are 2 things: it takes a lot of effort, and then the other thing is that even if you put in that effort, you don't want to fully back your own inside view. I don't think I would endorse that either.

So it's this one-two punch where developing your views about exactly how interpretability, adversarial robustness, control, or corrigibility fits into everything is a ton of work. You have to talk to a ton of people, and you have to write up a bunch of stuff. In the meantime, you're not getting money out the door while you're doing all this stuff, right?

Then, having done all this stuff, where are you going to end up? You're going to end up in a place where there are reasonable views on both sides, and it's a complicated issue. We probably want to hedge our bets and defer to different people with different amounts of the pot and so on.

I think people have a reaction that's very reasonable: “Okay, we're going to end up in a place where we've thought it through. It was a lot of work. It's still very uncertain. We still want to spread our bets. So why let it affect the decision that much? Why not just get to the point where you short-circuit all that, spread your bets, and lean on advisors?”

I have sympathy for that. Hopefully I represented that perspective reasonably well, but I feel like in my life, in my experience, having done the homework really qualitatively changes the details of the decisions you make in ways that I think can be really high-impact.

One thing that I'm able to do, having gone through the whole rigmarole of forming views, is work with researchers to find the most awesome version of their idea by the lights of my goals, pitch them on that, and co-create grant opportunities.

I think there's just something that I maybe won't be great at defending, but I feel like there are other nebulous benefits beyond that, and I really like operating that way.

Nathan Labenz

So in 2024, you actually took on responsibility for this whole portfolio, but I guess—

Ajeya Cotra

Yeah, late 2023.

Nathan Labenz

But I guess your personal philosophy of how to operate is somewhat in tension with how Open Phil as a whole is tending to operate, and with the way that—

Ajeya Cotra

—is in tension with making a large volume of grants in the short term.

Nathan Labenz

Right. So what did you end up doing in the role?

Ajeya Cotra

I think I ended up pursuing a compromise. One thing that just comes with the territory of this role is that there had been grantees we'd made grants to in the past who were up for renewal, and part of the responsibility of being the person in charge of this program area was to investigate those renewals and make decisions about whether we should keep the grantees on or not.

I tried to follow what an Open Phil canonical decision-making process would be for those grants. I also tried to pursue a barbell strategy for a while. On the one hand, there were renewals where I wouldn't hold myself to the standard of really understanding and defending the proposal on its technical merits, but would lean more on heuristics: this person seems aligned with the goal of reducing AI takeover risk; this person has a broadly good research track record, and so on.

I tried to make those grants relatively quickly. But I would also be trying to develop a different funding program, or some grants that I really wanted to bet on, where I would try to hold myself to that standard and really write down why I thought this was a good thing to pursue.

It turned out that the second thing basically turned into making a bet, from late 2023 to mid-2024, on AI agent capability benchmarks and other ways of gaining evidence about AI's impact on the world.

Nathan Labenz

So it's sort of the stuff that we were talking about earlier, where you're trying to get an early heads-up about whether AI is going to be really effective at operating as an agent. I guess in 2023, we were really unsure how that was going to go. It seemed like agents in general had been a bit disappointing, or hadn't progressed as much as I expected, or probably as much as you expected.

But at that point, it seemed like maybe by this point they'd be operating computers completely as well as humans, and you really wanted to know if that was the future we were heading for.

Ajeya Cotra

Yeah. So I launched this request for proposals. Open Philanthropy has done technical safety requests for proposals before, but this was by far the narrowest and most deeply justified technical RFP that we had put out at that time. We were looking for benchmarks that test agents, not just models that are chatbots, and these are the properties we think a really great benchmark would have. These are examples of benchmarks we think are good and not so good.

We had a whole application form that was, in some sense, guiding people—or trying to elicit the information about their benchmark that we thought would be most important for determining whether or not it was really informative. Mostly, this was just: be way more realistic and have way harder tasks than existing benchmarks. Even if you think your tasks are hard enough, they're probably not hard enough. There was a lot of push in that direction, so it was a very opinionated, very detailed, and very narrow RFP.

We ended up making $25 million in grants through that, and then another $2–3 million from the companion RFP, which was broader—all kinds of information, from RCTs to surveys about AI's impact on the world. I'm pretty happy with how that turned out. It was what you would expect from having a lot of effort poured into one direction.

If you were skeptical of this high-effort approach to grantmaking, you could argue that I could have put in way less effort and funded twice as much volume in grants across 10 different areas, picking up the low-hanging fruit in all those areas.

Nathan Labenz

So I guess halfway through 2024, you started feeling pretty burnt out, or you wanted to take a bit of a break. Why was that?

Ajeya Cotra

Yeah. I think throughout this, right around when I switched from doing mostly research to doing grantmaking, and especially when I was trying to ramp up this program area that had this more inside-view, more understanding-oriented approach to AI safety research, the person who had been running the AI team up to that point decided to step away and left the organization. They had been my manager.

I think I had a working relationship with my manager that involved a lot of arguing and discussing the substance of what I was working on. When my manager left, leadership was stretched more thin because someone in leadership was gone, and I think the people who remained on the leadership team didn't have as much context and fluency with all this AI stuff.

When I wrote up this big memo saying we should do AI safety grantmaking in a more understanding-oriented way, we should develop inside views, and here's why I think that would be good, I think what I wanted was for my manager or leadership to argue with me about the object level of that. I wanted there to be some sort of shared view within the organization about how good an idea this was, what the pros and cons were, and how much we wanted to bet on it.

I think that was just kind of unrealistic, given the other priorities on their plate and their level of context in this area. So I ended up having to approach it in a more transactional way with the organization. It was more like, “I want to do it this way,” and they were like, “Yeah, we don't know if that's the best way to do things, and we have some skepticism, but you can do that if you want.”

I felt kind of lonely because I think—and this is something I learned about myself over the course of trying to run this program, going on sabbatical, and reflecting on it—that I really like to be plugged into the central brain of the organization I'm part of. I didn't feel like I had a path to do that. Instead, what I had a path to do was stand up this thing, which I tried to do, but it just felt a bit tough going.

Nathan Labenz

It sounds like you were a bit on your own.

Ajeya Cotra

Yeah, I felt a bit on my own, and I'm not a very entrepreneurial person, I think. [laughter] I'm ambitious in some ways, but I just really have a high need for constantly talking to other people.

I tried to achieve that sense of team by hiring people under me to help me with this vision, but I think I was not very good at hiring and management. Partly, it was because this vision was pretty nebulous, and I probably needed to spend more cycles working out the kinks in it by myself and really solidifying what it was and what the realistic version of doing an understanding-oriented technical AI safety program would be.

It was very hard to hire because you kind of had to hire for someone who really resonated with that off the bat, even though it wasn't a very well-defined thing. That took a lot of energy.

With the people I was managing, I have always struggled—and in this case still struggled—with perfectionism in management. I have this long history of trying to get people to serve as writers who write up my ideas, and it never works for me because they don't do it just the way I want it. I'm myself a pretty fast writer, so working with a writer as their editor and getting their writing output to be something I'm satisfied with often ends up taking more time than doing it myself.

I found the same thing happened to some extent with grantmakers. At one point, we had a number of people who spent part of their time working on the benchmarks RFP. It's possible that I would have just moved through the grants faster if it were only me working on it, which is a bit tough.

I think this is a weakness or challenge that a lot of new managers go through, and I was going through that at the same time as feeling like some of the feedback and engagement I got from above me was much less than it was before. I had to prove this new way of doing things, and I thought—and still think—that there was a lot to the arguments I was making. But it was not a wild success when I took a swing at it by myself.

Nathan Labenz

So, in September last year, you decided to step away and take some time away from work, I guess after 8 years of working very hard full-time. What did you end up doing with that time?

Ajeya Cotra

It was a mix of things. I did a lot of life stuff. I found a new group house to move into—or started a new group house. That was cool. I invested more in trying to take care of myself. I started an exercise habit. I've fallen off that exercise habit now again, so we'll see.

I did a lot of reflecting on why this work situation ended up being so hard for me, and also on my journey through my career as a whole—what the patterns were and when things were hard for me.

I also jumped in and helped with some random projects that were going on. The Curve conference, which brings together AI skeptics, AI safety people, and people on all sides of the issue of AI's impact on society, was having its first iteration while I was on sabbatical. I was able to get more involved with that and try to be more helpful than I could have been if I had a full-time job, which was really cool.

I did some writing. Most of that writing hasn't been published, but it was still good for me to do. It kind of went by really fast, honestly. There was a lot of stuff to think about and a lot to do.

Nathan Labenz

What sorts of reflections did you have on your career so far, your motivation, and what had been difficult in 2023 and 2024?

Ajeya Cotra

Yeah. In terms of 2023 and 2024 specifically, I really do feel like I want to be an advisor and a helper to the central organization, and I had been that in many ways over the previous 6 years.

The transition to being more entrepreneurial—to feeling like I had a little startup making grants in my area, with the organization investing money in me but not necessarily a lot of attention—and not necessarily having a path to make arguments that then influenced things in a cross-cutting way was hard. I think that was an interesting thing to learn about myself: if I don't have that, I will still gravitate toward trying to meddle in everything else that's going on. If I don't have a productive path to meddle, I'll feel sad.

That was one big thing. Another big thing was how much depth I want. I do think I really want to get to the bottom of something. I'm always thinking about the counterargument and the counterargument to the counterargument, and the stuff I liked even when I was very young.

I really liked math tutoring, and I really liked math in general, because you could just dig and dig and get to an answer.

Nathan Labenz

And that's just inherently an uneasy fit with grantmaking—basically, investing or venture capital that Open Phil is engaged in, in a way. Yeah. So that was also interesting to reflect on. It was somewhat strange that, for my first 6 or 7 years at Open Phil, I did rather deep research even though we were a grantmaking organization. I just wasn't doing grantmaking.

Was that in part because Holden really wanted this deep research? He wanted to more deeply understand the idea, both personally, and he thought it was healthy for the organization.

Ajeya Cotra

Yeah, I think that's right. He had a lot of drive and demand for really figuring out timelines, really figuring out takeoff speeds, and exactly what our threat models are for whether AI could take over the world, and building that all up. I think he has a lot of the same instinct I have: it's really good to do your homework, and it's really good to have the response to the top 10 counterarguments and the response to those responses, and just really know your stuff. He was the driver of a lot of the work that I did.

If you rerolled the dice and Open Phil had been run by different leadership, it's probably pretty unlikely we would have gone as deep as we did into doing our own AI strategy thinking. The thought would have been, “Well, we should fund a place like FHI, or now Forethought, to do that stuff instead of us.”

Nathan Labenz

In your notes, you said that you spent a fair bit of time reflecting in this period about what it had been that you liked about effective altruism, I guess as an ecosystem and as a mentality, and what things you didn't like so much about it. Tell us about that.

Ajeya Cotra

Yeah. It's been a long time since you've talked about effective altruism on the show, so I'll just open with what it even is, which is this movement or idea that you should think explicitly, seriously, and quantitatively about how you can do the most good with your career or with the money you're donating. Different career paths and different charities you could donate to could differ by orders of magnitude in how much good they do.

If you're working on reducing climate change, it could be orders of magnitude more helpful to work on researching green technologies versus getting people to turn off their lights more or conserve electricity in their personal use. There's this ethos that if you're really taking this seriously and you really care about helping the world, you stop and think and do the math.

In the same way that if you had cancer, or your spouse had cancer, you would do the research and figure out what treatments had what side effects and what treatments had what success rates, and you would ask a lot of questions of the doctor, there's this ethos that that's what it looks like when you take something seriously. A lot of people, when they're doing good in the world, do what makes them feel instinctively good. There's a whole other approach where you respect the intellectual depth of that problem.

I was really drawn to this. I fell headfirst into the EA rabbit hole when I was 13, so it's been more than half my entire life that I've been extremely involved in this community and this way of thinking. I think there were maybe 3 big things that I really liked about this approach.

One is that EAs challenge themselves to care about people and beings that are very different from them, and very far away from them in time and space. Even the most, quote-unquote, vanilla EA cause area of global poverty, the vast majority of money given by individuals in rich countries to alleviate poverty goes to helping other individuals in rich countries, even though money could go much, much further overseas in countries where people have a much lower standard of living.

The reason people donate locally is that they feel more affinity for people who are closer to them and more similar to them. EA also has a lot of strains that challenge people to extend care to animals, to future generations that may live thousands or millions of years in the future, and to artificial intelligence too, if it can have consciousness and feel pain and so on. That was really appealing to me.

There was also a way of going about doing things that was very appealing to me. They were very nerdy and very intellectual. They were really thinking stuff through and almost innovating methodologically on how to figure out which charities are better than which other charities. There were lots of interesting arguments thrown around for this.

They were very transparent. There was a culture of open debate and admitting mistakes. GiveWell, an early pillar of the EA movement, had a mistakes page on its website where it discussed mistakes it had made. They were very honest and high-integrity in an interesting way that doesn't obviously follow from caring about other beings more.

For example, GiveWell refused to do donation matching because donation matching is usually a scam, where the big donor would have given that much anyway even if you hadn't made your donation. That whole package was really attractive to me. It hit a lot of psychological buttons for me at once and really felt like my people, and like the way I wanted to live my life.

Nathan Labenz

So there's being more compassionate to a wider range of beings, which I guess is still the case and probably still something you like about the effective altruist approach.

Ajeya Cotra

Yeah. And then there was the very high integrity about honesty—not allowing any chicanery whatsoever, or even a hint of chicanery.

Nathan Labenz

It was a fastidious and exacting level of integrity that other movements, even other pretty high-integrity movements, weren't aspiring to.

Ajeya Cotra

Even beyond what people were asking.

Nathan Labenz

Yeah. You're just proactively saying, “By the way, did you know donation matching is a scam? That's why we're not doing it, even though we would get more donations to help poor people.” It was interesting that this was such a natural part of the early EA movement, even though you're sort of giving up on impact.

Ajeya Cotra

Yeah, it's not necessarily implied. I mean, it could be a practical question of whether it is or not.

Nathan Labenz

As things evolved, you found that the second one, the intellectual depth, was lacking from your job. Were there other things that were changing that made you less enthusiastic?

Ajeya Cotra

Yeah, I think the intellectual depth was very much there in other parts of the EA ecosystem, especially AI safety and thinking through how exactly you would control early transformative AI systems and things like that. As I said, my heart was always pulled toward those kinds of questions, even though I worked at a grantmaking organization.

Nathan Labenz

Yeah, it feels like on some level you really were a more natural grant recipient rather than a grantmaker. You should have gone and done something to really go deep on some questions.

Ajeya Cotra

I think that if I had graduated college in 2022 instead of 2016, I probably would have done MATS, which is a program to upskill in ML and AI safety research, and then tried to join an AI safety group. I graduated college in 2016 and went to GiveWell. A big part of why I went to GiveWell at the time was that they had the most intellectual depth on the question of what the best charities were.

I think I'm naturally drawn to actually doing the research in some sense. In that sense, it was a mundane issue that my job—especially after Holden left, when the demand for that kind of research evaporated a little bit at the leadership level—was no longer a great fit. If I were to start over again, I probably wouldn't have applied to join Open Phil. I probably would have applied to join an AI safety group.

Then I think there's the third thing: this extremely, almost comically, high level of integrity that I really liked was also eroding over the years. When I think about why, I think that when a lot of the focus of the EA movement was convincing really smart people to donate differently, being unusually high-integrity was actually a really valuable and powerful asset.

People like me, and very wealthy people who were early GiveWell donors, really liked that GiveWell had a mistakes page and really liked that whole ethos and package. It helped them trust that the recommendations were real recommendations, that they weren't being spun or sold something like in the rest of the charity recommendation ecosystem.

Ajeya Cotra

But then when you move away from that being your primary method of change, when instead you've actually attracted quite a lot of funders and now you're trying to use that money and the talent that you've attracted to achieve things in the world, maybe things that involve a lot of politics, then being extremely transparent can be very challenging. Especially because donors want privacy, and if you're running a political campaign, you don't want your opponent to know exactly your strategy and the ways that you think you might have made mistakes. This is just not how most of the real world works.

Nathan Labenz

Yeah, it's not the case that the world's most impactful organizations are consistently incredibly transparent or even incredibly high-integrity.

Ajeya Cotra

Yeah. And so there was this tension between the goals, which I felt I should only care about—the goals of EA. This was what EA told me, and it kind of made sense to me: the point here is to help others as much as possible. The point is not to conform to an aesthetic or do things in a way that feels cleanest or prettiest.

At the same time, I think I was, to some extent, kidding myself about how much of my own motivation and my own attraction to the concept came from just the goals—just pillar 1 and altruism—versus pillars 2 and 3: that intellectual depth and intellectual creativity, and this crazy high level of openness and transparency, having absolutely nothing to hide, letting all comers come. I think for me, as a fact about my psychology, the latter 2 things were actually really important for my motivation, and over time they were just smaller and smaller features of what it was like to do EA—to try and pursue EA goals in my career.

Rob Wiblin

Yeah, I guess we should say, for people who don't know, that over this period the environment that Open Phil was operating in became a lot more challenging and a lot more hostile, I guess.

Ajeya Cotra

Yeah. For years, it had been funding all kinds of AI-related stuff, but as AI became a much bigger industry, it became apparent what sorts of concerns different people had. Its work, in some ways, started to clash with very large commercial interests, potentially, and also alternative ideologies that had different ideas about how things ought to be regulated or how things ought to go.

We're now in a world where there were people who would sit down and think, “How can I mess with Open Phil? What can I do to give these guys a terrible day? What did they publish that we could spread that would be embarrassing for them?”

Yeah. And in that kind of environment, where people just literally want to cause trouble for you, it's a lot less attractive to be maximally forthcoming about all of your internal deliberations and why you made all of your decisions. All of us would potentially be a bit more conservative in that kind of environment.

Ajeya Cotra

Yeah. Even before the latest round that started in 2023 of AI policy heating up, Open Phil compromised a lot on its initial wild ambitions for transparency. At the beginning, there was this idea that we would publish the grants we decided not to make and explain why we decided not to make them when people came to us for grants. We never did that. There's a reason most organizations don't do that.

For our earliest 2 program officer hires, we have a whole blog post that we wrote about their strengths and weaknesses as candidates, the alternatives we considered, and how confident we were that this would work out. We stopped doing that. There was a level of transparency that I still, in my heart, want, but it's absolutely insane.

I think the adversarial pressure that you mentioned makes it so that Open Phil, as an organization that funds a lot of this ecosystem, has a lot to lose. I think if we go down, a large number of helpful projects have a much harder time getting funding. We have to be a lot more risk-averse than many of our grantees, even though those grantees are also facing an adversarial environment.

I think the way many of them navigated it is to sort of fight back and explain their perspective and define themselves in the public sphere. My instinct is to do more of that—to say more and respond—but it's harder to do that from Open Phil's position for a number of reasons.

Yeah. So, over the years, a lot of people, usually critics, have said that effective altruism has something in common with religious movements. To what extent have you found that to be the case, and to what extent have you found that not to be the case?

Ajeya Cotra

Yeah. I think EA aspires to be, and very much succeeds in being, a lot more truth-seeking than the world's religions and a lot more truth-seeking than a lot of other communities and movements in the world. So, in that sense, I think there's a disanalogy that's extremely important.

I do think it's not a bad analogy in some ways, because for people who are deeply involved in the EA community, it provides a map of the good life. It's a vision of what it means to be good and have a good life. It's unlike a political movement in that it doesn't just have a set of policy prescriptions for the world, but, like many religious movements, it intersects with politics.

There are people who approach political questions, such as whether you should ban gestation crates for pigs, through the lens of their commitment to EA. It's not just a community or a social club. I think people get solace and friendship from their local community of EAs, like people do from their local church community, but it is more than that.

It is trying to say something about the sweep of the world and your place in it, and what it means to live a good and meaningful life. It intersects with politics, community, and a bunch of other things while not being exactly the same as any of them.

Rob Wiblin

Yeah, I would think a key way that it's not like a religion is that it feels more like a business to me in many respects, or like an organization—a startup or an organization that has a functional goal.

I guess that's a different aspect of it. Some people like the ideas; they like the blog posts but don't engage with the community whatsoever. I suppose for them it's going to be a different experience. There are people who like the community.

Actually, there are many people who participate in the kind of community of people who would say, “I'm involved in effective altruism,” but who are not that interested in the projects or necessarily even in the effort of helping people. So people sample the aspects that they like. For many of the people who work in organizations that have other people who would say, “I'm really into effective altruism,” it's much more pragmatic, I would say.

Speaker 1

Yeah, I think that is how it ends up manifesting for a lot of people, but I don't think that's really what EA is. I think it's a mistake to collapse EA into a set of 3 or 4 goals in the world: reducing suffering of animals in factory farms, improving quality of life for poor people in developing countries, plus AI safety.

I think in some ways people think of EA as a weird umbrella for those 3 things, and then those 3 things are basically professional communities pursuing a kind of well-defined goal. But I think EA is more like a way of looking at the world and a way of thinking about the good.

I think you can take an EA approach to cause areas that are in some sense more parochial than the big 3 EA cause areas. You can absolutely take an EA approach to US policy from the perspective of thinking about the welfare of US citizens and doing rigorous cost-effectiveness analysis of what policies actually help and don't help. A lot of people do.

And then I think there is EA as a generator of new cause areas that could get added to the canon. Right now there's a bunch of fertile ground around questions like, “Could EA be a force that helps society prepare for radical change by advanced AI?” AI safety is one big, important thing there, but there might be a range of other issues, and you might want to prioritize some of those based on your values and your sense of how things will play out.

Rob Wiblin

So, you're writing in your notes that, at least from your personal point of view, EA wasn't enough of a religion, or it wasn't as much like a religion as you might have liked. Explain that.

Ajeya Cotra

I think I'm someone who really benefits from structure and from emotional motivation and reinforcement. I also tend to socially conform a little bit, or I tend to try and achieve the ideal of the community I'm in. I think the ideal of my corner of the EA community is, as you said, to have a really impactful job, do a really good job at it, and work a lot of hours at it. That's the message you get from the community, and that's what I'm trying to do.

But I personally would have liked a bit more of a spiritual angle to the community. If you read my colleague Joe Carlsmith's blog, I think I get some of that existential reflection about our morality and our values, and this crazy thing that so many EAs believe: that in a matter of a decade or 2, we might be in an utterly transformed world that might be, relative to this vantage point, utopic or dystopic.

Just grappling with that—I think if there had been an EA church where every Sunday someone who's really good and thoughtful about these issues spoke about them and led a discussion about them.

Ajeya Cotra

I think that would have been very enriching for my life and probably ultimately made me higher impact. But that’s just not how the EA community is structured, and it’s sort of deliberately not structured that way, because, as a professional community, EA really wants to not care if people believe the deepest teachings and philosophical orientation. You really want to just be like, “If you’re doing great AI safety research, great. Keep doing AI safety research.” So the incentives of a professional community pull against what I might personally want here.

Rob Wiblin

Yeah. Do you think it sounds like you think that, while it might have been more appealing to you, it’s not actually necessarily better for things to go in that direction?

Ajeya Cotra

I mean, I guess, for me personally, I kind of like the more professional-community, limited aspect of it, because you kind of just want to be able to go home and not have to think about this stuff [laughter] all the time or have it necessarily—

Nathan Labenz

Whereas I want to go home and think about it in a different way.

Ajeya Cotra

I already go home and think about my work all day. I frequently have insomnia where I think about my work, and I just want to be like, instead of thinking about the next Google Doc I need to write or the next email I need to send, I would like to be—

Nathan Labenz

Thinking about your work in a more spiritual way.

Ajeya Cotra

Yeah, exactly. Yeah.

Nathan Labenz

I guess people have a range of views, but I guess it’s clear why many people have not embraced that or have been keen for a stronger division between this sort of thing, which can be very stressful.

Ajeya Cotra

Yeah, and it can be very dangerous and culty. I mean, there are a lot of reasons to worry about it, but I do think there is just a large contingent of EAs that are like me in wanting some sort of spiritual grounding. Joe Carlsmith’s blog is extremely popular with hardcore EAs. It’s not a generically popular blog—it’s reasonably popular—but there are a number of people who are like, “Oh, wow, this is really nourishing something in me that I didn’t realize I needed.”

Nathan Labenz

You wonder if there’s probably an age thing here a little bit as well. I guess I feel like, when I was younger, I noticed that the older people were less interested in this aspect of it. Now I’m in the kind of older class, and I’m like, well, I have my family to provide nourishment, and that’s absorbing a lot of time and energy that I kind of don’t have for attending church [laughter] or whatever else it might be. Do you feel like you had some sort of spiritual hole that was filled specifically by having a child, or were you always just not that interested in this?

Ajeya Cotra

Yeah, I think of myself as a deeply unspiritual person, so I think that wasn’t really a niche that I needed to have scratched. Earlier on, I was maybe more interested in the social scene, to make good friends and meet people. Having made more friends who I think of as like-minded and having a lot of common interests with, that’s kind of not as interesting either anymore. I’ve already got my friends, and now I’m just going to ride it out. [laughter]

Nathan Labenz

I actually thought of myself as an extremely unspiritual person and had a lot of disdain for spirituality when I was 20, so for me the age thing has gone the other way. I think I want more and more of a religion-shaped thing in my life as I age. When I think about why, I think it’s because, when I was 20, I had unrealistic aspirations for my worldly projects.

By that point, I’d already been an EA for 6 or 7 years, but I was just starting off trying to do EA things in the world. I had this sense that this is obviously correct, this is obviously great, everyone who’s good and reasonable will get on board with it, and we’ll just solve poverty and solve factory farming. I wouldn’t have exactly said this stuff, but I just had that inner vibe. I would go around being like, “Have you heard the good word about EA?” [laughter]

As I’ve just done things in the real world, I’m like, everything is very hard and slow. The feeling of doing my job, which involves writing these Google Docs and sending these emails, is just not automatically connected to my higher aspirations. There is a long grind, and there’s a lot of failure, so I think I have increasing demand for some separate thing that is specifically trying to reorient me mentally toward the bigger picture.

Ajeya Cotra

Yeah, so for me the bottom line there is that working on this stuff can be quite stressful and quite tiring, and I want to completely check out and stop thinking about it and just be with people and talk about other issues.

Nathan Labenz

So it’s just like I said: different strategies. I think I probably want some of both. I now live in a group house with a couple of little kids, which is really great. It’s good for that, but I find, unfortunately, that it takes a lot to pull my mind away. I watch TV and I’m thinking about other stuff in the background. [laughter]

So, I think during your sabbatical you considered going independent—becoming a writer or researcher, just doing your own thing—but in the end you decided to come back to Open Phil, at least for a while. Why was that?

Ajeya Cotra

Yeah, so toward the end of the sabbatical, I was planning on taking some time to just start a Substack and write about a bunch of stuff, including a lot of the stuff about EA that we were discussing and a lot of stuff about AI, and sort of see where it went. At that time, I honestly didn’t have a super-strong impact case for this, I think. I didn’t think it was crazy that it would be the highest-impact thing to do, but the reason I was doing it was just because I wanted this, and not because I could really defend that it was the highest-impact thing.

At that moment, after having gone through this whole journey, I was like, yeah, maybe I have more room in my life for making a career decision on the basis of not just impact. The reason I decided to stay was that, basically, while I was out, Open Phil was conducting a search for a new director to lead our GCR work—all our AI work and our bio risk work. This was the position Holden was in when he left in 2023.

Both of the top 2 candidates seemed really good to me, and I felt like someone new coming in could probably really use help from someone who’s not particularly running any given program area, doesn’t have a big team to worry about, and can just help that person develop contacts and figure out their strategy. Then it could be an opportunity for me to see if I could get the feeling of plugging in again that I had been missing for a while.

Nathan Labenz

And how did it go?

Ajeya Cotra

I think it went really well. Our director of GCRs is Emily Olsen, who’s also the president of Open Philanthropy, and I’ve been spending most of the last year—most of this year, 2025—just helping her in various ways, trying to understand: What have we funded? What’s come of that? What’s the AI worldview? What do we think is going to happen with AI? How is that informing our strategy? What are the strategies of the various subteams?

I work really, really well with her, and I’d actually been lonely at Open Phil almost the entire time I’d been at Open Phil, even though it got worse in 2023. Holden was really great at giving me a lot of bandwidth, which I’m really grateful for, and talking about object-level stuff with me. But Holden never ran a ship where he was like, “I’m doing this bigger project. Can you help me with this piece of it, and here’s how it fits in?”

Holden was always more like a research PI, where I was doing my own research project and he would talk to me about it a bunch and was interested in the results, but it was not integrated into a whole. Emily really does operate in more of an integrated way, where I’m doing stuff and I know she needs to know the answer and is going to do something with it, which is very cool and very novel for me as a way to work. It’s something that I always thought I would want.

She’s an extremely caring and thoughtful manager for me, who’s really good at eliciting work out of me. I noticed that I work more than I did right before I went on sabbatical, and it feels less hard. That’s just a sign that things are working.

Nathan Labenz

So you’re trying to decide what to do next—whether to stay at Open Phil or go into something less meta, or maybe something that will allow you to go into even more depth. How are you using the stuff that you’ve learned about yourself over the last few years to inform that decision?

Ajeya Cotra

Yeah, so, besides Open Phil, which is still a top candidate, I’m talking to—

I'm talking to 2 technical research organizations about potentially finding a fit there. One is Redwood Research; the other is METR. Redwood Research works on basically futurism-inspired technical AI safety research, and they've been best known for pioneering the AI control agenda. I think of METR as trying to be the world's early warning system for intelligence explosion. They're measuring all the different measures we want to be tracking to see if we're on the cusp of AI's rapidly accelerating R&D or of acquiring other capabilities that let AI systems take over.

Both of these missions are very close to my heart. They're both narrower than Open Phil, where I could, if I wanted to, dip my toes into absolutely everything that might help with making AI go well. But in exchange, they would let me go deep in a way that would probably be more satisfying for me, all else equal.

In terms of how I'm using what I've learned, I think I just—and this is so cliché, and it's something that if a 20-year-old version of me were watching this, she'd roll her eyes—but your extremely local environment, like the literal person you're reporting to, matters a huge amount. The 2 or 3 people you're going to be talking to most in your job, or just features like how much you're talking to people in your job versus working on your own, can make a transformative difference.

I found it interesting to reflect on all that stuff I said earlier about how EA has become a lot less transparent and a lot less inclined to prioritize maximal integrity at all costs. That still bothers me. Actually, the moral foundations of EA, which I think of as sort of utilitarian thinking, can take you down a long rabbit hole where it's very suspect in many ways, and we talked about this in some previous episodes.

Both of those things bother me a lot more when I'm also in a working environment that's locally hard for me. It's not like those issues aren't issues, but the salience of those heady, big-picture things versus extremely micro things, like what it feels like when you have a 1-on-1 with your manager, is very different. I think I had been underrating the mundane and the micro in how I had been thinking about my career up to now, and I'm trying to do trials. I'm actually in the middle of a work trial with METR as we're filming this episode, and that's what I'm paying attention to: How does the rhythm of the work feel? How do the people feel?

Nathan Labenz

Yeah. I guess the other generalizable observation is that Open Philanthropy's environment changed over the years. You were there for 8 years—

Ajeya Cotra

9 years now. Yeah.

Nathan Labenz

9 years, right? Yeah. The constraints that Open Philanthropy was laboring under in 2023 are very different from those in 2016. I guess it might have been a good fit for you to start with, but that doesn't necessarily mean it will be a good fit forever.

There was also a leadership change at Open Philanthropy. The person you were reporting to changed. Very often when that occurs, you see some other people leave as well, because they were in their roles primarily because of their very good working relationship with that person, or because they had strategic alignment with that person. Potentially, the CEO changing could have been a trigger for you to think, “Maybe this isn't so great anymore, and I should proactively start looking for something else.”

Ajeya Cotra

Yeah. I think that's possible. It sort of was true for me in both directions. Holden was very much a huge part of why I wanted to work at GiveWell rather than work in a number of other potential places or do earning to give, which I thought I was going to do at first. When he left, that coincided with a difficult period for me.

Now, with Emily in the position he was in before, it's again pretty dramatically changed what my work is and how it feels. It does seem like it's a big, transformative thing. If you're in an organization where there's a leadership change, I think it should probably be a trigger to think about what might be different about your role, your place, and what you're doing based on the different style, constraints, strengths, and weaknesses of the new leadership.

Nathan Labenz

It sounds like taking 4 months off was also a good call. I guess it stopped things when you were reasonably unhappy. It could have gotten worse if you hadn't done that, and it gave you breathing room to make good decisions.

Ajeya Cotra

Yeah, I think that's right. I'm very glad that I took the sabbatical. I'm also glad that I didn't leave. An especially salient alternative for me at the time that I decided to take 4 months off was to just leave and figure out what I wanted to do next.

I think it was good both for my impact and for my personal growth and satisfaction that I came back, helped Emily, and now I'm doing a proper job search. At the time that I left for my sabbatical, it was more about healing and reflecting, and not a focused search for a role.

Rob Wiblin

My guest today has been Ajeya Cotra. Thanks so much for coming on the 80,000 Hours podcast again, Ajeya.

Ajeya Cotra

Thanks so much for having me.

Nathan Labenz

If you're finding value in the show, we'd appreciate it if you take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website cognitiveevolution.ai or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a16z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And [snorts] thank you to everyone who listens for being part of the Cognitive Revolution.

Rob Wiblin

Coming back to effective altruism for a bit, you said we basically almost don't talk about effective altruism on the show anymore. It was a much bigger feature in the earlier years. The biggest reason for that, I suppose, is that now we're more AI-focused. AI is an issue that so many people are concerned about regardless of their broader moral values or broader moral commitments.

You don't have to be concerned about shrimp, or about beings very far away in time, to think it would be really good to do technical AI safety research, think about what governance challenges are going to be created by it, and so on. EA is a controversial idea, and I think at its core it's quite a controversial idea. Many people, even fully understanding it, would simply not agree with its prescriptions about how resources ought to be allocated.

Why bring along all that baggage when it's not actually decision-relevant for most people? Do you think we should talk about it more, or is that just a sensible evolution?

Ajeya Cotra

I think it kind of depends on the show's goals. My take is that it's correct and good that you don't need to buy into the whole EA package, with all of its baggage, to worry about misaligned AI taking over the world and to do technical AI safety research to prevent that; to worry about AI-driven misuse and to do research and policy to prevent that; or to generally worry about AI disruption and think about it.

There should be, and there is, a healthy, thriving “AI is going to be a big deal” ecosystem that does not take EA as a premise. At the same time, I think EA thinking and EA values probably still have a lot to add in the age of AI disruption.

I think it's going to be EAs, for the most part, who are thinking seriously about whether AIs themselves are moral patients, whether they should have protections and rights, and how to navigate that thoughtfully against trade-offs with safety and other goals. It's going to be EAs who, by and large, are still the ones who take most seriously the possibility that AI disruption could be so disruptive that we end up locked into a certain set of societal values.

We gain the technological ability to shape the future for millions or billions of years, and there's a lot to think about in terms of how that should go. There are a lot of degrees of extremity to the AI worldview. Even if you accept that AI is going to disrupt everything in the next 10 or 20 years, the people who are thinking hardest about the most intense disruptions are going to be disproportionately EAs, because EA thinking challenges you to try to engage in that kind of far-seeing, rigorous speculation, even though there are a lot of challenges with that and it's very hard to know the future. I think EAs are the ones who try hardest to peek ahead anyway.

Nathan Labenz

Yeah, digital sentience—worrying about AIs themselves suffering—is a good example.

Ajeya Cotra

Yeah, I would definitely make the prediction that effective altruism will loom large among the group of people working on that.

Nathan Labenz

For someone who isn't altruistic or isn't motivated by social impact, it's a bit unclear why you would go into that area. It's not particularly lucrative. It's not, at least yet, particularly respected. It's not super easy to make progress, and it's quite unconventional.

I think most people, most of the time in their careers, want to do something that's acceptable and that their parents will be proud of. It's a lot less clear that digital sentience will provide you with the kind of esteem or prestige—or safety and comfort—that many people want in a career.

Ajeya Cotra

So it's maybe natural that people who are altruistically motivated and also intellectually eclectic, willing to be avant-garde, are going to be more intellectually avant-garde and tolerant of quite a lot of philosophical reasoning and speculation.

In a sense, I think this might be what a healthy EA community is. It's an engine that incubates cause areas at a stage when they're not very respected, they're extremely speculative, and the methodology isn't firm yet. You kind of just have to be extremely altruistic and extremely willing to do unconventional things, and then it matures those cause areas to the point where they can stand on their own while also being a thing that many EAs work on.

I think digital sentience, and maybe the other things on Will and Tom's list, like space governance and thinking about value lock-in and stuff like that, are other candidates for EA to incubate the way it incubated worrying about AI takeover, basically.

Nathan Labenz

Yeah, I feel that less strongly in the case of the value lock-in thing, because many of the mechanisms there would be just ways that AIs—I guess you would get a power grab by people or a power grab by AIs, or somehow it undermines democracy or deliberation—in a way that makes it hard for society to adapt over time.

I think people are worried about that regardless of whether they're involved in effective altruism or would be very skeptical of it. There are some versions of the value lock-in concern that go through something else, kind of overtly scary and bad happening, like one person getting all of the power, and that's how that person's values get locked in, and that's how we get value lock-in.

But I think there's a whole spectrum of things that are almost like social media-plus-plus. It's sort of like, in this distributed way, this technology has made us meaner to each other and worse at thinking, and has allowed individuals to live in information bubbles of their own creation.

You can imagine AIs getting way better at creating a curated information bubble for each individual person that allows them to continue believing whatever it is they started believing, with superintelligent help preventing them from changing their mind. This might be something you think of as an important social problem for the long-run future, even if it doesn't happen via one person getting all the power. Power is still relatively distributed, but large fractions of society are sort of impervious to changing their mind.

So it's interesting that, in thinking about what niche EA can fill that others won't fill, the thing you were pointing to was not primarily actually altruism, although I guess that is a factor in terms of going into digital sentience, perhaps. It's actually a research methodology or a research instinct, which is being willing to be in that very uncomfortable space between just making stuff up and having firm conclusions that you can stand by because you've taken particular measurements.

For some reason, it feels like that is one of the most distinctive aspects of people who are passionate about effective altruism: being willing to try really hard to make informed speculation about how things will go, and neither just have it be a good story nor be too conservative to actually make hard predictions.

Ajeya Cotra

Yeah, absolutely. Even the tamest of EA cause areas, like global health and development, has a huge dose of this. If you look at GiveWell's cost-effectiveness analysis, they have to grapple with how the value of doubling one's income, if you make a very low amount of money, compares to a certain risk of death or the value of a certain painful disease you could have.

They have to try to get their answers based on surveys and weird studies people have done. It's not very rigorous in the end, and they have to form their judgments and spell out their judgments. I think the willingness to tackle questions like this and just be like, "Here's our answer," while recognizing that there is a lot to argue with, is very emblematic of EA organizations, including all the best AI safety EA organizations, like Redwood Research.

Nathan Labenz

Yeah, I guess more standard ways to approach these questions would be to just pick one slightly arbitrarily and then be really committed to it, or to be irritated at being asked the question and say that there's absolutely no way of knowing, or that there's no fact of the matter here whatsoever.

I don't know whether it's somewhere in the middle, but I think within EA there's a spectrum in terms of where in the middle you want to land. Everyone's looking at the person more speculative than them and thinking that they're just building castles on sand, and that this is not the way to do things. They're looking at people less speculative than them and thinking that they're just falling for the streetlight effect, ignoring the most important considerations, and not working in the most important area.

Ajeya Cotra

Yeah. So I guess for people who do have that mindset, an important message would be that people should take advantage of the fact that they have this unique mentality, or this reasonably rare mentality, and go into roles that other people probably won't fill because they feel too uncomfortable, or at least because they could reasonably think it's misguided. Other people aren't necessarily going to do this stuff.

Nathan Labenz

Yeah, I think that's right. I think it's interesting to think about: if you imagine EA as one piece of the world's response to crazy changes like AI, there's actually a case that EA should be heavily indexed on research.

The community has gone back and forth with how it thinks about this. At first, people were naturally attracted to research, so there was a huge glut of people who wanted to be researchers. Then there was a big push, including from 80,000 Hours and others, saying, "No, consider operations roles and policy roles and other things that aren't just research."

I think that was a good move at the time, but I wonder if we think about what EA's comparative advantage relative to the world is. Maybe that suggests that some of the people who are doing operations and policy, but who in their hearts just want to be a weird truth-teller thinking speculative thoughts, should consider going back and doing that again.

It's Crunch Time: Ajeya Cotra on RSI & AI-Powered AI Safety Work, from the 80,000 Hours Podcast | BidClub