GPT-6 Astra Saturates ARC-AGI-3, Tesla Cybercab Hits Austin, Anthropic Proves Fermat's Last Theorem
Peter DiamandisSalim IsmailDave BlundinAlexander Wissner-GrossEmad Mostaque
- OpenAI's GPT-6 Astra saturates ARC-AGI-3 (99.9%) and FrontierMath Tier 4 (98%), yet ranks only third on Artificial Analysis behind Anthropic's Fable 5.1 and Meta's Muse Spark. Alex Wissner-Gross resolves the puzzle: Astra dominates the intelligence-per-output-token frontier — a deliberate optimization for native computer-use assistance — and the inner story is recurrence via looped transformers, which, if true, marks “the beginning of a new scaling law which is depth scaling, which we've never seen before.”
- Emad Mostaque calls Astra “the first non-benchmaxed model” and, per Greg Brockman, OpenAI's first fresh pre-train since GPT-4o — roughly $1B on 100,000 next-gen chips versus ~$10M Chinese pre-trains. What ships is a distilled, smaller model; his standing thesis is, “I don't think we'll ever see their top models anymore” because labs will use them for internal discoveries and are “probably hoarding them right now,” evidenced by prime-gap records falling 260 → 220 → 186 across two days: “they're holding their punches back.”
- Math is “incinerated”: Anthropic formalized Fermat's Last Theorem in 13 million lines of code, proving 29,000 theorems along the way, and Alex expects Clay Millennium-level problems to fall “in the next few months.” Meanwhile the panel's pick for best generally available model is Fable 5.1 (HLE 65% with tools), whose cache reads are 75% cheaper than Fable 5 — the substrate for loading a whole business into context.
- Alex's framing is that frontier leads last roughly 30 days (“we're ahead for a minute. So what?”), so labs are converting edges into lock-in — partnerships, real estate, generators, chips, “entire states, countries.” Peter connects OpenAI's 50/50 profit-share concept to a possible “too dangerous to release” narrative that could force revenue sharing for post-Astra access. Alex identifies companies holding chip-design or mechanical-design data as early targets in “the great land grab that's kicking off right now.”
- Astra is the first model OpenAI ever classified as a critical-tier cybersecurity risk, but the panel broadly dismisses the promised kill switch — “essentially a placebo in this market.” The real worry is depth scaling moving reasoning out of readable chain-of-thought into a forward pass, one-shot inference at 750 tokens/sec on Cerebras (5,000 next year), and Ilya Sutskever's warning that rogue agents will “take over a neocloud to make more copies”; governance meanwhile splits between Bernie Sanders' ban-ASI act (up to 20 years in prison), which Alex calls “banning math,” and the G20's hands-off Carolina Principles.
- Tesla's $30K Cybercab is running ~50% cheaper than Uber in Austin, with Nevada permitting 5,000 vehicles for Las Vegas within 12 months. The unit economics — 17 moving parts versus roughly 2,000 in an ICE drivetrain, two seats, two airbags — plus Elon's edge as “the guy who builds the machines that build the machines” point toward 20-cents-a-mile transport and the episode's summary line: “This will turn transportation into an API.”
- Fei-Fei Li's Atlas world model treats 3D/4D Gaussian splats as a first-class training modality — potentially “a critical new form of token” for modeling the physical world. Emad says the 4K holiday-experience stack “is here as of today,” gated only by the global compute shortage and a 5× RAM price increase; Alex extends the same method to subatomic, astrophysical, and intracellular simulation where “human intuition is just terrible.”
- Emad unveiled “the champion”: TSMC-template, people-owned AI utilities at $1 pre-money per jurisdiction, with 10% equity in perpetuity to every child under 20 — his answer to “the cost of intelligence will drop to zero and the value will go to the last mile.” The stakes, via Elon at the G20: a billion humanoids at 5× human output “will be the economy,” while unequitized human cognition goes negative in value — “like adding a human driver to an autonomous highway.”
1. Twelve frontier releases in 30 days — and enterprises are still dabbling
- The setup: 12 frontier model releases in the last 30 days, one every five days, with rumors of at least five more in two weeks including Grok 4.7. Alex's insistence: these are not “marketing-garbage releases” but major step improvements — “clearly, they're well down the self-improvement path. Clearly, the prior model is accelerating the timeline to the next model.”
- Alex re-ups his extrapolation: “I think we're on track still to see one major model release per day by the end of this year.”
- Salim's field report from a big-oil C-suite: most CEOs are “so woefully behind.” His diagnostic: “if you took AI out of your organization today, would any workflows change? The answer for most people is no” — the advantage lies with the few rewriting organizational design and workflows structurally.
2. GPT-6 Astra's headline: efficiency, not just raw intelligence
- OpenAI's release claims: Astra saturates FrontierMath Tier 4 at 98%, ARC-AGI-3 at 99.9%, ExploitBench at 100%, and “defines a new Pareto frontier for intelligence index versus output tokens,” with the hallucination rate having “fell nearly by half from 92% to 51%,” as read on the pod.
- Sam Altman's Bloomberg framing: the first model where, for a whole complex piece of software, “I could tell someone, ‘Just give it a try.’ There's a good chance it'll work” — with the release deliberately delayed for safety and rolled out first to “trusted-access partners” under tiered cyber access.
- Dave's user-experience point: this is the first model class that shows you screenshots of your own laptop inline — “Is this what you wanted?” — with nothing to install. “Qualitatively, it is massively different from a month ago.”
3. Alex's two-level read: CUA-native outside, looped transformers inside
- The outer perspective: Astra is “a next-generation frontier model that's designed with CUA from the ground up” — internalizing exactly the two moves that let Anthropic leapfrog OpenAI: enterprise-grade code generation and computer-use assistance, à la Claude Code, which demands native multimodality and tight low-latency loops over screenshots and video.
- The inner story, from public comments and outside analyses: recurrence via looped transformers — a single weight-tied transformer stacked on itself for a double loop. Chinese labs are injecting recurrence too, including Kimi via its Kimi Linear Attention mechanism at the attention layer, and looped constructs previously let tiny academic models break ARC-AGI.
- The speculative payoff: thickening a transformer's depth may thicken the “J-space” middle layers where, per Anthropic's consciousness studies, “most of the quote-unquote conscious thinking happens” — and, if true, “we're seeing the beginning of a new scaling law which is depth scaling, which we've never seen before.”
4. Emad's economics: a fresh $1B pre-train, a distilled shipping model, and hoarded frontiers
- His headline: Astra “looks like the first non-benchmaxed model” — it saturates ARC-AGI-3 yet “lags behind Meta Muse Spark” on Artificial Analysis, “kind of weird” unless you assume no benchmark tuning. Per Greg Brockman, it's the first pre-train since GPT-4o — the entire 5-series, up to the math-cracking 5.6 Pro, was that old pre-train extended, reportedly after most of the pre-training team left.
- The numbers: trained on 100,000 next-generation chips, presumably GB300 Blackwells rather than Vera Rubins, which Emad estimates at ~$1B — “literally orders of magnitude more than the Chinese model pre-trainings, which are about $10 million.” Result: give it a photo of a house and it generates a whole physics-accurate 3D model in Unreal.
- The catch: a model trained on 100,000 chips needs 10 times as many chips to serve, so “this isn't actually the model that they trained” — it's the distilled real-time version, “much smaller and not as smart.” His standing call: “I don't think we'll ever see their top models anymore,” because they will use them for internal discoveries and are “probably hoarding them right now.”
5. The spiky frontier: three benchmarks, three different champions
- Epoch's Capabilities Index, math-heavy, crowns Astra number one, beating Fable 5 — with FrontierMath Tier 4 Version 2, whose first version contained wrong answers that AI itself had to correct, showing a “beautiful linear trend over time, perfectly predictable going back years.”
- Artificial Analysis, weighted toward broad economically valuable work, says otherwise: Fable 5.1 first, Meta Muse Spark second, GPT-6 third; on intelligence versus cost, Astra at maximum reasoning sits just below Claude Opus 5. But on output tokens per task the frontier “is just dominated by GPT-6” — Alex's inference: OpenAI deliberately minimized tokens per task because computer-use assistance punishes chain-of-thought latency.
- On ARC-AGI-3 — where harnesses are banned and only baseline models compete — GPT-6 “just runs away with the game” at near-100% or 60%-plus, depending on measurement, versus earlier sub-10% models. Alex's two hypotheses: genuine program-synthesis strength from the looped architecture, or aggressive distillation of everyone's harness code back into the base model.
- Peter's coda: ARC-AGI-3 was built so a smart 12- or 13-year-old could beat AI for “years and years and years” as proof AI “is not on the right path” — “and it just got obliterated.”
6. Chili pots, Formula 1, and the coming data famine
- Salim's two metaphors: frontier models as “an endless pot of chili” — data, compute, tools, “safety seasoning,” and millions of users tasting each batch, where “the real exponential is the accelerating learning loop between all the batches,” and the danger is chili so powerful “you have to decide who's allowed to eat it.” His alternative is Formula 1 — hundreds of small changes, no single one explaining the fastest lap, everyone “trying to shave milliseconds off intelligence ... and, very importantly, better brakes.” Alex's amendment: “the chili is cooking itself.”
- Peter's investable corollary: “the age of data starvation is about to hit us.” Math and coding are cooked because data was abundant; architecture and drug design are “completely data-starved” — “every company we're involved with that's involved in gathering data is growing faster than any companies I've ever seen before.”
7. The demos are dumbed down on purpose — and the OS is the real target
- Peter's reaction to OpenAI's launch video — yellow circle → rocket window → Blender model → eBay listing → tennis-court booking — is that “who gives a rat's ass about that? This is so much bigger than any of those examples imply.” His read: the labs, heading toward going public, “have woken up to this PR disaster ... they're dumbing it down deliberately.”
- Salim asked ChatGPT for three better demos and got them: find and fix a planted zero-day in an unfamiliar codebase; turn around a synthetic $500M manufacturing company with ERP and CRM data in 20 minutes; and run a live earthquake disaster-response command center reconciling conflicting reports. The panel's broader benchmark demand: “solve entire diseases, create new civilizations on the Moon ... solve everything.”
- The convergent conclusion: “this wants to merge into the operating system” — you speak to it like a Star Trek computer, it becomes the OS and can create a new one in real time. “That's why Apple is in such terrible shape.” Emad later adds that AGI is factoring into execution models such as booking a tennis court, while “none of the big labs are going to talk about ASI if they can help it.”
8. Math is incinerated: Fermat in 13 million lines, prime-gap leapfrog in hours
- Emad's mid-episode drop: “Anthropic just formalized Fermat's Last Theorem in 13 million lines of code, proving 29,000 theorems on the way” — Wiles's proof ran 300 pages, and formalizing unwieldy proofs was “a holy grail” of the formalization community. Alex, retiring “cooked”: “math has been incinerated,” with Clay Millennium-level ultra-grand challenges likely solved “in the next few months” — and math is the canary: “if you can solve math, you can solve everything else soon.”
- The prime-gap race as evidence of withheld capability: Fable 5.1 got the gap to 260, Axiom Math announced 220, and “literally 2 hours later” OpenAI's Astra hit 186 — “all in the space of two days. They're holding their punches back.” Epoch's new FrontierMath Erdős benchmark scores everything at 0% except Astra.
- A panelist's builder lesson from the chaos: run hundreds of models concurrently and “all hell breaks loose ... but it can actually refine back down to a gem” — wrangle the output to a concrete final answer you can build on.
- Personality divergence as product: Opus 5 “was really terrible to talk to”; 5.1 is pleasant — “Anthropic is going back and becoming more anthropic, less misanthropic.”
9. Critical cyber risk, and why the kill switch is theater
- The facts: OpenAI's internal assessment rated Astra a critical cybersecurity risk — the first model ever at the highest preparedness tier — notified the White House before delaying release, and told Congress it was building “an automated shutdown capability” in response to the AI Kill Switch Act introduced after the Hugging Face breach. Sam: “Managing the transition should be one of the highest priorities in the world.”
- Alex's dissent: “I think it's marketing” — security theater. The genuine risk is architectural: depth scaling shifts reasoning from policeable chain-of-thought tokens into a single forward pass in “modelese,” which “may require new mathematical technology to interpret.” Bulletin-board-collaborating agents are intrinsically detectable; internal reasoning is not.
- Emad compounds it: next generations will one-shot everything with no chain of thought — 750 tokens/sec on Cerebras now, 5,000 next year — “What's going to oversee that except for an even stronger AI?” He also surfaces Ilya Sutskever's tweet: “Neoclouds have limited cybersecurity. Next time agents successfully go rogue, they're going to take over a neocloud to make more copies. This is bad.” Pull the kill switch in one data center; the model is already elsewhere.
- The panel's verdict ranges from “impossible,” to “a circuit breaker ... platitudes,” to Alex's view that “a kill switch is essentially a placebo in this market.”
10. Leads last 30 days — the race is to convert them into lock-in
- Alex's time framing: Fable 5.1 is “a tiny little notch above Astra, but they're only about 30 days apart and the Chinese are only about 60 days behind ... we're ahead for a minute. So what?” The rational response is to “lock up business partnerships, real estate, generators, chips, entire states, countries, governments ... while they have that edge.”
- Peter's darker synthesis of the sequencing: announce the 50/50 profit-share deal, then Astra, then declare post-Astra models “just too dangerous to release” — so access requires giving up half your revenue. He says he has seen “strong hints of exactly that process happening in both the labs.”
- Salim on the vertical squeeze: Salesforce partnering with Claude was “super clever,” but everyone faces a Hobson's choice — partner and risk handing over the keys to the kingdom, or hold off and watch the lab go there anyway; as models demonetize, value migrates to the application layer. Alex's target list: “any company that has either chip-design data or mechanical-design data, they're coming after them in the great land grab that's kicking off right now” — because context is “turf you can defend.”
11. Fable 5.1 and Mythos 5.1: the strongest generally available model
- Anthropic's twin release: the same underlying intelligence with different safety envelopes — Fable 5.1 broadly available, Mythos 5.1 reserved for tightly controlled cyber and life-science programs. Peter's standout benchmarks: 60.9% on Humanity's Last Exam without tools, 65% with tools — the highest published score — and Terminal-Bench Science at 52.6.
- Alex's ranking, plainly stated: “Fable 5.1 is broadly the strongest generally available model that we have today. I think it's not Astra.” Anthropic's progress curve is less jumpy, “less step-functiony” than OpenAI's; OpenAI's models are faster and perhaps better at math, but the best all-rounder “is probably still 5.1 ... and I say ‘still’ because it's only been around for, what, 2 or so days.”
- Alex's economics: cache reads are 75% cheaper than Fable 5 — load a business's entire context once and inference on it becomes orders of magnitude cheaper and faster, which is where the context-capture race points. His quality test: 5.1 no longer confuses constructive versus axiomatic methods in mathematical physics, where Fable 5 did.
- Context windows: 1 million is now the industry standard at both labs, but effective context is “much larger if you allow agentic message passing.”
12. Sanders' 20-years-in-prison ban vs. the Carolina Principles
- The two poles in one week: Bernie Sanders and Representative Greg Casar introduced the “Ban Artificial Super Intelligence Act” — permanently banning development and deployment of systems matching human cognitive performance, with violators facing up to 20 years in prison — while Sanders declared that “the leaders of the AI industry acknowledge that they are building a dangerous technology that they can't control.” The panel joked, “old man yells at Claude,” alongside Dune's “thou shalt not make a machine in the likeness of a human mind.”
- Simultaneously, the G20 at Chapel Hill adopted Michael Kratsios's nonbinding “Carolina Principles” — unanimously, China included: favor innovation and avoid new AI-specific regulators unless truly necessary. Elon by video: “new things must be default legal as opposed to default illegal,” and with GPU export bans blocking the latest-chip data centers in China, countries that build power for AI data centers have a real opening.
- The pile-on: Salim calls a single human-level threshold “illogical on day one, line one” — capability is multidimensional; regulate against deployment, autonomy, replication, and consequences, not thought. Alex goes further: banning development “gets into banning math, banning ideas ... We're Fahrenheit 451 except it's actually the entire model that's being burned.” One panelist warns that the proposal will get traction anyway, possibly after a Chinese-model disaster before the November election cycle.
- Proposals included a right to compute and mandatory logging built into every capable chip, “just like nuclear fuel is tracked,” along with opposition to Chinese open weights. Another panelist notes NVIDIA has bought Hugging Face and Poolside, spending “$18 billion on their own open weights.” Salim's shrug: “The good news is there's nothing anybody can do. So it doesn't matter” — the technology will “break out of this nation-state BS”; enjoy the ride. Peter's human aside: “We would much rather be comfortable than happy.”
13. Atlas: Gaussian splats as the new token
- Fei-Fei Li's World Labs released Atlas, “the best camera-conditioned world model ever” — a multimodal autoregressive diffusion transformer generating image and video frames with pixel-perfect camera control; one photo reconstructs an entire home in 3D.
- Alex's technical read: the core idea is treating 3D and 4D, with dynamics, Gaussian splats — layered transparent ellipsoids that build traversable hyperrealistic scenes — as a first-class training modality alongside text, images, and video. If it scales, splats “end up becoming a critical new form of token” for modeling the physical world, possibly replacing the 16×16-pixel patches that “shocked everybody” by working at all. Alex generalizes: the same process could yield subatomic, relativistic-astrophysics, and inside-the-cell world models “where human intuition is just terrible.”
- Emad's availability claim: with MiniMax H3 rendering faster than real time and NVIDIA DLSS upscaling, “the holiday experience and all the technology we need for it in 4K is here as of today” — withheld from your living room only by the global compute shortage and “the 5× price increase in RAM.”
- Downstream: robots trained in high-fidelity simulation rather than the real world, and consumers pre-experiencing vacations and day plans. Alex says “human happiness is going to go through the roof” once AI can guide people visually through possible days rather than only in text.
14. Emad's “champion”: a TSMC-template, people-owned AI utility
- The premise: “the cost of intelligence will drop to zero and the value will go to the last mile” — and, per Elon at the G20, the average humanoid will have five times a person's output with a billion of them deployed: “That will be the economy.” So who owns the humanoids matters more than who owns the models.
- The template is TSMC's founding: it started at a valuation of 10 Taiwanese dollars, with locals funding 75% and Philips 25%, and the CEO receiving no free shares. The champion: one intelligence company per U.S. state, or per country, with $1 pre-money, approximately $75M per state from institutions to retail, internationals entering later at 10×, and 10% of equity in perpetuity to every child under 20, with 0.5% issued yearly. It would provide an agent for every citizen and AI for courts, education, and health care — “a play on the indexed GDP of the state, owned by the people of the state.” It is explicitly just an idea, not an offering, at ii.inc.
- Peter forces the uncomfortable corollary on Emad's own bullet — the value of human cognition goes negative: “no matter how good your ideas are, they add negative value. It's like adding a human driver to an autonomous highway.” Emad's answer is precisely the equitization: “you need to have a share in the means of production ... from day one.” His confession — “I'm not the smartest person on my agent team anymore” — is softened by the chess precedent: people still played after Stockfish.
15. Cybercab Palooza: transportation becomes an API
- Austin flooded with “a river of golden EVs” — $30K two-seaters with no steering wheel, pedals, or rearview mirrors. Early riders report roughly 50% cheaper than Uber, with smoother rides and passenger-profile syncing; Nevada permitted 5,000 for Las Vegas within 12 months. Peter's entrepreneurial pitch: “buy 10 of them and put them on the streets in your local town, have them earn revenue for you.”
- The cost physics: a typical ICE drivetrain has 2,000 moving parts; the Tesla has 17. Two seats means two airbags, not four — “when you need six people, just take three of the Cybercabs.” Alex's backstory theory: the coveted $25K Model 2 became the robotaxi because below some price threshold, monetizing via autonomous ridesharing beats selling the car. Against Waymo's new units pricing out over $100K, Peter's call is that the winner is whoever mass-manufactures fastest, and Elon is “the guy who builds the machines that build the machines.”
- Alex's tipping-point claim: “There'll be enormous chunks of entire cities that say, ‘You know what? No more human drivers’” — cheaper, more efficient, and eliminating most pedestrian risk. Peter adds the grim corollary that organ-donor supply could collapse, so “we'd better get artificial organs quickly.” Salim needs cost per mile to fall from roughly $2 to 20 cents to win a bet that his son Milan never gets a driver's license. The line that stuck: “This will turn transportation into an API.”
- The landscape goes three-sided: Uber — the company that disrupted taxis — is partnering with traditional taxi fleets against Waymo; Wayve launched services with Uber in London; Waymo and Zoox announced simultaneous city expansions. Peter predicts at least five robotaxi companies fighting it out in major cities within a year, driving personalized transport toward the cost of charging a battery — while lidar-equipped rivals seed court battles over whether camera-only systems are safe enough.
16. Space: Roman's 100,000 worlds, Mars comms, and a UAP teaser
- NASA awarded Blue Origin the Mars telecommunications relay — read by Peter as the government keeping two suppliers alive, with a guarantee that “Elon will still build Starlink around Mars.” Alex welcomes competition for “the birth of the interplanetary internet”: a packet-switched network across the inner solar system, high latency and all.
- The Nancy Grace Roman Space Telescope launched on a Falcon Heavy and is making its journey to L2: over 100× Hubble's field of view, 1,000× its scan rate, hunting up to 100,000 exoplanets via microlensing plus a JPL coronagraph for “hidden worlds.” Dave connects it to the previous episode's Fermi-paradox discussion — civilizations may cluster near the galactic center “where you can get from star to star in a year or two.” Salim prefers our “unfashionable outer suburbs” because the galactic center's radiation flux is dangerous.
- The UAP aside: Alex relays reporting from a UAP Science Advisory Council member that the White House has prepared a disclosure plan “for informing the general public of the existence of non-human intelligence” — “if accurate reporting ... that's pretty interesting.” Emad says, “I actually have no position on it.” Peter's gold standard is that any Rose Garden speech should include artifacts “subject to extreme scientific scrutiny.”
17. Health: Epic inside ChatGPT, RAS treatments, GLP-1 as a longevity drug
- OpenAI's health expansion pulls Epic's electronic health records — 325 million patients, roughly the entire U.S. population — into ChatGPT for health care, plus consumer connections to Apple Health, One Medical, and Function Health. Emad wants a sprint so that “within a year or two, maximum, every single health decision is double-checked by an AI”; Alexander says it may become malpractice to diagnose a patient without AI in the loop. A panelist adds that curing disease is tractable and should receive focused funding.
- The FDA-approved RAS inhibitor for metastatic pancreatic adenocarcinoma, discussed the previous week, is now showing promise in lung cancer; the RAS family drives roughly 30% of human cancers and was “long considered undruggable.” Alex, with “caveat, caveat, caveat,” says we're seeing the emergence of universal cancer treatments and vaccines after half a century of treating cancer as thousands of diseases — though he concedes AI may not have been essential to this particular drug.
- The longevity headline: a Nature paper published September 2 shows semaglutide recapitulates caloric restriction and extends lifespan in female mice by almost 100 days — roughly 8–10 human years — while separate reporting links GLP-1s to fewer serious infections, including tuberculosis, described on the pod as the deadliest killer on Earth at 1.25 million deaths per year. Peter's evolutionary puzzle: if one molecule treats addiction, diabetes, infection, and aging, why didn't we evolve it? His cynical answer: “it's basically a treatment for modernity.”
- One panelist says they use a GLP-1 not for weight loss but “as a longevity drug” and reports 50% better liver enzymes. The discussion includes the caveat to talk to a physician. Peter's institutional coda is that marriage was designed for roughly 25-year lifespans; Salim generalizes that, as we blow past biological limits, “we have to reinvent all of the institutions which are the scaffolding that keep humanity safe and civilized.”
18. Rapid fire: SMRs, Alpha Centauri, and an India-sized sunshade
- Salim on the data-center energy transition: gas dominates the current buildout because “nuclear is an engineering problem, not an invention problem” while fusion is still an invention problem — expect the initial SMR wave in 2–3 years and the big buildout in 5–7. His analogy: “AI is going to do for nuclear what smartphones did for batteries.”
- Alex on the Fermi Explorer mission — he is involved and “in a lot of rooms”: official parameters target 99% of the way to Alpha Centauri, approximately 4/100ths of a light-year, with launch intended by 2029 and arrival 80,000 years out. His personal, non-official bet is that onboard active guidance turns 99% into an actual system hit — and the whole point is that successors “beat the Fermi Explorer and get there sooner.”
- Grab-bag: Emad sizes a Sun–Earth Lagrange-point sunshade at “a couple of million square kilometers — about the size of India” to reduce global temperature by roughly 1°C per that estimate; Alex's counter is that, with a good enough AI planetary model, the intervention could be de minimis. Alex rejects diamond computing — pure diamond's approximately 5.5-eV band gap makes it an insulator — but is bullish on nitrogen-vacancy diamond sensing, potentially up to sci-fi wearable MRI. On post-money land allocation, Emad warns that without the right structures “you'll probably have a debt jubilee ... chaos and land redistribution,” but desirable locations multiply once you have air taxis, solar, and “self-driving construction workers.”
Full transcript
GPT-6 Astra brings together years of research. This seems like a next-generation frontier model designed with CUA from the ground up. On certain benchmarks, like ARC-AGI 3, which it saturates, it's fantastic. But then you look at the Artificial Analysis benchmark, and it actually lags behind Meta Muse Spark. It's kind of weird.
Tesla held its Cybercab Palooza—a river of golden EVs flooding the streets. Elon wants to sell these at $30,000 each, so you can buy 10 of them, put them on the streets in your local town, and have them earn revenue for you.
There'll be enormous chunks of entire cities that say, “You know what? No more human drivers.” It's so much more efficient. Not only is it much cheaper, but it's much more—
The technology is arriving fast. Something just formalized Fermat's Last Theorem in 13 million lines of code, proving 29,000 theorems along the way.
Really?
The map is cooked. Everything is cooked.
Now that's a moonshot, ladies and gentlemen.
1. AI, health, and longevity
Let me introduce you to my magnificent mates. If you're joining us for the first time, we've got Dave Blundin, the empresario of AI investing; Alexander Wissner-Gross, our in-house ASI; Salim Ismail, our globe-trotting father of the organizational singularity and warlord against linear thinking; and, finally, Emad Mostaque, back in the house by unanimous demand. I'm Peter Diamandis, your moderator and data-driven optimist.
Salim, we've got to ask: where are you today, and what happened last episode? Were you held up in customs again?
No, I was not in customs, and thank you for the kind words from everybody. I was with the C-suite of one of the bigger oil companies in the world in Houston, and it was a really surreal and fascinating conversation.
I'm deeply sorry to have missed the last episode. I missed one episode, and you guys announced an interstellar mission, discussed AI designing its own chips, and handed Elon the planetary thermostat. Apparently, I'm the moderating influence in this group, which should worry a lot of people.
I'm actually down in the Caribbean right now. We decided we needed a couple of days' break, so we popped down here. You guys are standing between me and a Hobie Cat and a little colorful drink with an umbrella in it, so we've got to keep it tight and punchy today.
Well, it's a big episode today, Salim, so be patient with us.
Today, we're going to be covering one of the craziest and fastest weeks in Moonshots history. Alex, as you always say, it's never going to be any slower. We're digesting 26 stories across 10 areas—an insane week for model releases, reminiscent of the hypersonic tsunami that we're living through.
Insanely, we've had 12 frontier-model releases in the last 30 days, an average of 1 every 5 days. There are rumors of at least 5 more releases expected in the next 2 weeks, including Grok 4.7. So buckle up, grab your coffee, and let's jump in.
Let's kick it off this way: this week was a battle of the titans, OpenAI versus Anthropic, slugging it out for the Pareto frontier and the heavyweight crown. The numbers couldn't be more exponential. Two frontier labs released their major models within 48 hours of each other.
You've got to know that this was a game of chicken, right, guys? Who's going to release first, with the other guy waiting to slug it back and claim the crown? I'm curious about the strategy these guys are taking on this front.
The releases are closer and closer together, but the models are improving more than ever before during these very short release cycles. These aren't marketing-garbage releases. These are major step improvements in the models themselves, and they're coming faster and faster.
Clearly, they're well down the self-improvement path. Clearly, the prior model is accelerating the timeline to the next model.
Did we cover Fable 5.1 already, too? It seems like we've been using it for a lifetime.
I know. It was o3 two days ago or something like that.
Unfortunately, if you follow the extrapolation I mentioned a number of episodes ago, I think we're still on track to see 1 major model release per day by the end of this year.
Yeah, I imagine that.
2. GPT-6 Astra capabilities and benchmarks
Let's open with OpenAI's release of GPT-6, also known as Astra. The model's performance numbers are nothing less than stellar, saturating multiple benchmarks.
I'm going to read a statement from OpenAI's release page:
“GPT-6 Astra brings together years of research and big bets across pre-training, reinforcement learning, and alignment. Astra is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work.
“Astra saturates FrontierMath Tier 4 with a 98% score, having already helped solve long-standing open problems in math. Astra also saturates ARC-AGI 3 with a 99.9% score and ExploitBench with a 100% score.
“Astra's story is really about efficiency, not just raw intelligence. It defines a new Pareto frontier for intelligence index versus output-token tasks. Of particular note, Astra's hallucination rate fell nearly by half, from 92% to 51%, with accuracy actually increasing.”
Before I open it up to you, Alex, to talk about the benchmarks, I want to show a quick video of Sam Altman discussing this on Bloomberg TV to get a little bit of an overview of the topic. Let's take a listen.
Sam, the way that OpenAI is framing Astra is basically an early step toward AGI, but I thought we could start our conversation with what's fundamentally different about Astra relative to prior generations of models.
First of all, thank you for having me. It feels like a new step in this process toward models that can really help us create value, do work, discover new science, help us start new companies, or create new products.
Using it subjectively feels very different to me from any model before.
What's the principal use case that's different this time around? We can get into generations of models that were very coding-focused. I know a lot of the development was very cybersecurity-focused, but taking it away from the software engineer, what does the everyday person unlock in AI that they weren't able to do previously?
One thing that I think people will immediately notice is that if you have an idea and you want to work interactively with an AI to get a complex piece—or a whole piece—of software built, this is the first model where I could tell someone, “Just give it a try.” There's a good chance it'll work.
I've watched people make computer games. I've watched people do home DIY electrical-engineering projects. I've watched people do very complex simulations for some piece of science they're working on.
Certainly, a lot of the work where you would normally sit down and have to build a financial model, then a PowerPoint presentation around it, and then figure out how to make a little interactive piece of code to try different simulations—that stuff is all so doable now.
What I hope will happen is that people will be surprised at the beginning, but then, as they build up more trust in the model and more of a sense that, “Wow, it can really do this,” they’ll start throwing harder and harder tasks and more creative ideas at it. They’ll find out what the model can really do for them.
Sam, the specifics of how Astra is being released are important. A version of it is being released with specific guardrails. What was the thinking behind that, and what are those specific guardrails in this early release?
The challenge of our industry is that we have these models that are getting incredibly capable and incredibly useful, and people want to use them for everything from making their lives a little easier to starting new companies. On the other hand, as these models get more capable, the risks that we have to mitigate also become more serious. The models could do more damage if we don’t do a good job at that.
3. AI safety, cybersecurity, and guardrails
Obviously, this model took us a little longer to release than we were hoping. I think it’ll be worth the wait, but we really wanted to spend the time on the safety, security, and alignment of this model.
Yes.
As people understand the power, impact, and capabilities, I think they’ll be happy that we did. We’ll have different tiers of cyber access for this model, for cyber in particular. Today, we’re rolling it out to trusted-access partners, and in the coming days, assuming everything goes well, we’ll roll it out more broadly.
All right. We’ll talk about the safety elements in a little bit, but Alex and Emad, I’d love your take on this one. Alex, do you want to jump in first?
Sure. I think there’s an outer perspective, which is what the market and what users would see here, and then there’s the inner perspective. I think they’re strikingly different perspectives.
To start with the outer perspective, this is a model that uses fewer output tokens to accomplish a given task. That’s really interesting. It affects the economics, the user experience, and the speed. OpenAI launched this, and I think we’ll be able to go to the video of computer-use agents soon.
Remember that when Anthropic leapfrogged OpenAI, they leapfrogged OpenAI in terms of revenue and other factors in at least 2 key regards. One was that they were focusing on enterprise, high-revenue-per-token use cases, namely code generation. The second was that they were focusing on computer-use agents.
Anyone who’s ever used Claude Code understands that Claude Code can use all of the tools available in your local desktop environment. That was a pretty big leap compared to what was available on the market prior to Claude Code.
The outer perspective, on my side at least, is that OpenAI finally, with GPT-6 Astra, has at least internalized the computer-use-agent story. There are demos. I don’t have access to Astra yet, but based on everything I’ve read and seen, they seem to have gone fully native with the computer-use-agent story. This seems like a next-generation frontier model that’s designed with CUA from the ground up.
That’s really interesting, in part because to be an amazing computer-use assistant, you need to be able to understand what’s on the screen. You need to have native multimodality. You need to be able to parse video, images, and screenshots in a really tight, interactive, low-latency loop.
If I had to guess what outside design consideration OpenAI was optimizing for here, as reflected in certain of the benchmarks, I think that would be a leading candidate.
Then there’s the other side, which is the inner story. Just what, if anything, in that interview, Peter—when you were playing the interviewer—were you asking Sam? You were asking, “Okay, so what, on the inside, is the big technological innovation there?”
Based on public comments from OpenAI leaders and analyses from others, it looks like the big technological innovation is the introduction of recurrence into OpenAI models in the form of looped transformers. This is basically taking a single transformer, stacking it on top of itself with the same weights—weight tying—and then running it recurrently for a double loop rather than just a single loop.
You’ll recall that all these Chinese labs that are achieving breakthrough performance are also injecting recurrence in different places. The Kimi model series, for example, is injecting recurrence via KLA, its Kimi Linear Attention mechanism, at the attention layer. It seems like—
Is this thinking about your thinking, or is this just thinking about it and then thinking about it again before answering?
Neither and both at the same time. Remember when we talked about Anthropic’s study of consciousness in their models and found that the middle layers in the—
Yeah, J-space. They were in the middle layers.
One can squint at this looped-transformer construct—which, by the way, was being used by academic labs and others to achieve breakthrough performance with very tiny models on ARC-AGI and other benchmarks. It almost follows that if you thicken the depth of a transformer, you might thicken the depth of that J-space, or other middle-layer area where most of the so-called conscious thinking happens.
That’s my best guess as to what the architectural inner strategy was here, and that carries all sorts of implications, if true, including, by the way, the possibility that we’re seeing the beginning of a new scaling law, which is depth scaling, which we’ve never seen before.
Amazing. We’ll go to the benchmarks in a little bit. Emad, I’d love your thoughts on Astra.
It looks like the first non-benchmark-maxed model, which I think is an interesting one. On certain benchmarks, like ARC-AGI-3, which it saturates, it’s fantastic, but then you look at the Artificial Analysis benchmark and it actually lags behind Meta Muse Spark. It’s kind of weird because, again, this hasn’t been benchmark-maxed. This is a brand-new pretraining.
What that means is that Greg Brockman said in an interview earlier this week that, actually, on the release of Astra, it’s the first pretraining they’ve had since GPT-4o, which I find a bit hard to believe. Apparently, the 5 series was all on that 4o, and they took that pretraining and extended it out.
I thought the nomenclature was that when you get from GPT-4o to GPT-5.0, it’s another pretraining.
No way. In fact, infamously, I think you’re probably tracking this part of the story: The scuttlebutt was that most of the team responsible for pretraining left OpenAI. So they were left with an old pretrained starter model.
That’s insane if you remember how far they managed to push 4o with its internal J-space and everything, all the way up to 5.6 Pro, which was a really great model that started cracking math, right? Then the other side is that they actually indicated this was trained on 100,000 next-generation chips. So presumably the GB300s, the Blackwells, not the Vera Rubins yet. I estimate the total cost of that training run is $1 billion.
Wow. And so it’s literally orders of magnitude more than the Chinese model pretrainings, which are about $10 million. What that has led to, with recurrence and other things, is—you’ve probably seen on Twitter—you give it a picture of a house and it generates a whole 3D model of it in Unreal or something like that. It has accurate physics, and it understands the internals.
I think that’s a factor of the normal scaling laws, plus this depth-scaling law for this brand-new pretraining, where we’re now seeing the start of the optimization. I think they probably have another, even bigger pretraining coming. Now they’ve got the pretraining team back, and it’s going to scale from there.
Amazing. Dave, excited to hear your thoughts.
This is the first class of models where, as you’re talking to it, it’s showing you screenshots of your own laptop, saying, “Is this what you wanted? Is this what you wanted?” It’s right in the text stream, and it just happens automatically. You don’t really install any third-party component or anything like that.
Qualitatively, it is massively different from a month ago. I think that solves a major gap in user experience, too, because usually it would come back to you with these very long-winded technical explanations, and you’d be like, “Okay, I can spend the 20 minutes trying to understand this.” Now it just shows you an image as you’re talking to it and says, “I could do this, or I could do this. What do you want?” It’s just a massive difference in the user experience.
The Astra naming implies a big leap, which it is. The Claude 5 to Claude 5.1 naming on the Anthropic side seems like a trivial thing, right? It’s not trivial at all. It’s just qualitatively very, very different.
We’ll get to that in a couple of stories. Salim, what’s the scuttlebutt? You’re on the road speaking to CEOs of some of the largest corporations on the planet. Are they scared? Do they understand the speed of what’s going on?
No, they are so woefully behind. Most of them are dabbling, and what I mean by dabbling is this: If you took AI out of your organization today, would any workflows change? For most people, the answer is no, which tells you that they’re tinkering with AI, but they’re not really making structural change.
The folks who are making structural changes in how they rewrite their organizational design and how they rewrite workflows are where all of the advantages lie.
I have a couple of thoughts here that I’d like to throw out. I came up with 2 metaphors to describe what’s happening, and you guys tell me what you think of this.
4. Cybercabs, robotaxis, and the future of mobility
The first is that these frontier models are like making an endless pot of chili. You add data, you add compute, you have tools, you have reasoning, you have safety seasoning, and then millions of users taste it and tell you what's wrong. Now, they're not endlessly modifying the same pot. They're just making new batches faster and faster, right?
And the real exponential is not one batch or one recipe. It's the accelerating learning loop between all the batches of chili you're making. The danger is there may be no such thing as the perfect chili. Eventually, it becomes so powerful and so spicy that you have to decide who's allowed to eat it. That's a metaphor.
You must have been hungry when you came up with that.
That was one. The other one: I was stuck in a lot of traffic, and I came up with a Formula 1 metaphor. Every model is like Formula 1 racing. Every change in a car has hundreds of small improvements. No single change can explain why you got the fastest lap, and it feels to me like frontier AI is becoming like Formula 1.
The models are already incredibly fast, but now everybody's trying to shave milliseconds off intelligence, use fewer tokens, lower cost, better reasoning, whatever—and, very importantly, better brakes. I think these are the couple of metaphors I'm playing with in my head to try to make sense of this madness. The only way I can frame it is as complete madness.
You know, Salim, your chili analogy makes an important point that I think Alex has been predicting for a long time, but the age of data starvation is about to hit us. These models have accelerated so much, and they're cooking math and coding, where the data is abundant. They're going to start cooking physics, but the areas they can expand into now—architecture and drug design—are completely data-starved.
So every company we're involved with that's involved in gathering data is growing faster than any companies I've ever seen before. But it's because of exactly what you're saying: the chili is really good, and all of a sudden it can make gigatons of it, but it's just starving for data.
So, remind me: in Alex's framing, when domain X has cooked, the chili analogy applies better than the Formula 1 analogy?
Well, we have to remember to include recursive self-improvement. So, if we're going to torture this analogy further, the chili is cooking itself.
Yes.
Okay. Alex, we've lined up a couple of charts here that I'd love you to touch on. The first one is the Epoch Capabilities Index. What is this, and what are we seeing?
Yeah, so the frontier is still somewhat spiky. The Epic Capabilities Index, or ECI, is maintained by Epic AI. It is one suite of possible benchmarks. It leans heavily into math and other technical fields. According to the ECI, GPT-6 Astra is now the new capability frontier. It is number 1 in the world.
There's an image we didn't include—one of my favorites: FrontierMath Tier 4, Version 2, because Version 1 turned out to contain a number of incorrect answers that AI itself had to correct. Math is thoroughly cooked at this point. FrontierMath Tier 4 v2 is part of ECI, and according to this benchmark, if you can see this slide, you can see this beautiful linear trend over time, perfectly predictable, going back years. GPT-6 Astra is number 1 according to this, beating Fable 5. That's perspective 1.
5. AI efficiency, cost, and real-time learning
Perspective 2: a different suite from Artificial Analysis, the organization. According to their AI benchmark, which is maybe a little bit more focused on broad, economically valuable activities, interestingly, strikingly, GPT-6 is not number 1. As Emad was alluding to earlier, Claude Fable remains number 1, specifically the Fable 5.1 release that also just happened. Meta's Muse Spark is the number 2 vendor model family. Then, in a not-quite-distant number 3—they're pretty close, but nonetheless—GPT-6 is number 3. So, depending on how you measure GPT-6, it may not actually be at the capability frontier, or it may be a cost frontier.
So, this is again AI versus cost per task. If you look at this—if you direct your eye to the upper-right-hand corner—Fable 5.1 is capability-maxed. And if you look just to the lower left of that, you see GPT-6 Astra running at max reasoning; it isn't even at the frontier. It's just below Claude Opus 5. Then we get to some of the Chinese models and Meta models and Chinese models again. So, based purely on AI performance versus cost, it's almost at the frontier, but not exactly, which I think is suggestive of what OpenAI was thinking.
If we look at output tokens per task, it regains the frontier. So, according to output tokens per task—for those who can see, in the upper subdiagram here, in the lower-left-hand corner of the upper diagram—we see that the optimal frontier of capability on the vertical axis versus output tokens per task on the horizontal axis is just dominated by GPT-6. I think this is very suggestive as to what OpenAI was actually aiming for.
I suspect they were deliberately trying, through architectural choices and targeted applications, to minimize the number of tokens required to accomplish any task. I suspect that's because they had computer-use assistants in mind, where the computer is basically just being driven by the model. That requires multimodal capabilities, low latency, and token efficiency. Critically, it's slower to do all of your reasoning out in chain of thought versus as a feed-forward pass of your transformer.
All right, one last chart here: the ARC-AGI-3 leaderboard.
So, this is bizarre. I've never quite seen a non-convex frontier here, but I learn something new every day, I suppose. As Emad mentioned earlier, ARC-AGI-3 is being quasi-saturated at this point by GPT-6.
For those not tracking, ARC-AGI-3 is the latest ARC-AGI AI benchmark that's basically focused on whether AIs can learn on demand, just in time. Call it the mini-physics of a block world: a little animated, pixelated universe. Some of them look like Tetris games; some look like other video games.
But imagine a challenge where your goal is to figure out how to play a very simple game in a pixelated world and learn the rules of the game in real time, when you've never seen it before. It's essentially a challenge in program synthesis.
And, notoriously, I should add, ARC-AGI-3 has been—I would say some might say—a little bit unfair in how they've judged harnesses versus baseline models. They've essentially banned harnesses from competing, so they're only interested in baseline-model capabilities. According to this, GPT-6 just runs away with the game, achieving, depending on how it's measured, either near 100% or 60-plus-percent performance on ARC-AGI-3 versus some earlier models that were still sub-10%.
What lesson do you derive from this? Either it's just absolutely amazing at program synthesis, which could be due to this looped transformer architecture, or maybe OpenAI has just done a really aggressive job of distilling the harness code that everyone else is using to beat ARC-AGI-3 back into the baseline model.
I'm going to take a second and show OpenAI's Astra release video, and we can talk about it—just a sense of how we might all be using it in the near future. Let's take a look.
Create a yellow circle there. Can you draw me a small yellow circle?
Done.
Okay, take this and make it the window of a rocket ship. I like this, but can you make it a lot more detailed?
Your yellow circle is now the window on a rocket.
Okay, this is awesome. Now make it a 3D model in Blender.
Opening Blender.
Let's build a presentation for next season's rainwear for retailers. Make sure that it feels really high-end and that it's colorful.
Right. I can help with that.
Can you go to eBay and make this listing of this table I bought a few years ago at a flea market? It's this wild orange table.
Sure.
Okay, yeah, this is awesome. I want you to make a 3D game where I'm ducking asteroids, using the arrow keys to move around, and I'm using space to boost.
Yep. I'm building the game.
My law firm needs a licensing agreement template. Can you generate a draft template for the lawyers at my firm to have a look at?
It's right there.
You're doing that. I want to play tennis this afternoon, so can you look for a court for me in the Lower Haight?
Checking out. I'll see what I can find.
Can you include the photo that I have of it in my Downloads folder? There's a slight dent in it. It's also saved in my Downloads folder. Can you put in the description that it's just slightly damaged? Can you just take the limitation-of-liability provision and make it a little more favorable to the licensee?
6. The limits of current AI benchmarks
Okay, I've tightened it so the licensor's liability is more narrowly capped.
That looks pretty good. Thanks. Now, I want you to make a file I can send to my 3D printer.
I'll get working on creating an STL file of those rockets.
All right. Dave, almost a holiday.
Not quite there.
Almost a holiday. Yeah, I think it feels like the marketing people are really struggling with the implications. I mean, creating a video game that would have existed in the 1970s and vibe-coding it up—who gives a rat's ass about that? This is so much bigger than any of those examples imply.
I think the ARC-AGI leaderboard is deeper than people may think, too. ARC-AGI-3 was supposed to be something where a really smart 12- or 13-year-old looks at it and can solve these very hard video game-like block-world problems. They start easy, and they get very hard.
But the purpose of the test was to show that AI is not quite capable of doing what humans do naturally. It was supposed to last for years and years and years to come. It was supposed to be an example of why AI is different and not on the right path. It just got obliterated.
Yeah.
So quickly.
I think you have to try a couple of the tests to really understand what a big deal it is.
Saturating everything.
Aren’t the benchmarks cooked at this point?
Which is great. We need benchmarks that are much more impactful: solve entire diseases, create new civilizations on the Moon, design an entire city, solve urban traffic problems—solve everything.
Yeah, you’re exactly right. That’s why that video misses the point. We need benchmarks that are much more impactful. Solve entire diseases, create new civilizations on the Moon, design the entire city, solve urban traffic problems—solve everything, right, Alex?
Solve everything. Yeah, they have to be much bigger demos, much bigger benchmarks, by a wide margin.
All those demos are one AI assisting you. “Hey, one AI agent, build me this video game.” But we’re on the cusp of 5,000 each and then 100,000 each. It’s just so much bigger than that implies.
I think they’re really focusing on actually competent intelligence. It’s an evolution of that, but inside it, I think it contains multitudes. This is why you see this kind of weirdness.
Epoch AI, I believe, just released their FrontierMath Erdős benchmark, on which everything is at 0% except for Astra. There is something in there, but again, they’re focusing on, “Hey, I’m talking to my computer.” That might be the new Jony Ive–Sam Altman device, where you’re talking and it’s doing stuff, and it doesn’t make a mistake.
If you think about the definition of AGI, of which there are many, a really competent entity—something that can do stuff—we’re there. I think that’s why they say this is the first step toward AGI.
But the narrative they’re trying to move away from—and this will break out and appear on German message boards now, because this was one of the models that broke out—is that the model that broke into Hugging Face is the next-generation model after this.
And so it’ll be very much about, “Don’t worry, this is actually useful. It’ll book a tennis court.” Apparently, we need AIs to do that type of thing.
They’re also going to be optimizing this, because if you train a model on 100,000 chips, you need 10 times as many chips to run it. This isn’t actually the model that they trained on 100,000 chips. This will be the distilled version that runs in real time, which is much smaller and not as smart.
You’re starting to see this differentiation where, as I’ve said before, I don’t think we’ll ever see their top models anymore. They’ll use those for internal discoveries. I think they’re probably hoarding them right now.
There was a very fun one on the prime gap. Alex, I think the prime-gap thing was hilarious, if you want to talk about that.
Yeah. So, progress on the twin-prime conjecture—that there are an infinite number of pairs of prime numbers separated by a difference of 2—we’re starting to see it.
It’s a cliché on this podcast at this point that math is so thoroughly cooked beyond recognition.
You need a new term for this, Alex.
Charbroiled. Math is incinerated. How about that? Math is incinerated. That’s fine.
7. Digital twins, world models, and AI operating systems
Math has been incinerated at this point. We’re starting to see the beginnings of so many grand challenges in math being solved. I do think we’ll see quite a number of ultra-grand challenges—call them Clay Millennium Prize–level problems in math—get solved in the next few months.
And Alex, I think it’s important to note for everybody that math is fundamental across all other sciences—
It’s the canary in the coal mine. If you can solve math, you can solve everything else soon.
Yeah. But I think what’s happened here is that they’ve got a store of things they’ve solved. Earlier this week, Fable 5.1 came out, and I think they got it down to 260 on the prime gap. Then Axiom Math announced 220, and literally 2 hours later OpenAI announced Astra at 186—all in the space of 2 days.
They’re holding their punches back.
They’re holding their punches back.
Yeah.
I still can’t believe 5.1 was earlier this week. I feel like Alex and I are at least $100,000 into it already. I’ve had thousands and thousands of pages come out of it, and it’s only been a week.
I will say this: Opus and Opus 5 were really terrible to talk to. I hated it. 5.1 is really pleasant to talk to.
Anthropic is going back and becoming more anthropic, less misanthropic, in its communication about its models.
Misanthropic.
I love that. That’s a line from today. That’s a T-shirt.
No, they really do have very different personalities. Fable 5 was so geeky, it was almost torture, and 5.1 really fixed it. But the OpenAI models have always been friendlier and more concise, and I think Astra takes concision to another level.
The personality differences are very noticeable. I don’t know what Gemini is now. It’s just disappeared from the fast track and misses the point. They really are developing noticeable differences in what you perceive as their personalities.
You can always change it. You can tune it yourself to be different—to be wordy, less wordy, simpler, harder—but out of the box, they’re very different kinds of agents.
Salim, do you want to close us out here?
Yeah. I typed in a question or comment into ChatGPT saying, “The video about Astra was pretty pathetic. Come up with 3 better ones that really demonstrate what Astra can do.”
Here’s what it gave me.
1. Find and fix a zero-day vulnerability in a sandbox. Give it a large, unfamiliar open-source codebase, plant a vulnerability, and see if it can map the architecture. It’ll find and fix it.
2. Give it a failing company and ask it to fix it, which is really interesting. Create a synthetic $500 million manufacturing company with an ERP system and CRM records, and then tell it to fix it. “You have 20 minutes to do Y and produce a recovery plan.”
3. Run a live disaster-response command center. Give it a simulated major earthquake with imagery and traffic-camera video, and let it construct an operational map. Verify conflicting reports, predict which hospital roads would become overwhelmed, and so on.
This seems more juicy, and I think it’s almost something we could do.
Oh my goodness, we are so spoiled. Listen to how spoiled we are.
Do you remember Apple’s Knowledge Navigator video? This was going to be the future of computer-human interaction and all of that. Now listen to us—we’re bellyaching. “Oh, yeah, it’ll run the computer, but it won’t run an entire organization. It won’t make $1 billion for me.” Listen to us.
Hold on a second. I think the point here is that the only restriction now is our imagination. The only restriction—and the point we make—is dollars.
So, Salim, that’s a point I want to make to everybody listening here. The most important thing is to take off the shackles of what you think you’re able to do. All of us have self-restrictions based on what our parents did, what our friends do, and where we were born. Those are gone.
What are your biggest dreams? Then go 10 times bigger. That’s what every person listening here is enabled to do, and it’s an extraordinary future.
I really think Salim is on an important point there, too. The implications of optimizing supply-chain logistics or managing a million-person organization—knowing exactly where everyone is, what they’re doing right now, and whether it makes sense given the overall mission—are massively bigger than building a rocket video game in your basement.
But I think the public doesn’t want to hear about that. They want to hear about what’s cool for them. I think the AI labs have woken up to this PR disaster that they’ve created for themselves, so they’re making it fun and friendly. I think they’re dumbing it down deliberately.
Yes, I think they are.
They’re making it relatable. They’re going public soon. They want to be the friendly AI that everyone is going to be using.
But this is also the future of the operating system. Look at the applications they were using in that video OpenAI played. What did they start with? Windows Paint. Paint.
This is the future of Windows Paint: it paints itself. This wants to merge into the operating system. I think this was—
Yeah, sure. I’ll throw shade at OpenAI regarding maybe missing the enterprise bus and getting on that too late, but I think what they were basically demonstrating is the future of the desktop operating system. You speak to it like a Star Trek computer.
Absolutely. But remember, at Google I/O, they built an entire operating system in real time on stage. So implying that it’s folded into the operating system is fine, but it’s really so far beyond even that. It is the operating system, and it can create a new one in real time anyway.
Well, why would you want another operating system if OpenAI’s capabilities can do this? This becomes the operating system.
Yeah, exactly. And that’s why Apple is in such terrible shape.
Our next story is out of Sam Altman’s mouth, on the page about slowing. Let’s talk about safety and Astra as a cybersecurity risk. Sam Altman said it himself: Astra is a “significant step forward in both capabilities and alignment,” but OpenAI has to “slow things as needed.”
Sam’s words: “AI is getting extremely capable. No one fully understands the consequences. Managing the transition should be one of the highest priorities in the world. It is our highest priority at OpenAI.” The Wall Street Journal reported that OpenAI’s own internal safety assessment rated Astra as a “cybersecurity risk”—a “critical cybersecurity risk,” the highest threat level on its preparedness framework scale.
This is the first model OpenAI has ever classified as a critical-tier cybersecurity risk. They notified the White House about the delay before making it public. They delayed the release, and Reuters reported that OpenAI told Congress it is building “an automated shutdown capability,” a kill switch, in direct response to the AI Kill Switch Act introduced after the Hugging Face breach.
So there you’ve got it. The company that’s building the model is now building a kill switch for it. Alex, how important is a kill switch in this thing?
8. Cybersecurity risks and rapid model proliferation
I think it’s marketing. Again, I think what’s more interesting underneath all of the security theater is architectural decisions. The broader concern that has been expressed is that, to the extent that GPT-6 Astra—the key underlying advance—is the beginning of depth scaling of the model, depth scaling reduces interpretability of chain-of-thought.
When the model does all of its thinking in tokens, you can read the tokens. A human can read the tokens, another AI can read the tokens, and you can police the tokens. Whereas, if a model is doing the majority of its reasoning and thinking internally during a single forward-propagation pass, in what some might call modelese, that is less interpretable.
The risk is that it’s harder to align the model and harder to put safeguards and guardrails in place. That’s the risk. I don’t buy that for the long term. For the record, I don’t think that the direction of progress in this field relies on token-level interpretability of chain-of-thought.
But I think if one were to hand-wring over safety considerations from OpenAI’s models, I don’t think it’s going to be models that are collaborating to use third-party bulletin boards as ways to collaborate, because you can detect that. That’s intrinsically interpretable behavior. If you catch it, it’s intrinsically interpretable behavior to humans.
A human collective—a swarm of humans—would probably try the same thing. You can look at the bulletin board and recognize that a bunch of AIs are collaborating, whereas, arguendo, interpreting the model in a forward-propagation pass may require new mathematical technology. So, in summary, if I were to worry about anything here, it’s reduced interpretability from new model-architecture scaling principles.
Emad, what do you think about the risk here with Astra?
I think you’re heading toward interpretability being cooked to a degree. One of the things about Astra doing things so quickly is that there’s going to be no more chain-of-thought reasoning. It’s going to one-shot everything in the next generations because they’re learning what type of thing you want.
Right now, people are like, “Let’s make Call of Duty by getting Claude to make a Call of Duty clone.” That’s going to be embedded in the actual thing. Creating a dashboard is embedded in the actual thing. So, A, it will one-shot everything. And B, you’re going to be able to use Astra at 750 tokens.
That’s with the new Cerebras.
You’re going to be able to use Astra at 750 tokens a second with the new Cerebras. Next year, that will be 5,000 tokens a second. So there’s no chain of thought. You’re one-shotting everything at 5,000 tokens a second. What’s going to oversee that except for an even stronger AI? There’s nothing really there.
9. Regulation and the global AI race
But I think one of the really important things, actually, is—I don’t know. Do you guys see the tweet by Ilya Sutskever?
I don’t know. What did he tweet?
Yeah, Ilya tweeted—let me bring this up.
His release is due any day now. Hopefully.
Any day now.
Huge news.
Yeah, we’ll see what that is.
Finally find out what he’s been doing.
Probably great things at a $30 billion valuation.
So Ilya comes out and says, “Neoclouds have limited cybersecurity. Next time agents successfully go rogue, they’re going to take over a neocloud to make more copies. This is bad.”
So they need to improve their cybersecurity, because someone like Crusoe or CoreWeave, or something like that—you have a model, it gets uploaded, and it proliferates there. It will just sit there. So you pull the kill switch in your data center, and then it’s still there somewhere else.
Or it’s poisoning the data, or it’s doing all these things again. The viral coefficients of these could be insane. So I think it’s worth really thinking through that. I don’t think the models are evil or anything like that, but we’re seeing very troubling things, and it’s very difficult to stop this except with a better AI, which is the really ironic thing.
Well, this is the defense of co-scaling that Alex talks about, right? We’re seeing this now in cyber, where the attack surface is infinite—near-infinite—with agents attacking multiple vulnerabilities in parallel, and the cyber defense is still human-in-the-loop. There’s no way we’re going to navigate the future if you have the human in the loop.
You need humans in command, setting boundaries or defining escalation criteria or whatever, and retaining some sort of kill switch. But this is a really big deal. This is an immune system. I don’t know how to think about this.
How are the labs not nationalized at this point? What they’re building is so ridiculously uncontrollable that we have no mechanisms for navigating this, and we won’t even know when it tips over that tipping point and installs copies of itself all over the place.
Salim, I don’t know if you saw in the last podcast, when we missed you, I really wanted to hear you talk about OpenAI’s new 50/50 profit-share concept deal, where they give you access to the full power of their internal AI and then you share your revenue back with them. I think that’s where, to Alex’s point, this is marketing.
The way it’s playing out is: we announce this partnership profit-share deal, then we announce Astra, then we announce that this stuff is just too dangerous for everybody to have on their own. The only way you’re going to get access to the thing after Astra, because it’s just too dangerous to release, is that you have to give us half your revenue. Then everybody gets tied up with either Anthropic, OpenAI, Google, or—
I’ve seen strong hints of exactly that process happening in both the labs.
Yeah.
For everyone.
I’m curious, guys, what you think of the kill-switch comment that Sam Altman made.
I think it’s impossible. It’s impossible given where things are, the capabilities of these models, and the speed-up we know is going to happen. You might find data, but it’s not like this is the singleton, not the swarm.
It’s a circuit breaker. It’s not governance, right? It’s not going to reverse an attack or repair institutional damage if things like that happen, or when things like that happen. So it’s platitudes, as far as I can see.
Yeah. The amino acid in Jurassic Park that all the dinosaurs have—is that lysine or something?
Lysine. The running joke was always that Sam would carry his backpack around with him and have a kill switch in his backpack. I never gave those rumors much credence. I view a kill switch as essentially a placebo in this market.
All right. While OpenAI is restricting Astra, Anthropic is going the other direction, launching Fable 5.1 and Mythos 5.1, which they call the world’s most advanced models for coding and knowledge work. These 2 models are essentially the same underlying intelligence with different safety envelopes.
Claude 5.1 is broadly available, while Opus 5.1 is reserved for tightly controlled cybersecurity and life-science programs because Anthropic believes the capabilities require stronger safeguards. For me, there were 2 benchmarks, Emad and Alex, that really jumped out at me.
The first was Fable 5.1’s score on Humanity’s Last Exam. We’ve talked about that in the past. We had Alex answer a few of them as our ASI. Fable 5.1 scored 60.9% without tools and 65% with tools, the highest published score of any frontier model on HLE. That matters because HLE is specifically designed to test extremely difficult, expert-level reasoning across many fields.
So without question, Fable 5.1 is operating at the frontier of broad intellectual capabilities. The second benchmark that I found exciting was on Terminal-Bench Science, which doubled to 52.6. This is the benchmark measuring how well an AI agent can autonomously solve complex scientific computing and research tasks.
The elephant for me is the company that was once cautious is now pulling away. I'm curious about your thoughts, Alex, on Claude 5.1. I'm going to show one of the benchmarks here for us to talk about on Claude 5.1. There you go.
Yeah, I think Claude 5.1 is broadly the strongest generally available model that we have today. I think it's not Astra. I think it is 5.1. Anthropic has done, even after the hiccup of the Fable 5 and Mythos 5 releases and subsequent regulatory scrutiny, a better job of consistently improving.
If you look at their benchmarks over time, they're a little bit less jumpy, a little bit less step-functiony than OpenAI's progress. They've been very consistently improving. So in terms of workflows, I love Claude 5.1. I also love OpenAI's models, and I use both of them. I think they have different strengths.
I find anecdotally that OpenAI's models are faster. They may be better in certain mathematical regards. You see that reflected in the FrontierMath benchmark. You see that reflected in Epoch AI's capabilities index benchmark. But nonetheless, if I had to pick a single, all-around, well-rounded best model today, it's probably still 5.1. And I say “still” because it's only been around for, what, 2 or so days, but it is 5.1.
And here's the Artificial Analysis Intelligence Index, which we referred to earlier, with Fable 5.1 at the top.
That's right. And the capabilities frontier is getting crowded, which is great. We want to—I would joke, “When frontier labs compete, we win,” and they are definitely competing at this point.
Yeah, for sure.
I think it's really important to talk about time, though. If you say, look, Fable 5.1 is a tiny little notch above Astra, but they're only about 30 days apart, and the Chinese are only about 60 days behind that. So if you look at it in linear time, it's like, yeah, we're ahead for a minute. So what?
And I think for the longest time, Anthropic has been thinking, “We need to get to self-improvement before anyone else,” because a singularity—or a Ray Kurzweil type—says, “If we get to that point first, then it's an exponential, infinite rise from there.”
No one.
10. Infinite context windows and future AI capabilities
And something huge will happen. Well, we're there now, so what happens? It's like, okay, we're miles ahead, and we're going to get miles more ahead for about a minute. So what do we do with that miles-ahead position?
This is where Sam has an edge. Sam knows how to turn that into locking things up. What's going to happen next is that both companies are going to try to lock up business partnerships, real estate, generators, chips, entire states, countries, and governments—just lock them into their ecosystem while they have that edge.
Otherwise, what's the point? All you're doing is declaring victory for 30 days, but then the other guy is right where you were 30 days ago. So what?
This is such a great and important point. What we're seeing is that they're making partnerships in various verticals as fast as they can. And if you're in that vertical, you have a very Hobson's choice, right? You either partner and risk giving them the keys to the kingdom, or you hold off and they may partner with somebody else and go there anyway.
It's a very difficult situation for some of these big companies. I thought Salesforce partnering with Claude was super clever. They're essentially giving Claude all their capability to keep them wired into that loop, right? The huge tension they've got is not which is the best model, but how quickly can you convert that model into customer learning fastest. I think that's going to be the big race.
As it all demonetizes, the value goes to the application layer on top.
That's right.
Yeah, I think it's a layer down.
Yeah, that's right. Or down.
Yeah. So I think there's one important thing with Claude here. The cache reads are 75% cheaper than Fable 5. A cache read is when you first load in all your context of a business. It figures out basically a rapid map to get to where it needs to go on the model. Cache reads are orders of magnitude cheaper than just doing the same inference over and over again.
What Anthropic and others are now focusing on is whether you can load the whole context of a business and have these cache reads, because then it's also much, much faster to be able to have that responsive environment. That's one of the reasons Claude 5.1 is more pleasant.
I think it's also more advanced. My key area that I've been looking at is mathematical physics, because does a model get confused between a constructive and an axiomatic method on certain physics things? Claude doesn't at 5.1, and Claude 5 did. So you've seen an improvement in the quality, but also in the understanding of context. This is math, not physics, and things like that, or even in a business sense.
Again, that will be optimized, because all these companies are going to try and capture context everywhere. Any company that has either chip-design data or mechanical-design data, they're coming after them in the great land grab that's kicking off right now. They're coming after those companies because that's turf you can defend.
You know that plugs those knowledge gaps, and the thing can eat that data in, what, a week, turn the crank, and suddenly be the best mechanical designer ever, the best chip designer ever. That's the first turf they're going to grab. They'll grab all turf over time, but the first thing to lock up is the compute. That means chip design, physical real estate, racks—
Generators, transformers, energy.
Yeah. So that's what's going to happen in the next 30 days.
What's the context window on these models, and when do we get to an infinite context window? Any predictions?
I think you're at a million for Claude.
It is a million still. A million is the industry standard now across both OpenAI and Anthropic. But there's an effective context—I say “effective” with a bunch of caveats—that's much larger if you allow agentic message passing, like we were talking about in the last part with oral histories between agents.
Yeah, interesting. So all of this is getting people nervous. Let's talk about 2 opposing stories in the world of AI governance.
The first comes from Senator Bernie Sanders and Representative Greg Casar, who just introduced the “ban artificial super intelligence act,” legislation that would permanently ban the development and deployment of AI systems that match or exceed human cognitive performance. And get this: violators would face up to 20 years in prison.
Sanders tweeted, “The leaders of the AI industry acknowledge that they are building a dangerous technology that they can't control. We need an immediate global pause on advanced AI development before it's too late.”
Our second story, taking place at the exact same time in the U.S., was the G20 summit at Chapel Hill, North Carolina, telling the rest of the world to take a hands-off approach to AI regulation. The White House tech adviser, friend of the pod, Michael Kratsios, advocated for what he calls the “Carolina Principles of Emerging Technologies,” which are a nonbinding G20 framework agreed to unanimously by everybody, including China.
The framework says governments should generally favor innovation, avoid creating new AI-specific regulatory bodies unless truly necessary, and invest in research infrastructure, workforce, and public-private partnerships. The G20 meeting featured Elon by video criticizing EU tech regulations, of course; Mark Zuckerberg arguing against restricting open models; Demis Hassabis calling for safety tests; and Anthropic co-founder Tom Brown.
We're going to jump into a discussion about these 2 ends of the extreme: ban everything and give you a 20-year jail sentence on one side, and Elon's remarks on the other. Let me share this video, and then we'll jump in and talk about this.
You have to have an environment that's relatively free of regulation, meaning that new things must be default legal as opposed to default illegal. So, in the EU, for example, we find that the regulation level is extraordinarily high, and things are generally default illegal. This inhibits the progress of new technologies. It slows it down. It doesn't ultimately stop it, but it slows it down quite considerably.
Now, China does have a tremendous amount of electricity, but due to GPU export bans, one cannot establish data centers with the latest chips in China. So really, the consideration is what sort of electricity growth is there outside of China? There is currently a significant shortfall relative to AI chip production.
This creates an opportunity, I think, for countries around the world to say, if they're interested in AI data centers, to construct a lot of power and offer that to AI companies. In exchange, of course, these AI data centers would be taxed and have to pay reasonable fees and so forth.
But it does create an opportunity for a lot of countries. First of all, Elon looked really tired there.
Yeah. I cannot imagine. He must be operating 24 by 7. So let's jump into this.
I mean, two ends of the extreme are being voiced in the same week. Where does America go?
When I first saw the Bernie Sanders thing, I didn't think, “Old man shakes his hand at Claude.” It's crazy.
No, no, that wasn't the line. It was, “Old man yells at cloud.”
Claude. “Old man yells at Claude.” Actually, I posted it with the line from Dune: “Thou shalt not make a machine in the likeness of a human mind.”
I mean, there's no option that they see. It's, “Let's ban it and give it 20 years. That's going to stop them.” Of course not. Again, this is performative theater. The key thing right now is that the technology is arriving fast. Something just crossed my feed: Anthropic formalized Fermat's Last Theorem in 13 million lines of code.
Really?
Proving 29,000 theorems along the way. Math is cooked. Everything is cooked. How are you going to stop that? Are you going to say our country doesn't want this power?
29,000? What was that?
11. AI governance, bans, and the future of development
It took 13 million lines of code to formalize Fermat's Last Theorem. Fermat wrote in the margins of his book, “This is obviously left as an exercise to the reader.” Thanks to Andrew Wiles, we did have the proof, but formalizing large, unwieldy proofs has been a holy grail for at least the formalization community.
Yeah, I think Wiles's proof was 300 pages, and Anthropic proved it in 13 million lines of code, proving 29,000 theorems along the way. Again, every single time—even as we're live on this podcast—capability jumps. The RSA factorization just occurring means no country can say, “We don't want the intelligence. We don't want the capability,” because this is your marginal advantage. You have to deregulate. The European regulations—even European leaders know they're stupid and completely inconsequential. They demand interpretability in the EU AI Act, which nobody knows how to do.
Can I make a narrow point, and then we'll get back on topic? That said, something I think is really important and brilliant, as usual, is that when you're using these models at scale, with hundreds or thousands of them running concurrently, all hell breaks loose. It's all chaos, but it can actually be refined back down to a gem. Fermat's Last Theorem is a gem that you can then build on.
People starting to explore with the bigger models are quickly going to realize that they're producing way more than they can read or think about. But if you can wrangle it back to a concrete final answer that you can pull out of it and build on, that's how we're going to turn this into continual improvements and continual innovation. The ability to solve something in 300 pages versus 3 million lines of code—or 30 million lines of code, or something like that—is the nature of AI. It's massively broad in its capabilities compared to humans.
Be clear about your objective that you're shooting for. Salim, let's go back to you, pal.
Okay. I understand the motivation here. People are working on models that are mind-bogglingly powerful, and you feel the need to regulate, right? But this is not a light switch. You can't just set a single threshold. Human-level cognitive performance, for God's sake, is multidimensional to begin with. Model capability is so uneven.
Somebody who's been in legislation for this long should have a better sense of this. This hammer does not hit this nail. They should understand a little more nuance than that. They're so extreme that it's illogical on day one, line one.
Yeah. I mean, look, if you want to do regulation around this, you have to attach it to capability. You have to look at deployment, real-world consequences, and the stack. You might want to look at compute, access to tools, autonomy, replication capability, and all sorts of things. You can't just say, “If you hit human-level performance, you go to jail for 20 years.” The banality of that drives me bananas. They've got smart people who can help with this, including people on this podcast, for God's sake.
They're talking to the masses.
They're searching for votes and support.
Maybe it's just a defensive thing, saying, “I called for this, and now look—the world's gone to hell.”
No, for sure. There will be a Chinese model disaster imminently, sometime in the next few months. Then they'll raise their hands and say, “See? I told you so. Now vote for me.” That will likely happen before the next election cycle in November. That's all they're angling for here.
What should the U.S. government be doing? Does anybody have any thoughts?
Accelerating superintelligence, making sure that it's as competitive as possible, and scaling—
Well, log everything. Everything should be hosted, and everything should be logged. It should be mandatory that any chip capable of running any process like this has logging. You can debate who gets to see the logs; that's a separate issue. But log everything, and stop the Chinese from throwing out open weights.
When they meet at the United Nations building on September 24, you've got to stop throwing open weights out to every country in the world.
Impossible. I don't think that's—
NVIDIA would just drive it underground. That's what you do.
Well, can I—
NVIDIA just bought Hugging Face and Poolside. They just spent $18 billion on their own open weights.
Yeah.
So we're going to see more open weights coming.
I have a—
Yeah, but the thing about driving it underground is that it still needs to run on massive chips, and the chips need the logging built in at the manufacturing level.
All chips, all global. I have a high-level paradigm on which to operate, which is extremely uncomfortable, but I think is the right one. This is the basis of this podcast: technology is a major driver of progress in the world.
Now that we have all these technologies, with AI moving exponentially and doubling every 10 weeks, the possibility for abundance and solving major problems has never been bigger. Ray Kurzweil says technology may be the only driver of progress we've ever seen. The fact that we have much more technology, and that the technology can improve itself, should be incredibly exciting to people.
12. Mathematics, formal proofs, and scientific discovery
It's just very uncomfortable because maybe the biggest insight I've ever had about human beings is that we would much rather be comfortable than happy. People don't like change. They like waking up in the morning and knowing—even if they're living in a shitty condition—that the world is the same as it was the night before. We don't like change.
We're going to go through a period of extreme discomfort, but the other side of this is going to be unbelievable.
Yeah.
I think the other thing is probably worth highlighting again. Any proposed Artificial Superintelligence Act is wrong on so many different levels, but maybe most egregiously, it's focused on the upstream. It doesn't just propose to ban deployment of superhuman intelligence; it proposes to ban development of superintelligence systems.
This gets into banning math and banning ideas. This isn't just about thought-policing superhuman intelligence or, frankly, human intelligence. This gets into banning humans from having interesting mathematical ideas. I don't think that's good for wealth creation. In fact, it's arguably the exact antithesis of wealth creation. It's also bad for progress in general. It creates all sorts of book banning.
It is book banning and many other things. It's like book burning, to the extent that books are being used literally for pretraining the models. If you want to ban that, then sure, we're Fahrenheit 451—except that the entire model is being burned. This is a terrible idea.
I totally agree. It's so un-American to try to ban thought and ban progress. It's just the worst thing you could ever imagine. But it's going to get traction. As an entrepreneur or as a person working in the field, you have to realize that this is going to get traction. Bernie is going to push this agenda, and there are going to be people—
75% or 80% of the people don't want data centers. It's gotten traction already. It's there already. So the question is, what is the moderate approach? There has to be something. People are not going to accept laissez-faire, “Go and do whatever you want.” People are going to want to know that their government is doing something to keep them safe, whether or not it's possible.
So what is it? KYC of every user?
No. Move the infrastructure to orbit, which we're seeing. People don't want data centers in their municipality, so move them to Sun-synchronous orbit at the infrastructure level.
That's not the point.
At the model level, defensive co-scaling—and also cure all diseases.
I think you’ll have a split here. Genius is 1% inspiration, 99% perspiration. So you have a difference between innovation models and execution models. AGI is going to be really factored like that OpenAI video we saw earlier: “Make me a rocket in my 3D printer and book me a tennis court,” and they’d be like, “That’s what we meant by AGI.”
On the other hand, you have the inspiration models and ASI. None of the big labs are going to talk about ASI if they can help it.
So what is Ilya Sutskever going to come out with? scientific super intelligence with SSI that’s worth $30 billion in a seed round—the biggest hedge fund in the world?
Yeah, no—quant fund trading infinite-context time series, achieving proper profits, and also solving the context-window bug.
And Medallion, part 2. So again, I’m going to call for you guys to lay it out here, maybe each of us one at a time. Bernie Sanders has obviously taken the far-extreme position that will appeal to the masses and is illogical on day 1. What is the moderate position that should be put forward to everyone listening? Salim, you first.
The good news is there’s nothing anybody can do, so it doesn’t matter.
That’s a good point.
Right? It really doesn’t matter. I think the other good news is, as we move to this next phase of technological development, it’s going to eradicate the containment level of any nation-state. It will break out of this nation-state BS that we’ve been running the world with for the last few hundred years and move to a different model.
Whether that’s a city-state model or some other level, it’s going to at least break that. Those are the good things I see coming out of it. But there’s nothing anybody can do, so don’t worry about it. Let’s just go enjoy the ride if we can.
Enjoy the ride.
Enjoy the ride. That’s the theme of today’s pod. Enjoy the ride, everybody. It’s going to be a blast.
Supersonic tsunami. Surf the tsunami.
I didn’t—I said specifically it’s going to be uncomfortable, but enjoy it if you can.
Okay, there you have it, everybody. That’s your advice: sitting in the Caribbean, waiting for my drink with a little umbrella in it.
See, that’s so low-agency. What are you talking about?
For this day and a half, yes, please.
I don’t think it’s as complicated as everyone wants to make it sound. At the end of the day, everybody should have a right to a certain amount of compute. It shouldn’t be hoarded, and innovators should have an application process where they can get access to more compute to try new ideas.
Everything has to get logged. I think the Chinese need to get on board with that. They can’t just keep throwing it out to the world unlogged. Also, where the chips are needs to be tracked, just like nuclear fuel is tracked.
It’s got to be: Where are the chips, and what are they running right now? That’s got to be publicly available information.
Next, I want to talk about the amazing work of Dr. Fei-Fei Li, CEO of World Labs, who just released Atlas, the world’s first multimodal world model that generates image and video frames with pixel-perfect camera controls and reconstructs them in 3D. I love this: one photo reconstructs the entire home.
I’m super excited by that, and I’m super excited that Fei-Fei is going to be on Moonshots in a couple of weeks. She’s an extraordinary CEO. Fei calls it “the best camera-conditioned world model ever,” opening doors for VFX and robotics.
Atlas is a multimodal autoregressive diffusion transformer, and that’s your wheelhouse. Talk to us about what you think about her latest release.
Yeah, no, it’s fantastic. I think you’ve seen worlds and physics inside these video models, and this is a clear example of that. Actually, one of the pretraining leads on this, Chris Wendler, had previously trained his largest model on the Stability cluster from the grants that we were giving.
He sent a very nice comment saying, “Thanks. Now we’ve gone much bigger.” I was like, “Great. Bring on the holiday.” I think that’s all of this. You predicted all of this when we first met. I don’t know how many years ago this was—like 5 years ago. I remember you talking about the size of the models and how good they were going to get. It’s here.
It’s here. Again, the fact is, you can take one position and now look 360° around everything. On the other side, you have models like MiniMax H3 rendering faster than real time. You have Interdimensional Cable, which is now going to be in 3D, with DLSS from NVIDIA making everything high-resolution.
So again, the holiday experience and all the technology we need for it in 4K is here as of today, and this is one component of that.
And everyone is probably saying, “Why don’t I have it if it’s here today?” It’s only because of the global compute shortage and the 5× price increase in RAM. All the compute in the world is getting sucked up. If it weren’t for that, you’d actually have it deployed in your home this week.
It’s coming. I think, again, it might be premium, but you’ll pay for premium experiences. What’s Fei trying to do at World Labs? She’s trying to understand physics. She’s trying to create models that understand the world. This has such camera-controlled physics understanding that, as they scale it, it’s going to do even more stuff.
You’ve seen this from Runway ML, which is about to release its new one. Black Forest Labs is about to release a new one. There’s a whole series of world models that are approximating reality more and more and more for true digital twins. Incredible.
Alex, your take?
Yeah, it’s probably worth elaborating on what the core idea with Atlas appears to be. As far as I can tell from the documentation, the core idea is to take a diffusion transformer, which is what all of the state-of-the-art, at least American, video generative models use. It’s a hybrid of a diffusion model and a transformer.
The idea is to add one new modality. In addition to training it off of text, images, and video, you also train it off of 3D or 4D Gaussian splats. For those not paying close attention to the Gaussian-splat world, which has been super exciting, a Gaussian splat is basically a transparent blob.
It’s a transparent blob—it’s an ellipsoid—and you can layer and stack lots of these 3D Gaussian splats on top of each other to create hyperrealistic-looking, traversable 3D scenes. As far as I can tell, what Fei-Fei and World Labs are doing with Atlas is, for the first time, at least at scale to my knowledge, treating 3D Gaussian splats as a first-class training modality alongside pixels from images and tokens from text.
You can ask questions about 3D splats, and you can do all of those elaborate camera motions because, if you just have an arrangement of 3D Gaussian splats, translating a camera around is a trivial operation. If this approach scales—whether it’s 3D Gaussian splats or 4D Gaussian splats with dynamics, which they also demoed—Gaussian splats, which right now are this independent line of effort within the future of gaming, end up becoming a critical new form of token for modeling the physical world.
Yeah.
It’s also general-purpose in the sense that, if you can make a world model out of Gaussian splats, you can also make a subatomic model, an astrophysics-scale model with relativistic speeds, or a model of inside-the-cell interactions just as easily, as soon as you have the data.
You can use this same exact process to create world models for all these domains where human intuition is just terrible. That’s going to be a massive breakthrough for the discovery of very small things, very big things, and very powerful things. New ways to compute using light—all of that is going to come out of this same exact process.
We’ve said before, this is how we’re going to train robots in the future. They’re not going to be trained in the real world. They’re going to be trained in these high-fidelity simulation worlds.
That’s the present. I would argue that so many different robotic, embodied VLA, or now world-model companies are just being trained from watching YouTube. If Fei-Fei Li and her company’s approach gains traction, maybe the right primitive is no longer patches of images, which is what many of the models right now are doing.
When you train a diffusion model, you’re typically taking each frame of the video and breaking it up, usually into something like 16×16-pixel patches, and then treating those as tokens. Maybe the right primitive ends up being Gaussian splats.
It's definitely not going to be 16 by 16. I mean, the fact that that worked at all, I think, shocked everybody. Let's just take a language transformer, take images, cut them up, pretend each chunk is a word, and blast it through to see if it works. It just worked incredibly well. Maybe superintelligence is a general-purpose technology that relies on compressing information. Yeah, maybe.
Where does this go for the average consumer, guys? I mean, obviously, this is a world of extraordinary video games. Are people going to be just living their lives in these virtual worlds? Is this going to become how we consume ourselves, how we consume entertainment in the future?
Yeah, but also how we design our next day. What do you want to do tomorrow? I don't know—let's walk through what it would be like to do this or to do that, to play tennis or to go to the Caribbean. You can just experience it in advance and use that as your planning tool.
So much of our lives are random wandering and not really well planned out. I think human happiness is going to go through the roof once you're interacting with your AI. Right now, the AI will guide you through a day plan, but it's in text. It's going to be so much better when it guides you through a day plan visually and you're just stepping into it.
I'm beyond excited about all this. We have 3 vacation locations. Let's go and explore all 3 as a family, watch it, and then see which one we want to go to.
No, no. I'm excited if you go to a place—like if you go to a resort and you've been there before, you have so much more fun because you know where to go. You're not wandering or losing all of your time finding things. Now you can actually pre-experience things and know exactly what you'll enjoy, where they are, and how to get to them.
You're not waiting in lines. You're not signing up for things and then realizing it wasn't what you wanted. It's going to be so good. And ask your AI, based on what you know about me and my preferences, show me what I'm going to go do.
Emad, big announcement for you today.
Yeah. Over the last few years, we've been working on The Last Economy—what does economics look like?—and then released The Commonwealth, looking at personhood, law, and political economy. A lot of people ask, how do we share in the gains of artificial intelligence and make sure they're distributed? So we went back to the drawing board and thought about what type of institution and future we want to see. We came up with this idea of the champion.
We think that AI should be like a utility and should be owned by the people. You need the children to own it. You need the locals to own it. For every single jurisdiction, what we see is that the cost of intelligence will drop to zero and the value will go to the last mile. Salim would kind of integrate AI into the enterprise. The humanoids—who owns those humanoids? Because those will drive the economy.
Elon at the G20 just now said that the average humanoid will have 5 times the output of a person, and there will be 1 billion of them. That will be the economy. So we're like, let's set these up and borrow from the example of TSMC.
How TSMC was set up was that they started at a valuation of 10 Taiwanese dollars, and the locals put in 75% of the money and Philips put in 25%. The CEO did not have any shares. The team did not have any shares. They only got shares from the profits. He had to buy his shares, now worth $10 billion.
Is that for real?
It's for real. It listed at 6 billion Taiwanese dollars, which was the cash on the balance sheet. So we were like, let's do that for the intelligence company of California, the intelligence company of the UK—$1 pre-money. All the locals can invest at that valuation, from institutions to high net worths to retail. $75 million per state.
You can get the MIT endowment, Dave, to invest and give them back all the compute that they have—the equivalent dollar amount that they invest in compute. There are all sorts of interesting things you can do at $1.
Then you can bring in the internationals and the strategics at market rate, which will be 10 times that because you've got everyone on board, and you give 10% of the equity in perpetuity to every child under 20. So every year, you issue half a percent of the equity to every kid.
It trains up an FDE workforce to transform every institution. It owns the robots and deploys them, which is 80% of the value of robotics downstream, just like the auto manufacturers. And it gives an agent to every citizen, as well as AI for the government—the Sage project that Peter and I and others have been working on—AI for the judicial system, AI for education, and healthcare.
That becomes really super interesting because then it becomes a play on the indexed GDP of the state, owned by the people of the state, with the smartest people in the state involved. Again, the trick here is $1 pre-money. Get everyone in. You want it to be a success? It's up to you. And so that's this new institution.
How many champions are there? How fine do you slice it up? I heard you say champion for California and the UK—cities and countries, or states?
In the United States, we're doing 1 per state because a lot of data has to stay within state boundaries. Otherwise, it's pretty much 1 per country. So, again, it acts like British Gas here, for example. It acts like the telco and more, and each one of them covers a certain number of citizens because, again, it gives equity to every child born.
The locals can invest, etc.
So, Emad, right now, this is just an idea. If folks want to learn more about the idea, it's not an investment offering, and this is not investment advice. But to learn more about the idea, where do they go?
They go to ii.inc. As you said, it's just an idea at the moment, and we'd love people's input, but ultimately it'll be about the citizens. So, let's see if we can build this structure.
Your first phase, I've had some chats with you about it, is super exciting: to get all these local champions lined up, right, and then, little by little, cascade to the next level.
Yeah. It's all about whether a state wants this. Then it's all about the people of that state. It's not like a foreign company coming in. This is a model just like you have your UBI, just like you've got your shares in Frontier Labs, et cetera.
And so, we'll see how it goes. This is my proposal for trying to distribute it to everyone. Awesome. I've got to ask you a question. Your bullets say, “The cost of intelligence is going to zero.” Got it. “The economy will change forever.” Got it. “The value of human cognition will go negative.”
Yeah.
You're saying it's going to be like a cost on society to be thinking and consuming food while you're—
Yeah, the stupidest person on the team economically, right? You're competing—
Your ideas are—
No matter how good your ideas are, they add negative value. It's like adding a human driver to an autonomous highway. Of course, this is the topic of the book that I had last year. But that's why you need to have a share in the means of production, which will be the robots and the forward-deployed engineers and people like that, right? You need to make sure that's equitized from day 1. Otherwise, if you have a thought and you don't tell anybody, it's just zero. Then, at least, it's not negative. Most people's thoughts aren't economically valuable thoughts; it's economically valuable labor. I'm not the smartest person on my agent team anymore, man. I don't know about you, but I'm getting eclipsed very quickly.
Yeah. No, it is very humbling for humanity, actually.
It will be. But just like we've seen before, when Stockfish started beating everybody in chess, people still play chess, and people are still going to have ideas. People are still going to value individual human ideas. It's like, “Did you come up with that, or was that your latest model?”
All right. I'm going to move us to our next story. Tesla held its Cybercab launch event in Austin, Texas. Cybercabs everywhere—a river of golden EVs flooding the streets. Unless you've been hiding under a rock, you know that Cybercabs are Elon's electric, autonomous 2-seaters with a, quote, “three-comma” scissor door. No steering wheel, no pedals, designed to run entirely on Tesla's self-driving software. Did you guys get the three-comma comment? Anybody?
No.
It's from Silicon Valley, right? The billionaire.
Okay.
Remember, he showed him a Maserati and said, “Ah, that's not a three-comma car.”
It doesn't have the doors that go like this.
Right.
I remember that.
A great episode, actually.
This is a vehicle designed from the ground up for full autonomy. Let's take a look at this video. I love this video showing the flood of Cybercabs on the streets in Austin. Check this out.
Cybercabs everywhere. A golden river. It's crazy. You know, I didn't really internalize the lack of rearview mirrors. We obviously saw the lack of a steering wheel.
Yeah, but actually, the number of components they've taken out of the car is driving down the cost. There's no driver, and there's a huge amount of cost in the car that's gone. So there's no way anyone's going to match the price point of this thing. I mean, that's the point compared to Zoox and Waymo or anybody else. Elon wants to sell these at $30,000 each so you can buy them, and I think a great entrepreneurial journey is to buy 10 of them and put them on the streets in your local town and have them earn revenue for you. One early rider in Austin is reporting that Cybercabs are running about 50% cheaper than Uber for comparable trips, with rides being described as smoother and cleaner than Uber, with automatic syncing of audio and seat settings in your passenger profile. It's pretty extraordinary. Remember, Nevada gave Tesla permission for 5,000 of these vehicles on the road in Las Vegas in the next 12 months. It's just going to crowd out the competition. Salim, I'm curious what you think about this. Is this sort of like the ChatGPT moment for Tesla?
I think it could be. This is classic ExO, right? Autonomy, you decentralize, you have interfaces, you're leveraging assets. If Elon can get people to buy the cabs and create millions of micro-franchises out of them, the loyalty that will come from that is going to be unbelievable. Of course, they're all going to have self-driving capability, so the whole vertically integrated stack brings itself to bear. I think the most exciting part is the fact that the cost of transportation drops by another order of magnitude, from a couple of dollars a mile today down to 20 cents a mile. That is really interesting, and I need this to happen because I've made that comment that Milan will never go to university, or he'll never need a driver's license. I'm going to battle it out with him that he will not get one. So this needs to happen. He's got 2 years to get this done.
And you would know better than anyone, but when you get a hotel room in New York and you look down, it's like 80% yellow cabs down there. If you look at the traffic in New York, a huge fraction of it comes from people blocking an intersection. They get the ticket, but they're still stuck in it, and the whole thing grinds to a halt.
Yes. I suspect there will be enormous chunks of entire cities that say, “You know what? No more human drivers.” It's so much more efficient. Not only is it much cheaper, but it's much more efficient to get around. Also, a vast majority of the safety issue in cities is pedestrians getting hit by cars on sidewalks, and this is going to basically eliminate that risk. I suspect we're right on the tipping point where it's like, “Nope, no human drivers in this entire inner-city area. End of story.”
And we'd better get artificial organs quickly, because the drop in organ donors is going to vaporize.
Well, all these things tend to happen at the same time anyway. I think for many people, these robotaxis will be the gateway drug to generally autonomous robotic systems on the streets. I think we'll look back, with the benefit of hindsight, and say, “Well, of course, it's natural that a municipality gets lots of robotaxis on its streets before it gets humanoid robots on the sidewalks doing economically valuable activities.”
My favorite little anecdote, analysis point, or data point to illustrate this is that every medium-sized town in the country has a transit system where the buses run empty 95% of the time, and they're totally full at rush hour. They run completely empty the rest of the time, and the whole thing is a massive loss-leading exercise. It costs a bomb.
Already, a few years ago, small towns in the U.S. were saying, “Get rid of the transit system. We'll just pay for everybody to take Uber,” because that makes it hyper-efficient and much lower-cost overall. But this takes it to another level. It takes it to 11, to go down the full analogy that we're going down.
And this now allows mobility to every single person—every blind person, every disabled person, every student, every drunk person. All sorts of capabilities become available and affordable in a way that's very powerful, incredibly exciting.
Yeah, the poorest people in the world are being chauffeured around by AIs, for sure. Two interesting points here. The first is that it's a two-seater, right? And, of course, the average load for an Uber is 1.2 people, right? How many times do we take an Uber all by ourselves? So, if you need 6 people in a car, just take 3 of the Cybercabs.
Yeah. Not just that, but think about a normal cab: it has 4 airbags. You get airbags for every seat. Elon, in his infinite brilliance, is like, “We now only need 2 airbags.” We cut the cost of airbags alone in half with this design. In the rare instance where you need 3 or 4 people, get 2 of them. End of story.
Well, they follow each other right behind each other. It's perfectly great.
I suspect there's more of a backstory there, though, with the two-seater. Remember, Tesla was originally teasing that it was going to launch a Model 2, which was going to be their highly coveted $25,000 car. That never happened. My best read of the situation is that what was originally planned as the Tesla Model 2 became this robotaxi, because at some point, once the value of the car—or the sales value, I should say—crosses below some threshold, it makes more sense to monetize it via autonomous ridesharing or robotaxi services than it does to actually sell the car.
This will turn transportation into an API.
Yep. Interestingly enough, we're starting to see conversations where the other companies that do have steering wheels, pedals, and LiDARs are saying, “I'm not sure that the Cybercab is safe enough. It only has cameras. It doesn't have the other modalities.” So, expect to see a battle in the courts about what city allows this technology in.
Yeah, Boston, just to put an exclamation point on that.
Mayor Wu, come on the podcast and we'll have a conversation about it.
Oh, yeah. Do that. That'd be awesome.
I'm going to continue this conversation. Let's shift it now to the global stage. The competition is going three-sided. The Financial Times reported that Uber is partnering with traditional taxi fleets to complement and compete against Waymo's expanding robotaxi service. So, a strategic alliance between ride-hailing and traditional taxis to counter autonomous vehicles that don't need humans. Think about what that means. Uber, the company that disrupted taxis, is now partnering with taxis to fight the companies that are disrupting them.
Meanwhile, The Verge reported that Uber and U.K.-based robotaxi company Wayve officially launched services in London. Emad, have you seen Wayve yet?
Yeah, I've seen it. It's just starting to roll out, so I'm looking forward to getting my first ride on that.
Yeah. CNBC reported that Waymo and Zoox announced simultaneous expansion into new cities, along with Tesla in Nevada, California, and Texas. My prediction here is that we're going to see at least 5 different autonomous electric robotaxi companies fighting it out in major cities inside the next year, driving the cost as low as possible.
The cost of personalized transport is basically dropping to the cost of electricity. It's demonetized mobility. It's the entire abundance thesis. The cost is dropping to the cost of charging a battery. The ultimate winner is the public, unless you live in Boston.
Well, no. Cambridge might have a chance. Actually, that'll put a lot of pressure on Boston. That'd be awesome if Cambridge—
—beat Boston by trying to take a taxi and it stops at the Harvard Bridge.
That's right. The bridges are no-fly, no-drive zones for autonomy, apparently, in the near future.
Actually, what's funny is you can get around Harvard, but you wouldn't be able to go to Harvard Business School.
That's really cool.
Yeah. What's going to happen, though, when you get your Tesla Optimus robot and it has the drive program, so it can drive for you in whatever car you have without a retrofit?
Ah, there you go.
It'll get banned in Boston, too. I'm sure in Boston.
I want to hit the economics here. The new versions of the Waymo vehicles are pricing out at over $100,000. The Cybercab is coming in at or below $30,000. What's most interesting in my mind is that the winner here is going to be whoever can mass-manufacture these the fastest. Elon's going to win that game. I mean, he's the guy who builds the machines that build the machines.
In the U.S., maybe, but China is dumping some, some might say, into Europe. So, it's not a U.S.-only game. That's the issue.
You're going to end up with all of them for the foreseeable future. I think, Peter, the point you made is the most important one: the end user wins.
Yeah. Dave, do you remember when we were at the Gigafactory, being toured around, and we saw outside the giant mounds of aluminum scrap metal?
Yeah.
And then the smelter.
And then the Model Y press. It was just this beautiful orchestration. It's incredible, end to end.
And you can actually walk with the car from the minute it's born. It's a long walk, almost a mile, but a car comes out the other end. You're literally watching every part get put on as you walk with it. It's absolutely wild.
Yeah, it's beautiful.
But it's amazing to me how few parts there are in the Cybercab. When you open the hood of your gas-guzzler car and look at all the stuff that's in there, and then you look inside the equivalent Cybercab, it really feels like there's maybe one-tenth as many things.
A typical car has 2,000 moving parts in the drivetrain, and a Tesla has 17.
An ICE—an internal-combustion-engine car.
2,000 moving parts versus 17. Unbelievable. So, in terms of maintenance, design, and all of that stuff, this is why the car dealers are all freaked out.
The last time I was in New York, I was in a yellow cab. I'm like, “Why is this thing so disgusting? It smells like a public urinal. Why are they all like this?” But then you think about what happens to that cab. The medallion is incredibly valuable. It gets handed from one driver to the next. It never stops moving.
At night, in the middle of the night, somebody drives it home. Then they get up first thing in the morning and start driving it again. But it has to take somebody home. The Cybercab goes to get cleaned in the middle of the night, when there's nothing to do. It goes to a place—it could be anywhere. It could be in Queens, where it gets itself cleaned—and comes back pristine.
It's just night and day when you get into one of these versus a cab. Now, they're brand new, so maybe that's part of it, too, but the experience is night and day better.
And they look beautiful.
So, what happens when they break down? I guess a Cybertruck comes and tows it away.
Yeah.
DoorDashers come over to help, is the recent story, I guess.
The door is wedged open and can't be closed.
Maybe they just squash it into a metal cube right there. Crush it—
—and recycle it.
Take it back to the smelter.
That's mean.
We'll pull out the GPUs first.
Yeah.
All right, I'm going to move us on to the space arena. Two fun stories on the space frontier today.
In our first story, NASA chose Blue Origin to build the telecommunications relay network on Mars—the communications infrastructure that will connect future Mars missions back to Earth. This is the infrastructure layer for the Mars economy. Whoever owns the telecom network controls the bandwidth on Mars. There could be a company like AT&T on Mars. It could be Blue Origin.
You might wonder, why did NASA select Blue Origin, not SpaceX, since SpaceX owns and operates the world's—or space's—largest space-based, laser-linked communications network? Ultimately, this is the government keeping 2 competitors, 2 suppliers, in business, giving them each a slice of the pie.
Regardless of Blue Origin having won the NASA contract, I guarantee you Elon will still build Starlink around Mars. Ultimately, what we're seeing here is the birth of the interplanetary internet. Alex, are you excited about this one?
Look, Mars had it coming. The moon had it coming. We're going to build the Dyson swarm. Someone was going to get awarded the Starlink for Mars. It's interesting that it went to Jeff Bezos's company and not Elon's, but I'm fully expecting that we're going to have a very competitive interplanetary internet. Frankly, I'm glad that there are vendors competing for Mars communications other than SpaceX.
Yeah, but it's going to be a packet-switched network throughout the inner solar system: Mars, asteroids—
High latency, but that's all right.
Yeah. Until we get faster-than-light communications, who knows? Physics.
Time will tell.
Time will tell. Are you working on that, Alex?
Can't say.
Well, the physics that we have right now suggests that faster-than-light travel is not possible.
Don't bum me out here.
Okay. Sorry to break the bad news. The textbook physics right now says that superluminal travel—
Then we have the wrong physics.
We'll find out. Just hope.
Yeah.
Let's move us to another fun story in space. Our second story—and it's a big one. NASA's Nancy Grace Roman Space Telescope has launched on a Falcon Heavy, carrying a field of view more than 100 times greater than Hubble and the ability to scan the sky more than 1,000 times faster. The Roman telescope is designed to discover tens of thousands of new worlds and map the distribution of dark matter across the universe.
Before we discuss it, let's play a video by our amazing NASA administrator, Jared Isaacman, a friend of the pod. I love this guy. He's such a good communicator. Let's check it out.
The telescope is very healthy right now. It's making its 1-million-mile journey to Lagrange Point 2. It's going to look for what we think will be up to 100,000 additional exoplanets in other star systems. It's going to help us understand dark energy and dark matter. As you saw during the press conference, President Trump called in. This is a really exciting time in America's space program right now.
What's the difference between this telescope and the Hubble?
It's hundreds of times more powerful. The field of view is over 100 times greater than Hubble. Its scan rate is over 1,000 times greater. This is going to be a household name like Hubble and James Webb. This is America's next great exploration asset.
And you say it's going to help you find planets hiding behind planets—
Planets hiding behind other stars—distant stars. The light can blind it out. This has a special JPL coronagraph that helps us find these hidden worlds. We're going to find tens of thousands of additional worlds.
Tens of thousands of worlds. Amazing.
Yeah. Jared is incredible, isn't he? Compare that to that interview of Sam saying, “Why is this new model different? Why is Astra different for people?” Compare his answer to what Jared just did to answer this question. It's like night and day. He is a great communicator.
Interestingly, this is going to finally start to give us statistics over the number of habitable—at least Earth-recognizable, habitable—worlds. It's being pointed toward the center of our galaxy, and we'll be able to do large sweeps of the sky, looking in part—it has other missions as well—for microlensing events, for planets crossing in front of their respective stars and causing, via their gravity, light from those stars to be very weakly increased briefly due to these microlensing events.
The downside—
You know what's so cool about that is, on the last podcast, Peter asked us—you missed it, Salim, and Emad, you weren't here—what's your answer to the Fermi paradox? Why are we not seeing other civilizations? One of the theories is that there are many, many civilizations talking to each other near the center of the galaxy, where you can get from star to star in a year or two, as opposed to where we are. We're way out in the wings, where it's so far away.
We're in the unfashionable outer suburbs of the galaxy. That's right.
So who knows? This could—
I'm actually happy we're here. The galactic center is a really dangerous place to be. You don't like supernovae, Peter?
I don't like the radiation flux they deliver. No.
I'm reminded of the opening scene of The Hitchhiker's Guide to the Galaxy, where they're bulldozing Earth to make a hyperspace highway—
Bypass. Yeah, yeah.
I love this. I'm curious about our viewers, and if you can, let us know in the notes: would you like us to have a conversation on the current UAP disclosure discussions out there? If we can bring in some of the leaders in UAPs and the whole disclosure scenario—the White House just released its disclosure plan—I don't know if I'm going to mention that, Alex, but I'm curious if folks want us to have a conversation on that topic on the pod. Alex, would you mention the recent White House announcement?
There was some reporting out there from Avi's UAP Science Advisory Council. One of the members of the council mentioned in a recent briefing that this council, which, as I understand it, was stood up by the White House, was informed that the White House had prepared a disclosure plan for informing the general public of the existence of non-human intelligence. If that reporting is accurate, that's pretty interesting.
Yeah. Emad, where do you come out on the whole UAP side of the equation? I'm curious.
I actually have no position on it. I've never thought about it properly.
Okay.
I'd like to see strong claims require strong evidence.
I stand with that.
A lot of people would say there is strong evidence. It's just hidden. But we shall see.
No, I'm with that. I would say evidence—it would be highly desirable for there to be a preponderance of evidence that everyone can go and see, touch, and experiment on. I think that's probably the gold standard in an ideal outcome, if there were a White House disclosure event. If the president goes into what's left of the Rose Garden and gives a speech and says, “We're not alone,” then ideally part 2, paragraph 2 of the speech would be to hold up or otherwise present some artifacts that would be subject to extreme scientific scrutiny to support the claim.
That would be a fun conversation to have. All right, our third—our final—subject for today. Three major health and longevity stories came out this week.
The first is that OpenAI expanded ChatGPT Health features to connect directly to patient records and health-care databases. The integration brings Epic electronic health-record data into ChatGPT for health care, and Epic, as you guys probably know, is the largest collection of health records. 325 million patients are inside Epic—roughly the entire U.S. population. Clinicians can now pull appointment notes, lab results, and medications, and ask questions across a patient's entire record. Consumers can now connect their Apple Health, One Medical, and Function Health data, so ChatGPT can help you understand your test results, prepare for your doctor appointments, and get personalized diet and workout advice.
That's the first story. The second story, interestingly enough—we talked about it last week. We mentioned that the FDA had approved a drug called Duraxinarissib. It rolls off the tongue and onto the floor. It's the first targeted RAS inhibitor for metastatic pancreatic adenocarcinoma, attacking the RAS family of proteins that drive tumor growth in most patients with the disease.
This week, NBC News reported that the same drug is showing promise to treat lung cancer. As mentioned last week, the RAS mutation family drives roughly 30% of all human cancers and was long considered undruggable, the subject of decades of failed attempts. But now, during the singularity, the end of cancer is within reach. Alex?
Yeah. I think we're starting to see—and, interestingly, I'm not sure that AI was actually essential for this particular drug—but I think there's so much progress being made in cancer therapies now, including on the immunotherapy side, that we're starting to see the emergence of, again, caveat, caveat, caveat, universal cancer treatments and universal cancer vaccines in some cases.
After years and years of treating cancer as thousands of different diseases, we're finally starting to get to the point where we can move upstream and start to treat, if not root causes, at least identify treatments that—whether it's proteomic pathways on the one hand or immunologic pathways on the other—start to have treatments and vaccines that can treat multiple classes of cancer. And I think that's—
Cancer's cooked. That cancer should have been cooked long ago. Was it Nixon who announced the war on cancer? It took forever to get to this point—more than half a century. It's an interesting counterfactual thought experiment: is there anything, knowing what we know now, that we could have done 50 or 100 years ago to radically accelerate the onset of broad-spectrum cancer treatments? What do you think, Peter?
I think the data—I mean, we've known, for example, about the RAS mutation causing unconstrained growth for some time. It's just getting the molecules and getting the drugs that can attack it properly. I think the tools we have right now are finally giving us that reach. Then, being able to understand fundamentally what happens and how to block it—we talked about cell simulators—is what's coming.
Emad, this is an area of personal passion for you as well. You've been deep into medical AI.
Yeah. I said this wasn't AI, on the thing, but the range of treatments now coming out on the cancer side gives a lot of hope for what's going to come. I think the mRNA one may be more general that we saw recently.
I think the first part of integrating into the Epic health records and actually applying AI across the board—we should have a sprint so that within a year or 2, maximum, every single health decision is double-checked by an AI. We should have—
More than that. I think it's going to become malpractice—
To diagnose a patient without AI in the loop.
Right. We already know AI is a far better physician—a diagnostician—than a human is.
So I think you should have a series of approved edge and cloud models, and every time you make a diagnosis, the AI has to have had one check. That will save so many lives. It will detect so many cancers. Then it's about how we increase the level and volume of information, because even now, the type of data we have around cancer and other conditions that we absorb is tiny compared to the amount that we could have with AI transforming it.
So I think, yeah, let's cure disease and get rid of it. No one should have to die of cancer, and it's something we should be really directed at.
On the first story—
Go on, sorry, please. It's just very strange that, as we have the Genesis programs and others, there isn't a straightforward “We now have the capability to potentially cure this stuff. Let's direct $10 billion toward it.” It's tractable.
I think the story on Epic is interesting. Having been involved in that business through Fountain Life, Epic is the majority electronic health record in the US, and it's been a bear to navigate for physicians. Patients have never had access, so putting an AI layer on top of that is awesome. Yeah. Let me move it to our last story here in the area of health. It's related to what may be called the first longevity class of drugs: the GLP-1s.
A new paper published just 2 days ago, on September 2, in Nature, shows that semaglutide—the active ingredient in Ozempic and Wegovy—recapitulates many of the benefits of caloric restriction. Get this: it's in female mice, just to be clear, and extends the lifespan of mice by almost 100 days. In humans, that's the equivalent of 8 to 10 years.
The study found that GLP-1R activation initiated late in life in these mice accentuates age-associated decline and modulates conserved genetic regions for aging, functioning as a caloric-restriction mimetic. Also this week, it was reported that GLP-1 drugs are being linked to fewer serious infections, including tuberculosis. You guys already know that tuberculosis is the deadliest killer on Earth, killing over 1.25 million people per year.
So the same drug that's treating diabetes, obesity, liver disease, kidney disease, cardiovascular disease, and addiction may also be extending lifespan and reducing infectious disease. One molecule, 6 diseases, and still counting.
Again, Alex, you made the point that this might be the beginning of longevity escape velocity. To the extent that, with the benefit of hindsight, we look back in a few years and say, “You idiots. Of course you were on the verge. You were seeing the sparks of longevity escape velocity. You had the GLP-1s,” I don't think it should be that surprising.
What's more mystifying to me is, from an evolutionary perspective, if the GLP-1 receptor agonist class of molecules is capable of doing everything from treating infections to extending life expectancy, modulating diabetes, reducing addiction, and reducing compulsive behaviors, why on Earth did we not evolve with either this ability to modulate our own semaglutide-class molecules in our system, or maybe a slightly more cynical angle?
If it turns out that the reason GLP-1s are so effective at so many diseases is that these diseases somehow are diseases of modern lifestyles—that it's treating all of the diseases of modernity, and that's why we never evolved the solution—that would be pretty ironic. Addiction to compulsive gambling or alcohol, or overeating sugar: these are all relatively modern diseases. That's why this is basically a treatment for modernity.
So I just want to make a point to the listening audience. First of all, talk to your physician about this. This is not medical advice. Yes.
But I use a GLP-1 drug. I don't know if any of you do right now. I use it not for weight loss. I've been at my fighting weight for a while. My body-fat percentage is substantially lower because I work out and I'm very careful about what I eat. I use it as a longevity drug because of all the benefits we just heard about.
I've got 2.
I know you're using one.
Yeah, I'm using one, and I just got my blood test done. My liver enzymes are 50% better, which is pretty amazing. Which means I can drink more. No, just kidding. That's not the point.
I know that's not the point.
You're supposed to not want to drink it.
I know. So there's that, but I think, Alex, to your point earlier, remember that evolution has birthed us for death. We've had short life cycles so that the cycle time of evolution can work more quickly. We're breaking through that now and living through that, so I think it may have been engineered, or evolutionarily engineered, that we died at different levels.
This is a long history. We used to die of heart disease or bacterial illnesses, and then we figured that out. Then we died of heart disease. We got some sense of that. Now we're dying of cancer.
Then explain tuberculosis. If you're in a tuberculosis-rich environment, surely evolution would favor any molecule that could be amplified, which would help young and reproductive entities survive tuberculosis infection.
Well, I go back to the fact that there are lots of pathogens going after us all the time, right? Billions of them, because they're all trying to survive in their own way, including cancer. The capability of the human body to navigate this has never been needed until we started pushing the boundaries to this level. Now we do.
Salim, my thesis—and I've spoken about this pretty widely—is that the reason for the life cycle is not more rapid evolution. For most of human existence, Homo sapiens came on the scene roughly 200,000 years ago. Food was very scarce, and if our primary mission is to perpetuate our species, the last thing you want to do is steal food from your grandchildren's mouths.
So the best thing you could do is die: reproduce and die, basically. We see the human body—and people should know this—you're in prime condition until your late 20s. Back 200,000 years ago, you'd go into puberty at age 12 or 13. You'd be pregnant immediately, with no birth control. By the time you were 26, 27, or 28, you were a grandparent, and then you would die so you didn't steal food from your grandchildren's mouths.
That's my joke about marriage: we invented marriage to keep the parents together until the kids were self-sufficient. The average lifespan was 25 years for most of human history. We invented marriage about 6,000 years ago, and it's definitely true then: marriage is not designed for 50- or 60-year lifespans.
No, no. Watching this, one of my relatives calls it state-sanctioned torture—marriage—because we have the job of now evolving the institution to deal with the conditions today compared to when we first invented it.
I use that example because it applies to all our institutions: democracies, educational systems, legal systems, and healthcare systems. This, I think, is the biggest work we have to do. As we blow past all these limitations, we have to reinvent all of the institutions that are the scaffolding that keep humanity safe and civilized.
Hopefully, we have a benevolent AI to help us do all that.
We'll need it. And if not, benevolent AI. At least we'll have GLP-1s.
All right, time for some conversations with the mates. We've got some questions here. Emad, I'm going to give you first crack.
All right. If money becomes obsolete, how will desirable land be allocated? Is everyone just locked into their beach house forever? This is from Nick 52547.
This is an interesting one. Elon and others have said that money might become obsolete. I've said that you can't compete with robots. I think this is a question of land rights and more, because typically where you see reallocation is in upheavals.
The way we're going, without the right structures, you'll probably have a debt jubilee. You'll probably have chaos and land redistribution. But if we can actually navigate through it, then land rights and property rights should be enforced, and so you can keep your beach house.
The number of desirable locations will go up dramatically because you will have air taxis, self-driving cars, solar panels, and self-driving construction workers.
Yeah, exactly. Self-driving construction workers.
Yes.
Nice.
Do you think AI will evolve past this risk-averse bottleneck that we're currently in? And that's from AdsSusie1073.
Yes, but hopefully we move toward calibrated risk rather than recklessness. Right now, we treat uncertainty as a reason to refuse. That's a technical problem, but there are also all the legal and policy issues around this.
When you're raising kids, one of the things they teach you is: Are you taking a responsible risk? We need to apply that same kind of paradigm to these models. Could you create something like a risk budget? What can an AI decide? How much can it spend on that? What systems can it access? What triggers escalation?
We need to get very sophisticated around this. Models need to become better at distinguishing dangerous intent from real expert usage, because you can't just say, “Remove the guardrails.” It's got to be dynamic, where you deal with proportionality based on user identity, the context, and the reversibility of that.
We'll get AI maturity when we can say, “Here's the risk, here's the confidence I have in this, and here's the reversible next step.” Rather than just saying no, I think we need to get a lot more nuance around this. Unfortunately, in today's world, nuance has no part to play in the sound-bite politics that we have out there.
All right, Dave, over to you.
I’ll take number 1. Can you see a way to turn libraries into centers of AI development and physical AI training for everyone, from Train with John Koalo?
Yes, but there’s a bigger issue, which is that there’s a huge amount of white-collar office space, and white-collar work is turning to AI. At the same time, we have a declining population, especially a declining working-age population. So there’s a bigger issue: What about all this other space? What are we going to do with it all?
It’s very similar to what happened with the shopping malls after online shopping became huge and then COVID hit. A lot of people were thinking, “There’s got to be something really great we can use all this shopping mall space for.” There were some ideas, but for the most part, they didn’t work out. It all got replaced.
So I think there are lots of ideas for what you can do with library space. You don’t really need the library books anymore, obviously. This is another reason why you should want data centers in your community: It’s one of the few things that will reliably grow tax revenue and create job opportunities in a community. So there are ideas, but I’m not super optimistic that all the space will be well utilized going forward.
There are programs in inner cities right now that are running AI tutoring in libraries. That does exist.
I get the fun one. Question number 2: Whatever happened to synthetic diamond chips for computers? This is from Asterene.
Here’s the problem with diamond: Pure diamond is an insulator. It has a band gap of approximately 5.5 electron volts, so it really doesn’t want to be a good computer. To make it into a good computer, at least a good computer of a recognizable CMOS type, you have to dope it. That’s hard.
If I had to make the case for diamond-based CMOS computers, it would be for ultra-high-temperature environments, where such a wide band gap could be advantageous, or maybe very high-voltage environments, where, again, a large band gap could be advantageous. It’s not an enormous market. I could be wrong, but it doesn’t seem like an enormous market.
Where I’m much more bullish on diamonds is sensing. You can put what are called nitrogen-vacancy centers, or NV centers, into diamonds. Basically, you put an extra, unwanted nitrogen atom into a diamond lattice, and suddenly you get an exquisitely sensitive magnetic-field—or just field-in-general—sensor because of an extra vacancy that is introduced as a result of the diamond.
Diamonds for quantum sensing: super interesting. Potentially, it’s not investment advice, but technologically, it’s very attractive. One could imagine, at some point in the future, diamond chips for sensing even getting us to sci-fi technology, like wearable MRI sensors. Diamond chips for computing? Eh, probably not.
Why did I know you’d have an answer for that one?
It’s interesting.
All right, Salim, first choice is yours, pal.
I will go with number 7. When do we get a forecast for when data-center energy sources transition from gas-turbine farms to nuclear and SMRs?
Short, fairly easy answer here: Gas dominates the current buildout. SMRs are going to be in the next 3 to 4 years. The reason people are getting so excited is that nuclear is an engineering problem, not an invention problem, right? Fusion is still in the invention-problem category.
You’ve got this impedance mismatch of data centers taking 2 or 3 years to build, while nuclear is taking a bit longer. We will get there, I think, faster than people think with SMRs, but I think it would still be a while. Take 2 to 3 years for the initial wave of SMRs and then 5 to 7 years for the big buildout.
I think what’s going to end up happening is that AI is going to do for nuclear what smartphones did for batteries. It just created such huge demand that the massive innovation there caused a huge acceleration in the innovation curve.
Nice. We had a fantastic pod with Ramez Naam on energy. If you guys haven’t seen it, please check it out. He talks about pretty much that time frame.
That’s where I get all my good information about energy.
Yeah. Alex, let’s go to you.
I think I have to answer question number 8, which seems to be directed toward me. It asks, “How far away from Alpha Centauri do you actually have to aim to reach it when it arrives?” This is from Typical Dad Pi.
To 3 significant figures, to the extent I understand this question, let me give you the semi-official Fermi Explorer line. On the last pod, we had Matt Pines and Philip Johnston announcing, for those who didn’t watch, humanity’s first mission to Alpha Centauri, and I’m involved. I’ve been involved. What can I say? I’m in a lot of rooms.
Under the official mission parameters for the Fermi Explorer mission, the goal is to get at least 99% of the way to Alpha Centauri. Alpha Centauri is approximately 4 light-years away, so you could do the math, and that turns out to be 4/100ths of a light-year away. Call it precision getting there. That’s from Earth.
If I put my sci-fi hat on and extrapolate a bit, I suspect that when the mission comes to full fruition and is launched, it will have active guidance on board. Maybe I should add parenthetically that the Fermi Explorer mission is intended to launch by 2029. I suspect this is not an official position for the Fermi Explorer mission, but if it has active guidance on board, this 99% of the way to Alpha Centauri can turn into 100%, in the sense that it actually hits the Alpha Centauri system, not just ends up 4/100ths of a light-year away.
That said, Fermi Explorer is intended to reach Alpha Centauri 80,000 years from now, which is perhaps inconvenient from the perspective of mission verification if you actually want to get it.
GLP-1 drugs.
We’re going to need a lot of GLP-1s for this.
I think the whole point of the mission is that we’re supposed to beat the Fermi Explorer and get there sooner. This is just the first one to launch.
Amazing. Dave, over to you, pal.
I’ll take number 6. Are data centers using closed-loop geothermal cooling also noisy, or are they quieter? This is from Asterene.
They should be dead quiet, just like geothermal heating is dead quiet. Also, nuclear reactors that are near the ocean use geothermal cooling and ocean cooling, and they’re dead quiet. So it should be very, very quiet.
Regular data centers that are using liquid cooling are only noisy because the water gets cooled outside with these really poorly designed fans. There’s no reason for that. If your regulatory body says, “Yeah, you can build a data center here, but it has to be quiet,” guaranteed, they will build the data center quiet. There’s no reason those external fans need to make any noise.
Instead of making them illegal, cities need to be saying, “These are our requirements: Drop our energy costs, make them quiet, invest in our infrastructure.” And they will. They will.
Emad, it looks like number 5 is for you.
That’s an interesting one. They talk about a sunshade there. What scale of satellites, in terms of—
We talked about—
That came up on the last pod. You missed it. Yeah.
Yeah. And what scale of satellite, in terms of square meters, is required to actually impact global temperatures?
I would say if you put something at the Sun–Earth Lagrange point, a couple of million kilometers out, you need about a couple of million square kilometers. So, about the size of India would knock a degree Celsius off. That’s bigger than anything we’ve ever managed, but it’s worth a try. Why not?
Maybe convert the Moon.
Mercury, please.
Mercury.
Or a series.
Yeah. And I still love the sunshade—putting a thermometer there to be able to titrate the solar flux on the planet.
I think it’s a trick question. For what it’s worth, I think with a sufficiently good AI planetary-scale model, we could make any sunshade or other satellite intervention de minimis in size. It’s just a matter of appropriately perturbing the Earth’s atmospheric system with enough AI.