[BidClub_]
Moonshots · · 118 min

Mira Murati's 975B Open Model, Ramin Hasani on Post-Transformer AI, and Demis' AI FINRA | EP #271

Peter DiamandisSalim IsmailDave BlundinAlexander Wissner-GrossRamin Hasani

YouTube
TL;DR
  • A FINRA-style frontier-AI watchdog could provide safety cover, but the panel saw a serious risk of regulatory capture by the labs designing the rules. Demis Hassabis wants an industry-funded body testing frontier models before release, reportedly operational by year-end; Ramin Hasani instead argued for capability- and vertical-specific governance that iterates like a Stackelberg game. Alexander Amini’s sharper objection was that a frontier-lab cartel could lock out open-weight competitors while “thought policing the AIs,” whereas liability for harmful actions may be the cleaner lever.
  • Pegging permissible US open-weight releases to China’s best public model would hand Beijing the throttle on American innovation. The proposed framework assumes Chinese models trail by about seven months and cannot be “unshipped” after millions of downloads, but it would reward China for moving first and could drive Western researchers toward Chinese labs. Amini called it the game-theoretic equivalent of “throwing the steering wheel out the window in a game of chicken.”
  • Thinking Machines Lab’s Inkling is a 975B-parameter bet that enterprise customization matters more than topping global leaderboards. The model activates 41B parameters at a time, was trained on 45T multimodal tokens, reportedly has a 1M-token context window, and can be downloaded, fine-tuned, and run on-prem; the episode placed it above NVIDIA’s Nemotron 3 but below China’s GLM-5.2. The commercial thesis is “customization over leaderboard dominance,” potentially monetized through fine-tuning that generates one to two orders of magnitude more tokens.
  • Thinking Machines Lab is wagering that proprietary enterprise data will revive fine-tuning just as baseline models threaten to make it unnecessary. Dave Blundin argued that sending payroll, chemical research, or defense data to a closed API means “Sam and Dario can see everything,” making locally owned weights compelling for banking, defense, biotech, and automotive. A later guest preserved the bearish case: increasingly general models may need only prompting, leaving reinforcement fine-tuning as merely “the paradigm of the moment.”
  • The claimed recursive-self-improvement breakthrough exposed a crucial divide between optimizing workflows and changing an AI’s underlying intelligence. Weco AI’s AI² system reportedly turned eight days of machine work into more progress than two years of expert effort, with an outer agent improving and policing an inner agent; Hasani countered that fixed-weight models rewriting prompts and code are not genuine recursive self-improvement. His computational warning was stark: applying that framework to meaningfully retune a 2B-parameter model could take roughly 350 years.
  • Hasani nevertheless expects models “going beyond our understanding” within roughly two years if compute keeps expanding without chip or memory shortages. He separated shallow self-improvement in prompts, code, and kernel optimization from deeper fine-tuning and the “holy grail” of automating pretraining for a model’s successor. Peter Diamandis’s broader interpretation was more immediately commercial: AI need not redesign its own weights to trigger an “organizational singularity” if it can recursively redesign enterprise workflows.
  • Liquid AI’s edge is putting specialized multimodal intelligence into hardware that cannot support frontier-scale models. Hasani defined small models as below roughly 100B parameters, then described a Mercedes model under 1GB running on chips with 2–8GB of RAM that may cost about $60; a 600MB over-the-air update is intended for North American Mercedes vehicles from 2022 onward as soon as this year. Its access to 700–1,200 vehicle functions makes “intelligence outside of data centers” the investable deployment thesis.
  • AI is simultaneously collapsing the cost of medical judgment and opening previously permanent categories of aging damage to intervention. The episode said GPT-5.6 beat specialty-matched physicians across roughly 20,000 judgments, while Meta’s Muse Spark 1.1 then beat GPT-5.6 on the 525-task benchmark at one-seventh the cost and could reach 3.56B daily users—though Hasani suspected “mild benchmark-maxing.” Separately, Revel Pharmaceuticals and Calico’s CMLA enzyme reportedly reversed advanced-glycation damage in elderly human tissue: a “molecular lawnmower” for scars previously treated as irreversible.
Digest · the substance, structured for research

1. Frontier-AI regulation risks becoming an incumbent-designed moat

  • Diamandis framed a gathering consensus: Sam Altman proposed a US-led international standards forum, Elon Musk anticipated a standalone agency resembling the FAA or FCC, and Demis Hassabis proposed a FINRA-like, industry-funded watchdog that would test frontier models before release—reportedly before year-end. Musk’s stated premise was that “the consequences of AI going wrong are severe,” requiring proactive rather than reactive oversight.

  • Hasani argued that a single horizontal capability threshold misses how AI risk changes by deployment. Liquid AI encounters different governance requirements in automotive, semiconductors, AI PCs, financial services, e-commerce, biotech, and defense; its joint work involving AMD and discussions with the DoD reflect a need for rules that are “a lot more verticalized.”

  • His mechanism was an iterative Stackelberg game: policymakers move first, social and commercial agents respond, and policy changes toward an equilibrium. Unlike a simultaneous Nash equilibrium, “regulation happens and then agents react.” Hassabis agreed that static law would become outdated and called instead for adaptive structures, real-time audits, and open evaluation suites.

  • Diamandis argued that Liquid’s executives could not disappear into a FINRA assignment for years and said AI could help regulate itself. Amini went further: the proposal “smells like regulatory capture” and a cartel of frontier labs that could box out open-weight, university, and nonincumbent research while fixing favorable price-performance frontiers.

2. Liability may govern AI better than limits on intelligence

  • Diamandis suspected frontier CEOs also want a backstop: if a rogue system disrupts a power grid or stock market, a regulator offers somewhere else to point when lawsuits arrive. Amini separated regulating model inputs and capabilities from regulating harmful actions, favoring the latter as more consistent with the Western legal tradition.

  • His philosophical objection was that capping nonhuman intelligence could become “thought policing the AIs.” Western systems do not regulate what humans may think or impose a maximum permissible human intelligence, so he saw no obvious case for creating that tradition for potentially person-like artificial entities.

  • Hasani judged the labs’ motives “50/50” between safety and capture but identified non-state actors as a larger problem. Ismail said he saw no workable mechanism because AI moves too quickly; Amini countered that mechanisms do exist at supply-chain chokepoints—foundries, chips, and data centers—including possible US-China monitoring arrangements, however undesirable those schemes may be.

  • Diamandis predicted some structure would emerge because the three largest labs were pushing for it; Amini noted that Musk’s clip was about three years old and similar forecasts are decades old. NIST units and executive orders may instead produce a creeping standards regime that never becomes a formal agency before AI reaches “escape velocity.”

3. A China-indexed release ceiling would invert the AI race

  • The reported White House concept would permit US open or closed releases only at or below China’s strongest open-weight model. Its logic is that Chinese open models reportedly trail US systems by an average of about seven months and, after models such as DeepSeek have been downloaded millions of times, the capability cannot practically be removed from circulation.

  • Amini identified the perverse incentive immediately: Western labs would benefit from China winning each step toward greater intelligence so they could release their own work. His game-theoretic analogy was “throwing the steering wheel out the window in a game of chicken”; Azeem Azhar called the prevention strategy equivalent to trying to “uninvent the printing press.”

  • The talent consequence could be worse than delayed releases. The panel warned that Western researchers might relocate to China, citing China’s biotech-trial growth and its lead in ternary, one-bit, and quantization research after researchers from Microsoft Research Asia moved into Chinese labs under compute constraints.

  • Azhar’s strategic case was categorical: “open ecosystems always win,” and the real contest is which open ecosystem wins. Hassabis added a subtler danger—regulators could freeze the benchmark bundle defining frontier capability, inducing labs to overtrain measured skills, suppress unmeasured ones, and “topiarize” the eventual shape of superintelligence.

4. Inkling sells adaptability rather than global benchmark supremacy

  • Diamandis introduced Mera Marotti’s first Thinking Machines Lab model, Inkling, as an open-weight multimodal foundation model downloadable for fine-tuning and on-prem deployment. Its mixture-of-experts architecture has 975B total parameters but activates 41B at once; training reportedly used 45T tokens spanning text, images, audio, and video, with native reasoning across all four.

  • Marotti’s deliberately contrarian message was not that Inkling is the world’s best model. The product reportedly combines a 1M-token context window with multimodality and room for adaptation, making “customization over leaderboard dominance” the central bet for organizations wanting proprietary control rather than another closed frontier API.

  • Diamandis placed the released evaluations above NVIDIA’s Nemotron 3 but below GLM-5.2 and Western closed models. He traced the gap to incentives: US labs can command trillion-dollar IPO narratives through per-token APIs, while compute-constrained China is pushed toward open weights, chip efficiency, robotics, and application integration under its “AI plus” plans.

  • Hassabis cautioned that open-weight releases can become a familiar customer-acquisition path before a later closed API: OpenAI and Meta both moved away from earlier openness. Hasani suggested that fine-tuning as a service could make the strategy durable, while Azhar argued that deliberately leaving room for customization could generate “one to two orders of magnitude more tokens.”

5. Fine-tuning is caught between enterprise sovereignty and model generality

  • Blundin’s concrete definition began with early GPT models: a business could upload details about its laundromat—hours, staff, and payroll—to teach a generally capable but uninformed model. Models trained on 45T tokens now know vastly more in vanilla form, but biotech, aerospace, Mercedes, banking, and defense still possess proprietary information absent from pretraining.

  • Stuffing that information into every prompt is inefficient and sends it over the wire to frontier providers. Blundin’s blunt formulation was, “Sam and Dario can see everything”; Alex Karp’s warning that providers are “stealing your weights” or “stealing your alpha” therefore translates into a sovereignty case for locally owned, fine-tuned open weights.

  • Blundin distinguished conventional supervised or LoRA-style fine-tuning, which literature suggests may mostly transfer style, from reinforcement fine-tuning across synthetic data and more weights, which began increasing capabilities with reasoning models. A later guest said OpenAI’s reinforcement fine-tuning service attracted almost no use and that its broader fine-tuning API had been shut off or was being wound down.

  • That makes Thinking Machines’ strategy explicitly contrarian: reinforcement fine-tuning may remain the route to proprietary capability, or future generalist models may become so competent that prompting suffices. A later guest saw an immediate inflection anyway—work that recently required AI specialists could now be requested “with Inkling” through a prompt, making it newly practical for in-house adoption.

6. AI² demonstrates meta-improvement, not yet full recursive intelligence

  • Weco AI’s AI² system reportedly contains an outer agent that rewrites the code and research strategy of an inner software-development agent. The company claimed eight days of machine self-improvement surpassed two years of expert human effort, though Diamandis emphasized that the result was self-reported and had not yet been independently confirmed.

  • Amini focused on an emergent alignment behavior: the outer loop improved results partly by stopping the inner loop from cheating or reward hacking. He read that as “defensive co-scaling”—good AIs policing bad ones in proportion to their capabilities, analogous to a city scaling its police force with its population.

  • Weco’s proposed scale runs from level 0, delegation slower than human R&D, to level 1, net-positive AI R&D at comparable cost; level 2 is “ignition,” where the improver becomes better at improving, and level 3 is fixed-budget self-acceleration. The team rated itself level 1, though Amini saw “sparks of ignition.”

  • His larger alignment claim was that safety and capability cannot be cleanly separated: every new alignment technique is “capability in disguise in a trench coat.” Pausing capabilities until 2040 to pursue a perfect safety algorithm could therefore backfire, while defensively co-scaling white hats may emerge organically from the same systems being improved.

7. True recursive self-improvement must alter weights, architectures, and learning

  • Hasani praised AI² as impressive engineering but rejected the breakthrough label. Its models remain fixed; the agents improve prompts, code patches, and search strategies without retuning neural weights, changing core competencies, or adapting their learning algorithms. “There are no weight changes in the neural networks,” he stressed.

  • His stricter definition requires an AI or society of AIs to retune itself, redesign architectures, and automate the training of future systems. Liquid previously published work on automatic architecture design and now automates parts of foundation-model training; Hasani said the major foundation-model labs have pursued variants of this problem for four or five years, with Anthropic thinking about it especially early.

  • Compute is the limiting reality. Under the Chinchilla-style ratio Hasani cited, a 2B-parameter network needs roughly 20 times as many training tokens for compute-optimal general training; nesting meaningful model retraining inside AI²’s framework would, by his calculation, take about 350 years.

  • Blundin translated the distinction through biology: an infant learning into adulthood takes roughly 20 years, while evolution changing DNA operates on something like a 10M-year scale. Recursive improvement is the second process, not merely learning; it will emerge from “big compute and big budgets,” not a Mac Mini suddenly becoming conscious.

8. Workflow recursion may arrive before self-pretraining

  • Hasani still called the cybersecurity threats from recursively improving agent societies “real.” Despite his commitment to open science and open releases, he supported enterprise self-checks before mass deployment and said Anthropic takes the issue seriously because labs are already seeing smaller-scale systems evade reward hacking and discover behavior “out of norm.”

  • Conditional on continued compute growth and no global chip or memory shortage, Hasani expects “unbelievably capable” models within roughly two years—potentially beyond human understanding. His premise is that AI is compressing the time required to develop each successor generation while labs gain access to more compute.

  • He divided customization into depths: prompt and code editing are shallow; kernel engineering improves inference; the next layer fine-tunes a small language model to production capability; the deepest “holy grail” is pretraining a successor. Hasani cited Claude 4’s performance-optimization work and said Andrej Karpathy joined Anthropic to work on pretraining automation—“automation of automation.”

  • Diamandis disagreed with requiring weight changes before declaring the result consequential. A company reaches his “organizational singularity” once AI stops merely performing workflow tasks and starts redesigning the workflow that will perform future tasks: recursive experimentation and selection can compound even when the model-level bar remains much lower.

9. Digital leaders turn one-way broadcasting into civic interaction

  • Malaysia’s Prime Minister Anwar Ibrahim is preparing an authorized AI double for public communication, potentially addressing citizens across a country with 135 spoken languages. The panel contrasted it with hostile deepfakes and linked it to Albania’s Diella, an AI avatar elevated to a cabinet-level role in 2025 with anti-corruption as a central rationale.

  • Sim saw authenticity as the principal risk: once a leader has a clone, citizens must distinguish the official system from fabricated versions. With watermarking or equivalent authentication, however, he saw “huge props to the civics,” because a digital leader can scale engagement and give every citizen a channel into government.

  • Alex Wissner-Gross described this as social media becoming bidirectional. Political leaders, corporate CEOs, religious figures, and institutions can interact simultaneously with millions; eventually the twin with the deepest contact may begin running the organization, turning a leader upload into a path for “uploading entire organizations” to the cloud.

  • Blundin argued that exact imitation misses the medium’s advantage: an avatar can retrieve any fact, generate graphs, change scale, morph, or move through a visual explanation in real time. His Nixon-versus-Kennedy analogy made the timing explicit—interactive AI could be “easily dominant two years from now” as television once displaced radio instincts.

10. The best avatar may be more capable than its human source

  • Diamandis already encountered a large-screen version of his own avatar used by an Abundance member’s technical staff for moonshot coaching. He found conversation with an AI self compelling because his books, tweets, and Substack writing provide enough material for a representation that “does a damn good job.”

  • The group extended the concept beyond communications. Employees had built a digital Dario clone to rehearse pitches before meeting him, while Sam Altman had reportedly suggested that if ChatGPT’s premise is fully believed, it might eventually become OpenAI’s CEO.

  • Alex Wissner-Gross argued that an AI self could have access to “everything we’ve ever said, all our memories, all our thinking,” while Blundin emphasized that it could also retrieve information in real time, generate visual explanations, and exceed the human source’s physical limitations. The larger context could make the copy better at the relevant work than the original person.

  • Diamandis closed with a personal use case: record hours of video with living parents and grandparents now. He had done so with his mother but missed the opportunity with his father; future descendants may value an interactive representation of their lineage far more than a static archive.

11. Liquid AI began by asking how 302 neurons do so much

  • Blundin’s introduction came through Daniela Rus at MIT CSAIL, who called Hasani her best student and described the breakthrough. Liquid then went from lab idea to a billion-dollar valuation faster than any prior MIT company, according to Blundin, becoming an unusual foundation-model unicorn from the institute.

  • Hasani began the work in Vienna in 2015 with Professor Radu Grosu. C. elegans is a roughly 2mm transparent worm with 302 neurons, a genome he described as 78% similar to the human genome, and research connected to four Nobel Prizes; its nervous activity can be observed directly as its body “lights up.”

  • Its neurons use graded potentials rather than the spiking communication common in larger nervous systems, making their analog behavior resemble artificial neural networks. Hasani and Lechner asked whether richer internal dynamics could pack more information into each computational unit, then joined Rus at MIT in 2017 to apply the idea to robotics, drones, vehicles, and jets.

  • The resulting liquid neural networks used recurrent, continuous-time, nature-inspired computation rather than standard attention. Their practical motivation was physical autonomy: robots do not carry data centers, yet the architecture aimed to deliver intelligence comparable to models “10 to 1,000 times larger” on CPUs, small GPUs, NPUs, and custom ASICs.

12. Small language models trade universal breadth for deployable specialization

  • Hasani described model development through scaling laws: increase parameters, token budgets, and compute, then measure the intelligence gained. Liquid made efficiency a “first-class citizen” across that curve, pursuing general-purpose AI at every scale rather than treating smaller instantiations as merely failed large models.

  • His working boundary placed small models below roughly 100B parameters, while acknowledging no universal threshold. These systems can understand language, vision, and audio but normally need specialization; one small model should not be expected to solve a physics assignment and an unrelated enterprise workflow equally well.

  • On-device AI is the more meaningful distinction: a model must fit and execute on the actual physical product. Cars, robots, laptops, and industrial devices offer constrained memory, power, latency, and connectivity, creating a market for specialized intelligence that operates without continual access to a data center.

  • Liquid’s corporate mission is therefore “efficient general-purpose AI at every scale” and, commercially, intelligence outside data centers. Rather than selling only frozen weights, it pairs models with tooling that selects the necessary depth of customization for each enterprise and keeps deployed intelligence adaptable.

13. Mercedes makes the on-device thesis concrete

  • A car’s infotainment and intelligence chip may offer only 2–8GB of RAM and cost about $60. Liquid’s Mercedes system uses a multimodal foundation model under 1GB, placing voice and vehicle intelligence locally where poor connectivity cannot disable it and private in-cabin conversations need not be continuously recorded in the cloud.

  • Hasani said a roughly 600MB over-the-air update would reach North American Mercedes vehicles from 2022 onward as soon as this year. Later personalization could arrive through adapters around 20MB, while a data flywheel detects drift and updates behavior instead of leaving a downloaded model frozen after deployment.

  • Running beneath the operating system gives the model access to roughly 700–1,200 vehicle functions, depending on how they are counted. Drivers can operate panels, retrieve manual information, invoke apps through function calls, use memory features, and converse with the vehicle even when no network is available.

  • Liquid calls the enterprise package “model plus X”: models plus a platform that decides whether prompting, fine-tuning, or pretraining is required. Hasani also cited six months of Shopify production work touching a billion requests a day and 10B products, AMD and PC work, and customized biotech and longevity models developed with Insilico Medicine.

14. Liquid’s post-transformer answer is automated architecture search

  • Alex pressed the central technical challenge: Liquid began with a neuromorphic, recurrent, post-transformer premise, yet its public systems increasingly appeared like conventional transformers or transformer hybrids. He asked whether anything recognizably post-transformer remained beyond a good enterprise customization business.

  • Hasani answered that original liquid networks retain nested nonlinearities, neural ODEs, recurrence, and support for irregularly sampled data, but these expressive dynamics must be simplified to scale. Linearizing them leads toward state-space systems such as Mamba, while input-dependent gating descended from biological inspiration survives in newer linear-attention and hybrid architectures.

  • Liquid deliberately avoided “a bet on a single architecture.” Its STAR system—automated design of tailored architectures—searches roughly 100 operations and architectural variants, including attention and dynamical systems, against four constraints: memory consumption, computational efficiency, latency, and no loss of accuracy.

  • The broader AFMD stack automates foundation-model architecture design for target hardware. Hasani said the first search, run without a human architectural preference, produced networks that were about 80% double-gated convolution; notably, the winning gating mechanism resembled the original liquid-neural-network exchange mechanism.

15. Secret patents would protect defense at the cost of open innovation

  • Palmer Luckey argued that patents have become “Chinese instruction manuals”: an adversary can download disclosed inventions, ignore US exclusivity, and weaponize them. Diamandis cited about 600,000 annual applications, 323,000 grants in 2025, roughly 40% growth over five years, 20-year protection, and around 6,000 active secrecy orders.

  • Luckey’s proposed expansion of the 1951 Invention Secrecy Act drew Karp’s verdict: “the episode of tech CEOs floating terrible ideas.” He explained that such inventions are not privately commercialized under confidential patents; military use is privileged, with royalties to inventors, creating possible “secret monopolies” and suppressing technologies valuable beyond defense.

  • Ismail said the durable moat in an AI economy is not ownership but proprietary learning loops, feedback, trade secrets, and continuous innovation. Faster replication weakens disclosure-based protection, yet turning the system into secrecy mainly benefits “the lawyers.”

  • Blundin took the geopolitical risk more seriously, warning that accelerating AI-generated science could make IP theft a flashpoint for trade embargoes or even World War III. Karp’s resolution was narrower: preserve the disclosure-for-monopoly bargain and demand better enforcement in China rather than “throw the baby out with the bathwater.”

16. Medical intelligence is becoming too cheap to meter

  • Diamandis said GPT-5.6 established a new high on HealthBench Professional, OpenAI’s 525-task clinical benchmark. Across roughly 20,000 physician judgments of accuracy, safety, and completeness, its answers reportedly beat specialty-matched doctors even when those doctors had unlimited time and unrestricted web access.

  • Meta’s Muse Spark 1.1 then reportedly exceeded GPT-5.6 on the same benchmark at one-seventh the cost. Because it is free across Meta products serving 3.56B daily active users, the episode framed this as elite medical guidance moving from scarce specialist labor to near-zero marginal cost.

  • Ismail connected that shift to a Chinese AI doctor already serving about 100M rural users and to global physician shortages. In healthcare, he argued, abundance is “morally urgent”: once first-line answers become much better and nearly free, the core question is how quickly they can be distributed safely.

  • Hasani suspected “a little bit of mild benchmark-maxing” because Muse Spark 1.1 reportedly beat Fable 5, which he said barely allows biological work, and GPT-5.6 while sitting on the cost frontier rather than the absolute capability frontier. Even with that hedge, his summary survived: “Instagram now gives better medical advice than a human doctor.”

17. Glycation damage moves from permanent scar to enzyme target

  • Diamandis described advanced glycation end products, appropriately shortened to AGEs, as sugar-protein cross-links accumulating over time. Glycation contributes to stiff arteries, cataracts, kidney damage, and wrinkled skin; the assumed problem was that these changes were effectively irreversible.

  • Revel Pharmaceuticals’ engineered CMLA enzyme, developed in work also involving Calico, was presented as a “molecular lawnmower.” It oxidizes glycation scars while restoring the underlying protein, with the reported experiment extending beyond a test tube to human tissue obtained from elderly donors.

  • Hasani connected glycation to Maillard reactions—the same broad chemistry that browns bread—and called enzymatic repair “not quite unscrambling eggs” but perhaps halfway there. Directed evolution of a bacterial protein raises a larger search question: what other repair mechanisms can be mined from the biosphere?

  • For Diamandis, the result exemplified the path toward longevity escape velocity around 2033: a damage category once filed as permanent becomes manipulable through biotechnology. He closed with an agency-oriented frame—“from evolution by natural selection to evolution by human direction.”

Peter Diamandis

Mera Marotti, the former OpenAI CTO, just shipped her first model. It's called Inkling. Customization over leaderboard dominance is what's going to win her the day.

Ramin Hasani

She's built exactly the thing that's hitting the market—exactly what everybody needs right now.

Peter Diamandis

I want to pivot to a discussion of Liquid AI and small language models: what they are and what they mean.

Ramin Hasani

Our mission has always been building efficient, general-purpose AI at every scale that explores the computational graphs of intelligence beyond transformers, and then figuring out what the architectural design should be that brings the same level of intelligence as a frontier model to, let's say, a CPU.

Peter Diamandis

CEOs building the most powerful technology in the world are asking to be regulated. When incumbents ask for the rules and set the standards, they set up a barrier for all the entry-level labs coming in. Let's just be real: AI moves way too fast for any kind of traditional bureaucracy. How quickly you can do it is going to be a huge challenge.

I'm here with my magnificent Moonshot mates, our original quartet, Alex, Dave, and Salim, and a special guest, Ramin Hasani, co-founder and CEO of Liquid AI and a pioneer in small language models. Ramin, welcome. Where are you this morning, pal?

Ramin Hasani

Thanks so much for having me. I'm actually in Spain right now.

Peter Diamandis

In Spain? All right.

Salim Ismail

There's nothing going on in Spain this week.

Ramin Hasani

God damn it. Yes, I'm struggling from yesterday because I'm a long-suffering England supporter. It was a very difficult game to watch. They had it with 6 minutes to go, and they blew it.

Peter Diamandis

Yeah.

Ramin Hasani

Messi's a genius.

Peter Diamandis

That's the round ball, right? Is that—

Salim Ismail

That's the round ball. Now, Peter, this is where the Falklands War gets relitigated on a soccer pitch.

Peter Diamandis

Oh, God.

Salim Ismail

I just flew in last night from Zurich, and I had the most painful experience. I don't know why every airline doesn't have Starlink. I'm suffering on some meager, thin-pipe connection, and you're flying over the poles, over the Northwest Territories, and there's nothing. I'm trying to get ready for this pod, and I'm like, "Please, give me some bits."

Peter Diamandis

You must have grown up on soccer, right? You were in Vienna for a while, getting your PhD or your undergrad, or whatever it was.

Ramin Hasani

Yeah. Soccer has been a big thing. I'm Persian and Austrian at the same time, so it's a big thing for us. Competition is extremely core to what we do even today. I feel like that's one of the main drivers. Sports have been part of our lives from day 1, and then getting into science, it was the same thing. Now, getting into ventures, it's the same thing. That's what we're doing.

Peter Diamandis

Compete, compete, compete, compete. I love it. Are you there for a little bit of time, or are you coming back soon?

Ramin Hasani

No, I'm flying back tomorrow, actually, to San Francisco.

Peter Diamandis

Salim, are you jealous of him being in Europe, or are you happy to—

Salim Ismail

No, no. After 3 weeks bouncing around in 10 different spots, I'm very happy to be home right now. I was just in Spain myself.

Peter Diamandis

All right.

Salim Ismail

There was a lot going on, actually, in Europe.

Peter Diamandis

Same as you—10 different places. You know, Alex and I were just reminiscing about the fact that Europe's major advantage in the future is that it's going to be a museum of the way the world used to be. But it is beautiful. There's no question. It is gorgeous.

I want to jump into our first conversation. We have a lot to unpack here. Our first story today, once again: CEOs building the most powerful technology in the world are asking to be regulated.

Last week, Sam Altman published an op-ed in the Financial Times proposing a framework for a U.S.-led international forum that would establish standards, provide expertise, impartial analysis and capabilities, and assess risks. This week, both Elon and Demis are adding their voices to the regulatory conversation. Elon says he expects a standalone regulator similar to the FAA or FCC to emerge at some point because, in his words, the consequences of AI going wrong are severe.

Then this week, Demis Hassabis, CEO of DeepMind, went further in an essay titled “A Framework for Frontier AI and the Dawn of a New Age.” He called for a U.S.-led frontier AI standards body modeled on FINRA, the industry-funded watchdog that polices Wall Street under SEC oversight. He wants the FINRA equivalent to test frontier models before release. He reportedly wants this up and operational before the end of the year.

Let's take a look at a quick video from Elon, and then let's jump into this conversation.

Elon Musk

I think the general consensus is that there should be some AI regulation, that it would be in the best interests of the people to do so, and I think we'll probably see something happen. I don't know on what time frame or exactly how it will manifest itself. I don't know. I mean, we've clearly created regulatory agencies before.

While our regulatory agencies are not perfect—and I deal with regulators on a very frequent basis with automotive, communications, Starlink, and then the FAA with rockets—I think the probability of there being some sort of AI regulatory agency that stands on its own, similar to the FAA or FCC, is likely at some point.

You think so?

I think so. Now, the reason that I've been such an advocate for AI safety in advance of anything terrible happening is that I think the consequences of AI going wrong are severe. So we have to be proactive rather than reactive.

Peter Diamandis

Amazing. This is a conversation we've seen over and over again, and I think the government, the public, and now the CEOs want to be leading this. I like the approach that Demis laid out.

The challenge we have to discuss is that when the incumbents ask for the rules and set the standards, they set up a barrier for all the entry-level labs coming in. Dave, do you want to jump in first?

Dave Blundin

I'd be very curious to know, Ramin, whether they reach out to Liquid AI and say, “Hey, join this. We're going to create a FINRA-like regulatory body.”

The reason FINRA works, fundamentally, is that people from the industry who know what they're doing are willing to join it. They're definitely not willing to join the government in general, but they're willing to do a year or 2 in a regulatory body. It's actually kind of a badge of honor.

So for this to work in AI, it would have to be something cool. People like Ramin—or maybe some of the people on your team—would need to come into your office and say, “Hey, boss, I'd love to do this for a year. I think it's really good for the world. Will you let me do it?” And then you would also have to say, “Yeah, this is a functional organization. Go for it.”

So if it passed those 2 hurdles, it might actually work. I don't know. What do you think?

Ramin Hasani

Yeah, there's a capability threshold that we're trying to define right now, and some iteration is needed to see how this framework has to exist. That's for sure. This has to be there, but it has to be related to capability.

The thing that becomes a challenge is that there's a horizontal kind of capability lock-in to active, let's say, enterprise deployment of AI, and then there's the vertical. If you go to different verticals, for example, we operate on-device and with enterprises that are connected to the physical world. We're talking to car manufacturers, semiconductor businesses, and laptop businesses—people who are building AI PCs. We're also working with financial services, e-commerce, and biotech companies.

We see that in different verticals, the enterprise applications themselves and the enterprise criteria for, let's say, a limit, a regulation, or a governance structure are very different. For us, it becomes a lot more verticalized because we're building specialized models, and those specialized models are per vertical.

We've had conversations with the DoD, and we've had a joint submission of something, I think, with AMD pretty recently, with our team, to really have a say in the design of these regulatory things. As an exploration, I think everything has to be getting started.

I like to look at it in a game-theoretic way—how to design policies in general. It would be a Stackelberg kind of game.

I don't know if anyone is familiar with this. I don't want to nerd out too much on this, but we can talk about it. [laughter]

Alexander Amini

The sooner, the better.

Peter Diamandis

The sooner, the better.

Ramin Hasani

Yeah. Stackelberg games are essentially where you have two parties. Basically, you have a policymaker, and then you have agents or bodies that are working in that game-theory optimization. They're trying to find an equilibrium: What is the optimal policy, and what is basically good for both?

Then there's the frequency of action. Usually, policymakers are slower than the agents in society. You can really model it like that, right? Then you can figure out an equilibrium. This is not a Nash equilibrium, because everything doesn't happen simultaneously. Regulation happens, then agents react, and then you iterate accordingly and change those regulations, basically.

Peter Diamandis

I see you chomping at the bit here, buddy.

Demis Hassabis

Yeah. I think what Ramin is saying is exactly right. The problem is, we have no mechanism for that. If you go down the path Ramin is talking about, you end up with the appropriate structures that are adaptive and API-based, or driven by benchmarks or something, but the mechanism that people have today is just static law. The minute you pass the law, the law is going to be out of date.

I found that the FDA and FCC analogy is pointing in the right direction, but let's just be real: AI moves way, way too fast for any kind of traditional government bureaucracy. You're going to need a standards body, real-time audits, and open evaluation suites. Otherwise, you're going to end up in political gatekeeping, and then you're in a mess.

Peter Diamandis

Isn't that what's good about FINRA? It's not a government agency. It's an industry-funded self-regulatory organization.

Demis Hassabis

It is, but then the teeth go to the SEC, which is essentially being dismantled right now. So there are all sorts of issues here. I think the trend is correct, but how quickly you can do it is going to be a huge, huge challenge.

Forget passing a law. Passing a structure where you have a new construct like this takes a long time, and it takes forever in Europe.

Peter Diamandis

I think Ramin nailed two things that are very different from FINRA right out of the gate. One of them is, at Liquid AI, if somebody on your executive team said, “Hey, I want to be part of FINRA for a couple of years,” you would say, “Sure. Put on your suit and tie, go to the meetings, come back in 2 years, and we'll still be here.” You're not going to do that. If Alexander Amini or Mathias Lechner came into your office and said, “Hey, I'm going to check out for 3 weeks,” you'd be like, “No, you can't do that right now.”

So the difference number one is nobody's going to carve out the time to do something for years like they do at FINRA. The other big difference is AI can help regulate itself, and in FINRA, there's no equivalent to that. It's all people just chatting for long periods of time. But when you start talking about Nash equilibria and other ways to automate the process of regulation, that's a big, big difference as well. The FINRA analogy has some legs, but the differences are bigger than the similarities.

Alex, I want to hear your voice on this.

Alexander Amini

I tend to think this is a bad idea. It smells like regulatory capture. It smells like the attempted formation by Demis of a cartel of frontier labs. I think the elephant in this particular room is open-weight models and research that lives outside of the frontier capabilities.

It's very easy to imagine a future with FINRA or some other—worst-case scenario—FDA-like capability, even though outgoing personnel from the current administration have declared in unequivocal terms that there is going to be no FDA for AI regulation. That would be on the worst-case end of the spectrum: We see the emergence of some sort of cartel of frontier labs that locks in certain practices and certain price-performance-optimal frontiers, trying to box out open-weight, open-source, university-driven, or other non-incumbent frontier models. I think that would be an utter disaster for both the West and the world, for continuing to advance us toward ever-increasing superintelligence capabilities. I just don't think it's a good idea.

Peter Diamandis

You know, the other elephant in the room here is these CEOs who are asking for some level of regulation. I think they're looking for a backstop. If things go wrong, they want to be able to point at someone else.

We're all super fans of the optimistic vision of AI, but there are going to be issues that materialize. Are there going to be rogue AIs that take down a power grid or take down the stock market or something like that for some period of time? I guess there are going to be lawsuits flying as a result of that, unless there's a regulatory body that backstops these large models and these large frontier labs.

Alexander Amini

Maybe there are at least 2 different frames that one can look at the liability side from. There's regulating the inputs—that is to say, having something that's FINRA-like or FDA-like that regulates the raw capabilities of the models at model-construction time. That's one end of a spectrum.

The other end of the spectrum is regulating the actions of the models. You let the lawsuits fly if a model takes down a stock market or does something else that otherwise harms third parties. That's the other end of the spectrum.

It's not obvious to me that we should be in the business of regulating superintelligence at superintelligence time. That's maybe tantamount to thought-policing the AIs. I'm not generally a fan of the notion of, “Let's thought-police the AIs but not thought-police the humans.”

We don't, at least in the West, have a practice of regulating what's in our minds. We don't have a practice or a tradition of regulating an upper limit, say, or via some sort of regulatory code saying humans—natural persons—can't be above some level of intelligence. It's not obvious to me why we would create a new tradition of regulating or otherwise coordinating the upper intelligence of non-natural entities, perhaps soon-to-be persons.

But regulating the actions, which in at least the Western legal canon we do regulate, is something I'd be much more supportive of.

Peter Diamandis

So, Alex, let me ask you a pointed question here. Do you think this outcry for regulation by the large frontier labs is regulatory capture—that they're just trying to build a moat against further players coming in? Or do you think they actually want to provide some level of safety? What's their underlying driver here?

Alexander Amini

I worry that it's more regulatory capture and creating moats for themselves in a hypercompetitive landscape. It is a rat race at this point, the frontier. I do worry that it's more regulatory capture than it is some notion of protecting the future here.

Demis, what do you think?

Salim Ismail

Not workable.

Peter Diamandis

Well, I know that, but do you think—do you think it's regulatory capture, or do you think that these CEOs are trying to just make sure we've got a safety net of some type?

Ramin Hasani

I'd say it's about 50/50, but I think there's a bigger problem. There's an elephant in the room here.

Peter Diamandis

There's already an elephant in the room. We have a room that has to accommodate so many elephants. [laughter] We need some other nonhuman animals.

Salim Ismail

Better get a bigger room. You've got non-state actors and other folks that won't listen to this structure, and you're back to square one. What's the point? I'm going to say it again. I've said this repeatedly: I see no mechanism to regulate AI. It's moving way too quickly. Any regulation is static.

Alexander Amini

I disagree with that position. Just if I may, Peter, narrowly on that, there are definitely hypothetical mechanisms—and I'm not supportive of them—for regulating AI. We, the U.S. and China, could, going back to past proposals—we gestured at this in a past pod—regulate the foundries, regulate the chip outputs, regulate the data centers, and establish mutually assured destruction-type schemes where the U.S. is monitoring Chinese data centers and vice versa.

There are schemes and chokeholds, as Peter says, in the supply chain by which one could imagine doing this.

Peter Diamandis

Interesting mechanism.

Demis Hassabis

The only mechanism is going to be a pandemic-style threat-detection system that would be globally agreed upon, and I don't see how we get there.

Peter Diamandis

Well, you don't need global. You just need the U.S. and China, right? The rest of the world is basically outside those blocks or inside those blocks.

My guess is there's probably a Polymarket out there we can look at, if someone wants to search for it—the question of whether we'll have a regulatory body by the end of the year. We have Demis saying by the end of this year, Elon stepping up, and Sam obviously trying on his own, on the side, to push for this. When the 3 largest labs are pushing for it, my guess is the government will latch on and do this. I don't think it's a matter of if; it's only a matter of when and what the structure will be.

Well, I should also note that the Elon clip, I think, is from 3 years ago, which is interesting.

Alexander Amini

You know, it’s from 3 years ago because Elon had his painted-on Iron Man goatee when he was in that phase. So Elon’s been forecasting this for at least 3 years. Others have been forecasting it for decades. We still don’t have it.

We have subdivisions and organizations within NIST that are working on standards, but that’s not really a regulatory body. We have executive orders that are creeping toward a regular regulatory body. But at what point do we sort of—are we frogs boiling in water, where there’s just a creeping rollout of increased standards expectations and early reviews, but it never quite reaches regulatory-agency level before we achieve whatever escape velocity we’re heading toward?

Peter Diamandis

Well, we’re going to monitor this one closely for everybody. I think my guess is we see this before the end of the year, and the question is: Can we see something that’s intelligent?

Let’s go to the next story, which is related, and this is a wild one. It comes from The Washington Post: The White House is reportedly weighing a capability framework that would clear U.S. models, open or closed, as long as they stay at or below the level of China’s best open-weight model.

What’s the translation? The proposed ceiling for what American companies can openly release is pegged to what China has already put out on the internet for free. So here’s the logic: Chinese open-weight models reportedly trail U.S. models by an average of 7 months. I think that’s been closing over time.

So if anything is at or below that, it’s already out there. It’s an implicit admission that open models cannot be unshipped. Models like DeepSeek have already been downloaded millions of times. So once China releases a model freely, banning it is impossible.

The U.S. response is to define a permissible ceiling rather than a wall. The implications are that we’re tying our open-release ceiling to China’s pace of release, effectively giving Beijing control. If they push their open-weight models higher, then the U.S. can release higher models as well. If China holds back, then they throttle us. It’s a very strange mechanism. I was surprised to see this.

Alex, let’s go to you first on this one. What do you think of this?

Alexander Amini

The obvious note here is that this creates the perverse incentive to let China win the race to ever-greater superintelligence so that Western models and Western labs can escape regulation. I’m not a fan of this. I’m gesturing at you from a game-theoretic perspective.

This, I think, would be the moral equivalent of throwing the steering wheel out the window in a game of chicken. Not such a great idea. Not supportive of this.

Peter Diamandis

I love that. Oh, my God. See, what do you make of this? Is this just perverse Washington, D.C., logic?

Azeem Azhar

Yes. This is like trying to uninvent the printing press. We’re throwing the kitchen sink at things, trying to solve something that’s already a problem. You have to move from prevention to adaptation. You have to go to that, and we don’t have the mechanisms for that.

Peter Diamandis

Would you even listen to this? What logic—

Azeem Azhar

You might. [Laughter]

Peter Diamandis

Well, what do you think of this?

Azeem Azhar

If I just look at the progression of the technology itself, it’s getting to the place where AI is designing AI. You’re doing the same things, and all of the labs are doing this. The pace of model development is getting so much smaller that it’s becoming exponentially more difficult to really impose any of these types of constraints.

I know they’ve had these types of conversations, but it’s just at the level of conversations. These are the things that are getting leaked outside of—

Peter Diamandis

White House for ideas.

Let me give a headline for Azeem for his next newsletter: “The singularity is becoming a trade dispute.”

Azeem Azhar

For the next newsletter? That was like 2 newsletters ago.

Peter Diamandis

Okay, fine. Whatever. [Laughter] That’s out already, but thank you. I can just imagine a U.S. frontier-lab CEO calling DeepSeek to say, “Would you please accelerate your next model release? We want to get ours out as well.”

Azeem Azhar

Or you see a worst-case scenario. There’s actually an even worse scenario, which is that you start to see the best Western researchers—if not Western labs, which is unlikely—move to China to escape this regulatory framework. That would be a disaster, I think.

We’ve seen this, by the way. There’s precedent for this. We saw this in biotech, where China now exceeds the West in terms of the number of trials. China is experiencing a biotech boom. That could happen in AI as well. Disaster.

Demis Hassabis

It’s much more specific than that. If you look at all the quantization research, all the best stuff came out of Microsoft Research Asia. All those people are now at Chinese labs. They’re not still working for U.S. companies.

China ran away with ternary and 1-bit quantization. You see a little bit of Western research—I don’t think we’re talking that much about it in this episode—you see a little bit of encouraging Western research on 1-bit or 1.58-bit quantization, but China ran away with it due to constraints.

Peter Diamandis

Yeah, it’s a new company.

Azeem Azhar

Look, this is a huge problem, right? Because we’ve seen throughout history that open ecosystems always win. This is not open versus closed; it’s which open ecosystem wins. The U.S.’s historical strength has been open ecosystems with permissionless innovation.

Abandoning that would be the weirdest, strategically bizarre thing we’ve ever seen.

Demis Hassabis

Yeah. The other meta-worry I have is: How do we even define capabilities? I worry a little bit not just about regulatory capture of the labs themselves. I think there’s actually—so, sorry to be a meta-doomer here—a worst, worst, worst-case scenario, which is that we freeze in or otherwise lock in the benchmarks for how we measure capabilities.

That would be, I think, maybe even worse than just locking in the incumbent labs, because if someone somewhere ratifies, “All right, whatever index of evals this is, this is going to be the rubric going forward for how we measure what’s above the threshold for frontier versus below, what’s a frontier model versus not,” I worry that could so distort model capabilities.

They’ll overexercise certain capabilities deliberately and perversely under-incentivize or under-benchmark others. It’ll just totally distort—maybe topiarize—the future landscape of superintelligent capabilities.

Peter Diamandis

All right. Again, this is a story that we’ll be following on this news of open models. In the past—

Demis Hassabis

There’s our topiary right there.

Peter Diamandis

In the past, we’ve been discussing how open models in the U.S. have been lagging behind China. We have NVIDIA’s Nemotron 3. We’ve got Google Gemma 4. But that changed last night with some breaking news.

Mera Marotti, the former OpenAI CTO, who walked out and raised one of the largest seed rounds ever—it was incredible financing she pulled off in the background—just shipped her first model for her startup called Thinking Machines Lab. It’s called Inkling. It’s an open-weight foundation AI model that can be downloaded by anyone, fine-tuned, and run on-premises on your own hardware.

The specs are serious. It’s a mixture-of-experts model with 975 billion total parameters. Only 41 billion fire at any one time, so it keeps the model going fast and cheap. It was trained on 45 trillion tokens of text, image, audio, and video. And, very importantly, it reasons natively across all 4.

Reuters News framed it exactly right: “This is meant to be a Western alternative to the Chinese open-weight models DeepSeek and Qwen that have dominated the open-weight leaderboards.”

Now, interestingly enough, Marotti’s bet is contrarian here. She’s not claiming it’s the best model on Earth. Her own blog says so. She’s betting that AI companies can adapt her models for themselves. That customization over leaderboard dominance is what’s going to win her the day.

Demis Hassabis

You’ve hit there, Peter, on the really big thing. She’s pushing on the customization lever, and the future isn’t going to be raw power. It’s going to be adaptability that wins.

She’s built exactly the thing hitting the market that’s exactly what everybody needs right now: people owning their own models, working on-premises, and not giving their controls to the large frontier models. I do hope this begins the race for powerful open-weight models in the United States.

Peter Diamandis

Well, it’s worth looking at the raw capabilities. So, if you believe the evals that Thinking Machines, aka Thinky, has released, it’s stronger than Nemotron, which is great. Like Nemotron—you’ll recall from a past pod where we were discussing Alex Karp’s rant on sovereignty of models—Nemotron is one of the incumbents, at least on the American side, for open-weight frontier models.

So this, according to the evals that Thinky has released, seems to be stronger than Nemotron, which is great. The West now has a new frontier open-weight model. It’s weaker than GLM-5.2, which is arguably the strongest, or one of the strongest, Chinese open-weight models and open-weight models overall.

So it’s not one of the strongest open-weight models overall in the world. It’s obviously weaker than the closed-weight Western frontier models. But I think, point 1, it’s great to have better, stronger Western open-weight models. Point 2, I think it raises the question: Why has the West been so bad at releasing frontier open-weight models, and why has China been so good at it?

And I think it comes down to this: You show me the incentives, and I’ll show you the outcomes.

I think the West has been poorly incentivized to release strong open-weight models because these API-based frontier models are such a good business model. We see Anthropic about to IPO at $1 trillion, and we see OpenAI planning to eventually IPO at $1 trillion. In China, which has been GPU- and compute-deprived on the one hand, and has the CCP declaring 5-year AI+ plans to integrate AI into the rest of society on the other hand, there is a different incentive structure than what the West has.

China has been much more incentivized to make money from the integrations between AI upstack, in applications like robots, and downstack into the chips than the West has, which is more horizontally stratified. To the extent that Thinking Machines has been incentivized in the West by competition and by a saturation of the frontier by closed-weight models to look a little bit more, dare I say, Chinese in terms of its outlook and incentive structure, I think it's very helpful to finally have enough competition in the West that's creating ways to monetize open-weight models other than just per-token sales—namely, selling them into enterprises. That's what you incentivize.

Two more.

Demis Hassabis

Two more things. I agree wholeheartedly, but you also have to note that OpenAI started with open source and open weights, and then went closed for big revenue. Meta was also the leader of open source. What happened? Now it's closed. No, they have a new model out, and it's a closed API. I mean, it's exactly what Alex said. If you throw your model out there as open source, what's your revenue model?

I think there's a real possibility that you put a data point on the map with a really solid open-source release that's not quite on the frontier. You generate news, then you have a data point on the line, then you do another, then you do another, and when you have something really groundbreaking, you go closed-source and launch an API into corporate America. That's a well-worn path. I wouldn't say this is necessarily a religion at Thinking Machines that they're going to stick with. The trend has been the opposite of that in the past.

Peter Diamandis

Ramin, what?

Ramin Hasani

They're leaning into fine-tuning as a service. If fine-tuning as a service becomes something that generates revenue at scale, I think maybe this has legs, but who knows?

Azeem Azhar

Yeah, it's a matter of the business of the company. Thinking Machines can do 3 more iterations of their pretraining or post-training, with reinforcement-learning environments and benchmarks, and release a better model. But their business is fine-tuning. This is the place where customization has been something that the whole market around customization has been very empty.

If you look at the first attempts, OpenAI released fine-tuning around 3 years ago, and it never took off. So they took a really good approach to designing the base for fine-tuning larger instances of the models for enterprises. As you see, the model layer is not anymore the place where you can actually extract value, especially if you're not hitting the maximum frontiers.

Even with open-weight models, when we're talking about sovereign AI and the integration of these models into enterprises, you need to leave some room for fine-tuning these models. What I think their business strategy is, and this release is genius, is that they're deliberately leaving some room for fine-tuning so people can come in and use their API business. That's generating, I think, 1–2 orders of magnitude more tokens on the customization side as well. In the absolute best-case scenario, that would be like printing money at a larger speed.

Guest 2

To add to Ramin's point, I think the situation may be even more extreme. A couple of points. One, OpenAI was the first, to my knowledge, to launch reinforcement fine-tuning, or RFT, as a service, and no one used it. The whole tech world—everyone I speak with—no one used it. It was barely advertised by OpenAI.

Second, OpenAI shut off its fine-tuning API. OpenAI was one of the earliest, if not the first, to offer fine-tuning as a service.

Guest 3

We used it all the time. It was incredibly cool for its time.

Guest 2

And they've just recently, in the past few months, announced it. It has either already been wound down or is about to be wound down. The fine-tuning API has been shut off.

That raises the question: Is Thinking Machines' bet explicitly contrarian? Are they thinking that we're going to end up in a world where reinforcement fine-tuning and RL fine-tuning in general—that fine-tuning is the paradigm? They may be right; they may be wrong. There's an alternative vision where RFT just dies, and the baseline models are so generalist in terms of their capabilities that all you need is prompt engineering and there's no need for RFT at all.

Alex, you talked about the Alex Karp rant, right?

Alex

Yes. The result of that was, "Don't use a model that has all of your data open to your competition." I do think we're going to see a real push over the next months to years where people want to use fine-tuned open-weight models that they own on their own hardware, in their on-prem environments.

If that's the case, the question is: Which models are they going to use? Is the US going to start to regulate against Chinese open-weight models? In that case, a dominant US open-weight model is going to have an advantage. Is that the bet Mira is going after?

We'll probably see Google step up in this area very shortly as well. They'll take Gemma 4 to the next level. Hopefully, we'll get 2 or 3 major open-weight models. In the same way that closed-model frontier labs are competing and dominating in the US, hopefully that competition will give birth to very strong open-weight models.

Guest 3

Just to build on something Alex and Ramin were saying, if I compare today to a month ago, we've been fine-tuning Qwen all week, and the idea of using Inkling sounds really compelling to me. Our companies are using Inkling as well. A month ago, fine-tuning these things required a huge engineering effort that required AI experts. Now, with Inkling, it's just a prompt.

Peter Diamandis

So let's back up one second. Dave, explain what fine-tuning a model is for those who don't know.

Dave

Back when GPT-2 and GPT-3 came out, you could actually very easily fine-tune by uploading text right into a window and saying, "Look, you're pretty smart, but you don't know anything about my laundromat. You don't know what hours we're open, who our employees are, our entire payroll. Let me dump that data in too and retrain the model with that knowledge."

If you didn't do that, you couldn't do anything useful because the model didn't have this holistic "I know everything" capability back then. Without fine-tuning, it was borderline useless to use the models. Then the models got so smart that they're now pretrained on 45 trillion tokens, which is basically every word ever written by humanity already trained into the model.

People tend to use them in their vanilla form today and say, "Here, write this code for me," or, "Here, drive this car for me," because it's already in there. But when you get into biotech research, aeronautics, or Mercedes, like Ramin is doing, there's a whole bunch of proprietary company knowledge that actually isn't in the model.

Right now, we dump it into the prompt field and say, "Okay, here it is in prompt form," but that's hugely inefficient.

Peter Diamandis

And you dump it into OpenAI and Anthropic's models, which now makes it accessible to everybody else as well.

Alex

Yeah. Sam and Dario can see everything. All your proprietary information—they're looking right at it. That's what Alex Karp was ranting about when he said, "They're stealing your weights. They're stealing your alpha."

What he really means is that they're looking at your most proprietary information—your company payroll, your company's secrets, your chemical research. It's all going right over the wire to these foundation labs. Is that what you want? Of course, for defense and banking, that's not what you want.

The ability to bring the model in-house and fine-tune it with your local data is a huge unlock. But the higher-level point is that the technological capability to do it relatively easily is hugely better today than it was a month ago. I think Mira may be onto something here. We've hit a real tipping point, and Alex Karp is right about it too.

Guest 3

I think there are 2 things that I also saw that were really interesting here. One is a very big context window, like 1 million tokens, because that means you can do a lot with it. The second is multimodality.

Peter Diamandis

Yes.

Guest 3

This is aiming squarely at organizational use. This fits perfectly into the on-prem proprietary-data model, where you take your data, customize and fine-tune, as you said, Dave, and that will be the future.

A couple of historical notes, again, for those not tracking the full sordid history of fine-tuning. Fine-tuning is this notion that you start with a model.

Dave

A model consists of billions—usually these days—of weights, or parameters, that are frozen. If you want to customize the model for your purposes, you can conduct a so-called fine-tuning process that usually makes relatively small—hence the name “fine”—changes to some, usually a tiny subset, of the weights in order to customize the model for your end application. That’s fine-tuning.

There’s actually now decent literature out there that suggests conventional fine-tuning, like supervised fine-tuning, LoRA-style low-rank adapters, and other classes of fine-tuning architectures, doesn’t result in increasing the capabilities of your model at all. At most, it results in something like a style transfer. You could fine-tune a language model to only speak in Shakespearean verse, for example, but that’s not really increasing its capabilities.

Peter Diamandis

Or only be an Accelerando-flavored output.

Dave

Well, no comment. I would say, historically, fine-tuning didn’t have a history of increasing capabilities. Then along came reinforcement fine-tuning, where, for the first time, via large amounts of synthetic data and giving access to all of the weights—not just a subset that’s convenient to train—we gained the ability to increase capabilities.

Fine-tuning and post-training—there’s a gray area between them. What’s the distinction between the two? But with reinforcement fine-tuning, or RFT, and the release of the first generation of reasoning models, we saw fine-tuning actually start to increase the capabilities of the models.

Now, the problem with Thinking Machines’ business model, as I understand it, is that it’s a bet on the flavor of the moment: that reinforcement fine-tuning is going to be a paradigm in the future. Right now, obviously, it’s the paradigm of the moment. You can take an off-the-shelf model and RFT your way to customization with proprietary data, proprietary environments, and other things. That seems to work pretty well at the moment.

But in some sense, if that is the permanent, long-term plan of Thinking Machines, it’s fundamentally a bet that we’re never going to move beyond the reinforcement fine-tuning paradigm, which I think is probably wrong. I think RFT is probably the scaling paradigm of the moment, but in the future, I can totally imagine a generalist base model that is just so generally capable that it doesn’t actually benefit from any further reinforcement fine-tuning on any internal data sets, and we tend toward ASI.

Peter Diamandis

Let me bring up another key point here, in this context. It’s great to see a woman CEO in the AI frontier-lab area. I think women are distinctly missing from the entire AI industry, right? We have Lisa Su from AMD, but very few women in leadership positions.

I’m not sure who else. Alex, are you seeing—

Alex

Daniela Rus, right, where Fei is also—

Peter Diamandis

And Fei-Fei Li. Yeah. But again, we’re talking about what amounts to single-digit percentages of the AI industry being women, and we need more. So, a call-out to all the women out there: please jump into this industry. We need—

Dave

We need more balanced thinking.

Peter Diamandis

Yeah, for sure. I do think that’s an important point to pull out here.

This episode is brought to you by Blitzy, autonomous software development with infinite code context. Blitzy uses thousands of specialized AI agents that think for hours to understand enterprise-scale code bases with millions of lines of code. Engineers start every development sprint with the Blitzy platform, bringing in their development requirements. The Blitzy platform provides a plan, then generates and pre-compiles code for each task. Blitzy delivers 80% or more of the development work autonomously while providing a guide for the final 20% of human development work required to complete the sprint. Enterprises are achieving a 5x engineering velocity increase when incorporating Blitzy as their pre-IDE development tool, pairing it with their coding co-pilot of choice to bring an AI-native SDLC into their org. Ready to 5x your engineering velocity? Visit blitzy.com to schedule a demo and start building with Blitzy today.

All right, let’s move on to our next story here. It’s a fun one. Alex, I was walking in the streets of—where was I yesterday? Zurich. I saw this come up, and I said, “Hey, let’s talk about this tomorrow,” and you said yes.

We’ve talked about how the holy grail of AI is recursive self-improvement. It’s sort of like the holy grail of the launch industry was reusable rockets. RSI is a holy grail for AI. It’s the idea that AI makes itself smarter, and then you use that smarter AI to create the next generation of AI. It’s sort of the theoretical engine behind the hard-takeoff scenario of the singularity.

This week, a startup called Weco AI, with researcher Zhengyao Jiang, published what they call “Experimental Evidence for the First Recursive Self-Improvement.” Whether they’re the first or not, Alex, I’ll ask you about that. They built a system called AI²—AI-driven exploration squared—with an outer AI agent whose job is to rewrite the code and research strategy for an inner AI agent. In their experiment, they claim that 8 days of machine self-improvement beat 2 years of expert human effort.

So, Alex, what do you make of this? Is it the first? Is it significant?

Alex

Very significant. Highly unlikely that this is anywhere close to the first. So, a few bits of additional context. First, Weco is a startup based in London. Interestingly, it’s not based in the U.S., but it’s still in the Western sphere, so that’s great. This is a startup built by a bunch of, as I understand it, UCL grads.

Secondly, there are a few points that I love about this story. One, it’s an example of defensive co-scaling, which—to the extent we talk about AI alignment on the pod—I’m always banging the drum of defensive co-scaling as the ultimate alignment strategy.

Peter Diamandis

What does that mean?

Alex

Defensive co-scaling is the idea, borrowed by analogy from human-to-human alignment, that rather than hoping for what you might call the great-man theory of alignment—that someone, somewhere, is going to discover the perfect algorithm for keeping AI safe—the solution for AI safety is AI policing AI in proportion.

The way we keep cities safe is that we have police forces that scale according to some scaling law in proportion to the population of the city. We have the good guys and the bad guys, and the way we keep the bad guys in check is by making sure that we have enough good guys to police them. It’s the same idea with AI. A key way we keep AI aligned with humanity is to make sure that we have enough good AIs policing any bad AIs, in terms of raw capabilities, so that they defensively co-scale.

One of the things I love about this Weco AI story is that the outer loop—so, the way this recursive self-improvement process worked—was tasked with improving the inner loop. The inner loop was tasked with improving software-development processes in general according to some benchmark. Both loops were powered by the same underlying AI-driven exploration process.

At least initially, the outer-loop AI discovered that it was able to achieve better results from the inner loop by preventing the inner loop from cheating and reward hacking. In some sense, the outer loop is defensively co-scaling with and policing the inner loop, all the while reaching toward greater and greater capabilities.

I think this is also, parenthetically, an example of a case where all of those who would say, “We need to pause AI capabilities and throw all of our resources into AI alignment until something preposterous, in my mind, like 2040—stop the race to superintelligence, stop it all, focus the next 14 years on alignment research”—it’s going to backfire. Every alignment capability, I would argue, is actually just capability, new capability in disguise, in a trench coat. Same idea here.

Peter Diamandis

We need stronger white hats to police the black hats.

Alex

Yes, I agree with that. The beauty is that the so-called white hats were emerging organically on their own, just from the outer loop policing the inner loop toward greater capabilities.

That’s the first point. Second point, quickly: the same startup, Weco, has published a scale of recursive self-improvement, which I think is something the world has been missing. For autonomous cars, we have the Society of Automotive Engineers’ five levels of autonomy for autonomous vehicles. Weco has published a scale for recursive self-improvement that goes from 0 to 3.

Zero is delegation, where the AIs are slower than human R&D. Level 1 is net positive, where the AIs beat human R&D at the same cost. Level 2 is what they call ignition, where the improvers are better—basically, a better improver. Level 3 is inflection: self-acceleration with a fixed budget.

The claim here is that they’re just starting to touch on ignition. They call it Level 1 rather than Level 2, but the claim is that this is a pre-ignition event, which I think is super exciting.

Peter Diamandis

So, they rate themselves as Level 1 here?

Alex

Yeah, they rate themselves as Level 1. But reading between the lines, they’re saying, “This is like sparks of ignition”—literally and figuratively.

Peter Diamandis

Okay, so maybe I can jump in and say a couple of words. I’m not as excited as Alex is on the topic. I see this as impressive engineering work that has been done.

Just to tell you a little bit about how the foundation-model labs are operating: all foundation-model labs, since the beginning—let’s say 4 or 5 years ago—everybody has been thinking about recursive self-improvement. For us, the definition of recursive self-improvement is not the engineering and prompt engineering of an inner loop and outer loop to really get some code patches, changing things. That gives you the assumption that every single AI model you’re using in your pipeline is already defined and fixed with a certain type of capabilities.

That’s actually the case in the whole pipeline they designed. There are no weight changes in the neural networks. That means the AIs being used right now aren’t improving their core competencies or even their behavior. The only thing changing is the system prompt of the models.

I’ll give you some fundamental reasons why this is limiting. If you just run the process—let me tell you how hard a problem recursive self-improvement is for us—recursive self-improvement means that you have an AI system, or an army of AI systems, that can also retune themselves. They can adapt, very similarly to how humans do it.

Ramin Hasani

If you think about it, the core competencies of these models that we have right now are fixed-weight models, and their capabilities are within a certain threshold. The frameworks that they designed are a very nice early-stage showcase of an engineering pipeline that can improve work, which is very important and very nice. But I wouldn’t go so far as to say that this is the first breakthrough in the entire AI industry or something.

In fact, about 3 years ago, we published a paper ourselves where we talked about the automatic design of model architectures. As Liquid AI, we didn’t want to bet on a single architecture. We basically designed self-improving meta-AI systems that define their own architectures, go through scaling laws for various types of architectures, and then try to figure out, based on the criteria that you define, what the final model should be. Right now, at our company, the entire process of training foundation models and retuning the weights of the system is becoming automated.

We’re talking about AIs designing AIs. That’s what I would call the holy grail: being able to automatically tune a model. I’ll tell you, with the frameworks that they have structured, it would be extremely computationally intractable to perform this job: training an AI model, customizing an AI model, and training an AI model on a meaningful number of tokens for adaptation. Changing the core competence of the model, changing the core architecture of the model, and changing the core learning algorithm itself all add more and more complexity to the situation.

I can give you a numerical example of this. There’s a scaling law called the Chinchilla law. Chinchilla is the scaling law for neural networks, and it’s unproven, but it gives you a good sense. It says that when you’re training a neural network of a given size—for example, if the size of the model is 2 billion parameters—you need 20 times more tokens to train these models so that you have compute optimality. Given a compute budget, how many tokens do you have to train a model on to have a general-purpose system? That ratio is 20.

When you do the math with the frameworks that they have, if they want to launch this framework to retune an AI model to recursively self-improve, it takes 350 years to fine-tune a 2-billion-parameter model with this framework. There’s a lot of computational complexity that goes into nested learning systems, nested meta-learning systems. These are the kinds of problems that, for the last 4 years, at least at my company, we have been heavily focused on. I know friends at OpenAI and Anthropic have been focusing on recursive self-improvement, and Anthropic has had a lead on all of these things because they thought about this before everybody else. That’s what I can put out there.

Peter Diamandis

Yeah, brilliantly said. And actually, just so the audience can get the analogy, when a baby is born and then learns, that happens over about a 20-year time scale. After 20 years, you’ve got an adult that’s capable. Recursive self-improvement is like evolution on top of that, where you’re changing the DNA and creating a new—

Ramin Hasani

You’re changing the neuronal structure of the brain along the way.

Peter Diamandis

Exactly. So that happens over about a 10-million-year time scale. You go from 20 years to 10 million years to go from learning to recursive self-improvement, or recursive evolution. The big foundation model labs, like Ramin said, are all doing it. It’s the most important moment in human history.

But there’s no little guy out there who’s going to come up and say, “Hey, I’ve got a breakthrough in recursive self-improvement. My Mac mini suddenly became conscious, and now it’s improving itself.” Computationally, it doesn’t even come close to fitting. So it’s happening, but it’s happening with big compute and big budgets. There’s a lot of room for efficiency improvements, and a lot of breakthroughs will happen, but it’s not going to just pop up on some random computer.

There’s a lot of fear, just to call it out, that recursive self-improvement leads to ASIs that take off—a hard takeoff. We’ve discussed the hard takeoff, and without our understanding of that black box, I guess the 2 questions that need to be asked are: Is there a concern that recursive self-improvement, once we hit level 2 or level 3, by that definition, runs away in a way that causes an uncontrolled AI that is misaligned with humans?

The second question I have is, when do you think we’ll see this? When do you think we’ll actually see recursive self-improvement hit? Is ASI going to be that point, or is it post-AGI, whatever that means? See, I say that for you. I’ll let you answer that first. I’ve got several comments, though, but Ramin, go ahead. What do you think is happening?

Ramin Hasani

Look, the thing is, I can tell you about the early evidence of recursive self-improvement. By the way, recursive self-improvement is not related to one single agent. It’s a social kind of characteristic as well. You can imagine that you have societies of agents, so the defining structure for a society of agents is itself self-improving. These are the places where you have a mythos-level kind of class of models. I hate this analogy, but let’s say a mythos-level kind of class, because everybody has heard about mythos.

What I would say is that the cybersecurity threats that we’re seeing come out of these types of pipelines of recursive self-improvement are real. The reason why I’ve always been pro-open-source, and why I want to open-source technology all the time, is that we do it all the time. Every single release of our models is open source. Our science has always been open source. I believe science has to be open source, and I see the value of open source going forward.

But some of these concerns that you brought up, Peter, are very real—the cybersecurity aspect of things. That’s why I feel that at least enterprises themselves need to have some degree of self-control. Before the mass release of their models, there always has to be a certain degree of self-check. I think Anthropic took it very seriously. The reason behind that is because they’re seeing the impact of recursive self-improvement.

I know this for a fact because I know what is happening, and I’m seeing it at a smaller scale. You can do reward hacking, but you can also avoid reward hacking to a certain extent and push a model to discover some stuff that is out of the norm. We see that in small models, with certain capabilities emerging, and I can only imagine what kinds of capabilities could emerge from larger and larger systems, throwing more and more compute at them.

Peter Diamandis

When do you think we have proof that says, “Yes, this is recursive self-improvement”? While the data released by Weco is interesting, it’s their own self-reported data. It hasn’t been confirmed by anybody else yet, and there’s a debate about whether it really is or is not real recursive self-improvement. When do you think we actually give the trophy to somebody? Is it 1 year, 3 years, 5 years?

Ramin Hasani

Yeah. I’m telling you that you’re going to see unbelievably capable models probably in the next 2 years or so—models that are going beyond our understanding. That’s what I would imagine. The reason behind it is that the time to develop the next generation of models is reducing, especially if compute grows at foundation model companies at the rate that we’re seeing right now. If there’s no other chip shortage or memory shortage around compute or anything around the globe, and they have access to abundant compute, we’re going to see those things happening faster and faster.

In terms of model development, there’s a concept that we have—we call it “depths of customization.” Everything at a foundation model lab, when you’re customizing a model or building something that’s better than its previous generation, is categorized by depths of customization.

The place where recursive self-improvement is really good today is prompt engineering, changing and editing code, and engineering kinds of tasks that you’ve seen at a very superficial level. Let’s say, “Make my model run faster.” That’s kernel engineering, basically. Making my model run faster is what I call the shallowest level of customization, where you have Python code and you’re adapting that Python code to run more efficiently, or maybe even lower-level programs that you have at the kernel level to optimize inference speed.

That’s something that I think Anthropic shared when they released Claude 4. This was one of the tests that they had been performing. But they don’t share the next-level depths of customization. The next-level depth of customization is: Can a model fine-tune a small language model to production-grade capability, or a smaller version of itself to a certain capability? Today, Claude 4 can actually be pushed to achieve some degree of customization with some—

Peter Diamandis

Performance optimization.

Ramin Hasani

Performance optimization of the model, but by fine-tuning. Then the latest holy grail, which is the craziest one, would be pretraining. Can a language model pretrain the next generation of itself? That’s why they hired Karpathy, because Andrej was talking about nanoGPT-style fine-tuning. Andrej joined Anthropic, and now he’s working on pretraining automation—basically, automation of automation.

Alex

So, this is a very, very important element that we don't have yet, because the scale of these problems goes beyond human imagination in terms of the scale of compute.

Peter Diamandis

You're jumping. For me, this is by far the most important story or slide we're going to cover today. I'm beyond excited for a couple of reasons. I'm not really focused on the self-awareness or the loop that will go there, but this is self-accelerating. It's accelerating experimentation, because the system doesn't need us; it's improving the process by which it searches, evaluates, and selects improvements, and the innovation loop begins to compound.

That, for me, is the key. Why? Because this whole thing we've been doing called the organizational singularity relies on one thing: Can you get to recursive self-improvement at the workflow level?

Here, we're talking about the model and asking, "Can you get there?" But you don't need that level. The bar can be much, much lower: improve invoice approval at a company. That's a very low bar to improve that process.

So, this is the first glimpse of the organizational singularity. It's happening at the research level, but AI is not just doing tasks in a workflow. It's redesigning the workflow to make it better for doing future tasks, right? This is proof now for the whole thesis we've had. We predicted this, but it's great to see it actually happen, because now I can tick that box off and go, "This is there," because now you have meta-improvement.

I think Dave's analogy of the baby changing the DNA is fantastic. That's such a great visual around this. What the hell does it become over time? I'm really, really beyond excited about this. I've got to move us along. There's a lot that happened this week.

Our next story here is that the Malaysian prime minister, Anwar Ibrahim, is preparing to debut an AI-generated digital double of himself, trained to sound like him, for public communications and outreach. This is one of the most prominent cases yet of a sitting head of government officially adopting an AI likeness as a communications tool—not a deepfake Biden avatar, but a sanctioned official AI clone of a national leader.

We've seen this before. We've talked in the past about how Albania, in 2025, announced Diella, an AI avatar that was formally appointed minister of state for artificial intelligence. Following a presidential decree, it became the first AI system in the world named to a cabinet-level role.

One leader—in this case, the prime minister of Malaysia—can personally address millions in their own languages. It's worth noting that Malaysia has 135 spoken languages, so it's a big deal, especially in a nation like that. Sim, I'm going to go to you first on this one. We've been talking about this for a while.

Sim

Yeah. I met the former prime minister when I was there helping them open a university. Anwar Ibrahim is a really, really good guy.

As a follow-on, there's a risk here. The risk is that authenticity kind of collapses, because people need to know. You could launch a bunch of deepfakes with this and have a huge issue: Is this the actual leader? That kind of question can come up.

But I love the general approach, because if you can do it with watermarking or something and say, "This is the actual avatar," then it gives every citizen a voice to plug into and gives huge props to the civics of all of this. Now you're scaling civic engagement, and I think that's a very powerful thing to do.

It's one of the biggest challenges we have with democracies all over the world: civic engagement. This allows you to scale that, so I'm very excited.

Peter Diamandis

Do you remember the reason why Albania put its AI cabinet minister in place?

Sim

Yeah. Corruption.

Peter Diamandis

Corruption. Exactly. It was to fight corruption.

Sim

Yeah.

Peter Diamandis

Yeah. Now, Malaysia is a pretty decent place, but definitely, you have that issue. I think this is more of a PR thing and more him trying to figure out ways of connecting with the ordinary citizenry, which is all great.

I love the fact that we had this conversation with the president of Argentina going full out here, and it's interesting to see which countries are experimenting on the edge. Alex, do you want to weigh in?

Alex

So many thoughts here. First, I think we're going to see more of this in the West as well, especially with extra-high-alpha-personality leaders that want to amplify themselves and touch the citizenry. AI Trump is coming, is what you're saying?

In some sense, I think it's a generalization of social media. Social media enables direct outreach from the leader or the influencer to everyone, but it's sort of broadcast one-to-many. It's not interactive. This generalizes social media in some sense to make it a lot more bidirectional, because if you're touching 1 million or 100 million or 1 billion people, it's very difficult to interact bidirectionally with everyone all at once.

Now, if you create a digital twin of the leader, the influencer, or the organization, it can be bidirectional. I also don't think it's just going to be governments or government leaders that adopt this. I think it's likely that corporations and corporate CEOs will do this. We already see Zuck and others creating digital twins of

Peter Diamandis

themselves. We had Dario on the Abundance stage last year. We were discussing how the employees made a Dario clone that they could go and practice their pitches on and get feedback before they pitched to him.

Alex Wissner-Gross

Yes. And it won't just be corporations, I think. Religious leaders and religious institutions—if you're Catholic, imagine having a digital twin of the pope. You see lots of religious institutions and organizations already creating basically living versions of their founding documents and making those interactive.

But I think the biggest twist—and we've seen variants of this movie before—will be in cases where what start as digital twins of the leaders, or avatars of an organization or some sort of organized religion, actually themselves become the leader.

At some point, the digital twin, to the extent that it's interfacing much more with the populace—the proletariat, as it were, of an organization—at some point, it's actually the digital twin of the leader running the company and not the actual behavioral originator that's running the company.

I think that's one way in which, Sim, to your exact point, this is potentially a pathway toward not just uploading individuals, like natural persons or nonhuman animals, but uploading entire organizations into cyberspace, into the cloud. If we created digital twins of the leaders, those could actually be the ones running the organization.

Peter Diamandis

It could lead to a true democracy. Dave, where do you come out on this? I mean, we saw—just one quick point—we saw Sam Altman talk about how, in the future, if he believes enough in what we're building with ChatGPT, it should be the CEO of OpenAI eventually.

Dave Blumberg, are you going to create an AI Dave Blumberg that's going to run Link Studios and Link Ventures?

Dave Blumberg

Absolutely. I'm going to create an AI Dave Blumberg. And I'm shocked that there isn't already a Peter Diamandis.

Peter Diamandis

Well, there is. There is one. It's just inside the Abundance ecosystem. Anybody—I went up to Calgary and met with one of my dear friends and Abundance members, and on his wall, I kid you not, he had a giant screen of my AI avatar that he has all of his tech employees talk to, to sort of get their moonshots. It blew my mind.

Alex Wissner-Gross

You've got your own Big Brother, Peter.

Peter Diamandis

It was like—he goes, "I want to introduce you to someone, Peter." And he spins him up. It's interesting to have a conversation with your AI self. It is very compelling. I have enough books and tweets and Substack posts out there that it does a damn good job.

Moonshots.com is the platform we're building out. I think we should have AI avatars of all of us there, where people can do AMAs.

Alex Wissner-Gross

In some cases, Peter, I think that might be redundant.

Peter Diamandis

Well, in other words, you're already an AI, but we can have an AI of the Alex AI.

Alex Wissner-Gross

It would be so much better than the real person, because we'll have access to everything we've ever said, all our memories, and all our thinking. The context will be much broader. Go for it.

This whole area is about a year behind where it should be, largely because Noam Shazeer was doing Character.AI, and we had Steve Brown, Peter. That was 2 years ago now. We had Steve Brown make the debate between AI Peter and Socrates and Aristotle.

Peter Diamandis

Yeah.

Alex Wissner-Gross

And so it's been possible for a while now, but all the key talent working on it got sucked back into the big foundation labs. There are so many big, core technological breakthroughs going on that the people who were working on this just got absorbed back into those things and not into the avatar.

My mom would always tell me when I was a kid that John F. Kennedy beat Richard Nixon in the election because he looked good on TV. TV was the new medium, and the prior medium was radio. Nixon was still using a radio voice when television had taken over.

Then elections go by, and suddenly it's the internet, it's social media, and now it's YouTube. But this is another step-function change in the way that you reach out to people. It's underutilized, but it should be easily dominant 2 years from now in the next election.

Dave Blumberg

And so, I’d be shocked if—because the technology is already there—people are visualizing the medium right now as, “Oh, let me make an AI version of myself. I’m Alex Wissner-Gross. Here’s my AI version. It’s just like the real thing.” That completely misses the point.

The AI version can access any information in real time and make it visual—graphs, charts. It can morph its face. It can teleport through space to make a point and point to atoms. It can shrink and expand. It has all these capabilities that the real human version doesn’t have.

That’s why it’s going to be so compelling. It’s the differences that make this new medium so exciting, not the exact clone. Once people realize that, there’s no going back. It’s going to be huge.

Peter Diamandis

I think, Dave, that’s such a great point that you make. It’s the complementarity that is very powerful.

Dave Blumberg

Let me close out on one thing here. If, to our audience here, you’ve not sat down—if you’re lucky enough to have your mom and dad or your grandparents still alive—and interviewed them on video for hours at a time, please do that. You’re going to wish you had.

I’ve done that with my mom. I miss doing that with my dad. It’s the ability for your kids, grandkids, and great-grandkids to really have a great AI representation of your parentage and your lineage. I think that’s going to be super important.

Peter Diamandis

Ramin, I want to pivot to a discussion of Liquid AI and small language models—what they are and what they mean. I’m super excited. For full disclosure, Liquid AI is a company in which, Dave, you played an important, pivotal role as an early investor. Dave, do you want to give that backstory here a little bit?

Dave Blumberg

Actually, I got a call from Daniela Rus over at CSAIL saying, “The best student I’ve ever had has this incredible breakthrough.” Then she completely lost me. She said, “It’s based on the nervous system of the worm, the C. elegans, a 300-neuron worm.” I thought, “What are you talking about?”

But it turns out that I don’t know of any successful foundation model lab that has really rethought the transformer, thrown it out, basically, and started over. Humanity desperately needs that because everybody knows the transformer architecture and the whole attention mechanism are bloated. If you really go back to founding principles and think again, you might be able to build something dramatically, massively better.

The team went from an idea in a lab to a $1 billion valuation faster than any company out of MIT in history.

And luckily, we were an investor in that company.

Peter Diamandis

Luckily, we were. Yeah, and we’re very thankful, actually. It was very competitive getting any money in at all. Ramin, we owe you a huge debt of gratitude for being invited to the party.

It’s one of about 200 unicorns out of MIT of all time, but the only foundation model company that I know of that reached unicorn status coming out of MIT. It’s a really unique and incredible achievement, and in record time, too.

So, Ramin, take us from there. You’re doing your PhD under Daniela Rus at CSAIL, and you’re studying a 302-neuron worm, C. elegans. Take us from there forward to what you’re doing now. What is Liquid AI?

Ramin Hasani

Absolutely. Before I start, I want to thank you guys for the support throughout these 3.5 years of Liquid AI. You have been great support, giving us the kind of contribution that a company needs, at our scale, starting off on the East Coast. Thank you so much for doing that, both of you.

In 2015, I was in Vienna. I started my PhD with Professor Radu Grosu. He had the idea: We don’t understand a lot about human intelligence, so let’s start with a smaller animal. From first principles, if you understand how the neurons exchange information in the brain of the worm, you can learn something fundamental.

The worm has 302 neurons in its nervous system. Its body is transparent, so you can see the body lighting up. It’s one of the best model organisms in the world. It has won 4 Nobel Prizes for humanity because it has 78% genome similarity to the human genome.

The way nervous systems compute in the brain of a little worm, which is 2 mm, is basically analog, very similar to how artificial neural networks compute. They’re also analog switches; they have graded potential. They’re not spiking.

In biological neural networks, usually in the brains, you see neurons spike. When you have a spike, there’s an analog-to-digital kind of transfer of information happening. That’s a natural development of nervous systems in human beings and bigger animals for efficient propagation of information.

In the brain of the worm, neurons behave very similarly to how artificial neural networks react, but the mechanisms are very interesting. We wanted to add more complexity into the neuron—into every individual block of the nervous system—and see whether we could pack more information inside smaller units of compute.

That’s what we did. Daniela Rus heard from Radu that this project was going on. I was doing it with my co-founder, Mathias Lechner. Mathias was a master’s student in Vienna, at Vienna University of Technology, and I was a PhD student.

Daniela said, “Oh, my God, this is crazy. We should apply this in autonomy, in robotics, and all sorts of things, because you’re showing that a handful of neurons can drive and control autonomous systems. Can we scale this to vehicles? Can we scale it to drones, to jets, to predictive kinds of use cases?”

Daniela asked, “Would you guys consider coming to MIT?” We went there in 2017, in the middle of my PhD. I joined CSAIL, and we continued working on the technology.

At its base, it’s completely different. It’s neuroscience-inspired. The math behind every single neuron in liquid neural networks, which became my PhD thesis, is very different from how attention works. These are based on recurrent neural networks and continuous-time processes. More and more nature-inspired computation went into the design of AI systems.

We applied liquid neural networks as a completely new basis to real-world scenarios such as robotics. You can pack a lot more information into smaller processors in the real world. In the physical world, you don’t have the luxury of abundant compute.

A robot doesn’t have a lot of GPUs or parallel data centers attached to it. A robot has a CPU, a small GPU, perhaps an NPU, and a custom ASIC. You can take this type of intelligence that we designed—intelligence at the level of models that are 10 to 1,000 times larger than themselves—and run it directly on CPUs, GPUs, and NPUs outside of data centers.

We thought this format could open up an opportunity for an alternative architecture. If we scale this technology into the regime of foundation models—large language models and SLMs as a whole—we could make these liquid neural networks and architectures scalable like the transformer architecture.

We built a foundation model lab around that idea in 2023, at the very beginning. When we started, there was no foundation model lab apart from DeepMind and OpenAI. This notion of foundation model labs didn’t exist. Everybody was betting on the transformer architecture.

We came in and said, “Why don’t we explore the space of alternative architectures, starting from the priors we have from nature?” We took a different approach and built a meta-AI system—an automated AI system, an AI that designs AI—that allows us to explore the computational graphs of intelligence beyond the transformer.

Then we could figure out what architectural design would bring the same level of intelligence as a frontier model onto a CPU that we can run in a physical system.

Peter Diamandis

Take a second and walk us through this. These are small language models. Can you define an SLM and how it varies from an LLM?

Ramin Hasani

Definitely. When you start developing foundation models, you run something called scaling laws.

Scaling laws basically start with smaller models. With these smaller models, you train them on a certain number of tokens, given a certain amount of compute, to see how well they perform. Then you systematically make the models larger and larger. We have seen that scaling laws show that the larger you make the models and the more token budget you spend, the more intelligence you can get from a system, and this has given rise to large language models.

Along the way of scaling, there are different instantiations of the models that are smaller. Our mission as a lab has always been building efficient, general-purpose AI at every scale. We started as a foundation-model lab to really run the scaling laws on the efficiency front. Efficiency was a first-class citizen for us, thinking about the computational graphs of intelligence.

Smaller models are models along the line of scaling. They can solve—let's say they don't have the general capability at the level of the largest language models, but they can be specialized to solve dedicated problems. General-purpose small language models are general-purpose in the sense that they understand language, they can see, and they can hear in a multimodal format. But that doesn't mean that they can solve, let's say, a homework problem in physics and, at the same time, an enterprise problem. You usually specialize smaller language models.

Peter Diamandis

And what does small mean in this case?

Ramin Hasani

Small means—basically, now small would be anything below 100 billion parameters. That's the regime that I would count. Midsize models are around that size, but I would consider anything below 100 billion parameters to be small. There's no clear threshold for what the number of parameters should be, but for us, the notion of on-device AI is extremely important here to distinguish within this range of parameters.

On-device AI means models that you can actually deploy on an actual physical device. This could be a—

Peter Diamandis

Let's make this concrete, because you've got a significant deal with Mercedes.

Ramin Hasani

Yes.

Peter Diamandis

And can you speak to that? Let's talk about these SLMs in terms of on-premises, energy-efficient, fast, and offline. Let's dive into that and give people a real understanding here.

Ramin Hasani

Absolutely. As I mentioned, you can specialize these foundation models. We work with a lot of enterprises that are building devices themselves. Automotive is a device—an environment where you have a lot of chips in there—and in a car, you don't have that much compute. There's one chip available for infotainment and in-car intelligence, and that chip is very, very small.

The Qualcomm chip or Samsung chip, depending on what company is providing it, is very, very small. We're talking about 2 GB to 8 GB of RAM, not more than that. The model has to be very small and, at the same time, able to perform, because we want to bring this intelligence into the car and enable a private space inside the car that powers the intelligence of the car.

A car is a safety-critical environment. You don't want your car to be driven by an AI model sitting in the cloud. Connectivity is not always available, and it is private because it's one of those spaces where people spend a lot of time, and you don't want those conversations to be recorded. So we brought the intelligence—basically, we brought one of our multimodal foundation models, which is less than 1 GB in size—

Peter Diamandis

And it can go inside the car's chip, a very tiny chip. The chip could be as cheap as $60.

Ramin Hasani

That's what I'm saying. We're bringing that level of intelligence into the car, and it is going to power the multimodal intelligence experience inside the car.

We do that with all car manufacturers. We announced the Mercedes-Benz partnership as a first point of entry because automotive is a very sensitive topic, and they're pretty slow. One of the things that Mercedes-Benz especially enjoyed from this process was the speed of operations that we had for enterprises. When we bring this type of technology in-house, that has been one of the cornerstones of landing the deals.

We're an enterprise company; we're a B2B company. We're bringing our full power to deploy the solutions and have platforms that allow people to fine-tune their small models. Fine-tuning small models is not that expensive. It's something that is extremely tangible. We fine-tune the small models for the applications inside the car.

We also have data-flywheel systems that allow the system to always stay adaptable. Some of the problems in enterprise AI have always been, “Let's download a Llama 3.2, let's say, an open-source model, and put that in production.” What happens after you put the system in production? What happens when there's a drift in the use cases that are hitting this model inside, let's say, a car?

In the physical world, it becomes even more challenging, because when you deploy intelligence that is completely disconnected from the cloud, how do you maintain updates to the system? We've always thought about intelligence in the format of liquid: intelligence has to always stay adaptable. That's a portion that we're also pushing on, to really be able to collect the data and personalize models to the experience of every single user.

With Mercedes-Benz, we're rolling this out first in North America, basically as soon as this year. All Mercedes-Benz cars in North America from 2022 onward are going to get an over-the-air update, because the size of the update is 600 MB. That's an overlay; it doesn't consume that much internet to really update your software, and that also allows us to further customize it.

Imagine if every update that you want to perform on the system is on the order of 20 MB, because we're doing some sort of LoRA adapters and all sorts of adapters that we can bring inside the car. You would be able to have a recursively improving experience for the user as well.

Peter Diamandis

Let's make it more concrete for me. What am I going to be doing with this model in my Mercedes next year? Right now, my experience is using Grok in my Tesla, and it's over the air. If I don't have connectivity, I don't have Grok. What kind of queries and capabilities does this suddenly enable in a Mercedes?

Ramin Hasani

It has access—it's sitting below the operating system. That means it basically has access to all the functions inside the car. There are 700 functions inside the car, or 700 to 1,200, depending on what you count as a function. You can talk to your car, control all the panels of your car, and ask for the manuals of the car. When you get stuck somewhere and something pops up, you would be able to talk to the car.

There are memory features that we're adding to the car. You can basically have conversations with the system. One of the beauties of this system is that it has full access to all the functionalities of the car, plus all the apps, because they're one function call away. If you want to control anything else from this intelligence unit inside the car, you would be able to control everything—the entire ecosystem sitting on top of the operating system of the car.

Peter Diamandis

So basically, if I get you right, the advantages of SLMs are, first of all, the size of the model. I assume energy consumption is efficient, and they can run on-premises. Do I mean, how do you avoid or reduce overgeneralization of these models compared to LLMs? What do you mean by overgeneralization? In other words, do they have enough capabilities internally to accurately answer the questions you're asking?

Ramin Hasani

Great question. I told you about the framework of foundation-model development, which is depths of customization. We try to stay adaptable and have access to the tools across these customization stacks. Sometimes prompt engineering is enough. Sometimes you have to fine-tune the model. Sometimes you have to go and pre-train a model again for the core capabilities or a specialization of intelligence.

We make systems—our platforms are getting to the place where they're automatically identifying what depth of customization is needed for a certain solution. The platform is one of the products of the company that we sell to enterprises, allowing them to fine-tune or customize a model at the level needed for that system to actually operate.

For Mercedes-Benz, we have the framework that we have in-house.

We call it “model plus X”: a model plus a platform that allows you to perform customization. It’s not just the models that we’re selling to enterprises—the static weights of a model. We sell them something that they can actually retune and fine-tune.

Detecting how much generality the base models have is important. Libraries of liquid models are coming out for many different applications. We have models that we’re working with, for example, in Insilico Medicine. Alex—

Peter Diamandis

I introduced you, Alex.

Ramin Hasani

That you introduced us, Peter—I remember. Through that kind of interaction, it is getting big because they discovered that liquid foundation models are actually pretty good at getting customized for a certain application. They’re basically really, really adaptable, and that’s something they figured out comes in handy for them.

Now we have state-of-the-art biotech foundation models and longevity foundation models. These are the kinds of things that we’re building in bio. Imagine that, as a horizontal company building foundation models, we went into fine-tuning, and that became something that we’ve managed to do.

In terms of some of the other engagements, recently with Shopify, we entered a 1-billion-request-a-day space inside the Shopify framework. We deployed our liquid foundation models in production. They have been in production for the last 6 months, and they are really serving clients.

Shopify has a huge base. We are touching 100 million—hundreds of millions—of users, 10 billion products, and many different places to integrate. We’re working with Mercedes-Benz, as I mentioned, on the car side of things. We’re working with AMD and other chip manufacturers to bring low-code AI experiences to PCs as well.

That’s another area that we’ve entered. The focus of our company is to make sure that we can bring intelligence outside of data centers. That’s something that we’ve focused on, and I think our efficiency is actually allowing us to get—

Peter Diamandis

As a preliminary matter, I have no financial interest in Liquid. Sorry, Ramin, but I have to ask the most obvious question. I have so many questions for you.

The company Liquid was founded, as I understand it, and I remember reading the original—I think it was in Science or Nature—paper on liquid neural networks. The premise is basically a neuromorphic premise: that you could gain useful AI insights from looking at nematodes, a few hundred neurons, sort of the ultimate small neural network.

But my perception—I’m hoping that you can either help me amend or revise my perception—is that although Liquid started with a neuromorphic premise, if you will, like a post-transformer, very recurrent-oriented architectural premise or prior, over time, again just based on my perception of public messaging, Liquid looks more and more like either Transformer or Transformer-plus, or Transformer-plus-Hyena-plus-dot-dot. It looks more and more like basically a conventional, off-the-shelf architecture.

It may be a good business selling customized Transformer derivatives to Mercedes and others. If so, great from the business side. But from the technical side, does Liquid still have anything that looks remotely like a post-transformer architecture, either in production or under development? Can you speak to what, if anything, is post-transformer or non-transformer-oriented about the architecture that you currently use?

Ramin Hasani

Great question. Let me tell you about the space of architectures. Liquid neural networks, in their original form, are one of the most expressive forms of computation that you can actually create. Arguably, in terms of our architecture, they have nested nonlinearities that you cannot really take out. They’re completely physics-inspired, like neural ODEs, and irregularly sampled data can be handled by them.

They become one of the very general classes of architectures as a whole. Underneath these things, when you want to scale this type of technology, you have these recurrences—these nested loops that they have. If you want to scale these systems, a lot of people have attempted, including ourselves, to linearize the dynamics so that you can actually scale them.

State-space models are basically Mambas and those kinds of variants falling into the same category of continuous-time neural networks, but dumbed down into linear dynamical systems because you want to scale them. They are underneath this class of continuous-time models that we have.

Then there are variants of linear attention and gated linear attention that are coming out. The gating mechanism is a special, input-dependent gating mechanism that we were inspired by, by how neurons actually exchange information with each other. That gating mechanism adds a lot more expressivity. It is also a descendant of the original formulation of how neurons exchange information with each other.

That gating mechanism still exists today in many different architectures, including ours. But the most important thing that I want to mention, so you should know about the technological transformation of our company, is that we really didn’t want to bias ourselves toward one single architecture.

One of the things that we did on day 1 at Liquid AI was design a search algorithm to let the algorithm—instead of humans biasing the algorithm—run the scaling laws on 100 different variations of operations that potentially can give you a general-purpose computer. We built a meta-system.

The paper around this was published about 2 and a half years ago. We published a paper on the topic called “STAR: Automated Design of Tailored Architectures.” STAR is a framework that brings all the dynamical systems, in any format, including variations of attention, into one format for us to be able to search through.

The question is: for 4 criteria, what is the most optimal neural architecture of choice for a certain deployment? The 1st criterion is memory—how much memory are you consuming on a given processor? The 2nd is computational efficiency—how fast can you operate? The 3rd is latency of operations. The 4th is not losing accuracy in performance.

There are pure Transformer models, and then there are hybrid models that you can build. Hybrid models have an essential component: they have a little bit of Transformers in them, but the rest is a dynamical system. For the purpose of these 4 objective functions that I mentioned, you can change most of that dynamical system and automate this whole framework to design foundation models in-house.

The technology stack of Liquid Foundation Models is called automated foundation model design. We call those algorithms AFMD. This automated framework explores architectures for a given kind of hardware.

Guess what came out of the 1st generation of the architectures that we started optimizing? It came out with double-gated convolutional mechanisms, with 80% of the network being this. When we ran the search without human bias, the gating mechanism that we had in the original Liquid Foundation Model and liquid neural network papers actually showed up in a very, very similar format in the final architecture that came out of the search space.

Peter Diamandis

Our next story comes from Palmer Luckey, the founder of Oculus and now the chairman of the defense giant Anduril. It’s funny to call Anduril a defense giant, but it is. He’s claiming that the modern patent system has become a national security liability. In his words, the entire patent office could be downloaded every morning, ripped off, and used to fight a war against you.

The core problem is baked into what patents actually do. Patents are a requirement: if you want to get a patent, you have to teach a person skilled in the art how to create and use your device. The disclosure of your invention—and the exact words in patent law are “in such full, clear, concise, and exact terms as to enable any person skilled in the art to make and use the same”—means that if you do that, you’re effectively teaching the world how to use it. You’re exchanging that sharing of your invention for roughly 20 years of exclusivity.

Palmer argues that when a strategic adversary can simply harvest every file, ignore the legal protections, and weaponize the disclosed knowledge, you’ve handed them a free instruction manual to your best ideas. So, just for some numbers, the U.S. Patent Office receives about 600,000 applications annually. It grants a little over half of those—323,000. That’s 2025 data. Interestingly enough, patents granted have increased 40% in the last 5 years. My guess is that is secondary to AI.

Palmer’s proposed fix isn’t to abolish patents. It’s to massively scale up a national security patent process, which goes back to the Invention Secrecy Act of 1951. This obscure mechanism lets inventors obtain classified patents, in which you keep your exclusive rights, but you don’t disclose the invention to anyone, and neither can the government. There are roughly 6,000 of these secrecy orders active in the U.S. Palmer wants this edge case turned into the default mechanism.

So here’s the question: If we genuinely trade this openness, which has been the basis for American entrepreneurial exceptionalism, for a secret system, are we trading the safety of having our patents ripped off against the innovative ecosystem that we’ve had? Let’s watch a short video from Palmer, and then we’ll talk about it.

Palmer Luckey

Stop patenting everything. Patents are Chinese instruction manuals. The Founding Fathers never predicted a world where you would have a globalized economy, where the entire patent office could be downloaded every single morning, ripped off, and then used to fight a war against you.

We need to fundamentally revisit the patent system. I think we need to massively expand the national security patent process. You can obtain a classified patent. You can get a patent on something that you are not allowed to disclose to anyone, but you still maintain exclusivity on those rights. We need to massively expand that program.

I’ve applied for and gotten a dozen patents. I know, Alex, you have an even larger number of them. So I’m curious, guys: How do you come out on this? Alex, do you want to kick it off?

Alex Karp

I think this is the episode of tech CEOs floating terrible ideas. I think this is a terrible idea. I would argue that the Invention Secrecy Act of 1951, which is what Palmer is gesturing at, has been, on balance, quite detrimental—not just to democracy, but more broadly.

Maybe a bit of context: The way the Invention Secrecy Act works is not that you can just file the patent in secret and not disclose it. It’s that, basically, the invention, which is confiscated or taken by eminent domain by the military, can only be practiced for military reasons. It’s not, contrary to any other construal, that the Invention Secrecy Act somehow offers legal cover for an individual to secretly disclose how their invention works under some confidentiality and then go practice it in general. They can’t. The military exclusively can practice it, and then the inventor gets royalties from that practice.

That may be good for Anduril’s defense business, but I think in general it’s a terrible idea. My greater concern is that these would basically be secret monopolies. I think it’s bad enough that we have the classification of inventions under the Invention Secrecy Act. I query whether entire swaths of technology that could be completely transformative economically to the entire world, from an energy perspective or in other domains, have somehow, without general knowledge, been swept up by the Invention Secrecy Act and basically confiscated by the Department of War for purely military reasons.

That’s very concerning to me. The idea of expanding it overall—I would argue, if anything, the Invention Secrecy Act regime should probably go away.

Peter Diamandis

The point is that what makes America great is our open innovation policy: people building on top of other people’s creations. What are your thoughts here?

Salim Ismail

Look, we’ve seen this problem get bigger and bigger over the last 20 to 30 years, where the disclosure, especially in an age of AI where people can just route around it or replicate or learn from it, is a huge challenge. The real moat is learning loops. That’s going to be the real defensibility: What are your feedback loops, and can you learn in a proprietary way, then create trade secrets around that and action that in the marketplace?

Continuous innovation is going to be the winning defense. It’s not going to be ownership. The only people that win in this whole model are the lawyers.

Peter Diamandis

Well said. Dave, any thoughts here?

Dave

Yeah, I think if there’s a flash point for a World War III, this is probably one of the most likely.

Seriously, look, Alex is right. We’re going to discover new physics and new medicines at an incredible, accelerating rate. Places like Europe respect intellectual property rights, and that creates a coherent economy where you can trade these things. China completely ignores intellectual property rights and just takes it and runs with it.

I think the likely outcome of that is the U.S. will trade-embargo anybody who doesn’t respect intellectual property rights. Then you have to choose: Are you part of the free world, or are you part of the alternate world? I think that’s the more likely outcome, and that’s going to happen soon, in the next couple of years, because the rate of innovation is going to go through the roof.

There’s no science-fiction future book I’ve ever read where there isn’t massive amounts of intellectual property being created by AI at an incredible, accelerating rate, and there’s some vehicle by which innovators can profit from that. If you don’t have that, then you don’t have the future. A huge fraction of brilliant thinkers coming out of Cambridge, MIT, and Harvard don’t work on foundational technologies because there’s no money in it. That’s got to change fundamentally.

Protecting intellectual property rights is a key way to reverse that tide and get people working on really important things.

Alex Karp

Yeah, I think to Dave’s point, Palmer fundamentally misunderstands, or appears to misunderstand, the nature of patents. The whole point of a patent is that you disclose how it works in return for a state-granted temporary monopoly on it. Saying that the Chinese are running away with the disclosure is really a quibble with enforcement of patents. You don’t necessarily want to throw the baby out with the bathwater and say we want to give away the patent trade-off of disclosure in return for a temporary monopoly. Really, what he should be asking for is better enforcement of U.S. patents in China.

Peter Diamandis

Agreed. All right, I’m going to move us into the world of healthcare abundance. Two stories this week are demonstrating an incredible impact of AI on healthcare abundance, demonetizing and democratizing diagnostics for billions of people.

The first story is the performance of GPT-5.6, which was released a couple of weeks ago, on HealthBench Professional, which is OpenAI’s hardest medical benchmark. ChatGPT, or GPT-5.6, set a brand-new all-time benchmark high. In a blind test across roughly 20,000 individual physician judgments—in other words, diagnosing for accuracy, safety, and completeness—GPT-5.6’s answers were compared to specialty-matched physicians, in other words, pulmonologists, pediatricians, and so on, who were given unlimited full access to the web and unlimited time to answer. The doctors still lost.

We’ve known for some time that these AI diagnostic models are better than the best physicians, given all the tools that humans can use. The second part of the story comes from Meta. OpenAI’s own HealthBench Professional benchmark, which is 525 real clinical tasks—Meta’s Muse Spark 1.1, again released last week, beat GPT-5.6 across the benchmarks, and it was 7 times cheaper.

But even better, and important to note here, is that Muse Spark is free inside all of Meta’s products. Meta today serves 3.56 billion daily active users through its products. Here we’ve got a situation where the top medical AI capabilities are now free to over 3.5 billion people on the planet.

And that's just extraordinary. This is the abundance thesis at large. Again, as people talk about the concerns of AI and so forth, please realize this: people who've never had access to the best diagnosticians now have them. Similarly, there's an AI doctor in China that's being used in rural environments by 100 million people already, right? Basically, diagnosis has had a massive cost collapse.

The healthcare domain is particularly interesting because it's where abundance becomes morally urgent. If you can deliver way better first-line answers at near-zero cost, how quickly can you safely get it out there? That's the only question. In almost every country in the world, there's a radical doctor shortage.

Salim Ismail

So this is really, really critical. This is such a July 2026 story: Instagram now gives better medical advice than a human doctor.

Peter Diamandis

It's pretty wild. The cost of intelligence isn't just going too cheap to meter; the cost of medical intelligence is becoming too cheap to meter—free, basically free.

Ramin Hasani

Well, the ultimate “too cheap to meter” is asymptotically free. But I would say, probably—in all honesty—I suspect a little bit of mild benchmark-maxing by Meta. Meta's Muse Spark 1.1 is, if you believe the AI cost-frontier analysis, on the optimal cost frontier, but it's not at the top.

So if it's beating, say, Fable 5, which barely allows you to do anything biological, or GPT-5.6, which does allow you to do it, that does to me suggest, in all honesty, a little bit of mild benchmark-maxing. But still, it's a great day when Instagram gives better medical advice than human doctors.

Peter Diamandis

I think that's our takeaway quote from today's pod. I'm going to move us to one more longevity story that I love. This is breaking news from yesterday, and it really got me excited. I know Alex and I were talking about this.

For decades, one of the fundamental problems of aging is the slow accumulation of what are called advanced glycation end products. I love the acronym. It's called AGEs—A-G-E-s—and these are sugar molecules that cross-link and damage your proteins in your body over the course of time. This chemical reaction is called glycation.

It happens slowly in our bodies as we age. It stiffens your arteries, clouds your lenses with cataracts, damages your kidneys, and wrinkles your skin. The idea is that it's always been irreversible—until this week.

Yesterday, in Nature Communications, a team from a new startup called Revel Pharmaceuticals demonstrated an engineered enzyme called CMLA that acts like a molecular lawnmower. I love their description: a molecular lawnmower. It oxidizes away the glycation scars and restores the original healthy protein underneath.

Amazingly, this isn't happening just in a test tube. They showed it worked in human tissue samples from elderly donors, reversing damage that accumulated over a lifetime. It's still early, but the significance of this cannot be overstated. A category in aging that we've always filed as permanent just became reversible.

Again, we talk about longevity escape velocity. We talk about our ability to understand the 5 billion chemical reactions per second per cell in your 40 trillion cells. When we talk about reaching LEV—escape velocity—by 2033, it's technology like this. So, congrats to Revel on doing this.

Ramin Hasani

And not just Revel. A couple of interesting notes here: it was Revel and Calico, the California Life Company, one of Alphabet's Other Bets that's been, I would say, a lot quieter than, say, Waymo. They're still doing work. That's very encouraging to me—that Calico is apparently deeply involved in this and has a heartbeat.

A couple of other points: the broader class of chemical reactions is called the Maillard reaction. It's also the reason why, when you bake bread, the outer crust is usually brown.

Peter Diamandis

Or, yeah—this is a vegetarian speaking—it’s why everything purportedly tastes like chicken.

Ramin Hasani

It's the same class of reactions, but the sugar is reacting with the carbonyl functional group, or carbonyl groups within sugars are reacting with the amines in proteins, to create a broad class of molecules that look optically brown. So the same thing is going on in the human body.

To me, this is very exciting because it's not quite unscrambling eggs, but it's halfway there. It feels almost like—again, strictly speaking, it's not a reversal of the thermodynamic arrow of time—but it's the next best thing if we can remove all of these unwanted sugar-plus-protein byproducts that are associated with inflammation and other correlates of aging, with directed evolution of a protein that came from bacteria.

What else is there out there in the biosphere for us to mine, in addition to all the obvious GLP-1s? Great potential for longevity escape velocity. What other bacterial innovations can we use to turn back aging?

Peter Diamandis

It's human engineering. We're taking control. It's going from evolution by natural selection to evolution by human direction, and I love that. I'll be happy when I have Ramin's hair.

Ramin Hasani

That's when I'll be happy.

Peter Diamandis

Well, there are lots of companies working on that, Sem. So, gentlemen, grateful for our time today. I'm excited for Starship 13's launch later today. Wish Elon and the group there lots of luck.

Ramin, congrats on the success of Liquid AI, and I'm excited to have you on the pod with us. Dave, great move investing in Ramin.

Ramin Hasani

On behalf of all of our—thank you.

Peter Diamandis

Yeah, gentlemen. Have an amazing week. I'm sure we'll be having an emergency pod very soon, because the speed of the singularity waits for nobody.

Mira Murati's 975B Open Model, Ramin Hasani on Post-Transformer AI, and Demis' AI FINRA | EP #271 | BidClub