[BidClub_]
Moonshots · · 128 min

Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272

Peter DiamandisSalim IsmailDave BlundinAlexander Wissner-GrossEmad Mostaque

YouTube
TL;DR
  • Kimi K3 put a Chinese open-weight model directly on the frontier, challenging the premise that capability leadership requires a closed US lab and its capital base. Moonshot AI’s 2.8 trillion-parameter multimodal model jumped 17 leaderboard places, ranked first in frontend coding and six other domains, and became the third point on Artificial Analysis’s cost-performance frontier behind Claude 5 and GPT-5.6 Solve Max. With weights expected around July 27, Alexander Wissner-Gross called the development “great for competition.”

  • The panel’s most consequential technical conclusion was that K3 contains “no magic”: recognizable transformer architecture, better data, relentless engineering, and optimized execution were enough. Emad Mostaque compared the process to Chinese EV manufacturing, noting that Moonshot remained on H800s while designing around Huawei and Alibaba chips. Peter Diamandis’s stronger—and more speculative—read was that the GPT-2 speedrun’s 99% cost reduction now implies “a 1% cost version” of models built in multibillion-dollar Western data centers.

  • K3 threatens foundation-model valuations, but the speakers sharply disagreed on how much revenue actually moves. Salim Ismail estimated that regulation plus open-weight substitution could erase 75% of a trillion-dollar lab’s value, because “frontier intelligence is now a totally perishable asset” with a shelf life measured in weeks. Blundin and Wissner-Gross pushed back: enterprises will still pay heavily for the best model, reliability, support, security, and frontier performance unless K3 becomes 2x, 3x, or 10x better—not merely close.

  • US chip controls may have accelerated the efficiency innovations now pressuring American labs, while China is explicitly treating open source as geopolitical infrastructure. The panel argued that constrained compute forced better quantization, data mixtures, kernels, and hardware-aware architectures; American inference providers may subsequently run K3 10x more cheaply on newer Nvidia and AMD hardware. Mostaque’s summary of China’s position was categorical: “We are going to fully back open source as a public good for humanity.”

  • For enterprises, the durable moat moves above the model into model-swapping architecture, proprietary data, verification, and workflow integration. Ismail argued that procurement cycles cannot keep pace with releases, so value accrues to interfaces capable of replacing models continuously. The practical recommendation was to evaluate Kimi K3 and Inkling internally, fine-tune on proprietary data, and treat generated code like human code: sandbox it, test it, scan it, and retain accountability rather than asking whether any model deserves blind trust.

  • Quantization could make today’s frontier capability local, persistent, and radically cheaper faster than model benchmarks imply. The cited Bonsai 27B result compressed a phone-scale model to 6 GB with a stated 5% accuracy loss, or 4 GB with 15%, while moving from 16-bit to ternary yielded a claimed 5x speedup; Samsung’s NanoQuant reportedly went below one effective bit per weight. Mostaque forecast K3-level capability in 16 GB of RAM by the end of next year, while Blundin projected 100x–10,000x raw-compute improvement within three years—potentially a million-fold when multiplied by algorithmic gains.

  • AI forecasting reaching statistical parity with human superforecasters could reshape markets, insurance, management, and individual decisions—but prediction becomes reflexive when institutions penalize people for ignoring it. Wissner-Gross imagined “hyperforecasting” connected to capital markets, able to model humanity’s next action before humanity takes it; Diamandis argued that much senior-management expertise then “essentially evaporates.” Mostaque supplied the warning: premiums and liability may punish anyone who ignores Dr. AI, making the central question not whether advice is accurate, but “whose grace” controls it.

  • The infrastructure opportunity broadens rather than contracts: cheaper models increase silicon demand, edge intelligence, robotics, and eventually orbital compute. The panel viewed semiconductors and inference providers as beneficiaries even if closed-model margins compress, while warning that 70 kg humanoids claimed to punch four times harder than Mike Tyson need safety rules before entering homes. Sam Altman and Elon Musk’s orbital-data-center positions sounded adversarial, but Mostaque found little numerical disagreement: limited deployment could reach a few percent of compute by the end of the decade, with the economic crossover later.

Digest · the substance, structured for research

1. Kimi K3 put open weights on the frontier

  • Peter Diamandis framed K3 as an “AI Sputnik moment”: 2.8 trillion parameters, a 17-place jump over the prior Kimi model, first place in frontend coding, and first-place rankings across brand and marketing, reference-based design, data analytics, consumer products, simulations, and content creation.

  • The promised full-weight release around July 27 is the strategic hinge. If delivered, organizations could download and run the model on-premises rather than sending proprietary data, intellectual property, or “proprietary alpha” through a US provider’s API.

  • Wissner-Gross added an important corrective to the shock narrative: Moonshot says Kimi held open-weight state of the art during nine of the previous 12 months. K3 was dramatic, but “over the past year it’s been basically Kimi all along.”

2. A recognizable transformer was enough

  • Wissner-Gross’s architectural read was blunt: “There’s no magic in it.” K3 remains recognizably transformer-like, using familiar mixture-of-experts improvements and Moonshot’s version of linearized attention rather than an unseen post-transformer breakthrough.

  • That makes its proximity to GPT-5.5 Max on the task-cost frontier more provocative. If a published, understandable architecture gets this close, Wissner-Gross asked, “What are the American frontier labs spending their money on?”

  • Mostaque located the differentiation primarily in data and execution. Kimi had long “felt a bit different” and led writing benchmarks; K3’s multimodality helps it understand varied inputs and produce unusually strong frontend experiences, from personal sites to games.

  • His manufacturing analogy carried the argument: Chinese models resemble Chinese EVs—known ingredients assembled efficiently, fully featured, and consumer-friendly. Moonshot was still using H800s while shaping K3 for future Huawei and Alibaba hardware, turning constraints into an engineering discipline.

3. The cost-performance duopoly became a free-for-all

  • On Artificial Analysis’s scatter plot, Wissner-Gross placed Claude 5 at maximum capability and cost, GPT-5.6 Solve Max second on the Pareto frontier, and Kimi K3 third. A frontier once described as an OpenAI–Anthropic duopoly now includes Meta, xAI, and Moonshot.

  • Blundin called “Sputnik” an understatement because open weights let any sufficiently capable corporation or government catch up without routing through a US lab, then fine-tune for a vertical use case that might outperform the general model in that domain.

  • The chart’s apparent endpoint also bothered him: benchmarks saturate at 100%, but intelligence does not. His preferred model is nested S-curves—one technology plateaus while constructing the next technology that begins another exponential ascent.

4. Software efficiency may matter more than training budgets

  • Diamandis and Wissner-Gross discussed the Keller Jordan speedrun around Andrej Karpathy’s nanoGPT: hackers repeatedly reproduced GPT-2 faster and more cheaply until its original training cost fell by roughly 99%.

  • In Diamandis’s interpretation, K3 is the first evidence that similar ideas scale to the frontier. He extrapolated that a $16 billion Colossus 2 facility training a 10 trillion- or 20 trillion-parameter model could face “a 1% cost version” built through better kernels, mixture-of-experts design, and software optimization.

  • Training data remains another large overhang. Diamandis argued that indiscriminately ingesting 20–30 trillion internet tokens includes enormous amounts of low-value or counterproductive material; pruning the dataset while retaining its intellectual difficulty lowers FLOPs without proportionally lowering intelligence.

  • Mostaque’s sharper economic analogy was pharmaceuticals: America funds costly R&D and charges premium prices, while generics reproduce the useful product dramatically more cheaply. Diamandis agreed that a 99% generic-style reduction fits AI better than the one-third-price car analogy.

5. Perishable intelligence moves the moat above the model

  • Ismail’s core call was that “frontier intelligence is now a totally perishable asset.” With leadership lasting weeks, an enterprise cannot evaluate vendors, conduct an RFP, convene committees, and deploy before several newer generations appear.

  • The value therefore migrates to an architecture that can swap models without rebuilding the organization around each release. In Ismail’s ExO vocabulary, that interface layer becomes more durable than any particular set of frontier weights.

  • Blundin connected that architecture to enterprise sovereignty: banks, governments, and industrial companies can create internal models around proprietary data rather than placing their future entirely in Anthropic’s or OpenAI’s hands.

  • Mostaque preserved the counterweight: buyers still pay IBM and non-Chinese suppliers for mission-critical support. US labs can grow revenue through reliability, forward-deployed engineers, integrated tools, and accountability even while open-weight substitution increases.

6. Foundation-lab valuation became the episode’s central disagreement

  • Mostaque initially guessed that regulation had already halved the hypothetical trillion-dollar value of a US frontier lab by delaying releases, and that open-weight competition could halve it again. Ismail later gave the AMA estimate of roughly $250 billion for a trillion-dollar OpenAI—a 75% haircut.

  • Diamandis contrasted Moonshot’s displayed $20 billion valuation with approximately $1 trillion each for Anthropic and OpenAI, suggesting public markets might have cut the US labs by 30% immediately. Both figures were explicitly presented as speculative marks, not observed transactions.

  • Mostaque was less bearish on near-term revenue: organizations will still distinguish between cheap substitution and an accountable supplier. His longer-term concern was vertical integration—once customers can own frontier-class models, those customers become the labs’ competitors, encouraging labs to move into their applications.

  • Wissner-Gross rejected the premise that K3 necessarily shrinks aggregate valuations. His companies would not divert frontier workloads merely for parity; Moonshot might need a 2x, 3x, or 10x advantage over Fable 5. Cheaper intelligence can also expand demand through Jevons-paradox-style effects.

7. Recursive self-improvement crossed the line earlier than policymakers thought

  • Moonshot’s claim that K3 designed kernels and a chip for its next generation made the system feel “AGI-ish” to Mostaque: the important unit is no longer static weights, but an ecosystem in which the model improves the machinery running and succeeding it.

  • Blundin argued that policymakers incorrectly waited for visibly Einstein-level intelligence. Recursive self-improvement only requires a model capable of improving its kernel enough to gain 10x speed; the faster successor then gets another chance to improve itself.

  • In his chronology, Opus 4.8—not Fable 5—had already crossed that threshold and could help China create Kimi K3. “The little spark was enough to ignite a flame,” he said, and the flame can become a fire and then “a sun.”

  • Ismail joined this to the law of accelerating returns: vacuum tubes helped design transistors, which enabled integrated circuits. Multiple reinforcing S-curves now operate inside AI simultaneously, making simple restraint policies poorly matched to the system’s dynamics.

8. China is treating open source as geopolitical infrastructure

  • Mostaque cited Xi Jinping’s World AI Conference speech as a commitment to support open source “as a public good for humanity” rather than stopping frontier releases. Chinese model approval, according to conversations he relayed, had fallen from roughly 60 days to about one week.

  • The incentives are domestic and geopolitical: tools can raise the effective capability of a billion people, robots can offset demographic pressure, and Chinese-trained intelligence can become embedded in critical systems worldwide.

  • A Chinese-backed regulatory initiative involving Brazil and parts of Asia and Africa looked to the panel like a new AI-focused Belt and Road. Mostaque summarized the inversion as “the Chinese Communist Party saving American capitalism from itself.”

  • The speakers worried that Washington might respond with model restrictions, disclosure requirements, “anti-token laundering,” or “know your prompter” rules rather than offering better US open weights abroad.

9. Blocking Chinese weights would be possible for companies, not for information

  • Wissner-Gross sketched an indirect route: require public companies to disclose Chinese-model use and subject it to SEC scrutiny, potentially through a FINRA-like, industry-funded frontier-AI body. Compliance costs could make K3 economically unusable for major corporations without erasing it from the internet.

  • Mostaque’s pushback was practical: once weights are released, mirrors, peer-to-peer networks, other jurisdictions, and VPNs make suppression porous. The main result would be denying American researchers, startups, and security teams capabilities available everywhere else.

  • His analogy was “denying Americans cheap insulin”—protecting expensive incumbents while generics spread abroad. The broader risk is that US regulation hobbles domestic capitalism as China encourages open development.

  • Wissner-Gross allowed one narrow exception: if a US court found that Moonshot obtained the model through copyright infringement, illegal trace distillation, or another proven violation, blocking distribution could be justified. Without that proof, US labs should study K3 and leapfrog it.

10. Export controls accelerated the efficiencies they sought to suppress

  • Mostaque argued that the Nvidia embargo created exactly the incentive China needed to exploit algorithmic, computational, and hardware efficiencies that were already available. The controls “irritated” enough to stimulate innovation without preventing competitive training.

  • Mostaque also compared the strategy to gradual escalation in Vietnam: neither decisive restraint nor open competition, but a middle course that produced the worst outcome. Quantization research triggered by scarcity now becomes permanent global knowledge.

  • Wissner-Gross said K3 used roughly the same total compute as Ling, yet Moonshot achieved roughly 2.5x better “data-to-intelligence conversion” through architecture and data mix. He treated that as visible evidence of learning under constraint.

  • Mostaque then argued that K3 has roughly 50 billion active parameters against nearly 3 trillion total and was optimized for Chinese chips. US providers with newer Nvidia and AMD systems could eventually serve it 10x more cheaply, with Mostaque forecasting a 10x–100x price decline after optimization for architectures such as Vera Rubin.

11. Open weights expand the semiconductor opportunity

  • Blundin rejected the market’s initial tendency to sell semiconductor companies alongside software labs. Cheap, customizable models increase total inference and training demand; the software value stack changes, but silicon becomes “more in demand than ever before.”

  • Quantized models also unlock chips and fabrication capacity unable to manufacture a GB300-class part but perfectly capable of inference on lower-precision models. Previously marginal compute becomes economically useful.

  • Mostaque pointed to American open-model inference companies—including Fireworks, Modal, and Baseten—as likely beneficiaries. In the AMA, he cited valuations of $17 billion for Fireworks and about $10 billion for Modal and Baseten, saying their recently raised capital would fund aggressive K3 optimization.

12. Stable Diffusion supplies the adoption playbook

  • Mostaque compared K3 with Stable Diffusion: restricted image generators were somewhat better, but blocked likenesses, proprietary IP, and many user-controlled applications. An open alternative generated 100–200 million downloads and an ecosystem that accelerated generative media.

  • The same logic applies when a closed model downgrades conversations involving biology or even philosophy. A customizable model at a fraction of the price becomes the substrate for tools that no single vendor would authorize or prioritize.

  • Mostaque used Mira Murati’s new open-source effort as evidence that frontier insiders believe an enterprise-controlled pathway can catch up. In the discussion, Tinker was described as the corporate fine-tuning path, while Inkling was also referred to as the model being tuned internally.

  • Ismail’s simple deployment prescription was two internal installations—Kimi K3 and Inkling—fine-tuned on company data. The proprietary learning loop, not the initial weights, becomes “the proprietary gold that you do not want to lose.”

13. Talent policy matters, but the Moonshot founder story was more nuanced

  • Diamandis used Moonshot founder Yang Zhilin’s Carnegie Mellon PhD to argue for stapling a green card to every US doctorate: America educates exceptional researchers and then allows them to build strategic companies elsewhere.

  • Wissner-Gross complicated that account with chronology. Yang began his CMU PhD in 2015, founded China-based Recurrent AI about a year later, and returned after graduating in 2019 despite reported offers from Google, Facebook, Huawei, and others. This was not clearly a case of America refusing to retain him.

  • The more actionable counterfactual, Wissner-Gross argued, concerns startup domicile: US incentives might have persuaded Yang to incorporate Recurrent AI domestically, after which Moonshot could also have remained American.

  • Ismail supplied the systemic number: roughly 70% of elite AI researchers are not US citizens, led by Chinese, Indian, Taiwanese, and UK talent. Mostaque added that about 80% of Chinese students return, while Indian graduates overwhelmingly stay, reflecting China’s stronger startup ecosystem as well as immigration friction.

14. Model releases are approaching continuous versioning

  • Diamandis counted 13 frontier releases since mid-April, about one every 10 days, versus eight during 2025 at one every 50 days and six during 2024 at one every 60 days.

  • Mostaque fitted an exponential to those intervals and got daily frontier releases by January if the trend held. At that cadence, named launches lose meaning and model infrastructure becomes continuously versioned.

  • Musk’s cited update added competitive pressure: a two trillion-parameter model, “better than our 1.5 trillion in every way,” would finish initial training the following week and might exceed Kimi while retaining speed and token efficiency near the 1.5 trillion model, called Grok 4.5.

  • The panel expected use cases to replace benchmarks as the compelling evidence. A marginal intelligence score feels abstract; a model recreating a game or producing a browser-based simulation of an Apple desktop makes the increment tangible.

15. Near-zero creation cost makes taste and purpose scarce

  • K3 demonstrations suggested that the friction of launching a personally imagined game is approaching zero. Mostaque preserved the caveat: one-shotting a game is not equivalent to building its marketing, customer support, community, and operating ecosystem.

  • Blundin saw a deeper organizational problem: once executives can prompt almost anything, the hard question becomes “What do we want?” Companies rarely had to articulate their purpose under conditions of nearly unlimited production capability.

  • Diamandis identified taste, imagination, and understanding public demand as the new constraints. His distinction was that passion is something one loves doing, while purpose is something one loves doing that also benefits the world.

  • The episode’s playful call to build an outro game resolved into a useful phrase: not a first-person shooter, but a “first-person solver” in which players cook problems rather than enemies.

16. Quantization puts serious intelligence in a phone

  • Mostaque described Prism ML’s Bonsai 27B, built on Qwen3-27B, as the first 27 billion-parameter-class model running entirely on a smartphone. He recalled it as roughly GPT-5-class, explicitly hedging that comparison with “from memory.”

  • Ternary quantization reportedly reduced the model to 6 GB with a 5% accuracy loss, or 4 GB with a 15% loss. The model can run offline, is smaller than some games, and supplies what Mostaque called a “1,020-IQ buddy” in a pocket.

  • Lower precision also increases speed: moving from 16-bit to three-valued weights yielded a claimed 5x improvement. A Tencent model from the former WizardLM team reportedly pushed a roughly 300 billion-parameter system to binary operation on a DGX Spark or a big MacBook, with about 5% performance loss.

  • The strategic consequence is persistent edge autonomy: vehicles, robots, factories, and consumer devices can make local decisions without connectivity or dependence on a remote provider.

17. Sub-one-bit models open new computing substrates

  • Ternary weights reduce core operations to multiplication by 1, 0, or −1. Diamandis’s point was that once the required arithmetic becomes this simple, AI no longer depends exclusively on conventional GPU-style multiply-accumulate machinery.

  • Mostaque said the most compressed Bonsai result was around 1.125 effective bits per weight, but sparsity, quantization, and low-rank factorization can go lower. Samsung’s NanoQuant had already crossed below one effective bit, with mainstream adoption naively extrapolated within a year.

  • Mostaque’s endgame forecast was unusually precise: 0.78 effective bits per weight. He also projected that distilling open K3 weights into smaller dense models, training at four bits, and casting to ternary or binary could yield K3-level capability in 16 GB of RAM by the end of next year.

  • Etching mature ternary weights directly into custom photonic silicon could eliminate much data movement. Mostaque forecast a 100x intelligence-cost reduction by the end of next year; Blundin projected 100x–10,000x raw-compute gains within three years, potentially approaching a million-fold with algorithmic improvements.

18. Superforecasting could automate judgment—and reshape its target

  • ForecastBench’s latest result, as presented by Diamandis, made several AI systems statistically indistinguishable from four human superforecasters on Brier score, where lower error is better.

  • Wissner-Gross highlighted Cassie—short for Cassandra—as the leading system, founded by a former British intelligence officer. His thought experiment connected “hyperforecasting” to algorithmic trading: models could predict humanity’s next collective action before humanity takes it.

  • That reflexive loop could make the efficient-market hypothesis newly powerful. Once forecasts move capital, the forecast helps produce the action it anticipated, creating what Wissner-Gross called “the ultimate market-efficiency outcome.”

  • Diamandis translated the result into corporate structure: budgeting, hiring, product launches, and investments are forecasts. If AI reproduces decades of executive judgment without equivalent human bias, much senior-management expertise “essentially evaporates,” leaving purpose and objective-setting as the human role.

19. Accurate advice creates liability, dependence, and control

  • Mostaque connected generative-AI mathematics with psychohistory’s idea that large populations can be modeled like gases using diffusion-style equations. He preserved Foundation’s limiting conditions: sufficiently large populations, no landscape-changing technological discontinuity, and ignorance of the prediction.

  • Deployment breaks the last condition. A healthcare, driving, business, or matchmaking recommendation changes behavior; insurers may raise premiums when someone ignores Dr. AI or refuses the approved autonomous-driving path.

  • Blundin preferred forecasting as a personal coach. His Homer Simpson chain—beer, couch, channel surfing, poor sleep, neglected family—illustrated how people drift into outcomes without choosing them; AI could expose an alternate path and its likely consequences.

  • Mostaque’s warning closed the loop: whoever controls the adviser can steer “vast waves of humanity.” If society is “watched over by machines of loving grace,” people need to know “whose grace that is,” especially when disobedience becomes more expensive.

20. Data-center backlash is disconnected from the cited scale

  • The panel compared 17 billion gallons of US data-center water use with 531 billion gallons for golf-course irrigation since 2024—a stated 31x multiple—and roughly 1 trillion gallons for California almonds, about 60x the data-center figure.

  • Additional comparisons sharpened the mismatch: Amazon warehouses reportedly occupy 10x more US land than all data centers, while Mostaque estimated roughly 600 gallons of water per Big Mac across 2 billion annual McDonald’s burgers.

  • Blundin’s concern was not the current water claim but the moving target: once rebutted, opposition may switch to another grievance and demand a moratorium, echoing nuclear policy. Diamandis traced that fear partly to decades of dystopian AI imagery.

  • Diamandis added that orbital compute could use closed-loop coolants, but expected objections to migrate toward atmospheric pollution or satellite decay. “The complaint will move on to something else.”

21. Humanoid combat is both an engineering test and a safety warning

  • China’s estimated 150 humanoid-robot companies are using spectacle to accelerate interest, including viral MMA-style fights. Mostaque acknowledged that combat stress-tests balance, impact resistance, recovery, locomotion, and latency in an exceptionally demanding environment.

  • Wissner-Gross found the spectacle disturbing, invoking the robot “flesh fair” in Spielberg’s A.I. and worrying that it establishes violence as a prior for future embodied intelligence. He would rather see robot competitions around ironing, coding, or another productive task.

  • The military implication was harder to dismiss: more autonomous edge models could place similar humanoids in PLA infantry. Competitive sports may advance the technology, just as Formula racing advances vehicles, even if audiences continue to prefer human drama.

  • Mostaque supplied the immediate safety case: EngineAI T800 robots weigh about 70 kg and were claimed to punch four times harder than Mike Tyson. With only 11,000 Unitree humanoids made so far but a forecast of 11 million annually within a few years, torque, autonomy, and household-operation rules cannot wait.

22. Orbital data centers are a timing dispute, not a destination dispute

  • Sam Altman called orbital data centers economically “ridiculous” today, citing launch cost and the difficulty of repairing GPUs, and said they would not matter at scale this decade. Musk’s answer was that SpaceX would begin launching them in two years.

  • Wissner-Gross saw a conflict of interest: OpenAI had shifted Stargate from owning data centers toward leasing terrestrial capacity, while other labs aligned with SpaceX infrastructure. He expected OpenAI’s rhetoric to change within two or three years as its position changed.

  • Mostaque found less disagreement in the actual numbers. With only about three and a half years left in the decade, orbital compute could reach a few percent of total capacity by the end of the decade and still cross terrestrial economics later.

  • Starship Flight 13 illustrated the enabling operational maturity: two of 33 Raptor engines failed to ignite at T−0, yet the system safely halted, unloaded propellant, replaced hardware, and targeted another attempt within days. Diamandis viewed the reported 5% SpaceX stock decline as missing the engineering achievement.

23. K3 is not yet automatically the cheapest or safest choice

  • Mostaque priced K3 at about $15 per million tokens, versus DeepSeek at $1, Sonnet at $20, Opus at $40, and Fable at $60. He estimated Chinese providers could already earn 80%–90% margins despite less-efficient chips.

  • K3 currently uses roughly twice as many tokens for the same task as GPT-5.6, while GPT-5.6 reportedly uses 37% fewer than 5.5 or Fable. Mostaque nevertheless forecast K3’s cost falling 10x–50x over the next few months as inference specialists optimize it.

  • On trust, Ismail’s answer was simply “no”—but he would not trust consequential human-written code without review either. The scalable system is model-generated code plus sandboxes, automated tests, security scans, proportionate permissions, and human accountability.

  • Blundin raised the possibility that the model could be benchmark-maxed; the open weights were expected to reveal within roughly two weeks whether that was true. Mostaque’s early use suggested a genuinely distinct model: not the best mathematician or cyber attacker, but unusually original in frontend, gaming, consumer, and entertainment work, potentially because of its multimodality.

Peter Diamandis

Today we put out the bat signal and called for an emergency pod because America just experienced an AI Sputnik moment. But more on that in just a moment.

We have the full quintet with us here today: Alex Wissner-Gross, Dave Blundin, Sal Khan, and Emad Mostaque. Gentlemen, welcome. Thanks for getting up early, wherever you might be—or midday in your case, Emad. I was up at 4:00 a.m. this morning, the benefits of jet lag, but I could have used another hour of sleep.

A lot is happening, and I appreciate everybody's time here.

Emad Mostaque

European siesta.

Peter Diamandis

Yeah, I've got a workout scheduled right after this.

A lot is going on. Before we get started, I want to personally say thank you to all our subscribers and viewers. I've had a chance—I don't know if you guys did recently—to watch and read the YouTube chat, and all I can say is, we love you guys, too.

Our mission here is delivering the news, and we spend an ungodly amount of time reviewing everything. Sal Khan, Alex, and Emad, I got your texts this morning: “Let's add this. Let's add that.” So much is going on.

Sal Khan

I have to say, all the memes of Alex explaining J-space are awesome. Keep memeing Alex every time you can.

Peter Diamandis

Yeah, for sure. And some great appreciation.

Alex Wissner-Gross

See if you can figure out my J-space.

Peter Diamandis

Yeah. Well, can we look inside? We'll be able to see. We're going to get a readout, and I see a lot of love for you in the comments as well. There's some wonderful people out there.

What's incredible is that most YouTube videos are just a kind of flamethrowing festival, and ours are completely the opposite. It's really amazing.

Dave Blundin

Kudos to you, Peter.

Peter Diamandis

Well, no, just absolute gratitude. I appreciate the fact that everyone—all of our subscribers and viewers—takes the time to listen to the podcast. We're constantly spending so much time with our entire team and the entire Moonshot Mates group, really trying to assess what's going on and deliver it.

We have these emergency pods. Gentlemen, shall we jump into the first story? It's a big one.

Dave Blundin

I'll just note that if we do enough of these emergency pods, at some point it turns into Moonshots Daily.

Peter Diamandis

Yeah. Or continuous. I still think moving into an Airbnb together and just turning on the camera—

Dave Blundin

It's going to happen.

Peter Diamandis

We've just had a Sputnik AI moment that's waking up the US frontier labs like a quadruple-espresso shot. Kimi K3 released yesterday, shocking the AI world with the largest open-weight model ever, and it went straight to number 1.

A little backstory: Kimi K3—and Kimi is from Moonshot AI, a Chinese lab. Over the last year, they've climbed the leaderboard. They put out K2, K2.6, and K2.7, each one closing the gap against Anthropic and OpenAI. This week, they didn't just close the gap; they jumped the fence.

Overnight, they released Kimi K3, and it's a monster: 2.8 trillion parameters. You have to remember the context here: China is doing this while under US export controls intended to starve them of the most advanced NVIDIA chips. That's a big deal I want to discuss with you guys.

They've completely engineered around the compute wall, and K3 jumped 17 places from the previous Kimi model, blasting past Claude Fable 5 to land at number 1 on the front-end code arena. K3 has also ranked number 1 in 6 other domains: brand and marketing, reference-based design, data analytics, consumer products, simulations, and content creation.

The full model weights are set to drop around July 27, which means anyone on Earth will be able to download and run this on their own premises. How big a deal is this, Alex?

Alex Wissner-Gross

I think it's great for competition. Let me first, as a preliminary matter, point out some things that have perhaps been slightly less obvious in the coverage—the meltdown, if you will, over K3.

The first is that, as Moonshot AI points out, they claim that in 9 of the past 12 months, Kimi models—the Kimi model series—have held state-of-the-art status among open-weight models. So, if that claim is indeed true, over the past year it's been basically Kimi all along. I think that's very interesting.

Secondly, taking a look at the published architecture—we haven't actually seen the open weights yet, but they're promised later this month—there's no magic in it, and that's pretty striking. One can imagine that behind the scenes at Anthropic or OpenAI, as Sam Altman continues to tease, there's some post-transformer architecture lurking behind the scenes and achieving all of these performance breakthroughs.

But taking a look at the published K3 architecture, there's no magic. It's still essentially a transformer. They've made, obviously, a number of innovations, but they're well-understood innovations concerning how they do mixture of experts and how they linearize attention. They have their own special Kimi brand of linearized attention, but it's still basically a recognizable transformer-like architecture.

And I think the fact that a recognizable transformer-like architecture can almost match GPT-5.5 Max on the task-cost frontier—which we should probably throw up a slide for you—I think that's—

Peter Diamandis

Pretty striking. That does raise the question: What are the American frontier labs spending their money on?

If you can just use a transformer to get this close—not like it's already on the cost frontier, but you can get close, like third place on the total state of the art for overall AI performance—what the heck are the American labs spending all of their money on?

So I derive great comfort, at minimum, in knowing that the transformer architecture is still alive and kicking.

Emad Mostaque, your analysis here, because you've been tracking this. We've been going back and forth on WhatsApp together.

Emad Mostaque

I think Kimi has been at the top of various benchmarks. Again, you can pick and choose. They've had the largest open-weight models out of China regularly ever since they first kicked off a year and a bit ago.

As Alex said, the architecture isn't anything super novel. There are improvements, like their Muon scaling that they did with UCLA, and other things. They've actually been releasing breadcrumbs of all of these parts.

I think what's key here is the underlying data. Kimi has always been a model that felt a bit different. That's why it was always at the top of the writing benchmarks, for example.

What they've done here seems extraordinary. When GLM came out, it was a fantastic model. It wasn't quite up to frontier, but it was text-only. K3 is actually a multimodal model, so it can have all sorts of inputs and understand things, which is one of the reasons it's so good at front end, although we wouldn't have expected it. Again, it's number 1 in front end versus everyone.

I think this comes to something I've said before, which is that building great, solid models is cutting-edge manufacturing. You'll have algorithmic improvements, and there are all sorts of things coming, but why are Chinese EVs better than Fords? This actually feels like the same thing, right?

It's engineering, but it's also about execution. The number 1 car here in the UK last month was the Jaecoo J7, or the new Land Rover, as it's been known. It comes fully loaded, full-spec, for like 50K—a third of the price. This actually feels like something very similar.

They've known what the ingredients are, the raw materials. They're now putting them together in an incredibly consumer-friendly way, and they're just executing that manufacturing process with what they have. When you look at the architecture internally, they're still on H800s. They're a couple of generations behind on the NVIDIA chips, but they've built it to take advantage of Huawei and Alibaba's next-generation chips, which you can see from the static shapes and all sorts of other things as well.

They're just relentlessly going at the engineering and the usability, which is why the front-end code, I think, is where they're standing out. They're asking, “How can we make it have the most amazing outputs—a personal website, a game, and other things?” Whereas the US labs are maybe looking in other directions and focusing a little bit on different things.

Peter Diamandis

Yeah. I mean, one question real quick: We've always talked about whether we need another breakthrough beyond LLMs to get to AGI. Does this give you comfort that we don't need another breakthrough to really move forward?

Emad Mostaque

Again, it comes down to the definition of AGI, right?

Peter Diamandis

Yeah, of course. Don't get me started.

Emad Mostaque

Six years ago, Peter. It was 6 years ago.

Peter Diamandis

I mean, I guess the question is, there's plenty of headroom still to progress these models.

Emad Mostaque

Well, I think you have the base model here, right? But then you've got all these amazing harnesses that are coming out, and the way that you're using the model to go back on itself. One of the things that's in the Kimi blog post is that it actually designed a chip for itself for its next generation, and it designed its own kernels for running as well.

So you move from this model weight to this whole ecosystem that the model itself builds. That feels AGI-ish, right? That feels like recursive self-improvement. That feels like the ability to learn and adapt new skills dynamically by changing itself. So I think, for most definitions of AGI, we probably don't need something new to optimize and make it super-efficient. There are various ways, even with what we know, that it could be more efficient than what we have here. We just don't have quite enough compute for it, and new architectures could push us even further.

Peter Diamandis

You could say, “Attention is still all you need.” I like that. So, Alex, we've thrown up here the performance charts, and we see Kimi K3 sort of topping the charts in a multitude of places.

Alex Wissner-Gross

If we could throw up the AI scatter plot, I think that's probably the most instructive one. So this is from the Artificial Analysis Intelligence Index, and this is, of all the charts at this point, my favorite one because it actually shows the cost per task as defined by AI versus the performance frontier.

One can mentally look at this, for those who can't see it. We see the frontier as a jagged frontier going from lower left to upper right. In the upper right, we see maximum cost per task and maximum overall score. That's still Claude 5. Riding the Pareto frontier down and to the left from that, we see number 2 on the frontier is still, as of a few days ago, GPT-5.6 Solve Max. And now, for the first time, Kimi K3 is number 3.

It's on the frontier. It's number 3 both in terms of raw capabilities and also the third point on the optimal cost-performance frontier. And I think that's totally striking. We went from a world where, as we mentioned a couple of pods ago, there was this OpenAI–Anthropic duopoly to now it's a free-for-all between Meta and xAI on the American side, joining the upper end of the Pareto frontier. And now China's Moonshot AI is number 3 on that Pareto-optimal frontier.

Peter Diamandis

Amazing. Dave, let me pull you in here. What are your thoughts?

Dave Blundin

Well, you know, Peter, you called it a Sputnik moment. If anything, that's an understatement of the implications of this. We had that Alex Karp rant on the podcast last week where he was saying, “Look, as a large enterprise or as a government, you can't just throw all of your proprietary weights, your proprietary alpha, all of your intellectual property over the wall to Anthropic and make that the basis of your whole future.”

But he didn't give you a roadmap to move forward. Here we are just a week later, and it's suddenly a free-for-all. As Alex was saying, it's a free-for-all where anyone who reads these weights has the ability to get very close to the frontier and then fine-tune for any vertical use case beyond the frontier. And so it gives everybody in the world—every corporation, every government in the world—a way to catch up to the frontier without going through the U.S. AI models.

So, Sputnik—Sputnik times infinity, essentially. The thing I don't like about this particular chart is that the left axis goes to 100%, and when you chart it out over the next 2 years, it looks like an S-curve. We're in this really steep part of the curve right now, but it implies that we get to 100% and then we've achieved the end.

But this is actually an exponential where intelligence goes to infinity. So the benchmark saturates, but intelligence itself goes to infinity. And so now it's really clear. Just for everybody, the way this works typically is nested S-curves, right? One particular technology tops out, but it builds the next technology that then begins its exponential ascent and so on.

Peter Diamandis

Exactly. Exactly. Right, let me just say one other thing. Alex and I have spent a lot of time working on this Keller Jordan speedrun. We talk about it a lot. It's a way you take a GPT-2-class model. You can find it online very easily—just look on GitHub. Look up “Keller Jordan speedrun.” And it's a whole bunch of hackers and AI researchers who are continually trying to take GPT-2, way back, 5 years ago—

David Blakely

In the form of Andrej Karpathy's nanoGPT, in particular.

Peter Diamandis

Exactly. And they try to recreate it faster and cheaper, faster and cheaper. If you look at the innovations in that repo, they've been able to cut the original cost of creating GPT-2 by 99%. So now it's 1% of the original cost.

Everyone kind of doesn't pay attention to it because it's GPT-2. Up until today, it wasn't clear whether those same ideas would apply at frontier scale. Now it's really clear that when Elon Musk takes his $16 billion Colossus 2 data center and builds a 10 trillion- or 20 trillion-parameter model for billions of dollars, there is a 1% cost version of creating effectively the same thing. Nobody knew until Kimi K3 whether that was going to work or not, and now it's really clear that it does work.

So we're looking at 100×-type innovations in the software stack, in kernel optimization, in mixture-of-experts architectures. These fundamental breakthroughs that come out of China are giving them a model at 1% of the cost. I think Emad gave a great analogy to the car: You can get a virtually identical car for about a third of the price. Here we're talking about less than 1% of the price to create the equivalent product.

So, Sputnik—yeah, that's the understatement of the century. This is just—and that's why we're on the emergency pod today.

Yeah, Salem, jump in.

Salim Ismail

I have 3 points to make. I think it's not so much that Kimi's beaten everything. It's the fact that frontier intelligence is now a totally perishable asset. The shelf life is weeks now for anybody that gets to the very edge, and any enterprise or government interested in that very latest cutting-edge frontier model doesn't have time to actually evaluate it, do an RFP, look at other models, have a committee internally, and think about whether to deploy it. Now you're 3 generations ahead in the model anyway.

So all the value comes in the architecture that can swap models, right? And that's going to be the next layer. We call that interfaces in our ExO world. That's going to be where all the value resides going forward.

Peter Diamandis

Yeah. Amazing. I love—

David Blakely

We need a new term for that. Maybe the Frontier Liberation Front.

Peter Diamandis

Let me say one other thing for the hypergeeks out there. Emad said the Muon optimizer, but he said it very quickly. Anyone who's an enthusiast should look that up as well, because one of the reasons this is happening is that when we built these original, very large-scale models, we took 20–30 trillion tokens from around the internet—every word ever written by humanity—and just dumped it into the training set and said, “Here, AI, become intelligent given all of this information.”

But when you look under the covers, the vast majority of that information is Taylor Swift's concert coming up and their wedding. It's a whole bunch of stuff that doesn't actually drive the intelligence of the model significantly.

David Blakely

The opposite, in fact.

Peter Diamandis

Yeah. Yeah, it's very true. A lot of those tokens might actually slow down the training, not accelerate it. And so, purely by pulling out the garbage and stripping down the training set to the relevant subset, you can still tax the model just as much, but it reduces the number of FLOPs—the amount of computation that the model is doing—to get to the same level of intelligence.

I don't think we're anywhere near done with that problem yet. So you can expect more 10× improvements to come out of just the Muon optimizer process and the process of stripping down the training data set.

I threw up this tweet from a guy named Alaric that I found fascinating. For those not viewing this, it says, from Anthropic, quote: “Fable is an agentic coding superweapon capable of developing cyber and bioweapons at unprecedented speed and scale. We cannot in good faith release it without guardrails.” Right? This was the conversation a month ago.

And China comes back and says, “Laughing my ass off. Here's Fable, but open source. Good bleeping luck.” So I am curious: How do you guys think about the fact that we were so constrained because of the guardrails, and here's an open-source equivalent of Claude?

Emad Mostaque

Well, the frontier labs have a major problem. They've got 3 fundamental, massive constraints that they can't get around. One is compute and the availability of chips, and all the electricity and power that's needed. The second is frontier open-source models that are as good as, or in many cases substitutable without much notable difference. The third is you've got the government coming down on you, going, "We need to check before you release anything."

I would make a thumb-in-the-air guess: the $1 trillion that OpenAI might have been worth shrank by about 50% when the government said, "We have to review all these models," because now it's going to take time to get things out. I think this crashes it by another 50%. I would put the finger-in-the-air value of these frontier labs at about a quarter of what they were 3 months ago.

Peter Diamandis

If I don't have to spend the money for the API calls and I can just use Kimi K3 on my on-prem, why would I spend the money? Are they going to be hit by massive reductions in revenue?

Emad Mostaque

Yeah, I think there are a couple of things here. Number 1 is reduction in revenue. Why do people pay for IBM? Why do they pay for non-Chinese cars for mission-critical things? I think having on-call entities where you know things aren't going to go wrong will still sustain them for a while. So I think revenues will still go up for OpenAI and others, and this is why they built these forward-deployed engineering companies as well. I think they've still got a way to go.

But you have the substitution effect again. This is just like Chinese industrial substitution: Why can't America build industrial things? Why do you have to buy Chinese? Sometimes you buy Chinese; sometimes you buy American. I think we'll see that for at least another year, but then it gets difficult with the cyberattack, security-theater kind of things that we've had. I've maintained that we would get to this point.

What does it mean? It means the only form of thing that you can actually do is cyber defense. This must be the absolute biggest category in VC right now. If you're a talented Stanford or MIT grad, build a cutting-edge cyber defense startup that goes into every other company and says, "Let's use this technology to defend against what's inevitably coming."

The proliferation of these capabilities is going to increase, but not quite as fast as we think, because what actually happens—and we've done some tests around this—is that GPT-5.6, the cyber version Fable, and so on are trained on lots of CVE and cyber data. The Chinese models don't actually have that much of that, so they're not that great. But someone can train that data into them if they have it.

We'll probably see cyberattack-capable open source emerge in a quarter or 2. So there'll be a bit of a lag there, but definitely for the types of big adversaries, it's going to get a bit crazy.

Peter Diamandis

Dave, you want to jump in?

Dave Blundin

Yeah, for sure. I think we glossed over recursive self-improvement there. Peter, you asked the question of whether this is the tipping point. From the view of the U.S. government, we always knew it was going to be too late, right? It just moves too slowly.

But the view was, look, when we get to a model that's capable of building itself and building the next model, we're not going to let that go out to everybody in the world so they can catch up overnight. There's never been a product in the history of manufacturing like a car. If you have your state-of-the-art car and you give it to a foreign government, they can't use it to make a better car. But AI doesn't work that way.

If you have state-of-the-art AI and you give it to a foreign government, they can use it to actually catch up to you and create state-of-the-art AI. That became clear to the government a month to a month and a half ago, that Babel 5 was over that line, and so they stopped it. But the reality is that Opus 4.8 was over that line. People in China could use Opus 4.8 to create Kimi K3.

That recursive self-improvement line was actually crossed earlier than Llama 5, and that's going to be obvious to the world now. All you need to do is have an AI that's capable of improving its own kernel. It doesn't have to be—this is a point I've made on a podcast months ago—at Einstein-level intelligence. All it has to be able to do is improve its own kernel and get a 10x step up in speed, which nobody perceives as being true AGI, but that's all it needs to accelerate itself by 10x.

Then the 10x-smarter, or 10x-higher-parameter, model will have some higher level of intelligence. A lot of people in academia were saying, "Well, look, we're getting diminishing returns with the parameter count, so a 10x-faster model won't natively be 10x smarter." But that turned out to be wrong. We're seeing slowing, but we're not seeing flattening of the intelligence curve.

All the evidence now is that if you boost the raw speed by another 10x, you're going to see genius-level AI. Then that genius-level AI will boost its speed again. I think when we look back on this in history, we'll say right around Opus 4.8 was the point where the little spark was enough to ignite a flame, and then a flame can become a fire, and then a fire can become a sun. That's the way we'll look back on this moment in time.

So the cat is definitely out of the bag. The current U.S. policy of constraining the next model has no way of containing the global and corporate proliferation of frontier AI. Do you think the U.S. government starts a strategy of constraining, in some fashion, Chinese open-source models from being used in the U.S.?

Emad Mostaque

Well, in 2 weeks these weights are supposed to be open-weight, open-source, and then—

Peter Diamandis

We'll see. If they're rational at the White House right now, they're spending every minute in a debate: Do we negotiate with China immediately and not release those open weights? I really doubt they'll move quickly enough. I'm sure they'll—well, I'm not sure. We'll see what happens in 2 weeks. Fascinating.

Emad Mostaque

Can I merge 2 ideas here?

Peter Diamandis

Yeah, of course. Please.

Emad Mostaque

Peter, you talked about exponentials and the law of accelerating returns. I think it's worth drilling into that, because if you connect that to what Dave just said, this is why we've been saying forever and a day on this podcast that this is unstoppable.

Ray's original observation was that once you have an information-based paradigm, you just keep hopping across multiple technologies. So we had vacuum tubes, relays, and then transistors in computing. At some point, you can only fit so many vacuum tubes into a room, but that architecture was used to design transistors. Transistors were used to design integrated circuits, and you get these nested S-curves.

What Dave is talking about is that as these architectures—the various pieces of the puzzle—get all these reinforcing loops inside them, each of those is like an S-curve that starts accelerating the collective, and it's unstoppable. There is no limit to where this goes. This is why people are so freaked out about the upper-end limit of this. It's so important to connect those 2 dots.

Peter Diamandis

Yeah, for sure.

Emad Mostaque

If I just say something, Peter, I'm following on from Dave. There was an important speech by Xi Jinping a couple of days ago—yesterday, God, time flies—at the World AI Conference in Shanghai, where he basically said, "We are going to fully back open source as a public good for humanity, and we're not going to regulate and stop it."

This is their plan. It's great for China for a variety of reasons: the fact that they have a billion people whose IQ is about to increase by having these tools, the fact that they need robots to solve their demographic problem, and the soft power from putting a Chinese-educated brain—a Tsinghua graduate—into every critical system in the world.

But they're going to keep on doing that because they actually have a regulator. From talking to some of the Chinese labs, it used to take 60 days for a model to be approved. Now it's like a week.

Peter Diamandis

Amazing. Xi also announced a regulatory body that they've created, which includes Brazil and different parts of Asia and Africa. I don't know if you guys saw that.

Emad Mostaque

I saw that. It means, obviously, the new Belt and Road is now focused on AI coming out of China. It's a bizarre future where the Chinese Communist Party is saving American capitalism from itself.

Peter Diamandis

It's so true. Let's also note that Yang Zhilin was a CMU graduate—

Dave Blundin

And we could have given him a visa to stay.

Peter Diamandis

Yeah, we're going to get to that story in a second. This is an interesting chart here that shows the valuation. Kimi's valuation, or Moonshot AI's valuation, is at $20 billion, compared with Anthropic at $1 trillion and OpenAI at basically $1 trillion as well. If they were public companies today, I think you would have seen a 30% stock-valuation drop.

I'll ask again: What are the American frontier labs doing with all of their capital?

Dave Blundin

Yeah. What are they spending their money on?

Emad Mostaque

Actually, if you go into the buildings and talk to them, and you have any idea at all, they'll give you the capital. They're desperate for more smart people to help because they're trying to deploy and change the world at this insane pace that no one has ever experienced before. They want to deploy that capital much more quickly than they can find smart people who have good ideas to use the capital.

But it's a great point. You're sort of saying it in an accusing way: “What are you guys doing with your capital?” But no one in the history of the world has ever had this much money pour into their building this quickly, with no prior business experience. We're talking about CEOs who have never run a company before. They're trying, but seriously, can any human being really rise to the occasion of AI that quickly?

My point there, though, is that if you're smart and you have good ideas, get into those buildings and propose your ideas. This applies to XPRIZE, too. They are desperate to move that money out the door into something productive that gives them a sustainable barrier to entry.

Peter Diamandis

I also think that the Frontier Labs are asking themselves that question and asking the U.S. regulatory apparatus that question. Anthropic regularly sends out smoke signals accusing various Chinese frontier labs of distillation attacks. Maybe, in Anthropic's public mind, that's how the Chinese labs are able to do it: through distilling and capturing reasoning traces.

Honestly, looking at Kimi K3's performance, I'm not at all convinced that Moonshot is achieving its performance purely, or even substantially, through distillation attacks on Claude. It just doesn't smell right.

Emad Mostaque

No, no, I totally agree. I think, though, that there's a tendency to underweight or undervalue the existence proof: just the knowledge that a highly scaled transformer running with a Muon optimizer and simplified data works gives you a much more refined road map. You don't have to copy, you don't have to cheat, and you don't have to steal every trace. You just have to know that the formula works, and that cuts your R&D costs by 90–95%.

I think it's just that simple. There's nothing sneaky or cheaty about it. It's just knowing you're on the right path.

Peter Diamandis

I have the greatest value-creation idea for ourselves ever: in 9 days, when they drop their open-source weights, we release an open-source model called Kimi 4 under the Moonshots podcast name and [laughter] IPO it.

Emad Mostaque

Okay.

Peter Diamandis

So you're saying, what's better than 1 Moonshot? Moonshots plural.

Emad Mostaque

Well, why not copy the copiers? Let's go. [laughter]

Peter Diamandis

That's good. Actually, I've got a good analogy for you, Emad. Why do Americans pay more for drugs than everyone else? All the R&D happens in America. You pay the premium, just like token premiums. Then what happens? You have generics elsewhere.

Emad Mostaque

Yeah, that is a good analogy because that's like a 99% cost cut. It's much more akin to AI than to cars. That's a great analogy.

Peter Diamandis

You know, our friend Gavin Baker wrote a brilliant post—you can find it on X—about the implications of this for businesses.

Emad Mostaque

It's essentially a must-read. Absolutely. But essentially, all businesses—all stocks other than the foundational AI labs—are huge beneficiaries of this. Then the foundational labs are, like you said earlier, asking, “What's your future? What's your revenue model? Why are you worth $1 trillion? I don't quite get it.”

You should see a really big reshuffling of valuations in the next week based on that observation. Any corporation that has its technical act together—there aren't very many of those—but if you're a bank that happens to be a very good bank with brilliant IT and technical skills, or you have great partners and great vendors, you now have a clear road map to controlling your own destiny with your own AI, your own JPMorgan AI.

I suspect the markets will react to that if you put your hand up and say, “Hey, we have a way to do this internally. We know how to do this with our partners, or however you get it done.” This is why we call it the organizational singularity.

Peter Diamandis

We're still seeing everybody who's using or trying to use Fable 5 getting downgraded every time they mention biology or something that's potentially on the edge. Why would you tolerate that?

In 9 days, what do we see? As soon as it's available, I'm going to upgrade. I'm running Kimi K2.7 on my Mac Studios; I'll upgrade it to Kimi K3. Everybody will. Do we start to see the wholesale U.S. entrepreneurial base of capabilities on K3?

Emad Mostaque

Well, I can think of an analogy for this, which is Stable Diffusion. When we released Stable Diffusion—God, 4 years ago, time flies—you had these really restricted image generators that were a bit better, but they were restricted and had all sorts of arbitrary restrictions because, obviously, it's a bit dangerous to have it. You couldn't have likenesses. There was no way to get IP in there, even if it was your own IP, and there were more restrictions.

What happened? 100 million, 200 million downloads and a whole ecosystem built around that, which accelerated generative media, as you said. Why are you going to have this model when I can't even talk about philosophy with it? It downgrades me, right? When you can have the fully open variant of it at even a fraction of the price, which you can then customize, a whole ecosystem will build around this and other models. It has already been doing so, and that's a real danger versus being locked into a single vendor.

Peter Diamandis

Which is why I think the labs will go vertically integrated. All their customers are now going to be their competition, and they're going to be like, “Okay, I'm going to take you all on.”

That directly ties to Mira Murati and Inkling. Are we going to talk about that story, too? That's huge this week.

Emad Mostaque

We talked about it in the last pod, which was 2 days ago. [laughter]

Peter Diamandis

Okay.

Emad Mostaque

Mira just released Inkling, which is fantastic to see from a U.S. open-source lab. But the question is, how many more will we get?

Peter Diamandis

How many more open-source, shocking AI Sputnik moments are we going to see? We have a lot of Chinese labs pursuing this beyond just Moonshot.

Emad Mostaque

Yeah. Tinker is really telling about where things are going to go because it's designed for you to pick it up as a corporation and fine-tune it within your corporate walls to whatever your use case is. If you're a biotech lab and you're researching, and you don't want everybody to see your proprietary data, you take Inkling and tune it internally.

The reason that's telling is because Mira Murati came from OpenAI. If she didn't believe that pathway was viable, she wouldn't have started Thinking Machines Lab on that thesis. It tells you that the people inside the best frontier labs believe that this process can catch up to the frontier. You combine that with Kimi K3 proving it, and it's a different world next week.

The other thing that was weird in the market at the end of the week is that things started to reshuffle pretty dramatically toward the end of the week. In the downdraft, the semiconductor companies also came down, but they're actually going to go the other direction. This is the point Gavin Baker was making: this drives up the need for silicon, not down. It changes the whole software landscape tremendously, but silicon is going to be more in demand than ever before and completely sold out, as we know.

Peter Diamandis

Can we talk a second about the NVIDIA embargo that we put in place for China? Here we see the highest-performance models. Was the whole NVIDIA regulatory embargo unnecessary? Did it do what we've always done before, which is spark China's need to develop its own capabilities with Huawei?

Emad Mostaque

Of course that's what happened. The embargo only incentivized the Chinese frontier labs to develop and cultivate new efficiencies that, by the way, were always there. To Peter's point earlier about the nanoGPT speedrun, there's this enormous overhang that isn't fully exploited in terms of leveraging algorithmic, computational, and hardware efficiencies to train larger and more capable models.

All these export controls do, I think, is incentivize the Chinese labs—which are already feeling plenty of demand pull to compete with Western frontier models—to leverage those efficiencies sooner. Maybe, on balance, although it's superficially bad for the West now that we've incentivized this new generation of much more efficient Chinese frontier models, in the end I think it's net good for not just the world but also for the U.S. to have this fire lit underneath them by Chinese competition that's much more efficient, more capital-efficient, more weight-efficient, and probably more bit-efficient.

This is all a net positive as long as the U.S., in my mind, does not set up or fall into some ultimately protectionist regime of trying to prevent what may be construed as Chinese superintelligence dumping on the U.S. Exactly.

Peter Diamandis

As long as we avoid that, it's great.

Emad Mostaque

Exactly what happened. That's exactly right. I think the U.S. learned a really important lesson in the Vietnam War. Because that's over 50 years ago now, it's been forgotten again, and then you have to be reminded again.

In the Vietnam War, it was really clear that either you go to war and win quickly, or you don't. What you don't do is send in a few troops, then send in a few more, and then creep in. Nothing good comes of that at all.

The embargo of chips on China was totally harebrained because it was enough to irritate but not enough to actually work.

Salim Ismail

It's just the worst-case scenario, and it sparked, exactly like Alex said, a huge amount of quantization research, which is critically important and underdiscussed. That allows faster performance on cheaper chips, and those innovations don't go away. That's going to be around forever now.

Alex

I think we've got something completely self-contradictory, but it has some interesting outcomes, so the total amount of compute used for Kimi K3 is the same as Ling.

Peter Diamandis

Wow.

Alex

You can tell that because of the amount of dense weights, and we roughly assume about twice the number of tokens trained because we don't have that. We were like, "But how does that work?" Well, you look at their architecture, and it's a 2.5-times data-to-intelligence conversion through the advantages in the data mix that they have, because they've had to operate under these constraints.

And we see that because the first model isn't as good as the second model. And for Ling, you're going from a 1-trillion-parameter model to a 300-billion-parameter model about to be released, which actually has better performance. So you see this with the labs, and these labs have had to deal with the constraints.

Emad Mostaque

But here's something really interesting, I think. If you look at that slide Alex loves and we put it up on the screen, what they've had to do is optimize their inference for Huawei Ascend 910 chips, for the new Alibaba chips, and others—64 nodes in 1—because this is a big model. You're going to have to buy another Mac Studio or 2, Peter, to serve this; it needs about 2 TB of RAM.

You see where Kimi K3 is there. That's because they can only use Chinese silicon to run it. They don't have Blackwells; they don't have Vera Rubins. Vera Rubins and Blackwells are designed for these really large, sparse models, because it's 50 billion active parameters against 3 trillion total.

American companies like Modal, like Fireworks, like Baseten will be able to serve this model 10 times cheaper than their Chinese competitors because they have access to the NVIDIA and AMD big chips. So, like I said, it's a bit ironic where the development R&D suddenly has gone there, but there's going to be a 10- to 100-times price drop once this is optimized for the next generation via Rubin.

Peter Diamandis

Well, that also, Emad, you're saying essentially the same thing, but that also unleashes a bunch of chips that aren't currently in circulation. They're underpriced, and it unleashes a bunch of fabs that can't make a GB300 but can make an inference-time chip that'll run the cheaper Chinese—or the lower-granularity Chinese—model. A lot of compute capacity gets unleashed through that same process you just described.

Salim Ismail

If I could go up a level and go a little woo-woo, right? We've had this mantra in the internet world—(laughter)—that information wants to be free. Basically, intelligence also wants to be free.

Essentially, we've gone over the course of evolution from biological intelligence, where you had evolution built in—recursive improvement—and then we broke through that to individual intelligence, to the intelligence of a species, to collective intelligence like markets or networks. Now we have AI, which can scan across all the data to create a whole other level of intelligence. This is not stoppable. Any entity or domain or government or whatever that tries to constrain it always, always, always, always fails.

It's just a fundamental law of nature that you cannot constrain this; it's just not possible. Why people bother is what really blows my mind. It's a very scarcity mindset to try and think about it this way. The faster we get to better intelligence, the faster we get to abundance, and the faster we don't need to fight over anything.

Peter Diamandis

I can't disagree with you, Salim. I wrote an entire paper arguing that intelligence manifests in the physical world as maximizing future freedom of action. So here's to the Right Frontier Liberation Front.

Now you remind me of the Monty Python thing where there's the People's Front of Judea and the Judean People's Front.

Salim Ismail

We need T-shirts.

Peter Diamandis

This episode is brought to you by Blitzy, autonomous software development with infinite code context. Blitzy uses thousands of specialized AI agents that think for hours to understand enterprise-scale code bases with millions of lines of code. Engineers start every development sprint with the Blitzy platform, bringing in their development requirements. The Blitzy platform provides a plan, then generates and pre-compiles code for each task. Blitzy delivers 80% or more of the development work autonomously while providing a guide for the final 20% of human development work required to complete the sprint. Enterprises are achieving a 5x engineering velocity increase when incorporating Blitzy as their pre-IDE development tool, pairing it with their coding co-pilot of choice to bring an AI-native SDLC into their org. Ready to 5x your engineering velocity? Visit blitzy.com to schedule a demo and start building with Blitzy today.

Let me bring up a related subject to the story here that I have a pet peeve about. It's this one. The founder and CEO behind Moonshot AI, Yang Zhilin, didn't learn his craft in Beijing. He earned his PhD at Carnegie Mellon, one of the best computer science programs in the world, in Pittsburgh.

We basically trained him up—we admitted him, trained him up—at one of our best institutions, and then, when he gets his PhD, he doesn't get a green card. He goes through the hassles of trying to get a visa, and he goes back to China and builds Moonshot AI there. Let's, for a moment, talk about this. I've stated publicly so many times that I think when anybody gets a PhD, they should get a green card stapled to the back of it. Why are we sending the most brilliant people who come here to get educated back home, whether it's to China, whether it's to India, whether it's to Brazil? Why don't we enable them to stay here and build? Gentlemen, comments on that.

Alex

Okay, so I did some research on this, and I think the story is not what it seems to be. A little bit of chronology first. Yang Zhilin, according to my research, starts his PhD after undergrad in China—starts his PhD at CMU in fall 2015. Then, approximately 1 year later, he founds a startup while a PhD student at CMU.

The startup is named Recurrent AI. Where is Recurrent AI based? It's based in China; it's not based in the US. 1 year into his PhD program, he starts a Chinese AI startup while still doing his PhD at CMU. That's interesting, and that's a problem.

This also runs counter to the sort of narrative of, "Oh, we wouldn't staple his visa or whatever, and then he goes back to China." No, actually, 1 year into his American PhD program, he starts a Chinese AI startup. Then he graduates in 2019, is my understanding.

My understanding is he had offers from Google, Facebook, Huawei, and others upon graduation in 2019, but he goes back to China because that's where his startup, Recurrent AI, was actually incorporated a few years earlier. I don't think necessarily this is the case where either the US was unwilling to retain him or even President Trump, somehow through some policy, was driving away this particularly talented Chinese graduate. He started his company during the tail end of President Obama's term in China.

Peter Diamandis

And Alex, I appreciate the deeper dive that you did. Thank you for that. The point still stands. You and I have seen this so many times, right, at Singularity University. Dave, you may have seen this at MIT.

The fact of the matter is, a lot of the most brilliant students aren't given the opportunity to stay and develop here. Dave—or actually, Emad, what are your thoughts on that, being someone not in the US?

Emad Mostaque

If I can just give my 2 cents, I completely agree with it, and here's this crazy thing: the math and the numbers are all there. What is the value of a PhD staying in America? It's actually quantifiable, and there have been multiple studies on that. Dave does a great job, obviously, of converting them into startups, into innovation.

The other thing that shoots us in the foot is that American companies can't invest in Chinese companies because of regulations and other things as well. Some of them are Chinese, but look at the trouble Benchmark got into for investing in Manus, for example. So I think it's 2-fold, but I completely agree that if you've created or contributed to creating a valuable asset, most foreigners stay in America after they do their PhDs, but too many don't have a very direct path, despite the math proving that they will add value to the American economy.

Peter Diamandis

Yeah. I mean, another point just to make here is that the AI race isn't only about chips and compute. It is about people. Key people are still driving the greatest value, at least for the moment.

Alex

I would argue it's not just people, but also, to my earlier point, it's about where the startups get domiciled. There's an alternative world where he, through whatever immigration-oriented regs, was deterred from starting his first AI startup in China while still an American PhD, and we incentivized him to start Recurrent here in the US.

I think there's maybe an alternative counterfactual world where Recurrent was American, and then its arguably intellectual successor, which is Moonshot AI, also remained domiciled in the US, and then he followed his own startup to stay here.

Emad Mostaque

Yeah. One thing that came out of the story is that when people come from India to get educated in the US, they overwhelmingly stay. When people come from China, about 80% of the time they go back, and it's just a difference in the local economy.

Going back to India to start your company is a nonstarter. It's just so unlikely to catch on. But going back to China—it's a thriving ecosystem, with lots of support—so going back to China to start your company is actually not a bad plan for a lot of people. I didn't realize that until this report came out.

But, as Peter was alluding to, we have tons of friends from MIT who came from China. I don't want to put them all in one bucket because there's a really clear distinction to me between people from China—Hong Kong, Taiwan, whatever—who come over and don't really align with the Chinese Communist Party at all. In fact, they kind of hate it.

Then you've got Chinese people who come over for an education. In one case at BU, a very good friend of ours is the dean of computer science at BU. There was a massive crisis because there was a concerted effort by the CCP to plant specific students into BU to gather specific knowledge. They were given tasks: "You have to go study this, learn it, and then send it back."

They didn't know what to do at BU. It's like, these are effectively trained spies who got into our PhD program, but we weren't ready for it. What are we supposed to do? They want to be highly ethical, so they don't want to just dismiss the students. I don't know how they resolved that.

So, that's a very different thing from the bulk of Chinese students who don't align with the CCP and just want to thrive in the world. They're happy to start their company here or anywhere else.

Peter Diamandis

And don't forget, when we looked at the frontier labs—I mean, originally, in the early days of xAI, for example, and at Meta—about 50% of their research staff, their research PhDs, were Chinese Americans. An extraordinary number. The Chinese, every year in the Math Olympiad, are at the top of the scoreboard. There's an incredible wealth of capability here that I think most companies desired to retain inside.

Salim Ismail

Yeah, 2 things here. One is, the asymmetry of the talent, I think, is the really important part here. I made this point a couple of podcasts ago: 70% of the elite AI researchers are not US citizens. They're, in order, Chinese, Indian, Taiwanese, and from the UK.

And so, that's a huge problem. Stapling in a green card is the easiest thing we could do, with zero friction, to give them incentive to stay here and build here. The US has a massive asymmetric advantage for the rest of the world. It was better to build here than anywhere else in the world, and that's starting to become less true.

That's why people are going back to China, going increasingly back to India, even to do things despite the friction that exists trying to do something in India. That is the part—the failure of the US to fix immigration is one of the biggest problems this country has right now.

Peter Diamandis

Amen. All right, I'm going to move us forward here. I just want to put up this slide. Since mid-April, we've seen 13 new frontier models launched, an average of 1 every 10 days.

Just comparing this to 2025, we had 8 frontier releases over the course of a year, 1 every 50 days. A year earlier, in 2024, we had 6 releases, 1 every 60 days, and it doesn't seem to be slowing down. Then, Emad, you sent me this morning this tweet from Elon. Thank you—I put it up here.

This is Elon's tweet: "Our 2-trillion model, which is better than our 1.5 trillion in every way, will finish initial training next week. It might be able to exceed Kimi, but with speed and token efficiency close to our 1.5 trillion, aka Grok 4.5."

So, I mean, this is the number one piece of evidence that we're living in the singularity. The speed at which this intelligence is accelerating is insane.

Emad Mostaque

Peter, it gets better. If you take Salim's list of frontier models and the dates and regress an exponential curve to the predicted frequency, or time period, between model releases—which I did just as an exercise—you find that, at the present rate, we're going to get to daily frontier-model releases by—wait for it—January.

By January, we're going to see daily new frontier-model releases if this exponential trend continues, which basically implies continuous versioning.

Peter Diamandis

So, I guess the question is, what does that really mean? What does it mean to have a new release if it's a continuous process?

Emad Mostaque

I mean, maybe it means that we'll have to do our daily Moonshots episodes about something other than point releases from the frontier labs. We'll need something new to talk about because it'll just be updated in the background.

Peter Diamandis

Like my son Jet said, "Okay, so another release, a little bit better. I mean, Dad, come on. What's really new here?"

Emad Mostaque

Well, actually, yeah. We'll see later in the pod some use-case demos, but I think those will take over because it's much more exciting when you see a tick up in the intelligence. You're like, "Yeah, so what?" It's, "Look what it made."

Peter Diamandis

That's what really gets people's attention. Let's take a second and just look at that because I skipped over it. But I think one of the things that's interesting here—and I'll just play these—is what we're seeing with gaming, like recreating your favorite game. On the right-hand side of the equation, we're seeing a browser-based web app simulating an Apple desktop. I think this is—we haven't talked about the implications for the gaming industry, which is huge, right?

Emad Mostaque

I'll stop that noise.

Peter Diamandis

You took over your computer there.

Emad Mostaque

It won't stop now. But what was fun in the last 24 hours was seeing everybody show their use of Kimi K3. It's impressive.

Peter Diamandis

Everybody becomes a creator. Everybody becomes a maker.

Emad Mostaque

One warning: one-shotting a game is very different from building the entire ecosystem, the customer service, and the marketing that goes around with it, et cetera. So you really have to be passionate about that domain. But the friction of getting a game launched for your personal interest, your fascinations, or a particular type of game that you want is near zero now, and that becomes really interesting.

Peter Diamandis

Yeah, exactly. It's mentally taxing because, if you take it to the limit—which is very soon—I can one-shot-prompt to create anything. And then you're sitting with your corporate exec staff saying, "Well, what do we want?"

Dave

Well, we've never had the ability before. We never really think this way. So then you have to stretch your brain: What's the purpose of our organization in the first place?

Peter Diamandis

Here's a thought. Historically on this pod, we've done calls to action to submit outro music videos. What about a call to action to submit an outro video game that people have just casually created?

Emad Mostaque

That's cool.

Peter Diamandis

Yeah. So, going to your point, Dave, I think having taste, having imagination, and understanding what the public wants—I think these become the scarce elements. For entrepreneurs out there, as you're seeing this capability, I think the entrepreneurial mindset and the ability to imagine something even greater... What happens when you're unleashed in what you can make, right?

Dave

Yeah, and visualizing happiness is something we're not used to trying. A lot of people don't manage their own happiness particularly well because they have to suffer through their daily job. They have to suffer through whatever mosquito bites and geography—it's just, you have no choice. Given choice, what would you do? That's so liberating for the mind.

But because we're not used to thinking that way, we're not ready for it. There must be an infinite number of things. The one that's easy for everybody is medicine and biotech: at least I want to be healthy. That's an obvious one. But what about all the other things that make humanity happy? Have we really thought through what we could voice or prompt tonight?

Emad Mostaque

Godlike. We are godlike in our abilities.

Peter Diamandis

And it's the name of your book.

Emad Mostaque

Yeah. Well, I mean, the idea is being a creator or a maker, right, versus a consumer or a taker. I saw your eyebrows go up. [laughter] What?

Peter Diamandis

No, I'm just agreeing with all of this. I think this is such a magical time to be alive. Everybody listening to this podcast, please think up some business idea, project, impact project, whatever, and use AI to go build it.

Emad Mostaque

Yeah. I just want to briefly—Nick Bostrom speaks about this a bit in Deep Utopia. Peter, you and I speak about this quite a bit in Solve Everything. I'll just outright suggest, folks, if listening—and I'm speaking just for myself—I would love to see an outro video game that you casually create. Maybe something in the theme of the Moonshots pod, since evidently we've completely solved and cooked music-video creation.

Dave

It'd be a first-person shooter game where we get to take aim at AGI—

Peter Diamandis

Oh no, please. No, no, no. Ideally, a nonviolent outro video. Nonviolent.

Dave

Civilization tech tree. That's what you need.

Peter Diamandis

I want to hit one point. Everybody watching and listening here, you have 2 options when you hear about this extraordinary ascent of Kimi K3. Fear might be one, and the other might be, “Oh my God, what an extraordinary time to be alive.” Hope, excitement, and an abundance mindset.

Rather than fear, realize you are being unleashed: your creativity, your ability to do whatever you want, and your ability to create.

Emad Mostaque

Your passion.

Peter Diamandis

Your purpose, right? Find what Sal and I talk about so much: your massive transformative purpose.

Just to distinguish between the 2, a passion is something you love doing. A purpose is something you love doing that actually benefits the world. If you can connect with that and realize that you can do it without any background, I think this is one of the most important things. You don't have to be a computer scientist. You don't have to be an expert. You have to be purpose-driven.

If you use these tools, you can make a dent in the universe. You can improve humanity at an awesome scale. That's what entrepreneurship is.

Emad Mostaque

3 steps.

Peter Diamandis

Yeah.

Emad Mostaque

Read Alex and Peter's paper, Solve Everything. Pick the biggest problem you dare to pick. Go download the Organizational Singularity cloud skill, which is free, and start building.

Peter Diamandis

Yeah.

Emad Mostaque

Awesome.

Peter Diamandis

Yeah. And a note for our production team, too: It's so cheap and easy now to do things like Alex suggested—to make a video game. We should collect and post some examples for the audience so that they can say, “Oh, that's what Alex was talking about.” Just a little roadmap is all people need. It can be this long.

If the audience doesn't send in amazing Moonshots-oriented video games as outros, I promise I will create a cyberpunk FPS. But it'll be a nonviolent FPS, if you can imagine that, oriented around—what are you shooting? You'll be tickling bunny rabbits.

Emad Mostaque

No, no. Okay. It'll be a cyberpunk FPS where we're cooking every problem. How about that?

Peter Diamandis

It's a first-person solver, not a first-person shooter. Love it. Oh God, that's great.

Emad Mostaque

You've got to do the tickling bunny rabbits, too, though.

Peter Diamandis

Okay, I'll tickle bunny rabbits. Fine. I'm going to move us to our next story.

And Emad, this is one you sent over the transom that I added here. If Kimi K3 is the frontier going big—trillions of parameters in a data center—this story is about frontiers going small, small enough to fit on your smartphone.

Bonsai 27B—B for billion—is the work of Prism ML. It's a U.S.-based AI startup out of Caltech. It's run by Babak Sabeti, backed by Khosla Ventures, Cerberus, and Google. It's the first 27-billion-parameter-class model to run entirely on a smartphone, not a stripped-down version. It's built on Qwen3-27B. Emad, tell us about this. Why is it important? We just talked about small language models with Liquid AI on our last pod.

Emad Mostaque

Yeah, this is one of Dave's favorite topics: quantization. Prism ML—and actually Tencent, which I'll talk about in a second—have had massive advances in being able to take a model that's been trained in a 16-bit architecture or an 8-bit architecture, like Kimi, which is basically 8-bit or 4-bit, and take it down to ternary, which is 3 values of information rather than binary.

Peter Diamandis

So ternary is 3 values, approximately 1.58 bits. 1.56. Yeah, 3 values, you geeks, you know.

Emad Mostaque

This is really, really important.

Peter Diamandis

Pay close attention, geeks. Really important topic.

Emad Mostaque

This is again the accuracy thing. What Prism managed to do is get the model down to ternary, which basically means—I think it was 6 GB for the model. This 27B model is really performant. I think it's basically GPT-5 class, from memory.

Peter Diamandis

Wow.

Emad Mostaque

With a 5% drop in accuracy, they managed to get it down to 6 GB, and with a 15% drop in accuracy, down to 4 GB.

Peter Diamandis

And you can get the accuracy back, too, by expanding the size of the network a little bit. Sorry—

Emad Mostaque

There are various things you can do. This is a big deal because it means you have a 1,020-IQ buddy that can work on your smartphone.

Peter Diamandis

Again, without an internet connection—

Emad Mostaque

Without a connection, it's smaller than a video game. You literally have this level of intelligence in your pocket all the time.

Peter Diamandis

But it's live. You can download the weights right now. You can run it on your smartphone.

Emad Mostaque

Exactly. Then away you go. But this is the super-interesting thing: When you reduce the bits, it also increases the speed. From 16 bits down to 3 bits, it's a 5-times improvement in speed.

There was another article, or another release, which is Tencent's latest model. This is actually the old WizardLM team, who had to leave Microsoft because Microsoft wouldn't give them compute. It's very ironic. They managed to get binary compression, taking it all the way down for their H3 model.

It's now the best on a DGX Spark or a big MacBook. You can take a 300-billion-parameter model—the size of the new model that's coming out of Inkling—and run it on binary with a 5% drop in performance.

Peter Diamandis

Wow. Yeah. And now I'm going to—sorry, Emad. This is so important. I'm only going to say this once on the pod because we're investing in a lot of companies that are working on exactly this. I don't want to tip it too much, but the implications of what Emad just said have one more step.

If I imagine that all of this intelligence under the covers—the computation going on—has always been matrix multiplications, I have a number, I multiply it by another number, and then I add 2 of those together. It's called a MAC: a multiply-accumulate. If one of those numbers is just 1, 0, or -1, I think we can multiply a number by 1, 0, or -1 pretty damn efficiently now.

That's the efficiency Emad's talking about, but it also opens the door for new ways to compute. What Alex has been saying for a while, if anyone listens through, is that we're going to discover new physics, but also new substrates on which we can compute. We're going to discover that computation is possible virtually anywhere—in crystals, in liquids—but the computation we're looking for is simply 1, 0, -1.

So it really narrows the focus on where we look for these computing substrates that'll take AI to the next level. What we're envisioning right now in the Dyson swarm is a bunch of GPUs from NVIDIA sitting in a satellite, with a solar panel and a radiator that's only going to last a couple of years. Something very different is going into space, something much more Star Trekkian, with crystals and holograms and things that are capable of doing the exact same thing.

You heard it here first, guys. So how efficient does this get? How compressed does this go?

Emad Mostaque

Oh my God. I normally agree with Peter on that. This is something I think about quite a bit.

The most quantized Bonsai model we were just talking about is approximately 1⅛—1.125—effective bits per weight. But you could ask the question: Is 1 bit per weight the limit? The answer is no. We can go below 1 effective bit per weight.

How do we do that? We do that with sparsity and quantization and low-rank factorization. By the way, that's what we're starting to see from some of the labs, in particular Samsung. For obvious reasons, Samsung wants to be able to host highly capable, frontier-class models on its own edge devices, like smartphones.

Just in the past 2 months, Samsung published a model called NanoQuant that breaks the 1-bit, 1-effective-bit-per-weight barrier. So it's sub-1-bit, which I think we're going to be talking quite a bit more about in the future, using a variety of tools.

This is my extrapolation episode. I went through the exercise of extrapolating frontier quantization out, and naive extrapolation finds that sub-1-bit quantization is going to go mainstream sometime in the next year.

Peter Diamandis

And then, overwhelmingly likely, photonics—the speed of light and photonics—will be the way we're computing in the future. Emad, you made a comment a couple of weeks ago that said we're going to get frontier-level capability running on a normal MacBook in 18 months, right? This is essentially the path you're talking about.

Emad Mostaque

Yeah, sorry. Please, go ahead.

Peter Diamandis

Go ahead.

Emad Mostaque

If you look at what NVIDIA did with their last NeMo series, they took the big model and actually distilled it, with logits, as they're called, down to a smaller dense model.

What's the difference between a 27-billion-parameter dense model and these really big sparse ones? When you have the model weights, you can actually do proper distillation, which is a bit different from the reasoning traces. What you're going to see is models like Kimi get distilled down to perfect datasets for smaller models that'll be trained at 4-bit and then cast down to ternary or binary, or even lower in terms of the bit weights.

When you look at Qwen Max versus Qwen3-27B, you can extrapolate what the sizes of these models will be as you move dense and go through the whole process. You end up with a model that works on 16 GB of RAM by the end of next year, at the level of Kimi K3. You can even extract all the knowledge out of Kimi K3 because it will be open source.

Peter Diamandis

Wow. So that means every vehicle, every robot, every manufacturing device, and every device in the world has its own built-in persistent intelligence and can make autonomous decisions at the edge for whatever task it can. So this decentralizes capability at the most infinite level.

Emad Mostaque

And it could go even 1 step further when you get down to ternary or binary. Actually, ternary is better for many things. You can build custom photonic silicon or even etch onto the silicon itself. The 0 is just—it doesn't have a path on it. So you can etch the model weights once they're good enough.

That leads to an actual increase in the total speed, and you don't need to use the smaller silicon anymore. So the cost of intelligence is going to drop by 100× anyway by the end of next year, just due to the new chipset.

Alex

Speed-running Star Trek. And I think this is what Dave was talking about earlier: as we move potentially to ternary or even sub-1-bit, it's far more ergonomic to adopt post-CMOS-type architectures underneath. There's plenty more room at the bottom.

Emad Mostaque

To answer the question, my bet there is that 0.78 will be the bottom, so I'm going to put that as a marker today.

Peter Diamandis

Okay, we can do our end-of-year predictions on that one. That's a really specific number. Do you want to go around quickly and ask everyone what their favorite quantization endgame is? Emad, it sounds like you have a bizarrely specific one.

Emad Mostaque

I'll post the details of that soon. We'll let everyone else have a think about it first, and then on a future episode.

Dave

I was going to say that the most likely forecast, based on everything Emad and Alex just said, is that we're expecting 100–10,000× within 3 years on just the raw compute through quantization and new compute methods. And that's multiplicative with the other algorithmic improvements. It's really hard to forecast, so realistically, a million times.

Peter Diamandis

Pause. Let's pause there 1 second, Dave, and just let folks absorb that for a moment. We've seen this incredible speed in performance and intelligence, and we're about to see what 10,000×—or, you know, adding algorithmic improvements—gets you: a million. What does that feel like over the course of the next 3 years?

Dave

3 years. One thing it feels like for sure is that AI is doing things that you really desperately want, but when it explains to you what it did, you just can't keep up. I'm already feeling this with Fable 5. I've got so many Fable 5 agents running, and the outcomes are exactly what I want, but it's like, “Well, what did you do?” And I can't get through it all.

Peter Diamandis

I had this conversation with Ray: the point at which AI is asking and answering questions that you can't even grasp.

Emad Mostaque

Yeah, I know that's very soon. So, to tie back to our Kardashev conversation, the idea of slowing it down is nutty. There's no regulatory concept of slowing it down that makes any sense. All we need now is some kind of global inspection and global partnership to monitor it, and then just take advantage of all the abundance that's going to come from it: all the new medicines, all the new capabilities, all the global happiness.

It's imminent. We just need to unleash it. Don't slow it down, but inspect everything. This whole mechanistic interpretability is going to become the most important thing that anyone can work on, and we just need global transparency and full throttle.

Peter Diamandis

I think this is one of the most important podcasts we've ever had, guys.

Emad Mostaque

Mind-boggling Sputnik moment.

Peter Diamandis

Sputnik moment.

All right. I'm going to move us along to another fun story, one that I love talking about. It's called “Predicting the Future.” There's a guy named Philip Tetlock. He's a political psychologist at the University of Pennsylvania who authored a book called Superforecasting: The Art and Science of Prediction.

After identifying what he called a group of superforecasters—ordinary folks who, through disciplined reasoning, consistently outpredict even CIA analysts with classified information—this is scored on what's called a Brier score, where lower is better. Now, the benchmark that pits AI against these superforecasters is called ForecastBench, and it's been tracking a steady year-long climb as models close the gap. We've talked about this before on the pod.

Well, the newest numbers have just come in, and according to the Forecasting Research Institute, for the first time, several AI models are now statistically indistinguishable from 4 superforecasters. So the implications are that if an AI can forecast novel events at a superforecaster level, then every decision that we make—in insurance, investing, policy, geopolitics, and corporate strategy—gets a cheap, tireless, superhuman adviser that's always on.

I find this fascinating. The data is out there, and an AI can gather it and make predictions. So, at the end of the day, every political decision is going to be modeled this way. Every investing decision is going to be modeled this way, and this becomes the differentiator. So who wants to jump in on this one?

Alex

I'll jump in. I absolutely love this to pieces. First, a few additional pieces of context. The number-1 AI superforecaster is from a British startup named Cassie, short for Cassandra, who, of course, made predictions but wasn't listened to. Interesting. It's founded by a British intelligence officer who served in Afghanistan and then advised the British government, and then formed this, in part inspired by superforecasters.

What I think is really interesting, though—we've spoken, when we've talked about these sorts of stories in the past, about Isaac Asimov's psychohistory and other riffs. I want to try a new riff here, which is an interesting thought experiment: what happens when hyperforecasting—not just superforecasting—is connected to capital markets?

What happens when the AIs—which are already AI algo traders, already completely dominating by volume public securities markets—have better internal autoregressive models of humanity than humanity does of itself? That's, in some sense, the same sense in which large language models were trained off the autoregressive task of predicting the next token of Internet text better than humans can. And now LLMs can predict, at least from a perplexity perspective, the next token I'm going to say in this sentence probably faster than I can generate it myself.

What happens when these hyperforecasters are able to generate the next actions by humanity collectively faster than humanity can take them? That's sort of the ultimate market-efficiency outcome, where literally—I think capital markets will be where this is maximally interesting, where the prediction is actually preemptively shaping the action of the market. And I think those who were so dismissive of the efficient-market hypothesis—I think the EMH is going to be crowned king of the capital markets once hyperforecasters like this are ultimately plugged in, which seemingly is imminent.

Peter Diamandis

I think this leads to wisdom. I think this is one of the most important things, and I've written a Substack on this. I've talked about it in the past. If you think about when you go to a wisdom council and ask, “What should I do?” you go to that wisdom council because they've had so many experiences in life; they can tell you, “Go this path—it's not going to succeed. Go down this path—you have a higher probability.”

So imagine a world in which everything's being simulated to the point where an AI can tell you what is the maximal path to take for world peace, to find your spouse, or to determine how to answer your kids. If you can literally simulate society on this level, we have a godlike support structure to help us navigate the decades ahead.

If I make this practical at an organizational level, think about most high-level management capabilities: budgeting, hiring decisions, product launches, investments in various things. Each of those is essentially a forecast, but you never predict—you never record the probability of that or score the accuracy of that. Once you have AI forecasting that approaches that capability, this means senior management essentially evaporates, because most senior management is there because they have deep expertise.

If you're the head of supply chain for BMW because you ran supply chain for Spain or you ran supply chain for that engine over decades, you built up experience to manage that domain. Once that judgment—which is hard to quantify—can be reproduced by an AI system without your biases, which are inevitable in human systems, that essentially wipes out all senior-management expertise. So now you need to focus even more on purpose, what you're trying to accomplish, and the objectives you have, et cetera. It completely changes the game for senior management in any company and any government.

Yeah, Emad.

Emad Mostaque

Yes. It's a topic close to my heart. In my bestselling book, The Last Economy, I actually describe how the mathematics of generative AI can apply to economics. Soon, we'll have a paper coming out that derives all of economics from the same math of generative AI—every single equation. It's kind of crazy, but one of the nice things here is—

Peter Diamandis

Even the incorrect ones.

Emad Mostaque

Even the incorrect ones, it shows them as limits and why they're incorrect, which is fantastic. But one of the interesting things in psychohistory—in Isaac Asimov's Foundation, he says that entire groups and populations can be modeled like gas. The equations of gas are the equations of diffusion models, which turn out to be better than humans at prediction. We're going to release a whole bunch of studies around that on economic prediction, where they're outperforming.

But then this raises something very interesting. You know, Peter, you said the wisdom—you know, Salim, you said no senior management. The way these models will start entering is through second opinions: medicine, business, and policy. But then the liability profile is going to go crazy.

Peter Diamandis

Matchmaking.

Emad Mostaque

Well, matchmaking, yeah. We have some dark things there, like Black Mirror and other things. But think about it this way: if you make a decision not approved by Dr. AI, your insurance premium goes up like that. If you drive and don't drive according to FSD in a few generations, your insurance premiums go up like that.

And that recursion is something that's super interesting because in Foundation, you had 3 requirements for psychohistory to hold. One of them was that the population is sufficiently large, and that can be like driving a car or entire economies. The next thing is a lack of technological advances of sufficient levels—the technological stagnation—because that can change the entire landscape of what's new. And the final thing was ignorance. [Laughter]

Alex just mentioned these things coming into the market and changing it, but these things coming into a healthcare decision, a government decision, or a company decision actually change the way it's like, “Hey, you're my match made in heaven according to the AI. How can you argue against the AI?” Worst pickup line ever right now, but who knows? In a few years.

Peter Diamandis

This is very meta, Emad—the sort of reflexivity in economics, I think many would call it. If the best predictor ends up being named after Cassandra and no one believes it, you can slice the irony with a knife. Dave, have you seen any startups in this area?

Emad Mostaque

They can make the money. It's okay.

Dave

No, shockingly no. Safe Superintelligence, SSI, may be a version of this, but they're keeping it in-house and launching it toward markets and printing money internally.

But the version I'd love to see very soon—I think a huge amount of human unhappiness comes from consumerism and consumer marketing. Homer Simpson comes home at 6 p.m., cracks open a beer, lies down on the couch, and starts channel-surfing. Then Naked and Afraid is on, and he ends up watching it until he falls asleep on the couch. He wakes up the next morning with a hangover, having not brushed his teeth, kicks the dog, and ends up with unhappy kids.

That chain of decisions is so bad, but there's no explicit decision to live that life in that chain, right? You just reacted to the beer ad, and then you went down this chain. I think AI is going to be an incredible coach to say, “Hey, dude, you know what? What if you take this alternate path, and here's the outcome you're going to get to?”

That, to me, is forecasting used correctly, just for changing: are we anywhere near optimal? The answer is no. If you objectively look at your life, nobody's near optimal. But with a little AI assistance, you can get on a much better path.

What we do right now is massive consumerism. You're reacting to billboards. You're reacting to TV ads. It's telling you that you think you need certain things, and people tend to get sucked into these pathways. I think we can get out of those pathways with AI.

Peter Diamandis

Dave, that's brilliant. Just to say, first of all, there is a rumor out there that Ilya's SSI is going to release something very shortly. I think everybody's feeling the pressure to release. We saw that with Mira coming out, so it'll be interesting to see.

But the point you made is brilliant: are these labs actually pulling their punches, holding on to this capability to generate revenue on their own? If you had this super-forecasting capability in the markets today, you would do that.

I remember having a conversation with Eric Schmidt, who said, “Listen, if Google wanted to maximize its income, it knows exactly which companies are going to have a stock bump in the fourth quarter because everybody's Googling this product or that product. We have advanced information about where the sales are going to be and which products are going to peak. But if we could only do that once, then we'd be shut down.”

It'll be interesting to see if these companies—and, Alex, you and I have talked about this—the notion in Solve Everything that the greatest money, the greatest income, these frontier labs are going to make is going to be as they solve scientific breakthroughs: superconductivity, age reversal, and so forth.

Alex

Exactly. And maybe just a footnote on the Google story: I've had this conversation with Google executives many, many times over the years. I totally agree with the premise that if Google were to attempt stock trading based on arguably insider or unfiltered insider information passing through the query stream, that's a one-and-done type shutdown scenario.

But there are other things that Google, hypothetically, could be trading besides public securities that wouldn't necessarily have the blowback. For example, again hypothetically, foreign exchange rates.

Peter Diamandis

Yeah, and I think that you have to be careful here, though. I think there's the market side of things and maybe—maybe not—I will launch a hedge fund based on our own stuff, but there's the moral side of things. Maybe not. Maybe not invest, okay?

Emad Mostaque

Of course, but at any rate, look, there's the moral side of these things. It's fantastic that we can optimize ourselves, but who controls these models and the advice they give can control vast waves of humanity, and there needs to be a real discussion about this.

We're going to rely on these far too much. And again, how can you debate it in just a few years' time? It'll be more expensive not to do this. You will be penalized for not listening. And if we're all watched over by machines of loving grace, we need to know whose grace that is. Again, that discussion needs to start now.

Peter Diamandis

Yeah. See? Just beer. Homer drinking beer advised by AI was not on my bingo card for this episode. That's all I'm going to say. [Laughter]

Welcome to the health section of Moonshots brought to you by Fountain Life. You know, my mission is to help you use the latest technologies, including AI, to not just do your work at home, teach your kids, but to help you live a long and healthy life. I'm here today with an extraordinary physician, the chief medical officer of Fountain Life, Dr. Don Mucalem. Let's talk about cancer. I know from the member database we have at Fountain Life that members come in thinking they're healthy. It turns out 3.3% of them have cancer in their bodies that they don't know about.

Dr. Don Mucalemon

That's right. The majority of cancers that we screen for aren't necessarily the ones that are taking lives when found at a late stage. We know that when cancer is found early, the chances for cure are much higher. We know it's much easier to treat a cancer when found early versus when found late.

What we're finding in our members is that over 3.3% were found to have these cancers that otherwise wouldn't have been found or detected.

Peter Diamandis

Yeah. It's interesting. People don't feel cancer until stage 3 or stage 4. If you don't know what's going on inside your body, it's like driving your car with your eyes closed. And so, when members come through Fountain, how do they detect cancers?

Dr. Don Mucalemon

We're doing full-body MRI, and we also do early cancer-detection screening. This is very, very important, and these are not typical tools used in the conventional care setting when it comes to prevention.

This is a hard thing because currently these are not studies that insurance would yet be covering. But the goal is to collect these numbers, do the research, and work hard to democratize wellness.

Peter Diamandis

Yeah. So at the end of the day, you can know what's going on inside your body. It's your obligation to know.

So check out Fountain Life. You can go to fountainlife.com/pater to get access to the latest technology to help you detect cancer at the very beginning, at stage one when it is curable, before it gets to stage three or stage four in your world of hurt.

So, Emad, you sent me an article—a chart. I just put this up here right now. This is our constant debate, and we're seeing this again across data-center wars in the United States. Data centers are sucking up electricity, driving up the cost for consumers, and also water. It's one of the loudest criticisms of AI right now: data centers are guzzling drinking water to cool their servers.

This week, this particular chart that I'm showing made the rounds, and it pairs 2 figures. On one side, every data center in the entire U.S., according to Lawrence Berkeley National Laboratory, is consuming 17 billion gallons of water on-site. But what it shows is that American golf courses have soaked up 531 billion gallons of irrigation since 2024. That's 31 times as much.

The posters I'm going to start seeing on the sides of the highways are going to say, “Forget data centers.”

We must ban golf courses immediately.

Emad Mostaque

Yeah. Where’s Peter? Where’s the Chinese influence campaign to get America to shut down its golf courses?

Peter Diamandis

Yeah, I tell you, I don’t see it anyplace. But here’s the shocking piece of data: besides golf courses, California almond farming alone consumes 1 trillion gallons of water—60 times all the data centers combined.

Emad Mostaque

I have one other stat—

Peter Diamandis

Please.

Emad Mostaque

Amazon warehouses occupy 10 times more land in the U.S. than all the data centers combined.

Peter Diamandis

Yeah.

Emad Mostaque

So it’s such a drop in the bucket compared to everything else in terms of land usage and water usage. The hue and cry is such completely non-data-driven garbage. It’s unreal.

Peter Diamandis

Well, exactly. That’s the concern, because the water use is such a nonissue. It’s such a joke. But if we take that head-on and say, “Guys, don’t worry about water,” the angry crowd is going to move to something else equally irrational.

The underlying problem doesn’t go away. The next issue is going to be something semi-insane—this is completely insane, but something semi-sane and still wrong. That’s going to create a populist movement. The word “moratorium”—let’s just stop. What kind of a decision, what kind of governance, is “Let’s just stop”?

But if you look at the history of nuclear and a whole bunch of other things, that’s the actual outcome we get. And so, David—

Dave Blundin

I mean, this is the pandemic of fear that I keep on speaking about, and that I’m very concerned about. There’s an underlying sense that AI and robotics are going to combat humanity, that they’re going to be our foes.

Again, I’ll just go back to it: I blame, to some degree, Hollywood. With all the dystopian movies out there, if all you see is negative visions of the future, you’re going to want to shut it down. What do you want to shut down? How can you shut down AI? Well, you can shut down the data center in your state.

Peter Diamandis

Yeah. Also, that elephant in this particular room is the Dyson swarm. If all the compute moves to sun-synchronous orbit, you can do closed-loop liquids, including water and other coolants, there. It’s not like it’s going to be consuming, on the margin, additional water.

And to Dave’s point, the complaints—which may or may not be, in part, the result of an influence operation from a foreign state actor—will move to something else. It’ll be very low-Earth orbit: SpaceX Starlink and other competing Dyson swarms are polluting the atmosphere with their decay, or something else. The complaint will move on to something else.

Did you hear the rant about the Starship rocket launches earlier? It was Falcon, actually—the pollution from the Falcon launches. Elon was just like, “Oh my God, I’m going to vomit.”

David Friedberg

It was like 0.00001% of all emissions from any form of rocket launch. He’s like, “But you have to actually answer these questions.” It’s driving him nuts.

I hope those individuals who are complaining have thrown away their smartphones, don’t use GPS, and are basically going back to subsistence farming.

Peter Diamandis

Yeah. As Elon likes to say, let them shake their fists at the sky.

Emad Mostaque

I have a fun stat. I was doing some numbers around the water thing. It’s about 600 gallons of water per Big Mac, and McDonald’s sells 2 billion burgers a year. So it’s about twice the total amount of water that golf courses use.

Peter Diamandis

So that I can get behind. Okay. So what you’re saying, Emad, is the Chinese influence operation should also be shutting down American Big Macs.

Emad Mostaque

Well, there you go. It’d be a big stab to the heart of America.

Peter Diamandis

That’s right. Definitely improve the health of America as well. Shall we move to one of our favorite conversations: humanoid robots?

Emad Mostaque

This is so cool.

Peter Diamandis

Yeah. China, as we’ve discussed before, has gone all-in on humanoid robots. It’s a national priority. Companies like Unitree and others are racing to commercialize.

In the last report—and, Alex, we’ve talked about this—there were 150 humanoid robot companies in China under development. Part of their strategy is spectacle, and it’s something you’re trying to bring, Alex, to America. They’ve been staging public robot combat events, literally MMA-style.

We’ve got a video to show. Let me pull this up here. Here’s a recent MMA match that went viral on the internet, and it’s a beautiful thing.

Emad Mostaque

Just so cool. [screaming] These are only going to get better.

Peter Diamandis

You’ve got to watch the full video. The way the fight ends is epically awesome.

Yeah. One of the robots kicks the other robot’s head off. Remember Rock ’Em Sock ’Em Robots?

Emad Mostaque

Yeah, yeah. [laughter]

Peter Diamandis

As a game, as kids. This goes viral. There’s a lot going on in the robot world. We just saw all of the workers at Hyundai start to strike because they don’t want robots brought onto their assembly line. That was fascinating. Alex, take it from here.

Alex Wissner-Gross

A few thoughts on this. I have thoughts on many different levels. One is mild horror. If anyone’s seen Steven Spielberg’s movie A.I.—without spoiling it too much, I think Steven would call it the dark sandwich at the center of the movie, the Flesh Fair, where humanoid robots are tortured and abused for human entertainment—I think that’s utterly horrifying.

At one level, I’m mildly horrified that humanoid robots, no matter the extent to which they’re being teleoperated here, are setting an inductive prior or bias for future, more autonomous embodied intelligences to be basically trying to kill or otherwise physically abuse each other for human entertainment. I’m concerned about that.

But one level deeper, now imagine that these robots are more autonomous, that they’re running algorithms at the edge, so they’re much more encapsulated. Now imagine that these humanoids are in the Chinese PLA infantry.

Peter Diamandis

Yeah.

Alex Karp

I think that’s the future that we are almost certain to find ourselves in. The West needs to catch up in humanoids. That’s why I’ve supported ProRL, which, Peter, you were gesturing at. It ran its first humanoid robot mini-marathon in America in the Boston Seaport a number of months ago.

The West needs something like this—hopefully less violent and more economically productive. I’d love to see people cheering on humanoid robots competing to iron clothing or perform some economically productive task, and not just kicking each other’s heads off.

Peter Diamandis

You prefer the humans to be doing that in the MMA matches?

Alex Karp

I’d prefer no one to be doing it. I’m not a fan of MMA. I think it’s destructive to humans, and I worry about the message that we’re sending to the future light cone by having robots do it instead of humans.

I’d rather see people in a cage competing, if they must compete at all, to do something positive, not negative.

Peter Diamandis

Coding, like a cage-match coding competition—

Alex Karp

If anything.

Peter Diamandis

Or just sitting there. Okay.

Emad Mostaque

A couple of thoughts. My normal commentary around kickboxing is that it’s not the greatest marketing demo for humanoid robots. But I will acknowledge something here: this is an unbelievably demanding engineering environment.

You’ve got a stress test. It’s stressing balance, impact resistance, recovery, locomotion, and latency. There are 20 things that they’re doing. It’s kind of incredible to watch them navigate that. Of course, a 4-armed robot would beat a 2-armed robot. [laughter] I’ll just leave it at that. There we go.

This is competitive. We’re going to see this go to competitive sports. We’ll see a version of the World Cup with robotics. The question is whether people will watch that or not.

Peter Diamandis

Yeah. I’ll say that the real test is whether a human being can make that penalty shot under pressure at that point in the game. Watching England implode the other day was really devastating for me, but still, I think people would much rather watch people in that environment rather than robots.

Sports is going to thrive for many, many decades to come. Formula racing pushes the edge, and I think when we start to see robotic sports, it’s pushing the edge. I think the point you just made is important: we’re going to see this happening in a competitive fashion.

I can’t wait to see Figure versus Optimus. I think that will be a fun competition, whatever form it takes.

Emad Mostaque

Yeah, I think these robots are a little bit different, though. You’ll probably first see the Real Steel-type teleoperated robots, because robots can’t actually respond fast enough if you look at the latency of a VLA model.

Peter Diamandis

This is impressive with some teleoperated flying kicks, but why aren’t they doing kung fu? When will robots do kung fu? That’s when you move to things like edge silicon, when you move to teleoperation. I think that’ll be the next stage that comes next year.

But I think there’s a bigger issue that I have with this. Although I love fighting robots and I can’t wait to see Gundams and all that, these robots are EngineAI T800s. They weigh about 70 kg and they punch 4 times harder than Mike Tyson.

So they could legitimately kill someone—us fleshy humans.

Emad Mostaque

Robots like that should not be allowed on the streets, and there’s no regulation against that. They could be in the PLA—the People’s Liberation Army—or whatever, but robots are about to enter our households. Who here has a 1X robot on order? Come on—it’s coming. They will be walking around very soon.

We need to have regulations about the safety of these things, what the torque is on them, how they operate, and so on, because they represent a real threat to individuals. They are machinery. Beyond that, you have embodiment and other issues. We need to have the discussion of what that looks like when they are autonomous, because these things are delivering themselves by pushing a button on the door, ringing your doorbell.

The final thing is that Unitree has only made 11,000 humanoid robots in total. We are literally at the very start of this.

A few years from now, it will be 11 million a year instead of 11,000. We’ve got to have this discussion fast as well. Lots of talking to do.

Peter Diamandis

Yeah. This is the work you and I were doing in terms of how governments counsel their policymakers around these areas, and it’s happening at blinding speed.

Emad Mostaque

Crazy.

Peter Diamandis

Yeah. All right, I’m going to move us to the most important conversation we always have, which is the Dyson swarm. [laughter] Let’s take a look at a video from our friend Sam Altman.

“I honestly think the idea, with the current landscape, of putting data centers in space is ridiculous. It will make sense someday, but if you just do the very rough math of launch costs relative to the cost of power we can generate on Earth, to say nothing of how you’re going to fix a broken GPU in space—and they do break a lot still—unfortunately, we are not there yet.

“There will come a time. Space is great for a lot of things. Orbital data centers are not something that’s going to matter at scale this decade.”

All right, we have the continuing MMA battle between Elon and Sam.

Emad Mostaque

Yeah, fascinating. I’m curious about reactions here. Alex, I’ll go to you first.

Alex

I think there’s an obvious conflict of interest. We saw similar messaging from Masayoshi Son regarding the lack of purported promise for orbital data centers.

Remember, OpenAI has retreated from its own data centers. Remember Project Stargate? Project Stargate has been rebranded from OpenAI owning and operating its own data centers to just leasing terrestrial data center capacity from others. OpenAI is delaying its own IPO.

One has to look at OpenAI’s messaging here and say that perhaps it’s not even in a financial or operational position at the moment to lean into orbital data centers, the way Anthropic, in its collaboration agreement—which was announced with SpaceX’s xAI for use of Colossus and Colossus 2—is far likelier to move toward orbital data center–based compute.

I think the crossover is going to happen. Elon’s messaging regarding when this crossover is going to happen is 2 to 3 years. You see other analyses that suggest the unit economics for orbital versus terrestrial data center costs are going to cross over sometime by the early 2030s. I’m not sure which is the case, but either way, I think there is an obvious conflict of interest.

Just as we were discussing with Philip Johnston, barring some surprising left turn, I expect that OpenAI’s tune is very conveniently going to change on ODCs sometime in the next 2 to 3 years—right on time.

Peter Diamandis

And, of course, Elon’s response to this is, “We’ll be launching them in 2 years.” So just stay tuned and watch.

Emad Mostaque

Well, I think anyone listening to this video would say, “Okay, Sam says space data centers make no sense. Elon says they make sense. The 2 guys hate each other.” But if you actually listen closely to Sam’s words, they don’t disagree at all.

Sam is saying that space data centers will not be meaningful this decade. There will come a time, but this decade has only 3 and a half years left. If you look at Elon’s forecast of his launch rate, they actually agree. They’re just hating on each other all the time, and it seems that way in this phrasing, but the truth is pretty clear: They both have the same numbers.

So Alex is right. They’re going to space, and it’s going to take a while. I think a couple of percent of all compute will be in space by the end of the decade, because we’re building out on land as quickly as we can, too.

Peter Diamandis

But then the lines cross.

Emad Mostaque

Yeah.

Peter Diamandis

Yeah. You know, Alex, you and I were going back and forth texting while the Starship Flight 13 attempt was being made a couple of days ago, and it’s been rescheduled. When this podcast comes out, we’ll be seeing the next launch attempt of Starship Flight 13 on Monday of this coming week.

That launch was thwarted at T-minus 0.

Alex

First time I’ve ever seen that, by the way.

Peter Diamandis

Yeah. Here’s the point: 2 of the 33 Raptor engines on the booster stage of Starship did not ignite, and they’re going to be replaced.

By the way, SpaceX’s stock dropped 5% on news of that failed launch, which is kind of ridiculous. The point people need to realize is that was an amazing demonstration of technology. The fact that you could shut down at T-minus 0, safe the vehicle, and unload the methane and liquid oxygen—I was part of the space industry in the ’90s, before it was a space industry, and those vehicles would have exploded on the spot. They would have failed on the spot.

The ability we have to control them at that level of detail is evidence of the extraordinary engineering that SpaceX has done.

Alex

I thought that was the most interesting part: how quickly the system diagnoses the problem and returns. It would have taken months and months to do this, fix it, recover everything, and replan another launch. You’re like, “Yeah, problem. Shut it down, redo it. Oh, we’re starting Monday.” It’s amazing.

Peter Diamandis

Yeah, extraordinary.

Emad Mostaque

Yeah. Because I think if you’re serious about superintelligence, with what we know, you have to have a space play. OpenAI is going to buy Planet Labs or something like that, and then the tune will change.

Peter Diamandis

All right, I’m going to go to some AMA questions. Emad, you had suggested I post questions to X, and we have a number of questions coming about Kimi from our X audience. Let me go ahead and show these, and let’s dive in.

Emad, I’m going to give you first crack. Which of these questions do you want to answer?

Emad Mostaque

I think number 4 is probably an interesting one: Given Kimi K3’s lower token efficiency, is it actually as cost-effective as advertised compared with Solo Fable?

Kimi K3 is an expensive model relative to the other Chinese models. DeepSeek is now $1 per million tokens. Kimi K3 is $15. Sonnet is $20, Opus is $40, and I think Fable is $60. But that’s because they’re actually making money.

When you back out the numbers from the Chinese models and the chips they’re running on, they’re probably making 80% to 90% margins now. That’s with their Chinese chips, which aren’t that efficient for running this.

We will see the cost of K3 drop by 10 to 50 times, I think, in the next few months as it gets optimized. Right now, it uses twice the number of tokens for the same task versus GPT-5.6—a frontier model that uses 37% fewer tokens than 5.5 or Fable.

Again, we’re going to see that drop because everyone and their dog is going to optimize the crap out of this. You’ve seen Fireworks just raise at a $17 billion valuation. Others, like Modal at $10 billion and Baseten at $10 billion, are the inference providers of open-source models. They’ve all raised $1 billion that they’re now going to spend to optimize the Chinese model and make it more efficient and run it.

American labs that do the inference side of things are going to optimize the crap out of this. We will see it catch up.

Peter Diamandis

All right. By the way, I welcome the guests to lean in on these questions. Selene, you want to go next?

Salim Ismail

Given that I made the comment about number 1—how much could Kimi K3 devalue U.S. frontier models?—I’ll stick with my original estimate of about 75%: 50% from the U.S. regulating the front end, and then you’ve got a lack of compute on the supply side, plus frontier open-source models kind of within a release, barely, of where you are.

That bleeding edge is such a perishable thing. I would say a 75% drop. So if OpenAI is worth $1 trillion, I’d put it at $250 billion.

You still have a very valuable business, because now the competitiveness is about reliability, security, integrated tools, and ease of deployment. But the actual frontier cutting edge becomes one ingredient among the whole thing.

Peter Diamandis

I would not want to be inside these frontier labs right now. It must be a frenetic code-red, 24/7.

Emad Mostaque

It is a total rat race. I have so many friends at the frontier labs—friends who are jumping, hypothetically, from one frontier lab, Google, which is nowhere at this point, missing in action, to other frontier labs. It is a total rat race.

Peter Diamandis

Yeah, it’s crazy. Dave—

Dave

You have a choice for me.

Peter Diamandis

No, pick one. You’ve got 2 and 3, I think.

Dave

Okay, I’ll take 2: What does the release of Kimi K3 do to the open-source versus closed-source race? Will this force the large companies to provide more open-source products? I think they’re implying more open-source products.

Yeah, it’s a total game changer in the sense that anyone with resources can build an internal model that’s tailored to a specific use case and then use it as a defensive moat.

Dave Blundin

I don't think the large US model providers will go open source. I think they're committed to their pathway. So if you were talking to Anthropic right now, they would say, “Look, Kimi has caught up for a week, but Fable 5.1 is coming out in just a few weeks.”

When you look at the all-important enterprise use cases—white-collar automation, drug discovery—people are going to use the best model no matter what. If you're using an AI to design a car or a rocket, a slight improvement in the design has a massive payoff. So you're going to use the best of the best of the best model. The Anthropic guys are going to scramble to stay a step ahead and keep their price point nice and high. The cost of the model itself is so small compared to the benefit that people will pay the price.

Peter Diamandis

So it does create, like Alex was saying, the rat race is incredible, but people aren't going to switch to Kimi unless it's proprietary data they want to keep in-house and they want to tune their own, or Kimi actually bypasses Anthropic—which it hasn't done. It's only caught up, or not even quite caught up.

All right, Alex, number 3.

Alex Wissner-Gross

All right, number 3 asks—and I think these questions seem to all be variations on a theme—but it asks, “How can US models—I think this means US frontier model providers—continue to justify their massive valuations if China can leapfrog with an open-weight model at less than half the token cost?”

I don't think the premise is quite accurate. There are so many elements, so many layers to superintelligence, and quite frankly, superintelligence itself, as it fully develops, I think is far larger than the total GDP of the entire world anyway. There's an enormous amount of pie that can be sliced.

To the extent we're talking about, say, Google, which, as I was mentioning earlier, seems to be MIA at this point on the frontier—I can't find a single top Google model at this point on the cost frontier for capabilities—what does Google do? Well, they can continue to race, obviously, in terms of capabilities. But if I'm Google, I'm thinking, “Yeah, I want to become a hyperscaler.” I mean, Google obviously is a hyperscaler, but a hyperscaler provider to other frontier labs. That's one obvious venue of differentiation.

We've seen that approach vector from SpaceX AI itself, which has now signed deals with Anthropic. We're seeing it with Meta, interestingly, which, on the one hand, is offering Spark 1.1 and, on the other hand, in the past 2 days, just as we were going to air, it was announced that Meta is exploring selling $10 billion of compute to Anthropic. So differentiating by going down-stack and offering your compute to other, more competitive providers—whether Western, usually Anthropic, sometimes OpenAI, or Chinese models in a self-hosting model—that's one area.

You can also go up-stack. You can try to vertically integrate and offer applications that are benefiting from the commoditization of their complement, namely the model layer. I also think the premise that valuations somehow are going to net shrink just because Kimi K3 exists now is completely fallacious.

We saw that incorrect thinking happen with the original DeepSeek shock, which was at the time also branded as a Sputnik moment. We saw a bit of a hiccup in capital markets at the time, but, as always, Jevons paradox kicks in, and we see the value of chip stocks ultimately increase, not deflate.

We also see that it's open. It's open-weight, so there's absolutely nothing in Kimi K3 that OpenAI, Anthropic, and other Western frontier labs can't immediately reappropriate for their own internal models.

Peter Diamandis

You don't think that the amount of revenue these labs are going to make gets reduced as people start to use Kimi K3 for their work instead of API calls?

Alex Wissner-Gross

No. For example, I spend, and my portfolio companies spend, an extraordinary amount on, let's say, Anthropic and OpenAI. To my knowledge, my expectation is Moonshot would have to release a 2×, 3×, or 10× better model than, say, Fable 5 to have a massive diversion of that spend.

Right now, what K3 buys, to the extent it's legal—query how much longer K3 will be legal to host within the US—but assuming it remains legal and regulatory-uninhibited, all it results in is greater in-house self-hosting. But it's not at the top of the frontier. To Dave's earlier point, Fable 5 is ahead at the moment; if you're trying to solve the frontier of problems, K3 is not causing you to divert your spend.

Peter Diamandis

Well, let me hit that point you just made, Alex, and ask you and the other mates a question here. Do you think it's possible that some legal policy in the United States prevents US companies from downloading K3? It's going to be on the open internet. It's going to be available through a multitude of sources beyond Hugging Face. Can it be shut down in the US?

Alex Wissner-Gross

It can effectively be shut down. This is not prescriptive, and I'm not a fan of this policy, but I think it can effectively be shut down by requiring that every public corporation disclose any use of Chinese open-weight models and subjecting them to scrutiny.

As we were going to air, the latest—we talked in the last pod about Demis's proposal to create a FINRA-like entity that would regulate the frontier. Well, guess what? The reports are that the present administration is actually running with a proposal like that and is planning to, or at least exploring, creating a FINRA-like agency to regulate frontier AI that would live under the SEC, because the SEC already has statutory authority to operate FINRA-like, industry-advised-and-funded entities. So it's a natural place organizationally.

Peter Diamandis

Yeah. Self-regulated governance, a.k.a. regulatory-capture cartels under the SEC. I think it's completely plausible, albeit highly undesirable, that we get, sometime in the future, an SEC suborganization that looks like FINRA and basically makes it completely economically infeasible for corporations of any size, especially public corporations, to actively use Chinese open-weight models.

Any other comments on this?

Emad Mostaque

I've got a comment on this. I mean, this is ridiculous in terms of trying to limit the use here, because once you release the weights, you can mirror them across jurisdictions. You can use peer-to-peer networks and VPNs. All you're going to do is deny American researchers, startups, and security experts access to those models, while the rest of the world goes ahead building on those models. I don't think there's a viable approach. I mean, this is the same—

Peter Diamandis

Yeah, please.

Emad Mostaque

This is the same as denying Americans cheap insulin. I mean, it's again regulatory capture, right? Like, why can't you have generics? Because, again, you have the regulatory-capture point.

There's Operation Gold Eagle, I think they're calling it, to approve access to frontier models. You will have anti-token-laundering regulations. You will have know-your-prompter regulations. The US government has really realized that this technology is about to break through, and I think they're a lot more worried about it than China is.

You look at that Xi Jinping speech. I would urge everyone to check it out. They're full-on open source: “We're going to do this.” America doesn't know what it's going to do, but, as you said, there's a real chance that they might hobble American capitalism. Oddly, China's encouraging capitalism.

Peter Diamandis

CCP saves American capitalism from itself. That's a crazy future. The world is so weird.

All right. Let's go back to you, Salim, on the next question.

Salim Ismail

Which one?

Peter Diamandis

Some of these are a little bit duplicative.

Salim Ismail

Yeah. I'll take number 5. Would you trust Kimi K3 to write your code for you without oversight or review?

The answer is no, but I wouldn't trust a human being to put consequential, untested code into production either. The question is not whether we trust the models; it's whether we trust the development system around them. AI-generated code needs to be run in a sandbox, pass automated tests and security scanning, and go through all sorts of things before it goes into production.

Then you do proportionate permissions based on the use case and the potential impact. Whatever the workflow is that AI is running, you're still going to need human review at the highest level and for the highest-consequence inputs. A lot of the routine can be automated, but the scalable model is not AI with no oversight; it's machine-generated plus verification plus human accountability combined. That's going to give you the real power.

Peter Diamandis

All right. Iman.

Emad Mostaque

Yeah. What role, if any, did distillation play in K3 development? They distilled data clearly from Opus and others, but, to be honest, using Kimi K2.5 and Kimi K3 now quite intensely, it feels different.

I think they did a lot of their own data creation based in part on distillation, but everyone's distilling from each other right now. The one area where it's clear that they've had a big leap ahead is in front-end development. Again, this isn't the best mathematician in the world, although it's quite a good general model. It's not the best cyber attacker from our benchmarks, but they've done something original and new on the front-end, consumer-entertainment side of things, which I think is really interesting. Although that might also be because it's a multimodal model.

Peter Diamandis

Mhm. Mhm. Dave.

Dave Blundin

Number 7: What are the reasons why Kimi K3 might not be as good as advertised, or why we shouldn't use it?

The scenario where it's not as good as advertised is if it's benchmark-maxed and, in 2 weeks, the open source will be out.

Alex MacCaw

We'll have beaten it to death. We'll know the answer if they benchmark-maxed it, so we're going to find out. I think it's unlikely that it's benchmarked to the point where every company in America, every company in the world, should be saying right now, “We need a crash program with our best possible adviser to decide: Are we going to do our own model on our own on-prem hardware, or are we going to use Anthropic, OpenAI, or Google and just trust that API?”

But we need to decide whether tuning and training on our own proprietary data gives us a long-term competitive advantage. And so there's going to be a desperate shortage of good advice on this, and vendors and McKinsey consultants, and you've got to grab those resources quickly. ExO consultants, make seed-stage investments, get your network together, find out who can answer that question for you internally, on your business and your use case, quickly, and then commit to the path.

And you can do something internally and still use the APIs, but if you don't start down the path of evaluating Kimi K3 on your own, you can't really come back to it later. So I think everybody's got to just get going on this question. We'll know in a couple of weeks, though, whether it was benchmarked to hell or not. But I think it's very, very likely that the open-source path is a viable path for every U.S. and world company and government.

Peter Diamandis

Can I just add to that real quick?

Emad Mostaque

Yes, of course. Very simple suggestion for every company: implement 2 installations, Kimi K3 and Inkling. Fine-tune your own internal data, because that learning loop is going to be the proprietary gold that you don't want to lose. And start there.

Peter Diamandis

Alex, why don't you close us out here? You've sort of answered number 6 already, but perhaps you could expand on it.

Alex

I'll say something new. Question 6 asks, should the U.S. move to block loading the weights of the next Kimi release onto Hugging Face? I'll give a conditional answer. I think that if some party, presumably in the U.S., can prove to a competent court that the next Kimi release—presumably a reference to this Kimi release—was somehow obtained or derived illegally, maybe through copyright infringement or illegal distillation of traces or something like that, that would probably be grounds for blocking its release in the U.S.

But if no one can prove that Kimi's parent, Moonshot, did anything otherwise wrong in creating it, no, I don't think the U.S. should be blocking its release in the process. I think, if anything, quite the opposite. I think every U.S. frontier lab should be closely scrutinizing it and learning whatever they can so that we can leapfrog it.

And I would like to see far more outward pressure from U.S. labs creating the best-in-the-world open-weight and open-source models, so that it's not the CCP with their new Belt and Road for AI initiative blanketing the world—some would even say dumping superintelligence on the rest of the world, or the so-called Global South. It should be the U.S., the cannon of freedom, the arsenal of freedom, that's also the arsenal of superintelligence, showering the rest of the world with open-weight and open-source superintelligence—not China.

Peter Diamandis

Showering the rest of the world—I love that. And remember, we're moving toward intelligence that's too cheap to meter, but 1,000,000 times more available and more powerful than ever before. Everybody listening, I'm grateful on behalf of the Moonshot Mates here for your time. We're going to be putting this out more and more often as we're starting to see the release dates move from months and weeks to days. There's no time to sleep during the Singularity.

Gentlemen, what's in store for the week ahead? Emad, I'll go to you next.

Emad Mostaque

Yeah, just getting a whole bunch of research papers ready to release. So finally, it's going to be exciting.

Peter Diamandis

Again, acceleration for Intelligent Internet, your company?

Emad Mostaque

Yes.

Peter Diamandis

Incredible.

Salim Ismail

Tuesday, I have my next Meaning of Life session at 7:00 p.m. Eastern.

Peter Diamandis

Alex, are you coming up? What's going on with you?

Alex

I'm so focused at this point on literally solving everything. I'll say large swaths of the sciences at this point, I'm convinced, are so thoroughly cooked. More to come on that subject. Peter, you and I wrote “Solve Everything” about it, but now it's actually coming true.

Peter Diamandis

I'm excited. You're going to be doing an AMA with my Abundance community coming up. That's going to be a fun deep dive. And, of course, we're going to have you during the Moonshots gathering on September 25. In fact, all of us will be here. Emad, you're joining us in L.A. in September.

Emad Mostaque

Yeah, it's going to be fun to have all of us together again for the full day.

Peter Diamandis

Dave, this has got to be the most exciting time to be in Link Studios.

Dave Blundin

Oh, my God. Yeah. I think that discussion we had of quantization on this podcast that Emad kicked off—I think that now vaulted to my new best piece of media ever recorded, passing Leopold Aschenbrenner. I have to go back and listen to that again in slow motion.

Also, we had Vlad Bulović from MIT.nano in this week. He's going to advise and help us on our new startup working on photonic computing, and he gave us a whole roadmap of people I need to meet next week. So we're looking to add 2 MIT people to our Princeton team to work on just the photonics, quantized photonics side of the equation.

So I'll be working on that next week. But I think I can take that video we shot earlier and use it as a recruiting tool. It was just so freaking brilliant. You guys are incredible.

Peter Diamandis

I love you guys so much. What a great week. We'll see what breaks tomorrow. Over the weekend—emergency pod. We need emergency pods every day by January.

Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272 | BidClub