[BidClub_]
Invest Like the Best · · 87 min

The Two Harvard Dropouts Who raised $800M to take on NVIDIA

Patrick O'ShaughnessyGavin UbertiRob Wachen

YouTube
TL;DR
  • The core thesis is stated flat-out in the opening seconds: "inference is going to be the biggest market in the world. Whoever produces the most tokens is going to be the most viable company in the world." Etched's founders argue the entire semiconductor stack is "built on buffer" — general-purpose assumptions (like EDA defaults that assume chips run at freezing temperatures) that a pure inference company can strip out, compounding "20% here, 50% there, 2x here" into a system they claim is not 10% but "10x better" than incumbent AI chips.
  • Two technical bets carry the whole design. For prefill: low-voltage inference — exploiting Dennard scaling (power is quadratic in voltage) to run at "under half the voltage of any other AI chip," solving thermal throttling before adding flops. For decode: cluster-scale memory — a fully custom interconnect above layer-2 ethernet that cuts chip-to-chip latency "by more than a factor of 5x" versus Blackwell's ~4,000 ns hops, letting the whole scale-up cluster's HBM and SRAM act as one pool.
  • Execution style is the moat as much as the architecture: extreme vertical integration ("the best vendor is no vendor"), a self-built rack, a Taiwan factory, and "pre-fetching" — 700 FPGAs running the full chip, racks pre-shipped to customer data centers without chips, thermal mock chips validating cold plates. Result: silicon back to inference-in-a-rack in 40 days, versus a "very famous AI chip company" that publicly took 10 months.
  • The near-death moment was capital: in early 2024, with $15M in the bank and a $100M twelve-month spending need, "every major investor in the valley passed immediately." They scraped to a $103M Series A of soft commits, with support from a debt provider, Synopsys emulators on multi-year terms, and TSMC — won over at a conference dinner by a then-22-year-old Gavin talking tensor math with a senior VP. Patrick, an early investor, breaks the fourth wall: "you kind of have to damn the base rate."
  • The structural edge over hyperscaler silicon is existential focus: "Google won't fail if TPUs fail... OpenAI won't fail if Jalapeño fails" — and Rob says lab/hyperscaler chips carry lower "FB8 × FB8" FLOP density than the "Blackwood B300" (as heard) because "they don't have to go take the risk." Supply-wise Etched claims it's additive, not zero-sum: 4nm and different HBM than Rubens' 3nm, so customer conversations are "2 gigawatts," not either/or.
  • The futurist calls are aggressive and dated: Rob predicts 2027 is when agents outnumber humans in knowledge work, that productivity will be measured in "agents per megawatt," and that a trillion-dollar single data center is "a matter of time." Gavin's model thesis: math gets cheaper faster than memory, so future models should burn compute lavishly — billion-token context, "every book ever written in its short-term memory," giant distributed MoE "brains."
Digest · the substance, structured for research

1. The consensus said 21-year-olds can't build chips — the answer was that the industry is "built on buffer"

  • Patrick opens with the diligence verdict: semis companies are built by 40-50-year-olds who've shipped multiple chips; "two 21-year-olds are not going to do this." Gavin's response is that the objection encoded stale constraints: "there's a certain level of naivety required to think that you could build a chip better than every other AI chip ever built... and we had the naivety."
  • The load-bearing insight: the entire semiconductor and data-center stack is "built on buffer" — every layer from EDA tools to power modules is general-purpose. His concrete example: timing sign-off "corners" default to chips running at freezing temperatures — "I've never seen an AI data center with ice in it" — when inference chips never run below 80°C. Killing dead constraints compounds: "20% here, 50% there, 2x here."
  • Rob's filter for early believers: heuristics people versus truth-seekers. The specimen is Mark Ross, ex-CTO of Cypress Semi (sold for $9B), who said "No, you can't. It will not work" — then demanded a white paper and a functional simulation, was surprised twice ("Huh, this works"), advised them to raise at least $3M (they raised five), and progressed from advisor to half-time advisor to full-time CTO.

2. Two bets, one per phase of inference: low-voltage prefill and thermal-first flops

  • The product framing: not a chip but a full rack-scale inference solution — chip, power delivery, boards, interconnect, and "really, the production is the product." Inference splits into prefill (reading text to set the model's KV cache) and decode (generating tokens), which they disaggregate across separate clusters — Patrick's gloss, "loading the gun and then firing it," which Gavin accepts.
  • On prefill, the metric that matters is real flops, not headline flops: MFU on GPUs runs 20-50%, and you can't reach 100% anyway because chips thermally self-throttle. So "if I just add more flops to a GPU today... it's just going to thermal throttle" — the thermal problem must be solved before adding flops.
  • The mechanism is Dennard scaling — power is quadratic in voltage ("if I 2x my voltage, my power goes up by 4x"). When they asked dozens of semiconductor veterans how to run below GPU voltages, the answer was "you can't" — dissatisfying, since "Bitcoin miners run at under a quarter of the voltage of GPUs." Their new power-delivery mechanism, low-voltage inference, runs "at under half the voltage of any other AI chip," and Gavin says: "we think all AI chips in the future are going to be low-voltage chips."

3. Decode is a memory game — and the right question is cluster bandwidth, not chip bandwidth

  • Gavin's reframe: people ask "how much memory bandwidth is on your chip?" when they should ask how much bandwidth is on your full scale-up cluster. On Blackwell, chip-to-chip hops run about 4,000 nanoseconds point-to-point, so an 8x tensor-parallel setup delivers "way, way less than an 8x improvement" in tokens per second per user.
  • Etched built a totally custom interconnect stack — "everything above the second-layer ethernet, built it full custom" — cutting latency "by more than a factor of 5x," so the SRAM and HBM of the whole cluster works as a single pool ("cluster-scale memory"). As world size scales, time per token drops proportionally.
  • The unsentimental diagnosis of incumbents: "all these architectures were built before ChatGPT" — flop organization, voltage domains, power planes, packaging, board design, interconnect all look different when designed for modern workloads. And the physics headroom is enormous: the mathematical latency limit is speed of light, "two, three nanoseconds" — against 4,000 today. "There's a lot of room at the bottom."

4. Why this is the decade's bottleneck: tokens haven't had their economies-of-scale moment

  • Gavin's macro case: real AI exists, and the constraint is now concurrency and speed — "it's just not possible for a billion people to use these models concurrently," and today paid AI plans reach only a few million users, "1/1000 of the global population." Faster decode also compresses wall-clock time: an agent task that takes a year of inference-time compute becomes a month.
  • Patrick's signature analogy: with iPhones, "more money does not really buy a better iPhone" — billionaire and average American buy the same one. Tokens aren't there yet: "a general-purpose system, relatively small one, is kind of handcrafting these tokens... like they made screws back in the Renaissance." He wants iPhone-grade economies of scale for token-making.
  • The practical stakes: certain products (coding models) are unusable below a tokens-per-second threshold, so the choice today is "shut off a bunch of the world from using this stuff, or everyone's going to get a worse experience" — hence the pressure for new hardware, "from the wafer to the watt... from the transistor to the token."

5. Origin stories: a tumor GPT-4V caught instantly, and a 17-year-old kernels engineer

  • Rob's motivation is visceral: stage-four bone cancer at the end of sophomore year of high school, under 30% survival odds, two years of chemo and learning to walk again. When GPT-4V shipped, he uploaded a pre-diagnosis photo of the bump on his back; the model immediately said "this could be a tumor, get an MRI" — "that took me six months" with real doctors. Then the kicker: a "you're all out of image credits today" notification. "Holy crap... we clearly don't have the infrastructure to serve it."
  • His second angle: running the Prod incubator (early companies included Cursor/AnySphere), he watched every startup spend its raise on compute and concluded software COGS "is not going to be zero anymore... it's going to be a function of inference" — a "decade march for inference to become the biggest market in the world."
  • Gavin's path: kernels development at Xnor at 17 (too young to sign a contract), then ExaNous (bought by Altera for $200M) and OctoML (bought by Nvidia for hundreds of millions — "Octo" as heard). The lesson from kernels work: "the math is relatively easy... what matters is data movement" — which is exactly what cluster-scale memory attacks.

6. The robotics template: two people, no outreach, just win

  • In FTC robotics, Gavin and partner Sanford abandoned the standard 20-person team model and its documentation/outreach culture: "we did nothing else besides build a robot that scored the most points," redesigning every 3 months. They held the world-record high score at one point and ranked third in the world by OPR.
  • The translation to Etched, teased out by Patrick ("you chose to have no communications" — "exactly like the robotics team," Gavin admits): velocity, velocity, velocity — "you win by shipping" — and the conviction that "you can do the best product in the world with far fewer people" than incumbents' 20,000.

7. "The best vendor is no vendor": building the rack, the factory, and the chip simultaneously

  • Rob's extension of "the best part is no part": vertical integration from chips to boards, cold plates, interconnects, and production itself — "we're the only startup right now that's building its own rack as well as its own chips... at the same time." They hired the rack lead who "built all of Nvidia's HGX and DGX systems, which is like 80% of the revenue" (Bryan Loyler, as heard).
  • Before silicon returned they built thermal mock chips with the exact expected hot spots, over-pressurized cold plates until they blew up — "we haven't had a single leak since" — and ran a Taiwan factory with cloned test stations in the office, a 2 MW data center on the floor, and 24/7 day/night development shifts.
  • On where integration stops: economies of scale draw the boundaries — "natural boundaries are on the chip side, on the bottom, and the model layer at the top. And we'll fill the whole gap between." They're deliberately not building data centers today: "that doesn't actually help us get more capacity online."

8. Talent: "legends" plus "chips on shoulders" — and you have to be "sick in the head" to join

  • The bimodal philosophy: for unsolved problems, find the literal best person in the world via "project-based recruiting" — map every hardest technical problem ever solved, find who did the zero-to-one, and persist: "the amount of people who say yes after the first conversation is pretty low, but after the 20th conversation is surprisingly high." Brian's value once landed: "that's a billion-dollar lesson I learned" — pointing to lessons learned before they're repeated.
  • The other mode — "chips on shoulders put chips in data centers" — is exemplified by Sanford, asked to build a cold plate in a week ("any thermal engineer would think you're totally naive... these things take months") and de-risking a key power question anyway. The magic is pairing: naive first-principles risk-takers who "don't know where the bodies are buried" plus scale veterans, working together.
  • Rob's self-aware recruiting pitch, verbatim: "You kind of have to be sick in the head to join our company... move to San Jose for the semiconductor company run by two, what, 24-year-olds now... against the biggest companies in the world... with a design that they're saying is not going to be like 10% better, but 10x better. Something must be wrong with you to do that." His worry: as specs go public and the company gets consensus, the contrarian self-selection filter may weaken.

9. Spending money to go faster: Bangalore, pre-fetching, and 40 days versus 10 months

  • When a physical-design vendor fell a year behind near tape-out, both obvious options cost a year — so option three: ship dozens of top engineers to Bangalore for 6 months (Gavin lived there 4.5 months), first in at morning, out at 1 a.m., with 12-hour US handoffs at 8 a.m. and 8 p.m. daily. Other chips at the same stage with the same vendor "still haven't taped out today."
  • The economic logic for burning cash: "the biggest risk is not taking risk" — with over a billion dollars of daily category revenue, "every day we don't ship, we're just leaving tons of opportunity on the table." Hence "pre-fetching": 700 FPGAs running the full chip and a dozen models on the real inference stack, racks shipped to customer data centers without chips to bring up software, the production line ready before silicon landed.
  • The payoff, as told: a "very famous AI chip company" took 10 months from silicon back to inference in a rack — publicly disclosed to investors. Etched did it in 40 days, "because by the time the chip came back, everything was boring." Over half the company lives next to the office; on the shift pay question, Rob answers: "the invisible hand does wonders."

10. The two darkest weeks: 50 picoseconds, quitters, and "the puzzle begins"

  • FPGAs verify digital but not analog logic — so when silicon returned, a back-pressure failure across a clock-domain crossing produced wrong attention results. The only fix: align two clock signals on-chip to within 50 picoseconds — "50 trillionths of a second" — on every chip, 2 billion times a second. "We had people quit... people literally were like, this problem is unsolvable, and best of luck, guys."
  • The method: "step one is, okay, let's assume the problem is solvable" — a drifting mechanism to phase-shift clocks by picoseconds, then lock them. It worked in about 2 weeks — "a very scary 2 weeks" — and Gavin's generalization: the moment things feel hopeless "is the most important time to be investing effort."
  • The companion war story: first wafer sort at 2-3 a.m. with TSMC on the phone, every die on the screen turning red instead of green. The silicon-validation legend leans back: "The puzzle begins." And the experimental philosophy behind it all — 30 board experiments, three worked, all "worth their weight in gold": "people come to me and say, Gavin, almost none of your experiments works. And I'll say, I only got to get lucky once."

11. The near-death fundraise: "every major investor in the valley passed immediately"

  • Early 2024, pre-Series A: architecture proven, but the physical-design stage alone costs at least $40-50M, giant MoE models meant building the whole cluster, and the bank held $15M against a $100M need over the next 12 months. Gavin's honest memory: "you're sitting in that moment and you think, holy crap, we can't afford this... how hard is it to go back to Harvard?" A 30-page technical memo took 100 hours; every major investor passed — "two kids that just finished Harvard, haven't taped out a chip... everything's going to be training... this could all be a bubble." The biggest semiconductor Series A fundraises were around ~$40-50M.
  • Survival mode: a $30M "ramen to tape-out" floor plan backed by a debt provider, then calling everyone — "we need a hundred million dollars... do you know somebody that wants to take an aggressive bet?" — until a board meeting where a spreadsheet read $103 million in soft commits: "we all look at each other and we say, we're going to take it." Nearly half a dozen rounds since, "many of them from those investors just doubling and tripling down."
  • Suppliers believed before the money did: Synopsys extended emulators on multi-year terms ("basically a big loan"), and TSMC signed on after a SEMI conference dinner where 22-year-old Gavin — the only speaker under 30 — sat by chance next to a senior TSMC VP, both math majors, talking per-tensor model mechanics on paper. Next day's email: "Gavin, want to work with Etched? Find a way to make it happen." Rob's broader TSMC verdict: the tech is the best, but "the real value is all in the service" — they ran a yield experiment on their own dime at Etched's suggestion, then rolled it across the line. "If I go to the steel works plant and say change the composition of the steel, they'll say screw you. Not TSMC."
  • Patrick breaks the fourth wall — "I'm a big Etched investor... I'm incredibly biased" — and reflects on contrarian investing: experts "laid out in very logical terms why this wasn't going to work," and the lesson is "you kind of have to damn the base rate... there's always the index fund." Gavin's observation on who did believe: pure market/team believers, or the ultra-technical — high-frequency trading firms who audited everything down to the RTL; "if you were anywhere in the middle, you just wouldn't understand it."

12. The software bet: kernels-first, under 100 models, and an existential edge over hyperscaler silicon

  • Three years ago the choice was graph compilers (work out of the box, poor performance) or kernels-first programming; Etched went kernels-first, betting "there was going to be under 100 models that actually mattered" — no arbitrary PyTorch, CUDA, or ONNX support. Skipping the compiler "saved us a tremendous amount of time"; the only early believers were HFT firms — "they all hate compilers too" — dozens of whose engineers have since joined.
  • The bet is aging well as AI eats kernel-writing: tooling is designed for how models will use it, and in one internal experiment "Codex actually got GPT-OSS running from scratch just based off of our docs, completely by itself... overnight."
  • The market-structure argument comes from a hire poached mid-recruitment (an architect at a frontier lab's chip project "Uno reverse card"-recruited into Etched within a week): "It fundamentally is not existential for my company for this product to win. Google won't fail if TPUs fail. Meta won't fail if MTIA fails. Microsoft won't fail if Maya fails. And OpenAI won't fail if Jalapeño fails." Rob's corroborating data point: lab and hyperscaler chips show lower "FB8 × FB8" FLOP density than the "Blackwood B300" (as heard) — "they don't have to take the risk. They just have to build a similar enough product and not pay the Nvidia tax."
  • Supply is framed as positive-sum by design: first-gen sits on 4-nanometer with different HBM, versus Rubens on 3-nanometer, so scale deployers see "2 gigawatts," not a swap. The closing discipline: think supply chain at design time, because "if you have the most performant product and you can't produce it, you're just a podcast."

13. The future: giant distributed brains, agents per megawatt, and trillion-dollar token factories

  • Gavin's model thesis starts from "machines don't think like people think" — airplanes don't fly like birds. For chips, loading data is expensive and math is cheap, and "math gets cheaper at a rate that is faster than memory gets cheaper" — so future models should burn compute: many parallel copies, gigantic experts spanning racks, billion-token context. "I would love to talk to a machine that was able to attend to every book ever written in its short-term memory." Rob adds the near-term architectural theme of dynamism — per-token, per-user control of compute and memory — which pre-ChatGPT hardware handles with "blunt force."
  • The wall-clock argument for speed, via Noam Brown (now an angel): 6-month agent tasks can't even be evaluated before the next model ships; and just as no single human can build a rocket, agent work will need teams — "maybe that's 10, maybe that's a million" — demanding colossal shared memory and flops. Cursor's agents built a browser from scratch in a week; "that will soon happen in under an hour."
  • Rob's boldest dated calls: inference on a "global march" to a majority of global GDP (may take more than 10 years), productivity re-denominated as "agents per megawatt," and "this is the second-to-last year a majority of the workforce is going to be human — in 2027 there's going to be more agents doing knowledge work than humans." On a trillion-dollar single data center: "Absolutely. It is a matter of time" — economies of scale don't stop at $40B fabs, and the same holds for "plants that make tokens."
  • The closing framings for Patrick's "smart alien": Gavin — thinking is valuable, every company runs on it, and someone must build the roadmap for "the future quadrillion-parameter models for a billion people all at the same time." Rob — the cost of producing intelligence is so far below its value that "we are in a many-year, probably many-decade supply shortage of these tokens." The episode ends with Rob's answer to the kindest-thing question: at 16, asked to choose surgery (live, maybe never walk) over radiation, then needing radiation anyway from one of the few machines in the world, in Boston — and both parents dropping everything to move there with him.
Gavin Uberti

We know inference is going to be the biggest market in the world. Whoever produces the most tokens is going to be the most valuable company in the world. We had people quit.

Rob Wachen

Yeah.

Gavin Uberti

People literally were like, “This problem is unsolvable. Best of luck, guys.”

Rob Wachen

You kind of have to be sick in the head to join our company. You’re going to convince your family to move to San Jose for the semiconductor company run by two, what, 24-year-olds now, going against the biggest companies in the world with a design that they’re saying isn’t going to be 10% better, but 10x better.

We need $100 million to do this. If we do this, we think this could be one of the most important companies of all time. You’re sitting in that moment thinking, “Holy crap. We can’t afford this.” And I’m looking at it like, “Huh. How hard is it to go back to Harvard?”

1. Why Nobody Believed Etched Would Work

Patrick O'Shaughnessy

All right, gentlemen. It’s been 3 years or so, Gavin, since you and I last did this, which is nuts. At the time, I was just wildly intrigued by your story and what you were going to build. I didn’t know a lot about chips, and I was considering investing in the company, so I was calling everyone I could conceive of who could give me an opinion or something.

At the time, basically, the consensus was that these kinds of companies are not built by young people. In the semiconductor world, the best companies are founded by 40- or 50-year-old people who have had a whole career’s worth of experience, have learned all the problems, and have shipped multiple chips. Two 21-year-olds are not going to do this. It’s just not going to work.

It was indicative of a theme, which was that nobody believes in us. That’s obviously changed a lot now. You just walk the halls and talk to the people who have chosen to come work here, but in the early days, it felt like this was something that you had to face down. What was that like, facing that down, where a set of incumbents and an industry’s worth of people and investors and everyone else sort of didn’t believe in you? What did that anneal in you to build the company the way that you are? What was the impact of that?

Gavin Uberti

I think there’s a certain level of naivety required to think that you could build a chip better than every other AI chip ever built and build a company to do it way faster than has ever been done. We had the naivety. There were many times when we would say, “Why isn’t this possible?” and really push on it.

It turns out that everybody’s answers are extremely siloed to a set of constraints that aren’t true anymore. The reality is that the entire semiconductor and data center industry is built on buffer. What I mean by that is that every part of the stack—from the EDA tools to the power modules to the circuit boards to the chip design and standard cells—is built to be general-purpose for everything, not just in the data center, but for IoT on the edge and so forth.

When you have a specific use case that you’re really trying to design for, you can change the constraints a lot. I’ll give you a very simple example, which is one of the things you care a lot about: the clock speed of your chip. It’s proportional to the throughput of your system.

When you’re doing sign-off for different timing and figuring out what clock speed you’re actually going to be able to run at when you tape out your chip, there’s this concept called corners, which is what temperatures you’re going to be able to run at that clock speed. The default configurations for a lot of these EDA tools assume that you’re going to be running your chips in freezing temperatures. I don’t know about you, but I’ve never seen an AI data center with ice in it.

We can feel pretty confident that our chips don’t need to run at full speed at 0°C. In fact, they’re never really going to be running below 80°C anyway. Just by knowing that’s a constraint that doesn’t matter, we can make a ton of changes throughout the entire system. That’s a very simple one, but there are many more that get you 20% here, 50% there, 2x here, and these compound to a system that can be radically better for inference.

Rob Wachen

Well, I think you found two kinds of people. There are some folks who went purely on heuristics: “Okay, young founders, they claim they can go beat the biggest company in the world on performance. It cannot happen. There is nothing you could say to me that would make me change my mind.”

But there are also people out there who were, of course, skeptical but willing to say, “I’ll spend the time, I’ll do the work, and figure out whether this is actually possible.” For example, one of our earliest supporters was this guy, Mark Ross.

Mark was a very prestigious semiconductor expert. He used to be CTO at Cypress Semiconductor, which sold for $9 billion. When we met him, we were just a couple of guys in a dorm room. We came to him and said, “Hey, you want to go build hardware for inference? We think we can be much faster than NVIDIA.” Mark was like, “No, you can’t. It will not work. But if you want to convince me, you should write a white paper, build a functional simulation, and show me.”

After a lot of very long nights, we went back to Mark and said, “Here’s a simulation. What do you think?” He was like, “Huh, this works. But to go do a company like this, you’ll need a large amount of capital. You’ll need at least $3 million even to get you started.”

We went ahead and raised $5 million and raised a lot more after that. He was again surprised but got more involved. Then he became an advisor, then a half-time advisor, and eventually our full-time CTO, as he saw more and more of the development progress.

In general, this has been a filter that really heavily filters out folks who want to be right regardless. It attracts folks who want to be very truth-seeking and say, “Sure, I’m skeptical, but I will go ahead and work through the numbers myself. And if I can figure out why this is possible, well, let’s go build it.”

Patrick O'Shaughnessy

The specifics that you’ve made bets on and the way that you’ve built the system are immensely interesting to me. So many people are trying to do this now—build new chips that will do a better job of serving inference at massive scale—and the world is interested in the research approaches and the different architecture approaches that people are taking to building a new AI chip.

I’d love you to start by describing what this thing is and what it does, but maybe more interestingly and more importantly, the process that you went through to decide what bets to take and what technologies to invent. Compare and contrast those with what you’ve seen the rest of the marketplace try to do.

Gavin Uberti

Yeah, I think to start with the product, we’re not just building a chip. We’re building a full inference solution, and that means a rack. That means the chip, the power delivery into the chip, the board in which it sits, and the interconnect that allows the chips to talk to each other. That means the production for this massive volume of racks. Really, the production is the product.

When we think about how we get our advantage, there are two key parts of running inference: prefill and decode. We have two key techniques that match both of these things. Prefill is reading in a huge volume of text, and decode is then using that data to generate output tokens.

When you run prefill, your key job is not to predict tokens. You already know the text. Your job is to get the model’s memory, what we call the KV cache, into the right state. Then you can run decode with that same KV cache.

What we will often do is what we call PD disaggregation, or prefill-decode disaggregation. You will have one cluster of servers running these prefills. You’ll then transfer those model memories, those KV caches, over to the decode cluster and use that cluster to generate the next tokens.

Patrick O'Shaughnessy

So it’s sort of like loading the gun and then firing it, if I think about it in super-simple terms.

Gavin Uberti

Yeah, you got it. It’s getting the models to remember the right things and then using those things to do tasks.

Rob Wachen

Generally, people think about this market a bit lazily, or they say, “Are you a prefill chip or are you a decode chip? If you’re a decode chip, are you an HBM chip, an SRAM chip, or a 3D DRAM chip? Are you using optics or using copper?”

When we started this, we just wanted to understand why extremely smart people were working on these different directions. We seriously looked at architectures like having a bunch of DDR memory in a shared memory pool and looking at advanced packaging to basically break out of the shoreline. We looked at things like, “Are there ways to put memory dies on top of compute dies?”

In doing so, we realized that there’s no free lunch. Everything has a trade-off. With 3D DRAM, you have a thermal issue, you have a supply-chain issue, you have to figure out hybrid bonding, you have to figure out the FLOPs, and now you’re a decode chip.

So we went through everything, both on the prefill and the decode side. In doing so, we realized there are a few design spaces that nobody had seriously tried to explore because they had never been done in AI chips. We asked ourselves, what are the actual metrics that are going to matter the most?

On the prefill side, the thing that matters is FLOPs and FLOPs density. People talk about FLOPs often as a headline number, but in reality, you should care about the FLOPs you’re getting when you’re running real workloads. There’s this concept called MFU, or model FLOPs utilization, which is, for every peak FLOP advertised, how many cents on the dollar are you actually getting? On GPUs, you often get somewhere between 20% and 50%, depending on the workload.

Gavin Uberti

Actually, you can probably not run at 100% because you have a thermal issue. Whereas, as you increase the FLOP utilization, you have more transistors going on and off, you draw more power, and the chip will self-regulate and actually lower its clock speed to make sure it doesn't overheat.

As we looked at inference, we said, if we want way more FLOPs because we want to run at way higher throughputs, we fundamentally need to solve the thermal problem before we even think about adding FLOPs to the chip. If I just add more FLOPs to a GPU today or another AI chip, I'm not actually going to get more performance because it's just going to thermal throttle.

Fundamentally, the essence of that is this concept of Dennard scaling, which is that power is quadratically proportional to voltage. So, if I 2x my voltage, my power goes up by 4x. If I cut my voltage in half, I cut my power down by a quarter.

So, we asked ourselves, how could we run voltages lower than GPUs? We talked to a lot of people about this. We flew out to Silicon Valley after dropping out and basically asked the dozens of people in semiconductors at all these different chip companies how they did it. The answer we got was, “You can't. You can't run at voltages lower than GPUs.”

This was very dissatisfying because there were many different industries of chips that run at voltages lower than GPUs. Bitcoin miners run at under a quarter of the voltage of GPUs, so this is obviously physically possible. The question is, are there issues with GPU architectures that make them unable to run at these voltages?

When we looked at the problem for a long time, we were able to create a new mechanism for running at much lower voltages, a new type of power delivery that we call low-voltage inference. We think all AI chips in the future are going to be low-voltage chips. They're going to have to cram way more FLOPs into the same silicon area and, without thermal throttling, run at way lower voltages. So, that's prefill.

For decode, it is all a memory game. More memory bandwidth means you can load the model faster, load the KV cache faster, and serve more tokens per second per user. We think people ask the wrong question here. People often ask, “How much memory bandwidth is on your chip?” You should be asking how much memory bandwidth is on your full-scale-up cluster.

What we're able to do is add way, way more bandwidth and a much lower latency from chip to chip to our interconnects. It allows us to serve models at this much higher speed because you can use the SRAM and the HBM from the full-scale-up cluster as a single pool. That's our second key technical bit, what we call cluster-scale memory.

On GPUs today, the cluster memory bandwidth is often very badly utilized because the time to hop from one GPU to another is extremely long. For example, on Blackwell chips, it can be about 4,000 nanoseconds to go point to point. That means that if you go ahead and go to an 8x tensor-parallel setup, you will get way, way less than an 8x improvement in your tokens per second per user.

What we did was build our own totally custom interconnect stack. We took everything above Layer 2 Ethernet and built it fully custom. We can go far better in terms of latencies and bandwidths this way, too. We can go ahead and cut this by more than a factor of 5x, and that allows us to use the memory of other chips much more effectively. As you scale the world size, your time per token goes down proportionally.

Patrick O'Shaughnessy

That's not that surprising, given all these architectures were built before ChatGPT. If we're trying to build a chip for modern workloads, it's going to look very different. The way we organize our FLOPs, the way we do our voltage domains, and the way we do our power planes are going to look super different. The way we do the packaging is going to look super different. The way we do the board design is going to look different. On the decode side, the way we connect everything is going to look very different.

Gavin Uberti

We're now bringing forward our first generation of this low-voltage inference technology, which is running at under half the voltage of any other AI chip.

2. Why Inference Is the Bottleneck

Patrick O'Shaughnessy

If you zoom all the way out, why is this so important? Why is the delivery of much higher throughput, much lower cost per token, better tokens per watt—all of these metrics—the universe is going to start talking about more and more? Everyone knows the supply side of the equation is a big problem right now. Why is this, in the bigger picture, looking at a decade, the bottleneck in the technology world?

Gavin Uberti

Well, I think it comes down to productivity. We are at this extremely interesting moment in the history of civilization where there is real artificial intelligence—not sci-fi stuff, but models that can solve problems that most humans can't. It's going to create new scientific discoveries, instant access to medical care, and instant access to education.

Now it's just about how many people can use this at the same time, how many products can service this at the same time, and also the speed of doing different tasks. When you think about wall-clock time, if we can take an agent that can run at a certain model quality and could take a year to solve a certain task using inference-time compute, if you have way faster decode speed, you can compress that into a month. The amount of scientific innovation and the amount of actual proliferation of technology will happen much faster.

The second part is concurrency. Today, it's just not possible for a billion people to use these models concurrently. Ultimately, some people are going to get downgraded, some people's models are going to be slower, and some people just won't be able to access the hardware.

A few years from now, there are going to be giant models serving billions of users. We're very much in the early innings of AI today. With paid plans, there are only a few million users in the world using paid plans for AI models, so we're at 1/1,000 of the global population actually using this stuff.

3. The Future of Models, Agents, and Intelligence

If you want to serve at giant scale, a lot of things change. One of them is the number of chips that communicate together. People usually think about this in the context of training. You have these giant training clusters, and you have Colossus with over 100,000 GPUs that are all networked together.

On the inference side, today people usually think about it as an 8-chip cluster, or maybe just NVL72 as the scale-up domain. But very quickly, this is going to become thousands of chips and tens of thousands of chips. The way to get the most performance there—the time between sending data from one chip to another—that primitive matters way more than is getting credit for right now.

When we think about optimizing memory bandwidth for the system, you have to think about how fast these chips can communicate together. If they can only communicate really quickly with themselves and very slowly with other chips, you're not going to actually be able to serve giant models at 10,000 or 20,000 tokens per second.

We need multiple orders of magnitude of infrastructure built out throughout the entire stack, from the power—from the wafer to the watt, from the transistor to the token—to actually bring this stuff to the world.

Patrick O'Shaughnessy

I think that when you look at most other goods, like the iPhone, for example, you've gotten to these economies of scale. As a result, more money does not really buy you a better iPhone. If you're a billionaire or you're just the average American, you buy the same phone.

Tokens aren't like that yet. We're still in the very early days, where a general-purpose system, a relatively small one, is kind of handcrafting these tokens, like they made screws back in the Renaissance. I want to live in the world where you have the same economies of scale for token-making that you do for making, say, iPhones or cars or anything else.

I think that is one of the huge unlocks that allows a huge group of people to use the best-quality models. I think that economies of scale have made capitalism very—I don't know—fair.

Rob Wachen

I think that allows you to go ahead and have the same product in many, many different hands, and you're able to serve way more users on a single scale-up cluster. It allows you to get closer to that point for token serving, too.

Patrick O'Shaughnessy

Yeah, and also, certain products aren't usable if they're slow.

Rob Wachen

Yeah.

Patrick O'Shaughnessy

So, if you want to serve coding models and you want people to actually use them, there's a certain number of tokens per second you need to hit. The question is: while maintaining that per-token speed, how many users can I serve at the same time?

You can basically decide, "I'm going to shut off a bunch of the world from using this stuff," or everyone's going to get a worse experience. Fundamentally, you need to find ways to push out the curve, and that's why there's such pressure for new hardware.

I'd like to take some time to step back and hear both of your stories for how you came to this idea and this company, and then walk through what it's been like to build it. I think in so doing, we'll understand the system that you've built for the company itself, which will then be able to power subsequent generations of products like this one for this crazy inference future that we're staring down.

Rob, maybe starting with you, just take it however far back you want. What I'm curious about in your personal story is that the very first thing I ever heard from either one of you was your personal story many years ago, which really blew me away. I'm most interested in your motivation, ultimately, for being here doing this thing.

Rob Wachen

It starts back in high school for me. I've been very unlucky and lucky at different points in life. This was one of the tougher times.

At the end of my sophomore year of high school, I got injured at a martial arts tournament. The next day, I couldn't walk for some reason. They thought there was something wrong with my SI joint or something. I went through physical therapy and did different types of scans, but they couldn't figure it out. Eventually, they found this big bump on my back in an MRI and told me it was a tumor.

It was stage 4 bone cancer, and I was told I had under a 30% chance of survival. It was a 2-year, crazy chemotherapy, surgery, and learning-to-walk-again experience. When you go through something like that, it really changes the Overton window of human experience and makes you appreciate what actually matters. You also ask yourself, "What are you going to do if you have the chance to live?"

If you actually want to get through something like that, you need to be hoping for something. I always knew I wanted to do something very impactful if I had the chance to get through it. It took me a couple of years to figure out what that was going to be.

At the same time, as I got to college and met a bunch of other people building cool tech, I got extremely excited by AI models, especially once GPT-3 came out. I was like, "Wow, this is the first model that can kind of speak English." These things are going to get really smart.

What happened was, when GPT-4 came out, there was GPT-4V, which was the first model with image uploading. I went through my camera roll and found a picture of my back with this bump on it before I was diagnosed. I said, "Hey, ChatGPT, pretend you're an expert doctor. A patient comes in and says they have this bump on their back. What could it be?"

It immediately said, "This could be a tumor. You should get an MRI immediately. Go to the doctor." I just kind of sat there still. It was—

Patrick O'Shaughnessy

This took 6 months.

Rob Wachen

Yeah, that took me 6 months. Yesterday, this feature wasn't there; today, it's here. I went to show my parents, and I got this notification saying, "You're all out of image credits today. You need to get a Pro plan."

I was like, "Holy crap. This is going to change everything." We clearly don't have the infrastructure to serve it, and there are very few things you can work on that can actually bring this technology to the world at scale faster.

There are plenty of people who are super smart working on models. The fabs seem maybe unreachable to work on, but it seemed like the hardware was all designed before ChatGPT. Every GPU, every TPU, and every AI chip that was serving these models was fundamentally built before this and retrofitted to serve these modern models.

There's going to be an entire new wave of architectures that come out. What are more exciting things to work on than bringing this to everybody?

A very different angle at the same time: I was running a startup incubator called Prod, which has incubated a bunch of different companies. Some of the earliest ones were Cursor and AnySphere, which merged, and Recur and Hatchpoint through it, along with a handful of others.

At the time, as these models were getting smarter—it was 2022—I was realizing all of these companies were spending all the money they raised on compute. I had this realization as I was working on some of my own stuff: "Oh my God, all the products I want to build are going to cost tens of millions of dollars a year in inference."

This is not going to be tenable. The cost structure of every software company—the COGS—is not going to be zero anymore for an incremental user. It's going to be quite high, and it's going to be a function of inference. The OPEX of an already-built business is also going to be inference as people use more and more coding agents.

4. Gavin and Rob’s Origin Stories

Fundamentally, it seems like inference is going to be really important, and it feels like we're on a decade-long march for inference to become the biggest market in the world. When you think about that, 10 years from now, there are going to be these giant projects where fundamentally nothing in that data center has been designed today. We should go pick something to work on, and that's kind of how it got started.

Patrick O'Shaughnessy

Yeah, but I'm really excited for you to go back about as far—probably early in high school, maybe even earlier—and tell me your favorite milestones on the timeline that ultimately led to your ambition to drop out of Harvard and start this company.

Gavin Uberti

My first job ever was at a company called Xnor, where I did kernel development. I was 17.

Patrick O'Shaughnessy

Yep.

Gavin Uberti

A 17-year-old can't sign legally binding contracts, so I had to do a traditional NDA. They sat me down and said, "Gavin, don't share this information."

Xnor was one of the only companies that saw, “Hey, maybe this is a good trade.” ExaNous got bought by Altera for $200 million. I did the same thing at OctoML, where they got bought by Nvidia for hundreds of millions of dollars.

When you do this sort of kernel work, what you realize is that the math is relatively easy.

Patrick O'Shaughnessy

Okay.

Gavin Uberti

But to get high-speed decode, I think what matters is data movement. Almost all the work that you do is optimizing how you move data around a single chip or across multiple chips.

That's why we went ahead and built this cluster-scale memory technology. We bring that interconnect time way, way lower. You can do way more movement and, as a result, get a much faster time to generate each subsequent token.

You can build these crazy things Rob was talking about, doing a year's worth of work in a month, or more than that in the future.

Patrick O'Shaughnessy

Can you talk about the competitive drive that's evident in some of the high school competitions that you participated in and won?

Gavin Uberti

We did a couple. For example, I was very active in FIRST Tech Challenge robotics, and I was lucky to have a very talented partner, Sanford.

For a long time, we were part of a traditional school team, where it was about 20 guys all working together, as is often typical of FIRST Tech Challenge. The goal is to build a robot that scores the most points, along with a bunch of other things.

In FIRST, they put a lot of emphasis on collaborating with other teams, trying to do really good documentation, and trying to inspire others to do the same thing. Sanford and I decided that, rather than do it this way, we were going to win. We did nothing else besides build a robot that scored the most points.

Patrick O'Shaughnessy

As a 2-person team.

Gavin Uberti

Rather than a 20-person team.

Patrick O'Shaughnessy

That you were much, much smaller than almost every other team in the competition.

Gavin Uberti

We figured that if we were going to specialize, if we were going to go out and win the damn games really well, we wouldn't need to advance based on the quality of our documentation or our outreach. We were just going to go win.

And so we did. We brainstormed as a 2-person team, built a robot, and decided we were going to redesign it every 3 months. And we did. We actually had the world record for the highest score during this competition at one point.

We were rated by OPR third in the world for software development, and it was a damn good machine.

Patrick O'Shaughnessy

What from that episode can I translate as an analogy into how you built Etched the company?

Gavin Uberti

When you're thinking about how you want to do a full rack-scale product like this, there are a couple of key ideas. One of them is velocity, velocity, velocity: you win by shipping. You're not going to go out and win by having the best outreach or the best communications.

Patrick O'Shaughnessy

You chose to have no communications.

Rob Wachen

Exactly. I’m just realizing how exactly like the robotics team this is. There are many ways you can win in business, but we’d rather focus on just building the best product.

Patrick O'Shaughnessy

Mhm.

Rob Wachen

Similarly, we think we can do it with a lot fewer folks. If you’re willing to just focus on product, product, product, and parallelize relentlessly, you don’t need 20,000 people like the big companies have. You can build the best product in the world with far fewer people.

Rob Wachen

Yeah, there's a saying that the best part is no part. I think for us it's also the best vendor is no vendor. As much as possible, we want to vertically integrate the entire product, both because we get more performance, but we can move way faster. So, everything from the chips to the boards, to the cold plates, to the interconnects, to even the production, we want to do all of it as in-house as possible. We're actually, I think, the only startup right now that's building its own rack as well as its own chips. And we did it all at the same time. A couple years ago was the last time we were public. At that point we just started building our rack team, and we brought over Bryan Loyler, who built all of Nvidia's HGX and DGX systems, which is like 80% of the revenue. And we said, "We're going to build the rack at the same time." We actually went through multiple iterations of the rack before the chips even came back. Before the chips came back, we made thermal chips that had the exact same hot spots as we expected our chips to have so we could build the cold plates, we could over-pressurize them and blow them up. We haven't had a single leak since our chips came back with the cold plates because we already validated them. Yeah, we have a factory in Taiwan. We have a few dozen people out there. We built a clone of a bunch of the test stations in our office. We have a 2 MW data center on this floor and we did 24/7 development cycles. People are doing day shifts and night shifts to actually get the hardware up and running as quickly as possible. So, it's that extreme vertical integration and extreme parallelization of the schedule that lets you get products to market way faster.

Patrick O'Shaughnessy

Hmm. If you think about building the early team and what it required as 2 young guys building this company, there are lots of very talented young entrepreneurs out there, perhaps building something of this scope or magnitude for the first time in a long time. All of whom could probably benefit from the lessons that you’ve learned getting very sophisticated, talented people to come join you, even after careers at the other great companies. If you were teaching this as a class—here’s how to get elite talent when you’re young and inexperienced and naïve—what would the syllabus be?

Rob Wachen

We have a pretty bimodal talent philosophy. It starts with what we call the legends. When we’re trying to solve an incredibly hard technical problem and generally do something that hasn’t been done before, we need to find the very best person in the world. Often, the number 1 guy in the world versus the number 10 guy versus the number 100 guy is a huge difference in whether it’s actually possible to solve the problem.

We created a system we call project-based recruiting, where we map out all of the hardest technical problems across all industries that anyone has ever had to solve. We look at temporality: Who are the people who did the 0 to 1? Who was in charge, quote unquote? Who actually did the work? We talk to as many people as possible, and then we just track it. You’d be surprised by the number of people who say yes after the first conversation being pretty low, but the number of people who say yes after the 20th conversation being surprisingly high.

Patrick O'Shaughnessy

You really have to keep at them. When you hear no from somebody who really is the best in the world, that really means, “Hey, should you come back when you have a few more milestones proven out?”

Rob Wachen

Yeah. I think one of the most convincing things to see is, “Hey, we make bold claims.” When you hit those again and again and again, that is really belief-inspiring.

When we decided we wanted to build a rack and not just a chip, we were looking at this and saying, “How many products have actually shipped at scale for a rack-scale system that actually has the power density we’re trying to solve?” We said, “If we were going to wave a magic wand, what would the best possible person in the world look like?”

We’d say, “If we could find somebody who started at NVIDIA and built the entire rack team through all their different generations, learned all this different stuff, but is still scrappy, still understands the startup culture, but has seen scale, that would be the best possible person.” So, we mapped all of the different teams that related to all of the different rack-scale products at NVIDIA, and we found 3 people that we thought could fit the bill.

We talked to all of them. Two of them had just retired, and one of them was planning to do one more generation for NVIDIA and then retire. His name is Brian. Over time, we convinced him to join.

Brian started the HGX and DGX team at NVIDIA, which was a majority of NVIDIA’s revenue—tens of billions of dollars a quarter. The other 2 guys ended up investing, by the way. When you have somebody like that, they just know what good looks like because they’ve seen it.

There were so many times where we’d talk to Brian and he’d just point to us and be like, “That’s a billion-dollar lesson I learned. Billion-dollar lesson I learned.” That just saves us cycles. You pair someone like Brian with somebody like Sanford.

Patrick O'Shaughnessy

Do you have a name for them? Brian’s a legend. What’s Sanford?

Rob Wachen

You know, we say, “Chips on shoulders put chips in data centers.” Sanford and Gavin, in high school, were world robotics champions. Sanford was finishing his senior year of college, and we called him up a couple of years ago and said, “Hey, can you come check out what we’re doing? We need some help on the platform side.”

He came for a week, and we said, “Can you build the cold plate this week?” If you asked any thermal engineer, they would think you were totally naïve, right? These things take months to do. To be clear, they do. But you can make real progress in a week if you put your mind to it and think it’s possible.

He built a contraption in a week that actually de-risked a pretty key power question we had. You put those 2 together, and they’ve done incredible things. One is not possible without the other because you need the extremely driven people who just keep asking why and don’t know where the bodies are buried to take tons of aggressive risks. Then you need the people who’ve seen scale and still have the startup scrappy mentality to help them along the way.

Patrick O'Shaughnessy

It’s really the legends plus the raw, somewhat naïve, first-principles-type talent. But it’s the combination. It’s not just that you have both in the company; it’s that they’re working together.

Rob Wachen

That’s right.

Patrick O'Shaughnessy

If I think about that funnel, is there anything else more interesting to say about how much better you’ve gotten at recruiting and why those metrics keep getting better?

Rob Wachen

One of the shocking things is, I wish we weren’t being so contrarian, but it kind of self-selects.

Patrick O'Shaughnessy

Right.

Rob Wachen

If you’re the kind of person who is somewhat opportunistic, you’re going to go join whatever the hot company is, or go ahead and do due diligence. You will not come work here.

Patrick O'Shaughnessy

Right.

Rob Wachen

It’s one of the things I worry about as we announce more and more of the product and its specs: We may lose some of this if we’re not very careful. You kind of have to be sick in the head to join our company.

If you think about it on paper, it’s like you—a person who is probably a very accomplished engineer, making a good amount of money; it’s liquid, it’s predictable somewhere else—you’re going to convince your family to move to San Jose and live in this apartment on this housing program for the semiconductor company run by 2 24-year-olds, pre-product, going against the biggest companies in the world in the most supply-constrained environment ever created, with a design that they’re saying is not going to be 10% better, but 10x better. Something must be wrong with you to do that.

People are just wired differently here. They really want not to prove people wrong who don’t believe, but to prove people right who do believe. They just take it personally. It’s really fun to find those people, and frankly, the nature of the company makes it very easy to weed out the people who aren’t like that.

5. Taking Huge Risks to Move Faster

Patrick O'Shaughnessy

One of the very first things you and I talked about, Rob, was that I started asking about Sohu, which is the name of the first product here. You said we could talk about that in great detail, but the thing you should know is that what we’re really focused on is building a machine that can, at scale, produce these things and generations of them as efficiently and at the highest possible quality levels.

So, we want to build the company—or the machine that is the company—that will produce this thing and subsequent things. I’d like to talk about a few principles or cornerstones of the company.

One of them we've talked about—we've alluded to some of them. You've said velocity, you've said vertical integration. These have become more popular topics. Parallelization is something maybe that we should talk about. But I'm especially interested in your guys' willingness to take huge risks to go faster. Maybe tell your favorite story about why this is the philosophy and what it's allowed you to do that maybe other companies haven't done.

Rob Wachen

There are a number of stories here, but one of my favorites is there was a time when we were getting close to taping out the chip. We realized, "Wait a minute, one of our vendors is way, way behind schedule." We had 2 very bad options. One option was to keep the current vendor and push our timelines out by on the order of a year. Another option was to switch vendors, start over, and also push our timelines out by a year. Neither of these was a good option. So we had to go look for option number 3.

What that was, was we figured out they're all in Bangalore—the team that's actually going and doing the work. We went out and shipped dozens of our top engineers across the world to Bangalore for 6 months. I was there as well. I lived in Bangalore for 4 1/2 months personally. Every morning, we'd walk across the crazy-busy Bangalore streets into the office. We'd be the first ones in. We'd build a wide variety of tools, both things like auditing a huge amount of the code that was going in, building a bunch of tools as well to make this go even faster, and making sure we're making the right design decisions on the spot, right there. No 12-hour back-and-forth. We could decide immediately. Then, at 1:00 a.m., we'd walk back through the now-empty Bangalore streets and do it all again the next day.

We still had a bunch of the team in the US. We ran these 12-hour-on-each-side handoffs, where we had a 24-hour cycle. At 8:00 a.m. and 8:00 p.m. every day, we'd all get on Zoom, share all the data, and say, "When I wake up, this must be done. We must get this chip out." It was extremely intense. At the same time, we saw other chips at the same stage as us with that same vendor that ended up taking years and still aren't out today. They still haven't taped out today. It's that level of extreme urgency that's required to bring products to market.

Patrick O'Shaughnessy

What is the key to doing this well? This has become a trope of Elon, mostly—that his special skill, and others who seek to emulate him would try to do this too, is figuring out what the binding constraint is and just flooding the zone personally on that thing, which is kind of like going to Bangalore or something. It seems like this is a central tenet of the business and of any business that's going to do this kind of vertical integration. What's the key to doing that well? Again, what have you learned about that specific act?

Rob Wachen

For me, I think there were 2 key tricks to this. The first one is that you can't build a chip alone. It's got to be a team problem. Your most important job is to go get great people to go with you and great people to be inspired and excited to do crazy things like this. It is a huge ask to say, "Hey, guys, uproot your lives for 6 months, or in one case, 12 months." We had sent one guy out well ahead. It sucks. But we're lucky to have team members who are in it for the right reasons.

I think the second big thing, too, is being able to make decisions very fast. One of the worst things is when there's a factory or a vendor who's waiting for you to make some call and is then just stalled. This happens all the time, even for very small things. So send folks, delegate a big amount of responsibility to them, and say, "Make a reasonable call. It's okay if you're wrong every now and then. But I would much, much rather be right most of the time and give an answer immediately than wait every time for the perfect response. Speed wins."

Patrick O'Shaughnessy

What about spending money to go faster? There's this learn-by-doing thing, which has become so interesting, and as the world has gone away from software and towards more hardware again in the world of technology, we've outsourced so much of the learn-by-doing by shipping stuff overseas and effectively just being the idea guys here in the US. It seems like that's obviously reversing, and you've adopted this way of learning by doing—you want to be in that iteration learning loop. Part of that is willingness to spend and take risks with dollars. Can you talk about that a little bit?

Rob Wachen

I think there's a great quote: "The biggest risk is not taking risk." It's very similar here. Every day there's over a billion dollars of revenue in this category, and a lot of it's inference. So every day we don't ship, we're just leaving tons of opportunity on the table. Your willingness to spend money should be extremely high if you can get a very clear ROI out of it.

Gavin Uberti

We have this concept that we call pre-fetching, which is when you're waiting for one thing to get done, and you know you're going to do other things once you have it. Are there ways that you can parallelize the entire schedule? For example, we know our chip is going to come back on a certain date. We want everything possible that could be done without the chip to be done before the chip lands.

This costs a lot of money. It means that we want to build our entire software stack beforehand. We actually shipped racks to customer data centers without our chips in them, with all of the networking, all the CPUs, and all the storage set up, so we could bring all that data center software up before the chips came back. It meant that we took over 700 FPGAs and put the entire full-radix chip on an FPGA cluster and ran a dozen different models with our full inference stack on them before the chips came back.

It means that we built a thermal chip to mock the thermal profile of our chip and built cold plates based on that before the chip came back. It means we had the entire production line ready. It means we did many revisions of the circuit board. The entire product was ready to go before the chips came back.

And this is what it gives you. There is another very famous AI chip company that took 10 months to go from getting their silicon back to having it running inference in a rack. This was publicly announced to their investors, and it was a really big deal. We were able to do it in 40 days.

It's because by the time the chip came back, everything was boring. The software was already written. The rack was already there. The production line was already set up. We were just getting everything together. You don't always catch everything. You make some tweaks on the fly, and then off you go.

Patrick O'Shaughnessy

Although in that particular case, too, that was a big part of it. Also, I think the shift made a big difference, too.

Rob Wachen

We went out and literally had a day shift and a night shift. There were team members who would come in at 10:00 a.m. and leave at around midnight.

Patrick O'Shaughnessy

Yeah.

Rob Wachen

You would come in at midnight and leave at 10:00 a.m. You're running around the clock to get to those 40 days. Yeah, I mean, over half the company lives next to the office, so it makes it easier to do that type of thing.

Patrick O'Shaughnessy

You pay them to do that, right? Pay them extra? Do you still do that?

Rob Wachen

Yeah, the invisible hand does wonders. It works for me, too. We're both there.

Patrick O'Shaughnessy

I'd love to take one big step back and talk a bit about just the broader ecosystem here. The amount of shortages on the supply side, the exposure of risks in the global system, and the supply chain around this stuff has become everyday Wall Street Journal front-page news. The stocks that people are watching and investing in and excited about—if you think about the memory stocks, these were boring, commodity-like nothing burgers 5 years ago. Now they're at the center of global attention.

If you just assess the global, connected supply chain that's required to make stuff like this possible—just riff on it. What scares you? What's working well? What needs to change? What do you hope you change by virtue of how you build this thing? What's your assessment of this story right now?

Rob Wachen

I think that one of the most undervalued pieces of the supply chain story is that, in almost none of these cases, do you buy something and never talk to the vendor again. You have to collaborate. That is the most important part to being successful, I think, in chips with TSMC or with memory vendors. You need that partnership.

I think that for TSMC in particular, people don't understand why it is so valuable. People look at the technology, and the technology is the best in the world. But for me, the real value is all in the service. TSMC customer service is way better than I have seen at any other company in any other industry.

It's the kind of thing where if you say, "Hey, we're trying to improve your yields by making this change," you can make them a recommendation, and they'll run the experiment on their own dime, in our case, to see if they could actually get the higher yield. When we found that we were right and the experiment worked, they moved it over to the rest of the line. That kind of thing just doesn't happen in most industries.

If I go to the steelworks plant and say, "Hey, I want you to change the composition of the steel," they'll say, "Screw you." Not TSMC. That is why they are number 1 and why they're going to win.

Patrick O'Shaughnessy

One of the things that matters a ton is power availability and time to power.

Gavin Uberti

And the problem is, the more power you want, the more shortage there is. It's actually very similar to chip clusters. Why is Colossus charging $12 an hour for Blackwells? It's because they're the only place you can buy 20,000 of them at once, right? Why is the 500 MW data center so hard to find? It's the exact same reason.

One of the things you need to think about is: How do we get way more juice out of each megawatt? People are looking throughout the entire stack, whether it's just improving the PUE, but also entirely new hardware to get the most tokens per megawatt to solve this problem. But fundamentally, building new buildings is hard. It's much easier to go from 100 MW to a gigawatt, and then from a gigawatt to 10, and 10 to 100. We are pushing the limits of what's possible on these timelines. So, there are a lot of people trying to scale in their data centers as much as trying to scale them out.

Patrick O'Shaughnessy

Yeah, I mean, one of the interesting things about a system like this is what it replaces. If I think about a rack like this versus, I don't know, a set of Blackwells or Rubins or whatever's coming next, how should I conceptualize that? It's not just watts; it's also physical space, to your point. Cerebras talked about this in their recent earnings call: This is literally the problem. There's literally no space to put the systems. How should I conceptualize what this represents or replaces in terms of other units of compute?

Gavin Uberti

Here's how customers think about deploying models generally. When I'm building a data center or I'm building a cluster, it's not in the abstract of, "Oh, I like these chips, and this is the power and footprint and so forth." It's, "I have a real production workload I'm trying to serve. For my product to be useful, there's a certain speed I need to serve at. For certain products, it's really fast, and for certain products, it's really slow. Whatever the speed is, this is my speed."

The question is: In a given amount of power, how many users can I serve while guaranteeing that speed? So, another way to put it is, this is what's called interactivity: What is my throughput? We are just finishing the early innings of the AI infrastructure boom, where people really just cared about speed. GPUs were not able to reach a lot of the speeds of other types of chips, like all these SRAM chips—thousands of tokens per second—and that enabled tons of new use cases that got people very excited.

There's an entirely new wave of AI chips, us being one of them, that are all going to be able to hit these speeds. The question then is: If you're hitting these speeds, what is the number of users you can serve at the same time? By proxy, if I have a 100 MW data center, how many software agents can I run at the same time? When people are doing that evaluation, our hardware is generally going to be able to get you an order of magnitude more concurrency at a given level of interactivity. That directly translates into tokens per watt, tokens per dollar, and all the things people care about when they're actually serving these giant mixture-of-experts models at scale.

Patrick O'Shaughnessy

There are these now-famous interactivity curves, right? Not many people publish them, but you can see a Blackwell curve, and you can see an AMD curve, which is a little bit worse than Blackwell's, and it's still an $800 billion company. So, if you think about what the impacts are of shifting that curve—not just a little bit further out, but much further out—

Gavin Uberti

Yeah.

Patrick O'Shaughnessy

What are the things that most excite you about what this will enable?

Gavin Uberti

I want to go out and solve some of the hardest problems, and I want to go solve these in much less time. There were things growing up that I was not sure I'd be able to live to see. For example, the unit distance conjecture was one of the things I thought about in college.

Patrick O'Shaughnessy

Yeah.

Gavin Uberti

And I was not sure I'd see that proven in my life. This was done by an AI model, and it was done over a long period of time. But if you're able to run the same model 10 times faster, you can shrink the time to go have these breakthroughs.

There are a huge number of other problems in math like this as well that I worry will take 100,000 years to prove. You can either have a much smarter model or a model of the same intelligence running much faster. You can then shrink that time, and I can see it. It's so cool seeing these breakthroughs get made. I am so, so excited to see much more of this happen.

Patrick O'Shaughnessy

I think too often people think about tasks and applications and stuff in these very short time horizons. Doing a chat and having it be 50% faster is nice, but it's not game-changing. As these agents go to longer and longer time horizons and the models get more and more capable, you're going to see gigantic bodies of work that would take months of compute.

When we think about this in wall-clock time, if you talk to a pre-training researcher at a lab, they'll tell you that wall-clock time is often one of the most important things that matters. What wall-clock time means is the time from starting your run to finishing it, to actually getting data back. If you can shrink this time from a 6-month run to a 2-month experiment, you're going to be able to do many more iterations, and people will make changes to the model architectures to actually improve the wall-clock time.

It's very similar here in terms of how we think about the use cases. The exciting part about super-low-latency decode is that wall-clock time on long-horizon tasks becomes much shorter. A year-long compute build would now take a month, and that month-long compute build will now take 3 days, and that 3-day compute build will now take 7 hours, and so forth and so forth.

That's the thing that I think is really hard to internalize, because the models are just getting capable enough to do this stuff. I thought it was really cool months ago when Cursor announced that they had a bunch of coding agents build an entire browser from scratch in a week. Totally nuts. That will soon happen in under an hour. There are going to be many of those types of things that happen with these massively parallel agents all working on a given task.

What are the ultimate limitations of these systems? Is it just a physics question? How many times faster and cheaper can we get theoretically? How do you think—

Gavin Uberti

There's a lot. There's a lot of room at the bottom, as they say. If you think about chip-to-chip latencies on an NVIDIA product, you're looking at 4,000 nanoseconds to go from one chip to another. We'll be able to do much better than that.

Patrick O'Shaughnessy

What's the mathematical limit?

Gavin Uberti

It's the speed of light. You can do it in just a handful, like 2 or 3 nanoseconds.

Patrick O'Shaughnessy

And they're at 4,000?

Gavin Uberti

4,000 today. There is a lot of room at the bottom. The same sort of thing applies to power efficiency. We're able to shave a huge amount by bringing the voltage down by so much, but you could go lower. You could go much, much lower. It's very challenging, but when I think about 20 or 30 years in the future, I think it's inevitable.

6. Kernels, Compilers, and the AI Stack

The same is true for economies of scale and cluster scale-up. For a long time, 8 chips was the biggest scale-up domain. Then they had the NVIDIA NVL72, bringing it to, well, 72. But you can be way, way bigger. You look at a fab, for example: You have a $40 billion single monolithic building with only a handful of lines running through it. You could have the same kind of thing for some futuristic mega-cluster—a $40 billion or $100 billion giant mega-token factory serving one or a handful of models for a massive number of users to get that same economies-of-scale effect. Same model, massive number of people.

Patrick O'Shaughnessy

You mentioned kernel engineering and that being your first job. That has emerged from being something nobody had ever heard of in their lives to now something that you hear about all the time—the importance of it to eke more performance out of the bare, raw metal. When will that just be something that AI does entirely as well? Are humans still the best kernel engineers? Are they doing it with the assistance of AI systems? How far down will humans still be in the loop of designing these things? When will that go away?

Gavin Uberti

Today, it's all very hybrid. The best kernels are still written by human-AI collaborations. Any AI model is built with these fundamental primitives, like matmuls, convolutions, chip-to-chip operations, and collectives. Making these overlap and making them really fast matters enormously.

It's a kernel designer's job to figure out: Where can I overlap? How do I allocate memory? How do I verify that, if there's some issue like a retransmit, it doesn't stall the whole pipeline? These things are very challenging, but they can make your overall performance 3% or 4% better per optimization, and you can do so many.

When we thought about our software stack, we wanted to see where the puck was going to be. 3 years ago, there were kind of two ways you could build software. One of them was to invest heavily in graph compilers. These things are not very performant, but they work out of the box. They don't require a human to come in and tweak all the kernels.

But we went the opposite direction. We are kernel-first programming, and that means that, for a long time, it did not work out of the box. But if you were a kernel expert, you could get incredibly, incredibly high performance.

Rob Wachen

And the thing about this is that now, as the coding models get better and better, they're doing more and more of the kernel-generation task. When the models keep getting smarter, they'll eventually do all of it. They will become superhuman. So, we're going to build for where the world is going.

Even today, we think about our profiling tools or debugging stack from the perspective of how the model will use these tools, more than how humans will use these tools. We sometimes run experiments internally, and we had Codex actually get GPT-OSS running from scratch, just based off our docs, completely by itself.

Patrick O'Shaughnessy

Wow.

Rob Wachen

And they did it, I think, overnight. We think about game selection a lot. What we mean by that is making sure we're investing our energy in the right bets, because regardless of what you choose to work on, it will take tremendous effort.

One of the things that we started with was the explicit decision not to build an arbitrary graph compiler, not to support arbitrary PyTorch, not to support arbitrary CUDA, and not to support arbitrary ONNX graphs. Instead, we envisioned a world where there were going to be under 100 models that actually mattered, and they were all going to look very similar from the underlying mathematical perspective. We were going to build primitives using physics that would accelerate these as much as humanly possible, and we were going to allow the most sophisticated customers to have direct access to the hardware and do whatever they wanted.

That has saved us a tremendous amount of time by not having to build a compiler, and it has allowed us to get much more performance. Funnily enough, when we started, a lot of people dismissed this idea, and the only people who took us seriously were in high-frequency trading because they all hate compilers, too. They all write their own kernels, and we've had dozens of people from high-frequency trading join the team because they saw this philosophy, too.

Patrick O'Shaughnessy

Mhm. What are the limits to vertical integration? How do you know where to draw the line? I'm starting with this question to talk a bit about the broader market. The circumstances of the broader market are really interesting to me, where the vast majority of AI chips get bought by a very small set of customers. Many of those customers are themselves trying to design their own AI chips.

Gavin Uberti

Yeah.

Patrick O'Shaughnessy

OpenAI announced Jalapeño. It seems like this very funny circumstance where the most valuable thing in the world all kind of flows through a couple of chip makers and a couple of chip buyers. They all seem to be thinking about doing each other's job. Then you've got the circumstance where these things go in a data center, and then you've got neoclouds and inference providers in this other part of the stack. You've got model builders and providers.

I can imagine a world where, because you have the best hardware, you design models and build data centers—you leak outside of your current vertical. So how do you think about where to draw the lines for the business? You have a saying that production is the product.

Gavin Uberti

Uh-huh. Ultimately, what matters here is that we know inference is going to be the biggest market in the world. Whoever produces the most tokens is going to be the most valuable company in the world. So all the decisions we make are about how we get the most token capacity online as possible.

Part of that is building a really good product that has way more throughput, that can run at way better latencies and so forth, so we can, per chip we make, get way more tokens online. Another part of it is not doing parts of the stack unless we absolutely have to in order to get to giant scale.

There are parts that we decided to do because it was absolutely required to get the scale, like building the rack instead of just building the chips, and doing a CM model instead of a JDM model. But there are parts of it that are kind of noise to us right now. We're not going and building our own data centers today. That doesn't actually help us get more capacity online.

In general, our customers are actually making power and moving their clusters around to get our chips online because they're such high-throughput. If there was a world where other things were constrained, we would totally go and integrate with them, but the reality is we're just purely focused on getting as many tokens online as possible.

Rob Wachen

I think it just comes down to economies of scale again: at certain parts of the stack, there are huge economies of scale, and at others there aren't. For example, in designing models, there are huge economies of scale there. For chip fabrication, same story. But if you think about building some small metal part inside of that rack, there's not that same effect. We think the natural boundaries are on the chip side, on the bottom, and at the model layer at the top. We'll fill the whole gap between.

Gavin Uberti

A few weeks ago, there was a guy who was running a next-generation AI chip for one of the frontier companies, and he was trying to recruit one of our architects. This person actually kind of did an Uno reverse card and started recruiting the person trying to recruit our guy, and within a week we hired him.

I was going on a walk as we were finalizing the offer, and I was like, "Why are you leaving this super-important project? Why are you deciding to join?" His answer was super interesting, which was, "It fundamentally is not existential for my company for this product to win. For Google, with TPUs, their revenue comes from search."

Patrick O'Shaughnessy

Google won't fail if TPUs fail.

Gavin Uberti

That's right. Meta won't fail if MTIA fails. Microsoft won't fail if Maia fails. OpenAI won't fail if Jalapeño fails. Ultimately, this is our product. It is completely unsurprising that the best chip in the world is built by a company that only builds that chip. It's Nvidia.

Patrick O'Shaughnessy

Right.

Gavin Uberti

And for us, it is completely existential for us to get as much token capacity online as possible. And it recruits a set of talent and recruits support from suppliers and from customers that view it with the level of intensity that we do.

Rob Wachen

Look at the raw FLOP stats. If you compare any of these chips built by labs or by the hyperscalers, the FLOP density for, say, FP8 × FP8 is lower than the Blackwell B300.

Gavin Uberti

Yep.

Rob Wachen

And that makes sense because they don't have to take the risk. They just have to build a similar-enough product and not pay the Nvidia tax.

Patrick O'Shaughnessy

As I think about you guys building the solution, the process of doing so is solving a sequence of really hard challenges. What has been the single episode that was the hardest to overcome?

Rob Wachen

When we were designing the chip, we built this massive FPGA cluster to verify the full chip workloads. FPGAs are digital entities. You can test digital logic, but not analog logic. It turns out that when the chip came back, we began to see issues in our attention datapath producing incorrect results.

We realized, wait a minute, there's a problem with the backpressure logic across a clock-domain crossing that is failing. This is going to cause the chip to produce wrong results. It is very, very hard to solve.

We realized there was one and only one way to solve it: we had to line up 2 clock signals on our chip to within 50 picoseconds. That's literally 50 trillionths of a second. We had to get the signals aligned to this super-small granularity and do it on every chip 2 billion times a second.

Gavin Uberti

A lot of people said this was impossible. We had people quit.

Patrick O'Shaughnessy

Yeah.

Rob Wachen

People literally were like, "This problem is unsolvable, and best of luck, guys."

When you have a problem like that, step 1 is, okay, let's assume the problem is solvable. How would it be solved? First, we realized what we had to be able to do was find a way to move our clock phase by a picosecond, 10 picoseconds. And we had an idea: what if we had these 2 clocks and had them just a little bit apart from each other?

We figured out that if we could figure out the phase, and then use a drifting mechanism to wait for just the right amount of time to get those 50 picoseconds always lined up, we could do this extremely reliably and then lock the phases exactly where they had to be. We could guarantee this would never happen.

People were, I think, somewhat blown away that this worked, and that it worked as well as it actually did. But we made it work.

Patrick O'Shaughnessy

How long did that take?

Rob Wachen

This was actually about 2 weeks.

Patrick O'Shaughnessy

It was a dark 2 weeks.

Rob Wachen

It was a very scary 2 weeks. But it was the kind of thing where, when that kind of thing happens, that is the most important time to invest effort. That is the hardest time to do it, when you feel like things are hopeless. But the sooner you solve that problem, the sooner you can get back to building and scaling our production to mass volumes.

I think a lot of our story is, as Gavin says, assume it is possible. Assume it is possible to have a chip with way more FLOPs on it. Assume it is possible to have a system with way lower latency between chips. Assume it is possible to create a shared memory pool that can run at way higher bandwidth. How would one do it?

A lot of the time when we do experiments, we'll do dozens of experiments and all of them will fail. But we only need 1 to work. There were multiple times—I mean, Gavin, I think you were leading the charge—during our chip bring-up with, I think, 30 different board experiments, and 3 of them worked.

And all three of them were worth their weight in gold.

Gavin Uberti

This is one of the things that people come to me and say, “Gavin, almost none of your experiments work.” And I’ll say, “I only have to get lucky once.”

7. Raising $100M to Survive

Patrick O'Shaughnessy

So one idea for one of these stories that I’m asking about—difficult moments in the company’s history—is around the ability to raise capital to fund the thing. I think when you started it, you knew you’d need capital, but you did not know you’d need the quantum of capital that you’ve ultimately raised and are spending to build the solution, and you hadn’t raised money before.

These are all new things, right? And there were moments where it was really, really difficult, because I was there; I saw it. There were moments where it was extremely difficult to raise the money that you did, without which the company would not exist. It would have died.

And like many great stories, there were many near-death moments, but money specifically in this new world—this isn’t software. You don’t just need a little bit of money. Maybe you could tell the story about the true hardest part about raising money early on, before you had something that you could show people and be so proud of, and performance that you could show them and blow their socks off.

It was just you guys talking about an idea. Talk about the early difficulties raising money, because it was pretty hardcore.

Gavin Uberti

We’ve had some intense moments. It reminds me of probably early 2024, before we raised our Series A. We were at this point where we had done enough of the architecture and enough of the design that we knew the chip architecture was sound. We had to go build it.

There was a lot more to do. We were ready to go into what’s called the physical design stage. We needed to sign an agreement with a physical design vendor, which will cost you at least $40–50 million.

And then we had this realization, as the models were getting bigger and bigger and these giant MoE models were coming out, that we were going to need to build the entire cluster—not just the chip. We were going to need to build boards, build interconnects, build cold plates, and figure out all of the networking and everything. This was going to cost a lot more than the $15 million we had in the bank.

Patrick O'Shaughnessy

And you’re like, man, that was scary. You’re sitting in that moment and you think, “Holy crap, we can’t afford this.” And I began looking at, “How hard is it to go back to Harvard?”

Rob Wachen

At the end of 2023, we put together this memo. We spent around 100 hours on it because we had no idea how people were going to believe us when we asked for the amount of money we were about to ask for.

It was 30 pages, extremely technical and in-depth, covering all the different things we needed to build, all the milestones we needed to hit, how the market was going to evolve, all the new use cases, the cost per token, and all this modeling.

And then we went and talked to investors, and every major investor in the Valley passed immediately. They were just like, “Okay, two kids that just finished Harvard, haven’t taped out a chip, no test chip. Inference—who knows if this is going to be a big market? Everything’s going to be training. The models still hallucinate. This could all be a bubble.”

At the time, the biggest semiconductor fundraises for a Series A were around $40–50 million. We were looking at this and tallying the bill. We were like, “We think we’re going to spend $100 million in the next 12 months. If we really want to do this, if we want to actually get to scale and actually get the performance we’re talking about, this is going to be extremely capital-intensive. How the hell are we going to pull this off?”

Gavin Uberti

I think one of the key ways we got started in this process was that we thought to ourselves, “What is the cheapest possible way we could do this?” And we decided, “Well, if I made almost nothing—”

Patrick O'Shaughnessy

Yeah.

Gavin Uberti

And if I ate nothing but ramen, then we would go ahead and spend basically just the money for the mask, let it rip, and make the tape-out, and that would be that.

Patrick O'Shaughnessy

Right.

Gavin Uberti

And if so, we could do it on $30 million—an obscenely low number. Then we actually went out and got a debt provider. They were willing to lend us the money we needed to get us across this barely ramen-to-a-chip threshold.

From there, I think it was just a matter of catalyzing a series of other steps: “Hey, maybe we can go do one more thing, one more thing, one more thing.”

Rob Wachen

Yeah, so we’re at this moment where we’re like, if we really want to build this company—because we’re not going to half-ass it. We’re not going to go do a test chip and spend years on it and let the entire AI market boom while we could be building the product. If we’re going to do it, we’re going to go all the way.

We’re going to need to find a way to get $100 million. I remember Gavin and I were sitting down in the office in Cupertino late at night, just looking at each other, and we’re like, “Could we cut $500,000 here? Could we cut $100,000 here? How long could we convince everyone not to take a salary?” The math was not going to close. We really needed to solve this.

There was a period of a few weeks where you just go into survival mode, and you call every person that could possibly know an investor. You’re like, “We need $100 million to do this. If we do this, we think this could be one of the most important companies of all time. Do you know somebody who wants to take an aggressive bet, somebody who wants to believe in us? Here’s all the information. We’re an open book. Here’s the team. They’re great people. We’ve been working super hard. We’ve done these things in record time, but we have these 100 things to go. Do you want to do this?”

The snowball starts, and you get $1 million here and $2 million here, and you’re like, “Okay, we’re not going to run out of money this month.” You get a $5 million check, a $10 million check, and you’re like, “Okay, maybe I can buy those FPGAs.”

And the snowball happened. We were very lucky that we ended up putting it all together. We had a board meeting. I showed you the spreadsheet, and we looked at it, and it was $103 million. These were all soft commits. We all looked at each other and said, “We’re going to take it.” And that was the Series A.

Luckily, I think it’s been much easier since then, and we’ve raised almost half a dozen rounds since then, many of them from those investors just doubling and tripling down. That’s allowed us to get to market so quickly. This rack would not be possible had we not been so aggressive.

Patrick O'Shaughnessy

I also think suppliers deserve a little bit of a commendation here.

Gavin Uberti

Yeah.

Patrick O'Shaughnessy

TSMC was willing to work with us back before we’d raised any of the $100 million.

Gavin Uberti

Yeah.

Patrick O'Shaughnessy

This was back when it was still really, really scary. Synopsys actually went ahead and let us get some of their emulators on extremely favorable terms, where we pay over many years.

Gavin Uberti

Yeah.

Patrick O'Shaughnessy

Yeah, basically a big loan. And it takes a lot of belief from your partners to go do this. But I do know that you came out with this very strong team, and all the folks who back you are not in it just out of pure financial incentive. They believe.

Why did TSMC believe, do you think?

Gavin Uberti

This is a great story. Even before you joined in, there was a conference, a SEMI event. I was one of the only young CEOs in semiconductors, and I think that’s kind of a novelty. They asked me to come in there and speak.

I get to the SEMI event, and I am the only speaker there—and the only person there—under 30. I was 22 at the time.

Patrick O'Shaughnessy

22.

Gavin Uberti

So, I go up and speak. There was a speaker's dinner afterward, and by pure luck, I happened to sit next to a very senior TSMC VP. It was a very nice dinner. The former CEO of Arm was there. It was very bougie; everyone was in a suit.

I'm there with this VP, and it turns out we both studied math in college. We both got a little piece of paper and began talking in great detail about how modern AI models work at the actual, tensor-by-tensor level. The guy just gets it. We were talking about, "How do you run this very efficiently? Why is memory such a critical technology to make this work?"

The following day, I get an email from TSMC saying, "Gavin, want to work with Etched? Find a way to make it happen." They've been a great partner ever since.

Patrick O'Shaughnessy

Crazy.

It's amazing to think about some of the tropes. I obviously should break the fourth wall here: I'm a big Etched investor. I've been involved for a long time, and I think the absolute world of you guys, so I'm incredibly biased in this conversation. I'm trying to ask questions that are broader and interesting and could be objections to what you're doing, and we'll keep doing that.

But it's so interesting to me that when you read about investing, everyone cites this idea of being contrarian and right as the quadrant that makes all the money. It sounds really nice, but contrarian means everyone else thinks you're stupid. When you get immediate no's from literally everybody, it's a fascinating quadrant to exist in before you become consensus.

Gavin Uberti

What was it like for you? I'm super curious.

Patrick O'Shaughnessy

Well, it's interesting. At the time, it was, by a lot, the largest first check that I'd written. Suffice to say, I'm not a math expert, a semiconductor expert, or really an AI expert at the time.

It was much more about believing in the concept of this market potentially being huge. You had made very, very clear bets on how the future was going to look and positioned the company to attack those things in a hardcore way. The two of you, and what I felt about you, were the majority of the reason why we made the bet when we did, in 2023 or whatever it was. At the time, it was the biggest.

I think the same thing you said about naivete applies to investing as it does to maybe building a semiconductor startup. I didn't know what I didn't know. When I called experts, they were basically like, "This is stupid." They laid out in very logical terms why this wasn't going to work and why it was such a low-probability bet.

I think one of the things I've learned from it is that you have to damn the base rate. If you invested on base rates, you should do something other than—

Gavin Uberti

There's always the index fund.

Patrick O'Shaughnessy

Yeah, there's always an index fund, exactly. So it's actually never been scary for me. Most of that is probably because I don't know there's a lot I don't know. If I knew more about what you guys have done and the difficulty, I probably wouldn't have done it.

I don't know what that says about maturing as an investor. Maybe I don't want to know a lot more and have some of that healthy naivete. I don't know.

Gavin Uberti

Funny. I think a lot of the traditional semiconductor funds missed the entire AI chip space—all the AI chip companies. All the coding experts missed all the coding companies. I think it's very hard to realize that the constraints have changed.

When you've looked at tape-outs for 20 years and seen so many of them not work on the first try, or the second try, or the third try, you couldn't even run a workload. You totally forget that EDA tools are way better, that FPGAs exist today in a way that they didn't before, and that all the types of validation you can do today just weren't possible before.

For us, a lot of our believers were on 2 sides of it. They were either believers in the market and the team, or they were building chips today and were extremely technical, like the high-frequency trading firms. They would literally audit everything, from the microarchitecture and the RTL to the board designs, the schedule, and the software stack.

We would sit down with 10 of their people who built their own chips, and they would ask us such detailed questions that we were wondering, "Are they going to build the chip?" It was really on either of those sides. If you were anywhere in the middle, you just wouldn't understand it.

Patrick O'Shaughnessy

In the investing world, they often talk about variant perception: something that you see or believe that others don't.

Gavin Uberti

Yeah.

Patrick O'Shaughnessy

I think I've invested, I don't know, 5 or so times in that, and every time you do it, the stakes get bigger and bigger. It does get a little scarier and scarier. Because you guys have been so quiet in the marketplace, I think it's very easy to dismiss you. As the stakes get bigger and bigger, those dismissals are harder to hear.

Gavin Uberti

Yeah.

Patrick O'Shaughnessy

I do think betting on something that you see, when what you hear from the outside world is very different—that perception gap equals opportunity.

Gavin Uberti

Exactly.

Patrick O'Shaughnessy

The last thing I would say is that the accumulated evidence of you guys and your team's ability to solve seemingly impossible problems is one of the most interesting things a company can have. It's binary: companies do this or they don't.

Gavin Uberti

That's the thing. It's a big advantage of people who have been here for a long time. You get some new joiners who are scared shitless. You see a thing like this, and there are old-timers who have been here for all of 2 years, smoking cigars in the trenches.

Patrick O'Shaughnessy

Another one.

Gavin Uberti

Yeah, there's definitely a find-a-way mentality. If you're here, you're here because you assume it's possible, so we can't be saying it's impossible. Everything is solvable, and we're just going to work at it until we figure it out.

There's a favorite story I have about this guy who's kind of a legend in silicon validation who joined our team. We were doing the early stages of what's called wafer sort. When your chips are coming out of the fab, they go out on these wafers, and you have this thing called a probe card that attaches to the wafer before you dice it, with these probe pads. You send electrical signals to basically test which chips are good and bad.

When I dice the wafer into a bunch of chips, I can package them and only package the good ones. We go through our first wafer, and it's around 2:00 or 3:00 a.m. because we're doing it with TSMC over the phone in Taiwan. We have the screen with the wafer that's all gray, and each chip is gray. As you start running the patterns, the squares are supposed to turn green or red.

They all turn red. We're like, "Fuck. This is really bad." Everybody's like, "Guys, take a breath." He leans back and says, "The puzzle begins."

I'm like, you have to have the attitude of, "Yes, you will go out and stare into the abyss. You will go see scary things, and we'll solve them."

Patrick O'Shaughnessy

When did you see the first green square?

Gavin Uberti

It wasn't that day. But in the moment, you're like, "I have worked for years for this. I put my life on the line. I've asked my family to stake everything on it, and then it's red." It's extremely scary.

There's a certain type of person who's just addicted to that feeling of fear and solving it, and we are lucky to have a lot of those people here.

Patrick O'Shaughnessy

If you think about applying all of this earned know-how from these last several years and now thinking ahead to Gen 2, Gen 3, and beyond—

Gavin Uberti

Sure.

Patrick O'Shaughnessy

What will you be doing most differently as a result of everything that you've learned? Just from a conceptual standpoint, how will you attack designing and producing this next one based on what you learned doing it the first time?

Gavin Uberti

It took us a while to get to the primitives that we think are really what matters for scaling inference. We tried a bunch of things early on, from compilers that would turn different models into FPGAs, to burning weights in silicon, to splitting your HBM into KV cache and weights, and all of these different things.

There were a lot of cycles of learning until we got to the point where we realized that fundamentally, if you want to run the majority of tokens in the world, you need to do 3 things. You need to build a chip with the most FLOPs in a given power budget. You need to build a chip that has the lowest latency between chips, so the biggest scale-up domain possible. And you need to produce as much of it as possible.

I think probably in the first half of our journey so far, we learned the first 2. That informed the design a lot, and it informs a lot about the bets we're making in the future with low-voltage inference and cluster-scale memory.

But the production part, I think in the past year, has made it extremely obvious how much people want to deploy this stuff if you can have it available today. The best ability is availability. If I have 1,000 chips today, someone's going to use them. We need to build a chip that's not just way better than what's been built before; it needs to be available at many-gigawatt scale.

We need to be able to build a product that is producible at gigawatts per month in the limit. As we think about that, a lot of the design decisions we're making with the next generation, which you've seen already, are just about simplicity. That means removing tons of parts, trying to assemble and disassemble the thing again and again, and learning how to make the cycle times as quick as possible in production. We need to make sure it's going to be reliable, serviceable, and producible at gigantic scales.

Patrick O'Shaughnessy

What about other problems in the ecosystem that are outside of your control, such as capacity at the leading nanometer nodes at TSMC or availability of HBM4 memory, or some of these other things where everyone is fighting for a scarce unit of capacity or whatever? How do you face up against those realities when you're trying to produce as much as humanly possible?

Gavin Uberti

The people deploying the most compute in the world do think about supply as a bit zero-sum. There are only so many wafers being produced on a given nanometer node, at a given fab, right? And there's only so much memory being produced.

And that's why, actually, for our first-gen product, we built it on a different supply chain than Rubens. We're on 4-nanometer; Rubens are on 3-nanometer. We're on different HBM than Rubens, and so forth.

So it actually is not a zero-sum thing. It's a positive-sum thing where more is more. So often, when we're talking to people deploying at scale, it's not a decision between a gigawatt of a GPU and a gigawatt of us. It's 2 gigawatts. And I think, as much as possible, you need to think about supply chain early in the design decisions, because if you have the most performant product and you can't produce it, then you're just a podcast.

Rob Wachen

That's the other big thing about vertical integration, too. With certain things, like the chips and the memory, you have to go ahead and partner. For most of the other stuff, those are also very highly in-demand components. And the more that you build yourself, the more stuff you can go do on top of what the world can currently build. It is not, “Oh, you're taking availability from somebody else.” You're adding way, way more. And I think that's how you win.

Patrick O'Shaughnessy

One of the things we really haven't talked at all about is the models themselves, which is kind of crazy—the things behind all of this demand. Anything interesting that you would say about the way that you see models progressing based on what we've seen so far? I guess I'm more interested in how you, as thinkers about hardware, think hardware might impact where the models themselves go in the future?

Gavin Uberti

One of the most important ideas that we believe in is that machines don't think like people think. You look at airplanes, for example. Airplanes don't fly like birds fly. When you think about how mechanical devices have to work, it's often very different.

And in much the same way, for people, storing data and loading memory is very cheap for neurons, and doing math is relatively expensive. It is the exact opposite for chips. Generally, loading data is very expensive, and doing math is very cheap.

As time goes on, I think you'll end up finding that math gets cheaper at a rate that is faster than memory gets cheaper. This is due to this fundamental limit on any kind of DRAM device. You should go ahead and think about: How can I make my model use a huge, huge amount of compute? What if I had, for example, many copies running at the same time? What if I activated a huge number of experts? What if I had gigantic experts that go ahead and run on multiple server racks at the same time? That is how I think you'll build models that are the next generation of intelligence.

And context, too. There's been a lot of work on very efficient inference. What if I don't load the full context in the memory? Most of the time, I think that makes a lot of sense. You want to go build a superintelligence. Why can't it go look at a billion tokens of context? Why can't it spend a huge amount of compute to go ahead and read all of it super fast? I would love to be able to talk to a machine that was able to attend to every book ever written in its short-term memory. And I think we're going to get to a point where you can.

Rob Wachen

A theme in models right now is this focus on something called dynamism, which is the ability to control the level of computation and memory spent at a per-token or per-user level when doing attention, as well as the ability to dynamically, in your chip, on the fly, send data to other chips for different MoE models doing certain types of operations.

The reason is fundamentally that, as we're scaling context length, as we're scaling model size, and as we're scaling the amount of computation per user, we're looking for ways to be more efficient. The first thing is, as Gavin says, mixture-of-experts architecture is where maybe we don't need every parameter being used for every token. But maybe there are things where, even at a token level, we can say, “Well, this token needs this context from this other token.” They can share that memory, so we don't have to have the overhead of using the memory as much.

Maybe this token is really important, so we should spend more compute; we should have longer context on that token. So hardware that really accelerates these types of very dynamic computations is extremely important. And, as you can imagine, current hardware that was designed before those types of architectures has lots of overhead in doing them.

You basically end up in these really bad worlds where you have inefficient hardware at doing this dynamism. Therefore, you can't run it very well, or you have these very blocky architectures that are kind of applying blunt force to many different tokens that all need more or less computation.

Patrick O'Shaughnessy

I have 2 questions about the future. We've talked a lot about what you've built so far and how you've built it. The first is about the new ways that people might start using these systems—the raw technology—over longer runtimes, things of this nature.

When inference gets much cheaper, faster, and more accessible, there's more total supply, and it's better, what are the things that you think people will use that capacity to do that are the most interesting and exciting to both of you?

Gavin Uberti

Yeah, there was a viral tweet by Noam Brown where he said that as these models are having longer and longer time horizons, they can do tasks that take, say, 6 months, and there's often not enough time to go and evaluate them for such a long period of time because by that point you'll have a new model out that you'll want to go evaluate instead. And with technology like our cluster-scale memory, you can go ahead and run that 6-month job much faster.

But there's a second piece of this, too. Talking to Noam about it—he's now an angel as well—it's not just the time; it's also the number of people or agents who are working on this. If you're trying to evaluate, can a human build a rocket? You will find that the answer is no. No one person can go and build a rocket. Instead, you have to go and put a team together. And I believe the same thing will be true of agents, too.

If you want to ask, can an agent go out and build some crazy futuristic piece of software, you'll probably need a very large team. Maybe that's 10; maybe that's a million. You have to go and have this enormous amount of that colossal-scale memory to go ahead and have that very short time per token and a huge amount of FLOPs to be able to run that whole fleet.

Rob Wachen

I'm going to be a little futuristic. I firmly believe we are on a global march of inference becoming a majority of global GDP. It may take more than 10 years, but it's going to happen.

Right now, we measure productivity as a society as GDP per capita. But really, it's going to look much more like agents per megawatt, or it may be agents per gigawatt by then. And while we're being futuristic, I think this is the second-to-last year where a majority of the workforce is going to be human.

I think in 2027 you're going to see that there are going to be more agents doing knowledge work than humans. And it's going to be extremely interesting to see what happens. You could imagine a world where, for countries, a majority of their energy ends up going into data centers doing inference, and the energy efficiency of those data centers basically governs how many agents and, therefore, how big their workforce is.

So you're going to see, like Gavin is saying, right now we have 1 agent, or a team of 5 to 10 agents, working on group projects for a couple of days. So you can do pretty cool stuff because they're smart, but it's not going to be civilization-scale.

What happens when you have countries that can have literally 1 billion concurrent agents—like 1 billion people in the workforce—working 24/7 concurrently on the same stuff? I mean, it's just kind of unfathomable what's going to happen. And it's going to be the biggest proliferation of technology humanity's ever seen.

Patrick O'Shaughnessy

I think, well, when you have these huge amounts of demand, you get this idea of economies of scale again.

Rob Wachen

Yep.

Gavin Uberti

When you think about people, I have a brain. I'm not using the whole thing all at the same time. Only part of it is going to be active, and this is the way healthy brains work. And for MoE models, it works much the same way. In an MoE model, only a small fraction of the parameters is being used for any given token at any given moment.

But if you have a large number of users on a piece of hardware, you can kind of take that brain, cut it up into many different experts on many different servers, and run a huge amount of volume through it.

You'll have a bunch of different pieces of traffic. You'll have many of them using each part of the brain at any given point in time, and you'll have to make the cost per thought, the cost per token, way, way lower. So, I think you're going to end up with these giant-scale distributed brains, where the real form factor is a big data center with a bunch of chips, a huge amount of FLOPs, and a huge amount of scale-up interconnect.

Patrick O'Shaughnessy

Do you think we'll see a $1 trillion individual data center?

Gavin Uberti

Absolutely. It is a matter of time. It's like asking, will you see a $1 billion fab, or a $10 billion fab, or a $100 billion fab? It is inevitable that the economies of scale don't stop at, “Oh, $40 billion is the magic number for fabs.” No. The cost per wafer keeps going down as you keep spending more money. And the same thing will be true of, say, plants that go out and make steel, or plants that go out and make tokens.

Patrick O'Shaughnessy

A very smart alien lands on Earth and wants to know from each of you how you would frame up this opportunity that you guys are tackling. What do you say to them?

Gavin Uberti

Frame it as: thinking is really valuable. Every company in the world runs on thinking. And we're entering this really unique moment in time where you have machines that can think almost as well as, and soon as well as, and soon better than, the best humans can.

Building these machines is going to be a huge opportunity. But more important than that, the way in which you run this kind of thinking is going to be very, very different as demand goes higher and higher and higher and higher. There's a unique moment right now to build a new set of solutions, a new roadmap for how you run the future 10^15-parameter models for 1 billion people all at the same time on a gigantic scale-up cluster.

Rob Wachen

We are in a new era of intelligence where the cost of producing intelligence is dramatically cheaper than the value of the intelligence, so we are in a many-year, probably many-decade supply shortage of these tokens. Basically, any chip or any system that can produce tokens is likely to be extremely valuable, and you should find some part of the supply chain of the token—it can be everything from model training down to what we're doing in the silicon and otherwise—to spend time on and push the frontier. And the companies that are the largest are, frankly, going to be the companies that produce most of the global supply of tokens and own the majority of the supply chain of that token.

Gavin Uberti

And importantly, the people who build systems that, as they get more and more chips put together, get cheaper. The way you want this to scale is not that, oh, if I want to go serve 10 times more tokens, I buy 10 times more servers. It must be some solution where, if I serve 10 times more tokens, then I get some economies-of-scale benefit with my chip's scale-up memory tech that allows me to not charge as much as 10 times more for those next tokens.

Patrick O'Shaughnessy

What a ridiculously exciting future that you guys are building to enable. When I did this with Gavin last time, I asked him my traditional closing question, so this time I'll ask you. What is the kindest thing that anyone's ever done for you?

Rob Wachen

During my cancer treatment, there was a big decision I had to make. The doctors came to me and said, “It's time for you to decide. Do you want to get surgery, or do you want to get radiation? Here's the trade-off. If you get surgery, you're more likely to live, but you have to assume you'll never be able to walk again. If you get radiation, you'll be able to walk again, but there's not the same probability that you'll live. You may die. What do you want to do?”

I was 16, and my parents said, “You have to make this decision for yourself.” I thought about it for a long time and decided, “I'm going to do the surgery.” I got the surgery. One of the things they do when you get a tumor resection is something called a necrosis analysis, where they look at all the different cells and say, “Is the cell dead or alive?” Because if you have a bunch of cancer cells that are alive, you have a problem.

They looked at it and said, “You know, you usually want 98%–99% necrosis for us to say you're in the clear. You're below that. You should go get radiation.” There were only a few machines in the world that could actually do the type of radiation I needed. One of them was in Boston. I was in a wheelchair, and I needed to move to Boston for multiple months. Both of my parents decided to move out, drop everything they were doing, and live with me. I'm eternally grateful.

Patrick O'Shaughnessy

Beautiful. Thanks, guys. Amazing conversation.

The Two Harvard Dropouts Who raised $800M to take on NVIDIA | BidClub