# The Future of Frontier Model Architectures with Walter Goodwin, Founder & CEO of Fractile

No Priors · 2026-10-02 · 36 min · https://www.youtube.com/watch?v=OpeCP4wCxkA

## Transcript

Walter Goodwin

Now we're trying to create one chip. If you open up an NVIDIA system, there are 6 to 9 custom chips designed by NVIDIA that together create something extremely powerful. So, the gap already exists. It would be great to achieve such productivity that we could start building our own answers in the same way.

The idea of having a dynamic set of bets at any given time that you hope are consistent and ready to run is a huge advantage if you can build such a system and mechanism. This is the equivalent of advanced models in the chip space: if you find a way to structurally secure a 3–6-month lead, you will win in all these implementations.

Sarah Guo

With me today is Walter Goodwin, founder and CEO of Fractile, a full-stack AI chip company. We'll talk about what it means to be a full-stack company, the technical bets they're making, why they're so focused on memory bandwidth, their predictions for future high-end architectures, and the structure of the chip market today and tomorrow between NVIDIA, AMD, in-house developments, and this new class of accelerators. Congratulations, Walter. Walter, thank you very much for coming.

Walter Goodwin

It's very nice to be here, Sarah.

Sarah Guo

You founded this incredibly interesting company called Fractile. Can you give us an overview of what you do?

Walter Goodwin

Fractile is a chip company. We're building very, very fast inference chips for the largest models in the world. This focus on speed is what we've had from the very beginning.

We founded the company in the summer of 2022, and it seems like we started to see 2 things at that time. First, of course, was the emergence of foundation models trained on data from the internet that generalized across everything around them. Secondly, several wise people said that we need to find a way to channel more computing power into these models during inference.

We need to find a way to take what was in AlphaGo, where you have a capable neural network that becomes superhuman at scale, and apply that to, for example, language modeling. So our main challenge over the last 4 years has been to find a way to create chips that can handle these huge models at the same time and run them much, much faster than current chips.

We need to do it in a way that scales—to models that go beyond today's capabilities and to extremely long contexts. This is exactly the mission we are carrying out.

### The Chip Landscape Now

Sarah Guo

In the 4 years since your company was founded, the chip landscape has become much more interesting. Where would you place yourself in the overall landscape of GPUs and inference accelerators?

Walter Goodwin

There is now an incredible variety of options available. If you look at this whole AI ASIC space, one of the things that really strikes me is the large number of relatively similar chips.

This is a certain structural feature of the industry, especially if you look at the efforts of hyperscalers, because Google started this with TPU more than 10 years ago. You see a paradigm where, while there are a number of in-house chips—Google with TPU, Meta with MTIA, Microsoft with Maia, and OpenAI now with Jalapeño—these chips are ultimately developed and shipped in partnership with a relatively small number of so-called ASIC companies.

These are development and supply companies. Broadcom is the largest of them, a company with a market capitalization of $2 trillion. They help others implement these projects.

So when you look at this apparent “luxury of choice” in the spectrum of what exists today, you see that there's actually a lot in common between these platforms. All of these chips have HBM, a type of DRAM that is common to NVIDIA GPUs, AMD GPUs, and all other ASICs.

You see the same bets on tensor cores to perform matrix multiplication and the same advanced packaging from TSMC. So I think one of the things we're seeing in the industry is that there's still a relative lack of effort that spans the entire silicon technology stack and tries to create fundamentally new capabilities.

It's a structural thing. There aren't many teams that have decided, “We're going to build from what you might call the architectural level, the front-end design level, which is mostly like writing code, but also all the way down to what we call physical design, process engineering, foundry interaction—all the stuff that makes you fly to Taiwan or Korea every week.”

This is something that still tends to fall on the shoulders of a relatively small number of companies.

### Common Handoffs From Architecture-Focused Players

Sarah Guo

For people who don't work in the chip industry, what is the general abstraction or handoff process between what architecture-focused players do and companies like Broadcom?

Walter Goodwin

I think if you're, say, Google and you're developing a TPU, you have people on your team who have a very good understanding of the workloads they're trying to accelerate.

Looking back 10 years, this was a world where you built things like tensor cores—a special circuit that the architect came up with that does matrix multiplication extremely well, because you realized that matrix multiplication is the vast majority of floating-point operations in the models that you run.

That's probably the architect's prerogative. It's someone who really understands how the chips behave but also, ideally, has a very deep understanding of the target workloads.

What you'll then see inside those same organizations is a group of pretty smart people who transform this into essentially a description of how this chip should work at the circuit level. A lot of this is what we call front-end design.

Mechanically, it still looks like writing code on a computer, and it's that same code that's then passed to a player like Broadcom. This type of front-end design, which describes the basic idea of the chip and its logic, ultimately becomes what gets sent to TSMC. It's literally a bitmap of sorts.

We have this file type called GDSII, and it's literally information about where to place the metal layers and where to put each individual transistor. This is the complete layout.

You go through this process of synthesizing the RTL into a set of circuits that are then eventually assembled. A big part of the complexity here is, in the case of Broadcom, ownership of the analog intellectual property that's used for chip-to-chip connectivity.

It is also this physical placement that is specific to a particular process node at TSMC. We're talking about 3 nanometers, 5 nanometers, and so on. This part, to this day, is generally an outsourcing activity for all of these projects.

### Full Stack Approach and Team Setup

Sarah Guo

You describe what seems like a very large project to create a new chip from start to finish. In terms of workload, you work with a few top players to characterize it and make sure you can service them. What do you think the Fractile team looks like today? How can you do this as a startup?

Walter Goodwin

It's funny. There are many industries where you take a cascade approach to things, and then suddenly a new way of thinking emerges—more agile, to borrow a term from management philosophy.

For us, as a full-stack company, we have a team that encompasses a very deep understanding of workloads. We're actually trying to get ahead of the curve in a lot of places and see where we could change the architecture of the model to better align with our rates.

### Fractile’s Most Important Technical Bets

What do we think the scaling-law vector is for the specific bets we make? Then we can advocate for it to our customers and partners. It also helps us form a range of bets for a chip that won't necessarily be in mass production for 1–2 years.

This becomes a really important part of the capacity that we have to build. When we look at this ability to create a very flexible loop, it means that inside Fractile we have front-end designers, our own physical design team, and our own back-end implementation team.

We do our own work on advanced packaging, for example.

Sarah Guo

And that's not a huge amount of staff, is it?

Walter Goodwin

Fractile has about 150 people today. We have limited resources, so to speak, in each of these sectors, but it allows us to create a much more flexible closed loop.

I think this is becoming increasingly important, given the pace at which this industry is moving. It's a constant catch-up in workloads, because in the field of AI chips, like in no other industry before, you need to make the right bets.

This requires a lot of skill and a little bit of luck, and we need to act very quickly. Therefore, such a structural model, where we do everything ourselves, is very different from previous approaches, where there was a stage of work transfer.

That is, you reach a certain level and then hand over the matter to another partner. To some extent, you depend on how this other partner behaves.

Sarah Guo

What do you think are the most important technical bets the company has made?

Walter Goodwin

For us, I think one of them is the direction we chose. Four and a half years ago, we were at the very center of the topic of inference, but I still felt that for 2 years after that, we were explaining to the world what inference was.

Some of the inference chips today, including some that have recently hit the market, were training chips just 18 months ago. There was a real avoidance of the idea that this small, so to speak, final marginal cost—marginal cost has 2 meanings, right? This sounds like a small cost, but it's actually a cost you pay every time you deploy these models.

So first of all, I think the very bet that we were about to find ourselves in an era of mass deployment was crucial. There was also a bet on speed. That's what really determined a lot of our subsequent actions as an architectural response.

In terms of architecture, we've come a long way. I think for the first 2 years of the company's existence, we were working on an SRAM chip similar to Groq or Cerebras.

We once noted that SRAM is a memory with extremely high bandwidth. It sits on the same silicon die as your logic, so you have extremely high bandwidth between the compute unit, where the mathematical calculations of the models are performed, and your model weights or KV cache during deployment. This is exactly what allows for speeds of thousands of tokens per second in these language models.

But I think one of the things that we started to worry about toward the end of 2023, and certainly in 2024, was the scalability of this approach. I think there are 2 things that are growing with AI today. One of them is, of course, the model parameters. But the other one—and that's what made us worried about this architectural approach—is the increasing length of context, which was becoming an increasingly significant part of how we saw these models being deployed.

That's what has driven us over the last few years toward what we think is a more exciting bet: working more closely with memory manufacturers, as well as our logic foundry partners, to find ways to access higher-capacity memory with extremely high bandwidth. So over the last few years, we've been involved in these kinds of experimental projects to move away from SRAM, looking at ways we can provide, for example, much, much higher bandwidth for DRAM.

That became a very interesting bet for us, because it allowed us to create a platform that will be fully operational in the second half of next year, and that combines the scalability of higher-capacity, lower-cost DRAM found in GPUs or TPUs with all the speed advantages of Groq or Cerebras chips.

I think that's particularly important because, if you look at where speed is really making a difference today, there's a version of this that's like a faster chatbot. But I think it's like when Henry Ford asked a rhetorical question about what people would want: they would ask for faster horses, not a car. A faster chatbot is the “faster horse” in the world of rapid data output.

The ability to take a model with many trillions of parameters and comfortably run it at speeds of many thousands of tokens per second is, in our opinion, a fundamental advancement in AI's ability to create long-lasting agents and radically speed up their performance.

So today, there is a certain tantalizing discrepancy between the properties of the fast-data-output chips we have: they have extremely high memory bandwidth but incredibly low capacity. If you look at the technical details of how they are actually used and implemented today, they don't do long-context processing. For this, you still fall back to using the GPU.

So there's this annoying discrepancy: in the very area where we're almost ready to take these things and make them much faster, we also don't have the technical capability to do so. In a way, like many technical challenges, it boils down to a somewhat mundane observation: we need chips with a certain special, elusive property—extremely high memory bandwidth, so we can load weights and state thousands of times per second, but at the same time, cost-effective memory.

When you look at what it looks like to run data-center-scale inference for thousands of users, the economics really come down to the cost of 1 gigabyte of memory that you're using. So that was another key focus for us: discovering this fundamentally new building block, which is finding a way to get aggressively high bandwidth from the world's cheapest memory, which is DRAM.

Sarah Guo

You said a few things that I think go against conventional wisdom. One is the bet on where workloads are going, and 2 is the idea that traditional chipmakers can think about delivering chips in generations, every year or a little faster. Big architectural changes happen slower than this update rhythm.

### Workload Predictions and Compressing the Chip Design Cycle

But you ambitiously stated in Fractile that you believe the speed of change in the architecture and technical solutions you implement can be much higher than that. Can you talk a little bit about what will allow this to happen? I think people today understand the physical constraints of the supply chain better than they did before. What kind of flexibility can there be in this area?

Walter Goodwin

I think there's some nuance here, but at any given point in time for an AI chip, you want to have a chip that's specifically targeted to your workload and that you already have in large quantities today. But if you had a chip like that, you know that in 6 months you would want to get something else, because these workloads change very quickly.

These days, we have a new model coming out about every 2 weeks. This is what you were talking about. Common sense suggests analyzing these models and looking for what they have in common. Fortunately, they have enough in common. It is always the case that the new large language model desperately needs significantly more memory bandwidth to run faster.

And yet these models remain autoregressive with a very small batch size during text generation, which again creates a fundamental trade-off between bandwidth efficiency, cost, and speed of service. So there are certain things that these models have in common and that they need, and memory bandwidth is one of them.

But then, as you say, these evolutions happen, and especially if you look at the advanced Chinese open-source models, the nature of the attention mechanism changes every few weeks. The level of sparsity of the MoEs you use, the sparsity of attention itself. Even greater shocks may await us. More fundamental shifts may occur.

So I think there is a certain duality between the boundaries of the physical and financial worlds and workload requirements that change at the speed of software, although there is a clear desire to release new platforms much faster. I think there are certain fundamental obstacles to this.

At Fractile, we are very excited to be introducing AI into our chip development process. I believe that the ability to control the entire problem-solving cycle allows us to fundamentally rethink the approach to our work.

If you think about a classic law in computer science—Amdahl's law—everything I can parallelize becomes very, very fast, and what I can't does not become as fast. I think it's a similar situation when you're in a long chain of organizations: there's no immediate need or benefit to radically changing processes for a chip design startup that's focused on the front end, even if it's aggressively implementing AI to reduce lead times from, say, 12 months of development to a tiny fraction of that time, if the next stage still has the usual bottlenecks and the usual pace of work.

I think that, to some extent, these bottlenecks will always exist. We all have the same production cycles at the facilities of our foundry partners. From the time you send them the chip design to the time you receive it, it takes 3 to 5 months, even in a rush-order scenario. These are probably the same inherent delays in this area.

When you get that chip back, I think it's vital to look at how these chips become financially profitable, because they have a certain payback period. You really have to have a 3-to-5-year amortization window for this chip for it to be a sound financial decision.

So there are 2 things that are in tension, and I guess I believe 2 somewhat contradictory things at the same time. The first is the enormous value in being able to take the overall chip development cycle and compress it as much as possible. But I think we don't lose the need to make very thoughtful and subtle architectural bets, because the chip that you ultimately build still has to have a useful life of more than 3 years.

For me, the ability to compress the chip development cycle over time means, essentially, the ability to make more attempts to achieve a goal. You want to be constantly ready to take a certain flagship platform and say, “Yes, this is it. This is truly the flagship, and it is what we will scale.” But then it's a 12-to-18-month ramp-up of production, and you expect it to be durable. You expect it to continue to bring value to your customers.

So I think there's a caveat here: I don't buy into the idea that we're going to end up releasing a fundamentally new chip every few weeks just because we've reduced that time. I believe it is the physical world. There is a certain limit to the power of data centers. There are delays in installing these things. And you need to finance the very basic silicon technologies, so there has to be a payback period.

But I think we can get to a world where the smaller this lag—the gap between seeing and realizing a bet at scale—the more you can get tremendous value out of it. I think that's also an area where Fractile is different. It's because the decisions you make about what to build are the right ones.

Sarah Guo

That's right. I think you want to be able to have more irons in the fire. One of the things that I think is so exciting about the growth of AI is that, as we become more productive, there are 2 responses in the economy. One is, “Oh no, we'll have less work. We may lose jobs.” Another is that we will be able to do more things.

Walter Goodwin

Being at Fractile, we're trying to build 1 chip. We know we're up against competitors that have—if you open up an NVIDIA system—6 to 9 custom chips that NVIDIA has built together to build something very powerful. So there is already a gap there.

It would be great to be productive enough to start creating your own answers in the same way. I think the idea of having a constant set of bets at any given time that you hope are deeply aligned and ready to go is important. It's a huge advantage if you can build such a machine and such an engine. The 6-month gap that this can consistently give you over your competitors is the wedge that allows you to move forward.

And we see this at Fractile. This is the equivalent of the leading edge in the chip industry: if you find a way to structurally secure a 3- to 6-month lead, you win all those implementations.

Sarah Guo

There’s a certain meme about CEOs who say, “Oh, AI is great,” and then redesign the organization to use it. Instead of 1 programmer, we will have 6 guys, and everyone will be happy.

### Architect Intent to Output Bottlenecks and Accelerating Trials

Walter Goodwin

Yes. I think there is some truth to this. I was talking to 1 of the top 3 CEOs of semiconductor companies, and he wouldn’t let me make this prediction on air. But I asked him how soon we could go from the chief architect’s idea to a fully fledged GDSII file, and he said, “10 years.” It won’t happen right now.

Sarah Guo

What is your view on this? Where do you expect to see the first impact in your organization?

Walter Goodwin

A heuristic that has been successful over the last few years is to always question your logical assumptions and then divide the timelines by 4. There’s a world where 10 years seems like a perfectly reasonable time frame, but I would divide that by 4 and probably take a little bit off. I think we will have a space where prototyping will be happening throughout the process in the coming years.

What’s interesting about chip design is that there are still intermediate loops that are common solutions to NP-hard problems. If you really want to get that GDSII file, today you have to use a bunch of layout tools that do the placement and routing with traditional algorithmic methods that run for days on end.

I think this is a very interesting point. If we go back to the workloads, which also fascinate us, many complex tasks look similar in a world where much of the thinking can now be automated. This is a lot of intellectual work, but there is also a certain internal delay in everything you do. I think we see this in the use of AI to drive the development of AI models today. So, RSI works.

We now have extraordinary intelligence guiding our experiments, but we’re limited by the experiments themselves. We’re limited by the computational time. I think the situation is a bit similar with chip design. Maybe if you take the question from the architect’s point of view to GDSII, it’s a bit like the RSI question, in the sense that the things that prevent us from moving extremely quickly are that there are a lot of things in this process that are really computationally expensive in a more conventional sense.

Sarah Guo

Do you think we’ll generally see this for all such problems, and that they won’t be replaced by surrogate models?

Walter Goodwin

Well, I think there are definitely these kinds of simulation models. I think the driving approximations are all the models we see today for finite-element analysis, thermal evaluations, and so on. I think there’s a huge payoff from this kind of work precisely because of Amdahl’s law: as we accelerate intelligence, we accelerate the period between these experiments, between these trials. They become a bottleneck.

It is much more important to accelerate these tests as well. We’re looking internally at what’s actually at the core of some of these algorithms, and whether there’s a rough approximation of how to do some of that.

I think chip design isn’t going to change at final approval for quite some time. It’s incredibly valuable to have Cadence and Synopsys, which have worked with TSMC and other foundries for decades, create the final signoff that confirms that this product meets the DRC and LVS requirements. It complies with the rules set by the foundry. That’s where this kind of work is extremely valuable.

But could you have some sort of rough placement algorithm in the meantime that would get you almost to the final design and help you iterate faster? I think this is where people should be doing more work, because otherwise it becomes a bottleneck.

The other part, and I think this is true for RSI and AI experiments as well, is that as the levels of intelligence in the thinking that drives the experiment get higher and higher, and as that experiment becomes a bottleneck, I think we think more. That is generally the Fractile workload thesis.

That’s why we’re so passionate about complex problems, and that’s partly why we see Fractile as a chip that lets us accelerate complex problem-solving in general. I think it becomes almost a moral imperative to think more deeply before every experiment that you run.

I think maybe it was Beren Millidge on Dwarkesh’s podcast recently who said that maybe you could think for 100 years before you run the experiment, and then spend another 100 years of human-equivalent time using these models on the results of that AI experiment, because the experiment itself in the middle is quite expensive and takes some real time.

You run into this also in things like chip design, where there are fundamental real-time bottlenecks for more conventional algorithms. We will generate many reasoning tokens before we move on to deploying a specific scheme.

### Workload Bets on Model Architectural Shifts

Sarah Guo

You bet on workloads and then work closely with design partners on architectural and model shifts, hopefully a little bit ahead of time so you can actually plan something. What do you see?

Walter Goodwin

The most important thing we do, of course, is this pursuit of speed: maximizing memory bandwidth to be able to run these models faster.

One of the dualities that I think you have to embrace as a chip company is creating something that does equally well across the spectrum of workloads that come up within the idea of a “hardware lottery,” where they’ll train on a GPU or an XPU with HBM. That’s why it’s great that this is a chip that copes exceptionally well with today’s highly autoregressive MoE transformers, running them much faster and more efficiently.

What’s interesting about creating chips with fundamentally new capabilities is that you can also explore whether there are gravitational forces that you can apply to a new landscape of models yourself because of these properties. For Fractile, this idea of having 25 times more bandwidth per chip than an HBM-based chip gives us the opportunity to explore, for example, ideas like scaling laws for bandwidth.

We traditionally think of scaling laws as FLOP scaling laws. We get better and better performance the more FLOPs we put into these models during training and inference. If you look at the known landscape of ideas today, we already see some scaling laws for bandwidth.

The first is models based on a mixture of experts. It is common knowledge that ideally we would make them increasingly sparse. For the same level of intelligence, you’ll save a bunch of operations if you go from 1-to-16 sparsity in your MoE to 1-to-128 or 1-to-256. But one of the obstacles to that is that on modern GPUs and XPUs with HBM memory, it becomes extremely difficult to efficiently serve such models.

You often encounter bandwidth bottlenecks. As a result, you get a very low level of compute utilization. So there are areas where, even while creating solutions for the modern world, you can find ways to unlock greater opportunities.

With modern LLMs, it’s actually very similar to the attention mechanism. There are forms of attention that are less demanding on bandwidth. They are usually more demanding on compute resources for a given level of intelligence.

One of the things that inspires us is improving another property: memory bandwidth, which hasn’t really scaled well on chips lately. Over the past 20 years, we have scaled computing power 1,000,000 times. Memory bandwidth has increased by about 40 times over the same period.

By scaling this threshold, you get the opportunity to save more on something else. We can reduce the number of computational operations we use for these models to achieve a certain level of intelligence. This is a real multiplier of overall performance for such models.

### The Future of AI Chip Players and Market Structure

Sarah Guo

I’ll refrain for now from asking whether this changes people’s ideas about how they should train the model. Last question for you regarding the market structure: it’s not at all clear to me what a major AI player or hyperscaler will buy in 5 years, or what chips they will consume—from Nvidia or AMD, or new classes of accelerators—given that at least 4 of these players have their own developments. How do you understand this?

Walter Goodwin

I think what we’re seeing today is a behavior where anyone implementing a solution at scale is trying to leverage as many individual platforms as possible. I believe that the real need for diversity of supply will remain as computing becomes existentially important for these players.

Today, they joke that the main goal of such in-house developments is to reduce the price people pay to Nvidia. There may be some truth to this, because these developments are architecturally quite similar. I would say that this is not a bet on providing fundamental capabilities that other chips don’t have.

In my opinion, it’s a game of: can we build something here, or can I buy something from AMD and then get a better price from Nvidia? It’s also a game of collective power and a certain level of control.

When you see the transition to chips that really enable new capabilities, that’s where I see a real need for what we’re doing, for anyone who wants to deploy AI at the frontier. They need a solution to run these models orders of magnitude faster.

For leading laboratories, this is a window of opportunity for premium-class intelligence, which is their main reason for existence. Otherwise, we would all be using Kimi models all the time. As you feel this pressure from open-source software, it becomes even more important for those working at the frontier to embrace all aspects of this concept.

The best model weights, as well as the fastest deployment, allow you to perform the most logical inferences in the shortest time. I think this premium segment of the chip market, where speed is optimized above all else, is becoming an extremely important part of this world.

One of the notable features of the dynamic between a chip supplier and a leading lab is that, if you look at Fractile, we try to align our work with the needs of the leading edge as much as possible. We also try to be a little bit like some of those teams. We have people who deeply analyze these workloads and so on.

One of the arguments we sometimes need to make is explaining why we believe that the world will support independent chipmakers for cutting-edge technologies for decades to come. Why isn't all of this concentrated exclusively within these laboratories?

I think it comes down to the asymmetric game that all these labs and developers are playing, because they are taking a huge risk if they bet entirely on one hardware card. Let's say I'm the first lab, and I'm completely focused on my own proprietary silicon. Then a second lab makes a computational breakthrough that provides significantly higher efficiency at the same level of intelligence, but it only runs on their chip.

I could just disappear for those 9 months while I try to implement enough of these chips to get that fivefold increase in computational efficiency myself. Therefore, there is an urgent need for these companies to use the same platforms. Now they are playing different games, trying to compete at the level of models.

I believe that making radically different bets at the silicon level is an irrational and very dangerous step for those who are working at the limit of their capabilities.

### Conclusion

Sarah Guo

Well, I think that's where we'll end. Thank you very much, Walter.

Walter Goodwin

Thank you very much, Sarah.
