前沿模型架构的未来:Fractile 创始人兼 CEO Walter Goodwin
Fractile 的核心押注是,前沿模型推理受限于内存带宽和内存成本,而不只是算力。 Goodwin 希望把 DRAM 的容量和经济性,与 SRAM 系统所代表的“每秒数千个 token”结合起来,从而支持长上下文智能体以及拥有数万亿参数的模型。
Goodwin 认为,今天 AI 芯片的多样性,很多只是品牌包装下的架构相似。 NVIDIA 和 AMD 的 GPU、云厂商 ASIC 以及 TPU,普遍依赖 HBM、tensor core 和 TSMC 先进封装。Fractile 的差异化在于,试图在整条技术栈上推动创新,而不是完成前端设计后将物理实现外包。
Fractile 在判断模型规模和上下文长度会超过 SRAM 承载能力后,放弃了最初的 SRAM 架构。 SRAM 能提供极高带宽,却缺乏具备成本效益的容量;即便是高速推理芯片,处理长上下文时也可能不得不回退到 GPU。公司转而追求“从全球最便宜的内存”DRAM 中获得极高带宽的平台,预计明年下半年全面投入运行。
公司约 150 人的垂直整合团队,旨在缩短工作负载研究、架构、物理设计、封装和晶圆厂协作之间的反馈周期。 Goodwin 认为,能否在部署中胜出,可能取决于能否结构性地锁定 3 至 6 个月的领先:“如果你能找到一种结构性锁定 3 至 6 个月领先的方法,所有这些落地都会由你赢下。”
AI 可能大幅压缩芯片开发周期,但无法消除制造时间或芯片的经济约束。 晶圆厂周转仍需 3 至 5 个月,量产爬坡需要 12 至 18 个月,而一款具备财务可行性的芯片需要 3 至 5 年的摊销周期。因此,设计提速意味着手上同时推进更多“押注”,而不是每几周交付一款全新处理器。
Fractile 声称,单芯片带宽可达到 HBM 架构的 25 倍,进而可能改变哪些模型架构具备经济性。 Goodwin 以混合专家模型的稀疏化为例:从 1:16 向 1:128 或 1:256 推进,可能大幅节省计算量,但当前加速器会受带宽瓶颈制约。过去 20 年,算力大约增长了 100 万倍,而内存带宽只增长了约 40 倍。
Goodwin 预计,即便前沿实验室开发自研芯片,也会继续向多家硬件供应商采购。 自研芯片可以带来议价能力、供应多元化和控制权,但排他性依赖风险极高:如果竞争对手在算法效率上实现 5 倍突破,实验室可能在重新设计硬件的 9 个月里陷入停摆。独立平台依然有价值,因为它们让实验室能够在模型层面竞争,而不必把生存押在单一路线的架构上。
1. 全栈整合是 Fractile 对 ASIC 同质化的答案
Fractile 成立于 2022 年夏季,起步时基于两个相互关联的判断:基础模型会具备广泛的泛化能力,而增加推理时算力可以让语言模型获得类似 AlphaGo 的规模化收益。公司的使命由此确定为,让全球最大规模的模型“快得多、快得多”,包括在极端上下文长度下运行。
Goodwin 的行业判断是:表面上的“奢侈选择”掩盖了共同的底层基础。Google 的 TPU、Meta 的 MTIA、Microsoft 的 Maia、OpenAI 的 Jalapeño、NVIDIA 和 AMD 的 GPU 以及其他 ASIC,最终往往都由少数几家 ASIC 公司参与开发和交付。在品牌之下,它们普遍共享 HBM、面向张量的矩阵乘法和 TSMC 先进封装。
传统流程从了解工作负载的架构师开始,经过 RTL 和前端设计,最终由专业团队把逻辑综合为 GDSII——即发送给晶圆厂的晶体管及金属层布局。Broadcom 等开发及供应链服务商会参与项目落地;Goodwin 认为 Broadcom 是其中最大的一家,市值达 2 万亿美元,服务内容包括模拟互连 IP 和针对特定制程节点的物理布局。
Fractile 把工作负载分析、前端设计、物理设计、后端实现和先进封装都放在约 150 人的团队内部。这样做的优势,是能够形成“灵活得多的闭环”,从而选择那些可能还要 1 至 2 年才能量产的架构。
2. Fractile 放弃 SRAM 路线,转向可扩展的 DRAM 带宽
Goodwin 表示,第一个决定性押注其实是推理本身:最近进入推理市场的一些芯片,在 18 个月前还只是训练芯片。部署成本听起来可能只是边际成本,但“每次部署这些模型都会发生”,因此推理经济性具有结构性的重要性。
在最初 2 年里,Fractile 走的是类似 Groq 或 Cerebras 的 SRAM 路线。由于 SRAM 紧邻逻辑电路,能够提供足够带宽,快速搬运权重或 KV cache,并实现每秒数千个 token 的输出。
到 2023 年末并延续至 2024 年,参数规模和上下文长度持续增长,改变了 Goodwin 的判断。SRAM 的容量成为瓶颈;高输出速度的系统无法处理长上下文任务,因此仍不得不依赖 GPU。
修订后的架构转向高容量、低成本,同时具备极高带宽的 DRAM。Goodwin 所说的“更快的马”之分很关键:更快的聊天机器人只是渐进式改进,而让拥有数万亿参数的模型以每秒数千个 token 运行,可能催生长时间运行的智能体,并大幅提升其性能迭代速度。
3. 更快的设计带来更多押注,而非更多“用完即弃”的芯片
Sarah Guo 的反驳值得保留:芯片代际本来就大致每年到来,重大的架构变化则更慢,而物理供应链也施加着无法规避的约束。那么,Fractile 如何可信地宣称架构迭代速度可以像软件一样提升?
Goodwin 承认其中存在张力。新模型大约每 2 周出现一次;注意力机制、MoE 稀疏化和注意力稀疏化也可能每隔几周就发生变化,尤其是在中国先进开源系统中。但共同需求仍然存在:高内存带宽,以及极小 batch size 下的自回归生成,由此形成带宽效率、成本和服务速度之间的权衡。
有些延迟是物理层面决定的:即便加急,晶圆厂周期仍需 3 至 5 个月;旗舰平台扩产需要 12 至 18 个月;而芯片必须拥有超过 3 年的有效寿命,才能支撑 3 至 5 年的摊销周期。Goodwin 表示,他“不相信”每几周就能出现一款全新的芯片。
压缩设计周期买到的是选择权。Fractile 希望维持一组动态调整、方向一致且随时可以启动的押注,以便选出能够长期成立的旗舰方案并扩大规模,同时与由 6 至 9 颗定制 NVIDIA 芯片组成的 NVIDIA 系统竞争。最终目标,是可重复地领先竞争对手 3 至 6 个月。
4. AI 让架构跑得比物理更快
Guo 转述了一位半导体 CEO 的私下估算:从首席架构师提出想法,到最终完成 GDSII,可能需要 10 年。Goodwin 的经验法则是先质疑其中的假设,把这个时间除以 4,再多减去一些。
智能只是整个闭环的一部分。布局布线要处理 NP-hard 问题,相关算法可能运行数天;有限元、热分析和其他仿真也会形成各自的瓶颈。这本质上是一个类似 Amdahl 定律的问题:即便 AI 可以指导模型研究,研究速度仍可能受实验时间限制。
Goodwin 预计,近似方法和类似代理模型的工具能够加速迭代,但短期内不会取代 Cadence 和 Synopsys 的最终签核。两家公司已经与 TSMC 及其他晶圆厂合作数十年,用于验证设计是否符合晶圆厂规则下的 DRC/LVS 要求。他更深层的主张是,应当在每次高成本实验前投入远多得多的推理,甚至在一次真实测试前后进行相当于“100 年”人类思考的分析。
5. 带宽可以重塑模型,也能保住独立芯片市场
Fractile 必须服务于今天的“硬件彩票”——那些为 GPU 或 XPU 类系统训练的模型——同时利用其声称的 25 倍带宽优势,对未来架构施加“引力”。Goodwin 将这一机会称为带宽扩展定律,与人们熟悉的 FLOP 扩展定律并列。
混合专家模型最能说明这一点。在智能水平相同的情况下,将稀疏度从 1:16 提高到 1:128 或 1:256,可能减少计算量,但 HBM 系统很难高效服务这类模型。类似地,一些注意力机制对带宽的要求较低,却需要更多算力才能达到同等智能水平。
这种失衡有其历史原因:过去 20 年,算力大约增长了 100 万倍,而内存带宽只增长了约 40 倍。抬高带宽上限,可能减少达到给定智能水平所需的操作数量,而不只是更快地执行现有工作负载,从而成为整体模型性能的乘数效应。
Goodwin 预计,大型 AI 公司和云厂商会继续保留多家供应商,以获得事关生存的供应保障,并保持对 NVIDIA 的议价能力。他认为,市场会保留一个把速度置于一切之上的高端细分市场,尤其是在开源模型压力下;否则,所有人都会一直使用 Kimi 模型。单押自研硬件则存在非对称风险:如果竞争对手在不兼容的设计上实现 5 倍效率,一家实验室在追赶期间可能“消失整整 9 个月”。共享的独立平台能让前沿实验室继续在模型质量和推理速度上竞争。
完整逐字稿
Now we're trying to create one chip. If you open up an NVIDIA system, there are 6 to 9 custom chips designed by NVIDIA that together create something extremely powerful. So, the gap already exists. It would be great to achieve such productivity that we could start building our own answers in the same way.
The idea of having a dynamic set of bets at any given time that you hope are consistent and ready to run is a huge advantage if you can build such a system and mechanism. This is the equivalent of advanced models in the chip space: if you find a way to structurally secure a 3–6-month lead, you will win in all these implementations.
With me today is Walter Goodwin, founder and CEO of Fractile, a full-stack AI chip company. We'll talk about what it means to be a full-stack company, the technical bets they're making, why they're so focused on memory bandwidth, their predictions for future high-end architectures, and the structure of the chip market today and tomorrow between NVIDIA, AMD, in-house developments, and this new class of accelerators. Congratulations, Walter. Walter, thank you very much for coming.
It's very nice to be here, Sarah.
You founded this incredibly interesting company called Fractile. Can you give us an overview of what you do?
Fractile is a chip company. We're building very, very fast inference chips for the largest models in the world. This focus on speed is what we've had from the very beginning.
We founded the company in the summer of 2022, and it seems like we started to see 2 things at that time. First, of course, was the emergence of foundation models trained on data from the internet that generalized across everything around them. Secondly, several wise people said that we need to find a way to channel more computing power into these models during inference.
We need to find a way to take what was in AlphaGo, where you have a capable neural network that becomes superhuman at scale, and apply that to, for example, language modeling. So our main challenge over the last 4 years has been to find a way to create chips that can handle these huge models at the same time and run them much, much faster than current chips.
We need to do it in a way that scales—to models that go beyond today's capabilities and to extremely long contexts. This is exactly the mission we are carrying out.
1. The Chip Landscape Now
In the 4 years since your company was founded, the chip landscape has become much more interesting. Where would you place yourself in the overall landscape of GPUs and inference accelerators?
There is now an incredible variety of options available. If you look at this whole AI ASIC space, one of the things that really strikes me is the large number of relatively similar chips.
This is a certain structural feature of the industry, especially if you look at the efforts of hyperscalers, because Google started this with TPU more than 10 years ago. You see a paradigm where, while there are a number of in-house chips—Google with TPU, Meta with MTIA, Microsoft with Maia, and OpenAI now with Jalapeño—these chips are ultimately developed and shipped in partnership with a relatively small number of so-called ASIC companies.
These are development and supply companies. Broadcom is the largest of them, a company with a market capitalization of $2 trillion. They help others implement these projects.
So when you look at this apparent “luxury of choice” in the spectrum of what exists today, you see that there's actually a lot in common between these platforms. All of these chips have HBM, a type of DRAM that is common to NVIDIA GPUs, AMD GPUs, and all other ASICs.
You see the same bets on tensor cores to perform matrix multiplication and the same advanced packaging from TSMC. So I think one of the things we're seeing in the industry is that there's still a relative lack of effort that spans the entire silicon technology stack and tries to create fundamentally new capabilities.
It's a structural thing. There aren't many teams that have decided, “We're going to build from what you might call the architectural level, the front-end design level, which is mostly like writing code, but also all the way down to what we call physical design, process engineering, foundry interaction—all the stuff that makes you fly to Taiwan or Korea every week.”
This is something that still tends to fall on the shoulders of a relatively small number of companies.
2. Common Handoffs From Architecture-Focused Players
For people who don't work in the chip industry, what is the general abstraction or handoff process between what architecture-focused players do and companies like Broadcom?
I think if you're, say, Google and you're developing a TPU, you have people on your team who have a very good understanding of the workloads they're trying to accelerate.
Looking back 10 years, this was a world where you built things like tensor cores—a special circuit that the architect came up with that does matrix multiplication extremely well, because you realized that matrix multiplication is the vast majority of floating-point operations in the models that you run.
That's probably the architect's prerogative. It's someone who really understands how the chips behave but also, ideally, has a very deep understanding of the target workloads.
What you'll then see inside those same organizations is a group of pretty smart people who transform this into essentially a description of how this chip should work at the circuit level. A lot of this is what we call front-end design.
Mechanically, it still looks like writing code on a computer, and it's that same code that's then passed to a player like Broadcom. This type of front-end design, which describes the basic idea of the chip and its logic, ultimately becomes what gets sent to TSMC. It's literally a bitmap of sorts.
We have this file type called GDSII, and it's literally information about where to place the metal layers and where to put each individual transistor. This is the complete layout.
You go through this process of synthesizing the RTL into a set of circuits that are then eventually assembled. A big part of the complexity here is, in the case of Broadcom, ownership of the analog intellectual property that's used for chip-to-chip connectivity.
It is also this physical placement that is specific to a particular process node at TSMC. We're talking about 3 nanometers, 5 nanometers, and so on. This part, to this day, is generally an outsourcing activity for all of these projects.
3. Full Stack Approach and Team Setup
You describe what seems like a very large project to create a new chip from start to finish. In terms of workload, you work with a few top players to characterize it and make sure you can service them. What do you think the Fractile team looks like today? How can you do this as a startup?
It's funny. There are many industries where you take a cascade approach to things, and then suddenly a new way of thinking emerges—more agile, to borrow a term from management philosophy.
For us, as a full-stack company, we have a team that encompasses a very deep understanding of workloads. We're actually trying to get ahead of the curve in a lot of places and see where we could change the architecture of the model to better align with our rates.
4. Fractile’s Most Important Technical Bets
What do we think the scaling-law vector is for the specific bets we make? Then we can advocate for it to our customers and partners. It also helps us form a range of bets for a chip that won't necessarily be in mass production for 1–2 years.
This becomes a really important part of the capacity that we have to build. When we look at this ability to create a very flexible loop, it means that inside Fractile we have front-end designers, our own physical design team, and our own back-end implementation team.
We do our own work on advanced packaging, for example.
And that's not a huge amount of staff, is it?
Fractile has about 150 people today. We have limited resources, so to speak, in each of these sectors, but it allows us to create a much more flexible closed loop.
I think this is becoming increasingly important, given the pace at which this industry is moving. It's a constant catch-up in workloads, because in the field of AI chips, like in no other industry before, you need to make the right bets.
This requires a lot of skill and a little bit of luck, and we need to act very quickly. Therefore, such a structural model, where we do everything ourselves, is very different from previous approaches, where there was a stage of work transfer.
That is, you reach a certain level and then hand over the matter to another partner. To some extent, you depend on how this other partner behaves.
What do you think are the most important technical bets the company has made?
For us, I think one of them is the direction we chose. Four and a half years ago, we were at the very center of the topic of inference, but I still felt that for 2 years after that, we were explaining to the world what inference was.
Some of the inference chips today, including some that have recently hit the market, were training chips just 18 months ago. There was a real avoidance of the idea that this small, so to speak, final marginal cost—marginal cost has 2 meanings, right? This sounds like a small cost, but it's actually a cost you pay every time you deploy these models.
So first of all, I think the very bet that we were about to find ourselves in an era of mass deployment was crucial. There was also a bet on speed. That's what really determined a lot of our subsequent actions as an architectural response.
In terms of architecture, we've come a long way. I think for the first 2 years of the company's existence, we were working on an SRAM chip similar to Groq or Cerebras.
We once noted that SRAM is a memory with extremely high bandwidth. It sits on the same silicon die as your logic, so you have extremely high bandwidth between the compute unit, where the mathematical calculations of the models are performed, and your model weights or KV cache during deployment. This is exactly what allows for speeds of thousands of tokens per second in these language models.
But I think one of the things that we started to worry about toward the end of 2023, and certainly in 2024, was the scalability of this approach. I think there are 2 things that are growing with AI today. One of them is, of course, the model parameters. But the other one—and that's what made us worried about this architectural approach—is the increasing length of context, which was becoming an increasingly significant part of how we saw these models being deployed.
That's what has driven us over the last few years toward what we think is a more exciting bet: working more closely with memory manufacturers, as well as our logic foundry partners, to find ways to access higher-capacity memory with extremely high bandwidth. So over the last few years, we've been involved in these kinds of experimental projects to move away from SRAM, looking at ways we can provide, for example, much, much higher bandwidth for DRAM.
That became a very interesting bet for us, because it allowed us to create a platform that will be fully operational in the second half of next year, and that combines the scalability of higher-capacity, lower-cost DRAM found in GPUs or TPUs with all the speed advantages of Groq or Cerebras chips.
I think that's particularly important because, if you look at where speed is really making a difference today, there's a version of this that's like a faster chatbot. But I think it's like when Henry Ford asked a rhetorical question about what people would want: they would ask for faster horses, not a car. A faster chatbot is the “faster horse” in the world of rapid data output.
The ability to take a model with many trillions of parameters and comfortably run it at speeds of many thousands of tokens per second is, in our opinion, a fundamental advancement in AI's ability to create long-lasting agents and radically speed up their performance.
So today, there is a certain tantalizing discrepancy between the properties of the fast-data-output chips we have: they have extremely high memory bandwidth but incredibly low capacity. If you look at the technical details of how they are actually used and implemented today, they don't do long-context processing. For this, you still fall back to using the GPU.
So there's this annoying discrepancy: in the very area where we're almost ready to take these things and make them much faster, we also don't have the technical capability to do so. In a way, like many technical challenges, it boils down to a somewhat mundane observation: we need chips with a certain special, elusive property—extremely high memory bandwidth, so we can load weights and state thousands of times per second, but at the same time, cost-effective memory.
When you look at what it looks like to run data-center-scale inference for thousands of users, the economics really come down to the cost of 1 gigabyte of memory that you're using. So that was another key focus for us: discovering this fundamentally new building block, which is finding a way to get aggressively high bandwidth from the world's cheapest memory, which is DRAM.
You said a few things that I think go against conventional wisdom. One is the bet on where workloads are going, and 2 is the idea that traditional chipmakers can think about delivering chips in generations, every year or a little faster. Big architectural changes happen slower than this update rhythm.
5. Workload Predictions and Compressing the Chip Design Cycle
But you ambitiously stated in Fractile that you believe the speed of change in the architecture and technical solutions you implement can be much higher than that. Can you talk a little bit about what will allow this to happen? I think people today understand the physical constraints of the supply chain better than they did before. What kind of flexibility can there be in this area?
I think there's some nuance here, but at any given point in time for an AI chip, you want to have a chip that's specifically targeted to your workload and that you already have in large quantities today. But if you had a chip like that, you know that in 6 months you would want to get something else, because these workloads change very quickly.
These days, we have a new model coming out about every 2 weeks. This is what you were talking about. Common sense suggests analyzing these models and looking for what they have in common. Fortunately, they have enough in common. It is always the case that the new large language model desperately needs significantly more memory bandwidth to run faster.
And yet these models remain autoregressive with a very small batch size during text generation, which again creates a fundamental trade-off between bandwidth efficiency, cost, and speed of service. So there are certain things that these models have in common and that they need, and memory bandwidth is one of them.
But then, as you say, these evolutions happen, and especially if you look at the advanced Chinese open-source models, the nature of the attention mechanism changes every few weeks. The level of sparsity of the MoEs you use, the sparsity of attention itself. Even greater shocks may await us. More fundamental shifts may occur.
So I think there is a certain duality between the boundaries of the physical and financial worlds and workload requirements that change at the speed of software, although there is a clear desire to release new platforms much faster. I think there are certain fundamental obstacles to this.
At Fractile, we are very excited to be introducing AI into our chip development process. I believe that the ability to control the entire problem-solving cycle allows us to fundamentally rethink the approach to our work.
If you think about a classic law in computer science—Amdahl's law—everything I can parallelize becomes very, very fast, and what I can't does not become as fast. I think it's a similar situation when you're in a long chain of organizations: there's no immediate need or benefit to radically changing processes for a chip design startup that's focused on the front end, even if it's aggressively implementing AI to reduce lead times from, say, 12 months of development to a tiny fraction of that time, if the next stage still has the usual bottlenecks and the usual pace of work.
I think that, to some extent, these bottlenecks will always exist. We all have the same production cycles at the facilities of our foundry partners. From the time you send them the chip design to the time you receive it, it takes 3 to 5 months, even in a rush-order scenario. These are probably the same inherent delays in this area.
When you get that chip back, I think it's vital to look at how these chips become financially profitable, because they have a certain payback period. You really have to have a 3-to-5-year amortization window for this chip for it to be a sound financial decision.
So there are 2 things that are in tension, and I guess I believe 2 somewhat contradictory things at the same time. The first is the enormous value in being able to take the overall chip development cycle and compress it as much as possible. But I think we don't lose the need to make very thoughtful and subtle architectural bets, because the chip that you ultimately build still has to have a useful life of more than 3 years.
For me, the ability to compress the chip development cycle over time means, essentially, the ability to make more attempts to achieve a goal. You want to be constantly ready to take a certain flagship platform and say, “Yes, this is it. This is truly the flagship, and it is what we will scale.” But then it's a 12-to-18-month ramp-up of production, and you expect it to be durable. You expect it to continue to bring value to your customers.
So I think there's a caveat here: I don't buy into the idea that we're going to end up releasing a fundamentally new chip every few weeks just because we've reduced that time. I believe it is the physical world. There is a certain limit to the power of data centers. There are delays in installing these things. And you need to finance the very basic silicon technologies, so there has to be a payback period.
But I think we can get to a world where the smaller this lag—the gap between seeing and realizing a bet at scale—the more you can get tremendous value out of it. I think that's also an area where Fractile is different. It's because the decisions you make about what to build are the right ones.
That's right. I think you want to be able to have more irons in the fire. One of the things that I think is so exciting about the growth of AI is that, as we become more productive, there are 2 responses in the economy. One is, “Oh no, we'll have less work. We may lose jobs.” Another is that we will be able to do more things.
Being at Fractile, we're trying to build 1 chip. We know we're up against competitors that have—if you open up an NVIDIA system—6 to 9 custom chips that NVIDIA has built together to build something very powerful. So there is already a gap there.
It would be great to be productive enough to start creating your own answers in the same way. I think the idea of having a constant set of bets at any given time that you hope are deeply aligned and ready to go is important. It's a huge advantage if you can build such a machine and such an engine. The 6-month gap that this can consistently give you over your competitors is the wedge that allows you to move forward.
And we see this at Fractile. This is the equivalent of the leading edge in the chip industry: if you find a way to structurally secure a 3- to 6-month lead, you win all those implementations.
There’s a certain meme about CEOs who say, “Oh, AI is great,” and then redesign the organization to use it. Instead of 1 programmer, we will have 6 guys, and everyone will be happy.
6. Architect Intent to Output Bottlenecks and Accelerating Trials
Yes. I think there is some truth to this. I was talking to 1 of the top 3 CEOs of semiconductor companies, and he wouldn’t let me make this prediction on air. But I asked him how soon we could go from the chief architect’s idea to a fully fledged GDSII file, and he said, “10 years.” It won’t happen right now.
What is your view on this? Where do you expect to see the first impact in your organization?
A heuristic that has been successful over the last few years is to always question your logical assumptions and then divide the timelines by 4. There’s a world where 10 years seems like a perfectly reasonable time frame, but I would divide that by 4 and probably take a little bit off. I think we will have a space where prototyping will be happening throughout the process in the coming years.
What’s interesting about chip design is that there are still intermediate loops that are common solutions to NP-hard problems. If you really want to get that GDSII file, today you have to use a bunch of layout tools that do the placement and routing with traditional algorithmic methods that run for days on end.
I think this is a very interesting point. If we go back to the workloads, which also fascinate us, many complex tasks look similar in a world where much of the thinking can now be automated. This is a lot of intellectual work, but there is also a certain internal delay in everything you do. I think we see this in the use of AI to drive the development of AI models today. So, RSI works.
We now have extraordinary intelligence guiding our experiments, but we’re limited by the experiments themselves. We’re limited by the computational time. I think the situation is a bit similar with chip design. Maybe if you take the question from the architect’s point of view to GDSII, it’s a bit like the RSI question, in the sense that the things that prevent us from moving extremely quickly are that there are a lot of things in this process that are really computationally expensive in a more conventional sense.
Do you think we’ll generally see this for all such problems, and that they won’t be replaced by surrogate models?
Well, I think there are definitely these kinds of simulation models. I think the driving approximations are all the models we see today for finite-element analysis, thermal evaluations, and so on. I think there’s a huge payoff from this kind of work precisely because of Amdahl’s law: as we accelerate intelligence, we accelerate the period between these experiments, between these trials. They become a bottleneck.
It is much more important to accelerate these tests as well. We’re looking internally at what’s actually at the core of some of these algorithms, and whether there’s a rough approximation of how to do some of that.
I think chip design isn’t going to change at final approval for quite some time. It’s incredibly valuable to have Cadence and Synopsys, which have worked with TSMC and other foundries for decades, create the final signoff that confirms that this product meets the DRC and LVS requirements. It complies with the rules set by the foundry. That’s where this kind of work is extremely valuable.
But could you have some sort of rough placement algorithm in the meantime that would get you almost to the final design and help you iterate faster? I think this is where people should be doing more work, because otherwise it becomes a bottleneck.
The other part, and I think this is true for RSI and AI experiments as well, is that as the levels of intelligence in the thinking that drives the experiment get higher and higher, and as that experiment becomes a bottleneck, I think we think more. That is generally the Fractile workload thesis.
That’s why we’re so passionate about complex problems, and that’s partly why we see Fractile as a chip that lets us accelerate complex problem-solving in general. I think it becomes almost a moral imperative to think more deeply before every experiment that you run.
I think maybe it was Beren Millidge on Dwarkesh’s podcast recently who said that maybe you could think for 100 years before you run the experiment, and then spend another 100 years of human-equivalent time using these models on the results of that AI experiment, because the experiment itself in the middle is quite expensive and takes some real time.
You run into this also in things like chip design, where there are fundamental real-time bottlenecks for more conventional algorithms. We will generate many reasoning tokens before we move on to deploying a specific scheme.
7. Workload Bets on Model Architectural Shifts
You bet on workloads and then work closely with design partners on architectural and model shifts, hopefully a little bit ahead of time so you can actually plan something. What do you see?
The most important thing we do, of course, is this pursuit of speed: maximizing memory bandwidth to be able to run these models faster.
One of the dualities that I think you have to embrace as a chip company is creating something that does equally well across the spectrum of workloads that come up within the idea of a “hardware lottery,” where they’ll train on a GPU or an XPU with HBM. That’s why it’s great that this is a chip that copes exceptionally well with today’s highly autoregressive MoE transformers, running them much faster and more efficiently.
What’s interesting about creating chips with fundamentally new capabilities is that you can also explore whether there are gravitational forces that you can apply to a new landscape of models yourself because of these properties. For Fractile, this idea of having 25 times more bandwidth per chip than an HBM-based chip gives us the opportunity to explore, for example, ideas like scaling laws for bandwidth.
We traditionally think of scaling laws as FLOP scaling laws. We get better and better performance the more FLOPs we put into these models during training and inference. If you look at the known landscape of ideas today, we already see some scaling laws for bandwidth.
The first is models based on a mixture of experts. It is common knowledge that ideally we would make them increasingly sparse. For the same level of intelligence, you’ll save a bunch of operations if you go from 1-to-16 sparsity in your MoE to 1-to-128 or 1-to-256. But one of the obstacles to that is that on modern GPUs and XPUs with HBM memory, it becomes extremely difficult to efficiently serve such models.
You often encounter bandwidth bottlenecks. As a result, you get a very low level of compute utilization. So there are areas where, even while creating solutions for the modern world, you can find ways to unlock greater opportunities.
With modern LLMs, it’s actually very similar to the attention mechanism. There are forms of attention that are less demanding on bandwidth. They are usually more demanding on compute resources for a given level of intelligence.
One of the things that inspires us is improving another property: memory bandwidth, which hasn’t really scaled well on chips lately. Over the past 20 years, we have scaled computing power 1,000,000 times. Memory bandwidth has increased by about 40 times over the same period.
By scaling this threshold, you get the opportunity to save more on something else. We can reduce the number of computational operations we use for these models to achieve a certain level of intelligence. This is a real multiplier of overall performance for such models.
8. The Future of AI Chip Players and Market Structure
I’ll refrain for now from asking whether this changes people’s ideas about how they should train the model. Last question for you regarding the market structure: it’s not at all clear to me what a major AI player or hyperscaler will buy in 5 years, or what chips they will consume—from Nvidia or AMD, or new classes of accelerators—given that at least 4 of these players have their own developments. How do you understand this?
I think what we’re seeing today is a behavior where anyone implementing a solution at scale is trying to leverage as many individual platforms as possible. I believe that the real need for diversity of supply will remain as computing becomes existentially important for these players.
Today, they joke that the main goal of such in-house developments is to reduce the price people pay to Nvidia. There may be some truth to this, because these developments are architecturally quite similar. I would say that this is not a bet on providing fundamental capabilities that other chips don’t have.
In my opinion, it’s a game of: can we build something here, or can I buy something from AMD and then get a better price from Nvidia? It’s also a game of collective power and a certain level of control.
When you see the transition to chips that really enable new capabilities, that’s where I see a real need for what we’re doing, for anyone who wants to deploy AI at the frontier. They need a solution to run these models orders of magnitude faster.
For leading laboratories, this is a window of opportunity for premium-class intelligence, which is their main reason for existence. Otherwise, we would all be using Kimi models all the time. As you feel this pressure from open-source software, it becomes even more important for those working at the frontier to embrace all aspects of this concept.
The best model weights, as well as the fastest deployment, allow you to perform the most logical inferences in the shortest time. I think this premium segment of the chip market, where speed is optimized above all else, is becoming an extremely important part of this world.
One of the notable features of the dynamic between a chip supplier and a leading lab is that, if you look at Fractile, we try to align our work with the needs of the leading edge as much as possible. We also try to be a little bit like some of those teams. We have people who deeply analyze these workloads and so on.
One of the arguments we sometimes need to make is explaining why we believe that the world will support independent chipmakers for cutting-edge technologies for decades to come. Why isn't all of this concentrated exclusively within these laboratories?
I think it comes down to the asymmetric game that all these labs and developers are playing, because they are taking a huge risk if they bet entirely on one hardware card. Let's say I'm the first lab, and I'm completely focused on my own proprietary silicon. Then a second lab makes a computational breakthrough that provides significantly higher efficiency at the same level of intelligence, but it only runs on their chip.
I could just disappear for those 9 months while I try to implement enough of these chips to get that fivefold increase in computational efficiency myself. Therefore, there is an urgent need for these companies to use the same platforms. Now they are playing different games, trying to compete at the level of models.
I believe that making radically different bets at the silicon level is an irrational and very dangerous step for those who are working at the limit of their capabilities.
9. Conclusion
Well, I think that's where we'll end. Thank you very much, Walter.
Thank you very much, Sarah.