[BidClub_]
The a16z Show · · 33 分钟

与 Google、Cisco 和 a16z 共建 AI 的现实世界基础设施

Raghu RaghuramAmin VahdatJeetu Patel

YouTube
TL;DR
  • Amin Vahdat 认为,AI 基础设施建设的规模将是“互联网时代的100倍”,而 Jeetu Patel 则称其是互联网、太空竞赛和曼哈顿计划“三者合一”。 Patel 表示,地缘政治、经济、国家安全和速度等多重要求同时到来,使这轮建设规模被严重低估。

  • Google 已有7代 TPU 投入生产,而7年前和8年前的 TPU 仍保持100%利用率。 Vahdat 表示,电力、土地改造、许可审批和供应链交付才是约束;即使承诺投入数万亿美元,买方也可能在3–5年内无法“兑现这些支票”。

  • 电力稀缺已经决定数据中心建在哪里,以及它们如何互联。 Patel 表示,企业仍处在必要的机架重构和功率密度升级早期;Cisco 已推出“scale-across”系统,让相隔一段在文字稿中被记为“8 900 km”的设施仍能作为一个逻辑数据中心运行。

  • 下一代计算栈将从芯片到软件协同设计,并将在5年内变得“面目全非”。 Vahdat 称这是“专业化的黄金时代”:对特定计算任务,TPU 每瓦效率可达到 CPU 的10–100倍,但即使是最优秀的团队,也需要2½年才能将专用架构从概念推向量产。

  • 网络正在同时成为稀缺算力的首要瓶颈和放大器。 计算与通信之间的负载可能摆动数十乃至数百兆瓦,但成本最高的网络容量可能只有5%的时间真正需要,这仍是一个尚未解决的架构与利用率问题。

  • 推理效率正在提升10倍乃至100倍,但用户会通过要求更聪明的模型和更长的自主运行来消耗这部分收益。 “每美元智能”在上升,但总成本仍可能增加;因此,prefill、decode、强化学习和推理原生基础设施会带来不同的硬件、内存、延迟和网络权衡。

  • AI 辅助工程正从演示走向此前看似经济上不可能的迁移项目。 Google 曾估算从 Bigtable 迁移到 Spanner 需要“7个员工千年”,因而放弃;Vahdat 表示,AI 后来已协助 Google 全代码库完成指令集迁移。Patel 希望 Cisco 的25,000名工程师能在1年内实现2–3倍生产力,同时警告,今天被否定的工具必须在4周内重新测试。

  • Patel 对创业者的警告是,围绕他人模型搭建的“薄包装”不会拥有持久壁垒。 他更看好与产品反馈紧密耦合的模型,以及在产品自有模型和基础模型之间动态路由;Vahdat 则预计,未来12个月真正带来变革的将是 agents,以及能提升生产力的图像和视频系统,而不只是更好的文本聊天。

摘要 · 为研究而整理的核心内容

1. 需求增速已经超过行业部署资本的能力

  • Vahdat 的历史类比是绝对性的:“我从没见过这样的事情。我相当确定,没有人见过这样的事情。”互联网建设规模巨大;AI 将是“互联网时代的100倍”,上行空间可能相当。

  • Google 已有7代 TPU 投入生产,但7年前和8年前的设备仍保持100%利用率。客户只能接受现有产能,而许多用例因为没有空间而被拒之门外。

  • Patel 表示,企业仍处于重建数据中心、适应更高单机架功率需求的早期阶段,超大规模云厂商和 neo-cloud 则走得更快。由于任何单一地点的电力都很稀缺,数据中心正在“有电的地方”建设,而不是把电力引到既有站点。

  • Vahdat 接受行业将投入数万亿美元的预测,但怀疑行业能否迅速“兑现这些支票”。电力、土地改造、许可审批和供应链交付可能在3–5年内限制部署;硬件可以及时采购,但空间和电力资产的折旧周期为25–40年。

2. AI 正在迫使全栈重构,而不是回归大型机

  • Raghu Raghuram 的挑衅是,NVIDIA 正在让大型机卷土重来。Vahdat 反驳称,一个由16,384块 GPU 组成的池,或一个拥有9,000块芯片的 TPU pod,仍然可以动态切分——一个工作负载可能需要256块芯片,另一个可能需要100,000块——因此 scale-out 软件仍将存在。

  • Vahdat 预计,硬件到软件的整个技术栈将在5年内“面目全非”。Google 上一轮转型是将通用集群与 Bigtable、Spanner、GFS、Borg 和 Colossus 结合起来;这一轮同样需要硬件和软件协同设计。

  • Patel 将这一逻辑延伸到组织层面:覆盖“从物理学到语义”的供应商必须“像一家公司一样工作”,同时保持开放生态,而不是建成封闭花园。他指出,公司与超大规模云厂商之间需要深度设计合作,甚至在产品签约前就要开始。

  • Vahdat 称这是“专业化的黄金时代”。对特定计算任务,TPU 每瓦效率是 CPU 的10–100倍,但即使是最优秀的团队,也仍受限于从概念到量产的2½年周期。他预计,随着当前节奏放缓,专业化会进一步加深,因为节省的电力、成本和空间太过可观,不容忽视。

  • 他还表示,地理位置将塑造架构:中国的7纳米芯片可以与充足电力和工程资源结合,而其他地区可能更偏好功耗更低的2纳米设计。因此,监管框架和区域扩张可能催生不同的架构。

3. 网络必须承受极端峰值,同时避免资产闲置

  • Vahdat 表示,建筑内部的带宽正成为首要瓶颈,而网络本身的功耗相对较低:“在这里多花一点,就能在那里得到更多。”已知的通信模式还可能允许采用不需要完整通用性的设计,而不必依赖全功能 packet switch。

  • 更棘手的问题是突发性。公用事业可能看到数十乃至数百兆瓦的功率需求暂停,用于网络通信,随后又恢复到计算;大规模预训练可能只有5%的时间需要庞大网络,而最新芯片随后又可能迁移到其他站点。

  • Patel 称网络是“放大器”:每节省1千瓦的数据包传输功耗,就能把这部分电力重新分配给 GPU。他预计,行业将采用 scale-up、scale-out 和 scale-across 网络,以及推理原生基础设施,而不是简单把训练系统改作推理用途。

  • Cisco 已推出用于 scale-across 网络的 silicon、芯片和系统,使两座相隔一段在文字稿中被记为“8 900 km”的数据中心可以作为一个逻辑数据中心运行。这一设计反映了电力稀缺所造成的设施分散。

  • 他的竞争警告异常直接:如果系统只是围绕 Broadcom 搭建的包装,“那你面对的就是一个将极具掠夺性的垄断”。Cisco 的战略依据,是在高消耗规模下拥有芯片选择和供应多样性。

4. AI 编程收益真实存在,但采用需要持续复测

  • Google 曾计算,从 Bigtable 迁移到 Spanner 将消耗“7个员工千年”;机会成本最终导向了“Bigtable 万岁”的结论。如今,Vahdat 指出,AI 已协助 Google 全代码库完成从 x86 到 ARM 的指令集迁移,使其不受未来 RISC-V 等架构限制。

  • 他还表示,Google 内部已在 AI 协助下完成 TensorFlow 到 JAX 等迁移,速度提升了多个数量级,同时承认仍有部分任务超出这些工具的能力范围。

  • Patel 表示,使用 Codex Cloud、Cursor 和部分 Windsurf 进行代码迁移,尤其是通过 CLI 调试,以及从零构建前端,都取得了很好的结果。老代码更难处理,尤其是基础设施栈更底层的部分。

  • Patel 表示,挑战更多是文化重置,而非技术问题。他要求150名杰出工程师假设工具会在6个月内显著变强,并在4周内重新审视失败案例,而不是搁置6–9个月。他希望 Cisco 的25,000名工程师能在1年内实现2–3倍生产力。

  • 销售准备、法律合同审查和产品营销也取得了不错的结果。Patel 表示,第一版基于 ChatGPT 的竞品分析草稿,通常已经优于从一张白纸开始。

5. 持久产品将把模型、反馈和路由绑定在一起

  • 在被要求不要只预测模型会变好时,Vahdat 表示,模型和 agent 框架已经“好得令人害怕”。让工作“长时间保持相当正确”的能力,可能在未来12个月带来变革。

  • Patel 给创始人的建议是:不要“围绕属于别人的模型搭建薄包装”。持久产品会让模型与产品紧密耦合,通过产品反馈持续改进模型,同时让基础模型承担其他任务,并在不同模型之间动态路由。

  • Patel 还表示,Cisco 在被视为传统公司或过气公司后,正在重新获得势头。他承诺将在 silicon、网络、安全、可观测性、数据平台和应用层持续创新,并邀请创业公司合作。

  • Vahdat 预计,图像和视频将重复文本模型从新奇玩具走向实用工具的路径:2½至3年前,文本模型还只能用来写有趣的俳句;下一阶段,图像和视频模型应成为生产力、教育和学习工具,而不再只是制造新奇内容。

Jeetu Patel

The good news is infrastructure is sexy again, so that's kind of cool. This is like the combination of the build-out of the internet, the space race, and the Manhattan Project all put into one. There's a geopolitical implication, an economic implication, a national security implication, and then just a speed implication that's pretty profound.

Amin Vahdat

I think it's easy to say I've seen nothing like this. I'm fairly certain no one's seen anything like this. The internet in the late '90s and early 2000s was big, and we felt like, “Oh my gosh, I can't believe the rate of the build-out.” This makes it—10x is an understatement. It's 100x what the internet was.

Raghu Raghuram

What better time and place to talk infrastructure? We were back in the green room, and just as the first question was getting answered, I got cut off, so this could be an entire repeat, for all I know. Anyway, let's go.

The first question is similar, so, firstly, welcome to both of you, and thank you for being here. I hope you'll have a great day and a half as well. Both of you have been in the industry for a while, and both of you have lived through many infrastructure cycles. Have you seen anything like this cycle from your vantage point—not from an investor vantage point, but from your internal vantage point, where you're responsible for building things and planning for things and so on? Either of you—where do you want to start? Amin?

Amin Vahdat

Sure. I think it's easy to say I've seen nothing like this. I'm fairly certain no one's seen anything like this. The internet in the late '90s and early 2000s was big, and we felt like, “Oh my gosh, I can't believe the rate of the build-out.” This makes it—10x is an understatement. It's 100x what the internet was. I think the upside is as big as the internet was—same thing, 10x and 100x. Nothing like it.

Jeetu Patel

I'd agree. I don't think there's any prior to this in size, speed, and scale. The good news is infrastructure is sexy again, so that's kind of cool. It wasn't sexy for a long time.

This is like the combination of the build-out of the internet, the space race, and the Manhattan Project all put into one. There's a geopolitical implication, an economic implication, a national security implication, and then just a speed implication that's pretty profound. None of us have ever seen it at this size and scale.

On the other hand, I think we're grossly underestimating it. The most common question I ask right now is, “Is there a bubble?” I think we're grossly underestimating the build-out. I think there's going to be much more needed than what we're putting our projections toward.

Raghu Raghuram

So, where are we, do you think, in the CapEx spend cycle? More importantly, what are the signals that you use internally in your thinking? You have to plan data centers 4 or 5 years in advance. You're buying nuclear reactors and whatnot. How do you think about the demand signals as well as the technology signals? And Jeetu, do the same thing for you, but from the point of view of enterprise and neo-clouds, et cetera. Amin?

Amin Vahdat

We're early in the cycle, certainly relative to the demand that we're seeing. For our internal users, we've been building with TPUs for 10 years, so we now have 7 generations in production for internal and external use. Our 7- and 8-year-old TPUs have 100% utilization.

Jeetu Patel

Oof.

Amin Vahdat

That just shows what the demand is. Everyone would, of course, prefer to be on the latest generation, but they'll take whatever they can get. This tells me that the demand is tremendous, but it also tells me who we're turning away and which use cases we're turning away. It's not like, “Oh, yeah, that's kind of cool.” It's, “Oh my gosh, we're actually not going to invest in this, and there's no option, because that's where we are on the list.”

It's the same with many of you in the room. We're working with many of you in the room, and many of you are telling me directly—and thank you—we need more, earlier.

The challenge here, though, is that we're limited by power, we're limited by transforming land, we're limited by permitting, and we're limited by backup delivery of lots of things in the supply chain. One worry I have is that supply isn't actually going to catch up to demand as quickly as we'd all like. I heard the previous session's discussion of the trillions of dollars that we're going to be spending, which I think is accurate. I'm not sure that we're going to be able to cash all those checks. In other words, you all have some money, but you can't spend it all as fast as you want. I think that's going to extend for 3, 4, 5 years.

Raghu Raghuram

Wow. How do you deal with the depreciation cycles that are involved there? Does the demand curve and the depreciation cycle curve match up?

Amin Vahdat

Fortunately, we buy just in time. The nice thing is that just-in-time applies to the hardware. The depreciation cycle for the space and power is more like somewhere between 25 and 40 years, so we have benefits there.

Jeetu Patel

If you think of the networking side and look at both enterprise and hyperscalers, as well as neo-clouds, I think the story is quite different. The enterprise is pretty nascent in its build-out of true infrastructure. If you assume that 100% of the data centers, at some point in time, will need to get reracked and will need a very different level of power requirement per rack compared with what used to be there in traditional data centers, I just don't think that enterprises are far enough along.

Maybe a few enterprises that are at super-high scale might be there, but I don't think enterprises are far enough along. Hyperscalers and neo-clouds are a completely different story.

To Amin's point about the scarcity of power, compute, and network being the 3 big constraints in this thing, I would say right now that because there isn't enough power in a single location, data centers are being built where the power is available rather than power being brought to where the data centers are. That's why you're seeing a lot of projects being built out all throughout the world.

The other point, though, is that the lion's share of the constraints we're going to have, I think, are going to persist for a long period of time. As data centers are being built farther and farther apart, 1, there's going to be a huge demand for scale-up networking so that you can have a rack that gets more and more networking for scale-up. The second is that you're going to have a lot of demand for scale-out, where you have multiple racks and clusters that need to get connected together.

We just launched a new piece of silicon, as well as a new chip and a system for scale-across networking, where you might have 2 data centers that act as a logical data center and could be up to 8 900 km apart. You'll see that just because there isn't going to be enough concentration of power in a single location. You'll have to have different architectures that get built out.

Raghu Raghuram

Actually, that brings us to the next topic that I wanted to discuss: the future of systems and networking, and so on and so forth. Google brought about the first—or at least the first large-scale—scale-out commodity servers in production for the web revolution, and now NVIDIA is bringing back the mainframe in a different form.

What do you think happens next? Is this the new style of coherent, cluster-wide computing that we need, with shared memory and all sorts of things, or do you think the pattern changes again?

Amin Vahdat

I don't think we're quite back to mainframes, in that it is still the case that people are running on scale-out architectures across these pools. In other words, whether you have GPUs or TPUs, you're not necessarily saying, “Hey, that's my GPU supercomputer.” You're saying, “I've got 16,384 GPUs,” and maybe I'm going to go grab some subset.

Now, I've got uniform, all-to-all connectivity in many cases, which is fantastic. It's the same with TPUs. It's not like I say, “I have a 9,000-chip pod, and I have to make my job fit on that.” Maybe I actually only need 256. Maybe I need 100,000. So I do think that software scale-out is still going to be there.

I'll note 2 things, though. First, you're absolutely right that, say, about 25 years ago at Google and other places simultaneously, there was really a transformation of computing infrastructure. The notion that you would scale out on commodity PCs—essentially the same ones that you could buy off the shelf, running a Linux stack—and that's what you would do for disk, compute, and networking: you all take it for granted, but this was radical. There were many people who thought that this was a terrible idea that wasn't going to work.

I think the exciting thing about this moment right now is that we're going to be reinventing—I'm not saying Google—we are going to be reinventing computing. Five years from now, whatever the computing stack is, from the hardware to the software, it's going to be unrecognizable.

And by the way, there was this co-design, because if you think about it, I'll use Google examples because I know those best: Bigtable, Spanner, GFS, Borg, and Colossus. They were hand-in-hand co-designed with the hardware—the cluster-scale-out architecture.

You wouldn't have done the scale-out hardware if you didn't have the scale-out software.

The same thing is going to happen in this moment. So I think that, actually, the mainframe is going to look very different.

Raghu Raghuram

Okay.

Jeetu Patel

Yeah, I do think there'll be this extreme demand for an integrated system. Right now, we are very fortunate at Cisco, where we do everything from the physics to the semantics—you think about the silicon to the application. Other than power, one of the constraints is how well integrated these systems are and whether they actually work with the least amount of lossiness across the entire stack.

That level of tight integration is going to be super important. What that means is that the industry will have to evolve into working like one company, even though we might actually be multiple companies that do these pieces. When we work with hyperscalers like Google or others, there is a deep design partnership that goes on for months and months before we even do the deal. Once the deal is done, of course, there's a tremendous amount of pressure to make sure that you're moving pretty fast, but I think the industry's muscle of making sure that you operate in an open ecosystem and not be a walled garden is going to get important at every layer of the stack.

Raghu Raghuram

Yep, completely agree. So let's talk about disaggregating the stack a little bit. One of the most interesting topics is processors, right? Clearly, there is an amazing vendor producing an amazing processor that has massive market share today, right? We see startups all the time doing all sorts of processor architectures. You've got an amazing processor inside your fortress. What do you think happens next in processor land?

Amin Vahdat

Yeah, we're huge fans of NVIDIA. We sell a lot of NVIDIA products and chips. Customers love them. We're also huge fans of our TPUs.

I think the future is really exciting, and it's not that I think we've hit the point of, okay, there's TPUs, there's GPUs, there's whatever Trainiums or something else. We're really seeing the golden age of specialization. That's my observation.

If you look at it, a TPU—I'll use that example again because I know it best—for certain computations is somewhere between 10 and 100 times more efficient per watt than a CPU. It's this watt that really matters. That's hard to walk away from, right? Ten to 100 times.

And yet we know that there are other computations that, if you built even more specialized systems for—not just a niche computation, but computations that we run a lot of at Google, right? For example, maybe for serving, maybe for agentic workloads—would benefit from an even more specialized architecture.

So I think that one bottleneck is how hard it is and how long it takes to turn around a specialized architecture. Right now, it's forever. For the best teams in the world, from concept to live in production, the speed limit is 2.5 years. I mean, that's if you nail everything, right? And there are a few teams that do. But how do you predict the future 2.5 years out for building specialized hardware?

A, I think we have to shrink that cycle. But B, at some point when things slow down a little bit—and they will—I think we're going to have to build more specialized architectures, because the power savings, cost savings, and space savings are just too dramatic to ignore.

And this will actually have a really interesting implication for geopolitical structures as well. If you think about what's happening in China, China doesn't make 2-nanometer chips; they make 7-nanometer chips. But they have an unlimited amount of power and an unlimited amount of engineering resources. What they can do is optimize on the engineering side, keep the 7-nanometer chips, and make sure they give people an unlimited amount of power.

We might have a different architectural design where you have to get extremely power-efficient. You don't have as many engineers as you might employ in China, and you can actually go to 2-nanometer chips. Those might be power-efficient in some ways, but they might have thermal losses in other ways. There are a whole bunch of things that have to get factored in on the architecture, which will get more specialized even by geo and by region.

Then, depending on how the regulatory frameworks evolve, how that geo then expands—like if China expands to different regions in the world, you will have a very different architecture take place than if America expands to different regions in the world. So this is a very interesting kind of game theory exercise to go through on what happens in the next 3 years in tech in general. No one knows right now. That's the beauty of the world that we live in.

Raghu Raghuram

Yeah, yeah. So we'll soon be measuring systems by engineers per token in addition to watts per token.

Jeetu Patel

Engineer per kilowatt. Engineer per kilowatt. In the US.

Raghu Raghuram

Networking, right? Obviously, you alluded to it—scale-up, scale-out. In your case, you mentioned scale-across. So it seems to me that networking is also going to get reinvented in a fairly significant way. What are the leading signs that you're seeing—the signals that you're seeing—of the direction networking is going to take?

Amin Vahdat

Yeah, networking is going to need a transformation, for certain. In other words, the amount of bandwidth that's needed at scale within a building is astounding. It's going up. The network is becoming a primary bottleneck, which is scary.

More bandwidth translates directly to more performance. And given that the network winds up being a small power consumer, the delivered utility you get per watt is a superlinear benefit: spend a little here, get way more there. So that side is absolutely there.

I'll put in a plug here: for these workloads, we actually know what the network communication patterns are a priori. So I think this is a massive opportunity. Do you then need the full power of a packet switch when you actually know what the rough circuits are going to be? I'm not saying you need to build a circuit switch, but there is an optimization opportunity.

The other aspect here is that these workloads are incredibly bursty. We've written about this: power utilities notice when we're doing network communication relative to computation at the scale of tens and hundreds of megawatts. Massive demand for power stops all of a sudden to do some network communication, and then bursts back to computing. So how do you build a network that needs to go at 100% for a really short amount of time and then go idle?

And then the same, actually, for the scale-across use case, which we're absolutely seeing: you don't run large-scale pretraining across all your wide-area data-center sites 12 months of the year. This is a problem I think about a lot. Let's say you build the latest, greatest chips in these 3 data-center sites. How long are you going to be there before you migrate to the latest chips in 3 other sites? And then what do you do with the network that you left behind?

People are going to run jobs on them, but you're not going to need nearly the network capacity that you did for large-scale pretraining. So the shift of needing massive networks for 5% of the time—I don't know how to build a network like that. So if any of you do, please let me know.

Jeetu Patel

We're trying to figure it out. It actually is a fascinating problem.

If you think of power as the constraint and compute as the asset, I think the network is going to be the force multiplier, because if you have low latency and low performance and high energy inefficiency, then every kilowatt of power you save moving the packet is a kilowatt of power you can give to the GPU, which is super important.

The other thing is, when you think about scale-up versus scale-out versus scale-up across, you also need to consider that, especially for inference versus training, there are different things that get optimized. You might optimize for latency much more on training runs; you might optimize much more for memory on inferencing.

There are architectural considerations to look at, and I also feel like the way networking will evolve is, rather than it being a training infrastructure that then gets applied to inferencing, you might have inference-native infrastructure that gets built over time. There are good considerations to look at in how all of the architectural components are moving.

Strategically, one of the biggest things happening in networking from our vantage point is that if you're just a wrapper around Broadcom, then you've got a monopoly that's going to be a very predatory one. One of the big reasons Cisco is super relevant is that you don't just have a Broadcom world, with people wrapping Broadcom in their systems around Broadcom; you will actually have a choice of silicon. That choice and diversity of silicon is going to be super important, especially for high-volume consumption patterns.

Raghu Raghuram

So, last question on the system, since you brought that up, and then we’ll move to use cases. Inference—both of you have mentioned it. You talked about it in the context of the processors, and you just started talking about the architecture. Are you deploying specific architectures for inference today, or is it still shared workloads?

Amin Vahdat

We are deploying specialized architectures for inference, and I think it’s as much software as hardware. The hardware is also deployed in different configurations, is the way I would say it. The other aspect of inference that is becoming really interesting is reinforcement learning, especially on the critical path of serving, because latency just becomes absolutely critical. How you would build your system and how you would connect it to one another—and, of course, networking plays a key role there—becomes increasingly interesting.

Raghu Raghuram

Are there singular choke points that, if removed, would accelerate the thousandfold reduction in the cost of inference that we need, or is it just a natural curve that we’re riding down?

Amin Vahdat

We’re seeing massive reductions. I mean, 2 things here. First, maybe many of you are familiar with this: prefill and decode in inference look very, very different. Ideally, you would have different hardware, because the balance points are different. That’s 1 opportunity that comes with downsides we can talk about.

What I would say, though, is that maybe something people don’t realize is that we’re actually driving massive reductions in the cost of inference—10× and 100×. The problem, or opportunity, is that the community and the user base keep demanding higher quality, not better efficiency. As soon as we deliver all the efficiency improvements we’re looking for, the next-generation model comes out, and its intelligence per dollar is way better, but you still pay more and it costs more relative to the previous generation. Then we repeat the cycle.

It’s almost like the longer the reasoning that you have, the more impatient the market gets. For example, if you have a 20-minute reasoning cycle, like with Deep Research, you could have autonomous execution for about 20 minutes, and that was interesting. Now you have most of the coding tools that can go up to 7 hours to 30 hours of autonomous execution. When that happens, there’s actually a greater demand for saying, “Compress the time down.” It’s kind of a self-fulfilling prophecy: you need to have more performance because you’ve been able to go out and do things for a longer amount of autonomous time. It’s almost a never-ending loop where you’ll need to have more performance for inference in perpetuity.

Jeetu Patel

Yeah, though inference intelligence per dollar is a business-model metric, so it’s not just a processor capability.

Amin Vahdat

No, it’s end-to-end, absolutely.

Raghu Raghuram

Yeah. So, let’s change topics and talk about actual usage. Both of you have massive organizations. Where are the key wins that you’re getting today by applying all the AI that’s available to you? Then we’ll talk about what your customers are doing, but I’m actually curious about what you’re doing internally, within the teams.

Amin Vahdat

Coding is the obvious one, and that’s picking up increasing traction and capability. We just, in the last couple of days, published a paper that showed how we applied AI techniques to do instruction-set migration. In other words, we had a fairly massive migration from x86 to ARM, making our entire codebase at Google—a very, very large codebase—sort of instruction-set agnostic, including to future RISC-V or whatever else might come along.

Raghu Raghuram

Your entire codebase—you’re going to make it agnostic?

Amin Vahdat

The entire codebase, because we want and need all of our codebase to be agnostic.

Raghu Raghuram

That’s a crazy-ass project.

Amin Vahdat

Yeah, it was. The motivation for this was that, a few years ago, we had this amazing legacy system called Bigtable and then a new amazing system called Spanner. We decided to tell the company, “Hey, everyone needs to move from Bigtable to Spanner.” Bigtable was amazing for its time, but Spanner was better.

The estimate for doing that migration for Google was 7 staff millennia.

Raghu Raghuram

How much? How much?

Amin Vahdat

Seven staff millennia. We had a new unit that we had to invent to describe it. It wasn’t like we were making it up or that people were being lazy; this is what it was.

Raghu Raghuram

It’s endearing that they came up with that, though.

Amin Vahdat

And you know what we decided? Long live Bigtable. I decided it just wasn’t worth it.

Raghu Raghuram

Yeah.

Amin Vahdat

Honestly, the opportunity cost was too high. We have these sorts of migrations—TensorFlow to JAX, for example. We’ve actually, again somewhat privately but not too secretly, effected this internally with AI assistance, many factors faster. There are other tasks that the tools probably aren’t quite up to whatever standard they need to meet yet, but the area under the curve is getting bigger and bigger and bigger.

We’re seeing probably 3 or 4 really good use cases, and then we’re seeing some use cases that aren’t working yet.

Jeetu Patel

What is working? Code migration is working relatively well so far. We use largely a combination of Codex Cloud, Cursor, and some Windsurf, so code migration tends to work pretty well. Debugging, oddly enough, has actually been very, very productive with these tools, especially with CLIs.

Front-end, zero-to-one projects tend to do extremely well. The engineers are super productive. When you go to older code, especially further down in the infrastructure stack, it’s much harder to make that happen. The challenge we have in orienting our engineers on this is much more of a cultural reset problem than a technical problem.

If someone uses something and says, “This isn’t working,” you can’t put it back on the shelf and say, “This doesn’t work for another 6 or 9 months.” You have to come back to it within 4 weeks and see if it works again, because the speed at which these tools are advancing is so fast. You almost have to get your mental model aligned with where the tool is going.

I was with 150 of our distinguished engineers today, and what I had to urge them to do was assume that these tools are going to get infinitely better within 6 months. Make sure that your mental model reflects where that tool is going to be in 6 months, and think about what you’re going to do to be best in class in 6 months, rather than assessing it for where it is today and then putting it aside for 6 months. Assuming that it’s not going to work for the next 6 months is a big strategic error.

We’ve got 25,000 engineers, and I’m hoping that we can get at least 2× or 3× productivity within a very short amount of time, within the next year. We’ll be able to see if that happens.

A couple of other big areas where we’re starting to see good results are sales preparation going into an account call—really good—and legal contract reviews, which are actually much better than we had thought. The last one isn’t super high inference volume, but it’s product marketing. I think the first ChatGPT take on competitive analysis is always better than what any product marketing person comes up with on their own. We should never start from a blank slate; just start from ChatGPT and go from there.

Raghu Raghuram

We could be talking about this topic for a long time, but they showed me the 2-minute warning. I want to focus on 1 last question here. We’ve got a lot of founders here building amazing companies. What is the most interesting development they should look forward to in the next calendar year, or the next 12 months? What would you say is coming from your company, and what’s coming from the industry, if you look into your crystal ball?

Amin Vahdat

To build on the point, these models are getting more spectacular by the month, and there’ll be a bunch of really exciting ones from whatever companies you like, including ours.

Raghu Raghuram

Well, then I forgot to say: you’re not allowed to say the models will get better.

Amin Vahdat

Everybody knows the models are going to get better, but they’re getting scary good—that’s the part I would say. The agents that get built on top of them and the frameworks for making that happen are also getting scary good. The ability to have things go quite right for quite long over the coming 12 months is going to be transformative.

Raghu Raghuram

Do you want to leak any aspect of your roadmap for the next 12 months?

Amin Vahdat

Not right now.

Raghu Raghuram

Okay. You too?

Jeetu Patel

I’d say the big shift in what I would urge startups to do is: don’t build thin wrappers around models that are other people’s models. The combination of a model working very closely with the product, and the model getting better as there’s feedback in the product, is going to be super important.

You’re going to need foundation models, but if you just have a thin wrapper, I think the durability of your business will be very, very short-lived. That would be something that I would urge you to consider. I think an intelligent routing layer of some sort—that says, “I’m going to use my models for these things, I’m probably going to use foundation models for other things, and dynamically keep optimizing”—will be a good way for the software development life cycle to evolve. Cursor does that pretty well.

What you should expect from Cisco is—truth be told—for the longest time, people thought Cisco was a legacy company.

They were a has-been. I think in the past year, hopefully you've paid attention. I think there's a level of momentum in the business, and there's a spring in the step of the employee base.

So, you should expect, like I said, from the physics to the semantics, in every layer from silicon to the application, a fair amount of innovation in silicon and networking and security and observability and the data platform, as well as applications, from us. We're excited to work with the startup ecosystem, so if you ever feel like you want to work with us, make sure that you reach out to us.

Were you going to say something?

Amin Vahdat

One aspect that I want to highlight about the models is where we were with, let's say, text models 2½ or 3 years ago. They were fun: “Write me a haiku about Martine.” They did a great job. Now they're amazing.

I think what's going to happen in the next 12 months is that the same thing is going to happen with the input and output of images and video to these models. To the extent that, even for images, you can imagine them as productivity and educational tools—not just, “Okay, here's Martine as Superman on black asphalt,” but using them for productivity gains and learning—I think it's going to be really, really transformative.

Raghu Raghuram

Awesome. On that note, we'd like to end this session. Thanks for a great conversation, and thanks, Jeetu.