那块可能解锁 AGI 的芯片
Naveen Rao 的核心判断是,AI 面临的能源墙需要改变计算机,而不是只把今天的架构继续做大。 他说,美国数据中心已经消耗全国电网约4%的电力,而一些估算认为,未来10年还需要新增400 GW。瓶颈正变成物理问题:“当前范式尽管已经足够优秀、也把我们带到了今天,但无法把我们带到那个水平。”
Unconventional AI 的起点是物理学习理论,而不是传统芯片路线图。 Rao 希望使用自身动力学就能执行有用智能的电路,而不是依赖多层数字抽象:Bornstein 将这一想法概括为“智能就是物理学”(“Intelligence is the physics”),Rao 表示认同。人脑功耗约20瓦,而松鼠或猫可能只需约0.1瓦——如此巨大的效率差距,足以让人重新审视第一性原理。
模拟计算并非要取代数字计算,而是针对可能受益于物理动力学的随机、随时间变化的工作负载。 Rao 关注扩散模型、流模型和能量模型,因为它们的动力学可以写成常微分方程,并有可能映射到物理系统上。他的目标是打造与传统计算并行的“智能基底”,对需要精度的问题仍保留数字计算。
Rao 认为动力学可能补上当前 AI 缺失的因果性,但他反复称 AGI 这一论证“很不严谨”。 他的直觉是,由具有真实时间演化的元素构成的系统,比那些基础中不包含这类动力学的系统,更有可能理解因果关系。他说,今天的模型已经包含智能、也非常有用,但距离 AGI“还差得远”(“nowhere close to AGI”),因为它们仍会犯基础错误,也还不像是在与一个人共事。
商业化的检验标准,是 Unconventional 能否在5年内找到一种类似智能的物理范式,并实现大规模量产。 Rao 认为,TSMC 将是原型开发和最终规模化所必需的合作伙伴;他判断 Google 拥有完整的内部能力,正围绕 TPU 持续进行低风险改进,而 Nvidia 已经建立起占据主导地位的编程平台。Unconventional 正试图构建“比矩阵乘法更好的基底”,但 Rao 没有预设双方必然直接竞争,也为合作留下空间。
执行风险极高,但 Rao 认为这不是一次盲目跳跃,而是有多项局部验证支撑。 人脑提供了存在性证明,超过40年的学术研究显示出这一方向的潜力,动力系统和神经科学理论则提供了工程师可以组合并“打磨”的零件。首个原型可能是规模最大的模拟芯片之一,甚至可能是有史以来最大的模拟芯片,因此,理论、系统、算法、模拟电路和数字电路等领域的人才都构成这一命题的核心。
1. 公司从物理学习出发,而不是从芯片规格出发
Rao 立即纠正了这个前提:Unconventional 严格说“并不是一家芯片公司”(“not a chip company per se”)。公司的起步工作是理论研究——从第一性原理追问学习如何在物理系统中发生——因为他认为,基本未变、已有80年历史的计算架构可以被重新设计。
从无线硬件和实时视频压缩,到神经科学博士,再到 Nervana、MosaicML 和 Databricks,Rao 的经历让跨越边界变得顺理成章。在 Rao 过去的定义里,真正的全栈工程师需要理解硅器件、逻辑、架构、底层软件、操作系统和应用,而不只是 JavaScript 和 Python。
因此,硬件和软件对他而言并不是天然边界;边界只存在于人们选择在哪里划线、哪些东西由自己配置。核心问题是,一项能力最终会在哪里被消耗,然后再让解决方案与问题的规模相匹配。
2. 数字计算赢在可扩展性,模拟计算保留效率优势
数字计算机用固定比特表示数字,以精度误差换取一台能够模拟所有可用算术表达的问题的通用机器。早期模拟系统效率很高,却因制造波动无法规模化;真空管可以可靠地表示高电平或低电平,即便其中间状态难以准确刻画。
Bornstein 将1945年 ENIAC 使用的18,000根真空管,与一些大规模训练系统所使用的 GPU 数量作比较。
风洞体现了另一种路径。工程师不再通过数值方法近似流体动力学、始终留下误差,而是搭建一个物理模拟体,让底层物理规律直接对过程进行建模。
智能可能尤其适合这种模式:神经网络具有随机性和分布式特征,但今天运行在精确、确定性的计算基底上。人脑也展示了一个惊人的效率目标:人脑功耗约20瓦,而松鼠或猫可能只需约0.1瓦。
Bornstein 用“智能就是物理学”(“Intelligence is the physics”)概括生物学上的一面;Rao 表示认同,并解释说,神经动力学直接由化学扩散和神经元的物理属性介导,中间没有操作系统或 API 将计算与物质隔开。
3. 能源稀缺让架构效率成为硬约束
Rao 说,美国拥有全球约50%的数据中心容量,数据中心约占其电网用电量4%。他补充称,2025年夏季,美国西南部开始出现关于停电的新闻报道;如果这一比例升至8%或10%,约束将明显恶化。
发电能力可以扩张,但基础设施昂贵且建设缓慢。Rao 援引的估算显示,未来10年需要新增400 GW 电力容量,而当前扩张速度“每年大约4 GW”;Bornstein 补充说,即便发电量足够,也可能压垮主要建于1970年代的输电电网。
Bornstein 将这一努力描述为人类调动“物种级资源”来发明未来。Rao 说,随之而来的供给缺口意味着必须重新思考问题,同时也主张继续建设更多发电能力。
Rao 拒绝将问题简化为数字与模拟二选一。数字计算仍适合确定性的数值问题;模拟动力学则可能适合跨多个输入进行检索和总结,形成补充传统计算的智能基底。
4. 物理动力学可能在接纳复杂现实的同时保留精度
Bornstein 讲述了一个故事:Steph Curry 曾搭建特殊的追踪系统,以确认篮球击中的是篮筐中央,而不只是最终穿过篮筐。但在比赛中,位置、防守球员、球鞋、场地、篮球的黏性以及出汗的双手,会让每一次输入都独一无二。
大脑能够整合这些变量,同时产生极其精准的行为。这种“模糊、分布式输入产生精确行动”的组合,正是 Rao 希望智能物理基底能够处理的问题类型。
Unconventional 会从当前的模型家族出发,而不是将它们全部抛弃。扩散模型、流模型和能量模型尤其值得关注,因为它们包含动力学,有时可以表示为常微分方程,并可能映射到物理电路的时间演化。
Transformer 仍然有价值,因为它们让 GPU 的计算构件发挥出了极高效率,但 Rao 认为,其参数化方式“没有自然法则”可言。他预计 Transformer 参数空间与其他参数空间之间可以相互映射,认为 Transformer 可能只是依靠“海量参数”来获得结果。
5. 时间与因果性是通往 AGI 的推测性桥梁
当被问及这一路径是否能推动 AGI 时,Rao 的回答刻意保留余地:“坦白说,我确实认为会”,紧接着又说“这说法很不严谨”(“this is hand wavy”)。他的直觉是,一个包含时间和因果性的基础,可能优于不具备这类动力学、或只在数值层面表示时间的基础。
Bornstein 指出,数学系统往往可以在时间上逆向运行,而物理世界至少在人类的感知中通常并非如此。Rao 认为,从拥有真实时间演化的基本单元出发构建系统,可能会产生真正理解因果关系的系统。
Rao 从幼儿身上看到了一条存在性线索:幼儿似乎理解事件会以因果方式展开,而人们也知道,向手臂发出某个特定指令,就会产生某种特定动作。他猜测,大脑天生由因果性基本单元构成,但没有声称自己知道其中的具体机制。
当前机器已经拥有智能,也能提供有用工具,但 Rao 说它们“距离 AGI 还差得远”(“nowhere close to AGI”)。它们仍会犯“愚蠢的错误”(“stupid errors”),人与其互动也还不像是在和另一个人共事。
6. 制造规模与组织广度决定理论能否兑现
Rao 设定了两个里程碑:5年内找到一种类似智能的范式,然后在这一期限内从制造角度实现规模化。如果无法制造10 million 个器件,这项技术就无法解决全球能源问题。
因此,TSMC“绝对会成为合作伙伴”。Rao 说,Google 拥有完整的内部能力;根据他能从公开信息看到的情况,Google 正围绕 TPU 持续进行低风险改进,以服务自身业务。Nvidia 已经建立了所有人都在其上编程的平台;Rao 无法判断 Nvidia 最终会成为竞争对手还是合作伙伴,只能确定 Unconventional 想要的是“比矩阵乘法更好的基底”。
他的信心来自人脑这一存在性证明、超过40年的学术研究、研究人员制造的概念验证器件,以及来自神经科学和动力系统的持续发展理论。伟大的工程,最终就是把不完美的零件组合起来:正如 Bornstein 所说,“这个东西不太合适——把它磨一下,做到合适为止。”
最初几年,公司会以实践型研究实验室的方式运作:先证明一个想法可行,再让制造方面的顾虑关上大门。Rao 预计组建一支混合信号团队,成员覆盖理论研究者、模型专家、系统架构师,以及模拟和数字电路工程师。首个原型可能是规模最大的模拟芯片之一,甚至可能是有史以来最大的模拟芯片。
Rao 偏好早期创业公司式的广泛职责和高自主权文化,因为狭窄的专业分工可能难以适应变化。领导者应提升组织的自主性,在人们对某种路径充满热情时为其让路;每个人都应对好结果和坏结果负责,包括承认:“好吧,我搞砸了。”他的动力来自这样一种信念:改变计算机,可以让 AI 无处不在。他“完全不是 AI 悲观主义者”(“the opposite of an AI doomer”),并将 AI 视为人类的下一次进化。如果 Unconventional 成功,他说,“这个世界很长一段时间都不会忘记这件事。”
I think AI is the next evolution of humanity. I think it takes us to a new level and allows us to collaborate and understand the world in much deeper ways.
Naveen Rao is here, an expert in AI. Naveen Rao is probably one of the smartest guys in this domain. He sees things well before anybody else sees them. You had a lot of success doing Nervana, MosaicML, and Databricks. Why start a new chip company now?
First off, it’s not a chip company per se. Most of what we’re doing, at the beginning, is really looking at first principles of how learning works in a physical system.
NVIDIA, TSMC, Google: are these potential allies for Unconventional AI, or are these competitors?
I think TSMC is absolutely going to be a partner. Google has everything internally. NVIDIA, of course, built the platform that everyone programs on today. So, are we going to be at odds with NVIDIA going forward? I don’t know. We’ll see what the world looks like, but there could be a world where we collaborate.
Has anyone called you crazy yet for doing this?
Oh, yeah. Plenty of people.
Our guest today is Naveen Rao, co-founder and CEO of Unconventional AI, which is an AI chip startup. Prior to that, Naveen was at Databricks as head of AI and co-founder of 2 successful companies: MosaicML, in the cloud-computing world, and Nervana, doing AI chip accelerators before it was cool. We’re here reporting from NeurIPS. Great to have you on the podcast, Naveen. Welcome.
Thanks. Thanks for having me.
So, you were kind of at the vanguard of thinking about what the proper hardware is for running AI workloads.
Absolutely. When you have a hammer, everything’s a nail, I suppose. The early part of my career was really about how to take certain algorithms and capabilities, shrink them, make them faster, and put them into form factors that make those use cases proliferate, like wireless technology or video compression.
You couldn’t do video compression in real time on a laptop back then. There just wasn’t enough computing power, so you actually needed to build hardware to do those kinds of things. The early part of my career was all about that. Then I went back to academia and did a PhD in neuroscience, so you still look at it like, “Hey, can I make something better that’s more efficient?”
And so, you sold Nervana to Intel.
Yeah.
And then founded MosaicML, which is a cloud company. It’s interesting to sort of cross domains like that. I would argue MosaicML was really a software company. How did you make that decision, and why do you think you have these diverse interests?
I think I was—I guess you would call it an OG full-stack engineer. “Full-stack engineer” means something different now than it did back then. I think back then it meant someone who understood potentially devices, like silicon; how to do logic design; computer architecture; low-level software, maybe OS-level software; and then applications. That was a full-stack engineer, and I had actually touched all those topics.
To me, it was very natural to think across these boundaries. Software and hardware aren’t really natural boundaries; they’re just where we decide to draw the line and say, “Okay, this is something I configure, or I don’t.” It’s about where the world is going to consume something, where the problem is, and then right-sizing and figuring out the solution to go and hit it.
Now, “full-stack” means I know JavaScript and Python.
That’s right.
You’ve had a lot of success doing both of those things, and at Databricks. Why start a new chip company now?
It is kind of crazy. It’s one of these things. First off, it’s not a chip company per se. Most of what we’re doing, at the beginning, is theory and really looking at first principles of how learning works in a physical system.
The reason to go back and do this is purely out of passion. I think we can change how a computer is built. We’ve been building largely the same kind of computer for 80 years. We went digital back in the 1940s, and in undergrad in the 1990s, when I learned about the thermodynamics of the brain—the brain’s 20 watts of energy and the kind of computations that can happen inside the brain and neural systems—I was just blown away. I’m still blown away by it.
I think we haven’t really scratched the surface of how we can get close to that. Biology is exquisitely efficient. It’s very fast, and it right-sizes itself to the application at hand. When you’re chilling out, you don’t use much energy, but you’re still aware of other threats and things like this. Then, once a threat happens, everything turns on. It’s very dynamic, and we really haven’t built systems like this.
I’ve been in the industry long enough to know that we have to have an incentive to build things. You can’t just say, “Hey, I want to build this cool thing,” and therefore go build it. Maybe in academia you can do that, but in the real world, I can’t. Now it’s exciting because those concepts are super relevant. We’re at a point in time where computing is bound by energy at the global level, which just was never true in all of humanity.
For those of us who aren’t experts, can you describe the difference between digital and analog computing systems? Why do you think the architecture has evolved the way it has, becoming more digitally focused over the decades, as you said?
Very simply, digital computers implement numerics with some sort of estimation. In a digital computer, a number is represented by a fixed number of bits, and that has some precision error and things like this. It’s just a way we implement the system. If you make it enough bits—64 bits, for example—you can largely say that maybe the error is small and you don’t have to think about it.
The digital computer is capable of simulating anything that you can express as numbers and arithmetic, so it became a very general machine. I can literally simulate any physical process. All of physics—we try to do computational physics, right? I have an equation, and I can then write numeric solvers that deal with those imprecisions in the number of bits.
This became computer science, the entire field now. We went in that direction very early on because we couldn’t scale up computation. It’s actually an interesting comparison if you look at that time. If you look at the papers and things, they actually looked very similar to today in terms of scaling up GPUs.
Analog computers were actually some of the first computers, and they worked really well. They were very efficient, but they couldn’t be scaled up because of manufacturing variability. Someone said, “Okay, you know what? I can make a vacuum tube behave as a high or low very reliably. I can’t characterize the in-between very well, but I can say it’s high or low.” That was where we went to digital abstraction, and then we could scale up.
ENIAC, which was built in 1945, had 18,000 vacuum tubes.
Wow.
So, 18,000 is kind of similar to how many GPUs people use now for large-scale training, right? And so, in digital computers we have transistors. Just to make it concrete, what kind of substrates are you talking about for analog computers?
Analog computers can do lots of different things. Wind tunnels are a great example of an analog computer, in a sense. I have a race car on a track or an airplane, and I want to understand how the wind moves around it. In theory, you can solve those things computationally. The problem is you’re always going to be off. It’s very hard to know what the real system is going to look like, and doing things with computational fluid dynamics accurately is pretty hard.
So, people still build wind tunnels. That’s actually modeling; that’s an analog computer. I think we still have lots of reasons to build these analog-type computers.
In the situation we’re talking about, we can actually build circuits in silicon to recapitulate the behaviors of neural networks. What we’re doing today is more specified than what we were doing 80 years ago, in a sense. Back then, we were trying to automate generic calculations, which were used to calculate artillery trajectories, finances, and maybe some physics problems, like going into space. Those require determinism and specificity around the numbers and computations.
Intelligence is a different beast. You can build it out of numbers, but is it naturally built out of numbers? I don’t know. A neural network is actually a stochastic machine.
And so why are we using a substrate that is highly precise and deterministic for something that's actually stochastic and distributed in nature? We believe we can find the right isomorphism in electrical circuits that can subserve intelligence.
That's a pretty wild idea, isn't it? Maybe unpack it one level deeper, because I totally agree with you. Computers for decades have been the complement to human intelligence, right? My brain isn't really great at computing an orbital trajectory.
That's right.
And I don't want to burn up on reentry. A computer can help us with this incredible degree of precision. We're now going in the opposite direction, right? We're actually trying to encode more fuzziness into computer systems. Go a little bit deeper on this idea of analog, and why intelligence is a good fit for analog systems.
The best examples we have of intelligent systems in nature are brains. It's often been said that human brains run on 20 watts of energy. That is true, but if you look at an animal brain, they're generally extremely efficient. A squirrel or a cat is using something like a tenth of a watt, so there's something there that we're still missing.
Not to say that we understand all of it, but part of what I think we're missing is that we have lots of abstractions in a computer that are quite lossy. In a brain, the neural network dynamics are implemented physically.
So there is no abstraction. Intelligence is the physics. They're one and the same. There's no operating system, API, or anything like that. A visual stimulus, for instance, directly activates an actual neural network and produces some semantic response.
Exactly. Those things are mediated by chemical diffusion and the physical properties of the neuron—the physics itself. So I think it's absolutely possible to build something that's much more efficient by using physics in an analogous way. That is 100% true. Whether we can do it and build a product out of it is really the question we're asking here at Unconventional.
Is part of the idea that now is the right time because AI is both a huge and a unique workload?
Yeah, absolutely. It's interesting. Just to give you some statistics, the US has about 50% of the world's data center capacity, and today we put about 4% of the US energy grid into those data centers. In 2025, we started to see news articles about brownouts in the Southwest during the summer. Imagine what happens when this goes to 8% or 10% of the energy grid. It's not going to be a good place to be.
Can we build more power? Absolutely—we should. Building power generation is very hard and expensive, and it's infrastructure. It takes time. You can only bring online so many kilowatts or gigawatts per year—something on the order of 4 gigawatts per year. By some estimates, we need 400 gigawatts of additional capacity over the next 10 years to power the demand for AI.
Wow.
So we have a huge shortfall, and we really need to rethink this. The 15-year-old sci-fi nerd in me says, “Wow, we're mobilizing species-scale resources to invent the future.” We are. Then there's the practical side: even if we add 400 gigawatts of production capacity, our 1970s-era transmission grid is probably going to melt under the load. There are very serious infrastructure hurdles to this.
It's hard to get a lot of humans to act together. That's just the reality, and that's what has to happen to solve these problems. What trade-offs do you think this entails—the path you're pursuing versus the mainstream digital path now?
I don't actually see it as digital or analog. It doesn't work like that. I think there are certain types of workloads that are amenable to these analog approaches, especially workloads that can be expressed as dynamical systems.
Dynamics means time. They have time associated with them. In the real world, every physical process has time. In the computing world—in the numerical computing world—we actually don't have that concept. You simulate time with numbers.
Actually, simulating time is very useful for certain problems. I think we should still build those things, and we should still have those capabilities for the problems we need to solve that way. But for problems where, as you said, things are a bit fuzzier—trying to retrieve and summarize across multiple inputs—that's what brains do really well. They can take in tons of data and formulate a model of how those things interact. Sometimes those models can be extremely accurate. Look at an athlete.
Alex Honnold, who climbed El Capitan. Just think about the precision that's required. It still scares me every time I
It's insane, right? If he slips—if he's off by a millimeter in some places—he dies. That's true for every top-level athlete and anyone who's at the Olympics.
Steph Curry—the story is that he set up a special tracking system so he could make sure the ball was hitting the middle of the rim, not just going through.
The level of precision these guys achieve with a noisy neural network is actually quite high. Neural systems can do a lot of precision under certain circumstances. What's interesting about these situations is that Steph Curry, when he shoots a ball, is never going to shoot it under ideal circumstances in a game. It's always a unique input, and there are a lot of different input variables coming at you: where the other players are, precisely where you're standing, whether your shoes are different, whether the surface is a little different, whether the ball is tackier, or whether your hands are sweaty.
There are so many inputs, and we put them all together and integrate them while still producing very accurate behavior. Brains are exceptionally good at this. That's a set of problems that is very useful to solve, and now we're approaching those problems. But it doesn't mean we don't still use computational substrates to do actual computation. This is an intelligent substrate.
What types of AI models or data modalities do you expect your hardware will be well suited for?
We're obviously starting with the state of the art today: transformers and diffusion models. They work and do really good stuff, so we shouldn't throw that out. Diffusion models, flow models, and energy-based models are actually pretty interesting because they inherently have dynamics as part of them. They're literally written as ordinary differential equations.
That makes it possible to ask: Can I map those dynamics onto the dynamics of a physical system in some way that's either fixed or has some principled way of evolving? Then can I use that physical system to implement the model and do it very efficiently with physics? That's the nature of what we're doing. We will be releasing some open-source work and other things around this to let people play around.
Transformers are a big innovation because they made the constructs of a GPU work extremely well. That doesn't mean it's wrong, but I don't think there's anything natural—there's no natural law—about the parameters of a transformer. A transformer's parameters are a function of the nonlinearities and the way the whole thing is set up with attention. There will be some kind of mapping between transformer parameter spaces and these other parameter spaces. Transformers, I think, have used lots of parameters to accomplish what they do.
I have to ask: Since you mentioned energy-based models, and Yann LeCun has been writing quite a lot about this, do you think pursuing these sorts of paths gets us closer to AGI—whatever AGI means?
Honestly, I do. The reason I feel that way—and again, this is hand-wavy; I'm going to be really honest—I don't—
That's why I'm putting quotes around it. I think the discussion is necessarily hand-wavy.
It's got to be, because we just don't know. My intuition says that anything where the basis is dynamic, with time and causality as part of it, will be a better basis than something that isn't.
We've largely tried to remove that. A lot of times you can write math down that's reversible in time and things like that, but the physical world tends not to be, at least the way we perceive it.
So can we build out of elements of the physical world that do have time evolution? I think that's the right basis to build something that understands causation. I do think we'll have something that's better and will give us something closer to what we really think is intelligence.
Yes, we have intelligence in these machines. I don't think they're anywhere close to AGI because they still make stupid errors. They're very useful tools, but they're not like working with a person, right? I think most people at that—
That's actually really interesting.
So the sort of thing that's missing in AI behavior—which I think a lot of us see, that there's something missing but can't quite put a name to it—it sounds like you're arguing that part of that is a real sense of causality. And that training, and a more dynamic sort of regime, may impart this kind of apparent understanding of causality better than what we have now.
Yeah. Again, hand-wavy, but yes. Look, you have kids—little kids—and you see them. Children kind of innately understand causality in some ways: this happened, then that happened. You can say it's reinforcement learning or whatever; that's some part of it, but there's something innate that we understand about causality. In fact, that's how we move our limbs and all of that. I know that if I send a certain command to my arm, it'll do something. So I think there's something innate about the way our brains are wired, built out of primitives that do understand causation.
Put Unconventional AI in the context of the broader industry for me. Nvidia, TSMC, and Google—are these potential allies for Unconventional AI? Are they competitors? How do you think about it?
Yeah, a couple of things that we set out to do when we were starting this company were to see if we could find a paradigm that's analogous to intelligence within 5 years. At the 5-year mark, we should be able to build something that's scalable from a manufacturing standpoint. You can think about building a computer out of many different things, but if it's not scalable from a manufacturing standpoint, we can't intercept this global energy problem. We need somebody to say, “Okay, go build 10 million of these things,” right?
I think TSMC is absolutely going to be a partner going forward. I met with them recently, and we want to work closely with them to make sure we get what we need, get fast turnaround times to prototype, and all of that. Google, Nvidia, Microsoft—all these guys are at the forefront of where the application space is. Obviously, Google has everything internally, and I think they're working on lower-risk but continual improvements for their hardware with TPUs.
With TPUs, you mean?
With TPUs? Yeah. From what I can see, just publicly, it makes total sense. They have a business to run, and they're trying to make their margins better. How can I do that with all the tools I have in front of me?
Nvidia, of course, has built the platform that everyone programs on today. So, are we going to be at odds with Nvidia going forward? I don't know. We'll see what the world looks like. We're trying to build a better substrate than matrix multiply. There could be a world where we collaborate on such solutions, and we're open to all of these things.
Where do you personally get the motivation to get up in the morning and build this company? You've had a lot of success in your career, in startups. What's exciting about this to you?
I don't know. It's a weird thing. If you haven't worked in hardware, it's hard. I've been fortunate to work in hardware and software, and I love writing a bunch of software, hitting compile, and seeing it work. That's a good dopamine hit. But when you work on a piece of hardware and turn that thing on, that's a big dopamine hit. It's like celebration—jumping up in the air, high-fiving. It's a different thing, and you sort of live for these moments.
When I was at Intel, I was one of the only execs who would go to the lab when the first chip would come back. I'm like, “I want to see it turn on, see what happens.” Sometimes you turn it on and you see a little puff of smoke come—
That's not good.
But you want to be there; you want to be part of the moment. I think that's part of it. For me personally, we have this opportunity now that we can really change the world of computing and make AI ubiquitous. I'm the opposite of an AI doomer. I think AI is the next evolution of humanity. I think it takes us to a new level, allows us to collaborate, understand each other, and understand the world in much deeper ways.
Totally agree.
Every technology has negatives, but the positives to me so far outweigh them. The only way we're going to get to ubiquity is that we have to change the computer. The current paradigm, as good as it is and as far as it's taken us, is not going to take us to that level.
I think that's such a great way to say it. AI actually can help us understand each other better, help us understand ourselves better, and understand the natural world better.
I don't think it's at all what some of the doomers think of as replacing human experience.
That's a short-term thing. There will be bumps along the way, right? Technology does that. That's what happens when you've seen too many sci-fi movies.
That's right. But look at Star Trek.
Yeah, yeah, yeah. Totally. Totally. It's great.
This is a really big swing, right? This is a very ambitious company. What gives you confidence that it's going to work, or has a reasonable shot of working?
There's a number of data points. Of course, like I said, the brain's an existence proof. But there's also 40-plus years of academic research showing a lot of promise here. People have built different devices, albeit not with the latest technology or with professional engineering teams, but they have built proofs of concept that actually show some of these things work.
We've also, from a theory standpoint—both from neuroscience and from pure dynamical-systems and math theory—started to understand how these systems can work. So I think we now have pieces at different parts of the stack that show, “Hey, if I can combine these things the right way, I can build this.” That's what great engineering is all about: exploiting this thing that someone else built for something else, exploiting that thing, and then—
Engineers are kind of the opposite of theorists: “Well, all right, that thing doesn't quite fit. Sand it down and make it right.” So we've got to do a little bit of that right now, and then we can build something and put it all together.
Yeah.
That's awesome. Has anyone called you crazy yet for doing this?
Oh, yeah. Plenty of people at this point.
Is it everybody?
Well, I'm used to this at this point. My family has called me crazy. I was called crazy going back to grad school years ago, when I had a very good career in tech. So it's fine. I think you need crazy people to go out and explore. If you think about humanity coming out of Africa, the crazy people who went out—
We would be lost without crazy.
You need some crazy in there. So it's okay. I'm fine with that.
What kind of people are you looking to bring onto the team? It's a very ambitious goal. Who should be interested in joining you?
Yeah, I think some of the traditional-ish—when I say traditional, over the last 5 years this field of AI systems has evolved—people who are really good at taking algorithms and mapping them very effectively to physical substrates. Those folks who understand energy-based models, flow models, gradient descent, and different ways—this kind of thing is what we need there.
We need theorists who can think about different ways of building coupled systems, how I can characterize the richness of dynamical systems, and relating that to neural networks. So there is a theory aspect of this. Then there are folks who are kind of at the system architecture level: “All right, here's what the theory says. This is what I can really build. How do I bridge that gap?” And then there are the people actually physically building this stuff—analog circuit people, and digital circuit people, too. We're going to have a mixed-signal team here. So that's the whole stack.
The stack is hard because these are all things that no one's really pushed to that level. When we build this chip, our first prototype, it's going to be probably one of the larger, maybe the largest analog chip people have ever built, which is kind of weird. The first time you do something, things don't usually work the way you think they—
So you can get in on that Cerebras–Jensen game where they were each pulling the biggest possible wafer out of an oven.
Something like that. Yeah, exactly right.
Put a few vacuum tubes on top for effect.
Yeah, we could. I need blinking lights.
Yeah, exactly.
We're not going to have cool heat sinks. It's going to be super—it's going to be cold. You don't need big heat sinks, you know? So I hope they make something that looks interesting here.
This is a funny time for top AI people, right? You have the option, if you want to start a company, of a lot of venture capitalists who would probably fund you. If you want to get a cushy job at a big company, you can get a very cushy job and do some interesting things.
Yeah.
Or, you know, people can join a startup like Unconventional AI that has a lot of the nice aspects people look for in AI careers and is taking super-big swings.
I’m just curious. You’ve been on all sides of this. Do you have any advice for younger people starting out in their careers, or how do you think about this?
I think you get such a breadth of experience from working at a startup at the beginning of your career that it will pay dividends later on. The reason I can think across the stack is because I did all those things very early in my career. I built hardware, I built software, and I built applications.
In big companies, it’s not anyone’s fault. It’s just the way it is. You get hired to do a thing, and you do that thing over and over again. You get really good at doing that thing, and that’s fine. You need people who are really good at doing specific things. But if you want to be prepared for change in the future, being really good at one thing is probably less valuable than being slightly good at a lot of things.
Yeah, that’s interesting. Is it fair to say Unconventional is sort of a practical research lab? Is that the kind of culture you’re going for?
Absolutely. Yeah. I mean, the first few years, it really is open-ended. I don’t want to close doors. I’m really specific about this. I always try to bring the conversation back when people say, “Oh, that’s going to be hard to manufacture.” Stop. Don’t think about that. Will it work? First, come up with existence proofs. Then we go back and try to engineer it, with all the trade-offs therein.
But if you make those trade-offs up front, you don’t go into a good place. So yes, we are really thinking wide open, but with an eye on the future, who we are building a product—
And to your point, it takes not only people with diverse skill sets, but people with high agency to try new things, learn new things, and integrate across the stack.
Yeah. I think what I’ve done really well across the companies I’ve built has been going after hard problems, which lends itself to smart people wanting to come in and try to solve them. They see a challenge—it’s like climbing Mount Everest—but then giving them agency.
I look at it like, what decisions can I make as a leader to increase the agency of the organization overall? Me making a top-down-style decision may be globally better for the company in the short term, but I think long term we’ll do better if more people have agency and can try more things out.
Personally, I like to find ways to get out of the way when I see people who are very passionate about trying something. It’s like, “Okay, you really want to do this. That makes sense. Go for it.” Then you own it. You own both the good and the bad, right? Agency to me is like, you’ve got to be able to say, “Okay, I screwed up. No, this wasn’t right.” That’s okay, too. But give people the room to do that.
Anything else you want to say before we wrap up?
I think this is an opportunity to do something that will be felt for generations. To me, that’s what gets me up in the morning. You can go work on a product and make a tweak, and people will use it. That’s great, but in 5 years, many times people forget those things.
But if we are successful here, the world will not forget this for a very long time. This will be written in history books. I feel like those opportunities are rare.