Speaker 1
This is the greatest force for reducing the technology divide the world has ever known. It will have an impact on the GDP of every country in the double digits in the coming years. Nobody’s going to do this for you. You’ve got to do it yourself.
It’s up to organizations, enterprises, and countries to build what they need. The stakes at play are basically the equivalent of modern digital colonization. AI isn’t just computing infrastructure; it’s also cultural infrastructure.
Today, we’re talking about sovereign AI, all things national infrastructure, and open source. Let’s start with the first question I usually get from nation-state leaders: Is AI actually a general-purpose technology?
In the history of humanity, we’ve had maybe a handful of these—22 24. Economists call these specific technologies that accelerate economic progress broadly across society: electricity, the printing press. The question everybody’s asking right now is, Is that the right way to think about AI, or is AI just another important but ultimately narrow technology?
Arthur Mensch
It is a general-purpose technology because it basically revisits entirely the way we are building software and the way we are using machines. In the same way that the internet was a general-purpose technology, AI is a general-purpose technology.
It allows us to build agents that are doing things on your behalf, and in that respect it can be used in any vertical of the industry. It can be used for services and public services. It can be used to change the lives of citizens, for agriculture, and obviously for defense purposes. It covers everything that a state needs to worry about.
In that respect, it’s very natural that any state makes it a priority and finds a dedicated national AI strategy.
Jensen Huang
By the way, everything Arthur said is 100% correct. It is also exactly the reason why everybody’s given up, and it’s precisely wrong. The reason for that is this: If it’s a general-purpose technology and one company can build the ultimate general-purpose technology, then why should anybody else do it?
That is the flaw, but that’s also the mind trick to convince everyone that this is a technology that intelligence is only something that a few people ought to go build. Everybody ought to sit back and wait for it.
I would advise that everybody engage AI, and that this is not for the privileged. Intelligence is for everyone. It is not just a few companies in the world who should build it. Everybody should build it.
Nobody’s going to care more about Swedish culture, the Swedish language, the Swedish people, and the Swedish ecosystem than Sweden. Nobody’s going to care about the ecosystem of Saudi Arabia more than Saudi Arabia, and nobody’s going to care about Israel more than Israel.
Despite the fact that the technology is general-purpose—and absolutely true, how could intelligence not be general-purpose?—it is also hyper-specialized. The reason for that is because, let’s face it, I don’t think I’m waiting around for a general-purpose chatbot to be an expert in a particular area of disease. I still think that I would prefer to have somebody who is hyper-specialized in that field to fine-tune, train, and post-train an AI model that’s going to be specialized in that.
Arthur Mensch
It’s the general-purpose technology, the same way your programming language is the general-purpose technology. In addition to that, it’s also a culture-carrying technology.
I think what that means is that there is infrastructure. There are chips that obviously not every country is going to build. There are general-purpose models, like base models that are a compression of the web, that are eventually going to be open source and can serve as the right basis for constructing specialized systems.
Beyond that, I think it’s up to organizations, enterprises, and countries to build what they need. The way to make it work is to take a general-purpose model—an open-source model, for instance—and get the knowledge you have specifically, or ask your citizens or employees to distill their knowledge into the systems and agents that are going to be working on your behalf.
As you do that progressively, those agents become more accurate and better at following the instructions and specifications that the country or an enterprise may have when it’s building AI systems.
You need vertical experts, cultural experts, or people with a certain national agenda to partner with technological companies that can expose the open-source infrastructure in a way that is easy to use and easy to specialize. I think that’s where the frontier lies. It’s a very horizontal technology, but to make anything useful out of it, you need a partnership between the horizontal providers and the vertical experts.
Speaker 1
Unlike previous general-purpose technology waves in history, like electricity or the printing press, how is this one different? If I’m a nation-state leader and I’m trying to understand the right framework for thinking about AI in my country, should I think about it like digital labor? Should I think about it as something else?
Arthur Mensch
I think it’s similar to electricity in the sense that it will have an impact on the GDP of every country in the double digits in the coming years. That means that from an economical point of view, every nation needs to worry about it.
If they don’t manage to set up infrastructure and their own sovereign capacities in the right place, that means this is money that might flow back to other countries. That’s changing the economic equilibrium across the world. In that sense, it’s not very different from electricity, which, 100 years ago, if you weren’t building electricity factories, meant you were preparing yourself to buy it from your neighbors.
At the end of the day, that isn’t great because it creates dependencies. In that sense, it is similar. What is fairly different, I think, is two things.
First of all, it’s an amorphous technology. If you want to create digital labor with it, you need to shape it. You need to have infrastructure, talent, and software, and the talent needs to be created locally. I think this is quite important.
In contrast with electricity, this is a content-producing technology. You have agents that are producing content—producing text, images, and voice—and interacting with people. When you’re producing content and interacting with society, you become a social construct. In that respect, social constructs carry the values of either an enterprise or a country.
If you want those values not to disappear and not to depend on a central provider, you need to engage with it more profoundly than you would engage with electricity, for instance. Would you agree with that, Jensen?
Jensen Huang
A couple of ways to think about it: Your country’s digital intelligence is not likely something you would want to outsource to a third party without some consideration. Your digital intelligence is just now a new infrastructure for you—your telecommunications, your health care, your education, your highways, your electricity. Now there’s a new layer. This new layer is your digital intelligence.
It’s your responsibility to decide how you want this digital intelligence to evolve and whether you want to outsource it so that you never have to worry about intelligence again, or whether this is something that you feel you want to engage with, maybe even control and shape into a national infrastructure.
Of course, it has all the things that Arthur said: AI factories, infrastructure, and so on. There’s another way you could think about it: It’s your digital workforce. This is a new layer, and you’ve got to decide whether the digital workforce of your country or your company is something that you decide to outsource and hope it evolves the way that you would like it to, or whether it’s something you want to engage with, maybe even decide to control, nurture, and make better.
We hire general-purpose employees all the time. We hire them out of school. Some of them are more general-purpose than others, and some of them are more intelligent than others. But once they become our employees, we decide to onboard them, train them, guardrail them, evaluate them, continuously improve them, and make the investment necessary to turn general-purpose intelligence into superintelligence that we can benefit from.
I think that second layer—thinking about it as a digital workforce—is important. In both cases, it contributes to the national economy. In both cases, it contributes to social advancement. In both cases, it contributes to the culture. I think that in both cases, a country needs to play a very active role in it.
So I think it’s back to your original question about sovereign AI and how to think about it. Yes, it is definitely a general-purpose technology, but you have to decide how to shape it.
Your country’s digital data belongs to you. Your national library and your history, insofar as you want to digitize them, belong to you. You could make them available to everybody in the world. You could also make them available to companies, researchers, and institutions in your own country.
These are all amorphous things. They’re very soft ideas, but they do belong to you. They belong to you in the sense that this is where you came from. They belong to you in the sense that you can decide how to put them to use for the benefit of your people. They belong to you in the sense that it’s your responsibility to shape their future.
Sovereign AI is your responsibility.
Speaker 1
There are several other types of assets that nation-states fund and protect: the military, your electricity grid. If I’ve understood the criticality of AI infrastructure and sovereign AI, do I now have to take control of every part of the stack? Jensen mentioned the digital workforce. How should I think about that?
Arthur Mensch
I think it’s a very good analogy that you need an onboarding platform for your AI workforce. That means you need to be able to customize the models and pour the knowledge sitting in your national libraries into the models, so that suddenly they speak your language better.
You need to be able to get your systems to know about your laws, so that suddenly the guardrails that are set when you’re deploying AI software are compliant with everything that you’ve done before.
That onboarding platform requires you to customize, guide, and evaluate the systems. Then, when you notice that certain things need to be improved, you need to fix and debug them. That’s the platform that we’re building.
Once the custom systems are made, it’s important to be able to maintain them yourselves. That means being able to deploy them on your own infrastructure and being able to ask your technological partners to potentially disappear from the loop.
Jensen Huang
Your IT department is going to become the HR department of your digital workforce. They’re going to use these tools that Arthur describes to onboard AIs, fine-tune AIs, guardrail them, evaluate them, and continuously improve them.
That flywheel will be managed by the new, modern version of the IT department. We’ll have a biological workforce and a digital workforce, and it’s fantastic.
Nobody’s going to do this for you. You’ve got to do it yourself. That’s why, even though we have so many technology companies in the world, every company still has its own IT department. I’ve got my own IT department. I’m not going to outsource it to somebody else.
In the future, IT departments will be even more important to me because they’ll be helping us manage these digital workforces. You’re going to do this in every country and every company within those countries.
The space that Arthur is describing—taking this general-purpose technology and fine-tuning it into domain experts, whether they’re national, industrial, corporate, or functional experts—is the future. It’s the giant future space of AI.
Speaker 1
You both said something that I want to make sure I’m understanding correctly. You called it a soft concept, like your culture, and you said there are a bunch of norms that the training data has that you customize the models on.
It sounds like you’re saying that, unlike compute, storage, and networking—which have rules, algorithms, and laws that are very specific—you’re talking about norms. That exactly means it’s soft versus rules, which are harder.
Arthur Mensch
There are different things that you want to incorporate into your AI systems. There are some elements of style and knowledge that you’re not going to enforce through strict rules, but that you can enforce through continuous training of models, for instance.
You take the national library, and you take preferences and distill them into the models themselves. Then you have a set of rules and a set of policies if you’re in a company, and those are strict.
Usually, the way you build it is that you connect the models to the strict rules and make sure that every time the model answers, you verify that the rules are respected.
There are really 2 things. On one side, you’re pouring and compressing knowledge in a soft way into the models. On the other side, you’re making sure that you have a certain number of policies and rules that are strictly enforced and have 100% accuracy.
Jensen Huang
On one side, this is soft. This is preference and culture. Somebody’s preference is multidimensional. You know what you prefer; many times it’s implicit.
There are so many features that define my preference. It takes AI to be able to precisely comply with the description that Arthur was giving just now. Could you imagine if a human had to write this in Python or C++? You’d have to describe every one of these things and capture every one of these things: Based on this, I prefer that, but if you did that, I prefer that other thing.
The number of rules would be insane, which is the reason why AI has the ability to codify all of this. It’s a new programming model that can deal with the ambiguity of life.
Speaker 1
It sounds like you’re saying AI isn’t just computing infrastructure; it’s also cultural infrastructure. Is that right?
Arthur Mensch
Yes. It’s about making sure that your cultural infrastructure and the human expertise in your company or country makes it to the AI systems.
Culture, reflection, and values—we were just talking about how each one of these AI models and AI services responds differently to the type of questions you ask, because they codify the values of their service or their company into each one of their services.
Jensen Huang
Could you imagine this amplified at an international scale? For me, this is an inherent limitation of centralized AI models, where you’re thinking that you can encode some universal values and some universal expertise into a general-purpose model.
At some point, you need to take the general-purpose model and ask a specific population of employees or citizens what their preferences and expectations are. You need to make sure that you’re specializing the model in a soft way and in a hard way, through rules, culture, and preferences.
That part is not something that you can outsource as a country. It’s not something that you can outsource as an enterprise. You need to own it.
Speaker 1
Then is it an exaggeration to say that, if it is cultural infrastructure and I don’t own sovereignty over it, the stakes at play are basically the equivalent of modern digital colonization?
If you’re saying that you’ve got to think about AI as almost like your digital workforce, and another country or somebody who isn’t my sovereign nation can decide what my workforce can and can’t do, that’s a problem.
Arthur Mensch
Some of it is universal. It is possible for certain companies to serve nations, societies, and companies around the world because it’s basically universal. But it can’t be the only social fabric. It can’t be the only layer. It can’t be the only digital-intelligence layer. It has to be augmented by something regional.
Jensen Huang
I think McDonald’s is pretty good everywhere. Kentucky Fried Chicken is pretty good everywhere. Starbucks is fine everywhere. But you still want the local style and local taste that augment that.
You want local cafés, the mom-and-pop restaurants, because they define the culture. They define society. They define us.
I think it’s terrific that you have Walmart everywhere, that you can count on it everywhere. I think that’s fine, but you need to have local taste, local style, local preference, local excellence, and local services.
Let me swing in another way. In the context of our digital workforce in the future, we’ll have some digital workers that are generic. They’re just really good at doing basic research or something basic. They’re good at a college or graduate level, and they’re useful for every company. It’s unnecessary for me to create something new.
I think Excel is pretty good. Microsoft Office is universally excellent. I’m perfectly fine with that. It’s a good reference architecture and a good base.
Then there are industry-specific tools and industry-specific expertise that are really important. We use Synopsys and Cadence. Arthur doesn’t have to, because they’re specific to our industry, not his.
We probably both use Excel. We’ve both probably used PDFs, and we’ve both used browsers. There are universal things that we can all take advantage of, and there will be universal digital workers that we can take advantage of.
Then there will be industry-specific and company-specific systems. Inside our company, we have some special skills that are very important to us and that define us. They’re highly biased, if you will, toward doing the things that I need them to do—highly guardrailed to do very specific work and highly biased toward the needs and specialties of our company.
We become superhuman in those areas. Your digital workforce is going to be the same, and AI is going to be the same. There will be some systems that you take off the shelf. The new search will likely be some AI. The new research will probably be some AI.
Then there will be industrial versions of AI that we might get from Cadence and others. We’ll have to grow our own using Arthur’s tools. We’ll have to fine-tune them, onboard them, and make them incredible.
Arthur Mensch
I very much agree with this vision of having a general-purpose model, then some layer of specialization for industries, and then an extra layer of specialization for companies and countries. You will have a tree of AI systems that are more and more specialized.
To give a concrete example of what we recently did, we released a model called Mistral Small in January. It’s a general-purpose model. It speaks all of the languages and knows about most things.
Then we took it and started a new family of specialized models that were specialized in languages. We added more language data in Arabic and in Indian languages, and we retrained the model. We distilled this extra knowledge that the initial model hadn’t seen.
In doing that, we made it much better and more idiomatic when it speaks Arabic and when it speaks languages from the Indian subcontinent. Language is probably the first thing you can do when you’re specializing a model.
For a given size of model, you can get a model that is much better if you choose to specialize it in a language. Today, our 24-billion-parameter model, called Mistral Saba, is a model tuned on Arabic and is outperforming every other language model that is 5 times larger. The reason for that is that we did the specialization.
That’s the first layer. If you think of a second layer, you can think of verticals. If you want to build a model that is not only good at Arabic but also good at handling legal cases in Saudi Arabia, you need to specialize it again.
There’s extra work that needs to be done in partnership with companies to make sure that your system is not only good at speaking a certain language, but also good at understanding the legal work that is done in that language.
It’s true for any combination of vertical and language that you can think of. If you want to have a medical diagnosis assistant in French, you need to be good at French, but you also need to understand how to speak the French language of physicians.
Those 2 things are very hard to do as a general-purpose model provider.
Speaker 1
If this is true, and what you’re describing is real—that I need the capabilities to customize this AI layer on my local norms and local data, which is fairly sophisticated from a technical-capability perspective—how would you advise a big nation to think about the stack we’re talking about?
There are the chips, the compute, the data center, the models that sit on top, the applications, and then ultimately what you were describing as the AI nurse or the AI doctor. How would you advise someone who’s a smaller nation differently?
Arthur Mensch
I would say you need to own—or, I guess, buy and set up—the horizontal part of the stack. You need the infrastructure, the inference primitives, the customization primitives, and observability. You need the ability to connect agents to models, to connect models to sources of information, and to connect them to real-time information.
Those are primitives that are fairly well factorized across the different countries and enterprises. Once you have that—these are things that can be bought—you can start working and building from those primitives according to your values, your expertise, and your local talent.
The question is, Where is the frontier between what is horizontal and what is vertical? If you’re a small enterprise or a small country, you should probably buy what is horizontal. What is vertical and specific to you is definitely something that you need to build.
Jensen Huang
You have to get it in your head that it’s not as hard as you think it is. First of all, the technology is getting better, so it’s easier. Could you imagine doing this 5 years ago? It was impossible. Could you imagine doing this 5 years from now? It’ll be trivial.
We’re somewhere in the middle. The only question is, Do you have to do it? The truth of the matter is that I hate onboarding employees, and the reason for that is because it takes a lot of work.
Once you set up an HR organization, a leadership and mentoring organization, and processes, your ability to onboard employees becomes easier and systematically more enjoyable for everybody involved. In the very beginning, it’s hard. Setup is always hard. Setup is always hard. This is no different.
The only question is, Do you need to do it? If you want to be part of the future—and this is the most consequential technology of all time, not just of our time—digital intelligence: How much more valuable can it be? How much more important can it be?
If you come to the conclusion that this is important to you, then you have to engage with it as soon as you can. Learn it and learn along the way. Just know that it’s getting easier and easier all the time.
The fact of the matter is, if we tried to do agentic systems even 3 years ago, it was incredibly hard. Agentic systems are a lot easier today, and all of the tools necessary for curating data sets, onboarding digital employees, evaluating them, and guardrailing all of those digital employees are getting better all the time.
The other thing about technology is that when it becomes faster, it’s easier. I had the benefit of seeing computers from their earliest days, and the performance of the computers was so frustratingly slow. Everything you did was hard.
These days, the types of things we do are magical because they’re fast. Whether it’s motivated by your institutional need to engage with the most consequential technology of all time, or by the fact that it’s getting better all the time and therefore isn’t that hard, I think the number of excuses is running out.
Speaker 1
Let’s talk about that for a second, because change is hard. I have an endless list of things I’m facing. If I’m a nation-state leader, I’m dealing with increasing geopolitical risk. I don’t know who my allies are. Elections are coming. There are any number of things I have to deal with.
But now, let’s say I understand that this is important. You both spend so much time talking to nation-state leaders who are thinking about the risks of adopting AI too fast. You’re right that the zeitgeist has shifted based on the Paris AI Action Summit. It seems like there’s a tone of optimism now, more than there was a tone of pessimism a year ago.
What are the most common questions you get from nation-state leaders when they’re asking you about risks and how to think about them?
Arthur Mensch
I’ve heard several questions. One of the risks is seeing your population start getting afraid of the technology for fear of it replacing them. That is something that can actually be prevented if we collectively make sure that everybody gets access to the technology and is trained in using it.
The skilling of the various citizens in the population is extremely important. We need to state that AI is an opportunity for them to work better and show the purpose of it through applications and things they can actually install on their smartphones, as well as through public services.
We’re working, for instance, with the French unemployment system to connect job opportunities to unemployed people through AI agents that are being actioned by human operators within the agency. That’s an opportunity for people to find a job more effectively.
That’s part of what can make sure that the population understands the opportunity and the fact that AI is really just a new change for them to adopt, the same way they had to adopt personal computers in the 1990s and the internet in the 2000s.
The common aspect with these changes is that you need people to embrace the technology. I think the biggest problem that nation-states may have is seeing AI increase the digital divide, which is already relatively big. But if you work together and do it in the right way, we can make sure that AI is actually reducing the digital divide.
Jensen Huang
AI is a new way to program a computer. By typing in some words, you can make the computer do something, just like we did in the past. We used to type words and make computers do something, and now you talk to it.
You can interact with it in a whole lot of ways. You can make the computer do things for you much more easily today than before. The number of people who can prompt ChatGPT and do productive things, just from a human-potential perspective, is vastly greater than the number of people who can program in C++.
Therefore, we have closed the technology divide. It’s probably the greatest equalizer we’ve seen. It is, by definition, the greatest equalizer of technologies of all time.
You still need citizens to know about it. The fact is, there are more people who program computers using ChatGPT today than there are people who program computers using C++. In 4 short years—in fact, 3 years—that is a straight-up fact.
This is the greatest force for reducing the technology divide the world has ever known. It’s just a matter of perception. What Arthur is saying is that, through perception, people may say, “I don’t know what I’m talking about, and I don’t know how to talk about it.” But the fact of the matter is, it is not stopping.
The number of people who are actively using ChatGPT today is off the charts. I think it’s terrific. Anybody who’s talking about anything else apparently isn’t working.
People realize the incredible capabilities of AI and how it’s helping them with their work. I use it every single day. I used it this morning. Every single day I use it, and I think deep research is incredible.
The work that Arthur and all of the computer scientists around the world are doing is incredible. People know it. They’re picking it up.
Speaker 1
Let’s talk about open source for a bit, because both of you have talked quite publicly about the importance of open models in the context of sovereign AI.
At DeepMind, you were part of the Chinchilla scaling laws, which were openly published. Your co-founder Guillaume created Llama, and last year NVIDIA and Mistral worked on a jointly trained model called Mistral NeMo. Why are open models such a big part of your focus?
Arthur Mensch
It’s a horizontal technology, and enterprises and states are eventually going to be willing to deploy it on their own infrastructure. Having this openness is important from a sovereignty perspective. That’s the first point.
The second important point is that releasing open-source models is a way to accelerate progress. What we saw during our early careers, when we were doing AI between 2010 and 2020, was an acceleration of progress because every lab was building on top of each other.
That disappeared with the first large language models from OpenAI in particular. Spinning back up that open flywheel—where I contribute something, another lab contributes something else, and then we iterate from that—is the reason why we created Mistral.
I think we did a good job at it. We started to release models, and then Meta started to release models as well. Chinese companies like DeepSeek released stronger models, and everybody benefited from it.
Coming back to Mistral NeMo, one difficulty of creating AI models in an open way is that this is more of a cathedral than a bazaar setting when it comes to open source. You have a large span of work to do to build a model.
What we did with the NVIDIA team was mix the 2 teams together, have them work on the same infrastructure and the same code, have them work through the same problems, and combine their expertise to build the same model.
That was very successful because NVIDIA brought a lot of things we didn’t know. I think we brought things that NVIDIA didn’t know. At the end of the day, we produced something that was, at the time, the best model for its size.
We really believe in these collaborations, and we think we should do them more and at a higher scale—not only with 2 companies, but probably with 3 or 4. That’s the way open source is going to prevail.
Jensen Huang
I completely agree. The benefit of open source, in addition to accelerating and elevating basic science and the basic endeavor of building general models and general capabilities, is that open-source versions also activate a ton of niche markets and niche innovation.
All of a sudden, health care, life sciences, physical sciences, robotics, and transportation were activated as a result of open-source capabilities. The number of industries that were activated is incredible.
Don’t ignore the incredible capabilities of open source, particularly in the fringe—the niche but mission-critical markets where data might be sensitive. It could be mining or energy. Who’s going to create an AI company to mine energy? Energy is really important, but the mining of energy is not that big of a market.
Open source activates every single one of these markets: financial services, health care, defense. You pick your favorites. Anything that is mission-critical and requires you to do your own deployment, potentially on the edge, and anything that requires strong auditing and the ability to do a thorough evaluation can benefit from open source.
You can evaluate a model much better if you have access to the weights than if you only have access to APIs. If you want to build certainty around the fact that your system is going to be 100% accurate, I don’t think you should be using a closed-source model.
Arthur Mensch
You have to connect it into your flywheel. How are you going to connect your local data? You have to connect it to your own local data and your own local experience. The more you use it, the better it becomes.
You can’t do that without open source.
Speaker 1
Let’s say I’m a nation-state leader. I’ve been considering open source, and I’m starting to hear things like, “Open source is a threat to national security. We should not be exporting our models because these open models actually give away a ton of nation-state secrets.”
More importantly, the bad guys can use these open models, too, and so this is a threat to security. Instead, what we should be doing is locking down development among 2 or 3 labs that have the infrastructure to get licenses from the government, do training, and implement the right safety and certification.
I’ve certainly been hearing that a lot. How should I think about that versus what you’re telling me, which is that open is better for mission-critical industries?
Arthur Mensch
Collaboration between labs is going to be critical for humanity’s success. If one state decides to lock things down, the only thing that’s going to happen is that another state will take the leadership.
Cutting yourself off from the open flywheel is just too high a cost for you to maintain competitiveness. If you do that, this is a debate that has occurred in the United States. If there’s some export control over weights, this is not going to stop any country in Europe or Asia from continuing its progress.
They will collaborate to accelerate that progress. I think we just need to embrace the fact that this is a horizontal technology, very similar to programming languages. Programming languages are all open source, so I think AI just needs to be open source in that respect.
We’re glad to see that at the AI Action Summit that occurred last week, this was very much on the agenda: the realization that we could accelerate together by being more open about the way we build the technology.
It’s great to see that open source has a lot of good days ahead of it.
Jensen Huang
It is impossible to control software. If you want to control it, somebody else’s software will emerge and become the standard, just as Arthur mentioned.
The question is, Is open source safer? Open source enables more transparency, more researchers, and more people to scrutinize and work with the technology.
The reason every single company in the world and every cloud service provider is built on open source is because it is the safest technology of all. Give me an example of a public cloud today that’s built on an infrastructure stack that isn’t open source.
You start from open source and customize it. The benefit of open source is the contribution of so many people and the scrutiny. You can’t just put any random stuff into open source. You get laughed off the internet. You’ve got to put good stuff into open source because the scrutiny is intense.
Open source provides all of that great collaboration to accelerate innovation, elevate excellence, ensure transparency, and attract scrutiny. All of that improves safety.
Speaker 1
In a sense, you’re saying it’s partly more secure because, as we’ve seen with open-source databases, storage, networking, and compute, you get mass red-teaming. The whole world can help red-team your technology versus just a small group of researchers inside your company. Is that roughly the right way to think about it?
Arthur Mensch
Exactly. By pulling a lot of organizations together to come up with a technology that they can all use and specialize in their own domains, you’re forcing the technology to be good for every one of them.
That means you’re removing biases and really making sure that the general-purpose models you’re building are as good as possible and don’t have failures. Open source, in that respect, is also a way to reduce the number of failure points.
If, as a company, I decide today to rely fully on a single organization and on its safety principles and red-teaming organization, I’m trusting it a little too much. Whereas if I’m building my technology on open-source models, I’m trusting the world to make sure that the basis on which I’m building is secure.
That’s a reduction of failure points, and that’s obviously something that you need to do as an enterprise or as a country.
Speaker 1
We’re going to transition a little bit now into company building, which is something a lot of people are excited to hear from both of you about.
Let’s start with you, Jensen. You’ve remarked that NVIDIA is the smallest big company in the world. What enabled you to operate that way?
Jensen Huang
Our architecture was designed for several things. It was designed to adapt well in a world of change, either caused by us or affecting us.
Technology changes fast, and if you overcorrect on controllability, then you underserve a system’s ability to become agile and adapt. Our company uses words like “aligned” instead of words like “control.” I don’t know that I’ve ever used the word “control” in talking about the way that the company works.
We care about minimum bureaucracy, and we want to make our processes as lightweight as possible. All of that is so that we can enhance efficiency, enhance agility, and so on.
We avoid words like “division.” When NVIDIA was first started, it was modern to talk about divisions. I hated the word “divide.” Why would you create an organization that’s fundamentally divided?
I hated the words “business units.” Why should anybody exist as one? Why don’t you leverage as much of the company’s resources as possible?
I wanted a system that was organized much more like a computing unit—like a computer—to deliver an output as efficiently as possible. The company’s organization looks a little bit like a computing stack.
What is this mechanism that we’re trying to create, and in what environment are we trying to survive? Is this much more like a peaceful countryside, or is it much more like a concrete jungle?
The type of system you want to create should be consistent with that. The thing that always strikes me as odd is that every company’s org chart looks very similar, but they’re all different things. One is a snake, another is an elephant, another is a cheetah. Everybody is supposed to be somewhat different in that forest, but somehow they all have the same exact structure and organization. That doesn’t make sense to me.
Arthur Mensch
I agree that it feels like companies have personalities, despite the fact that they’re sometimes organized similarly. Obviously, we have a lot of things to learn. The company isn’t even 2 years old.
One challenge we have with Mistral, and I think our competitors have the same challenge, is that this is one of the first times that a software company is actually a deep-tech company driven by science.
Science doesn’t have the same timescales as software. You need to operate on a monthly basis, and sometimes you don’t know exactly when something will be ready. On the other hand, you have customers asking, “When is the next model coming out? When is this capability going to be available?”
You need to manage expectations. The biggest challenge, and I think we’re starting to do a good job at it, is managing the hinge between product requirements and what science is able to do.
Jensen Huang
Research and product.
Arthur Mensch
Research and product. You don’t want the research team to be fully dedicated to making the product work. You need both to work together.
I think we’ve started to do a good job of making sure that you have several frequencies in your company. You have fast frequencies on the product side, iterating every week, and slow frequencies on the science side, looking at why the product is failing in certain domains and how you could fix it through research, new data, new architectures, or new paradigms.
That’s fairly new. It’s not something that you would find in a typical SaaS company, because this is inherently a science problem.
Speaker 1
NVIDIA is one of the most successful companies, over a 30-year timeline, to figure out a way to keep science and research ahead of the rest of the world. Whether it was CUDA back in 2012, which was fundamental systems research, or Cosmos today, which is state of the art in how simulation should work, you’ve harmonized exactly what Arthur just described.
Is that right for you?
Jensen Huang
We harmonized that inside our company. We have basic research, applied research, architecture, development, and multiple layers of each.
These layers are all essential, and they all have their own time clock. In the case of basic research, the frequency can be quite low. On the other hand, all the way to the product side, we have a whole industry of customers counting on us, so we have to be very precise.
Somewhere between basic research and discovering surprises that nobody expects, on the one hand, and being able to deliver predictably on what everyone expects, on the other hand, we manage these 2 extremes harmoniously inside our company.
Speaker 1
There are so many fascinating things about this market, but there’s one in particular that I want to call out. Both of you have customers who are also your competitors, and those competitors are huge, highly capitalized technology giants.
NVIDIA sells GPUs to AWS, which is building its own chips, Trainium. Arthur, you’re training models that you sell through AWS and Azure, which have funded labs like Anthropic and OpenAI.
How do you win in an environment like this, and how do you manage those relationships? We talked about company building internally, but now I’m curious externally. How do you survive in this situation?
Arthur Mensch
Jensen said it well: You give up control, but you work on alignment. Despite the fact that sometimes certain companies can be competitors, you may have aligned interests, and you can work on specific agendas that are shared.
You have to have your own place.
Jensen Huang
Obviously, these cloud service providers aren’t working with Arthur because they already have the same thing. They want to do the same things. It’s because Arthur and Mistral have a position in the world that is unique to Mistral and add value in a particular place that is unique.
A lot of the conversation we’ve had today involves areas where Mistral, its work, and its position in the world make it uniquely good. We’re different. We’re not just another ASIC, and we can do things for the cloud service providers that are not possible for them to do themselves.
For example, NVIDIA’s architecture is in every cloud. In a lot of ways, we have the first onboarding path for amazing future startups. By onboarding to NVIDIA, they don’t have to make a strategic or business commitment to a major cloud.
They can go into every cloud, and they could even decide to build their own system if they like, because the economics may turn out to be better for them at some point, or they may want access to capabilities that we have that are more protected within the clouds.
Whatever the reasons are, in order to be a good partner to somebody, you still have to have a unique position and a unique offering. I think Mistral has a very unique offering. We have a very unique offering, and our positions in the world are important to even the people we compete against.
When we’re comfortable with that and comfortable in our own skin, we can be excellent partners to all of the cloud service providers. We want to see them succeed.
I know it’s a weird thing to say when you see them as a competitor, which is the reason we don’t see them as competitors. We see them as collaborators who happen to compete with us as well.
Probably the single most important thing that we do for all the cloud service providers is bring them business. That’s what a great computing platform does. We bring people business.
I remember when Arthur and I first met. We sat down in London at a late-night restaurant and sketched out the plan for his Series A. We were figuring out why he needed so much capital for the Series A, which in hindsight was remarkably efficient.
I think the Mistral Series A we put together was $500 million, relative to other companies that had to spend multiple billions to get to the same place. I asked him what chips he would like to run on, and I don’t think he even looked at me. He looked at me as if I had asked an absurd question, as if there could be any answer other than NVIDIA—other than H100s.
Speaker 1
What is the philosophy that led you to invest so deeply in startups and founders so early on, even before anybody knew about them?
Jensen Huang
There are 2 reasons. The first reason is that I rarely call us a GPU company. What we make is a GPU, but I think of NVIDIA as a computing company.
If you’re a computing company, the most important thing you think about is developers. If you’re a chip company, the most important thing you think about is a chip. All of our strategies, actions, priorities, focus, and investments—100% of it—is aligned with the attitude that is developer-first.
It’s about the computing platform first. Another way of saying that is ecosystem. Everything starts there, and everything ends there. GTC is a developers conference, and all of our initiatives inside the company are developer-first.
The second thing is that we were pioneering a new computing approach that was very alien to the world of general-purpose computing. This accelerated-computing approach was alien, counterintuitive, and rather awkward for a very long time.
We’re constantly seeking out and looking for the next incredible breakthrough, the next impossible thing to do without accelerated computing. It’s very natural that I would find and seek out researchers and great thinkers like Arthur.
I’m looking for the next killer app. That’s a natural intuition and instinct of somebody who’s creating something new. If there’s an amazing computer-science thinker we haven’t engaged with, that’s my bad. We’ve got to get on it.
Speaker 1
From a computing perspective, what are the most significant trends you see on the horizon? In particular, for an audience of people who might be prime ministers, presidents, or ministers of IT in some of the world’s fastest-growing markets, trying to understand where computing is going, how would you guide them?
Arthur Mensch
We’re moving toward workloads that are more and more asynchronous—workloads where you give a task to an AI system and then wait for it to do 20 minutes of research before returning.
That’s changing the way you should be looking at infrastructure because it creates more load. I guess it’s a bull case for data centers and for NVIDIA.
As I said at the beginning of this episode, none of this is going to happen well if you don’t have the right onboarding infrastructure for the agents. You need a proper way for your AI systems to learn about the people they interact with and to learn from the people they interact with.
That aspect of learning from human interaction is going to be extremely important in the coming years. There’s another aspect around personalization: having models and systems consolidate a representation of their users so they can be as useful as possible.
I think we’re in the early stages of that, but it’s going to change quite profoundly the interaction we have with machines. They’ll know more about us, know more about our tastes, and know how to be as useful as possible to us.
To close on that, as a prime minister or a leader of a country, I want to think about education and making sure I have a local talent pool that understands AI well enough to create specialized AI systems.
I want to think about infrastructure, both on the physical side and on the software side. What are the right primitives? What is the right partner to work with that is going to provide the platform for onboarding?
Those 2 things are important. If you have the infrastructure and the talent, and if you build deep partnerships, the economy of your state is going to be profoundly changed.
Jensen Huang
The last 10 years have seen extraordinary change in computing: from hand-coding to machine learning, from CPUs to GPUs, and from software to AI. The entire stack and the entire industry have been completely transformed, and we’re still going through that.
The next 10 years are going to be incredible. Of course, the industry has been wrapped up in talking about scaling laws, and pretraining is important and continues to be. Now we have post-training, and post-training includes thought experiments, practice, tutoring, coaching, and all of the skills that we use as humans to learn.
The idea that thinking, agentic, and robotic systems are now just around the corner is really quite exciting. What that means for computing is very profound.
People are surprised that Blackwell is such a great leap over Hopper. The reason for that is because we built Blackwell for inference, just in time. All of a sudden, thinking is such a big computing load.
That’s one layer: a computing layer. The next layer is the types of AI systems we’re going to see. There are agentic AI systems and informational digital-worker AI systems, but we now have physics AI making great progress, and physical AI is making great progress as well.
Physics AI, of course, involves systems that obey and understand physical laws, atomic laws, chemical laws, and all of the various physical sciences. We’re going to see great breakthroughs there. That affects industry, science, higher education, and research.
Then there’s physical AI that understands the nature of the physical world—from friction to inertia, cause and effect, and object permanence. These are basic things that humans have as common sense but most AI systems don’t.
I think that’s going to enable a whole bunch of robotic systems with great implications for manufacturing and other industries. The U.S. economy is very heavily weighted toward knowledge workers, while many other countries are very heavily weighted toward manufacturing.
I think it’s important for prime ministers and leaders of countries to realize that the AI systems they need to transform and revolutionize their vital industries—whether those industries are focused on energy or manufacturing—are just around the corner. They ought to stay very alert to this.
I would encourage people not to overrate the technology. Sometimes, when you over-admire a technology, you overrate it. You don’t end up engaging with it because you’re somehow afraid of it.
Don’t do that. Some of the things we said today about AI closing the technology divide are genuinely worth recognizing. This is of such incredible national interest that you have the responsibility to engage with it.
Speaker 1
Anyhow, exciting times ahead. That was incredible. Thank you both so much for making time. If they want to go learn more and figure out how to partner with the two companies, call us.
Jensen Huang
You’re kidding me. You can call us.
Speaker 1
Yes. Am I right? Start with listening to this podcast and then giving them a speed dial. Put their numbers in the show notes. Jensen, NVIDIA.com. Job done. You heard it here. We’re very risk-on, too.
Jensen Huang
I can attest to that.
Speaker 1
Thank you so much for listening to the a16z podcast. If you’ve made it this far, don’t forget to subscribe so that you are the first to get our exclusive video content, or you can check out this video that we’ve hand-selected for you.