Harry Stebbings
Roman, I am so excited for this, dude. I think Nebius is one of the most unbelievable, incredible stories in terms of what we've seen over the past few years, but also, holy shit, what an exciting few years we have ahead. Thank you so much for agreeing to do the show.
1. Why AI Infrastructure Is Not a Bubble
Roman Chernin
Thank you for inviting me. I'm glad to be here.
2. The Biggest Threat to Nebius Isn't Competition
Harry Stebbings
I would love to start with a question that I think is at the top of a lot of people's minds: where are we at in the inflection point for AI infrastructure? A lot of people are seeing the capital going in and saying, "Oh, it's a bubble," while a lot of people are saying, "It's just the start." How do you think about it? Are we at an AI infrastructure bubble moment right now?
Roman Chernin
No, I don't believe it's a bubble. Define the bubble. Do I believe that we will need tens or hundreds of times more to build? I thoroughly believe that. I'm probably biased. I probably wouldn't be in the business we're in if I didn't believe that.
I think we're just at the beginning of this amazing moment when Jensen calls it "useful AI." We're just at the beginning of real adoption. Honestly, we have maybe 1 use case that works out of so many use cases, and the 1 use case that works—coding—started working maybe a few months ago.
Let's put it in perspective: we're just a few months from the moment when we got maybe the first use case that works at scale, and we are starting to see it applied here and there. I think we'll see many, many more use cases and much more adoption.
What we see today is that if you take every company in the world, maybe outside of the fastest-moving startups—before the show, we were speaking about who is moving fast enough and who is not fast enough—there may be some exceptions. But if you take any company in the world today and look at AI adoption there, you will actually see that they are starting to use AI in the first 1% of the volume and the first 1% of the use cases.
3. The Real Impact of Open Source on OpenAI & Anthropic
If you take any large company, even a pretty advanced technology company, you will see that they are just starting. I take from that the fact that we're only at the beginning. Even if you don't believe what Musk says about everything in the future space and so on, practically, from an enterprise-adoption perspective, it's just the first step.
Harry Stebbings
We're completely aligned, but it's a very boring discussion if I just say, "I agree with you on everything." My question, on the back of the fact that we've seen coding work for the last 6 to 12 months, is this: there is a question of whether we will move to open-source models hosted locally, because the cost will be too significant for some of these enterprises to bear, and whether we're going to see that shift happen soon.
If we do that, is it damaging both to the providers—OpenAI and Anthropic—and to Nebius? Why is that perspective wrong?
Roman Chernin
First of all, I think it's not in the future; it's already in the present. What we see in a lot of examples at the moment is that when our customer or product builder gets to scale, they start looking for ways to improve the economics, accelerate growth, and so on. That's when many of them start looking at alternative models.
The best way to build today is obviously to build on the frontier models from great providers like OpenAI, Anthropic, and Google, because they provide the best capabilities in the world. But when you figure out the use case, start seeing adoption, and see the customer-data loop, you may find a cheaper—or even not cheaper, but higher-quality—way to serve the same use case.
You may not need the best universal model in the world. You can create a specialized model that, in your particular case, will work even better. That's when you may consider shifting from frontier closed models to open source.
The most important characteristic of those models is not just that they are open source, but that they are tunable and trainable. You can take them, post-train them, and create a specialized model that, in your particular case, may work better. That's what we see all over the world, across use cases.
But why doesn't it hurt Anthropic and OpenAI? In reality, they move to the next frontier. Going back to the point we discussed, there are so many unsolved tasks, or tasks that don't necessarily have a limited budget to be solved.
Every time we find a way to solve some task more efficiently—and we saw it with DeepSeek a year ago, and we continue to see it now—we start solving more complex tasks at the same time. This is a continuous journey, I believe. You always push the frontier, and you always have more complex tasks to figure out how to solve.
When you figure out how to solve them, you can go down and reduce the price or improve the quality. But we have so many unsolved tasks that Anthropic, OpenAI, and all the other frontier-model providers still have such an unaddressed market in front of them that they continue to grow exponentially.
Harry Stebbings
Do you buy that? These companies are priced to perfection in a lot of cases, at $1 trillion. If the value that they create is eroded, and they're constantly playing a game of leapfrogging from value to value to value while open source continuously eats behind them, you've got to find a lot of problems continuously, dude. That's a hard life to live.
Roman Chernin
Actually, most people are concerned about the other side: will we have a strong enough open-source and specialized-model environment to build this floor?
I think we're at such an early point of adoption, and we have so many unsolved problems, that it's just a matter of the total pie. I think there is enough space to solve so many tasks in the future that there is enough pie for both frontier capabilities and highly tuned models for specific use cases, as well as the whole world of open-source or specialized models that we can build on top of them to gain economic and performance advantages when we know what we need.
You said that every time we have a cheaper model, it hurts the business. My favorite anecdotal story about that is the DeepSeek moment. If you remember, about 15 months ago there was this DeepSeek moment. I remember that Nebius stock went down 40% in 1 week or so—in February, I think it was February or March 2025.
The anecdotal story is that the same exact week, we probably had the best week in sales. We were pretty early in our story, but that was the best commercial week in the history of the company, because so many people figured out that they could run inference in their production workloads with DeepSeek and the economics would work.
At the same time, Cursor started growing. I think they were the first to really benefit from tuning those models for coding and so on. Every time we get intelligence cheaper—the same unit of intelligence cheaper—we're not reducing consumption; we're increasing consumption.
4. Jevons Paradox: Why Cheaper AI Creates More Demand
We can solve more complex tasks with the same budget, or we can finally solve economically viable tasks that we already knew were solvable, but where the economics didn't work and we couldn't scale. I think it's quite fascinating to observe these economic improvements.
Harry Stebbings
Speaking of Jevons' paradox—producing more and that yielding more demand—where are you not moving fast today that you would like to be moving faster?
Roman Chernin
Everywhere. When we think about how we build a company, we talk about it in 4 dimensions. One dimension is capacity: how many megawatts, gigawatts, and GPUs we deploy.
We are an infrastructure company. We need to be large. If you're not large enough, nobody needs you to exist. This is the physical-world expansion. The team is doing an amazing job, but it's never enough. You want to move as fast as possible, and there are a lot of complications in the real world that prevent you from moving fast enough sometimes.
To launch a new data center, you need to go through the entire supply chain, regulatory processes, fiber and water, and everything that happens in the real world. That's 1 dimension. Another dimension is the product. You want to move fast enough to address new types of workloads and new types of customers coming to the market.
Think about it. We started, as an industry, in this AI journey with the people who first of all built the models. Those were companies like OpenAI, hyperscalers, large labs, and so on. What they need from you as an infrastructure provider is bare compute: just throw infrastructure at them. We see a lot of these large bare-metal deals on the market, and we do them as well, but this is only the first layer of what we build.
5. The Four Layers of AI Infrastructure Explained
The first layer is scaled physical infrastructure that customers like Meta and Microsoft, in our case, can consume at large volumes. The second layer is what we call multi-tenant cloud. It is still addressing research-heavy teams, but now we have hundreds or thousands of teams that don't want to deal with physical infrastructure. They want to deal with managed infrastructure—classical infrastructure as a service, in cloud terms.
You have storage, compute, networking, virtualization, and a good environment with APIs, observability, security, and everything that normal teams expect a cloud to have. You log in, get your cluster provisioned, and can start training or run inference if you need to. You can manage your application or workflow yourself, but have the infrastructure figured out for you.
If the first layer speaks in megawatts—and literally, if you read announcements, someone signed a large deal with Meta, Microsoft, or OpenAI, people speak in megawatts there—it's like you deliver the megawatts of compute. When you speak about this managed cloud, people speak in GPU hours, because this is the key unit you sell: the efficient hours you spend on compute, with storage and complementary services. You still buy managed compute.
6. If Nebius Had 10x More Capacity Tomorrow
Then the next layer that we're working on is managed inference, when people don't want to think in terms of GPU hours. They don't want to figure out B200s versus H200s versus B300s, what is better for a particular workload, or manage which LLM or SLM to deploy themselves and do all the optimizations. Our product, called Nebius Token Factory, is a managed inference platform. Again, this is a new type of customer—mostly people we call vertical AI companies or enterprises.
These are people who actually build products. They don't build models; they build products on top of the models. This is to your point about specialized and open-source models, when they need to shift from Anthropic, for example, or diversify the models they use. Again, this is a new primitive that we provide, or a new kind of entity that customers need. Now we speak in tokens. It's not that you pay for GPUs; you consume tokens, and you can build your applications without thinking in terms of the clusters underneath.
But I also don't think this is the final stage of where we're going, because now people are building agentic applications and agentic workflows. When you build an end-to-end agent, you may not even think in terms of a particular model or a particular number of tokens that you want to generate. You want the end-to-end task to be efficiently executed and provide the expected outcome.
The magic that the platform can make is to think for you about which model is better to use in a particular call. Do you need to go to the smarter model, or can you ask two times within the same inference budget? You can request two lighter models, get less-smart tokens, and then have a judge model choose the best result. You can also determine what size of context you should have, and so on. This is the next layer, when a developer may not even think in terms of particular types of tokens, but in terms of end-to-end execution of their task.
Harry Stebbings
So that's layer 4, a direct competitor to OpenRouter.
Roman Chernin
What we would love to bring at that level is the same thing we do on the layers below: the optimization engine. You can build your agent in so many kinds of open-source or proprietary tools, but when you need to scale it, you start thinking about the economics and reliable execution—repeatable execution.
It's not just a model-choice problem or an outcome problem; it's a system problem. You need to make it reliable, repeatable, and economically viable. That's probably where Nebius could create value. In the same way that we don't tell people how to build their applications, we just say, “If you need this model to work for you with these economics, we will help you optimize it.”
The same is true here. If you need this agent to run end to end with this budget and this quality, maybe we can help you optimize it. Again, just to make sure, this is a little bit speculative—thinking about what's next. It's not what we already have, but this is where we see our customers evolving and where we think we could create the next layer of the product offering.
Harry Stebbings
I love this, and I have all of these notes. I want to go through the 4 pillars that you said there. You said, number 1, capacity.
Roman Chernin
Yeah.
Harry Stebbings
If you had 10 times the capacity today, what would be different? Could you sell it overnight?
Roman Chernin
Yeah, it's a good question. Not overnight, but we would definitely have demand for that. I think the key question for us isn't whether we have demand, but how we actually build a portfolio of demand, because you have so many customers in this market that you can balance between.
Again, going back to the 4 layers of the product, you can sell bare metal, managed cloud and infrastructure, inference, and maybe in the future some new layers of product. I think what we're trying to do is build a quite diversified portfolio of customers. We believe that the higher up the stack we move, the more value we can potentially create for customers.
Actually, the higher up the stack we move, the larger the population of customers we can serve. At the bare-metal level, you have maybe a dozen customers in the world that you can work with. At the managed-infrastructure level, there are hundreds; on inference, there are thousands; and on agentic applications, there will be tens of thousands of new developers building them.
Harry Stebbings
Right. Okay. On the customer portfolio, I love that for the capacity. You want to be big enough that you're meaningful, but not too large that the business relies on them. With that delicate balance, where do you settle on what revenue concentration with Meta or Microsoft you're happy with?
Roman Chernin
It's a great question, and I would say it's a main question of our business—not even Nebius, but the product category. We've always said publicly, and to our investors and customers, that the long-term strategy of Nebius is to serve as diversified a portfolio as possible. We do our best to have many customers that we work with.
If, in reality, you're serving a dozen customers in the world at the level of Meta and Microsoft, which are super advanced and have their entire software stack, they literally need only physical infrastructure. They bring everything with them, deploy it on your infrastructure, and run it. You have tiny, tiny additional value that you can provide them above the physical infrastructure.
By the way, satisfying them with what they need in physical infrastructure is quite a challenge, because you can imagine that they are quite demanding. They need the most scaled infrastructure in the world that exists. Sometimes people say it's a commodity, but it's not really a commodity at that scale. Nothing is a commodity when it comes to real scale.
Again, to your point, this is quite a small population of customers that you can work with, and you don't necessarily need all the full-stack software to work with them. We intentionally built—and from day 0 of Nebius, we were building—this software stack because we thought it was much more beneficial for us and for the world to have someone who can support customers not only on this physical-infrastructure layer, but beyond.
Harry Stebbings
For the long-term protection of the business, do you not have to build the full stack? Otherwise, you become the capacity provider to these mega-players, which will make a shit ton of money. But you're incredibly concentrated and very vertically focused.
Roman Chernin
Yeah, I think so. Again, we don't know where the world will end up. In a world of infinite demand, you may sustain, even long term and midterm, selling these bare-metal contracts.
But the more competition you have for customers on the demand side, the more you can be picky—even with the customers you work with—and work with customers that appreciate the value of the platform we built. There are different customers in the world: someone is more obsessed with price, someone is more obsessed with quality, and someone really wants to have a much more advanced platform because they want to concentrate and focus on their platform or product and not spend time on infrastructure.
Harry Stebbings
Yeah. Before we move to number 2, being product, just staying on capacity: given the insufficient supply of capacity today, if you doubled pricing, would you see any change to demand?
Roman Chernin
It's a difficult question. We actually raised prices just a couple of months ago.
And we still have fair pipeline pressure, let’s say, on supply. Again, we don’t really know where the balance is, and I’ll tell you why. It’s not only us being greedy and wanting to get as much money as possible, with people in the shortage still having to pay to some extent. It works like this: people need compute to build, but then there is a point—and especially, it’s less so in training, because in training it’s a one-off cost—where the economics doesn’t work. If you believe that we’re moving to inference, and inference is the cost of serving the customer, there is a level where the economics of our customers’ products doesn’t work. If they work, they can grow, and then we can grow with them.
It’s not just a supply-and-demand situation with absolutely elastic prices. They are elastic to some extent, but we also want to be meaningful and thoughtful about what our customers need. By the way, it’s not only the GPU-hour cost; it’s all the optimizations you do, all the real cost—we call it TCO, total cost of ownership—that you incur. This is partially why we build the software platform. I’m sorry to come back to product again and again, but you want to speak about capacity, and people are too obsessed with capacity. Capacity is important, but people are too obsessed with the nominal price of capacity.
You can price a GPU at $3, $4, or $5. Depending on the use case and the quality of the platform, it can create completely different outcomes for the customer in real cost. How long does it work? What is the effective, uninterrupted time that you can run there? If you talk about inference, how many tokens can you extract? We see all these optimizations happening that change the price of the tokens by an order of magnitude. People speak so much about the cost of a particular GPU, but if you do the right thing with the model, you can change the price by many times over.
This all should work together as a system, not just as—again, if you speak about raw infrastructure, then you can manage only the price. But if you build the platform and provide a high level of service to the customer, then you can extract much more economics, not only from the infrastructure cost structure, right?
Harry Stebbings
If we move to that second layer, moving away slightly from capacity to GPU hours, the product itself—multi-tenant—what is the main question that you ask yourself within that segment? If in the first layer, capacity, it’s how much revenue concentration we have, what is the big question in that layer of value?
7. The Shift from Training to Inference and Agents
Roman Chernin
What does the customer need? You speak with a lot of product founders, and this is the same: What does the customer need? How do customers evolve in their needs? Where is demand moving? We see all this transition from training to inference, from just using the models to building agents, and from mostly AI labs consuming AI compute to enterprises coming into the game. All the time, if we want to be relevant, we need to follow the changes, and this is the main question we ask ourselves in the product: What should we build, what do customers need, and what is Nebius’s value—what is the value we need to create? Because, again, we are a small company, we cannot build everything, and we need to be very precise about what we can do better than others and where the value is that we should focus on, given how customers are evolving.
Harry Stebbings
What changes are you seeing in customer needs that you’re not seeing discussed much in public?
Roman Chernin
Everybody’s talking about moving from training to inference. I think it’s a very 30,000-foot view, because this move means that people are actually building specific products, and in those products they have their economics and their trajectory of growth. It’s not just that the same GPU is being used for other purposes. I think it brings new requirements: you need to build your inference platform, and you need to help your customers not only run inference, but understand where the model they’re running inference on comes from.
Everybody is taking open-source models and fine-tuning them. So how do we help them? And then, when they run them, they generate a lot of data. How do we help our customers, when they’re already running their application and their inference, collect the data, curate it, and then use it to improve the model or the application that they run? It’s this flywheel analogy: you run inference, you generate data, you can observe this data, then you can improve the model that you run and continue to improve the quality of the end product.
I think there are a lot of pieces, both at the system level and at the AI-magic level, if you want. I think the most fascinating moment for me is that what we see is that the barrier to building is going down. We see more and more customers—builders—coming to the market who are not necessarily AI researchers or inference engineers. The value that companies like Nebius can create is actually to lower the barrier to building AI-enabled products and AI-enabled applications that really work, and hide from the developer all the complexity of infrastructure and some of the complexity of AI, like how you tune the model or how you optimize the inference. It’s also a very research-heavy area, and we can just let people focus on their customers and use case, by the way, the same way they do with closed ecosystems like Anthropic’s and OpenAI’s.
Harry Stebbings
You mentioned the word differentiation, and one thing I was discussing with my partner before is a theme we have to discuss, and it’s within these layers. You’ve spoken extensively about product buildout and the importance of building the product underneath capacity. When people look at you versus other neoclouds—when we look at you versus CoreWeave—you both run GPUs, you both have NVIDIA relationships, and you both have Meta as a customer. What’s the difference?
Roman Chernin
I don’t like comparing with others. The principles we build on are full-stack; we call it full-stack integration. You can think about it as full-stack down and full-stack up. Full-stack down is that we’re really deep in the physical world: we build data centers, we build racks and servers, and we build the platform. When you control these kinds of things downstream, you can move faster, squeeze more cost, and provide more economically viable solutions for customers.
Then your vertical integration upstream is actually what we spoke about: product, and how you can follow the customer’s needs and customer segments, not be limited by the small population of people who just need infrastructure, but really serve enterprises and product companies and meet them where they need us. This is, I think, what we do differently. And how it’s showing up, I would say, is, again, less concentration in the business and a more diversified customer portfolio. We believe in better long-term positioning as we go to enterprises, where we believe eventually a lot of demand will come from.
Again, now most of our segment is AI natives working with AI natives, but we have a huge market of enterprises, existing companies, and someone needs to serve them. They will not buy raw compute; they will need platforms and tools. They will need us to respect their legacy and be able to work with their more complex environment. They’re not nimble; they have data to migrate and systems to integrate. That’s the big game, and I think that, for us, it’s the main direction to move.
Harry Stebbings
You mentioned the third layer of the four-pillar stack being managed inference. For people who don’t understand, how do you think about this layer, and how would you explain it to them?
Roman Chernin
Yeah, very simple. You built your product on whatever you call your—where you write code.
Harry Stebbings
I’m actually an OpenAI investor.
Roman Chernin
Code. Okay, good enough. You built your great product with OpenAI. You cracked the use case, started growing, and have amazing traction. The only problem may be that you don’t have enough margin, or you want to start applying the data more aggressively and tune the behavior of the model, and you cannot do that in the closed ecosystem. So you go to the internet and read that there are a lot of great open-source models that, on the benchmarks, are close to OpenAI, and you think, “Oh, great. It will be 10 times cheaper, inference is cheaper, I can tune those models, I can apply my data, and my product will be better and my growth will accelerate.”
So you go, you take the weights from Hugging Face, you take some engine to run it, like vLLM or SGLang or something, and then it doesn’t work. Because to really extract the value you expect, you need to do optimizations, deploy it in a proper way, and not just have one GPU for token generation or a one-host setting. If you have a large product, you run it on hundreds or thousands of GPUs already. You need all the orchestration, caching, and observability. Your customers ask you, “How does it work?” and so on and so forth.
8. How Token Factory Cuts AI Costs by 70
By the way, you had all of that on OpenAI because this is the production service for you. You don’t think about infrastructure when you work with OpenAI; you just subscribe to the plan you need and pay for whatever end result. That’s where you need the product, Token Factory. Token Factory gives you managed inference with open-source or specialized models. You can run an existing open-source, vanilla open-source model, or you can tune the model and deploy your own weights, and then we’ll take care of all the rest. We’ll apply all the optimization techniques, manage better economics for you, and it will be reliable. You don’t need to think about the next 100 GPUs, where you will find them, and so on and so forth.
It’s a service. It’s like a managed service.
Harry Stebbings
With Token Factory, you run on 60 open-source models. You said before about cutting inference costs by up to 70% through optimization. Can I ask a dumb question? How do you actually make a token cheaper?
Roman Chernin
It’s not magic. You take the model, some baseline model, and then you can optimize it for the particular scenarios that you have. You can distill the model, make a smaller model that works with the same quality, do speculative decoding, optimize caching, and so on.
You take the model, and out of this model you actually build a system that, in your particular case, works with your requirements and optimized economics. By the way, one of the things that I think is also important for customers using managed platforms like Token Factory is that the models are changing every week or every month.
Right today, maybe minimax 3 was released, and there is Ultra that was announced and released. This happens every few weeks. Every time a new model is released, it may work better on some benchmarks and maybe not on other benchmarks, and you want to have flexibility.
You want someone to support you in experimenting and actually adopting the best new models for your use case every time they come online. Platforms like ours abstract away all the work that you need to do to change from one model to another, benchmark all of them, and so on.
You can be sure that you’ll be on the frontier every time something new is happening. It will be in the platform, you’ll be able to test it, and if it works better for your use case, you’ll be able to switch. It will all be smooth and transparent for you.
Harry Stebbings
Does the pace of model development sustain? You said every couple of weeks. I would argue, respectfully, that it’s every couple of days. Does that sustain in 5 years’ time? Are we seeing that level of iteration?
Roman Chernin
I don’t know. There’s a good chance that we’ll continue to see a lot of niche models show up and improve. I’m a believer that we’re quite far from the wall, and we’ll see a lot of model improvement happening.
I think what we’ll also see is many more new modalities and specialized models coming into the game. We speak about these frontier LLMs, but there’s an entire world of life-science models, robotics, world models, video models, image models, and so on. They all have their own use cases as well.
We’ll see more and more small, specialized models for particular use cases that are very much optimized. Just this morning, I spoke with a team here in Israel that develops a cyber defense foundational model—a model optimized to build cyber defense agents.
They don’t start from scratch. They take one of the open-source foundational models, but then they train it for the particular case, optimizing for the quality and latency needed in cyber defense use cases.
I think we’ll continue to see a lot of specialized models, both pre-trained and post-trained, that still need optimized inference and optimized infrastructure around them to let customers use them.
Harry Stebbings
Going back to Token Factory, token costs, and token usage, what are you seeing that you don’t think other people are talking about enough? What has shocked you recently?
Roman Chernin
I think everybody is speaking about the same thing: how fast it’s growing. When we see the trajectories of companies like Anthropic, Cursor, and Cognition in coding, and now we’re starting to see it in other verticals as well—healthcare examples and financial use cases—I think it’s quite amazing.
What’s interesting is to see how non-AI startups are moving. We have Revolut as a customer. When we started working with them, I think 99% of their inference budget was in closed models from OpenAI.
They started to crack some of the use cases, and some of them didn’t work for them economically. They practically couldn’t replace humans or enhance humans in the use cases they wanted to address. They started moving to open-source models, but it didn’t move fast for them because they had to spend time building the entire engine internally in the company.
First of all, they were focusing on evaluations. I think this is something that people underestimate: how important it is to build the foundation for an improvement and experimentation engine.
As a company and as a team, you need to understand what is good for you. You solve a use case and it works, but then you want to change the model. How do you know that you’re not ruining the quality? You need metrics, you need an evaluation mechanism, and you need to have this CI/CD process established for AI development.
What we see with many customers, like Revolut, is that they have these foundational investments that they need to make in understanding how to evolve the models and how to safely integrate them into their production processes.
When they solve these foundational problems, they start growing exponentially. I wouldn’t underestimate how fast those customers can grow when they build the system that lets them ship fast.
Shipping fast means they know how to evolve and they know how to make decisions. This is something that we see across a lot of customers. They have what you can call foundational investments, or a cold-start problem: how to start shipping.
But when they solve it, they start to grow exponentially. They can use different models, build many more products inside the company, and so on. I think this is something where, when you look from the outside, you say, “They’re not growing. They started small, they’re taking time, and so on.”
But if the company has a strong team, they build this foundation and then they start growing exponentially. I think we’ll see a lot of explosive growth in enterprises—in digital companies, cloud companies, and cloud-native companies like Revolut, Shopify, and Booking.com.
When they solve this cold-start problem and build the system for how to ship, their AI adoption will grow like crazy.
Harry Stebbings
How much more do you think Revolut will pay you in 3 years’ time?
Roman Chernin
I don’t know. I don’t want to speak about that. But I can say that, in total, I think they grow multiple times over. They grow like this.
We all see these AI companies reporting ARR growth. For them, it’s not ARR; it’s their budget. The most advanced companies are growing their AI budgets—not this fake-or-not-fake, all-this-maximizing kind of race, but in terms of how they do it in the production workload.
They’re growing at the same pace as these AI-native companies report. Their AI consumption is growing in line with their ARR. Companies like Revolut are growing along the same exponential trajectory.
Harry Stebbings
I always push back on people who proclaim that open source would be a credible threat to the largest model providers. I say, listen, the biggest enterprises want reliability. They want security, and most of all, they want ease. They don’t want to be tinkering around with all the architecture and everything beneath the surface.
What you’re telling me is that you’re able to provide all of that, allowing them to move away from those providers and have a cheaper, better experience because you take away the plumbing. Correct?
Roman Chernin
Yes, but again, my point is that it’s not about closed models versus open-source models. It’s not about whether they’re reliable or not reliable. The work of companies like Nebius is to make it possible, as you say, not to think about the plumbing if you want to use alternative models.
But I think it’s about capabilities. Closed-source frontier models are great, and they’ll become even better. They’ll solve so many problems that we haven’t solved yet, and we have such a diversity of use cases that we want to solve.
There will be a market for the smartest models in the world, the fastest models in the world, and the models in between—smart enough but cheap enough. As a customer, you’ll be able to pick the right source of tokens for each particular task.
Going back to the agentic-layer point, maybe it won’t even be the customer’s task to choose which model to call. It will be the engine that knows all the capabilities of all the models underneath.
When you go to OpenAI and do research, you don’t think in terms of how many loops you want it to make. You don’t think about when it should go to an LLM and when it should go to search. You don’t think about which prompt it should call. It’s happening automatically.
You give it a task, there’s a reasoning engine that decides how to run the task, and you get the result. I think a lot of enterprise cases and agentic tasks will be solved in the same way.
It won’t be you, as a developer focusing on the customer’s needs, who has to orchestrate all these tokens and models. We’ll need all the models: the smartest ones for the most complex intelligence, and the fast models that can do quick iterations.
We’re not even speaking about all the modalities and what we’ll need in the physical AI world. My point is that we’ll have enough need for different models.
What we need to do as an infrastructure company is help, to the extent we can, developers feel comfortable using all the capabilities that models provide. Because, as you rightly said, it’s not about model capabilities.
9. Sovereign AI, Europe, and the Future of Model Building
Harry Stebbings
It’s not only about model capabilities. It’s about plumbing: getting them working, getting them optimized, and getting them reliable.
When we look at the explosion of models and the specialization of models, like you said, and how many will be built and the depth across different use cases, sadly, the one thing that is quite clear is that Europe does not have anywhere near the model buildout that we’ve seen both in the US and in China. How important do you think it is that nations have their own sovereign models?
Roman Chernin
It looks like the world is divided, whether we like it or not. I think that having good-enough foundational models available for the big parts of the world is important.
I think here in Europe, or at least in this part of the world, we should think about how we have enough capabilities available here. We’ve had a lot of conversations over the last couple of years about sovereignty and this whole sovereign AI agenda, and I think it was too concentrated around megawatts and power rather than what we have on the builder layer.
Megawatts will come. I think what we at Nebius have always said is that we will build infrastructure. Companies like us will build infrastructure if we have demand, and demand is coming from the builders. What we need to care about here is having more great companies like Lovable, Black Forest Labs, and likely Mistral.
We need enough people investing in research and enough people investing in products. Then they will create enough demand, and there will be enough of a flywheel to have good-enough models if we need them. I think this is something that we should care about.
Harry Stebbings
Where is the most interesting area to invest today? I’m giving you 4 options: infrastructure, horizontal model, vertical model, or application layer.
Roman Chernin
We build infrastructure, so we’re quite happy here. I think it’s a good place to be in the current world.
10. Competing Against Hyperscalers with 10x More Capital
Even though, to some extent, we are building kind of the easiest part—not in a way that it’s easy, because it’s complex execution—we kind of know what’s needed, and our customers help us understand what’s needed. I think the most amazing people in this industry are those who take a risk to go and build end-user products, in my view. They actually drive most of the growth here: people who take the real risk of building something people would need or not need. I think these are the heroes of our AI journey.
Harry Stebbings
Speaking of heroes of AI journeys, before I do a show, I go and speak to—I’m very fortunate. You mentioned earlier that I’ve interviewed some big people, and I go and speak to some of those big people.
A theme that did come up when I was speaking to them was their relationship with NVIDIA. Is a marriage a marriage if one has more power than the other? How do you think about the power dynamics in a relationship with NVIDIA when they have so much power?
Roman Chernin
We look at this in a very simple manner. We just need to build what we build. We need to build our product, we need to tell our story, and then the rest will complement it.
I think what’s most fascinating about NVIDIA is that it’s still, to a big extent, an engineer-driven company. The best thing you can do to get respect from NVIDIA—it’s my read, and they may have a different point of view—is to have NVIDIA’s engineers respect your engineers. You will have the right foundation for the relationship.
I think we’ve managed to prove again and again that we know what we build and that we have a strong engineering team. I think they see it and respect it. We have a lot of engineer-to-engineer relationships on a hardware level, on the software layer, and on the inference platform layer.
The better NVIDIA’s engineers think about you, the better the relationship and partnership it enables. We may be wrong thinking this way, but that’s what we see we can do. We just focus on being reasonable and focused on long-term value.
It sounds fluffy. Everybody says it. But just do your fucking job at the end of the day, right?
Harry Stebbings
I’m going to title this, Roman: “Just do your fucking job.”
Roman Chernin
No. What else can we do? We’re in such a race, and we can just do our best to do our work better. I think that’s that, yeah.
Harry Stebbings
“Just do your fucking job.” I just—I know it’s funny. I like it. But what’s the hardest part of just doing your fucking job today?
Roman Chernin
Four dimensions: build scale, build product, work with customers, and capital. It’s actually like 3 dimensions at first. We discussed scale and product; the third is customers.
We are in the field business. We like to say that cloud is a post-sales business. When you sell, you sell the promise, and then you need to satisfy the customer. Working with customers, covering the customers, and having this strong customer-facing engineering team—the FDE team—is the third dimension.
Go talk to your customers. Make sure that they know you and that you know them. This is the third dimension.
The fourth, the most boring but also the most exciting, is capital. We’re in a capital-intensive game, and we’re competing with the most capitalized companies in the world.
Harry Stebbings
If I gave you an unlimited budget, what would you do differently?
Roman Chernin
Build faster. That’s very easy.
Harry Stebbings
Build what faster?
Roman Chernin
Data centers, and fill them with GPUs. Just build faster. Our capex program this year is 2025 billion. Our competitors, the hyperscalers, have 8 times more. If I had 10 times more capital, I would just build more data centers, fill them with GPUs faster, and serve more customers.
That’s what we started with: what would I do if I had 10 times more supply? I would move faster.
Harry Stebbings
Gavin Baker said, I think quite intelligently, that permitting, regulation, and the delayed buildout of data centers have actually helped, because if I enabled you to build 10 times the data centers today, it would actually create the glut.
Roman Chernin
Yeah, it’s actually a great question. Our investors sometimes ask us what the main bottleneck is, and the main bottleneck is everything. But you need to look at this from a time-span perspective.
In the next 6 months, capital cannot help you. 6 months is too short a time. You have what you have, and you need to deliver. In the next 12 months, you can accelerate something, but again, it’s more about capacity constraints. In the next 12 months, we can accelerate something with capital or with execution. But in 24 months, you definitely can unlock so many things.
We’re not building one data center. It’s also important to understand that we’re building a portfolio of capacity. The more execution power and capital we have, the more things we can do in parallel and unlock.
That’s why we do what we do. We secure power and land, then we build data centers, and then we fill them with GPUs. Every next stage requires more capital, but we do as much as possible in advance to make sure that when we’re at the next stage, we already have power secured. When we have enough capital to deploy in GPUs, we’ll have data centers that are up and running.
It’s phases of investment, and again, the bottlenecks are different from a different time-span perspective. Obviously, if you have more capital, you can move faster—not in 6 months, but in 18 or 24 months, for sure.
Harry Stebbings
Can I ask you, when you think about the data center buildout that we’re seeing, there’s more and more public angst toward AI? Eric Schmidt is getting booed off stage, not because of the content but because of the AI innovations. We’re seeing public resentment toward data centers. I think 40 out of 100 now are not being built when they go through planning and approvals. How do you think about and reflect on that internally?
Roman Chernin
This is the environment we need to work in. Again, there are two sides to the thing. One is how we think pragmatically as a business. That’s what I said: we think about it as a portfolio of projects. We need to make sure that we’re oversubscribed, if you want, so that if one data center is delayed, we will still deliver enough capacity to our customers.
Most of the customers are not locked into one physical location. It’s a cloud. We can build in different places and then bring the workloads where we have capacity. That’s the pragmatic side of things.
What we obviously see is that communities and local authorities require companies like us to work closely with them, explain and show what we do, work with them on their concerns, and address them. This is the reality.
You can compare it to when Uber started growing. In many places, there was pushback: “What’s happening? It’s something new. We didn’t expect it to move so fast.” I think you go and work and explain. It’s just a part of your duty to engage and work with the new communities that become dependent on you. They have concerns, and sometimes they have concerns because they’re not educated enough. Sometimes they have rational concerns that you can address.
And the same: do your job.
Harry Stebbings
Do you think you’ve done a good job at it so far?
Roman Chernin
We come from a place where we always think that we didn’t do enough.
I think we made quite a bit of progress in the places where we started building. Historically, we had more experience in Europe. Now, probably 70–75% of the new capacity that we’re building in the midterm is in the U.S., so we’ve built a lot of presence on the ground to communicate with those local communities in the U.S. We try to do the best job.
We need to do better, always, but we’re moving.
Harry Stebbings
Can you help me on another one? We laughed earlier when we talked about space. Data centers on planet Earth are a very difficult logistical buildout. Data centers in space—I love technology, I’m an optimist. I hope it is that fucking nuts.
Roman Chernin
I think everything we see is fucking nuts. So many smart people are now working to make it happen. My view is very simple: so many smart people are working on this, so most likely, I may be less pessimistic that we’ll see—I don’t know whether we’ll build more in space than on Earth in 3 years.
I’m humble enough to say that so many smart people are trying to solve this task and bring compute to space, so why wouldn’t I believe it will happen? I think there are still a lot of challenges and a lot of things to figure out.
But if someone had said to us even 3 years ago that we would build multigigawatt data centers and that it would be large, interconnected compute clusters, would you have believed it? I didn’t think like that, and we’re here. It’s routine.
Harry Stebbings
I want to do a quick-fire with you. I say a short statement, and you give me your immediate thoughts. What job does not exist today that you think will be very common in 5 years’ time?
Roman Chernin
One thing that is obviously happening is that we’re democratizing what people call being a developer right now. Each of us can be a developer. What I mean by being a developer is converting an idea into some digital asset.
I hope that, again—we have to be optimists here—I hope that this democratization of building, letting each of us be a builder, will open up so many opportunities that we don’t even imagine yet. When we give millions of new people, tens of millions of new people, the ability to convert their ideas into something that works very easily, we will see a lot of new businesses and a lot of new ideas coming to life.
They will create a lot of new work that we don’t even think exists. It’s like a second-order effect of all this democratization of building.
What is challenging and what will need to be changed—and I think it’s as risky as it is an opportunity—is how education will change. Now, when everybody has access to intelligence, what should people learn? You definitely don’t need them to learn the facts. Everything is available. All the knowledge is kind of available.
How do you really train people to think when they don’t need to think so much? How do you teach people to continuously change? Many professions will not be stable. How do you help people find themselves in a changing environment and actually think and learn new concepts constantly?
I think this gives a lot of new opportunities, but it also creates a lot of risks.
Harry Stebbings
You mentioned that you have 2 teenage daughters. What do you advise them as they’re entering the workforce in the next 10 years? What do you advise them?
Roman Chernin
What I literally tell them is that I think 2 things will be needed. I don’t know what will be needed, but I’m sure that 2 things will be needed.
One is being able to communicate with people with empathy, with empathic communication. Understand humans, communicate with humans, and be empathetic.
The second is creativity—all the art. I hope that art, in a way, will exist. I think that all the hard skills that I thought would be needed 10 years ago, when I thought the most important things they needed to learn were math and engineering, I’m now far from that belief.
I’m quite happy they’re much more focused on soft skills than I was when I was a kid. Again, being able to communicate with humans, understand humans, be empathetic to humans, and have this creativity—being able to try new things and be creative.
I think these 2 things, if you can help your kids develop them, will mean that in 10 years they will be in demand.
Harry Stebbings
There’s a question of how you teach creativity, but I completely agree with you. The big finish: complete this sentence. The biggest threat to Nebius is not competition but…
Roman Chernin
Consolidation in general. I think the main threat for Nebius as a business is that the world will become too consolidated. Again, as we discussed, we try to be diversified. We try to solve the problems of different customers and have different customers on different layers.
If you end up in a world where, I don’t know, 3–5 super models, super companies, or super empires control the world, then Nebius, or companies like Nebius, will be needed only to help them serve their needs on the physical layer.
In general, I think consolidation is our main threat. The more democratized the world, the more diversified the world, the more we’re needed as a business.
Harry Stebbings
Do you think that’s likely? We’re seeing the concentration of value in fewer and fewer players. We’re seeing the opposite of diversification.
Roman Chernin
I hope it will not happen. As a business, I think it’s better for us, and for humans as well. For you and me, I hope the world will remain quite diversified in different manners, and I’m optimistic here.
I think there are so many people who want to build something independently. There are a lot of people with the need to try things and build new things, and that organically creates this pressure and organically creates a more diversified world. Hopefully, it will remain that way.
Harry Stebbings
Penultimate one. Leo Aschenbrenner is a famous investor right now and has a huge cult following. He recently disclosed a very large position for him: 5.3% of the company. I think it’s 15% of his portfolio. How do you guys sit internally? Are you like, “Yeah, go, Leo”?
Roman Chernin
I wouldn’t say that we didn’t notice it. Obviously, everybody noticed it. The stock jumped, and it was big news all around.
Again, I think we take it as a justification of what we do. Then you get this justification and say to yourself, “Okay, those people give you credit that you will execute.”
I come back again and again to the fact that what we do is a post-sale business. Every time we sign a deal, every time someone invests in us, they give us credit and the opportunity to deliver. Then go back to your job and deliver.
I think we’re in such an emotional market as well that you should keep yourself down to earth. Remember that all this growth, all these credits that customers give you—it’s an opportunity to deliver. Go and do your job.
Harry Stebbings
You’re such an Israeli. Americans would be like, “Yeah, go!”
Roman Chernin
I think I’m Russian in this way. Russians always know that you need to look at things very pragmatically. Russians always have these faces, like they always expect something will happen, and you need to be ready. You need to be ready.
I think it’s a really important part that comes from our CEO and founder. You wake up, and it’s a new customer, a new day; you need to deliver. Nothing is guaranteed. You need to concentrate on the work.
I know how much effort the team is putting in to make things work, how much depends on every day’s dedication, and how fast the market is moving. To stay relevant, you need to continue moving at the same pace—or try to move at the same pace—with the market.
Again, on a romantic note, I would say that we could celebrate a little bit more, but we just don’t have time to use the opportunity to actually say kudos to the team. I don’t think we celebrate enough, and I think it’s right—we’re not relaxed—but I think we could celebrate a little bit more and give the team more respect for how much has been done.
It was not easy, it’s still not easy, and it will not be easy. But, yeah, never stop. We cannot stop. It’s like a shark: you’re alive when you move, right? This famous thing. So we have to move.
Harry Stebbings
On that note, I cannot thank you enough for joining me and for putting up with my very meandering questions. You’ve been fantastic, Roman. A really huge thank you.
Roman Chernin
Thank you. Too kind to me.