# Rethinking Legacy Data Infrastructure with Eon Co-Founders Ofir Ehrlich and Gonen Stein

No Priors · 2026-08-27 · 35 min · https://www.youtube.com/watch?v=YscDZpVF4CQ

## Transcript

Elad Gil

Ofir, Gonen, thank you so much for joining me today. It’s great to see you.

### What Eon Does

Gonen Stein

Absolutely. Thanks for having us.

### Autonomous Security Threats

Elad Gil

One thing that you’re doing at Eon is—actually, why don’t you give a quick overview of Eon and what it does? I think that will set the context for how we think about AI, data, models, and fine-tuning models. There’s a whole stack built on top of different types of data sets, so maybe we can start with what you do, and then walk through how the world is shifting relative to the enterprise data stack.

Gonen Stein

Sure. At a high level, we’ve created a new data foundation that runs in the cloud. We provide multiple capabilities that allow customers to first map and classify their data across their environment, across multiple hyperscalers, and identify what they have, where they have it, what’s sensitive, what’s not sensitive, and so on and so forth.

Then we provide the ability to easily ingest that data from all these different sources—structured and unstructured data—into this data foundation. The data foundation provides a very cost-effective way of both maintaining the data for protection and recovery and making sense of it. It allows customers to very easily access it, query it, search through it, and apply their AI models and LLMs on top of that data, which is ingested from a variety of sources.

### Data as Moat

Elad Gil

My sense is that your starting point was really a backup, data recovery, and protection service. Along the way, you realized that if you have all this data from a backup perspective, and you have all of a customer’s history over time, you can start using that for interesting applications. What are some of the directions where you’re seeing customers take this full history of data that you have or represent?

Gonen Stein

As you mentioned, when we started, I said, “I’m this crazy person starting a non-AI company in an AI world,” and the AI tailwind became absolutely insane. It made data the most important thing that an organization has. If you think about it, models, compute, and everything else are relatively ephemeral, with almost zero switching cost. Those are an important part of the infrastructure for the industry.

But if you’re a company—whether you’re a hotel chain, a food chain, or a technology company—it doesn’t matter. The most valuable thing that you have is actually your data. You’re seeing more and more companies finding this out. Just 2 days ago, you saw Google buy something from bankrupt Spirit Airlines. They didn’t buy airplanes; they bought the data. They bought the data for $10 million because they think it’s very important. They’re using that to train models.

Elad Gil

I think the rumor, too, is that the other bidder on that data set was Mercor, in terms of the bankruptcy bid process. It’s interesting: You had multiple companies in the AI world bidding on a bankrupt airline’s enterprise data set, which is fascinating, yes.

Do you think we’ll be seeing a lot more of that in the future? Do you think we’re going to be seeing these out-of-bankruptcy data buys?

Gonen Stein

We’ve seen it for multiple use cases. That’s what’s really cool about it. You see Mercor and other companies continually trying to buy data. If you’re a tech CEO today, I can tell you that you constantly get questions: “Are you willing to sell your data?” I hear it all over, and it seems that’s going to be a significant trend as we go.

I’m hearing about labs going through Wall Street and trying to buy data from hedge funds to understand how to map and analyze companies. You see data that was accrued throughout the years by companies, which was usually on tapes, usually sitting on a shelf collecting dust, and all of a sudden this becomes very important.

You’re seeing companies realize, first, that what I have today that differentiates me from anyone else is my data. This data is gold, and I can leverage it to get more value for my company and continue building my business as AI comes in and flattens the playing field. It seems that everyone—even large and small companies—can have basically the same playing field, and the only real advantage that a company has today is, of course, its people, but also the data that they’ve accrued, because everyone has access to all of those cool new tools. It’s become a moat.

Elad Gil

People have been saying “data is the new oil” for a long time, and I was always a little skeptical of that statement. But now, with post-training and reinforcement learning, companies like Applied Compute and others are starting to provide services where you can fine-tune models or open-source models against specific data sets. It seems like people are trying to optimize these things for their own use cases.

### Training Agents with Good Data

In the case of something like Spirit Airlines, is it customer support for building an airline app? What do you think they’re actually going to do with this information? Is it something else? Is it internal documents? I’m curious: What is the reinforcement learning for? Is it a customer support agent?

Gonen Stein

If you’re trying to build agents today, you can’t just build them in a lab. You need to train them on new data—on some training data—and it’s very hard to find very good data sets. Harvard just released a legal data set a few days ago, but you don’t find too many good data sets that don’t look like synthetic data and can actually be used to represent the real world.

Spirit Airlines can be used both as an airline company and as a large enterprise, a place where lots of people work: the hierarchy, middle management, top management, and workers working together. When you look at what other public data sets are out there, there aren’t a lot of them. Seriously, I’m speaking with companies and asking, “What kind of data do you have? What do you train on?” I’m looking for real company data, for example. The Enron data is out there in public, and people are actually using that as real data from a company to understand how a company works.

The reason is that it’s very hard to find data that helps you work in the real world. Anytime you see someone building an agent or building a new application, most of them don’t really work. You have to go to the real world and actually interact with real-world companies in order to build something significant.

You can do that when you go to customers. You can buy data and train in-house, so when you first release a product, you don’t have to first interact with customers during that initial interaction. I think you’re going to see more and more of that, both by creating synthetic data in new, innovative ways and by getting existing data, whether it’s real data or somehow masked. Think about it: It contains sensitive information like PII, financial information, and so on and so forth. You want to be able to build real-world things on top of that.

Google, obviously, has been in the travel space for a while, right? They want this type of data. They’re already monetizing it. This allows them to understand it, train on it, and monetize it even further. It’s a unique situation that people obviously want to take advantage of, and I think we’re going to see more and more of that in these situations. Regardless, customers who have existing data want to be able to unlock that data as well.

### Data is the New Oil

Elad Gil

What sort of tooling are you building at Eon to allow people to make use of their data for AI applications? How are you thinking about this problem yourselves, and what sorts of tools are your customers asking for?

Gonen Stein

Let’s go back to the problem statement and why there are so many tools for data and processing. Why do you need new tools? Isn’t it solved already? There have been so many great companies throughout the years, and everyone understands that data is important.

To put it this way, back in the day, every data team could find its own data, decide what projects it had, and get data and do something with it—very tactical. They were using, I don’t know, some great companies: Fivetran, dbt, Monte Carlo, and all the data tools that exist in order to fulfill their tasks.

For some of the data, they didn’t even know it existed. It was locked. Why was it locked? Because there are multiple business-unit owners across the same company.

Let’s say you’re a data team leader in a company, and you’re based in San Francisco—or we’re here in New York—and both of us are different business-unit leaders. Now there’s this thing called AI, and even the boss is playing with ChatGPT. The CEO, the C-suite, the board, and the shareholders understand that AI is real.

Ofir Ehrlich

So they're coming to you, and they tell you, “Elad, we have a lot of data in the organization. We now realize data is the new oil. We can actually activate it with the new tools that we have today. We couldn't do something with the data before, make it useful, and use AI for that because it's valuable for us and because it's cool. What can you do?”

So you say, “Great, I've done this thing before. I know all of those new, cool things that are coming out every day in Silicon Valley. I can just leverage them.” The problem is: Where's the data?

And so you come to us, and we are business-unit leaders. If you even know us—maybe you don't—but let's say that you somehow find your way to me. I'm a leader of a business unit. I probably have data, and somehow you convince me to give you access to my data.

Now, I don't know what data I have. I have a lot of people working for me, and they have data in multiple systems going back 20 years. Some of them are systems that no one really understands, and they contain production data because it includes sensitive information. You know, there's always this server that no one knows what it's doing but is connected to the power, whether virtually or physically, that everyone's afraid to turn off because we don't know what's in there.

So we have all of that, and let's say that somehow I know what's in there. Now I need to bring in engineers and potentially compromise security and compliance and production uptime to extract the data, just to give it to you and store it in a very inefficient manner. It's very hard.

We understood that there's a problem with how this works because we have different incentives. You were tasked with doing that; I'm tasked with making sure my systems work, and I'm tasked with making sure that data is intact. No data is running away. I don't accidentally have the salary of the CEO inside my data, and it's actually going to be used for training or post-training by you.

So we at Eon solve it in a very different way. We can help you—not me, you, the data-team leader—find all the data that's in the organization in a very simple way, understand what it is, classify it, map it, understand the context, layer a semantic layer on top of that, and then be able to continuously bring all the data for me that is relevant without compromising production or security and compliance.

We're actually keeping an audit, and because the data is classified, I know that I'm not accidentally going to share sensitive information with you that you shouldn't have in your data. We can do it in a cost-efficient and performant way. So you can actually do it from all over the place, bring it to you, and actually use it.

Elad Gil

So it sounds like there are 3 or 4 things that you're solving for. One is that you're aggregating lots of historical and current data for people. Number 2 is that you're able to mask personally identifiable information or other fields that they don't necessarily want shared, or set permissions on top of that. And then, third, it sounds like all this can then be exposed to AI models for their uses or applications.

Ofir Ehrlich

Yeah, and the key point to that is that customers already have this data. That's kind of the ironic thing: customers today already have this data. It's kept in their environment in different forms, but it's locked, not accessible, and usually very, very expensive, right?

We're able to take what customers already have, convert it into this new data-foundation format that's stored much more efficiently, provide the mapping, classification, and access control, and connect it into the AI workflows.

Elad Gil

How do you think about security? There's been a lot of news recently about labs where they'll have agents escape sandboxes and do all sorts of things. There may be broader things afoot in terms of why that's happening beyond just the agent capabilities—who knows how these things are set up or configured? Sometimes it's a little bit uncertain whether there's that much thought in how people are approaching these things.

Fundamentally, there's a lot of discussion about AI security. How do you think about that in the context of the enterprise stack—what people should or shouldn't do, and how CISOs should be thinking about all this?

Gonen Stein

Yeah. Up until now, the concerns came from human threats, right? This is not new: customers would come to us and say, “Hey, we were exposed by this ransomware attack.” During our time at AWS, a very large customer was impacted by ransomware. We thought they were completely protected using our technology—the disaster-recovery service that we managed there—and we learned, unfortunately, that the customer thought they were protected.

They weren't protected because they didn't map, classify, and tag their resources properly, so the environment wasn't protected. 60% of the environment was exposed by ransomware. That's one of the reasons why we decided to launch Eon and solve that pain point around human threats such as ransomware.

We need to be able to detect when that happens, look for irregular write patterns and entropy changes and things like that, protect against it, and then also allow customers to recover in a granular fashion and very quickly. What we're seeing now, on steroids, is that the same type of threat is coming from non-human actors—from AI agents that essentially have legitimate access to the environment, with legitimate permissions into such-and-such databases.

All of a sudden—and this now happens very rapidly—a table is dropped. Fortunately for us, it's a very similar methodology in terms of detecting that, protecting against that, and allowing customers to recover, but the velocity of that happening is extreme.

Yeah, something I noted is that 6 months ago, no one would even discuss that with me. But a few months ago, pretty much every person I meet, every leader in a company, tells me either they are afraid of that happening to them or it personally happened to that person who was speaking with me, which is crazy. You see it all over the place.

You see real fear from AI: “I no longer decide what's really running on my data. I no longer understand. I need to be prepared for both external threats, because all the new models make it much easier for attackers to come to me and attack me, but also from the inside, with agents I actually approved running in my environment.”

Ofir Ehrlich

So it's a very, very tricky time. We need to assume breach, whether it's malicious or not, and need to be able to handle it and act accordingly. It's a very weird situation today.

### How Agents Change the Enterprise Stack

Elad Gil

Yeah. How do you think about the broader enterprise stack and agents? The current stack really evolved around people, or humans, asking very defined analytical questions. So we have warehouses, dashboards, ETL pipelines, BI, and agents may behave differently and more dynamically. They may be able to reason over much larger sets of data. They may have access to SAP, SaaS apps, historical data, and a variety of other things, and then act.

What do you think changes in terms of how you store, access, and interact with data in the context of the agentic world? What else do you think needs to change? Do dashboards go away? What shifts?

Gonen Stein

I actually think we'll see more dashboards, because this will be the only way to figure out what they have going on in the world. First, coding agents have started to write most of the code that's running in the world. That's indirect, but agents are also activating other agents, which would activate other agents, and trying to keep track of the non-human identities becomes almost impossible.

There are so many actors inside the organization, and it's very hard for a human to understand the chain of responsibility. This is part of what you're seeing in the proliferation of cybersecurity companies. How many cybersecurity companies do you see in NHI—non-human identity—right now? An infinite amount, and there's a reason for that: it became the number 1, number 2 problem right now.

In addition to that, the second thing is endpoint. You see endpoint security, which looked like it solved the problem. There are so many great companies around it, and just a few years ago, when endpoint was a completely different problem with EDRs—

Now, with everything that's happening, you see people running agents today on their laptops, and the agents are sometimes connected to other networks. They're connected to things like OpenClaw, connected to your WhatsApp, but also to your internal network and to other applications. You see, it's very hard for the VP of IT and for the CISOs to understand what they should do.

On the one hand, they want to—they're being pushed by the board and by the CEO: “Enable AI in my organization now. Don't block me. You can't block me.” On the other hand, it's so scary. I mean, every person—don't even think of technical people; think of non-technical people—building something with, let's say, Lovable or any other software for themselves, putting company data there.

They're not even aware of things like security or compliance, or who is going to use this data, and they're all using all of those new, cool things. So maybe the agents they're building are using other agents, and they're not technical enough to even understand what it means. So it creates a complete set of actors inside an organization, not bound by the rules of the organization and not necessarily running within the premises of the organization, but handling sensitive data—

Ofir Ehrlich

which is the property of the organization, could be exposed to the world, could be incorrect, could be incorrectly used, and becomes a big problem.

Gonen Stein

It's a good thing and a bad thing that everyone inside an organization can become a builder. Whether you're a social media manager, a finance person, or you're in legal or finance—

Ofir Ehrlich

Mhm.

Gonen Stein

So, it's amazing, but we live in very interesting times from that perspective.

### Re-imagining Data Infrastructure

Elad Gil

How much of the existing data infrastructure do you think survives all this? There's all the ETL and data-engineering infrastructure that people have been building and deploying over the last decade. Does that stick around? Does it shift? Does it change? How quickly does all this upend?

Ofir Ehrlich

You see, there's a strong, compelling event to pretty much change everything, because what's called the plumbing today is very limited, and everyone built a solution for their own set of problems. So think about what happens now. Gonen goes downstairs after recording this podcast and really wants coffee. He goes to the store, buys coffee, and puts it on his credit card.

Gonen Stein

Now there's a transaction, and this is written in some database somewhere.

Elad Gil

Okay. Right.

### Ofir Ehrlich and Gonen Stein Introduction

Ofir Ehrlich

Today, someone needs to extract the data, put it somewhere, and that's it. Someone else, at some point, takes this data and processes it in some other way, and that's it. There's no connection with all of that stuff, and every person is very different. They don't have the context of what happened before, and the reason is very simple: It wasn't so important before to have all the context and all the data in the organization, because you could only do with the data the things you really intended to do to begin with. So you had a single purpose in your mind when acting on the data.

Elad Gil

Mhm. Mhm.

Ofir Ehrlich

Today it's very different. Today you understand that if you're able to smartly collect and clean all your data, make sure you store it in an efficient manner, and activate it efficiently, you can let a team go wild with all the data they have. The more data they have, the more high-quality data they have, and the more context they have on that data, the team handling it can create wonders and think of things that were unimaginable.

Let's say that there's one person in an organization who has a list of all the people in New York who love burgers, and another person in the organization has a database of all the people in New York who love pizza. They don't know that they can find a list of all the people in New York who love both burgers and pizza because they didn't work together.

If you use it for data posture and for new capabilities, you can actually do wonders with that. You can start asking your data intelligent questions. You can start using that for your own purposes—just something that you couldn't do before.

You're seeing companies, first, collecting tons more data than before. The amount of data being ingested is absolutely insane, especially compared to earlier. We see trends continuously, both from us and from other companies, in the data. You see data growing out of proportion. There's so much of it, and so much of it is being generated by those new agents. There's a lot of noise in the data. There's a lot of value in the noise as well. So you need tools that are able to both understand data from multiple locations, clean the noise, and make sure all of this data that's being created is actually usable.

Elad Gil

Mhm.

Ofir Ehrlich

And it doesn't apply with the old tools. Some of them were incredible. Fivetran was an incredible company, dbt, and so on and so forth, but they were very specific tools for that purpose. So this creates a very interesting brave new world.

You've seen companies like Databricks, one of the most incredible companies on the planet, in my opinion, looking at, "I have more and more data coming in. I don't necessarily know where it is. I'll help you catalog the data and make use of that," but it's an aftereffect. You already have the data; now you need to process it. But they are reinventing themselves all the time because they understand that more and more data is being generated by agents, and they thought, "The way I see it, if you can't beat them, join them. We'll build our own agents. We'll build our own databases. We want to take charge of how data is being used and how data is being created."

Gonen Stein

And it's completely different from how other people used it just 3 or 5 years ago.

Elad Gil

Yeah. So the goal is really to enable this culture of builders and the culture of agents, with the ability to automatically help them understand what's there, automatically help them ingest the data without having to build manual pipelines for each and every application that's being built, and then also help them maintain control over the data that's created. Makes sense?

Gonen Stein

And with every piece of data that you have, there's another problem. Right now, lots of data is amazing, but it's scattered, which is sort of a problem. Then you need to access it, and you need to pay for storage and, of course, tokens. We're not in the time of token-maxing anymore. We're trying to get value from every token that we have because it becomes more and more and more expensive.

So you want to be very wise. I don't want to say, "Don't pay millions." Pay millions, and even more than that, if you need to, but get the value that you can from actually doing so. So it's very expensive and very lucrative; let's make it as inexpensive as we can.

### Cloud vs. AI Era Shift

Elad Gil

I guess the other thing that you guys have really lived through is the cloud transition. Prior to Eon, you started a company called CloudEndure that was acquired by AWS. At AWS, you really saw that migration from on-premises to the cloud at a huge scale, in terms of that big, generational shift that had happened before this. How would you compare this infrastructure change to what's happening with AI right now? What do you view as the cloud era versus the AI era, and what are the takeaways or lessons that you can apply across them?

Ofir Ehrlich

Yeah, I think, again, it's like that, but on steroids. Even before we sold our last company, CloudEndure, to AWS, we supported similar large-scale enterprise migrations with other hyperscalers, with Azure and GCP, where our product was integrated as an OEM into the console. Very large enterprises were moving thousands, tens of thousands, or hundreds of thousands of servers, and then we saw those modernized further in the cloud.

After we sold to AWS, we did that as part of AWS Application Migration Service. But that's kind of where it ended, and it required a lot of work and effort, both from a technology side as well as from the human side. What we're seeing now in this crazy world of AI and agents is that those transformations are happening way faster, and customers are losing control to a point where that's becoming an inhibitor, not an enabler.

They're stopping. They're pausing because they're afraid that things might break, that data might leak, that IP might leak out, and so they're looking desperately for this level of understanding of what's happening and control.

Gonen Stein

So it became so insane and so fast. One of the reasons is that cloud, in my opinion, is somewhat abstract because it's very hard to explain what it means. Cloud is basically just someone else's computer, but who knows what it is? It's hard to explain the cloud to my grandmother.

AI—everyone understands AI. Everyone lived through the ChatGPT moment when we all learned what AI could do: "Oh my God, this is incredible." So they're getting pushed by C-levels, by CEOs, by the board, and by the shareholders: "Use AI for the business, otherwise we're irrelevant."

### How AI is Changing Companies

You see people doing it both for the value that you get from AI and from the fear that you get from AI. You see new trends. For the first time in many years, you see how companies consume software in a brand-new way. What's happening with forward-deployed engineers used to look like something services companies such as Palantir were doing. No one really understood what it meant, and now everyone's doing that.

Ofir Ehrlich

Now it seems that when you come into a large legacy enterprise, they really want to adopt AI because they have to. The problem is that they don't know how to do it. They understand that their processes are very long; some take a year or 2 or more, but they need to have it now.

The only way they can actually get it deployed and become AI-enabled much faster is by letting strong engineers who understand what they're doing come in with the tools that they've created in top Silicon Valley startups and sometimes larger companies, and transform those organizations. You see them shrinking sales cycles, and you see companies growing really fast because of that.

You also see companies buying really fast, especially the new companies, using product-led growth to buy AI infrastructure really, really fast, which actually helps them build agents, because everyone now wants to build agents. In the past, I was arguing that for the majority of things, PLG doesn't work, especially for dev tools, because the world is very fragmented and people don't want to move so fast.

Now it has become super hot. Look at companies like Cognition, for example, an incredible company that was able to first go through a PLG motion—we use that approach at Eon—and then through the FDE motion, going to banks and saying, "We'll replace the engineering that you don't want to do with our engineers, making you focus on the things that you do want to do."

Gonen Stein

So, leveraging on all fronts, it became super, super, super interesting. The world is changing so much. One other really interesting way that companies are leveraging AI is that they're very slow to adopt AI, but there are really great companies—for example, Long Lake—that say, “Instead of you adopting AI, I know how to do it more efficiently. If I can buy the company and transform that into an AI company, we can all win.”

We can capture arbitrage, make higher margins, and do it more efficiently. This is a really radical new way for those companies to actually start using AI and become more efficient. We speak about this as a revolution, but I think we just started. Most companies still don't use AI; most companies are still at the beginning of this journey.

They all understand that something is happening. They understand that data is important. They understand that their existing processes are somewhat mundane, and they need to do something about it. But it's scary, and you have to do it. So, it's a fascinating thing to see.

Elad Gil

It's a fascinating evolution.

Ofir Ehrlich

What's going on right now is how companies consume AI software, how companies transform into being more modern, and how much we're being pushed to do that. I think that, eventually—I know it's a very wild ride—but I think everyone is going to go through it. The world, in my opinion, is going to be better because of that.

Elad Gil

Amazing. Well, thank you so much for joining me today. Very interesting, wide-ranging conversation on data and AI. I really appreciate it.

Gonen Stein

Thank you. Our pleasure.

Ofir Ehrlich

It was a pleasure.

Elad Gil

All right.

Gonen Stein

Thank you.
