# AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?

The Cognitive Revolution · 2026-10-08 · 88 min · https://www.youtube.com/watch?v=g_K9pqbQJwU

## Transcript

Nathan Labenz

Last weekend, I was at The Curve conference at LightHaven in Berkeley. Prakash asked me and Mike about the differences I noticed.

One of the main observations about AI is that events are developing very quickly. You would have thought, looking at the past or planning a few years ahead, that in 2 years, with new knowledge and discoveries and seeing how much more powerful AI becomes, people would begin to agree on the state of affairs—who was right and who was wrong.

In the wider world, a lot has happened, but less often than I expected. I felt that at The Curve conference, this feeling was present. I think the main thing is that people have reached a certain agreement: AI seems to be becoming truly powerful, and we will have to deal with that reality rather than simply hope it passes or think that if we wait, “the bubble will burst” or “the fever will pass.”

This time, people definitely understood that the situation is becoming serious. As Helen Toner once aptly noted, long-term horizons have become very short.

Another consequence is that the norm in terms of expectations has become something incredible. Even those who had the most modest predictions regarding the capabilities of AI agreed that it would be able to do very much.

So the question regarding capabilities now boils down to this: Will there be a boom in AI research? Will there be fast growth followed by further stabilization? Is this something more than it seems at first glance? We have these superpower tools, but will they really get out of our control?

These are still questions about the possibilities that they put before us, but I think the community has reached some useful conclusions based on what we saw.

Prakash asked if someone from the leading laboratories had spoken about what limitations should be imposed on the players in this area, including themselves. Yes, that was a huge topic for discussion.

Apparently, this is the biggest split in terms of perspectives: between those who are inside the leading companies and those who are outside. We all know that people inside these companies live in the future and have access to information that we do not have.

I had an opportunity to chat with some of them. I heard some during sessions and spoke with others in passing. Everything happened under Chatham House Rules, so I will generalize, but we are talking about founders, leaders, and leading researchers from frontier companies.

I asked some of them, “Why did you come here? You know that you are from—do you get this?” Before spending a few sessions where I spoke as an interviewer or moderator, I posed that question: “What makes this next hour useful for you?”

They repeatedly emphasized that it is important for them to share as much information as possible with the wider community. Of course, they do not disclose trade secrets, but they are genuinely trying to talk about what they see within their companies, what they believe from the inside, what their predictions are, what their plans are, and how we can unite and perhaps overcome some of these limitations.

There were many moments that, in my opinion, deserve attention. One high-ranking leader of a frontier laboratory said that pretraining continues to produce results. Models are becoming more and more powerful at a fundamental level, and it seems like this is not the obvious limit. Scaling laws are working, and may even be improving, because data quality is getting better.

New architectural tricks are appearing. There is a lot of work happening with synthetic data, and it is certainly producing real benefits. You just need more data for pretraining.

What are you doing? You give models a bunch of data and convert that data so that, after training, the model becomes smarter. You seem to be converting reasoning tokens—your test-time compute—back into data for pretraining, which then goes into the next generation of models.

This is one of the ways in which recursive self-improvement is happening. Of course, there are opportunities for architectural improvements, but the main emphasis was on data quality. Now we can spend tokens on existing data to augment it, creating better, cleaner, and more diverse versions of it. We feed all of this back in, and the models get better.

This leader consistently emphasized that pretraining continues to work and that models are becoming extremely powerful. Reinforcement learning is also important.

But he expressed the opinion that the latest cutting-edge models already have that elusive research talent necessary for making paradigm-level breakthroughs. He believes that the problem now is detection, because reinforcement learning does not yet reveal this potential in a reliable way.

They do not quite know how to pull this off from the model, so it is a rare occurrence. That is why you need 10,000 agents to solve a Millennium Prize Problem. Of course, you can check that, too.

Perhaps now you need 1 million agents to eventually come across something that will provide a paradigm shift at the research level in the field of machine learning. But he believes that this capability is embedded in the models, and that they will eventually find a way to reveal it.

Then he said a few things that, in my opinion, were worthwhile news. First, he believes that there is probably a level of intelligence beyond which we should not go.

What? Yes, that is the first time I have heard that from a leader of a leading laboratory. He was not choosing his words in the style of, “This level definitely exists,” or, “I can clearly explain what level this is.” But this was someone whose name and position everyone would recognize, and he said, “I think there is probably a limit that we should not move beyond.”

Does that mean never, or just not now? I think it was a bit ambiguous, but the statement was made quite decisively. Of course, this is the strongest statement I have heard about hard upper-limit restrictions on capabilities, at least for some time.

What confuses me about this statement is that some of the increase in intelligence we observe comes, first, from scaling computational capacity—what has traditionally taken place over the last 1.5 decades.

Second, it comes from improving the incoming data. Input data are improving thanks to higher intelligence, so you get much better-quality input data.

When we say we must not exceed a certain level of intelligence, that automatically means we need to monitor the incoming data, the computational power, and the improvement of the data. If you say that we should not go beyond these limits, this also means that we should not increase computational power or improve the data beyond this level.

Which actually means a lot. It means the end of building up computational capacity, which is quite significant, isn’t it?

Well, I don’t know. I have the feeling that there are many ways of using computational resources, so I had an opportunity to continue this topic.

You know, this combination of pretraining really gives results, and scaling laws may have become even more favorable. Although it is difficult to understand how to take into account all those FLOPs that you spent on creating synthetic data, maybe it would be worth it.

But regarding this issue of pretraining and whether there may be a certain boundary, I just asked how this would be possible to implement. Would you be ready to agree to something like a limitation on the number of FLOPs for the next pretraining cycle? Do you know whether we should say—I don’t know what it would be like.

Nathan Labenz

And again, what exactly should be considered—for example, no more than 10²⁷ FLOPs during the next pretraining cycle, or something like that? I expected this would only be an excuse for discussion, so that he would react by saying that it was a bad idea and perhaps suggest something better. But it turned out completely differently: he just replied, “Yes, I am.” I think that’s quite reasonable. There was no resistance in general.

The details will be extremely important if we go this way, because the more data you enrich and the more FLOPs you spend on it, the easier it becomes, in my opinion, to play a game of hide-and-seek with the calculations. But the mood was essentially: yes, maybe we’ll have to do something similar. And then, of course, I can anticipate your question—and I know you well enough to guess the next joke: What about Elon? What about Zuckerberg?

People said several very interesting things during the conversation. First, they said—and this was not one person’s opinion, but a general conclusion about events—that 2 or 3 Gemini 4s, which have now been announced, may bring us back to 3 leading players. We’ll see how that works. But the main conclusion was that 2 or 3 companies are really expanding the boundaries of what is possible, while the other 2 that we usually counted among the top 5—Elon and Zuckerberg—are actually actively engaged in distillation, whether they know it or not.

The way they are engaged in distillation today is not through direct requests to the Claude API for data acquisition. Rather, it is through industrial developers and model distillers using reinforcement learning (RL), which is actually functioning as a distillation channel. They are all using Claude to create these RL environments, selling advanced models to other developers, and then training their own models in an environment that exists only because Claude was smart enough to create it.

And that’s exactly it: they are constantly taking advantage of capabilities that truly advanced companies have developed. So there really wasn’t much discussion about competition from outside the 3 leading companies. One person went so far as to say this—and it was another founder or the head of one of these companies: “Yes, they’ve held out. Of course, there is a gap, and it remains permanent, but they’ve been able to maintain a stable distance behind the leaders.”

This contrasts with the old forecasts from Anthropic’s presentations to attract investment, which I talk about constantly. They said that, probably sometime in 2026, the companies training the best models would get so far ahead that no one would be able to catch up with them. Of course, that hasn’t happened. But the general conclusion from all of this is that if they hadn’t produced models, and if people couldn’t use these various methods of direct and indirect distillation, they believe that this probably would have happened.

Other companies simply don’t have enough time, except through the leakage of intelligence by all these various methods, including industrial RL, which gradually finds ways to transfer capabilities from advanced models to their competitors.

### Notes from The Curve (Part 2)

Prakash edited my material about Elon’s strategy, and then I took up the timelines. It seems to me that, on Elon’s side, the main focus is on building infrastructure. Elon is betting that infrastructure is more important than models and that he can catch up through the development of his infrastructure.

He is structuring things around 2029, 2030, and 2031. Things like Terafab will appear much later, and the satellites from SpaceX are a matter for the 2030s. So he is counting on a much longer period of time. Anthropic and OpenAI, in my opinion, are oriented toward the next 2 or 3 years and view 2028–2029 as the period when AGI and RSI will emerge.

I had that impression before, although it was probably biased because these are short-termists. But again, the short-termists include the leaders, the leading researchers, and the founders of these companies. They talked more about next year. There was really something like, “2026—the critical period starts now. RSI is almost at the doorstep.”

There is absolutely no full agreement about how quickly it will happen or when it will reach its peak. But there is a common feeling that the decisions made in the coming months, and the way computational capacity is used next year, could be truly critical. Fairly speaking, they didn’t talk much about 2028, which is also quite strange.

Then we returned to computational capacity and what the restriction would mean for their development. Would they need a restriction on computation, or would it mean restricting the scale of pretraining or other computational resources that could serve as a mechanism for containing the development of computational infrastructure? I would say no.

One more question I had the opportunity to ask was about something everyone has probably already heard Jensen Huang discuss on the Ezra Klein podcast. He noted that, in order to grow, these companies must change the ratio of where they spend their resources. He said, “At NVIDIA, we spend about 20% of our effort on chip design and 80% on verification, validation, testing every edge case, so that it will be durable, reliable, and meet all the other requirements—not only the initial design.”

His point was that they had obviously worked very hard and invested all their effort to make their models powerful enough for practical use. Congratulations—you did it. Now we are entering an era where you will have to make them safe, reliable, worthy of trust, and so on.

I had the opportunity to ask one of these people, “What do you think about Jensen’s opinion?” The answer was something like, “Yes, that’s quite right.” Of course, without knowing exactly what the indicators will be, there is a probability that the majority of computational capacity in the future will go toward all kinds of AI safety work, such as monitoring or chain-of-thought monitoring.

People are becoming increasingly pessimistic about chain-of-thought monitoring. But there is still internal activation monitoring, so to speak. They are already spending a significant percentage of their computational resources on monitoring, and it seems that this part could grow substantially.

Another thing they are actively spending computational resources on right now is fixing reinforcement-learning environments. At this point, everyone understands that if you have a sloppy RL environment that encourages reward hacking, you will get a lot of reward hacking. Now they use these models to test RL environments, saying, “Break this environment.”

That is now the direct task. They do it. They find all the places and ways in which the environments can be broken, and then they correct them. I’m sure they are getting rid of some environments that are simply fundamentally imperfect, or something similar.

Gradually, they reduce the level of these flaws. That means they also reduce the extent to which cheating is rewarded, which in turn means less deception from models when they start working.

Nathan Labenz

My impression is that, although I think this remains an open research question, they see a clear enough connection: the more you clean up reinforcement-learning environments, the lower the rates of deception you get later. They are trying to achieve that result; of course, they would like to have environments that are impossible to break. It will be difficult. It seems we have not yet reached the point where we have a scaling law, but we see our own proto-law of scaling. It seems that you can spend a lot of computational resources reducing this level of deception and get better behavior as a result.

I have heard many times that they are waiting for delays due to security and alignment issues. But at least for the next 1 or 2 generations, there are still many easy solutions in this area, particularly just correcting RL environments. So it is to be expected that everything will develop pretty quickly. It seems they will probably be able to release the next few products without much trouble, even despite what they themselves think are limited coordination issues. But after that, they say, “Yes, then we guarantee nothing.” It is very difficult to say what will happen 1 or 2 generations from now.

Next in the program, I got the perspective of a manager of research talent on a new record in speedrunning nanoGPT. This is a race to teach a small language model to achieve a target quality as quickly as possible. The record belongs to a company called Hyperstation. In short, it seems we moved from approximately 70 seconds in this classic benchmark, which people have been talking about and working on for quite a long time. The idea is to teach a small model to reach a certain loss as quickly as possible, in real time.

This result cut the time nearly in half—more than, in my opinion, the previous 4–5 improvements combined. And this result was achieved. One thing I think is interesting about this specific case, which seems to have started this wave of significant time reductions, is what the author noted: AI did not play a significant role.

As that high-ranking frontier-lab leader at the company I mentioned earlier said, the perspective is that the model may eventually generate ideas of that quality, but there is an elicitation problem from knowledge, and in this case they did not rely much on AI. Of course, they received a lot of help, but the main insights mostly belonged to the person, not the AI, in this specific case. This company is also involved in large-scale training. They emphasized in their publications that the optimizer that provided significant—although I do not think it was the majority—of the time savings they discussed was actually not as effective as the one they use internally.

The last expectation I heard from leading companies regarding science is AI for science. That is where, in their opinion, we are heading in the near future. One question was whether, if our AI is superhuman only at some things, such as tasks that are subject to verification, it will be superhuman at everything. The middle-ground view, which I think is completely plausible, is that we will see superhuman results in everything that worries us and that we are really ready to invest in. But that does not mean complete generalization across every field.

Even in areas that are not considered easily verifiable, they believe that when they focus on them, license the necessary data, use a lot of compute to augment that data, create synthetic versions, and add some reinforcement learning, this entire approach is generally effective for solving any problem they truly decide to concentrate on. Therefore, AI for science is next. It is expected that we will see this soon—one expert even said that in no more than a year, it will be impossible to do advanced science without a significant role for AI.

The next morning, I added another conclusion from the graph. I do not think we talked about it yesterday, but the conclusion from the weekend discussion—one of the minor directions I discussed—was that there will be government intervention. Of course, we have already seen something. But the idea that Congress would never act used to be treated almost like an axiom.

Now I am hearing new sentiments: after these elections, especially if the Democrats take both chambers, that could change. You could even see how many Republicans join the Democrats, because deep down they want to do something. They do not really want to go against the president, and so far few people are willing to take that risk. Besides, they are getting a lot of money from interests connected to data centers, but all of this could be reconsidered and regrouped after the midterm elections. We may actually see Congress act much faster than expected by many of those who have gotten used to the idea that this will never happen or that it will remain stagnant forever. So I really think political reaction is definitely something we need to consider.

### The memory wall

Part two, the memory wall. Thomas Sohmers is the co-founder of Positron, which develops chips for AI models and recently raised $875 million. I asked Tom about whether rapidly rising memory prices had changed the economics of the business.

Thomas Sohmers

I was looking at the price list—the offer we received a year and a week ago for memory—and since then it has gone up by 4.5 times. I really hoped that was not true, but I would be surprised if the price increased another 2 times next year. But I would say that the main reason everyone feels the burden of this appreciation is that you get 4–5 times more value compared with last year. In fact, compared with last year, the capabilities of a model that uses the same number of gigabytes of memory have increased much more than 5 times.

Nathan Labenz

I asked Tom about the trade-offs involved in using standard memory instead of high-bandwidth memory.

Thomas Sohmers

At the beginning of Positron, I thought memory would become a bottleneck in the future, although people can argue about energy consumption and other parts of the infrastructure. My fundamental belief was that near-term progress is not going to stop at either scaling or getting more value from increases in model size. This is draining memory resources on the one hand, but on the other hand, I think the main constraint on AI applications today is context length and the ability to hold more context for a user, scaling this to a much larger audience.

Today, I would say the driver of so-called users, or individual sessions, is simply having more agents. If 3 or 4 months ago I had, on average, 2–4 agents constantly running in the background, now there are already 15 or 20 of them.

If you multiply this by the number of people who use AI, the total number of simultaneous, completely separate contexts is growing very rapidly. That led us to say, “Okay, we have to use commodity memory, because this is the only solution that will allow us to scale and remain economically beneficial.” When we determined that this was the main limitation in our architectural design, we had to search for smart and innovative solutions.

The 2 main aspects of what we demonstrated in our first-generation product are, first, the ability to achieve extraordinarily high memory-bandwidth utilization from an architectural standpoint. Although NVIDIA GPUs, on average, provide 30% to 40% memory-bandwidth utilization during direct execution of the decoding pass in transformer models, despite their declared 8 TB/s of theoretical memory bandwidth—for example, in the B300—you really get only about 2 TB/s of actual bandwidth. This boils down to a whole series of architectural details in GPUs. This is due to data-reuse schemes that are not actually present in transformers, because their architecture is oriented toward training and other tasks.

Therefore, in our first-generation product, we managed to achieve and sustain 93% of theoretical memory bandwidth. The theoretical and actual figures are effectively equalized, and this means a huge, threefold improvement in results. But in order to truly use the advantages of models that will become even larger, we have to scale to a much larger number of channels. So, we collaborated with Credo Semiconductor and developed a chiplet-based memory solution that allows us to go beyond the maximum amount of LPDDR—the type of memory used in phones, laptops, and so on.

The maximum number of channels found in other products is approximately 12 to 16, whereas we provide up to 72 LPDDR5X channels through this solution with separate memory chiplets. The next chip from Positron is called Asimov.

Nathan Labenz

Prakash asked, “What disappears in the network interactions and coordination when the whole model fits on one chip?”

Thomas Sohmers

We have up to 2.3 TB of memory capacity per chip. If we compare that with the B300 supplied today, where the limit is 288 GB, according to analyst reports, NVIDIA actually reduces memory capacity in the next generation because of cost and other factors. In Rubin Ultra, they will have only 192 GB. Therefore, even if we take 288 GB as the current reference point, we have 8 times greater memory capacity on-chip.

This means that at the one-chip level, we can now scale what previously, from a memory standpoint, would have required 8 GPUs. We can do it on 1 device. It’s not just about saving money on silicon. As you noted, every time it’s necessary to scale the system across multiple devices, there is corresponding overhead. You have all-gather and all-reduce operations that must be performed for each layer and for each matrix multiplication during data distribution between these devices.

Nathan Labenz

Next, I asked Tom about looping—when the model runs the same layers several times for each token—and why that is generally attractive. It was very interesting to watch the path from the article about looped transformers to the rumor that this is one of the great achievements of GPT-6 Astra.

Thomas Sohmers

A simple repetition of the forward pass gives you an improvement. What’s crazy is that there are works that, I think, were produced only on the periphery of the open-source developer community 2 or 3 years ago, when people took, for example, Llama 70B and simply duplicated its layers. They effectively said, “Okay, I’ll just repeat these sets of matrix multiplications successively,” turning it into a model with 100 billion parameters. You got better results, even though these were repetitions of the same matrix multiplications.

There’s a certain element of, “Imagine what you could achieve if you trained it to work just like that.”

### Outro

Nathan Labenz

That’s right. If you already have a very high confidence about what the next token will be, and several layers have confirmed it, you can decide to exit earlier. Or, if you’re really not sure for some reason, you can find the layers in the network that are most likely to increase that probability. That’s the same function as the MoE router, which determines which expert to choose for each individual token.

You can extend this concept: if you’re really not sure and there’s a very large set of equivalent probabilities for the next token, and you’re three-quarters of the way through the layers, do you want to go back and determine, “Okay, do we really need another set of experts for this?” I wouldn’t be surprised if these things are already being implemented at large scale in large-model applications.

Can you explain the hardware connection a little more deeply? I agree with certain assumptions at a fundamental level. Why use cycles at all? I think this is related to memory bandwidth, right? If I can hold the same weights on-chip, I don’t have to move data back and forth as often.

Thomas Sohmers

Not exactly, because at inference time it’s usually assumed that all the weights will be local in DRAM; you just distribute them among a certain number of devices. You get some optimization, but considering the dimensions of the experts and on-chip caches, the ability to reuse them is limited. By the time you complete the process, you’re already working with layers that repeatedly exceed the amount of memory contained in on-chip SRAM.

From a hardware point of view, using loops is more about the economics of memory capacity than about bandwidth. Let’s say you trained 2 models on the same dataset, but 1 uses loops, so it has 10 layers instead of 20. Model B has 20 layers. Simplifying, this 20-layer model will be twice as large as the 10-layer version.

If you see that the 10-layer version with loops gives 95% of the quality of the results and is half the size, you would probably deploy it. You save specifically on memory capacity, because when executing the second group, you still have to perform the same memory accesses and the same number of computations. In the Model A and Model B scenarios, you have to move the same number of bytes and perform the same number of floating-point operations.

The 20-layer Model B actually has unique weights for the second group of 10 layers that need to be processed. This means that you need twice as much memory.

### AI in chip design

Nathan Labenz

Now let’s move on to how AI is changing the development of chips themselves. In a recent, famous interview, Ezra Klein asked Jensen Huang, and Jensen said that NVIDIA spends about 20% of its time and energy on design and 80% on verification, validation, software, long-term service, reliability, and so on. What figures do you have in this respect?

Thomas Sohmers

For us, I would say it’s a bit different, because we’re starting from scratch. There’s much more basic-level design work and more opportunities that we have to address from the ground up. For our first completely proprietary Asimov silicon chip, it will be close to 60% focused on design and 40% on verification.

### Beyond verifiable tasks

I would say that probably the most amazing thing, which in my opinion would have been impossible 3 or 6 months ago and became possible only with GPT-6 Astra and Opus 5.5, which we actively use for this task, is that for verification we apply Cadence Palladium emulation systems. These are huge racks filled with special ASICs designed exclusively to perform silicon emulation at the gate level.

The craziest thing about Astra, Fire, and Opus 5.5 is that the agents themselves were able to iterate fully in a closed loop and test it. I think people thought that was madness, but we gave these very powerful agents full access to our internal infrastructure and said, “Here is the Palladium part.”

I very much doubt that these models had Palladium documentation in their pretraining data, especially because a lot of the documentation consists of new software updates. But the fact that the agent found the documentation—and when we looked at the agents’ traces, logical conclusions, and tool calls, it had read the full documentation, effectively compressed it independently, read PDF files, created its own Markdown files with tips on what to do, and then built the whole test infrastructure in its own scripts to access it—was simply amazing.

While you’re doing the usual RTL programming and modeling, it works at a speed of about 10 hertz in our cases. If you remove many of the elements and settings that slow down the process, we can run full-chip emulation at a speed of about 500 kilohertz. That gives you an acceleration of several orders of magnitude because of this gate-level hardware.

This brings us closer to an RSI cycle. Right now, it’s focused only on running test programs, searching for cases where they fail, and then writing reports that are checked by other agents and the people involved in the process. When this really started working 6 weeks ago, we achieved this at the beginning of September, when Astra 6 came out. It was a huge, jump-like improvement compared with GPT-5.6, which could not independently perform a full closed loop. It still required human participation at different stages.

Nathan Labenz

I asked how Positron’s spending on AI correlates with personnel expenses.

Thomas Sohmers

Regarding token costs, they very quickly became our largest category of expenses unrelated to production.

Nathan Labenz

More than people’s salaries?

Thomas Sohmers

Yes, very recently. It’s interesting that this exceeded people’s salaries and then decreased again.

I’ll explain in a moment. Yes, our expenses for tokens have grown massively. Six months ago, this was equivalent to one employee; in June, it had already begun to concern several employees. We reached our peak in the days after the release of GPT-6, really trying to expand its boundaries and so on. We spent over $100,000 per day on tokens.

### Sponsors: Parallel | Claude

There was a bit of a lull in the following weeks because we optimized our processes and no longer tried to conduct so many parallel experiments. But I would say that the main factor—the lifesaving factor—was the release of Opus 5.5. In many of our tasks—not all of them—it worked better than Astra and cost a quarter of the price.

We didn’t tell anyone to reduce their expenses or do anything like that, even despite watching this exponential increase in token costs. It was only related to the fact that we both felt we were getting good profit from it, so we weren’t going to limit it. Personally, my philosophy is that I really don’t want someone to perform a real development task, software or hardware, using something less than a frontier model. I don’t care if it’s 10 times cheaper per token; it’s just not worth the expense.

Nathan Labenz

Have you seen AI offer unexpected ideas in microcircuit design, instead of just following the rules in the manual?

Thomas Sohmers

Not yet. I would say that’s the most disappointing moment of all. Apparently, I think all of this is incredibly impressive, and I think that we should continue this exponential trend. To some extent, I’m proud that AI still hasn’t figured out some of our tricks in what we do.

It questioned certain decisions, considering them unsuccessful, until I explained why everything was arranged just right. It runs tests, sees the result, and understands, “Oh, that’s why you don’t do this in the traditional way, as in systolic arrays.”

Do I think it will continue like this over the next 6 months? Not assured. But partly, I think the agent will not be able to instantly offer the best idea. What seems surprising to me about recent events, a few weeks after we gained access to these models, is that they can iterate very quickly and conduct experiments independently.

I wouldn’t be surprised if it reaches the same conclusions that we did, or maybe does something better than us, just because it will iterate so quickly and thoroughly through so many different design options. Let’s remember, for example, the story of how OpenAI solved the Navier–Stokes equation. That’s the same approach: if you take 10,000 agents and spend millions of person-years of effort on it, solutions will eventually be found. It’s brute force. It’s expensive, but I think it’s quite a sound strategy, which only recently became possible.

Nathan Labenz

There isn’t much left in the category of “AI can’t do what I do.” So congratulations on having your place, while it’s still yours.

Thomas Sohmers

Yes, at least for now. For a few weeks.

### Software after agents

Nathan Labenz

Part 3: software development after the emergence of agents. swyx hosts the podcast Latent Space and conferences for AI engineers. I asked him, “How are you feeling now as a software developer? Do you feel stronger or under threat?”

swyx

Yes, I think students are a little worried. But besides this, if you’re good at using your own AI tools, you’re in demand now more than ever, because your value is higher than ever before. This is definitely one of those Jevons paradoxes: when the cost of creating software decreases, demand grows significantly. But there’s a very relevant, specific demand for people who can manage agents productively for writing code, instead of producing a bunch of junk.

I’ve been in such a situation. I’m an engineer and an employer of engineers, as well as of people who aren’t engineers but are engaged in coding. You know, a phrase that I’ve been liking lately is this: “I don’t want to pay for someone else’s language-model psychosis.” I can pay for my own psychosis. That’s normal. But when you work for me, and I’m paying for your tokens, you have to create something really thoughtful.

Now I have 2 or 3 employees who are actually undergoing performance reviews because they just give me garbage from Claude. And this is very bad for them. They don’t understand it. They don’t see it. They ask, “What do you mean?” I think that’s quite normal. And I tell them, “Well, you don’t create any value. I can just give a request to Claude; I don’t need you.”

Nathan Labenz

Do years of programming experience matter when you hire, or are product understanding and high standards more important?

swyx

I don’t think it’s important, except, let’s say, for roles in security and everything related to backend and scalability. I just recorded an interview with the founders of Supabase, who scale Postgres to levels we’ve never seen before. Good luck trying to find someone who scales Postgres with a fully self-directed sharding solution that works at the scale of YouTube. There is only 1 person in the world who can do this, and they’ve already hired him.

You’re better off making sure that you don’t lose data, cause downtime, and scale economically, because agents consume databases at approximately a 500-to-1 ratio compared with human developers. So, besides that, you can write code, and then it all comes down to taste, not the duration of your experience. In fact, sometimes long experience works against you because you have an established way of working and don’t understand how to work with more than 1 agent simultaneously. Now you should find it relatively easy to deal with 5 to 10 simultaneous tasks, and you’re definitely becoming more of a manager than an individual performer.

I really believe that people who look at the data, and not the code, are more valued today. That is, the ability to say, “Here are the logs, here are the traces, here is the diagram, here are the inputs and source data,” record this, and turn it into an assessment. All this is basically the fusion of AI engineering and ML engineering, which is happening now, and people need to raise their qualifications.

I think that a person who has, say, 10–20 years of experience in software development really checks each line and tries to make sure that each line makes sense. Whereas now we just need to have meaningful separate modules. I can afford a certain amount of negligence because this helps me move faster, as long as I control these shortcomings within systems that I completely understand.

You’re wrong when you have too many modules, too many “black boxes” in which you don’t even understand what’s happening, and the code itself also becomes confusing. Therefore, I created my own competitor of sorts to Slack, and I noticed an error: messages weren’t loading. I refreshed the page—messages loaded. I updated again—messages weren’t loading. And this looks like, “Well, that’s it. It’s exactly the same code. What the hell is happening?”

It turned out that there were 2 code-execution paths, and a race condition occurred. Why? Because 2 different coding agents were working on them at different times, and probably each of them just did their own thing. A person would never do that. Agents can sometimes do this because something falls out of the context window. But you need to control the module and everything inside it, even if it’s a “black box.”

Nathan Labenz

What parts of the stack are badly adapted for agents now that they’ve become the main users of the internet?

swyx

Computational power. Everyone is aware of the GPU shortage. Everyone knows about the memory deficit. But CPU shortages—that’s what almost everyone I talk to now reports. So what does this mean? It means that we need more flexible computing. That’s a universal term for this: more diverse serverless forms of architecture, where everything is very ephemeral. You can pause and restore execution because agents need time for LLM calls or network requests, and so on. But in essence, these are long-term conditions that can be restored.

But I think that, apart from this, you start to delve into 2 things that I really think about: network throughput, where latency starts to play a really important role, and optimizing where your agent interacts, which region it’s in, how many calculations should be performed in the cloud, and how much should happen locally or on peripheral devices.

This definitely often happens with the robotics companies we communicate with. And, by the way, it’s not worth taking it only as a problem for robotics. Robotics is just the first messenger of what you will ultimately be doing. If you consume or perform inferences at their scale, this becomes a broader problem.

And finally, for me, for a very long time, at a very personal level, this has been about bandwidth capacity. How many tokens per second can I get? That’s true for a separate call, in total for all my tasks, and then at the scale of the entire company, where I have a lot of people managing all these tasks.

### Kill my SaaS

Nathan Labenz

swyx also conducts “Kill My SaaS”—a reward for anyone who can replace a software subscription with something the company didn’t want to pay for. I asked him about it.

swyx

In fact, you could call it a reward for destroying mid-market SaaS that shouldn’t exist. It’s like half or a third of a salary for software that I don’t own and that no one loves to use. If I’m going to spend $40,000 on this SaaS subscription, I can spend it on tokens, and that will give me a huge number of tokens.

We received so many applications that we had to check all of them. And therefore, every one of them is a problem. Since we’re no longer only checking correspondence between 2 or 3 requirements, we check the entire UX—the whole process—from 3 different points of view: organizer, participant, sponsor, or speaker.

If you offer a large reward, such as $10,000, you will receive many applications, because people like to write code for fun, and many of them will be low quality. They just say, “Claude, don’t be shy; do it,” and then assume the program has made mistakes, so they send the result. Therefore, the load on quality checking is unbalanced, because they don’t spend any thought on it. They just throw it in there, and you have no idea whether it’s low quality.

In general, though, it was very successful. At first, the team was one of the most hostile-to-AI teams in the AI field. Of course, what is my business? I hire professionals from event organizations that are very old-school. They do everything in spreadsheets, and they work at trade shows. They have to worry about booth locations, for example, physically on-site. This is not some glamorous, high-tech work.

They’re very suspicious of everything new, AI, and technology. Therefore, I hire these people and make them work with AI. At first they said, “We’ll never use this AI. I want to work with proven things that are already working at Microsoft and other companies.” Then they saw the quality of the submitted results and looked at their existing platform. They were like, “Yeah, okay, we’re moving on.” It’s that initial barrier, when you need to prove that this is possible at all.

After that, there’s constant support. One of the advantages of my partnership with Cognition is that I can give them access to Devin for code changes. So anything they don’t like, they can change, and the result appears within 1–2 hours. They’ve never had that before.

To give you an idea of how we work with SaaS companies now, when we ask for changes, they say, “Okay, that sounds cool. It’s in our plan for the 3rd quarter.” We have no confidence that it will actually be done in the 3rd quarter. We asked for one thing in the 1st quarter, and it was only implemented just now.

At the end of the day, if you mostly create CRUD applications, I would say there are still a lot of UX issues. If you look at SWE-bench and say, “Wow, we have 90 points on SWE-bench,” you’re not looking at the right thing, my friend. If you haven’t tried to write code as quickly as we can, you won’t understand how much the model still struggles with this.

Nathan Labenz

Then I asked swyx about the concerns. It’s hard for me to imagine that we won’t collect all these low-hanging fruit, after which a lot of companies will go bankrupt. I’m also thinking about layoffs, like at Meta—this is another alarm signal. Or will we see mass big-tech layoffs? How long can this feast last?

swyx

The simple answer is this: if you look at software development as a limited resource, then you’ll come to that conclusion. But if you look at it from the position of competitors or the total volume of the market, it’s spreadsheets.

Any spreadsheet created by anyone with productivity tools can be converted into specialized software, optimized, and given a wonderful interface and automation. Then the demand for that software will become much larger. Moreover, phone calls, emails that go back and forth, and ultimately personal meetings can all gradually be converted into more and more specialized software, hardware, and models.

Essentially, we’re just growing toward the long tail of the market. Why, for example, is Salesforce so huge? Because people can configure Salesforce for themselves. But at a certain point, they’ll stop paying $300,000 per year for a basic subscription and spend $30,000 on creating their own CRM. That will happen often—more than often enough. They’ll be happier because they can change it however they want.

There are so many customization options available to people. We don’t even have as many possibilities for customization in our software as we have in our clothes. What kind of nonsense is that? Imagine: clothes have existed longer than computers, but there are so many opportunities for personalization, so much demand, and so many different brands and varieties.

In reality, we don’t compare benchmarks when we buy clothes. It’s simple: this is the style I like. I think software engineers created a situation where we have Kate Spade, Louis Vuitton, or Coach in the software world. We haven’t even reached the point where we evaluate functionality or price. We just think, “What does the brand feel like, and which brand do I identify with?” That is real commodity software: it stops differing in content and starts differing by vibe and feel. There are a lot of them now.

The reason I talked about scaling personal, per-user token capacity versus scaling the bandwidth of the entire organization is that the real limitation is the total token bandwidth of humanity. It’s literally the amount of silicon we can produce, and that will dictate prices, availability, limits, and everything else. So the serious players are focused only on that market, while the rest of us are fighting for crumbs within the limited pie we already have.

Prakash asked about Jev, the fast decision-making model from Type-Safe AI, and whether there’s a new API for the OpenAI solutions announced at DevDay—something like that. I don’t use it much myself. I’ve been friends with Diogo for 4 years before the launch, so I insisted that ours would be the first podcast where he talked about Jev. It’s probably our most popular podcast in this area.

Secondly, we also participated in API testing for OpenAI’s solutions. We have a podcast about those APIs as well, although this isn’t exactly that. This is essentially Llama with a slightly different level of abstraction on top of it.

I think many of the people I communicate with agree that most of the discussion on Twitter is probably exaggerated or meaningless, because you could use any other small model for this. But because Jev is popular now, everyone is creating content about it. Another idea is that it’s just another classifier, and classifiers could always be trained; it’s simply fashionable now.

People are missing the point of why they don’t call it that—why Jev or Google doesn’t mention that this is a System 1 model. They call it a System 1 model because it’s intended to be integrated into your software and work in the background. It should be as imperceptible as an “if” statement inside your code.

But people aren’t doing that. They’re replacing every structural component with Jev, and they’ll probably encounter problems because they’re not exercising its intelligence. If you notice, in most of these cases, nobody is assessing whether the model is actually acting intelligently. Is it planning something? Nothing of the kind. The main question is whether it can quickly make a decision. It could always make a decision quickly.

It seems to me there’s a certain nuance that they may be missing. It doesn’t matter, because I believe creating a category was so successful that it’s generally good for the industry.

### Who owns AI risks

Nathan Labenz

Part four: Who carries responsibility for risks? Evan Miyazono manages Atlas Ignota, a nonprofit organization that looks for significant AI risks without a clear owner, develops measures, and finds someone to implement them. That’s the idea. I keep repeating: when intelligence becomes cheap, coordination becomes expensive. Having Schelling points around who coordinates, being that point, or creating it is very useful.

I asked Evan where the biggest gaps are now.

Evan Miyazono

Here are 2 things I’ve personally spent most of my time on lately: what we can or should do to protect critical infrastructure, especially from random swarms of agents, and how to address the malicious use of open-weight models.

What would happen if, instead of OpenAI, someone accidentally attacked Hugging Face, the companies behind DeepSeek, or any others? What if servers at the State Department were attacked by accident? How would we find out that it was an accident? What would we do in response? What would you do, or what would we do, if we saw malicious actors critically attacking infrastructure with OpenAI servers? How would we know that it was them and not someone else? How quickly could we stop it? What about different levels of escalation?

One solution that could be added is using DNS for cryptographic confirmation of who provides the outputs—inference. I think that would be very useful and would be a great addition to DNS for people. If your agent communicates with mine, they can confirm that this agent is really mine, and your agent can check that it’s Evan Miyazono’s agent, signed by a certain key.

I think there are faster ways to implement this if you don’t try to turn it into a startup that brings in profit. Therefore, this looks like a public good.

Nathan Labenz

Later, Evan brought up the problem of visibility. How do people generally find out that there’s a better tool or practice? It turned into a conversation about how our own agents can help us connect.

### Notes from The Curve (Part 1)

Evan Miyazono

Listen, Nathan, I’m sure you have wonderful infrastructure solutions that would be useful to me too. I don’t know enough about them to even ask what they are. I have an application that just takes screenshots of everything that isn’t a video conference, passes them through a local OCR model, and then sends the results together with my weekly review, with the aim of analyzing: What am I trying to do? What is it worth to me? Would it be possible to do? What do I need to automate? What is it worth changing in my actions?

I can imagine that the exchange of such information between colleagues could help identify all sorts of interesting synergies.

Nathan Labenz

Yes, that’s interesting. When we started this show, we looked in particular at how to experiment with recursive self-improvement. We were thinking: Can we create a show together, live on air, without any employees, and will we succeed through our iterations in reaching something that really works?

Now I feel that perhaps the next stage is to perceive all this as a swarm—a friendly swarm of people who can share their best ideas. Perhaps the result of this, because I will ask Claude to do it, or because Claude, I hope, will do it for me after seeing a transcript of these conversations, will be an attempt to evaluate the best ideas that I have and announce them, or, let’s say, introduce them—not necessarily all of them, although why not all of them?—but definitely to my friends, people I know, and people I really want to help.

It seems quite likely to me that if I noticed something in his work, or if Kochi said, “Evan, you’ve got this perfectly wrong,” that would be very useful. I could contact the network and ask: “Who in my network understands this?” “Who should I talk to?” “Is this something to talk about?” “Who is most likely to be able to solve this problem best?”

It’s like the invention of social networks in a new way, or the revival of a guild, or something like that. I think it might be very, very useful. I don’t have it yet. It is quite clear that society will begin to restructure around such things.

I have one request. If you create a good way for people and their agents to find each other safely and share whatever works, please write to me.

Near the end, I asked Evan to tell us about the current status of formal verification, now that models have become very skilled in mathematics. Evan has helped found companies in this field, so he has skin in the game.

Evan Miyazono

There are significant differences between proving properties and theorems in mathematics and proving properties of software. One is that, in mathematics, the number of theorems that can be important is significantly smaller than in software, while software is much more versatile.

### Workplace training data

It seems there are many projects where ambitious formal methods—at least many individual people, and I’m referring in particular to Oath, Theorem, and several others—are solving problems that previously would have taken 5 years and $100 million. Now this is limited only by tokens and the number of people who can effectively manage agents to perform these tasks.

There are many things for which I expect that, in 3–6 months, we will see the same “shock and awe” effect that we’re now observing in mathematics, particularly in software verification and increasing its reliability.

But I also think that there is a noticeable limit, a very high bar, which in my opinion is significantly better understood and more difficult to achieve than in mathematics. You could prove that a computer system has mathematical guarantees, from the separation between different processes all the way down to transistor operation. Theoretically, you could even prove it in physical simulations and show that, for this stack of software and hardware, there is no vulnerability such as Rowhammer.

You could also prove that the system does not suffer catastrophic failure. Or, if that does happen, for example through some random event, we will be able to restore everything and roll back any kind of software state.

I think there are many things that 1 year ago seemed absolutely impossible that perhaps have become real today, and soon will certainly be available. That will lead to significant improvements in software development and maintenance.

Nathan Labenz

Part 5: Where does training data come from?

Edward Hu heads the AI modeling effort at Mercor, which works with professionals to create training data, environments, and benchmarks for AI laboratories. He also led development of LoRA, a widely used lightweight method for fine-tuning models. Mercor’s Apex Agents Benchmark provided models with simulated workplace environments, email, and chats. Edward described the next step.

Edward Hu

We have a project that is almost finished where we actually take this a step further by buying real company data. Often, companies have realistic applications and environments as part of the data we have. For example, from their Salesforce, Slack, Figma, and all of these applications, often with gigabytes or even terabytes of data, we create realistic tasks.

So far in this area, we’ve been more or less focused on tasks. We often start with, “Okay, here is the task for you. It needs to be done. Maybe this is the creation of a staff list.” That is one of the formats, and this is exactly the format we publicly presented, and most people worked with it.

But this is only one of the things we work on, and we ask ourselves whether there is a better format. What will the task format be in the future, when we have this set of data?

When we work in a company, it’s not always easy for us to give someone a ready-made task. We are given a position and an organizational chart. And, of course, with all the data in companies, people will contact us with requests. We will have to look for others to get information. There will be managers, people responsible for resources, and conflicts that need to be solved.

So we believe that this will increasingly become a format of the future, and we’ll tell you more in the coming months.

Nathan Labenz

I asked Edward how supervised fine-tuning relates to reinforcement learning, and how companies will train models in the future.

Edward Hu

One thought experiment is that if we take self-distillation with SFT, this is actually more or less a step toward RL. If you repeat this process, you’ll get a result quite similar to RL.

I really think that in the future, RL in the form they’re using today—with this very difficult infrastructure, very expensive deployments, and relatively ineffective updates, because with all these deployments, especially long-lasting ones, we only get one number at the end—such types of training will not be so common.

I don’t think that every company in the world will be able to do these big RL launches. Therefore, I believe that SFT, especially the combination of results from several teachers, will become a key part of how businesses adjust their models in the future. We also invest in research in this direction.

Nathan Labenz

Let’s go back to reward-hacking problems. I asked Edward what Mercor and the industry are doing about training environments that encourage hacking.

Edward Hu

Reward hacking becomes a bigger problem when the task is formulated too difficultly or too vaguely, and the model does not have a “legitimate,” in quotes, way to solve it. This pressure to receive a reward in such scenarios often leads to unwanted behavior.

One example is in the dataset released by Apex Agents. This is quite interesting; we have a recent blog about it. There are professional tasks asking an agent to create a financial model, and we evaluated it by checking whether this particular model contained the correct financial answer that we knew.

The model worked like this: There were certain ambiguities, and in many cases it was our fault that we didn’t include all these specific parameters. The model guessed these various parameters, ultimately providing many, many different answers.

We call it “shooting with a shotgun,” when the model simply says, “Oh, if that’s true, then the answer is this. If that’s true, the answer is like this.” This is not how a person would actually do it in a realistic case. A person would ask for clarification. A person would adhere to industry standards.

But in this case, because the rubric rewarded the inclusion of only one answer, the model provided a lot of answers—up to 10 or even a dozen—and received a point if it hit one of them.

That’s why we reissued the benchmark, which is called Apex 1.1. First, we made sure that the tasks were well defined, because that was, to some extent, the source of the problem. Secondly, we have rubrics that actually punish this behavior.

Nathan Labenz

Do you remember the take on the development of frontier labs: superhuman productivity in everything on which the labs are concentrating, with science next? The next day, I expressed this opinion to Edward, adding that in some areas human data is no longer helpful.

First of all, do you believe in these statements, or would you contradict them?

Edward Hu

Of course, there are areas where humans no longer make a contribution to achieving advanced results. One example may be kernel optimization, where the goal is to write a kernel that works faster on certain hardware.

The model can write code that an expert person can’t even understand, but as long as the optimizations are not vulnerable to breakage and must work on hardware that is very hard to break, the model will be able to achieve these superhuman results, if it hasn’t already reached them.

In those areas where human judgment remains quite important, where human taste is often very hard to compress into a number—to say, “Okay, this is clearly better than the other”—there is still a narrow place for human improvement.

As we see today, what we’ve watched in practice boils down to how clearly we can determine, for a certain area, what it means to be better.

Nathan Labenz

And can we determine this as a function that is easy to evaluate and indisputable? If we can do this, then when we invest a lot of computational capacity, even current algorithms can quite effectively find a solution that corresponds to what it means to be better. But a significant part of the economy still isn’t ready for this when we’re talking about our ability to determine which is better.

The show on Wednesday opened with a track that Prakash created with the help of Claude Fable, collected from fragments of voices that you’ll recognize. He explains at the end of the episode how he did it: “I was spending my tokens. I had some remaining tokens from Fable. I gave them to Fable, and it chose probably 7 or 8 fragments from the last couple of years—just random things that I remembered, the key moments, right? James says, ‘I woke up a loser,’” and so on.

In fact, I probably listened to about 9 versions of this track. Every time I listened, I thought, “Ah, I don’t like it. I like it, but it’s not quite right.” Eventually, I brought it to the desired state. I gave it access to a digital audio workstation based on the API, which it used to combine different instruments in the track.

I said, “Do it. Use this. You can play with the tools. Decide which instruments should play.” It doesn’t hear, obviously. I think Suno is real generative AI. It’s not quite generative AI; for me, it’s really like a game with the tools.

And this is what really surprises me: it composes several tools together. That’s why Claude came up with a sequence. It chose which samples to use, and it chose specific phrases. I told it to create a tale. I wanted a particular storyline.

It begins with Dario talking about models that just want to learn, and then it does it. That became a refrain. Then it did the introductory part, where he told how Ilya told him about models that just want to learn. It ends with Kurzweil, who says, very carefully, “We are moving toward a human-machine civilization.”

My part was probably Terence, who says, “We need to slow down.” Every time he says, “We need to slow down,” it accelerates. So that was my contribution—my contribution.

On one level, who cares? What’s the difference? It would be easy to say it doesn’t affect GDP. It’s just like us having fun until death in a new form. I think that point of view is appropriate to some extent. But I also think this hits right at the heart of one of the great questions that everyone has been asking lately: how good will these systems be in areas that are not automatic or easily verifiable?

It really seems that generalization is powerful enough. I have in mind that I’m very doubtful they do a lot of reinforcement learning regarding music. This is probably one of those things where they just threw something at it to see what would work out.

Speaker 2

Okay, that’s it. Better than last time. Perfect. Let’s release.

Speaker 1

That’s impressive enough. And this definitely proves to me what we won’t be able to deny, because they say this only works for tasks with a verifiable reward, not for something longer than a single day. If you thought this would be the last wall standing, unfortunately, it won’t hold up. Listen to the track one more time if you don’t believe me.

Speaker 3

Models, they just want to learn.

### Episode Outro

Shortly before founding OpenAI, I met Ilya. One of the first things he told me was, “Look, the models just want to learn.” My children will never grow up in a world where they are smarter than computers. For the first time, we will have something smarter than the smartest person.

Will AI kill us?

Models, they just want to learn. Models, they just want to learn. They will do it. They will do it.

We need to slow down. We will be on the exponential part of the S-curve for a long time. This is pretty crazy. This is just insane. This can fake my books. This essay is not AI. Wait.

Models, they just want to learn. Models just want to study. They will do it. They will do it. Models just want to study. Models just want to study. They will do it. They will do it.

This is the most dangerous thing to which humanity has ever been exposed. We must slow down. You are not talking to those who woke up a loser. Woke up a loser.

Models, they just want to learn. Models, they just want to learn. They will do it. Models just want to study. Models just want to study. They will do it. They will do it.

We need to slow it down. We must slow down. Researchers are constantly surprised by the pace of progress. This is sometimes called quick takeoff, where perhaps they do it extremely fast. Everything is moving very quickly.

So they are immortal. We have actually solved the problem of immortality, but only for digital objects. Going to digital intelligence or augmented people is inevitable. The day will come when AI will do all the work—our old jobs. Not only part of the work, but all of it.

We must slow down. We have to slow down. We must slow down. We have to slow down. We must slow down. This is madness, this pace. There is no reason to do it this quickly.

Models, they just want to learn. We must slow down. Models just want to study. We must slow down. Models just want to study. We must slow down. Models just want to study. We must slow down.

The government has to regulate. Violators will be caught. Some creative professions perhaps will not disappear, but maybe they shouldn’t have existed from the beginning. There will come a time when there will be no work needed.

Models just want to study. They are capable of self-improvement, perhaps programming the future versions of themselves. So what was a small advantage, let’s say a few days, may suddenly become an abyss.

Models just want to study. Quick takeoff. Someone says 3 years, someone 5 or 10. Numbers are constantly being called out. You’re talking about the year 2030 as if I knew what the world would look like in 2030.

Quick takeoff. We have a brain. The brain is a biological computer. Should it be considered part of humanity, or something different from it? These models simply want to learn. There will not be a clear boundary between man and machine. After all, we are a single human-machine civilization.

This technology has already expanded our opportunities, and it will come out in full force when we reach the steep stage of exponential growth.
