[BidClub_]
Latent Space · · 80 分钟

Amp:皇帝没穿衣服

Quinn SlackThorsten Ball

YouTube
TL;DR
  • Amp 之所以存在,是因为 Quinn Slack 和 Thorsten Ball 认定,工具调用型 agent 与 Cody 是两种不同的产品,也是两套不同的经济模型。 从 Claude 3.7 开始,他们移除限制,赋予模型工具和权限,最终得到一种既无法塞进 20美元订阅、也无法适应 Sourcegraph 企业版发布节奏的产品。独立品牌让 Amp 可以颠覆母公司,同时重置客户预期:「唯一重要的事情,就是做出最好的 coding agent。」

  • Amp 的经营前提是:如今 coding 工具里的任何显性赢家,都可能在 6到12个月内过时。 Copilot 的注意力先让位于 Cursor 和 WinServe,随后 Amp 和 Claude Code 登场,开发者又开始拿 Codex 与 Claude Code 比较;Quinn 说,没有任何工具长期保持统治地位。因此,Amp 优先优化反应速度——每天大约发布 15次——而不是守护今天的界面、功能集或合同。

  • 早期商业信号显示,即便异步 agent 尚未将用量成倍放大,高价值的重度用户市场已经存在。 Amp 的月环比增速超过 50%,有些只有 2到3人的团队,年化支出却达到数十万美元;Amp 还表示自身保持正毛利率。Thorsten 预计,全天候、并发运行的后台 agent 会创造出「多出 10倍或100倍」的 agent,进而可能同时主导推理需求和软件产出。

  • Amp 的 8人核心团队,像是在 Sourcegraph 成熟基础设施之上运行的个人项目。 工程师直接向主分支提交代码,跳过正式代码审查,自己长期使用 Amp,并在一天之内持续发布;底层平台则由既有的安全和云团队负责稳定。Thorsten 用冰山作比:水面上可见的部分可以「融化并重建」,水面下的结构仍然可靠。

  • 模型正在变成实现细节,但眼下针对具体模型的优化仍然决定成败。 Amp 在产品中隐藏模型选择,沿着「模型的纹理」调校系统提示词和工具,并组合 Anthropic、OpenAI 与 Google 的模型,而不是把某一家供应商视为永久赢家。Amp 内部团队中约有一半已经不再把 Sonnet 作为主要模型;更快的模型则可能开启新的工作流,而不必让人感觉那只是更便宜的「中档」agent。

  • Amp 的反共识产品立场是:许多流行的 agent 功能制造的是复杂性或虚假信心,而不是持久能力。 提示词增强器无法凭空创造缺失信息;MCP 服务器可以暴露 50个底层工具,却同时消耗上下文、不可预测地失败;自定义子 agent 链条既昂贵又难以调试;压缩则会让长对话变得「模糊」。其核心比喻是「木制脚手架」:当下一代模型吸收今天的权宜之计后,脚手架就应当自行消失。

  • 下一个瓶颈将从代码生成之外转向上下文、测试、审查和 agent 原生基础设施。 Thorsten 仍然反对「把思考外包出去,却不把打字外包出去」;Quinn 则认为,定向测试子 agent、可运行的 Playwright 环境、统一的机器可读日志、已部署的预览,以及让并行 agent 快速完成环境定位,都是尚未被充分开发的空间。他们瞄准的终端市场,不是让中位数开发者提速 30%,而是那些离问题足够近、能够创造全新软件,或自己替代一款年费 25万美元产品的重度用户。

摘要 · 为研究而整理的核心内容

1. Amp 的诞生,是为了颠覆 Cody

  • Thorsten 于 2月回到 Sourcegraph,随后围绕 Claude 3.7 与 Quinn 开始实验:给模型工具、宽泛的文件系统和编辑器权限,然后「让它彻底放开去做」。在旧金山那一周的最初几天里,两人反复出现的反应只有一句:「Quinn,你得来看看。太疯狂了。」

  • Cody 建立在 RAG、助手侧边栏和 Sourcegraph 平台之上;Amp 的工具调用型 agent——由模型、系统提示词、工具及工具提示词共同定义——需要完全不同的用户行为、定价方式和发布机制。它不能简单地变成 Cody 的另一个模式,否则会让那些被大型长期合同绑定的客户无所适从。

  • 新身份有意将 Amp 与 Sourcegraph 分开:有些用户甚至不知道它是 Cody 的创始团队做出来的。Thorsten 认为这种切割「太好了」,因为团队可以改变产品,而不必背负旧产品的品牌、客户合同和既有假设。

  • Quinn 的表述非常绝对:「唯一重要的事情,就是做出最好的 coding agent。其他什么都不重要。」他紧接着补充说,目前还没有人做到这一点;每一个当下的领导者都在跑步机上,下一步技术就可能让上一个产品失效。

2. 波动性是 Amp 的核心商业假设

  • Copilot、Cody 的聊天和 RAG、Cursor 与 WinServe,再到 Amp、Claude Code 和 Codex,这条已观察到的更替路径表明,开发者的忠诚周期可能只有 6到12个月。Thorsten 反复强调的公理是「一切都在变化」,而且每次重大模型发布都可能带来变化。

  • 年初,Cursor 还是市场公认的王者,被称为「有史以来最大、增长最快的网站」;几个月后,人们就可能说某个新工具让 Cursor 「看起来像 GitHub Copilot」——陈旧、无聊、面向企业。就在大约 2年前,Copilot 也还是业界最先进的产品。

  • Amp 的应对方式不是预测最终界面,而是让品牌、预期、架构和团队都具备快速转向的能力。Thorsten 认为,即便规模化进展顺利,也必须保留迎接新技术「掀翻桌布」并改变整个品类的准备。

  • 目前 Amp 的增速超过 50%月环比,有些只有 2人或3人的团队,年化支出已达到数十万美元。但更大的潜在断点,是异步 agent 在后台并发工作,可能产生 10倍或100倍数量的 agent,并主导推理需求和软件产出。

3. 小型产品团队,建立在稳定的企业业务之上

  • Sourcegraph 已盈利的既有产品、收入和客户信任,为 Amp 的实验提供资金。正是这种信任,让 Amp 可以暂缓可预测定价、用户自主选择模型,以及各种安全、合规和法务勾选项——这些事项原本都可能吞掉 9个月时间。

  • 核心团队大约 8人。目前仍没有正式的代码审查流程,直接向主分支提交代码,每天发布约 15次,并默认由弄坏 CI 的人——或者当时还醒着的人——负责修复。

  • Thorsten 形容,每天大约有 15次判断发生在「个人项目式的胶带拼接模式」与「Google 就是这么做的模式」之间。真正的能力,在于知道哪些标准有价值,哪些标准只会把一个可能很快需要重建的产品硬化成化石。

  • Sourcegraph 的安全和基础设施团队承担云端运营、扩展等持久性问题。这样一来,Amp 的应用和 UX 层就可以像冰山尖端一样运作:它可以「融化并重建」,而水面下的平台仍保持稳定。

4. 编辑器和终端都还没有赢

  • Amp 最初做成 VS Code 扩展,是因为 marketplace 负责分发和快速更新,同时 agent 可以贴近选区、诊断信息和编辑器光标。团队刻意避免深度集成,以便让产品表面保持小巧且可替换。

  • Claude Code 颠覆了 Thorsten 对开发者会拒绝 CLI 的判断。终端天然提供 SSH、目录、环境变量、标签页、窗格和多个并发会话;独立桌面应用则必须一点点重新发明这套运行环境。

  • Amp 在重做 CLI 两次后发现,内部用户在 CLI 与 VS Code 之间呈现 50/50分布。编辑器更自然地呈现图表、图片和丰富控件;终端则可以跨编辑器和远程环境使用。Thorsten 暂时认为,年轻开发者可能更偏向编辑器。

  • 对话发生前 3周,Thorsten 还认为 Amp 会痛苦地放弃 VS Code 扩展。随后他重新考虑:未来可能出现基于网页、手机或 WhatsApp 的异步 agent,它们或许需要不同界面;而一个未公开的项目也让他觉得 Amp 应该保留 VS Code UI。现在他们会同时保留两种客户端,但「总有一只手悬在按钮上」,随时准备砍掉复杂性。

5. 目标不是赢下今天的 UI 交互

  • Alessio 追问,把 Amp 嵌入 Cursor 这类 AI 原生编辑器是否很别扭,因为 Cursor 自己的「修复」和聊天按钮也在争夺注意力。Quinn 的回答是,如果开发者在 6个月或12个月后会以不同方式交互,那么今天赢下这场特定交互,「坦白说……并不真的重要」。

  • 用户会提出 BYOK、模型选择器、100美元订阅,或每天支出达到 30美元后的封顶机制。Amp 的回应是:一个只在 2%时间里被使用的功能,是否值得拖慢客户真正看重的 agent 改进速度?

  • Alessio 问,市场上是否存在「中档 coding agent」。Thorsten 承认,今天确实有人愿意为更便宜的产品付费,但如果用户 6个月后又转向下一款产品,这种持续流失无法支撑可持续的软件业务:「如果真是这样,所有软件商业模式都行不通。」

  • Amp 唯一一次正面输给竞争对手,是对方将一款产品 100%折价 2年、另一款产品折价 85%。Amp 拒绝跟进;它更愿意服务那 20%需要最大算力和能力的可信重度用户,让 Copilot 或 Cursor 覆盖另外 80%,从而保留继续「大胆而疯狂」的许可。

6. 最好的 harness 必须具体,但模型终将退居幕后

  • 过去每发布一个新模型,就宣布它立即可用的仪式,已经失去意义,因为实际表现取决于系统提示词、工具和周边脚手架。Thorsten 可以在 10分钟内接入 Gemini 2.5,但如果没有沿着「模型的纹理」进行调校,用户就会错误判断它真正的能力。

  • 主持人说,Amp 用 Sonnet 4 驱动 agent,用 o3 驱动 Oracle。Thorsten 表示,Amp 不会在产品中展示模型名称,只是因为用户提出要求才将其写入文档;他预计,模型版本会越来越多地隐藏在交互模式之后。

  • 一个能力不如 Sonnet 4 的模型,速度可能仍然快 10倍。在受限工作流或不同 harness 下,这种速度可以让用户感受到完全不同的交互,而不是一个打折的「中档模型」——抽象层级更接近于选择任务,而不是选择引擎。

  • Amp 已经组合使用 Anthropic、OpenAI 和 Google 的模型,并接近推出一个快速的开源子 agent。Thorsten 提到 Qwen3-7B-Chat 和 Kimiko 等开源模型;不同模型家族各有强项,因此坚持单一供应商越来越像人为制造的纯洁性。

7. 模型供应商的领先地位仍极不稳定

  • Thorsten 说,Claude 3.5 Sonnet 于前一年 6月发布后,用户花了 8到9个月才发现工具调用型 agent 能做什么。Claude Code 及相关工具把这种能力展示出来后,大量资金涌入模型和 harness,而相关收益到 2025年9月才开始显现。

  • Amp 内部团队中约有一半以非 Sonnet 模型作为主要工作流;过去,他们只有在测试时才会不情愿地这么做。Thorsten 认为,如果假设投入模型研发的人才和资本最终什么都做不出来,只能达到 Sonnet 的水平,那将是「一个极度悲观的结论」。

  • GPT-5 Codex 的 API 访问延迟,并没有让 Thorsten 假设供应商在恶意封锁。他相信模型公司关于首发版本能更快收集数据的说法,也愿意给予它们 Amp 希望从客户那里获得的同等迭代空间。

  • 他对供应商经济学的判断更尖锐:在「1个月或2个月内」,新闻周期可能会宣称,基础模型公司投入数十亿美元资本开支并激进招聘,最终却丢掉领先地位,引发资金和人才出逃。他没有直接预测这一幕,只是以 Anthropic 在 Sonnet 之前的逆转为例,说明它确实可能发生。

8. 持久产品会使用注定消失的脚手架

  • 相比「harness」,Thorsten 更喜欢「脚手架」这个词:模型周围搭起的木制结构,等更好的模型能够记住更多内容、可靠编辑或吸收某个权宜之计后,脚手架就会被拆掉。在单独的 apply 模型上投入 3个月,可能会因为版本「0.7或0.8」已经能够直接完成这些编辑而白费。

  • 这种逻辑会消灭传统的 6个月路线图。设计、产品和工程组成的三角关系变成一个未定义的区块,因为没人知道模型最终能做什么;今天规划的功能,交付前可能就已经针对了一个消失的能力边界。

  • 更深层的复杂性来自非确定性。Thorsten 把 LLM 比作「一台奔腾时代电脑里来自外太空的这个球体」:假设你要求 Word 加粗文字,它却偶尔把文字改成斜体。一个看起来 98%或95%有效的工作流,仍然可能在剩下的场景里失败。

9. 流行的 agent 功能经常制造虚假信心

  • 最清晰的失败模式,是把 vibe coding 变成「把思考外包出去,却不把打字外包出去」。工程师仍然需要理解应用、架构、部署和目标结果;否则,数小时看似高效的生成,最后只会产出没人理解的意大利面代码。

  • 提示词增强器从前提上就不成立:有用的提示词包含用户脑中的信息,而另一个 LLM 不可能「凭空把这些信息变出来」,也无法在没有获得这些信息的情况下缩小搜索空间。

  • 用户自行配置 MCP,往往会引入痛苦的身份验证、可能多达 50个底层工具、巨大的 token 开销、延迟、费用,以及模型未经训练去选择的动词。自定义子 agent 链条会放大同一问题:更多上下文边界、模型调用和非确定性,以及更难调试的输出。

  • Alessio 的 GPT-5 实验把批评落到了实处:模型提出 10个能让自己成为更好工程师的工具,完成了实现,却在重复任务中一个都没用,因为它判断自己不需要这些工具。今天看起来聪明的工具,可能只是在解决模型本身已经学会绕过的限制。

10. 上下文质量胜过拯救一段疲惫的对话

  • Amp 包含压缩功能,但 Thorsten 的手又一次「悬在拆除按钮上方」。把 50轮来回消息压缩,并不能恢复干净的上下文;它只会牺牲信息保真度,让线程变得「模糊」,尤其当对话长度本身记录了反复失败时。

  • Thorsten 建议,在重启前,先把失败对话中真正有价值的学习内容写回一个规范的 Linear issue。Quinn 警告,模型可能不知道哪个细节是承重的,因此 Amp 鼓励使用大量短线程,并严格控制每个上下文里进入了什么内容。

  • 工具输出抽象也有同样风险。把冗长的 Bazel 输出隐藏在摘要之后,可能破坏「所见即所得」,让模型获得与用户不同的证据;而且,随着 CLI 为 agent 调整输出,永久维护一个包装器也可能很快过时。

  • Thorsten 的「顺着模型纹理走」原则同样适用于记忆。大多数人使用 coding agent 还不到 3个月,而模型团队需要时间做微调,也只能运行有限数量的实验——其中可能有一半会做错。因此,持久性上下文失败最终可能需要训练来解决,而不是再增加一层本地工作流。

11. 测试和可观测性正在成为面向 agent 的基础设施

  • 测试仍是最具持久性的工具之一,因为它能闭合反馈回路。目前,即便有不错的 agent.md 测试指令,Amp 的 agent 也可能只有 90%–95%的执行概率;它有时会调用错误命令,或以错误方式转义命令。

  • Quinn 认为,定向测试天然适合做成子 agent:利用编辑器位置和项目结构,只运行相关测试,而不是整套嘈杂的测试集。规模化后,可以为标准 Go 测试、Vitest 和流行框架提供优化过的确定性加模型行为,再用更强的通用模型覆盖长尾场景。

  • Playwright 是更大的现实瓶颈。大多数使用 agent 的网页开发者,要么没有 Playwright,要么身份验证配置糟糕到 agent 根本进不去应用;一个能够对运行中的应用执行自然语言端到端流程的并行子 agent,可能显著改善验证。

  • 代码库会围绕 agent 进行适配:把浏览器、客户端、数据库和应用输出合并成一份 JSON-lines 日志;记录 HTTP 和推理交互,让测试暴露真实的响应结构;将 agent 集成进框架,让失败的应用能够主动提供上下文。Quinn 的问题已经从人类可读性转向:「怎样才能针对 agent 的消费方式优化它?」

12. 外层循环会奖励重度用户,而不是中位数用户的普及

  • 并行 agent 带来的第一个问题是定位,而不是轻松获得杠杆。Quinn 说,他会忘记自己启动过哪些 agent,也无法舒服地同时控制 5个任务;Thorsten 把理想工作流比作一个人在中央公园同时对弈 10张棋桌——坐下、理解局面、走一步,然后继续下一桌。

  • Amp 团队用 Amp 编写了自身代码库的 80%–90%。团队跳过正式审查,但要求有人检查 agent 的输出,因此每段代码至少经过 1个人审阅;GitHub 的 reviewer 模型无法体现这样的情况:两个 agent 产出代码,而人类已经把它检查了 3遍。

  • Quinn 不认为「非技术」是关键分层。一名可能只上过 1门计算机科学课程的董事会成员,曾用 Amp 替代一款年费约 25万美元的软件,因为他能够毫不含糊地说明需求;目标用户是有动力的重度用户,而不论这个人过去是否把自己定义为 coder。

  • Thorsten 没有固定的 coding-agent 评测套件,尽管他过去曾搭建并推动过评测。代码库、提示词和任务的覆盖面太广,这项投入与 Amp 的速度不兼容;相反,每天 15次发布和高强度 dogfooding,可以在当天暴露提示词回归。他们的终点不是赢下今天的基准测试,而是 Quinn 所说的「coding agent 之神」;Thorsten 称,造出它将是编程的「最终 Boss」。

Speaker 1
Speaker 1

Everyone, welcome to the Linear Space podcast. This is Alessio from Sourcegraph Labs, and today there’s no Swyx. He’s in Europe with an AI engineer, but I’m joined by Quinn and Thorsten from Sourcegraph. Welcome.

Quinn Slack

Thanks.

Thorsten Ball

Great to be here. Great to be here.

Speaker 1

We already had “Origin of Sourcegraph” with Beyang and Steve. This was when you launched Cody. Today, I guess Cody is a brand that’s past, and now you have Amp. Let’s maybe start there. Obviously, Quinn, you’re CEO of Sourcegraph. Thorsten, what’s your role or title? How do you describe what you do?

Quinn Slack

CEO is much easier.

Thorsten Ball

I’m not going to name my internal title, but I’m—

Quinn Slack

The dictator of Amp.

Thorsten Ball

Yeah, that’s the internal title. I’m the lead engineer and one of the creators of Amp.

Speaker 1

Were you part of the thumbs-up, thumbs-down decision on the Cody brand? How did you decide? How did you get to Amp? Let’s tell that story.

Thorsten Ball

I’ll start. You can jump in. Basically, I came back to Sourcegraph in February, and this was when Claude 3.5 and Claude 3.7 happened, too. Quinn and I started hacking on, “What if we just take Claude 3.7, give it tools, and let it go nuts?” No constraints, and none of the other stuff that we had in Cody, which works for Cody. We thought, “Let’s just start trying this out.”

I remember the first week in San Francisco, I would stand up in the middle of the room and say, “Quinn, you’ve got to see this. This is crazy.” Then he was like, “Okay, let me try this,” and we went off from there.

We realized relatively quickly that it’s a different kind of product. Cody was very much one of the first of its kind, with RAG and assistant panels, an assistant sidebar, and so on. But with a tool-calling agent—which I define as a model, a system prompt, tools, and tool prompts that go along with it—that you give a lot of permissions to, so it can actually see and interact with the file system or your editor, it’s a different thing.

We realized we had to handle this differently. We had to reset expectations and tell users that it’s a different thing, that they had to use it differently in some sense, and also that we couldn’t make it work with a $20 subscription. Back then, that was seen as an offensive thing to say.

Speaker 1

And we were charging money?

Thorsten Ball

Yeah, yeah. But now people are paying hundreds of dollars per month, which I’ve been saying every day for the last 2 weeks: that’s still crazy to me, how far we’ve come.

This is just how it started: “Okay, this is a different thing.” We were astonished, surprised, and amazed by what these models can do, so we decided to reset expectations and tell a new story.

We have enterprise customers for Cody, but they have expectations. We have contracts, and these are large, long-running contracts. You can’t just say, “Guys, here’s a new mode. It costs however many dollars more, it works completely differently, and you need to use it in a different way.”

In order to avoid this, and to avoid being disrupted, you create the new thing that kind of disrupts the business on its own.

Speaker 1

Yep. That’s it. I don’t know. Want to add?

Quinn Slack

Yeah. The only thing that matters is building the best coding agent. Nothing else matters, because if you can build that, it’s way bigger than anything else that came before. To be clear, nobody has built that yet. We’re getting better and better, but you’ve seen this treadmill of tools that you use as a developer.

First, it started with Copilot. Then Cody—we were really good at chat RAG. Then Cursor and WinServe showed that IDE forks and partially agentic tools could get better and better. Now you have the next generation: Amp and Claude Code. You’re already seeing people say, “Codex is better than Claude Code.” There hasn’t been any tool that has stuck with developers for more than 6 or 12 months or something.

Thorsten Ball

Six months, yeah. We saw that firsthand. We’re now on our second iteration, and we’re able to move so much faster because it has a totally different name and a totally different brand. Some people don’t even know that Sourcegraph, or the people behind Cody, made Amp. That has been so good.

I don’t know how, if you had an AI tool that was relevant 9 or 12 months ago, you can even bring the same brand and the same customer contracts along with you and make a good product. It’s so liberating to be able to say, “Totally different.”

On the technical level, Cody was, or is, a Sourcegraph product, so it works with the Sourcegraph platform. That means you’re tied to the release cycle of the Sourcegraph platform. Sourcegraph is in the cloud—we have cloud versions of Sourcegraph, but also on-premises versions for some customers. It’s a completely different game.

With Amp, we basically said, “Let’s not do this.” Let’s build something that allows us to ship 15 times a day, and that’s what we’ve been doing over the last 6 months. We’re still doing this, and it’s a game changer.

Anybody who’s done this knows that internally and externally, you need to reset expectations that this is a new way of building software. Having a new project with a new way to do it is, I think, better than trying to get the old one to move in this new way, because that would take longer.

Speaker 1

Are there any numbers that you share about developers, like Amp usage overall?

Quinn Slack

It’s growing really fast. It’s growing more than 50% month over month, and a lot faster in some weeks. What we’ve seen, too, is a huge change in who’s using it. We have teams with 2 or 3 people that are on annual run rates of hundreds of thousands of dollars.

We also made a decision not to try to go to every single developer in an enterprise, which we had done with Cody. We pick off the people who want to move as fast as we want to move and who want to stay at the model-product frontier, like us.

Thorsten Ball

Yeah. It’s all about being able to move really fast. I think that the way agents work today, most of them are used in your editor or CLI interactively. You have 1 agent at most running with you at all times. That’s going to be blown up with async agents, when they’re running 24/7 concurrently in the background. Then you’re going to have 10 or 100 times as many.

That’s going to dominate inference, and it’s going to dominate the output you get. Amp is growing really fast, but it’s more about how we get to be the first ones with that 10- to 100-times improvement. Everything is about how we can move fast and learn along the way. It just so happens that we have positive gross margins along the way.

Quinn Slack

I would say that’s one of the biggest axioms we have with Amp: we don’t know where this ride is going, but what we do know is that it’s changing every few months.

At the start of the year, Cursor was the king, and the biggest, fastest-growing site of all time. Now, if you were to ask a lot of developers, “What do you think is the dev-tool king?” I don’t think they would name Cursor first.

Then, I think a couple of months later—or maybe a couple of months before—someone in sales said, “Blah, blah, blah makes Cursor look like GitHub Copilot,” like it makes it look old, boring, and enterprisey.

Just think about this: Copilot is not that old. It was state of the art maybe 2 years ago or something. Now the world has changed completely, and we know that this isn’t over yet. The changes are still coming.

From an engineering and business perspective, priority number 1 is to position yourself in a way that you can react to these changes. Position your product, your expectations, and your technical codebase in a way that lets you react to these things as fast as possible. Everything else flows from there.

Everything else we’ve done is based on this idea that everything can change with the release of another model or something.

Speaker 1

How are you doing it internally from a team perspective? Obviously, you have a lot of customers already on the Sourcegraph product. There’s this tension between going into founder mode and burning the bridge on some of the old use cases, versus having a smaller team and a dictator for a new product. How does that look from a company-building perspective?

Quinn Slack

When you have a really popular, successful product that’s highly profitable and funds a lot of this craziness, you’re able to do this with customer trust. There are a lot of things on Amp that we do—no consistent pricing, no user model choice, no checking off all the boxes that security, compliance, and legal want, which takes 9 months.

We’re able to get away without doing that stuff because we have that customer trust. That has been a big thing.

It requires you to totally change how you think about an existing business. It’s not a way to sell through that same channel to those same users. It’s a way to use that trust and that revenue to fund the crazy stuff that you have to do.

It’s something that we deal with all the time. We’ve got really smart developers, and yet it’s hard for people to throw away everything they’ve learned about how to build software development tools. In some cases, it’s been really refreshing to have people who have only ever been at tiny, one-person companies.

Thorsten Ball

Yeah, yeah. And they come here, and they have no preconceived notions about how you do planning or anything like that. That is great, because you can throw all of that out the window.

This was radical in some sense: when we started, it was Quinn and I working on main—no code reviews, nothing, just pushing. It was like a personal project. I think we’re both experienced engineers, so everybody owns their stuff. You push, and if you break CI, you go and fix it, or, if the other person is awake, they fix it or something.

It seems like when you move this fast and ship this often, throughout the day there are 15 decisions where you have to flip between duct-tape personal-project mode—move fast—and the “this is how they do it at Google” mode. It requires a certain expertise, or it requires being free from the thinking of the last 15 years of always doing things like Google does. The base assumption behind the whole Google thing was always, “Oh, we found product-market fit. Now we have a product. Let’s scale this up,” right?

Every company I ever worked in was based on this assumption: this is the product, so let’s make it proper and engineer it up. But now, with these changes, what’s ingrained in Amp is the understanding that even if it scales up, we have to be prepared for somebody to pull the rug out and for a new technology to come out and shift everything. We have to be prepared for this. Again, it all flows from this.

Now, in our development mode, the team is super small compared to, I guess, other companies, but I think we’re around 8 people on the Amp core team. We still don’t do formal code reviews. We still push to main, we still ship 15 times every day, and we dogfood this as much as possible.

It turns out that in a fast-moving environment like this, this beats a lot of other things. Fast feedback loops, using the product yourself, and dogfooding it—using the product to build the product—beats a lot of established processes. We can get away with it because we can dogfood it.

Quinn Slack

And, yeah, how has it been internally received?

Thorsten Ball

I think we have the luxury of making use of the infrastructure that we already have. For example, we have a fantastic security team, right? The security team comes in and says, “Guys, let us take care of the security stuff for Amp,” so that is fine. I’m like, “Cool, I don’t have to worry about this.”

Then we have infrastructure people: “Guys, let us take care of how to run this in the cloud.” Cool, I don’t have to worry about this. I can concentrate on the client or the UX.

This is a nice spot to be in, where we can move fast but use platform teams to make sure that it doesn’t break, that it scales up, or whatever, while the tip of the iceberg can melt and be rebuilt, basically, while the thing beneath the waterline is stable.

Yeah, not the greatest analogy, but I think there’s a distinction between platform stuff that does work and the UX or product application layer. You want to be able to tear that thing down and rebuild it as fast as possible. I think that’s what we’re doing.

Speaker 1

One thing is, you get a separate team, and then the other thing is: how do you put that team to work, right? If you look at the coding agent space, I mean, obviously, you started with Cody, and I think there was maybe a thesis behind it. Then you had the rise of Claude Code. Now you have Codex CLI, which is trying to catch up. I would say they’re maybe a little behind on the UX and all of that, but they obviously have billions of dollars to train a custom model, so that kind of weighs a lot of the options.

How did you decide about the structure? You have both a plugin for IDEs, so I use Amp in VS Code and Cursor, but I can also go in the CLI and use Amp. Was that an easy choice? Was there a lot of discussion about whether we should just do one of the modes? Supporting both is obviously more work, right? A lot of these products don’t support both. So, what was that initial design choice of the structure of the product? Then we’ll dive into the models as well.

Thorsten Ball

We started with a VS Code extension because it was the easiest thing to get off the ground. When you have a VS Code extension, you have a marketplace, you can ship it, and you can update it 15 times every day. You don’t have to think about updating stuff. You’re also next to the editor.

Looking back, it’s been 6 months. The editor might be dying, or you might do a lot of coding outside the editor. Back then, it sounded much more radical than it does right now. So we started with, “Let’s explore this.” Having the thing next to your editor is a good place to start. You can see the cursor, you can make selections, and whatnot.

But from the start, we really didn’t want to have a deeply integrated thing. It was always, “Let’s keep the feature small. We’ve got to be able to move fast.” Then we built up the CLI on the side as a different client, which also gives us the ability to abstract the core and client stuff. That’s a nice boundary to have.

To be 100% honest, we were also surprised by how many people were fine with using a CLI for Claude Code, for example. If you had asked me 6 months ago, I would have said, “No way, a CLI tool?”

What we realized is that a CLI is not just a UI. Sure, it’s a UI, but it’s also a CLI program. That means you can run it over SSH, in any other editor, in multiple split panes, or in multiple tabs. If you want to do this in VS Code, you have to rebuild a lot of stuff. You have to rebuild the way you switch between conversations. SSH works out of the box with VS Code, sure, but you’re still tied to it.

We had an internal experiment with a desktop application—a standalone application. It turns out, yes, that’s great for having multiple agents, but you also have to reinvent everything that a terminal gives you for free. If I use Ghostty, iTerm, the VS Code terminal, or whatever, I can press Command-N or Command-T, and I get tabs and splits, with different environments per tab. You can `cd` into directories, you can set environment variables—you get this for free.

If you do it in a desktop application, then you run into the issue that people see with a lot of the asynchronous agents. You run a task, set the environment variables, switch to the directory you need to be in, figure out what’s in the path, and so on. You have to do this beforehand, whereas in the terminal you get it for free.

That’s the short version: we started with VS Code because it was easy and gave us a lot of feedback. We could concentrate on the stuff that matters and not worry about distribution, which VS Code takes care of. Then, with the emergence of CLIs, we noticed that there were big improvements and other advantages to them.

We rebuilt the CLI twice, and now we have a really nice tool with our own framework. One interesting thing is that our VS Code extension has a lot of advantages over the CLI. For example, it’s easy to display diagrams, images, and a bunch of other things. We can press Command-Return to submit messages, and so on.

It turns out that we had an internal poll last week at our company meetup, where Beyang was asking who uses the CLI and who uses VS Code. It was a 50/50 split. It’s very strange that it falls out like that. There’s not a clear winner, and both have advantages.

Speaker 1

Do you cut that data based on the level of the engineer or maybe the specialty—front end versus back end? How do you segment that, or do you just take it?

Thorsten Ball

We haven’t really segmented it. If I had to guesstimate, there’s also a generational divide. I would say younger people—those younger than 25—find the terminal old, and they’re much more inclined to use the stuff in the editor.

But, yeah, we don’t have any fancy segmentation. I think, not to sound too dramatic, one of the other guiding principles we’ve had from the start with Amp is that whenever somebody asks, “What’s the data on this?” or “Do we have analytics on this?” it’s like, “Well, did you look for it yourself? Did you try it out? Did you talk to customers?”

We constantly talk to customers. That beats a lot of other stuff. So, we don’t have any segment analysis of who uses what, where, and how.

Speaker 1

I use both.

Thorsten Ball

This idea that everything is changing applies to this. We looked at it, saw the way things were going and how much more flexible a CLI was, and about 3 weeks ago we said, “We think probably—it’s painful—but we will kill the VS Code extension for Amp.”

I laid that out, and I didn’t like it, but it seemed like that’s how things were going. Then you think about asynchronous agents, which probably need to be on your phone and on the web, or maybe you use WhatsApp to interact with them. That’s a whole other mode of interaction.

If it’s on the web, that’s like the VS Code UI, not the terminal UI. Then there’s this other thing that we’re planning on doing that I can’t share more about, but that also makes me think, “Well, actually, we really need to keep the VS Code UI in.”

So, we’re keeping it. It’s definitely adding some more complexity, but there are a lot of things we can do to reduce that and simplify it.

But there's always a hand hovering over the button: Can we get rid of this? Can we shed weight? Can we reduce complexity so we're again in the spot where, if a new model comes out, we can react quickly? Sure, it's good engineering, and there's not a lot of duplication, but updating 1 client is still faster than updating 2 clients.

So there's this constant tension between what's the most minimal thing, the most minimal product that we can have. To pick some other examples, there are a lot of niceties you can do in VS Code. A common example is that you can hover over a diagnostic and then say, “Fix this,” or whatever. People would ask me, “Can you add an Amp ‘Fix this’ button?” But Amp also knows about your selection and the diagnostics; it can see all of this. You can just ask, “Fix this for me,” and if you type 3 words, it will usually do it.

That's something where it's like, well, you can already do it. It's a nicety, but let's remove the surface area. Let's remove this other thing that we have to backport, keep working on, or whatnot. That's a tiny example, but there are 500 of these where we say—

Speaker 1

But how do you think of that when the IDE is already an AI IDE? I use Cursor, right? There's already Fix and Chat that pops up, and they obviously want that button to go to their chat, versus you guys being on the left side and saying, “Yeah, just do this here.”

Do you feel that, in a way, the VS Code extension is more for people who aren't using these AI-first tools and are using the features like most people? I'm sure GitHub is eventually going to have something good to put in VS Code. How much do you think about the VS Code extension just being maybe a stepping stone to the thing you can't talk about, and then the bifurcation of the TUI versus the fully asynchronous experience, where you're not looking at anything?

Quinn Slack

We're not trying to maximize our revenue or our user adoption literally today, with the state of today's models and today's tools, because everything is changing so fast. We're not trying to fight Cursor over who gets to have users fix things with our AI or their AI. Frankly, it doesn't really matter to us. I don't think that interaction is a really important way that people will be interacting with AI in 6 or 12 months. I don't think we learn anything from that, so we just said we're not going to do it.

Some users have definitely asked for it. The other thing is that we have to figure out what users actually want. They say they want a lot of things. In the case of customers, a lot of times they'll say they want a lot of things. They'll say they want bring-your-own-key, model choice, a subscription for $100 a month, or pricing that locks users out if they spend more than $30 in a day.

But what we've actually seen is that they want the very best coding agent. Not everyone, but we're focused on the ones who want the very best coding agent. When we tell them how one of those things will slow us down, that starts a conversation where they'd rather not have something they might use 2% of the time if it means the tool is worse.

It feels like we alone among the entire industry are being really honest and really bold about that. I am really concerned, just for the rate of progress overall, that a lot of these other tools that are great—like Claude Code, Codex, Cursor, and so on—have forgotten what made them great and what made them grow so fast, which is building the very best product. They've built them in a way that's too overfit to the current capabilities, and so they're just going to peak and then have a slow fall.

Zero of the software business models work if that happens. You need to have growth into the future. We're trying to do what's best for our business, but I also think we're trying to push the whole industry to be radical about the changes that are coming.

Speaker 1

When you said “the best coding agent,” I'm always wondering: Is there a market for the mid-level coding agent? I think model choice is a great example. Why would you want model choice? Pricing, I guess, is the only thing that people bring up, but to your point, you're already paying engineers a lot of money.

The cost of Sonnet 4 versus Sonnet 3.5 is minimal compared with the $150,000, $200,000, or $300,000 that you pay employees once you include taxes, benefits, and all of that. I think we're in this part of the market where people aren't maxing out these things.

Thorsten Ball

There's absolutely a market today—literally today. Someone will pay a monthly fee for that cheaper AI product today, but they're not going to be paying that in 6 months. It's going to be a different product. They'll be paying for something else. If you have that much churn as a product, you simply cannot build software that way.

But a lot of people get tempted by that, and they hear a lot of users ask for it. 6 months ago, it was still the game of, “A new model got released,” and then everybody would tweet out that it was already available in their editor or extension. I think that's kind of over. People realize that the benchmarks are one thing: This is the best model, but it turns out it's not the best model in this editor. It feels different in this editor.

The whole “the models are the thing”—I don't want to say that's over, but it's becoming less important. People are also waking up to the fact that it's not just the model. It's the system prompt, the tools, the harness, and the scaffolding around the model.

I can give you the choice to use Gemini 2.5 in Amp, but without the system prompt being tuned to it, without going with the grain of the model, it doesn't make a lot of sense. The models are trained in different ways, so you want to optimize the tool and everything around it for the specific model. Without that happening, you get the wrong signal.

I can drop you into a new model right now and have it available in 10 minutes, but that's not what you're after, right? You want the best possible version of this model in this tool.

Speaker 1

And that's become more important, rather than the model selectors and whatnot. Why do you mention the models at all? You have Sonnet 4 for the agent and o3 for the Oracle.

Thorsten Ball

We don't. We don't show them in the product. We don't mention them at all. We put them in the manual—we have an owner's manual—because people kept asking us.

Quinn Slack

Even then, why does it matter that they ask? If you want to change it tomorrow, you have to tell people that you changed the model. Where do you think we are on the slope of, “You guys should forget all about what model is even running and what the difference is”?

Thorsten Ball

I think we're going toward a future where the model will become an implementation detail, in some sense, and we'll end up on a different abstraction layer. You asked when you would use a mid-level model. When you put it like this, it sounds obvious: Who wants to use the worse version of the better version? But we're thinking actively about this.

There are models that might not be as smart as Sonnet 4 as the main agentic driver, but they might be 10 times as fast. That doesn't mean you think, “Well, now I need to go fast, so I have to use this.” But I think there are different modes of working in your day-to-day work. Using this model in a different harness or configuration can be another way to get things done, versus talking to an agent in a back-and-forth.

We've seen this with planning modes, where people use different models, but it's still pretty clear that it's a different model and whatnot. I do think it will be pushed more and more into the background, and people will choose or have different ways to interact with models. The specific model or its version won't be as visible anymore.

Speaker 1

I know Cody was using StarCoder for inline edits, at least, as I said publicly, so I'm not leaking anything. Does it still seem interesting to you to figure out whether there's something in open source that we can use and maybe fine-tune to make better, or are you still thinking—

Thorsten Ball

Yeah, absolutely. We just want to be at the cutting edge, and that's maybe on the back burner.

First, it took people 8 or 9 months to figure out what Claude 3.5 Sonnet was capable of from when it was released last June. Around the time we were building Amp, Claude Code came out, and you realized that a tool-calling agent was incredible. At that moment, everyone—all the smartest people in the world—also realized that billions of dollars had gone into training new models and harnesses based on that.

Now it's September 2025, and we're reaping the benefits of all that investment. You have so many more models coming out. You have open-source models like Qwen3-7B-Chat and Kimiko, and they're moving so fast. You have xAI's models, you have GPT-5, and we're still figuring out how to use these things.

It would actually be an incredibly pessimistic outcome if all those smart people and all that money weren't able to build anything better than Sonnet.

In our internal team right now—and this could change—we have about half of our internal team using a different model other than Sonnet as their main way of using Amp. And that's a huge change. In the past, we had done that only to test and begrudgingly, but now we're using it.

There's a different way of interacting with an agent that's not the linear chat transcript. It means you don't feel like you're getting a cheaper mid model; you feel like this is a different way of interacting, where that speed is really beneficial and it's more constrained. So, things are changing so fast.

Speaker 1

Is GPT-5 Codex only being available in Codex? Does that make you nervous about future availability of cutting-edge models? Does that put more emphasis on figuring out maybe an open-source strategy?

Thorsten Ball

They make it available to API customers; it's delayed. And if they were doing that, I really think that, for the most part, I take these model houses at their word. They wanted to get it out to their first-party product as quickly as possible because they honestly need to gather more data, and they're iterating in public.

I would love it if all the model houses perfectly coordinated with us before they released anything, but I know that would slow them down, and I don't want to slow them down like that. In the same way, we want our customers to give us grace and help us iterate in public.

Speaker 1

I think there's an interesting dynamic in the market. When Cursor switched from Sonnet to GPT-5, the default model that was, you know, $200 million of revenue for Anthropic kind of went away and moved on to GPT-5. So there's kind of like, okay, we're all friends now, you know, but maybe later that's going to change.

Thorsten Ball

Another thing is that if you're building an agent and you're not at one of the model houses, you can use multiple models from different providers, right? Which is what we do. When you use Amp, you're using a model from Anthropic, a model from OpenAI, and a model from Google. We're also very close to shipping a fast open-source model that we can use as a different sub-agent in there, too.

When you put it like this, it seems silly to say we only use one model of this family because they all have different strengths and weaknesses. I think we are 1 or 2 months away from a possible news cycle where the foundation-model companies have spent billions of dollars in CapEx and hired like crazy, and now they're no longer the best in this realm, and there's a huge stampede away from them. That's very possible, and I'm not saying anything new. Just imagine last May, when people were counting Anthropic out before Sonnet came out. Things change so fast here.

Speaker 1

Yeah, yeah, yeah. I think OpenAI, obviously, with Jony Ive and some of that, is moving more in a consumer fashion as well. So it's been interesting to see the big push on Codex. I would have imagined them to go more toward education. I know they have a lot of big enterprise contracts, like ChatGPT Enterprise, that kind of thing.

You guys, I think, are in a good spot because you have both the Sourcegraph trust, like you said, but also Amp. I see a lot of great stuff on Twitter. People are like, “I just put all my Amp agents running, I come back, it's great.” I think it's now on that wave of, okay, this is one of the best tools out there. If you're a serious engineer, you should probably use Amp at least in some capacity, and then make your own choice.

How difficult is it to think about what goes in your harness versus what people should build? You have custom commands. You've done a great job on the tooling, where people can put executables in as tools instead of having to define an MCP server. How much of it is, “Hey, we're just giving you the tools,” versus how much do you want to be opinionated with things like— I think of compacting conversation as maybe one of the key commands that people have. In Claude Code, you can give a custom prompt to compact. What's that discussion like?

Thorsten Ball

The main assumption—again, everything is changing, and we've got to be able to move fast—is that you want to build a scaffolding around the model. I don't use the picture of a harness often. What I use is a scaffolding: a wooden scaffolding that, if the model gets better or you have to switch it out, falls away.

Like the bitter lesson, you want to embrace that a lot of stuff might fall into the model as soon as the model gets better, right? Because then it can remember more, whatever, or it doesn't. Why invest 3 months in a separate apply model when the next generation—the 0.7 version, 0.8, or whatever version of this model—can now do all of the edits on its own?

That's the bigger thing. With that in mind, we really try to restrict a lot of the features that we add around the model. You can do a lot of stuff. We could be busy all day adding things to our clients and whatnot, making a product more complicated, but we don't want to. So that's the first thing.

The other thing is that we're living in strange times from a product-development perspective. Basically, I think the old triangle of design, product, and engineering is changing. It's not a triangle anymore. I don't know what shape it is. It's a block, whatever, but it's not a triangle anymore.

The reason for this is because you can't build a roadmap. You can't say, “This is what we're going to build in the next 6 months.” People don't know yet how these models can be used to their full extent. Everybody's figuring this out on the go.

The third thing—which we just talked about over coffee before coming here—is that we're now in a world where the only UI people basically have is a text UI, and you can use this in the wrong way. The example I used earlier was if you buy Jira, for example, but use it for your shopping list, Atlassian is happy about this, but that's not what they built the product for, right? You can use it in the wrong way and still get results.

The problem with LLMs and a lot of the models is that you can use them in the wrong way, and it looks like you're getting results. You can use OpenAI ChatGPT to look up serial numbers or technical specifications for a camera or something, and it will tell you this, but it might be wrong. Ninety-nine percent of the time, or 98% or 95% of the time, it might work, but in 5% it might not work.

Having nondeterministic LLMs as the heart of your product is something unprecedented that we have in software, I think. With that in mind, a lot of the features that we see involve people building elaborate workflows: “I have my custom slash commands, and they trigger custom sub-agents, and they in turn trigger custom MCP tool calls on the backend, which again another model is doing inference on and taking the input and blah, blah, blah.”

I think a lot of this will—and has—resulted in hangovers where people realize, “Oh, this looks like it's a deterministic workflow. It looks like it does the thing that I want it to do, but actually I can't use it if it only does it 98% of the time.”

That's something we're really conscious of. Everybody's experimenting, and everybody's sharing their experiences—the thread-boy tweets about what to prompt where and how—but you have to be super strict about not giving users a false sense of what the product can do and how reliable it is, because I think it's dishonest in some way, and it doesn't lead to good results.

Just as an example, over the last 3 months, I would say we're ahead of the curve using Amp internally. We're ahead of the mainstream by maybe a month or 2, where we've tried a lot of this stuff and then realized, “Oh, this wasn't the best use of our time or the tokens.” Now you see a lot of other people waking up to this.

Armin Ronacher, the Python developer from Austria, has done a lot of good stuff with Claude Code and shared a lot of his learnings. You could see that the way he tweeted was super excited: “I can now do this and this and this.” Then a month later it's like, “Oh, maybe having 8 remote-control agents that I control with my phone and let them run for 20 hours, maybe that's not as productive as I thought it would be.” And, yeah, it's something that we're super conscious about.

Speaker 1

What are those things? What are the failure modes that you heard from customers where it's like, “Hey, we tried Amp and it just didn't work at doing XYZ”? Is there a collection of those that you guys use as almost like a North Star as you keep building?

Thorsten Ball

I think one of the things is the whole vibe-coding stuff, where people just use it and they're like, “Hey, I spent $10 in tokens and it didn't build me the whole app,” or something.

The failure mode is outsourcing the thinking but not the typing, which I think should be the opposite. You still have to know engineering, you still have to know how to program, and you still have to know your application and its architecture, how it's deployed. Then basically use the agent to do the work that you would have done, but you have to know what the desired outcome is, or whatnot.

That's a common one where people just hand off the wheel: “Agent, you go and write this for me.” Then it turns out a couple hours later, “Oh, actually, nobody understands this. It's spaghetti code.”

Quinn Slack

Amp is different from the products it competes against. We've had 1 head-to-head loss with Amp, where we lost against the usual players. The reason why is that 1 of them discounted their other product 100% for 2 years. The other one discounted it 85% for 2 years, which is just crazy.

We wouldn't want to do that, because are we really going to learn from that? And then, how's it going to be used? Can it be used in a different way? Usually, the way that we might lose is that there's some other product that would go to 80% of the devs in a company as the base layer. Sometimes that's Copilot or Cursor.

Amp is more expensive, it's more powerful, and they'll give it to that 20% of devs that they trust more. In a previous world, any software company would say, "Oh, no, we need to get 100%. We don't want our competitor getting in there." But actually, that means that we're able to focus even more on being bold and crazy, because all those devs can always fall back to Cursor or Copilot. We actually really like that kind of deal.

Yeah. The other thing there—I think a bunch of questions already touched on this—is that, talking about segmentation, the market, or the ideal user, everything is changing. What we tried to do is build a tool for people who are at the frontier, or at least curious about it and want to figure out how to use these agents in the best possible way. That's based on the assumption that if you build for the mainstream user—mainstream sounds like, I don't know, it sounds bad, but what I mean is if you build a product for somebody who does not know what a good prompt looks like, you will fall behind right now.

You'll spend time and resources building stuff like the prompt enhancer, but then you'll end up building this and miss the next step change that might happen. So, the way we think about it is that we build for the people who already get that a lot of stuff is changing, but we want to leave the door open. If you're open to learning new things and want to learn how to use AI and agents in your workflow, please come with us. We're happy to have you. But if you're skeptical and you think prompt engineering—that's a term—I don't care about this, we're not building a product for you right now because we would fall behind.

Yeah, so prompt enhancer—that's a feature that doesn't actually work. The theory behind it is nuts, because what helps LLMs is not tricks in phrasing your prompt in a certain way; it's fundamentally information that you have in your head that you can bring into the prompt. If you don't have that in a prompt enhancer, an LLM cannot magically conjure that up. It cannot narrow the search space for you.

Custom sub-agents—the way we disqualify that as something we want to build at this point is that you look at all of the tokens you're sending to the model, and there are so many more. It's so much more convoluted. We don't think that these models are trained in a way that would support this use case, and the output of this going in here is so much harder to debug.

MCP is another thing. MCP has done a great job of getting products to expose the verbs that agents might want to interact with, although in most cases they do not actually get the right verbs exposed. But as a user-facing technology, it is such a common failure mode where a user will go and add in some MCP servers. Auth is a huge pain, but let's say they get over that hurdle; then they have 50 tools exposed that are often too low-level in granularity. It takes a ton of tokens in the model, makes everything slower and more expensive, and the tools are often misused.

It's just not a good experience. So there are all of these things that we've said no to, and other tools are bringing them in and saying yes to all these things. I think it feels like they're making progress in the meantime, and people retweet and talk about how they're able to do these amazing things.

But the simplest example that seems so obvious—and frankly, it confounds me that more people don't do this—is to make it so that my Google Docs, Notion, Linear, and GitHub Issues are all accessible to my agent. The vast majority of developers who use Amp or Claude Code or anything else don't have all those context sources set up. That seems like such a slam dunk.

Before we would do that—and we've built that; we ripped it out—we'd have to get an answer, even for our own usage, as to why we're not doing that. It's frankly still puzzling to us, but we're not going to touch that until we get confident about it.

And to come back to the example you mentioned, compact: we have this in the product, but again, the hand is hovering over the "rip it out" button. I think compact is such an alluring thing where people think, "Oh, I ran out of context, I hit that button, and now I'm back to the start." But you lose signal, you lose data, and it's something where—are the models really good enough? Is compacting good enough to really gloss over this so that the user doesn't have to worry about it?

Or is it something where you would have to somehow make it clear to the user that, "Hey, look, your conversation has 50 messages back and forth. If you hit compact, this is all going to become blurry. You're going to compress it, and you lose signal, you lose fidelity, and then you put it in a new context. Now, are you sure this is the right trade-off?" Some users are, but again, it's strange times because now we have this thing at the heart of our software, this orb from outer space that can sometimes do whatever it wants. It's strange to build on top of this, and it's strange to educate your users about this.

Imagine the end of the '90s PC era. You had to build Microsoft Word, and then say, "At the heart of this new personal computer, the Pentium 3, whatever, there's a weird orb from outer space, and sometimes if you bold text in Word, it actually makes it italic." But that's the situation we're in. That's the fact: it doesn't always bold the text; sometimes it underlines it if you reach 150 tokens or 150,000 tokens or something. How do you teach this to the user?

Speaker 1

Yeah, and we're in the church of context engineering at the Amp office. When we had Jeff on the podcast, they talked about the context drought paper that they did, and they specifically mentioned that encoding—for example, showing previous failures—was not helpful at all.

Speaker 1

Mhm.

Thorsten Ball

To the agent. And so, I think when you're compacting a conversation, there's almost—if you have a long conversation, it usually means something went wrong along the way and you had to go back and forth on a bunch of things that didn't work, and you're keeping those in. I've been trying to figure out what that's going to look like.

In my mind, it's almost like if you take the idea of Linear, which I use and give to my agents, just because then I have a canonical prompt for 1 issue. Often, you have to restart—

Quinn Slack

Yep.

Thorsten Ball

Because it just goes too far down the wrong path. A lot of people don't restart. A lot of people just try to keep going. Yes, that's bad. In that case, what can you take from that conversation as a learning and put it back in the upstream issue, so that the issue is either more descriptive or has more information?

That is not compacting, but it's almost like how you would do it as an engineer: you're doing it in your mind, right? You get an issue, then you start working, and then you update your mental model. That doesn't really work for agents, but people are not doing this small increment in the initial issue.

Quinn Slack

I would say in this case, you still cannot outsource your thinking, right? In this case, I don't think you can expect a model right now to say, "Out of this conversation, this is the most important thing. Let me put this back in the Linear thing." Maybe if you phrase it like this and automate it like this, and it's always a perfect conversation, maybe it works. But I think in this case, you still have to be mindful of the context.

What we encourage users to do, for example, in Amp, is to start a lot of small threads and be really strict about what goes into context and what doesn't. The other thing that touches on this is where a lot of CLI tools, for example, have super-verbose output. Basil—sorry to call this out—I’m not a big Basil fan, but you could just call Bazel out for its super-verbose output.

The natural assumption is, "Oh, let's hide this from the user. Let's abstract this away and summarize the output or whatever—just the exit code or something." Then you get into this dangerous territory where "what you see is what you get" is not true anymore. In the context, what you see is some other thing in the context, and that could lead to issues.

But for me, the meta thing here, too, is that everything is changing. That means we're seeing this: CLI tools right now are also adapting to being used by agents, so they're changing their output, too. If you focus on the fact that Basil will always be verbose and build something for this issue, you might be outdated in 6 months, where somebody is like, "No, no, no, we have a Bazel agent wrapper now. This is not an issue anymore."

Speaker 1

Interesting. Yeah. One model that I have is if you are relatively on the cutting edge of using agents and there's some persistent problem like this, it feels kind of out of band.

Quinn Slack

This includes how the model itself updates its memory or updates the Linear issue. The model needs to be trained to do that better. If it's something like your own coding conventions, that's different. But if it's something fundamental that feels out of band from the agent, the model needs to be trained to deal with memory better, or to accept that it might have an incorrect view of its own history if you go back and edit it.

We're feeling these pains right now because people have only been using agents and coding tools for a matter of months. Most people have been using them for less than 3 months. If we're only feeling them now, it takes a little bit of time for a team at a model lab to fine-tune one of its really big models. They've got other big models and new revisions being trained, and they can only fit a certain number of experiments like this in. They're probably going to get half of their approaches wrong, so you can only do so much. That's Thorsten's idea of going with the grain of the model.

You've seen this, I'm sure, where a lot of users go through this lesson: “Let me just add this MCP server that does everything I want it to do,” and then 2 days later, it doesn't use it. It never calls the tools. It's like, “Yeah, it wasn't trained to do this.” You can sense that the model labs have different philosophies. Anthropic, from what I can tell, is working or training a lot toward using memory—storing information—whereas ChatGPT, obviously, has its own OpenAI approach. So if you give it a memory thing, it might use it. But then you have the issue that if I give it another custom-made MCP that we built internally, and our processes don't map to anything OpenAI or Anthropic have seen or been trained for, it won't be used and you won't get good results. It's super strange, right?

Speaker 1

Yeah, I wrote this article for the GPT-5 release about models self-improving for coding. I basically asked GPT-5, “What are tools that would be useful to you to be a better software engineer?” It was like, “Well, you know,” and gave a list of about 10 tools. I said, “Okay, implement them,” and it wrote all the tools.

Then I asked it to do the same task I'd done before, but with those tools. It went through the old task, and I asked, “Which of the tools did you use?” It said, “Oh, I didn't use any of them.” I asked, “Why did you not?” It said, “To be honest, I don't really need the tools. I can just do this task.”

I think that's a good metaphor for the trend of the models: they're going to use less and less of these custom-made tools to fix today's issue. The things we can bet on—and I'm curious to hear your thoughts—are that they'll always have some sort of test runtime. I don't think there will be a world in which the model doesn't run tests and say, “I'm sure this is going to work.” The other one is that there will always be some sort of infrastructure as code to handle the deployment side. Whenever there's a runtime issue, they'll need to understand where they're running. You can put them in a box, and having an actual Dockerfile and whatnot is helpful for explaining what they have access to.

What do you think are other things that you don't expect the model to have built in, but that you still want to expose to it? We can assume it's going to test, and we can assume it's going to have some definition of its environment. Are there other things that come to mind?

Quinn Slack

I think tests are a big one, and there are many different kinds of tests. We had sub-agents in Amp, among the first to come out with this conception of sub-agents: a separate context window and a separate, more curated set of tools. I think there's a lot of potential to take a tool like test. Right now, you invoke it through the Bash tool with some complex invocation. Too often, it'll run all of your tests, which is noisy and takes a long time. If you're in your editor and have something nice set up, you can hit a hotkey and it'll only run the tests you need at your cursor. Giving the LLM a tool like that seems to have a lot of potential.

That could even potentially be a smaller, fine-tuned model for that task. It could be multiple models based on what projects or stack you're using, and that could eliminate a lot of the confusion. Even with good agent.md guidance about how to run tests, I still see with Amp—and I think we've tried to make this really good—that it only gets it right maybe 90% or 95% of the time. Sometimes it'll run the wrong tests or won't escape them correctly. I think we can eliminate that with a sub-agent. There's so much more potential to go deep in areas like that. For every language, it's a little bit different, so you have to handle all those cases.

Thorsten Ball

Do you feel like that will just be built by each company on its own, or do you think there's a single default that you guys are going to build for that, one that's going to be effective for most codebases and test structures?

Quinn Slack

This is where scale helps, and we have a lot of scale. Increasingly, we're able to see that in this framework, for example, the standard Go unit-test package is easy. Vitest in JavaScript is easy. Once you start getting into more of the long tail, it might have to just fall back to a really good model. But I think we could probably make something optimized for some of these more popular unit-testing frameworks.

It's a combination of deterministic and nondeterministic stuff. Right now, in my VS Code, I can hit Command-T if I'm positioned in a test file inside one of those test blocks, and it'll only run that one. Even that is a benefit. Now I'm mostly bottlenecked by Playwright.

Thorsten Ball

Yeah, it just takes a long time, man. But the crazy thing is that the vast majority of developers who are building web applications with coding agents do not have Playwright. And if they have it, it's set up in such a shitty way that it cannot really log into their app. They don't have any pattern for that.

So even something like that is another example of a sub-agent that can go and try this basic end-to-end testing flow, described in natural language, with the running application. Wouldn't it be great if they could also do it in parallel? There are all these ways that you can improve.

Quinn Slack

That's a great example. Touching on this, we have coding agents: they are productive and add value. We cannot assume that everything around the agent—in developer tooling or codebases—will stay static. People are already adapting their codebases to be better used by agents, or adapting their tooling to be better used by agents, with more descriptive help text or whatever it is.

I don't know, we should have a counter for it, but everything is changing. I'm saying this again, but we cannot build right now with, “This is the tool that's going to stick around,” given that all of the codebases, processes, and developer tools will stay the same. We have to assume that this stuff will change, too. We have to stay nimble, so we have to make short bets, or small bets, and try to get ourselves forward in small steps, while always being reactive to this stuff.

Again, let's not use Bazel, but I think Playwright is a good example of how the feedback loop is incredibly important to working with these agents: the agent can see whether what it's doing is actually working. What we've seen people do now is, instead of having the client log, browser log, and database log, have 1 unified log, because then it's easier for the agent to look at it and make sense of it. It turns out it doesn't have to be nicely formatted; it can be verbose. You can just have JSONL outputs and whatnot, because the agent can understand them much better than a human can.

I think that's a little preview of more things that we will see. You'll be like, “Wait a second, this is not made for human consumption anymore. How can we optimize this for agentic consumption?” Then maybe the game changes.

Thorsten Ball

There are some things that we now get—for example, in my pytest suite, I have a hook to record HTTP calls. Especially for inference, you can't really mock it. When we do a classification or things like that, you just need to see what happens. We save the whole interaction, and then the model can see what the API returned in much more detail and reference it in the future.

So when you add a new feature, it can look at the test and see what the API usually returns. It's like, “Okay, it's going to have that key,” and things like that. I think there's more of that to be done.

There was maybe also a time when having console logs was really bad. I think now there may be a console log that's only funneling not to the actual console in the browser, but to some way for the agent to see all the details of everything that's happening. What I haven't figured out is how you instrument that. You can't put a whole bunch of console logs that go somewhere else in the code, because then you're also polluting the model's context window, right? So you need some other way to do it. But I think the more you're logging, the more the model can self-iterate.

Quinn Slack

And you just described 5 approaches that seem absolutely worthwhile to explore to improve how coding agents work.

Let somebody do it. We can do some of it at Sourcegraph Live, but we cannot do all of it, so somebody help.

But again, the world around us is also changing. José Valim, the creator of Elixir and, you know, a contributor to Phoenix—I cannot remember the name—but basically, they have new framework tooling out for Phoenix. It is about, well, what if you build a framework for an agent, too? What if the agent is integrated into the framework?

Then, if the application fails to run, you can ask the agent, which has access to all of the context. That is going to happen more and more, I think. A lot of developers will build things because they are fed up with copying and pasting stuff around, so we are going to see this in developer tooling.

Thorsten Ball

It is true, but what we have said is that we are especially building Amp for the people who do want to move. That has been so liberating. I think that is the great thing about what you are seeing in the market today: you have all these companies that are so AI-first and just use it and integrate it.

Then you go on Hacker News and it is, “I have never gotten a single good result from AI.” I am like, “Well, obviously that is not true.” The extreme is definitely true, though. To me, that is kind of the thing. The people who are spending $100,000 a year on Amp with 2 people are obviously getting value. It is not like they love burning money.

But for the people who are negative, that is not always true, because it is easy to be negative and it does not cost anything to leave a bad comment. What is going to be the thing that forces the rest of the market to say, “Whatever, man. I just get on Amp and make that work”?

Quinn Slack

They just have to see this work once or twice. We have been in developer tooling for a long time with Sourcegraph, and it has always been hard, for the last 10 years or so, to get a company to adopt a developer tool that does not immediately fit into their codebase. The codebase is the standard. Everything else has to adapt to our codebase, our processes, and whatnot.

What we are seeing now with agents is that as soon as somebody has seen what they can do, they have such a multiplying effect and bring so much value that people are willing to adapt the codebase to them. This is the first time in how many decades that people are saying, “Maybe our codebase is wrong. Maybe we should change the way we develop code to make more use of this.”

I think people have to see this, and then the agents will pull them along. The value this brings will pull them along.

I was on the board of a company called Launchable, which was founded by Kohsuke Kawaguchi, who built Jenkins. The idea behind Launchable was, “Instead of running all of your tests, we will use machine learning to figure out what tests are impacted by your pull request and just run the small subset of them.”

Thorsten Ball

Yes. I think what we found then—the company got bought by CloudBees—was that a lot of companies would go in there and say, “How can we trust it? Let us do a proof of concept.” Then you do the proof of concept, and it works great for the subset. But is it going to work for the whole test suite?

Then you go through the whole process. With coding, for some companies, they see it work on 1 task and think, “It is worth trying on every task.” Then there is another subset of companies that say, “It works a little bit on the front end, but it does not work on my Java service back there, so I am not going to use it at all.”

I have not quite figured out what the market pressure will be to make those people move along. But, as you said, for some people it needs to work once. Maybe for other people, it has to be 1 task that they always feel.

My 1 task that I always use is this kind of dream product we have built, which is an MCP playground and tester. I have a task that is basically, “Add yellow mode.” It lets a user toggle between auto-running and approval mode, which sounds easy, but it is actually quite hard work without LLMs. You have to stop inference to approve a tool and then run it again.

Every model was failing until GPT-5-Codex and Codex CLI. That was the first time I got it in 1 shot. It made the whole thing. I wonder if everybody should build 4 or 5 tasks that are like, “If you can actually do this end-to-end, then I am in.”

But I feel like people are still in denial that this is going to work. They do not want to have the conversation at all. If you look at that early-adopter and laggard chart of technology adoption, there is a reason why the early adopters are the tiny little start of the curve—3%.

So many of these arguments are people saying, “What if we made a product for the early adopters, but somehow made the laggards adopt it early too? Why are we not going after that big market? It is the vast majority of the area under the curve.”

It is like, “Because they fundamentally do not want what you are building. Maybe they should. Maybe they are going to realize that, but you are not going to make them realize it. If you waste your time trying to make them realize it, you are going to be trounced by people like us who are only focused on the early adopters.” It is a total mindset shift.

Quinn Slack

And if you are just focused on building something for early adopters, and you literally do not care—you set up your entire business and product so that you do not have to care about the laggards—you can do a much better job. That is what we are experiencing now.

Let us talk about the outer loop, because I think that is the next step, at least for me. I think the coding agents themselves degrade on a task-by-task basis, but then there is pull request review, where GitHub is so slow and clunky. It is ordered by file, whereas I think we should get to a world that is more semantic.

It is like, “These are really the 50 lines of code that matter to look at.” Everything else is—

Thorsten Ball

It is fine. You can skim through it. Speaker 1

How do you think about that, especially when you think about async agents? There should be an easy way to spin them up, which I think is fairly clear, but I am not sure there is yet an easy way to catch up on what they are doing.

When I used Conductor, like VIP Kanban, I spun out 5 or 6 of them. I worked on them and jumped between them, and then my wife said, “Let us have dinner.” We had dinner, and I went back and thought, “What the fuck is going on here again? Which one is doing what?”

It is hard to see at a high level what each of them is working on and where it is getting blocked. Have you seen anything that works there? Have you been thinking about building any tools in that space?

Quinn Slack

I agree. I feel this too. With our internal experiments, for example, there is this idea of, “I just spun up 10 agents, they work, and I control them.” I think TV is doing this, and he has a whole workflow around it that seems to work for him.

But for me, I guess I am a one-tasker. I cannot do this. I cannot control 5 agents at the same time. When I do it asynchronously, I realize that I need to be really strict about how I review what they have done and that I do not jump between them.

It is also about making sure that you do not miss anything. I have spun up so many agents and then have not checked back on them because I forgot that they were running. That is something you need to build into the product. But I do not think it is figured out. There is still so much to do.

Thorsten Ball

Yes, it is wide open. We think of it right now like this: if you are playing chess, you can play on 1 board at a time, or you can be one of the people in New York City’s Central Park who plays against 10 different tables at once. They sit down in front of a table, get oriented, make a move, and then go.

That is what we are trying to build. It turns out that even if you have a coding agent running in your editor or in the CLI and it makes a big diff, you still have to understand it. That becomes even more important when you have a lot running in the background.

Yeah. We want to make it easier to orient yourself to what changed. There’s a lot of stuff outside the realm of coding agents that would help, like having a deployed preview consistently available so you could just click through it. And then we want to make it faster for you to make a move and get on with your next thing.

You want a UI where, at a glance, you can see—I don’t know what it is yet—but in the console you can see what the agent actually did without having to read through the summary: “Finally, we have it,” and blah blah blah, stuff like this. But to come back to your question about the outer loop, I think—and if Beyang were here, he would talk for a long time about this, because he’s passionate about it—the inner loop has changed a lot. You know, write, test, review, and whatnot; now you review a lot more code.

For me, for example, we don’t do any formal code reviews on the Amp team, but it doesn’t mean that code isn’t reviewed, because we use Amp to write 80% to 90% of our codebase. That means everybody should review the code that the agent wrote, so it’s reviewed by at least 1 person, right? And that’s not reflected at all in GitHub yet. GitHub is still based on this other mode, where you tag somebody. But then it’s like, “Well, I actually went through 2 agents to produce this code, and I reviewed it 3 times. Do I now tag 5 other people?” Right now we’re stuck in this mode where people would say yes, but I don’t think it’s going to hold that much longer.

The other thing I noticed is merge conflicts. I used to have very few because I knew what I was working on, and if I was doing multiple tasks, I knew how one was going to impact the other, and I was going to build toward it. With agents, especially when you run them in parallel, they just start changing whatever is convenient to them, and across the agents they’re changing the same thing.

So one thing we’ve been thinking about building is better cross-agent orchestration of these changes. I built, for the GPT-5 post, a task manager that’s CLI-first. Basically, any agent can append the files it’s touching, and then it can read what files other agents are touching and see those diffs to implement them back.

But then I think the question is, “Well, maybe what they’re doing now doesn’t end up being the final thing, and now you’re wasting all these tokens reviewing all these changes before you review.” I think at this point the question is: Is Git well designed for this future world that we’re going into? Everything is back on the table.

Maybe 5 years ago, there were a couple of YC companies doing, “Oh, we’re a new version control system.” And I was like, “Look, man, I’m not really interested in listening at this stage.” Same with programming languages. When Chris Lattner even started working on Mojo, I thought, “Okay, because of AI, I understand why you need to build a superset of Python.”

And I think now with agents it’s maybe clear why TypeScript should win, because type-checking is very good for the model to do self-improvement. What are the other things? I think the interesting flex here is that people assume coding agents need to meet the bar of writing the exact same kinds of software to the exact same standard. That is not necessarily an assumption that end users and consumers will apply.

If they have software that’s much faster, cheaper, and much more personalized, and if they can conjure it up on their own, then, yeah, you’re going to tolerate it if the loading state of this thing doesn’t quite work correctly. So changing user demands and standards is an interesting thing that you can flex here. What do you think about that?

Speaker 1

We’ve been thinking about enterprise software moving more toward user-generated content. Expenses are a great example: Why are there so many expense tools when the core action you’re taking is to take 1 line of an expense and tag it with different things? But then you have to set up all these categories and whatnot, versus just generating it for my company and for each team separately, because they have different needs.

To me, it feels like more and more of that will become true. Then the real value is what the underlying data store or data stores are that you’re feeding into this. I know some enterprises are already building internal versions of Lovable, basically, where each employee can create a simple tool, connect it to internal data stores, and they might be the only users of it. Nobody else uses it.

I’m curious how you guys think about this. I know that bold that new, for example, now has Claude Code integration. Where do you see the line moving between software engineers building software—obviously Amp is a great tool for that—and going more upstream, where nontechnical people can also plug into the code and build things on top of it? It feels very different, but also very similar in terms of the challenges you need to solve for.

Thorsten Ball

I think this idea of “nontechnical” is the wrong way to look at it. There are always going to be people who are good at unambiguously specifying what they want out of a computer. We’ve had non-coders, including 1 of our board members, who built something with Amp that replaced a $250,000-a-year piece of software he used for a lot of their internal fund tracking.

He may have taken 1 computer science class. He hasn’t really coded, but he’s a really smart guy, and he knows how to unambiguously specify what he wants to his CEOs, certainly, and now to a computer as well. So if you can get people like that a tool that’s really powerful, they don’t think of themselves as nontechnical people. I think that’s just such a bad mindset.

We want to build for the power user, and if that person has not been a coder but can pick it up really quickly, that’s great. Again, we’re completely focused on the people who know how to and want to get the very best out of this, and who want the agent to win. They aren’t trying to be like, “Oh, nah-nah, hey, it didn’t do this thing. Tell me when it does.”

Quinn Slack

Yeah, we had this at the start a lot with people. Whenever you have an AI tool, I think there’s a natural tendency among engineers to get it in a gotcha moment. “Oh, yeah, I asked it this and it didn’t know this.” Are you trying to get something out of it, or are you trying to get it to fail? It’s not worthwhile to build for somebody who doesn’t want to fail.

If you fast-forward to how the world is going, you’re seeing already, over the last few years, companies really slowing down their growth in engineering headcount. This is a global phenomenon. You’re seeing engineers, like here on the Amp team and at other companies that are using agents really heavily, cutting out the middleman.

They’re putting the people who are building the product closer to the customer, because you can hear an idea from a customer literally in the meeting, kick off an agent to go and build it, and then have a first draft of it. Overall, the person who’s using the coding agent is getting so much closer to the problem. They’re also going to share more in the rewards from solving the problem because, without needing to share the profits with everyone else, there’s naturally more to go to them.

If you fast-forward this, it’s not that the firm or big companies are going to completely go away, but you’re going to have people who have an incredible vision in their head, are so close to the problem, and have an incredible incentive to go solve that problem. Equip them with a coding agent. If you can build a coding agent that those people want, that is way better and more valuable, and you’re creating more value.

You’re allowing more new things to be created in the world than if you were building a coding agent for the median developer that makes them 30% better. So that’s who we’re targeting, and I don’t think that will necessarily look like vibe coding. Vibe coding is this really unproductive thing to discuss, because everyone has a different definition of it, and too often it’s having the agent write code with poor feedback loops and poor quality control. I don’t think that’s valuable.

But it’s giving that person the ability to build something truly great really fast when they’re so incentivized, and they will have every desire for it to work well. And I know we’re getting close to time, but there are a couple of things I want to touch on. So, Thorsten, I was reading through your blog. You left Sourcegraph a year and a half ago, and then you joined back.

Thorsten Ball

Good job, Quinn, bringing him back home.

Quinn Slack

Thank you, Thorsten. When you wrote a post about leaving, 1 thing you wrote was that when you first joined in 2019, 1 thing that Quinn told you was, “Hey, Sourcegraph is your playground. You have skills and talents, and I want you to use those skills to move the company forward.”

How do you take this idea of the power user getting close to the customer and how people are going to build teams overall? There used to be engineering and product, like you were saying—the triangle. That’s kind of going away. What are the types of people that you think are going to be most successful? How should people think about structuring teams? Obviously, you’re doing this with Amp in a way, right? You’re building a subteam and subproduct within a larger company.

Any tips that you have for other founders and executives? Thorsten is incredible, and Amp would not exist in any way without him. He has a strong internal constitution for how he uses it and what’s real and what’s not, and it’s so easy to get carried away with the hype and the possibilities, especially when you see a lot of other smart people getting carried away by it.

Thorsten has this incredible ability to stay grounded. With everything changing so fast and this being such a hype cycle right now, that’s really important. Also, just this first-principles thinking: how we’ve completely rethought how we build everything in Amp based on how we should actually do it, rather than what has come before.

Thorsten is the rare person who’s been at bigger companies, who’s seen how Sourcegraph builds enterprise software—not the Google way, but in a different way—and has taken the parts of it that work and not the parts that don’t. All of that combined with someone who’s an incredible engineer, writer, and communicator is a really powerful combination.

So find those people. When he rejoined, I said that he was the dictator. That made him feel really uncomfortable, as you can see. I hope you cut to his face. But that’s exactly what you have to do: put so much trust in people like that.

That also shows everyone else in the company that they can do crazy stuff. They can go way beyond, take it to the extreme, and make mistakes, and that’s still okay. We’re not trying to build something that’s going to go really big in its current state. Amp is growing incredibly fast, but the most important thing is that we’re building the coding agent god—that thing in the future—and that’s something that we’re all in search of.

None of the mistakes or successes in the month-to-month timeframe really matter. It’s all about getting ourselves on the right trajectory, and you’ve got to do crazy stuff. Equipping Thorsten to do crazy stuff and take the ideas that he has and scale them up with all the reach that Sourcegraph has—that’s been my goal.

On the first-principles thinking, how do you think about that? There’s the world of evals and there’s the world of vibes. How do you approach it? How do you look at the product and say, “Okay, this is good, this is bad, and this is what we need to improve”? Is there something formal that you guys use internally, or is it mostly you as the dictator directly?

Thorsten Ball

Two-part answer. I think the first part also answers the other question a little bit. What I’ve seen become more important, or the shift I’ve seen, is that I said the triangle was the product manager, designer, and engineer. I think as an engineer—or any of the three—you now need to know a lot more about the other parts.

I don’t think there’s anyone anymore who just types out a spec or turns a product PRD into code. You need to be aware of business, be aware of product, and have some taste for software. Otherwise, I think the value of your work will diminish over time, because the pure typing out of code—for most code, with exceptions like John Carmack and whoever else—the value will diminish.

We’ve already seen it. Compare the value of a GitHub contribution chart today to what it was 2 years ago. To come back to the second part, vibes and whatnot: I think we don’t have any set evals. We don’t, and we haven’t had them.

This was controversial up until a week ago—or 2 weeks ago, I think—when Boris from Anthropic said they don’t have evals for the coding agent either. But we don’t, and we haven’t had them.

I’ve built evals before. I’ve fine-tuned models before. I know that they’re good. I love evals. I was addicted to LLM-as-a-judge. I wrote about LLM-as-a-judge. But for a coding agent that’s supposed to work in many different codebases, with many different types of prompts and many different types of tasks, it’s a time investment that we cannot afford with everything changing and having to stay fast.

If you ship 20 times a day, you will get a lot of good feedback. I swear to you, I could tune my system prompt a little bit, and by this evening people on our team would say, “Why does it call this tool so often? What’s going on? What did we ship?” That’s incredibly valuable feedback when people talk with the product and use it all day.

How do I make these calls? I don’t know. I think it’s experience. I think about software a lot. I love using software. I listen to a lot of business podcasts and read a lot about business. I listen to a lot of software podcasts and read a lot about software.

Then I try to project: What does the business need? How can we get growth to 10×? How can we get our users to 10×? How can I use my engineering capabilities to serve as a function of the business and reach those goals? How can I organize the team, or get the team, to help me reach those goals or reach them together?

It’s hard to explain, but I feel like this year, truly here at Sourcegraph, everything I’ve learned over the last 15 years of my career is coming together. All of the hours spent listening to the Acquired podcast help me as much as reading Hacker News for hundreds of hours and writing code for thousands of hours.

With code now being a tool that you can wield much more easily, much faster, and much more often, I think it’s become much more important to think about how you want to wield it, when, and for what reason.

I think the fact that it’s hard to explain is a great explanation for why you just cannot create these things in one shot, because there are a lot of implicit preferences.

Speaker 1

Awesome, guys. Anything to wrap up? Are you hiring? Who should reach out to you? Requests for startups? What should people build that’s going to be helpful to you guys?

Thorsten Ball

All of the above. I don’t know. We’re always interested in talking to fellow engineers who are interested in agentic programming and figuring new stuff out. We want to hear from them about what works and doesn’t work.

We’re always willing to hire people with exceptional talents who are fully committed to this and realize that programming is changing a lot.

Quinn Slack

We’re just so excited. We feel like kids in a candy shop.

Thorsten Ball

Yeah, that’s true. We get to go build the future of coding. It feels like the final boss.

Speaker 1

Nice. Thank you guys for coming on. This was fun. Thank you.