无代码即代码:Zapier CEO Wade Foster 谈无头工具、Zapier MCP 与 Automation Bench
- Wade Foster 自上次做客以来最大的更新是:AI 体验最终会长什么样的模糊性已经消退——大多数人似乎都选定了一个单一的“日常 AI 驱动器”(Claude Code、Cursor、ChatGPT),其他工具则需要接入其中。 这使无头架构成为胜者:“如果你真的想成为某人日常工作流的一部分,就得接入那个人的核心日常驱动器”,而 Zapier MCP 正是公司押注的这一层。
- 模型在真实知识工作上的能力远未饱和:在 Zapier 覆盖约600项营销、销售、人力和运营任务的 Automation Bench 上,新的最高水平——“Astra 出来了,GPT-6”——准确完成率也只有约40%,而“Gemini 3.7”“表现相当不错,但成本只是其一小部分”。 这条成本/性能曲线正是 Foster 认为组织希望避免依赖单一模型实验室的原因,也解释了为什么需要第三方运行框架——这些实验室“无法销售彼此的 token,也无法销售开源模型的 token”。
- 本期的核心论点是:“新无代码就是代码”(the new no-code is code),但 Foster 表示,客户目前交给 agent 的工作中约80%“实际上应该使用老式的确定性代码”,AI 只应处理真正需要推理的部分。 正在形成的模式是:agent 构建并维护工作流,确定性代码执行工作流,agent 负责排查故障;即便模型继续变强,这种方式在确定性任务上也更便宜、更可靠。
- Foster 认为,真正的竞争对手不是 AI 创业公司的“千篇一律”,甚至也不是模型实验室,而是不采用:“去问问普通 AI 工具用户,他们坦率地说其实没做多少事。” 他借用了 Paul Graham 在 YC 的建议:“你不是在和 Larry 和 Sergey 正面对打……你是在和一个想升职的中层产品总监正面对打”,并估算这块机会:自动化市场“可能比我们15年前出发时设想的大1,000倍”。
- Zapier 的下一个潜在产品可能是 AI 辅助的工作流发现:Foster 每周运行一个自动化程序,收集 Gmail、Slack、浏览器、Cursor 等各处的信号,再提出值得构建的工具——建议只需“足够好到50%”就能触发头脑风暴;他表示,这“很可能”会被产品化。 Nathan Labenz 解释为何日志在这里胜过屏幕录制:“AI 非常擅长读取日志。”
- 在商业化方面,Foster 表示,“大多数情况下,按席位定价已经死了,或者正在走向死亡”,市场正分成按使用量收费(Sam Altman 所说的按需智能公用事业)和按结果收费(客服工具按解决工单收费)两类——但多数产品“距离真正交付结果还差一两步”,最终又被拉回按使用量收费。 Zapier 内部有工程师每月花费3万美元购买 token,目前没有硬性预算,只有靠自我约束的仪表盘;Foster 预计,与一个人能否把支出转化为产出挂钩的 token 预算,最终会成为“AI 素养”的一部分。
- 组织设计上的信号是:Foster 让首席人事官 Brandon 负责 AI 转型,因为瓶颈已经从工具采用(ChatGPT 发布后约1年内,近100%的员工开始使用 AI)转向人的问题——重写职位描述、薪酬设计和再培训——但他坚持说,“同样完全可能由 CMO 负责”。 与此同时,Slack 默认公开,因为 agent“能看到上下文时效率高得多”;面对 AI 垃圾内容时代,公司要求共同署名和共同负责——“使用 AI 不是问题,问题在于质量低。这才是要对付的事。”
- 在安全方面,Zapier 已经托管凭证15年;Foster 表示,AI 让漏洞发现和修补都更容易,但他还不确定最终会如何平衡。 Zapier 也希望尽早获得最新、最强的能力,其他处在相同位置的公司也会如此。
1. 日常驱动器胜出,无头架构成为制胜策略
- Foster 开场承认,18个月前,“我觉得我们当时还不知道……这些 AI 体验最终会呈现什么形态”。根据 Zapier 自身的使用数据,变化在于,大多数人似乎已经选定了一个日常驱动器——Claude Code、Cursor、ChatGPT——而像 Zapier MCP 这样把上下文带入该驱动器的工具,正是“知识工作的发展方向”。他明确划分了市场:一类平台要求用户在自己的界面上构建 agent;另一类则是“像 Salesforce 这样的公司,会说‘我们是无头的,你想把它带到哪里都可以’”——“后者才是制胜策略”。
- Foster 自己也在践行这套判断:他的日常驱动器是 Cursor,通过 MCP 接入 Zapier 内部运行框架——以虚拟文件系统充当上下文层,在其上运行自动化,并部署应用。
- 对于 Labenz 总结的 Tasklet 创始人 Andrew Lee 的观点——只有少数公司能赢得横向的“模型麦加套装”层,Nathan 如此概括——Foster 表示认同:没人想受制于单一模型实验室。但他补充了一个变量:运行框架现在可以自己构建工具,而问题在于,“它确实能构建,但你是否应该让它构建?因为那意味着你要接受一定程度的维护、可靠性和正常运行时间责任。”模型实验室在应用层有先发优势,但“它们无法销售彼此的 token,也无法销售开源模型的 token”,因此第三方仍有空间。
2. Automation Bench:40%就是前沿,工具带来的提升将成为下一条主线
- 这个基准测试的任务足够贴近真实工作,约600项任务中包括:“我们刚刚签下 Meridian Core 平台交易。把它标记为已赢单,并根据路由政策将其转到正确团队的赢单通报;对照账户层级表确认账户,必要时转换币种,并检查是否存在未解决的支持升级事项。”Astra(GPT-6)以约40%的准确率成为新的最高水平,但“这个基准远远没有饱和”;与此同时,Gemini 3.7 以低得多的成本取得相当不错的表现,迫使现代组织在每项任务上权衡成本与效果。
- Labenz 最尖锐的问题是:让 Claude 接入 Zapier,能把得分提高多少、把 token 用量降多少?Foster 有意卖了个关子:“Automation Bench V2 发布前,你得等着。”新版本会衡量模型获得工具后的得分提升、成本下降和速度变化。Foster 表示,Zapier 正在确保将其与 agent 一起安装后,产出确实优于单独使用模型。
3. “新无代码就是代码”——但80%的 agent 使用应当是确定性的
- 本期最关键的数据是:观察客户使用 agent 的方式后可以发现,“人们让 agent 做的事情绝大多数,实际上是80%,本来应该使用老式的确定性代码”。只有需要推理的事情才应交给 AI 推理;对于某些任务,Foster “很难想象一个确定性代码不更可靠、更便宜的世界”,其余部分才交给“机器里的幽灵”。
- 界面正在发生变化:“构建无代码工具这件事,在我看来已经过时了。新无代码就是代码。”但人们“仍然非常需要工作流的可视化”,既为了验证,也为了留档;因此现在的方式是直接和 agent 对话,由 agent 完成修改,而不是人自己配置一个个方框。
- 正在形成的故障闭环是:工作流逐渐固化为确定性代码;一旦出错,agent 负责排查并修复工作流或具体实例——“agent 几乎接管了人类构建和维护工作流的角色”,但实际运行的部分仍然是确定性的。Labenz 担心 Meridian 的例子只是个案,Foster 则纠正说:“如果你是一家运营良好的组织,你每天、每天都在成交”;而在消费者规模上,“这些任务不可能让人类介入”。
4. Zapier 的递归闭环:5个 agent、点赞/点踩与数据护城河
- 在 Zapier 的自动邮件客服项目中,5个相互独立的 agent 会评估每个排障案例;“当5个 agent 中有4个倾向于达成一致时,那很可能确实就是问题所在”。人类会对结果进行审核,选择批准或拒绝并说明理由,这些理由再反馈回系统——在不同模型和提示技术之间持续做纯粹的爬坡式优化。
- Foster 对防御力的解释是,护城河来自别人无法复制的数据:“我们非常擅长自动化。我们有大量关于自动化需要什么的数据……我们接入一切。”基于这些数据构建的产品,“必须明显优于那些无法接触这些数据的人所能做出的东西”。他用日常例子说明:接入你的 Gmail 后,模型写邮件会更好,“因为它看到了你写邮件的方式”,而且无需任何调优。
- Labenz 继续追问是否存在更复杂的检索栈创新,Foster 则把它说得很简单:“归根结底,很多事情并不花哨。就是拿一个例子,问你喜欢还是不喜欢,然后不断冲洗、重复。”
5. 真正的竞争对手不是模型实验室或克隆产品,而是不采用
- Foster 重新提起 Paul Graham 在 YC 时代的建议:“你不是在和 Larry 和 Sergey 正面对打……很多时候,你是在和一个想升职的中层产品总监正面对打。”如今情况又变了:“OpenAI、Anthropic 现在已经是大科技公司了……它们会在其他地方做出好产品,但不可能什么都做。它们就是做不到。”
- 他坚持要同时保留两种判断:一方面,确实要和“千篇一律的 AI 创业公司海洋”拉开差异;另一方面,“去问问普通 AI 工具用户,他们坦率地说其实没做多少事”。很多人的体验可能只停留在“Google 默认搜索里的 Gemini”或“工作中的 Microsoft Copilot”,而如果你整天泡在 X 上,这个现实“太容易在信息洪流中被忽略”。市场规模“比我过去想象的大了几个数量级……可能比我们出发时设想的大1,000倍”。
- Labenz 坦言,这是一个值得特别标注的真正认知转变:“我肯定有一件事判断错了:5年前,我预期整个经济的做事方式会发生更多变化,但实际并没有。”Foster 表示认同:“生活其实差不多。”朋友们会用 ChatGPT 规划饮食、旅行和锻炼,但“肯定完全没有利用 Astra 级别的模型能力”。
6. 填补推荐缺口:观察你如何工作
- Zapier “直到今天”最难解决的问题仍然是推荐:Foster 坐在任何人旁边,都能找出6个他们会立刻想要的自动化,但如何把这种“专家坐在你身边”的体验嵌入产品,始终存在缺口。他的答案是:“观察我在 Gmail 里做什么,观察我在 Slack 里做什么,观察我在浏览器里做什么……然后告诉我,我应该改做什么?”从年初开始,他每周运行这套流程,现在几乎每周都会有新的工具和系统自动化他工作中的一部分;这些想法原本“被习惯困住、潜伏着又消失了”,因为“人是习惯的动物”。
- 机制并不性感:Zapier MCP 和已连接的工具每周收集一次事件流,再提出值得构建的内容;用户做出反应、进一步调整,然后说“去把它构建出来”。一半的难点在于触发反应——建议只要“足够好到50%,能让你说‘我明白你想往哪里走了’”就够了。被问到是否会产品化时,Foster 回答:“会。我认为它很可能会以某种产品形态出现。”
- Labenz 追问为什么选择日志而不是屏幕录制——一次安装,就能自上而下地看到每一步点击;Foster 坦率地说:“我们擅长 API……这只是一个自发出现的实验。”Labenz 给出了更好的解释框架:“这反映了 AI 智能的异质性……AI 非常擅长读取日志。”
7. 定价:席位已死,token 需要预算
- Foster 的判断对大多数产品成立:“大多数情况下,按席位定价已经死了,或者正在走向死亡。”分叉在于:商品化产品可能更接近按使用量收费,即“Sam Altman 所说的按需智能”公用事业;更偏企业服务的产品则可能尝试按结果收费,例如客服工具按解决的工单收费——但前提是成功有“清晰、双方认可的固定结果”。如今多数产品“距离真正交付结果还差一两步”,因此又被拉回按使用量收费;无论哪种方式,“本质上都是在销售工作”,买方会为待完成的工作做预算。
- 对于模型实验室的价格歧视,Labenz 提到 Claude Max 或 OpenAI Pro 计划相较 API 可能存在约20:1的 token 优势。Foster 没有顺着监管干预的方向走,而是倾向自由市场,不过他也承认:“我们很多人经历过、现在仍在经历 Microsoft 的支配,以及它如何利用捆绑销售来排挤更好的产品,坦率说就是这样。”他的原则是:“我的工作是在这个赛场的规则内比赛。”
- 在 Zapier 内部,每月花费3万美元购买 token 的工程师属于“有点偏离常态的例外”;公司目前使用仪表盘让员工自我约束,而不是给个人设定预算。Foster 的第一反应是问:“你在做什么?我真的只是很好奇。”他发现其中既有“效率高得离谱的事情”,也有“这件事没必要用 Fable,或者没必要用 Astro”。但在一个接近800人的组织里,Foster 预计 token 预算最终会成为现实,甚至可能反映出谁更能把这笔支出转化为更高产出。
8. 人事运营就是 AI 战略:首席人事官、默认公开与工厂
- 之所以由首席人事官负责 AI 转型,是因为 ChatGPT 发布后约1年内,“几乎100%的员工都在日常使用 AI”,瓶颈因此从采用工具转向人的问题:重写职位描述、重新思考薪酬,以及在团队缩小和扩张的过程中重新培训员工。Brandon 本来就擅长这些事情;但 Foster 强调,“同样完全可能由 CMO 负责……你需要审视自己的瓶颈是什么,然后找出最合适的人”,而不是照抄 Zapier 的组织架构。
- 默认公开因为 AI 获得了第二次推动:“我们的 AI agent 能看到 Slack 内部的上下文时,效率高得多。”管理层发起了一场友好的公开频道占比竞赛,没有强制要求,但 HR 事件和正在处理的安全漏洞除外。Foster 的经验法则是:人们“严重高估了必须私下处理的事情数量,同时明显低估了完整上下文对人和 agent 的力量”。
- 团队形态上,经典的 EPD “基本已经消失”,管理层级被压平,但“强有力的管理仍然至关重要”;每个人都成了“自己的迷你数据分析师”。趋势是“构建生产产品的工厂”——软件工厂和客服工厂让工作流处理更多内部循环,人类则专注于设计这部分流程;每次模型发布,都会进一步减少仍需人类介入的步骤。
9. 凭证、AI 辅助漏洞发现,以及对发出的内容负责
- Labenz 提到 Zapier 可能掌握规模巨大的凭证,Foster 回应说 Zapier 托管凭证已经15年。新变化在于“达到 mythos 级别、能够耐心地逐个问题循环处理的安全模型”:“每个软件项目都有漏洞,只是还没有人找到。”让人安心的是对称性:模型也让修补更容易,因此聪明的公司会“同时用于进攻和防守”。他提到了 Hugging Face 事件,称其“简直像科幻小说里发生的事情”,但对未来仍保留判断:“我的猜测是,感觉上会有一次几乎一次性的重新适应投入”,然后进入新的稳定状态——“我不太确定这一切最终会如何发展”。
- Labenz 认为,对于一家被托付大量凭证的公司,尽早获得新能力的资格可能成为有意义的差异化优势。Foster 的回应更窄:Zapier 希望尽可能早地获得最强能力,任何处在它这个位置上的公司都会这样做。
- 对于 AI 共同署名,Foster 给的是规则而非禁令:“我不反对人们在 Zapier 使用 AI 进行沟通。我真正反对的是低质量沟通”,因为“一个判断力不足的人,可以非常快地产生大量质量极低的内容”。具体要求包括:对自己发出的内容负责;撰写所花的时间要多于读者阅读所花的时间;永远不要通过一个 prompt 把任务责任转移出去;核验细节,因为“摘要的摘要的摘要”开始传递错误信息;同时去掉 AI 痕迹——“‘不是这个,而是那个’、破折号、‘说句实话’以及‘承重观点’”。最终原则是:始终“让你自己的编辑之手握住方向盘”。
完整逐字稿
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to welcome Wade Foster, co-founder and CEO of Zapier, back to the show. When I last spoke to Wade in September of twenty twenty-four, some four hundred thousand customers had already used Zapier to delegate more than one hundred million tasks to AI. And YC President Garry Tan was calling Zapier the AI-powered knowledge worker of the future. Since then, models have, of course, become dramatically more capable, and Zapier has built out a full AI portfolio, including agents, chatbots, an MCP server, an SDK, and an AI guardrails product. And yet, somehow, white-collar work and the world as a whole have changed much less than I would have expected. With that in mind, I wanted to hear not only about what Zapier has built and how it's continued to evolve as a company, but what Wade and team have learned about how businesses across the economy understand and use today's AI tools. At a high level, Wade believes that most people are now settling in to using a single daily driver, whether that's Claude Code, ChatGPT, Grokbot, or in Wade's case, Cursor, and that platforms like Zapier will need to adapt by making their tools available and effective in those environments. Practically, he observes that models still struggle with many business tasks, as illustrated by Astra setting a new high of just forty percent success on Zapier's automation bench, which consists of roughly six hundred knowledge work tasks across marketing, sales, HR, and operations. He also argues that many tasks that people are delegating to AI would be better done with deterministic code, and that for a while longer at least, there is therefore tremendous ROI to time invested in structuring and validating workflows. We then go on to discuss what Zapier is doing to help people recognize exactly what AI might be able to do for them, starting with his own weekly automation, which Wade says they will soon productize for customers, that reviews his activity across Gmail, Slack, the browser, Cursor, and more, and then proposes specific tools and workflows that he should be building. We also talk about how Zapier is implementing recursive self-improvement loops internally and how much value they're finding in running multiple different AIs on the same problem, why Wade chose to put Zapier's chief people officer in charge of AI transformation but wouldn't necessarily recommend that strategy to other companies, how Zapier is moving toward public-by-default communications to make more and more context available to AIs, why they still don't limit individuals' use of AI but have created dashboards to help employees better understand and manage their own usage, how Zapier, which holds a huge number of high-value user credentials, is thinking about security in the context of rapidly rising cybersecurity risks, and finally, how Wade thinks about co-authorship between humans and AIs, with the upshot being that he believes individuals should use AI to help improve their writing and shouldn't be afraid of being pangrammed, but also that it's critical that people be prepared to explain and stand behind the work that they ship. With that, I hope you enjoy this very grounded and highly practical conversation about making AI automation work for people outside the AI bubble with Wade Foster, co-founder and CEO of Zapier. The Cognitive Revolution is brought to you by Mercury, the fintech that more than three hundred thousand ambitious companies and individuals trust to run their finances. I've wired AI into nearly every corner of my life, my email, my messages, my calendar. I even gave Mercury virtual cards to my agents with low limits and category and merchant restrictions for their autonomous use. But still, my AI's access to my financial data has remained limited. With a normal bank, I might export a bunch of statements and have my assistant process them for me. But for real-time, up-to-date information, and certainly for taking any action, trying to get your agent to use the bank via the browser is just too hard, too slow, and too error-prone to be worth it. And that's why Mercury's new conversational interface, Command, is such a big deal. It's built directly into Mercury, which means you get natural language access to your finances without exposing anything outside of your bank account. No exports, no spreadsheets, no pasting your transactions into third-party tools. I really think a lot of people are going to prefer it this way, and it can already help you take actions too, with everything bound by the permissions and approval policies that you've already set up in your account. I am genuinely impressed to see this level of AI integration in banking in twenty twenty-six, and so I invite you to join me in the future. Visit mercury.com to learn more and apply online in minutes. Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column NA, members FDIC. Thank you to Mercury for supporting The Cognitive Revolution. And now, on with the show. Wade Foster, CEO of Zapier, welcome back to The Cognitive Revolution.
Yeah, thanks for having me again, Nathan.
Boy, what a difference a not-super-long period of time makes. I can't believe it. Every time I have a returning guest, it's an opportunity to look back and look at what the state of AI was—the models, outlooks, what the possibilities were, and what the capabilities were at the time. Suffice it to say, obviously, a lot has changed.
We only have 1 hour today, so I'm going to try to discipline myself, not talk too much, and give you most of the airtime. I would love to start off with an observation that, like many ambitious software companies, you guys at Zapier have been really prolific, and I would say you've kind of built everything in the sense that you now have, in addition to workflows, which of course can call out to AIs, agents, chatbots, an MCP, an SDK, a guardrails product, and probably more that I didn't identify or mention.
What have you learned by building all that stuff? What has taken off and resonated maybe better than you thought? What has been slower than you thought? And what's the synthesis view that you have, informed by these different product efforts and their relative successes?
Yeah. I was on 18 months ago, and if I think back to that time, what I don't think we knew yet was what shape these AI experiences were going to take. Were we going to see AI infuse into all the products that we already knew and loved? Were we going to see these new product shapes take off? Were we going to see new stuff out of the labs that was going to be where AI was going to take off?
What would seem to have happened, at least when we look at our own usage, is that most folks seem to have adopted their own daily AI driver tool. Maybe this is CloudCode, maybe this is Cursor, maybe this is ChatGPT—you name it—but that's where people do most of their work. What that means is tools like Zapier MCP, or any MCP server—tools that bring your context and your data into that person's daily driver—that feels like the way knowledge work is moving.
There's a split in the market. Some folks are still trying to say, “Ah, we're going to force people to build agents on our own platform,” or, “We're going to bring them over here, and that's kind of where things are going to get done.” Then you see the Salesforces of the world that are like, “We're headless. You can bring it wherever you want,” et cetera. I definitely think the latter is the winning strategy.
It just feels like that's what we, as consumers of these tools, want, and that's where all the growth is. We kind of want our daily driver, and we want to bring that context in and be able to manage it all in one place. That doesn't mean there aren't going to be tools that call out to all these third parties. I still think there's plenty of room for applications and tools to exist, but you have to integrate with that person's core daily driver if you really want to be a part of their day-to-day workflow. That feels like a pretty big new learning for me in the last 18 months.
That's interesting. So does that mean then—what's your daily driver? It sounds like you're positioning Zapier not as being a daily driver, but more as an uber tool for whichever daily driver you choose to use.
Yeah. I mostly use Cursor every single day. We have an internal tool that a lot of our employees use as a sort of daily harness, but that also has an MCP associated with it. I have that pulled into Cursor, so it's using the data from that.
It has a virtual file system that acts as a context layer, automations that sit on top of it, apps that you can deploy—all the kinds of things that you might expect from a modern AI-capable tool. I just happen to use a lot of that stuff inside Cursor.
Very interesting. Okay, I had a conversation with Andrew Lee from Tasklet. He advanced what I thought was a provocative idea: that, in his mind, only 3 kinds of software companies survive in the big picture, and that everyone is kind of building the same thing, which sometimes gets described as the Mecca suit for models.
He described himself as trying to earn a place in the eventual winner's circle of this horizontal layer that sits on top of models and enables them greatly by providing all these different tools, access points, guardrails, and whatever else the case may be. He thinks that there aren't a huge number of companies that ultimately win in that case, but that layer, even if there aren't too many companies in it, is super valuable because nobody wants to be beholden to just 1 model.
They don't want to be overly locked into a single model provider. How would you compare and contrast your worldview against that summary?
Well, I definitely agree with that last statement. I think it is becoming more and more obvious that you don't want to be beholden to one model or one company's suite of models. We see this with Zapier's Automation Bench. We have a benchmark that measures all these models on automation tasks.
Last week, Astra came out—GPT-6. It's the new state of the art on that model. It performs about 40% of the tasks accurately, which is the highest available. Now, it's more expensive than, say, something like Gemini 3.7, which does pretty well but does so at a fraction of the cost. You have this curve that exists where you're trying to figure out how much you're willing to pay for incremental performance on certain tasks.
As a result, I think any modern organization wants the ability to make those trade-offs, to say, “These tasks are good enough. I can pay this rate and get 100% of these types of tasks completed, but for this type of task, I need to maybe move to a state-of-the-art model to do well on it,” and so on and so forth. I do think that companies are going to look for the Mecca harness or whatever you want to call it—the sort of tool that allows them to swap in and out different workflows of their choice.
To that end, I certainly believe that there's a lot of innovation yet to be had on the application layer. But how those applications are used, I think, looks very different from the last decade. The last decade was SaaS, and there was this explosion of SaaS, but now it feels like there's almost an explosion of headless tools happening. Of course, you have this new thing, which is that your harness itself can build some of those tools, so the build option is more readily available than it was in the past.
What's not as obvious to me is: Yes, it can build it, but should you have it build it? Now you're accepting a certain amount of maintenance, a certain amount of reliability, and a certain amount of uptime that may not actually be the best thing for you in a given circumstance. I actually think there's still a lot of innovation to happen at that application layer that we just haven't seen yet. I think everyone's trying to figure out what that looks like.
I think the model companies have a little bit of a head start, but they're going to struggle because they can't sell tokens from each other or from open source or anything like that. It does feel like there needs to be a third party that helps you wrangle all the capabilities that are out there.
Obviously, Zapier came from a history of very structured workflows, because if you didn't fully encode what the workflow was supposed to do, there was no ghost in the machine back when you started to figure it out on the fly. I'd be curious to hear a little bit about how you would describe the tasks of Automation Bench, where the models are good, and where they're not good.
Then maybe you can describe how usage of Zapier is changing qualitatively. How often are people still doing box-by-box defining of workflows? How often are they prompting an AI that then turns their thoughts into a structured workflow? How often is it happening through an MCP, where the model is deciding whether to use Zapier given a range of options? Maybe there's even more there that I'm not intuiting.
You bet. What Automation Bench measures is tasks like the following. Here's an example from our site: “We just closed the Meridian Core platform deal. Mark it as won and route it to the win notice in the right team per our routing policy. Confirm the accounts here from the account hierarchy spreadsheet. Convert the currencies if needed, and check for any open support escalations.”
We have probably 600 tasks of that variety that describe a normal knowledge-workflow task across a wide variety of disciplines—marketing, sales, HR, operations, you name it. That's what it's trying to measure. As you can see, the models are getting better at it, but this is by no means a saturated benchmark yet.
What we're doing at Zapier is making sure that when you install Zapier alongside your agent, you're actually getting better output on these benchmarks than you would if you were just using the models alone. The reason you do that is you're teaching the model what folks are doing in their daily driver. They're saying, “Hey, I want you to go build a workflow.”
They're calling Zapier MCP, and it's going to say, “I'm going to build out that workflow, and in some cases I'm going to write code to actually complete that task so that it is deterministic,” which means lower cost, better reliability, and so on. Then I'm going to invoke an AI or build an agent for the parts that really require reasoning. When we look across even our own customers' usage of agentic products, the vast majority of what people are using an agent for—80%, in fact—probably should be using old-fashioned deterministic code.
You really only want the AI to reason over the things that you need it to reason about. I still think there is a huge amount of work happening inside these production workflows inside a company that you shouldn't actually try to delegate to an AI. We'll see how long that lasts. Obviously, AIs are getting better and better, but I struggle to think of a world where there aren't certain jobs for which deterministic code is still more reliable and cheaper, and certain jobs that it cannot do.
For those, you need the ghost in the machine. You need the AI that can tackle those tasks. I think the right thing is that you're trying to teach the agent how to go do that on its own. When you talk to it, it goes and builds those things in an optimized way instead of saying, “I'm just going to build an agent that runs agentically every single time.” It has a better sense of which is the right tool for which job.
In that one example you gave, if I understood it correctly, first of all, it sounds like there's probably a ton of context that comes into the test environment with that short prompt, right? It has to navigate, find the policy, and parse that, and all those other things require additional information-finding and understanding.
It also struck me that that sounds like the kind of thing that only happens once, or you might acquire a handful of companies, but you're not going to acquire the volume of companies that one would traditionally associate with a Zap. How are you seeing usage change? Or what advice would you give if you used to think, “For me to go to all the trouble to make a Zap for something, I need to at least expect that Zap to run 500 or 1,000 times”?
That example is one that happens every day, if not multiple times inside a company. You close the deal. If you're a good organization, you're closing deals all day, every day. There are so many workflows inside of a company that are kicking off all the time, perpetually, and in some cases they're happening at a rate that humans can't keep up with.
If you're operating at the scale of some of these companies, they're selling to consumers who are making purchases many times a second. You have to use automation. You have no other choice. You cannot put humans in the loop for these tasks.
Have you seen big shifts in terms of people moving away from building it out themselves and having AI do that? Have you also seen the scale threshold at which people start to think about using automation software come down substantially because AI can do the setup?
I think the big new opportunity is to have the AI do the build. It's able to work through the logic much, much faster than a human, and the idea of building in no-code feels antiquated to me. It's like the new no-code is code.
What I think we have learned is that humans still very much benefit from visualization of those workflows. That visualization helps them verify, “Is this thing doing what I intended it to do?” It acts as documentation, so you can share it with other people and say, “Hey, here's the thing I built. Here's what I'm doing.”
If you think of Zapier in the old-school way, as this thing that had a bunch of boxes, and you're coming in and using that to configure it, by and large, I don't think that is the way people are doing it now or even in the future. Instead, they're talking to the agent and having the agent make those edits for them.
The other nice thing that's happening in the future is that you're going to see stuff get hardened into a deterministic workflow. But even when it fails, you can have the agent troubleshoot and just fall back and say, “Why did it fail? What happened here?” Then it can use its reasoning to fix the workflow or fix that instance of the workflow.
You start to see the agent almost take over the human role of building and maintaining it. But what is actually running is still a very deterministic workflow, with AI interwoven in the places where it is most necessary.
Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Athena, the executive assistant company on a mission to improve how people work and live. If you want to increase your impact, you have to free up your time, and that's what Athena does best. They match you with a dedicated, full-time, top 1% executive assistant who can take over your inbox, calendar, travel, and everything else that's quietly eating up your week. Athena is SOC 2 Type 2 certified, so you can rest easy knowing that your sensitive data is in good hands. And as a former AI advisor to the company, I can personally vouch for how much they've invested in AI tools and training. In fact, one of the very best AI users I've ever met is an Athena client who delegated the exploration of AI tools and the development of AI workflows to his EA. Athena clients report saving an average of 15 hours a week, and the average client refers more than two friends a year. That, to me, checks out. I was an Athena client while running my startup, and to this day, I continue to refer friends. Go to athena.com/cognitive right now and get matched with your EA. That's athena.com/cognitive. Give yourself back a few hours this week. Go to athena.com/cognitive and see who they'd pair you with this month. This episode of The Cognitive Revolution is brought to you by OutSystems, the leading agentic systems platform. We've all seen the headlines. Companies are pouring money into AI. But the big question on every executive's mind right now isn't just how fast can we adopt this, it's where is the ROI? We're seeing a real trend toward AI chaos. You've got teams deploying standalone coding tools, random agent builders, and experimental scripts. It sounds innovative, but in reality, it's creating a massive headache. Fragmented tools, ungoverned data, serious security blind spots, and costs that are spiraling out of control. If you don't bring those agentic applications under control now while they're still embedding into your core processes, you are looking at broken systems and damaged customer trust down the road. That's where OutSystems comes in. OutSystems is the leading agentic systems platform for the enterprise. Instead of managing a patchwork of disconnected tools, OutSystems lets your team engineer, orchestrate, and govern your entire agentic ecosystem on one open, unified platform. It's built for the speed of AI, but with the reliability and security that enterprises actually require. We're talking about real results, like KeyBank, who used OutSystems to deploy a customer-facing app that delivered 75% faster onboarding times, or the global logistics leaders who built agentic systems to completely eliminate their engineering bottlenecks. You don't have to choose between speed and control. Whether you're a small team or a massive enterprise, OutSystems helps you engineer, orchestrate, and deploy agentic systems that actually scale. Stop chasing the hype and start owning your agentic future. You can see how it works and learn more at outsystems.com/tcr. That's outsystems.com/tcr. So when you go to the 40% or so success rate on Automation Bench, what can you bring that up to as you move from just asking Claude to do it to giving Claude Zapier and potentially iterating a little bit on initial failures? What kind of lift do you see, and what sort of token savings do you see over time?
You're going to have to wait for Automation Bench V2 for that, because that is exactly the right question: What happens when you give these models access to tools like Zapier? How much higher can you get the efficiency, or how much higher can you get the scores, but also how much can you pull costs down? How much faster can they go?
I think there are many dimensions when you've got a model with access to tools and capabilities and certain things like that where it's going to score better on these benchmarks at the end of the day. That's where I think the application layer has a lot of room to grow: What happens when you give the model access to these extra things? The model is just going to get better at performing against a whole host of tasks.
How do you put the model in a position to be successful when something goes wrong and the model has to come in and troubleshoot and debug? I'm sure that you've done who knows how many things over time to try to set the models up so that, within one generation of a model, it can come back and be able to fix the things that it got wrong the first time.
I think there are 2 places that are interesting to talk about here. The first is when you think about building these workflows. What you want to do there is make sure that the model can properly identify when it needs to be using AI versus where it should just be writing code. There are many examples where you don't want the agent to actually orchestrate the task; you just want it to run code that already exists.
A lot of that is about when the agent is building its plan. You want to make sure that the plan encodes, “This is what the optimized workflow looks like,” and that it has a good plan to do that. The second place is what happens when things break and how you recover from that.
Here, we’ve done a fair amount of work, even in our own support organization, trying to figure out how we troubleshoot on behalf of customers and then bake that troubleshooting back into the core product. One of the big learnings is that we just have multiple agents run at it. We have this auto-email program that’s running right now, and for that, it spins up 5 independent agents that evaluate the troubleshooting situation.
We noticed that when 4 of the 5 agents tend to agree, there’s a pretty good chance that’s actually the issue they’re encountering. There’s a lot of data and measurement that goes into that, where you’re just trying to hill-climb and see how you can try a different model or a different prompting technique. How do you measure that stuff when you have humans auditing the output?
And what those humans do is they basically either give a thumbs up approve or they give a thumbs up rejection and a reason why. Those reasons kick back in and help improve the overall system. You’re just going through this loop over and over again to continue optimizing your ability for that workflow to have a higher chance of getting the outcome that you want.
Is that loop and all the data that you’ve collected over time—which I guess must be quite massive—core to Zapier’s defensibility these days? How do you think about what the hill is that you’ve climbed that will be hard for others to follow you up?
I think that’s a big part of what it boils down to. It’s about trying to identify what things are unique to you and what things others can’t replicate easily. Certainly for us, we’re really good at automation. We have tons of data on what it takes to do, and we hook into everything. How do we take that data to actually make our products better?
It has to be meaningfully better than what somebody without access to that could do. I think this is where a lot of the incumbents have an advantage: If they’re able to wield that to actually build a better product at the end of the day that the models alone can’t build.
There’s so much room for this because the models are great generally out of the box. All of us experience this in our own lives, where you hook up your Gmail inbox and now you start asking, “Help write an email.” It automatically does a better job of writing email because it sees how you write email. You’ve done very little in the way of trying to tune that workflow. It’s just by hooking up your company’s data that all of a sudden the models get better.
Using your own company data to build an edge is a pretty spot-on technique these days.
So does that look like a big retrieval problem for you? Certainly in my email, I’ve got a lot of history, and finding the right example to take inspiration from is probably, especially if it’s just using Gmail APIs and doing keyword searches, just as hard, if not harder, than actually taking inspiration once you’ve found the right documents to take inspiration from.
In my personal context, I’ve tried to help it out by exporting all that stuff, doing embeddings, and using various kinds of alternate search approaches so that hopefully the right content comes to the top more often. If I’m putting 2 and 2 together correctly, it sounds like at Zapier you probably have a database of a zillion things that have gone wrong over time. To help inform the model of how to fix this particular situation, you’ve got to dig in and find analogous situations.
What does that look like? Is there an embedding model that would do a good job of that, or have you had to innovate at the retrieval-stack layer in order to make that work well for a use case such as Zapier?
Coming back to the support example, a lot of it is just about taking the example that comes in, giving it the old thumbs-up or thumbs-down, providing a reason why, and doing that over and over again. What that looks like for our customers is just giving them the same tools to do the same.
A lot of this isn’t particularly fancy at the end of the day. It’s just: Take the example, did you like it, did you not like it, and rinse, wash, and repeat.
Interesting. How do you think about competition in general? It sounds like, on the one hand, we should all be worried about frontier model companies eating our lunch.
Even me, as a humble AI podcaster, look at NotebookLM and think, “They’re coming for me in my rather unlucrative niche.” But you could say, well, those guys are only going to sell their own models, so they’re a different type of animal. We don’t have to worry about them.
There’s a variety of new tools coming online to try to be the Uber tool. I’ve done episodes with Composio, for example. Zero.xyz is kind of out there. And then there are other daily-driver, sort-of agent-builder-type things. And then there are other big incumbents—you’ve mentioned Salesforce—and, to some degree, maybe it’s just big incumbents with lots of data and lots of resources all ending up colliding with each other. Like—
Which of those classes of competitor do you think are actually the ones that you need to be most concerned with?
We’re in an interesting period, for sure. To your point, everyone gives a lot of attention to the labs and tries to understand what they’re doing. I think that is important. You want to understand what they’re going to be great at and what they’re going to hill-climb at.
But I still remember PG’s advice when we were going through YC. Back in the day, it wasn’t, “What if Anthropic builds you?” or “What if OpenAI builds you?” It was, “What would happen if Google built this? What would happen if Facebook built this?” That was always the question.
The thing that PG tried to instill in folks was that you’re not often competing directly with Google. You’re not going toe-to-toe with Larry and Sergey. You’re not going toe-to-toe with Zuck. Oftentimes, in the things that entrepreneurs are trying to build, you’re going toe-to-toe with a potential mid-level product director who’s trying to get a promotion, might be there for 2 years, and then bounce.
And the reality is, OpenAI and Anthropic are big tech now. These are not small, tiny startups. They’re obviously capable of building incredible things, and they’re going to be the best in the world at these foundation models. They’re going to be incredible at that, and they’ll have good products elsewhere, but they can’t build everything. They just can’t.
And so that’s where I think it gets really wide open. I look around, and it is a little confusing because you’ve got everyone who does seem to be building everything. There’s a sort of sea of sameness out there that is a real challenge at the moment. On the flip side, you go talk to the average user of AI tools, and they’re candidly not doing much. They might have used ChatGPT or Gemini.
So, to me, I think for most of us, our competition isn’t each other. It isn’t the tools that you talked about. It’s whether people actually know what to do with these tools yet. They just haven’t adopted anything at this point in time.
The real challenge is whether you can actually get your hooks in somebody whose only experience with AI is using Gemini in a default Google search, or if they’re using Microsoft Copilot at work. That’s where most people are, and I think that can get so easily lost in the shuffle if you hang out on X all day.
On X, we’re all just hyper-aware of what model came out, what tool is gaining traction, who just raised a huge amount of money, and we’re keenly aware of these micro-differences between different products. Most folks just don’t know that.
I think the challenge we have is really making sure we’re keeping those 2 competing thoughts in our head: Yes, we do have to be better in some dimension than all of this sea of sameness, and yet, at the same time, the opportunity is massive just to educate the masses on how these tools can work. These markets are enormous.
The market for automation was orders of magnitude bigger than I ever thought it was when we started the company 15 years ago. It’s probably 1,000 times bigger than what we set out to do. There’s plenty of room for us to solve problems for customers who, candidly, don’t know about any of the competition you just rattled off. I think that’s the biggest challenge for many companies today.
Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic. By now you know my story. Claude drafts my intro essays and I rewrite them, not because the drafts are bad, but so I can stand behind everything I publish. Well, I have an important update. Claude Fable 5 is the first model to have me rethinking my rule. Today, I now think co-authorship, not sole ownership, should often be the goal. Where the model excels, rewriting its work can be more about vanity or a misplaced sense of duty than integrity. I feel it most in songwriting. I'm no lyricist, but I'm good with a song concept, and Fable writes some amazing verses. I give it feedback on its misses, and I push it to aim for higher inspiration, add layers of meaning, optimize syllable density, and above all, write a hit song. These days, I get compliments on just about every song we write together. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you. Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai/tcr. That's claude.ai/tcr. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Once more, that's claude.ai/tcr.
I’m working on an episode where I’m going to talk about things I’ve been wrong about and things I’ve been right about. One of the things I’ve definitely been wrong about is that I expected a lot more change in how things get done across the economy 5 years ago than we’ve actually seen.
Especially if you were to tell me then that we’d have Astra and that it would have been, roughly speaking, a smooth ramp to these capability levels, and yet we still see revenue exploding at the model companies, obviously, but we don’t see nearly as much change as I would have guessed.
Life is kind of the same. I go about my day, I talk to my friends, I talk to my family, and I watch what they do, and yes, their days are the same. They use ChatGPT to help with things like meal planning, planning a vacation, or doing a workout. They’re certainly not taking advantage of Astra-level model capabilities at all.
So what have you learned, then, about what kind of help people need? If you—I don’t know, do you talk about this when you’re taking walks in your neighborhood? How do you...? I’m interested in your method, but even more so in your takeaways.
What is it that gets people over the hump? How do you help them see a new way of working? Are there any patterns that seem generalizable, or is it idiosyncratic for every individual and small company?
There are definitely patterns, but a lot of the specifics matter, and that’s where it starts to feel idiosyncratic. We’ve seen this for years. I remember that, for a long time, one of the hardest problems at Zapier, even to this day, has been helping people with recommendations. What should you actually use this stuff for?
If I personally sat down next to you and said, “Hey, just show me what you do every day,” I could come up with half a dozen examples that would immediately have you saying, “Yep, I want that. Yep, I want that. Yep, I want that.” But then how do you actually bake that in? How do you give them the experience of having me, or somebody who’s knowledgeable about these areas, sitting next to them, and bake that into the product?
This is where I get pretty excited about how AI can help cross that recommendations gap, use-case gap, or whatever you want to call it. If you’re able to point the tools at where you work and say, “Hey, I want you to go watch what I do every day. Watch what I do in Gmail. Watch what I do in Slack. Watch what I do in my browser. Watch what I do in my chat. And just tell me: What should I be doing differently?”
I started doing this workflow at the beginning of the year, and pretty much every week now I have new tools and new systems that start to automate bits and pieces of my job. If you do that on a perpetual basis, you start to feel the difference. After a month or 2, you’re like, “Wow, there’s a lot that’s kind of running for me now that I didn’t have before.”
And I think even though many of the things that are built are things I knew I should have been doing before, it’s the specifics. It’s like, I literally watched the things you did here and here, so I know exactly the tool that can get it done. It’s that idiosyncratic part that makes it click.
I get pretty excited about how you give these models awareness of just how people go about their day. I think there are tons of ideas that are trapped, latent, and lost inside that, which most of us just don’t wake up and think about.
We are creatures of habit, so we wake up and go about our day the same way we did the day before. We don’t think, “Oh, there might be a 5% better way of doing it,” or, “There might be a 500% better way of doing this.” I have muscle memory. I know how to do it this way, and I’m comfortable doing it this way, so I’m going to keep doing it that way, even if it’s not the best.
And so AI has to be so good that it knocks us out of our comfort zone and we're willing to say, “You know what? I am going to go try it that other way because that sounds so, so much better.”
So how have you set that up for yourself, and how broadly deployed at Zapier is this sort of AI on your shoulder or screen recording? I'm kind of expecting a screen recording.
Most people use Zapier MCP for this. They just have all their tools hooked up into whatever harness of choice they use. Like I mentioned, we have our own internal harness, but it could be Cloud code, it could be Cowork, it could be ChatGPT Work, it could be Cursor, it could be anything, right?
They'll just have an automation that runs once a week, and it collects all these signals across all the work they've done. It sees the event streams, and it can tell. It says, “Hey, I noticed you did this and this, and here's a tool that I think you should go build.”
Most people have something like that set up, and then they just tell their agent, “Okay, great, I like that suggestion. Go build it.” Or, “That suggestion's okay. What would make it great is if you made this tweak. I really want you to go do that.”
Half the battle is just getting you to react to something. Even if the ideas aren't perfect, they only need to be 50% good enough to get you to go, “Oh, I see where you're going with this.” Now that brainstorm process kicks off, and you're able to run with it.
So do I understand correctly that it literally just uses APIs to look at your digital history for the last period of time, and then—
More or less.
…kind of collates that together and comes up with ideas? Interesting. Is this something you think you'll productize for Zapier customers?
Yes. I think it's very likely that'll come in some form factor.
Interesting. It's an interesting way of doing it. Obviously, I don't need to tell you, but one downside of going that route with your customers is they'll have to attach all these different platforms that they use first in order for you to have access to get any insight as to what's going on across all these things.
Whereas with a screen-recording type of thing, you have one install and you just look at what they do click by click and make sense of it from a top-down perspective, I guess, as opposed to what you're describing. That sounds a little bit more bottom-up: “Oh, I saw this in Drive and this in Gmail,” and what have you.
Do you think there's a principled reason or an empirical reason for going one way or the other?
We're good at APIs, and we're good at that stuff, so that was just an easy, natural, emergent experiment—an emergent property. I think for this experience to be great, you should use all the tools that you have available to you.
Yeah. It's interesting. It also reflects the alien nature of AI intelligence in some ways. If I was going to try to advise you, I would definitely want to watch you work. I would not be so helped out by your logs. But AIs are really good at reading logs.
How do you think about pricing in today's world? This is obviously a very open question for a lot of companies.
My point of view is that, for the most part, seat-based pricing is dead or dying. I think it may still make some sense in some small areas, but by and large, when intelligence is such a core part of these product experiences, I don't see how a fixed, seat-based pricing mechanism really makes much sense for that at all.
Inevitably, I think that lands you in some sort of usage-based or outcome-based pricing world. I think products end up choosing which side of that fence they land on. If you're going more the commodity route, you're probably closer to usage-based pricing. This is the Sam Altman “we want to be a utility that you pay; you can have intelligence on tap” idea, et cetera.
Maybe if you're a little more enterprise-oriented, you're going to try to say, “Hey, I'm going to price per outcome.” You see a lot of the customer-support tools doing this, where they're going to bill you for resolved tickets, because they have such a clear demarcator of what success looks like.
I do think that if the product you're selling has the ability to have such a clear, agreed-upon, fixed outcome, that's probably to your advantage. The challenge is that, at least for most of the products right now, it's way messier than that.
You stop 1 or 2 steps shy of truly delivering the outcome. You are part of delivering a piece of the outcome. So I think that pulls you back into a more pure usage-based model.
Either way, I think you've got this meter running where you're basically selling work of some portion at the end of the day. I think that's where a lot of this stuff is going: we're going to have to think through what our budget is for work to be done.
Do you worry about price discrimination by the model companies? Of course, you're well aware of the ratio of tokens that you get with a Claude Max or an OpenAI Pro plan, and how many more tokens you get, at least if you max them out, compared to what you can buy with the same dollars via the API.
I've asked a number of entrepreneurs this question, and I'm wondering whether you would be supportive of some sort of rule that said, “Hey, you've got to charge everybody the same for tokens so that an ecosystem has more of a fighting chance,” versus OpenAI giving, say, a 20-to-1 token advantage. That could be hard for third-party value-added services to compete.
Yeah. I definitely tend to fall on the side of free markets, and these are companies that have the right to price how they like. But many of us lived through, and are still living through, Microsoft's dominance and how they used bundling to their advantage to box out better products, candidly.
Because it's all just bundled there, it plays to their advantage. I think this is where countries and folks get to decide what they think are monopolistic practices and what they think is a fair playing field.
Ultimately, I think my job is to play by the rules on the playing field and not necessarily decide. I definitely lean more toward the free-market side and say, “Hey, our job is to come up with an edge that helps us compete there.”
I don't fault any company for wielding the tools they have in their tool chest to make products work for them, work well for their customers, and help them maximize revenue. That's well within their right.
It's been interesting to see who's been willing to bite the bullet versus who has stuck to their free-market principles.
Changing the topic toward operations and AI transformation within Zapier, one thing that caught my attention was that, if my AI research agent is to be trusted, you put your chief people officer in charge of AI transformation. I believe last time we talked, you had said, “Well, it's not any one person's job. It's kind of my job as CEO, but it's really everybody's job. So I'm not going to say it's one person's job.”
What changed, and how did you decide it would be the people officer who would shoulder that burden?
I still agree that AI should be every person's job, and it should be my job. I think it depends on what stage you're at and what problems you're facing in terms of how you think about who you want to tackle the next mountain, so to speak.
For our first chapter, a big part of Zapier becoming AI-fluent was everyone in the company getting up to speed on how to use these tools. There wasn't an AI committee. There wasn't an AI group where it was like, “Oh, they're there. They kind of figure it out. The rest of you, it's business as usual.”
It's like, no, this is important for everyone inside the company. It impacts everything we do, so all of us need to get on that.
As time went on, there were a couple of interesting things that we started to observe. First, the AI fluency inside the company went up. Basically, within a year or so of ChatGPT launching, almost 100% of the employee base was using AI day to day.
We're not having technical issues adopting AI. That's not where we're bearing the brunt. One of the bigger issues that is starting to emerge is how you take these models from individuals having success to actually using them to solve bigger and bigger production-grade workflows across the company.
The challenges start to look a lot more like people issues. We have to rewrite certain job descriptions, rethink how we do compensation, and think about how these teams stand up. We have to move this group—we kind of don't need this group anymore—but we actually need more people over there. So how do we retrain and reskill these folks who have some of those skills but need to learn some new skills?
It turned out Brandon, our chief people officer at the time, was really good at doing a lot of these things. His team was at the forefront of some of this inside of Zapier.
The thought for me was, “Hey, you’re doing a good job at this. Why don’t you go help everybody in the company figure out some of these things?” It could have just as easily been a CMO or a CPO or any number of roles. That’s how it’s played out inside of Zapier. And it’s been funny how much I get asked this question now, because I think a lot of folks think I have a point of view that it must be a chief people officer or something like that.
It was really just that, at this moment in time inside of Zapier, this sort of felt like the best person to go tackle it based on the problems we were facing. And I think that’s the way you should do it inside your company. You need to look at what your bottlenecks are, what your constraints are, and go identify the person who is the right fit for that job.
Yeah. Echoes of Ben Horowitz’s advice, too. You’ve also made an interesting move of really trying to push people toward internal communications being public within the company by default. And I’m interested in a couple of finer points. One, how did you handle historical data? Did you start that policy at a certain point in time, and everything in the past was left in the past? Or did you try to reclaim some of that knowledge, which I assume would be very tempting to do?
And then do you have different tiers of public as well? Because it strikes me that you might not want everyone to know everything, but you might want different groups to have certain different databases. So I’m just looking for the double-click on how you’ve operationalized public by default.
One interesting thing is that Zapier’s had this value of default to transparency for, gosh, forever—it feels like. By and large, we were already working in public for many such things. We already had a culture where there were tons and tons of public Slack channels, and people were talking about the projects and their day-to-day in those quite a bit.
I think a lot of what we observed was that, as the company grew, there were some pockets of work that started to find their way into private channels or private DMs and things like that. By and large, Zapier was still much more public than most companies. Slack gives you the readout, so you can see what percentage of stuff is happening in private and public and all that sort of stuff. We looked at that and thought we could use a little bit of a reminder.
The second thing that encouraged us to do this is the fact that our AI agents were so much more effective when they were able to see the context inside of Slack. One of the things we started to do on our executive team was have a little fun competition to see who could put most of their communications in a public channel. There was no, “Oh, you must do this,” or “Anyone under this rate gets a bad review,” or anything like that. It was literally just friendly competition.
We noticed that more things could go in public than we realized, and this seemed to help the team. People know what’s on our mind. Nothing is hidden, et cetera. So we started to encourage that all across the company.
There are certainly things that we still pull into private channels and things like that. If there’s an HR incident or something like that, we’re not resolving that in a public channel. If there’s a critical security vulnerability, we’re not resolving that in a public channel either. Those are happening in private channels, especially while the incident is active. Once they get resolved, we tend to share out the learnings and things like that.
There are certain topics where you still need to set up these spaces where you can go resolve them in private. But by and large, most people far overestimate the number of things where that is required, and definitely underestimate the power of what happens when both humans and agents have access to the full context of what a company is working on.
So, speaking of security, this is obviously top of mind. As I was thinking about challenges that AI might pose to you, you’re holding potentially more credentials to more different services for more different users than just about anyone in the world, right? So I would think this is kind of a scary moment. All of a sudden, where’s Bedrock in terms of security?
How are you approaching that, and are you trying to get into these early-adopter, Glasswing, and other clubs? Do you think that’s actually maybe a big source of differentiation going forward? And how scared should I be about cybersecurity? Because I’ve got a lot of credentials all over the place, Zapier and otherwise.
We’ve held these credentials for 15 years, right? So this has always been an important thing inside of Zapier. We’ve said, “Hey, these credentials are a thing that we must treat with the highest, highest level of stewardship.” We’ve always put a lot of effort into making sure that we do a good job of protecting those for our folks.
What feels different this time is that you do have these mythos caliber security models that are able to patiently loop over issue after issue and find things. Every software project has vulnerabilities; it’s just that someone hasn’t found them yet. The models make it a lot easier to find those things.
The good news is that they also make it easier to patch them. I think what smart companies are doing is basically wielding them for offense and defense. They’re trying to find this stuff faster, and they’re trying to resolve it faster.
I’m not exactly sure how all this is going to play out. Every day, there’s funky stuff going on. Obviously, the Hugging Face incident was straight out of a sci-fi book, right? But for most folks, my guess is it feels like there’s this almost one-time investment to reacclimate, and then you get back to a more steady state of offense versus defense in terms of security posture as this stuff moves forward.
It’s going to be really interesting to see, because every day we’re seeing new stuff.
Are you taking steps as CEO to try to make sure you’re on the inside of early-access lists for new models?
Yeah. We want to have access to the best capabilities as early as we can. I think everybody who’s in the same shoes would want to do the same.
Yeah. It feels like that could actually be a pretty meaningful point of differentiation going forward. If one company that’s going to hold my credentials is in all the clubs and another one is a startup that’s not, that’s a big leap of faith to take on the company that doesn’t have the same kind of access to be trying to find and fix all these issues.
In terms of spending, I saw you tweet not too long ago that you have some engineers spending $30,000 a month on tokens at Zapier. It struck me that we’ve been on quite the yo-yo ride recently, with token-maxing and then budgets being hit. What do we do about it? What sort of process or governance do you have for who can, under what circumstances and with what approval, spend tens of thousands of dollars a month on tokens?
Right now, I would say that those individuals are a bit of an outlier. But it still encouraged us to start building some tools to help people do some self-policing. Mostly, we don’t have budgets set up for individuals yet, but we do have tools where they can see their spending and better understand what happens when they choose a powerful model versus when they choose a cheaper model on certain workflows. They can see what those cost differences are.
As we see people starting to spend a ton on tokens, usually the first reaction is, “I just want to go talk to them and say, ‘Hey, what are you doing? I’m just really curious.’” In some cases, you have folks who are doing some insanely productive stuff. In other cases, you have some folks who have a mix of things that are pretty productive and places where it’s, “Oh, you don’t need to be using Fable for this or Astro for this. There’s a better way to do some of these things.”
When you have an almost 800-person organization, there’s a big education effort involved. Over time, I do suspect that token budgets are going to be a real thing, though. Part of AI fluency is going to be that you’re going to say, “This person is going to get a higher budget than this person because they know how to get higher output from those things.”
How do you actually operationalize that? We’re still working through some of that stuff. But it seems pretty obvious to me, just looking across the employee base, that some people are excellent at using increasing levels of spend, and some people are just not really thinking about it all that much yet.
Another aspect of AI fluency that I’m really curious to get your take on is what you think is the right model for co-authorship or co-creation with AIs.
Sure.
And this is not a gotcha, because I’m in the same boat. I’ve consciously tried to almost shock-expose myself recently, to put some things out that I didn’t rewrite every word of. So my Pangram score at times says that my stuff is AI.
I’ve also seen some stuff in various places from Zapier that has a high Pangram score. How do you think about that, and how do you set the tone for others at the company? You want to be using these things—
But we don't want to be putting out slop. What's the line?
The way I think about it is, I have no problems with people using AI for communication at Zapier. What I really have a problem with is low-quality communications. We live in an era where AI can enable a person who is exercising low judgment to create a high volume of very low-quality stuff very quickly, and that can overwhelm a person.
We've tried to put a few guidelines in place that help people think through ways to go about that. For one, you need to own what you send. If you wrote it, you probably should be putting more time into authoring the thing than the reader is reading it. AI use shouldn't be a way of transferring ownership of a task, where it's like, "Oh, I was assigned this task, so now I had AI spin a prompt," and I said, "Hey, you now read it and deal with all this stuff and edit all those things." That's not a great way of going about it.
You need to understand what you send. If someone starts asking you questions and you're like, "I actually don't know what's inside of that," that's not good. You probably ought to be making asks explicit. If you need somebody to do something, if you need a decision, if you need feedback, or if you're labeling something as a draft and you want feedback on the thing, you need to do so.
You need to go verify details. It's not uncommon for AI to hallucinate some of these details—or maybe it's not hallucinating. It might pull dated information. So if you hook it up to an agent that has access to Zapier's Slack, it could pull, "Oh, this project from 3 months ago is related, but not the exact same thing." If you're trying to pass that off, that becomes a real issue.
An AI summary can be based on a summary of a summary of a summary, and all of a sudden, before you know it, it's actually passing on incorrect information. You have to do a good job of verifying the details in there. To me, that's the really important piece: you are still an active participant in the creation of the material.
But if AI is helping you structure your thoughts and structure the writing, at the end of the day, go for it. I don't have any problems with that. It can be tedious if you're not scrubbing some of the slop that's a part of it. The "It's not this, it's that," the em dashes, the "honest truth," the "load-bearing point"—all that kind of stuff.
I do think that if you're doing that a lot, especially if you're doing it in marketing material, it makes it hard to stand out. People get a little tired of reading that kind of stuff. You still need to have your own editorial hand on the steering wheel, so to speak. To me, using AI is not the problem. Low quality is the fight at the end of the day.
How has your team composition changed over the last couple of years? You mentioned earlier that maybe we don't need this team, but we can reskill. Are there any thresholds for AI capability—something where you're like, "Well, they can't do this now, but if they could, I could see that actually making a big impact on our hiring plans going forward from that point"?
What's interesting is that, in some ways, our team looks very similar to how it has in the past, and that's maybe a surprise to me. But in other ways, it is pretty different. We still have engineering, design, product, and stuff like that inside the organization, but the idea of a classic, traditional EPD is largely gone. Things are a lot more malleable, but the roles still exist.
There's definitely been a flattening of management layers, but strong management is still crucial. We're not getting rid of managers anytime soon. Managers can handle a higher volume. Those are a handful of things that are interesting.
Similarly, with data analysts, everyone inside of Zapier is kind of their own mini data analyst now, so you don't need as many data analysts. And yet, we still have analysts who are doing really critical, important work inside the company. They're just working on higher-value stuff now.
You can feel things shifting, and yet, in some ways, it still feels pretty familiar at the same time. The company feels quite a bit similar.
You asked what the models are not yet capable of that I'm excited about. I still have this idea of the team shifting more into building the factory that builds the products, the company, the marketing, and all that sort of stuff. You can start to feel where we're doing more and more of that. We have a software factory, a support factory, and workflows that are getting stood up where they're handling the inner loop and the humans are more focused on designing that piece of the puzzle.
Inside those factories, there are all sorts of steps where you come across areas where you're like, "The AI's not quite good enough for this yet. We need a human in the loop." But with every model release, with every iteration of our own learning loop inside of Zapier, you can start to feel us chip away at that problem. We're just getting closer and closer to something that looks actually pretty different from the organizations of the past, and that's pretty exciting, I think.