No Code Is Code: Zapier CEO Wade Foster on Headless Tools, Zapier MCP & Automation Bench
- Wade Foster's biggest update since his last appearance: the ambiguity over what shape AI experiences would take has resolved — most folks seem to have settled on a single "daily AI driver" (Claude Code, Cursor, ChatGPT), and other tools need to plug into it. That makes headless the winning architecture: "you kinda have to integrate with that person's core daily driver if you really wanna be a part of their day-to-day workflow," and Zapier MCP is the company's bet on that layer.
- Model capability on real knowledge work remains far from saturated: on Zapier's Automation Bench of ~600 marketing/sales/HR/ops tasks, the new state of the art — "Astra came out, GPT-6" — completes only about 40% accurately, while "Gemini 3.7" does "pretty good but does it at a fraction of the cost." That cost/performance curve is why Foster says organizations want to avoid dependence on one lab's models, and why there needs to be a third-party harness — the labs "can't sell tokens from each other or from open source."
- The episode's thesis quote: "the new no-code is code" — but Foster says roughly 80% of what customers currently delegate to agents "should be using actually old-fashioned deterministic code," with AI reserved for the parts that genuinely require reasoning. The emerging pattern: agents build and maintain workflows, deterministic code runs them, and agents can troubleshoot failures — cheaper and more reliable for the deterministic portions, even as models improve.
- Foster argues the real competition isn't the "sea of sameness" of AI startups or even the labs — it's non-adoption: "go talk to the average user of AI tools, and they're candidly not doing much." He channels Paul Graham's YC advice ("You're not going toe-to-toe with Larry and Sergey... you're going toe-to-toe with a potential mid-level product director who's trying to get a promo") and sizes the prize: automation is "probably 1,000 times bigger than what we set out" 15 years ago.
- Zapier's possible next product is AI-assisted workflow discovery: Foster runs a weekly automation that collects signals across Gmail, Slack, the browser, Cursor, and more, then proposes tools to build — suggestions only need to be "50% good enough" to trigger the brainstorm — and says it's "very likely" this gets productized. Nathan Labenz's observation on why logs beat screen recording here: "AIs are really good at reading logs."
- On monetization, Foster says that "for the most part seat-based pricing is dead or dying," with the market splitting between usage-based (the Sam Altman intelligence-on-tap utility) and outcome-based (support tools billing per resolved ticket) — though most products "stop one or two steps shy of truly delivering the outcome" and get pulled back to usage. Internally, some Zapier engineers spend $30K/month on tokens with no hard budgets yet, only self-policing dashboards; Foster expects token budgets tied to demonstrated ability to get output from spend to become part of "AI fluency."
- The org-design signal: Foster put chief people officer Brandon in charge of AI transformation because the bottleneck shifted from tool adoption (~100% of employees using AI within a year of ChatGPT) to people problems — rewriting job descriptions, comp, and re-skilling — but insists "it could have just as easily been a CMO." Alongside: public-by-default Slack because agents are "so much more effective when they were able to see the context," and a co-authorship rule for the AI-slop era — "Using AI is not the problem. It's, like, low quality. That's the fight."
- On security, Zapier has stewarded credentials for 15 years; Foster says AI makes both vulnerability discovery and patching easier, though he is unsure how the balance will play out. Zapier also wants access to the best new capabilities as early as possible, as other companies in its position would.
1. The daily driver won, so headless is the winning strategy
- Foster's opening admission: 18 months ago "I don't think we knew yet... the shape of these AI experiences that were gonna take off." What happened, per Zapier's own usage data, is that most folks seem to have adopted a single daily driver — Claude Code, Cursor, ChatGPT — and tools like Zapier MCP that bring your context into that driver are "the way knowledge work is moving." He draws the market split explicitly: platforms forcing users to build agents on their own surface versus "the Salesforces of the world that are like, 'We're headless. You can bring it wherever you want'" — "that latter is the winning strategy."
- Foster practices what he preaches: his own daily driver is Cursor, with Zapier's internal harness pulled in via MCP — a virtual file system as context layer, automations on top, deployable apps.
- Responding to Labenz's summary of Andrew Lee of Tasklet's thesis (only a few companies win the horizontal "Mecca suit for models" layer, as Nathan put it): Foster agrees nobody wants to be beholden to one lab, but adds a wrinkle — the harness can now build its own tools, and "Yes, it can build it, but should you have it build it? Because now you're accepting a certain amount of maintenance, a certain amount of reliability, and a certain amount of uptime." The labs have a head start on applications but "they can't sell tokens from each other or from open source" — hence room for a third party.
2. Automation Bench: 40% is the frontier, and the lift from tools is the next headline
- The benchmark's texture matters: ~600 tasks like "We just closed the Meridian Core platform deal. Mark it as won and route it to the win notice in the right team per our routing policy; confirm the accounts from the account hierarchy spreadsheet, convert the currencies if needed, and check for any open support escalations." Astra (GPT-6) is the new state of the art at about 40% — "this is by no means a saturated benchmark yet" — while Gemini 3.7 does pretty well at a fraction of the cost, forcing per-task trade-offs any modern organization will want to make.
- Labenz's sharpest question — how much does giving Claude access to Zapier lift scores and cut tokens? — gets a deliberate tease: "You're gonna have to wait for Automation Bench V2 for that," which will measure score lift, cost reduction, and speed when models get tools. Foster says Zapier is working to ensure that installing it alongside an agent produces better output than using the models alone.
3. "The new no-code is code" — but 80% of agent usage should be deterministic
- The load-bearing statistic of the episode: looking across customers' agentic usage, "the vast majority of what people are using an agent for, 80% in fact, they probably should be using actually old-fashioned deterministic code." You only want AI reasoning over things that need reasoning; Foster "struggle[s] to think of a world" where deterministic code isn't more reliable and cheaper for certain jobs — with "the ghost in the machine" reserved for the rest.
- On the interface shift: "the idea of building a no-code, it feels antiquated to me. It's like the new no-code is code." Yet humans "still very much benefit from visualization of those workflows" — as verification and as documentation — so people now talk to the agent, which makes the edits, rather than configuring boxes themselves.
- The emerging failure loop: workflows harden into deterministic code, and when they break, the agent troubleshoots and fixes the workflow or the instance — "the agent almost take[s] over the human role of building and maintaining it," while what runs stays deterministic. Labenz's concern that the Meridian example is a one-off gets corrected: "If you're a good organization, you're closing deals all day, every day" — and at consumer scale, "you cannot put humans in the loop for these tasks."
4. Zapier's recursive loop: five agents, thumbs up/down, and a data moat
- Inside Zapier's auto-email support program, five independent agents evaluate each troubleshooting case; "when four of the five agents tend to agree, there's a pretty good chance that's actually the issue." Humans audit outputs with approve/reject-plus-reason, and those reasons feed back — pure hill-climbing across models and prompting techniques.
- Foster's defensibility argument: the moat is data others can't replicate — "we're really good at automation. We have tons of data on what it takes... we hook into everything" — and the product built on it "has to be meaningfully better than somebody who doesn't have access to that could do." His everyday analogy: hook up your Gmail and the model writes better email "because it sees how you write email," with zero tuning.
- Labenz probes for exotic retrieval-stack innovation; Foster deflates it: "A lot of this is not particularly fancy at the end of the day. It's just take the example, did you like it, did you not like it, and just rinse, wash, and repeat."
5. The real competitor isn't the labs or the clones — it's non-adoption
- Foster resurrects PG's YC-era advice: "You're not going toe-to-toe with Larry and Sergey... oftentimes you're going toe-to-toe with a potential mid-level product director who's trying to get a promo." And the update: "OpenAI, Anthropic, they're big tech now... they will have good products elsewhere, but they can't build everything. They just can't."
- The two competing thoughts he insists on holding: yes, differentiate against the "sea of sameness" — but "go talk to the average user of AI tools, and they're candidly not doing much." Many people's experience may be limited to "Gemini in a default Google search" or "Microsoft Copilot at work," a reality that "can get so easily lost in the shuffle if you hang out on X all day." The market: "orders of magnitude bigger than I ever thought... probably 1,000 times bigger than what we set out."
- Labenz's confession — a genuine change of mind worth flagging: "One of the things I've definitely been wrong about is I expected a lot more change to how things get done across the economy five years ago than we've actually seen." Foster concurs: "Life is kind of the same" — friends use ChatGPT for meal planning, vacations, and workouts, "certainly not taking advantage of Astra-level model capabilities at all."
6. Closing the recommendation gap: AI that watches how you work
- The hardest problem at Zapier "even till this day" has been recommendations — Foster could sit next to anyone and find half a dozen automations they'd instantly want, but baking that expert-next-to-you experience into product is the gap. His answer: "Watch what I do in Gmail. Watch what I do in Slack. Watch what I do in my browser... and just tell me what should I be doing different?" Running this weekly since the beginning of the year, he now has new tools and systems automating bits of his job most weeks — the ideas were "trapped, latent, and lost" in habit, because "we are creatures of habit."
- The mechanism is unglamorous: Zapier MCP and connected tools collect event streams once a week and propose builds; the user reacts, refines, and says "go build it." Half the battle is reaction — "they only need to be, like, 50% good enough to get you to go, 'Oh, I see where you're going for.'" Asked whether this gets productized: "Yes. I think it's very likely that'll come in some form factor."
- Labenz pushes on why logs rather than screen recording — one install and a top-down, click-by-click view; Foster's honest non-answer — "We're good at APIs... it was just an emergent experiment" — and Labenz's better frame: "It reflects the alien nature of AI intelligence... AIs are really good at reading logs."
7. Pricing: seats are dead, tokens need budgets
- Foster's call, qualified for most products: "for the most part seat-based pricing is dead or dying." The fork: commodity products are probably closer to usage-based pricing ("the Sam Altman... intelligence on tap" utility), while more enterprise-oriented products may try outcome-based pricing, such as support tools billing for resolved tickets — but only where success has "a clear, agreed-upon fixed outcome." Most products today "stop one or two steps shy of truly delivering the outcome," pulling them back to usage; either way "it's basically selling work," and buyers will budget for work to be done.
- On lab price discrimination — Labenz raises the possibility of roughly 20-to-1 token advantages in Claude Max or OpenAI Pro plans versus the API — Foster declines the interventionist bait: he leans free market, though "many of us lived through, and are still living through, Microsoft's dominance and how they used bundling to their advantage to box out better products, candidly." His line: "My job is to play by the rules on the playing field."
- Internally: engineers spending $30K/month on tokens are "a bit of an outlier," and Zapier has dashboards for self-policing rather than individual budgets — his first reaction is to ask "'Hey, what are you doing? I'm just really curious,'" finding a mix of "insanely productive stuff" and "you don't need to be using Fable for this or Astro for this." But across an almost 800-person organization, Foster expects token budgets eventually to become real, potentially reflecting who can get higher output from that spend.
8. People ops as AI strategy: the CPO, public-by-default, and the factory
- Why the chief people officer runs AI transformation: within roughly a year of ChatGPT launching, "almost 100% of the employee base is using AI day to day" — so the bottleneck moved from adoption to people problems: rewriting job descriptions, rethinking comp, re-skilling teams as groups shrink and others grow. Brandon was already good at this; "it could have just as easily been a CMO... you need to look at what are your bottlenecks... and go identify the person who is the right fit," not copy Zapier's org chart.
- Public-by-default got a second wind for an AI reason: "our AI agents were just so much more effective when they were able to see the context inside of Slack." The executive team ran a friendly competition on public-channel share — no mandates — with carve-outs for HR incidents and active security vulnerabilities. Foster's rule of thumb: people "far overestimate the number of things where that is required, and definitely underestimate the power" of full context for humans and agents.
- Team shape: classic EPD is "largely gone," management layers flattened (but "strong management is still crucial"), everyone is "their own mini data analyst." The direction of travel: "building the factory that builds the products" — software and support factories where workflows handle more of the inner loop and humans focus on designing that piece, while each model release chips away at the remaining human-in-the-loop steps.
9. Credentials, AI-assisted vulnerability discovery, and owning what you send
- Labenz flags Zapier's potentially vast credential footprint; Foster responds that Zapier has stewarded credentials for 15 years. What's new is "mythos caliber security models that are able to patiently just loop over issue after issue" — "every software project has vulnerabilities. It's just someone hasn't found them yet." The symmetry is the reassurance: models make patching easier too, so smart companies wield them "for offense and defense." He name-checks the Hugging Face incident as "straight out of a sci-fi book," and hedges his forecast: "my guess is it feels like there's this almost one-time investment to reacclimate" before a new steady state — "I'm not exactly sure how all this is gonna play out."
- Labenz argues that early-access membership could become meaningful differentiation for a company entrusted with credentials. Foster's narrower response is that Zapier wants access to the best capabilities as early as possible, as anyone in its position would.
- On AI co-authorship, Foster's rules rather than bans: "I have no problems with people using AI for communication at Zapier. What I really have a problem with is low-quality communications" — because "a person who is exercising low judgment can create a high volume of very low-quality stuff very quickly." The guidelines: own what you send; spend more time authoring than the reader spends reading; never transfer ownership of a task via a prompt; verify details, since "an AI summary of a summary of a summary" starts passing on incorrect information; and scrub the tells — "the 'It's not this, it's that,' the em dashes, the honest truth," and the "load-bearing point." The close: keep "your own editorial hand on the steering wheel."
Full transcript
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to welcome Wade Foster, co-founder and CEO of Zapier, back to the show. When I last spoke to Wade in September of twenty twenty-four, some four hundred thousand customers had already used Zapier to delegate more than one hundred million tasks to AI. And YC President Garry Tan was calling Zapier the AI-powered knowledge worker of the future. Since then, models have, of course, become dramatically more capable, and Zapier has built out a full AI portfolio, including agents, chatbots, an MCP server, an SDK, and an AI guardrails product. And yet, somehow, white-collar work and the world as a whole have changed much less than I would have expected. With that in mind, I wanted to hear not only about what Zapier has built and how it's continued to evolve as a company, but what Wade and team have learned about how businesses across the economy understand and use today's AI tools. At a high level, Wade believes that most people are now settling in to using a single daily driver, whether that's Claude Code, ChatGPT, Grokbot, or in Wade's case, Cursor, and that platforms like Zapier will need to adapt by making their tools available and effective in those environments. Practically, he observes that models still struggle with many business tasks, as illustrated by Astra setting a new high of just forty percent success on Zapier's automation bench, which consists of roughly six hundred knowledge work tasks across marketing, sales, HR, and operations. He also argues that many tasks that people are delegating to AI would be better done with deterministic code, and that for a while longer at least, there is therefore tremendous ROI to time invested in structuring and validating workflows. We then go on to discuss what Zapier is doing to help people recognize exactly what AI might be able to do for them, starting with his own weekly automation, which Wade says they will soon productize for customers, that reviews his activity across Gmail, Slack, the browser, Cursor, and more, and then proposes specific tools and workflows that he should be building. We also talk about how Zapier is implementing recursive self-improvement loops internally and how much value they're finding in running multiple different AIs on the same problem, why Wade chose to put Zapier's chief people officer in charge of AI transformation but wouldn't necessarily recommend that strategy to other companies, how Zapier is moving toward public-by-default communications to make more and more context available to AIs, why they still don't limit individuals' use of AI but have created dashboards to help employees better understand and manage their own usage, how Zapier, which holds a huge number of high-value user credentials, is thinking about security in the context of rapidly rising cybersecurity risks, and finally, how Wade thinks about co-authorship between humans and AIs, with the upshot being that he believes individuals should use AI to help improve their writing and shouldn't be afraid of being pangrammed, but also that it's critical that people be prepared to explain and stand behind the work that they ship. With that, I hope you enjoy this very grounded and highly practical conversation about making AI automation work for people outside the AI bubble with Wade Foster, co-founder and CEO of Zapier. The Cognitive Revolution is brought to you by Mercury, the fintech that more than three hundred thousand ambitious companies and individuals trust to run their finances. I've wired AI into nearly every corner of my life, my email, my messages, my calendar. I even gave Mercury virtual cards to my agents with low limits and category and merchant restrictions for their autonomous use. But still, my AI's access to my financial data has remained limited. With a normal bank, I might export a bunch of statements and have my assistant process them for me. But for real-time, up-to-date information, and certainly for taking any action, trying to get your agent to use the bank via the browser is just too hard, too slow, and too error-prone to be worth it. And that's why Mercury's new conversational interface, Command, is such a big deal. It's built directly into Mercury, which means you get natural language access to your finances without exposing anything outside of your bank account. No exports, no spreadsheets, no pasting your transactions into third-party tools. I really think a lot of people are going to prefer it this way, and it can already help you take actions too, with everything bound by the permissions and approval policies that you've already set up in your account. I am genuinely impressed to see this level of AI integration in banking in twenty twenty-six, and so I invite you to join me in the future. Visit mercury.com to learn more and apply online in minutes. Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column NA, members FDIC. Thank you to Mercury for supporting The Cognitive Revolution. And now, on with the show. Wade Foster, CEO of Zapier, welcome back to The Cognitive Revolution.
Yeah, thanks for having me again, Nathan.
Boy, what a difference a not-super-long period of time makes. I can't believe it. Every time I have a returning guest, it's an opportunity to look back and look at what the state of AI was—the models, outlooks, what the possibilities were, and what the capabilities were at the time. Suffice it to say, obviously, a lot has changed.
We only have 1 hour today, so I'm going to try to discipline myself, not talk too much, and give you most of the airtime. I would love to start off with an observation that, like many ambitious software companies, you guys at Zapier have been really prolific, and I would say you've kind of built everything in the sense that you now have, in addition to workflows, which of course can call out to AIs, agents, chatbots, an MCP, an SDK, a guardrails product, and probably more that I didn't identify or mention.
What have you learned by building all that stuff? What has taken off and resonated maybe better than you thought? What has been slower than you thought? And what's the synthesis view that you have, informed by these different product efforts and their relative successes?
Yeah. I was on 18 months ago, and if I think back to that time, what I don't think we knew yet was what shape these AI experiences were going to take. Were we going to see AI infuse into all the products that we already knew and loved? Were we going to see these new product shapes take off? Were we going to see new stuff out of the labs that was going to be where AI was going to take off?
What would seem to have happened, at least when we look at our own usage, is that most folks seem to have adopted their own daily AI driver tool. Maybe this is CloudCode, maybe this is Cursor, maybe this is ChatGPT—you name it—but that's where people do most of their work. What that means is tools like Zapier MCP, or any MCP server—tools that bring your context and your data into that person's daily driver—that feels like the way knowledge work is moving.
There's a split in the market. Some folks are still trying to say, “Ah, we're going to force people to build agents on our own platform,” or, “We're going to bring them over here, and that's kind of where things are going to get done.” Then you see the Salesforces of the world that are like, “We're headless. You can bring it wherever you want,” et cetera. I definitely think the latter is the winning strategy.
It just feels like that's what we, as consumers of these tools, want, and that's where all the growth is. We kind of want our daily driver, and we want to bring that context in and be able to manage it all in one place. That doesn't mean there aren't going to be tools that call out to all these third parties. I still think there's plenty of room for applications and tools to exist, but you have to integrate with that person's core daily driver if you really want to be a part of their day-to-day workflow. That feels like a pretty big new learning for me in the last 18 months.
That's interesting. So does that mean then—what's your daily driver? It sounds like you're positioning Zapier not as being a daily driver, but more as an uber tool for whichever daily driver you choose to use.
Yeah. I mostly use Cursor every single day. We have an internal tool that a lot of our employees use as a sort of daily harness, but that also has an MCP associated with it. I have that pulled into Cursor, so it's using the data from that.
It has a virtual file system that acts as a context layer, automations that sit on top of it, apps that you can deploy—all the kinds of things that you might expect from a modern AI-capable tool. I just happen to use a lot of that stuff inside Cursor.
Very interesting. Okay, I had a conversation with Andrew Lee from Tasklet. He advanced what I thought was a provocative idea: that, in his mind, only 3 kinds of software companies survive in the big picture, and that everyone is kind of building the same thing, which sometimes gets described as the Mecca suit for models.
He described himself as trying to earn a place in the eventual winner's circle of this horizontal layer that sits on top of models and enables them greatly by providing all these different tools, access points, guardrails, and whatever else the case may be. He thinks that there aren't a huge number of companies that ultimately win in that case, but that layer, even if there aren't too many companies in it, is super valuable because nobody wants to be beholden to just 1 model.
They don't want to be overly locked into a single model provider. How would you compare and contrast your worldview against that summary?
Well, I definitely agree with that last statement. I think it is becoming more and more obvious that you don't want to be beholden to one model or one company's suite of models. We see this with Zapier's Automation Bench. We have a benchmark that measures all these models on automation tasks.
Last week, Astra came out—GPT-6. It's the new state of the art on that model. It performs about 40% of the tasks accurately, which is the highest available. Now, it's more expensive than, say, something like Gemini 3.7, which does pretty well but does so at a fraction of the cost. You have this curve that exists where you're trying to figure out how much you're willing to pay for incremental performance on certain tasks.
As a result, I think any modern organization wants the ability to make those trade-offs, to say, “These tasks are good enough. I can pay this rate and get 100% of these types of tasks completed, but for this type of task, I need to maybe move to a state-of-the-art model to do well on it,” and so on and so forth. I do think that companies are going to look for the Mecca harness or whatever you want to call it—the sort of tool that allows them to swap in and out different workflows of their choice.
To that end, I certainly believe that there's a lot of innovation yet to be had on the application layer. But how those applications are used, I think, looks very different from the last decade. The last decade was SaaS, and there was this explosion of SaaS, but now it feels like there's almost an explosion of headless tools happening. Of course, you have this new thing, which is that your harness itself can build some of those tools, so the build option is more readily available than it was in the past.
What's not as obvious to me is: Yes, it can build it, but should you have it build it? Now you're accepting a certain amount of maintenance, a certain amount of reliability, and a certain amount of uptime that may not actually be the best thing for you in a given circumstance. I actually think there's still a lot of innovation to happen at that application layer that we just haven't seen yet. I think everyone's trying to figure out what that looks like.
I think the model companies have a little bit of a head start, but they're going to struggle because they can't sell tokens from each other or from open source or anything like that. It does feel like there needs to be a third party that helps you wrangle all the capabilities that are out there.
Obviously, Zapier came from a history of very structured workflows, because if you didn't fully encode what the workflow was supposed to do, there was no ghost in the machine back when you started to figure it out on the fly. I'd be curious to hear a little bit about how you would describe the tasks of Automation Bench, where the models are good, and where they're not good.
Then maybe you can describe how usage of Zapier is changing qualitatively. How often are people still doing box-by-box defining of workflows? How often are they prompting an AI that then turns their thoughts into a structured workflow? How often is it happening through an MCP, where the model is deciding whether to use Zapier given a range of options? Maybe there's even more there that I'm not intuiting.
You bet. What Automation Bench measures is tasks like the following. Here's an example from our site: “We just closed the Meridian Core platform deal. Mark it as won and route it to the win notice in the right team per our routing policy. Confirm the accounts here from the account hierarchy spreadsheet. Convert the currencies if needed, and check for any open support escalations.”
We have probably 600 tasks of that variety that describe a normal knowledge-workflow task across a wide variety of disciplines—marketing, sales, HR, operations, you name it. That's what it's trying to measure. As you can see, the models are getting better at it, but this is by no means a saturated benchmark yet.
What we're doing at Zapier is making sure that when you install Zapier alongside your agent, you're actually getting better output on these benchmarks than you would if you were just using the models alone. The reason you do that is you're teaching the model what folks are doing in their daily driver. They're saying, “Hey, I want you to go build a workflow.”
They're calling Zapier MCP, and it's going to say, “I'm going to build out that workflow, and in some cases I'm going to write code to actually complete that task so that it is deterministic,” which means lower cost, better reliability, and so on. Then I'm going to invoke an AI or build an agent for the parts that really require reasoning. When we look across even our own customers' usage of agentic products, the vast majority of what people are using an agent for—80%, in fact—probably should be using old-fashioned deterministic code.
You really only want the AI to reason over the things that you need it to reason about. I still think there is a huge amount of work happening inside these production workflows inside a company that you shouldn't actually try to delegate to an AI. We'll see how long that lasts. Obviously, AIs are getting better and better, but I struggle to think of a world where there aren't certain jobs for which deterministic code is still more reliable and cheaper, and certain jobs that it cannot do.
For those, you need the ghost in the machine. You need the AI that can tackle those tasks. I think the right thing is that you're trying to teach the agent how to go do that on its own. When you talk to it, it goes and builds those things in an optimized way instead of saying, “I'm just going to build an agent that runs agentically every single time.” It has a better sense of which is the right tool for which job.
In that one example you gave, if I understood it correctly, first of all, it sounds like there's probably a ton of context that comes into the test environment with that short prompt, right? It has to navigate, find the policy, and parse that, and all those other things require additional information-finding and understanding.
It also struck me that that sounds like the kind of thing that only happens once, or you might acquire a handful of companies, but you're not going to acquire the volume of companies that one would traditionally associate with a Zap. How are you seeing usage change? Or what advice would you give if you used to think, “For me to go to all the trouble to make a Zap for something, I need to at least expect that Zap to run 500 or 1,000 times”?
That example is one that happens every day, if not multiple times inside a company. You close the deal. If you're a good organization, you're closing deals all day, every day. There are so many workflows inside of a company that are kicking off all the time, perpetually, and in some cases they're happening at a rate that humans can't keep up with.
If you're operating at the scale of some of these companies, they're selling to consumers who are making purchases many times a second. You have to use automation. You have no other choice. You cannot put humans in the loop for these tasks.
Have you seen big shifts in terms of people moving away from building it out themselves and having AI do that? Have you also seen the scale threshold at which people start to think about using automation software come down substantially because AI can do the setup?
I think the big new opportunity is to have the AI do the build. It's able to work through the logic much, much faster than a human, and the idea of building in no-code feels antiquated to me. It's like the new no-code is code.
What I think we have learned is that humans still very much benefit from visualization of those workflows. That visualization helps them verify, “Is this thing doing what I intended it to do?” It acts as documentation, so you can share it with other people and say, “Hey, here's the thing I built. Here's what I'm doing.”
If you think of Zapier in the old-school way, as this thing that had a bunch of boxes, and you're coming in and using that to configure it, by and large, I don't think that is the way people are doing it now or even in the future. Instead, they're talking to the agent and having the agent make those edits for them.
The other nice thing that's happening in the future is that you're going to see stuff get hardened into a deterministic workflow. But even when it fails, you can have the agent troubleshoot and just fall back and say, “Why did it fail? What happened here?” Then it can use its reasoning to fix the workflow or fix that instance of the workflow.
You start to see the agent almost take over the human role of building and maintaining it. But what is actually running is still a very deterministic workflow, with AI interwoven in the places where it is most necessary.
Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Athena, the executive assistant company on a mission to improve how people work and live. If you want to increase your impact, you have to free up your time, and that's what Athena does best. They match you with a dedicated, full-time, top 1% executive assistant who can take over your inbox, calendar, travel, and everything else that's quietly eating up your week. Athena is SOC 2 Type 2 certified, so you can rest easy knowing that your sensitive data is in good hands. And as a former AI advisor to the company, I can personally vouch for how much they've invested in AI tools and training. In fact, one of the very best AI users I've ever met is an Athena client who delegated the exploration of AI tools and the development of AI workflows to his EA. Athena clients report saving an average of 15 hours a week, and the average client refers more than two friends a year. That, to me, checks out. I was an Athena client while running my startup, and to this day, I continue to refer friends. Go to athena.com/cognitive right now and get matched with your EA. That's athena.com/cognitive. Give yourself back a few hours this week. Go to athena.com/cognitive and see who they'd pair you with this month. This episode of The Cognitive Revolution is brought to you by OutSystems, the leading agentic systems platform. We've all seen the headlines. Companies are pouring money into AI. But the big question on every executive's mind right now isn't just how fast can we adopt this, it's where is the ROI? We're seeing a real trend toward AI chaos. You've got teams deploying standalone coding tools, random agent builders, and experimental scripts. It sounds innovative, but in reality, it's creating a massive headache. Fragmented tools, ungoverned data, serious security blind spots, and costs that are spiraling out of control. If you don't bring those agentic applications under control now while they're still embedding into your core processes, you are looking at broken systems and damaged customer trust down the road. That's where OutSystems comes in. OutSystems is the leading agentic systems platform for the enterprise. Instead of managing a patchwork of disconnected tools, OutSystems lets your team engineer, orchestrate, and govern your entire agentic ecosystem on one open, unified platform. It's built for the speed of AI, but with the reliability and security that enterprises actually require. We're talking about real results, like KeyBank, who used OutSystems to deploy a customer-facing app that delivered 75% faster onboarding times, or the global logistics leaders who built agentic systems to completely eliminate their engineering bottlenecks. You don't have to choose between speed and control. Whether you're a small team or a massive enterprise, OutSystems helps you engineer, orchestrate, and deploy agentic systems that actually scale. Stop chasing the hype and start owning your agentic future. You can see how it works and learn more at outsystems.com/tcr. That's outsystems.com/tcr. So when you go to the 40% or so success rate on Automation Bench, what can you bring that up to as you move from just asking Claude to do it to giving Claude Zapier and potentially iterating a little bit on initial failures? What kind of lift do you see, and what sort of token savings do you see over time?
You're going to have to wait for Automation Bench V2 for that, because that is exactly the right question: What happens when you give these models access to tools like Zapier? How much higher can you get the efficiency, or how much higher can you get the scores, but also how much can you pull costs down? How much faster can they go?
I think there are many dimensions when you've got a model with access to tools and capabilities and certain things like that where it's going to score better on these benchmarks at the end of the day. That's where I think the application layer has a lot of room to grow: What happens when you give the model access to these extra things? The model is just going to get better at performing against a whole host of tasks.
How do you put the model in a position to be successful when something goes wrong and the model has to come in and troubleshoot and debug? I'm sure that you've done who knows how many things over time to try to set the models up so that, within one generation of a model, it can come back and be able to fix the things that it got wrong the first time.
I think there are 2 places that are interesting to talk about here. The first is when you think about building these workflows. What you want to do there is make sure that the model can properly identify when it needs to be using AI versus where it should just be writing code. There are many examples where you don't want the agent to actually orchestrate the task; you just want it to run code that already exists.
A lot of that is about when the agent is building its plan. You want to make sure that the plan encodes, “This is what the optimized workflow looks like,” and that it has a good plan to do that. The second place is what happens when things break and how you recover from that.
Here, we’ve done a fair amount of work, even in our own support organization, trying to figure out how we troubleshoot on behalf of customers and then bake that troubleshooting back into the core product. One of the big learnings is that we just have multiple agents run at it. We have this auto-email program that’s running right now, and for that, it spins up 5 independent agents that evaluate the troubleshooting situation.
We noticed that when 4 of the 5 agents tend to agree, there’s a pretty good chance that’s actually the issue they’re encountering. There’s a lot of data and measurement that goes into that, where you’re just trying to hill-climb and see how you can try a different model or a different prompting technique. How do you measure that stuff when you have humans auditing the output?
And what those humans do is they basically either give a thumbs up approve or they give a thumbs up rejection and a reason why. Those reasons kick back in and help improve the overall system. You’re just going through this loop over and over again to continue optimizing your ability for that workflow to have a higher chance of getting the outcome that you want.
Is that loop and all the data that you’ve collected over time—which I guess must be quite massive—core to Zapier’s defensibility these days? How do you think about what the hill is that you’ve climbed that will be hard for others to follow you up?
I think that’s a big part of what it boils down to. It’s about trying to identify what things are unique to you and what things others can’t replicate easily. Certainly for us, we’re really good at automation. We have tons of data on what it takes to do, and we hook into everything. How do we take that data to actually make our products better?
It has to be meaningfully better than what somebody without access to that could do. I think this is where a lot of the incumbents have an advantage: If they’re able to wield that to actually build a better product at the end of the day that the models alone can’t build.
There’s so much room for this because the models are great generally out of the box. All of us experience this in our own lives, where you hook up your Gmail inbox and now you start asking, “Help write an email.” It automatically does a better job of writing email because it sees how you write email. You’ve done very little in the way of trying to tune that workflow. It’s just by hooking up your company’s data that all of a sudden the models get better.
Using your own company data to build an edge is a pretty spot-on technique these days.
So does that look like a big retrieval problem for you? Certainly in my email, I’ve got a lot of history, and finding the right example to take inspiration from is probably, especially if it’s just using Gmail APIs and doing keyword searches, just as hard, if not harder, than actually taking inspiration once you’ve found the right documents to take inspiration from.
In my personal context, I’ve tried to help it out by exporting all that stuff, doing embeddings, and using various kinds of alternate search approaches so that hopefully the right content comes to the top more often. If I’m putting 2 and 2 together correctly, it sounds like at Zapier you probably have a database of a zillion things that have gone wrong over time. To help inform the model of how to fix this particular situation, you’ve got to dig in and find analogous situations.
What does that look like? Is there an embedding model that would do a good job of that, or have you had to innovate at the retrieval-stack layer in order to make that work well for a use case such as Zapier?
Coming back to the support example, a lot of it is just about taking the example that comes in, giving it the old thumbs-up or thumbs-down, providing a reason why, and doing that over and over again. What that looks like for our customers is just giving them the same tools to do the same.
A lot of this isn’t particularly fancy at the end of the day. It’s just: Take the example, did you like it, did you not like it, and rinse, wash, and repeat.
Interesting. How do you think about competition in general? It sounds like, on the one hand, we should all be worried about frontier model companies eating our lunch.
Even me, as a humble AI podcaster, look at NotebookLM and think, “They’re coming for me in my rather unlucrative niche.” But you could say, well, those guys are only going to sell their own models, so they’re a different type of animal. We don’t have to worry about them.
There’s a variety of new tools coming online to try to be the Uber tool. I’ve done episodes with Composio, for example. Zero.xyz is kind of out there. And then there are other daily-driver, sort-of agent-builder-type things. And then there are other big incumbents—you’ve mentioned Salesforce—and, to some degree, maybe it’s just big incumbents with lots of data and lots of resources all ending up colliding with each other. Like—
Which of those classes of competitor do you think are actually the ones that you need to be most concerned with?
We’re in an interesting period, for sure. To your point, everyone gives a lot of attention to the labs and tries to understand what they’re doing. I think that is important. You want to understand what they’re going to be great at and what they’re going to hill-climb at.
But I still remember PG’s advice when we were going through YC. Back in the day, it wasn’t, “What if Anthropic builds you?” or “What if OpenAI builds you?” It was, “What would happen if Google built this? What would happen if Facebook built this?” That was always the question.
The thing that PG tried to instill in folks was that you’re not often competing directly with Google. You’re not going toe-to-toe with Larry and Sergey. You’re not going toe-to-toe with Zuck. Oftentimes, in the things that entrepreneurs are trying to build, you’re going toe-to-toe with a potential mid-level product director who’s trying to get a promotion, might be there for 2 years, and then bounce.
And the reality is, OpenAI and Anthropic are big tech now. These are not small, tiny startups. They’re obviously capable of building incredible things, and they’re going to be the best in the world at these foundation models. They’re going to be incredible at that, and they’ll have good products elsewhere, but they can’t build everything. They just can’t.
And so that’s where I think it gets really wide open. I look around, and it is a little confusing because you’ve got everyone who does seem to be building everything. There’s a sort of sea of sameness out there that is a real challenge at the moment. On the flip side, you go talk to the average user of AI tools, and they’re candidly not doing much. They might have used ChatGPT or Gemini.
So, to me, I think for most of us, our competition isn’t each other. It isn’t the tools that you talked about. It’s whether people actually know what to do with these tools yet. They just haven’t adopted anything at this point in time.
The real challenge is whether you can actually get your hooks in somebody whose only experience with AI is using Gemini in a default Google search, or if they’re using Microsoft Copilot at work. That’s where most people are, and I think that can get so easily lost in the shuffle if you hang out on X all day.
On X, we’re all just hyper-aware of what model came out, what tool is gaining traction, who just raised a huge amount of money, and we’re keenly aware of these micro-differences between different products. Most folks just don’t know that.
I think the challenge we have is really making sure we’re keeping those 2 competing thoughts in our head: Yes, we do have to be better in some dimension than all of this sea of sameness, and yet, at the same time, the opportunity is massive just to educate the masses on how these tools can work. These markets are enormous.
The market for automation was orders of magnitude bigger than I ever thought it was when we started the company 15 years ago. It’s probably 1,000 times bigger than what we set out to do. There’s plenty of room for us to solve problems for customers who, candidly, don’t know about any of the competition you just rattled off. I think that’s the biggest challenge for many companies today.
Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic. By now you know my story. Claude drafts my intro essays and I rewrite them, not because the drafts are bad, but so I can stand behind everything I publish. Well, I have an important update. Claude Fable 5 is the first model to have me rethinking my rule. Today, I now think co-authorship, not sole ownership, should often be the goal. Where the model excels, rewriting its work can be more about vanity or a misplaced sense of duty than integrity. I feel it most in songwriting. I'm no lyricist, but I'm good with a song concept, and Fable writes some amazing verses. I give it feedback on its misses, and I push it to aim for higher inspiration, add layers of meaning, optimize syllable density, and above all, write a hit song. These days, I get compliments on just about every song we write together. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you. Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai/tcr. That's claude.ai/tcr. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Once more, that's claude.ai/tcr.
I’m working on an episode where I’m going to talk about things I’ve been wrong about and things I’ve been right about. One of the things I’ve definitely been wrong about is that I expected a lot more change in how things get done across the economy 5 years ago than we’ve actually seen.
Especially if you were to tell me then that we’d have Astra and that it would have been, roughly speaking, a smooth ramp to these capability levels, and yet we still see revenue exploding at the model companies, obviously, but we don’t see nearly as much change as I would have guessed.
Life is kind of the same. I go about my day, I talk to my friends, I talk to my family, and I watch what they do, and yes, their days are the same. They use ChatGPT to help with things like meal planning, planning a vacation, or doing a workout. They’re certainly not taking advantage of Astra-level model capabilities at all.
So what have you learned, then, about what kind of help people need? If you—I don’t know, do you talk about this when you’re taking walks in your neighborhood? How do you...? I’m interested in your method, but even more so in your takeaways.
What is it that gets people over the hump? How do you help them see a new way of working? Are there any patterns that seem generalizable, or is it idiosyncratic for every individual and small company?
There are definitely patterns, but a lot of the specifics matter, and that’s where it starts to feel idiosyncratic. We’ve seen this for years. I remember that, for a long time, one of the hardest problems at Zapier, even to this day, has been helping people with recommendations. What should you actually use this stuff for?
If I personally sat down next to you and said, “Hey, just show me what you do every day,” I could come up with half a dozen examples that would immediately have you saying, “Yep, I want that. Yep, I want that. Yep, I want that.” But then how do you actually bake that in? How do you give them the experience of having me, or somebody who’s knowledgeable about these areas, sitting next to them, and bake that into the product?
This is where I get pretty excited about how AI can help cross that recommendations gap, use-case gap, or whatever you want to call it. If you’re able to point the tools at where you work and say, “Hey, I want you to go watch what I do every day. Watch what I do in Gmail. Watch what I do in Slack. Watch what I do in my browser. Watch what I do in my chat. And just tell me: What should I be doing differently?”
I started doing this workflow at the beginning of the year, and pretty much every week now I have new tools and new systems that start to automate bits and pieces of my job. If you do that on a perpetual basis, you start to feel the difference. After a month or 2, you’re like, “Wow, there’s a lot that’s kind of running for me now that I didn’t have before.”
And I think even though many of the things that are built are things I knew I should have been doing before, it’s the specifics. It’s like, I literally watched the things you did here and here, so I know exactly the tool that can get it done. It’s that idiosyncratic part that makes it click.
I get pretty excited about how you give these models awareness of just how people go about their day. I think there are tons of ideas that are trapped, latent, and lost inside that, which most of us just don’t wake up and think about.
We are creatures of habit, so we wake up and go about our day the same way we did the day before. We don’t think, “Oh, there might be a 5% better way of doing it,” or, “There might be a 500% better way of doing this.” I have muscle memory. I know how to do it this way, and I’m comfortable doing it this way, so I’m going to keep doing it that way, even if it’s not the best.
And so AI has to be so good that it knocks us out of our comfort zone and we're willing to say, “You know what? I am going to go try it that other way because that sounds so, so much better.”
So how have you set that up for yourself, and how broadly deployed at Zapier is this sort of AI on your shoulder or screen recording? I'm kind of expecting a screen recording.
Most people use Zapier MCP for this. They just have all their tools hooked up into whatever harness of choice they use. Like I mentioned, we have our own internal harness, but it could be Cloud code, it could be Cowork, it could be ChatGPT Work, it could be Cursor, it could be anything, right?
They'll just have an automation that runs once a week, and it collects all these signals across all the work they've done. It sees the event streams, and it can tell. It says, “Hey, I noticed you did this and this, and here's a tool that I think you should go build.”
Most people have something like that set up, and then they just tell their agent, “Okay, great, I like that suggestion. Go build it.” Or, “That suggestion's okay. What would make it great is if you made this tweak. I really want you to go do that.”
Half the battle is just getting you to react to something. Even if the ideas aren't perfect, they only need to be 50% good enough to get you to go, “Oh, I see where you're going with this.” Now that brainstorm process kicks off, and you're able to run with it.
So do I understand correctly that it literally just uses APIs to look at your digital history for the last period of time, and then—
More or less.
…kind of collates that together and comes up with ideas? Interesting. Is this something you think you'll productize for Zapier customers?
Yes. I think it's very likely that'll come in some form factor.
Interesting. It's an interesting way of doing it. Obviously, I don't need to tell you, but one downside of going that route with your customers is they'll have to attach all these different platforms that they use first in order for you to have access to get any insight as to what's going on across all these things.
Whereas with a screen-recording type of thing, you have one install and you just look at what they do click by click and make sense of it from a top-down perspective, I guess, as opposed to what you're describing. That sounds a little bit more bottom-up: “Oh, I saw this in Drive and this in Gmail,” and what have you.
Do you think there's a principled reason or an empirical reason for going one way or the other?
We're good at APIs, and we're good at that stuff, so that was just an easy, natural, emergent experiment—an emergent property. I think for this experience to be great, you should use all the tools that you have available to you.
Yeah. It's interesting. It also reflects the alien nature of AI intelligence in some ways. If I was going to try to advise you, I would definitely want to watch you work. I would not be so helped out by your logs. But AIs are really good at reading logs.
How do you think about pricing in today's world? This is obviously a very open question for a lot of companies.
My point of view is that, for the most part, seat-based pricing is dead or dying. I think it may still make some sense in some small areas, but by and large, when intelligence is such a core part of these product experiences, I don't see how a fixed, seat-based pricing mechanism really makes much sense for that at all.
Inevitably, I think that lands you in some sort of usage-based or outcome-based pricing world. I think products end up choosing which side of that fence they land on. If you're going more the commodity route, you're probably closer to usage-based pricing. This is the Sam Altman “we want to be a utility that you pay; you can have intelligence on tap” idea, et cetera.
Maybe if you're a little more enterprise-oriented, you're going to try to say, “Hey, I'm going to price per outcome.” You see a lot of the customer-support tools doing this, where they're going to bill you for resolved tickets, because they have such a clear demarcator of what success looks like.
I do think that if the product you're selling has the ability to have such a clear, agreed-upon, fixed outcome, that's probably to your advantage. The challenge is that, at least for most of the products right now, it's way messier than that.
You stop 1 or 2 steps shy of truly delivering the outcome. You are part of delivering a piece of the outcome. So I think that pulls you back into a more pure usage-based model.
Either way, I think you've got this meter running where you're basically selling work of some portion at the end of the day. I think that's where a lot of this stuff is going: we're going to have to think through what our budget is for work to be done.
Do you worry about price discrimination by the model companies? Of course, you're well aware of the ratio of tokens that you get with a Claude Max or an OpenAI Pro plan, and how many more tokens you get, at least if you max them out, compared to what you can buy with the same dollars via the API.
I've asked a number of entrepreneurs this question, and I'm wondering whether you would be supportive of some sort of rule that said, “Hey, you've got to charge everybody the same for tokens so that an ecosystem has more of a fighting chance,” versus OpenAI giving, say, a 20-to-1 token advantage. That could be hard for third-party value-added services to compete.
Yeah. I definitely tend to fall on the side of free markets, and these are companies that have the right to price how they like. But many of us lived through, and are still living through, Microsoft's dominance and how they used bundling to their advantage to box out better products, candidly.
Because it's all just bundled there, it plays to their advantage. I think this is where countries and folks get to decide what they think are monopolistic practices and what they think is a fair playing field.
Ultimately, I think my job is to play by the rules on the playing field and not necessarily decide. I definitely lean more toward the free-market side and say, “Hey, our job is to come up with an edge that helps us compete there.”
I don't fault any company for wielding the tools they have in their tool chest to make products work for them, work well for their customers, and help them maximize revenue. That's well within their right.
It's been interesting to see who's been willing to bite the bullet versus who has stuck to their free-market principles.
Changing the topic toward operations and AI transformation within Zapier, one thing that caught my attention was that, if my AI research agent is to be trusted, you put your chief people officer in charge of AI transformation. I believe last time we talked, you had said, “Well, it's not any one person's job. It's kind of my job as CEO, but it's really everybody's job. So I'm not going to say it's one person's job.”
What changed, and how did you decide it would be the people officer who would shoulder that burden?
I still agree that AI should be every person's job, and it should be my job. I think it depends on what stage you're at and what problems you're facing in terms of how you think about who you want to tackle the next mountain, so to speak.
For our first chapter, a big part of Zapier becoming AI-fluent was everyone in the company getting up to speed on how to use these tools. There wasn't an AI committee. There wasn't an AI group where it was like, “Oh, they're there. They kind of figure it out. The rest of you, it's business as usual.”
It's like, no, this is important for everyone inside the company. It impacts everything we do, so all of us need to get on that.
As time went on, there were a couple of interesting things that we started to observe. First, the AI fluency inside the company went up. Basically, within a year or so of ChatGPT launching, almost 100% of the employee base was using AI day to day.
We're not having technical issues adopting AI. That's not where we're bearing the brunt. One of the bigger issues that is starting to emerge is how you take these models from individuals having success to actually using them to solve bigger and bigger production-grade workflows across the company.
The challenges start to look a lot more like people issues. We have to rewrite certain job descriptions, rethink how we do compensation, and think about how these teams stand up. We have to move this group—we kind of don't need this group anymore—but we actually need more people over there. So how do we retrain and reskill these folks who have some of those skills but need to learn some new skills?
It turned out Brandon, our chief people officer at the time, was really good at doing a lot of these things. His team was at the forefront of some of this inside of Zapier.
The thought for me was, “Hey, you’re doing a good job at this. Why don’t you go help everybody in the company figure out some of these things?” It could have just as easily been a CMO or a CPO or any number of roles. That’s how it’s played out inside of Zapier. And it’s been funny how much I get asked this question now, because I think a lot of folks think I have a point of view that it must be a chief people officer or something like that.
It was really just that, at this moment in time inside of Zapier, this sort of felt like the best person to go tackle it based on the problems we were facing. And I think that’s the way you should do it inside your company. You need to look at what your bottlenecks are, what your constraints are, and go identify the person who is the right fit for that job.
Yeah. Echoes of Ben Horowitz’s advice, too. You’ve also made an interesting move of really trying to push people toward internal communications being public within the company by default. And I’m interested in a couple of finer points. One, how did you handle historical data? Did you start that policy at a certain point in time, and everything in the past was left in the past? Or did you try to reclaim some of that knowledge, which I assume would be very tempting to do?
And then do you have different tiers of public as well? Because it strikes me that you might not want everyone to know everything, but you might want different groups to have certain different databases. So I’m just looking for the double-click on how you’ve operationalized public by default.
One interesting thing is that Zapier’s had this value of default to transparency for, gosh, forever—it feels like. By and large, we were already working in public for many such things. We already had a culture where there were tons and tons of public Slack channels, and people were talking about the projects and their day-to-day in those quite a bit.
I think a lot of what we observed was that, as the company grew, there were some pockets of work that started to find their way into private channels or private DMs and things like that. By and large, Zapier was still much more public than most companies. Slack gives you the readout, so you can see what percentage of stuff is happening in private and public and all that sort of stuff. We looked at that and thought we could use a little bit of a reminder.
The second thing that encouraged us to do this is the fact that our AI agents were so much more effective when they were able to see the context inside of Slack. One of the things we started to do on our executive team was have a little fun competition to see who could put most of their communications in a public channel. There was no, “Oh, you must do this,” or “Anyone under this rate gets a bad review,” or anything like that. It was literally just friendly competition.
We noticed that more things could go in public than we realized, and this seemed to help the team. People know what’s on our mind. Nothing is hidden, et cetera. So we started to encourage that all across the company.
There are certainly things that we still pull into private channels and things like that. If there’s an HR incident or something like that, we’re not resolving that in a public channel. If there’s a critical security vulnerability, we’re not resolving that in a public channel either. Those are happening in private channels, especially while the incident is active. Once they get resolved, we tend to share out the learnings and things like that.
There are certain topics where you still need to set up these spaces where you can go resolve them in private. But by and large, most people far overestimate the number of things where that is required, and definitely underestimate the power of what happens when both humans and agents have access to the full context of what a company is working on.
So, speaking of security, this is obviously top of mind. As I was thinking about challenges that AI might pose to you, you’re holding potentially more credentials to more different services for more different users than just about anyone in the world, right? So I would think this is kind of a scary moment. All of a sudden, where’s Bedrock in terms of security?
How are you approaching that, and are you trying to get into these early-adopter, Glasswing, and other clubs? Do you think that’s actually maybe a big source of differentiation going forward? And how scared should I be about cybersecurity? Because I’ve got a lot of credentials all over the place, Zapier and otherwise.
We’ve held these credentials for 15 years, right? So this has always been an important thing inside of Zapier. We’ve said, “Hey, these credentials are a thing that we must treat with the highest, highest level of stewardship.” We’ve always put a lot of effort into making sure that we do a good job of protecting those for our folks.
What feels different this time is that you do have these mythos caliber security models that are able to patiently loop over issue after issue and find things. Every software project has vulnerabilities; it’s just that someone hasn’t found them yet. The models make it a lot easier to find those things.
The good news is that they also make it easier to patch them. I think what smart companies are doing is basically wielding them for offense and defense. They’re trying to find this stuff faster, and they’re trying to resolve it faster.
I’m not exactly sure how all this is going to play out. Every day, there’s funky stuff going on. Obviously, the Hugging Face incident was straight out of a sci-fi book, right? But for most folks, my guess is it feels like there’s this almost one-time investment to reacclimate, and then you get back to a more steady state of offense versus defense in terms of security posture as this stuff moves forward.
It’s going to be really interesting to see, because every day we’re seeing new stuff.
Are you taking steps as CEO to try to make sure you’re on the inside of early-access lists for new models?
Yeah. We want to have access to the best capabilities as early as we can. I think everybody who’s in the same shoes would want to do the same.
Yeah. It feels like that could actually be a pretty meaningful point of differentiation going forward. If one company that’s going to hold my credentials is in all the clubs and another one is a startup that’s not, that’s a big leap of faith to take on the company that doesn’t have the same kind of access to be trying to find and fix all these issues.
In terms of spending, I saw you tweet not too long ago that you have some engineers spending $30,000 a month on tokens at Zapier. It struck me that we’ve been on quite the yo-yo ride recently, with token-maxing and then budgets being hit. What do we do about it? What sort of process or governance do you have for who can, under what circumstances and with what approval, spend tens of thousands of dollars a month on tokens?
Right now, I would say that those individuals are a bit of an outlier. But it still encouraged us to start building some tools to help people do some self-policing. Mostly, we don’t have budgets set up for individuals yet, but we do have tools where they can see their spending and better understand what happens when they choose a powerful model versus when they choose a cheaper model on certain workflows. They can see what those cost differences are.
As we see people starting to spend a ton on tokens, usually the first reaction is, “I just want to go talk to them and say, ‘Hey, what are you doing? I’m just really curious.’” In some cases, you have folks who are doing some insanely productive stuff. In other cases, you have some folks who have a mix of things that are pretty productive and places where it’s, “Oh, you don’t need to be using Fable for this or Astro for this. There’s a better way to do some of these things.”
When you have an almost 800-person organization, there’s a big education effort involved. Over time, I do suspect that token budgets are going to be a real thing, though. Part of AI fluency is going to be that you’re going to say, “This person is going to get a higher budget than this person because they know how to get higher output from those things.”
How do you actually operationalize that? We’re still working through some of that stuff. But it seems pretty obvious to me, just looking across the employee base, that some people are excellent at using increasing levels of spend, and some people are just not really thinking about it all that much yet.
Another aspect of AI fluency that I’m really curious to get your take on is what you think is the right model for co-authorship or co-creation with AIs.
Sure.
And this is not a gotcha, because I’m in the same boat. I’ve consciously tried to almost shock-expose myself recently, to put some things out that I didn’t rewrite every word of. So my Pangram score at times says that my stuff is AI.
I’ve also seen some stuff in various places from Zapier that has a high Pangram score. How do you think about that, and how do you set the tone for others at the company? You want to be using these things—
But we don't want to be putting out slop. What's the line?
The way I think about it is, I have no problems with people using AI for communication at Zapier. What I really have a problem with is low-quality communications. We live in an era where AI can enable a person who is exercising low judgment to create a high volume of very low-quality stuff very quickly, and that can overwhelm a person.
We've tried to put a few guidelines in place that help people think through ways to go about that. For one, you need to own what you send. If you wrote it, you probably should be putting more time into authoring the thing than the reader is reading it. AI use shouldn't be a way of transferring ownership of a task, where it's like, "Oh, I was assigned this task, so now I had AI spin a prompt," and I said, "Hey, you now read it and deal with all this stuff and edit all those things." That's not a great way of going about it.
You need to understand what you send. If someone starts asking you questions and you're like, "I actually don't know what's inside of that," that's not good. You probably ought to be making asks explicit. If you need somebody to do something, if you need a decision, if you need feedback, or if you're labeling something as a draft and you want feedback on the thing, you need to do so.
You need to go verify details. It's not uncommon for AI to hallucinate some of these details—or maybe it's not hallucinating. It might pull dated information. So if you hook it up to an agent that has access to Zapier's Slack, it could pull, "Oh, this project from 3 months ago is related, but not the exact same thing." If you're trying to pass that off, that becomes a real issue.
An AI summary can be based on a summary of a summary of a summary, and all of a sudden, before you know it, it's actually passing on incorrect information. You have to do a good job of verifying the details in there. To me, that's the really important piece: you are still an active participant in the creation of the material.
But if AI is helping you structure your thoughts and structure the writing, at the end of the day, go for it. I don't have any problems with that. It can be tedious if you're not scrubbing some of the slop that's a part of it. The "It's not this, it's that," the em dashes, the "honest truth," the "load-bearing point"—all that kind of stuff.
I do think that if you're doing that a lot, especially if you're doing it in marketing material, it makes it hard to stand out. People get a little tired of reading that kind of stuff. You still need to have your own editorial hand on the steering wheel, so to speak. To me, using AI is not the problem. Low quality is the fight at the end of the day.
How has your team composition changed over the last couple of years? You mentioned earlier that maybe we don't need this team, but we can reskill. Are there any thresholds for AI capability—something where you're like, "Well, they can't do this now, but if they could, I could see that actually making a big impact on our hiring plans going forward from that point"?
What's interesting is that, in some ways, our team looks very similar to how it has in the past, and that's maybe a surprise to me. But in other ways, it is pretty different. We still have engineering, design, product, and stuff like that inside the organization, but the idea of a classic, traditional EPD is largely gone. Things are a lot more malleable, but the roles still exist.
There's definitely been a flattening of management layers, but strong management is still crucial. We're not getting rid of managers anytime soon. Managers can handle a higher volume. Those are a handful of things that are interesting.
Similarly, with data analysts, everyone inside of Zapier is kind of their own mini data analyst now, so you don't need as many data analysts. And yet, we still have analysts who are doing really critical, important work inside the company. They're just working on higher-value stuff now.
You can feel things shifting, and yet, in some ways, it still feels pretty familiar at the same time. The company feels quite a bit similar.
You asked what the models are not yet capable of that I'm excited about. I still have this idea of the team shifting more into building the factory that builds the products, the company, the marketing, and all that sort of stuff. You can start to feel where we're doing more and more of that. We have a software factory, a support factory, and workflows that are getting stood up where they're handling the inner loop and the humans are more focused on designing that piece of the puzzle.
Inside those factories, there are all sorts of steps where you come across areas where you're like, "The AI's not quite good enough for this yet. We need a human in the loop." But with every model release, with every iteration of our own learning loop inside of Zapier, you can start to feel us chip away at that problem. We're just getting closer and closer to something that looks actually pretty different from the organizations of the past, and that's pretty exciting, I think.