[BidClub_]
Moonshots · · 82 min

Which Industries Survive AI, The New AI Benchmarks, and the 2026 Recursive Learning Timeline | #218

Peter DiamandisMatthew FitzpatrickSalim IsmailDave BlundinDr. Alexander Wissner-Gross

YouTube
TL;DR
  • AI will disrupt documentation-heavy sectors first, not flatten every industry at once. Fitzpatrick puts media, legal services, and business-process outsourcing in the blast zone, while oil and gas or real estate retain much of their underlying function; the competitive question is whether “the startups get distribution before the big companies build the technology,” especially in banking, where much of the application footprint is over 20 years old.

  • The missing enterprise-AI asset is not another broad model score but task-level benchmarks. Public coding benchmarks show models improving 50% to 100% across many dimensions over three years, yet businesses need “accuracy or human equivalent on a specific task”—from claims processing to title insurance—and an “80% accurate very smart deployment” still carries too much production risk. The discussion points toward thousands of narrow benchmarks, whose ownership could create visibility and leverage in neglected industries.

  • The practical 2026 playbook is to follow value, pick two or three use cases, and put an operator—not the technology department—in charge. Fitzpatrick recommends reaching a working prototype in roughly a month, tying it to metrics such as CSAT, inventory days, stockouts, or cost per call, and asking: “Would you bet your annual bonus that whatever use case you deploy works?” For the first project, he favors an outcomes-based RFP to a third-party vendor.

  • Full autonomy is the wrong initial architecture for most enterprises. Klarna reportedly said its AI handled 2.3 million calls a month, did the work of 700 full-time agents, and could save $40 million annually, only to reverse course 8 to 12 months later; Fitzpatrick’s lesson is to route routine work to agents while escalating complex refunds, system write-backs, and customers seeking human contact. “Human in the loop is going to be a feature, not a bug.”

  • General-purpose models may commoditize model-building, but company context, custom evaluation, and data security remain enterprise concerns. Fitzpatrick expects enterprises to tailor frontier models rather than pre-train their own, with swappable models sitting beneath company-specific documents, workflows, and benchmarks. The key is separating genuinely proprietary information—such as Jane Street’s trading data—from ordinary back-office data, using on-premise or small language models where warranted instead of treating everything identically.

  • The strongest deployments begin with narrow data assembly and already show measurable operating gains. Invisible combined 750 SwissGear tables, expanded overall inventory coverage about 30%, and doubled the number of SKUs with reliable predictions in a couple of months; it also built player-movement models for the Charlotte Hornets, a HIPAA-compliant control tower for Lifespan MD, and underwater-drone decisioning for the U.S. Navy. The repeated pattern is specific data, a bounded decision, and an observable result.

  • The episode’s central disagreement is whether recursive self-improvement soon eliminates specialized human feedback. Wissner-Gross gives a “two to three years max” conservative outer bound for an AI researcher matching or surpassing human model researchers, while Fitzpatrick argues that tacit expertise, company-specific work, new modalities, languages, robotics, and multi-step reasoning keep creating fresh evaluation needs. He leaves open the possibility that useful training frontiers could run out 10 to 15 years from now, but argues that human expertise and human-in-the-loop work remain important for a long time.

  • The 2026 architecture call is multi-agent, multimodal, and simulation-heavy, with government process automation among the largest potential beneficiaries. Fitzpatrick expects task-specific agents orchestrated by an LLM, more audio/video/image interaction, and “mirror worlds” or RL gyms that test systems before real deployment; cited estimates suggest AI-assisted permitting could cut energy and data-center timelines 50%, while licensing, benefits, and compliance cycles might shrink 70%.

Digest · the substance, structured for research

1. AI disruption will concentrate where documentation is the product

  • Fitzpatrick rejects the premise that every sector will be hit equally. Media, legal services, and business-process outsourcing produce large volumes of knowledge work and documentation—the activities most directly exposed—while Wissner-Gross carefully narrows his own claim: “Knowledge work as we currently know it” is cooked, not necessarily knowledge workers or entire companies.

  • Oil and gas or real estate look materially different because their core functions remain. Fitzpatrick argues that deciding which apartment or office building to buy will work much as it did five or six years ago, even if AI improves supporting analysis; companies should identify the portions of their operations that can genuinely change rather than declare the whole business “AI-first.”

  • The competitive race is whether “the startups get distribution before the big companies build the technology.” Banking embodies the tension: much of its application footprint is north of 20 years old, while newer fintechs such as Revolut can build differently without the same modernization burden. Fitzpatrick does not claim to know which side wins.

  • Capability is a separate constraint from industry exposure. A 50-person company may lack even a CTO, while established IT teams may not possess relevant skills; even knowing Python can leave gaps. Fitzpatrick’s advice is unsentimental: companies unable to hire or develop the expertise should rent it through partners rather than pretend every firm can build internally.

2. Measurable baselines determine where adoption can move safely

  • Mortgage underwriting advanced because banks can back-test decisions against a statistically valid baseline and check whether the resulting credit decisions work without redlining. Contact centers offer similarly legible measures—time per call, cost per call, and CSAT—which should have made them unusually favorable territory for AI.

  • Legal work splits between advice and commodity production. Fitzpatrick expects high-end counsel to persist for a large M&A transaction, while routine NDAs and standardized documents compress sharply; Diamandis notes that venture documents repeatedly carry a $50,000 legal cap, contain “about eight knobs,” and somehow run to “$49,999.99” despite being nearly identical.

  • Klarna became the cautionary example. The company reportedly said its AI handled 2.3 million calls monthly, replaced the work of 700 full-time agents, and would save $40 million a year; roughly 8 to 12 months after becoming a flagship agentic-success story, it announced a return to human contact-center agents.

  • Fitzpatrick offers hypotheses, not inside knowledge: some customers simply demand another person, while refunds and other non-first-line cases require complex write-backs into source systems. His confusion is architectural—the sensible design was always a changing mixture of agents and humans, not “all humans, all agents, back to all humans.”

3. Enterprise AI needs a value-led operating system, not a strategy document

  • A CEO facing the board’s “What’s your AI plan?” should “follow the value.” Fitzpatrick would select two or three levers capable of materially moving the business—customer service, FP&A forecasting, inventory management, or digital marketing—then take only one or two into a real pilot rather than invite unfocused experimentation.

  • Generative AI reverses the older machine-learning deployment pattern. A prototype can be running in about a month rather than after months of construction, but reliability emerges through intensive testing and validation afterward. A strategy deck is not progress; the decisive test is, “Would you bet your annual bonus that whatever use case you deploy works?”

  • For a first implementation, Fitzpatrick recommends an RFP to a third party compensated according to outcomes. An internal team may lack experience and cannot be held to the same “you get paid if it works” standard; an outcome-linked contract transfers some execution risk while forcing everyone to define success in advance.

  • Citing an MIT report that only 5% of enterprise models had reached production, Fitzpatrick identifies organization as a central failure point. Put the best operator in charge, outside the technology organization, and assign a business KPI: CSAT or call time for service, inventory days and stockouts for forecasting. Otherwise, “a thousand flowers bloom” into unaccountable science projects.

4. Narrow benchmarks become the control layer for enterprise AI

  • Broad public benchmarks, especially coding benchmarks, remain useful indicators of general progress; by Fitzpatrick’s reading, models improved 50% to 100% on most observable dimensions over three years. Their limitation is relevance: a business does not need abstract intelligence so much as “accuracy or human equivalent on a specific task.”

  • A contact-center benchmark should compare AI agents with the company’s own expert agents across representative calls. Claims processing needs its own human-equivalence set. An “80% accurate very smart deployment” may still be unacceptably risky, so each workflow requires a custom eval built around its actual errors, thresholds, and escalation conditions.

  • Diamandis hears an entrepreneurial opening: practitioners who understand both AI and a neglected domain such as title insurance can define and broadcast the benchmark before anyone else. In his framing, declaring credible ownership of an unclaimed evaluation category can make someone “an instant star,” because the benchmark may be harder to create than the post-trained model.

  • Invisible already builds customer-specific benchmarks around individual tasks. A generic sales agent cannot simply be purchased like conventional SaaS; it must learn the company’s products, knowledge corpus, selling method, and “way of speaking,” then be tested against a local eval that determines whether the tailored behavior is actually good.

5. Generalist models win the base layer, while enterprise context stays local

  • Wissner-Gross invokes the “infamous BloombergGPT moment”: Bloomberg possessed proprietary financial data and sought domain-superior performance, yet frontier labs’ generalist models reportedly leapfrogged the project within months. His challenge is whether internal data and post-training have a durable future if general models keep absorbing specialized capabilities.

  • Fitzpatrick distinguishes building a language model from adding company context. He does not expect individual institutions to pre-train their own compute-intensive LLMs; he expects them to tailor leading models using private documents, preferred outputs, and workflows, with an enterprise layer designed so newer foundation models can later be “dropped in.”

  • A law firm’s preferred M&A documents illustrate what the public model cannot know. That information still requires local tailoring, post-training, or evaluation. Ismail extends the argument: the trade-secret knowledge of “how do we do things?” may become a company’s most valuable edge, increasing the need for protection between internal data and the broader AI world.

  • Fitzpatrick resists treating every byte as sacred. Banks, hospitals, and trading firms may keep sensitive information on-premise or use small language models, but Jane Street’s trading data and its back-office forecasting data are not equally proprietary. The workable policy classifies data by actual competitive or regulatory sensitivity rather than refusing every external model.

6. Data readiness means assembling the minimum viable truth

  • An agent built on fragmented customer and product records “is going to break by definition.” Fitzpatrick therefore starts every use case with its inputs, but he does not endorse a five-year attempt to perfect the entire corporate data lake—the path many large organizations have already pursued without making every repository accurate, accessible, and coherent.

  • Credit underwriting may need five or six central categories: the credit itself, market conditions, company financials, the security of the credit, and related variables. It does not require every record across the commercial bank. The practical question is, “What data do I need for this specific use case?” followed by focused remediation.

  • Generative AI also elevates information that never lived in a system of record: video, images, free text, and other unstructured files. These are assets that people have not historically tried to master, so the first step is identifying the task and making the relevant data ready.

  • Blundin’s Vestmark example sharpens that distinction: reconciliation records show how an account ended up reconciled, not what the employee did to resolve it. An AI assistant can first observe and accelerate the workflow, then produce the human-feedback or tuning data needed for automation. Another bank CIO faced 300 guarded customer databases, each owned by a different product silo.

7. Domain deployments work when the decision and data are tightly bounded

  • For the Charlotte Hornets, Invisible fine-tuned computer-vision models over single-point footage from multiple college and international venues. Conventional statistics capture points, rebounds, or plus-minus; the model instead measured player movement, spacing, and who creates space across inconsistent camera angles, giving draft evaluators evidence on characteristics and player fit that transactional stats miss.

  • Lifespan MD began with data architecture, not automated diagnosis. Invisible’s Neuron platform assembled a HIPAA-compliant, multi-tenant view of patients, providers, and practice performance, enabling questions such as which longevity tests men aged 35 to 50 use most. Patient data remains at individual practices while clinicians and central operators receive only the access appropriate to them.

  • Fitzpatrick calls clinical decisioning a murkier target than administration. The U.S. spends roughly $13,000 to $14,000 per capita on health care versus $2,500 to $3,000 in Germany or Canada, with something like 30% to 40% going to administration. The nearer opportunity is removing scheduling and paperwork while making physicians “even more empowered.”

  • Other projects widen the pattern without changing it: with SAIC Vantour and the U.S. Navy, Invisible worked on decisioning around sensor-rich underwater-drone swarms; at SwissGear, it combined 750 tables, expanded overall inventory coverage about 30%, and doubled the number of SKUs receiving reliable forecasts in a couple of months. The inventory work was aimed at minimizing stockouts and excess inventory.

8. Human feedback is becoming more expert, not simply disappearing

  • Invisible’s Meridial business trains models, while its enterprise side builds custom applications. Wissner-Gross asks whether an AI researcher capable of constructing models, datasets, and benchmarks would erase the need for a marketplace of human ML contributors; Fitzpatrick replies that predictions of RLHF’s imminent disappearance have persisted for five years without matching deployment reality.

  • The work is changing from commodity “cat dog, cat dog labeling” toward PhD- and master’s-level evaluation, controlled RL environments, simulations, and RL gyms. Fitzpatrick says studies and operational experience favor pairing synthetic and human data, particularly for multi-step reasoning where hallucinations can compound across the chain.

  • His best specimen is deliberately narrow: a model studying evolutions in 17th-century French architecture, in French, still needs qualified humans to validate it. The same logic applies to a new legal dataset, where an associate or M&A lawyer must determine whether the resulting document is genuinely comparable to expert work.

  • Wissner-Gross’s pushback is about efficiency, not whether feedback has any value. Reinforcement fine-tuning may require fewer human hours than large annotation workforces, and an RL environment can scale once built. Fitzpatrick concedes the form will evolve but notes that post-training feedback is a small share of total compute cost and among its most valuable inputs.

9. Recursive self-improvement may outrun corporate absorption

  • Wissner-Gross gives “two to three years” as a conservative outer bound for an AI researcher as good as—or stronger than—the human researchers building ML models. He expects recursive self-improvement capabilities to accelerate in 2026 while corporations move “at a snail’s pace,” leaving implementation, change management, and workflow bottlenecks as the binding constraint.

  • Fitzpatrick leaves open that machines might exhaust useful training frontiers “10, 15 years from now,” but sees no near-term shortage: new languages, modalities, robotics tasks, legal subfields, and company-specific workflows continuously create narrower evaluation demands. General competence does not manufacture private precedent or tacit expertise that exists only in experienced people’s heads.

  • Sales is his counterexample to total generalization. The best sellers’ patterns are often undocumented, and a market full of 500 email-based SDR vendors may make genuine interaction scarcer and more valuable. “The human-touch elements become more and more important,” particularly where trust, exception handling, or authenticity drives the outcome.

  • Wissner-Gross frames the commercial opportunity as the gap between capability and adoption: companies such as Invisible can be “the lubricant” between frontier labs and enterprises. In his tests for contact centers, 80% of people preferred the AI, while the dissatisfied 20% could “torture the whole thing to death,” creating lucrative work in routing, evaluation, data, and exception handling.

10. AI-native challengers will redesign flows rather than automate job boxes

  • Ismail’s Canon thought experiment replaces departmental thinking with a functional flow: marketing, retail sales, registration, ink replenishment, predictive repair, upselling, and accounting become one AI-managed printer lifecycle. The product reports its own state and triggers the next action, potentially leaving humans “90% out of the loop.”

  • He calls today’s employee-by-employee automation “radio over TV”—like putting radio announcers on television to read the old scripts instead of adapting to the new medium. An AI-native company starts from the desired function and reconstructs the entire operating system, while incumbents tend to attach assistants to inherited roles and handoffs.

  • Diamandis returns to distribution as the incumbent’s advantage. His proposal is to invite global AI entrepreneurs to pitch how they would disrupt the company, fund the best five, grant them data and access, then acquire or take a majority stake in the ventures that work and make them the new company. The objective is “innovation on the edge” followed by displacement of the legacy center.

  • Ismail points to Apple’s small, secret teams as the organizational model, citing roughly 18 groups examining different industries and patiently iterating until an entry is ready. Large operators may struggle to attack their optimized core, but their knowledge of adjacent industries can support AI-native ventures aimed at neighboring markets.

11. Multi-agent teams, multimodality, and mirror worlds define 2026

  • Fitzpatrick’s first architecture call is the multi-agent team. Rather than one agent making every decision, enterprises train task-specific agents to high accuracy and place them under an LLM that orchestrates the broader logic. Contact centers are a natural example.

  • The second call is a multimodal leap. Audio, video, and images should become much larger parts of how people engage with models, moving interaction beyond historically text-heavy interfaces. The earlier basketball, health-care, and drone examples show why the enterprise data surface is already multimodal even when its software interface is not.

  • The third is the “mirror world” or RL gym: simulated environments and digital twins in which a coding, contact-center, or manufacturing system can execute tasks and be tested before touching production. Simulation supplies repeatable tests for systems whose real-world mistakes would be costly or dangerous.

  • Avatars are another likely 2026 development. Fitzpatrick says training one from his public statements would not be difficult and notes sports-related avatar work already underway; he expects people may prefer conversing with recognizable people via avatars over generic chatbots, making synthetic personalities a more natural part of interaction.

12. Human work shifts toward physical presence, authenticity, and new categories

  • Asked for the “last expert standing,” Fitzpatrick names broad categories rather than three precise occupations. Seismic work, drilling-site operations, real-estate selection, physical trades, and jobs built around human interaction survive longer because their function is not simply searching documents. Nearer-term disruption remains concentrated in BPO, legal services, and media.

  • He does not equate task disruption with falling employment. Media’s economics and channels changed, yet Substack, Medium, blogs, and other formats created more media entrepreneurs. Fitzpatrick cites estimates that roughly 25% of each U.S. high-school class eventually enters a field that did not exist during high school, while 20% of U.S. employment now consists of digital-ecosystem jobs.

  • Wissner-Gross offers three competing candidates for the final roles: politicians, because they make the laws; the greatest physicists and mathematicians, as the peak of human intellectual work; or occupations where customers demand an authentic human counterparty. Diamandis summarizes the third hypothesis as “tastemakers will dominate.”

  • Government may provide the clearest social return. Fitzpatrick cites a study suggesting AI-assisted permitting could cut energy and data-center implementation timelines 50%, and an OECD estimate that licensing, benefits approval, and compliance cycles could shrink 70%. He identifies project management and timelines for spending and infrastructure deployment as a simple, high-value use.

Peter Diamandis

Most of the public focus today has been on the large public benchmarks for things like coding. I think the problem is, though, Matt Fitzpatrick, CEO of Invisible Technologies.

It sounds like your position is that we need thousands of new narrow benchmarks to capture perhaps every labor category and every industry vertical. That is an interesting second part of this, which is:

Salim Ismail

We're going to see the largest disruption ever in 2026 from companies that don't make this change.

Matthew Fitzpatrick

There are many sectors where the structure of what the industry does is going to change. If you think about knowledge work—the production of large amounts of documentation—these technologies are very disruptive.

Peter Diamandis

What are you seeing most companies get wrong on their mission to implement AI?

Matthew Fitzpatrick

You've got 2 different challenges.

Peter Diamandis

In today's episode, we're going to be discussing why all companies need to become AI companies in 2026, how they do that, and what happens if they don't. We'll discuss whether big legacy companies can even make such a dramatic change and how they can best do it.

We'll go over some fun and meaningful AI use cases. I think they're going to get you excited about what you can do, and we'll dive into some predictions from our guest for 2026.

Today, joining us is a friend of the pod, Matt Fitzpatrick, who spent more than a decade at McKinsey, rising to the position of global head of QuantumBlack Labs. I love that name, QuantumBlack. It's so cool. He led the firm's AI software development R&D and global AI products.

A year ago, Matt joined as the CEO of Invisible Technologies, a company started by a brilliant friend of mine, Francis Pedraza. For those of you who don't know Invisible, the company is a modular AI software platform that uses AI training and provides AI training for most of the large language-model providers out there. It builds custom workflows and agents for enterprises. The company anchors its work in creating clean data and human-in-the-loop delivery to ensure measurable business results.

Matt, welcome. Good to have you here. Hey, Matt.

Matthew Fitzpatrick

Thank you for having me. We've got DB, AWG, and Salim. Almost happy holidays, guys. It feels like we're on this pod every other day. I think we should just move into a large podcast house, and we will have fully documented the singularity.

Dave Blundin

I'm really looking forward to hearing from Matt because on Thursday we have to do our predictions for next year, and Matt is going to give us a ton of insight today. One of my predictions out of the gate is that enterprises are going to move super stupidly slowly compared to AI capabilities. Matt is the world-leading expert on the intersection between AI and enterprise, so I cannot wait for this.

Peter Diamandis

You cannot cheat this way, Dave, and use Matt's predictions as yours.

Dave Blundin

I can't?

Peter Diamandis

No.

Dave Blundin

I'll be there. But everybody listening to this pod will know that.

Peter Diamandis

All right. Take good notes, if nothing else. Are you guys ready to jump in?

Matt, I'm going to kick it off with a broad question. Within the past year, we've heard from every company out there and every CEO that we're going to be pivoting to become an AI company. Salim, in the last pod, you said something like, "We're going to see the largest disruption ever in 2026 from companies that don't make this change." And I think, Alex, the term you used is that they're going to be cooked if they don't.

Dr. Alexander Wissner-Gross

I said knowledge work is cooked. Not knowledge workers, not companies. Knowledge work as we currently know it.

Peter Diamandis

So you don't think that companies are going to be cooked if they don't make the transition to AI?

Salim Ismail

I think we're going to see many more companies over time, and many more smaller companies as well.

Peter Diamandis

We're going to dive into that.

In an earlier episode, we pointed out that when you thought you were at product-market fit and scaling a SaaS company, you're toast because everything needs to be rethought now given AI. This is now applying to big companies as well.

Matt, the question to kick this off is: Can every company truly become an AI company, and how? Which companies and industries do you think need to disrupt themselves now before they become basically irrelevant? It's a softball question to kick us all off here.

Matthew Fitzpatrick

Peter, save the hardballs for me, Matt.

Always. I think your second question relates to your first question in some ways. I don't think, based on all the data that has come out on this so far, that all industries are going to be impacted equally by this.

There are some sectors and areas where you're going to see materially different impacts. I think areas like media, legal services, and business process outsourcing are sectors where the structure of what the industry does is going to change.

If you think about knowledge work—the production of large amounts of documentation—these technologies are very disruptive. Where I think the hype has been a bit overblown is in sectors like oil and gas or real estate. The function of what they do is going to stay pretty consistent.

I think most of the good analytics of how job dynamics will change over the next couple of years will get at this. The decision about which apartment building or which office building to buy is going to function pretty similarly to what it did 5 or 6 years ago.

I think the question is which parts of your business can really change with AI. It's not all of them, and some sectors will be more or less affected.

The second part of your question—can everyone actually become an AI company?—is also an interesting one. There aren't that many people who know how to build these sorts of models or deploy these sorts of models well.

One of the big challenges is whether you have the expertise in-house to do this. How do you think about adjusting the operating functions of your company to do it? Is it the same team you have in your IT function doing it now?

Particularly, Peter, I know the set of folks that you and I have spoken with in the past—small businesses. If you're a 50-person company, it's hard to deploy a lot of this stuff at scale if you don't even have a CTO in-house.

I think there's a mix of whether your industry is going to fundamentally change and what the actual core competencies your company has to implement it are. I think what you're going to end up finding is—

Peter Diamandis

So, do you end up bringing a chief AI officer into your company? Are you going to bring that capability in, or are you basically renting it?

Part of the other thing that's going on right now, which we've talked about on the pod a lot, is that your competition isn't really the large multinational. It's the AI-native startup that came out of nowhere and reinvented itself from the ground up as an AI-first company, right?

Like, down and dirty, Matt: Which happens first? I can get a mortgage by talking to an AI and get it done in under an hour, or we're walking on Mars with our own 2 feet? Which of those 2 things is going to happen in the real world first?

Matthew Fitzpatrick

The way I've heard the question asked is: Do the startups get distribution before the big companies build the technology? I do think that will be the tension in a lot of ways.

I think there's a lot of big, established companies that are going to figure out how to do this really well. If you take a sector like legal services, I do think the big law firms will figure out how to use a lot of this over time.

I think there are sectors where banking is a really interesting one to look at right now. If you look at the age of the application footprints in banking, most of the tech that exists in banking is north of 20 years old.

You do have a bunch of very fast-moving newer fintechs that are approaching it in different ways, companies like Revolut. I don't know how that plays out, but I do think that becomes the question in a lot of ways: Which moves faster, the emerging entrants or the modernization of the existing companies?

Peter, to hit on what you were asking as a second part of that: Do you buy or rent? I think that's something you've got to be really honest with yourself about as a company.

The idea that everyone can buy, or everyone can hire people to do this, is challenging. The challenge of trying to adapt an existing IT function to do this is that many of the skill sets people hire for—even things like knowing Python—have gaps in them.

I think the answer that most companies I've seen who don't have the resources in-house are coming to, through a very direct push, is that they're finding ways to rent or buy this externally and partner with folks that can allow them to do it.

Peter Diamandis

I think legal and accounting are really cool case studies, and I know you know more about this—your McKinsey time and QuantumBlack. You're the guy understanding and parsing all of this. But they're really cool because they can be replaced by a startup, like Harvey.

Matthew Fitzpatrick

Dave, two things I'd say about that. I think one challenge of implementing GenAI in the enterprise setting is having a statistically validatable baseline to compare against.

As an example, if you take something like mortgage underwriting, it has made huge progress in a very positive way. The percentage of mortgage underwriting that's now done by a very guardrailed and very effective set of algorithms developed by the banks is pretty high because they can back-test and say, "This is a correct credit decision that has no redlining or anything else."

But if you think about a domain like this, the reason contact centers have been one of the use cases where we've seen a lot of adoption is that you do have a clear baseline. You have time per call, CSAT, cost per call—you have a set of metrics you can compare against.

Something like, "Let me generate an investment memo," which is different in format at every firm—it could be 10 pages versus 40 pages, and the content is different—has made it harder for folks to build baselines. I do think that's why legal services are interesting: there are certain areas of legal where those baselines are clear. You can look at what documents are really good for, like an ISDA agreement.

Where I think you're going to see this in a lot of different segments is that the high end of that market still persists in a really differentiated way. If you're doing a large M&A transaction, you're still going to want a really good lawyer's advice.

Where it changes, I think, is in the more basic, "Produce an NDA" type of work. I think that's going to be one of the shifts again: really good human guidance is going to persist forever. It's the basic commodity information that right now a lot of people are paid probably excessive amounts of money to produce.

Dave Blundin

Yeah, well, the NDA is pretty extreme. But I'll tell you, the venture fundings that we do—we do tons of these every year—the term sheets always say that the company we're investing in will bear the cost of the legal work, capped at $50,000. The documents are freaking identical every single time. There are about 8 knobs, and you could store all the combinations on the smallest thumb drive in the world.

I'm like, how is this $50,000? It always runs up to $49,999.99. It's like, wow, what a miracle.

That to me feels like it would be on the mid-to-hard end of the scale, yet it's still so doable. An NDA is a no-brainer. Mortgages are no-brainers.

Matthew Fitzpatrick

I completely agree. I think what's been interesting, though, is how slow the actual adoption curve has been in contact centers. Contact centers should have had generally measurable CSAT scores, and people don't really like most contact-center interactions. The general customer feedback you get is pretty unhappy, and that's been true for a decade.

Peter Diamandis

Yeah, I guess technology would have—let's talk about the whole Klarna thing, actually. I know you're an expert on this. The Klarna thing has been really interesting to watch. Wait, tell us the story. What is the Klarna thing?

Matthew Fitzpatrick

Well, I was not involved in Klarna, but I can say at least what I know from reading about it and what my hypothesis would be.

Basically, Klarna announced that they were going to move entirely toward a fully end-to-end agentic contact center. By the way, the interesting thing was that at that time, they were the most frequently cited example of agentic success in deployments. Then, about 8 to 12 months later, they announced they were rolling the whole thing back and moving entirely back to human contact-center agents.

I found the entire evolution interesting because, if you think about how these systems should be defined and deployed, a multi-agent system should have an orchestration of the types of calls. You'd have a set of validations on which calls could go well or badly, and you'd have some sense of where you need escalations to human agents.

You would never want to move to doing everything agentically. This is a theme in this whole area: you're never going to want to do everything agentically. You're going to want humans in the loop in almost every industry and on almost any topic.

If these models are trained on precedent data, you can train them really well to continue that logic. But you're going to want humans for things where you don't really have precedent data, or where you need them to work through complex situations for which you don't have enough historical information.

I found the entire structure of how the change happened quite confusing because you would always want to keep a contact center as a mix of humans and agents, and then evolve the mix based on the topics. The whole movement from all humans to all agents and back to all humans was confusing, I think.

Peter Diamandis

Salim, you are presented with a question.

Salim Ismail

I just wanted to give out some details here. The Klarna situation was that they rolled out an AI to handle customer-service calls, and the claim was that in the first month it did the work of 700 full-time agents, handled 2.3 million calls a month, and was projected to save them $40 million a year.

They were proudly saying this was month 1, and it was only ever going to get better from there. When I saw that, I thought, "Okay, if I were doing that, this sounds like a PR exercise more than anything real, because you'd never put that out in the first month. You'd wait a couple of months to see what exactly happened."

Matt, you may be able to give a little more color on why they rolled it back in the end. Did they find the hard cases were too many? Was the exception handling too much? Or was it a cultural backlash? What was it exactly that had them undo the whole thing?

Matthew Fitzpatrick

I don't know, in the sense that I haven't worked with Klarna. But you hear a variety of different pieces of feedback on why folks have struggled in contact centers.

One reason is that there are cases in which humans just want to talk to another human. So, some of the PR around saying, "We're moving to only agents," has its challenges.

Two, a lot of the challenges—and where contact centers are most sensitive—is with non-first-line call-resolution topics. It's not something like, "Check your balance." It might be something like, "Process a refund," which is pretty complex. You have to write back to the source systems.

It was surprising to me how quickly they rolled that out, and I wonder how well it was able to deal with some of the more complex functionality in that example.

Peter Diamandis

Right. You go from level 1 to levels 2 and 3 very quickly on those support calls, and then you do not want an AI dealing with you.

Can we get back to the main question here? 2026 is coming up. If you're listening to this in 2026, it's here now. Here's the question: you're a medium-sized or large-sized company, and your board of directors has just said to the CEO or CTO, "Guys, what's your AI plan? What are you doing?"

We're seeing that over and over again. What is their first reaction typically, and what should they do? I want to get some of the fundamentals here because I want to serve our listener base in that fashion.

Matthew Fitzpatrick

If you're that CEO, you've got 2 different challenges. One is, what are the things I should focus on? And two is, who should do them? Do I have those skills in-house?

The first thing I'd start with is making sure you answer the first question. I do think this is a question of following the value. I'd go down a list. I would not start with letting 1,000 flowers bloom. I would start by identifying 2 or 3 things that, if you do them well, materially move the needle for your business.

Maybe it's customer service. Maybe it's forecasting in your FP&A function. Maybe it's inventory management. There are definitely 2 or 3 things that almost any business on Earth, even a small company, has. Digital marketing is probably another one that you see pretty frequently.

You focus on 1 or 2 of those and make sure you get to a pilot stage in those areas.

Meaning, not a strategy document. I do think the one thing that anyone who's spent real time in this space will tell you is, if you take the paradigm of how machine learning is deployed—where you spend months and months building something, and then it works, and you can underwrite statistically that it works—this is kind of the exact opposite paradigm. You can get a prototype up and running in a month, but you have to do a lot of testing and validation to make sure you can trust it. And so, it is really a function of making sure you can get something up and running, and testing and validating.

You know, Peter, the question I always ask is: would you bet your annual bonus that whatever use case you deploy works? And that's a complicated thing. If it's, let's say, generating a claims-processing review, and you have to do 10,000 of them, most companies don't know how to say whether that works or it doesn't. And so, just to summarize, make sure you have a list of the 2 or 3 things that move the needle. Make sure you get to a proof of concept in one of them. And I would probably do that first use case as an RFP to a third-party vendor that gets compensated based on results.

Peter Diamandis

Yeah. And I say that very specifically because I think if you do it in-house, the odds are the in-house team has not had a lot of experience with this. And so, you also can't hold them accountable in the same way: you get paid if it works. And so, I do think tying it to outcomes limits your risk. I mean, that's still the business model for Invisible, right? You're paid by money saved.

Matthew Fitzpatrick

Correct. We do outcomes. Yeah, outcomes in various ways.

Peter Diamandis

Yeah. Alex, I want to bring you into the game here. Much appreciated. So, maybe just as a preliminary matter, for full disclosure, I have no financial interest in Matt's company, Invisible. I do have a number of questions, though. First question, maybe pulling the thread on testing. One of the things that we talk about here on the pod all the time is benchmarks, the importance of benchmarking. I'm curious, given that—

Dave Blundin

We talk about that constantly, Alex. That is all we talk about. We talk about nothing else. That is all we talk about. Oh, wait. Maybe that's you. Okay.

Dr. Alexander Wissner-Gross

Given that's all we talk about, as Dave just mentioned, and given that Invisible is also in the business of training so many models, what benchmarks do you think most need to be brought into existence in the world? What's most missing? What are the top 3 benchmarks you'd like to see summoned into existence?

Matthew Fitzpatrick

Yeah, look, I think—and you've seen a bunch of these start to get publicized in the past couple of months—but most of the public focus to date has been on the large public benchmarks for things like coding. And I think those are very useful as metrics for whether the models are improving broadly. That is why you've been able to see, by any standard, if you look at our last 3 years, the models have improved 50% to 100% on most dimensions that you can look at.

I think the problem is, though, if you think about enterprises or small businesses, your benchmark for most cases is not a broad-based, accurate cognitive benchmark. It's accuracy or human equivalence on a specific task. And so, what I think you're going to see more and more need for is custom evals on highly specific topics. If you go back to the contact-center example, the benchmark you'd want to build if you're going to roll this out for a contact center is a series of expert agents that are in your contact center, how they perform, and then how the AI agents perform similarly. The same applies to claims processing. Basically, most businesses are going to have to get comfortable with doing what's called an eval or a custom benchmark for the tasks they're trying to modernize.

Because an 80% accurate, very smart deployment is not enough—there's still too much risk in that rollout framework. And so, I think the way that we think about benchmarking will evolve from broad-based benchmarks to hyper-specific benchmarks.

Peter Diamandis

I freaking love that because I can immediately see 10,000 listeners right now who just found a calling in life based on what you said. All this benchmarking within any of these domains is really, really hard to figure out unless you know the space. Take title insurance. What's the benchmark for successful AI in title insurance? Somebody in that industry listening to this pod right now is going to be like, “You know what? I was an early adopter of AI, and I know this space inside and out. That's my benchmark to own.”

And if you declare yourself the owner of it and then broadcast the benchmark, the evidence so far is you become an instant star. Nobody's grabbing topic ownership in all these areas, and if you just get there first, you become an instant star.

Dr. Alexander Wissner-Gross

I completely agree with that. In this era of post-training as a commodity, if you own the benchmark, often the benchmark is the hard part, and you can leverage existing resources to post-train an off-the-shelf model. I am curious, though, Matt, maybe following up on this. It sounds like your position is that we need thousands of new narrow benchmarks to capture maybe every labor category and every industry vertical. Assuming that's correct, is that something that Invisible is working on, can be working on, or should be working on?

Matthew Fitzpatrick

Yeah, we do spend quite a bit of time working on that. In fact, a lot of the time, what we're building is customer-specific benchmarks for an individual task. That is a lot of what we think about: how to test equivalence for a given task.

And I think one of the things that folks have not fully realized is, let's say you take a really high-performing LLM and you want to tailor it to your individual context. That process of actually fine-tuning it off of your data.

I think one of the challenges is that people were hoping this would be a SaaS buyers' paradigm, meaning I could just buy something off the shelf that would solve everything I needed. So, I wanted to buy a sales agent; I wouldn't have to do anything. I could just take in a sales agent that would sell well. And the reality is that's pretty hard to do. You need to actually train it up on your specific knowledge corpus and your information.

The way we would think about it is: you take the LLM, or you take an agent that's been trained for sales, and then you fine-tune it off of your specific company information, your products, the way you sell, and your way of speaking. Then you have to build an eval or a benchmark against that to say whether this is performing well or not on that task.

Dr. Alexander Wissner-Gross

Well, a quick follow-up question, if I may, because there was the sort of infamous BloombergGPT moment, where Bloomberg was sort of in quasi-competition with the frontier labs. They had a wide variety of internal proprietary data sets. Their original plan—this is now sort of an infamous episode from 1 to 2 years ago—was to offer their own proprietary frontier model, basically, but, critically, pre-trained and/or post-trained off of their internal data sets.

The plan was to achieve superb performance in the financial domain because they had all the data, or a lot of data, that were not broadly available to the general public. But what actually happened is the generalist models offered by the frontier labs, which were trained basically off of the internet and more or less publicly available data sets, within a few months leapfrogged BloombergGPT.

And so, I guess the moral of that parable, in my mind, is: how far do you think we can really get with proprietary data sets and proprietary benchmarks before the generalist models completely wipe the floor with them?

Matthew Fitzpatrick

Sorry, to clarify, I'm saying you use an LLM. The process I'm describing of actually fine-tuning a large language model for your specific context is basically adding more context. Most of the LLMs offer a paradigm where you can do this, where you can add your knowledge corpus and train it to be more specific to your individual context.

I don't think you'll see individual institutions building their own LLMs. I think that's a very compute-intensive, very difficult thing to do. I think you'll see them tailoring the large language models to their context.

Dr. Alexander Wissner-Gross

Sure. To be clear, I wasn't asking whether you think every institution is going to get into the business of pre-training its models. I was rather asking whether you think post-training—which is inclusive of supervised fine-tuning, reinforcement fine-tuning, and a variety of other post-training—has a long-term future.

Or will, maybe in 1 to 2 years, we just use a pre-trained plus post-trained generalist model off the shelf and not need any internal benchmarks or any internal data sets for post-training?

Matthew Fitzpatrick

Well, I think there are clearly going to be use cases where you are going to need the context of the individual company, right? If you take the law-firm example, there are documents that a company has on how they want their future-state documents for M&A agreements to look, right? And the LLMs are not going to have that information.

So, at some point, you are going to have to see the post-training layer happening at the enterprise. And what we're seeing more and more is there are ways to design that layer so that, as new models evolve, you can drop those in. We are seeing more and more folks experiment with that.

Speaker 1

So, they’re using all the new technology that’s being rolled out.

Dave Blundin

I think, in fact, what’s going to happen is that, over time, that edge in data is going to be the most valuable part of any company. Is that trade-secret type of “How do we do things?” At some point, it may leak into the public models.

Speaker 2

Like, if you used OpenAI, right? If you use any of the frontier models connected to them—I remember we were talking to Replika, et cetera—people are using it, and then the data is going straight into the cloud, right? That’s kind of dangerous. They’re going to have to solve that layer in a very powerful way.

That’s one of my predictions to forecast, et cetera: we’re going to need to see a layer of protection between company data and the broader AI world.

Peter Diamandis

Matt, I want to make this a little more tangible. I know you can’t talk about the work you’ve done with the hyperscalers, but you’ve identified, I think, 5 or 6 cases where you can speak publicly about it. If you don’t mind, maybe we can toss a few of those in and talk about them as concrete examples.

Since Alex made his no-financial-involvement statement, I will say I’m a proud advisor and am conflicted, in a positive fashion, supporting what Matt and Francis are doing. Do you want to pick one of those? I loved the example on the basketball court. Can you speak to that one?

Matthew Fitzpatrick

Yeah, sure. We worked with the Charlotte Hornets on fine-tuning custom computer-vision models for draft preparation. In their case, they wanted to look at the spatial movement patterns of players on a very broad scale, across single-point cameras at a whole host of different college universities and international locations.

We fine-tuned a custom computer-vision model to specifically look at the movement patterns they were interested in before the draft. That was a big part of their draft evaluation.

Peter Diamandis

In English, you basically took the video and were able to use models to evaluate every player based on the video, to see how well they performed in every different—

I’m not a sports guy, so it’s the—

Speaker 2

Yeah, that’s becoming clear here, actually.

Speaker 1

Yeah, yeah. It says the salt economy. [Laughter]

Matthew Fitzpatrick

Sure. If you take typical NBA stats, there are things like points, rebounds, and what’s called plus-minus, which is one ratio that’s often used. It’s the amount you score versus the amount you give up when you’re in the game.

But those are mostly transactional stats. What they don’t look at is the movement patterns of the players: who creates space and where people are positioned at any point in time. That’s actually a lot of the most interesting data.

If you go back to some of the original baseball analytics that Billy Beane did for the A’s, it’s the movement patterns of players and who is in the best spacing, right? There are companies that do this in very consistent formats, such as on the same court. What we’ve been able to do is perform that analysis over many different camera angles and many different stadiums, very quickly.

That uses custom computer-vision models. We’re effectively able to take a single-point camera and understand the movement patterns of players in many different environments.

Peter Diamandis

And how do the Hornets use this? For team selection? Player selection?

Matthew Fitzpatrick

Draft selection—to understand which players fit the characteristics they were looking for.

Peter Diamandis

Fascinating. It’s a complicated problem, too, because chemistry between players matters. It’s not just about finding the best player; the chemistry between players matters, too. It gets infinitely complex, and it’s a cool little case study.

Gavin Baker was saying recently that, in fantasy football leagues all over the country—which I used to love before I ran out of time—

Now you have an agent doing it for you and having fun.

Dave Blundin

That’s exactly the point.

Salim Ismail

We’re now obsoleting human sports leagues, replacing them with robot sports leagues and esports.

Peter Diamandis

Yes. Very 21st century, not 20th century.

Dave Blundin

That’s right. Great fantasy-football players are losing all over the place because the AI agent is tracking a huge amount of more detailed data. If you look at the video footage, somebody might be making it up and down the court very slowly. Nobody’s going to notice that, but the AI will notice it in a heartbeat. Then that just goes into the great model. It’s really a cool little case study.

Matthew Fitzpatrick

Since you asked a little bit about how a traditional business might be thinking about doing this, I’ll give a slightly different example, which is Lifespan MD. Peter, I think this one will resonate with you in particular. It’s a concierge—

Peter Diamandis

I know Chris, who runs it.

Matthew Fitzpatrick

Yeah. Lifespan MD is a concierge-medicine business. You can think of it as having a network of practices both internationally and in the United States, all of which have very different sets of data on their patients and practice information.

The thing I always start with in any AI use case is that you have to get the data right. Before you can even start with AI, you have to make sure that you have the structured and unstructured data together that you want.

The first thing we’re doing for them on our data platform, Neuron, is creating a HIPAA-compliant, multitenant cloud instance where we bring together all the patient and provider data that’s of interest. We start to bring a 360-degree view of both the patient and the practice.

You can start to think of things like which longevity-focused tests male patients between 35 and 50 are using most frequently. You can start to think about patient outcomes that are really interesting. If you want to understand practice performance, or where you have certain patients who are not compliant or not as interactive, it’s effectively a control tower to understand everything that’s going on across that footprint of practices.

Then I think the area where generative AI has become more important for that is chat agents, where people can ask questions—knowledge-management systems that allow them to interrogate and ask questions of all the key data from all of those practices.

One of the key challenges is that, obviously, in health care, you have to be extremely careful about which data is stored locally at the practice versus how it’s brought centrally. The HIPAA-compliant, multitenant cloud is one of the key components of that. It makes sure that no patient data leaves the premises of the individual practices, while doctors can access certain things and certain practice metrics are organized centrally.

Peter Diamandis

I heard the coolest thing this week. It’s a quality-assurance company that has invented “Talk to Your Defect.” It’s just the coolest concept. The defect actually has a personality, and you can ask it questions about itself, like, “Where did you originate?”

I can totally imagine what you just said in health care being “Talk to Your Illness.” You have a conversation with it: “Where did you come from? How do I treat you? Are you getting better or worse if I do this thing?” It’s talking back to you with a personality. It’s the coolest idea ever, isn’t it?

Salim Ismail

I think it’s amazing. It’s one thing with the defect. It’s a little awkward when you say, “Here’s the bacteria you’re talking to.”

Peter Diamandis

Well, I just mean the defect is real. “Talk to Your Illness” maybe gets a little weird. I don’t know what voice you would give it. A Voldemort voice or something.

Tell me, how do I kill you? How do I dispatch you? [Laughter]

Matthew Fitzpatrick

Well, Dave, one thing I’d note there, too, is that I think there’s a question—and I get asked this often—of how sectors evolve. Peter, you asked earlier how sectors evolve. I think the question of whether the decision-making around individual patient care changes with generative AI is much murkier.

I think the easier place to start, and where it would be very interesting, is that the United States, as an example, spends about $13,000 to $14,000 per patient per capita on health care, compared to $2,500 to $3,000 per capita in, say, Germany or Canada. Something like 30% to 40% of that is administrative cost, and that is not an administrative cost that anyone wants to bear.

This is something where I actually think the idea Lifespan MD is pursuing is not to change the standard of care, but actually to make the physician even more empowered—to take all of the really painful administration and scheduling and make that the part they don’t have to deal with anymore. AI should do a huge amount of damage in those areas.

Peter Diamandis

Exactly. What are you seeing most companies get wrong on their mission to implement AI?

Matthew Fitzpatrick

I think it’s a couple of different things. The first one is a lack of focus on data as the starting point. I think the challenge is that if you tried to build an AI agent on fragmented customer and product data, it’s going to break by definition.

You have to be in a place where the data you’re going to feed into the models is clear and working. That’s been one major challenge.

Peter Diamandis

Do you think most companies—as a whole, in the medium and large size—have clean data? How long does it take a company to get its data into a format and to a level of fidelity that’s useful? Is this a hard lift or an easy lift?

Matthew Fitzpatrick

It depends. If you take the paradigm of, “I’m going to put everything in a data lake and get everything right,” that can take 5 years.

And the reality is that most big companies have spent half a decade trying to get all their major data schemas in order. But I think if you start with the question, “What data do I need for this specific use case?”—let’s take credit underwriting. To do that well, you need a set of data around the credit itself and the market. You probably have 5 to 6 core data variables you need: the core financials of the business, the security of the credit, and all those kinds of core pieces.

But you don’t need every piece of data across the entire commercial bank to be right. You need the core elements for that use case. And so I think companies that are focused on the exact data they need to get right have done pretty well.

But I do think that trying to get all the data right—I mean, you’ve also seen the enterprise for a long time, Peter. If you asked any Fortune 1000 company to look at their full data repository and determine how much of it is accurate, working, clear, and accessible right now, very few companies have that.

So I do think it’s important to be very tactical about what data you need. The other thing I think, for generative AI in particular, is that a lot of the most important data is non-system-of-record, unstructured data. It’s things like images, videos, and text files. Those are just not things that people have tried to master historically. And so I think the first step in this is saying, “What is the thing I’m trying to solve, and how do I make sure I have that data ready?”

Peter Diamandis

Yep. One thing I see a lot of—I had a long board meeting this morning with a company that’s very AI-forward in portfolio accounting, a company called Vestmark. And the data, for account reconciliation, for example, is abundant. But it doesn’t tell you what the person actually does. It just tells you how it was reconciled.

So now the path to success is first the AI assistant, which helps accelerate you through the day, but it also knows what you’re actually doing. Then that accumulates, and then that becomes the RLHF for the training or tuning data. Because what you’re trying to do is, “What are you doing, guys?” And that’s not really represented in the data.

But a lot of times you go talk to a bank or an insurance company, and they’re like, “Our data is our advantage. Go ahead, bomb it into the neural net and train it.” You’re like, “I don’t even know what that means. I’m just going to throw all the terabytes of spreadsheet data in and see what happens? That’s going to go Clippy on you?”

Well, you have all sorts of other issues as well. I was talking to the CIO of one of the biggest banks in the world, and they have 300 different customer databases. Three hundred: one for mortgages, one for loans, one for this. The mortgage people don’t want to tell the loans people about their customer data, so they guard it jealously. It’s a total disaster for the poor CIO. Fascinating. Alex—

Dr. Alexander Wissner-Gross

I think these are all very interesting points. I’d like to, if I may be so bold, jump up several levels and speak a little bit more about the business model of Invisible. My understanding—correct me if I’m wrong, Matt—is that there’s an element of the business, I think it’s called Meridial, that serves as a sort of marketplace for ML freelancers, if I understand correctly.

And I’m curious. I think, in my mind, one of the many elephants in the room in this conversation is that we’re arguably on the edge of recursive self-improvement. All of the frontier labs, more or less, I think would agree with the assertion that we’re nearing the point where you could have an AI researcher, where you just turn over computer resources to that AI researcher, and the AI researcher does as good, if not a better, job than the human AI researchers who work for the frontier labs.

If that is indeed the case, surely one of the several elephants in this room—but given limited time, let’s focus on this one—is that the need for a marketplace of ML freelance researchers to train models evaporates entirely as we start to reach the point where AI researchers can build custom models off of custom datasets and custom benchmarks for each client. Doesn’t it evaporate entirely?

Matthew Fitzpatrick

Yeah. So, as you said, we have 2 sides of our business. One side, Meridial, is where we train all the large language models. Then, on the enterprise side, we build basic custom applications for enterprises.

Look, I think there has been a 5-year evolution where folks have consistently said that, at some point, you will not need reinforcement learning from human feedback to validate and test models. And I think the challenge of that logic is a couple of different things.

One, the spectrum of expertise—if you take language, multimodality, extreme expertise on things like computational biology—and then the fact that a lot of these are reasoning tasks, you do need human feedback on almost every different sort of agent you want to roll out. There’s a whole host of studies on this showing that pairing synthetic and human data together is stronger, but you do need human feedback in some form.

And so I think the nature of RLHF is changing. I think you’re moving more toward things like RL gyms, controlled environments, and simulations. I think you’re starting to see much more of the expert work being done by PhDs and master’s-level researchers. So it’s less what I’d call commodity “cat, dog, cat, dog” labeling.

But if you say tomorrow you’re going to train a model to figure out different evolutions in 17th-century French architecture in French, you are going to want RLHF to do that, to validate it. And I think you’re seeing that over and over: as the models move more and more into very specific areas, there is more and more RLHF needed for them.

Dr. Alexander Wissner-Gross

That’s interesting. Maybe I’ll share my intuition, and then I’d be curious to hear what you’re seeing in your version of the ground truth. My intuition, my impression, is that we’re seeing greater and greater data efficiency.

And pardon me, I mean, RLHF was obviously very fashionable over the past 3 years. Maybe it went through peak fashion, if you will, and then we saw the rise of reinforcement fine-tuning mechanisms that are far more data-efficient and maybe even more human-time-efficient.

If you have to just build an RL environment, arguably that’s, per human hour involved, probably a lot more time-efficient than staffing out to some so-called developing-country folks to, as you say, do “cat, dog, cat, dog” supervised fine-tuning or some other RLHF-type mechanism.

Surely I’m projecting. My intuition is you’d see more data efficiency, not less, and therefore the amount of time, effort, and money expended on RLHF—or any sort of mechanism—even if we buy your assertion that we’re seeing hyper-parochialization of lots of different tasks and each of them is going to need artisanal annotation, surely there is a competing force, which is increasing data efficiency from algorithmic efficiencies like reinforcement fine-tuning. What are you seeing?

Matthew Fitzpatrick

Yeah. People have been arguing that for 5 years, but I think at least what I’ve seen on the ground is that, given the accuracy that you want, if you think about a reasoning task that involves a several-step leap and you think about the risk of hallucinations, it is more useful to have human feedback involved in that in some form, all right?

And so I don’t think that means—if you think about it, in some ways, RLHF happens after all the pretraining compute cost—it’s a pretty small percentage of the total cost in training. And it is some of the most valuable feedback.

As you see more and more specific agents being trained for specific tasks, take legal services as an example. If you get a new legal-services dataset, which is interesting, and you want to train a model off of that, you’re going to want to see some sort of comparable equivalent, whether it’s an associate or an M&A lawyer equivalent, where you actually test if it works.

Now, is it possible that at some point, 10 to 15 years from now, you run out of things to train on? Possibly. But actually, if you take the number of languages and modalities, robotics is probably the next frontier of this in some ways. RL gyms, contact centers—there’s a lot.

We are, as a company, a full believer in—I talked about it on the enterprise side, too—that human-in-the-loop is going to be a feature, not a bug, for a long, long time. And I think the entire red herring of the enterprise, for example, is that autonomous agents will do all of this with no human in the loop. I actually think you’re going to need more and more humans at every step.

Peter Diamandis

Alex, you’re saying that the level of intelligence of these agents, as we pass through AGI and get to ASI, is such that they’ll figure it all out as well as any human and replace that human in the loop. What’s your timing on that?

Dr. Alexander Wissner-Gross

That was exactly my question, Peter. So my timeline, if I had to spitball—of course, this is not the predictions episode, so don’t hold me to it. Hold me to my predictions in the next episode—is approximately 2 to 3 years as a conservative outer bound for some element of recursive self-improvement, where we get an AI researcher that’s as good, if not stronger, than the human researchers for building ML models, as a conservative outer bound.

Now, 10 to 15 years? 2 to 3 years max. That’s the outer, outer edge. But I also believe Matt’s totally right that 2026 is going to be the year of recursive self-improvement, with capabilities growing at crazy exponential rates and corporations moving at a snail’s pace compared to what they could be doing.

And it’s all going to be stuck, bottlenecked, log-jammed, and it’s going to frustrate the hell out of Google and OpenAI.

And companies like Invisible are the lubricant that's going to actually get it from point A to point B. But that Clippy use case is a really good one. In our tests for contact centers, 80% of people massively prefer the AI. But the 20% who don't like it would rather torture the whole thing to death, make it better, or repeal the entire thing. There are probably 8 ways to fix that quickly.

Peter Diamandis

Yeah, but it's not going to come from Google, and it's not going to come from OpenAI. It's going to involve data that isn't in the natural data set. If you told me 2 years ago that everyone in the world would know what RLHF stands for, and that there would be 3 people who are multibillionaires from building RLHF companies walking around saying, “That's not even a thing,” I would have laughed. Oh, wait—now it's not only a thing; it's massive in scale.

There'll be new terminology in 2026 for many of these other bottlenecks: the AI can do it, but for whatever reason, the bank isn't doing it, and the contact center isn't doing it. Those bottlenecks are going to be so lucrative for companies like Invisible to just plow through.

I can't answer the specific question of whether your workforce is going to involve the distributed workforce that you just described. What was it called, Alex? Or Matt? It's called Meridial. Meridial. Yeah, so there is a really healthy debate over whether Meridial is a key part of this, whether a network of even more agents is a key part of this, or whether 2026 is the transition year between the 2. It's going to be a really interesting footrace between those 2 different approaches.

But I think that's it. Dave, I think you put your finger on it. That really is what I'm asking, which I think is a distinct question from whether there's value in supervised fine-tuning or reinforcement learning with human feedback going forward. Of course there is. What I'm really asking is how much of that can come from AI bootstrapping it in the near future versus needing human inputs.

What I'm saying is, think about a balance between generalizability and hyper-specificity. I agree with you on generalizability. I don't actually think RLHF is important even now for that. But where it gets more complicated is when you want to train off specific tasks. So let's take the insurance-claim example that I mentioned earlier. You're going to generate a 10-page insurance claim, and you could apply this to any enterprise use case and many consumer use cases.

In that world, an LLM is producing an outcome and is fine-tuned off a specific company's data, but you need a way to actually say, at that point, does this produce an output comparable to what a human doing this task before was doing? When I mentioned custom benchmarks earlier, that's the process by which you do that. You actually do need human-equivalence testing. You need a human to provide a comparable data set and say, “This looks good,” or, “It doesn't.”

You just don't have precedent data to train that off of in any of these LLMs because the human input is not there. Now, again, that's going to keep going down to more and more specific tasks. If you take legal services, take it by language, take it by topic, take it by document type, there's human feedback required for all of that.

I don't want to put too fine a point on it, but I want to make sure that those in this episode who want to drink the bitter pill with the bitter glass of water for The Bitter Lesson are drinking it. I'm curious, Matt, to understand how you see this. Surely there's a wave of generalism that, over time, maybe we can finesse what the appropriate time scale is. It sounds like maybe your timelines are a little bit longer than mine, but would you at least agree with the premise that, over time, even specialized skills end up getting subsumed by generalist models? Or do you think that's just never going to happen? Will we always—or by “always,” I mean on time scales of 10 to 15 years, which is a pretty long time scale—just have generalist models that are always specially fine-tuned?

Matthew Fitzpatrick

I don't think all expertise—all specialized expertise—is going to go away. Again, if you think about a lot of the information that specific experts have, there's no training data available for that. It's stuff that sits in people's heads; it's experience.

I'm aware of many of the narratives that human expertise becomes less important. Again, we're a company that actually thinks the human-touch elements become more and more important. Take sales, for example. Many of the best-selling patterns, and many of the people who have done that best—there is no information you can train on from what they do. They live in human interaction.

In a world where there are 500 companies selling email-based SDRs, I think human beings become more important in that world. I don't actually think specialization goes away. I think the shift is that expertise becomes more and more important in many different areas. I think the human loop stays really important.

But if you take a contact center—and Alex, I understand the theory of what you're saying—but we're 4 to 5 years into this, and if you look at the number of US contact centers that have migrated to using agents, it's a pretty small percentage.

Can I ask you the Jane Street question? It's really burning a hole in my pocket, too. It's really clear that stock picking is moving to AI at warp speed.

Dr. Alexander Wissner-Gross

And the reason is that there are no barriers. You're just placing a trade that's already automated, so—

Peter Diamandis

And that's the bellwether to me. It's a great benchmark. More money. More money.

Dave Blundin

Yeah, and also almost all of the volume on public-equities markets has long since been dominated by algos. So this happened decades ago. It started with rapid trading, so the quants were already there. Now that it's moving to fundamental analysis, it's the same mindset. That's one of the reasons it's taking off.

Like Peter said, you're making more money. Okay, let's just keep going. There's nobody who's saying, “But I'm going to lose my job.” It's like, “No, we'll just pay you more. Let's just go.” So it's a really interesting bellwether. But within that world, they're struggling because the data is so proprietary. Mhm.

Dr. Alexander Wissner-Gross

And it's looking more and more likely that these self-improving, massive foundation models are going to get to superhuman IQ this year—this year being 2026. The context window is getting massive, and the recursive chain-of-thought reasoning is getting really good. So you can actually feed it data without having to retrain it and have it achieve the job.

If I take that mindset from Jane Street and move it over—now I'm a mechanic and I'm trying to fix a car, diagnose what's wrong with it, and I have audio and sensor data—great, easy use case. But am I going to put that data into the LLM API and transmit it to OpenAI, where they can accumulate it, and then, if they decide later they want to be a garage, they can have all my data? Or am I going to run some kind of walled-off model?

A garage mechanic's maybe not the best example. That's why I chose Jane Street, because they're never going to take their proprietary data and give it to OpenAI. But in the middle ground, you have banks, insurance companies, hospitals. How are they going to deal with this? It's easy now. Sometime in 2026, it becomes easy, but the data is proprietary. That's my only reason for having a competitive advantage. I don't want to give it over to the API.

Matthew Fitzpatrick

Yeah, look, I think you're seeing that there are definitely sectors, many of which you just named—banking and health care—where people are deciding to keep their data on-premise, or they're using things like small language models for those sorts of reasons. I think you may continue to see that as a trend.

I think one mistake folks often make is that not all data is proprietary. Take the Jane Street case: maybe their trading data is proprietary, but their back-office forecasting data might not be, and their back-office finance data might not be. I think one thing is being clear about the data that you need to keep proprietary and around which you do want to take more security measures, and then what data you say, “Look, I'm going to be very careful as a company, but this is data that isn't as proprietary.”

I think that sort of balance is similar to what we discussed with contact centers. The idea of “I will not give anything to the LLM, but I'll keep it all in-house” doesn't make sense either. But I do think that's a paradigm you're seeing more and more.

Salim Ismail

Yeah. So I want to change tack a bit, if that's okay. I actually do agree that we'll automate, but I think we'll automate in a way that's different from this discussion. Let me give an example.

Let's say I'm Canon and I'm selling home printers. Right now, I have a bunch of people doing marketing, content development, and brand management; salespeople to sell to Best Buy and so on; online folks; post-purchase staff getting the customer to try to register the dang printer; repair-support and technical staff; and accounting folks in the company.

You could get all your job functions managed by AI, right? So you've got pockets of people doing different functions across the board.

If I was going to build an AI-native printer sales company, then I might think about having all of those things automated completely with AI. Then you're not human-centric, but function-centric across those. The printer could report when it's running out of ink, and you ship it a new thing. It tells you when there's a problem with it or a problem coming up, and you alert your repair staff, saying, “Hey, this guy, maybe we can upsell him a printer.”

You essentially automate all the functionality with AI, and you leave the human 90% out of the loop almost completely because you've automated the core functionality. What I'm seeing right now is what I used to call “radio over TV.” When you first had television, we took radio announcers and put them on TV to read radio scripts. We didn't adapt for the medium.

I think what I'm seeing right now is we're automating what human beings are doing at each of those functions, but surely, over time, we're going to automate the functional flow and then get rid of the human beings completely. AI-native, AI-first, right? Not to mention getting rid of the printers. Well, that's a separate question. I'm just using that example.

Peter Diamandis

Who's going to be doing any of the printing? Let's leave that part aside just for the moment. I think you're absolutely right, Salim. This is where a young AI-native company reimagines an entire field and has zero legacy and zero friction in coming forward. The question, as Matt said at the beginning, is: Do they have the distribution?

This is where a large company—Canon, in this case—should actually be investing in entrepreneurs. One of the things you and I talk about a lot of times is, if I'm a large company and I don't know what to do, I would basically hold a competition and ask young AI entrepreneurs around the world to come forward: How would you disrupt my company? Give me a pitch. Then I would pick the best 5 of them and fund them.

I would say, “We're going to fund you to disrupt us, and then we're going to give you access to our data, to everything we have. Ultimately, we're going to buy you or buy a majority stake in you, and we're going to make you our new company.” This is the innovation on the edge, the displacement of the core, et cetera—whatever you want to call it.

You're a medium-sized or large-sized company. I'm not going to focus on the startup right now. What do you do in 2026? You're going to have to do something. You're going to have pressure from your board, from your shareholders, from Alex.

Dr. Alexander Wissner-Gross

From just competition.

Peter Diamandis

So you've got to do something. What I heard you say so far, Matt, is: Number 1, you've got to get clean data. You need to make sure you understand what your data situation is. Number 2, you should pick 2 or 3 areas—call them benchmarks—where you're going to run experiments. It's not a proposal or an idea. It's actually: Run it. Actually do it—run an experiment to see how it works.

Then pour money on the things that do work, and have an expanding, increasing circumference around the company's major revenue engines. How do you think about that? Walk us through a few more steps.

Matthew Fitzpatrick

One of the things that has been a topic of conversation here is, given all the improvements in the models and what Salim was walking through about the potential to clean-sheet and design a company from scratch, why has that been so much harder? There was this MIT report that came out saying that 5% of enterprise models right now make it to production, right?

I think there's a starting question: Given all this tech excitement, why has that been so much harder? It's not the technical challenges that we talked about. It's the data and the focus on which priorities to look at.

I think the other 2 big ones are the organizational structure through which you pursue those initiatives. Particularly, the advice I give everyone is: Do not locate this in your technology organization. Take your best operator—your best ops person—give them an operational KPI, and track it to that. Make sure it's a really clear operational KPI.

We talked a bunch about contact centers. You should have an operational person there lead it around CSAT score, time per call, or whatever the core metrics are that you're looking at. That should be your guide. If you want to take something like inventory forecasting, you should do it around inventory days, stockouts, and all those kinds of metrics.

If you have a clear sense of which operational person is leading it, how they're marshalling resources around it, and you have a clear KPI, you're going to make progress if you focus on a couple of different things. I think the failure mode has been that you let a thousand flowers bloom, none of them have an operational metric, and you end up with a science-project dynamic.

Peter Diamandis

Yes, exactly. That's exactly right. If you walk in, a thousand flowers bloom. You walk in and say, “I am going to give you a million genius-level people for free. Do something.” It fails.

It's like, “Here's a million people for free, and they're all geniuses.” It fails for that same reason. It's like, “I didn't think of an idea, so I said, ‘A thousand flowers, just go bloom.’ I couldn't think of anything, so maybe you will.” How's that going to work? I've seen that. You're exactly right. It's just so sad.

Salim Ismail

We go even further. We basically say: Not just take the operator and put them outside the organization and let them build something from scratch at the edge. Otherwise, you're getting encumbered by all the internal rules and bureaucracies, and that gets slowed down a huge amount. Then it fails for legacy reasons.

Peter Diamandis

Yeah, it's not the company skunk works; it's the Apple Macintosh team. Apple is actually a master at this. What Apple will do is form a small team that's very disruptive. They'll put them at the edge of the company, keep them secret and stealth, and say to them, “Go disrupt another industry.” Whether it's watches, retail, or whatever.

At last count, I think they have 18 teams looking at different industries to think about. When they think it's ready to disrupt, they go into it and patiently iterate. The Apple Watch, for example.

This is the model I think we're going to see many other companies take on, where you do this. If you think of any operational company, the insights they have on all sorts of adjacent industries are incredible. It's very hard to disrupt their own industry because they're probably pretty optimized for it, unless you come with the AI startup, but they can really disrupt a lot of the edge cases and a lot of the industries around them. So I expect them to launch AI-native startups that go into adjacent industries and attack some of their neighbors.

Matthew Fitzpatrick

Nice. We worked with SAIC Vantour and the U.S. Navy on building intelligence for an underwater drone swarm of unmanned underwater vehicles. Think of it this way: If you have a series of drones and enormous numbers of sensors on each of those drones, you need to understand the movement patterns of those different drones.

In each case, you see an object underwater. What do you do? Do you engage? Do you step back? Do you move with other drones? That whole movement pattern and decisioning for underwater unmanned vehicles is what we worked on: fine-tuning a model to do that, training it, and looking at all the movement-pattern data.

Again, this is one of those interesting things about drones: They are autonomous, and so thinking about how those movement patterns evolve in complex environments is very, very tricky to do. But you also have lots and lots of interesting sensor data to do that.

I think one that anchors more on the human decisioning side is SwissGear—like Swiss Army, the luggage brand. Similarly, I actually think this is, Peter, one that a lot of folks in the audience may relate to in some form. They had an enormous mix of different data tables around products, customers, et cetera, that they couldn't really bring together for inventory forecasting.

We used our data platform, Neuron, to bring together 750 tables really quickly and then optimize the forecasting to look at both minimizing stockouts and optimizing which inventory to hold. If you get inventory forecasting right, it's probably one of the major issues for most big and small businesses: You minimize lost revenue, and you make sure that you don't hold lots of excess inventory.

It's one of the hardest things to do, particularly if you have a 6- to 8-month order cycle time. And so that was something we partnered with them on, and I think it was a great outcome. We ended up expanding their overall inventory coverage by about 30% and basically 2x-ing the number of SKUs with a reliable prediction. And again, that was done in a couple of months.

Peter Diamandis

All right. So later this week, my Moonshots mates and I are recording our 2026 predictions. We'll have Emad back, and we'll be talking each of us will provide two predictions for 2026. We'll have our top 10 from the Moonshots podcast. It's going to be fun. Uh it's going to be a battle. Uh we're going to ask our listeners to vote on which predictions they like best. I mean, of course they're all going to vote for Alex's, but hey. Uh Matt, uh talk to us about what you see coming in 2026.

Matthew Fitzpatrick

Yeah, I think I'll call out a couple, and we've just done a bunch of research on our 2026 predictions. So I won't say all of them, but I'll call out a couple.

I think one of the first ones I would anchor on is multi-agent teams. I think one of the challenges—and it's inherent in a lot of what we've discussed here—is that if you're a large enterprise or medium-sized company implementing a use case, you won't necessarily have one decisioning agent that does everything. You'll train task-specific agents for individual tasks, usually orchestrated by an LLM. What that allows you to do is pinpoint the accuracy on those specific tasks, and then use the broader logic set of the LLM to make sure they all work together properly.

I think that's been an architecture that's been discussed pretty broadly for a while, but I think we're just starting to see the green shoots of more and more folks having success with that. Contact centers are a good example, so I think that's a big one that I would call out.

I think the second one I'll call out is the multimodal leap. More and more, video, images, and audio are going to become a bigger and bigger part of how people engage with these models. Audio is probably one of the most interesting. I do think the way you'll be able to speak to them, interact with them, and visualize them is going to be a really interesting moment for 2026. And I don't think that will all be text-based like it has been historically.

Peter Diamandis

Mhm. And then maybe one other thing to talk about. No, I was going to ask Alex for feedback. Go ahead, but finish up, Matt.

Matthew Fitzpatrick

Yeah, so the third one I'll call out, because we've talked about it a couple of times on this episode, is what we call either the mirror world or RL gyms. I don't actually think that's a well-understood concept for many folks in the audience, but think of that as creating simulated environments or digital twins for tasks you might want to test. Maybe that's a coding environment; maybe that's a contact center, as we've used that a couple of times. But it allows you to simulate a series of function calls, tasks, or environments by which, if you're going to train a model or a task, you can test how it's going to work in a manufacturing environment before you roll it out to your actual physical world. And I think that's more and more an interesting topic for both model builders and the enterprise.

Peter Diamandis

I want to go around and maybe ask some final questions of Matt. Alex, do you want to kick us off?

Dr. Alexander Wissner-Gross

Yeah. I think the most interesting crux of what we're discussing here is: What is the future of human expertise? For that matter, does human expertise have a future? And assuming it does, what's the half-life of the value of human expertise?

To put that question to Matt, what do you think, of all the forms of human expertise, all of the labor categories and job roles that exist in the economy today, will be the last 3 of those job roles or forms of expertise to disappear or ultimately succumb to AI? What are the last 3 to survive?

Peter Diamandis

Last expert standing. Okay.

Matthew Fitzpatrick

That's right. I'll go back to where I started the episode. I think a lot of the commentary on mass shifts ignores the actual function of jobs in society today. So let's take sectors, for example: oil and gas. A lot of the functional expertise—geoscience, if you look at seismic engineers, people on oil and gas sites drilling—that is a human function. Real estate is another example. You can go down a whole list of different areas.

I think there are sectors where you're going to see more disruption in the near term. I call out a couple of them: BPOs and legal services. I think media is a fast-changing area. But I'm also not exactly sure that those lead to negative—meaning, have negative employment consequences.

If you take media, it's a really interesting one. 5, 6, 8 years ago, I think media as a category really struggled in a lot of ways, with paid media as an example. And you've actually now seen, in the last couple of years, Substack, Medium, and all these blogs become much more interesting. You have way more media entrepreneurs. And so you've changed the function of society, and where the money is coming from changes, but it has not changed total employment.

And look, I understand a lot of the skepticism that says AI is going to radically change everything, but I think if you look at American society for the last 100 years, it's something like 25% of every high school class goes into a field that did not exist when they were in high school. And the reason that persists is people go into the working world understanding the tools they have, thinking about what they can create from that.

One of my favorite statistics, which I saw The Wall Street Journal report a couple of weeks ago, is that 20% of U.S. employment right now is in digital ecosystem jobs. And something like 9% of U.S. citizens are full-time social media influencers. It's mind-boggling to me.

But again, this is the changing nature of work, and I think that pattern will persist. I think the core of what will change is that the process of looking up information across multiple systems and documents is going to become less valuable. But I think all the jobs that involve human interaction and physical work—

Peter Diamandis

Optimistic.

Matthew Fitzpatrick

I was actually on another panel with someone who runs a recruiting company. They were saying that job profile, I think, will 2x, 3x, 4x over the next couple of years. And so that will have pretty interesting implications for the education system and everything else. But I don't think it will all be displacement; I think we'll see an evolution.

Peter Diamandis

I mean, the humanoid robot electrician and plumber. Alex, very quickly, what are your 3 last-standing human roles here? Or your last 3 standing?

Dr. Alexander Wissner-Gross

So I'll present multiple competing hypotheses.

Peter Diamandis

Briefly.

Dr. Alexander Wissner-Gross

One hypothesis is that it's the politician, because they have to make the laws. Another hypothesis is that it's the greatest intellects: the physicists or mathematicians. Even though, as we talk on the pod, math and the sciences are all getting solved, on the one hand, they're still perhaps—to the extent that that represents the culmination of human intellectual accomplishment—maybe the greatest intellects will be the last to be automated.

There's another school of thought that says, "No, it's the roles that involve the greatest need for human authenticity." Because even though it's not actually a capabilities question, people nonetheless demand human contact, or something to that effect. And so it's going to be the highest-touch job roles, where people just want to know that there's a human counterparty on the other side of the interaction. So that's a set of 3 hypotheses.

Peter Diamandis

Tastemakers will dominate. That's authenticity. Bucket number 3. Saylor said that word for word, actually. Yeah, we had that enjoyable sunset conversation on his boat. Salim, do you want to go next with a closing question for Matt?

Salim Ismail

I think you covered some of the industries that are going after it. You guys have done some government work. Where in government functionality do you see the biggest opportunity for AI automation, efficiency, et cetera?

Matthew Fitzpatrick

Yeah, everywhere. Look, I actually think this could be one of the really positive trends for society. I saw a study recently that AI-assisted permitting could cut energy and data center project implementation timelines by 50%.

Think about housing. One of the biggest challenges right now for housing development in the U.S. is NIMBY regulations and how complex it is to build housing because of the myriad of different regulations and zoning constraints by location. The OECD came out with a report that AI could shrink public-sector process-cycle timelines by 70% on licensing, benefits approvals, and compliance.

To me, the simplest thing that AI can do is project management and timelines related to all spending and infrastructure deployments.

Peter Diamandis

This would be a really positive thing for society, in my mind. Amazing. Good question, Salim. Dave, why don't you close us out on the questions here?

Dave Blundin

I've got so many, but I'll pick the best. First, Matt, how many hours of video footage will there be of you 1 year from today compared to 1 year ago? Because I know we saw each other in Riyadh a few weeks ago.

And I know that you are the thought leader in this whole bottleneck of AI getting into the enterprise. It feels like what we're doing right now. The footage of you that's out there right now is all this CNBC, Bloomberg-type, 5-minute format. But here we're getting your real thoughts. It's just so much better. How many hours can we count on 1 year from today?

Matthew Fitzpatrick

Well, look, I think as of 12 months ago, I had done almost no interviews of any kind. So this job has been fun on that front. What I enjoy about the podcast format is that it does allow you to talk about some more complex topics. I particularly find podcasts like this really interesting. So hopefully many more in the year to come.

Dave Blundin

Well, I'm hoping for at least a 10X on that. And then my follow-up question is: the avatar version of you that's also out there talking—is that a 2026 thing, do you think, or when?

Matthew Fitzpatrick

Yeah, it probably happens in 2026. I don't think it would be that hard to train an avatar off of my public statements. So I think that'll be interesting. We are actually working in the sports space on the topic of avatar training. And I think it is an interesting space where you could imagine a lot of different areas where, rather than a chatbot interaction, people want to speak to people they know via an avatar. I actually think that will become a more natural part of society and a pretty interesting one, actually.

Dave Blundin

I totally agree. I just think the timeline could be as soon as 2 months, as far as I'm concerned.

Peter Diamandis

What makes you think it's not an avatar we're speaking to right now, Dave?

Dave Blundin

That's a good question.

Matthew Fitzpatrick

That seems very human, actually. I don't know. The best ones are.

Peter Diamandis

The orbs behind you kind of give it away.

Dave Blundin

Yeah, they are pretty strange. That's not real.

Peter Diamandis

Matt, where do people find you? Where do people find Invisible? Who should go to Invisible to check out what you do and how you do it?

Matthew Fitzpatrick

Sure. So we have 7 offices now: New York, San Francisco, Austin, Texas, D.C., London, Poland, and Paris. I'm probably the easiest to find. We have an office right off Union Square, which is where I'm at least half the time when I'm not on the road.

In terms of who should come to us, and from the listener base in particular, any mid-cap or enterprise company that knows there is potential in its business, that knows AI can transform it in a positive way, and is struggling to bring all the pieces together. I think that is the main thing I would say. There is no doubt: the technology Alex is asking about has made an enormous step change over the last couple of years. The hard thing is actually the change management, the operationalization, the metric tracking, and the evaluation.

It's kind of bringing together—I think it's the difference between the... Our founder, Francis, has an idea: you have all the components to build a cake, but you don't have a cake. What we do is we actually bake the cake in the end. We build you something that works. We make AI work, and we use all the modern tools to do that.

Peter Diamandis

Amazing. And the website?

Matthew Fitzpatrick

invisible.tech.ai.

Peter Diamandis

All right. Thank you, Matt. Salim, Dave, AWG, I'm going to see you guys in a couple of days for our 2026 predictions. Make them brilliant. It's going to be fun. All right.

Speaker 1

To benchmark for tracking benchmarks.

Speaker 2

All right. No, that's not the one I'm going to talk about. Okay.

Peter Diamandis

All right, guys. Have a great day.

Which Industries Survive AI, The New AI Benchmarks, and the 2026 Recursive Learning Timeline | #218 | BidClub