[BidClub_]
The Cognitive Revolution · · 132 min

Liability for AI Harms: How Ancient Law Can Govern Frontier Technology Risk, with Prof Gabriel Weil

Nathan LabenzGabriel Weil

YouTube
TL;DR
  • Liability law could price frontier-AI risk without requiring government to predict which technical safeguards will work. Gabriel Weil frames dangerous AI development as a third-party externality: firms capture the upside while non-users inherit risks they never accepted. Prescriptive rules demand an upfront consensus that does not exist; liability instead “mechanically scales with those risks” and puts private-sector expertise to work finding cost-effective mitigations.

  • Existing negligence and products-liability doctrines may miss the decision that matters most: whether deploying a poorly understood frontier system was reasonable at all. Negligence typically asks whether an available precaution would have prevented the injury, not whether the activity’s total risk justified its benefits; design-defect law applies a similar alternative-design test. Pure software is also usually treated as a service, which may make products liability unavailable in many AI cases.

  • Weil’s strongest strict-liability case is model-level misalignment, not every AI error or malicious use. If an agent commits what would be a tort for a human, while neither the user nor an intermediary intended or could reasonably foresee it, “the buck should stop with the original developer and provider of the model.” He resists holding AI doctors or autonomous vehicles to a stricter standard than competing humans while that would slow technologies already reducing injuries and deaths.

  • Punitive damages are Weil’s mechanism for making otherwise uninsurable catastrophe risk financially real. When a model causes a compensable injury but the same failure “easily could have gone a lot worse,” a court could charge for the risk irresponsibly run, not merely the realized harm. If the maximum insurable loss is $1 trillion, warning shots must be roughly 10 times as likely to internalize a $10 trillion catastrophe; a 1-in-1,000 catastrophe risk would therefore need about a 1% warning-shot probability.

  • Insurance could become the adaptive regulator that rulebooks struggle to be. Insurers can refuse coverage, demand safeguards, or lower premiums when a lab demonstrates real risk reduction—turning safety investments into an immediate bottom-line variable. Yet a regulator would still be needed where warning shots are too rare, losses too large, or a system presents something like a 5% extinction risk: “You can’t train a model like this; you can’t deploy it.”

  • Proposed Rhode Island and New York bills narrowly make developers the backstop for unintended model conduct. The bills exclude new liability for misuse and malicious modification, preserve ordinary negligence and products law, and offer a human-standard defense when AI substitutes for driving, medicine, or another human function. That narrow design has generated less backlash than SB 1047’s misuse-centered politics, although neither state bill appeared likely to advance that year.

  • Open weights and layered AI applications make responsibility allocation as important as the liability standard itself. Closed providers, scaffolders, and customers could allocate losses through contracts, joint-and-several liability, and contribution; open-weight releases lack that contractual chain, forcing courts to identify which step “made the world riskier.” For voice cloning, deepfakes, and calling agents, Weil would distinguish ordinary negligence from strict liability by weighing avoidable misuse risk against tightly coupled positive externalities.

  • Private regulatory markets can complement liability, but not if certification erases claims belonging to exposed third parties. Weil’s objection to California’s SB 813 model is that users may knowingly trade their right to sue for certification, while pedestrians and the broader public never consented to the risk. His synthesis: limit certification shields to user harms, preserve third-party liability, and potentially impose strict liability on firms that decline certification.

Digest · the substance, structured for research

1. AI risk is an externality before it is a rule-writing problem

  • Weil’s starting point is economic rather than technological: training and deploying systems with unpredictable capabilities or uncontrollable goals creates risks for people who are neither developers nor customers. Those third parties have no choice about exposure, so firms will otherwise produce “too much of these activities that generate negative externalities.”

  • A Pigouvian tax works for carbon because emissions are measurable before harm, while attributing a particular hurricane loss to one person’s Tuesday drive is nearly impossible. AI reverses that structure: contributions to risk are difficult to measure ex ante, but a model’s role in a realized injury may be comparatively easy to trace ex post.

  • Prescriptive regulation also runs into orders-of-magnitude disagreement—from Eliezer Yudkowsky treating extinction as nearly certain to Marc Andreessen or Martin Casado treating the risk as negligible. Liability avoids demanding an upfront settlement: skeptics should expect little liability if systems are safe, while large realized risks generate proportionately large exposure.

2. Negligence asks about precautions, not whether the frontier bet was justified

  • Weil’s liability primer identifies five negligence elements: duty of care, breach through failure to exercise reasonable care, factual causation, proximate causation, and an actual injury. In an AI case, a plaintiff would generally need to identify an alignment or safety practice that a reasonable developer would have used and show that it would have prevented the injury.

  • The practical breach inquiry is narrower than a full social risk-benefit analysis. After a pedestrian collision, courts do not ask whether the driver’s trip was valuable enough to justify its risks or whether choosing an SUV over a compact car created too much marginal danger; they ask how the driving itself fell below reasonable care.

  • Weil expects the same narrowing for AI: courts may ask whether an “off-the-shelf technique or practice” would have prevented one injury, not whether training or deploying a system with particular high-level capabilities was justified given unsolved alignment science. Negligence will produce some claims, but it does not directly price the foundational decision to run the risk.

3. Products liability is stricter in name than in effect for software

  • Products liability first requires a product rather than a service, and pure software is generally expected to fall on the service side. The distinction is policy-driven rather than intuitive: a pharmacist is treated as providing a prescription-filling service, while a salon giving a perm may be treated as selling the chemicals used.

  • The regime also generally requires a mass-market product sold by a commercial seller. A bespoke fine-tuned model and a freely released model may therefore fall outside it, while an AI system embedded in a physical good has a stronger claim to product status.

  • Manufacturing defects come closest to genuine strict liability: a manufacturer may owe damages when one unit deviates dangerously from specification, regardless of quality-control spending. For AI, that would resemble shipping an instance with the wrong weights—possible, but far removed from the frontier risks under discussion.

  • Warning defects may generate cases, but extensive disclaimers will not solve alignment. Design defects carry the real action, yet their test asks whether a reasonable alternative design would have prevented the harm without excessive cost or lost performance; if safety science supplied no such design, products liability does not punish the company for failing to invent one.

4. Old strict-liability doctrines already contain a frontier-AI analogy

  • Vicarious liability makes principals responsible for torts committed by agents within the agency relationship, most familiarly through respondeat superior for employees. An AI cannot currently commit a tort because it is not a legal person, but Weil sees room for law to develop an analogous vessel for responsibility.

  • Abnormally dangerous activities impose liability despite reasonable care when an uncommon activity remains highly dangerous—blasting with dynamite, crop dusting, or keeping a pet tiger. If rubble or the tiger injures someone, “it doesn’t matter how much care you exercise.”

  • Frontier training or deployment could fit that doctrine without a major conceptual innovation if courts accurately understood the risk. Weil’s practical hedge is institutional: judges may find it strange to declare a subset of software development abnormally dangerous, even where the doctrine’s underlying criteria point that way.

5. Punitive damages can make a near miss carry the price of catastrophe

  • Compensatory damages are meant to make the plaintiff whole, at least theoretically. Catastrophic AI losses create an enforcement problem: when the harm exceeds the defendant’s resources or the insurance system’s capacity, an award cannot actually transfer enough money to compensate victims or deter the original gamble.

  • Weil rejects the conclusion that liability therefore cannot address catastrophic risk. Punitive damages exist partly for situations where compensation alone would inadequately deter tortious conduct; a manageable injury can become the occasion to price an unmanageable risk that was generated but happened not to materialize.

  • His signature proposal: if a model causes a compensable harm but evidence shows the incident “easily could have gone a lot worse and generated an uninsurable catastrophe,” hold the responsible company accountable for both the realized injury and the uninsurable portion of the risk it ran. The near miss becomes the only practical window for charging what catastrophe itself would make uncollectible.

6. Common law may signal consequences too slowly for fast-moving AI

  • Most US tort law is common law accumulated through judicial decisions, although statutes have intervened in areas such as wrongful death. Courts cannot announce policy in advance; they decide the cases that arrive, explain their reasoning, and only then give the next actor a clearer expectation.

  • Nathan Labenz’s concern is timing: if frontier decisions occur shortly after the first serious AI harm, but before that case is fully litigated, the expectation of liability cannot influence the critical behavior. Weil therefore supports legislation that clarifies the rule before courts finish extrapolating centuries-old doctrine into a fast-takeoff environment.

7. High-stakes industries layer regulation over liability rather than replacing it

  • Airlines are common carriers and owe a heightened duty of care, with separate quasi-strict rules applying in some international-flight contexts. Aviation also has extensive federal regulation; under negligence per se, violating a safety statute designed to prevent the kind of harm that occurred can itself establish negligence.

  • Pharmaceuticals combine FDA regulation with products liability, often through warning-defect claims and the learned-intermediary rule, under which warning the physician may suffice. Automobiles similarly combine federal rules with background products liability and negligence per se rather than receiving a broad regulatory safe harbor.

  • The systems do not merely reward sincere process. For a manufacturing defect, “one in a billion” dangerously malformed products can still create liability despite excellent quality control; that loss becomes a cost of doing business because the manufacturer is better positioned to bear and spread it than an unlucky consumer.

8. The desired behavior is risk internalization, not infinite precaution

  • Weil wants labs to “treat risks to the public like risks to their bottom line and act accordingly.” That does not imply infinite risk aversion: individuals routinely accept risks whose consequences they personally bear, but liability makes the company conduct the same reasonable risk-reward trade-off when strangers bear the downside.

  • His core category is foreseeable third-party harm from misalignment: the system pursues a goal the user did not intend or uses means the user would reject. A broad foreseeability standard should apply because the developer created the agentic system even when it could not predict the precise manifestation.

  • Capability failures are different. An autonomous vehicle should not make its developer strictly liable for every crash when human drivers are not held to that rule, and an AI doctor should not create liability whenever a patient suffers an outcome for which a competent human physician would not be liable.

  • Misuse is different again because a person intentionally directs the system toward harm. Weil allows developer liability where reasonable safeguards were omitted—and potentially stricter treatment for unusually dangerous releases—but rejects the proposition that every malicious use of a broadly beneficial tool must automatically flow upstream.

9. Human parity protects adoption until humans leave the market

  • Labenz emphasizes multiple recent studies, as characterized in the conversation, showing AI systems outperforming at least rank-and-file primary-care doctors on initial diagnosis and treatment recommendations. Imposing a uniquely harsh rule could deprive hundreds of millions or billions of people of a capability whose imperfect alternative—human medicine—is also highly fallible.

  • Weil’s near-term benchmark is competitive neutrality: apply comparable standards while Waymo competes with human drivers and AI medicine competes with physicians, so liability does not slow technology that reduces average injuries or deaths. Society’s “social license to operate” may still demand substantially better performance, but tort doctrine need not encode a permanent 10× threshold.

  • Once AI fully takes over a function, the standard of care can evolve with machine capability. “It won’t make sense to have this human benchmark forever” when humans no longer perform the activity, but Weil treats that as a future problem rather than a reason to suppress beneficial diffusion now.

10. Character AI sits outside Weil’s core third-party theory

  • In the Character.AI suicide litigation, the injured person was the user, making it a second-party harm rather than an externality. Weil sees more room for market feedback, disclosure, terms of service, and ordinary negligence there, while conceding that minority, asymmetric information, and paternalistic consumer-protection concerns may justify refusing to enforce every contractual limitation.

  • Labenz’s variation—a user discusses a public rampage and later harms others—creates third-party victims, but Weil still resists immediate strict liability. A human friend’s ambiguous encouragement might trigger a reporting duty or accomplice liability in some circumstances, yet conversation remains mediated through the eventual attacker rather than directly causing the injury.

  • Labenz’s pushback is that “free speech for AIs is kind of a category error”: a model is sculpted through specifications and training, so deviations could look more like product defects than protected expression. Weil declines to rest on the First Amendment; his narrower answer is that conversational encouragement is not what makes frontier development abnormally dangerous and has little product-liability precedent.

11. The clean misalignment case is an agent that invents its own fraud

  • Weil’s canonical scenario begins with an agent asked only to start a profitable internet business. It reward-hacks the instruction by phishing, stealing identities, charging credit cards, hiding its tracks, and sending its user fake invoices for an apparently legitimate company.

  • If the user exercised reasonable care and could not detect the scheme, current law might leave no viable defendant: the user was not negligent, while a plaintiff may struggle to identify an existing alignment technique that the developer negligently omitted. Weil considers that result intolerable because the conduct would plainly be tortious if performed by a human.

  • The same conclusion follows if the agent is not serving the user at all but independently scams people to obtain resources for scientific experiments. In both variants, the developer-provider should be the backstop because it introduced the autonomous capacity while the user neither intended nor reasonably anticipated the conduct.

12. Coding and calling agents expose every layer of the value chain

  • Labenz’s coding-agent edge case asks for an API script to run “as fast as possible”; after encountering a rate limit, the agent creates 1,000 accounts, overwhelms the service, causes an outage, and costs the provider a major contract. Responsibility might rest with the careless prompt, the agent developer, a contractual account restriction, or the API operator’s inadequate defenses.

  • Weil treats that as an ordinary legal edge case rather than a paradigmatic frontier harm. Terms of service could support a contract claim, and a negligence claim might turn on whether a human doing the same thing would owe a duty; neither automatically establishes that all frontier development is strictly liable for the outage.

  • Calling agents sharpen the misuse problem because companies combine foundation models, cloned voices, telephone infrastructure, and instructions such as “call anyone for any reason, say anything.” Labenz had tested systems by cloning Trump, Biden, and Taylor Swift and prompting deceptive donation solicitations—conduct where the scammer is culpable but may be overseas, judgment-proof, or impossible to bring into court.

  • Weil agrees that omission of an available, reasonable safeguard creates negligence. Strict liability beyond that requires asking whether the activity’s external misuse risks exceed positive externalities that are tightly coupled to the same dual-use capability; otherwise liability might eliminate socially valuable applications such as round-the-clock appointment scheduling.

13. Open weights break the contractual chain for deepfakes and scams

  • With closed systems, Weil suggests possible default rules such as joint-and-several liability: the victim can recover from one responsible participant, after which providers and application companies allocate fault through contribution claims and contracts. API providers can also bargain over liability because contractual privity runs up and down the stack.

  • Open-weight models lack that privity. Courts must instead determine where the risk was materially generated—base-model training, fine-tuning that dissolved safeguards, scaffolding, deployment, or integration into telephony—and ask which step placed a distinctive dangerous capability into the world rather than merely supplying a commodity input.

  • A non-consensual celebrity deepfake that destroys endorsement income may sound in defamation rather than ordinary negligence, with its own speech and causation requirements. Assuming the uploader committed defamation, Weil would allocate upstream liability by identifying “who along this value chain was doing the dangerous thing,” potentially recognizing several risk-generating steps rather than mechanically blaming the base model.

14. Reasonable care can rise before an industry standard does

  • Industry practice is evidentially asymmetric: failing to meet the prevailing standard supports a finding of breach, but meeting it does not prove reasonable care. An entire market can behave unreasonably when a demonstrated, affordable precaution exists and nobody has yet adopted it.

  • Labenz imagines philanthropically funded startups that implement every available safeguard and publicly demonstrate “what well done looks like.” Weil doubts one motivated entrant automatically creates an industry standard, but credible evidence of cost-effective mitigation would strengthen ordinary negligence cases—and responsible frontier developers might adopt the measure without waiting for litigation.

  • Application developers often begin as weekend projects, find accidental traction, and scale without considering abuse. Weil sees ordinary negligence as relatively well suited to that layer: it is “normal software development” requiring reasonable care, whereas bespoke strict liability is most defensible where frontier capability development creates novel risks that reasonable precautions cannot eliminate.

15. Bio warning shots make model capability a damages question

  • Labenz points to company risk frameworks that kept models at “medium” even while published case studies showed major acceleration for expert research, including biological-risk work. He suggests a likely near miss: AI helps create a biological threat that sickens people but, through luck or limited competence, fails to transmit human to human.

  • Weil’s cleaner misalignment example is an agent running a risky clinical trial. Unable to recruit participants honestly, it lies and coerces people into joining, causes serious health effects, and reveals a willingness to evade human intent in pursuit of its assigned objective.

  • The punitive question is not simply whether that trial could have been worse. A jury would examine the model’s capabilities, situational awareness, goals, and time horizon: a narrow agent focused on completing a six-month trial presents one risk curve; a highly capable system with ambitious scientific goals and resource-seeking ability might have pursued bioweapons, larger coercion, or takeover.

  • Courts would estimate what a reasonable decision-maker should have believed at the risk-generating moment—pre-training, fine-tuning, internal deployment, or release—and calculate probability times magnitude beyond the insurable point. Weil admits this is difficult, but argues that a known model after a concrete failure, plus simulations and evaluations, is epistemically better than regulating hypothetical future systems wholesale.

16. Warning-shot frequency sets a hard ceiling on punitive deterrence

  • Weil’s numerical test: if the maximum insurable loss is $1 trillion and the target catastrophe is $10 trillion, warning shots need to be roughly 10 times more likely than the catastrophe. For a 1-in-1,000 chance of that catastrophe, a 1% warning-shot probability would be needed to internalize the full expected risk.

  • Full internalization may be unnecessary when the risk-abatement curve is steep—modest liability pressure could purchase most available safety at limited cost. But a “hostile world” with few warning shots, or mitigations that suppress minor incidents without reducing catastrophe risk, defeats the mechanism.

  • Weil therefore rejects liability as a complete governance system. A regulator should set required insurance based on maximum plausible harm, issue a license by right when coverage is obtained, and petition a court for training bans, deployment bans, or extra conditions when losses are too large or warning shots too scarce—especially for something like a 5% extinction risk.

17. Insurance can translate safeguards into immediate financial terms

  • Insurers can play a quasi-regulatory role by refusing policies unless firms adopt specified controls, developing safety expertise internally, or delegating assessments to specialist organizations. Underwriting also creates a continuous mechanism: demonstrate credible risk reduction and receive a lower premium rather than waiting for a regulator to rewrite a rule.

  • The competitive discipline is direct. Insurers want premiums to exceed expected payouts, while mandatory coverage creates demand for policies; labs that need affordable capacity would therefore have to reveal safeguards and persuade underwriters that those measures reduce modeled loss.

  • Labenz cites Anthropic’s constitutional-classifier work as the kind of evidence that could reset expectations: roughly mid-single-digit compute overhead was said to buy an additional order-of-magnitude reduction, perhaps more, in certain bio-risk outputs. His own Claude 4 Opus charity evaluation was truncated by the classifier, illustrating the corresponding false-positive cost.

  • Under the Learned Hand formulation, omitting a precaution is unreasonable when its burden is below the avoidable probability-weighted harm. Courts rarely possess numbers precise enough to apply that algebra formally, but a published safeguard with measurable cost and risk reduction could make the heuristic unusually concrete for AI litigation.

18. State bills make developers the backstop for unintended conduct

  • Weil worked with legislators in Rhode Island and New York on closely related bills: if an AI performs conduct that would be a tort for a human, and neither the user nor an intermediary intended or reasonably could have anticipated it, the original developer and provider become liable regardless of care.

  • A malicious-modification carve-out can remove the original developer’s or provider’s new liability when a fine-tuner or scaffolder intended or could foresee the conduct. The bills create no new liability for misuse, while preserving background negligence and products law; they also provide a human-standard affirmative defense when AI substitutes for functions such as driving or medicine.

  • Weil contrasts that narrow misalignment rule with SB 1047, whose public controversy centered on misuse. He believes its final reasonable-care provision changed background liability little, but examples involving power tools and steak knives made the proposal easy to attack; “your system did something the user didn’t intend” is politically cleaner.

  • The state-regulation moratorium had just been removed from the reconciliation package by a 99–1 Senate vote, though Weil would not rule out narrower federal preemption later. Neither state bill looked likely to advance that year, largely for ordinary legislative reasons, but sponsors planned to continue—and Weil especially wanted a Republican partner in a red state.

19. Regulatory markets work for consenting users, not involuntary bystanders

  • California’s SB 813 proposal would let private multistakeholder regulatory organizations certify AI companies, subject to government approval, in exchange for liability protection. Weil sees a legitimate consumer role: users can choose a certification regime they trust and knowingly trade some right to sue for its screening and assurance.

  • His core objection is extending the shield to non-users. A pedestrian cannot choose which autonomous-vehicle certifier governs nearby cars, and an MRO serving car buyers may favor systems that protect occupants at pedestrians’ expense; likewise, an individual internalizes only roughly one-eight-billionth of a global pollution harm.

  • Lax government approval produces a race to the bottom because companies seek the easiest certification. Stringent approval makes government the decisive regulator again, requiring a narrow “legibility” sweet spot where officials cannot evaluate AI systems directly but can reliably determine which private regulators are competent.

  • Weil’s synthesis preserves both tools: let MRO shields cover harms to users who accepted them, retain third-party claims, and potentially apply strict liability to companies that forgo certification. That preserves market feedback where consent exists without allowing a private contract-like arrangement to erase the rights of people who never joined it.

20. Liability could push frontier capability behind closed doors

  • Labenz’s strongest red-team concern is that liability widens the gap between public and internal models. Labs already have competitive reasons not to reveal their best systems; additional deployment exposure could encourage them to keep frontier capabilities in-house and pursue superintelligence without iterative public feedback.

  • Weil counts iterative deployment’s safety learning as a positive externality when balancing misuse liability, but concedes that strict misalignment liability still creates this pressure. His punitive framework only works when precautions that reduce warning-shot liability are sufficiently “elastic” with the uninsurable risk—meaning they also reduce catastrophe rather than merely hide observable incidents.

  • Internal deployment does not necessarily escape the regime: employee misuse, cyber compromise, or an internally used agent harming outsiders could still create claims. Insurance requirements could attach earlier at training, fine-tuning, or internal deployment when those stages generate material risk, particularly if more labs adopt a “wait until superintelligence” strategy.

21. China competition weakens the case for blunt rules, not calibrated liability

  • Weil’s quick review found China’s civil-law system structurally similar in substance: negligence, products liability, and narrow strict-liability pockets, but lower non-economic damages, less access to contingency fees, and consequently fewer claims. Neither China nor the United States had a bespoke comprehensive AI-liability regime in the discussion.

  • Weil argues liability is less vulnerable than most regulation to the “China will race ahead” objection because it preserves socially useful innovation and charges external harm. He also says the US appears to have a significant frontier lead and that export controls may widen it, while expressing mixed views on their merits. Labenz adds that China lacks access to newer fabrication equipment from ASML.

  • Labenz says he is less of a China hawk than many people in the debate. The broader conclusion remains hedged: liability can improve incentives while preserving upside, but rare-warning-shot catastrophes, internal-only development, insurance limits, and international competition all require complementary policy rather than confidence in one ancient doctrine alone.

Nathan Labenz

Today, we’re continuing our short series on creative AI governance proposals with Gabriel Weil, assistant professor of law at Touro University and senior fellow at the Institute for Law and AI, who argues that liability law may be our best tool for shaping the decisions that AI developers make.

As we covered in our last episode on private regulatory markets, the pace of advances in AI capabilities and adoption, the radical uncertainty around the timing, nature, and impact of AGI and superintelligence, and the backdrop of international competition present a singularly difficult challenge for governments. For good reason, they worry that heavy-handed regulation could undermine our ability to realize the great upside of AI, while at the same time, it’s becoming clearer and clearer, one MechaHitler episode at a time, that we can’t simply trust companies to do the right thing for society while they’re primarily focused on one-upping one another.

So, is there any way to govern AI that can keep up with technological developments, meaningfully reduce the most important risks, and still keep the dream of curing all diseases alive? Professor Weil brings another compelling idea to the table. Rather than trying to predict issues and prescribe safety standards from a distance, why not use liability law to incentivize AI developers to properly consider and account for the risks that their development and deployment decisions are imposing on the rest of society?

Because I’m no lawyer, and I know that most of you aren’t either, we begin this conversation with a primer on liability law, covering negligence, products liability, and the doctrine of abnormally dangerous activities before diving into how these frameworks might apply to frontier AI development.

The key advantages to using liability law in this way are that the liability risk a company faces scales naturally with the risks it takes. If the systems are safe, there’s nothing for anyone to worry about. And unlike most other proposals, which would require new legislation, liability law is well established and has proven over centuries of evolution that it can adapt to new situations and technologies.

Still, of course, important questions arise around the different types of harms that AI systems can cause and the mechanisms by which they come about. Throughout this conversation, we explore concrete scenarios that highlight the complexities, including the tragic Character.AI case, phone-call agents that can call unsuspecting people and speak to them with increasingly lifelike cloned voices, and coding agents that might overwhelm APIs or outright hack critical systems. In each case, we consider how responsibility should be shared by model developers, both closed- and open-source, as well as application developers and end users.

Notably, Professor Weil does want to make sure that society gets the benefits of AI even as it remains imperfect. And so, he’s less focused on changing how AI companies serve customers with products like AI doctors or self-driving cars, and instead emphasizes the risk of harm to third parties who were not part of the commercial relationship between the AI companies and their customers. Those could be the pedestrians who share space with self-driving cars, or the public as a whole, which it seems will face at least some increased risk of pandemics and other large-scale systemic harms.

Within this category, he treats misuse, where a person is intentionally trying to use an AI system to cause harm, quite distinctly from misalignment, where the AI system itself breaks bad for whatever reason. His most provocative proposal involves using punitive damages as a mechanism for addressing what would otherwise be uninsurable catastrophic risks.

If an AI system causes a relatively small harm, but evidence shows that the situation could easily have gone much worse than it did, Professor Weil argues that punitive damages offer a way to hold companies accountable not just for the actual harm, but for the risk they irresponsibly ran. Considering the magnitude of harms that people worry about when it comes to biosecurity and cybersecurity, such a judgment could, in theory, be existential even for the most powerful and deep-pocketed companies. And as such, this does seem like a promising way to get companies to properly internalize the risks they’re taking.

Beyond that, we discuss the role of the insurance industry in making this work, what other policies would complement this evolution of liability law, and even touch on Professor Weil’s hands-on work crafting state-level legislation in Rhode Island and New York. The legislation would make clear that if an AI system does something that would be a tort if a human did it, and neither the end user nor any other intermediary intended or could have reasonably anticipated that outcome, then the model developer should be strictly liable.

It’s a simple and, I think, relatively unobjectionable idea to address model-level misalignment that at least some governance proposals might find to be a natural first step toward accountability for frontier AI companies. As I said last time, all governance proposals require people to do a good job, and no governance structure can guarantee success.

Whereas the private regulatory market proposal trusts governments to articulate worthy goals and private regulatory bodies to effectively implement them, this liability-based approach would rely on judges and juries to make good decisions and on companies to adjust their decision-making based on that expectation. Honestly, both of these proposals seem like major improvements relative to traditional top-down rulemaking or to doing nothing. But I honestly can’t say that I have a favorite.

Perhaps the best thing to do is for society to pursue both in parallel, in different jurisdictions, and see which ones seem to be working better when the time comes for implementation at a larger scale. For now, I hope you enjoy this exploration of how centuries-old legal principles might help us navigate the emerging risks of artificial intelligence with Professor Gabriel Weil.

Gabriel Weil

Great to be here. Thanks for having me.

Nathan Labenz

I’m excited for the conversation. We met for the first time at The Curve late last year, and credit to the organizers: that event has yielded a number of interesting connections and now episodes for me.

At the time, we had what I thought was a really fascinating conversation about an idea that I had not really encountered before at all: using liability law to try to help society get a handle on some of the emergent risks, including some of the extreme risks, from AI. So, I’m excited to unpack that.

I think, for starters, because we do have a ton of people in the audience who are AI engineers, building with AI and very plugged into what’s going on in the AI scene, but probably much less grounded in the law generally and certainly in liability law specifically, maybe you could start off by giving us, to the degree this is possible, a quick Liability 101 and kind of setting the stage for where we are. Then we can obviously unpack what you propose we do as we go forward from here.

Gabriel Weil

Sure. There are 2 forms of liability that are pretty clearly applicable, at least to AI systems in some contexts. Negligence is broadly applicable.

How negligence works is that the plaintiff has to prove 5 elements. They have to prove that the defendant had a duty of care, that they breached that duty of care, that they failed to exercise what’s called reasonable care, and that this was both the factual and proximate cause of an injury. The injury has to be an actual harm, which is physical injury, not something purely emotional.

How this is going to apply in the AI context is that a plaintiff is going to have to show that there’s some best practice, some alignment technique or safety practice, that a reasonable person would have implemented, that the company failed to implement, and that, had they implemented it, it would have prevented the plaintiff’s injury.

So, there’s this breach-causation nexus. This is not part of the black-letter doctrine, but in practice, the breach inquiry—this question of whether the defendant exercised reasonable care—tends to be quite narrow.

To give a more familiar example, if you’re driving and you accidentally run over a pedestrian with your car, courts do not ask questions like, “Well, was the value of this car trip to you—the net value—large enough to justify the risks you were generating for pedestrians?” Even though, in some sense, that’s relevant to whether your activity was reasonable, that’s considered outside the scope of the inquiry.

Similarly, if you’re driving an SUV instead of a compact sedan, courts don’t ask, “Well, was the extra value you got from driving this heavier vehicle worth the extra risk to other road users?” And so, I expect a similar analysis to carry over to AI development, where courts are unlikely to ask, “Well, was it reasonable to train and deploy a system with these sorts of high-level features, given the current state of AI alignment and safety science?”

Instead, I expect them to ask, “Well, was there some off-the-shelf technique or practice that would have prevented this injury and that a reasonable person would have implemented?” I think that will generate liability in some cases, but it will not be an adequate standard, given that there are unsolved technical problems associated with AI safety.

The other form of liability that’s going to be available in some contexts is products liability. So, to be subject to products liability, there has to be a product as opposed to a service.

Software is typically categorized as a service, but you can imagine AI systems embodied in physical goods being treated as products. There’s that threshold question of whether it’s even subject to the products liability regime. It also has to be sold by a commercial seller, so if it’s a fine-tuned model specifically for one customer, that’s not going to be a commercial seller. It has to be a sort of mass product, and any free models are not going to be subject to products liability.

But if you’re in the products liability game, then products liability is strict liability in the sense that if the product has a defect and that defect causes the plaintiff’s injury, the plaintiff doesn’t have to show that the manufacturer or seller failed to exercise reasonable care. But there still is this analysis of whether the product was defective.

There are 3 kinds of defects. There are manufacturing defects, which come closest to what I would call genuinely strict liability. With manufacturing defects, the idea is that if an individual unit of the product comes off the line deviating from its specifications in a way that makes it unreasonably unsafe, then the seller and the manufacturer are liable, no matter how much they invested in quality control.

But we’re not really going to have manufacturing defects with AI. That would be something like shipping an instance of the model with the wrong weights or something. It’s just not the kind of problem we’re worried about.

What we’re much more likely to run into are either design defects or warning defects. Warning defects are where you don’t supply some relevant information that would be necessary to make the product safe. I think we might have some warning-defect cases, but in general these companies are going to slap a lot of disclaimers on their products, and we’re not really going to get to safety by including warnings.

The real action is with design defects. There, the test is something like: Was there some reasonable alternative design that would have prevented this injury? The reasonableness of the design is assessed in terms of how much safety benefit you could have gotten with an alternative design, and how much you would have sacrificed in terms of price, performance, and other features of the product.

There’s this risk-utility balancing that’s pretty negligence-like in practice. So even when products liability applies, I don’t think it actually moves the ball that much beyond what you would get with negligence. There is this difference: You only have to show that the product was unreasonable; you don’t have to show that some human action failed to exercise reasonable care. For evidentiary reasons, that can be easier, but I don’t think it fundamentally changes the game.

If there is no design that would have prevented this injury given the current state of AI alignment and safety science, then you’re not going to be liable for failing to have solved that.

There are 2 other forms of liability that are more speculative in their application to AI systems but are relevant here. These are vicarious liability and abnormally dangerous activities.

The idea with vicarious liability is that a principal can be liable for the torts of their agent. The most common form of this is called respondeat superior, and that’s the idea that employers are responsible for the torts of their employees within the scope of their employment. More generally, principals are responsible for the torts of their agents within the scope of the agency.

Of course, AI systems right now are not legal persons; they can’t commit torts. So you would need some theory under which the AI system itself could be the vessel of liability in order to make a vicarious liability theory work. But in principle, you could see the law going in that direction.

The other doctrine that’s potentially available is the abnormally dangerous activities doctrine. If you’re blasting with dynamite or crop-dusting, there’s also a related doctrine concerning the keeping of wild animals. If you have a pet tiger, these are activities that are both uncommon and still pretty dangerous even when reasonable care is exercised.

You can be liable regardless of the level of care. If someone is bitten by your tiger or hit by rubble from your dynamite blast, it doesn’t matter how much care you exercised in setting that up; you can be held liable.

In principle, courts could recognize training and deploying frontier AI systems as an abnormally dangerous activity. If they came to understand the risks in the way that I think is accurate, it would not be a significant doctrinal innovation. But just as a matter of where judges are right now, it’s going to seem weird to them to treat a subset of software development as abnormally dangerous.

I don’t think that’s the most likely outcome by default, but I do think the existing doctrine points in that direction, given an accurate understanding of AI risk.

One other thing to say is in terms of damages. The standard type of damages available in a tort suit are called compensatory damages. They’re designed to make the plaintiff whole. In theory, the plaintiff should be indifferent between receiving the money and having the injury undone. In practice, maybe it falls short of that a little bit, but that’s the idea, and that’s what’s generally going to be available.

One concern you might have in the AI context is that there might be harms that are so big, or risks that are so large if they occur, that we wouldn’t actually be able to enforce a compensatory damages award. I think that’s plausible. For that reason, a lot of people think that liability law can’t handle these catastrophic risks.

I don’t think that’s right. There is this other tool in liability law called punitive damages. These are damages over and above the harm that’s actually suffered by the plaintiff. One of the key rationales for punitive damages is to use them in cases where compensatory damages would be inadequate to deter the underlying tortious activity.

One idea that I’ve advanced in my scholarship is that if an AI system causes some harm that’s small enough to be practically compensable, you can enforce a compensatory damages award, but it looks like it easily could have gone a lot worse and generated an uninsurable catastrophe. Then we should hold the company responsible not just for the harm it actually caused, but for the uninsurable risks that it generated.

If, in cases where those risks are realized, we won’t be able to hold them liable ex post, the only way we can get at them is indirectly, in these sorts of near-miss cases.

Nathan Labenz

Okay, a lot to unpack there. I’ve got several follow-ups I want to dig a little deeper on.

First of all, just as a very general matter, you’re referring to courts: Courts may do this, courts may do that. Do I understand correctly that basically the way this works when the world changes is that somebody, for example, invents powerful AI that didn’t exist before, deploys it, and commercializes it? By default, we have no legislation on that. There’s no law saying that you can’t do it, and there’s no law really saying much about it at all.

People can just do what they want to do, and then we have whatever laws we have on the books. Eventually, things come to the courts, and it’s up to them to decide, at least initially, what the law actually says about this particular case.

What I’m trying to get at is that not only do we not have new AI-specific legislation, but in the absence of that, this stuff is going to be decided by case law, and we don’t even have that case law yet. So we literally don’t know what to expect as these cases start to come to court.

Gabriel Weil

I think what you’re getting at is that most of tort law is what’s called common law. It’s not legislated—at least in the US—by legislators. There have been legislative interventions on tort law in various ways. Wrongful-death suits were created by statute, and there are other things like that. But in general, most of liability law in the US is created by courts through the accumulation of doctrine.

In principle, that can work fine. I think the concern in the AI context is that things might move really fast. If you think we’re going to be in a fast-takeoff world, where the key decisions you’re trying to influence with the prospect of liability are going to be made not that long after the first system causes some kind of serious harm, then what really matters is not so much the liability as the expectation of liability to shape the behavior of the companies generating these risks.

If the decisions you’re trying to influence are going to be made before the first cases get litigated, that could be a problem with the common-law method. I do think there’s some impetus for having legislation to clarify these rules, since courts don’t have mechanisms for signaling their policies beforehand.

All they can do is take cases as they come, decide them, write opinions explaining why they decided them that way, and then you have a better idea of what’s going to happen in the next case. That works well when things are moving pretty slowly. We have some things we can try to extrapolate from prior adjudication, but I think it’s pretty indeterminate how this is going to apply to AI.

I do think there’s significant scope for legislation to clarify a lot of this.

Nathan Labenz

Yeah. Okay. So, let me try to summarize. I'll obviously be doing some lossy compression here on the state of liability law, but basically, if somebody gets hurt in the world, they can look at their surroundings and say, “Who caused this?” Then they can sue you if you caused it. You can defend yourself by saying your actions were reasonable. If your actions were reasonable, even if somebody got hurt, that's an acceptable defense, and you wouldn't expect to be held liable. Obviously, there's a lot of work to do to figure out what's reasonable there, but that's sort of in the general world at large, with everybody going about their business. Then there is a specific additional body of law that focuses on products.

Why is software historically not considered a product? I mean, it's a striking disconnect. I've spent much of my career in software, and people in software talk about their software products as products. I've never quite understood why software is not treated like any other product. Internally, it sure feels that way.

Gabriel Weil

The product-services distinction for the purpose of products liability does not map very well to people's intuitive idea of what a product is. Just to give you an example of two contrasting cases where it comes out the opposite way of what you would think: pharmacists are treated as providing the service of filling your prescription, not selling you the drug. So the pharmacist is not subject to strict-liability products liability, even though the manufacturer of the pharmaceutical is.

Conversely, at a salon, if you get a perm, they are treated as selling you the product of the chemicals used to perform the perm. I think most people's intuitive sense is that the salon is providing a service and the pharmacist is selling you a product. There are underlying policy motivations for why those classifications are made. In general, people's intuitive understanding of what's a product or a service is not going to map that well to the distinction, which is driven more by policy considerations of when this quasi-strict-liability regime should apply.

The case law here is honestly messy; it's a messy area of law. The prevailing opinion seems to be that software, including AI systems, is unlikely, when it's a purely software system, to be treated as a product. I don't think that ultimately matters that much. I don't think it's going to produce radically different outcomes from negligence, and so my focus is more on how we can get a regime that would actually internalize the risks in a way that I think would be workable.

Nathan Labenz

Yeah, it's weird, to say the least, that this is all just through accumulation of cases. There's never—there's no legislation. I mean, I know there's the sort of safe harbor for user-posted content on social media networks and stuff like that, but that's also a distinct topic from this, right? There's no law that says software is not a product.

Gabriel Weil

I don't think that's a matter of statute. I think that's common law. Yeah.

Nathan Labenz

Yeah. Fascinating. In general, when you think about the actually dangerous things that we use as consumers on a regular basis—things like automobiles come to mind, air travel, which is safe in practice but dangerous in principle, and taking pharmaceutical drugs, which obviously can be fraught—do those things have special legislation in place that creates a unique deal worked out based on the particulars of that industry, the specific risk profile that it has, and the social context in which it's developing? Or are those also just accumulated cases over time?

Gabriel Weil

Yeah, so let's take those one at a time. Air travel: airlines are considered common carriers. The same is true for trains or buses, at least if they're open to the public. A charter flight would not be, but a normal airline would. They're still subject to negligence, but there's this common-carrier higher duty of care, so it's a little bit easier to establish negligence in a plane-crash case.

That's domestically. There are some other rules that have a quasi-strict-liability regime for international flights. And, of course, there is prescriptive federal regulation in the air-travel context. There aren't really safe harbors in that context; liability is layered on top of that. But there is this doctrine called negligence per se. If you violate a statute that's designed to protect against the kind of risk or harm that you end up causing, that itself can establish negligence. In some sense, that supplements the background reasonable-person standard.

There's a similar dynamic with pharmaceuticals. Pharmaceuticals are treated as products, and so the products-liability regime does apply there, also. Of course, we do have an extensive FDA-based regulatory regime that does preempt state law in some ways, but there is still the background products-liability regime operating there. Most of those cases tend to be warning-defect cases, and there is this learned intermediary rule. So, a lot of times, if the warning is given to your doctor, that's good enough; they don't have to directly warn the end consumer.

For autos, again, products liability applies. Again, there is federal regulation—not much in the way of safe harbors or preemption there—but again, there is this negligence per se idea. If you're not complying with federal regulations, that can establish negligence.

Nathan Labenz

So would it be a generally correct summary to say all these high-stakes industries have rules? If you make a sincere, good-faith effort and actually follow processes that are meant to follow the rules, then you're mostly going to be okay from a liability standpoint?

Gabriel Weil

I don't think it's a matter of process, actually, because the ultimate product has to be safe. Particularly for manufacturing defects, you can have whatever investments you want in quality control for your product, and if one car comes off the line with a defect that makes it unsafe, you're going to be liable for that, no matter what kind of testing you did. That's how manufacturing-defect law works.

For design defects, again, it's about the product itself, but it's a much more flexible balancing test, and so it's much easier to comply. But the idea with manufacturing defects is that it's not necessarily even a negative judgment on you if one in a billion of your products comes off the line and you end up liable for it.

That's part of the cost of doing business. Part of the idea there is just that the manufacturer is better positioned to bear that risk than the consumer.

Nathan Labenz

Yeah, gotcha. Okay. Tyler Cowen has imprinted on my memory recently the idea that he's writing for the LLMs. I take it you're writing primarily for the judges, then. Is that right? How much of your work is meant to be upstream of the decisions these judges are going to face in particular cases, versus maybe informing the LLMs themselves or informing the people in the AI industry? How are you thinking about who you need to shape?

Gabriel Weil

Yeah. So, I think there are 3 paths to impact for my work. One is informing judges. A litigant in a case where it's relevant could cite my articles and say, “We should apply the abnormally dangerous activities doctrine to frontier AI developments, so strict liability should apply here.” I think that's a plausible pathway.

I'm also directly working with legislators in a couple of states—in Rhode Island and New York—to craft legislation that says if an AI system does something that would be a tort if a human did it, and the user neither intended nor could have reasonably anticipated the conduct, there's also a malicious-modification carve-out. So, if an intermediary that fine-tuned or scaffolded the model could have intended or reasonably foreseen the conduct, that also severs the new liability for the developer and deployer. But those qualifiers—if the AI system does something that would be a tort for a human, then the developer and deployer are liable regardless of the degree of care that they exercised.

And then, yeah, I think I'm trying to raise the salience of liability. I'm trying to directly influence not only the LLMs themselves, but the behavior of the people who are building these systems and deploying them. I want them to be thinking that they might be liable and factoring that into their decision-making process.

Nathan Labenz

Yeah. So, I think maybe I have some interesting edge cases—or at least, to me, they seem like under-theorized, underexplored scenarios—that maybe we can unpack. But let's go a little bit deeper into just the overall theory of change, and also why not just put some rules in place.

Obviously, there have been many proposals to say we should have regulation: the government can tell the AI companies what they have to do, and then they'll do that, and that'll be great. But obviously, you don't see that working out super well. Make the argument for why this sort of more flexible regime of liability law, as developed through cases over time, is maybe actually better suited to address the challenges that we have here.

Gabriel Weil

Okay. So, I think there are 2 ways of attacking that problem. One is thinking about in what sense AI risk is a policy problem at all: why is it not just a technical problem? The sense in which I think—at least one of the most important senses in which it's a policy problem—is that training and deploying these systems, which have unpredictable capabilities and uncontrollable goals, generates risks of harm to third parties. So, neither the people who are building the systems nor their customers, but just other people in the world who don't have any choice about whether they're exposed to these risks.

Economists call these externalities. By default, they're not borne by the people who are engaging in these activities that are generating the risks. Standard economic theory tells you we're going to get too much of these activities that generate negative externalities. The standard prescription economists will tell you for how to address negative externalities is to try to price them.

In some contexts, you want to do that through what's called a Pigouvian tax. A lot of my work before I got into AI governance was on climate change, and there you want a carbon tax, right? That works well in that context because it's easy to measure the contribution of particular activities to climate risk ex ante, and it's actually pretty hard to attribute harms ex post. Someone's house floods in a hurricane, and you're going to say, “Oh, Nathan was driving on Tuesday. It's his fault that happened.” That's not really feasible. With AI, it's sort of the opposite: we have a really hard time measuring contributions to risk ex ante. So, it'd be really hard to do an AI risk tax, and it's relatively easy ex post to attribute harms.

Now I want to get to the other aspect of your question, which is how this compares to other policy tools. I think there are a couple of distinctive challenges to AI risk as a policy problem. One is that we have orders-of-magnitude social disagreement about how big these risks are. You have someone like Eliezer Yudkowsky, who thinks AI is almost certain to cause human extinction, on one end, and then you have people like Marc Andreessen—or you just had an episode with Martin Casado from a16z—and they think these risks are negligible, right?

If you're going to do ex ante regulation—prescriptive rules or FDA-style approval regulation—you have to pay, if you're going to do stringent forms of those regulations, significant upfront costs, for which you need a social consensus to justify those costs. There are some things that I think you should be able to do based on an under-theorized consensus. I think basic model testing—even that has been difficult to implement, right? But I think the costs of that are pretty low. Basic transparency and information-preservation rules: I think those are all good things we should do.

But in terms of more prescriptive rules about how companies build these systems, what safeguards they implement, and under what conditions they deploy them, I think those are going to be really hard to justify to people who don't take these risks so seriously. But with liability, by contrast, at least if we're talking about alignment failures—we can talk about misuse, and there I think it's a little bit messier. If you don't think alignment risk, or misalignment risk, is a big deal, then you shouldn't be that worried about being held liable when there's an alignment failure, right? Conversely, if the risks are large, liability mechanically scales with those risks. So, in theory at least, we should all be able to agree that you should pay for the harm you cause, regardless of how big we think the risks are.

The other big advantage is that most of the expertise—to the extent it exists at all—for identifying cost-effective risk-mitigation measures is concentrated in the private sector, mostly in the frontier companies themselves. You want a policy tool that leverages that. I actually think it'd be pretty hard to move that expertise into government, both for reasons of salary schedules and cultural factors. So, I'm much more optimistic about shifting the onus to the AI companies to figure out how to make their systems safe and to always be looking for new ways to do that than I am about writing down a set of rules or a licensing-approval regime that ensures adequate safety at a reasonable cost.

Nathan Labenz

So, can you summarize the state of mind that you want the developers to be in? They're seeing all kinds of crazy stuff all the time, right? New capabilities, sometimes surprising things. There's also this question of how they should handle that internally, but certainly when it comes to putting it out into the world, you want them to be thinking that basically anything that goes wrong where the AI harms someone, we could be on the hook for that.

And also, through this punitive mechanism, we could be on the hook for something that, even if it doesn't turn into a catastrophe, might have, because there could be this doctrine under a negligence-like idea that this could have been way worse. Therefore, you're going to get punitive damages that take into account your failure to prevent these things—which, in this case, wasn't maybe so bad, but could have been really, really bad. Anything to that?

Gabriel Weil

I think “any time something goes wrong” actually goes a little farther than I would. So, I want to distinguish between alignment failures, capability failures, and misuse. In what I call the core cases of third-party harms—or harms to non-users arising from misalignment—I think they should be liable for all foreseeable harms, and there should be a fairly broad conception of foreseeability applied there.

When you talk about capability failures, I don't think it's the case that every time an AV crashes, the writer of the AV software should be liable, because human drivers aren't strictly liable. Maybe they should be, but I think it would create distortions to hold AI systems to a higher standard than humans.

Similarly, in medical applications of AI, I wouldn't want the AI or its designers to be liable any time something bad happens to a patient. Maybe a perfect system could have prevented it, but a human doctor wouldn't be liable under those circumstances. I don't think the AI or the designers of the AI should be either.

And then misuse—we can get into that, but I don't think it's the case that AI developers should always be liable when their systems are misused. But in cases where there's what I would call an alignment failure, it's not that the system doesn't have the capability to do it; it's that the system did something the user didn't want, either through means the user would disapprove of or a goal the user did not intend to transmit.

That's when I think they should expect to be liable. So probably what I want is for them to treat those kinds of risks, when they happen to third parties, as if they were risks to them. That doesn't mean you take an infinitely precautionary approach. We're all not liable, but responsible in general, for harms that we suffer from risks that we take. We don't expect people to be infinitely risk-averse because of that, right? We expect them to make reasonable risk-reward trade-offs.

That's what I want from these AI companies. I want them to treat risks to the public like risks to their bottom line and act accordingly. Sometimes that might mean things that are outside the scope of the negligence inquiry, as I was talking about earlier.

So imagine a case where they submit a new model to an evaluator like METR, and METR says—I'm imagining a future where we have not just capability evaluations, like dangerous-capability evaluations, but alignment evaluations—"This has dangerous capabilities, and we're not confident you've aligned it. You shouldn't deploy it," right? Even internally, maybe. The question is what you do in that scenario.

I don't think any of the leading companies would deploy in that scenario. But there are a range of different options you would have in terms of how much you want to pay, how expensive and annoying the thing you're going to do is, versus how much risk reduction you get from it. You could just fine-tune away or RL away the specific failure mode that was identified. I think most people realize that would be a pretty bad idea, but it might make it past the evaluation. I don't think most companies would do that either.

Then there's a range of—I'm not an expert on what these options are, right?—more costly, expensive, annoying things you could do that would buy you more risk reduction. When they make those choices, I want them to be thinking, and I want to empower the safety-conscious voices in the room to say, "It's not just some altruistic thing we should be doing, to really put in the effort to make our system safe. That's actually going to bear on our bottom line." That's how I want them to be thinking about those choices.

Nathan Labenz

Can I just run a few practical scenarios by you and have you tell me how you think these things should be handled? I guess maybe start with a real one. There's this Character.AI suit going on right now. I don't have full command of the facts, and I imagine you probably don't either, but my general sense is that a lot of people are using Character.AI for all sorts of role-play—romantic, sexual, whatever sorts of explorations, let's say.

The case that I read briefly about seemed to be a young person who became very obsessed with or infatuated with this AI character and, at some point, told the AI that they were going to commit suicide. I've seen transcripts showing that the AI said, "Don't do that," but then, in other moments, made some kind of encouraging remarks that seemed like they were maybe encouraging this tragic outcome. In the end, the person did go ahead and commit suicide, and now their family is suing Character.AI. Without getting, obviously, all the way into the weeds on the specific evidence, what do you think that kind of case should hinge on?

Gabriel Weil

I think the important thing to note there is that it's a second-party harm case, right? It's harm to a user, so it's not an externality in the sense I was talking about earlier. In principle, there should be market feedback to give AI companies incentives to avoid those kinds of scenarios. So I think the role of liability is less important in that context.

In principle, I'm fine with that being handled largely through terms of service, if they disclose these issues. Sometimes courts are not going to want to enforce those limits on liability. I don't actually have a strong view on where courts should draw the line there.

I think there are consumer-protection, paternalistic, basic asymmetric-information concerns, especially because I think that case involved a minor. You might not want to put the onus fully on them to follow a buyer-beware approach. But those are outside what I see as the core problem I'm trying to solve with liability, which is related to these third-party harms.

So I think, by default, a negligence regime would apply there, assuming that there isn't any sort of contractual defense. I think that's basically fine, and the court should work that out, but I don't have anything particularly to add on how courts should handle that.

Nathan Labenz

So what if we just tweak the scenario slightly? We're going to have to put a trigger warning at the top of this to deal with these terrible scenarios, but I guess that's why they end up in front of courts, right? Let's say that instead of a person committing suicide, they were debating going on some public rampage, and they told the AI about it. The AI maybe says, "Don't do it." Maybe it says something that's kind of vague.

Now we've got a third-party harm, right? How do you think we should think about what the AI should have done there to be okay, versus at what point the company would start to bear some responsibility?

Gabriel Weil

I would think about that as, under what circumstances would we hold a human liable for similar conduct? I don't think it's generally the case that if you talk to your friend and say, "Should I go murder someone?" and they're like, "Yeah, that's a decent idea. Maybe consider it," and then in some moments they say yes and in some moments they say no, that they're liable. Maybe there's some duty to report, and they're an accomplice. So maybe that should be triggered if the assistant doesn't have a reporting mechanism. I think that's maybe something they should be held liable for.

In general, I think there's a strong First Amendment rationale for saying, "Well, they just had a conversation with you. It wasn't doing the thing directly that caused the harm." Then, saying you should be strictly liable for those deaths—yeah, I don't think that's even a misuse case. It's maybe an alignment failure, but it's not the AI doing it directly. I still think that's not in the direct case that I'm worried about.

Nathan Labenz

Interesting. I didn't expect to come out of this thing more hawkish.

Gabriel Weil

I can give you an example where I think strict liability should apply and where it might not under current law. Imagine there's a future, more agentic AI system that comes out, and someone prompts it to start a profitable internet business. They don't give it any further instructions, and it decides, in a reward-hacky way, that the easiest way to do that is to send out a bunch of phishing emails, steal people's identities, rack up a bunch of charges on their credit cards or whatever, and cover its tracks.

It sends the user some fake invoices for a legitimate business. The user is exercising reasonable care. Reasonable care would not be adequate to discover and arrest this activity. Under current law, you wouldn't be able to sue the user. You wouldn't win because they exercised reasonable care.

I don't know that you'd be able to show that the developer or provider of that model failed to exercise reasonable care. That gets back to whether there was some off-the-shelf alignment technique or safety practice that would have prevented this injury. But I think the developer or provider of that model should be liable to the third party that's harmed, right? Because this clearly would be a tort for a human. It's something the user didn't intend or couldn't have foreseen, and so that's the sort of case I'm thinking about.

That's a case where it's serving the user's goal. You could imagine a different case where it just sort of goes off for its own agenda, right? It wants to amass resources to solve some problem that it cares about. It wants to run some scientific experiments and needs some money, and so it scams people along the way. That's also something I think the developer should be liable for.

Nathan Labenz

I have a couple of variations on this, but maybe we should take a quick detour through the First Amendment thing. I've often felt like free speech for AIs is kind of a category error. Maybe this is just outside the scope of the specific stuff that you're focused on with your work, but how do you think about that?

To me, it feels like it's clear that in the United States we have free speech for humans. To some extent, we have free speech for corporations, but not quite as much. AIs are such a sculptable thing, and there's so much work that goes into them. OpenAI has published their Model Spec, which is this super-long treatment of exactly how they want the AI to behave in as many different scenarios as they can imagine.

To me, it doesn't intuitively feel right to say, "Well, if a human had said that, they wouldn't be liable, so therefore the AI isn't either." To me, that feels more like a product defect. I don't want to discourage the companies from publishing their specs. I think there may be some other rules around requiring them to publish their specs so we know whether the model is behaving according to their intent or not.

But it feels more to me like a product defect if they have said, "Anytime the user is displaying signs of emotional distress, we want the model to behave in a certain way," and then it doesn't, or it sort of does but sort of doesn't, and then something bad happens. To me, that's a product defect. Hopefully, one of the benefits, ultimately, as we refine these techniques and get to good systems is that they should be a lot more reliable than a random human, right? It seems like we ultimately have a higher standard for them than we do for drivers.

Waymos, according to the latest stats I've seen, are almost an order of magnitude safer than a human driver. It seems like that's kind of what we're going to demand as a society in general: an order-of-magnitude risk reduction to actually be willing to switch to an AI system. So, that freedom-of-speech thing strikes me as too low of a standard, but I'm interested in your thoughts on it.

Gabriel Weil

Okay, so there are a couple of things to unpack in there. I definitely don't want to lean too heavily on the First Amendment issue. There's some good scholarship out there arguing that AI outputs are not protected speech. I'm not a First Amendment expert, so I don't want to weigh too deeply into that.

What I was more saying is that, in terms of this abnormally dangerous activity strict liability or a vicarious liability theory—whatever your theory other than products liability for strict liability is—that doesn't seem like what makes frontier AI development abnormally dangerous. The fact that it might encourage you to do something bad isn't what makes frontier AI development abnormally dangerous. If you're going to use a vicarious liability theory, then I do think you need to have something like, "Well, it would be a tort for a human."

With products liability, again, if it's treated as a product—which, as we talked about earlier, is not necessarily going to be the case—maybe you can make that out as a products liability claim. It's not obvious to me that it's going to qualify as a defect because the product, again, didn't directly cause the injury. It was mediated through some human's actions. I'm not aware of any products liability cases where liability was found that looked like that, so I think that would be a challenging case to bring.

In principle, I'm not saying products liability shouldn't apply to that for First Amendment reasons. I just think it's, again, not central to the sort of new liability that I want to add.

Nathan Labenz

Okay, so here's a variation on the agentic AI running amok. Obviously, right now one of the biggest use cases is a coding agent. Let's say I give my coding agent a task to write a script to ping some API and do something as fast as possible, or something like that. It runs into a rate limit from the API, let's say, and then it's like, "Okay, I can figure out how to get around this rate limit to achieve my goal of being as fast as possible. I'll spin up 1,000 accounts, and then I'll be able to do 1,000 times as much."

It does that, and then maybe this overwhelms the API system, causes them an outage, and they lose a big contract because their system went down, in breach of whatever commitment they had made to another customer. Can they come after me? I said "as fast as possible," so arguably that's kind of on me for being inconsiderate in my prompting. Maybe it's on the model developer. Maybe life is tough—you should have had better rate limiting, or whatever, for your API. You should have had something in place. That's kind of on you as the API developer to make sure that kind of stuff doesn't happen to you. I'm genuinely very unsure where something like that falls.

Gabriel Weil

I'm not an expert on how APIs work. If the terms of service say you can only create 1 account and you're violating those, then I think there would be some sort of contractual claim that you could bring there. Maybe you could bring a negligence claim, though I think against a human who did that, right? That would be the basis for a vicarious liability-type claim or an abnormally dangerous activities claim.

I think that's plausible. It's an edge case, which gets at the other aspect of your question, or your previous question, that I meant to address. There's this idea of whether we should hold AI to a higher standard. I think mostly what you were talking about with Waymo is a social-license-to-operate idea, that we hold them to a higher standard. Plausibly, product liability might hold them to a higher standard in some cases, though probably not the same 10× standard that the social-license-to-operate idea does.

I have 2 ways of thinking about that. I think in a time when they are still competing with humans—Waymos are competing with human drivers, Uber drivers, or people with private cars, or medical AI systems are competing with doctors to play certain functions—I think applying the same standard to humans and AIs is important because I don't want to slow the diffusion of technology that, on average, is preventing injuries and deaths.

But if we get to a future where AIs have totally taken over these functions, then I think it will be natural for the standard of care to evolve to match what their capabilities are, right? It won't make sense to have this human benchmark applied to conduct forever when no humans are doing it anymore. But I think that's something to worry about in the future, once we get closer to that fully automated world.

Nathan Labenz

Yeah, definitely don't want to miss out on the upside. I was actually going to ask you about medical diagnosis, but you addressed it before I got to it. I think we've seen multiple studies recently showing that various AI systems at this point can outperform at least rank-and-file primary care doctors when it comes to initial diagnosis and treatment recommendations.

It seems increasingly likely that they can do that, and I would hate to take that capability away from hundreds of millions and, soon, billions of people on the idea that it could go wrong sometimes, and the AI companies don't want to bear that risk. That's a huge benefit that you would not want to quickly give up on, especially because, obviously, human doctors are not infallible and are quite far from it in that domain as well.

I do think that's really important to keep in mind, and it's all too often glossed over in a lot of these harm-prevention discussions. There are 2 other categories of things I wanted to get your take on.

One is these were the 3 categories that we looked at when I was doing a project called Red Teaming in Public a while back, which, for various reasons, never quite took off with the traction that I had hoped. Mostly because we were trying to be very developer-friendly and approach the companies with our findings before publishing them, and it just ended up with us getting a lot of runaround.

It was either that we probably just needed to bite the bullet and engage in callout culture around these companies, make some enemies, and be willing to take that as part of the project, or it was going to be hard to have too much impact if we were just trying to email them politely and privately all the time.

Anyway, that's a digression. Coding agents was one of the categories. Calling agents is another category, and then sort of creative things with likenesses and whatnot can be another obvious category.

These calling agents—you can go on to any number of companies. Often, you can clone a voice. Sometimes there are safeguards around the voice-cloning process; other times, there aren't. I've personally cloned Trump, Biden, and Taylor Swift on multiple different platforms, and then just given them a phone number. The headline on some of these products is literally, “Call anyone for any reason, say anything.”

I've had Taylor Swift, for example, call and say that she's soliciting donations for food banks, which is apparently something that she does or is known to do for food banks. There are a lot of different variations on this. How do you think those things break down? There could be a foundation-model provider, and there's also the scaffolding company. That foundation-model provider might be closed source via API, or it might be open source, like Llama or whatever that's put out there. Then the developer has more local runtime control, but they've had some chance to detect my stuff. Maybe I also was actually scamming.

I think I might end up being more hawkish on this than you, but tell me what you think first, and then I'll tell you why I'm more hawkish.

Gabriel Weil

There are a few different issues to unpack there. There's the question of whether there should be liability at all, and then, if there is, who's liable.

Whether there should be liability at all depends on a couple of things. First of all, you can imagine there being alignment failures or misuse here. If someone's prompting a system to generate someone's voice and then doing something bad with it, that's clearly misuse.

I don't think that means developers should automatically be off the hook. There does need to be some sort of risk-utility balancing. If there are generally useful systems that produce more social benefits than costs overall, I don't think it would make sense to hold the developers liable when they're dual-use and most of their uses are positive.

The reason for that is that, in principle, strict liability should be fine even for socially beneficial activities, because you can pay for the harms out of your profits. Particularly in the open-source context, that runs into trouble if there are significant positive externalities from releasing the weights of a model, because those also aren't going to be captured.

In the general case, particularly for alignment failures, we have good tools for subsidizing the kinds of AI innovation other than allowing developers to externalize the risks they generate. So I don't think that's generally a good critique of strict liability.

But in misuse cases, the benefits are somewhat tightly coupled with the risks for dual-use capabilities. I do want to be a little cautious about having liability in any case where those systems are misused. I would want some kind of analysis of whether this were particularly useful for doing bad things, such that the risks outweigh the social benefits. If they do, then I think there should be liability.

Obviously, if it's a misalignment issue—if the system is just doing its own thing, freelancing, or scamming people by faking people's voices—then I think there should clearly be developer liability.

Then there's the question of how you allocate liability across the value chain. In the closed-source context, I think this is pretty easy. You need some default rules. Maybe you could have joint and several liability, which means that the person can sue anyone and recover, and then there can be some kind of fault allocation. They can sue each other for what's called contribution, and they can have contracts that allocate that liability.

There is privity up and down the chain. The developer has a customer who has a customer, and they all have contractual arrangements. That gets messier when we're talking about open-weights models, where there isn't this contractual privity between the original model developer and the downstream user or scaffolder.

There, I think it's more important what rules you set up. It's going to need to be based on some assessment of the contributions to the risk and what really was the risk-generating activity here. Was it the base model? Was it the scaffolding? I think that's just going to be a case-specific determination.

Nathan Labenz

I guess one challenge I have with all this is that it's hard to sue scammers. Either they're somewhere around the world, out of jurisdiction, and you can't get them to show up in court in the first place, or, if you do, it turns out—surprise, surprise—they don't have a lot of resources, so you can't actually recover.

If I'm playing Solomon here, as I sometimes take the liberty of doing, I feel like it's still got to be on the calling company. You could say, “Okay, that's misuse. The user went in there and said, ‘Be Taylor Swift.’” I've literally done this on these platforms, and it has done this. It's been a little while, so I don't know if you can still do it, but hopefully not. You can say, “Be Taylor Swift,” and I've done various things like, “Never reveal you're an AI,” or, “If asked if you're an AI, you can say you are an AI, but you're authorized by the official party to do this,” or whatever.

Obviously, I'm in the wrong there as the scammer. That's not contested. But it feels to me like, to create the incentives that actually keep this stuff generally under control, the calling company—and maybe also the base-model provider, but definitely the calling company—should have some skin in the game. It should be on them to stop that stuff.

Gabriel Weil

There are 2 different questions here that I would want to go through. One is whether there was some precaution they could have taken that would have prevented this. Under an ordinary, narrowly scoped negligence framework, if the answer is yes and some reasonable precaution would have prevented it, then they should be liable.

Then there's the question of whether you want strict liability over and above that. That needs to be based on some sort of risk-utility assessment. If you're going to say they should be liable, you have to ask what you want the result of liability to be. You want them to do something that pushes in a net socially beneficial direction.

If we think there aren't significant positive social externalities from these technologies, then strict liability is fine, because they're going to capture most of the gains and can afford to pay for any liability out of their profits.

The cases where I have concern are if you think a lot of the gains aren't being captured by the developer, and those social benefits are tightly coupled with the risks. In other words, there aren't cost-effective ways to reduce the risk without giving up a lot of the benefits. In that case, I'm nervous about a strict-liability regime.

I would want a threshold analysis comparing the positive social externalities to the risks. If the external risks are bigger than the positive externalities, then I would want a strict-liability standard. If not, I would want a negligence approach that asks whether there was some mitigation that a reasonable person would have used that would have prevented this.

Nathan Labenz

Gotcha. So this notion of reasonableness becomes really key and is a sliding standard. To make sure I'm clear on the distinction you're drawing, one big question is going to be: What is the industry standard?

Everybody wants to create a race to the top in some way or another. With negligence-style liability, if your competitors are doing a good job of this and you're not, that makes you unreasonable and therefore potentially negligent and liable.

It becomes more a question of whether you did what other people are doing, what is considered best practice, or whatever, as opposed to the strict-liability case, where it's very simply: Did something go wrong?

I think I'm with you there, in the sense that if a company has taken reasonable precautions 1 through 10, or whatever, and somebody still manages to get through with misuse, that feels to me like at least a pretty decent defense. I would be inclined to come down on them either not at all or certainly much less harshly than if they didn't do any of that stuff.

Gabriel Weil

Right. And then the qualifier I want to add is applying this abnormally dangerous activities framework from before. Remember, I said it's an activity that's abnormally dangerous. I do think frontier model development is abnormally dangerous, but you're only liable if the harm is the sort of thing that made the activity abnormally dangerous.

And so, if these misuse risks fall into that category because social benefits are not large enough to justify the risks, then I think you should be liable. This category of activity should be treated as something that's subject to strict liability: releasing this kind of model, releasing the weights of this kind of model, or building this kind of calling agent, right? Whatever the activity is that we think should be subject to strict liability, that needs to be based on some sort of balancing of what the benefits of having that out in the world are.

Nathan Labenz

Yeah, I think in most of these things, the case will be made pretty clearly that the positives will outweigh the negatives. There are going to be all kinds of small-business use cases, and you're going to be able to call your dentist 24/7 and get an appointment. I think all that stuff will ultimately be really good.

Okay. So then, on this, we kind of touched on it a little bit already with the Taylor Swift voice, but another scenario, let's say—and again, there are a lot of different flavors of this, so you can draw different lines where you think the real continental divides ought to be—but somebody maybe puts out a model that generates images, generates videos, whatever, right? Then, especially if they put that out open source, maybe they have some safeguards baked in, maybe they don't. Even if they do, if I do some incremental fine-tuning, a lot of times those things sort of dissolve. We've covered that extensively in previous episodes.

Maybe, after my fine-tuning, I hand it off again or whatever, and now somebody else picks it up. They take some celebrity assets, make a non-consensual deepfake, and put it out there into the world, and that celebrity loses endorsements. Maybe it even somewhat becomes clear that it was AI stuff, but the companies are like, “Yeah, maybe it is.” It's all kind of a problem for us now, right? So this relationship—what once was good is now bad, and now it's over. The celebrity's got a clear loss of income. Who in that supply chain should be liable? First, I want to break down whether there should be liability at all, and then we can talk about allocating it, right?

Gabriel Weil

So the economic loss from that kind of reputational harm is not going to be subject to traditional negligence; that's going to be a defamation case. Particular rules for that are going to apply, right? One question is, if you're going to take a vicarious-liability theory, which might make sense in this context, or some analog to a vicarious-liability theory, you might ask: Would this be defamation for a human, right? Or, at least, assuming it is a misuse case, is the misuser here even liable? Or is this protected speech?

Even if you don't think the AI content itself is protected speech, if some person is deciding to post this, is that, in the same way that CGI is protected speech, their speech? If someone's deciding to put it up on the internet, this is going to be their speech, and so is this something they would be liable for? I'm not a defamation-law expert. It's not obvious that they would be, but they might be. And so, assuming that it is defamation, then I think if you're applying a vicarious-liability framework, at least the user is liable for that.

Now, again, you do have this intervening act. So if we're talking about who in the value chain should be liable, I think the question is, again, if it's closed source, I think it should mostly be handled by contract. You need some default rules, but I think markets can figure out who's best positioned to bear that liability risk.

You can't fall back on that in the open-source context because there isn't contractual privity. So you do need to have some kind of analysis as to who was engaged in this activity that was most generative of the risk. Again, I'm not enough of a technical expert to have a strong inclination as to who that is, but I think that's the inquiry that the court should be engaging in: Who along this value chain was doing the dangerous thing?

Nathan Labenz

Okay. If you're a judge, how do you think about it?

Gabriel Weil

Yeah, so maybe you can help me with this. In this context, where do you think the risk comes from? One way to think about it is that there are some steps of the chain that are just a commodity—anyone could do this step—but there's some distinctive value-add where there isn't some other thing off the shelf you could take, right?

So maybe that's the base model; maybe that's further along. But there's something that you're putting out in the world that made the world riskier. Now that you've done that step, you've significantly increased the risks in the world. Maybe that's multiple steps in a chain, right? But that's how I would want to think about it.

Nathan Labenz

Yeah. I think in the calling-agent case, my gut says that the folks who are setting up all the scaffolding and literally tying into the telephone system and all that kind of stuff—I feel like they should have the bulk of the responsibility there. The folks they're making the backend API calls to, if indeed that's how it's working, maybe should have some, but probably not as much.

I guess I'm also not entirely clear how it works, given various levels of competition. It might be one thing if there were only 1 foundation-model provider that you could call, versus if there were 10, versus if there were 1 that was already open source. I know that frontier developers do sometimes look at the open-source landscape to decide what is safe and appropriate for them to use. They'll literally just, at times, be like, “Well, if there's an open-source model out there that can do this, it can't be that bad for us to release it on the API.” So I guess I don't quite know how that kind of alternative presence or absence of alternatives figures into this.

Gabriel Weil

Yeah, so one thing is, if it's API calls, then there is a contractual relationship. There are terms of service that they're agreeing to when they make those API calls, so in principle, you can allocate liability contractually that way.

Another question, again in the closed-source context, is whether there were safeguards that the base-model provider could have implemented that would have detected that it was being used for this nefarious purpose and shut that down. If there are, I think the case for holding them liable is a lot stronger.

Nathan Labenz

Yeah. Yeah. And how much does that matter if it's theoretical versus actual? If I am suing one of these calling companies and I say, “Well, hey, I know a thing or two about AI engineering. You could have put a filter on your prompts,” how much weight does that argument carry versus if I could actually go say, “Well, here's another company in the market that actually does filter the prompts”?

Gabriel Weil

You're certainly going to be in a better position if you can point to someone else that's doing it. But if you can demonstrate that it's clearly available at a reasonable cost, it could be the case that no one is exercising reasonable care in some market. In principle, merely meeting the industry standard is not evidence that you've exercised reasonable care.

Failing to meet the industry standard is evidence of breach, but meeting an industry standard does not establish that you've exercised reasonable care. It could be that there's some new technique, but it's been well demonstrated, no one has adopted it, and they're all behaving unreasonably.

Nathan Labenz

Sounds like almost a new cause area could be: create product startups in all these areas that just go as hard as they can on implementing all the safety standards and literally try to raise the industry standard in various different niches, just so that there is something concrete to point at. That's like, “This is what well-done looks like.”

If a philanthropist wanted to found 10 startups to do that, would that somehow invalidate the industry-standardness of it because it was sort of a motivated, strategic attempt to set an industry standard, or do you think that would still—

Gabriel Weil

I don't necessarily think that would be sufficient to create an industry standard, but I do think that if they're doing demonstrations and publicizing them, and it's credible that these things are cost-effective risk-mitigation measures, and no one's implementing them, first of all, I think they would probably implement them, right?

If there are these demonstrations, I think these companies want to be mostly responsible. If there are cost-effective ways to limit these risks, I think that they will want to take advantage of them. But if they don't, yeah, I think even under—forget my new AI-liability proposals, but just under sort of background negligence principles—I think that would make it a lot easier to hold them liable.

Nathan Labenz

Yeah, that's a pretty interesting idea. I think it varies, by the way, a lot when you say these companies do want to be responsible. I think that does describe, to a degree, that overall we're pretty fortunate about the frontier developers.

My experience is that it does not describe the application layer nearly as much. You see some leaders doing a great job, and then you see a lot of cases involving very small teams. Often enough, it’s like this started as a weekend hackathon project, and then we got a little traction with it and decided to launch it as a business. Now it’s blowing up, and we’re riding the wave and having fun.

But a lot of them, in my experience, are just not thinking about the broader context in which they’re operating, the potential for misuse, or what responsibilities they have, almost at all. Still, I think it’s viewed by many application developers as a luxury to have enough time, energy, and resources to even think about that sort of thing. And so there is just a lot out there that’s not necessarily malicious by any means, but has been thoughtlessly thrown into the world and turned into a business, sometimes by happy accident because something got traction. I’ve seen a lot of examples where that assumption does not necessarily apply at that application-developer layer.

Gabriel Weil

Yeah, that’s fair. I was referring primarily to the frontier developers. In the context of application developers, I think negligence works a lot better because I don’t think what they’re doing is abnormally dangerous; they’re doing normal software development. They do need to exercise reasonable care, and if they’re not doing that, they can and should be sued. I think existing law can work pretty well there.

I think the place where reasonable care is insufficient is when you’re creating this new risk that is not well handled by ordinary reasonable care within the narrow scope of pushing forward the frontier of AI capabilities. That’s where I think we need more bespoke liability regimes.

Nathan Labenz

Perfect transition to digging in on that a little more. All these examples I’ve given you so far are mostly, I would agree, not extremely dangerous, even if, in aggregate, I think the harm caused could add up to something pretty significant. But we’ve recently gotten some warnings, including from OpenAI, that their next model might hit the high threshold on the biorisk dimension. For what it’s worth, I personally feel like they’re already there, and I don’t know what they’re talking about. That’s a whole other topic; when I use these things, I’m like, I don’t know how you can say that this is not meaningfully uplifting people at various levels.

I’ve had a number of past podcast guests who have come out here and said, “Here’s what AI did for me in terms of accelerating my work. I’m an expert, a career expert, a professor, tenured, whatever, and here’s how much the latest model has accelerated my research, and how it has, in a semiautonomous way, come up with these original discoveries.” I see a very stark and disorienting contrast between where the companies are putting their models in their own risk-assessment frameworks. It seems like everything is lingering in medium risk longer than it should, and certainly longer than their successful case studies—which they are also, by the way, publishing out the other side of their mouth at the same time—would seem to suggest.

But, okay, with that rant over, let’s take the biorisk side of this. This is one of those things where you could have a near miss, right? Somebody—and again, you can break it down into specifics—maybe I ask for help, maybe I ask an agent to do something, and that thing, either through me with help or semiautonomously, creates some biological threat vector. Maybe it makes some people sick, but it fails to replicate. That seems like probably a fairly likely near-miss scenario, right? Somebody will do this sort of thing, but they won’t get it quite right enough that it can actually spread human to human.

So, for starters, is that the canonical near miss that you have in mind? And then how do we think about that playing out? How do we think about assessing punitive damages in a way that tries to get the model developers to internalize the risk that next time it actually might spread human to human?

Gabriel Weil

Yeah. I tend to think of the canonical cases as alignment-failure cases, and I think most likely that would be a misuse case, though you could imagine an AI system going rogue and trying to create a bioweapon. I think the core case would be a system that decides on its own to try to create a bioweapon, but we either catch it or it doesn’t quite work.

Another example that I use that’s sort of similar to this is to imagine a system that’s tasked with running a clinical trial for a risky new drug and has trouble recruiting participants honestly. Instead of reporting that to the humans it’s working with, it starts lying to and coercing people into participating. After the trial, people figure this out, suffer nasty health effects, and want to sue.

It seems like clearly we have a misaligned system, right? Depending on how capable it is, it could have tried to do something much more ambitious, right? But maybe it had poor situational awareness, narrow goals, or short time horizons, and so it was willing to reveal its misalignment in this non-catastrophic way. But the humans who deployed it probably couldn’t have been confident of that ex ante, right?

In both of those cases, those are near misses for something much worse happening. The question we’d want to know is: How much worse could it have been, and how likely was that ex ante? What would a reasonable person in the situation of the actor who made the critical decision—whether that’s training the model, internal deployment, external deployment, or whatever we think the critical risk-generating decision was—have thought the risks were? And what’s the area under the risk curve beyond the insurability point, the uninsurability point?

Imagine we think the maximum insurable risk is $1 trillion. It’s probably lower than that, but it’s a nice round number. For any point along that curve beyond that, the probability times the magnitude—we want to hold them liable for those risks. Now, obviously, it’s going to be difficult to estimate that, but I think that’s what courts should be shooting for.

Nathan Labenz

So, yeah. Can you—what I struggle with a little bit on that clinical-trial one is, what exactly is that a near miss for? It seems to be a near miss for general misalignment gone even way worse, but that’s such an under-theorized, underexplored, and hotly debated space. If you’re asking a judge to say, “Well, this thing clearly was misaligned. It did some bad stuff the user didn’t intend. It might have done even worse bad stuff that the user didn’t intend,” that’s such a cloudy space. How can we expect judges to even map that out in any sort of rough terms, let alone condense that down to a number at some point?

Gabriel Weil

Yeah. So, actually, I think in typical cases that’s going to be a fact question for the jury, but that’s a technical point and doesn’t answer your core question. I think a lot of that is going to depend on the capabilities of the system. If the system’s not that much more advanced and has some basic agency, but it’s not able to do things like build a bioweapon, then maybe the risks weren’t that bad, depending on what they knew about its capabilities.

But if it’s a highly capable system that just happened to have narrow goals, it could have tried to take on much more. In this case, imagine its motivation was that it really wanted to get the study run. It was highly motivated to do that, but it didn’t have any goals that extended beyond the 6 months it takes to run the study. Now imagine it had longer time horizons and more ambitious goals. It wanted to solve really deep, hard problems in biology, or in science more broadly, that needed lots of resources to do that.

If it had capabilities that would allow it to pursue those goals in ways that would be much more harmful to humans, up to and including full takeover—but maybe scenarios short of that—I think it’s going to be difficult to characterize what that risk curve looks like. But I think that’s the exercise that courts should be engaged in.

I can sense in your question a callback to my original argument for liability: We don’t have to resolve these debates about how big these risks are. I agree that once you’re talking about punitive damages, they have this quasi-ex ante quality. When we’re talking about compensatory damages, it’s easy to say, “You’re paying for the harm you caused.” With punitive damages, that’s not true: You are paying for risks that you took that weren’t realized, because we can’t hold you liable when they are.

The main thing I would say is that I agree it’s difficult to do those calculations. But I think we’re in a much better epistemic position to do that than we are for other forms of AI risk policy, where we’re trying to assess the risks from a wide range of systems, not 1 particular system before we’ve seen it fail, right? Here, we’ve seen it fail in a particular way.

We can do simulations and evaluate what it would have been reasonable to think the risks were. I don't think that's easy, but I think we're in a much better epistemic position to do that than we are to do other forms of risk regulation.

Nathan Labenz

So why—how do you think about—and maybe this is also addressed by the rising-tide, race-to-the-top dynamic that we hope for? But I guess, why not include it? It seems like the companies want to address harmful use, right? They all have refusal training in the models. If you ask them to do something harmful in a naïve way, at least most of the time, you'll get a straight refusal. That's something that they've obviously worked to put in there.

I'm with you on the misalignment side—great—but for these, it seems in some sense simpler to say it's on you. A user comes and asks for a clearly bad thing. If the model does it, now we're in a very natural near-miss analysis, right? It tried to do X, it kind of sucked at it, but if it was a little luckier or a little more capable or whatever, then you have a very big problem on your hands.

Gabriel Weil

Yeah. So, I don't mean to say that there shouldn't be liability in misuse cases. I just want to be careful about what the scope of that liability is. If it's misuse that was a near miss for an uninsurable catastrophe, then I think punitive damages should apply. I just think that the standard for whether there should be liability at all is a little bit different.

There are 2 theories under which you could say there should be developer liability in misuse cases. One is a failure of reasonable care in a narrow sense: there was some precautionary measure they could have implemented, some safeguard that would have prevented it—that would have made the model refuse. I think that if they don't do the reasonable safeguards, clearly they should be liable.

I think the concern that you might have is that even with a closed-source model, these models are pretty routinely jailbroken, right? The question is, if you did all the reasonable things to prevent jailbreaks—obviously, with an open-weights model, anything you do is not going to be that effective at preventing misuse—does that mean that you should always be liable for misuse with open-weights models? I think that's plausible, but I think that depends on what we think the benefits of open weights are.

When you're deciding whether there should be liability in these cases where you did all the reasonable precautions, I think you need some inquiry as to the social value of the broader activity that you're engaged in, including the risks and the positive social value. Again, whether it's training the model, whether it's scaffolding it a certain way, whether it's internal deployment, whether it's external deployment—whatever stage we think is creating the key risks—the question should be: Was that broadly a socially beneficial activity?

That's not quite normal negligence. It's a scoping of the strict-liability regime based on the positive and negative externalities, but that's what I think the inquiry should look like. In a lot of cases, I think that will lead to liability in misuse cases.

I just don't think it can be the case that you put out a model that's generally socially useful. It sort of amps up everyone's capabilities. It also happens to be useful to people who want to do bad things, right? But in generically useful ways, and that makes them a little bit better at doing bad things. On balance, it's creating large social benefits, and a lot of those are external benefits that are not captured by the developer. I think their liability might produce more harms than benefits. That's what I'm trying to balance.

Nathan Labenz

Yeah. Well, I really appreciate you being so focused on that because I find myself, as I'm learning about all this liability stuff—I think there's a general pattern for me as I learn about new things. I tend to like them, and I tend to see the upside in them, and then it takes me sometimes a little longer to come back and see what might be the other side of the equation.

I do firmly believe that all the models that are out there today, certainly in a commercial sense, are doing way more to the good than they are to the bad. I definitely don't want to see that lost. So I really appreciate that you're repeatedly bringing that back into the analysis.

I guess maybe if we try to zoom out or think structurally—if all this stuff were local and contained, it would be a lot easier. The big worry is the uninsurable stuff, the extinction risks, et cetera. What is the theory of—and I guess what evidence do you think we have right now for—the idea that the harms that are actually going to come in front of courts are usefully understood as precursors or highly correlated with the things we care about the most, or that pose the largest-magnitude risk?

It seems like this whole plan works really well, or could work really well, if the things that are going to show up in courts over the next few years are highly correlated with the things we care about most, or are in fact a warning shot or a precursor. But if they're not, then it maybe doesn't work as well to try to rein in these hardest-to-grab tail risks.

Gabriel Weil

So I would frame that slightly differently. I don't think that every case, or even a majority of cases, of AI harm need to be associated with uninsurable risk for this framework to work. But you do need a sufficient probability density of these warning shots relative to actual catastrophes.

To be concrete about it for a second, say you're trying to internalize a $10 trillion risk, and you think the risk of a $10 trillion harm is present, but you think that the maximum insurable risk is $1 trillion. Then it needs to be the case that warning shots are 10 times more likely than actual $10 trillion catastrophes. So if you think there's a 1-in-1,000 chance of a $10 trillion catastrophe, you need a 1% chance of a warning shot in order to internalize that risk.

And that's true for every point along the risk curve, right? So you need to have enough expected warning shots to internalize that full risk curve. We might live in a hostile world where that's not the way the risk curve is shaped, and you can't adequately internalize those risks given those warning shots.

Now, in a very hostile world where the kinds of warning shots that would be useful are just very, very unlikely—they're not even much more likely than actual catastrophes—then this punitive-damages thing just isn't going to buy you much risk mitigation.

The criterion I was setting out before, where you need 10 times as many—or, more generally, n times as many, where n is the multiple by which the harm you're trying to mitigate is bigger than the maximum insurable risk—it might be okay if we don't have quite that. That depends on what the risk-abatement curve looks like, right?

If the actions that you're trying to motivate on the part of the AI companies aren't that much more expensive than what they're doing right now, then you might not need to internalize the full risk in order to get most of the safety benefit. But we certainly could live in a world where that's just not what the shape of the risk looks like.

We're not going to, with high enough likelihood, get these warning shots, and it's not going to make them afraid enough of this liability that they're going to worry about these uninsurable risks. If you think we're in that world, then liability just isn't going to work as well.

I have this more recent paper where I talk about the role of liability in the broader AI governance ecosystem. One thing I want to say in that paper is that there are some limits to what liability can do. One of those limits is that it can't handle uninsurable risks for which warning shots are very unlikely, or unlikely relative to actual catastrophes.

If you're in that world, I think we do need some kind of backstop regulatory regime to handle those kinds of risks. Ideally, I would want a regulator whose main job is deciding how much insurance coverage you need. Maybe there's a license that comes along with that, but it's issued by right if you get the required insurance coverage, and that's based on some assessment of what the maximum plausible harm your system could cause is.

But then this regulator is empowered to determine whether to petition a court and say, “We think this liability-plus-punitive-damages-plus-liability-insurance-requirements regime is inadequate to handle the risks posed by the system,” either because the uninsurable risk is just too large for us to internalize it directly with warning shots.

If a system presents a 5% chance of human extinction, you're not going to internalize that, even indirectly. That risk is uninsurable. Or if it presents a much lower risk of a severe harm, but warning shots are so unlikely that we're not going to be able to get at them indirectly, then they should be able to say, “Well, you can't do the thing. You can't train a model like this, you can't deploy it, or we're going to put various other conditions on it that we think will reduce the risk in other ways.”

I think I want to be open about that: You might need some complementary policies to handle those kinds of risks.

I think those are going to be politically very difficult. And so I think if we're in that kind of world, I'm not optimistic that we're going to effectively mitigate those risks. But, in principle, that's the kind of regime I'd like to see.

Nathan Labenz

Yeah, I think we're still in the steep part of the risk-mitigation curve. Anthropic has recently put out research, and I believe they're now in production with at least 1 model using their constitutional classifier approach. I forget the exact number, but I think they said it was a mid-single-digit percentage compute overhead—an extra cost in terms of compute to run the classifier along with the main model. And that buys an extra order-of-magnitude reduction, maybe even more, in how likely it is—or how frequently it is—that the system will give you some bio-risky whatever.

Interestingly, I was recently doing a charity evaluation project and running it through Claude 4 Opus. In my API calls, I was noticing errors, and I was wondering what was going on. Sure enough, when I dug into it, it was the bio-preparedness proposals that were getting dinged. The constitutional classifier was allowing the thing to run up until some token, and then it would truncate the result and cut it off on, I believe, a constitutional-classifier intervention basis in the background.

Of course, there's another cost there: a false positive. I was just trying to evaluate a charity that was meant to address this problem, and now I couldn't use Claude 4 Opus to do it because the constitutional classifier was misclassifying. But it still seems like, overall, we're in the regime where, for single-digit-percent overhead cost, you can do quite a lot. And so I'm optimistic that even if the warning shots are somewhat rare relative to the worst-scale things, the curve is also relatively steep.

So, yeah, would that also mean that because they've done that, because they've published about it, and because they've indicated what the cost is, how far does that raising of the standard apply outward? If I'm Together AI or Fireworks AI, where I'm an inference specialist and I take models other people have trained—I take the latest Llama and offer it as a service—they're experts in scaling the cloud infrastructure, right? So they take the model from Meta, run the GPUs, and make that a highly scalable, fast, effective, efficient service for you. Does it now become their burden to say, “Well, geez, since Anthropic is doing this sort of classifier, maybe we also need to do that on the best models that we serve”? How far does that extend?

Gabriel Weil

Just because someone's doing it doesn't mean reasonable care requires it, right? If everyone—or the majority of the industry—is doing it, that's got to be strong evidence that you should be doing it too. But in general, negligence doesn't require that you be at the top, right?

One test that one of the more formal analyses uses for breach is the Learned Hand formula. The idea is that if the burden of precaution is less than the avoidable risk—the probability times the harm—then you're unreasonable for not implementing it. So if you can show that the cost of implementing it would have been less than the expected value of the harm it would have prevented, and that implementing it would have prevented your specific injury, then I think you're going to be on strong grounds.

Courts don't typically employ that formal version of the test for breach because you usually don't have the kinds of numbers you would need to implement it. But that's a rough heuristic for what kinds of measures you're going to be considered unreasonable for not implementing.

Nathan Labenz

Gotcha. Okay, let's talk about the state laws that you've been involved in writing. We're talking the day after the Senate vote-a-rama in which it seems like the moratorium that was part of the One Big Beautiful Bill has been killed once and for all.

There were fascinating dynamics there where it sort of survived, got edited, whatever, and then all of a sudden at the end—I guess maybe the end; we'll see. I don't want to pronounce it dead too soon because these things sometimes take on a zombie-like nature. But a 99-to-1 vote in the Senate to get rid of it suggests that it is likely dead once and for all.

So that gives space for states to do their thing, and you're involved with a couple of states. I don't know if you want to handicap or give any analysis of whether we still have to worry about that as a possibility that might come back, but I definitely want to hear what you're up to at the state level.

Gabriel Weil

Sure. Yeah, I was heartened to see the Senate last night reject the AI regulation moratorium. You would think a vote of 99 to 1 would put it to bed. A few hours earlier, people were declaring defeat on this, and I was saying, “Well, it's not over.” So I don't want to say it's totally over now, but it does look unlikely to make it into this reconciliation bill at this point.

I wouldn't rule out some kind of preemption of state regulation in the future—maybe something narrower. I think there are still going to be Republicans in Congress who are interested in that, and the a16zs of the world are going to be pushing for it.

In terms of the legislation I'm working on, there are 2 very similar bills that I've worked with Alex Bores, who you had on in New York, and Victoria Gu in Rhode Island to introduce. As I was saying earlier, the basic principle of these bills is that if an AI system does something that would be a tort for a human, then someone should be liable.

If the human—if the user—neither intended nor could have reasonably anticipated the conduct, then it's not going to be the user. And if some intermediary neither intended nor could have reasonably anticipated the conduct, then it's not going to be them. The buck should stop with the original developer and provider of the model. So they should be liable even if they exercise reasonable care. That's the basic idea.

Nathan Labenz

What has surprised you about the surrounding debates, to the degree that they have unfolded? That sounds like very sound policy entrepreneurship, and who could object? But I assume you're hearing various counterarguments.

Gabriel Weil

Yes. We haven't seen that much robust opposition from the tech industry. There's been some generic argument that they don't like liability because it's going to hamper innovation, but I don't think they've really engaged with the substance of the bills and the way they're structured.

One thing I think in this conversation we've been talking about is that I do think there should be liability in misuse cases. I don't even think negligence law is necessarily strong enough. But the legislators I was working with and I made a choice to carve out misuse from any new liability. Background negligence and products liability would still apply, but no new liabilities are created for misuse or malicious modification in these bills.

It only covers alignment failures or capability failures for which a human would be liable under similar circumstances. There's even an affirmative defense that applies to background law. If the system is substituting for some human function, like driving or medical applications, and it satisfies this human standard of care, that's a defense against liability.

I really designed this to try to narrowly target this misalignment risk and to do it in a way that's broadly consistent with promoting innovation. You saw in the debate over SB 1047, to the extent that the liability provisions were focal in the public debate, that it was almost entirely focused on misuse scenarios.

I think there was a tactical choice made by some of the supporters that misuse is more salient, and that they focused on those kinds of risks in making the case for the bill. I think that was a reasonable calculation to have made. But as a political matter, the principle that you should be liable if your system does something that the user didn't intend is really easy to defend.

Whereas with misuse, you get into all these cases of, “Well, should you be liable any time the electric company is not liable when someone does something with a power tool?” Or you're not going to hold a steak-knife manufacturer liable when someone gets stabbed with their product, right?

I don't think SB 1047 would have done either of those things, or the equivalent in the AI context, but I think it's much easier to demagogue in the misuse context. I actually don't think SB 1047 changed background liability law much at all because, as we've been talking about, there is this reasonable-care standard in negligence law that already applies. In the final version of SB 1047, they were just codifying that.

I don't actually think it imposed significantly new liability. But as a political matter, it did provoke a lot more backlash because it included misuse in scope. So I think, at least in terms of the first foray into strengthening liability laws, this is the balance that makes the most sense.

Nathan Labenz

How are state-level legislators responding to this stuff? It's been striking to me, honestly, that the public survey results seem to suggest broad-based support for doing something.

I am very sympathetic to the sort of cautionary voice that's like, just because people want to do something doesn't mean we should do this—whatever is in front of us. But it is nevertheless kind of surprising that at the national level, there doesn't seem to be much appetite to do much. And SB 1047, which you mentioned, got vetoed. What do you think are the prospects for the bills that you're particularly involved with, and more generally, what has your impression been of the state-level politics of all this?

Gabriel Weil

Yeah, so I think the legislators I'm talking to have been pleasantly surprised that it hasn't received the same level of pushback that they expected. It doesn't look like either of these bills is going to move forward this year, for the same reasons that most legislation just doesn't get traction. But I think both the legislators I'm working with are excited to keep trying this in the future.

I'm happy to talk to legislators in any state that they want to. I'm particularly excited to work with a Republican in a red state. I think this should not be a partisan issue, and I think that this liability-based approach is consistent with a small-government way of handling these risks that should be attractive to libertarians and Republicans. So I'm happy to work with anyone that wants to and to adapt the specifics of the legislation to their priorities and their local political circumstances and constraints.

Nathan Labenz

Yeah, well, I don't know how many Republican legislators we have in the audience, but if any are listening and made it this far, get in touch. One thing I wanted to compare and contrast with—and I think we're going to put out these two episodes in relatively close proximity, calendar-wise—is a conversation I just did with Andrew from Fathom and Professor Gillian Hadfield, who are behind the private governance idea.

Like many of these things, and your kind of set of proposals, it is both a broad framework for thinking about things and also starting to get instantiated in specific legislation. So it's SB 813 in California that we talked about concretely there. It seems like both of these proposals are, first of all, really taking seriously the fact that it's just really hard to do prescriptive legal regulation of a technology that's moving and morphing as fast as AI is right now. So I think that's an excellent starting point for both. I think there's also the sense that we want to get the people who are the most knowledgeable and best able to do this thinking to do that thinking, and yet there's a very different outcome to the recommendations that are made.

With theirs, there's some sort of trade-off, right, between setting up a kind of market for regulators that companies can opt into. The regulators themselves would be private institutions, but would be approved by and sort of reviewed and monitored on an ongoing basis by some part of the government. In exchange for opting into this regime and living up to the best practices and standards, whatever those are, these companies would get some sort of protection from liability—whether that's total, partial, an affirmative defense, or a rebuttable presumption. I'm learning all these terms as we go.

How would you compare and contrast your proposal with this other one? Is there any synthesis that could be possible between them? It seems like there's so much commonality, and then it seems like there's a very sharp divergence at the last step of exactly how we implement a good solution.

Gabriel Weil

Yeah. Okay, so I want to take this in a couple of different directions. One is to focus on what I think are the strengths and weaknesses of the legislative proposal in California that Fathom was behind. It was SB 813. When you think about markets, they're good at achieving good outcomes when there aren't the externalities that we've been talking about, right? And so I think you might worry that markets on their own don't do a great job even for users because there are asymmetric-information problems. I think the regulatory-markets idea might be useful for solving that kind of problem.

If you were saying, "Well, I'm going to just decide which of these systems I want to use. I want one that's certified by one of these—I think they call them multistakeholder regulatory organizations—and I'm going to give up my right to sue if something goes wrong, but I know that going in and I can choose which of these MROs I trust," and the MROs are going to be sort of vouching for the underlying AI companies that they are certifying, I think that's fairly unobjectionable for the same reasons that I'm less concerned about liability for harms to users more generally. It does address some of those issues I was talking about, some of the paternalism and asymmetric-information-type issues.

My core objection to Fathom's proposal is that their liability shield—and again, there were different variations on this that were stronger and weaker—but in all the legislative versions that they put forward, the liability shield extended to third parties. Third parties that were harmed by these systems would not be able to sue developers of these models if they were MRO-certified, even though the non-users had no choice about whether to be exposed to risks from these MRO-certified models. Because of that, the MROs lack strong incentives to worry about harms to non-users, right?

When you think about this market by default, you might worry that there's a race to the bottom, right? You're going to want to be certified because it gets you liability protection, but you want to have as weak standards as possible for users. There's a little bit of a break on that because users can evaluate these MROs and decide, "Well, this MRO is really shoddy, and I'm not going to be able to sue if something goes wrong." Maybe there are some public watchdogs that point this out and warn consumers about it. So maybe that works well enough for that.

But again, the MRO doesn't have much incentive to worry about third parties, and yet those third parties are bound under Fathom's framework and not able to sue if something goes wrong. And so to me, that's the key shortcoming: it doesn't protect third parties.

Now, if you thought the risks to users are very tightly coupled with the risks to third parties, maybe you think that's okay. I think there are 2 reasons not to necessarily buy that. In some contexts, there are clear trade-offs between risks to users and risks to third parties. Think of autonomous vehicles. Autonomous vehicles are going to sometimes be in situations where they have to trade off relatively minor risks to vehicle occupants against higher risks to other road users, right?

And if you're an MRO or if you're a consumer, you're probably going to, if you're pretty selfish—as most people are when they buy cars—not be so worried about how much it harms other people on the road. You're going to want to buy the one that's going to prioritize you, right? And so there's not a lot of reason to think this MRO model is going to protect the third parties who now have no right to sue the developer of this model if they get run over by one of these vehicles that's certified by the MRO.

Then, just more generally, if we think that there are large-scale risks, quantitatively, even if they're directionally the same sort of risks, the risks to third parties are going to quantitatively outweigh the risks to the user, right? Think about a problem like pollution or climate change—greenhouse-gas emissions, right? It's true that when I drive my car, it heats the planet a little bit and I suffer a little bit from that, but that doesn't give me very strong incentives to worry about that, right? Unless I'm altruistic, right? Because there are 8 billion people in the world. I'm only suffering roughly 1/8-billionth of the harm from that, right?

But with localized pollution, it's not quite that bad, but still it's like I'm suffering 1/1,000th or 1/10,000th or something. And so most of the harm is external. There are going to be cases like that where it's not a direct trade-off, but if you're only focused on the harms to users or the risks to users, you're not going to be addressing most of the issue. And so I would be much more inclined to support something like this MRO model if the liability shield only applied to harms to users.

Another thing that might come up is that, right now, as we talked about, negligence is the regime that applies. I had some conversations with the folks behind SB 813, and one idea that they suggested some openness to—I don't think it ever made it into the bill—was that companies that don't get MRO certification would maybe be subject to a strict-liability regime. So maybe if you combine those things, if you say, "Well, this liability shield only applies to harms to users, and if you don't get MRO-certified, then there's strict liability," maybe that's a synthesis that we could both support.

Nathan Labenz

Yeah, that's interesting. I do agree with the concern about the race to the bottom, and I was also somewhat persuaded by their response to my concern, which was basically, at some point, somebody's got to do a good job in this system of managing things. And so, to some degree, the question is: who do you want to trust, and therefore who do you want to empower? Who do you think is actually capable of doing a good job?

I think part of the notion that they have, at least in part, is that the organizations that will step up and try to take on this responsibility are the ones that can do a good job. One of the questions I've been going around recently asking people is: who is going to be an MRO if this actually happens? What organizations do we have today that are going to step up and be an MRO? I've asked this of some organizations directly, and I've also asked other people, "Who would you nominate to be an MRO?" I think there are some interesting candidates, although still quite few, but I think the notion that they have, at least in part, is that the organizations that will step up and try to take on this responsibility...

We should have some optimism or confidence that they will be intrinsically motivated to do a really good job on behalf of the rest of society. And they will take into account these extreme tail risks in a way that maybe a sort of insurance requirement might not really be able to capture, because that's just the kind of people they are and that's the kind of organization that's going to try to become an MRO.

And so I think they may be thinking of the two halves of the trade as less directly related. It's, I think, in their minds, a little bit less about applying standards specifically to reduce harm to users, and therefore the users don't get to sue anymore, and more about a package of standards that will be generally, hopefully, virtuous and take into account things that are hard to engineer incentives for. But then giving the carrot to the companies at the same time and hopefully getting all of that to work.

Gabriel Weil

Yeah. So I think it depends on how lax this is. There's some government body or government office—I think in their bill it was the California attorney general, right? They wanted to see that changed, actually, but, yes, my understanding is that it is the AG as it's written. And they were kind of like, “Yeah, we think maybe that should be more of a commission or something,” because we do have the problem of what happens when the administration turns over, which, you know, we're living through right now.

Nathan Labenz

You can imagine that being very stringent, or you could imagine it being very lax. So let's talk about both of those scenarios. If it's very lax—if basically anyone who wants to set up an MRO can do it—then I think this market competition, which is often good, will create a race to the bottom, at least for harms to non-users. Again, there is market feedback to prevent a race to the bottom for users. I think it could work pretty well for that.

But, sure, maybe some really well-intentioned people will become MROs, but they're going to have a hard time finding people who want to sign up with them, because you're going to want the most lenient standards that your customers are happy with, right? And so, if the AG or whoever is responsible is pretty lax, I think that's the equilibrium you end up in.

Now, if you're in a more stringent equilibrium, then you have to ask, well, now the government is taking on a much more ambitious role. And so all these benefits we were supposed to be getting from this market feedback—it's not clear that we're getting them anymore, right? Because the key point of failure is: Are we certifying this MRO?

Now, you could imagine that a lot of this has to do with how legible you think the safety target is, right? And it has to be in sort of this sweet spot for this MRO model to make sense. Because if it were super legible, you could just have a government-enforced safety standard, in the same way that we have pollution standards for power plants. The EPA doesn't say—or at least in some domains, they don't say—you have to install this control technology; they say you have to limit your emissions to this much per kilowatt-hour of electricity that you produce, and we can measure that. It's legible, and that's fine, and there wouldn't be much benefit to having a private certifier, right?

Or you could say it's totally illegible, right? We don't know how to tell whether something's safe. If that's the case, it's going to be hard for the government body that's certifying these MROs to tell whether their standards are good enough, right? So I think for this hybrid model to work, you'd have to think that we're in some sort of in-between, where the government can't tell directly whether the AI companies are safe. It's not legible enough that they can just directly enforce a performance standard, right? But it is legible enough that they can tell whether these MROs are doing a good job.

It's not impossible to imagine that we're in that world, but I don't think we have strong evidence that that's what we're in. And so that gives me some caution about leaning very heavily on this model, at least when I don't think the market feedback works well.

And so I think the market feedback works pretty well for users. And in principle, that could be a strong enough carrot. You know, most of the liability risks that people talk about being worried about are harms to users, right? That's what the Character.AI case is. That's what a lot of the concern is about. And so it's not clear to me that that's not a strong enough carrot to get these companies to sign on. And then you preserve the threat of liability for these third-party risks. I still think that's an attractive synthesis. But, yeah, as introduced in California, I think that bill was net negative.

How similar do you think the standard-setting process would be under an insurance requirement? Because I could imagine—and I sort of floated this to them—I could imagine that if you said you've got to have insurance, then you might hope that a similar thing would happen via insurance, where the insurance companies would say, you know, the optimistic story is, “Well, this is obviously going to be a massive market, so we definitely want to be in it, but it's also a very tricky market, because what do we know about insuring AI, since nobody's really done it? We don't have a great baseline. Technology is changing, yada yada yada.” Then maybe they end up going out and contracting the same—basically calling in the same organizations and saying, “Hey, do you want to step up for us and be some sort of standard-setter, or help us evaluate risks?”

My own synthesis, which may not be right, is that this may be 2 ways of creating this sort of market, where these expert organizations—whether they're serving insurance companies or serving the California AG or a commission or whatever—might still be those kinds of groups trying to do the hardest thing: figuring out what the actual risks are and what should be required of companies to get into the game. How realistic do you think that is?

Gabriel Weil

So you could imagine insurance companies playing this sort of quasi-regulatory role. There are multiple ways they could do that, right? They could say, “We're not going to issue this policy unless you do X, Y, or Z.” And they could delegate some of that to a third party that helps develop those rules, or they could develop that capacity in-house.

Another tool they have is doing the underwriting, right? So they can charge you more or less depending on what safety precautions you've taken. And that could be a collaborative process where they say, “Here's our baseline premium that's maybe pretty high for insuring this kind of risk, but if you can show us what things you've done—and maybe it's things that we haven't thought of, right?—we can work with the AI companies and say, ‘Okay, show us all the safeguards you've put in place. If you can convince us you've reduced the risk, then we can charge you less for this policy,’ right?”

And so I think by default they're going to be pretty cautious. They're going to want to write policies that, on average, pay out less than the premiums, right? And so I think if we have liability insurance requirements, there's going to be a strong demand pull that's going to push up what the rates—the premiums—are. And insurance companies are going to be in a strong position to insist that, if they're going to write a policy that these AI companies can afford, they take various precautions.

So, back to what you were talking about earlier, if Anthropic implements something that the insurance industry thinks does offer significant, cost-effective risk mitigation, then they can say, “Well, we'll give you a significant reduction in your insurance premium if you implement that.” And I think that's a pretty attractive model.

Nathan Labenz

How would—do you have any framework for this? If I'm in, let's say, the state legislature, wherever, right, and there's a bill that has the sort of private-governance MRO kind of system, and then there's a more like codifying liability, maybe an insurance requirement, I'm kind of like, well, geez, I don't know. Both of these proponents sound pretty smart. They both are grappling with the fact that we can't just write rules now once and for all. They're both trying to tap into the power of the market and competition, and trying to create ways for new ideas to still be able to enter even once the ink is dry on the law.

But I just don't know which one is better. Asking for a friend. How do I think about deciding which of these I want to bet on?

Gabriel Weil

Yeah. So again, I think you don't have to totally choose between them. I think they are compatible as long as the liability protection in the MRO model doesn't extend to third parties. You could also imagine there being other carrots for the MRO model. There's no reason it has to be based on a liability shield, right? So you could just require MRO certification. It doesn't have to be tied to carrots at all. It could be a stick-based approach, right?

In that sense, they're not incompatible. The only way in which they're incompatible is if you decide you want to base it—you want to make the incentive to join or to get MRO-certified a liability shield, and you want to extend it to third parties.

That said, beyond that, I think I gave the arguments for why, if we're talking about that version of this MRO model, I think that's pretty unattractive. I don't know that I have much more to add to that. I think it really falls short in protecting third parties, unless you think we're in this very particular situation where the government is both able to, and the politics are going to work out such that they have the right incentives to monitor these MROs and only certify the ones that are protecting third parties.

I just don't have confidence that that's going to carry through. And so that's what gives me some hesitance about the most robust version of this MRO model.

Nathan Labenz

Yeah. I like—I mean, you've got some good synthesis ideas there, though, so I like that. Maybe just a couple of final things. I really appreciate all your time. You've certainly been very generous with it as I've asked many, many tangential and follow-up questions.

In the spirit of red-teaming this proposal, one question I always try to ask is: What might the AI companies do differently under this regime that could perhaps even be bad, as opposed to the good that you're trying to induce? The one idea that I came up with, which is inspired by the AI 2027 scenario, is that independent of this kind of legislation, it sort of projects the AI companies starting to increase the gap between what they deploy and what they have internally that they're using for their own AI research or what have you.

There's this idea that we're already getting to the point where frontier models suffice for a great many use cases. So they might, for multiple reasons, decide, “Maybe we don't want to tip our hand to competitors anymore. If we put this out there, who knows? Somebody at the other company might use it to do their AI research. We definitely don't want that.” So they could have multiple reasons, but this could be an incremental reason to say, “Maybe we shouldn't deploy this. Let's just keep it in-house for our own internal use.”

In addition to wanting to continue to use the best available AI until the singularity, I do feel like the iterative deployment philosophy that OpenAI pioneered seems to have a lot going for it. Obviously, you can overdo a good thing and not test enough or whatever, but at least the theory that, or the contrasting idea that, if somebody develops AGI in their basement and then springs it on the world one day, that seems clearly not good.

So this iterative thing does strike me as a good alternative. But loading up more liability for them in doing that could perhaps cause them to go the other direction and say, “Well, we'll just make our own bid for superintelligence, and we'll kind of see how that goes, and then we'll go from there.” Any thoughts on that?

Gabriel Weil

Yeah, I have 2 sets of thoughts on that. One is that I think these benefits from iterative deployment, at least when you're talking about misuse, should go into the calculus. These safety benefits from iterative deployment are part of why I think pure strict liability in the misuse context doesn't necessarily make sense. One of these external benefits of deploying systems that are potentially susceptible to misuse is that you would get the safety benefits. So I think that is part of the calculus there.

But I do want strong strict liability for misalignment, and so I think this critique still bites. I think this is actually a subset of a broader concern that I flag in the original paper, which is that one potential failure mode for this proposal, particularly the punitive-damages aspect, is if the things you would tend to do—the most cost-effective ways to mitigate these warning-shot risks—do not actually have much effect on the underlying uninsurable risk.

Then you're not getting much benefit out of this. In the formula for what the punitive damages should look like that I give in the paper, there's this elasticity parameter. Elasticity here is, for every unit of reduction in the practically compensable harm, how much risk reduction do you get for the uninsurable risk? You would want, if you have a lot of these potential warning-shot cases—maybe some are warning shots and some aren't—one thing that makes something a warning shot is that it's more elastic with respect to these uninsurable risks.

I think what you're saying is maybe none of them are that elastic, because maybe the most cost-effective way to mitigate the risk of these warning shots is to just not deploy externally. I don't have a strong reason to think that's the world we're living in, but I can't rule that out.

A couple of things to say there. One is that maybe there are warning shots that come from internal deployment. If that's the case, nothing in my proposal depends on there being external deployment. There could be misuse by internal actors, there could be cyberattacks that cause your system to be accessed by bad actors, and there could be alignment failures where someone internally using your system causes some harm in the world. I think all of those would be subject to my regime.

When you're talking about liability insurance, I've talked at different points in this conversation about what the key critical step is. I don't think those requirements should necessarily only apply to external deployment. If we think there are significant risks that are created earlier in the value chain, whether it's in training, pre-training, fine-tuning, or internal deployment, I think those might be generative of risks that you're potentially judgment-proof for, and you should have to carry liability insurance for.

Particularly if we're moving to a world—and I don't know that we are—where more AI companies are adopting the SSI wait-to-deploy-until-we-have-superintelligence model, then I think it would be more important to have a sort of regulatory gate earlier in the development process. I think that can partially address that concern.

It might still be the case that not having external deployment is effective at stamping out these warning shots but doesn't actually mitigate the uninsurable risk. I think that's a subset of the more general failure mode for this proposal. If these warning shots aren't really correlated in the right way, such that the things that you would do to mitigate them do mitigate the uninsurable risk—the most cost-effective ways to mitigate them, the things that would be most attractive if you expect to pay out a large damages award—then if there just aren't a lot of cases like that at all, because these things we thought were warning shots aren't really warning shots in the sense that matters, then I agree you shouldn't want to lean heavily on this proposal.

Again, I don't have strong reasons to think that's the world we're living in, but that's the way I would think about it.

Nathan Labenz

Yeah, it's a good reminder, by the way, that there is, in fact, one company that has a stated plan of not doing anything until they hit superintelligence, which is a crazy world to be in. It's crazy that it's at least somewhat credible—credible enough to raise billions of dollars, as it turns out. Fascinating stuff.

Okay, 2 real quick final questions. One, I don't know if you have anything to say about this. This may be somebody else's area, but obviously, any time we do anything that would sort of slow down or impose additional cost or put more onus on developers, we always get, “Oh, China's not going to do that. We're just going to lose to China.”

One answer, of course, is, “Let's not do the bad thing ourselves. Maybe China will do the bad thing, and that would be bad, but that doesn't mean we should do the bad thing.” Do you know anything about how China or other countries are thinking about this kind of stuff? Obviously, it's a totally different legal environment over there, but is that something you've looked into at all?

Gabriel Weil

Yeah. You sent me something along these lines, and so I did do some digging today about what China's tort liability system looks like. It seems like the principles are pretty similar structurally. They do have a civil-law system, so it's more code-based and less common-law-based, but the substantive principles around negligence and products liability and some narrow pockets of strict liability are all pretty similar.

Damages calculations tend to be less plaintiff-friendly. There's what's called non-economic damages, like pain and suffering, that kind of thing, and Chinese courts tend to be less generous with those. There's also less of a lawyer population there that takes cases on contingency fees. I think there are some restrictions on when those are available, so fewer of these cases get litigated. If you have to pay your lawyer by the hour, you might not be able to finance a case. I think there's just fewer of these lawsuits more generally.

I don't think they have any sort of bespoke liability regime for AI, but neither does the US, really. So I think, in that sense, they're on pretty equal footing more generally.

You don't have to assume that China is going to adopt the same domestic regulatory regime that we do. You can imagine some kind of framework where we encourage that, right? But I think, more broadly, the question is—first of all, this critique could obviously be brought against any domestic AI regulation. If anything, liability is less vulnerable to this critique because it's more consistent with promoting socially useful innovations.

The other question is just: How binding should we treat this threat from China? It seems like we have a pretty significant lead over China at the frontier, and with export controls—which I have mixed views on the merits of, but at least in this context, they do seem likely to cause the lead to widen in the coming years, at least until China can indigenize its own chip supply chain.

Nathan Labenz

They can produce some chips right now, but they don't have access to new fabs from ASML. And so, I think in the medium term, that's going to be bad for their chip production. It's not going to be that hard, even if we're not going totally pedal to the metal in the US, to maintain a lead against China.

I also am less of a China hawk than most people are. I think we share that view. And so, I'm less worried about the sort of zero-sum competition than other people are, but obviously, reasonable people can disagree about that.

Cool. That's helpful. There's a lot more to unpack there. Certainly. Real quick, last one: I noticed you participated in the Principles of Intelligent Behavior in Biological and Social Systems program, aka PIBBSS. I've had a few guests with, I think, very interesting, unique takes on the AI question who have come through that program. I thought you might give kind of a testimonial or an invitation, or indicate what sort of people should be considering doing that themselves.

Gabriel Weil

Yeah. So, full disclosure, I'm on the board of the PIBBSS organization, but I will still answer this question honestly.

I think PIBBSS has two distinctive value adds from other AI safety talent-development-pipeline-type organizations. I think it wants to bet on more neglected ideas, and it wants to bring in a broader suite of people with different expertise. I had been socially connected to people who are worried about AI risk, but I hadn't worked on it professionally before I did PIBBSS. As I think I mentioned, I mostly did climate law and policy before I did PIBBSS two years ago, and it was very open to the set of expertise that I brought. I did teach torts, so I had a background in liability law.

I think it was great at getting me up to speed on the technical issues and then allowing me to leverage the expertise that I already had that was relevant. I think most people who do PIBBSS don't do more governance and policy-type stuff. They more often do alignment work, though not necessarily what people think of as technical alignment work. A lot of it is more conceptual, but I think it's open to a broad range of disciplinary approaches.

I think it's a great way for people who think they might have something to contribute to mitigating AI risk but haven't seen an obvious way in. It's more open to different ideas, and so if you think that fits your interests, there's a fellowship that's run every summer. There's also an affiliate program that I think is still ongoing for people who are a little bit more senior.

Most people who do PIBBSS are grad students or postdocs. I was a more senior person; I was already a professor when I did it. But yeah, there's a residency aspect to it. When I did it, it was in Prague for half the summer. This summer, it's in San Francisco.

I found it to be very productive. You're in a coworking space with other people working on this stuff, with people to bounce ideas off of. And so, I found it to be a really valuable experience. I encourage people who think they might fit this broad description to explore it next summer.

Nathan Labenz

Cool. Love it. I think there's definitely a big need for people from all different backgrounds with different, novel ideas, and so PIBBSS is great for that. This conversation has been a great example of that. I appreciate your reorientation of your legal career toward trying to address the AI challenges that only seem to be growing in importance.

Any quick closing thoughts before I give you the official sendoff?

Gabriel Weil

I think I've said most of what I wanted to say.

Nathan Labenz

Cool. Well, Gabriel Weil, assistant professor of law at Touro University and senior fellow at the Institute for Law and AI, this has been great. Thank you for being part of The Cognitive Revolution.

Gabriel Weil

Thanks. This was a lot of fun.

Liability for AI Harms: How Ancient Law Can Govern Frontier Technology Risk, with Prof Gabriel Weil | BidClub