# How AI agent liability works when the agent gets it wrong

> AI agent liability rarely stops with the vendor. Learn how tort, contract, product rules, evidence, insurance, and practical caps divide the loss.

An AI agent can send the refund, merge the code, reject the applicant, or place the order. It cannot sign the complaint, fund the settlement, or explain to a regulator why nobody owned the decision. The first rule of AI agent liability is simple: the law looks past the agent to the people and companies that designed, sold, configured, deployed, supervised, or benefited from it.

The harder question is which one pays. There is no universal answer, because the same bad action can trigger a customer claim under a contract, a negligence claim from an outsider, a statutory enforcement action, and an indemnity dispute between vendors. Founders need to separate those paths before negotiating a cap. A cap that handles service credits but ignores a discrimination claim or a corrupted production database is decoration. This is a practical operating framework, not legal advice for a specific jurisdiction.

## An agent does not absorb liability

The company that gives an agent authority normally owns the immediate operational result, even when a model vendor supplied the reasoning engine. Courts and regulators do not need to treat software as a legal person to reach a familiar defendant. They can ask who owed the duty, who controlled the system, who made the representation, and whose business received the benefit.

Calling the output autonomous does not break that chain. A purchasing agent acts through credentials issued by the buyer. A support agent speaks inside the seller's service process. A coding agent changes a repository because an employer connected it and chose the approval rule. Agency law may bind a company when its system appears authorized to transact. Negligence law may reach a company that exposed foreseeable harm without reasonable controls. Contract law usually puts the customer's direct claim against the supplier named in the agreement.

Responsibility and liability are different. Responsibility is the operating assignment inside a company: a named person owns the agent, its permissions, and its shutdown decision. Liability is an enforceable duty to pay or provide another remedy. A vendor can accept operational responsibility for fixing a defect while excluding most consequential damages. A deployer can own daily supervision yet still recover from a vendor whose warranty was false. Blurring these terms produces contracts where everyone is responsible and nobody has budgeted for the loss.

Start with a map of legal actors, not a diagram of model calls. Record the deployer, the application vendor, the model provider, integration contractors, data suppliers, affected customers, and foreseeable outsiders. Add the employer or regulated professional whose duties cannot be handed to software. The agent belongs in the technical record, but it is not the empty chair at the legal table.

## Four legal theories can reach the same error

Most agent failures fit existing causes of action, although facts and local law decide whether a claim succeeds. A novel system does not require a novel theory. The useful exercise is to run one failure through four routes and see which defendants, damages, and defenses change.

Contract is usually first in a business dispute. Did the agent perform a promised service, meet an accuracy commitment, follow instructions, protect confidential data, and stay inside an agreed use? Warranties, service levels, disclaimers, acceptance terms, indemnities, and the liability cap then shape the remedy. A model hallucination by itself is not a breach. The breach is failure to deliver the contractual result or honor a contractual control.

Negligence matters when someone owed a duty of reasonable care and caused foreseeable harm by breaching it. The disputed conduct may be poor testing, excessive permissions, missing review, ignored warnings, weak monitoring, or failure to stop after known incidents. A professional negligence theory can apply where licensed or specialized judgment remains with a doctor, lawyer, accountant, or other professional. Saying that an agent recommended the action rarely removes that person's duty.

Product liability becomes relevant when software causes the kind of personal injury or property damage covered by the governing regime. It does not automatically cover every wrong invoice or lost profit. The EU Product Liability Directive 2024/2853 expressly treats software, including AI systems and software delivered as a service, as products under a regime that does not require proof of fault. It applies to products placed on the market or put into service after 9 December 2026, once member states transpose it. The directive also provides evidence disclosure and rebuttable presumptions where technical complexity makes proof excessively difficult. That makes traceability a litigation asset, not merely an engineering preference.

Statutes create separate exposure. Privacy, consumer protection, employment, credit, intellectual property, discrimination, and sector rules can attach to a decision regardless of the contract's damage definition. The United States Federal Trade Commission has put the point bluntly in its AI enforcement work: using AI does not create an exemption from existing law. A contract can shift the bill between commercial parties, but it cannot authorize deception or prevent a regulator from acting.

## The claimant determines who pays first

Who suffered the loss determines which agreement matters and whether any agreement matters at all. Founders often inspect the model vendor's terms and conclude that exposure is capped. That conclusion covers only a narrow path: a claim between parties bound by those terms, for a type of loss the cap lawfully reaches.

Suppose an agent reads a forged email, changes a supplier's bank details, and pays the wrong account. The deployer's bank or insurer may reimburse part of the transfer under separate rules. The deployer may sue the application vendor for breach. The vendor may seek indemnity from an integration contractor. A real supplier may claim the original invoice remains unpaid. Each arrow has its own duty, evidence, exclusions, and cap. There is no single pot marked "AI loss."

Now change the facts. A recruiting agent rejects an applicant on an unlawful basis. The applicant did not sign the deployer's software agreement, so its limitation of liability cannot normally control the applicant's statutory remedy. The deployer may face the external claim first because it made the employment decision. It can pursue the vendor later only if the warranty or indemnity covers the failure and the deployer complied with notice, use, and terms governing control of the defense. Cash timing matters: a right of recovery three years later does not pay today's counsel or regulator response.

Treat the chain as two layers. External liability covers duties owed to customers, employees, applicants, consumers, regulators, and other outsiders. Internal allocation covers who reimburses whom among the deployer, application vendor, model provider, and integrator. Contracts are strongest in the second layer. They influence the first through required controls and insurance, but they do not rewrite an outsider's rights.

Resellers and open source components complicate recovery without making liability vanish. The party visible to the claimant may have no direct agreement with the developer that introduced the defect. An application vendor may incorporate a model, an orchestration library, and a community tool whose licenses disclaim warranties. Those disclaimers can shape claims among contributors and integrators, but the company that assembled and sold the working service still needs to answer for its own promises and controls. The revised EU product regime makes a related policy choice: noncommercial free and open source software sits outside its scope, while a manufacturer that integrates such software into a commercial product can remain exposed for the finished product.

Choice of law and forum also change the path. A Delaware clause between two vendors does not necessarily control an employment claim in another country or a regulator acting under local statute. Arbitration may keep a commercial dispute private, yet it may not bind an affected consumer or government agency. Build the actor map by jurisdiction and claim, then ask counsel where mandatory rules override the paper. Copying one global cap into every customer order ignores the part of liability that moves with the affected person.

One more distinction saves arguments: an agent error is not always a model error. A correct model output can cause loss because the workflow mapped it to the wrong account, retrieved stale data, skipped an approval, or retried an irreversible action. If the incident record stops at the prompt and response, the company may blame the most visible component while missing the component that legally and technically caused the harm.

## Contracts allocate the bill but not the law

A useful AI clause allocates a specific failure, a specific control, and a specific remedy. Broad language saying that one party is "responsible for AI" leaves the expensive questions unanswered. Put the operating terms in a schedule that engineering, procurement, security, and counsel can all test.

The schedule should define the permitted use, forbidden decisions, data classes, tool permissions, approval thresholds, logging fields, retention period, evaluation method, change notice, incident notice, suspension right, and exit assistance. State which party supplies each control. Tie a warranty to facts within that party's control. The deployer can warrant its input rights and configuration. The vendor can warrant that the service materially follows documentation and that it will not silently grant a new tool permission.

A compact allocation can look like this:

```text
Use: Draft supplier refunds up to USD 500.
Action: Human approval is required before payment.
Vendor control: Emit amount, payee, evidence IDs, model version, and confidence flag.
Customer control: Restrict payment credentials and maintain the approver list.
Change: Vendor gives 30 days' notice before removing an approval control.
Incident: Notice within 24 hours after confirmed unauthorized payment.
Evidence: Preserve the action record and relevant configuration for 24 months.
Remedy: Suspend automated payment action without terminating read-only use.
```

That fragment prevents a common failure. The sales order says the agent "assists with refunds," the production credential permits any transfer, and each party assumes the other created the limit. A numeric authority boundary plus a named credential owner turns an assumption into a test. The numbers and periods above are examples, not universal defaults. Set them from the actual transaction and legal retention needs.

Indemnity and limitation clauses do different jobs. An indemnity assigns defined third party claims and usually includes defense procedures. A cap limits covered monetary exposure between contracting parties. Exceptions may apply to confidentiality, data protection, intellectual property, fraud, deliberate misconduct, or indemnity obligations, but local law decides what can be limited. Avoid an unlimited exception for "any AI issue." It is too vague to price and encourages both parties to relabel ordinary defects as AI failures.

Read the dependency contracts before promising an indemnity downstream. A model provider may cap all recovery at a small fee, exclude outputs, require use of specified safeguards, or reserve rapid changes to the service. The application vendor can still offer broader protection, but it is then underwriting the difference. Price that gap explicitly instead of assuming a claim will flow neatly through the supply chain. Require notice soon enough to preserve upstream rights, and make cooperation duties possible within the evidence and access each party actually controls.

Defense control deserves its own sentence in the agreement. The indemnifying party usually wants to select counsel and approve settlement, while the protected party needs a right to participate when reputation, licensing, customer relations, or injunctive relief is at stake. State who pays defense costs, whether those costs reduce the cap, when consent can be withheld, and what happens if the indemnifying party refuses the tender. Many apparent indemnities fail at this procedural layer rather than on the definition of an AI error.

## A liability cap needs a loss model

Set the cap from credible loss paths, not from the vendor's annual fee by reflex. A fee multiple is easy to negotiate because the number already exists. It may still be unrelated to the authority given to the agent. A USD 2,000 monthly tool that can release USD 5 million has a cheap subscription and an expensive permission design.

Build one row for each material action: send money, delete data, publish content, change production, disclose records, approve a benefit, or bind the company. For each row, estimate the maximum amount per action, actions before detection, direct restoration cost, required notice or response cost, and plausible third party harm. Do not force a precise probability when the company lacks data. Use scenarios and state the assumptions.

Then choose three contractual buckets. The ordinary cap handles routine breach, rework, and service failure. A higher cap can cover exposures that both sides can insure or control, such as certain confidentiality, security, or indemnity claims. Uncapped exposure should be rare and confined to conduct the law will not permit a party to limit or conduct the parties deliberately exclude from protection. This structure is often called a supercap, but the label matters less than the listed claims and the wording that stops double recovery.

Test the proposal with a short calculation:

```text
Maximum authorized action          USD 25,000
Actions possible before detection  x        4
Direct reversal and response       + USD 40,000
Modeled operational loss           = USD 140,000
Proposed ordinary cap              = USD  60,000
Unfunded contractual gap            USD  80,000
```

The result does not prove that USD 140,000 is the correct cap. It exposes a decision. The company can raise the cap, lower the agent's authority, shorten detection time, require a reserve, or buy insurance. I prefer reducing authority before paying for a larger promise. Controls prevent loss; a larger cap starts an argument about reimbursement after the damage exists.

Also model aggregation. A cap for each claim may fail if one configuration error creates thousands of transactions, while an aggregate cap may exhaust after an unrelated incident. Define related claims, the cap period, defense costs inside or outside the cap, service credits, and whether indemnity payments erode the same amount. If those mechanics remain vague, the headline cap tells the board very little.

## Evidence decides whether the cap matters

A company cannot enforce allocation or defend its conduct if it cannot reconstruct the agent's action. Logs need to prove authority, context, control, and causation. A transcript alone rarely does that because an agent acts across models, retrieval systems, tools, policies, and human approvals.

Record an immutable event identifier, time, tenant, agent and workflow version, model identifier, policy version, tool call, redacted arguments, identifiers for data sources, approval request, approver identity, approval result, final side effect, retry count, and rollback status. Preserve the exact configuration or a content hash that retrieves it. Keep access records for the log itself. Sensitive prompts should not become an uncontrolled second database, so store only what the investigation and retention rule require.

A normalized incident record can use this shape:

```json
{"event_id":"evt_0187","agent_version":"refund-12","policy_version":"pay-4","tool":"issue_refund","amount":750,"approval_required":true,"approval_id":null,"result":"blocked","evidence_ids":["invoice_391","case_882"]}
```

The useful output is "blocked," not a paragraph saying the model was uncertain. A control should fail closed when the approval record is absent. For a completed action, record the external transaction identifier so finance and counsel can connect the agent event to the actual loss. Test that chain during normal operations. An audit log that nobody can query during an incident is a storage bill.

Evidence also changes bargaining power. The EU Product Liability Directive's disclosure and proof provisions can penalize a defendant that withholds relevant evidence and can ease a claimant's burden in technically complex cases. Elsewhere, ordinary discovery, regulatory requests, and rules allowing an adverse inference create similar practical pressure even when the legal test differs. Logging everything forever is the wrong response. Use a documented retention schedule, legal holds, access controls, and deletion that actually runs.

## Approval gates must match irreversible harm

Human approval reduces risk only when the reviewer has information, time, authority, and a real chance to stop the action. A button that invites automatic approval transfers blame on paper while leaving the machine's pace unchanged. Review design belongs in the liability model because it decides what the company could reasonably prevent.

Classify actions by reversibility and impact. Drafting, searching, and proposing can often run automatically inside controlled data boundaries. Sending a message, changing a record, or opening a pull request may need sampled review or a low numeric limit. Paying money, deleting production data, making a regulated eligibility decision, or accepting legal terms should require explicit authority and stronger separation. Some uses should remain prohibited when the company cannot supply a competent reviewer or explain the decision.

Do not require approval for every tool call. That popular recommendation feels safe and fails in practice because reviewers habituate to noise. Place the gate immediately before the irreversible side effect, show the proposed action and evidence, and make rejection easy. Batch approval deserves suspicion: one click that releases 500 individually risky actions has not preserved human judgment.

The EU AI Act makes this separation concrete for systems within its scope. It distinguishes providers from deployers and assigns duties along the value chain. For AI systems classified as high risk, it requires measures for human oversight and expects deployers to assign oversight to people with competence, training, authority, and support. The regulation is a compliance regime, not a general compensation statute. Still, ignoring a relevant compliance duty can become powerful evidence in a later negligence or product case.

Measure override rate, blocked actions, time to review, repeated approver exceptions, and incidents discovered after approval. Do not reward approvers for throughput alone. If almost every request is approved in seconds, either the gate is unnecessary or the review is fictional. Fix the workflow instead of treating the signature as a shield.

## Insurance closes only named gaps

Insurance pays only when the policy language reaches the claimant, event, damage, and time period. Buying a cyber policy and declaring the agent covered is careless. Agent errors can look like technology service failures, privacy events, security incidents, media claims, professional mistakes, employment decisions, bodily injury, property damage, crime, or unauthorized funds transfer. Those categories often sit in different policies or exclusions.

Ask the broker and coverage counsel to map each modeled action to a policy, insuring agreement, limit, retention, exclusion, and notice rule. Technology errors and omissions coverage may address failure of a technology service. Cyber coverage may address certain privacy and security events. Crime or funds transfer coverage may matter when deception triggers payment. General liability may matter when physical harm occurs. Directors and officers coverage may enter a governance dispute. None of these descriptions guarantees coverage. Endorsements and definitions decide.

Give the insurer a truthful description of autonomy, transaction limits, production access, human review, model changes, and prior incidents. A vague application can create a later fight over misrepresentation. Check whether the policy treats contractors and model suppliers as insureds, vendors, or neither. Check exclusions for contractual liability, voluntary payments, fines and penalties, intellectual property exclusions, war or infrastructure language, and consent requirements before settling or hiring incident firms.

Match contract mechanics to coverage mechanics. If the customer contract demands notice within 24 hours but the vendor waits for monthly incident review, the indemnity may fail before insurance begins. If defense costs sit inside a small cap, litigation can consume the recovery. If the supplier must carry a limit, require evidence of coverage and notice of cancellation, but do not mistake a certificate for the policy.

Keep a funded retention and an operational reserve for exclusions and timing gaps. Insurance should finance a defined residual risk after permissions, approvals, monitoring, and contracts have done their work. It cannot make an unsafe deployment reasonable.

## A risk plan engineers can execute in 30 days

A founder can get control of AI agent liability in 30 days by narrowing authority, assigning owners, and producing evidence before rewriting every vendor agreement. The work needs one accountable executive, one technical owner per agent, security or privacy input where relevant, and counsel for jurisdiction and contract decisions.

1. During days 1 through 5, inventory every agent that can call a tool or influence a consequential decision. Record its owner, purpose, users, data, credentials, action limits, jurisdictions, vendor chain, and shutdown method. Disable orphaned agents and credentials immediately.
2. During days 6 through 10, select the five loss paths with the highest credible impact. Walk each from input through side effect. Mark where a person can intervene, how fast the company detects failure, which outsider can be harmed, and which contract would fund recovery.
3. During days 11 through 15, reduce permissions and install gates at irreversible actions. Separate drafting from execution, use dedicated credentials, set transaction limits, block missing approvals, and test rollback. Run one failure drill with a forged instruction or stale record.
4. During days 16 through 22, make the evidence record complete. Capture versions, policies, tool calls, approvals, and external transaction IDs. Confirm that an incident lead can retrieve one event without help from the engineer who built the agent.
5. During days 23 through 30, reconcile the customer contract, vendor terms, insurance, and incident plan against the same loss table. Escalate gaps that exceed the company's reserve or risk tolerance. Put renewal reviews and reviews after model changes on the calendar.

The deliverable is a single page register for each material agent, backed by evidence rather than a policy deck. It should answer: what can this agent do, what is the largest irreversible action, who approves it, how soon will we know it failed, who receives notice, and where does the money come from? If the team cannot fill one field, that is an operating task with an owner and date.

A Team & AI Audit from oleg.is can put the team and cost side of that decision on a clock of five business days while leadership keeps legal advice with qualified counsel. The service has a fixed USD 5,000 fee and guarantees at least USD 50,000 a year in identified savings or the audit is free. Savings do not excuse unsafe authority, but they can fund the controls, reserve, and senior ownership that a serious agent program needs.

Do not wait for an AI-specific court opinion that matches your stack. The agent can already move money and data under ordinary law. Cap that authority first, preserve the evidence, and then negotiate the financial cap with a loss model in hand.
