top of page

What Is an AI Incident? Building a Response Plan Before You Need One

Writer: Matt Lazarus
Matt Lazarus
Sep 28
6 min read
Isometric illustration on a dark navy background of a central AI core fanning output lines out to many small terminals, with three discoloured lines travelling well beyond the others before any alert ring appears.
The characteristic AI incident is a wrong answer travelling quietly to many people well before anything raises an alarm.

Most organisations deploying AI have spent their governance effort on the front end. There is an acceptable use policy, a vendor assessment, perhaps a risk register with a dozen entries and an owner against each. All of it describes what should happen. Almost none of it describes what happens at eleven o'clock on a Thursday night when an assistant has been confidently giving three hundred staff the wrong leave entitlement for a fortnight.

 

That gap matters because AI failures do not look like the incidents your existing plans were written for. There is rarely an alert or an obvious moment of compromise, just a slow accumulation of plausible, wrong, unlogged answers until someone acts on one. By the time it surfaces, the question is not only how to stop it but how far back the damage runs.

 

An AI incident response plan closes that gap. It is not a separate discipline bolted onto security - it is your existing incident process extended to cover a class of system that fails quietly and without the evidence trail your responders are trained to look for. This guide sets out what counts as an incident, who owns it, how you detect one, and what the first hours look like.

 

Key Takeaways

 

  • AI fails without alarms. The characteristic incident is a confident wrong answer acted on repeatedly, not a breach your monitoring will flag.

  • Ownership must be named in advance. AI incidents cross security, legal, data and the business owner, so an unassigned incident stalls at exactly the wrong moment.

  • Retention decides your blast radius. Without prompt and output logs you cannot establish who was affected, which turns a contained issue into an open-ended one.

 

What Counts as an AI Incident?

 

An AI incident is any event where an AI system produces or enables an outcome that causes harm, breaches an obligation, or would fail scrutiny if it became public. That definition is deliberately broader than a security incident, because most AI failures involve no attacker and no compromised system - the technology worked exactly as designed and still produced an unacceptable result.

 

In practice they cluster into a handful of shapes. An assistant surfaces a document the user was never meant to see. A model produces confident, fabricated detail that reaches a customer or a regulator. Staff paste sensitive information into a consumer tool. An agent with write access takes an action nobody sanctioned. A model's behaviour shifts after a vendor update and degrades a process that depended on it.

 

Writing that list down matters more than perfecting it. Staff escalate what they recognise as an incident, and "the assistant gave me a wrong number" does not feel like one. If your definition does not name plausible-but-wrong output as reportable, you will not hear about it until the consequences arrive on their own.

 

Why Doesn't Your Cyber Incident Response Plan Cover This?

 

Because a cyber plan is built around unauthorised access, and most AI incidents involve entirely authorised access producing the wrong result. The triggers, the evidence and the containment steps all assume an intrusion that, in these cases, never happened.

 

Consider detection. Security tooling watches for anomalous access patterns and exfiltration. An assistant answering a permitted query from a permitted account with an over-permissioned index generates none of those signals, which is the mechanism behind most oversharing incidents - a point we covered in how Microsoft 365 Copilot surfaces files you forgot existed. Nothing was breached. The permissions were simply wrong and the tool was efficient.

 

Containment differs too. Isolating a compromised host is well understood. Deciding whether to disable an assistant four hundred people now depend on, while you work out whether the fault sits in the model, the retrieval layer, the source content or the prompt, is a continuity judgement needing someone with authority to make it fast. If the plan does not say who, the default is an argument.

 

Isometric illustration on a dark navy background of an incident timeline running from a sealed evidence block through a containment barrier to a widening scope cone over a grid of user tiles.
Preserve first, contain second, then establish scope: reversing that order destroys the record of who was actually affected.

Who Should Own an AI Incident When It Happens?

 

A single named accountable owner, supported by a standing cross-functional group. The common failure is not that nobody cares, but that four functions each hold one necessary piece and none can act alone, so the first hour goes to convening rather than containing.

 

The four seats worth naming before you need them:

 

Role

Owns

The question they answer

Business system owner

The affected process and its users

What decisions were made on this output, and can we pause?

Technology or data lead

The model, retrieval layer and permissions

What actually caused it, and how far back does it go?

Legal, privacy or risk

Obligations and disclosure

Does this trigger a notification, and to whom?

Communications

Internal and external messaging

What do staff and affected parties need to be told?

 

For most mid-market organisations these are existing people wearing an additional hat, not new hires. What matters is that the hat is assigned in writing and that the accountable owner has authority to suspend a system without escalating further. Where the internal capability genuinely is not there, a fractional data and AI team can hold the technical seat until it is.

 

How Do You Detect an AI Incident in the First Place?

 

Through deliberate feedback channels and sampling, because passive monitoring will not find these. The overwhelming majority of AI incidents are first noticed by a user who thinks an answer looks wrong, so your detection strategy is largely a question of whether that person has somewhere obvious to say so.

 

Three mechanisms carry most of the weight. First, in-tool reporting: a visible way to flag an output as wrong, routed somewhere monitored rather than a mailbox nobody reads. Second, periodic sampling: a human reviewing a small set of real interactions each month against known-good answers, which catches gradual drift no single user would report. Third, outcome monitoring on the process itself - if an assistant supports credit decisions, the credit outcomes are your smoke detector.

 

All three depend on logging you configure deliberately: prompts, retrieved sources, outputs, user and timestamp, retained long enough to reconstruct a period rather than a moment. Retention decides whether you can answer "who else got this answer?" in an afternoon or not at all, and it is the most consequential technical decision in the plan.

 

What Are the First Steps When an AI System Goes Wrong?

 

Preserve the evidence, then contain, then establish scope - in that order. The instinct is to fix the system immediately, but a well-meaning correction to a prompt or an index can destroy the record of what the system was doing when it failed, and that record is what tells you who was affected.

 

A workable first sequence:

 

  1. Preserve logs, the current configuration, the prompt or system message, and the retrieval index state before anyone changes anything.

  2. Contain by restricting access, reverting to a prior configuration, or suspending the system if the harm is ongoing.

  3. Establish the window: when did the behaviour start, how many interactions fall inside it, and which of those look affected.

  4. Assess consequence - not just what was said, but what was decided or communicated as a result.

  5. Notify internally, including the accountable executive, on a defined timeframe rather than when convenient.

  6. Record the sequence contemporaneously, because a reconstructed timeline written a week later is worth considerably less to a regulator.

 

Steps three and four are where organisations discover whether their earlier architecture decisions were sound. If you cannot tell which users received the affected output, every user is potentially affected, and a contained incident becomes an enterprise one. This is precisely the capability an AI Data Readiness Audit tests for before deployment rather than during a crisis.

 

When Do You Have to Tell Someone Outside the Organisation?

 

Whenever an existing obligation is triggered, and AI does not create a separate reporting regime so much as reach into several at once. There is no single AI notification threshold in Australia today. There are privacy, contractual, sector and directors' obligations that an AI failure can each independently engage.

 

The realistic triggers: personal information disclosed to people not entitled to see it, which may engage the notifiable data breaches scheme; contractual commitments to clients about how their data is handled or whether AI is used on it; sector rules in financial services, health and education that apply regardless of the technology involved; and material operational impact that a board must weigh. Individual entries in your AI risk register should already map to these, which is what makes the register useful in the moment rather than a compliance artefact.

 

That judgement is genuinely difficult and should not be made for the first time under pressure. Decide in advance who assesses notification obligations, on what timeframe, and what legal advice is available out of hours. Organisations working out who to phone at that point invariably notify late, and lateness is what regulators reliably take an interest in.

 

A Plan You Have Not Rehearsed Is Only a Document

 

The distance between a written incident plan and a working one is a single tabletop exercise. Take a plausible scenario - the assistant has been giving incorrect entitlement advice for two weeks and someone has now acted on it - and walk your named owners through it for ninety minutes. You will find the gaps immediately, and they will not be the gaps you expected.

 

What it typically exposes is not a governance failure but an architectural one: logs never retained, permissions nobody could enumerate quickly, no way to establish which users saw what. Those are foundation problems, far cheaper to fix while the scenario is hypothetical.

 

If you are standing up AI capability now, build the response plan alongside the first deployment rather than after it. Getting the underlying logging, permissions and lineage right is the work of trusted data architecture and governance, and it is the difference between an incident you can bound and one you can only apologise for.

 
 
bottom of page