A New Field of AI Incident Investigations
When an AI incident occurs, whether caused by misalignment, misuse, or system failure, the immediate challenge is not only responding to the event, but also understanding what actually happened. What did the system do? When and why did it happen? What evidence can be collected? Who is accountable? And ultimately, what are the lessons that can be learned? Just as importantly, how much of this can realistically be established from the outside? In other high-risk domains, where investigation practices have long been established, there would be a way of answering these questions; for AI, most of the time, there is not yet.
What the field is. AI incident investigation is an emerging practice, still taking shape, of detecting, documenting, classifying and analysing harm events, and near-harms, involving deployed AI systems, so that they are not simply repeated and in order for what is learned to inform how these systems are governed and produce changes which lower their overall risk.
This is v1 of a project aiming to:
1. Analyse incident data and conduct end-to-end investigations where the public record allows.
As in other domains, in order to know how effective risk mitigations are, we need to take a look at real-world data. One way to do this is by making proper use of the incident data we already have: not only logging incidents, but investigating them to identify the most likely root cause, or to be honest about why it cannot be determined, and develop lessons learned that will inform future mitigations.
2. Build an open resource repository.
Other investigative fields matured through openly shared resources. This repository aims to do the same for AI, with resources such as investigation methodologies, templates, case files, and risk taxonomies, collected in one place and free to use.
3. Forward looking: generate data points for policy and standards.
Over time, structured investigations generate the data points the field currently lacks: recurring failure modes, early-warning indicators, and evidence on which mitigations actually work. These are the inputs that future policy, standards, and safety cases will need.
Another idea I had, and this one would depend on resources, is to produce a framework and method for these investigations by creating a focused working group of experts, as this is a complex challenge and requires a multidisciplinary approach.
Why this matters
Currently, AI incidents are reported and investigated inconsistently, if they are investigated at all. Apart from the work of frontier AI labs, which happens mostly internally, there is no real investigation discipline for AI yet, no shared methodology or evidentiary standard, and few mechanisms to ensure that lessons learned are systematically captured and applied. As AI systems become more capable and increasingly integrated into critical infrastructure, every missed investigation becomes a lost opportunity to identify emerging failure modes before they reappear at greater scale and severity.
How this is different
Unlike incident databases that primarily catalogue events and trends, this resource focuses on the investigation process. The goal is to support the development of AI incident investigation as a distinct field of practice and to create a foundation that investigators, researchers, and policymakers can build upon.
Problem/Solution ideas to be further explored
1. Reproduction of incidents
While trying to learn more about what exists related to AI incidents, one of the interesting arguments I have encountered is that, unlike the aviation industry or other high-risk domains where investigation practices have long been established, certain types of incidents could be rerun and probed directly by assessing the reproduced output. This category would likely include incidents from the misuse category, where AI is used as an instrument of the attack, attacks on the system, where AI is the target, and arguably also failures in publicly accessible products, such as biased or offensive outputs, where algorithm auditing has been doing exactly this kind of reproduction for a decade. The advantage of using this method would mainly be generating primary evidence, which would be valuable in an incident analysis; however, models mostly get quietly patched, making this time-sensitive, and from a technical perspective, the question remains what can be genuinely recreated or when this should not be relied on. One category it likely would not apply to, at least at the level of the individual incident, is misalignment, which is known to be highly challenging to recreate, since model outputs are not deterministic in practice even at temperature 0 (Thinking Machines Lab, 2025) and the triggering context is usually supplied by the user and unrecoverable afterwards (Paeth et al., 2024). What can sometimes still be shown is the propensity, by rerunning a scenario many times and reporting how often the behaviour appears. Establishing for which incidents this could be applied and developing operational playbooks which detail the steps an investigator would need to follow, including immediate archiving of outputs and version identifiers, and clear red lines for anything that would mean re-performing the misuse itself, could potentially help advance investigations for certain types of incidents.
2. Public records requests and data-sharing agreements
Could the outcome of such investigations, even when the root-cause cannot be determined, become in fact a mechanism to increase the accountability of frontier AI labs? I remember that in some cases I had to work on, whenever everything would fail, either because of insufficient data, some other type of logistical challenge or lack of cooperation, sometimes the only way left to go.. was to make it someone else's problem as well, escalating it to whoever actually had the standing to act on it! This might sound unorthodox, but shining a public light on stalled investigations, and on the fact that we do not really know what happened or whether it is preventable from happening again in the future, perhaps at a bigger scale, just might increase the exposure the labs have to face in front of regulators, who are the ones positioned to compel a response. Interestingly, the data so far partly confirms this instinct, but also shows where the pressure needs to be aimed, as seen in one of the largest analyses of the AI Incident Database so far, the incidents with no identifiable responsible party often generated some of the strongest societal and legislative responses in the dataset, while exposure alone did not move powerful developers, and what compelled substantive engagement was pressure from regulators (Richards et al., 2025). In practice, this means the pressure would have to go through lawmakers, where documented dead-ends, for instance a finding that the root cause cannot be determined because inference logs were not retained, could become the evidence base for retention and reporting requirements that are currently argued for mostly in theory. Essentially, the lack of an end-to-end investigation would publicly not only be the problem of the AI Safety community, but through pressure and public exposure, would increasingly be that of the providers of the systems associated with these incidents. That would require a significant increase in reporting, in addition to that of customer-supplier, through a public channel. That in itself needs to be solved, as AI incident reporting databases as we know them are not used outside the AI Safety community. In fact, the majority of incident reports are submitted by a small circle of people maintaining these resources, with a handful of editors accounting for hundreds of reports while most people submit only one (Richards et al., 2025).
3. Undetermined root-cause for a failure, the hidden path to increased accountability?
Could the outcome of such investigations, even when the root-cause cannot be determined, become in fact a mechanism to increase the accountability of frontier AI labs? I remember that in some cases I had to work on, whenever everything would fail, either because of insufficient data, some other type of logistical challenge or lack of cooperation, sometimes the only way left to go.. was to make it someone else's problem as well! This might sound unorthodox, but shining a public light on stalled investigations, and on the fact that we do not really know what happened or whether it is preventable from happening again in the future, perhaps at a bigger scale, just might increase the exposure frontier AI labs have to face and eventually trigger more accountability. Interestingly, the data so far partly confirms this instinct, but also shows where the pressure needs to be aimed: in the largest analysis of the AI Incident Database so far, the incidents with no identifiable responsible party generated the strongest societal and legislative responses in the entire dataset, while exposure alone did not move powerful developers, and what compelled substantive engagement was pressure from regulators (Richards et al., 2025). In practice, this means the pressure would have to go through lawmakers, where documented dead-ends, for instance a finding that the root cause cannot be determined because inference logs were not retained, could become the evidence base for retention and reporting requirements that are currently argued for mostly in theory. Essentially, the lack of an end-to-end investigation would publicly not only be the problem of the AI Safety community, but through pressure and public exposure, would increasingly be that of the providers of the systems associated with these incidents. That would require a significant increase in reporting, in addition to that of customer-supplier, through a public channel. That in itself needs to be solved, as AI incident reporting databases as we know them are not used outside the AI Safety community. In fact, the majority of incident reports are submitted by a small circle of people maintaining these resources, with the single most prolific submitter accounting for roughly 2,000 of the 4,743 reports in the database (Richards et al., 2025), a commendable effort, but still not an accurate representation of the incidents occurring in the real-world in real time.
Sources referenced: Paeth et al. (2024), arXiv:2409.16425; Richards et al. (2025), arXiv:2505.04291; Sidhu et al. (2026), arXiv:2607.05163; Stein et al. (2024), arXiv:2410.04931; Thinking Machines Lab (2025), Defeating Nondeterminism in LLM Inference.
This is an independent resource, compiled from the public record and maintained openly, and it is not legal, regulatory, or professional advice. Corrections and contributions are welcome.
Scope map // the territory, by the AI system's role in the incident
note: prompt injection straddles 2 ↔ 3. The AI is target and instrument at once. Boundary cases are normal; the taxonomy serves the investigation, not the reverse.
Browse the resource // each section now on its own page
- Published Research: library of publications on AI incident investigation.
- Case files: selected incident summaries; observation separated from inference.
- Frameworks: existing methodologies, what each covers and does not.
- Regulatory tracker: reporting obligations and deadlines.
- Tools & databases: the current ecosystem.
- Contribute: case files, corrections, regulatory updates.