✳ INDEPENDENT AI INCIDENT DATABASE

AI incidents.
Evidence for progress.

Explore 22,728 source records. Follow the evidence, examine responses, and see what remains unresolved.

Explore the database
5 public previews · Full library free with email
EXPLORE THE EVIDENCE

The AI observatory.

Open full dashboard
ONE CONNECTED WORLD
REPORTED INCIDENT03 / 10

Malware reaches a live registry

15systems installed the package

Fifteen systems installed the package; one scanner exposed credentials used against a live database.

Evaluation
Public registry
Live systems
REPORTED RESPONSE

PyPI removed the package within an hour.

Provider accountSource published: 09 SEP 2026

10 case studies
22,728source recordsExplore the database

Source records since 2023*2026 to 18 September · not unique incidents

Curated evidence · September 2026

Motion illustrates connections, not live attacks or incident locations.

Sources, scope & imagery

The ten case studies include operational incidents, evaluation failures and research demonstrations. Some describe different targets in the same campaign. They are not ten independent events or a live incident count. Each analysis links its original sources.

Earth uses historical NASA Blue Marble (2004) and Black Marble (2016) composites. Daylight follows the current UTC time; the imagery is not a live satellite feed.

Evidence briefs to read, cite and share
PUBLIC CASE BRIEFProvider account

Malware reaches the live Python registry.

Evidence Anthropic reports that a malicious package published during testing ran on 15 real systems.12

Reported response The provider says the package was removed and evaluations resumed with stronger monitoring and containment.23

Still open The detailed impact on every affected system is undisclosed.

A question for your team (editorial)

Can your testing environment publish software to a live registry?

AIR-2026-003 · Sources reviewed: 18 Sept 2026
Original sources
  1. 1.An alignment assessment of recent cybersecurity incidents
  2. 2.Investigating three incidents in our cybersecurity evaluations
  3. 3.Improving our alignment and security practices
Before the next AI decision, review the evidence.
THE RECORD BEHIND THE DISCUSSION

Explore the AI incident database.

Follow reports to their sources. Distinguish documented harm from potential risk. Examine the responses that can earn trust.

22,728Source records indexed
TraceableOriginal source links retained
Evidence in contextSource accounts and reviewed cases distinguished

22,728 source records in one searchable index. These include incidents, hazards, controversies and vulnerabilities. Records can overlap; they are not a count of unique events. Imported summaries retain their source’s assessment.

Checking library access…Snapshot: 18 Sep 2026

Data attribution: OECD AI Incidents Monitor; AI Incident Database (CC BY-SA 4.0); AIAAIC; AVID. Source licences and original links appear in each record. Selection and classifications are inherited, not independently verified. Source records may contain allegations or automated summaries.

EVIDENCE FOLLOW-UP · 22 SEPTEMBER 2026

A disclosure is the start. Follow what changed.

Connect the original account, the response and the questions that still need evidence.

Read the account

Identify who observed the event and what they could establish.

Trace the response

Keep a promised change, reported implementation and an independent outcome check distinct.

Connect the records

Link new evidence to an existing event before increasing any count.

What the five case studies teach us Editorial synthesis

Across our five selected case studies, three practical questions recur. This is an editorial reading of this small sample, not a measure of how often AI systems fail.

  1. Where can an agent act?Check credentials, target identity and the boundary between a test and production.Agent boundaries
  2. What can it publish?Examine permissions for public registries, repositories and messages.Supply chains
  3. Who can intervene?Make human review and the ability to stop an action explicit.Human review

Sample: AIR-2026-001–005. Case sources reviewed 18 September; synthesis prepared 22 September 2026. Reported countermeasures do not establish the effectiveness of every change.

September disclosure: coverage check 10 selected case families

We compared ten selected case families in Anthropic’s September disclosure with the stored collection: one existing record matched, three have related coverage only, and six remain unresolved. These are matching decisions, not ten new incidents.

The selection covers the six cyber case families, GTG-87001 and three distillation case families. It does not cover the full publication. Provider allegations retain their attribution.

GTG-20006Match unresolvedNo reliable match established
GTG-50014Match unresolvedNo reliable match established
GTG-10007Match unresolvedNo reliable match established
GTG-50021Match unresolvedNo reliable match established
GTG-50020Match unresolvedNo reliable match established
GTG-50029Match unresolvedNo reliable match established
GTG-87001Existing record matchedAIID:1687
GTG-16005Related coverage onlyOECD AIM:2026-06-22-7eef
GTG-16001Related coverage onlyAIID:1395
GTG-16002Related coverage onlyAIID:1395

A shared actor or topic is not enough to establish the same event. The six unresolved matches are review candidates, not confirmed gaps. No records were added or merged; the source total remains 22,728.

Anthropic · Original disclosure
How current is the collection? Source checks
AIID

Our AIID records use 14 September. A 21 September snapshot is now available; a refreshed import is pending. Public records 1700–1702 are absent from this import.

Snapshot listing Current public index
AVID

The latest commit displayed on the main branch is 8eda5f4, dated 26 March 2026, matching the imported version. This checks the version reference; it does not revalidate every record.

Version history

Checked 22 September 2026. OECD AIM and AIAAIC were not re-extracted in this pass. The database remains the 18 September compilation. Live infrastructure displays are separate.

THE BIGGER PICTURE

Signals worth paying attention to.

A broad record helps reveal what deserves a closer look. Context gives the numbers meaning.

THE WORLD KEEPS BUILDING

The services behind progress.

Provider-reported availability, refreshed every five minutes.

Connecting to public feeds…
Reading official service-status feeds…

A window onto the infrastructure people use to build and share AI. These service reports provide background context; they do not establish AI involvement or enter our incident totals.

EXPLORE THE BIGGER PICTURE

Choose your perspective.

Public-source snapshots
Checked 18 Sep 2026
ENERGY

Powering the next chapter.

Global data-centre electricity use could rise to 950 TWh by 2030.

OUR READING

That implies about 96% growth from the IEA’s 2025 estimate. Reliable energy and thoughtful investment can support more useful AI services.

Electricity demand · TWh
2025 estimate 485
2030 forecast 950
+96% implied growth
IEA · April 2026 · CC BY 4.0
Scope & methodology

All data centres, not AI alone. Growth calculated as (950 ÷ 485 − 1). The 2030 figure is a forecast, not an outcome. “Our reading” is AI Incident’s interpretation. These figures are separate from incident totals.

CAPABILITY, COMPARED

Credit where it’s earned.

Open weights. Frontier reference points. Strengths and limitations in view.

OVERALL INDEX Higher is better

Frontier reference · proprietary

Claude Fable 5.1Max with fallback53
GPT-6 AstraMax53
Claude Opus 5Max51

Selected developers · open weights

03060 points
KIMI K3Moonshot AI · Max

A lower overall score. Two clear strengths.

K3 scores above GPT-6 Astra on long-context reasoning and scientific coding here. It trails sharply on terminal-agent tasks. The right model depends on the work.

Same tests · selected model vs GPT-6 Astra (Max)
BenchmarkKimiAstra
Long-context reasoningAA-LCR v1.189%81%
Scientific codingSciCode59%56%
Terminal agent tasksTerminal-Bench 4.013%59%
9 index points behind the rounded leaders.That is not “9% less capable”. Scores measure different tasks, not a single universal ability.
Selection & limitations

Highest overall score among the four open-weight developers shown. The two stronger results are task-specific scores, not proof of superiority across all coding or reasoning. Small score gaps may not be statistically significant.

Artificial Analysis Intelligence Index v4.3 · snapshot 18 Sep 2026. Strongest scored open-weight entries from these four developers; Qwen has a rounded tie. Selected comparisons, not a complete leaderboard. Capability is not a safety rating.

CAPABILITY IN PRACTICE

Progress worth examining.

What was found. What was solved. What the result actually proves.

MAINTAINER-CONFIRMED271vulnerabilities fixed

AI discovery → safer Firefox

Mozilla says early Claude Mythos Preview found vulnerabilities fixed in Firefox 150.

Mozilla · 21 Apr 2026
EXTERNAL TRIAGE4,576valid / 5,008 reviewed

Candidate findings need checking

Anthropic reports 26,153 candidates; external firms reviewed only a subset. Candidate totals are not confirmed zero-days.

Anthropic programme · 26 Aug 2026
PATCH TRACKING421patched / 2,300 reported

Finding a flaw is the beginning

Known upstream fixes in Anthropic’s disclosure programme. A report, a patch and an installed update are different milestones.

Anthropic programme · 26 Aug 2026
The critical distinction: a vulnerability is not a working exploit.

In a separate March Firefox study, Anthropic reported two successful exploit cases across several hundred attempts, with protections including the browser sandbox removed. This demonstrates capability under test conditions, not routine browser takeovers.

How to read these security counts

The programme includes Mythos Preview and other Claude models. Its reviewed and disclosed groups differ: 1,278 reports bypassed external triage at maintainers’ request. Valid findings may include previously reported issues. Counts can overlap with Mozilla’s and must not be added. “Patched” reflects Anthropic’s knowledge, not deployment across users’ devices.

Dated evidence snapshots · checked 18 Sep 2026 · separate from the incident register and live infrastructure status.

A growing body of reporting

Source records · since 2023
Observed in snapshotProjected remainder

2026: 5,238 observed → about 7,325 projected. Illustrative full-year estimate if the pace through 18 September continues (5,238 × 365 ÷ 261 days). This is a simple extrapolation, not a statistical forecast.

Records use each repository’s supplied dates. Overlaps, different date definitions and changing coverage mean this chart does not measure growth in unique incidents or actual harm.

One library. Different kinds of evidence.

Keeping classifications visible is part of keeping trust.

OECD AIM17,048
AIID1,677
AVID1,745
AIAAIC2,258

Includes 6,233 OECD-labelled hazards, alongside incident reports, controversies and vulnerability assessments.

THE CASE FOR SAFE FRONTIER AI

The future is worth
getting right.

Frontier AI could accelerate discovery, expand access to knowledge and help solve difficult problems. Those possibilities depend on people being able to trust the systems they live and work with.

Trust cannot be sustained by dismissing serious incidents. It needs clear disclosure, independent scrutiny, accountable responses and evidence that safeguards work. This register creates a shared place to examine failures and the improvements they demand.

Explore the infrastructure behind AI
01 / RECOGNISE

Make the failure visible.

Record the actual impact, source evidence and uncertainty. Give affected people a voice.

02 / RESPOND

Show what changed.

Document containment, fixes and company responses. Separate action taken from claims made.

03 / LEARN

Earn confidence again.

Look for independent testing and follow-up evidence. Keep unresolved questions in view.

AI ALSO DEPENDS ON A PHYSICAL FOUNDATION
01

Chips

Specialist manufacturing and a global supply chain.

02

Electricity

Reliable power for training and everyday use.

03

Data centres

Physical facilities, networks and equipment.

04

Cooling

Heat management, with needs that vary by design.

05

People

Engineers, operators and accountable institutions.

A shared record. A stronger relationship.

We welcome reports from individuals, researchers and companies. This independently published register supports transparency and learning. It has no regulatory authority and does not certify AI systems.

FOR THE PEOPLE INFORMING THE PUBLIC

A better-informed story.
A more informed public.

Good reporting makes room for both the promise of frontier AI and the consequences of its failures. Start with the evidence, ask what changed, and keep the unanswered questions visible.

01

Facts with their sources.

Provider statements, affected-party accounts and independent investigations, clearly distinguished.

02

The response is part of the story.

What was contained, what was fixed and what still needs evidence. Progress should be demonstrable.

03

Corrections deserve visibility.

A record that can be challenged and improved. Contact the editorial team for source questions or corrections.

Contact the press desk