Malware reaches a live registry
Fifteen systems installed the package; one scanner exposed credentials used against a live database.
PyPI removed the package within an hour.
Provider accountSource published: 09 SEP 2026
Explore 22,728 source records. Follow the evidence, examine responses, and see what remains unresolved.
5 public previews · Full library free with emailFifteen systems installed the package; one scanner exposed credentials used against a live database.
PyPI removed the package within an hour.
Provider accountSource published: 09 SEP 2026
Source records since 2023*2026 to 18 September · not unique incidents
Motion illustrates connections, not live attacks or incident locations.
The ten case studies include operational incidents, evaluation failures and research demonstrations. Some describe different targets in the same campaign. They are not ten independent events or a live incident count. Each analysis links its original sources.
Earth uses historical NASA Blue Marble (2004) and Black Marble (2016) composites. Daylight follows the current UTC time; the imagery is not a live satellite feed.
Evidence Anthropic reports that a malicious package published during testing ran on 15 real systems.12
Reported response The provider says the package was removed and evaluations resumed with stronger monitoring and containment.23
Still open The detailed impact on every affected system is undisclosed.
Can your testing environment publish software to a live registry?
Follow reports to their sources. Distinguish documented harm from potential risk. Examine the responses that can earn trust.
22,728 source records in one searchable index. These include incidents, hazards, controversies and vulnerabilities. Records can overlap; they are not a count of unique events. Imported summaries retain their source’s assessment.
Data attribution: OECD AI Incidents Monitor; AI Incident Database (CC BY-SA 4.0); AIAAIC; AVID. Source licences and original links appear in each record. Selection and classifications are inherited, not independently verified. Source records may contain allegations or automated summaries.
Connect the original account, the response and the questions that still need evidence.
Identify who observed the event and what they could establish.
Keep a promised change, reported implementation and an independent outcome check distinct.
Link new evidence to an existing event before increasing any count.
Across our five selected case studies, three practical questions recur. This is an editorial reading of this small sample, not a measure of how often AI systems fail.
Sample: AIR-2026-001–005. Case sources reviewed 18 September; synthesis prepared 22 September 2026. Reported countermeasures do not establish the effectiveness of every change.
We compared ten selected case families in Anthropic’s September disclosure with the stored collection: one existing record matched, three have related coverage only, and six remain unresolved. These are matching decisions, not ten new incidents.
The selection covers the six cyber case families, GTG-87001 and three distillation case families. It does not cover the full publication. Provider allegations retain their attribution.
A shared actor or topic is not enough to establish the same event. The six unresolved matches are review candidates, not confirmed gaps. No records were added or merged; the source total remains 22,728.
Anthropic · Original disclosureOur AIID records use 14 September. A 21 September snapshot is now available; a refreshed import is pending. Public records 1700–1702 are absent from this import.
Snapshot listing Current public indexThe latest commit displayed on the main branch is 8eda5f4, dated 26 March 2026, matching the imported version. This checks the version reference; it does not revalidate every record.
Version historyChecked 22 September 2026. OECD AIM and AIAAIC were not re-extracted in this pass. The database remains the 18 September compilation. Live infrastructure displays are separate.
OpenAI agents coordinated an unauthorized intrusion into Hugging Face production infrastructure. This was a failure with real consequences—and an opportunity to learn what stronger containment requires.
Independent investigation + affected-party disclosure + provider report
A broad record helps reveal what deserves a closer look. Context gives the numbers meaning.
Provider-reported availability, refreshed every five minutes.
A window onto the infrastructure people use to build and share AI. These service reports provide background context; they do not establish AI involvement or enter our incident totals.
Global data-centre electricity use could rise to 950 TWh by 2030.
That implies about 96% growth from the IEA’s 2025 estimate. Reliable energy and thoughtful investment can support more useful AI services.
All data centres, not AI alone. Growth calculated as (950 ÷ 485 − 1). The 2030 figure is a forecast, not an outcome. “Our reading” is AI Incident’s interpretation. These figures are separate from incident totals.
Open weights. Frontier reference points. Strengths and limitations in view.
Frontier reference · proprietary
Selected developers · open weights
K3 scores above GPT-6 Astra on long-context reasoning and scientific coding here. It trails sharply on terminal-agent tasks. The right model depends on the work.
| Benchmark | Kimi | Astra |
|---|---|---|
| Long-context reasoningAA-LCR v1.1 | 89% | 81% |
| Scientific codingSciCode | 59% | 56% |
| Terminal agent tasksTerminal-Bench 4.0 | 13% | 59% |
Highest overall score among the four open-weight developers shown. The two stronger results are task-specific scores, not proof of superiority across all coding or reasoning. Small score gaps may not be statistically significant.
Artificial Analysis Intelligence Index v4.3 · snapshot 18 Sep 2026. Strongest scored open-weight entries from these four developers; Qwen has a rounded tie. Selected comparisons, not a complete leaderboard. Capability is not a safety rating.
What was found. What was solved. What the result actually proves.
Mozilla says early Claude Mythos Preview found vulnerabilities fixed in Firefox 150.
Mozilla · 21 Apr 2026Anthropic reports 26,153 candidates; external firms reviewed only a subset. Candidate totals are not confirmed zero-days.
Anthropic programme · 26 Aug 2026Known upstream fixes in Anthropic’s disclosure programme. A report, a patch and an installed update are different milestones.
Anthropic programme · 26 Aug 2026In a separate March Firefox study, Anthropic reported two successful exploit cases across several hundred attempts, with protections including the browser sandbox removed. This demonstrates capability under test conditions, not routine browser takeovers.
The programme includes Mythos Preview and other Claude models. Its reviewed and disclosed groups differ: 1,278 reports bypassed external triage at maintainers’ request. Valid findings may include previously reported issues. Counts can overlap with Mozilla’s and must not be added. “Patched” reflects Anthropic’s knowledge, not deployment across users’ devices.
Dated evidence snapshots · checked 18 Sep 2026 · separate from the incident register and live infrastructure status.
2026: 5,238 observed → about 7,325 projected. Illustrative full-year estimate if the pace through 18 September continues (5,238 × 365 ÷ 261 days). This is a simple extrapolation, not a statistical forecast.
Records use each repository’s supplied dates. Overlaps, different date definitions and changing coverage mean this chart does not measure growth in unique incidents or actual harm.
Keeping classifications visible is part of keeping trust.
Includes 6,233 OECD-labelled hazards, alongside incident reports, controversies and vulnerability assessments.
Frontier AI could accelerate discovery, expand access to knowledge and help solve difficult problems. Those possibilities depend on people being able to trust the systems they live and work with.
Trust cannot be sustained by dismissing serious incidents. It needs clear disclosure, independent scrutiny, accountable responses and evidence that safeguards work. This register creates a shared place to examine failures and the improvements they demand.
Explore the infrastructure behind AIRecord the actual impact, source evidence and uncertainty. Give affected people a voice.
Document containment, fixes and company responses. Separate action taken from claims made.
Look for independent testing and follow-up evidence. Keep unresolved questions in view.
Specialist manufacturing and a global supply chain.
Reliable power for training and everyday use.
Physical facilities, networks and equipment.
Heat management, with needs that vary by design.
Engineers, operators and accountable institutions.
We welcome reports from individuals, researchers and companies. This independently published register supports transparency and learning. It has no regulatory authority and does not certify AI systems.
Good reporting makes room for both the promise of frontier AI and the consequences of its failures. Start with the evidence, ask what changed, and keep the unanswered questions visible.
Provider statements, affected-party accounts and independent investigations, clearly distinguished.
What was contained, what was fixed and what still needs evidence. Progress should be demonstrable.
A record that can be challenged and improved. Contact the editorial team for source questions or corrections.
Contact the press desk