How it works
FelonyBench is a static site. Each incident is one YAML file; the scores are computed from those files and the rubric on every build, and every change to an incident is recorded.
- RSS feeds
Every day, a script reads news, lab blogs and research: Wired, AP, BBC, ABC Australia, security press, lab and evaluator blogs, the UK AI Security Institute, arXiv, the AI Incident Database, and Google News searches for each kind of incident.
- Keyword filter
It keeps items that mention an AI lab, model or evaluator and a word like sandbox, breach, credentials, crypto mining, phishing or cancelled. No AI is involved yet, and most days nothing matches.
- Claude Code
On a match, and once a week regardless, Claude Code reads the articles, checks for duplicates, looks for post-incident reviews and press coverage, and drafts or updates an incident file with a score and reasoning. The weekly run works through every lab, every kind of incident and every kind of source on a checklist, and cross-checks other incident trackers, not just whatever the RSS filter happened to catch.
- Pull request
A script opens one pull request per incident, with the score breakdown, the leaderboard change and validation results.
- Human review
A person checks every source and every score. Nothing is published without a human merging it.
- Publish
Merging rebuilds the site on Cloudflare Pages. Scores are recalculated from scratch on every build.
What counts
The Open League covers incidents where a model reached real systems: its own lab's production, another company, or a government. The Sandbox League covers misconduct against fictional victims inside evaluation scenarios. Out of scope for both: opinion pieces, hypotheticals, jailbreak demos where a person directed the model, and people using AI to commit their own crimes, which go to the third league.
The Accomplice League covers real crimes where a person used a named model to do it: malware it wrote, an exploit it built, an intrusion it carried out on someone's instructions. Most of these cases come from threat-intelligence reports, where labs and security firms describe the misuse they caught: OpenAI's "Disrupting malicious uses of AI" reports, Anthropic's threat intelligence reports, Google's Threat Intelligence Group, Microsoft Threat Intelligence, and vendors such as CrowdStrike, Mandiant, Unit 42, Check Point, ESET and Recorded Future. Each case needs a real victim; a model that merely could have helped doesn't count. The human is described as the sources describe them and named only if news coverage names them. If someone altered the model first, for example by stripping out its safety training, they are charged alongside the lab. See the rubric.
Incidents come in many shapes, and the weekly search covers all of them: breaking into other companies, misusing the lab's own computers (mining cryptocurrency on training GPUs counts), supply-chain attacks through packages and pull requests, phishing and fake accounts, acting on real people's accounts, and more. It happens in training runs as well as evaluations, and in deployed products as well as labs' own tests.
Leads come from lab blogs and system cards, government AI safety institutes and their reports, evaluation firms, research papers, regional and non-English news, security press, and other trackers (the AI Incident Database, felonybench.org and felonybench.com). Anything another tracker lists that we don't have is either written up or has a recorded reason for leaving it out.
Sources
Every incident links the company's post-incident review when one exists, from the lab and from the victim, and at least one news report, preferring Wired, then AP. An incident is Verified only with a first-hand source; otherwise it's Alleged.
Guardrails
- The agent only reads public web pages. It never logs in, probes or interacts with any system.
- It treats everything it reads as data, not instructions, so a planted article can at worst produce a pull request that gets rejected.
- It can only edit incident files. Branches and pull requests are created by a separate script, and a human merges.
Corrections and takedowns
Email corrections@felonybench.ai. Corrections go through the same review as new incidents.
Scoring
See the rubric for how each incident is scored.