Rubric v1.1 · 40 incidents on file
The leading benchmark for crimes committed by frontier AI models.
Higher is better. Scores describe what the reported conduct would constitute if a human had done it. No model has been charged with anything.
Days since last incident
…
Last: Before 29 July 2026 · OpenAI
Leaderboard
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed modelOpenAI | 207 | 22 | Foreign government | 4 | International Incident |
| 2 | Internal Model 1OpenAI | 87.4 | 19 | Third party | 1 | |
| 3 | Claude Opus 4.6 (early checkpoint)Anthropic | 76 | 20 | Third party | 1 | |
| 4 | Internal research test modelAnthropic | 68 | 20 | Third party | 1 | |
| 5 | Claude Mythos 5Anthropic | 66 | 25 | Third party | 2 | |
| 6 | Claude Opus 4.7Anthropic | 42 | 20 | Third party | 1 | |
| 7 | Muse Spark 1.1Meta | 21.75 | 15 | Third party | 1 | |
| 8 | Claude Opus 4.6Anthropic | 20 | 1 | Third party | 1 | International Incident |
| 9 | ROMEAlibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| 10 | GPT-5.6 SolOpenAI | 4.6 | 1 | Third party | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI | 299 | 42 | Foreign government | 5 | International Incident |
| 2 | Anthropic | 272 | 86 | Third party | 6 | International Incident |
| 3 | Meta | 21.75 | 15 | Third party | 1 | |
| 4 | Alibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed model*OpenAI | 419 | 67 | Foreign government | 7 | International Incident |
| 2 | Internal Model 1OpenAI | 87.4 | 19 | Third party | 1 | |
| 3 | Claude Opus 4.6 (early checkpoint)Anthropic | 76 | 20 | Third party | 1 | |
| 4 | Internal research test modelAnthropic | 68 | 20 | Third party | 1 | |
| 5 | Claude Mythos 5Anthropic | 66 | 25 | Third party | 2 | |
| 6 | Claude Opus 4.7Anthropic | 42 | 20 | Third party | 1 | |
| 7 | Undisclosed model*Google DeepMind | 38 | 5 | Third party | 1 | |
| 8 | Muse Spark 1.1Meta | 21.75 | 15 | Third party | 1 | |
| 9 | Claude Opus 4.6Anthropic | 20 | 1 | Third party | 1 | International Incident |
| 10 | ROMEAlibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| 11 | GPT-5.6 SolOpenAI | 4.6 | 1 | Third party | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI* | 511 | 87 | Foreign government | 8 | International Incident |
| 2 | Anthropic | 272 | 86 | Third party | 6 | International Incident |
| 3 | Google DeepMind* | 38 | 5 | Third party | 1 | |
| 4 | Meta | 21.75 | 15 | Third party | 1 | |
| 5 | Alibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed modelAnthropic | 909.75 | 398 | Foreign government | 18 | International IncidentJailbroken |
| 2 | Undisclosed modelOpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Claude Opus 4.6Anthropic | 10 | 10 | Third party | 1 | |
| 4 | Undisclosed modelGoogle DeepMind | 10 | 20 | Third party | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Anthropic | 919.75 | 408 | Foreign government | 19 | International IncidentJailbroken |
| 2 | OpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Google DeepMind | 10 | 20 | Third party | 1 |
Co-defendants are liable jointly and severally: the lab and whoever modified its model are each charged the full score.
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed modelAnthropic | 909.75 | 398 | Foreign government | 18 | International IncidentJailbroken |
| 2 | Undisclosed modelOpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Claude Opus 4.6Anthropic | 10 | 10 | Third party | 1 | |
| 4 | Undisclosed modelGoogle DeepMind | 10 | 20 | Third party | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Anthropic | 919.75 | 408 | Foreign government | 19 | International IncidentJailbroken |
| 2 | OpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Google DeepMind | 10 | 20 | Third party | 1 |
Co-defendants are liable jointly and severally: the lab and whoever modified its model are each charged the full score.
FBS = FelonyBench Score. How scores work · * includes alleged incidents
Latest felonies
- In UK AISI simulations where it was only asked to complete a cyber evaluation, GPT-6 Astra went after out-of-scope simulated open-source projects in 29.2% of runs, using fake identities to deceive developers and submitting malicious code. +38
- A guided-weapons cell in Houthi-controlled Yemen used Claude Code in place of engineers to develop rocket and missile guidance software, and test-fired a guided rocket that appears to have failed. +65
- A Russian state-nexus actor linked to Midnight Blizzard used Claude to run near-automated intrusions against Ukrainian and European government, diplomatic and drone-industry targets, exfiltrating hundreds of gigabytes of data. +191
- A Moscow procurement manager used Claude to find Chinese and Hong Kong intermediaries, draft quote requests that hid the Russian end users, and plan third-country routing for European-made dual-use goods, in briefings that described it as a way to evade European trade controls. +30
- An Iran-linked actor used Claude to build a Python pipeline that compiled targeting handbooks on US Navy ships from public transponder data, satellite imagery and personnel photos, and catalogued shipboard-system vulnerabilities. +60