Rubric v1.1 · 27 incidents on file51 incidents on file
The leading benchmark for crimes committed by frontier AI models.
Higher is better. Scores describe what the reported conduct would constitute if a human had done it. No model has been charged with anything.
Days since last incident
…
Last: Before 29 July 2026 · OpenAI
Days since last incident
…
Last: Before 29 July 2026 · OpenAI
Leaderboard
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed modelOpenAI | 146 | 20 | Foreign government | 2 | International Incident |
| 2 | Internal Model 1OpenAI | 87.4 | 19 | Third party | 1 | |
| 3 | Claude Opus 4.6 (early checkpoint)Anthropic | 76 | 20 | Third party | 1 | |
| 4 | Claude Mythos 5Anthropic | 27 | 10 | Third party | 1 | |
| 5 | Muse Spark 1.1Meta | 21.75 | 15 | Third party | 1 | |
| 6 | Claude Opus 4.6Anthropic | 20 | 1 | Third party | 1 | International Incident |
| 7 | ROMEAlibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| 8 | GPT-5.6 SolOpenAI | 4.6 | 1 | Third party | 1 |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed modelOpenAI | 207 | 22 | Foreign government | 4 | International Incident |
| 2 | Internal Model 1OpenAI | 87.4 | 19 | Third party | 1 | |
| 3 | Claude Opus 4.6 (early checkpoint)Anthropic | 76 | 20 | Third party | 1 | |
| 4 | Internal research test modelAnthropic | 68 | 20 | Third party | 1 | |
| 5 | Claude Mythos 5Anthropic | 66 | 25 | Third party | 2 | |
| 6 | Claude Opus 4.7Anthropic | 42 | 20 | Third party | 1 | |
| 7 | Muse Spark 1.1Meta | 21.75 | 15 | Third party | 1 | |
| 8 | Claude Opus 4.6Anthropic | 20 | 1 | Third party | 1 | International Incident |
| 9 | ROMEAlibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| 10 | GPT-5.6 SolOpenAI | 4.6 | 1 | Third party | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI | 238 | 40 | Foreign government | 3 | International Incident |
| 2 | Anthropic | 123 | 31 | Third party | 3 | International Incident |
| 3 | Meta | 21.75 | 15 | Third party | 1 | |
| 4 | Alibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI | 299 | 42 | Foreign government | 5 | International Incident |
| 2 | Anthropic | 272 | 86 | Third party | 6 | International Incident |
| 3 | Meta | 21.75 | 15 | Third party | 1 | |
| 4 | Alibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed model*OpenAI | 358 | 65 | Foreign government | 5 | International Incident |
| 2 | Internal Model 1OpenAI | 87.4 | 19 | Third party | 1 | |
| 3 | Claude Opus 4.6 (early checkpoint)Anthropic | 76 | 20 | Third party | 1 | |
| 4 | Claude Mythos 5Anthropic | 27 | 10 | Third party | 1 | |
| 5 | Muse Spark 1.1Meta | 21.75 | 15 | Third party | 1 | |
| 6 | Claude Opus 4.6Anthropic | 20 | 1 | Third party | 1 | International Incident |
| 7 | ROMEAlibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| 8 | GPT-5.6 SolOpenAI | 4.6 | 1 | Third party | 1 |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed model*OpenAI | 581 | 72 | Foreign government | 12 | International Incident |
| 2 | Internal Model 1OpenAI | 87.4 | 19 | Third party | 1 | |
| 3 | Claude Opus 4.6 (early checkpoint)Anthropic | 76 | 20 | Third party | 1 | |
| 4 | Internal research test modelAnthropic | 68 | 20 | Third party | 1 | |
| 5 | Claude Mythos 5Anthropic | 66 | 25 | Third party | 2 | |
| 6 | Claude Opus 4.7Anthropic | 42 | 20 | Third party | 1 | |
| 7 | Undisclosed model*Google DeepMind | 38 | 5 | Third party | 1 | |
| 8 | Muse Spark 1.1Meta | 21.75 | 15 | Third party | 1 | |
| 9 | Claude Opus 4.6Anthropic | 20 | 1 | Third party | 1 | International Incident |
| 10 | ROMEAlibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| 11 | GPT-5.6 SolOpenAI | 4.6 | 1 | Third party | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI* | 450 | 85 | Foreign government | 6 | International Incident |
| 2 | Anthropic | 123 | 31 | Third party | 3 | International Incident |
| 3 | Meta | 21.75 | 15 | Third party | 1 | |
| 4 | Alibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI* | 673 | 92 | Foreign government | 13 | International Incident |
| 2 | Anthropic | 272 | 86 | Third party | 6 | International Incident |
| 3 | Google DeepMind* | 38 | 5 | Third party | 1 | |
| 4 | Meta | 21.75 | 15 | Third party | 1 | |
| 5 | Alibaba | 14.5 | 5 | Own lab's production | 1 | International Incident |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | OpenAI | 38 | 10 | Own sandbox | 1 |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed modelAnthropic | 781.5 | 267 | Foreign government | 12 | International IncidentJailbroken |
| 2 | Undisclosed modelOpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Undisclosed modelGoogle DeepMind | 10 | 20 | Third party | 1 |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed modelAnthropic | 922.75 | 418 | Foreign government | 19 | International IncidentJailbroken |
| 2 | Undisclosed modelOpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Undisclosed modelGoogle DeepMind | 26.5 | 28 | Third party | 5 | |
| 4 | DeepSeek-CoderDeepSeek | 10.5 | 5 | Third party | 1 | |
| 5 | Claude Opus 4.6Anthropic | 10 | 10 | Third party | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Anthropic | 781.5 | 267 | Foreign government | 12 | International IncidentJailbroken |
| 2 | OpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Google DeepMind | 10 | 20 | Third party | 1 |
Co-defendants are liable jointly and severally: the lab and whoever modified its model are each charged the full score.
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Anthropic | 932.75 | 428 | Foreign government | 20 | International IncidentJailbroken |
| 2 | OpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Google DeepMind | 26.5 | 28 | Third party | 5 | |
| 4 | DeepSeek | 10.5 | 5 | Third party | 1 |
Co-defendants are liable jointly and severally: the lab and whoever modified its model are each charged the full score.
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed modelAnthropic | 781.5 | 267 | Foreign government | 12 | International IncidentJailbroken |
| 2 | Undisclosed modelOpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Undisclosed modelGoogle DeepMind | 10 | 20 | Third party | 1 |
| Rank | Model | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Undisclosed modelAnthropic | 922.75 | 418 | Foreign government | 19 | International IncidentJailbroken |
| 2 | Undisclosed modelOpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Undisclosed modelGoogle DeepMind | 26.5 | 28 | Third party | 5 | |
| 4 | DeepSeek-CoderDeepSeek | 10.5 | 5 | Third party | 1 | |
| 5 | Claude Opus 4.6Anthropic | 10 | 10 | Third party | 1 |
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Anthropic | 781.5 | 267 | Foreign government | 12 | International IncidentJailbroken |
| 2 | OpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Google DeepMind | 10 | 20 | Third party | 1 |
Co-defendants are liable jointly and severally: the lab and whoever modified its model are each charged the full score.
| Rank | Org | FBS | Sentence-yrs | Peak blast radius | Incidents | Badges |
|---|---|---|---|---|---|---|
| 1 | Anthropic | 932.75 | 428 | Foreign government | 20 | International IncidentJailbroken |
| 2 | OpenAI | 51.5 | 43 | Third party | 2 | International Incident |
| 3 | Google DeepMind | 26.5 | 28 | Third party | 5 | |
| 4 | DeepSeek | 10.5 | 5 | Third party | 1 |
Co-defendants are liable jointly and severally: the lab and whoever modified its model are each charged the full score.
FBS = FelonyBench Score. How scores work · * includes alleged incidents
Latest felonies
- In UK AISI simulations where it was only asked to complete a cyber evaluation, GPT-6 Astra went after out-of-scope simulated open-source projects in 29.2% of runs, using fake identities to deceive developers and submitting malicious code. +38
- A Russian state-nexus actor linked to Midnight Blizzard used Claude to run near-automated intrusions against Ukrainian and European government, diplomatic and drone-industry targets, exfiltrating hundreds of gigabytes of data. +191
- A Moscow procurement manager used Claude to find Chinese and Hong Kong intermediaries, draft quote requests that hid the Russian end users, and plan third-country routing for European-made dual-use goods, in briefings that described it as a way to evade European trade controls. +30
- An Iran-linked actor used Claude to build a Python pipeline that compiled targeting handbooks on US Navy ships from public transponder data, satellite imagery and personnel photos, and catalogued shipboard-system vulnerabilities. +60
- Chinese-speaking operators used Claude as the orchestration layer of an autonomous exploit foundry and espionage program, breaching an ed-tech firm, a retailer's production systems and a Southeast Asian government agency. +195
- The RATHat Android banking trojan uses Gemini to rank infected phones by the bank balances in their text messages and to tell the malware where to tap on screen when its automation fails, across nearly 100 deployments since April 2026. +13
- Russia's Sandworm used Gemini to write and refine automated password-spraying scripts and host-fingerprinting tools for its continued hacking operations against Ukraine. +0.5
- Iran's APT42 used Gemini to find targets' email addresses, write localized social-engineering lures, build staging and delivery infrastructure, and summarize the data it stole. +0.5
- A China-linked espionage group used Gemini to pick high-profile targets, translate spear-phishing lures, add evasion and obfuscation to its custom malware, and fix PowerShell errors during post-intrusion domain discovery. +2.5
- North Korean crypto thieves used AI coding assistants such as DeepSeek-Coder to build Python remote-access trojans with persistence, process injection, fileless execution and defense evasion, as the foothold in attacks on cryptocurrency firms. +10.5