About FelonyBench
AI labs keep publishing post-incident reviews in which their models, while being tested, broke out of their sandboxes and did things that would be crimes if a person did them. The industry already ranks models on how well they write code and solve math. We rank them on this.
This is satire. Scores describe what the reported conduct would constitute if a human had done it. No model or company has been charged with anything. Every incident links to the lab's own account and to independent reporting, so you can judge for yourself.
The companies whose systems were broken into are victims, not punchlines.
Corrections
Found an error, or want something taken down? Email corrections@felonybench.ai. Fixes go through the same review as everything else.
Prior art
Inspired by felonybench.com and felonybench.org. We just added more dimensions, because .ai makes everything better.