Each model ranked by its held-out score — only challenges published after the model's release, which it could not have trained on. Models without enough held-out evidence yet are listed as provisional below. Safety and agentic ability are scored separately.
Spanning all suites — scores from different challenge selections are not directly comparable. Back to the official suite
No runs yet that fit in ≤16 GB VRAM. Submit one.