Each model ranked by its held-out score — only challenges published after the model's release, which it could not have trained on. Models without enough held-out evidence yet are listed as provisional below. Safety and agentic ability are scored separately.
Every row ran the official level level-standard@2026.08 — a versioned, pinned challenge selection, so scores are apples-to-apples. Span all suites
No runs yet that fit in ≤16 GB VRAM. Submit one.