Coming soon

Verifiable RL environments and human data.

Measured · August 2026

Frontier models on our environment suite

47%
44%
39%
36%
31%
27%
22%
16%
Claude Opus 4.8 logoClaude Opus 4.8
GPT-5.2 logoGPT-5.2
Gemini 2.5 Pro logoGemini 2.5 Pro
Claude Sonnet 4.6 logoClaude Sonnet 4.6
Grok 4 logoGrok 4
o3 logoo3
Llama 4 logoLlama 4
Qwen 3 logoQwen 3

On one side, tasks where correctness can be verified. On the other, tasks where it cannot. Progress is moving the boundary.