Models · The Decoder ·
Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
Anthropic's Claude Opus 5 scored 30.2% on ARC-AGI-3, surpassing the reported 7.8% record held by GPT-5.6 Sol. Benchmark developers said the model independently formulated reflection equations, which they linked to stronger logical reasoning.