Models · The Decoder ·

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

OpenAI says GPT-5.6 Sol scored 38.3% on ARC-AGI-3 using its API with retained reasoning and context compaction, versus Opus 5 at 30.2%. In the official test environment, GPT-5.6 Sol scored 7.8%.

Read the full story at The Decoder →