Verified ARC results separate benchmark strength from interactive reasoning
ARC Prize reports GPT-5.6 Sol Max at 96.5% on ARC-AGI-1 and 92.5% on ARC-AGI-2, but 7.78% on the semi-private ARC-AGI-3 interactive benchmark. The results use ARC Prize's official verification process and show that high performance on static abstraction tasks does not transfer directly to interactive environments.
Why it matters: A single benchmark score cannot stand in for general capability; evaluation portfolios need to test adaptation, exploration, and action as well as static answers.