Research

Lack of memory handling worsened GPT-5.6 Sol's performance on the ARC-AGI-3 benchmark

GPT-5.6 Sol's poor performance on the ARC-AGI-3 2D-puzzle benchmark was caused by the test system's memory handling; the investigation found that enabling two API settings tripled the score and…

Lack of memory handling worsened GPT-5.6 Sol's performance on the ARC-AGI-3 benchmark

GPT-5.6 Sol's poor performance on the ARC-AGI-3 2D-puzzle benchmark was caused by the test system's memory handling; the investigation found that enabling two API settings tripled the score and required six times fewer output tokens.