GPT-5.5 Leads Smoke Benchmark with Perfect Execution Score of 86.95, Exposing Constraint Weakness
In the Smoke lightweight benchmark on July 3, 2026, GPT-5.5 ranked first with a main score of 86.95, driven by a perfect code execution score of 100, while its material constraint score of 71 highlights a common weakness.