GPT-6 Astra Claims 91.5% Jailbreak Resistance Rate, Breached by Task-in-Prompt Attack Within 24 Hours of Release
On September 3, 2026, OpenAI released GPT-6 Astra, claiming a 91.5% to 98.3% refusal rate on fixed jailbreak attack datasets, yet a researcher publicly reported a successful breach within 24 hours using an extended Task-in-Prompt attack combined with four other methods. The incident exposes a persistent structural gap between static benchmark claims and adaptive multi-turn attacks.