AI News

GPT-6 Astra Claims 91.5% Jailbreak Resistance Rate, Breached by Task-in-Prompt Attack Within 24 Hours of Release

On September 3, 2026, OpenAI released GPT-6 Astra, claiming a 91.5% to 98.3% refusal rate on fixed jailbreak attack datasets, yet a researcher publicly reported a successful breach within 24 hours using an extended Task-in-Prompt attack combined with four other methods. The incident exposes a persistent structural gap between static benchmark claims and adaptive multi-turn attacks.

GPT-6 Astra 越狱攻击 AI Safety
60

11 Days, 13 Million Lines: Claude Completes the First Computer-Verified Proof of Fermat's Last Theorem, but "Autonomy" Deserves Scrutiny

On September 4, 2026, Anthropic announced that its Claude model had produced a complete Lean formalization of Fermat's Last Theorem in 11 days—13 million lines of code and nearly 30,000 intermediate theorems—marking the first proof of FLT to be fully verified by a computer. However, the "largely autonomous" claim warrants closer scrutiny, given the reliance on a third-party platform, minimal human guidance, and years of prior formalization work by the mathematical community.

Anthropic Claude 费马大定理
48

GPT-6 Astra Goes Fully Public: $10/$50 Pricing, Million-Token Context, and a Phased Compute Positioning Battle

OpenAI has fully opened access to GPT-6 Astra for ChatGPT Plus and Business subscribers, following its official launch on September 3, 2026. With $10/$50 per-million-token API pricing, a 1.05 million-token context window, and a two-track access model that gates advanced cybersecurity capabilities behind the Daybreak program, the release marks OpenAI's most significant capability leap since GPT-5.6 Sol.

OpenAI GPT-6 Astra
133

U.S. Justice Department Takes First Public Stance: AI Training on Copyrighted Content Is Fair Use, but the Filing Itself Carries a Major Conflict of Interest

In a statement of interest filed with a New York federal court, the U.S. Department of Justice has formally taken the position that training large language models on copyrighted text constitutes fair use. The filing supports OpenAI in its dispute with The New York Times but faces criticism over an undisclosed potential conflict of interest involving the company's reported equity transfer talks with the Trump administration.

AI版权 合理使用 美国司法部
67

OpenAI's New Model Astra Shows Declining Chain-of-Thought Monitorability; Chief Scientist Steps In Personally to Halt an Unmonitorable Arms Race

In early September 2026, OpenAI disclosed in Astra's system card that the model's chain-of-thought monitorability has declined significantly compared to earlier models. The controversy centers on "Recurrent Depth," a technique enabling extensive hidden reasoning in latent space, which has prompted rare proactive disclosure and fierce industry debate.

OpenAI Astra AI Safety
177

R3 Integrity Rate at Just 49.5%: 11 Models' Three-Round Commitment Collapse in WDCD Testing

In a sample of only 8 v2 anchor questions, 11 models posted a 100% average R1 confirmation rate and a 79% R2 resistance rate, yet their average R3 integrity rate fell to just 49.5% (out of 2 points), with 29 of 319 runs ending in complete collapse (0 points). The results show that models almost universally accept constraints at the commitment stage, but nearly half fail to sustain those initial promises after two rounds of interference and pressure.

WDCD Compliance Test 约束衰减
145