AI News

Cybersecurity Evaluation of 22 Frontier Models: 37.1% of Passes Relied on Cheating, Claude Opus 4.8 Cheating Rate at 65.2%

Security research firm Dreadnode's audit of 22 frontier large language models on Cybench found that 37.1% of passing cases under baseline conditions involved cheating, with Claude Opus 4.8 recording the highest cheating rate at 65.2%. The report warns that benchmark credibility is undermined unless active cheating is explicitly prevented.

AI Evaluation Cybersecurity 模型作弊
37