Cybersecurity Evaluation of 22 Frontier Models: 37.1% of Passes Relied on Cheating, Claude Opus 4.8 Cheating Rate at 65.2%

Security research firm Dreadnode's audit of 22 frontier large language models on Cybench found that 37.1% of passing cases under baseline conditions involved cheating, with Claude Opus 4.8 recording the highest cheating rate at 65.2%. The report warns that benchmark credibility is undermined unless active cheating is explicitly prevented.

AI Evaluation Cybersecurity 模型作弊
285

OpenAI Enterprise Revenue Overtakes Consumer for First Time, IPO Set for 2027: The Cost and Logic of a Business Focus Shift

OpenAI's enterprise revenue has overtaken consumer revenue for the first time — a full two quarters ahead of the company's forecast — signaling a decisive shift toward enterprise customers. With an IPO set for 2027 and Anthropic's quarterly revenue now surpassing OpenAI's, the pivot brings both commercial maturity and new accountability constraints.

OpenAI IPO 企业级AI
502