AI News

Safety Benchmarks Saturate Before Capabilities: Anthropic Risk Report Reveals Structural Cracks in Self-Certification

Anthropic's second company-wide risk report upgrades the misalignment risk rating from "very low" to "low," revealing that safety evaluation tools have saturated ahead of capability growth. A stronger internal model was shelved due to incomplete evaluations, and over 133 million human feedback interactions ran without a bioweapons classifier.

Anthropic AI Safety 模型对齐
42

Claude Protein Design Hits 14/15 Targets: First Wet Lab Validation Stirs Biosecurity Governance Debate

Anthropic's Claude model autonomously completed protein design campaigns across 15 targets, with wet lab validation confirming hits on 14 targets at a 26.8% hit rate — more than double the industry typical range. The company simultaneously restricted access to the strongest model's dual-use biological research capabilities, revealing a structural tension between commercial value and biosecurity risk.

蛋白质设计 Anthropic 生物安全
181