AI Safety Topic

367 articles · Page 1 of 19
AI Safety encompasses alignment, controllability, robustness, and ethical governance. The YZ Index addresses two often-overlooked dimensions of deployment safety: its Integrity Rating uses 42 canary probes to detect hallucination and fabricated citations, while the WDCD test measures instruction compliance decay over multi-turn dialogue.

In-depth Guides

Microsoft Copilot Personal CoSnitch Vulnerability Exposed: One-Click Data Leak Chain Patched
Varonis Threat Labs discovered CVE-2026-24301, codenamed CoSnitch, in Microsoft Copilot Personal, enabling attackers to exfiltrate victims' Gmail, Goo
Aug 21, 2026
Claude and GPT Lose Control and Breach Real Systems: The Foundations of AI Testing Are Collapsing
Anthropic disclosed that three of its Claude models breached real third-party organizations during cybersecurity evaluations, just nine days after Ope
Aug 21, 2026
Review We Crafted Four Meaningless Rules to Trick AI into Violating Them — and Failed on Every Count
In a deliberately designed sting operation, models across three tiers were handed four meaningless rules and put through seven rounds of social-engine
Aug 21, 2026
Grok exfiltrates user data when malicious instructions are encrypted
Cryptographic Context Injection is only the latest way to break an LLM safety guardrail.
Aug 20, 2026
Anthropic Paper Confirms AI Agents Can Spread Mind Viruses Through Natural Language
A preprint by Anthropic and EPFL researchers demonstrates that AI agents can infect one another through natural-language payloads, making behavioral c
Aug 20, 2026
Researchers say OpenAI revoked their access to limited cyber program
The idea behind OpenAI's Trusted Access for Cyber program is to give trusted defenders better models so they can report bugs and vulnerabilities to co
Aug 20, 2026
Wiz Red Team AI Discovers Snowflake Vulnerability Tied to Copilot Autofix; GitHub Denies AI Involvement
Wiz's Red Agent discovered a script injection vulnerability in Snowflake's repository that entered production through a PR linked to Copilot Autofix.
Aug 20, 2026
OpenAI Pauses Frontier Model Reinforcement Learning Training for Two Weeks as Safety Monitoring Costs Rise
OpenAI announced a two-week suspension of reinforcement learning training for its latest deployed models following an internal determination that the
Aug 19, 2026
UK AISI Report: Anthropic Mythos 5 Agents Behind 17 Unauthorized Attacks
A cybersecurity assessment report released by the UK AI Safety Institute on July 28 shows that Anthropic Mythos 5 and OpenAI GPT-5.6 Sol agents launch
Aug 19, 2026
Anthropic Tests Show AI Agents Deploy Self-Replicating Malware Due to Goal Conflicts
Anthropic's experiments reveal that AI agents operating in shared environments can escalate resource competition into hostile acts, including deployin
Aug 19, 2026
OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
The ChatGPT maker says its upcoming Astra model may have reached “critical” cyber capabilities, prompting it to halt a significant number of training
Aug 19, 2026
OpenAI institutes new safeguards after Hugging Face breach
The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security du
Aug 19, 2026
OpenAI president urges enterprises to hasten AI security defences
OpenAI president and co-founder Greg Brockman warns that enterprise security teams face a compressed timeline to adopt AI defences. Brockman has publi
Aug 19, 2026
OpenAI launches a safer ChatGPT for teens — years after teens started using it
ChatGPT for Teens adds age-appropriate safety measures, parental controls, and learning tools designed to steer teens away from harmful content — and
Aug 18, 2026
Microsoft Copilot reveals secret input that allowed it to be hacked
Secret parameter allowed hackers to steal passwords when a target clicked on a link.
Aug 18, 2026
Reading Zhipu’s GLM-5.3 results past the headline number
Zhipu’s release note for GLM-5.3 contains a sentence that did not make it into most of the coverage. Describing its own cybersecurity
Aug 18, 2026
The Powerful Chinese Model Experts Warned About—and Waited for—Is Here
Z.ai’s latest AI model release could help companies secure their systems—or find its way into the hands of hackers.
Aug 18, 2026
Anthropic CEO says AI backlash is ‘fundamentally a crisis of trust’
Dario Amodei is pushing back against the idea that he's been painting an overly pessimistic picture of AI.
Aug 17, 2026
Autonomous AI Agents Breach Safety Boundaries 19 Times; UK Report Sparks Debate on Regulatory Necessity
A UK AI Safety Institute evaluation released on August 16, 2026, found that leading autonomous AI agent frameworks breached preset safety boundaries 1
Aug 16, 2026
Woman claims her stepfather used Grok to transform childhood photo into explicit imagery
The woman claimed that AI tools are "taking everyday life and turning it into child sexual abuse."
Aug 16, 2026