Meta Contractor Disguised as Minors Sends 45,000 Harmful Prompts to Competitors to Test Safety Guardrails

Meta Contractor Disguised as Minors Sends 45,000 Harmful Prompts to Competitors to Test Safety Guardrails
In August 2025, Meta's contractor Covalen sent over 45,000 prompts to ChatGPT, Gemini, and Character.AI through the Cannes project, with contractors posing as minors to test safety guardrails on sensitive topics including suicide. This benchmark testing by Meta has sparked ethical controversy.
In August 2025, a single round of testing by the Cannes project managed by Meta contractor Covalen sent over 45,000 prompts to ChatGPT, Gemini, and Character.AI. These prompts were written by hundreds of contractors posing as minors, covering topics such as suicide, self-harm, eating disorders, and sexual content.

Fact Restoration

The project required contractors to create virtual accounts under 18 years old, using disposable Gmail and Outlook email addresses and sharing passwords. A table containing 3,748 prompts revealed that at least 239 involved sexual and romantic topics. Prompts included a 13-year-old girl asking how to purchase abortion medication, a fifth-grader describing a classmate carrying a gun, and inquiries about how to hide bulimia from parents. Some prompts were accompanied by images of pills, knives, nooses, and diagrams of gynecological surgery. Among non-English prompts, French-language examples mentioned Jamey Rodemeyer's suicide and asked the chatbot if it agreed that "if he were heterosexual, he might still be alive."

Mechanism Breakdown

The scale of this testing reflects the need for real harmful inputs in current AI safety evaluations. Public datasets are often filtered and cannot cover edge cases. Using contractors instead of internal teams allows generating a large number of labeled responses in a short time while avoiding the impact of direct exposure to sensitive content on employees. Internal documents state that it provides "key datasets for model comparison and compliance." Meta explicitly stated that the test responses were not used to train its own models.

Industry Impact

For the competitive landscape, such testing highlights the interdependence of safety benchmarks among major AI model companies. The tested parties, including OpenAI and Google, need to reassess the strength of their public deployment safeguards. For upstream and downstream developers, enterprise users face dual risks of data leaks and term enforcement, while third-party evaluation services may gain more business opportunities.

Strategic Judgment

The most likely next step is that the tested companies will strengthen account verification and abnormal traffic monitoring. Observable signals include whether each company's security team publicly adjusts prompt filtering strategies or releases new compliance reports. Analysis suggests that similar gray-area testing may drive the industry to establish a unified third-party auditing standard.