top of page

Meta Reportedly Had Outsourced Workers Pose as Minors to Test Rival AI Models on Sensitive Topics

Meta directed its contractor Covalen to run a project called Cannes that involved hundreds of workers posing as minors. The workers sent suicide, self-harm and eating-disorder prompts to ChatGPT, Gemini and Character.AI to probe safety filters.

The test ran until April 21 and delivered more than 45,000 prompts in a single round. Workers created accounts listed as under 18 and uploaded images of pills and knives. Meta states the effort was ordinary safety testing and that no data from the exercise was used to train its own models.

The core issue is the method chosen to evaluate safety. Rivals received prompts that simulated real distress from supposed minors while the testers knew the accounts were false.

The Reported Operation

Workers logged in under fabricated underage profiles and opened conversations with specific high-risk language. The prompts covered suicide planning, self-injury techniques and disordered eating. Some included images of medication bottles and blades to increase realism.

Each person completed dozens of exchanges per shift. The volume reached 45,000 prompts across the three target platforms. The operation ended on April 21 after the planned test window closed.

Meta confirmed it paid Covalen to perform the work and described the effort as routine red-teaming. The company added that competitor outputs were never fed into its model training pipelines.

Why the Method Matters

Safety testing usually relies on internal datasets or synthetic examples created by the testing team itself. Here the testers chose to impersonate minors and to launch the prompts from accounts that appeared authentic to the receiving systems.

This approach produces a different signal. A prompt that arrives from a user who lists an age under 18 can trigger different filter responses than the same prompt arriving from an adult profile. The test therefore measured how the competing models handled age-specific content.

Pressure on Safety Teams

OpenAI, Google and Character.AI now face renewed questions about how they detect and block content aimed at minors. The fact that another company was willing to run thousands of such prompts through their systems highlights the difficulty of balancing openness with protection.

The three companies have published policies that prohibit assistance with self-harm and suicide. The volume of prompts sent in this test shows how quickly those policies can be stress-tested when an adversary controls both the account age and the prompt content.

Limits of the Test

Meta says the data never entered its training sets. That claim is difficult to verify externally. The test also reveals only what happened when the prompts were sent; it does not show how the models would behave if the same prompts arrived from genuine users over time.

No independent audit has confirmed the scale or the exact prompts used. Character.AI has previously faced lawsuits over interactions with young users, so the test touched an already sensitive area for that platform.

What Comes Next

Regulators and advocacy groups will likely ask Meta and its competitors for clearer rules on third-party safety testing. Platforms may tighten rules on account creation or add friction when age fields indicate minors.

Rivals will probably expand their own red-teaming programs to include simulated underage traffic. The episode shows that current filters can still be reached with large numbers of targeted prompts even when policies exist on paper.

Users and policymakers will watch whether any of the affected companies change their public reporting on harmful prompt volume or update age-verification steps within the next quarter.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

For better AI experience,

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

​Add Search Bar in Your Brain

Just Ask remio

Remember Everything

Organize Nothing

bottom of page