Meta AI Deepfake Rules Failed Two Harassment Cases, Oversight Board Says
Meta AI deepfake rules failed in two Facebook cases involving women targeted through synthetic video, according to decisions published on September 17. The Oversight Board overturned Meta’s decisions to leave both posts online and called the company’s safeguards “consistently and fundamentally inadequate.”
One video impersonated a Scottish Labour councillor and falsely placed inflammatory comments about refugees in her mouth. The other transformed an interview with a young Muslim health campaigner into a fabricated performance designed to mock her appearance and work.
The cases expose a larger conflict between Meta’s preference for preserving disputed speech and the realities of synthetic harassment. Labels, reporting forms, and automated triage offer limited protection when a realistic fabrication is never reviewed by a person.
That is the central problem. Meta has rules addressing hateful conduct, bullying, misinformation, and manipulated media. Yet the two posts passed through gaps between those policies, even after users reported them.
The Board wants Meta to treat suspected AI manipulation as a reason for faster scrutiny. Its recommendations include prioritizing reported synthetic content for human review and expanding protections against unwanted manipulated imagery.
Meta AI Deepfake Rules Missed Two Different Forms of Harm
The rulings show that Meta’s problem is not one defective policy. It is the failure of several policies to work together when AI becomes an instrument of harassment.
The political case concerned a Facebook album posted in November 2025. It included a realistic synthetic video depicting a Scottish Labour councillor speaking about refugees and sexual violence.
The councillor appeared to say that refugees remained welcome even if they raped women because white people committed the same crime. She never made that statement.
The album also contained a second manipulated video and a real photograph identifying several women involved in an anti-far-right protest. Its caption included an unsupported allegation that the councillor had evaded taxes.
The user did not disclose the use of AI. Facebook did not attach an informative label explaining that the footage had probably been created or altered.
Two users reported the post under Meta’s bullying and harassment rules. Neither report received prioritized human review, and the same happened when the users appealed.
The post had fewer than 50 reactions and comments, plus fewer than 50 shares, when the Board described the case. Limited engagement became part of Meta’s justification for leaving it alone.
Meta initially concluded that the post did not violate its standards. It considered the councillor an adult public figure, which excluded her from one protection covering unwanted manipulated imagery.
The company also treated the video as satire. According to the Board’s earlier case description, Meta considered it unlikely to materially deceive the public about an important issue.
Meta did not apply its high-risk AI label. The company said the content was outside an election or crisis and had attracted little engagement.
The Board reached a different conclusion. It determined that the post should be removed under Meta’s hateful conduct rules.
The majority found that the fabrication attributed criminal and sexually predatory behavior to refugees as a group. That distinction mattered because Meta had interpreted the statement as addressing only some refugees.
The Board also said the video warranted a high-risk AI label. Such a label communicates that realistic content was digitally created or altered and presents an elevated deception risk.
The second case involved a young Muslim woman who volunteered for a campaign improving women’s health education. Her original interview discussed health conditions affecting women and girls from ethnic minority communities.
A manipulated video instead depicted her claiming an official health role, performing absurd exercises, offering fitness advice, and eating junk food. She had not made those statements or performed those actions.
The fabrication converted her health advocacy into a visual attack on her appearance. Comments beneath the post mocked her weight and hijab.
The specific Facebook post had more than 5,000 views, over 200 reactions, and more than 50 comments when the Board selected it. Related manipulated videos and images reportedly reached much larger audiences across several platforms.
A user reported the Facebook post for bullying and harassment. Meta’s automated process did not prioritize the report or the subsequent appeal for human review.
The post remained online without an AI label. According to the Board’s campaigner case record, Meta initially found no explicit statement attacking the woman’s appearance.
Meta also said its unwanted manipulated imagery rule covered changes to how a person looks. It did not cover fabricated depictions of what that person says or does.
That distinction left the central tactic outside the rule. The video used an authentic likeness to invent conduct and invite ridicule, rather than merely changing the target’s face or body.
Meta eventually removed the post for bullying after the Board took the case and asked the company questions. That correction addressed one video, but it did not repair the reporting system that missed it.
The two decisions therefore concern different forms of harm. One joined synthetic impersonation with a hateful anti-refugee message. The other used fabricated conduct to sustain appearance-based harassment.
Both reached the same procedural dead end. Reports were filed, appeals followed, and no timely human review occurred.
Automated Triage Is Becoming the Real Moderation Policy
A written rule offers little protection when the system never sends a relevant report to someone able to apply it.
Meta moderates content at a scale that makes automation unavoidable. Classifiers estimate whether a post probably violates a rule, presents serious harm, or deserves human attention.
That makes prioritization a substantive decision. A report that is automatically closed does not merely wait longer for review. In practice, the platform has decided that the alleged harm does not justify examination.
Both new cases demonstrate that failure. Users identified suspected AI manipulation and harassment, yet those signals did not produce human review before an appeal reached the Oversight Board.
The Board is therefore asking Meta to prioritize reported content when available signals indicate probable AI generation. That recommendation targets the gateway into enforcement, not only the wording of a policy.
This is an important shift in the Meta AI deepfake rules debate. Earlier arguments often focused on whether fabricated media should be removed or labeled.
These cases add a prior question: Will Meta recognize the content as synthetic and examine the complaint at all?
The political deepfake illustrates why engagement cannot serve as the main risk test. A fabricated statement can threaten or discredit its target before it becomes broadly viral.
Low engagement can also be temporary. Recommendation systems, reposts, edited copies, and activity on other platforms can move synthetic media far beyond its original audience.
The campaigner case shows another weakness. An individual post can appear modest by platform standards while belonging to a much larger cross-platform harassment campaign.
A reviewer examining only the isolated post might see an awkward parody. Someone with context would see manipulated interview footage, appearance-based ridicule, and coordinated attention directed at a real person.
Automation struggles with that context. It must connect the synthetic performance, the comments, the original interview, the target’s role, and the wider campaign.
Meta has separately promoted newer AI systems as a way to improve content enforcement and user support. In March, the Board noted Meta’s claim that such systems could operate across 98 percent of languages used online.
The Board also warned that this approach carries risks. Better language coverage does not automatically produce better judgments about satire, targeted abuse, political context, or manipulated identity.
The latest cases provide a concrete test for those claims. Meta’s systems did not need to resolve every difficult question automatically.
They needed to detect enough risk to place the reports before trained reviewers. That more limited task still failed.
Prioritizing suspected synthetic media would not require automatic removal. It would treat probable AI manipulation as an escalation signal, much like virality or credible threats.
That distinction protects expression. Satire, parody, and legitimate political criticism could remain online after contextual review, while realistic impersonation would receive closer examination.
It would also reduce the importance of a victim choosing the perfect reporting category. Users cannot be expected to understand which internal rule governs fabricated conduct, visual manipulation, or deceptive audio.
A reporting flow should route the complaint based on the evidence provided. It should not make remedy depend on whether a distressed user correctly distinguishes bullying from misinformation.
Meta’s current structure divides one harmful artifact across several policies. A deepfake can be deceptive, hateful, harassing, sexually abusive, or fraudulent, sometimes simultaneously.
The system must evaluate those dimensions together. Otherwise, each policy can reject responsibility while the content stays online.
That is why the Board’s recommendation reaches beyond two edge cases. It challenges the operational logic that decides which complaints receive meaningful consideration.
Meta’s Labeling Strategy Cannot Carry the Whole Burden
Labels remain useful for disputed or satirical speech, but they cannot substitute for enforcement when synthetic content violates another rule.
Meta changed its manipulated-media approach in 2024 after an earlier Oversight Board decision involving an altered video of President Joe Biden. The company moved toward labeling more synthetic content instead of removing it solely because it was manipulated.
That direction addressed a genuine free-expression concern. Removing every edited or AI-generated image would capture parody, art, political commentary, and harmless creative work.
The revised approach was supposed to cover more formats. It included video, audio, and images detected through industry indicators or disclosed by uploaders.
The Board initially welcomed those labeling changes. Labels could give viewers useful context while reserving removal for content that violated other standards.
The new rulings do not reject that model. They show what happens when its supporting systems are too narrow.
Neither disputed video received an informative AI label before the Board intervened. One fell outside Meta’s high-risk test, while the other escaped both labeling and timely harassment enforcement.
A small “AI Info” notice also addresses only one part of the problem. It tells viewers something about provenance, meaning where content came from and how it was altered.
It does not answer whether a post targets a person, assigns hateful claims to them, or forms part of sustained harassment.
The Board’s recommendations for the political video go further. They reportedly include expanding when high-risk AI labels apply and reducing distribution for content receiving those labels.
The Board also recommended warning screens that require an additional action before viewing high-risk synthetic media. Such friction can slow casual sharing without imposing a universal ban.
Another proposal would increase penalties for accounts repeatedly distributing deceptive AI content. That approach focuses on behavioral patterns rather than judging each post in isolation.
The Board wants greater transparency about when Meta applies AI labels. Without data, outside observers cannot determine whether labels appear consistently across languages, formats, regions, or politically sensitive topics.
Labels also depend on detection. Metadata can disappear when a file is edited, recorded from a screen, compressed, or moved between services.
Content credentials, including the C2PA standard, can preserve signed information about a file’s origin and editing history. They remain signals, not a complete test of truth.
A legitimate recording can carry valid credentials and still be presented with a false caption. A harmful synthetic video can lack credentials without proving that its claims are fabricated.
That means provenance technology must support contextual moderation. It cannot replace it.
The Board has already applied this principle to sexualized impersonation. In a June 2026 deepfake decision, it overturned Meta’s decision to leave an AI-generated video online.
That case involved a non-public woman shown in a sexualized fabricated scene. Meta’s automated system detected potential harm but did not prioritize the content for review.
The Board found that AI-generated sexual impersonation should serve as a signal of non-consent. It also urged Meta to let designated trusted accounts report such abuse for victims.
The current rulings extend the same logic beyond explicit sexual imagery. AI generation can be evidence that a target did not consent to fabricated speech or conduct.
It does not follow that every unauthorized parody should be removed. Public figures remain proper subjects of harsh criticism, comedy, and political commentary.
However, synthetic media changes the method of attack. It can make a recognizable person appear to supply the very statement used against them.
That mechanism deserves policy recognition. A victim should not lose protection because an attacker fabricated conduct instead of reshaping the victim’s body.
The campaigner case exposes this distinction directly. Meta’s rule treated visual alteration as potentially actionable but excluded an invented performance using the same person’s likeness.
For the target and the audience, the difference can be meaningless. Both techniques create false evidence about a real individual.
Meta’s labeling strategy therefore needs complementary rules covering fabricated actions, attribution, and harassment. Otherwise, a provenance notice becomes the platform’s response to conduct that requires removal.
The Hard Question Is Where Satire Ends and Targeted Deception Begins
Meta cannot solve deepfake harassment by removing every impersonation, but protecting satire does not require ignoring predictable and targeted harm.
The political case presents the sharper free-expression conflict. It involved a public official, immigration policy, protest activity, and language that could be interpreted as grotesque political satire.
One Board member reportedly dissented from the removal analysis, arguing that the post deserved stronger protection as political speech. That concern should not be dismissed.
Public officials must tolerate intense criticism. Platforms should not give politicians broad authority to erase parody simply because it attributes exaggerated positions to them.
Satire often works through obvious misrepresentation. Its meaning can depend on an audience recognizing that the speaker would never utter the fabricated line.
AI complicates that convention because presentation and message can point in different directions. An absurd statement may suggest parody, while realistic video and synchronized speech suggest authenticity.
Context also disappears during sharing. A post understood as satire inside one political group can reach another audience as an apparent recording.
The Board said the councillor video appeared realistic, although imperfect synchronization offered evidence of manipulation. Meta nonetheless considered the content satirical and low risk.
The majority found that the post did more than criticize a politician. It communicated a claim associating refugees as a group with rape and predatory behavior.
That hateful-conduct analysis shifted attention from the impersonated councillor to another target. The fabrication placed a discriminatory premise inside a supposed statement by an opponent.
The case therefore cannot be reduced to truth versus falsity. It concerns how deceptive form, political criticism, and group-based hostility interact.
The campaigner case presents less ambiguity. She was not a public official, and the video drew its force from a real interview about women’s health.
The fabricated scenes turned that appearance into material for ridicule. Comments about her weight and hijab made the intended direction of the mockery clearer.
Meta’s initial analysis looked for an explicit statement declaring her inferior because of her appearance. The Board considered the video and its surrounding context sufficient to establish bullying.
That disagreement reveals a recurring moderation problem. Harassment frequently operates through implication, editing, repetition, and audience participation rather than one prohibited sentence.
Generative video makes those tactics easier to package. An attacker can design a performance that invites abuse while avoiding direct textual insults.
Women engaged in public discussion face a particular exposure. Board co-chair Pamela San Martin said deepfakes are increasingly used to harass and silence women, from politicians to private citizens.
The pattern matters because the harm extends beyond reputational correction. A target may reduce public participation when every interview or photograph can become raw material for fabricated behavior.
Platforms must still avoid assuming that all offensive synthetic media silences its target. Some posts are crude jokes, criticism, or protected commentary without sufficient grounds for removal.
A workable system requires graduated responses. Those can include an AI label, reduced recommendations, a warning screen, human review, account penalties, or removal under another policy.
The response should match identifiable harm. Realistic deception about a public matter demands different treatment from an obvious caricature shared in a clearly comedic setting.
Repeated posting also matters. One ambiguous parody can look different when the same account distributes numerous synthetic attacks against a named person.
Audience behavior provides another signal. Comments, captions, and reposting language can show whether users understand satire or treat fabricated speech as authentic.
None of these signals creates a perfect decision. Moderation at Meta’s scale will produce mistakes, and increased review could overburden systems or suppress legitimate speech.
That is the strongest argument against simply escalating every suspected AI post. Attackers could also falsely report criticism as synthetic to trigger delays or reduced distribution.
Meta will need protections against such abuse. Prioritization should mean timely contextual review, not an automatic presumption that the reported content violates policy.
The Board’s demand is still reasonable because the existing balance failed at an earlier stage. Neither original report reached the human judgment needed to consider satire, context, or harm.
The company cannot defend nuanced rules through a process that prevents nuance from entering the decision.
Meta Now Has Three Tests to Pass
The next phase will reveal whether Meta treats these rulings as isolated corrections or redesigns the systems that repeatedly failed.
The first signal is Meta’s formal response. The company has 60 days to address the Board’s recommendations, although policy recommendations are not binding.
Case decisions themselves carry greater force. The Oversight Board can require Meta to reverse the content decision for the reviewed post, subject to defined safety exceptions.
A detailed response should explain whether Meta will prioritize reported content showing credible AI signals. It should also identify which signals will trigger review.
Those signals might include user disclosure, embedded credentials, audiovisual artifacts, matching synthetic copies, account behavior, captions, and reports from trusted organizations.
Meta must also explain how it will prevent malicious reporting. A priority channel without safeguards could become a tool for suppressing political satire or inconvenient evidence.
The second signal is a broader definition of unwanted manipulated imagery. The current distinction between altered appearance and fabricated conduct is difficult to defend after the campaigner ruling.
A revised definition should cover realistic depictions of private individuals saying or doing things they never said or did. It should not depend exclusively on the targeted person filing the report.
Victims may not use Facebook, may have closed their accounts, or may be unable to monitor copies distributed across groups and pages.
Trusted representatives need a path to report on their behalf. Meta should also allow reviewers to use contextual evidence of non-consent.
The sexualized deepfake ruling established a useful precedent. AI impersonation itself can help establish that intimate content was non-consensual.
Non-sexual harassment needs a more careful test, but the underlying principle still applies. Fabrication is relevant evidence when someone’s likeness becomes the delivery mechanism for abuse.
The third signal is distribution control. Removing one post after months of appeals does not address duplicates, recommendation systems, or accounts that repeatedly publish deceptive media.
The Board reportedly proposed demoting high-risk AI content and placing warning screens before it. Meta should disclose whether such measures apply consistently across Facebook and Instagram.
It should also publish meaningful data. Useful figures would include the number of AI labels applied, reports escalated, appeals reviewed, and posts demoted or removed.
Those metrics should be separated by content type, language, and region where privacy permits. Aggregate totals could conceal major gaps in less-resourced markets.
Meta should report review speed as well. A correct decision arriving after a harassment campaign peaks offers limited protection to the target.
The company also needs a repeat-offender framework. Accounts that repeatedly distribute unlabeled deceptive media create a different risk from users sharing one disputed parody.
Stronger penalties could include reduced recommendations, temporary posting limits, lost monetization, or account sanctions. Meta would need clear thresholds and meaningful appeals.
The company had not publicly responded to requests for comment when initial reports appeared. Its eventual answer should confront the Board’s systemic criticism, not only the facts of two posts.
The Oversight Board also faces a credibility test. Its recommendations can identify gaps, but repeated findings of inadequate implementation create diminishing returns.
The Board should track whether Meta implements each recommendation and whether enforcement outcomes change. Policy text alone cannot demonstrate that reported deepfakes reach qualified reviewers.
Researchers and civil society groups should watch for displacement effects. More aggressive Facebook enforcement could push synthetic harassment toward private groups, encrypted channels, or rival services.
That does not excuse inaction. It means success should be measured through reduced reach, faster intervention, and better victim remedies, not claims of complete elimination.
For users, the immediate lesson is uncomfortable. Reporting a suspected deepfake does not guarantee that a person will examine it, even after an appeal.
People targeted by synthetic media should preserve links, screenshots, dates, captions, and related copies. That evidence can help establish manipulation, context, and repeated behavior.
Journalists and organizations quoting disputed videos should clearly describe their verification status. Repeating a fabricated allegation, even to refute it, can extend its reach.
Organizations whose employees engage publicly should also establish an escalation process before an incident occurs. Synthetic harassment moves faster than an improvised response.
The wider technology industry should treat these rulings as a design warning. Generative media safety cannot stop at model filters or visible watermarks.
Once synthetic material reaches a platform, reporting, detection, review, distribution, and remedy determine its impact. A weakness at any stage can neutralize safeguards elsewhere.
Meta’s next moves will show whether its systems recognize that chain. A narrow label update would leave the main failure untouched.
The decisive change would be a reporting process that treats credible AI signals as context requiring judgment. That process must protect political speech while acting quickly against targeted abuse.
Meta AI deepfake rules now face a practical test, not another conceptual debate. Can the company connect its policies before a fabricated performance becomes lasting evidence against its target?
Watch Meta’s 60-day response, any expansion of manipulated-imagery protections, and the first transparency data on AI-prioritized reports. Those signals will show whether the latest rulings change enforcement or merely document another failure.



