top of page

SARAAB Deepfake Detection Model Claims 91% Accuracy, but Open Testing Comes Next

4 days ago
13 min read

Dubai’s SARAAB deepfake detection model reportedly reached 91% accuracy while analyzing videos as short as seven seconds. That claim gives the Dubai Electronic Security Center, or DESC, an unusually concrete position in the race to authenticate synthetic media.

The headline number is only the beginning of the story. DESC says it will release SARAAB through developer platforms, allowing outside researchers to inspect, test, and improve the system. That decision shifts the contest from a government announcement to a testable open-source project.

The main opponent is not another named detector. It is the persistent gap between strong benchmark results and reliable performance on unfamiliar, compressed, or deliberately altered videos. NIST research shows why generalization remains a defining problem for media forensics.

What Dubai Has Actually Announced About SARAAB

SARAAB matters because a government cybersecurity agency plans to expose its deepfake detector to outside scrutiny.

DESC officially unveiled SARAAB on September 16, 2026, during GISEC Global at the Dubai Exhibition Centre. The agency described it as an open-source artificial intelligence model developed entirely by an Emirati team.

DESC also called it the first open-source deepfake detection model created by a government entity in the Arab region. That description establishes the institutional first claimed by the agency. It does not establish that SARAAB outperforms every detector available worldwide.

The agency’s SARAAB announcement says the model will be distributed through developer platforms. AI engineers, cybersecurity specialists, and researchers will be able to test it and contribute to its development.

DESC did not publish the reported accuracy numbers in that initial announcement. Those details appeared in subsequent coverage based on information attributed to the center.

According to the published account, SARAAB reached accuracy of up to 91% in DESC’s evaluation. The model reportedly recorded 89% accuracy on data it had not encountered during training.

That second result is arguably more important than the headline figure. Testing on unseen data is intended to measure whether a detector learned transferable forensic signals rather than memorizing characteristics of its training set.

However, “unseen” can describe several different testing conditions. It can mean new videos from familiar generators, new people, a separate dataset, or output from entirely unfamiliar generation methods. Those conditions impose very different levels of difficulty.

The available public material does not identify the test datasets, sample counts, demographic composition, video sources, or manipulation tools. It also does not disclose class balance, false-positive rates, or confidence thresholds.

Without those details, readers cannot reproduce the 91% result or compare it fairly with another model. Accuracy alone can also mislead when authentic and manipulated samples are unevenly represented.

The report says SARAAB can assess clips as short as seven seconds. Other open-source models included in DESC’s comparison reportedly required at least 15 seconds of footage.

That capability addresses an important operational constraint. Many suspicious videos circulate as short clips on messaging services and social platforms. Investigators cannot assume that a longer, higher-quality original will remain available.

A short-input detector could help teams triage such material before conducting deeper forensic work. It could also fit workflows where delayed analysis allows a deceptive clip to spread faster than a response.

Still, a seven-second minimum does not reveal how performance changes with duration. A detector might accept seven seconds while producing its most reliable results on longer clips. Public evaluation should measure that curve rather than treating minimum input length as a complete performance claim.

The reporting also says SARAAB can identify manipulated facial regions with heat maps. A heat map is a visual layer that marks areas contributing to a model’s decision.

That feature can help analysts investigate why a clip was flagged. It does not automatically make the underlying prediction correct or fully explain the model’s reasoning.

SARAAB reportedly supports analysis of partially manipulated videos. In this scenario, most of a recording remains authentic while one segment or face is altered. Detecting that localized intervention is harder than classifying a fully synthetic clip.

It is also closer to a realistic attack. A fraudster does not need to fabricate an entire meeting when changing a request, statement, or payment instruction is enough.

Those capabilities make the Dubai deepfake detector interesting. Its eventual value, however, will depend on what outside evaluators can reproduce after the code, weights, and evaluation procedure become available.

Why the SARAAB Deepfake Detection Model Targets a Harder Problem

The real technical contest is between reusable forensic signals and artifacts that disappear when generation methods change.

Deepfake detection systems commonly learn patterns associated with synthetic or manipulated media. These patterns can include facial inconsistencies, temporal irregularities, blending artifacts, frequency signals, or relationships between consecutive frames.

The difficulty is that generation systems keep changing. A clue associated with one face-swapping pipeline may vanish in a newer diffusion-based workflow. Compression and editing can further weaken the remaining evidence.

A detector can therefore perform well on a familiar benchmark without remaining dependable in the wild. This problem is known as generalization, meaning performance on data or manipulation techniques outside the training distribution.

NIST’s work on the generalization problem illustrates the size of that gap. In one comparison, a baseline detector recorded an AUC of 0.89 on a known generator.

Its AUC dropped to 0.58 and 0.55 on examples from unfamiliar generation methods. AUC measures how well a classifier separates two classes across different thresholds. A result near 0.5 approaches random ranking.

Another evaluated system generalized much better, recording AUC values between 0.97 and 0.99 across the illustrated datasets. The comparison shows that failure is not inevitable. It also shows that architecture, training data, and evaluation design matter enormously.

SARAAB’s reported 89% accuracy on unseen data speaks directly to this challenge. Yet the result remains difficult to interpret without knowing what was unseen and how the test set was constructed.

A meaningful public evaluation should separate several dimensions. It should test unknown generators, identities absent from training, varied skin tones, different lighting, and multiple camera types.

It should also include common transformations. Social platforms often resize video, reduce its bitrate, modify color, strip metadata, or recompress the file several times.

Attackers can introduce those transformations intentionally. They can crop faces, add noise, overlay text, change frame rates, or record a manipulated video from another screen.

Each operation can erase forensic artifacts while leaving the deceptive meaning intact. A detector intended for operational use must be measured after those changes, not only against clean laboratory samples.

Partial manipulation creates another challenge. If only several seconds are fake, a model that averages signals across the full video might dilute the suspicious segment.

Frame-level or segment-level analysis can help localize the alteration. It also creates more predictions, increasing the chance of isolated false alarms.

Heat maps can guide a human analyst toward facial regions that deserve examination. However, attention maps are not equivalent to causal explanations. They can look persuasive even when a model relies on irrelevant patterns.

Outside researchers should therefore test whether SARAAB’s highlighted regions correspond to actual edits. They should also check whether harmless conditions, including makeup, filters, low light, or video conferencing effects, trigger similar patterns.

Short clips raise their own statistical problem. Fewer frames provide less temporal evidence, while platform compression can remove fine visual detail. Fast movement and occlusion further complicate facial analysis.

That makes the seven-second claim useful but demanding. Independent tests should report performance across several clip lengths, including seven, 15, 30, and 60 seconds.

They should also report latency and hardware requirements. A detector that processes a short clip accurately but takes several minutes on specialized hardware serves a different role from a real-time screening system.

The central question is not whether the Dubai deepfake detector can find examples that humans miss. It is whether its decision remains stable across the messy transformations that characterize real distribution.

SARAAB’s open-source plan creates a path toward answering that question. An accessible repository can expose preprocessing, model architecture, thresholds, and known limitations.

Open weights would permit independent red-team testing. A documented dataset and evaluation script would make the 91% claim reproducible rather than merely repeatable inside DESC.

The distinction matters. Repetition by the original team shows internal consistency. Reproduction by outsiders shows that the reported result survives independent methods and assumptions.

Open Source Changes the Balance Between Claims and Verification

Releasing SARAAB openly turns outside criticism from a communications risk into part of the development process.

Government agencies often procure detection products from private vendors. Those systems may provide a confidence score without revealing their training data, model behavior, or failure conditions.

Such tools can help with triage, but their opacity creates operational risks. Investigators might place too much weight on a score they cannot examine or reproduce.

DESC is pursuing a different route. By planning a release through platforms used by developers, the agency is inviting security researchers to test the model against cases its creators did not select.

That openness supports faster discovery of blind spots. Researchers can probe sensitivity to compression, demographic variation, adversarial changes, and new generation models.

It can also expose vulnerabilities to attackers. Once weights and preprocessing logic are public, an adversary can study the detector and optimize synthetic media against it.

This tension is familiar in cybersecurity. Open code enables inspection and collective improvement, while disclosure gives attackers information about defensive assumptions.

The decisive factor is maintenance. An open detector that receives regular evaluation, issue tracking, and model updates can respond to discovered weaknesses.

An abandoned repository offers transparency without durable protection. Its published benchmark becomes less relevant as generation systems and distribution patterns change.

DESC has said researchers will be able to build on and contribute to SARAAB. That promise will become measurable when the repository appears.

A serious release should identify its license, model weights, intended use, and prohibited uses. It should include versioned evaluation results and a clear reporting channel for security findings.

Dataset documentation also matters. Deepfake datasets can embed demographic imbalances, consent problems, or narrow assumptions about what counts as authentic media.

A model card should explain which faces, languages, video formats, and manipulation families were represented. It should state where the model performs poorly.

The release should also distinguish between research use and high-stakes decisions. A detector score should not independently determine whether a video is fraudulent, whether a person is lying, or whether evidence is admissible.

NIST operates the OpenMFC evaluation to provide shared testing for media-forensics systems. Its approach reflects why common evaluation conditions are essential.

When every developer selects different datasets and metrics, headline accuracy becomes difficult to compare. A shared challenge can test systems against held-back data under consistent rules.

SARAAB would benefit from that kind of evaluation. DESC’s comparison reportedly placed it above other included open-source models, but the names and configurations of those models have not been published.

That omission prevents readers from judging the baseline. A comparison against older, unmaintained detectors would mean something different from a test against leading systems configured by their own developers.

Independent benchmarking should compare SARAAB at the same operating points. False-positive rate, false-negative rate, precision, recall, and calibration all reveal information hidden by a single accuracy figure.

False positives deserve particular attention. A detector that incorrectly labels authentic footage can damage reputations, delay reporting, and create doubt around legitimate evidence.

False negatives create a different harm. They can give a deceptive video an undeserved aura of authenticity when a screening tool fails to flag it.

The problem becomes more serious when users treat “not detected” as “verified real.” Deepfake detectors usually assess whether known forensic signals are present. They do not establish the truth of every event shown.

Open documentation should make that boundary explicit. The best result would be a tool that supports trained analysts, not an automated judge that replaces verification.

This is where openness gives SARAAB a structural advantage over an unsupported black box. Researchers can test the boundaries and communicate them before organizations build consequential workflows around the model.

The release can also encourage regional research capacity. An Emirati-led project can bring additional institutions, datasets, and operational contexts into a field often shaped by North American, European, and East Asian benchmarks.

That benefit depends on responsible data practices and transparent evaluation. Regional representation should improve the evidence base without weakening consent, privacy, or documentation standards.

What 91% Accuracy Does Not Tell Security Teams

A 91% result is a promising research signal, not a universal probability that SARAAB will identify any deepfake correctly.

Accuracy measures the share of correct classifications in a particular evaluation. Its meaning depends on the dataset, labels, class balance, threshold, and definition of a correct result.

Consider a test set dominated by authentic videos. A model could achieve high accuracy by favoring the authentic label, even while missing many manipulated clips.

Precision asks how many flagged videos were truly manipulated. Recall asks how many manipulated videos the system successfully found. Both are important in a screening workflow.

The preferred balance depends on the use case. A social platform handling millions of uploads may prioritize scalable triage. A forensic team examining evidence may tolerate slower analysis to reduce damaging errors.

Calibration is also important. A well-calibrated 80% confidence score should correspond to correct predictions at roughly that rate across comparable cases.

An uncalibrated model can produce confident scores that do not map reliably to real risk. That behavior can mislead operators even when aggregate accuracy looks acceptable.

SARAAB’s reported results do not yet disclose these measures. They also do not reveal whether the team tested manipulated audio alongside video.

The public description focuses on facial manipulation in video. That scope matters because a clip can use authentic visuals with cloned audio, synthetic subtitles, or misleading edits.

Conversely, an authentic recording can be presented with false context. No pixel-level detector can determine whether the caption, date, location, or surrounding claim is truthful.

Deepfake detection therefore belongs inside a broader verification process. Analysts should examine the source, publication history, metadata, corroborating evidence, and behavior requested by the sender.

Content provenance provides another layer. The C2PA standard supports tamper-evident records describing the source and editing history associated with a digital asset.

The official C2PA explainer emphasizes that Content Credentials do not decide whether an assertion is true. They verify that provenance information is properly formed, trusted, and associated with the asset.

Provenance and detection solve different problems. SARAAB searches for forensic signs of manipulation after examining the media.

Content Credentials can record where a file came from and what declared changes occurred. Their absence does not prove that a file is fake, since many cameras and editing tools do not attach them.

Credentials can also be removed through conversion or platform processing. A detector remains useful when provenance is missing, while provenance can strengthen confidence without relying solely on statistical artifacts.

Human verification provides a third layer. Investigators can contact the purported source through a known channel, locate earlier copies, and compare the media with verified material.

These layers work best together. No single detector, watermark, credential, or analyst can cover every manipulation and distribution path.

The stakes extend beyond misinformation as an abstract problem. Impersonation fraud targets people through messages that appear to come from relatives, executives, businesses, or public officials.

The US Federal Trade Commission reported that consumers lost nearly $3 billion to impersonators during 2024. Its fraud guidance advises people to verify unexpected contacts using information they independently know is genuine.

That practice remains necessary even when automated screening improves. A 91% detector still implies errors within the conditions of its own reported test.

Real-world performance can be lower when the input differs from those conditions. Attackers also adapt once they learn which artifacts trigger detection.

Security teams should avoid turning the SARAAB deepfake detection model into a green-light system. A low-risk score should not authorize a payment, reset an account, or validate an executive request.

A better workflow uses the score to route cases. High-risk clips receive immediate review, while consequential requests require separate identity checks regardless of the detector’s output.

Organizations also need procedures for disputed results. People affected by a false flag should have access to human review and supporting evidence.

Logging model versions is essential. If SARAAB changes over time, investigators must know which release analyzed a clip and which threshold produced the result.

Those records help reproduce decisions and measure drift. They also prevent teams from presenting a current model’s behavior as evidence for an older assessment.

Three Signals Will Show Whether Dubai’s Detector Delivers

SARAAB’s credibility will rise or fall through publication quality, independent stress tests, and evidence of operational use.

The first signal is the actual open-source release. Researchers should look for usable code, downloadable weights, an explicit license, evaluation scripts, and documentation of the training and test conditions.

A repository containing only a demonstration interface would not provide enough access to reproduce the 91% figure. A complete release would let evaluators trace the path from raw video to final classification.

The model card should explain SARAAB’s intended users and supported formats. It should also describe its limitations, data governance, and expected performance across different input conditions.

DESC should publish which models appeared in its comparative test. Version numbers, preprocessing settings, hardware, and decision thresholds should accompany the results.

If those materials appear, the project’s core claim becomes auditable. If they remain absent, the announcement will carry less weight than the accuracy number suggests.

The second signal is independent testing against unfamiliar generators and degraded media. Researchers should challenge the model with new face-swapping systems, diffusion tools, filters, screen recordings, and repeated compression.

They should include videos from varied devices, lighting conditions, and demographic groups. Results should be reported at several clip lengths, especially the claimed seven-second minimum.

Partial manipulation deserves a dedicated test. Evaluators should measure whether SARAAB can find a brief altered segment inside a longer authentic recording without creating excessive false alarms.

Heat-map localization should receive separate scoring. A correct video-level label does not necessarily mean the highlighted facial region matches the true manipulated area.

The strongest evidence would come from shared or held-back evaluations that prevent developers from tuning directly to the answers. Results from NIST-style challenges would be more comparable than self-selected demonstrations.

Strong cross-dataset performance would reinforce DESC’s claim that SARAAB offers transferable detection. Sharp declines on new generators would show that the model remains tied to its original benchmark.

The third signal is responsible operational adoption. DESC or partner organizations should explain where SARAAB is used, what decisions it informs, and how humans review its output.

Adoption alone would not prove accuracy. However, a documented workflow can show whether the model processes real submissions, reduces review time, or helps localize suspicious segments.

Useful reporting should include error rates observed after deployment. It should explain how teams handle uncertain scores and how often analysts overturn model outputs.

Organizations should also disclose whether SARAAB runs locally or sends media to another service. That distinction affects privacy when users submit personal, confidential, or evidentiary videos.

A local open-source deployment could help agencies retain control of sensitive material. It would also place responsibility for security, updates, and model governance on the deploying organization.

These three signals create a clear standard. A complete release establishes transparency. Independent testing measures generalization. Responsible adoption reveals whether the model belongs in real investigative workflows.

Until then, the fairest assessment is conditional. SARAAB is a notable government-backed open-source initiative with a reported 91% accuracy result and a useful focus on short videos.

It is not yet a universal authenticity system. The published evidence does not establish its performance across every generator, platform, demographic group, or adversarial transformation.

That distinction should not erase the project’s value. It should guide the next phase of scrutiny.

Developers and security teams should watch the repository rather than the headline alone. They should reproduce the reported benchmark, test the seven-second claim, and publish failures alongside successes.

For ordinary users, the immediate lesson remains practical. Treat suspicious video as one piece of evidence, especially when it accompanies urgency, secrecy, or a request for money.

The SARAAB deepfake detection model can strengthen that verification process if its open release matches DESC’s commitment. The next question is whether independent testers can reproduce its results under conditions the original team did not control.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page