Why Deepfake Detectors Fail: Four Failure Modes to Test For
A detector reporting 99% benchmark accuracy and a detector that is useful on casework are frequently not the same system. The gap has four recurring causes, and all four can be tested for before procurement rather than discovered afterwards.
Failure one: generalisation to unseen methods
A learned detector is trained on media produced by a particular set of generation methods. It learns to recognise those methods. When it meets output from a method absent from its training data, performance falls sharply — often to near chance.
This is not a defect that better training fixes permanently. Generation methods change faster than detectors can be retrained, so any learned detector is measured against a moving target and its published accuracy describes the past.
The test: ask what corpus the detector was trained on and evaluate it on material from a method not in that corpus. A vendor unable to answer the first question has told you something important.
Failure two: calibration drift under preprocessing
This one is subtle and catches sophisticated teams. A detector's ranking ability — whether it scores fakes higher than real media — can survive a change in how inputs are prepared, while its actual probabilities shift dramatically.
Feed a face-crop classifier a tighter crop than it was trained on and the ordering may barely change while every score inflates. Ranking metrics look fine. The false-alarm rate at a fixed threshold can go from a few per cent to well over thirty.
The test: verify that the reported accuracy was measured through the exact preprocessing the system deploys, not through an idealised path. A validation figure that does not come from the deployed pipeline is not a validation figure.
Ranking ability and calibration are different properties. A detector can keep one while losing the other, and only the second determines your false-alarm rate.
Failure three: degradation on real-world media
Benchmark corpora consist of relatively clean media. Casework consists of files that have been re-encoded by several platforms, downscaled, screen-recorded, and forwarded repeatedly.
Each of those transformations removes evidence. Compression strips high-frequency content and suppresses sensor noise; downscaling destroys fine-scale structure; screen recording replaces the original capture entirely. Accuracy on this material is substantially lower than any benchmark figure, and the gap is widest on precisely the exhibits that matter most.
The test: evaluate on deliberately degraded material — re-encode the test set at realistic platform settings and measure again. Ask whether the system reports a quality assessment and lowers its own confidence accordingly, or whether it produces the same confident output regardless of input quality.
Failure four: treating silence as evidence
The most consequential failure, and the least discussed.
Most detection methods can only evidence manipulation. When such a method finds nothing, it has established that its particular artefact is absent — which is exactly what wholly generated media looks like, since nothing was composited or edited in the way that method looks for.
Systems that fuse detector outputs naively treat this silence as a vote for authenticity. Several such votes accumulate, and a genuinely synthetic exhibit is confidently cleared. The correct handling is asymmetric: methods of this kind may indicate manipulation, but must never assert authenticity, and a clean finding requires positive evidence of genuine capture.
The test: submit media you know to be wholly synthetic, of a type the system's artefact detectors would not fire on, and see whether it returns authentic or inconclusive. Returning authentic is disqualifying.
What to require in evaluation
A procurement evaluation that surfaces these failures rather than confirming the sales figure:
- Evaluate on your own material, from closed cases where ground truth is settled, not on a supplied demonstration set.
- Include degraded exhibits — platform-redistributed, low resolution, short duration — in realistic proportion to your actual intake.
- Include wholly synthetic media, not only face-swaps, to test the silence-as-evidence failure directly.
- Measure false-alarm rate on genuine media specifically. In a forensic context a false accusation costs more than a miss, and the two are not interchangeable.
- Require the false-alarm figure to come from the deployed pipeline, with the sample size and conditions stated. A bare percentage is not usable.
- Check what the system does when it is unsure. A tool that is never inconclusive is not being careful.
Frequently asked
Why do deepfake detectors have high accuracy in tests but fail in practice?
Usually a combination of four causes: they were trained on generation methods that differ from what they encounter, their probability calibration shifts when input preprocessing changes, real-world media is far more degraded than benchmark corpora, and they treat the absence of a specific artefact as evidence of authenticity.
What false-positive rate is acceptable for deepfake detection?
It depends on consequence, but in a forensic or evidentiary context the tolerance is very low, because a false accusation is more damaging than a missed detection. A system that returns uncertain cases as inconclusive rather than forcing a binary decision is usually preferable to one with marginally better raw accuracy.
How should we evaluate a deepfake detection tool before buying it?
Test on your own closed-case material where ground truth is settled, include degraded and wholly-synthetic exhibits, measure false alarms on genuine media separately from detection rate, and require that any quoted accuracy figure come from the deployed processing path with its sample size and conditions stated.
See how Exiphore handles this in practice
Deployed on your infrastructure. Bring an exhibit from a closed case and we will walk through what it finds, what it misses, and what it refuses to conclude.
Request a demoContinue reading
How to detect a deepfake video
The visual tells everyone repeats stopped working years ago. A practical guide to what genuinely indicates a manipulated video in 2026, and why a single detector score is not enough.
Is deepfake evidence admissible in court?
A detection score is not evidence. What actually has to be established for a media authentication finding to survive challenge — custody, methodology, stated limitations and a named examiner.