The newest place this page's questions get sharp: services that help institutions tell real from fake. 43% of UK adults met a deepfake inside six months (Ofcom), and the response — detection, attribution, law — is being built right now. Every layer of it is a user-research problem, and the oversight findings above (07–09) apply to it directly.
Finding 17
Nobody can eyeball it — and people think they can.
The experimental record is blunt: people cannot reliably detect deepfakes and overestimate their own ability to do so — a finding that has now replicated across a 56-study meta-analysis. In the UK, only 9% of adults feel confident they could identify one; the confident aren't the accurate. Any service assuming human spotting as a control has already failed.
Source: "Fooled twice", iScience, 2021 · 56-paper meta-analysis, 2024 · Ofcom
Finding 18
Detectors don't travel — and fail unlike us.
In Meta's Deepfake Detection Challenge the winning model scored 82% on familiar data and ~65% on unseen fakes; on 2024's in-the-wild media, leading open-source detectors lost roughly half their AUC. And the errors are complementary — a 2026 evaluation of 200 people against 95 detectors found humans miss polished fakes while models flag rough-but-real footage. That asymmetry is the case for human–machine teaming, designed and tested as such.
Source: Meta Deepfake Detection Challenge, 2020 · Deepfake-Eval-2024 benchmark · 2026 human-vs-detector evaluation
Finding 19
The fraud is operational, not hypothetical.
Engineering firm Arup lost US$25 million to a single video call in which the "CFO" and colleagues were all deepfakes — one employee, one meeting, real money gone. The threat model professional services now face includes the meeting itself being synthetic. Verification workflows, not vigilance, are the defence — and workflows are designable.
Source: Arup case, Hong Kong, 2024 — confirmed publicly
Finding 20
The UK is building the benchmarking muscle.
The Home Office's Deepfake Detection Challenge returned in early 2026 as a four-day live hackathon — 450+ people, 16 teams, INTERPOL and Five Eyes participants, hidden multimodal datasets — alongside a "world-first" evaluation framework for detection tools built with Microsoft. Top image-category entries reached an F1 of ~92% on the hidden set — with the honest caveat that challenge data, however hidden, is still curated. Benchmarking over vibes: exactly the shift the rest of this page argues for.
Source: Home Office / Accelerated Capability Environment · reported February 2026
Finding 21
Attribution is now a design problem.
The counter-play to detection's arms race is provenance, now an ISO standard (C2PA, ISO/IEC 22144) — and since May 2026, OpenAI and Google pair signed Content Credentials with the SynthID watermark, 100 billion files marked, so each layer covers the other's blind spot. But both layers bend: a screenshot strips C2PA clean, and researchers have scrubbed SynthID from ~90% of test images. Ofcom's stack — prevention → embedding → detection → enforcement — treats every layer as a testable intervention, which is what it is.
Source: C2PA / ISO 22144 · OpenAI & Google, May 2026 · ETH Zurich, 2026 · Ofcom, Deepfake Defences 2, 2025
Finding 22
The law caught up — on the sharpest harms.
Sharing intimate deepfakes became an offence under the Online Safety Act 2023; creating them became one on 6 February 2026 (Data (Use and Access) Act 2025 s.138, amending the Sexual Offences Act). The Crime and Policing Act 2026 then added a 48-hour takedown duty and banned supplying "nudification" tools. But there is no general offence of creating a deepfake — fraud runs through the Fraud Act, and electoral harms hang on a 1983 provision that only bites during regulated election periods. Specific, recent, and still moving.
Source: OSA 2023 · DUAA 2025 s.138 / SI 2026/31 · Crime and Policing Act 2026 · RPA 1983 s.106
Finding 23
The users aren't hypothetical — and they disagree.
Forensics units need court-defensible reasoning, not a score — models that classify well but can't explain themselves are brittle in real investigations. Platform trust-and-safety analysts get minutes per decision and drown at the edge cases. And for victims the service is often simply absent: anonymous accounts and VPNs end the police conversation. Same media, entirely different services — the service-design problem in one sentence.
Source: digital-forensics practice ethnographies · service-design research, 2024–26
Finding 24
The detection seat inherits the automation-bias problem.
A "97% fake" verdict is a psychological anchor: reviewers drift toward rubber-stamping, and under time pressure human anomaly-spotting drops 15–20%. The mitigation is the one findings 07–08 predicted: explainable outputs — a 2025 UK forensics framework pairs high-accuracy detection with Shapley-value heatmaps so an investigator can check why — plus explicit training on the bias itself.
Source: automation-bias studies, 2025–26 · UK digital-forensics framework, 2025
Finding 25
Labels backfire in measurable ways.
85% of UK adults want AI labels; only 34% have ever seen one. And labels carry side-effects: marking some synthetic content makes people trust unlabelled content more — the implied-truth effect — while a preregistered UK/US study (~5,000 participants) found an "AI-generated" tag cut perceived accuracy by just 2.7 points against 9.3 for "False". At platform scale, fatigue sets in and the warnings become wallpaper. Attribution UX is a research problem, not a checkbox.
Source: Ofcom, Deepfake Defences · preregistered label study, 2024 · implied-truth literature
Finding 26
The case against detection-first — taken seriously.
The steelman deserves its own card. Detectors permanently lag generators. At platform scale, even 99% accuracy means a million wrongly flagged genuine videos a day. The liar's dividend lets real evidence be waved away as fake. And cryptographic provenance risks becoming an elite signal that downgrades genuine footage from anyone outside corporate hardware chains — the citizen journalist most of all. The sceptics' conclusion is that the binding constraints are legal, procedural and human, not algorithmic — which is also, precisely, the argument for doing the human-factors research properly rather than buying another classifier.
Source: Deepfake-Eval-2024 · base-rate analyses · liar's-dividend literature · sceptical coverage of the UK framework, February 2026