UR × AI · September 2026

UR × AIstate of play

Where AI and user research actually stand: what the machine does well in the workflow, how you research AI systems themselves, what stays human — and the synthetic-media front where it all gets sharp. Twenty-six findings, each carrying its source.

Read this first

Compiled, not opined.

Built from a structured deep-research sweep of the public record — government evaluations, regulators, academic benchmarks, practitioner debate — then hand-checked and edited.

The claim

AI adoption is an evidence question. So this page treats the AI debate the way I treat a service: find the evidence, grade it, keep the disconfirming case.

01 · In the workflow

What AI does to research practice

The honest picture is neither the vendor pitch nor the ban. AI is a strong semantic sorter and a weak sense-maker — and the failure modes land exactly where qualitative rigour lives.

Finding 01

Deduction works. Induction doesn't — yet.

Applying a fixed, pre-agreed codebook, LLMs agree with human coders at F1 0.84 / Cohen's κ 0.78 — a reliable automated coding assistant. Asked to originate themes from raw data instead, performance degrades significantly: the QualAlign benchmark tested four inductive-coding systems across eight datasets and found none aligned well with expert analysis.

Source: peer-reviewed coding evaluations · QualAlign benchmark, 2025

Finding 02

85,000 responses — and 125 humans checking.

The UK government's own "Consult" tool (built by i.AI) processed 85,000+ responses to DWP's Pathways to Work consultation. Human analysts left 73% of AI-suggested themes unchanged. Scaled, deductive sorting is real — and it still shipped with 125 human reviewers attached. DSIT's ThemeFinder found the same shape: faster themes, not final reporting.

Source: Incubator for AI (i.AI) / DWP evaluation · DSIT, 2025

Finding 03

The disconfirming case is the first casualty.

Probabilistic models converge on the majority pattern, so the outlier participant — the one rigorous analysis is built around — gets smoothed away, and minority accounts are flattened into normative language. Below roughly 1,000 datapoints, clustering tools also manufacture redundant, over-broad themes. Most qualitative studies live far below 1,000 datapoints.

Source: 2024–26 academic evaluations · SARA framework study, 2025

Finding 04

Confident fabrication has measured rates.

Forced to estimate beyond the evidence, some widely used models fabricated confidently in over half of test cases — rates of 0.53 and 0.76 against 0.15 for the best tested. The dangerous failure isn't refusal; it's fluent invention. Diary studies fare worst: models collapse a four-week sequence into one flat summary, erasing the timeline the study existed to capture.

Source: model-evaluation benchmarks, 2025–26

Finding 05

Faster, then weaker: the dependency paradox.

With AI assistance, participants' analytical accuracy rose 21%. When the assistance was removed, they fell 15 points below their own starting baseline. The skill doesn't idle while the tool works — it atrophies. For research teams, that's the case for keeping deep reading in the loop, deliberately.

Source: MIT Media Lab

Finding 06

Hybrid survives scrutiny. Autopilot doesn't.

The UK government ran the head-to-head: an AI-assisted rapid evidence review finished 23% faster — and its drafts were judged "stilted", with "peculiar hallucinations", needing more revision than the human one. Final quality matched only because a researcher verified everything against source. Meanwhile ~80% of workplace AI tools go unmanaged: the real workflow is already hybrid, just unsupervised.

Source: Behavioural Insights Team for DSIT & DCMS, 2025 · industry workforce reports, 2026

02 · Researching the machine

Methods for AI products & oversight

When the product is probabilistic, "can users complete the task" stops being the question. The question becomes: do people understand it, can they catch it failing, and is their oversight real or ceremonial?

Finding 07

Oversight degrades on a schedule.

Under time pressure, high volumes and confident-looking interfaces, human review of AI output collapses into rubber-stamping — the automation-bias literature saw this coming decades ago. The result is a "moral crumple zone": formal responsibility assigned to people with no real agency. Nominal review can even legitimise unfair outcomes rather than catch them.

Source: human-factors literature · UK parliamentary evidence, 2024–26

Finding 08

Don't ask if they'd catch the error. Inject one.

Deliberate-error injection — seeding a system with known mistakes — measures the real catch rate instead of the claimed one. A counter-intuitive result: detailed feature-attribution explanations didn't beat a plain, calibrated confidence score at helping people spot errors. Explainability features have to earn their place empirically.

Source: error-injection & explainability studies, 2025–26

Finding 09

Red-team the experience, not just the model.

UX red teaming briefs participants to confuse, contradict and push the system until it breaks, while the researcher catalogues the failures. The working standard is blunt: if a user can't recognise and correct a model error within two conversational turns, the design has failed — whatever the satisfaction score says.

Source: UX red-teaming practice, adapted from AI safety, 2025–26

Finding 10

One good run proves nothing.

Agentic systems are non-deterministic: a 60% single-run success rate can mean ~25% consistency across repeated trials of the same task. Single-response evaluations and informal prompt A/Bs are invalid here — you evaluate whole workflow trajectories, repeatedly, and report the spread.

Source: agentic-evaluation literature, 2025–26

Finding 11

Wizard of Oz is back — and it works offline.

HCI's oldest trick is now the foundational method for multi-turn AI: humans simulate the system before it exists, down to injected errors and faked latency. And it degrades gracefully into secure settings — in air-gapped environments it runs on paper, with pre-generated responses on cards. No network, no code, real findings. High-stakes environments don't excuse teams from research; they change its form.

Source: HCI methods literature · secure-environment research practice

Finding 12

"Meaningful oversight" is now a legal test.

UK data law (as amended in 2025) gives people the right to demand human review of automated decisions; MOD's JSP 936 mandates meaningful human control across an AI system's lifecycle. Researchers supply the evidence that oversight is real: do reviewers understand the output, do they have time to check it, can they overrule it and survive? That's a research protocol, not a policy assertion.

Source: Data (Use and Access) Act 2025 · MOD JSP 936 · ICO guidance

AI Playbook for UK GovernmentGDS · February 2025 Advisory

Twelve principles for responsible use. The starting point, without legal force.

Algorithmic Transparency Recording StandardCross-government Mandatory

Departments must publish records of algorithmic tools — yet by January 2025 only 33 records existed across government. The gap between rule and practice is itself a finding.

GDS Service ManualService Standard Mandatory

AI-powered services still face the same user-research and assessment expectations as any other government service.

Code of Practice on AI & ADMICO · SI 2026/425 · in force May 2026 Statutory

Binds the ICO to produce the UK's first statutory AI code. Once final, courts must take it into account — departure will need explicit justification.

03 · What stays human

Graded on evidence, not sentiment

The usual defence of the researcher leans on "empathy" and asserts the rest. The evidence draws a sharper line — and one grade here goes against the profession's favourite story.

Finding 13Strong

Framing the question.

Models execute inside boundaries they're given. Deciding what matters to a decision — what "accurate enough" means for this call, which question addresses the actual strategic bottleneck, whether the study should run at all — stays human. The value of research is the decision culture it builds, not the artefact it ships.

Source: practitioner literature, incl. Erika Hall

Finding 14Strong

Recruitment — and catching the machine's bias.

Stanford analysed 3.4 million people through commercial AI screening tools: severe, systematic bias — identical qualifications, different outcomes by demographic group — repeated across employers using the same models. And human reviewers followed the AI's recommendation 90% of the time. Reaching under-served participants doesn't happen unless a human makes it happen.

Source: Stanford HAI, 2026

Finding 15Thin — honestly

Rapport. The evidence says: it's complicated.

Participants showed three times more expressions of joy with a human interviewer — and no measurable difference in willingness to disclose, including sensitive material. On stigmatised topics, people are sometimes more candid with the machine, because it can't judge them. Warmth is real; it isn't automatically data. A grade the profession should sit with, not argue away.

Source: Curtin University experimental study, 2026

Finding 16Strong

Reading distress. Noticing absence.

An AI interviewer tested with vulnerable participants couldn't reliably spot safeguarding risks — and some participants mistook the research bot for a support service. Separately, fed deliberately blank medical images, leading vision models confidently described findings that weren't there. Recognising distress, and recognising silence as data, remain human work.

Source: Nesta (UK) · vision-language model evaluations, 2025

Junior roles are seniorising

AI-exposed junior roles are 7× more likely to demand traditionally senior skills; the AI-skills wage premium jumped from 25% to 56% in a year. Execution is automating; judgement is appreciating.

PwC AI Jobs Barometer, 2026

The UK framework moved first

The 2025 update to the government's user-researcher capability framework added research leadership and stakeholder relationship management as named skills — advocacy in, pure execution out.

Government Digital & Data profession framework, 2025

Synthetic users, graded by the field

The field's own verdict, after the hype cycle: synthetic users "got the skepticism they earned" — a narrow role in desk research, not a stand-in for real human decisions.

Maze, Future of User Research report, 2026

04 · Synthetic media

Deepfakes, detection & attribution

The newest place this page's questions get sharp: services that help institutions tell real from fake. 43% of UK adults met a deepfake inside six months (Ofcom), and the response — detection, attribution, law — is being built right now. Every layer of it is a user-research problem, and the oversight findings above (07–09) apply to it directly.

Finding 17

Nobody can eyeball it — and people think they can.

The experimental record is blunt: people cannot reliably detect deepfakes and overestimate their own ability to do so — a finding that has now replicated across a 56-study meta-analysis. In the UK, only 9% of adults feel confident they could identify one; the confident aren't the accurate. Any service assuming human spotting as a control has already failed.

Source: "Fooled twice", iScience, 2021 · 56-paper meta-analysis, 2024 · Ofcom

Finding 18

Detectors don't travel — and fail unlike us.

In Meta's Deepfake Detection Challenge the winning model scored 82% on familiar data and ~65% on unseen fakes; on 2024's in-the-wild media, leading open-source detectors lost roughly half their AUC. And the errors are complementary — a 2026 evaluation of 200 people against 95 detectors found humans miss polished fakes while models flag rough-but-real footage. That asymmetry is the case for human–machine teaming, designed and tested as such.

Source: Meta Deepfake Detection Challenge, 2020 · Deepfake-Eval-2024 benchmark · 2026 human-vs-detector evaluation

Finding 19

The fraud is operational, not hypothetical.

Engineering firm Arup lost US$25 million to a single video call in which the "CFO" and colleagues were all deepfakes — one employee, one meeting, real money gone. The threat model professional services now face includes the meeting itself being synthetic. Verification workflows, not vigilance, are the defence — and workflows are designable.

Source: Arup case, Hong Kong, 2024 — confirmed publicly

Finding 20

The UK is building the benchmarking muscle.

The Home Office's Deepfake Detection Challenge returned in early 2026 as a four-day live hackathon — 450+ people, 16 teams, INTERPOL and Five Eyes participants, hidden multimodal datasets — alongside a "world-first" evaluation framework for detection tools built with Microsoft. Top image-category entries reached an F1 of ~92% on the hidden set — with the honest caveat that challenge data, however hidden, is still curated. Benchmarking over vibes: exactly the shift the rest of this page argues for.

Source: Home Office / Accelerated Capability Environment · reported February 2026

Finding 21

Attribution is now a design problem.

The counter-play to detection's arms race is provenance, now an ISO standard (C2PA, ISO/IEC 22144) — and since May 2026, OpenAI and Google pair signed Content Credentials with the SynthID watermark, 100 billion files marked, so each layer covers the other's blind spot. But both layers bend: a screenshot strips C2PA clean, and researchers have scrubbed SynthID from ~90% of test images. Ofcom's stack — prevention → embedding → detection → enforcement — treats every layer as a testable intervention, which is what it is.

Source: C2PA / ISO 22144 · OpenAI & Google, May 2026 · ETH Zurich, 2026 · Ofcom, Deepfake Defences 2, 2025

Finding 22

The law caught up — on the sharpest harms.

Sharing intimate deepfakes became an offence under the Online Safety Act 2023; creating them became one on 6 February 2026 (Data (Use and Access) Act 2025 s.138, amending the Sexual Offences Act). The Crime and Policing Act 2026 then added a 48-hour takedown duty and banned supplying "nudification" tools. But there is no general offence of creating a deepfake — fraud runs through the Fraud Act, and electoral harms hang on a 1983 provision that only bites during regulated election periods. Specific, recent, and still moving.

Source: OSA 2023 · DUAA 2025 s.138 / SI 2026/31 · Crime and Policing Act 2026 · RPA 1983 s.106

Finding 23

The users aren't hypothetical — and they disagree.

Forensics units need court-defensible reasoning, not a score — models that classify well but can't explain themselves are brittle in real investigations. Platform trust-and-safety analysts get minutes per decision and drown at the edge cases. And for victims the service is often simply absent: anonymous accounts and VPNs end the police conversation. Same media, entirely different services — the service-design problem in one sentence.

Source: digital-forensics practice ethnographies · service-design research, 2024–26

Finding 24

The detection seat inherits the automation-bias problem.

A "97% fake" verdict is a psychological anchor: reviewers drift toward rubber-stamping, and under time pressure human anomaly-spotting drops 15–20%. The mitigation is the one findings 07–08 predicted: explainable outputs — a 2025 UK forensics framework pairs high-accuracy detection with Shapley-value heatmaps so an investigator can check why — plus explicit training on the bias itself.

Source: automation-bias studies, 2025–26 · UK digital-forensics framework, 2025

Finding 25

Labels backfire in measurable ways.

85% of UK adults want AI labels; only 34% have ever seen one. And labels carry side-effects: marking some synthetic content makes people trust unlabelled content more — the implied-truth effect — while a preregistered UK/US study (~5,000 participants) found an "AI-generated" tag cut perceived accuracy by just 2.7 points against 9.3 for "False". At platform scale, fatigue sets in and the warnings become wallpaper. Attribution UX is a research problem, not a checkbox.

Source: Ofcom, Deepfake Defences · preregistered label study, 2024 · implied-truth literature

Finding 26

The case against detection-first — taken seriously.

The steelman deserves its own card. Detectors permanently lag generators. At platform scale, even 99% accuracy means a million wrongly flagged genuine videos a day. The liar's dividend lets real evidence be waved away as fake. And cryptographic provenance risks becoming an elite signal that downgrades genuine footage from anyone outside corporate hardware chains — the citizen journalist most of all. The sceptics' conclusion is that the binding constraints are legal, procedural and human, not algorithmic — which is also, precisely, the argument for doing the human-factors research properly rather than buying another classifier.

Source: Deepfake-Eval-2024 · base-rate analyses · liar's-dividend literature · sceptical coverage of the UK framework, February 2026

Genuinely unsettled

Where honest people still disagree

A state-of-play that claims everything is resolved isn't evidence, it's marketing. These six are live.

Open question 01

Can a machine make meaning at all?

A 2025 open letter in Qualitative Inquiry419 researchers, 32 countries — rejects generative AI for all reflexive qualitative analysis. Others argue for AI as a dialogic partner, not an autonomous analyst. The boundary between pattern-recognition and interpretation is still philosophy, not settled method.

Open question 02

Who owns a hallucinated insight?

When a team builds on an AI-synthesised "user need" that turns out to be confident conflation of thin evidence, where does accountability sit? The provenance tooling to trace a generated insight back to a specific human quote is not there yet.

Open question 03

What replaces inter-coder reliability?

Much AI-coding evaluation scores the model against a single human coder treated as ground truth — a flaw the field has now noticed. There is no accepted reliability metric for a non-deterministic coder. The statistics are still being invented.

Open question 04

Do AI labels calibrate trust — or erode it everywhere?

Mandatory labelling is rolling out at platform scale, yet the implied-truth effect and label fatigue point the other way: it is genuinely unresolved whether labels calibrate public trust or quietly degrade trust in all media while users learn to ignore them.

Open question 05

Are detectors biased — and against whom?

Independent evidence on detector performance across skin tones, accents and low-light footage is exceptionally thin. Whether deployment embeds systemic false positives against particular groups — in identity checks, triage, moderation — is an open question with high stakes.

Open question 06

Whose infrastructure is the synthesis on?

The UK hosts about 4% of world AI compute and controls about 0.1% of it, researchers told Parliament in 2026. For public-sector research on foreign-controlled models, sovereignty of the synthesis stack is an open strategic question — not a procurement detail.

How this page was made

Compiled September 2026. Checked by a human. Built to change.

A structured deep-research sweep of the public record — government evaluations, regulators, academic benchmarks and practitioner debate — reviewed against the originals, then edited and arranged by me. Extended in September 2026 with a synthetic-media sweep as that front moved. Sources are named on every card; nothing is linked that can't be verified, and nothing is claimed beyond what the source supports. When the evidence moves, the page moves.

250+public sources swept
26findings, each attributed
6questions left honestly open