Tech4Humanity AtlasGround ZeroCurrent ThemesFuture ResearchGalleryLive Q&ASearch

Ground Zero / CARE / CSO

GZ-STORY-017

Harm & Auditability

At 8:10 on Monday morning, Samira Patel watched two people receive the same AI advice and move in opposite directions. One became more capable. The other became quieter, less certain and increasingly dependent on the next prompt.

Samira Patel, a independent assurance chair in London, had been asked to decide whether Harm & Auditability belonged in a real programme or in another folder of attractive ideas. The immediate problem was practical: Current approaches to harm & auditability are fragmented, poorly calibrated or insufficiently measured, making it difficult to distinguish real benefit from substitution, novelty or surveillance effects. People affected by the decision could feel the consequences, but the organisation lacked a defensible way to separate benefit from novelty, substitution or hidden burden.

Instead of beginning with a product demonstration, Samira Patel put the research question on the wall: “Which audit designs surface real-world harm early enough for effective remediation?” That changed the room. The question did not assume the technology worked. It asked where, for whom and under what conditions it might work. The proposed hypothesis was equally testable: Continuous participatory audit with public response ledgers will detect and correct harm faster than periodic external review alone. Around it, the team marked assumptions in amber and evidence in blue. Almost everything was amber.

The first trial was designed as a scenario, not reported as a completed experiment. Participants would move through a baseline condition, an active comparison, a calibrated AI condition and a user-controlled condition. They would also face a failure or handback moment, because a system that works only when nothing goes wrong is not a safe system. The organising lens was the Capability–Access–Representation–Equity model applied to Harm & Auditability. Measures would cover immediate performance, understanding, confidence calibration, retained capability, burden, equity and the ability to recover without the tool.

In the scenario, the most revealing moment came after the apparent success. A participant completed the task quickly, then could not explain why the answer was sound. Another worked more slowly but spotted a subtle error and corrected it without assistance. Samira Patel refused to call either result decisive. The contrast showed why the study needed repeated measures, subgroup analysis, delayed follow-up and participant debriefing rather than a single productivity score. It also exposed the risk of measurement intrusion, which could make a polished outcome look more trustworthy than it was.

The team then tested the human boundary. Participants could reduce assistance, inspect sources, challenge recommendations and withdraw. People using non-dominant formats, different languages or accessibility settings were not treated as exceptions to be averaged away. Their experience was part of construct validity. The success gate remained demanding: Statistically and practically meaningful improvement; preserved or improved human skill and agency; acceptable burden; no disproportionate subgroup harm; reproducible performance; effective handback, correction and recovery. The failure gate was explicit too: No meaningful benefit; gains mask reduced understanding or skill; dependency, fatigue, stress or exclusion exceeds benefit; effects fail to transfer or persist; user control is ineffective; correction or handback fails.

By the end of the workshop, Samira Patel had no miracle result to announce. That was the point. The evidence state was recorded as “PARTIAL — structured hypothesis seed; empirical evidence not yet bound”. The story could therefore demonstrate the decision pressure and the proposed research design, but it could not claim efficacy, causation or generalisability. The next legitimate step was clear: Complete primary-source review for Harm & Auditability; appoint owner; define benchmark, comparison and measures; convene affected-user and expert review; pre-register protocol; establish handback, adverse-effect and recovery tests.

A commercial pathway still emerged, but in the right order. If the research survives pre-registration, ethical review, comparative testing and replication, harm & auditability assessment module could become a useful module for technology assurance. Services could include assessment, implementation support, training, certification and assurance. Policy makers could use the resulting measures to ask for evidence of agency, equity, handback and retained skill rather than accepting output uplift alone.

Later, Samira Patel returned to the original dashboard. It still showed green. Now the team understood that green was not a finding; it was an invitation to investigate. For Harm & Auditability, progress would mean building a traceable chain from question to protocol, participant experience, raw data, analysis, limitation and decision. The strongest closing lesson was simple: human capability must remain visible even when the machine makes performance look effortless.

Reflection

What did we learn?

Apparent improvement in Harm & Auditability must be separated from retained capability, agency, verification burden and recovery.

Why does this matter?

Harm & Auditability may materially affect human capability, independence, confidence, safety and productivity. Poorly designed assistance can create hidden costs even where short-term output appears to improve.

What research does this connect to?

CARE / CSO → Adoption and Equity; related register: Related subtopics within the same theme and adjacent themes.

What should happen next?

Complete primary-source review for Harm & Auditability; appoint owner; define benchmark, comparison and measures; convene affected-user and expert review; pre-register protocol; establish handback, adverse-effect and recovery tests.

Ground Zero scenario narrative — not an empirical finding.