Tech4Humanity AtlasGround ZeroCurrent ThemesFuture ResearchGalleryLive Q&ASearch

Human-AI Cognition & Performance / Human–AI Collaboration

SUB-0031 · Story

Escalation and Handover

Day 19, 16:40. Anika wrote one sentence in the margin of the protocol: “The improvement may not belong to the intervention.” Practice and emerging evidence suggest that escalation and handover is a distinct determinant of outcomes within human–ai collaboration, but current approaches are inconsistent.

In Germany, Anika's team at a regional health service had been asked to explore escalation and Handover. The immediate pressure was practical: current approaches to escalation and handover are fragmented, poorly calibrated or insufficiently measured, making it difficult to distinguish real benefit from substitution, novelty or surveillance effects. People could see activity, outputs and confident recommendations, but those signals did not establish that capability, safety or agency had improved.

Anika resisted turning the scenario into a success story too early. As a patient advocate, Anika knew that a memorable example can clarify a research problem, but it cannot validate a causal claim. The team therefore framed one answerable question: What information is required for safe transfer from AI to a person and back again? The story gave the work human stakes; the question gave it a boundary.

The working hypothesis was specific enough to fail: Structured handover packets containing goal, state, evidence and unresolved risks will reduce recovery time and errors. That wording changed the conversation. Instead of asking whether the idea sounded beneficial, the team had to compare conditions, define what improvement meant, and decide what evidence would count against the intervention. They also had to test whether a short-term gain concealed dependence, reduced understanding, new exclusion or a difficult handback when assistance disappeared.

The proposed study centred on workflow experiments, task-allocation trials, simulation, incident analysis. The design varied Primary variables: handover content, timing, urgency, channel and observed time-to-context, repeated work, omission rate, handover success. Subgroup and accessibility analysis were not treated as optional additions. A result that helped an average participant while predictably harming a smaller group would not satisfy the programme's definition of success.

During the imagined pilot, the most useful moment was not a dramatic breakthrough. It was a disagreement. One participant completed the task faster but reported less control; another moved more slowly yet retained the process after support was withdrawn. Anika asked the team to record both observations without choosing a preferred ending. They were scenario prompts, not findings, and they exposed why performance alone could not carry the evaluation.

The team built recovery into the protocol. Participants could challenge a recommendation, inspect relevant reasoning, pause the intervention and resume unaided. Failure scenarios tested changed conditions and incomplete information. Delayed follow-up asked whether any advantage persisted and whether people could still act independently. This made the study less theatrical and more useful: the system had to support correction and handback, not merely produce an impressive first result.

The unknowns remained visible: Effect size, causal mechanism, subgroup variation, optimal dose, long-term persistence, transfer beyond the test context, implementation cost.. The principal risks included diffused accountability, automation bias, hidden agent actions, deskilling. None could be resolved by the narrative itself. They required sourced literature, approved ethics and accessibility review, a pre-registered protocol, traceable evidence and reproducible analysis.

If the hypothesis is supported, the value could extend beyond one pilot in health and care. Target: improve outcomes relating to escalation and handover while preserving agency, skill, dignity, accessibility and sustainable human capability. The same evidence could inform product requirements, assurance services, training, procurement criteria and policy guidance. If the hypothesis is not supported, that result would still be valuable by preventing a weak approach from scaling behind attractive claims.

At the closing review, Anika replaced the original programme claim with a more honest sentence: “We know what must be tested next.” Integrated assurance protocol for escalation and handover linking immediate performance, retained human capability, agency, wellbeing, equity, longitudinal adaptation, safe handback and recovery. For the people represented by the story, progress would not mean a system doing more. It would mean a person remaining more capable when the system stepped back.

Reflection

What did we learn?: The scenario shows why escalation and Handover must be evaluated as a human-capability claim, not inferred from activity or short-term output. It also shows why assistance, burden, agency, subgroup effects, handback and recovery belong in the same evaluation.

Why does this matter?: Escalation and Handover may materially affect human capability, independence, confidence, safety and productivity. Poorly designed assistance can create hidden costs even where short-term output appears to improve.

What research does this connect to?: This subtopic sits within Human–AI Collaboration and draws on team cognition, organisational design, safety engineering and collaborative AI. Existing work provides useful foundations but rarely integrates individual differences, AI behaviour, long-term adaptation and measurable human outcomes in one programme. Related subtopics: Human–AI Task Allocation; Shared Decision Making; Human Oversight Design.

What should happen next?: Complete primary-source review for Escalation and Handover; appoint owner; define benchmark, comparison and measures; convene affected-user and expert review; pre-register protocol; establish handback, adverse-effect and recovery tests.

Research connection

Hypothesis: Structured handover packets containing goal, state, evidence and unresolved risks will reduce recovery time and errors.

Scientific uncertainty: Effect size; causal mechanism; subgroup variation; optimal dose; long-term persistence; transfer beyond the test context; implementation cost.

Variables: Primary variables: handover content; timing; urgency; channel; outcome variables: recovery; errors; duplicated effort; accountability; contextual and control variables: baseline capability; prior exposure; age; motivation; context; technology access; implementation fidelity.

Research methods: Workflow experiments; task-allocation trials; simulation; incident analysis; ethnography; decision audits; controlled handover tests; pre-registered analysis; active comparison; subgroup and accessibility analysis; delayed retention or longitudinal follow-up; adverse-effect capture; reproducibility testing; participant debrief.

Evidence: Validated instruments for time-to-context; repeated work; omission rate; handover success; pre-registered protocol; representative sample; baseline and comparison condition; raw and derived data; analysis code; consent and ethics records; subgroup results; limitations; authoritative primary sources; representative and accessible samples; documented comparison; analysis code; raw and derived data; subgroup analysis; delayed retention or longitudinal evidence; adverse-effect, handback and recovery records.

Frameworks: Goal–Role–Authority–Evidence–Handover model applied to Escalation and Handover, linking baseline capability, context, AI intervention, observable outcome, subjective burden, retained skill, handback and recovery.

Links: NIST AI RMF — https://www.nist.gov/itl/ai-risk-management-framework; ISO human-centred AI — https://www.iso.org/; OECD AI Principles — https://oecd.ai/; UNESCO AI Ethics — https://www.unesco.org/.

Commercialisation and public value

Products: Human-agent orchestration layer; oversight console; escalation router; responsibility map; decision receipt system; Escalation and Handover assessment module; Escalation and Handover intervention toolkit.

Services: Enterprise, education and consumer subscriptions; adaptive-assistance modules; analytics and assurance services; benchmark licensing; implementation support; training and certification; sector-specific human-performance solutions.

Industries: Decision support; service delivery; software development; operations; healthcare; finance; government; emergency response.

Government: Workers; team leaders; boards; risk officers; AI vendors; unions; customers; regulators; auditors; operations controllers; incident-response teams.

Policy: Accountability; duty of care; auditability; human override; worker consultation; delegated authority.

Future research: Complete primary-source review for Escalation and Handover; appoint owner; define benchmark, comparison and measures; convene affected-user and expert review; pre-register protocol; establish handback, adverse-effect and recovery tests.

Business opportunity: Develop and validate a minimum safe handover schema; package the evidence into research cards, implementation guidance, assessment instruments and reusable data assets.

Scenario narrative — not an empirical finding.