Tech4Humanity AtlasGround ZeroCurrent ThemesFuture ResearchGalleryLive Q&ASearch

Human-AI Cognition & Performance / AI Assistance Calibration

SUB-0009 · Evidence — SEEDED / PARTIAL — defensible non-empirical baseline; validation and replayable search outstanding

Optimal AI Assistance Thresholds

1. Hypothesis

Performance will follow an inverted-U curve, with the optimum varying by expertise and task novelty.

2. Experiment design

Design: multi-arm adaptive assistance experiment with delayed retention testing; baseline, active comparison, calibrated AI condition, user-controlled condition, failure or handback scenario and delayed follow-up Methods: multi-arm assistance-level experiments; longitudinal usage analysis; skill-retention tests; preference elicitation; A/B testing; pre-registered analysis; active comparison; subgroup and accessibility analysis; delayed retention or longitudinal follow-up; adverse-effect capture; reproducibility testing; participant debrief Independent variables: assistance intensity; user expertise; task novelty; error cost Dependent variables: accuracy; speed; learning retention; autonomy Confounders: baseline capability; prior exposure; age; motivation; context; technology access; implementation fidelity Measures: effect-size curve; delayed unaided test; reliance rate; satisfaction Success criteria: Statistically and practically meaningful improvement; preserved or improved human skill and agency; acceptable burden; no disproportionate subgroup harm; reproducible performance; effective handback, correction and recovery. Failure conditions: No meaningful benefit; gains mask reduced understanding or skill; dependency, fatigue, stress or exclusion exceeds benefit; effects fail to transfer or persist; user control is ineffective; correction or handback fails.

3. Seed result / current evidence

DEFENSIBLE SEED RESULT — NON-EMPIRICAL. The current evidence supports Optimal AI Assistance Thresholds as a testable research proposition. Problem basis: Current approaches to optimal ai assistance thresholds are fragmented, poorly calibrated or insufficiently measured, making it difficult to distinguish real benefit from substitution, novelty or surveillance effects. Directional expectation: If the hypothesis is supported, the intervention condition should improve accuracy; speed; learning retention; autonomy while avoiding material deterioration in independence, confidence calibration or delayed performance. Proposed observations: effect-size curve; delayed unaided test; reliance rate; satisfaction. Seed data profile: Evidence Strength 10/100; Confidence 25/100; Maturity 20/100; Overall Health 36/100; Novelty 78/100; Strategic Importance 95/100. Evidence boundary: No validated results yet.; experiments 0, studies 0, participants 0. This is suitable for protocol formation and baseline comparison, not as a finding of effect.

4. Seed conclusion

DEFENSIBLE SEED CONCLUSION — PROVISIONAL. Optimal AI Assistance Thresholds warrants structured testing because the CSV identifies a defined problem, falsifiable hypothesis, measurable outcomes and relevant literature foundations. The present position is that “Performance will follow an inverted-U curve, with the optimum varying by expertise and task novelty.” is plausible and decision-relevant, but unvalidated. Proceed to controlled testing against the stated success and failure conditions. Confirm, narrow or reject this seed after effect sizes, uncertainty, subgroup outcomes, adverse effects, persistence and handback performance are observed.

Prior-art search performed before starting

PRIOR-ART SEED BASELINE — PARTIAL. The CSV records these literature domains: Automation bias; cognitive offloading; adaptive automation; trust calibration; scaffolding and expertise-reversal effects. It also records: NIST AI RMF — https://www.nist.gov/itl/ai-risk-management-framework; ISO 9241 — https://www.iso.org/; OECD AI Principles — https://oecd.ai/; UNESCO AI Ethics — https://www.unesco.org/. Evidence register status: “Seeded; authoritative source register refreshed; empirical evidence not yet ingested”. This is defensible as a starting prior-art inventory, but not as proof of a completed systematic search because search dates, databases, exact queries, reviewer, result counts, screening decisions, claim mapping and a replayable receipt are absent.

Prior-art material named: Existing literature: Automation bias; cognitive offloading; adaptive automation; trust calibration; scaffolding and expertise-reversal effects. References: NIST AI RMF — https://www.nist.gov/itl/ai-risk-management-framework; ISO 9241 — https://www.iso.org/; OECD AI Principles — https://oecd.ai/; UNESCO AI Ethics — https://www.unesco.org/

Critical gap / next action

Create and attach a dated prior-art search log; lock the protocol; execute the proposed study; link raw data and analysis; then replace the results and conclusion placeholders with evidence-bounded findings.

Evidence classification: SEEDED / PARTIAL — defensible non-empirical baseline; validation and replayable search outstanding — provisional research record, not a validated finding.