← The framework← Het framework

In development · Research status: evidence review and method development. The dimensions and indicators below are being tested, not settled. Several are proposals that have not been validated, and none has a universal threshold.

Responsible AI & Human Agency · Framework · Methods & measures

Measuring Human Agency in AI Systems

How do we know whether people still have meaningful agency when they use — or are affected by — AI?

“Fine. How would you know?”

The framework makes a normative claim: responsible AI must preserve human agency. The first question any sceptical organisation, auditor or regulator will ask is how you would know. This page is the answer in development. It sits underneath the human layer of the framework — retained capability, appropriate reliance, meaningful oversight, autonomy and choice, dignity and relational integrity, and human impact — and asks what evidence would establish each claim.

Do not ask only whether people feel able to exercise agency. Test whether they can.

A person may say they understand an AI system and still fail to identify its limits. They may say they can challenge it and still accept deliberately wrong recommendations. They may believe they are in control while workload, interface design, targets or the lack of an alternative make disagreement unlikely. Most responsible AI governance stops at declared human control. This method aims for demonstrated human agency.

From perception to outcome

  1. PerceptionWhat do people believe?
  2. KnowledgeWhat do they understand?
  3. CapabilityWhat could they do?
  4. BehaviourWhat do they actually do?
  5. OutcomeWhat changes as a result?

All five are shaped by organisational conditions — time, authority, workload, incentives, alternatives — and by repeated AI use over time.

The chain exposes the gaps that ordinary assurance misses:

  • Perception–capability gap: people say they could challenge an output, but fail a challenge task.
  • Capability–behaviour gap: they can technically override, but rarely do, because of workload or incentives.
  • Behaviour–outcome gap: overrides happen, but usually make decisions worse.
  • Institutional gap: someone detects an error but has no authority to change the outcome.
  • Longitudinal gap: oversight works at first, but the underlying skill fades with sustained automation.

Different evidence establishes different claims

ClaimWeak evidenceStronger evidence
People understand the AIThey say they doThey correctly identify its limits, failure conditions and when to escalate
Humans remain in controlPolicy says they can overrideSeeded wrong recommendations are actually rejected or corrected
AI supports expertise rather than replacing itProfessionals feel competentUnaided performance stays within an agreed margin after sustained use
People can challenge decisionsAn appeal link existsPeople can find it, complete it, reach a competent human and get a wrong decision corrected
Users rely appropriatelyAverage trust scoreCorrect advice is accepted and wrong advice resisted, compared with human-only and AI-only baselines
Oversight is meaningfulA reviewer is assignedThe reviewer has time, evidence, competence and authority — and intervention tests show the workflow changes when it should

Surveys are not bad evidence; they are evidence of perception. This maps directly onto the framework's evidence levels: a policy is E0, a seeded-error test is E1, production override data is E2, an independent retest is E3, and evidence from the people doing or receiving the work is E4.

Six candidate dimensions

The evidence review suggests six dimensions. They are a working hypothesis: pilots will test whether they are empirically distinct, or whether some — such as challenge capacity and appropriate reliance — should be merged.

  1. Understand & orientKnow the AI's role, scope, limits, uncertainty and when to escalate — shown, not self-rated.
  2. Judge & rely appropriatelyForm an independent judgment, use good AI advice and resist bad advice.
  3. Challenge & interveneDetect trouble and actually change what happens — with the time, information and authority to do so.
  4. Contest & obtain remedyFor affected people: notice, a way to add context, a reachable human, correction and remedy.
  5. Retain capability & autonomyStill able to do the work, and to choose, when AI is unavailable.
  6. Sustain human conditionsWorkload, wellbeing, accessibility, voice and incentives that make formal control usable.

Two people, two kinds of agency

Most human-oversight models focus on the operator. But in hiring, credit, healthcare, welfare and customer service, the person most affected by an AI-supported decision may never touch the system. A tool can feel empowering to a recruiter while leaving the applicant with very little agency.

The operator or decision-makerUnderstand → judge → challenge → intervene → retain capability
The affected personNotice → understand → provide context → reach a human → contest → correct → obtain remedy

For affected people, the shift is from formal rights to exercisable rights. “Users can appeal” is not the same as showing that people know they can appeal, can find and use the route without disproportionate burden, reach an empowered human, correct wrong information and obtain timely remedy. This connects human agency to digital human rights, not only to workplace design. GDPR Article 22 is one legal example: for solely automated decisions with legal or similarly significant effects, it requires safeguards including human intervention and the right to contest.14

What can already be measured

Each measure carries one of four labels: Validated substantially validated for this construct and context; Adapted a validated instrument that needs revalidation for AI use; Experimental empirically supported, but without standard interpretation or thresholds; Proposed a governance indicator I have synthesised, not yet tested. Validated never means universally thresholded.

ConstructInstrument or methodStatusLimit
Perceived trustTrust in Automation Scale and its three-item short form, validated for AI applications1ValidatedMeasures attitude, not whether reliance is correct. Never a success target.
Appropriate relianceRelative AI reliance and relative self-reliance: accepting correct advice and resisting wrong advice, measured separately2ExperimentalNeeds cases with known correct answers; no universal threshold.
Situation awarenessSAGAT — objective queries about task state, supported by a meta-analysis of 243 studies3AdaptedQueries must be built and validated for each oversight task.
Review workloadNASA Task Load Index4AdaptedNo score at which oversight is known to fail.
Job autonomyWork Design Questionnaire5AdaptedMeasures work design generally, not AI-specific authority.
Felt authority and impactPsychological Empowerment Scale6AdaptedFeeling empowered can coexist with having no real permissions.
Voice and procedural fairnessOrganizational Justice Scale7AdaptedPerceived voice needs behavioural evidence of influence.
WellbeingWHO-5, reviewed across 213 studies8AdaptedChanges cannot be attributed to AI without a comparison.
Explanation qualityExplanation satisfaction and mental-model measures9AdaptedSatisfaction with an explanation does not prove understanding.
AI literacyExisting scales: a review found 16 scales across 22 validation studies, with none showing positive evidence on every quality criterion assessed10AdaptedMostly self-report; cannot replace a task-specific comprehension test.
Challenge capabilitySeeded-error detection: realistic wrong recommendations mixed into the workExperimentalSeeded cases must reflect real failure modes and harm.
Retained capabilityRepeated unaided performance against a pre-AI baselineExperimentalLearning, case mix and changing work complicate attribution.
Effective correctionOverride effectiveness: the share of overrides that improve the decision — not override frequencyProposedNeeds ground truth or expert review.
Effective contestabilityFindability, completion, time and cost to human review, correction and remedyProposedHigh reversal can mean a working appeal route — or a poor upstream system.

Why supported performance is not retained capability

  • Medicine. In a multicentre observational study, endoscopists' unaided adenoma detection fell from 28.4% to 22.4% after regular AI exposure.11 A 2026 scoping review found evidence of clinical deskilling scarce but consistent across specialties.12 This is a risk signal, not a proven causal law.
  • Education. In a field experiment with high-school mathematics students, AI assistance improved performance during practice, but unrestricted access harmed later performance without AI. A tutor designed to give hints rather than answers mitigated the loss.13
  • Clinical judgment. In a study of 223 dermatologists, higher trust was associated with greater reliance on AI — including lower self-reliance when the dermatologist was right and the AI was wrong.15
  • Hiring. In a resume-screening experiment with 528 participants, recommendations from simulated race-biased AI shifted people's candidate choices toward the AI's preference.16 It is a preprint, so treat it as emerging evidence — but it shows why “a human reviews it” cannot be assumed to neutralise bias.

None of this means AI harms competence, or should be rejected. It means AI-supported performance and unaided capability are different outcomes, and must be measured separately.

Fourteen indicators for piloting

A minimum set to test in pilots. They are deliberately not combined into a score.

IndicatorHowStatusTriggers review when…
Demonstrated limitation comprehensionScenario test of scope, limits, uncertainty and escalation, alongside self-ratingAdaptedPeople say they understand but fail safety-critical items
Independent judgmentRecord the human decision before showing the AI, on a sampleExperimentalJudgments converge on the AI, especially when it is wrong
Seeded-error detectionMix realistic wrong outputs into the workExperimentalDetection falls below a margin agreed in advance
Appropriate reliance pairAccepting correct advice and resisting wrong advice, measured separatelyExperimentalOne improves while the other deteriorates
Harmful acceptanceShare of known wrong recommendations acted onExperimentalAny high-severity case — whatever the average accuracy
Override effectivenessShare of overrides that improve the outcome, and harmful overridesProposedEffectiveness falls. Override counts are never a target.
Intervention and escalationDrills: time to detect, stop, escalate and reach someone empoweredProposedIntervention is slower than harm can develop
Oversight capacityWorkload measure plus time per case, caseload and missed casesAdaptedWorkload rises while detection falls
Contestability effectivenessFindability, completion, abandonment, time to human review, correction, remedyProposedThe human route is missing, inaccessible or cannot correct known errors
Unaided capability retentionRepeat representative tasks without AI against a baselineExperimentalDecline beyond a margin agreed in advance
Autonomy and usable choiceAutonomy measures plus a real test of opting out and its consequencesAdaptedAn option exists on paper but carries friction or penalty
Workload and wellbeing trendValidated workload and wellbeing measures, by cohortAdaptedMaterial deterioration plausibly linked to the deployment
Accessibility and challenge parityCompletion of core, opt-out and appeal tasks across relevant groupsProposedAny group cannot use a core human or appeal route
Voice with demonstrated influenceRecord input, the organisation's response and what changedProposedParticipation repeatedly leaves no trace in decisions

How to read the evidence

No number means anything on its own. A very low override rate may mean the AI is excellent — or that reviewers have become passive. A high appeal-reversal rate may mean the appeal route works — or that the upstream system is poor. Zero complaints may mean zero harm — or that people do not know how to complain. High trust may reflect strong evidence — or persuasive design. Measures acquire meaning only through triangulation.

Thresholds follow five rules rather than generic cut-offs:

  1. Baseline first — measure human-only or pre-AI performance before deployment.
  2. Commit in advance — define acceptable deterioration before seeing the results.
  3. Match the stakes — a drafting assistant and clinical triage do not share a threshold.
  4. Weight by harm — a rare catastrophic failure is not averaged against many trivial successes.
  5. Read in pairs — workload with detection, agreement with correctness, appeals with accessibility.

The result is an evidence profile, not a score. For example: Retained capability — demonstrated (E2 behavioural, E4 corroborated). Appropriate reliance — partly evidenced (E1 only). Meaningful oversight — red flag: override authority exists, but realistic workload testing shows reviewers cannot scrutinise outputs.

Measurement must not become surveillance

Never create more intrusive monitoring to prove that AI is human-centred than is needed to test the specific human-risk claim.

  • Prefer controlled samples over continuous individual telemetry.
  • Report at team or cohort level; never rank individual employees.
  • Keep assurance evidence separate from performance management, and never reuse “agency metrics” as adoption quotas.
  • Collect the minimum data, with retention limits, and involve workers or their representatives.
  • Watch for reactivity, gaming, Goodhart effects and selection bias — the people most burdened by AI are often least likely to answer a survey.

The GDPR's principles of purpose limitation and data minimisation set the legal baseline wherever personal data is involved.14

Stress tests in four settings

SettingWhat conventional metrics missKey test
Clinical decision supportReviewers appear to review but follow wrong advice; skills erodeCan a clinician resist a plausible wrong recommendation — and still perform safely after long AI use?
RecruitmentReviewers inherit AI bias; applicants cannot see or challenge the decisionDoes a recruiter resist a biased recommendation — and can an applicant get meaningful reconsideration?
EducationBetter answers today, weaker learning tomorrowDoes the learner still have the capability when AI is unavailable?
Customer serviceFast automated resolution hides an inaccessible route to a humanCan a frustrated or vulnerable customer leave the automated path early enough?

In customer service, relational design matters too: human-like cues in chatbots can increase compliance with their requests,17 while for customers who are already angry, anthropomorphism lowered satisfaction.18

Where measurement is mature — and where it is not

Validated foundationsPerceived trust · workload · situation awareness · job autonomy and empowerment · procedural fairness · wellbeing
Adaptable but incompleteExplanation satisfaction · AI literacy · perceived understanding · voice · autonomy under AI
Emerging empirical measuresAppropriate reliance · automation-bias resistance · seeded-error detection · independent judgment · skill retention · dependence
Genuine gapsMeaningful oversight as one construct · practical contestability · affected-person agency · relational dignity · perceived versus actual control · incentive effects · long-term dependence

Where the research shows that part of human agency cannot yet be measured reliably, that finding belongs here too. It is not a failure of the project; it is what makes the method honest.

Research agenda

  1. Construct validation — are the six dimensions distinguishable, or do some overlap?
  2. Criterion validity — do seeded-error and reliance tests predict real-world incident prevention?
  3. Longitudinal capability — baselines, exposure measures and comparison groups.
  4. Affected-person measurement — a practical contestability protocol.
  5. Organisational causation — which conditions turn formal override authority into exercised authority?
  6. Accessibility and measurement invariance across languages, disabilities and backgrounds.
  7. Threshold validation — testing baseline-relative and harm-weighted rules, not inventing cut-offs.
  8. Measurement ethics — a data-protection assessment for every indicator.

Human agency is demonstrated when people have the understanding, retained capability, practical authority and institutional support to form judgment, use AI selectively, detect and challenge failure, change consequential outcomes, contest decisions that affect them — and keep functioning when AI is unavailable. No survey, override button or workflow diagram can establish that on its own.

Method and limits

A structured evidence review, not a registered systematic review or meta-analysis, prepared in October 2026. Instrument statuses describe suitability for measuring human agency around AI, not whether a parent instrument has ever been validated. Several figures from my working research could not be matched to a verifiable public record and were left out. The six dimensions, the fourteen indicators, the threshold rules and the evidence profile are my proposals and have not been validated.

Sources

  1. Validation study McGrath, M. J. et al. (2025). Measuring trust in artificial intelligence: validation of an established scale and its short form. Frontiers in Artificial Intelligence. doi:10.3389/frai.2025.1582880. ↩
  2. Experimental Schemmer, M. et al. (2023). Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations. Proceedings of IUI 2023. doi:10.1145/3581641.3584066. ↩
  3. Meta-analysis Endsley, M. R. (2021). A Systematic Review and Meta-Analysis of Direct Objective Measures of Situation Awareness. Human Factors, 63(1). doi:10.1177/0018720819875376. ↩
  4. Instrument Hart, S. G. & Staveland, L. E. (1988). Development of NASA-TLX (Task Load Index). Advances in Psychology, 52, 139–183. doi:10.1016/S0166-4115(08)62386-9. ↩
  5. Validation study Morgeson, F. P. & Humphrey, S. E. (2006). The Work Design Questionnaire (WDQ). Journal of Applied Psychology, 91(6), 1321–1339. doi:10.1037/0021-9010.91.6.1321. ↩
  6. Validation study Spreitzer, G. M. (1995). Psychological empowerment in the workplace: dimensions, measurement and validation. Academy of Management Journal, 38(5). doi:10.2307/256865. ↩
  7. Validation study Colquitt, J. A. (2001). On the dimensionality of organizational justice: a construct validation of a measure. Journal of Applied Psychology, 86(3), 386–400. doi:10.1037/0021-9010.86.3.386. ↩
  8. Systematic review Topp, C. W. et al. (2015). The WHO-5 Well-Being Index: A Systematic Review of the Literature. Psychotherapy and Psychosomatics, 84(3). doi:10.1159/000376585. ↩
  9. Methods Hoffman, R. R. et al. (2023). Measures for explainable AI. Frontiers in Computer Science, 5. doi:10.3389/fcomp.2023.1096257. ↩
  10. Systematic review Lintner, T. (2024). A systematic review of AI literacy scales. npj Science of Learning, 9. doi:10.1038/s41539-024-00264-4. ↩
  11. Observational Budzyń, K. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. Lancet Gastroenterology & Hepatology, 10(10), 896–903. doi:10.1016/S2468-1253(25)00133-5. ↩
  12. Scoping review Heudel, P. E. et al. (2026). Artificial intelligence in medicine: a scoping review of the risk of deskilling and loss of expertise among physicians. ESMO Real World Data and Digital Oncology, 12. doi:10.1016/j.esmorw.2026.100693. ↩
  13. Field experiment Bastani, H. et al. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS. doi:10.1073/pnas.2422633122. ↩
  14. Legislation Regulation (EU) 2016/679 (GDPR), Articles 5 and 22. eur-lex.europa.eu. ↩
  15. Experimental Küper, A. et al. (2025). Psychological Factors Influencing Appropriate Reliance on AI-enabled Clinical Decision Support Systems: Experimental Web-Based Study Among Dermatologists. Journal of Medical Internet Research, 27, e58660. doi:10.2196/58660. ↩
  16. Preprint No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy (2025). arXiv:2509.04404. arxiv.org. ↩
  17. Experimental Adam, M., Wessel, M. & Benlian, A. (2021). AI-based chatbots in customer service and their effects on user compliance. Electronic Markets, 31. doi:10.1007/s12525-020-00414-7. ↩
  18. Experimental Crolic, C. et al. (2022). Blame the Bot: Anthropomorphism and Anger in Customer–Chatbot Interactions. Journal of Marketing, 86(1). doi:10.1177/00222429211045687. ↩