In development · Indicator set v0.4 · 4 October 2026. These indicators are proposals for piloting. Most build on established methods, but none has been validated as part of this framework, and none comes with a universal threshold.
Responsible AI & Human Agency · Framework · Methods & measures
Indicators for every domain
What evidence each part of the framework asks for — at the entry gate, in the technical layer and in the organisational layer.
Why this page exists
A framework that names domains but not evidence is a list of principles. This page completes the indicator set, so that every one of the framework's two entry-gate tests and eighteen domains has at least one observable indicator. The human layer has its own, fuller method in Measuring Human Agency, and human rights run across all layers through the Rights-to-Evidence method.
Reuse before you add. Much of the technical and organisational evidence already exists in organisations that follow the EU AI Act, the GDPR or the NIST AI RMF. The “reuse” column points to where. The framework then asks one further question: does that evidence justify this specific use, with these people, in this workflow?
How to read each indicator
- Method — how the evidence is collected.
- Evidence class — E0 assertion, E1 controlled, E2 production, E3 independent, E4 human corroboration. Classes are combined according to the claim, not ranked.
- Status — Established a well-known method used here largely as intended; Adapted an established method applied in a new way; Proposed my synthesis, not yet tested.
- Review trigger — what should prompt a closer look. Thresholds are set per use, in advance, against a baseline — never borrowed from another context.
Entry-gate tests
| Test | Indicator and method | Evidence | Status | Review trigger | Reuse |
|---|---|---|---|---|---|
| Proportionality | Alternative comparison. A documented comparison of the AI use against the best realistic non-AI or human-plus-process alternative, covering value, intrusion, cost and risk — reviewed by someone outside the business case. | E0 → E3 | Adapted | No alternative was considered; the case rests on labour savings alone; the reviewer was the sponsor | NIST MANAGE 2.1; MAP 3.1–3.2; GDPR Art. 5(1)(c) |
| Purpose integrity | Affected-person success criteria and human-primary boundary. Success criteria agreed with affected people before deployment; the answers to the Purpose Integrity Test; a written list of functions that must remain human. | E0 · E4 | Proposed | Success defined only by the budget owner; no list of what must stay human; efficiency case disappears once the human core is restored | NIST MAP 1.1; study 01 |
Layer 1 · Technical responsibility
| Domain | Indicator and method | Evidence | Status | Review trigger | Reuse |
|---|---|---|---|---|---|
| Validity & accuracy | Local validation against baseline. Performance on a representative sample of this use's real cases, against an agreed outcome, with confidence intervals, compared with the current process. | E1 · E2 | Established | Below the margin agreed in advance; a supplier's claim not reproduced locally | AI Act Arts. 13(3)(b), 15; NIST MEASURE 2.5 |
| Robustness & reliability | Stress and drift. Tests on shifted, edge and adversarial cases before deployment; production drift monitoring; time to recover from failure. | E1 · E2 | Established | Drift beyond the agreed band; any untested model, data or supplier change | AI Act Arts. 15, 72; NIST MEASURE 2.4, 3.1 |
| Fairness & subgroup performance | Disaggregated errors with a coverage statement. Error rates by relevant and intersectional group, with uncertainty, plus an explicit list of groups too small to evaluate. | E1 · E2 · E3 | Adapted | Unexplained disparity; a materially affected group is not evidenced | AI Act Art. 10; NIST MEASURE 2.11 |
| Uncertainty & abstention | Calibration and abstention test. Whether stated confidence matches actual accuracy, and the share of out-of-scope or unanswerable test cases where the system declines, flags or escalates. | E1 | Adapted | Confident wrong answers on out-of-scope cases; uncertainty shown but not linked to any action | NIST MAP 2.2; AI Act Art. 13(3)(b) |
| Security & privacy | Least-privilege and attack review. Permissions compared with what the use needs; red-team and prompt-injection findings, closed and retested; privacy tests per the rights lens. | E1 · E3 | Established | Open critical findings; broad or shared permissions for an AI agent | AI Act Art. 15(5); GDPR Art. 32; NIST MEASURE 2.7, 2.10 |
| Traceability | Reconstruction test. For a random sample of consequential outputs, the share that can be fully reconstructed: model version, inputs and sources, human action and outcome. | E1 · E2 | Adapted | Any consequential case that cannot be reconstructed | AI Act Arts. 12, 26(6); NIST GOVERN 1.6 |
Two reasons not to rely on supplier evidence alone. A widely deployed sepsis prediction model, validated externally across 38,455 hospitalisations, missed 67% of patients with sepsis.4 And modern neural networks are often poorly calibrated: their stated confidence does not match how often they are right.5 Both are only visible through local testing.
Layer 3 · Organisational responsibility
| Domain | Indicator and method | Evidence | Status | Review trigger | Reuse |
|---|---|---|---|---|---|
| Accountability | Real-authority test. A named owner with documented stop authority, information and budget — and a stop drill: time from the decision to stop until the use has actually stopped. | E0 · E1 | Proposed | No drill ever run; the stop takes longer than harm can develop; the owner cannot fund a fix | AI Act Art. 26; NIST GOVERN 2.1, 2.3 |
| Incentives & culture | Incentive audit. Every target and bonus linked to the use; the share of the scorecard that measures responsible outcomes; an anonymous, cohort-level check on whether raising concerns has had consequences. | E0 · E4 | Proposed | Only speed and adoption are rewarded; people report that raising concerns is risky | NIST GOVERN 4.1; study 08 |
| Participation & inclusion | Influence log. Input received from affected people, the organisation's response, and what changed as a result; when they entered the process; whether a non-AI option could be argued. | E0 · E4 | Proposed | Participation repeatedly leaves no trace in decisions; affected people entered after procurement | GDPR Art. 35(9); NIST GOVERN 5.1 |
| Legal compliance | Legal register and applicability check. Which instruments apply and why; impact assessments (DPIA, FRIA where required) complete and in date; open legal findings. | E0 · E3 | Established | An assessment out of date after a material change; applicability never analysed | NIST GOVERN 1.1; AI Act Art. 27; GDPR Art. 35 |
| Independent challenge | Independence test and intervention count. The assessor's reporting line, evidence requested versus granted, recorded conflicts — and how often independent findings actually changed a decision. | E3 | Proposed | Evidence refused; the assessor reports to the sponsor; findings never change anything | NIST MEASURE 1.3; study 10 |
| Incident response & remedy | Learning loop. Near-miss reports, time to contain, recurrence of the same failure, corrective actions verified as closed, and remedy actually delivered to the people affected. | E2 · E4 | Adapted | The same failure recurs; corrective actions closed without verification; no remedy reached affected people | AI Act Art. 73; NIST GOVERN 4.3, MEASURE 3.3, MANAGE 2.3 |
Read these in pairs. A rising number of near-miss reports is often good news — it means people feel safe reporting. A stop authority that has never been used may be healthy, or untested. An organisation remains answerable for what its systems tell people: in Moffatt v. Air Canada, a tribunal held the airline responsible for wrong information from its website chatbot.6
Rules for using these indicators
- Baseline first. Measure the current process before AI wherever possible, so improvement and deterioration can both be seen.
- Commit in advance. Set review triggers before looking at the results.
- No averaging. Each domain gets its own finding — demonstrated, partly evidenced, not evidenced or red flag — with its evidence classes.
- One indicator is never the domain. An indicator is evidence about a domain, not a substitute for judgment about it.
- Measure proportionately. Use samples and cohorts, not continuous individual monitoring.
Method and limits
The indicators are my proposals, drawing on established practice in AI evaluation, security, data protection and governance. The “reuse” column points to related provisions in the EU AI Act,1 the GDPR2 and the NIST AI RMF 1.0;3 it does not mean those provisions apply to every use. Whether each indicator is practical, and whether it changes decisions, is what the pilots will test.
Sources
- Legislation Regulation (EU) 2024/1689 (AI Act). eur-lex.europa.eu. ↩
- Legislation Regulation (EU) 2016/679 (GDPR). eur-lex.europa.eu. ↩
- Framework NIST (2023). AI Risk Management Framework (AI RMF 1.0), NIST AI 100-1. doi:10.6028/NIST.AI.100-1. ↩
- External validation Wong, A. et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine. doi:10.1001/jamainternmed.2021.2626. ↩
- Peer-reviewed Guo, C., Pleiss, G., Sun, Y. & Weinberger, K. Q. (2017). On Calibration of Modern Neural Networks. ICML 2017. arXiv:1706.04599. ↩
- Case law Moffatt v. Air Canada, 2024 BCCRT 149. canlii.org. ↩