← The framework← Het framework

In development · Indicator set v0.4 · 4 October 2026. These indicators are proposals for piloting. Most build on established methods, but none has been validated as part of this framework, and none comes with a universal threshold.

Responsible AI & Human Agency · Framework · Methods & measures

Indicators for every domain

What evidence each part of the framework asks for — at the entry gate, in the technical layer and in the organisational layer.

Why this page exists

A framework that names domains but not evidence is a list of principles. This page completes the indicator set, so that every one of the framework's two entry-gate tests and eighteen domains has at least one observable indicator. The human layer has its own, fuller method in Measuring Human Agency, and human rights run across all layers through the Rights-to-Evidence method.

Entry gate · 2 testsThis page
Technical · 6 domainsThis page
Human · 6 domainsHuman agency method
Organisational · 6 domainsThis page
Rights lens · across allRights-to-Evidence

Reuse before you add. Much of the technical and organisational evidence already exists in organisations that follow the EU AI Act, the GDPR or the NIST AI RMF. The “reuse” column points to where. The framework then asks one further question: does that evidence justify this specific use, with these people, in this workflow?

How to read each indicator

  • Method — how the evidence is collected.
  • Evidence class — E0 assertion, E1 controlled, E2 production, E3 independent, E4 human corroboration. Classes are combined according to the claim, not ranked.
  • Status — Established a well-known method used here largely as intended; Adapted an established method applied in a new way; Proposed my synthesis, not yet tested.
  • Review trigger — what should prompt a closer look. Thresholds are set per use, in advance, against a baseline — never borrowed from another context.

Entry-gate tests

TestIndicator and methodEvidenceStatusReview triggerReuse
ProportionalityAlternative comparison. A documented comparison of the AI use against the best realistic non-AI or human-plus-process alternative, covering value, intrusion, cost and risk — reviewed by someone outside the business case.E0 → E3AdaptedNo alternative was considered; the case rests on labour savings alone; the reviewer was the sponsorNIST MANAGE 2.1; MAP 3.1–3.2; GDPR Art. 5(1)(c)
Purpose integrityAffected-person success criteria and human-primary boundary. Success criteria agreed with affected people before deployment; the answers to the Purpose Integrity Test; a written list of functions that must remain human.E0 · E4ProposedSuccess defined only by the budget owner; no list of what must stay human; efficiency case disappears once the human core is restoredNIST MAP 1.1; study 01

Layer 1 · Technical responsibility

DomainIndicator and methodEvidenceStatusReview triggerReuse
Validity & accuracyLocal validation against baseline. Performance on a representative sample of this use's real cases, against an agreed outcome, with confidence intervals, compared with the current process.E1 · E2EstablishedBelow the margin agreed in advance; a supplier's claim not reproduced locallyAI Act Arts. 13(3)(b), 15; NIST MEASURE 2.5
Robustness & reliabilityStress and drift. Tests on shifted, edge and adversarial cases before deployment; production drift monitoring; time to recover from failure.E1 · E2EstablishedDrift beyond the agreed band; any untested model, data or supplier changeAI Act Arts. 15, 72; NIST MEASURE 2.4, 3.1
Fairness & subgroup performanceDisaggregated errors with a coverage statement. Error rates by relevant and intersectional group, with uncertainty, plus an explicit list of groups too small to evaluate.E1 · E2 · E3AdaptedUnexplained disparity; a materially affected group is not evidencedAI Act Art. 10; NIST MEASURE 2.11
Uncertainty & abstentionCalibration and abstention test. Whether stated confidence matches actual accuracy, and the share of out-of-scope or unanswerable test cases where the system declines, flags or escalates.E1AdaptedConfident wrong answers on out-of-scope cases; uncertainty shown but not linked to any actionNIST MAP 2.2; AI Act Art. 13(3)(b)
Security & privacyLeast-privilege and attack review. Permissions compared with what the use needs; red-team and prompt-injection findings, closed and retested; privacy tests per the rights lens.E1 · E3EstablishedOpen critical findings; broad or shared permissions for an AI agentAI Act Art. 15(5); GDPR Art. 32; NIST MEASURE 2.7, 2.10
TraceabilityReconstruction test. For a random sample of consequential outputs, the share that can be fully reconstructed: model version, inputs and sources, human action and outcome.E1 · E2AdaptedAny consequential case that cannot be reconstructedAI Act Arts. 12, 26(6); NIST GOVERN 1.6

Two reasons not to rely on supplier evidence alone. A widely deployed sepsis prediction model, validated externally across 38,455 hospitalisations, missed 67% of patients with sepsis.4 And modern neural networks are often poorly calibrated: their stated confidence does not match how often they are right.5 Both are only visible through local testing.

Layer 3 · Organisational responsibility

DomainIndicator and methodEvidenceStatusReview triggerReuse
AccountabilityReal-authority test. A named owner with documented stop authority, information and budget — and a stop drill: time from the decision to stop until the use has actually stopped.E0 · E1ProposedNo drill ever run; the stop takes longer than harm can develop; the owner cannot fund a fixAI Act Art. 26; NIST GOVERN 2.1, 2.3
Incentives & cultureIncentive audit. Every target and bonus linked to the use; the share of the scorecard that measures responsible outcomes; an anonymous, cohort-level check on whether raising concerns has had consequences.E0 · E4ProposedOnly speed and adoption are rewarded; people report that raising concerns is riskyNIST GOVERN 4.1; study 08
Participation & inclusionInfluence log. Input received from affected people, the organisation's response, and what changed as a result; when they entered the process; whether a non-AI option could be argued.E0 · E4ProposedParticipation repeatedly leaves no trace in decisions; affected people entered after procurementGDPR Art. 35(9); NIST GOVERN 5.1
Legal complianceLegal register and applicability check. Which instruments apply and why; impact assessments (DPIA, FRIA where required) complete and in date; open legal findings.E0 · E3EstablishedAn assessment out of date after a material change; applicability never analysedNIST GOVERN 1.1; AI Act Art. 27; GDPR Art. 35
Independent challengeIndependence test and intervention count. The assessor's reporting line, evidence requested versus granted, recorded conflicts — and how often independent findings actually changed a decision.E3ProposedEvidence refused; the assessor reports to the sponsor; findings never change anythingNIST MEASURE 1.3; study 10
Incident response & remedyLearning loop. Near-miss reports, time to contain, recurrence of the same failure, corrective actions verified as closed, and remedy actually delivered to the people affected.E2 · E4AdaptedThe same failure recurs; corrective actions closed without verification; no remedy reached affected peopleAI Act Art. 73; NIST GOVERN 4.3, MEASURE 3.3, MANAGE 2.3

Read these in pairs. A rising number of near-miss reports is often good news — it means people feel safe reporting. A stop authority that has never been used may be healthy, or untested. An organisation remains answerable for what its systems tell people: in Moffatt v. Air Canada, a tribunal held the airline responsible for wrong information from its website chatbot.6

Rules for using these indicators

  1. Baseline first. Measure the current process before AI wherever possible, so improvement and deterioration can both be seen.
  2. Commit in advance. Set review triggers before looking at the results.
  3. No averaging. Each domain gets its own finding — demonstrated, partly evidenced, not evidenced or red flag — with its evidence classes.
  4. One indicator is never the domain. An indicator is evidence about a domain, not a substitute for judgment about it.
  5. Measure proportionately. Use samples and cohorts, not continuous individual monitoring.

Method and limits

The indicators are my proposals, drawing on established practice in AI evaluation, security, data protection and governance. The “reuse” column points to related provisions in the EU AI Act,1 the GDPR2 and the NIST AI RMF 1.0;3 it does not mean those provisions apply to every use. Whether each indicator is practical, and whether it changes decisions, is what the pilots will test.

Sources

  1. Legislation Regulation (EU) 2024/1689 (AI Act). eur-lex.europa.eu. ↩
  2. Legislation Regulation (EU) 2016/679 (GDPR). eur-lex.europa.eu. ↩
  3. Framework NIST (2023). AI Risk Management Framework (AI RMF 1.0), NIST AI 100-1. doi:10.6028/NIST.AI.100-1. ↩
  4. External validation Wong, A. et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine. doi:10.1001/jamainternmed.2021.2626. ↩
  5. Peer-reviewed Guo, C., Pleiss, G., Sun, Y. & Weinberger, K. Q. (2017). On Calibration of Modern Neural Networks. ICML 2017. arXiv:1706.04599. ↩
  6. Case law Moffatt v. Air Canada, 2024 BCCRT 149. canlii.org. ↩