In development · Research status: method development. The Rights-to-Evidence method and the six evidence perspectives are proposals. Many of the measures below are adapted, emerging or missing — not validated AI-specific rights measures — and none comes with a universal threshold.
Responsible AI & Human Agency · Framework · Human rights & evidence
Measuring Human Rights in AI
What evidence would justify the claim that a human right is actually protected in practice?
A metric is not a right
Saying that AI should “respect human rights” is easy. The hard question is what would show that it does. Human rights affected by AI can be translated into measurable evidence — but not into a single metric, threshold or “human rights score”.
A metric can provide evidence relevant to a human right. It cannot, by itself, prove that the right has been respected.
Differential privacy, re-identification tests, subgroup error rates, accessibility conformance and autonomy scales each measure one specific thing. None settles the underlying rights question on its own. This is why the framework treats indicators as evidence about whether the conditions a right protects exist in practice — never as mathematical substitutes for the right.
Rights are normative constraints. Indicators are evidence. Triangulation supports findings. Findings support decisions. Continuous assurance tests whether the evidence remains true.
A lens across the whole framework, not a fourth layer
Human rights do not sit neatly in one layer. Privacy has technical and organisational dimensions. Non-discrimination runs through model performance, workflow and outcomes. Autonomy is largely human, but depends on organisational choices. Remedy needs organisational authority and the experience of the people affected. So in the framework, human rights and evidence form a cross-cutting assurance lens applied across the technical, human and organisational layers — not another box in the stack.
The Rights-to-Evidence method
Start from the right, not from whatever data is easiest to collect.
- RightWhich right may be affected?
- Protected human conditionWhat is that right actually protecting?
- Observable constructWhat would protection or impairment look like in practice?
- IndicatorsWhat could we observe or measure?
- EvidenceWhat supports the indicator?
- InterpretationWhat does the evidence establish — and what does it not?
- FindingDemonstrated · partly evidenced · not evidenced · red flag
- DecisionContinue · restrict · redesign · suspend · withdraw
This prevents a common error: picking a convenient operational number and declaring, after the fact, that it measures a human right. For privacy, for example: the protected condition is not being subjected to unnecessary, disproportionate or insecure processing and inference; one observable construct is model leakage; one indicator is how well a membership-inference attack performs; the evidence is a controlled adversarial test. What that justifies is a finding about one form of privacy leakage — not “privacy: 93% protected”.
Ten rights domains for AI assessment
| Right | The protected human condition AI assurance should examine |
|---|---|
| Privacy & data protection | People are not subjected to unnecessary, disproportionate, insecure or unexpectedly inferential uses of information about them. |
| Equality & non-discrimination | Groups do not systematically receive unjustifiably worse treatment, service, error burdens or opportunities. |
| Autonomy & choice | People keep a meaningful capacity to judge, choose, refuse or seek an alternative where the context requires it. |
| Dignity & recognition | People are treated as persons, not merely as scores, behavioural targets or objects of manipulation. |
| Accessibility | Disability or access needs do not stop people from using, understanding or challenging an AI-mediated service. |
| Participation & inclusion | People materially affected can influence design, success criteria, deployment or review — not only be consulted. |
| Contestability & remedy | People can discover AI involvement where relevant, challenge errors, add context, get meaningful human reconsideration and obtain correction or remedy. |
| Procedural fairness | Decisions are consistent, intelligible and impartial enough for the context, with voice, reasons and review where required. |
| Labour & workplace rights | AI does not bring unjustified surveillance, intolerable demands, discriminatory treatment or loss of protected working conditions and worker voice. |
| Expression & information | Where AI mediates speech or information, people keep the ability to seek, receive and impart information without unjustified restriction. |
These domains overlap. An inaccessible appeal route is at once an accessibility, contestability, procedural-fairness and possibly a discrimination problem. Indicators should be reused across rights rather than counted as separate points.
Six evidence perspectives
The UN human rights indicator method distinguishes structural, process and outcome indicators: from commitment, through effort, to results.1 For AI, I propose adding three perspectives — behavioural, experience and distributional — because sociotechnical systems can have every safeguard on paper while failing the people they affect. The first three are established; the last three are my proposed extension.
| Perspective | Question | Example: an appeal route | Cannot establish alone |
|---|---|---|---|
| Structural | Are the rules, authority and safeguards in place? | “There is an appeal process.” | That it works |
| Process | Are the safeguards being implemented? | “Appeals are processed.” | That outcomes are acceptable |
| Behavioural extension | Can and do people exercise the capability or right? | “People can successfully submit appeals.” | That every group can |
| Outcome | What actually happens to people? | “Wrong decisions are corrected.” | Why, or whether it is lawful |
| Experience extension | How do affected people experience it? | “People find the process understandable and usable.” | Legality or actual control |
| Distributional extension | Do burdens or benefits differ across groups? | “Some groups abandon or lose appeals far more often.” | Whether a disparity is unlawful |
Are we measuring the existence of a safeguard — or whether the safeguard actually protects people?
Process evidence must never be quietly converted into outcome evidence. “We did a DPIA”, “there is an appeal button”, “we comply with WCAG” or “our provider is ISO/IEC 42001 certified” can all be relevant. None shows that privacy, remedy, accessibility or rights are actually protected in a specific use.
Rights in practice
A right can exist formally and still be almost useless. The framework measures how far along this path a right actually reaches:
- Formally available
- Findable
- Understandable
- Exercisable
- Changes the process
- Corrects the outcome
- Remedies the harm
The same path applies to contestability, consent and choice, access to a human, data rights, worker voice, participation, accessibility and remedy. It is the difference between paper rights and operational rights.
Privacy: measure each property separately
Privacy is unusually measurable — and a clear illustration of why no single metric represents a right. The GDPR's principles cover lawfulness, purpose limitation, data minimisation, accuracy, storage limitation, security and accountability, with protection by design and by default.2
- Data necessity
- Did we need these data for this purpose? A documented, independently reviewed rationale for each attribute.
- Identifiability
- Can records be re-identified? Removing names is not anonymisation: sparse datasets can be linked using outside information,3 so de-identification is risk management, not a yes-or-no state.4 Pseudonymised data remain personal data.5
- Model leakage
- Can a model reveal whether someone was in its training data? Membership-inference attacks are an established test for this,6 reported with the attacker model, the baseline and retesting after mitigation.
- Formal guarantees
- Differential privacy gives a mathematically defined bound on how much an output depends on any one person,7 provided the parameters, accounting and implementation are right.8
- Rights in practice
- Are access, correction and deletion requests completed correctly and on time?
- Incidents
- What happens in production: unauthorised access, disclosure and recurring complaints.
A technical privacy guarantee is evidence about a defined privacy property. It is not proof that privacy as a whole is protected.
Differential privacy does not answer whether collection was necessary, whether a secondary use is lawful, whether retention is excessive or whether internal access is proportionate.
Non-discrimination: no single fairness metric
Fairness has several mathematically incompatible definitions. When outcome rates differ between groups, calibration and equal error rates generally cannot all hold at once.910 So choosing the fairness metric is itself a substantive decision about which errors and burdens matter most.
Broad-group averages are not enough. Commercial gender-classification systems showed large accuracy gaps once sex and skin tone were analysed together,11 and parity across large groups can coexist with unfairness across their subgroups — “fairness gerrymandering”.12 Where a group is too small to evaluate reliably, the finding is not evidenced — not an assumption of equality.
Matched-case field testing — randomly varying a protected characteristic while holding everything else constant — is an established way to detect differential treatment,13 and can be applied to AI-assisted recruitment pipelines. Sensitive data need not end the analysis either: research has shown fairness can sometimes be audited with encrypted protected attributes.14
How mature is the measurement?
Status labels: Validated established validity for the stated construct — never legal compliance on its own; Adapted established elsewhere and reasonably adaptable to AI; Emerging plausible and increasingly used, but not validated for AI; Missing no credible standalone measure yet.
| Area | Status | What is still missing |
|---|---|---|
| Formal privacy guarantees | Validated for defined technical properties | A route from technical guarantees to whole-right assurance |
| Privacy leakage testing | Validated · Adapted | Standard threat models and context-specific acceptance criteria |
| Accessibility conformance | Adapted, highly mature | AI-specific outcomes for conversational, adaptive and agentic systems |
| Subgroup performance | Adapted | A defensible link from statistics to substantive and legal equality |
| Intersectional fairness | Adapted · Emerging | Reliable small-group methods that protect privacy and keep statistical power |
| Matched-case discrimination testing | Validated method | Validation for complex, multi-stage and changing AI workflows |
| Procedural-justice perception | Validated construct, Adapted to AI | A link between perceived fairness and the actual ability to correct decisions |
| Workload and job quality | Validated construct, Adapted to AI | Longitudinal evidence that changes are caused by AI |
| Algorithmic-management exposure | Validated in gig-work samples | Validation in conventional employment, other sectors and cultures |
| Autonomy perception | Validated construct, Adapted | A link between felt autonomy, real choice and rights protection |
| Contestability effectiveness | Emerging | A validated end-to-end measure from notice to effective remedy |
| Participation and influence | Emerging | A measure that tells genuine influence from consultation theatre |
| Dignity and relational integrity | Missing as a universal measure | Behavioural and experience measures linked to the normative idea of dignity |
| Expression and information | Emerging | Ways to separate legitimate moderation from rights-impairing restriction |
Supporting instruments include WCAG 2.2 — whose own guidance says conformance does not guarantee usability for everyone and recommends testing with disabled users15 — the Copenhagen Psychosocial Questionnaire for workload, work pace and influence,16 the Algorithmic Management Questionnaire, validated across three samples of gig workers,17 and a validated work version of the basic psychological needs scale for autonomy.18
The hardest rights to evidence are not the least recognised ones. They are those with the greatest distance between principle and observation: dignity, meaningful autonomy, substantive participation and effective contestability.
Leading and lagging evidence
| Right | Leading evidence | Lagging evidence |
|---|---|---|
| Privacy | Necessity review; re-identification and membership-inference tests; access review | Breaches, disclosures, complaints, data-rights outcomes |
| Non-discrimination | Representative test sets; subgroup power analysis; matched-case tests | Outcome disparities, complaints, successful appeals |
| Accessibility | WCAG and assistive-technology testing; disabled-user task tests | Abandonment, failed transactions, exclusion complaints |
| Autonomy | A real alternative route; manipulation review; simulated refusal | Forced reliance, unreachable human alternative, reported coercion |
| Contestability | Appeal simulation; notice comprehension; reviewer authority test | Uptake, abandonment, overturn rate, time to remedy |
| Labour | Workload baseline; algorithmic-management exposure; worker consultation | Work intensification, strain, grievances, loss of influence |
| Procedural fairness | Matched-case consistency; quality of reasons; reviewer independence | Overturns, repeat errors, perceived injustice |
Zero complaints can mean zero harm — or no awareness, inaccessible procedures, fear of retaliation or no confidence that complaining helps. Denominators matter: “twenty successful appeals” means nothing without knowing how many people received an adverse decision, knew they could appeal, tried, and gave up.
Findings per right — nothing averaged
Findings are made right by right, never averaged into a percentage. A worked illustration:
| Right | Finding | Why |
|---|---|---|
| Privacy | Demonstrated | Necessity review, access logs, an independently replicated leakage test and no unresolved incidents |
| Non-discrimination | Partly evidenced | Good results for major groups; too few cases for two materially affected groups |
| Accessibility | Demonstrated | WCAG conformance plus assistive-technology and disabled-user task testing |
| Autonomy | Not evidenced | An opt-out exists in policy, but nothing shows declining AI is practical or penalty-free |
| Contestability | Red flag | An appeal exists, but people cannot add evidence and reviewers cannot change the outcome |
A hypothetical illustration, not an assessment of a real system.
Strong privacy cannot cancel an ineffective remedy. High accuracy cannot cancel an inaccessible service. And a policy declaring that humans remain in control cannot cancel evidence that they cannot actually intervene.
Rights-respecting measurement
The method used to measure a human right must itself respect human rights.
Measuring discrimination often needs protected-characteristic data, which brings its own privacy risk. The answer is neither “collect everything” nor “measure nothing”, but measurement proportionality:
- collect only what the specific assurance question needs;
- separate identity from analytical data where feasible, and aggregate where individual data are unnecessary;
- restrict access and retention, and use controlled environments for highly sensitive audits;
- consider differential privacy or cryptographic methods where scientifically appropriate;
- never publish small groups in ways that allow re-identification;
- never protect worker autonomy by individually monitoring every worker — measure at cohort level.
Stress tests in five settings
| Setting | Rights under most pressure | Escalate when… |
|---|---|---|
| AI-assisted recruitment | Non-discrimination, privacy, accessibility, procedural fairness, contestability | Unexplained subgroup disparity persists; affected groups are not evidenced; the application or appeal route is inaccessible; no substantive human review |
| Public eligibility decisions | Procedural fairness, remedy, non-discrimination, privacy, dignity | People cannot practically challenge loss of essential services; the burden of proving an algorithmic error falls on vulnerable claimants |
| Workplace AI | Labour rights, privacy, autonomy, dignity, participation | Unjustified work intensification or loss of influence; excessive monitoring; discipline driven by unchallengeable AI; retaliation for challenge |
| Customer-facing chatbot | Autonomy, dignity, privacy, accessibility, contestability | Users cannot reach a human for consequential disputes; the design misleads about competence or authority; sensitive data is elicited unnecessarily |
| Clinical decision support | Privacy, non-discrimination, autonomy, procedural fairness | Subgroup performance deteriorates; clinicians cannot detect plausible failures; patient review becomes ceremonial |
Three lessons hold across all five: baselines matter — without pre-AI measures, deterioration and improvement cannot be shown; denominators matter; and distribution matters at every stage — equal predictive accuracy can coexist with one group bearing longer appeals, heavier documentation, more monitoring or fewer human alternatives.
Research agenda
- A contestability funnel — seed correctable errors and measure, step by step, whether affected people notice, find the route, complete it, reach an empowered reviewer and get the error corrected.
- AI-specific autonomy — combine felt autonomy with observable refusal, correction and use of alternatives.
- Dignity and relational integrity — identify observable conditions that reliably predict affected people's experience. Not a “dignity percentage”.
- Privacy-preserving equality measurement and intersectional methods for small groups.
- Longitudinal workplace effects, and measurement invariance across languages, cultures and disabilities.
- Threshold validation — baselines, uncertainty and meaningful-change criteria instead of invented cut-offs.
Method and limits
A structured synthesis of human rights indicator methodology, legislation, standards and empirical research, prepared in October 2026. It is not legal advice. The Rights-to-Evidence method, the three added evidence perspectives, the rights-in-practice path and the maturity ratings are my proposals. Each source below was checked against its published record.
Sources
- UN methodology OHCHR (2012). Human Rights Indicators: A Guide to Measurement and Implementation (HR/PUB/12/5). ohchr.org. ↩
- Legislation Regulation (EU) 2016/679 (GDPR), Articles 5 and 25. eur-lex.europa.eu. ↩
- Peer-reviewed Narayanan, A. & Shmatikov, V. (2008). Robust De-anonymization of Large Sparse Datasets. IEEE Symposium on Security and Privacy, 111–125. doi:10.1109/SP.2008.33. ↩
- Government guidance NIST IR 8053, De-Identification of Personal Information. csrc.nist.gov. ↩
- Regulatory guidance European Data Protection Board. Guidelines 01/2025 on Pseudonymisation. edpb.europa.eu. ↩
- Peer-reviewed Shokri, R., Stronati, M., Song, C. & Shmatikov, V. (2017). Membership Inference Attacks Against Machine Learning Models. IEEE Symposium on Security and Privacy, 3–18. doi:10.1109/SP.2017.41. ↩
- Monograph Dwork, C. & Roth, A. (2014). The Algorithmic Foundations of Differential Privacy. Foundations and Trends in Theoretical Computer Science, 9(3–4), 211–487. doi:10.1561/0400000042. ↩
- Government guidance NIST SP 800-226, Guidelines for Evaluating Differential Privacy Guarantees. csrc.nist.gov. ↩
- Peer-reviewed Chouldechova, A. (2017). Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments. Big Data, 5(2), 153–163. doi:10.1089/big.2016.0047. ↩
- Peer-reviewed Kleinberg, J., Mullainathan, S. & Raghavan, M. (2017). Inherent Trade-Offs in the Fair Determination of Risk Scores. Innovations in Theoretical Computer Science (ITCS). arXiv:1609.05807. ↩
- Peer-reviewed Buolamwini, J. & Gebru, T. (2018). Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. Proceedings of FAT* 2018, PMLR 81. proceedings.mlr.press. ↩
- Peer-reviewed Kearns, M., Neel, S., Roth, A. & Wu, Z. S. (2018). Preventing Fairness Gerrymandering: Auditing and Learning for Subgroup Fairness. ICML 2018. arXiv:1711.05144. ↩
- Field experiment Bertrand, M. & Mullainathan, S. (2004). Are Emily and Greg More Employable Than Lakisha and Jamal? American Economic Review, 94(4), 991–1013. doi:10.1257/0002828042002561. ↩
- Peer-reviewed Kilbertus, N. et al. (2018). Blind Justice: Fairness with Encrypted Sensitive Attributes. ICML 2018. arXiv:1806.03281. ↩
- Standard W3C. Web Content Accessibility Guidelines (WCAG) 2.2 and Understanding Conformance. w3.org · conformance. ↩
- Validation study Burr, H. et al. (2019). The Third Version of the Copenhagen Psychosocial Questionnaire. Safety and Health at Work, 10(4), 482–503. doi:10.1016/j.shaw.2019.10.002. ↩
- Validation study Parent-Rocheleau, X. et al. (2024). Creation of the algorithmic management questionnaire: A six-phase scale development process. Human Resource Management, 63(1), 25–44. doi:10.1002/hrm.22185. ↩
- Validation study Olafsen, A. H. et al. (2021). The Basic Psychological Need Satisfaction and Need Frustration at Work Scale: A Validation Study. Frontiers in Psychology, 12. doi:10.3389/fpsyg.2021.697306. ↩