← Responsible AI & Human Agency← Verantwoorde AI & menselijke regie

In development · Working draft v0.1 · 4 October 2026. I am developing this framework in the open. It is a proposal, not yet verified clause by clause, piloted or validated — and it will change. Please do not use it to certify, approve or reject an AI system.

Responsible AI & Human Agency · Framework · In development

The Responsible AI & Human Agency Framework

From ten studies to one assessment method: how to establish whether an AI system is technically responsible while preserving human autonomy, capability, dignity and accountability.

What this note does

The ten studies in this series each answer one question. This note turns them into a single assessment method. It fixes six things before any workbook, questionnaire or pilot is built:

  1. What is assessed — the unit of assessment.
  2. The structure — an entry gate, three layers and a loop.
  3. The decision rules — red lines, weakest link and five outcomes, never a single score.
  4. The evidence standard — what counts as proof, and how strong it must be.
  5. Who assesses — independence, participation and conflicts of interest.
  6. What already exists — a map against law and standards, showing where this framework adds something and where it must not duplicate.

The framework assesses two things at the same time: the quality of the technology, and the quality of the human situation it creates — for the people who use it, the people affected by it and the organisation deploying it.

1 · The unit of assessment is a use, not a model

The same model can be proportionate in one setting and unacceptable in another. So the framework never assesses “model X”. It assesses an AI use, defined by six elements:

Purpose
What the use is for, and what it is not for.
System
Model, version, configuration, data, retrieval sources, tools and supplier.
Workflow
Where AI sits in the work, and where authority sits.
People
Users, people affected, and people who carry the work or the risk.
Organisation
Owner, incentives, oversight and remedy routes.
Context
Sector, legal regime, stakes and reversibility.

A change to any element — a new model version, a new population, a new target — is a reason to reassess. That is what links the method to continuous assurance.

2 · Architecture: a gate, three layers and a loop

The gate comes first because the layers can only tell you how to use AI responsibly. Only the gate can tell you whether to. A use that fails the gate is not assessed further — the outcome is do not use or human-primary.

3 · Eighteen domains, each traced to a study

Each layer has six domains. Each domain has one core question and at least one measurable indicator. The study column shows where the reasoning and evidence for that domain are set out.

DomainCore questionExample indicatorStudy
Entry gate
ProportionalityIs AI necessary, suitable and the least intrusive option for a legitimate purpose?Documented comparison against the best realistic non-AI alternative01
Purpose integrityDoes the activity keep the human good it exists to provide?Affected people's own success criteria, recorded before deployment01
Layer 1 · Technical responsibility
Validity & accuracyDoes it work for this task, population and setting?Local validation against outcome; error rate with confidence interval09
Robustness & reliabilityDoes performance hold under change and stress?Performance under drift, adversarial and edge-case tests10
Fairness & subgroup performanceWho does it serve worse?Error rates disaggregated by relevant groups and languages03
Uncertainty & abstentionCan it signal when not to rely on it?Calibration curve; tested abstention rate09
Security & privacyAre access, data and actions proportionate and protected?Least-privilege review; red-team findings closed10
TraceabilityCan an output and an action be reconstructed?Version register; complete logs for consequential outputs06
Layer 2 · Human responsibility
Retained capabilityCan people still do the work, and explain, challenge and correct outputs, without AI?Unaided performance before and after sustained use04
Appropriate relianceDo people accept correct outputs and reject wrong ones?Correct and harmful acceptance and rejection, measured on seeded cases09
Meaningful oversightDo reviewers have the information, time, competence and authority to disagree?Override rate and quality; review time per case against workload05
Autonomy & choiceCan people refuse, opt out or reach a human without penalty?Availability and use of a non-AI route; time to reach a human05
Dignity & relational integrityDoes the design misrepresent competence, agency or care?User testing of what people believe the system is and does09
Human impact & distributionWhere does saved time go, and who carries the cost?Workload, job quality and the share of gains reaching workers07
Layer 3 · Organisational responsibility
AccountabilityIs there a named owner who can answer, decide and stop?Owner with documented stop authority06
Incentives & cultureDo targets reward responsible outcomes, or only speed and adoption?Responsible-outcome measures in the same scorecard as adoption08
Participation & inclusionDid affected people have influence before the decision?Changes made because of affected-people input03
Legal complianceAre legal duties met — as a floor, not the finish line?Legal register; impact assessments completed and current02
Independent challengeCan someone without a stake in the deployment question it?Independent review with evidence access and escalation rights10
Incident response & remedyCan harm be reported, corrected and compensated?Appeal success rate; time to remedy; near-miss reporting volume06

Measuring people without surveilling them. Human indicators should be collected at team or cohort level, with consent, minimal data and a clear purpose — never turned into individual performance telemetry. A framework that protects human agency cannot be implemented through covert monitoring of the people it is meant to protect.

4 · Decision rules: no single score

Each domain receives one of four findings, together with its evidence level:

DemonstratedThe claim is supported at the required evidence level.
Partly evidencedSupported, but below the required level or with gaps.
Not evidencedNo adequate evidence either way.
Red flagEvidence of harm, or a red line is crossed.

Findings are combined by five rules:

  1. Red lines block. A prohibited practice, an illegitimate purpose, or a severe and irreversible rights or safety harm without effective remedy means do not use — whatever the other findings say.
  2. Weakest link, not average. Strong performance in one domain cannot offset a red flag in another. There is no total score to average out.
  3. Uncertainty raises the bar. The higher the stakes, vulnerability and irreversibility, the higher the evidence level required — and the more “not evidenced” counts against use.
  4. Layer 2 cannot be waived. A use with strong technical and organisational findings but unassessed human domains is not ready for scale.
  5. Every outcome is time-limited. Each decision names its conditions, its owner and its next review date.

The rules produce one of the five outcomes from study 01: do not use · human-primary · restricted pilot · use with strong safeguards · proportionate use. Reassessment can move a use in either direction.

5 · Evidence standard

The framework reuses the evidence levels from continuous assurance, so assessment and monitoring speak the same language:

  • E0 · Assertion — a policy, vendor or team says it is true.
  • E1 · Controlled test — pre-production or simulation evidence.
  • E2 · Production evidence — live outcome data.
  • E3 · Independent challenge — reproduced or stress-tested independently.
  • E4 · Human corroboration — confirmed by affected people, grievances or longitudinal human evidence.

Proposed minimums: E1 for any pilot; E2 for proportionate use; E3 for use with strong safeguards in high-stakes settings; and E4 for the human-layer domains of any use that affects people's rights, livelihoods or access to services. E0 alone never supports a positive finding.

The method also distinguishes four kinds of claim and never mixes them: legal requirement, voluntary standard, empirical finding and my proposed principle.

6 · Who assesses

RoleResponsibility
Use ownerProvides evidence and is accountable for the decision — but does not sign off alone.
Independent assessorTests the evidence, sets findings, has stop and escalation rights, and has no stake in the deployment's success.
Users and frontline professionalsSupply human-layer evidence: workload, reliance, capability and whether override is real.
Affected people or their representativesShape success criteria and purpose integrity before the decision, not only after deployment.
Worker representativesReview workplace effects, surveillance risk and the distribution of gains.
Legal, privacy and securityConfirm the legal floor and the red lines.

Every assessment records conflicts of interest — including commercial relationships with the AI supplier — and how they were managed.

7 · What already exists — and where the gaps are

The framework should not reinvent what law and standards already require. This indicative map shows which existing instruments address each area, and where none of them reaches.

AreaEU AI ActGDPRNIST AI RMFISO/IEC 42001OECD · CoECoverage
Proportionality & red linesArt. 5 prohibitionsArt. 5 purpose limitation, minimisation; Art. 35 DPIAMAP (context, purpose, benefits and costs)Cl. 6 planning, impact assessmentCoE Art. 16(4) moratorium or banPartial — no general necessity test for lawful uses
Purpose integrity————CoE Art. 7 dignity and autonomy (principle)Gap
Validity, robustness, securityArt. 15 accuracy, robustness, cybersecurityArt. 32 securityMEASURECl. 8 operationOECD 1.4; CoE Art. 12 reliabilityWell covered (for high-risk)
Fairness & subgroupsArt. 10 data governanceArt. 9 special categoriesMEASURE (harmful bias)Cl. 8 operationOECD 1.2; CoE Arts. 10, 17Well covered in principle
Transparency & explanationArts. 13, 50, 86Arts. 13–15, 22GOVERN, MAP—OECD 1.3; CoE Art. 8Well covered
Human oversightArts. 14, 26Art. 22GOVERN, MANAGECl. 8 operationCoE Art. 8Partial — requires oversight, rarely tests whether it works
Retained capability & dependencyArt. 4 AI literacy (indirect)———CoE Art. 20 digital literacy (indirect)Gap
Appropriate relianceArt. 14(4)(b) automation bias awareness—MEASURE (indirect)——Largely a gap
Relational integrityArt. 50 AI disclosure; Art. 5 manipulation———CoE Art. 15(2) notice of AIPartial — disclosure only
Human impact & distributionArt. 27 FRIA (some deployers)—MAP (impacts)Impact assessmentOECD 1.1 well-beingGap — no duty to measure who gains
Participation & inclusionArt. 27 FRIA (some deployers)Art. 35(9) views of data subjectsGOVERN, MAP—CoE Arts. 18, 19Partial — consultation, not influence
Accountability & remedyArts. 26, 73Arts. 5(2), 24, 77–82GOVERNCl. 5 leadership and rolesOECD 1.5; CoE Arts. 9, 14, 15Well covered
Incentives & culture——GOVERN (culture)Cl. 5 leadership—Largely a gap
Continuous assuranceArts. 9, 72, 73Arts. 32, 35(11)MANAGECls. 9, 10CoE Art. 16Well covered (for high-risk)

Indicative mapping at article, clause or function level, as of 4 October 2026 — my reading, not legal advice. Most AI Act high-risk duties apply from 2 December 2027. To be checked in full against each instrument before publication.

Existing instruments provide substantial coverage of technical risk, governance processes and certain rights impacts. Coverage is less systematic for retained human capability, appropriate reliance, relational integrity, the distribution of gains and organisational incentives — and for whether an activity still serves its purpose. That is where this framework aims to add value.

8 · Validation: three contrasting pilots

SettingWhy it tests the frameworkDomains under most pressure
Enterprise AI assistantLow stakes per task, high scale, strong adoption incentivesRetained capability, human impact & distribution, incentives
AI-assisted recruitmentHigh stakes for people who never chose the systemFairness, participation, meaningful oversight, remedy
Customer-facing conversational AIRelational design, vulnerable users, contested outcomesRelational integrity, appropriate reliance, autonomy & choice

Each pilot should report what the framework found that existing assessments did not, what it missed, how long it took, and whether any finding changed a decision. A framework that never changes a decision has not been validated — it has been rehearsed.

9 · What comes next

  1. Check the crosswalk clause by clause against each instrument.
  2. Define each indicator — method, data source, threshold and evidence level.
  3. Build the practitioner workbook and an organisational questionnaire.
  4. Run the three pilots and publish what changed.
  5. Publish version 1.0 as an open method, with a public log of changes.

Development path

  1. v0.1 · Working draft, published in developmentThis page — you are here. Architecture and decision rules proposed.
  2. Clause-by-clause verificationEvery crosswalk entry checked against the instrument's text.
  3. Substantive reviewDomains, indicators, thresholds and decision rules reviewed and revised.
  4. Three pilotsEnterprise AI assistant, AI-assisted recruitment, customer-facing conversational AI.
  5. Documented revisionsWhat the pilots found, what they missed, what changed in the framework.
  6. v0.9 · Public betaA verified and piloted method, open to challenge.
  7. v1.0 · Validated methodOnly once pilots show the framework can change real decisions.

Method, limits and independence

Independent and vendor-neutral. The framework is not affiliated with, commissioned by or endorsed by any employer, supplier, institution or standards body — including those referenced in the crosswalk.

Methodology disclosure. This framework is my own synthesis of the ten studies in this series, which draw on legislation, standards, regulatory guidance and empirical research. Referencing an instrument does not imply that its authors agree with this method. The architecture, domains, indicators, decision rules, evidence minimums and roles are proposals and have not been validated. The crosswalk is indicative.

Version status. All v0.x versions are developmental. They are not validated assessment standards and should not be used to certify, approve or reject an AI system.