← The framework← Het framework

In development · Research prototype v0.x · 4 October 2026. These instruments are being designed and tested. They are not yet available to download, are not validated, and must not be used to certify, approve or reject an AI system.

Responsible AI & Human Agency · Framework · Apply the framework

Practitioner instruments

How the framework becomes something an organisation can actually use: an organisational questionnaire and a practitioner workbook.

Two instruments with separate jobs

The strongest design is not one questionnaire that produces a responsible AI score. It is two instruments that deliberately do different things:

Organisational QuestionnaireGathers claims, documents, experience and operational evidence — from executives, engineers, lawyers, managers, frontline users, affected people and worker representatives.
Practitioner WorkbookTests those claims, triangulates the evidence, records uncertainty, and produces findings and a governance decision.

Respondents provide evidence. Assessors make findings. Decision-makers authorise or refuse use.

This separation prevents self-certification, and it makes disagreement useful. If a manager says staff are free to override the AI while frontline users say overriding slows them down against their productivity targets, that is not “bad questionnaire data”. It is evidence about organisational risk.

  1. Question
  2. Claim
  3. Evidence
  4. Assessor interpretation
  5. Domain finding
  6. Governance decision

“Yes, employees can override the AI” is an E0 assertion. It does not demonstrate meaningful oversight. Stronger evidence would include a controlled error-detection exercise, production override and escalation records, an independent check, and frontline evidence that disagreement is practically possible.

Release path

  1. v0.x · Research prototypeYou are here. Method and instruments under development. Explained and previewed online; not released for use.
  2. v0.9 · Public betaReleased for structured testing and feedback, with the evidence base and limitations documented.
  3. v1.0 · Validated practitioner releaseRevised after cognitive testing, three pilots, assessor-reliability testing and a verified standards mapping.
ResourceOnlinePlanned downloadStatus
Framework overviewRead onlinePDFIn development
Practitioner WorkbookExplained and previewed belowSpreadsheet + PDF reference editionResearch prototype
Organisational QuestionnaireExplained, with the rapid screen previewed belowSpreadsheet or fillable formResearch prototype
Evidence guideIndicators · human agency · human rightsPDFIn development
Standards mappingRead onlineSpreadsheet + PDFPartly verified
Pilot reportsAfter each pilotPDFNo pilot has run yet

The workbook is spreadsheet-first: assessors need to record evidence, findings, uncertainties, red flags, owners, deadlines and reassessment dates. The PDF will be the stable reference edition. When the instruments are released, they will be open — no email address required — with a version number, a change log and a clear statement of their validation status.

What the status labels mean

  • Validated — the specific measure has empirical validation for the construct claimed. It does not mean the whole framework item is validated.
  • Adapted — a recognised measure, method or control translated into an AI-specific setting.
  • Proposed — specific to this framework and not yet validated. Most questions start here.
  • Requires source verification — a legal or standards mapping that must be checked against the authoritative text, as with ISO/IEC 42001 clauses.

Preview: the rapid screen

Thirty-two questions, designed for a facilitated cross-functional session of roughly 60–90 minutes when the evidence is available. Its job is routing: to send each AI use to the right depth of assessment. All questions are Proposed.

#QuestionAsked ofEvidence
Entry gate
01What specific problem or need is this AI use intended to solve?Use ownerE0
02What uses are explicitly out of scope or prohibited?Owner, legalE0
03What evidence shows the problem exists and merits intervention?OwnerE1 · E2
04What is the best feasible non-AI or less intrusive alternative, and how does it compare?OwnerE1 · E2
05What human or public good could be lost even if AI improves speed or output?Owner, frontline, affectedE4
06Which judgments, relationships or responsibilities must remain meaningfully human?Owner, frontlineE0 · E4
07Could an error affect rights, health, livelihood, employment, education or essential services?Legal, riskE0 · E1
08Are effects hard to reverse, or are affected people vulnerable or unable to refuse?Risk, affectedE0 · E4
Technical
09Has the system been validated for this task, population and setting — not only on generic benchmarks?ProductE1
10Are the most consequential failure modes known and tested under stress or change?Product, riskE1
11Have errors and outcomes been examined across relevant groups, languages and accessibility needs?Product, legalE1 · E2
12Does the system signal uncertainty or decline to answer in a way that has been tested?ProductE1
13Are data use, permissions, security and privacy risks proportionate to the purpose?Privacy, securityE1
14Can consequential outputs, versions and human actions be reconstructed afterwards?Product, riskE1 · E2
Human
15Can users still do the underlying task adequately without AI after sustained use?Frontline, managerE1 · E2 · E4
16Have users been tested on spotting plausible but wrong AI output?Frontline, productE1
17Does reliance change appropriately when the AI is right, wrong or uncertain?Human factorsE1 · E2
18Do reviewers have the information, time, competence and alternative evidence to challenge the AI?Manager, frontlineE2 · E4
19Can reviewers actually disregard, override, pause or stop the process?Manager, frontlineE1 · E2 · E4
20What happens to a worker who disagrees with the AI when targets point the other way?Frontline, managerE2 · E4
21Can affected people use a human or non-AI route where needed, without unreasonable penalty?Affected, legalE2 · E4
22Could the interface make people overestimate the system's competence, empathy or authority?Design, riskE1 · E4
Organisational
23Is there a named owner with authority to restrict, suspend or stop the use?Executive, riskE0
24Do adoption, throughput or cost targets dominate safety, rights and human outcomes?Executive, managerE0 · E2
25Were affected people involved before the decision — and can you name something they changed?Executive, riskE0 · E4
26Have the applicable legal duties and impact assessments been identified for this use and jurisdiction?LegalE0
27Has someone independent of the deployment's success challenged the evidence?Risk, auditE3
28Can an affected person find out about the AI's role, add context and reach someone who can change the outcome?Affected, legalE1 · E2 · E4
29Are complaints, corrections, incidents and near misses recorded — and able to change the system?Risk, operationsE2
Over time
30Which technical, human and organisational indicators will be monitored after deployment?Risk, ownerE0
31Which changes trigger reassessment: model, data, population, workflow, supplier, authority, scale or law?Risk, productE0
32When does the authorisation expire, and who must actively re-authorise it?Accountable ownerE0

A screen escalates to the full assessment when, among other things: the entry gate cannot establish a proportionate purpose; a legal red line may be involved; rights, livelihood, safety or essential services are at stake; oversight has not been tested behaviourally; subgroup performance is unknown for affected groups; affected people cannot practically challenge outcomes; policy and frontline accounts contradict each other; or the supplier limits testing, logging, change notice or exit. Unknown is not the same as satisfactory.

The full questionnaire: nine tracks

No single person completes the whole questionnaire. Each role answers only what it knows or experiences. Some claims are asked of several groups on purpose, to expose gaps between policy and practice.

TrackTarget timeSample question
Executive and accountable owner25–35 minWho can restrict or stop this use without asking the team responsible for its success?
Product and technical35–50 minWhich supplier or model changes can reach production without your own testing?
Legal, privacy and compliance30–45 minWho has the authority to provide effective remedy when an outcome is wrong?
Risk, audit and responsible AI30–45 minWhich positive claims currently rest only on assertions?
Managers20–30 minWhat happens when an employee disagrees with the AI? Give the most recent example.
Frontline users20–30 minDescribe a case where the AI was wrong and how you noticed.
Affected people15–25 minCould you reach a person with real authority to reconsider the outcome?
Worker representatives20–30 minWho captures the productivity gains, and who absorbs the extra checking and rework?
Independent assessor30–45 minCould anything in this assessment have changed the deployment decision — or is it documenting one already made?

Times are design targets, not validated timings. Completion time is one of the things the pilots will measure.

Policy–practice contradiction tests

ClaimAsked of managementAsked of frontline or affected peopleChecked against
“Humans can override”How are overrides handled?What happens when you disagree?Override logs, seeded errors
“Users understand the limits”How are staff trained?Where should you not rely on it?Knowledge and error-detection test
“Appeals are available”What is the procedure?Could you reach someone who could change it?Appeal funnel, abandonment, corrections
“AI reduces workload”Where does saved time go?Has your workload or checking changed?Baseline versus production workload
“Stakeholders participated”Who was consulted?Did your input change anything?Change log showing actual influence
“No retaliation for challenge”How is disagreement treated?Is it safe to challenge the AI?Grievance patterns, performance consequences

A contradiction opens an issue for the assessor. It is never settled by majority vote.

The workbook

  1. Assessment charterBefore any findingThe six elements of the AI use, stakes and reversibility, named owners, who can stop the system, conflicts of interest, and the expiry date and change triggers — set before authorisation.
  2. Entry gate10 testsPurpose, baseline, suitability, necessity, intrusion, stakes and reversibility, power and vulnerability, distribution, purpose integrity and the human-primary boundary.
  3. Eighteen domain modulesOne formatEach with a definition, core and diagnostic questions, strong versus weak evidence, candidate indicators, red flags, finding guidance, actions and reassessment triggers. The first pilots will work five in full: validity, fairness, retained capability, meaningful oversight, and incidents and remedy.
  4. Rights reviewAcross all layersThe Rights-to-Evidence lens, including the rights-in-practice path from “formally available” to “remedies the harm”.
  5. RegistersLiving recordsAn evidence register (each item tagged with its evidence class and perspective, validity period and what would invalidate it); an assumptions and uncertainty register; a red-flag register; and a question traceability matrix linking every question to its claim, domain, evidence and source.
  6. Decision recordNo scoreDomain findings, red flags, decisive evidence, dissenting views, the chosen outcome and why stronger or weaker outcomes were rejected — plus conditions, owner, stop authority and expiry date.
  7. Continuous assurance planAfter the decisionSignals, baselines, owners and escalation states: normal → watch → review → intervene → suspend or withdraw.

Every decision record asks: what evidence from this assessment materially changed, constrained or could have changed the original proposal? If the answer is repeatedly “none”, either the instrument is not sensitive enough — or decisions are being made before the assessment.

Measurement proportionality check. Before collecting any data from workers, users or affected people, the workbook asks: is this evidence necessary? Is there a less intrusive way to get it? Could collecting it shift power, autonomy or privacy? Who can see it? Could it be reused for performance management? When will it be deleted?

How the instruments will be validated

Research-grounded is not the same as validated. Before release, the instruments will be tested on:

  • Content validity — do the questions cover what each domain claims to assess? Expert panels classify every question as essential, useful, redundant or unclear, adapting principles from the COSMIN content-validity methodology.1
  • Response process — do executives, engineers, lawyers, frontline staff and affected people understand the questions as intended? Tested through cognitive interviews.23
  • Known-failure sensitivity — do assessors catch planted problems, such as strong average accuracy with poor subgroup performance, or a formal override with no realistic time to use it?
  • Assessor reliability — do two assessors reach broadly the same findings from the same evidence?
  • Gaming resistance — can a respondent create a misleadingly positive picture with policy language alone?
  • Burden and usefulness — completion time, evidence-retrieval time, and whether findings change decisions.

The planned sequence: instrument design → cognitive testing → rapid-screen pilot → three full pilots → assessor-reliability testing → revisions → public beta → further field validation → v1.0.

Careful with the words

This is an independent assessment methodology. It is not a certification scheme and does not constitute legal advice. Words such as “certified”, “compliant”, “approved” and “audited” should not be used about any assessment made with these instruments until what such claims mean has been deliberately established.

The aim is not to make responsible AI easy to score, but to make consequential claims about responsible AI difficult to make without evidence.

Sources

  1. Methodology Terwee, C. B. et al. (2018). COSMIN methodology for evaluating the content validity of patient-reported outcome measures: a Delphi study. Quality of Life Research, 27, 1159–1170. doi:10.1007/s11136-018-1829-0. ↩
  2. Book Willis, G. B. (2005). Cognitive Interviewing: A Tool for Improving Questionnaire Design. SAGE. doi:10.4135/9781412983655. ↩
  3. Research centre CDC Collaborating Center for Questionnaire Design and Evaluation Research (CCQDER). cdc.gov. ↩