In development · Research prototype v0.x · 4 October 2026. These instruments are being designed and tested. They are not yet available to download, are not validated, and must not be used to certify, approve or reject an AI system.
Responsible AI & Human Agency · Framework · Apply the framework
Practitioner instruments
How the framework becomes something an organisation can actually use: an organisational questionnaire and a practitioner workbook.
Two instruments with separate jobs
The strongest design is not one questionnaire that produces a responsible AI score. It is two instruments that deliberately do different things:
Respondents provide evidence. Assessors make findings. Decision-makers authorise or refuse use.
This separation prevents self-certification, and it makes disagreement useful. If a manager says staff are free to override the AI while frontline users say overriding slows them down against their productivity targets, that is not “bad questionnaire data”. It is evidence about organisational risk.
- Question
- Claim
- Evidence
- Assessor interpretation
- Domain finding
- Governance decision
“Yes, employees can override the AI” is an E0 assertion. It does not demonstrate meaningful oversight. Stronger evidence would include a controlled error-detection exercise, production override and escalation records, an independent check, and frontline evidence that disagreement is practically possible.
Release path
- v0.x · Research prototypeYou are here. Method and instruments under development. Explained and previewed online; not released for use.
- v0.9 · Public betaReleased for structured testing and feedback, with the evidence base and limitations documented.
- v1.0 · Validated practitioner releaseRevised after cognitive testing, three pilots, assessor-reliability testing and a verified standards mapping.
| Resource | Online | Planned download | Status |
|---|---|---|---|
| Framework overview | Read online | In development | |
| Practitioner Workbook | Explained and previewed below | Spreadsheet + PDF reference edition | Research prototype |
| Organisational Questionnaire | Explained, with the rapid screen previewed below | Spreadsheet or fillable form | Research prototype |
| Evidence guide | Indicators · human agency · human rights | In development | |
| Standards mapping | Read online | Spreadsheet + PDF | Partly verified |
| Pilot reports | After each pilot | No pilot has run yet |
The workbook is spreadsheet-first: assessors need to record evidence, findings, uncertainties, red flags, owners, deadlines and reassessment dates. The PDF will be the stable reference edition. When the instruments are released, they will be open — no email address required — with a version number, a change log and a clear statement of their validation status.
What the status labels mean
- Validated — the specific measure has empirical validation for the construct claimed. It does not mean the whole framework item is validated.
- Adapted — a recognised measure, method or control translated into an AI-specific setting.
- Proposed — specific to this framework and not yet validated. Most questions start here.
- Requires source verification — a legal or standards mapping that must be checked against the authoritative text, as with ISO/IEC 42001 clauses.
Preview: the rapid screen
Thirty-two questions, designed for a facilitated cross-functional session of roughly 60–90 minutes when the evidence is available. Its job is routing: to send each AI use to the right depth of assessment. All questions are Proposed.
| # | Question | Asked of | Evidence |
|---|---|---|---|
| Entry gate | |||
| 01 | What specific problem or need is this AI use intended to solve? | Use owner | E0 |
| 02 | What uses are explicitly out of scope or prohibited? | Owner, legal | E0 |
| 03 | What evidence shows the problem exists and merits intervention? | Owner | E1 · E2 |
| 04 | What is the best feasible non-AI or less intrusive alternative, and how does it compare? | Owner | E1 · E2 |
| 05 | What human or public good could be lost even if AI improves speed or output? | Owner, frontline, affected | E4 |
| 06 | Which judgments, relationships or responsibilities must remain meaningfully human? | Owner, frontline | E0 · E4 |
| 07 | Could an error affect rights, health, livelihood, employment, education or essential services? | Legal, risk | E0 · E1 |
| 08 | Are effects hard to reverse, or are affected people vulnerable or unable to refuse? | Risk, affected | E0 · E4 |
| Technical | |||
| 09 | Has the system been validated for this task, population and setting — not only on generic benchmarks? | Product | E1 |
| 10 | Are the most consequential failure modes known and tested under stress or change? | Product, risk | E1 |
| 11 | Have errors and outcomes been examined across relevant groups, languages and accessibility needs? | Product, legal | E1 · E2 |
| 12 | Does the system signal uncertainty or decline to answer in a way that has been tested? | Product | E1 |
| 13 | Are data use, permissions, security and privacy risks proportionate to the purpose? | Privacy, security | E1 |
| 14 | Can consequential outputs, versions and human actions be reconstructed afterwards? | Product, risk | E1 · E2 |
| Human | |||
| 15 | Can users still do the underlying task adequately without AI after sustained use? | Frontline, manager | E1 · E2 · E4 |
| 16 | Have users been tested on spotting plausible but wrong AI output? | Frontline, product | E1 |
| 17 | Does reliance change appropriately when the AI is right, wrong or uncertain? | Human factors | E1 · E2 |
| 18 | Do reviewers have the information, time, competence and alternative evidence to challenge the AI? | Manager, frontline | E2 · E4 |
| 19 | Can reviewers actually disregard, override, pause or stop the process? | Manager, frontline | E1 · E2 · E4 |
| 20 | What happens to a worker who disagrees with the AI when targets point the other way? | Frontline, manager | E2 · E4 |
| 21 | Can affected people use a human or non-AI route where needed, without unreasonable penalty? | Affected, legal | E2 · E4 |
| 22 | Could the interface make people overestimate the system's competence, empathy or authority? | Design, risk | E1 · E4 |
| Organisational | |||
| 23 | Is there a named owner with authority to restrict, suspend or stop the use? | Executive, risk | E0 |
| 24 | Do adoption, throughput or cost targets dominate safety, rights and human outcomes? | Executive, manager | E0 · E2 |
| 25 | Were affected people involved before the decision — and can you name something they changed? | Executive, risk | E0 · E4 |
| 26 | Have the applicable legal duties and impact assessments been identified for this use and jurisdiction? | Legal | E0 |
| 27 | Has someone independent of the deployment's success challenged the evidence? | Risk, audit | E3 |
| 28 | Can an affected person find out about the AI's role, add context and reach someone who can change the outcome? | Affected, legal | E1 · E2 · E4 |
| 29 | Are complaints, corrections, incidents and near misses recorded — and able to change the system? | Risk, operations | E2 |
| Over time | |||
| 30 | Which technical, human and organisational indicators will be monitored after deployment? | Risk, owner | E0 |
| 31 | Which changes trigger reassessment: model, data, population, workflow, supplier, authority, scale or law? | Risk, product | E0 |
| 32 | When does the authorisation expire, and who must actively re-authorise it? | Accountable owner | E0 |
A screen escalates to the full assessment when, among other things: the entry gate cannot establish a proportionate purpose; a legal red line may be involved; rights, livelihood, safety or essential services are at stake; oversight has not been tested behaviourally; subgroup performance is unknown for affected groups; affected people cannot practically challenge outcomes; policy and frontline accounts contradict each other; or the supplier limits testing, logging, change notice or exit. Unknown is not the same as satisfactory.
The full questionnaire: nine tracks
No single person completes the whole questionnaire. Each role answers only what it knows or experiences. Some claims are asked of several groups on purpose, to expose gaps between policy and practice.
| Track | Target time | Sample question |
|---|---|---|
| Executive and accountable owner | 25–35 min | Who can restrict or stop this use without asking the team responsible for its success? |
| Product and technical | 35–50 min | Which supplier or model changes can reach production without your own testing? |
| Legal, privacy and compliance | 30–45 min | Who has the authority to provide effective remedy when an outcome is wrong? |
| Risk, audit and responsible AI | 30–45 min | Which positive claims currently rest only on assertions? |
| Managers | 20–30 min | What happens when an employee disagrees with the AI? Give the most recent example. |
| Frontline users | 20–30 min | Describe a case where the AI was wrong and how you noticed. |
| Affected people | 15–25 min | Could you reach a person with real authority to reconsider the outcome? |
| Worker representatives | 20–30 min | Who captures the productivity gains, and who absorbs the extra checking and rework? |
| Independent assessor | 30–45 min | Could anything in this assessment have changed the deployment decision — or is it documenting one already made? |
Times are design targets, not validated timings. Completion time is one of the things the pilots will measure.
Policy–practice contradiction tests
| Claim | Asked of management | Asked of frontline or affected people | Checked against |
|---|---|---|---|
| “Humans can override” | How are overrides handled? | What happens when you disagree? | Override logs, seeded errors |
| “Users understand the limits” | How are staff trained? | Where should you not rely on it? | Knowledge and error-detection test |
| “Appeals are available” | What is the procedure? | Could you reach someone who could change it? | Appeal funnel, abandonment, corrections |
| “AI reduces workload” | Where does saved time go? | Has your workload or checking changed? | Baseline versus production workload |
| “Stakeholders participated” | Who was consulted? | Did your input change anything? | Change log showing actual influence |
| “No retaliation for challenge” | How is disagreement treated? | Is it safe to challenge the AI? | Grievance patterns, performance consequences |
A contradiction opens an issue for the assessor. It is never settled by majority vote.
The workbook
- Assessment charterBefore any findingThe six elements of the AI use, stakes and reversibility, named owners, who can stop the system, conflicts of interest, and the expiry date and change triggers — set before authorisation.
- Entry gate10 testsPurpose, baseline, suitability, necessity, intrusion, stakes and reversibility, power and vulnerability, distribution, purpose integrity and the human-primary boundary.
- Eighteen domain modulesOne formatEach with a definition, core and diagnostic questions, strong versus weak evidence, candidate indicators, red flags, finding guidance, actions and reassessment triggers. The first pilots will work five in full: validity, fairness, retained capability, meaningful oversight, and incidents and remedy.
- Rights reviewAcross all layersThe Rights-to-Evidence lens, including the rights-in-practice path from “formally available” to “remedies the harm”.
- RegistersLiving recordsAn evidence register (each item tagged with its evidence class and perspective, validity period and what would invalidate it); an assumptions and uncertainty register; a red-flag register; and a question traceability matrix linking every question to its claim, domain, evidence and source.
- Decision recordNo scoreDomain findings, red flags, decisive evidence, dissenting views, the chosen outcome and why stronger or weaker outcomes were rejected — plus conditions, owner, stop authority and expiry date.
- Continuous assurance planAfter the decisionSignals, baselines, owners and escalation states: normal → watch → review → intervene → suspend or withdraw.
Every decision record asks: what evidence from this assessment materially changed, constrained or could have changed the original proposal? If the answer is repeatedly “none”, either the instrument is not sensitive enough — or decisions are being made before the assessment.
Measurement proportionality check. Before collecting any data from workers, users or affected people, the workbook asks: is this evidence necessary? Is there a less intrusive way to get it? Could collecting it shift power, autonomy or privacy? Who can see it? Could it be reused for performance management? When will it be deleted?
How the instruments will be validated
Research-grounded is not the same as validated. Before release, the instruments will be tested on:
- Content validity — do the questions cover what each domain claims to assess? Expert panels classify every question as essential, useful, redundant or unclear, adapting principles from the COSMIN content-validity methodology.1
- Response process — do executives, engineers, lawyers, frontline staff and affected people understand the questions as intended? Tested through cognitive interviews.23
- Known-failure sensitivity — do assessors catch planted problems, such as strong average accuracy with poor subgroup performance, or a formal override with no realistic time to use it?
- Assessor reliability — do two assessors reach broadly the same findings from the same evidence?
- Gaming resistance — can a respondent create a misleadingly positive picture with policy language alone?
- Burden and usefulness — completion time, evidence-retrieval time, and whether findings change decisions.
The planned sequence: instrument design → cognitive testing → rapid-screen pilot → three full pilots → assessor-reliability testing → revisions → public beta → further field validation → v1.0.
Careful with the words
This is an independent assessment methodology. It is not a certification scheme and does not constitute legal advice. Words such as “certified”, “compliant”, “approved” and “audited” should not be used about any assessment made with these instruments until what such claims mean has been deliberately established.
The aim is not to make responsible AI easy to score, but to make consequential claims about responsible AI difficult to make without evidence.
Sources
- Methodology Terwee, C. B. et al. (2018). COSMIN methodology for evaluating the content validity of patient-reported outcome measures: a Delphi study. Quality of Life Research, 27, 1159–1170. doi:10.1007/s11136-018-1829-0. ↩
- Book Willis, G. B. (2005). Cognitive Interviewing: A Tool for Improving Questionnaire Design. SAGE. doi:10.4135/9781412983655. ↩
- Research centre CDC Collaborating Center for Questionnaire Design and Evaluation Research (CCQDER). cdc.gov. ↩