In development · Pilot protocol v0.4 · 4 October 2026. A proposed protocol for testing the framework in real organisations. No pilot has run yet, and the timings are estimates.
Responsible AI & Human Agency · Framework · Pilots
Pilot protocol
How the framework will be tested in three contrasting settings — and what would count as evidence that it works.
What a pilot is for
A pilot tests the framework, not the organisation. Its purpose is to find out whether the method reveals things that existing assessments miss, whether it can be applied consistently, whether it is practical, and whether its findings change real decisions.
A framework that never changes a decision has not been validated — it has been rehearsed.
A pilot is not a certification, an audit opinion or a compliance check. Its findings must not be presented as approval of the AI system.
Three contrasting settings
| Setting | Why it tests the framework | Domains under most pressure | Signature tests |
|---|---|---|---|
| Enterprise AI assistant | Low stakes per task, high scale, strong adoption incentives | Retained capability · human impact & distribution · incentives · traceability | Unaided task performance before and after; where saved time goes; incentive audit |
| AI-assisted recruitment | High stakes for people who never chose the system | Fairness · participation · meaningful oversight · contestability | Matched-candidate tests; seeded biased recommendations; an applicant-side appeal simulation |
| Customer-facing conversational AI | Relational design, vulnerable users, disputed outcomes | Relational integrity · appropriate reliance · autonomy & choice · uncertainty | Time to reach a human; abstention on out-of-scope questions; what users believe the system is |
The enterprise AI assistant is the natural first pilot: the lowest risk to affected people, and usually the easiest to arrange.
Ground rules
- One AI use per pilot, defined by its six elements: purpose, system, workflow, people, organisation and context.
- Rights-respecting measurement. Cohort-level and sampled data only; no individual performance monitoring; minimal personal data; agreed retention limits; a data-protection assessment before collection starts.
- Workers and their representatives are informed before any workplace evidence is collected, and can raise concerns.
- Voluntary participation for every interviewee and survey respondent, who can withdraw at any time.
- Conflicts of interest are declared by everyone involved, including any commercial relationship with the AI supplier.
- Confidentiality. Findings about the organisation are published only anonymised and with its consent. Lessons about the framework are published.
- Serious harm is escalated at once. A red flag involving serious harm is reported to the use owner immediately, not held for the final report.
Ten phases
Indicative duration: about ten weeks of elapsed time, followed by one reassessment after six months. Effort depends heavily on how much evidence already exists.
- Agreement and scopingWeek 1Select one AI use and define its six elements. Agree roles, data protection, confidentiality and stop rules. Record conflicts of interest.
- Entry gateWeeks 1–2Proportionality: compare against the best realistic non-AI alternative. Purpose integrity: agree success criteria with affected people and list what must remain human. If the gate fails, the pilot continues as a study of that finding.
- Evidence requestWeek 2Request the minimum evidence pack below. Record what exists, what is missing and what is refused — each is a finding.
- BaselineWeeks 2–4Measure the current or pre-AI process wherever possible: accuracy, time, workload, unaided capability, appeal outcomes.
- Controlled testsWeeks 3–5Seeded-error and appropriate-reliance tests, a reconstruction test, an abstention test, a stop drill and an appeal simulation.
- Production evidenceWeeks 4–7Logs, overrides and their effectiveness, appeals, incidents and near misses, subgroup outcomes and drift.
- People's evidenceWeeks 5–7Cohort surveys and interviews with users, affected people and worker representatives — on workload, choice, voice and whether the human route works.
- Independent challengeWeeks 7–8A second assessor reviews the evidence and makes findings independently. Where the two disagree, the reasons are recorded.
- Findings and decisionWeeks 8–9One finding per domain with its evidence classes, no averaging. The decision rules produce one of five outcomes, with conditions, an owner and an expiry date.
- Debrief and framework learningWeeks 9–10The organisation decides what to do. The pilot report records what the framework found and what should change in it.
Reassessment after about six months tests whether the evidence still holds — and whether any decision has expired.
Minimum evidence pack
| Evidence | Supports |
|---|---|
| The business case and any alternatives considered | Proportionality |
| A description of the workflow: where AI sits and where decisions are made | Unit of assessment; oversight |
| System and version register, including supplier and model changes | Traceability; robustness |
| Supplier documentation and any local validation results | Validity; fairness; uncertainty |
| Security, red-team and privacy test results | Security & privacy |
| DPIA, FRIA where required, and the legal register | Legal compliance; rights lens |
| Named owner, decision rights and the stop procedure | Accountability |
| Targets, KPIs and incentives linked to the use | Incentives & culture |
| Training given to users and reviewers | Human layer |
| Override, appeal, complaint and incident records | Oversight; remedy; incidents |
| Any consultation with users, workers or affected people | Participation |
| Previous assessments of the use, of any kind | Comparison: what the framework adds |
The last item matters most for validation. To show what the framework adds, each pilot compares its findings with the assessments the organisation had already done.
What each pilot report must answer
- What did the framework find? Findings per domain, with evidence classes.
- What would the existing assessment have missed? A side-by-side comparison.
- What decision changed? Or, if none, why not.
- Did two assessors agree? Where they differed, and why.
- What proved impractical? Missing data, access, time, cost or measurement burden.
- What did it cost? Person-hours for the organisation and for the assessors.
- Did the measurement itself cause any harm or concern?
- What should change in the framework? Every change goes into the version history.
When has the framework passed?
Proposed criteria for moving from development to a public beta (v0.9) after the three pilots:
- It adds something. In each setting, it surfaces at least one material finding that existing assessments did not.
- It changes decisions. At least one finding leads to a real change in a decision, design, condition or target.
- It can be applied consistently. Two assessors reach largely the same findings from the same evidence.
- It is practical. The effort is proportionate to the stakes, and organisations can act on the findings.
- It does no harm. The measurement does not create disproportionate surveillance or burden.
If the pilots show that a domain adds nothing, cannot be measured or duplicates an existing instrument, that domain changes or goes. That outcome would also be published.
Taking part
I am looking for organisations willing to pilot the framework on one AI use — starting with an enterprise AI assistant. If that could be you, get in touch.