Responsible AI & Human Agency · Research · 2026Verantwoorde AI & menselijke regie · Onderzoek · 2026
How Do We Know AI Remains Responsible Over Time?
From one-time assessment to continuous human assurance.
This research page is currently available in English only. Research position as of 3 October 2026 — governance analysis, not legal advice.Deze onderzoekspagina is op dit moment alleen in het Engels beschikbaar. Onderzoekspositie per 3 oktober 2026 — governance-analyse, geen juridisch advies.
Eighteen months later
An AI system passes its legal review, technical testing and human-rights assessment. It is approved for deployment.
Eighteen months later, the model has changed. The workforce has adapted to it. Human reviewers rarely override its recommendations. Several near misses have occurred, but no serious incident has been recorded.
Is the original approval still evidence that the system is responsible?
A hypothetical illustration, not a description of a particular system.
The short answer
No. An AI system can change without its model changing. Users change, populations change, workflows change, data sources and retrieval indexes are updated, prompts and guardrails are edited, vendors upgrade models, attackers find new techniques, people learn to over-rely on the system, skills erode — and a use that was proportionate can become inappropriate as the organisation around it changes.
Responsible AI is not a state an organisation achieves. It is a claim that must remain supported by evidence.
That is why monitoring is not enough. Monitoring tells us what is happening. Assurance determines whether what is happening remains acceptable — and requires the authority to intervene when it does not.
Five words that are often used interchangeably
| Concept | Meaning | Why it is not enough on its own |
|---|---|---|
| Monitoring | Observing selected variables over time | Data can be collected without any decision changing |
| Evaluation | Testing a defined claim against evidence | Often periodic and narrow |
| Audit | Structured, more independent examination against criteria | Usually point-in-time and limited by access |
| Assurance | Building justified confidence that claims remain supported | Needs several kinds of evidence and explicit limits |
| Continuous assurance | Repeatedly refreshing that confidence and linking evidence to intervention, remedy and renewal | Needs organisational authority, not only instrumentation |
Continuous does not mean real-time. Some signals — a security compromise, a prohibited use — must be watched constantly. Others, such as skill retention or how productivity gains are shared, need longitudinal evidence. A useful idea is the evidence half-life: how long a finding stays credible before it must be refreshed. A static rules engine may allow long review intervals; a frequently updated foundation-model service in a high-stakes workflow needs re-evaluation after every material change.
The continuous assurance loop
- Purpose & approved use
- Legal & rights baseline
- Technical & security baseline
- Human-impact baseline
- Deploy within boundaries
- Continuous evidence collection
Complaints, near misses and incidents feed the independent challenge at any time.
Six assurance domains
The object being assured is not “model version X”. It is the model plus its data, prompts and configuration, retrieval sources, tools, infrastructure, human workflow, incentives, affected population, legal environment and suppliers.
| Domain | Question | Example evidence |
|---|---|---|
| Legality and rights | Is the actual use still permissible and proportionate? | Legal register, impact-assessment updates, complaints, regulatory change |
| Technical performance | Does it still work under current conditions? | Outcome metrics against baseline, calibration, drift, subgroup results |
| Security and privacy | Has the attack or privacy surface changed? | Vulnerability and adversarial tests, access logs, privacy incidents |
| Human capability and impact | What is happening to people because of the system? | Unaided skill tests, error detection, workload, accessibility, grievances |
| Governance and supply chain | Can the organisation still control what it deployed? | Vendor-change notices, ownership, audit rights, decision logs |
| Incidents, remedy and resilience | Can failures be detected, contained, corrected and learned from? | Incident and near-miss register, rollback tests, appeal and remedy records |
Accurate is not the same as responsible
A system can stay accurate while becoming irresponsible: technically stable output while workers lose the ability to challenge it; unchanged accuracy while targets rise until verification becomes impossible. In one multicentre observational study, the unaided adenoma detection rate in colonoscopy fell from 28.4% to 22.4% after endoscopists began routinely using AI assistance.1 The design cannot prove causation, but it shows why assurance must sometimes measure the human without the AI.
Do not create a single “Responsible AI score” in which serious failure in one domain is offset by success in another. Better productivity cannot cancel a rights violation; excellent accessibility cannot compensate for a critical vulnerability. Use separate dimensions, hard-stop conditions and evidence confidence instead.
Continuous human assurance
Most technical monitoring asks whether the AI is changing. Continuous human assurance asks whether people are changing because of the AI — and whether that change is acceptable. This is where the earlier studies in this series come together: capability and dependency, meaningful oversight, the distribution of human cost and the people affected who never chose the system.
- Retained capability
- Unaided performance before and after sustained use.
- Error detection
- Can reviewers catch deliberately seeded wrong outputs?
- Appropriate reliance
- Accepting correct AI and rejecting incorrect AI — high agreement is not automatically good.
- Fallback
- Performance loss during an outage or withdrawal.
- Workload and autonomy
- Verification burden, effective override and escalation.
- Contestability and remedy
- Appeals completed, time to outcome, corrections actually made — zero appeals may mean success, or an inaccessible process.
A human signal can matter even when no incident is recorded. If, after twelve months, most employees rarely verify the AI's output independently, that is an assurance finding — even though conventional incident monitoring might show nothing at all.
Evidence status, confidence and freshness
A green metric built on old or weak evidence is a common assurance failure. Every claim should carry three markers: its status, the strength of its evidence and when it was last refreshed.
- E0 · AssertionA policy, vendor or project team says it is true.
- E1 · Controlled testEvidence from pre-production or simulation.
- E2 · Production evidenceSupported by live outcome data.
- E3 · Independent challengeReproduced or stress-tested by an independent function.
- E4 · Human corroborationConfirmed by affected-user, grievance or longitudinal human evidence.
A dashboard can then say “Accuracy: within boundary · E3 · refreshed 12 days ago” instead of simply “Accuracy: green”.
| Assurance claim | Status | Evidence | Freshness |
|---|---|---|---|
| Accuracy in intended context | Within boundary | Production + independent sample | 18 days |
| Security | Review required | New supplier release | 2 days |
| Human capability | Evidence insufficient | Next proficiency exercise scheduled | 94 days |
| Accessibility | Within boundary | User and technical testing | 31 days |
| Incidents | One open investigation | Root cause pending | Live |
An illustrative ledger with invented values, showing the format — not data from a real system. The E0–E4 levels are my proposed taxonomy.
From signal to intervention
- NormalRoutine monitoring
- WatchA signal appears — enhanced observation
- ReviewThreshold exceeded or evidence stale — formal reassessment
- InterveneUnacceptable risk — restrict, redesign, remediate
- Suspend or withdrawUnresolved serious risk
Thresholds should come from a legal boundary, a safety tolerance, a statistical control limit against the validated baseline, an operational capacity limit (a review queue beyond which meaningful review is impossible), or a pre-agreed human-impact boundary. A generic “5% drift means red” rule should be avoided: five percent can be irrelevant for one metric and catastrophic for another.
Incidents, near misses and weak signals
The OECD's common reporting framework for AI incidents sets out 29 criteria to help compare incidents across contexts and assess emerging risks.2 Inside an organisation, the categories should be wider than harm that has already happened:
| Category | Example | Default response |
|---|---|---|
| Weak signal | Unexpected but benign behaviour; users describe confusion or workarounds | Log and trend |
| Control weakness | Missing log, stale evaluation, inaccessible escalation route | Fix the control and verify |
| Near miss | A wrong AI recommendation caught before it caused harm | Investigate as if the control might fail next time |
| Incident | Actual harm or material rights or security impact | Contain, investigate, notify where required, remediate |
| Critical incident | Severe or irreversible harm | Immediate suspension and executive or regulatory escalation |
Organisations need several legal clocks, not one generic deadline. Under the GDPR, personal-data breaches normally have to be notified within 72 hours where feasible.3 Under the EU AI Act, providers of high-risk systems must report serious incidents no later than 15 days after becoming aware, with shorter deadlines for the most serious cases.4 Sector rules may add more.
A system cannot certify itself
The people who benefit from a deployment should not be the only judges of whether it remains acceptable.
If the team responsible for an AI productivity target also decides whether the AI should be suspended, the conflict is obvious — the core finding of my research on accountability. Independent challenge needs evidence access, methodological freedom, freedom from disabling conflicts, escalation rights, protection from retaliation and the power to require reassessment. NIST's ARIA programme illustrates the direction: it combines model testing, red-teaming and field testing rather than relying on a model owner's own description.5
For foundation-model deployments, the supply chain matters most. Contracts should secure, in proportion to risk: notice of material changes, version identification, incident cooperation, evidence preservation, audit rights, the ability to freeze or reject an upgrade where feasible, remediation duties and credible exit support.
Withdrawal readiness: could you stop it tomorrow?
Responsible AI work spends great effort asking what must be true before deployment. The neglected counterpart is: what evidence would make us stop — and could we?
An exploratory exercise — not a validated audit. Your answers stay in this browser tab and are not stored or sent anywhere.
Answer the questions to see a reflection.
What real cases show
- Enterprise AI assistant — UK Department for Work and Pensions
- A government trial of an enterprise AI assistant involving 3,549 staff estimated an average saving of about 19 minutes a day. The evaluation itself notes possible selection bias, because participants were not randomly assigned.6 A successful pilot should become the baseline for longer-term assurance — of quality, rework, capability and dependency — not the final verdict.
- Sepsis prediction — external validation
- An external validation across 38,455 hospitalisations found an area under the curve of 0.63; the model missed 67% of patients with sepsis while alerting on 18% of all hospitalisations.7 Vendor performance claims cannot simply be inherited; local validation and continuing outcome monitoring are needed.
- Colonoscopy and retained skill
- Unaided adenoma detection fell from 28.4% to 22.4% after routine AI exposure.1 Human capability can drift while the AI remains technically effective.
- Hiring — iTutorGroup
- Software allegedly rejected older applicants automatically; the employer settled with the EEOC for $365,000.8 Monitor configured rules and decision outcomes, not only model performance.
- Customer chatbot — Moffatt v. Air Canada
- A tribunal held the airline responsible for incorrect fare information from its website chatbot.9 A contradiction reported by one customer should trigger a systemic correction.
Five different failure modes — and none of them would be caught by monitoring model accuracy alone.
What law and standards already require
- EU AI Act. For high-risk systems, risk management is “a continuous iterative process” throughout the lifecycle (Article 9); systems must allow automatic logging (Article 12); providers must run post-market monitoring (Article 72) and report serious incidents (Article 73).4 Following the AI Omnibus, most high-risk obligations apply from 2 December 2027 (Annex III) and 2 August 2028 (Annex I).10
- GDPR. Security measures must ensure ongoing resilience and be regularly tested (Article 32), and a data-protection impact assessment must be reviewed when the risk changes (Article 35).3
- NIST AI RMF includes post-deployment monitoring, appeal and override, incident response, recovery, change management and decommissioning.11 ISO/IEC 42001 institutionalises continual improvement in an AI management system.12
- WHO stresses lifecycle validation and post-market surveillance in health AI, while describing its 2023 overview as a resource rather than guidance or a regulatory framework.13
None of these on its own is a complete continuous-assurance method. Legal classification should set minimum obligations — not decide which human effects deserve measurement.
What to put in place
- Time-limit every material deployment approval, with conditions and a next review date.
- Keep a system and version register covering model, configuration, data, tools and retrieval sources.
- Give an independent function evidence access and the power to require reassessment.
- Measure human capability — baseline and longitudinal unaided proficiency, seeded-error tests.
- Track who gains and who carries the burden as use scales.
- Run a near-miss system that is safe to report into.
- Secure supplier update, evidence and audit rights.
- Practise rollback and withdrawal before you need them.
- Balance adoption incentives with outcome measures: if people are rewarded only for speed and adoption, nobody is rewarded for discovering that a system should slow down or stop.
- Report a “responsible intervention count”: how often assurance evidence actually changed a condition, workflow, target, supplier decision or design. A programme that never changes a decision may be very successful — or ceremonial.
Conclusion
Responsible deployment asks whether we should deploy. Continuous assurance asks whether we should still be deploying.
Responsible AI is not demonstrated by proving that a system was responsible when approved. It is demonstrated by maintaining a living chain of evidence that the technology, the people and the institution remain within justified boundaries — and by acting when that evidence changes.
The governance test is therefore: what evidence would cause us to change, restrict or stop this system — and have we built an organisation capable of acting when that evidence appears?
Method and limits
This is a research-led synthesis of legislation, standards, regulatory guidance, evaluations and empirical studies, prepared in October 2026. It is not legal advice; legal dates are described as of 3 October 2026 and should be rechecked. Longitudinal evidence on human capability under sustained AI use remains thin. The assurance loop, domains, evidence levels, intervention ladder, withdrawal-readiness questions and recommendations are my own proposals. I checked each source below against its official or published record.
Sources
- Observational Budzyń, K. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. Lancet Gastroenterology & Hepatology, 10(10), 896–903. doi:10.1016/S2468-1253(25)00133-5. ↩
- International framework OECD (2025). Towards a common reporting framework for AI incidents. OECD Artificial Intelligence Papers. doi:10.1787/f326d4ac-en. ↩
- Legislation Regulation (EU) 2016/679 (GDPR), Articles 32, 33 and 35. eur-lex.europa.eu. ↩
- Legislation AI Act, Articles 9, 12, 72 and 73. Art. 9 · Art. 12 · Art. 72 · Art. 73. ↩
- Evaluation programme NIST. ARIA — Assessing Risks and Impacts of AI. ai-challenges.nist.gov. ↩
- Government evaluation UK Department for Work and Pensions (2026). An evaluation of DWP's Microsoft Copilot 365 trial. gov.uk. ↩
- External validation Wong, A. et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine. doi:10.1001/jamainternmed.2021.2626. ↩
- Regulatory action U.S. EEOC. iTutorGroup to pay $365,000 to settle EEOC discriminatory hiring suit. eeoc.gov. ↩
- Case law Moffatt v. Air Canada, 2024 BCCRT 149. canlii.org. ↩
- Official guidance European Commission. AI Act — regulatory framework for AI. digital-strategy.ec.europa.eu. ↩
- Framework NIST (2023). AI Risk Management Framework (AI RMF 1.0). doi:10.6028/NIST.AI.100-1. ↩
- Standard ISO/IEC 42001:2023, Artificial intelligence — Management system. iso.org. ↩
- International guidance World Health Organization (2023). Regulatory considerations on artificial intelligence for health. who.int. ↩