← Responsible AI & Human Agency← Verantwoorde AI & menselijke regie

Responsible AI & Human Agency · Entry gate · Research · 2026Verantwoorde AI & menselijke regie · Toegangspoort · Onderzoek · 2026

When Should AI Not Be Used?

A proportionality test for deciding when to automate, augment — or preserve human judgment.

This research page is currently available in English only. Research position as of 3 October 2026 — governance analysis, not legal advice.Deze onderzoekspagina is op dit moment alleen in het Engels beschikbaar. Onderzoekspositie per 3 oktober 2026 — governance-analyse, geen juridisch advies.

The question before the questions

The fact that AI can perform a task does not establish that AI should perform it.

Most responsible AI work starts after the decision to use AI has been made. It asks whether the system is accurate, fair, secure, explainable and overseen. Those questions matter. But they skip one that comes first.

Responsible AI begins earlier: by asking whether using AI for this purpose, in this relationship and in this institutional context is necessary, proportionate and compatible with the human or public good the activity exists to serve.

That is why this study is the entry gate to the series. Its argument is simple, and it is not anti-AI: not using AI, keeping humans primary, running a restricted pilot and deploying with safeguards are all legitimate outcomes. A governance process that can only ever end in “yes, with controls” is not assessing anything. It is approving with paperwork.

The short answer

AI should not be used — or should be confined to a supporting role — when:

  1. the purpose is unlawful, manipulative, illegitimate or not important enough;
  2. comparable value can be reached by a less intrusive, less resource-hungry or more accountable method;
  3. errors are severe, hard to reverse or imposed on people with little power;
  4. meaningful human oversight cannot actually be provided;
  5. the system undermines dignity, agency, equality, privacy, professional duties or due process;
  6. automation displaces something that is the activity — care, recognition, learning, authorship, deliberation, responsibility or democratic legitimacy;
  7. the evidence of benefit is weaker than the direct, indirect and distributional costs.

This is close to what international frameworks already say. UNESCO's Recommendation puts proportionality and “do no harm” first: AI use should not go beyond what is necessary to achieve a legitimate aim.3 The NIST AI Risk Management Framework asks organisations to establish intended purpose, expected benefits and costs, and impacts on individuals, groups, organisations and society — before deployment.4 What is usually missing is a method for reaching the answer no.

The AI Use Proportionality Test

Flow diagram titled AI Use Proportionality Decision Model. Ten questions in sequence: legitimate purpose; legal or rights-based red line; evidence that AI is suitable compared with a human or non-AI baseline; necessary and least intrusive means; acceptable stakes, vulnerability, reversibility, privacy, discrimination and security; purpose integrity preserved; meaningful human oversight feasible; contestability, remedy, accessibility, fallback and exit available; distribution of benefits and burdens acceptable; continuous monitoring and stop conditions established. Failing the first two leads to Do Not Use. Failing evidence leads to Human-Primary or AI-Assist Only, or a Restricted Pilot. Failing later questions leads to a Restricted Pilot or Use With Strong Safeguards. Passing all leads to Proportionate Use. A monitor, learn and improve loop sends every outcome back for suspension or redesign if risks increase. Notes: efficiency alone is insufficient; uncertainty raises the burden of proof; the human-only baseline must also be tested. View the full decision model →
  1. PurposeWhat legitimate human or organisational purpose are we trying to achieve?
  2. ValueIs there credible evidence that AI materially improves that outcome against the best realistic alternative?
  3. NecessityCould a less intrusive alternative deliver comparable value?
  4. Human impactWhat happens to rights, autonomy, capability, dignity, equality and relationships?
  5. ControlAre errors detectable, contestable and reversible, under meaningful human authority?
  6. DistributionWho receives the benefits — and who carries the risks, work and consequences?
  7. Purpose integrityDoes automation preserve the human or public good the activity exists to provide?
Figure 1 — The AI Use Proportionality Test. Seven questions in sequence, with legal red lines checked first. The test does not produce a risk score; it determines what form of AI involvement, if any, is justified. My proposed model, synthesised from existing frameworks — not a statutory process or validated instrument.

Five legitimate outcomes

Do not use AIUnlawful or illegitimate purpose, severe uncontrollable harm, impossible oversight, no accessible remedy, or irreparable damage to the purpose of the activity.
Human-primary · AI-assist onlyAI adds value, but judgment, reasons, relationship or accountability must stay human.
Restricted pilotBenefits are plausible but unproven. Scope, population, actions and duration are bounded, with stop rules agreed in advance.
Use with strong safeguardsHigh-stakes use with demonstrated value, effective oversight, impact assessment, contestability, monitoring and rollback.
Proportionate useLow or controlled risk, demonstrated net value, reversible effects and a fair distribution of benefits and burdens.

Three rules hold the test together. Efficiency alone is not enough: cost or speed gains do not justify AI use without broader value and acceptable risk. Uncertainty raises the burden of proof: where stakes, vulnerability or irreversibility are high, missing evidence is a reason to restrict or not deploy — not permission to “learn in production”. And the human-only baseline must also be tested: compare AI against the best feasible human-plus-process alternative, not against perfect humans and not against today's avoidably poor system.

Purpose integrity: what disappears when AI succeeds?

Some activities produce two things at once: an outcome, and a human or institutional good that only exists because of how the outcome was produced.

ActivityProduces an outcome……and a good created by the process
CareTreatmentBeing attended to
EducationAnswersLearning, formation and authorship
LawDecisionsReason-giving, procedural justice and legitimate authority
ManagementAllocationRecognition, coaching and responsibility
Public consultationInformationDemocratic participation
Creative workArtefactsExpression, voice and attribution

Automation is purpose-undermining when it improves a measurable proxy — speed, throughput, consistency — while hollowing out the underlying good. So the governing question is not whether AI can simulate empathy, deliberation or authorship. It is whether simulation meets the duty the institution owes to the person.

Imagine an AI system that produces excellent personalised essays for students. On a conventional assessment, quality goes up, speed goes up and cost goes down: a success. But if part of the purpose was for the student to learn to reason and write, learning, independent capability and authorship all go down. The technology can optimise the output while undermining the purpose. (A hypothetical illustration.) This connects directly to my work on dependency and the value of friction.

The Purpose Integrity Test

  1. What good does the activity exist to provide?
  2. Which elements are constitutive of it, rather than merely instrumental?
  3. Would a reasonable affected person experience AI substitution as less recognition, care, voice or legitimacy?
  4. Which human capabilities will stop developing if the task is delegated?
  5. Who remains visibly responsible and able to give reasons?
  6. Can AI remove routine burden while preserving the human core?
  7. If the human core is restored through safeguards, does the efficiency case largely disappear?
  8. Would the organisation still automate if the affected person — not the budget owner — chose the success criteria?

If human presence, judgment or authorship is constitutive and cannot be preserved without eliminating the claimed benefit, the outcome should normally be human-primary or do not use.

Human-primary is a real governance category

The debate is usually framed as human versus AI. A more useful spectrum has five positions:

Human-onlyHuman-primaryAI-supportedAI-primaryAutomated

The question is not whether a human is “in the loop”, but where authority sits. In a human-primary arrangement, AI may do substantial supporting work while the moral or cognitive core stays human:

AI mayretrieve · compare · summarise · translate · flag · structure
A human mustinterpret · deliberate · contextualise · decide · justify · take responsibility

This is much stronger than putting a ceremonial reviewer at the end of an automated process. It specifies which functions must remain human, and makes using AI for those functions out of scope.

Six boundaries for AI use

BoundaryDecision postureExamples
ProhibitedDo not build, procure or deployPractices prohibited by law, such as harmful manipulation, exploitation of vulnerability, specified social scoring and emotion inference at work or in education
Presumptively inappropriateNon-use is the default; exceptions need compelling evidence and independent approvalAI presented as a human in intimate or trust interactions; automated adjudication without effective appeal; covert worker surveillance
High riskOnly with strong evidence, impact assessment, oversight, audit, remedy and the ability to suspendHiring, credit, benefits, clinical decision support, access to education, essential infrastructure
Human-primaryAI searches, summarises, translates or flags; a qualified human forms the judgmentDiagnosis, legal advice, employee evaluation, educational assessment, social-work decisions
Low-risk augmentationLightweight governance, privacy discipline and validationDrafting, transcription, routine translation, accessibility support, reversible admin
Insufficient valueDo not use, even if technically feasibleTrivial convenience that adds surveillance, cost, environmental burden, dependency or worse service

My synthesis of the EU AI Act's risk structure, human-rights law, UNESCO and NIST — not a legal classification.

What the law already says — as of 3 October 2026

  • EU AI Act — prohibitions. Article 5 bans specified practices, including harmful manipulation and deception, exploitation of vulnerabilities, specified social scoring, criminal-risk prediction based solely on profiling, untargeted facial-image scraping and emotion inference in workplaces and education.2 The first eight prohibitions have applied since February 2025; a ninth, on AI that generates non-consensual intimate content or child sexual abuse material, was added through the AI Omnibus and applies from December 2026.1
  • EU AI Act — high risk. Rules for high-risk uses in areas such as employment, education and biometrics apply from 2 December 2027; rules for AI embedded in regulated products from 2 August 2028.1 Until then, GDPR, equality, labour, consumer and sector law continue to apply in full. And legal high-risk classification is not a finding of ethical proportionality.
  • GDPR Article 22. People have the right not to be subject to decisions based solely on automated processing with legal or similarly significant effects, except under limited conditions and safeguards.16 In SCHUFA, the EU Court of Justice held that a score can itself be such a decision when a third party draws strongly on it.6 The Dutch Data Protection Authority's 2025 guidance on meaningful human intervention organises the question around people, technology and design, process and governance — a person who usually clicks “approve” is not enough.7
  • Council of Europe Framework Convention. The first legally binding international AI treaty requires each party to assess the need for “a moratorium or ban or other appropriate measures” for uses it considers incompatible with human rights, democracy or the rule of law (Article 16(4)).8 Non-use is written into the treaty as a governance outcome.

Legal proportionality asks whether an interference is lawful, pursues a legitimate aim, is necessary and strikes a fair balance. In the Dutch SyRI judgment (2020), the court accepted that fighting fraud is an important aim — and still found the legislation insufficiently transparent and verifiable, failing the fair balance Article 8 ECHR requires.5 A legitimate purpose does not justify any means.

What this means in practice

DomainAppropriate AI roleRestrict or do not use when…
Health and careDetection support, documentation, scheduling, translation, evidence retrievalNo accountable clinician; deskilling; opaque triage; relational care replaced
EducationAccessibility, formative feedback, practice, teacher supportAutomated consequential grading; substitution for learning or authorship; behavioural monitoring
LawSearch, disclosure review, issue spottingFinal advice without qualified verification; adjudication without reasons and appeal
Public servicesAdministrative support, anomaly detectionAutomated suspicion, debt or eligibility with weak evidence or inaccessible channels
EmploymentDrafting, job matching, administrationInferred emotion or personality; covert productivity scoring; AI-determined hiring, pay or dismissal
Customer serviceRoutine self-service with easy escalationAI-only channels for disputes or vulnerable customers; AI presented as human
Creative workIdeation, editing, accessibilityHidden substitution where authorship, consent or attribution is the point
Autonomous agentsBounded, reversible actions with least privilegeIrreversible actions, uncontrolled tool access, unclear authority, untraceable action chains

What real cases show

SyRI (Netherlands)
An important aim — fraud prevention — did not justify a system that was insufficiently transparent and verifiable.5 Purpose is necessary, not sufficient.
Robodebt (Australia)
A Royal Commission into the unlawful automated debt scheme made 57 recommendations, many aimed at public administration rather than technology.9 Automation failure is institutional, not only technical.
2020 exam grades (England)
After exams were cancelled, grades were first standardised by a statistical model; following public outcry, the regulator switched to grades submitted by teachers.10 Population-level consistency can conflict with individual justice.
Automated test vehicle (Tempe, 2018)
The NTSB found operator distraction to be the probable cause, with inadequate safety-risk assessment, ineffective operator oversight and insufficient attention to automation complacency as contributing factors.11 A human expected to take over at the last second is a weak safeguard.
Colonoscopy and retained skill
An observational study found unaided adenoma detection fell from 28.4% to 22.4% after routine AI exposure.14 Human capability is part of the cost-benefit equation.

Why non-use can also be irresponsible

Humans are biased, inconsistent, tired, slow, expensive and scarce. AI can improve translation, accessibility, pattern detection and capacity where expertise is missing. A systematic review of 19 economic evaluations of clinical AI found that AI often improved diagnostic accuracy and reduced costs — while noting that infrastructure, indirect costs and equity were often under-reported, so benefits may be overstated.13

But humans are also poor passive monitors. Automation complacency occurs in both novices and experts and is not overcome by simple practice.15 And when systems fail, blame tends to land on the nearest human operator — what Madeleine Clare Elish calls a “moral crumple zone”.12

Human primacy is justified not because humans are infallible, but because some decisions require context, moral responsibility, reason-giving, exceptions, care or democratic authority that prediction alone cannot discharge.

Escalation triggers

Redesign, independently review, restrict, suspend or reject when there is:

  • no credible baseline or representative evidence;
  • severe or irreversible effects, or vulnerable or captive populations;
  • a decisive, opaque score;
  • no accessible route without AI, or weak override, appeal or compensation;
  • humans systematically agreeing with AI, without independent testing;
  • a sudden model, data, vendor or agent-capability change;
  • burdens concentrated on less powerful groups;
  • material loss of purpose integrity;
  • no identifiable accountable owner.

A benefit at one level cannot erase harm at another. A faster benefits process may improve throughput for the organisation while citizens carry proof burdens, navigate inaccessible appeals and experience presumed suspicion.

Not a one-way ratchet

The proportionality decision is not taken once. Suppose a system was approved for use with strong safeguards. Twelve months later the model has changed, people rely on it much more, appeal rates have risen, unaided skills have declined and benefits are lower than predicted. Continuous assurance should send the system back through this test.

Use → Restrict → Redesign → Human-primary → Do not use

Every movement on that line must be possible. Responsible AI governance cannot be a one-way ratchet toward more automation.

Try the test on one AI use

Pick one specific use of AI — not “AI” in general — and answer honestly.

An exploratory exercise — not a validated assessment or legal advice. Your answers stay in this browser tab and are not stored or sent anywhere.

  1. Is the purpose specific, legitimate and important enough — not just “adopting AI” or cutting headcount?
  2. Are you confident the use crosses no legal or fundamental-rights red line?
  3. Is there credible evidence that AI outperforms the best realistic human or non-AI alternative here?
  4. Have you ruled out a less intrusive way to get comparable value?
  5. Are the effects on rights, autonomy, capability, dignity and equality acceptable — including for vulnerable people?
  6. Are errors detectable, contestable and reversible, with people who have real authority to override?
  7. Are benefits and burdens fairly distributed between the organisation, workers and affected people?
  8. Does the activity keep its human core — care, learning, authorship, reason-giving or responsibility — if AI is used?

Answer the questions to see which outcome your answers point toward.

What this asks of different roles

  • Boards and executives: approve purpose, risk appetite and prohibited-use boundaries; reward responsible non-deployment.
  • Business and finance: reject business cases built only on labour savings; price in rework, appeals, supervision, accessibility, exit and environmental cost.
  • Product and engineering: design for uncertainty, override, least privilege, logging, reversibility and safe failure.
  • Risk and compliance: integrate AI, privacy, human-rights, accessibility and equality assessments; keep independent stop authority.
  • HR and people leaders: keep employment judgment human-primary; never turn individual telemetry into an AI-adoption metric.
  • Procurement: require representative evidence, subgroup performance, change control, audit access, incident notification, rollback and exit rights.
  • Frontline professionals: keep the authority to disagree, to get a second opinion and to work without the system.
  • Worker representatives and affected communities: take part before procurement, not only after deployment.

Conclusion

The deepest mistake is to treat AI governance as a choice between automation and human fallibility. Both humans and machines have limits. The task is to decide which limits are acceptable, detectable, correctable and legitimate in a particular relationship.

Use AI only where its legitimate added value is evidenced, its intrusion is necessary, its burdens are fairly distributed, its failures are contestable and reversible, and it strengthens rather than empties the human or public purpose of the activity.

Where those conditions fail, not using AI is not technological conservatism. It is a governance outcome. Where AI clearly reduces error, scarcity or exclusion under effective human authority, refusing it can be just as disproportionate. A mature organisation must be able to reach — and reward — either conclusion.

Method and limits

This is a research-led synthesis of legislation, case law, regulatory guidance, inquiries and empirical studies, prepared in October 2026. It is not legal advice; legal dates are described as of 3 October 2026 and should be rechecked. The proportionality test, outcomes, boundaries, Purpose Integrity Test, escalation triggers and self-check are my own proposals, not validated instruments. Two claims from my working research could not be verified against a public source and were left out. I checked each source below against its official or published record.

Sources

  1. Official guidance European Commission. AI Act — regulatory framework for AI (application timeline and prohibited practices). digital-strategy.ec.europa.eu. ↩
  2. Legislation Regulation (EU) 2024/1689 (AI Act), Article 5 — prohibited AI practices. artificialintelligenceact.eu. ↩
  3. International recommendation UNESCO (2021). Recommendation on the Ethics of Artificial Intelligence. unesco.org. ↩
  4. Framework NIST (2023). AI Risk Management Framework (AI RMF 1.0). doi:10.6028/NIST.AI.100-1. ↩
  5. Case law Rechtbank Den Haag, 5 February 2020, ECLI:NL:RBDHA:2020:865 (SyRI). rechtspraak.nl. ↩
  6. Case law CJEU, Case C-634/21, SCHUFA Holding (Scoring), 7 December 2023. eur-lex.europa.eu. ↩
  7. Regulatory guidance Autoriteit Persoonsgegevens (2025). Handvatten betekenisvolle menselijke tussenkomst. autoriteitpersoonsgegevens.nl. ↩
  8. Treaty Council of Europe Framework Convention on Artificial Intelligence and Human Rights, Democracy and the Rule of Law (CETS No. 225), Article 16. rm.coe.int (PDF). ↩
  9. Public inquiry Royal Commission into the Robodebt Scheme (2023). Report. robodebt.royalcommission.gov.au. ↩
  10. Regulator statement Ofqual (2020). Statement from Roger Taylor, Chair, Ofqual. gov.uk. ↩
  11. Accident investigation National Transportation Safety Board. Investigation HWY18MH010 (Tempe, Arizona, 2018). ntsb.gov. ↩
  12. Peer-reviewed Elish, M. C. (2019). Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction. Engaging Science, Technology, and Society, 5, 40–60. doi:10.17351/ests2019.260. ↩
  13. Systematic review El Arab, R. A. & Al Moosa, O. A. (2025). Systematic review of cost effectiveness and budget impact of artificial intelligence in healthcare. npj Digital Medicine, 8, 548. doi:10.1038/s41746-025-01722-y. ↩
  14. Observational Budzyń, K. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. Lancet Gastroenterology & Hepatology, 10(10), 896–903. doi:10.1016/S2468-1253(25)00133-5. ↩
  15. Review Parasuraman, R. & Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors, 52(3), 381–410. doi:10.1177/0018720810376055. ↩
  16. Legislation Regulation (EU) 2016/679 (GDPR), Article 22. eur-lex.europa.eu. ↩