← Writing← Schrijven

Responsible AI & Human Agency · Research · 2026Verantwoorde AI & menselijke regie · Onderzoek · 2026

Can an AI System Be Legally Compliant Yet Still Irresponsible?

Why passing the legal and technical tests is not the same as deploying AI responsibly — and what to ask instead.

This research page is currently available in English only.Deze onderzoekspagina is op dit moment alleen in het Engels beschikbaar.

Six months later

Imagine an organisation introducing an AI assistant. The legal team has approved the deployment. The security assessment is complete. The system has passed its technical tests.

Six months later, employees increasingly accept its recommendations without checking them. Junior colleagues find it harder to complete tasks on their own. Managers celebrate the productivity gains, but nobody is measuring what people might be losing.

Has the organisation deployed AI responsibly?

This is a hypothetical scenario, not a description of a particular organisation.

The short answer

Yes, a system can be compliant and still irresponsible. Legal compliance establishes that a system meets the obligations that apply to a specific actor, use, jurisdiction and moment. It does not establish that:

  • the problem was a legitimate one to automate;
  • the technical objective reflects the real human need;
  • the deployment is fair, accessible and proportionate;
  • human oversight actually works in practice;
  • benefits and burdens are distributed acceptably;
  • affected people can understand, challenge and obtain remedy; or
  • the system stays responsible as users, data and circumstances change.

Compliance is evidence of rule-conformance. Responsibility is a continuing judgement about purpose, power, practice and consequences.

The model at a glance

Diagram titled Four layers of responsible AI governance: from principles to people, with continuous learning. Four stacked layers, each with the evidence it produces: 1. Legal compliance (laws, duties, prohibited uses, documentation, conformity; evidence: legal opinions and conformity records). 2. Technical assurance (validation, robustness, security, privacy, fairness testing, explainability, monitoring; evidence: test results and audit evidence). 3. Organisational ethics and governance (purpose, incentives, accountable owners, escalation, procurement, human oversight, deployment authority; evidence: decisions, ownership, culture and resources). 4. Actual human outcomes (benefits, error burdens, accessibility, dignity, autonomy, contestability, distributional effects, remedy; evidence: lived experience, complaints, outcome disparities and downstream effects). Between each pair of layers is a responsibility gap: laws met but real-world harm still possible; technically sound but misaligned or misused; well governed but unintended consequences remain. A feedback loop runs from real-world outcomes back to governance, testing and legal review, all set within the context of people, systems and society.
  1. Responsibility gap — laws met, but real-world harm still possible
  2. 2Technical assuranceValidation, robustness, security, privacy, fairness testing, explainability, monitoring.Evidence: test results and audit evidence
  3. Responsibility gap — technically sound, but misaligned or misused
  4. 3Organisational ethics and governancePurpose, incentives, accountable owners, escalation, procurement, human oversight, deployment authority.Evidence: decisions, ownership, culture and resources
  5. Responsibility gap — well governed, but unintended consequences remain
  6. 4Actual human outcomesBenefits, error burdens, accessibility, dignity, autonomy, contestability, distributional effects, remedy.Evidence: lived experience, complaints, outcome disparities and downstream effects
  7. ↺ Feedback loop Real-world outcomes inform governance, testing and legal review.
Responsible AI is not achieved through compliance alone. Legal, technical, organisational and human responsibilities must continuously inform one another. This is my proposed synthesis of existing governance approaches — a conceptual model, not an established or validated assessment framework.

Four layers that are often confused

Organisations often treat compliance, technical assurance, ethics and outcomes as alternative proofs of responsibility. They are complementary controls, and each answers a different question.

LayerWhat it can establishWhat it cannot establishTypical failure
Legal complianceThat identified conduct meets applicable law and enforceable dutiesThat the law covers every harm, or that the use is socially legitimate“Lawful, therefore acceptable”
Technical assuranceThat specified properties — accuracy, robustness, security, chosen fairness metrics — hold under stated test conditionsThat the target is the right one, or that tests represent real deployment“Passed the benchmark, therefore safe”
Organisational governanceThat values are translated into ownership, incentives, review gates and the authority to stopThat declared principles actually hold under commercial or political pressure“We have principles, therefore we act ethically”
Human outcomesWho benefits, who is burdened, how errors fall, and whether people keep dignity, agency and remedyOn their own, the causal explanation or a universally agreed fair allocation“Average outcomes improved, therefore no one was wronged”

These layers are not a ladder that ends at compliance. They form a feedback loop: what happens to people after release should update technical tests, organisational decisions and, where new harms expose gaps, legal interpretation. The US National Institute of Standards and Technology makes a related point: trustworthy AI is sociotechnical, and addressing its characteristics separately does not make a system trustworthy.1

Where the gaps open

Coverage gap
No law covers the harm, or the use falls below a legal risk threshold.
Specification gap
The system performs well on its metric, but the metric represents the wrong goal.
Translation gap
Principles exist, but nobody has the authority, resources or incentive to delay a launch.
Deployment gap
The model was tested correctly, but workflow, staffing, interface or automation bias changes its effects.
Evidence gap
People experience harm, but the organisation only measures technical failures or averages.
Supply-chain gap
Provider, integrator and deployer each control different risks and each assume someone else owns the rest.

Law: an essential floor, not a complete verdict

The EU AI Act entered into force on 1 August 2024. Prohibitions and AI-literacy duties applied from February 2025, general-purpose AI rules from August 2025, and most of the remaining Act from 2 August 2026. Following the AI Omnibus amendments, obligations for high-risk systems listed in Annex III apply from 2 December 2027, and for high-risk systems in regulated products from 2 August 2028.23

For high-risk systems, the Act requires risk management, data governance, documentation, logging, human oversight, accuracy, robustness, cybersecurity and post-market monitoring. That is substantial, and it narrows the gap this article describes. But conformity does not always mean independent scrutiny: for many Annex III systems, providers assess conformity through internal control. And even verified conformity shows that defined requirements were met, not that a deployment is necessary, proportionate or harmless. The Commission itself notes that the Act sets no specific rules for the large majority of AI systems, which it classes as minimal risk.3

Other law still matters. In SCHUFA, the EU Court of Justice held that a credit score can itself be automated individual decision-making under the GDPR when the bank receiving it draws strongly on that score to decide on a contract.4 What counts is the system's actual role in a decision, not whether it is labelled “advice” or “support”.

The Council of Europe Framework Convention on AI, opened for signature on 5 September 2024, frames AI around human dignity, autonomy, equality, privacy and remedy. The European Union is listed as a Party and many states as signatories.5 Its obligations work through states, not as a certificate for individual products.

Why can lawful still mean irresponsible?

  • Law is selective. A minimal-risk system can still manipulate attention, degrade work quality or exclude users.
  • Law is role-specific. A provider can meet its duties while a deployer chooses an inappropriate use.
  • Law is a snapshot. A system can comply at release and then drift, be repurposed or meet new misuse.
  • Law permits trade-offs. Legality often means a balance was allowed, not that everyone affected finds it fair.
  • Law lags evidence. Harms can become visible before classifications or standards change.

Technical assurance: evidence about a chosen claim

Technical assurance asks whether a system meets specified claims under specified conditions: validation, subgroup analysis, robustness and security testing, red-teaming, logging and drift detection. Frameworks such as the NIST AI Risk Management Framework1 and the ISO/IEC 42001 management-system standard6 bring discipline to this work. But certifying a management system is not proof that every deployment produces acceptable outcomes.

  • The proxy may be wrong. A system can accurately predict healthcare spending while failing to identify healthcare need (see the first case below).
  • Fairness is not one property. Common fairness criteria can be mathematically incompatible when groups differ in base rates.7 Choosing a metric is a policy judgement, not just an optimisation.
  • Explanation is not justification. Knowing why a system produced an output does not establish that its objective or use is appropriate.
  • People over-rely on outputs. The AI Act asks oversight design to address automation bias and to let people disregard, reverse or stop the system.
  • Lab performance may not transfer. A system that works in testing can behave differently inside real workflows and incentives.

Governance: where intent meets power

Organisational ethics is about what an organisation is willing to do, not just what it says it values. A credible responsible-AI function has access to product and deployment evidence, independence from delivery targets, influence over budget and launch decisions, a way to record unresolved dissent, direct escalation to accountable executives, protection for people who raise concerns and continuing contact with affected people.

Without those, principles risk becoming ethics-washing: an ethical appearance without the reasoning, implementation or authority behind it. The opposite risk is real too: ethics should not become an unaccountable veto. Its legitimate role is to make value conflicts explicit, structure participation and require accountable leaders to justify their decisions.

Human outcomes: the hardest and most important test

Outcome governance asks questions that technical tests cannot answer on their own:

  • Benefit — did the relevant human or public outcome improve, not just throughput or accuracy?
  • Burden — who absorbs false positives, false negatives, delay and the work of correcting errors?
  • Agency — can a person refuse, reach a human, or choose a non-AI route without penalty?
  • Capability — does the system build people's skills and judgement, or quietly erode them?
  • Contestability — is there notice, an intelligible reason, timely appeal and effective remedy?
  • Dignity and accessibility — does the system stigmatise, surveil or exclude?
  • Distribution — are benefits broad while harms are concentrated?

UNESCO's Recommendation on the Ethics of AI joins proportionality, do-no-harm, human dignity, oversight and multi-stakeholder participation in a similar spirit.8 Outcome evidence has limits — causality can be uncertain, qualitative harms resist aggregation and people disagree about fairness — but that is a reason to combine numbers, testimony and transparent judgement, not to stop looking.

This is where the question I care most about sits: responsible according to whom, and responsible for what? An assistant can raise productivity while reducing independent problem-solving. A recruitment system can meet its accuracy and fairness metrics while creating barriers for disabled candidates. A healthcare tool can improve average outcomes while leaving some patients less able to understand or challenge decisions about them. None of these shows up as a failed test.

Two cases, examined closely

A healthcare algorithm that measured the wrong thing

Researchers studied a widely used commercial algorithm that helped decide which patients received extra care. At the same risk score, Black patients were considerably sicker than White patients. The reason was the target: the algorithm predicted healthcare costs rather than illness, and unequal access to care meant less money was spent on Black patients with the same needs. Correcting this would have raised the share of Black patients receiving additional help from 17.7% to 46.5%.9

What it shows: a technically plausible, predictively useful proxy encoded an existing inequality. This is a specification and outcome failure. It says nothing about legal compliance either way — and that is the point: the test that mattered was not a legal one.

SyRI: authorised by law, and still unlawful

The Dutch System Risk Indication (SyRI) was a statutory instrument the government used to detect fraud in benefits, allowances and taxes. In 2020 the District Court of The Hague held that the legislation governing SyRI breached Article 8 of the European Convention on Human Rights, the right to respect for private life. The court found that the use of SyRI was insufficiently transparent and verifiable and did not strike the required fair balance between the public interest and the intrusion into private life.10

What it shows: a statutory basis was not the same as compliance with higher law, let alone legitimacy. Procedure and authorisation are not proof of responsibility.

Other cases worth knowing

  • iTutorGroup (United States). The EEOC alleged that tutor-application software automatically rejected women aged 55+ and men aged 60+; the companies settled for $365,000 with monitoring and other relief.11 Automation scaled an unlawful rule; it did not move accountability away from the employer.
  • Moffatt v. Air Canada (Canada). A website chatbot gave incorrect fare information, and the tribunal held the airline responsible for information on its own website.12 The deployer remained responsible for the customer-facing system.

None of these cases proves that a harmful system was legally compliant. They show something more careful: a statute, a vendor, technical plausibility or automation did not prevent harm — and courts or regulators later exposed the gap.

Seven decision gates

  1. Purpose — should this be automated? State the human or public outcome, the alternatives considered and why AI is necessary and proportionate.
  2. Rights and law — may we do it? Map jurisdictions, roles, affected rights, prohibited uses, sector law and required assessments.
  3. Evidence — does it work here? Test with representative populations, realistic workflows, subgroup results and known uncertainty.
  4. Organisation — who can stop it? Name an accountable executive, an operational owner, a technical owner and an independent escalation route, each with real authority.
  5. Deployment — can people use and challenge it safely? Test oversight, workload, accessibility, explanations, non-AI alternatives, appeal and remedy.
  6. Outcome — what happened after release? Monitor benefit, error burden, complaints, overrides, appeals and effects on people's skills and judgement.
  7. Renewal — is continued use still justified? Time-limit approval; reassess after material change; define rollback and exit criteria.

The board question is not “did it pass?” It is:

What exactly passed, against whose criteria, under which conditions — and what evidence would cause us to stop?

Fair objections

  • “Responsibility beyond law is subjective.” Partly true. That is why it should rest on published values, human rights, structured participation and documented reasoning — not private moral rules.
  • “Outcome accountability punishes innovation.” Governance should be proportionate. But uncertainty argues for staged deployment, monitoring and reversibility, not for no responsibility.
  • “No system can satisfy every fairness measure.” Correct — which is exactly why the choice should be explicit and accountable.
  • “Humans are flawed too.” Also correct. The fair comparison is often AI-assisted versus current practice. Yet scale, opacity and persistence can make automated harm different in kind.

Layered accountability

The mature position is not ethics instead of law, or outcomes instead of testing. It is layered:

  • law supplies enforceable floors;
  • assurance supplies inspectable evidence;
  • governance supplies authority and incentives;
  • the people affected, and what actually happens to them, supply correction.

The decisive test is not whether an organisation can defend a system at release. It is whether it can justify the purpose, evidence the safeguards, hear those affected, correct the consequences — and stop the system when continuing is no longer defensible.

Seven questions before deployment

An exploratory reflection exercise — not a validated audit, certification or legal advice. Your answers stay in this browser tab and are not stored or sent anywhere.

  1. Have we established that AI is appropriate for the problem we are trying to solve?
  2. Have we assessed applicable legal obligations and fundamental-rights risks?
  3. Have we tested technical performance under realistic deployment conditions?
  4. Have we examined potential effects on human autonomy, skills and decision-making?
  5. Have affected stakeholders had a meaningful opportunity to contribute?
  6. Can people challenge decisions and obtain effective remedies?
  7. Is someone authorised to suspend or withdraw the system if necessary?

Answer the questions to see a reflection.

This is the starting point for a Responsible AI & Human Agency framework I am developing, with evidence requirements, escalation criteria and validation across deployment contexts.

Questions this research opens

  • When does AI assistance become human dependency?
  • What does meaningful human oversight actually mean?
  • Who is accountable when AI makes a mistake across a supply chain?
  • Can we measure the human cost of AI efficiency?
  • When is withdrawal more responsible than further mitigation?

Method and limits

This is a research-led synthesis of legislation, case law, empirical research and standards, prepared in October 2026. It is not legal advice or a systematic review. I checked each source below against its original or official publication; the legal status of the EU AI Act and the Council of Europe Convention is described as of 3 October 2026 and may change. The four-layer model, the decision gates and the self-check are my own proposals.

Sources

  1. Framework National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). doi:10.6028/NIST.AI.100-1 · nist.gov. ↩
  2. Legislation Regulation (EU) 2024/1689 (Artificial Intelligence Act). eur-lex.europa.eu · implementation timeline (as amended): artificialintelligenceact.eu. ↩
  3. Official guidance European Commission. AI Act — regulatory framework for AI. digital-strategy.ec.europa.eu. ↩
  4. Case law Court of Justice of the EU, Case C-634/21, SCHUFA Holding (Scoring), 7 December 2023. eur-lex.europa.eu. ↩
  5. Treaty Council of Europe Framework Convention on Artificial Intelligence and Human Rights, Democracy and the Rule of Law (CETS No. 225). coe.int. ↩
  6. Standard ISO/IEC 42001:2023, Artificial intelligence — Management system. iso.org. ↩
  7. Empirical research Chouldechova, A. (2017). Fair prediction with disparate impact: A study of bias in recidivism prediction instruments. Big Data, 5(2), 153–163. doi:10.1089/big.2016.0047. ↩
  8. International recommendation UNESCO (2021). Recommendation on the Ethics of Artificial Intelligence. unesco.org. ↩
  9. Empirical research Obermeyer, Z., Powers, B., Vogeli, C. & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. doi:10.1126/science.aax2342. ↩
  10. Case law Rechtbank Den Haag, 5 February 2020, ECLI:NL:RBDHA:2020:865 (SyRI). rechtspraak.nl. English translation: ECLI:NL:RBDHA:2020:1878. ↩
  11. Regulatory action U.S. Equal Employment Opportunity Commission. iTutorGroup to pay $365,000 to settle EEOC discriminatory hiring suit. eeoc.gov. ↩
  12. Case law Moffatt v. Air Canada, 2024 BCCRT 149 (Civil Resolution Tribunal, British Columbia). canlii.org. ↩