Responsible AI & Human Agency · Research · 2026Verantwoorde AI & menselijke regie · Onderzoek · 2026
Why Do Responsible AI Commitments Fail Under Pressure?
Incentives, authority and the system that decides what responsible AI is actually worth.
This research page is currently available in English only.Deze onderzoekspagina is op dit moment alleen in het Engels beschikbaar.
Accountable, but not in charge
Imagine a responsible AI lead who finds a fairness problem two weeks before a product launch. She is accountable for responsible AI. But the launch date is tied to quarterly targets, the product team's bonuses depend on adoption, and the vendor contract was chosen mainly on price and speed. She can raise the concern. She cannot move the date.
The organisation has principles, a policy and a governance board. What it does not have is an incentive system in which her concern can win.
A hypothetical illustration, not a description of a particular organisation.
The short answer
Responsible AI commitments rarely fail because people do not care. They fail because the system rewards something else. Fifty years ago, Steven Kerr described the pattern as “the folly of rewarding A, while hoping for B”: organisations hope for one behaviour while formally rewarding another.1
In AI, the organisation hopes for safety, quality, inclusion and rights. It rewards speed, adoption, cost reduction and launches. When those collide under pressure, the reward usually wins — and the person holding “responsibility” without authority absorbs the risk.
Responsible AI is not only a set of principles. It is an incentive system. If the incentives point one way and the principles another, the incentives will decide what responsible AI is actually worth.
The Responsible AI Incentive System
View the full system →
↓ Pressure KPIs and bonuses · budgets · procurement scores · deadlines · career status · adoption targets
↑ Countervailing force liability and trust · regulation · escalation protection · incident learning · user feedback
What the research shows
Old lessons about incentives
- Rewarding A, hoping for B. Organisations routinely reward behaviours they say they do not want, while hoping for behaviours they do not reward.1
- Indicators get corrupted. Campbell observed that the more a quantitative indicator is used for social decision-making, the more it is subject to corruption pressures and the more it distorts the process it was meant to monitor.2
- Goals can go wild. Specific, challenging goals can narrow focus, encourage risk-taking and unethical behaviour, and crowd out learning.3
- Speaking up depends on safety. In a study of 51 work teams, psychological safety — a shared belief that the team is safe for interpersonal risk-taking — was associated with learning behaviour.4
What responsible-AI practitioners report
- Structure and culture decide effectiveness. Interviews with industry practitioners mapped how organisational structures support or hinder responsible-AI initiatives, and what a transition to better practice requires.5
- Policies, practices and outcomes decouple. AI ethics workers struggled to have ethics prioritised in an environment centred on product launches, found ethics hard to quantify where goals are driven by metrics, and lost relationships through frequent reorganisations. Individuals took on great personal risk when raising ethics issues — especially those from marginalised backgrounds.6
- Fairness work depends on individual advocates. In a co-design process with 48 practitioners, checklists were seen as organisational infrastructure to formalise ad-hoc processes and empower individual advocates — with organisational culture shaping whether that works.7
- Ethics gets absorbed into corporate logics. Research on “ethics owners” in Silicon Valley shows how ethics is institutionalised in ways shaped by the logics of the companies adopting it.8
- Audits need to be built in, not bolted on. An end-to-end internal auditing framework aims to close the accountability gap by embedding review across the whole development lifecycle.9
A warning from outside AI
The clearest illustration of incentives overpowering stated values comes from banking, not AI. In 2016, the US Consumer Financial Protection Bureau fined Wells Fargo $100 million after employees, “spurred by sales targets and compensation incentives”, opened accounts without customers' knowledge. By the bank's own analysis, more than two million deposit and credit-card accounts may not have been authorised.10
No AI was involved. That is the point: the mechanism is older than the technology. When targets are set without guardrails and people are rewarded for the number rather than the outcome, the number wins. AI simply moves faster and at greater scale.
Responsibility without authority breeds risk
Holding teams accountable without giving them real power to act sets people up to fail — and does not make AI safer. Real authority needs three things:
- Resources
- Time, budget and people to do the work — not responsible AI as an unfunded extra.
- Information
- Access to models, data, vendor documentation, incidents and outcomes.
- Protection
- The ability to raise concerns, delay a launch or escalate without career or reputational cost.
Protection is increasingly a legal matter too. The EU AI Act applies the EU Whistleblower Directive to the reporting of infringements of the Act from 2 August 2026.1112 Legal protection is a floor; a culture in which raising a concern is expected is the goal.
This connects directly to my research on distributed accountability and meaningful human oversight: accountability is only fair when it comes with matching knowledge, capacity and authority.
An incentive audit
For each incentive in the system, ask what it rewards, what you hope for, and how to realign it.
| Incentive | What it rewards | What we hope for | How to realign |
|---|---|---|---|
| KPIs and bonuses | Usage, revenue, cost reduction | Safe, fair, valuable use | Pair every adoption metric with a quality, harm or outcome metric |
| Release deadlines | Shipping on time | Shipping when ready | Gates with evidence criteria and a named authority to delay |
| Procurement scores | Price and speed | Trustworthy suppliers | Weight audit rights, documentation, incident cooperation and exit |
| Budgets and headcount | Doing more with less | Capacity to oversee | Fund oversight, testing and remediation as part of the business case |
| Career and status | Visible launches | Courage to raise concerns | Recognise people who stop, fix or improve systems |
| Performance pressure | Doing more, faster | Careful human judgement | Protect review time; never target override rates |
The responsible AI control loop
targets → design and procurement → deployment → human oversight → outcomes → incidents and feedback → target and control redesign
The loop matters because incentives drift. Targets set at the start shape design and procurement; what happens in use should, in turn, reset the targets. If incidents and feedback never reach the people who set targets, the loop is broken — and the same pressures produce the same harms.
What leaders can change
- Make the trade-off explicit. Write down which outcome wins when speed and safety collide — before it happens.
- Balance every pressure metric with a countervailing metric at the same level of the organisation.
- Give responsible AI a budget line and decision rights, not only an advisory voice.
- Protect the people who raise concerns, and track what happened after they did.
- Close the loop: route incidents, complaints and outcome data back to the people who set targets and incentives.
- Review incentives as rigorously as models. An incentive that reliably produces harm is a design defect.
If you want to know what an organisation really believes about responsible AI, do not read its principles. Read its incentives.
Method and limits
This is a research-led synthesis of organisational research, responsible-AI practitioner studies, regulatory action and EU law, prepared in October 2026. Most responsible-AI practitioner evidence is qualitative and drawn from large technology companies; incentive effects in smaller organisations and the public sector are less studied. The incentive-system model, the incentive audit and the leader actions are my own proposals. I checked each source below against its published record.
Sources
- Theory Kerr, S. (1975). On the folly of rewarding A, while hoping for B. Academy of Management Journal, 18(4), 769–783. doi:10.2307/255378. ↩
- Theory Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67–90. doi:10.1016/0149-7189(79)90048-X. ↩
- Review Ordóñez, L. D., Schweitzer, M. E., Galinsky, A. D. & Bazerman, M. H. (2009). Goals gone wild: The systematic side effects of overprescribing goal setting. Academy of Management Perspectives, 23(1), 6–16. doi:10.5465/amp.2009.37007999. ↩
- Field study Edmondson, A. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350–383. doi:10.2307/2666999. ↩
- Interview study Rakova, B., Yang, J., Cramer, H. & Chowdhury, R. (2021). Where responsible AI meets reality. Proc. ACM HCI, 5(CSCW1). doi:10.1145/3449081. ↩
- Interview study Ali, S. J., Christin, A., Smart, A. & Katila, R. (2023). Walking the walk of AI ethics: Organizational challenges and the individualization of risk among ethics entrepreneurs. FAccT '23, 217–226. doi:10.1145/3593013.3593990. ↩
- Co-design study Madaio, M. A., Stark, L., Wortman Vaughan, J. & Wallach, H. (2020). Co-designing checklists to understand organizational challenges and opportunities around fairness in AI. CHI 2020. doi:10.1145/3313831.3376445. ↩
- Ethnography Metcalf, J., Moss, E. & boyd, d. (2019). Owning ethics: Corporate logics, Silicon Valley, and the institutionalization of ethics. Social Research, 86(2), 449–476. doi:10.1353/sor.2019.0022. ↩
- Framework Raji, I. D. et al. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. FAT* '20, 33–44. doi:10.1145/3351095.3372873. ↩
- Regulatory action U.S. Consumer Financial Protection Bureau (2016). Wells Fargo Bank, N.A. enforcement action. consumerfinance.gov. ↩
- Legislation AI Act, Article 87 — Reporting of infringements and protection of reporting persons. artificialintelligenceact.eu. ↩
- Legislation Directive (EU) 2019/1937 on the protection of persons who report breaches of Union law. eur-lex.europa.eu. ↩