Responsible AI & Human Agency · Research · 2026Verantwoorde AI & menselijke regie · Onderzoek · 2026
When Does AI Assistance Become Human Dependency?
Evidence, mechanisms, design boundaries and governance — and why what people can do with AI is only half the question.
This research page is currently available in English only.Deze onderzoekspagina is op dit moment alleen in het Engels beschikbaar.
Faster every month
Imagine a junior analyst who starts using an AI assistant for first drafts of every report. Each month she is faster. Her managers are pleased; her reports are cleaner than ever.
A year later the assistant is unavailable for a week. She finds it surprisingly hard to structure an argument from scratch — and, more worryingly, she struggles to say which parts of last month's AI-assisted analysis were weak.
Did the AI make her better at her job, or only better at producing her job's output?
This is a hypothetical illustration, not a report about a particular person or organisation.
The short answer
Using an external tool is not dependency. Humans have always offloaded memory and calculation to writing, calculators, maps and colleagues. AI assistance becomes dependency when use of the tool starts to undermine the capabilities a person needs to:
- form an independent judgement;
- detect or correct the tool's errors;
- learn or transfer the underlying skill; or
- function safely when the tool is unavailable or wrong.
The research does not say “AI makes people less capable”. It describes a performance–capability trade-off whose direction depends on what is offloaded, how the interaction is designed and whether people keep exercising the skills that matter for learning, verification and fallback.
Performance — with AI
- Speed
- Quality
- Productivity
- Accessibility
Capability — without AI
- Independent judgement
- Skill retention
- Error detection
- Recovery ability
AI strengthens human capability when it removes unnecessary burden while preserving the chance to build mental models, practise consequential skills, receive feedback and exercise independent judgement. It creates dependency when it removes those chances while leaving the human nominally responsible for checking it or taking over.
1. What is dependency?
Four mechanisms account for most of the risk.
- Cognitive offloading
- Using external tools to reduce the mental demands of a task. Usually helpful in the moment; risky when the offloaded operation is needed later for learning, checking or fallback.1
- Automation bias
- Treating an automated output as a shortcut for vigilant judgement — following wrong advice, or missing what the system did not flag. It is more likely when checking is cognitively complex.345
- Out-of-the-loop skill loss
- When automation does the routine work, people lose the practice they need for the rare moments it fails.67
- Trust miscalibration
- Appropriate trust is not maximum trust: it is reliance that tracks what the system can actually do — including knowing when not to rely on it.8
Offloading is not the enemy. In the well-known “Google effect” experiments, people who expected information to remain available remembered where to find it rather than the information itself2 — a redistribution of memory, not proof of decline. The problem begins when the offloaded task is part of the expertise itself.
Lisanne Bainbridge described the structural trap in 1983: the more reliably automation handles routine work, the less opportunity people have to maintain the skill needed when it fails.6 A system that is wrong half the time keeps people alert. The harder problem is one that is right almost all the time.
A proposed working definition
There is no accepted numerical threshold between assistance and dependency. I propose calling AI use functionally dependent when all three conditions hold:
- The AI performs or strongly shapes an operation needed for the task.
- Repeated assistance materially reduces the person's independent capacity to perform, verify or recover that operation.
- The person's remaining responsibility still assumes that capacity.
Frequent calculator use is not problematic when mental long division is irrelevant to someone's role. Frequent autopilot use is different when pilots remain responsible for recognising failures and recovering the aircraft. The mismatch between responsibility and retained capability is what matters.
2. What does the research show?
Evidence exists for both trajectories — augmentation and dependency.
Human + AI is not automatically better
A preregistered meta-analysis of 106 experimental studies and 370 effect sizes found that human–AI combinations performed, on average, worse than the better of the human or the AI alone (Hedges' g = −0.23). Combinations did relatively poorly on decision tasks and better on content creation.16
Strong evidence of augmentation
- Customer support. Across 5,172 agents, an AI assistant raised issues resolved per hour by 15% on average, with the biggest gains for less experienced and lower-skilled workers — and evidence of worker learning and improved English fluency.17
- Professional writing. In an experiment with 453 college-educated professionals, ChatGPT cut time by about 40% and raised rated quality by about 18%, narrowing the gap between workers.14
- Skin-cancer diagnosis. Good-quality AI support improved accuracy beyond either AI or physicians alone, and the least experienced clinicians gained most.10
Strong signals of dependency
- The same diagnosis study: faulty AI misled clinicians across the whole spectrum of expertise, including experts.10
- Colonoscopy. In a multicentre observational study, the adenoma detection rate of standard, non-AI colonoscopy fell from 28.4% to 22.4% in the three months after endoscopists began using AI assistance. Observational design means other explanations remain possible, but it measures exactly the right thing: performance when the AI is absent.19
- Mathematics. Students with unrestricted GPT-4 access did far better during practice but worse than a control group once access was removed; a tutor design with safeguards largely avoided the penalty.18
- Navigation. Among 50 drivers, more lifetime GPS use was associated with poorer spatial memory during self-guided navigation; in 13 participants retested three years later, more GPS use was associated with steeper decline. The longitudinal sample was small.9
- Automated driving. A meta-analysis of 51 experiments found that non-driving tasks impaired takeover performance — being in the seat is not being ready to take over.13
Individual gain, collective narrowing
In creative writing, AI-generated ideas improved individual stories — especially for less creative writers — while making the stories more similar to one another.15 Each person can get better while the field gets narrower.
3. When does AI improve learning?
The key distinction is between scaffolding and substitution. Scaffolding gives temporary help that reveals or stimulates the reasoning a learner needs, and can later be withdrawn. Substitution delivers the finished product while bypassing the operation the learner needed to acquire.
The mathematics study shows the difference within a single tool: the same model, designed to answer or designed to tutor, produced different effects on later unaided performance.18 The customer-support study shows that AI can also propagate expertise: novices became faster and showed evidence of learning.17
A practical pattern where learning matters:
Human attempt → AI critique → human revision → optional AI example → independent retrieval later
Expertise has a non-linear relationship with risk. Novices often gain most from AI, but may lack the knowledge to recognise subtle errors. Experts have stronger models, but are not immune to being misled.
4. Which design choices make the difference?
| Factor | Tends toward augmentation | Tends toward dependency |
|---|---|---|
| Task | AI handles separable retrieval or routine processing | AI performs the reasoning humans must later verify or recover |
| Explanation | Shows evidence, uncertainty, provenance and alternatives | Plausible rationales that only make advice more persuasive |
| Order | Person forms a view before seeing the AI output | AI output is the starting anchor |
| Feedback | People learn when the AI was right or wrong | No ground truth ever reaches the user |
| Practice | Regular unaided work keeps skills alive | Months of automated operation with little manual practice |
| Time | Checking is resourced | Checking takes ten minutes; the workflow allows thirty seconds |
| Reliability | Visible limits and local reliability information | High apparent reliability with rare, consequential failures |
Two findings deserve emphasis. First, explanation is not safety: in experiments across three tasks, AI explanations did not improve complementary team performance; they increased acceptance of the AI's recommendation regardless of whether it was correct.11 Second, productive friction works but is unpopular: cognitive-forcing designs — such as asking for a judgement before showing the AI's — significantly reduced overreliance (N=199), yet participants rated them least favourably.12
That is the value of friction in practice: a small, well-placed effort can be a safety feature, even when people would rather skip it.
A design rule
Never automate away a capability that people are still expected to exercise when the system fails — unless the system deliberately keeps that capability alive.
5. How should dependency be measured and governed?
Most AI evaluation measures performance with the system switched on. Dependency is about what happens after repeated use — especially when assistance is removed, degraded or wrong. Establish a baseline before deployment; without one, later change is hard to attribute.
The instruments in this section are proposed assessment methods for further development — not validated scientific measures.
Two dashboards, not one
| Performance dashboard | Capability dashboard |
|---|---|
| Output quality | Unaided proficiency |
| Productivity and cycle time | Error detection when the AI is wrong |
| AI accuracy | Reliance calibration |
| User satisfaction | Independent verification behaviour |
| Adoption | Skill retention and transfer |
| Automation rate | Recovery under outage; expertise coverage |
Proposed measures
- Reliance discrimination = P(accept | AI correct) − P(accept | AI wrong). Someone who accepts everything can look very “trusting” but scores near zero.
- Skill retention ratio = unaided performance after sustained AI use ÷ unaided performance before.
- Recovery — error detection, time to stabilise and residual errors when AI is deliberately wrong, slow or absent.
- Process — does the person form a view before seeing the AI, check independent evidence, ask for provenance?
I would not combine these into a single dependency score until the indicators and weights have been validated. A single number would hide the weakest link.
Institutional dependency
Individual skill is only one level. Can an organisation become more capable with AI while its people become less knowledgeable? An organisation can raise productivity while losing internal expertise, independent verification capacity or the ability to operate without one provider. It can also do the opposite: use AI to spread specialist knowledge and speed up development. The customer-support evidence shows the second path is possible.17 Organisations should measure which path they are on — outage degradation, recovery time, how many people can do critical tasks unaided, verification capacity relative to AI output volume, and whether juniors still build independent competence with tenure.
For leaders
- Define a minimum residual capability: what must people still be able to do without AI, why, and how it will be tested.
- Train for AI failure literacy, not just prompting: plausible wrong answers, missing evidence, disagreement and escalation.
- Protect the expertise pipeline: if juniors bypass the hard work through which seniors learned, today's productivity becomes tomorrow's shortage of experts.
- Treat human-performance drift as seriously as model drift in post-deployment monitoring.
The standard
Do not measure human–AI success only by what people can accomplish with AI. Measure what they learn from it, what they can still do without it, whether they know when not to trust it — and how well the combined system survives when the AI is wrong.
Limits of the evidence
- Exposure is short. Most studies last minutes or hours; dependency may develop over months or years.
- Transfer is uncertain. GPS findings do not prove that generative AI causes broad cognitive decline; one mathematics setting does not settle law, medicine or programming.
- Neuroscience is easy to overstate. Reduced brain engagement while a tool does the work is not evidence of structural damage.
- Organisational memory is understudied. Most research looks at individuals, not staffing, training pipelines or vendor concentration.
Method and limits
This is a research-led synthesis of human-factors, cognitive-science, HCI, education, medical and economics research, prepared in October 2026. It is not a systematic review. I checked each source below against its published record or abstract. The working definition, dashboards, measures and institutional-dependency questions are my own proposals for further development.
Sources
- Review Risko, E. F. & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. doi:10.1016/j.tics.2016.07.002. ↩
- Experiment Sparrow, B., Liu, J. & Wegner, D. M. (2011). Google effects on memory. Science, 333(6043), 776–778. doi:10.1126/science.1207745. ↩
- Theory Parasuraman, R. & Manzey, D. H. (2010). Complacency and bias in human use of automation. Human Factors, 52(3), 381–410. doi:10.1177/0018720810376055. ↩
- Systematic review Goddard, K., Roudsari, A. & Wyatt, J. C. (2012). Automation bias: a systematic review. Journal of the American Medical Informatics Association, 19(1). doi:10.1136/amiajnl-2011-000089. ↩
- Systematic review Lyell, D. & Coiera, E. (2017). Automation bias and verification complexity: a systematic review. JAMIA, 24(2), 423–431. doi:10.1093/jamia/ocw105. ↩
- Theory Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775–779. doi:10.1016/0005-1098(83)90046-8. ↩
- Meta-analysis Onnasch, L., Wickens, C. D., Li, H. & Manzey, D. (2014). Human performance consequences of stages and levels of automation. Human Factors, 56(3), 476–488. doi:10.1177/0018720813501549. ↩
- Theory Lee, J. D. & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50–80. doi:10.1518/hfes.46.1.50_30392. ↩
- Observational Dahmani, L. & Bohbot, V. D. (2020). Habitual use of GPS negatively impacts spatial memory during self-guided navigation. Scientific Reports, 10. doi:10.1038/s41598-020-62877-0. ↩
- Experiment Tschandl, P. et al. (2020). Human–computer collaboration for skin cancer recognition. Nature Medicine, 26(8), 1229–1234. doi:10.1038/s41591-020-0942-0. ↩
- Experiment Bansal, G. et al. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. CHI 2021. doi:10.1145/3411764.3445717. ↩
- Experiment Buçinca, Z., Malaya, M. B. & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI. Proc. ACM HCI, 5(CSCW1). doi:10.1145/3449287. ↩
- Meta-analysis Weaver, B. W. & DeLucia, P. R. (2022). A systematic review and meta-analysis of takeover performance during conditionally automated driving. Human Factors. doi:10.1177/0018720820976476. ↩
- Experiment Noy, S. & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192. doi:10.1126/science.adh2586. ↩
- Experiment Doshi, A. R. & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28). doi:10.1126/sciadv.adn5290. ↩
- Meta-analysis Vaccaro, M., Almaatouq, A. & Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour, 8(12), 2293–2303. doi:10.1038/s41562-024-02024-1. ↩
- Field study Brynjolfsson, E., Li, D. & Raymond, L. (2025). Generative AI at work. Quarterly Journal of Economics, 140(2), 889–942. doi:10.1093/qje/qjae044. ↩
- Field experiment Bastani, H. et al. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS, 122(26). doi:10.1073/pnas.2422633122. ↩
- Observational Budzyń, K. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterology & Hepatology, 10(10), 896–903. doi:10.1016/S2468-1253(25)00133-5. ↩