← Writing← Schrijven

Responsible AI & Human Agency · Research · 2026Verantwoorde AI & menselijke regie · Onderzoek · 2026

Measuring the Human Cost of AI-Driven Efficiency

Minutes saved are not the final result. What happens to the time, the people and the value is.

This research page is currently available in English only.Deze onderzoekspagina is op dit moment alleen in het Engels beschikbaar.

Two hours saved — and still busy

Imagine a case manager whose AI assistant now drafts her reports. Each report takes half the time it used to. Across a week, she saves around two hours.

A month later her caseload has risen to match. The reports are faster, but every draft needs checking, the hard cases have not changed, and she is now expected to reply within the hour. On the dashboard, productivity is up. In her week, nothing is lighter.

Did the AI make her work better — or only faster?

A hypothetical illustration, not a report about a particular person or organisation.

The short answer

The human cost of AI-driven efficiency can be measured — but not with one productivity number. It needs a portfolio of economic, human, professional and societal outcomes, measured over time.

The evidence already rules out two simple stories. AI is not uniformly productivity-enhancing: controlled studies range from a 40% reduction in writing time to a 19% increase in completion time for experienced developers in familiar code.15 And efficiency does not automatically become human benefit: research using two decades of time-diary data associates higher occupational AI exposure with longer workdays and less leisure, especially where workers have little bargaining power.7

An AI productivity gain is meaningful only when output is quality-adjusted, displaced burdens are counted, human agency and capability are preserved, and the distribution of gains is visible and defensible.

Where does the saved time go?

RecoveryReduced workload, shorter hours, better wellbeing
DevelopmentLearning, mentoring, professional growth
PerformanceMore output, better quality, better service
IntensificationHigher targets, more monitoring, more workload

Illustrative outcomes, not measured findings.

The technology does not decide which path a saved hour takes. Decisions about staffing, targets, incentives and job design do. That makes the human cost of efficiency a leadership and governance question, not only a technical one. A school-randomised trial with 259 science teachers shows why the path must be recorded: teachers using ChatGPT saved 25.3 minutes a week on lesson preparation — about a third — with no detected difference in resource quality, and reported using the time for other teaching work or to reduce workload.10

The Human Impact Balance Sheet

Diagram titled Human Impact Balance Sheet for AI-driven Efficiency. An AI intervention — use case, technology and data, implementation, change management — leads to task changes, which are measured at eight levels: task, worker, job or role, team, organisation, labour market, distribution of gains, and society or public value. For each level, a green column lists benefits and opportunities (for example faster completion, reduced routine work, new skills, better collaboration, new services, new job categories, broader access, better public services) and a red column lists costs and risks (for example errors and rework, higher cognitive load and always-on expectations, loss of autonomy and surveillance, uneven workload, restructuring, displacement and inequality, concentration of value, environmental footprint). A note says a positive task result can create downstream costs at worker, job, team and societal levels. Feedback loops on the left — governance and accountability, worker voice, training and skill development, job and work redesign — feed insights back into the intervention. The bottom line reads: net human value equals quality-adjusted benefit minus displaced burdens minus distributional and external costs. View the full balance sheet →
  1. TaskFaster or more accurate — or more errors and rework?
  2. WorkerLess routine work — or more checking, vigilance and always-on pressure?
  3. Job / roleBroader, more autonomous — or narrowed and monitored?
  4. TeamBetter collaboration — or uneven load and less peer learning?
  5. OrganisationReal gains after implementation, checking and vendor costs?
  6. Labour marketNew roles and entry paths — or displaced apprenticeships?
  7. DistributionWho captures the value: workers, firms, customers or vendors?
  8. SocietyBetter public value — or eroded trust and external costs?
Figure 1 — The Human Impact Balance Sheet. AI effects travel from the task to the worker, the organisation and society. Net human value = quality-adjusted benefit − displaced burdens − distributional and external costs. My proposed assessment model — not a validated instrument or an accounting standard.

Three distinctions sit underneath it:

  • Task productivity is not worker productivity. Faster drafting can be offset by checking, integration or more volume.
  • Worker productivity is not organisational productivity. Licensing, training, security, governance and rework can consume the gain.
  • Measured output is not wellbeing. More cases handled can coexist with less discretion, poorer health or weaker relationships with the people served.

What the evidence shows

Each finding is labelled by design: experiments and field trials can support causal claims; observational and natural-experiment studies show associations.

SettingFindingWhat it means for human impact
Professional writing
Experiment
453 professionals: time fell 40%, rated quality rose 18%; weaker writers gained most.1Strong task evidence; short, self-contained assignments with little verification burden.
Customer support
Field study
5,172 agents: 15% more issues resolved per hour; biggest gains for newer and lower-skilled agents; evidence of learning.2AI can spread good practice — but output targets can later ratchet up.
Management consulting
Experiment
758 consultants: inside AI's capability frontier, 12.2% more tasks and 25.1% faster with higher quality; outside it, 19% less likely to be correct.3The “jagged frontier”: knowing when not to use AI becomes part of the job.
Software development
Field experiments
Three company trials, 4,867 developers: 26.08% more completed tasks; larger gains for less experienced developers.4Completed tasks are not the same as maintainable, defect-free value.
Experienced developers
Randomised trial
16 developers, 246 real tasks: early-2025 AI tools made them 19% slower, while they believed they were 20% faster.5Prompting, waiting and reviewing can outweigh generation speed — and perceived productivity can mislead.
Financial analysts
Natural experiment
AI access produced richer, more timely reports, but relative forecast accuracy fell when information-processing demands were high.6More information is not better decisions; attention becomes the bottleneck.
Working time
Observational
Higher occupational AI exposure is associated with longer work hours and less leisure, amplified where workers have limited bargaining power.7Gains can be captured by firms or consumers while workers absorb the time.
Mammography
Experiment
With incorrect AI suggestions, correct ratings by inexperienced radiologists fell from 79.7% to 19.8%; very experienced from 82.3% to 45.5%.8A “human in the loop” does not guarantee effective oversight.
Colonoscopy
Observational
Unaided adenoma detection fell from 28.4% to 22.4% after routine AI exposure.9A longitudinal warning about skill retention — not proof of deskilling.
Teachers
Randomised trial
259 teachers: 25.3 fewer minutes a week on preparation (about 31%), no detected quality difference.10Records explicitly how saved time was used.

Causal evidence is strongest for bounded tasks and structured workflows. It weakens as work becomes context-heavy, interdependent, long-term or hard to score.

Cognitive workload: removed, moved or intensified

AI can remove searching, transcription, repetitive drafting and formatting. It can also create a new chain of work:

frame → prompt → inspect → verify → reconcile with context → correct → document → decide → remain accountable

That chain hides real burdens: verifying sources, monitoring for plausible-but-wrong output, handling exceptions, switching tools, reconstructing reasoning you did not do yourself, and staying vigilant when errors are rare but serious. Across more than 100 experiments, human–AI combinations did not automatically beat the better of the human or the AI alone.11

This produces an efficiency paradox: each micro-action gets faster, throughput targets rise, more cases run in parallel — and the worker experiences more mental demand, not less. The saved time is real; so is the rebound work.

Who captures the gain?

RecipientHow value can reach them
EmployeesShorter hours, higher pay, less intensity, learning time, safer or more meaningful work
Employer / shareholdersMore output, lower labour cost, margin, faster scaling
CustomersLower prices, shorter waits, better quality, broader access
Technology providersSubscription and usage revenue, data and ecosystem dependence
Contractors and platform workersMore tasks — or lower rates, more monitoring and unpaid waiting
Society and environmentBetter services and innovation — or displacement, inequality and energy, water and infrastructure burdens

A productivity gain can be operationally real and still distributionally regressive.

Try the Balance Sheet on one deployment

An exploratory exercise — not a validated audit. Your answers stay in this browser tab and are not stored or sent anywhere.

Since this deployment began, how has each dimension changed?

Economic & operational
Quality-adjusted output
Errors, rework and verification
Distribution of financial benefits
Implementation and running costs
Human & professional
Cognitive workload
Work intensity and pace
Autonomy and professional judgement
Learning and skill retention
Wellbeing and job security
Collective & societal
Customer and public value
Accessibility and fairness
Environmental and infrastructure costs

Rate the dimensions to see a reflection.

The cross-level masking test

For every positive KPI, ask four questions:

  1. Where did the burden go?
  2. Who gained and who paid?
  3. Does the benefit last without the AI?
  4. What happened one level above and one level below?

A task-level improvement should be called conditionally positive, not successful, while any material dimension is negative, unknown or carried by a less powerful group. That is not the same as calling every trade-off unacceptable: judge harms by their severity, distribution, reversibility and context. Never collapse unlike dimensions into a single weighted score.

Decision rules for leaders

Reinvest saved time in

  • Recovery or shorter hours when intensity, after-hours work or burnout was already high.
  • Learning and supervised practice when AI removes apprenticeship tasks or unaided skill declines.
  • Service quality and relational work in health, education, social services and customer care.
  • Innovation and experimentation when routine work falls and people have genuine discretion.
  • More output only when quality, workload, autonomy, capability and recovery stay within agreed limits.

Redesign, pause or withdraw when

  • errors or rework erase the time saved
  • people cannot explain, challenge or override consequential recommendations
  • override exists but is routinely penalised
  • pace rises without recovery, staffing or pay
  • assisted performance rises while unaided capability falls
  • monitoring exceeds what service or safety needs
  • benefits concentrate centrally while workers carry the costs
  • high-stakes quality cannot be independently verified

Who owns which part

Boards and executives
Approve a value-allocation policy, risk appetite and stop criteria; review human outcomes alongside return on investment.
HR and people leaders
Own job quality, capability, wellbeing and transition — not only adoption and training completion.
Product and AI teams
Instrument verification, corrections, overrides and failure recovery; design for contestability.
Operational managers
Prevent target ratcheting by default; protect breaks, discretion, mentoring and escalation.
Finance
Report quality-adjusted net benefit after implementation, checking, training and human-impact costs.
Worker representatives
Take part before deployment, agree acceptable monitoring and review how gains are shared.

Whoever captures the efficiency gain should not be able to push verification, professional, health, transition or environmental costs onto others without disclosure and governance. This is the Accountability Chain applied to value.

Measuring without surveilling

Human-impact measurement can itself become surveillance. Use purpose limitation, data minimisation, aggregation, short retention, independent governance and worker participation. Do not infer mental states from behavioural traces, and never turn wellbeing surveys into performance scores. A practical rhythm is a baseline, a pilot, reviews at 30 and 90 days, then at six and twelve months — adapted to the risk and context of each deployment.

The evidence itself has limits: short time horizons, volunteer and adoption bias, fast-changing models, single-firm settings, inconsistent definitions and self-report error. The gap between perceived and measured speed in the developer trial is a reminder that subjective productivity should never stand alone.5

Conclusion

AI can reduce drudgery, spread expertise, improve service, speed up learning and ease burnout. It can also intensify work, automate managerial control, erode judgement, narrow jobs, lengthen the working day and shift returns away from the people carrying the burden. Both are credible, because outcomes are produced jointly by technology, task, organisational design, power and governance.

The Human Impact Balance Sheet should not sit underneath the productivity business case as an ethical appendix. It is the business case — completed.

Method and limits

This is a research-led synthesis of experimental, field, natural-experiment and observational research, prepared in October 2026. It is not a systematic review. I checked each source below against its published record or abstract. The Human Impact Balance Sheet, the masking test, the decision rules and the self-check are my own proposals.

Sources

  1. Experiment Noy, S. & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192. doi:10.1126/science.adh2586. ↩
  2. Field study Brynjolfsson, E., Li, D. & Raymond, L. (2025). Generative AI at work. Quarterly Journal of Economics, 140(2), 889–942. doi:10.1093/qje/qjae044. ↩
  3. Experiment Dell'Acqua, F. et al. (2026). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Organization Science, 37(2). doi:10.1287/orsc.2025.21838. ↩
  4. Field experiments Cui, Z. K., Demirer, M., Jaffe, S. et al. (2026). The effects of generative AI on high-skilled work: Evidence from three field experiments with software developers. Management Science. doi:10.1287/mnsc.2025.00535. ↩
  5. Randomised trial Becker, J., Rush, N. et al. (METR, 2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. arXiv:2507.09089. ↩
  6. Natural experiment Generative AI for analysts (2025). arXiv:2512.19705. Working paper. ↩
  7. Observational Jiang, W., Park, J., Xiao, R. J. & Zhang, S. (2025). AI and the extended workday: Productivity, contracting efficiency, and distribution of rents. NBER Working Paper 33536. nber.org. ↩
  8. Experiment Dratsch, T. et al. (2023). Automation bias in mammography: The impact of artificial intelligence BI-RADS suggestions on reader performance. Radiology, 307(4). doi:10.1148/radiol.222176. ↩
  9. Observational Budzyń, K. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. Lancet Gastroenterology & Hepatology, 10(10), 896–903. doi:10.1016/S2468-1253(25)00133-5. ↩
  10. Randomised trial Roy, P., Poet, H., Staunton, R., Aston, K. & Thomas, D. (2024). ChatGPT in lesson preparation — A Teacher Choices Trial. NFER / Education Endowment Foundation. nfer.ac.uk. ↩
  11. Meta-analysis Vaccaro, M., Almaatouq, A. & Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour, 8(12), 2293–2303. doi:10.1038/s41562-024-02024-1. ↩