Responsible AI & Human Agency · Research · 2026Verantwoorde AI & menselijke regie · Onderzoek · 2026
What Makes a Relationship With AI Trustworthy?
Why responsible AI requires appropriate reliance — not more trust.
This research page is currently available in English only. Research position as of 3 October 2026 — governance analysis, not legal advice.Deze onderzoekspagina is op dit moment alleen in het Engels beschikbaar. Onderzoekspositie per 3 oktober 2026 — governance-analyse, geen juridisch advies.
High trust is not the goal
High trust in AI is not, in itself, a sign of responsible AI.
If a system is unreliable, confidence in it is dangerous. If a system is demonstrably reliable but people reject it out of habit, distrust can cause harm too. So the goal is not to maximise trust. It is to make human reliance proportionate to what the system has actually shown it can do.
Many organisations treat trust as an adoption target: a satisfaction score to raise. This study treats it as a governance outcome. A trustworthy relationship with AI is one in which people know when to rely, when to question, when to stop — and who remains answerable when the system is wrong.
The entry gate asks whether AI should be used at all. This study asks about the quality of the relationship once AI is in the room. It sits between human capability (can I still think and act independently?) and meaningful oversight (can I actually intervene when I need to?).
Six words that are often confused
| Concept | Meaning |
|---|---|
| Trust | An attitude: the expectation that something will help achieve one's goals in a situation of uncertainty and vulnerability.2 |
| Trustworthiness | The properties and governance that warrant reliance: validity, reliability, safety, security, accountability, transparency, privacy and managed bias, in a defined context.1 |
| Reliance | Behaviour: accepting, acting on or deferring to AI. Reliance can happen without trust — because there is no alternative, no time or no expertise. |
| Appropriate reliance | Reliance that matches the system's demonstrated capability for this task, these stakes and this uncertainty.2 |
| Dependence | A structural inability to work, decide or get a service without the system. It can grow while trustworthiness stays the same. |
| Calibration | Alignment between reliance and actual performance. Both overreliance and under-reliance are failures. |
Under-reliance is real. People can abandon an algorithm after seeing it make a mistake, even when it outperforms them — the phenomenon known as algorithm aversion.3 The aim is not a universal midpoint between trust and distrust but discriminating reliance: accept, verify, reject or escalate, case by case, according to evidence and risk.
The trust gap
The central problem is the distance between how trustworthy a system seems and how trustworthy it has been shown to be.
View the full framework →
Perceived signals are cheap; demonstrated evidence is expensive. Fluent text, a confident tone, citations, a friendly voice and a trusted brand are all easy to produce. Validation in the real population, reliability under change, calibrated uncertainty and an incident record are not.
A widely deployed sepsis prediction model shows the gap clearly. In an external validation across 38,455 hospitalisations, it reached an area under the curve of 0.63, missed 67% of patients with sepsis and alerted on 18% of all hospitalisations.16 Wide adoption had signalled trustworthiness; independent evidence did not support it.
Transparency alone does not close the gap. The NIST AI RMF is explicit that a transparent system is not necessarily an accurate, privacy-enhanced, secure or fair one.1 Certification of a management system, such as ISO/IEC 42001, is a useful organisational signal — but not proof that a particular model is accurate, fair or safe in a particular use.14
What the evidence says about reliance
- Automation bias is well documented. A systematic review of 74 studies found it is shaped by experience, trust and confidence, workload, task complexity and time pressure — and can be reduced by training, emphasising user accountability and interface design.5 Complacency occurs in novices and experts alike and is not overcome by simple practice.4
- Explanations can increase blind acceptance. In studies where AI and humans performed comparably, explanations increased the chance that people accepted the AI's recommendation regardless of whether it was correct.7
- Confidence scores help — but are not enough. Showing a model's confidence helped people calibrate trust, but did not by itself improve joint decisions; that also depended on whether the human brought knowledge the AI lacked.6
- The wording of uncertainty matters. In a pre-registered experiment with 404 participants answering medical questions, first-person uncertainty (“I'm not sure, but…”) reduced agreement with the system and increased accuracy — reducing, but not eliminating, overreliance on wrong answers. More general phrasing had weaker effects.9
- Friction works, but people dislike it. Cognitive forcing designs reduced overreliance compared with simple explanations (N=199), but participants rated them least favourably, and they helped people who enjoy effortful thinking most.8 Designs that improve reliance may lower satisfaction — so satisfaction cannot be the measure.
The practical lessons: show uncertainty only when it corresponds to a validated quantity; connect it to an action (verify, abstain, escalate); in consequential cases, record the human's initial judgment before showing the AI's answer; and measure correct acceptance, correct rejection, harmful acceptance and harmful rejection separately — not “trust” or satisfaction alone. A miscalibrated confidence score is not safer than silence.
Anthropomorphism and relational honesty
Human-like design — names, voices, avatars, “I” language, empathy cues, humour, memory — can make systems more accessible and easier to use. The problem is not human-like interaction itself. It is misrepresentation. People readily attribute minds to non-human agents, especially when they are motivated to understand them or seek social connection.10 Warmth can therefore raise perceived competence and care without any change in actual competence or care.
Principle: social persuasiveness must not exceed epistemic justification. The more human an AI feels, the more carefully we should ask what users are being led to believe about it.
Governance should constrain:
- claims of feelings, intentions, care or understanding that users may take literally;
- emotional reciprocity designed to deepen dependency;
- persuasive warmth that hides uncertainty or commercial incentives;
- memory that is unclear about what is kept, for how long and how to delete it;
- design that shifts moral responsibility from the provider to “the agent”;
- human-like authority in professional or high-stakes advice.
The law already sets a floor. The EU AI Act requires that people are informed when they interact directly with an AI system, unless that is obvious,11 and the Council of Europe Framework Convention asks parties to ensure, as appropriate for the context, that people are told they are interacting with AI rather than a human.12 And when a chatbot is wrong, the organisation cannot hide behind it: in Moffatt v. Air Canada, a tribunal held the airline responsible for incorrect information from its website chatbot.15 “The AI said so” is not an accountability model.
Layered transparency, not maximum disclosure
Transparency has many objects: am I dealing with AI? What is it for — and not for? Where has it been shown to work? What role did it play in this decision? What is uncertain? Who is responsible, and how do I challenge the result? The OECD AI Principles ask for meaningful, context-appropriate information that lets people understand outcomes and challenge them.13
The design rule is to give users the decision-critical facts first, deeper evidence to professionals, auditors and regulators, and to respect legitimate privacy and security limits. Disclosure becomes trust theatre when it is vague, buried, non-actionable or presented as proof of safety. A raw technical dump can overwhelm people and shift responsibility onto them without giving them control.
Six assessment domains
For the framework, I consolidate trust and appropriate reliance into six domains:
| Domain | Core question | Red flag |
|---|---|---|
| Evidence | Has capability actually been demonstrated for this task, population and setting? | Only a vendor benchmark headline |
| Uncertainty | Can people tell where the AI's knowledge ends? | Confident output by default; no tested abstention |
| Reliance | Do people accept correct advice and reject wrong advice? | Override rates collapse; only satisfaction is measured |
| Relational integrity | Does the interface create misleading impressions of competence, agency or care? | Emotional cues that deepen dependency or hide limits |
| Human agency | Can people verify, disagree, opt out and reach a human? | Forced use; inaccessible alternative; no time to check |
| Accountability & remedy | Can errors be challenged, corrected and learned from? | “The AI decided”; apology without control change |
Trust can be appropriate at one level and misplaced at another. A worker may rightly trust a colleague's review process while distrusting the model underneath; a technically reliable system may be deployed by an organisation that discourages overrides. The NIST AI RMF treats AI as sociotechnical for exactly this reason.1 A worker cannot calibrate reliance if the organisation denies the evidence, time, skills, independence or authority needed to disagree.
The trust-calibration test
Not “how much do you trust this AI?”, but what level of reliance does the evidence justify? Answer for one specific use.
An exploratory exercise — not a validated assessment. Your answers stay in this browser tab and are not stored or sent anywhere.
Answer the questions to see which level of reliance your answers support.
Do not rely — evidence inadequate or consequences unacceptable.
Verify independently — useful assistance; consequential claims need confirmation.
Use as decision support — AI contributes; a qualified human forms the judgment.
Rely within defined bounds — validated task, population and conditions, with monitoring and fallback.
Delegate reversible low-risk actions — least privilege, logging, confirmation, undo and rapid revocation.
Greater accuracy does not automatically justify greater autonomy. The more an AI can act, chain actions and scale, the stronger the evidence, the narrower the permissions and the better the containment it needs.
The AI Trust Card
A trustworthy relationship needs a shared, honest description of the system. I propose a short public card for every AI system that people rely on:
- What this AI does
- Its intended role.
- What it does not do
- Explicit boundaries and unsuitable uses.
- What the evidence shows
- Validated capability and known limitations.
- When not to rely on it
- Known failure conditions.
- What to do when uncertain
- Verify, escalate, or use a human alternative.
- Who is responsible
- A named organisation or function.
- How to challenge an outcome
- Correction, appeal and remedy routes.
- Last evaluated
- Date, system and version.
A proposed format, not an existing standard. It complements, but does not replace, legal transparency duties and technical documentation.
When to act
- Redesign or strengthen warnings when confidence is uncalibrated, provenance is weak, users misread explanations, or human-like cues inflate perceived competence.
- Restrict or review independently when override rates collapse, workers lack real authority, subgroup performance diverges, human alternatives are inaccessible, or adoption figures hide exclusion.
- Suspend or withdraw when serious incidents recur, the system cannot reliably abstain, irreversible actions lack confirmation, audit trails are incomplete, remedies fail — or perceived trustworthiness materially exceeds demonstrated trustworthiness.
Trust repair takes more than an apology: timely disclosure, preserved evidence, correction and remedy, independent investigation where warranted, governance change, retesting and public evidence of improvement. That is the territory of continuous assurance.
What this asks of different roles
- Boards and executives: govern reliance and its side effects, not adoption alone; require evidence and stop authority.
- Product, engineering and UX: test calibration, abstention, provenance and recoverability; measure correct acceptance and rejection, not satisfaction alone.
- HR and managers: protect questioning, verification time and human fallback; do not reward blind adoption.
- Procurement: require evaluation access, incident duties, logs, audit rights and exit provisions.
- Frontline professionals: keep domain accountability; document disagreement and escalation.
- Providers: make limitations and incentive conflicts visible; accept responsibility for design choices.
- Educators: teach source evaluation, uncertainty and epistemic responsibility — not only prompting skills.
Conclusion
Trust should never substitute for evidence. Compliance sets a floor; a trustworthy relationship also depends on proportionality, human agency, inclusion and identifiable accountability.
The most trustworthy relationship with AI may involve confidence, cautious support, independent verification, deliberate friction — or refusal. Its quality is shown not by how smoothly people defer to AI, but by whether they can know when to rely, when to question, when to stop, and who remains answerable when the system is wrong.
Method and limits
This is a research-led synthesis of human-factors and human-computer-interaction research, legislation, standards and case law, prepared in October 2026. Many reliance studies use short laboratory tasks and convenience samples; longitudinal effects of dependency, organisational power and cultural context are under-researched. Several recent studies in my working research could not be matched to a verifiable public record and were replaced by established, peer-reviewed work on the same questions. Provider self-reports were not used as evidence. The framework, six domains, calibration test, AI Trust Card and escalation rules are my own proposals, not validated instruments. I checked each source below against its official or published record.
Sources
- Framework NIST (2023). AI Risk Management Framework (AI RMF 1.0). doi:10.6028/NIST.AI.100-1. ↩
- Peer-reviewed Lee, J. D. & See, K. A. (2004). Trust in Automation: Designing for Appropriate Reliance. Human Factors, 46(1), 50–80. doi:10.1518/hfes.46.1.50_30392. ↩
- Experimental Dietvorst, B. J., Simmons, J. P. & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126. doi:10.1037/xge0000033. ↩
- Review Parasuraman, R. & Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors, 52(3), 381–410. doi:10.1177/0018720810376055. ↩
- Systematic review Goddard, K., Roudsari, A. & Wyatt, J. C. (2012). Automation bias: a systematic review of frequency, effect mediators, and mitigators. JAMIA, 19(1), 121–127. doi:10.1136/amiajnl-2011-000089. ↩
- Experimental Zhang, Y., Liao, Q. V. & Bellamy, R. K. E. (2020). Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. Proceedings of FAT* 2020, 295–305. doi:10.1145/3351095.3372852. ↩
- Experimental Bansal, G. et al. (2021). Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance. Proceedings of CHI 2021. doi:10.1145/3411764.3445717. ↩
- Experimental Buçinca, Z., Malaya, M. B. & Gajos, K. Z. (2021). To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1). doi:10.1145/3449287. ↩
- Pre-registered experiment Kim, S. S. Y. et al. (2024). “I'm Not Sure, But…”: Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust. Proceedings of FAccT 2024, 822–835. doi:10.1145/3630106.3658941. ↩
- Theory Epley, N., Waytz, A. & Cacioppo, J. T. (2007). On seeing human: A three-factor theory of anthropomorphism. Psychological Review, 114(4), 864–886. doi:10.1037/0033-295X.114.4.864. ↩
- Legislation Regulation (EU) 2024/1689 (AI Act), Article 50 — transparency obligations. artificialintelligenceact.eu. ↩
- Treaty Council of Europe Framework Convention on Artificial Intelligence (CETS No. 225), Article 15. rm.coe.int (PDF). ↩
- International principles OECD. AI Principles — Principle 1.3, Transparency and explainability. oecd.ai. ↩
- Standard ISO/IEC 42001:2023, Artificial intelligence — Management system. iso.org. ↩
- Case law Moffatt v. Air Canada, 2024 BCCRT 149. canlii.org. ↩
- External validation Wong, A. et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine. doi:10.1001/jamainternmed.2021.2626. ↩