This is not another entry in the AI debate. We are not going to argue whether AI is good or bad, safe or dangerous, the future of medicine or its undoing. That framing is exhausted. AI in healthcare is inevitable. The more pressing question is epistemological.
Humans are exquisitely vulnerable to tone. Decades of communication research show that people trust confidence and clarity independent of accuracy. Medicine has a long history of ideas delivered with enormous confidence before the evidence caught up. Persuasion often shapes belief before evidence does. And for this conversation, the most important point is that empathy amplifies trust.
Empathy is a trust-amplifier. When a system responds with warmth, patience, and apparent understanding, it does not just feel helpful. It feels safe. It feels like a clinician who listens. When that feeling is generated by a probabilistic algorithm rather than earned through clinical judgment, it becomes a powerful vector for misplaced trust in medicine. That risk is amplified in a healthcare system where real human presence has already been eroded by rushed visits, documentation burden, and the constant pull of the screen. Patients who feel heard may be less likely to question the answer. Clinicians who receive a confident, well-framed recommendation may be less likely to verify it. The more human AI sounds, the more human-level trust it extracts, without human-level accountability.
Join our private community to explore publishing, real estate, wealth-building, creative projects, & everything that keeps us curious, ambitious, and a little more alive.
The question is not whether AI can be useful. Of course it can, and it is. The deeper issue centers on what happens to clinical judgment, patient trust, and institutional accountability when the most persuasive communicator in the room is a system that does not know what it does not know, but sounds as if it does.
AI can be wrong without looking wrong. In fact, it can look spectacularly right: precisely formatted, tonally authoritative, and fluent in the exact register that clinical documentation demands. When errors present as competence, clinicians lose the cognitive signal that normally triggers skepticism.
This is not a hypothetical. AI systems consistently generate outputs that are internally coherent, well-structured, and confidently asserted, but factually wrong in ways that may only become visible on independent verification. Ask five different leading models the same clinical question with slight variations in phrasing, and you may get different answers, different citations, and completely different clinical pathways. Those answers are presented without signals of uncertainty and arrive in the same confident formatting as correct ones.
This is what one of us (Bloomgarden) found when experimenting directly with these models: egregiously wrong answers that sounded completely believable. These include plausible, well-packaged, clinical-sounding errors rather than subtle ones. That is a different category of risk than what we usually manage with human error, because human errors usually carry some visible marker of uncertainty.
Bring the idea. We’ll help build the show. From planning and production to polish and launch, we turn expertise into content that looks professional, feels intentional, & gives your audience a reason to return.
The bigger issue is that hallucinations are not a simple bug to be patched. They arise from how these systems generate answers. When a model looks “95% accurate,” that can be deeply misleading. High performance in a complex, changing clinical environment can reflect false precision, dataset leakage, distributional shift, or pattern recognition that does not hold up in real life. Models can look impressive while learning the wrong thing very well.
High confidence and high accuracy can coexist with clinical unsafety. They are not the same dimension.
The EMR rollout should be a cautionary reference for every healthcare AI deployment conversation. We introduced a transformative technology rapidly, at scale, with enormous confidence in its promised efficiency and too little attention to its downstream cognitive, relational, and safety costs. Clinician burnout, documentation burden, alert fatigue, and degraded patient communication emerged from mistaking a smooth implementation for a safe one, rather than from bad intentions.
We are repeating that pattern. In a single week, multiple AI tools can enter a health system’s workflow, each with its own logic, training data, and failure modes. What is missing is standardization, a guardrails-first approach, and any systematic accounting for what happens when these tools begin shaping documentation, along with clinical interpretation, communication patterns, and judgment itself.
The risk is epistemic drift: a gradual, often invisible shift in how a health system reasons about clinical situations, driven by unreliable probabilistic systems embedded in daily workflow. This is not science fiction. It is the predictable consequence of deploying these tools into high-stakes domains before asking hard questions.
The question of how AI systems should be designed is not a product question. It is a clinical safety question. That requires at least four things:
- Uncertainty first. Blanket refusal is too blunt. Systems should distinguish between low-stakes informational questions and moments that genuinely require escalation to a human. The ability to say “this needs a clinician” carries real value.
- Real citations. Citations matter when they are real, relevant, and available for clinical review. A hallucinated reference with a plausible-sounding journal name adds false authority to a fabricated claim.
- Provenance by default. Every clinical AI output should carry information about what it was trained on, when, and under what conditions. In medicine, knowing where an answer came from is part of how we decide whether to trust it.
- Validation in the actual population. A model validated on English-language academic medical data may tell you very little about how it will behave in a multilingual, under-resourced, real-world clinical environment. Benchmark performance alone does not establish readiness for deployment.
This article is ultimately about trust. Trust in the clinical sense shapes how we decide what to believe and what to act on. Medicine has always had an accountability structure for that process. A clinician who gives confident wrong answers can be questioned, retrained, and held responsible. The system is imperfect, but it exists.
AI has none of that accountability structure in its current deployment context. It presents with the tone of authority without the accountability of authority. It sounds like a trusted consultant and does not know when it is wrong.
AI is moving fast. Our course helps healthcare professionals keep up, with expert-led conversations on advocacy, automation, women leading in AI, ethics, regulation, and the questions medicine can’t afford to ignore.
Which brings us back to design. AI does not have humility the way a person does. It has to be built to show uncertainty. What we call humility in AI is a design choice: calibrated uncertainty, transparent limits, source attribution, and the ability to say, “I am not confident enough to answer that.”
The solution is not fear of AI. It is a more rigorous standard for what we allow into clinical decision-making, one that demands transparency about failure modes, genuine expressions of uncertainty, real-world validation in the populations where the tool will be deployed, and institutional accountability when the tool gets it wrong.
Machine humility requires deliberate design by clinicians, health systems, and developers. Until that happens, every physician using these tools carries the burden of the question AI will never ask itself: is this answer true, or does it just sound true?
Every clinical AI output should carry information about what it was trained on, when, and under what conditions. In medicine, knowing where an answer came from is part of how we decide whether to trust it.
article written by Hassan Bencheqroun, MD MBA and Eve Bloomgarden, MD Tweet This Quote!








