What AI actually changes in a doctor visit
Most writing about AI in medicine is either a promise that diagnosis is about to be automated or a warning that it is dangerous. The reality in clinics right now is narrower and more useful than either: the technology is very good at a few specific jobs, unreliable at others, and its clearest current benefit has almost nothing to do with diagnosis.
The documentation problem it actually solves
Physicians spend a large share of the working day on documentation — often comparable to the time spent with patients, with more of it spilling into the evening. Electronic records were supposed to reduce that and mostly did the opposite, because they were designed around billing and compliance rather than around clinical thinking.
Ambient documentation tools listen to a visit, with consent, and draft the note. The physician reviews and corrects it. This is unglamorous and it is the change patients are most likely to notice, because it determines whether the person across from them is looking at a screen or at them.
That matters more than it sounds. A physician typing while you describe a symptom is a physician partly absent for it, and the details people volunteer tend to arrive in the gaps of a conversation rather than in answer to a direct question.
Where the technology is genuinely strong
Pattern recognition in images. Diabetic retinopathy screening, mammography support, and detection tasks in radiology and dermatology have produced results in defined settings good enough for regulatory clearance. These systems do one narrow task on one image type, which is exactly the shape of problem this technology is suited to.
Surfacing what a person would have to remember. Drug interaction checks, overdue screening flags, and abnormal trends buried across years of results. This is not intelligence so much as attention that does not get tired, applied to data volumes no one can hold in their head.
Drafting and summarizing. Condensing a long record before a visit, drafting a referral letter, turning a specialist note into plain language. Bounded tasks with a human reviewing the output.
Where it is not reliable
Producing confident wrong answers. Language models generate plausible text, and plausibility and accuracy are different properties. A fabricated citation or a subtly wrong dose is more dangerous than an obvious error precisely because it reads correctly. Any clinical use requires the output to be checked by someone who could have written it.
Performing outside its training distribution. A model validated on one population can degrade on another. Because these systems fail quietly rather than loudly, degradation is hard to notice without deliberate monitoring.
Inheriting bias from historical data. Models trained on past clinical data learn past patterns, including the inequitable ones. This is a well-documented problem and an active area of work.
Knowing what is not being said. A large amount of clinical information arrives as hesitation, or the symptom mentioned while standing up to leave, or the mismatch between what someone reports and how they look. None of it is in the transcript.
What it does not change
The structure of a clinical encounter is not primarily an information-processing problem. Deciding whether to treat, weighing a small risk against a real side effect, and understanding what a particular person is willing to live with are judgment calls made with someone, not for them.
Nor does it change accountability. A physician who accepts an AI-generated recommendation owns that decision entirely. Tools that make it easy to accept output without engaging with it are worse than no tool, and that is a design question, not a technical one.
Questions worth asking about it
If a practice uses these tools, a few things are reasonable to ask:
- Is a recording made, and is it retained? Ambient documentation involves capturing a conversation. Whether audio is stored, for how long, and who can access it are answerable questions.
- Is there a business associate agreement with the vendor? Any company handling identifiable health information on a practice’s behalf needs one. This is a requirement, not a courtesy.
- Can you decline? Consent to ambient documentation should be real, which means declining should be straightforward and should not change your care.
- Who reviews the output? The answer should be a person, before anything enters the record.
The realistic version
The near-term effect of this technology in primary care is less about diagnosis and more about removing work that never needed a physician doing it. That is a smaller claim than the field usually makes, and it is worth more than it sounds: attention is the scarce resource in a clinical visit, and most of what has eaten it over the past fifteen years was administrative.
Used well, these tools give some of that back. Used carelessly, they add a confident, tireless source of errors to a system that already struggles to catch them.
The same standard is worth applying to any clinical claim, technological or not: what outcome was measured, and compared to what. That question is what separates useful screening from testing that generates more anxiety than information, and it is the one worth asking of any medication being considered long term.