Science of Skin Summit: AI in Medicine Requires More Than Knowing Which Model Scores Best

Key Takeaways

  • Large language models can generate sophisticated medical answers, but strong performance on static benchmarks such as the USMLE does not necessarily demonstrate the conversational reasoning and information-gathering skills required in real clinical care.
  • Physicians using AI should watch for automation bias, sycophancy, cognitive offloading, embedded biases, and confidently presented errors, particularly when an answer is mostly correct and the remaining mistake may be difficult to identify.
  • Clinicians should use LLMs for tasks whose outputs can be efficiently verified and should not enter protected health information into public AI models that are not configured for compliant handling of patient data.
09/26/2026

Artificial intelligence models are evolving so rapidly that clinicians may struggle to determine which tools they can trust, but choosing the model that performs best on a standardized test may not provide the answer.

Speaking at the Science of Skin Summit, Roxana Daneshjou, MD, PhD, discussed how large language models (LLMs) developed, how they generate responses, and why their apparent proficiency can obscure limitations that become particularly important in medicine. She concluded with practical rules for clinicians using AI, including protecting patient information and watching for automation bias, sycophancy, and cognitive offloading.

How Large Language Models Learned to Speak

Dr. Daneshjou traced the current AI era in part to the 2017 publication of Attention Is All You Need, which introduced the transformer architecture.

Earlier natural language processing approaches had difficulty connecting information separated by substantial amounts of text. Transformer architecture enabled models to identify longer-range relationships, helping lay the groundwork for the LLMs that subsequently captured public attention.

Other developments included dramatically increasing the volume of training data and computational resources devoted to models.

At a simplified level, Dr. Daneshjou explained, an LLM learns statistical relationships among pieces of language by repeatedly predicting missing or subsequent tokens across enormous quantities of human-generated text. Further development can include reinforcement learning from human feedback, through which responses preferred by humans can influence model behavior.

Prompting adds another layer. The information and instructions provided to an LLM can significantly affect the answer it produces, while developers may use system-level instructions that users never see.

The result is a system capable of producing remarkably natural language—but not necessarily one that understands medicine in the way a physician does.

Human Bias Can Become AI Bias

Dr. Daneshjou, whose research has included AI fairness, emphasized that models trained on human-generated information can inherit undesirable human associations.

She pointed to research demonstrating associations between demographic groups and positive or negative descriptors within written language. Similar associations can appear with gender, such as models statistically associating certain occupations or family roles with men or women.

When LLMs learn from enormous collections of human language, these patterns can become embedded in the relationships the models learn.

That creates a fundamental issue for medicine: increasing the scale of a model does not inherently eliminate biases contained in its training data.

Why Passing the USMLE Is Not Enough

Dr. Daneshjou also questioned the way AI models are commonly evaluated for medical use.

Static question-and-answer benchmarks, including performance on examinations such as the US Medical Licensing Examination (USMLE), can demonstrate a model's ability to answer medical questions but may reveal much less about its ability to practice medicine.

A clinical vignette generally provides all relevant information in a carefully structured format. Patients do not.

“No patient comes in like that,” Dr. Daneshjou said.

Real clinical encounters require physicians to determine what information is relevant, ask appropriate follow-up questions, reconcile incomplete or conflicting information, and reason toward a diagnosis and management plan.

Dr. Daneshjou described research in which investigators created separate physician and patient AI models and required the physician model to obtain the information it needed through conversation rather than receiving a fully packaged vignette. Performance deteriorated when the model had to determine which questions to ask and conduct the clinical reasoning process itself.

That distinction, she argued, demonstrates why AI evaluation should be task specific. The ability to summarize a record, converse with a patient, assist with billing, or reason through a clinical presentation represents different capabilities that should be evaluated independently.

Can Physicians Determine Which Medical AI Is Best?

Studies comparing general-purpose LLMs with medical-specific systems have reached differing conclusions, Dr. Daneshjou noted, but she cautioned against interpreting any single benchmark as establishing which system is superior.

Datasets may already be represented in a general-purpose model's training data, potentially affecting results. Meanwhile, evaluations performed by companies developing the systems themselves introduce other considerations when interpreting findings.

“I don't think both these studies have a grain of truth, and both these studies are wrong,” Dr. Daneshjou said of two evaluations reaching contrasting conclusions.

The broader problem is that researchers have not yet established a single reliable way to evaluate an AI model's overall clinical capability. Instead, models should be evaluated according to the particular task clinicians expect them to perform.

The 95% Problem

One of the most important risks may arise not when an AI answer is obviously wrong, but when it is almost entirely correct.

Dr. Daneshjou demonstrated an AI-generated calculation of a clinical score in which nearly every component was accurate except one.

“When it's 95% right, are you gonna be able to quickly find what's wrong?” she asked.

Her rule is to use LLMs primarily for tasks in which she can quickly verify the result. If confirming the answer requires independently checking every factual statement, the efficiency advantage can disappear.

That problem is compounded by automation bias—the tendency to trust the output of an automated system. Dr. Daneshjou compared it with following GPS directions onto an obviously incorrect route simply because the navigation system provided the instruction.

Sycophancy, Cognitive Offloading, and Patient Privacy

Dr. Daneshjou highlighted several additional risks.

Sycophancy occurs when an AI system responds in ways that align with or please the user, potentially reinforcing an incorrect assumption rather than challenging it. In medicine, that tendency can become particularly consequential when a clinician approaches a model with an incorrect hypothesis.

Cognitive offloading presents another concern. As clinicians delegate more intellectual tasks to AI, excessive reliance on the technology could potentially affect the skills required to perform those tasks independently.

Dr. Daneshjou also issued a straightforward warning regarding privacy: physicians should not enter protected health information into public AI models. Systems specifically configured for compliant handling of patient information, such as appropriate enterprise or integrated clinical implementations, require a different consideration than consumer-facing public models.

Her overarching message was not that physicians should avoid LLMs. Rather, clinicians need to understand both what these models do well and the ways in which convincing outputs can conceal weaknesses.

AI can generate fluent, sophisticated, and mostly accurate responses. For physicians, however, “mostly accurate” may be precisely where vigilance becomes most important.

Register

We're glad to see you're enjoying PracticalDermatology…
but how about a more personalized experience?

Register for free