Adam Rodman, Laura Zwaan
Numerous studies have now shown that large language models (LLMs) have remarkable similarities to physician cognition in silico, showing both human-like—or even superhuman—performance on cognitive medical tasks as well as reflecting human biases.1–6 Reasoning models, which were publicly introduced in September 2024, have since become ubiquitous in commercial products such as ChatGPT (GPT-5, OpenAI) and Gemini (Gemini 3, Google). They use chain-of-thought processing during model inference, allowing them to decompose complex clinical scenarios into intermediate steps, verify logic, correct errors prior to providing output and drastically increase their performance on a variety of cognitive tasks. The recent piece by Wang and Redelmeier in this issue of BMJ Quality and Safety shows that these powerful models, similar to the base models that preceded them, continue to show human cognitive biases when tested on clinical vignettes.7
This finding is interesting but unsurprising since the models inherit both the high performance and the cognitive biases of clinicians. [...]
In conclusion, it is not surprising that LLMs have biases, nor is it clear that these are a drawback that needs to be mitigated. More fundamentally, we must move to a system where we explicitly study human-AI collaboration. How do we optimise interaction to trigger critical thinking? Where should LLMs be inserted in clinical workflows? And most fundamentally, does the human-AI dyad perform better than the human alone, even with an imperfect AI? LLMs are fundamentally a new class of technology in clinical decision support. We cannot evaluate them with the tools and frameworks of the past. We must embrace their messy, human-like nature and learn to reason with them. Only then will we use this technology to better care for our patients.