Tarek Souaid, Alexander J Ryu, Donna K Lawson, Lindsay Norgaard, Jennifer M Hovell, M Caroline Burton, Shant Ayanian
Background Large language models (LLMs) are entering clinical practice, but real-world use in hospital medicine is not well described. Methods We conducted a two-month pilot of a general-purpose LLM application in a hospital medicine division. Two cross-sectional anonymous surveys were administered - one at pilot initiation and one at pilot conclusion. Results Response rates were 71.4% (initial) and 48.6% (end of pilot). Reported use was high at both timepoints (18/20 [90.0%] vs 15/17 [88.2%]); median prompts per workday were 2.0 (IQR, 1.0-6.3) and 1.5 (IQR, 1.0-2.0), respectively. Net Promoter Score was +20.0 at initiation and +11.8 at pilot conclusion. Early use emphasized clinical reasoning and administrative tasks; later use shifted toward documentation, summarization, and research work. Efficiency and time savings were the dominant perceived benefit, whereas hallucinations, source transparency, and privacy concerns were the main risks. Conclusions General-purpose LLM applications may offer near-term value as supervised tools for lower-risk workflows, but clinical implementation should include structured governance, training, and ongoing evaluation.