科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Laryngoscope investigative otolaryngology2026-08-01

Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects.

James Anibal, Geetha Krishna Chaitanya Nama, Shrramana Ganesh, Yosef Nafii, Samantha Salvi Cruz, Veronica Daoud, Madeleine Zanin, Tram Le, Parsa Khorrami, Bradford J Wood, David Clifton, Jamie Toghranegar, Yael Bensoussan, Yael E Bensoussan, Olivier Elemento, Jean-Christophe Bélisle-Pipon, David Dorr, Satrajit Ghosh, Alistair Johnson, Phillip Payne, Maria E Powell, Anaïs Rameau, Vardit Ravitsky, Alexandros Sigaras, Shaheen Awan, Ruth Bahr, Donald Bolser, Kathy J Jenkins, Frank Rudzicz, Jennifer Siu, Stephanie Watts, Yassmeen Abdel-Aty, Toufeeq Ahmed Syed, James Anibal, Steven Bedrick, Isaac Bevers, Micah Boyer, Rahul Brito, Selina A Casalino, John Costello, Enrique Diaz-Ocampo, Mahmoud Elmahdy, Kenneth Fletcher, Alexander Gelbard, Karim Hanna, Bill Hersh, Lochana Jayachandran, Kaley Jenney, Andrea Krussel, Chloe Loewith, Tempestt Neal, Claire Premi-Bortolotto, Sarah Rohde, Samantha Salvi Cruz, Elizabeth Silberholz, Duncan Sutherland, Venkata Swarna Mukhi Talluri, Jamie Toghranegar, Kimberly Vinson, Claire Wilson, Madeleine Zanin, Theresa Zesiewicz, Robin Zhao

一句话结论 · In one sentence

LLMs may provide useful, explainable recommendations when presented with dataset distribution statistics and candidate profiles. In the future, this simulated scenario may be extended to align with conditions in emergency departments or other high-volume settings.

原始摘要(英文原文)· Original abstract
OBJECTIVES: Past studies have shown that many clinical machine learning models have performance limitations due to imbalances in the training data. For voice data generation projects, the origin of the problem may lie in the recruiting methods used during data collection efforts. This study introduces a generative AI pipeline for "dataset decision support", recommending recruitment decisions based on high-dimensional insights. METHODS: The publicly available GOSSIS-1-eICU dataset was filtered to create patient populations that were relevant to voice data generation projects. Lab results and vital signs from the electronic health record were also used to train a neural network for prediction of disease type. Prediction uncertainty estimates were included in the dataset as approximate indicators of health complexity. To select the best recruitment choice for addressing imbalances, an open-source large language model (LLM) was then instructed to assess dataset statistics and the characteristics of possible participants. Simulations were run in which the system constructed datasets of 250 patients. RESULTS: In over 90% of cases, the proposed system reduced categorical imbalances and widened continuous distributions when compared to randomly sampled counterfactual datasets (q-value < 0.05). Variables included race, age, BMI, sex, disease type, oxygenation status, co-morbidities, post-operative status, Glasgow Coma Scale verbal response score, the Acute Physiology Score III, prediction uncertainty, vital signs, and lab results. CONCLUSION: LLMs may provide useful, explainable recommendations when presented with dataset distribution statistics and candidate profiles. In the future, this simulated scenario may be extended to align with conditions in emergency departments or other high-volume settings. LEVEL OF EVIDENCE: 3.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects. — 科研速览 Science Skim