Fiete Gehrisch, Kürsat Kirkgöz, Antonie Willner, Faik G Uzunoglu, Marianne Sinn, Jan Bardenhagen, Mara R Goetz, Andreas Brandl, Felix Nickel, Anna Nießen, Thilo Hackert, Thilo Welsch
The GPT-5 configuration combining MDT-style collaboration with population-level in-context learning showed promising discrimination for PNC prediction after ATAAD surgery. However, the MDT layer did not produce a statistically significant incremental improvement in AUC over the corresponding single-agent settings, and its predictive contribution remains to be confirmed. The framework generated structured, role-specific rationales, supporting its further evaluation as a proof-of-concept approach for postoperative risk stratification. Larger multicenter studies are required before routine clinical implementation.
BACKGROUND: Large language models (LLMs) such as GPT-4 are being evaluated for their use as supportive tools in oncological treatment planning. However, in pancreatic cancer, current studies are confined to predefined question-answer formats, while studies specifically investigating real-world scenarios that benchmark LLM performance against multidisciplinary tumor board (MDT) decisions are lacking.
METHODS: This prospective comparative analysis evaluated treatment and diagnostic recommendations for patients with newly diagnosed or suspected pancreatic cancer between an MDT and GPT-4. Using MDT referrals, clinical data were entered into a clinical data matrix and submitted to GPT-4 for therapeutic and diagnostic recommendations. Outputs were assessed before and after additional prompting with 41 high-ranking abstracts relevant to pancreatic cancer care. The primary endpoint was the concordance of recommendations between the MDT and GPT-4 before and after literature-based prompting.
RESULTS: Between September 1, 2024 and March 31, 2025, 45 patients were enrolled. The overall concordance rate between the MDT and GPT-4 was 73.3% (κ = 0.64, p < 0.0001) and did not improve following literature prompting. Discordance most often occurred in complex clinical scenarios. Concordance was highest in cases of metastatic disease (90.0%) and in neoadjuvant settings (90.0%) while it was lowest in patients requiring additional diagnostic workup (50.0%).
CONCLUSIONS: GPT-4 demonstrated substantial agreement with MDT recommendations in patients with newly diagnosed or suspected pancreatic cancer. However, specific abstract prompting did not enhance the rate of concordance and GPT-4's limitations in individualized or complex contexts underscore the need for a cautious future integration into oncologic workflows.