H Rugo, A Forsythe, D Flora, S Glück, S Grieve, R Campden, R J Liu, M Musat, J M Rege, P A Kaufman, R L Coleman, P Tarantino, J O'Shaughnessy, L Schwartzberg
A human-conducted, AI-augmented living SLR integrated with guidelines and regulatory data can provide real-time evidence support. Future studies are required to evaluate its impact on physician workflows, clinical decision-making, and implementation in oncology practice.
BACKGROUND: Oncologists face increasing difficulty staying current with rapidly evolving clinical data, guidelines, and regulatory updates. Building and maintaining an annotated clinical trial evidence library is time- and labor-intensive. To address this challenge, we developed and validated a living oncology evidence platform (Living-OEP) for breast cancer (BC), powered by an agentic artificial intelligence (AI) system that supports daily, human-conducted, AI-augmented systematic literature review (SLR).
MATERIALS AND METHODS: The agentic AI system, incorporating GPT-4.1 and o3 (OpenAI) and Claude [Anthropic, PBC, San Francisco, CA] Sonnet-4 (Anthropic), was designed to emulate expert-led, Cochrane-compliant SLR workflows. Guided by a human-developed annotation manual, the system decomposes tasks, self-debugs, and validates outputs. Training data included 29 236 clinical trial abstracts across BC, lung, and prostate cancer, each annotated with four review and 32 extraction variables. Structured data were integrated with guideline-based treatment pathways, forming a real-time, evidence-linked OEP. Accuracy was benchmarked against 1997 human annotations. Living-OEP evidence quality was compared with AI chatbots (ChatGPT [OpenAI, San Francisco, CA], Perplexity [Perplexity AI, Inc., San Francisco, CA], Consensus [Consensus, Boston, MA], and OpenEvidence [OpenEvidence Inc., Miami, FL]) using six criteria across eight BC treatment scenarios.
RESULTS: The agentic AI review accuracy ranged from 95.1% to 97.2%. Extraction accuracy ranged from 51.1% to 99.4%, with three variables still undergoing iterative refinement. Compared with other AI tools, the Living-OEP provided more comprehensive and accurate evidence and linked all data to original publications and Food and Drug Administration labels.
CONCLUSIONS: A human-conducted, AI-augmented living SLR integrated with guidelines and regulatory data can provide real-time evidence support. Future studies are required to evaluate its impact on physician workflows, clinical decision-making, and implementation in oncology practice.