J. Creeden, M. Olivecrona, N. Benavent, R. Hladun Alvaro, L. Valero-Arrese, A. Carolina Torres, A. Soriano
Background. Comprehensive genomic reports in oncology contain complex molecular information that must be translated into concise summaries for treating clinicians. Manual summarisation is time-intensive, and robust evidence supporting AI-assisted summarisation systems remains limited. We evaluated whether AI-generated summaries were non-inferior to manually produced summaries in faithfully representing information from tertiary paediatric cancer genomics reports. Methods. QNOMX-VHIR-CPSP-001 Phase 1 was a single-site, non-interventional clinical performance study conducted at Vall d'Hebron Institut de Recerca under ISO 20916:2019. De-identified tertiary paediatric cancer genomics reports were summarised using both the AI-assisted system under evaluation and the standard manual workflow. Three qualified raters assessed both summary types against the source reports using the six-domain, five-point Quality Summary Index (QSI) in a blinded, counterbalanced crossover design with a minimum fourteen-day washout. Co-primary endpoints were composite content and presentation scores. Non-inferiority was assessed using a pre-specified Bayesian hierarchical model with a margin of 0.5 points and confirmed using a frequentist linear mixed model. Results. Thirty-seven of 38 planned cases contributed 74 paired evaluations. Mean content scores were similar between AI-generated and manual summaries (4.01 vs 3.97), whereas AI-generated summaries achieved higher presentation scores (4.39 vs 4.12). Non-inferiority was demonstrated for both co-primary endpoints: content +0.05 (95% CrI -0.18 to +0.28), presentation +0.28 (+0.10 to +0.47). Frequentist analyses were concordant. Accuracy was the only QSI dimension for which Bayesian non-inferiority was not concluded. Inter-rater reliability was low across all QSI domains. Conclusions. AI-generated summaries were non-inferior to manually produced summaries for both content and presentation endpoints in tertiary paediatric oncology genomics reports. Presentation scores were higher for AI-generated summaries, whereas content scores were comparable between approaches. Low inter-rater reliability limits confidence in the dimension-level estimates and is the main constraint on future validation studies. Clinical validity and workflow effect were not assessed.