Eugenio Ventimiglia, Rolf Gedeborg, Marcus Westerberg, Paolo Zaurito, Fredrik Jäderling, P. Stattin, Hans Garmo
BACKGROUND AND AIM: Magnetic resonance imaging (MRI) is crucial for prostate cancer (Pca) diagnosis, risk stratification, and treatment planning. However, large-scale observational studies require structured MRI data, which are often only obtainable from free-text reports. We aimed to extract information from narrative prostate MRI reports and to describe subsequent biopsy outcomes in a nationwide population-based cohort. METHODS: We identified 108,361 prostate MRI examinations in Prostate Cancer database Sweden with extended treatments and endpoints data (PCBase Xtend) performed in 2015-2023. A rule-based text recognition algorithm was created and used to extract Prostate Imaging Reporting and Data System (PI-RADS) score and prostate volume from free-text MRI reports. Extracted data were validated against manually extracted information in the National Prostate Cancer Register (NPCR). We examined biopsy rates and Gleason score according to PI-RADS, Prostate Specific Antigen (PSA) density, and calendar year. RESULTS: The proportion of reports with identifiable PI-RADS scores increased from 38% in 2015-2016 to 83% in 2022-2023, with excellent agreement with NPCR data (correlation coefficient r = 0.94). Extracted prostate volumes correlated well with those in NPCR (r = 0.88). Biopsy rates decreased for PI-RADS 3 lesions over time, particularly in men with PSA density < 0.15 ng/ml/ml, while the proportion of men with PI-RADS 5 lesions who underwent biopsy increased. Almost all prostate cancers in men with PI-RADS 3 lesions were Gleason 6 or 7 (3+4). Gleason 9-10 was almost exclusively found in PI-RADS 5 lesions. CONCLUSIONS: Automated extraction of information from unstructured MRI reports is feasible and accurate. The observed temporal trends reflecting increasing quality and standardization of prostate MRI support its use in large-scale epidemiological research.