Richard Sidebottom, Donna Webb, Des Cambell, Laura Satchwell, Emily Greenlay, Suzanne England, Tanja Gagliardi, Steve Allen, Bas Balhudin, Romney Pope, Elliot Elwood, Julie Scudder, Victoria Sinett, Christina Messiou, Mu Koh
Our three-layer framework combining programmatic extraction with iterative clinician validation, and limited manual curation transformed inaccessible real-world EPR data into an audit-ready research dataset at scale.
OBJECTIVES: To develop and validate a scalable, semi-automated framework for extracting high-granularity research data from legacy Electronic Patient Records (EPR), using a decade of family history breast screening as the exemplar.
METHODS: Our multidisciplinary team developed a three-layer architecture distinguishing raw EPR data, a context layer holding a structured patient journey, and analysis-ready output variables. The context layer was implemented in Structured Query Language (SQL) with explicit rules for cohort identification, exclusions, imaging-event linkage, and outcome derivation. Validation comprised a cohort inclusion audit and an independent patient-journey audit of 904 attendances.
RESULTS: The framework distilled 1,276,903 events in 7,781 women into a final cohort of 5,392 women comprising 26,483 screening attendances between 2010 and 2019. The inclusion audit found no missed cases. The journey audit returned seven errors (0.77%); four shared a systematic pattern of clinical recall with normal mammographic coding. Encoding this pattern as an additional SQL rule flagged 82 additional recalls and reduced the effective error rate to 0.33%. Screening performance (cancer detection rate 0.5%, recall rate 4.1%) reproduced the FH01 benchmark.
CONCLUSION: Our three-layer framework combining programmatic extraction with iterative clinician validation, and limited manual curation transformed inaccessible real-world EPR data into an audit-ready research dataset at scale.
ADVANCES IN KNOWLEDGE: Our study provides a practical approach to overcome the technical barriers and utilise EPR data at a scale not feasible manually. It demonstrates that semi-automated curation can benchmark clinical performance and validate new technologies like DBT in real-world settings where prospective data collection is absent.