科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cancer causes & control : CCC2026-09-20

Bridging the data latency gap: automated extraction of genomic biomarkers from unstructured clinical documents to support real-world oncology data.

Qianyun Luo, Rui Zhang, Nikitha Vobugari, Jane Y C Hui, Schelomo Marmor

一句话结论 · In one sentence

Automated extraction of genomic biomarkers represents a scalable approach to reducing delays in cancer data availability. Earlier capture of genomic information may support cancer registry modernization and improve real-world evidence generation in precision oncology.

原始摘要(英文原文)· Original abstract
PURPOSE: Real-world oncology data are essential for clinical research and precision cancer care. However, genomic biomarkers are often embedded in scanned, unstructured clinical documents requiring manual abstraction before becoming available in cancer registries, delaying real-world evidence generation. This study evaluated and compared three open-source optical character recognition (OCR) approaches, Tesseract, EasyOCR, and a hybrid implementation, to determine which best enables automated extraction of Oncotype DX recurrence scores and improves the timeliness and quality of real-world oncology data. METHODS: We evaluated the feasibility of automated genomic data extraction using 675 Oncotype DX reports from a Midwestern U.S. health system. EasyOCR, Tesseract, and a hybrid OCR approach were used to extract recurrence scores from scanned reports. OCR-derived values were compared with manually abstracted scores and local cancer registry data. Performance was assessed using agreement, precision, recall, F1 score, and processing time. Multivariable logistic regression was performed to identify factors associated with discordance between registry-reported and manually abstracted scores. RESULTS: The hybrid OCR approach demonstrated the highest performance, achieving 97% agreement with manual abstraction, precision of 0.997, recall of 0.972, and an F1 score of 0.984. Registry abstraction demonstrated comparable performance but required greater manual effort. Automated extraction substantially reduced processing time while maintaining high accuracy. Logistic regression showed registry discordance was largely independent of patient and tumor characteristics, with unknown progesterone receptor (PR) status as the only significant predictor. CONCLUSION: Automated extraction of genomic biomarkers represents a scalable approach to reducing delays in cancer data availability. Earlier capture of genomic information may support cancer registry modernization and improve real-world evidence generation in precision oncology.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Bridging the data latency gap: automated extraction of genomic biomarkers from unstructured clinical documents to support real-world oncology data. — 科研速览 Science Skim