科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cancers2026-08-23

Hybrid Lexical-Semantic AI Architecture for Automated Cancer Registry Coding for the Vet-ICD-O-Canine-1 System from Free-Text Veterinary Pathology Reports.

Vitória Souza de Oliveira Nascimento, Marcello Vannucci Tedardi, Guilherme da Silva Rogério, Katia Cristina Pinello, Maria Lúcia Zaidan Dagli

原始摘要(英文原文)· Original abstract
Background/Objectives: Free-text veterinary pathology diagnoses contain essential information for cancer registration but are difficult to convert into standardized ontology-based codes because of linguistic variability, contextual modifiers, and large ontology search spaces. This study evaluated a hybrid lexical-semantic architecture for the automated assignment of Vet-ICD-O-Canine-1 morphology codes. Methods: A retrospective single-registry benchmark included 211 diagnoses from the São Paulo Animal Cancer Registry. Of these, 190 contained sufficient morphological information for expert-reviewed reference coding, whereas 21 generic or insufficiently specified descriptions were retained as an exploratory challenge subset. Fuzzy lexical matching retrieved Top-10, Top-20, or Top-30 candidates from the complete 971-entry morphology ontology, followed by semantic selection using Claude Haiku 4.5 and structured JSON output. Performance and computational efficiency were compared to direct full-ontology inference. Results: Among the evaluated fuzzy metrics, token_set_ratio achieved the highest Top-30 reference-code retrieval rate of 89.5%. End-to-end exact-match agreement increased from 73.7% with Top-10 to 79.5% with Top-20 and 85.8% with Top-30 (95% CI, 80.1-90.0%). Top-30 generated non-null codes for 93.2% of the 190 evaluable diagnoses and achieved a conditional exact-match agreement of 92.1%. By contrast, the direct full-ontology baseline achieved 71.2% conditional exact-match agreement (42/59) among non-null predictions and 22.1% end-to-end exact-match agreement (42/190) when incorrect predictions, null outputs, and technical failures were considered non-concordant outcomes. Compared to direct full-ontology inference, Top-30 reduced input-token consumption by 92.4%, total token consumption by 92.2%, and inference cost by 91.3%, while avoiding the 118 API rate-limit failures observed with the direct baseline. Among the 21 insufficiently specified diagnoses, Top-30 returned null codes in 38.1% and non-null codes in 61.9%. Conclusions: Ontology-guided candidate reduction improved coding agreement, computational efficiency, and operational robustness within this retrospective single-registry benchmark. However, the reported performance estimates require confirmation in larger independent datasets, and an upstream data-sufficiency or abstention mechanism is needed before prospective operational deployment.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Hybrid Lexical-Semantic AI Architecture for Automated Cancer Registry Coding for the Vet-ICD-O-Canine-1 System from Free-Text Veterinary Pathology Reports. — 科研速览 Science Skim