科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-09-11· cs.CL

Meddies-PII: A Multilingual Framework for Personally Identifiable Information Extraction in Clinical De-identification

Linh Uyen Le, Christian Hoang, Huy Hoang Ha

原始摘要(英文原文)· Original abstract
Clinical de-identification relies on accurately identifying personally identifiable information (PII). However, manually annotated datasets are costly to construct, while existing synthetic alternatives often provide limited details about their generation process or rely on relatively simple synthesis strategies. We introduce Meddies-PII-Dataset, a corpus of one million synthetic clinical documents spanning seventeen languages and nine PII labels. The documents are generated using attribute-conditioned prompts and validated through thirteen deterministic gates that enforce structural and annotation consistency. To evaluate the dataset's utility, we train Meddies-PII-Model, a BIOES token classifier, and compare it with existing PII extraction systems using exact-match entity-level F1. Meddies-PII-Model achieves the highest performance among the evaluated systems on all reported benchmarks, with a mean F1 of 0.827 across fifteen external benchmarks, compared with 0.658 for the strongest baseline. Upon acceptance, we will publicly release the dataset, benchmark suite, model, generation framework, and evaluation code to support research on multilingual clinical de-identification.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Meddies-PII: A Multilingual Framework for Personally Identifiable Information Extraction in Clinical De-identification — 科研速览 Science Skim