Yixi Xu, Md Nasir, Valentina Matos-Romero, Tiane Chen, Ralph H Hruban, William B Weeks, Rahul Dodhia, Juan Lavista Ferres, Ashley L Kiemen
Per-class truncation improved independently rated image quality and subtype representation. Successful synthetic histology generation requires careful data curation, domain-specific oversight, and independent validation; clinical utility requires separate task-based evaluation.
BACKGROUND/OBJECTIVES: The training of diagnostic pancreatic pathologists is largely limited by the diversity of available pathology images, especially those of rare diseases or conditions.
METHODS: Using a cohort of seven pancreatic neoplasms, we developed an iterative, pathologist-in-the-loop workflow integrating scalable tile pruning, generative-model optimization, and postprocessing truncation. Two pathologists evaluated synthetic-image quality and subtype representation on a 0-3 scale; the independent pathologist was blinded to image source and truncation condition and also rated curated real training tiles.
RESULTS: Untruncated images had lower class-balanced FID than per-class-truncated images (6.64 versus 29.08) and higher recall and coverage, whereas truncation increased precision. The independent pathologist rated per-class-truncated images higher than matched non-truncated images (mean difference 0.80, bootstrap 95% CI 0.57-1.03). Quadratic-weighted Cohen's κ was 0.599 (95% CI 0.489-0.691); after grouping ratings as 0-1 versus 2-3, raw agreement was 80.7% (95% CI 75.0-86.4%).
CONCLUSIONS: Per-class truncation improved independently rated image quality and subtype representation. Successful synthetic histology generation requires careful data curation, domain-specific oversight, and independent validation; clinical utility requires separate task-based evaluation.