Zsolt Bedőházi, Zsófia Sztupinszki, Ragnar P Kristjánsson, Mikkel Werling, Stephen Hamilton-Dutoit, Kristina L Lauridsen, Lisa Ottander, Trine L Plesner, Peter Hollander, Ingrid Glimelius, Lene Sjö, Estrid Høgdall, Carsten Utoft Niemann, Klaus Rostgaard, Péter Pollner, Henrik Hjalgrim, István Csabai
Accurate stratification of Hodgkin lymphoma (HL) by immunologic/histological subtypes and Epstein-Barr virus (EBV) status is essential for epidemiological and translational research, yet large-scale testing is impractical and expensive. Digital pathology models that utilize routinely used hematoxylin and eosin (H&E) whole-slide images (WSIs) could close this gap. We developed and validated a hierarchical Vision Transformer pipeline that aggregates cell-, patch-, and region-level context to predict EBV status and the three most prevalent immunological/histological HL subtypes: nodular sclerosis (NS), mixed cellularity (MC), and nodular lymphocyte-predominant HL (NLPHL)-from H&E-stained WSIs, and additionally evaluated a standard attention-based multiple-instance learning (ABMIL) baseline for direct architectural comparison. The development pool comprised 1643 HL cases (1952 WSIs) from 18 Danish hospitals and was used for hospital-preserving 5-fold cross-validation; external validation was performed on an independent hold-out cohort of 458 cases (532 WSIs) from five hold-out hospitals. For subtype prediction, analyses were restricted to the 1560 cases belonging to NS, MC, or NLPHL. On the external EBV cohort ( N = 458 ) the hierarchical pipeline achieved an area under the receiver operating characteristic curve (ROC-AUC) of 0.73 (95% confidence interval (CI) 0.68-0.77), precision-recall (PR)-AUC 0.57 (95% CI 0.49-0.66) with recall (sensitivity) 0.74 (95% CI 0.67-0.80) and macro-F1 score 0.60 (95% CI 0.54, 0.65). For 3-class subtype prediction on 359 external cases, discrimination reached ROC-AUC 0.84 (95% CI 0.80-0.88) and PR-AUC 0.63 (95% CI 0.56-0.71) with a macro-F1 of 0.56 (95% CI 0.48, 0.64) and macro-recall 0.55 (95% CI 0.47, 0.63); residual errors were dominated by NS-MC confusions. An ABMIL baseline using the same patch embeddings achieved ROC-AUC 0.76 (95% CI 0.72-0.81) for EBV and 0.89 (95% CI 0.86-0.92) subtype prediction, outperforming the hierarchical model on both tasks. This multicenter study shows that both hierarchical and attention-based architectures can determine EBV status and major HL subtypes directly from routine H&E slides with externally validated performance across hospitals, whereas the finding that the simpler baseline outperformed the hierarchical model suggests that strong foundation-model embeddings combined with attention-based pooling may reduce the need for explicit multi-scale modelling in cohorts of this size.