Ziying Hua, Yushi Zhou
Deep models for oral squamous cell carcinoma (OSCC) detection from haematoxylin-eosin histopathology degrade sharply across laboratories, scanners and magnifications, yet the target centre is usually unavailable during training, making generalization to unseen domains the decisive requirement for reliable screening. We cast multi-centre OSCC classification as a domain-generalization problem and propose SICA (Style-Invariant Consistency Alignment), a single objective combining three established invariance mechanisms: style randomisation of channel-wise feature statistics, a stain-consistency term forcing two independently stained views of each image to agree in embedding and prediction space, and a class-conditional feature alignment that preserves the decision boundary. Two of them rely only on paired stain views, so SICA remains effective with a single source domain, where classical cross-domain alignment collapses to plain risk minimisation. On a reproducible leave-one-domain-out benchmark spanning a genuine cross-magnification shift and a controlled multi-centre stain-shift simulation, SICA is compared against nine baselines over eight metrics. Among the ten algorithms sharing a common ImageNet backbone it attains the best and most stable out-of-domain AUROC (90.6% on the multi-centre protocol), the best worst-case centre, and a balanced operating point better calibrated than empirical risk minimisation. With confidence intervals, paired tests and multiplicity correction, that margin over the strongest baselines proves consistent but not statistically decisive on five simulated centres. Under the same protocol, a frozen pathology foundation model closes part of the stain gap but not the scale gap, and the proposed objective improves it further (92.3% to 93.4% AUROC on identical frozen features), so better representations and explicit invariance are complementary. An audit of the public corpus uncovered 28 byte-identical images filed under contradictory labels, removed before re-running the benchmark under a group-aware, leakage-controlled protocol. Finally, on a genuinely independent 150-patient cohort we report a negative result: every method, ours included, falls to chance-level discrimination, because the binding constraint is field-of-view scale rather than stain. Stain invariance is therefore necessary but not sufficient, and a simulated stain-shift benchmark must not be read as an estimate of deployment performance.