Faseela Abdullakutty, Younes Akbari, Somaya Al-Maadeed, Ahmed Bouridane, Rifat Hamoudi, Suchithra Kunhoth
Accurate breast cancer subtyping guides treatment selection, yet
histopathology captures morphology without molecular state, while genomic
profiling captures molecular signatures without spatial context. Existing
fusion methods rely on concatenation, or on attention applied only after each
modality is encoded independently. This work identifies a scale-dependent
asymmetry in the direction of cross-modal conditioning: the direction that
performs best under limited samples is not the one that holds at scale, and
the reversal is traced to the capacity of the modulation pathway rather than
to the fusion principle. The comparison is carried out within a
hypernetwork-guided framework in which an auxiliary network maps one
modality to conditioning parameters that modulate the other's feature
representation, shaping features at the parametric level rather than the
decision stage; modulation is patient-specific rather than patch-specific.
Both directions are instantiated --- gene-to-image (HyperG2I) and
image-to-gene (HyperI2G) --- and trained under a label-aware MixUp strategy
that interpolates within-class samples across both modalities, preserving the
hard binary labels clinical decisions require. The framework is evaluated on
two paired TCGA-BRCA cohorts --- one limited-sample, one independently
assembled at scale --- under a single protocol spanning two whole-slide
representations, multiple visual backbones, and both conditioning directions.
On the limited-sample cohort, gene-to-image conditioning at its optimal
augmentation setting exceeds early fusion and both unimodal baselines, giving
the highest recall on the aggressive Basal/HER2 class of any configuration
evaluated, and an ablation favours intra-class over inter-class mixing. At
scale this ordering does not hold: image-to-gene conditioning sustains its
performance whereas gene-to-image does not, recovering only partially under
the full tissue bag and isolating the capacity of the modulation pathway as
the binding constraint. Direction and capacity of cross-modal conditioning,
rather than fusion depth alone, therefore govern how such frameworks scale.