C. Moreno Garcia, J. Mata Vazquez, V. Pachon Alvarez
The automatic detection of thoracic pathologies remains a challenge when the visual evidence present in the chest X-ray (CXR) is subtle, ambiguous, or practically imperceptible, and the available textual information is limited, especially in the early phases of patient care. To address this limitation, in this work a clinically motivated subset of MIMIC-CXR is taken as the study set, composed of four diseases associated with dyspnea presentations in the emergency department (heart failure, pulmonary infection, COPD/asthma, and pulmonary embolism), along with a control class. This selection allows evaluating the multimodal fusion of visual and textual information in a clinically diverse set where the contribution of each modality can vary depending on the considered pathology and the quality of the available information. In this work, a comparative study of different combinations of visual and textual encoders is presented, with the objective of identifying the most suitable configuration for multimodal fusion. Based on this evaluation, we propose Neural Gated Fusion, a GMU-based architecture that incorporates an adapted Gated Multimodal Unit to regulate the contribution of visual and textual representations, in contrast to strategies based on static feature concatenation or late decision-level fusion. This architecture uses a neural gate to dynamically regulate the contribution of the visual representation of the radiograph and the textual representation of the clinical report before the multi-label classification stage. The experiments conducted on 25,245 multimodal pairs show that the ViT + ClinicalBERT combination with modulated fusion achieves a macro ROC-AUC of 0.8545 and a macro F1-score of 0.6528, outperforming unimodal models, decision-level fusion, and static early fusion. The per-class analysis suggests that the designed fusion strategy especially improves performance in pathologies with limited visual evidence, such as pulmonary embolism, without compromising the detection of normal cases. These results suggest that the dynamic regulation between image and clinical context is an effective strategy for improving the robustness of multimodal thoracic pathology detection systems.