Souaad Hamza-Cherif, Nesma Settouti
Multimodal behavioral sensing may support depression screening, but evaluation on small clinical-interview datasets is particularly vulnerable to data leakage and model-selection bias. We present a leakage-audited trimodal framework evaluated on DAIC-WOZ (n=180, PHQ-8 ≥10), combining SBERT text embeddings, OpenFace facial-behavior descriptors, and COVAREP prosodic features. Participant-level partitioning is performed before augmentation, while decision thresholds and neural-model checkpoints are selected exclusively from internal validation data. A controlled five-seed experiment showed that a deliberately leaky full-pool MixUp construction, in which a retained development sample could include a held-out participant as its second parent, was associated with a 27-33 percentage-point increase in Macro-F1 across four fusion configurations. Under the participant-level leakage-free 5-fold protocol, trimodal late fusion achieved a Macro-F1 of 0.532±0.041, compared with 0.446±0.021 for Text+Imaging late fusion. Paired participant-level correctness outcomes also favored trimodal fusion (McNemar χ2=6.618, p=0.010), consistent with improved paired classification when the audio modality was included under the leakage-free protocol. On 86 participant-disjoint E-DAIC sessions, using a consistent PHQ-8-based outcome definition (PHQ-8 ≥10), trimodal late fusion achieved Macro-F1 = 0.621 and AUC = 0.676; this experiment is interpreted as within-family generalization rather than independent cross-corpus validation. Overall, the results show that leakage control can substantially alter both absolute performance and comparative conclusions, and that multimodal gains should be established through participant-level evaluation, modality-specific analysis, and reproducible model-selection procedures.