Xiao Zhang, Wei Zhang
The full model achieved an AUC of 0.842 ± 0.011, an RMSE of 0.091 ± 0.007 under the current task definition, and a Kendall tau of 0.730 ± 0.018. Relative to matched ablations, visual affect increased AUC by 0.041, time-aware recurrence reduced RMSE by 33.6%, and spatial constraints increased Kendall tau by 0.089. The RMSE unit must be standardized after the propagation-speed versus propagation-time definition is resolved, as marked in the proof.
INTRODUCTION: Disaster-response robots and other embodied agents can use social-media imagery as contextual evidence, but existing approaches typically analyze affect at the image level or model diffusion without isolating affect-specific information.
METHODS: We encoded disaster images into 128-dimensional continuous visual-affect embeddings representing visual intensity, fear, sadness, and negative valence. These embeddings initialized node states in a time-aware graph neural network whose edges were constrained by observed interaction links, temporal precedence, and spatial reachability. Multi-task heads predicted propagation scale, propagation speed, and spatial-ranking consistency, and a parameter-decomposition module separated affect-associated contributions from structure-associated residuals. Experiments used 4,450 reconstructed image-centered CrisisMMD samples, a common 70/15/15 split, five random seeds, matched baselines, and ablation tests.
RESULTS: The full model achieved an AUC of 0.842 ± 0.011, an RMSE of 0.091 ± 0.007 under the current task definition, and a Kendall tau of 0.730 ± 0.018. Relative to matched ablations, visual affect increased AUC by 0.041, time-aware recurrence reduced RMSE by 33.6%, and spatial constraints increased Kendall tau by 0.089. The RMSE unit must be standardized after the propagation-speed versus propagation-time definition is resolved, as marked in the proof.
DISCUSSION: Visual affect provides an interpretable complementary signal for modeling disaster-image diffusion. The findings do not establish causal effects, cross-platform generalizability, real-time device performance, or closed-loop robot benefit; the framework should therefore be regarded as an upstream social-sensing prior for embodied decision support.