Ibrahim Güler, Armin Kraus, Gerrit Grieb, Henrik Stelling
Background: Burn photography is a stringent testbed for visual artificial intelligence (AI) in medical imaging, since it permits quantitative and qualitative clinically relevant parameters to be elicited. It is therefore a useful proxy for current vision AI. Reducing a photograph to a semantic segmentation removes color, texture and anatomy while providing pixel-exact boundaries. How spatial assessment responds, and how far instance-level structure can be recovered from a map that does not encode it, is unknown. Methods: Three state-of-the-art multimodal large language models (MLLMs) each assessed 153 burn photographs five times under three input conditions: the photograph alone (IMAGE), its segmentation mask alone (MASK), or both combined (BOTH). Four items were elicited per image: burn presence, the burned body region, per-cell classification of a 3 × 3 grid, and the number of separate burns, a surrogate for instance-level recognition. Each model was compared across conditions, paired on the same images. Results: Body-region accuracy was 97.1-100.0% under IMAGE, 47.5-54.0% under MASK and 50.2-99.9% under BOTH. Grid-cell accuracy was 77.7-82.3%, 77.9-96.3% and 90.9-98.4%. On the burn count, exact-match was 60.5-66.3%, 61.3-96.6% and 84.7-97.4%, and no model significantly beat the trivial classifier under IMAGE, two of three did under MASK and all three under BOTH. Across 18 paired comparisons of BOTH against a single channel, BOTH was better in 11, worse in three and indistinguishable in four. Conclusions: Combining the two inputs improved spatial assessment in most comparisons and degraded it in others, consistently across neither models nor tasks. Additional visual information therefore does not deterministically improve model performance, and disruption remains possible.