科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Bioengineering (Basel, Switzerland)2026-08-28

Burn Extent and Fitzpatrick Skin Tone Assessment from Clinical Photographs: Systematic and Random Error in Multimodal Large Language Models.

Ibrahim Güler, Armin Kraus, Gerrit Grieb, Henrik Stelling

一句话结论 · In one sentence

Averaging repeated answers removes only the smaller, random component; the larger, systematic one persists and requires calibration against reference data before clinical use can be considered.

原始摘要(英文原文)· Original abstract
BACKGROUND: Burn extent guides triage, transfer and fluid resuscitation, yet its clinical estimation is imprecise and observer-dependent. Multimodal large language models (MLLMs) process clinical photographs without task-specific training, but their error has rarely been separated into systematic and random components or their performance across skin tones characterized. METHODS: Three state-of-the-art MLLMs (Gemini 3.1 Pro, GPT-5.6 Sol, and Fable 5) each assessed 153 burn photographs five times under an identical prompt. The tasks were as follows: burned proportion of the imaged field, against an expert-guided pixel-wise segmentation (tolerance ± 10 percentage points, pp); burned percentage of total body surface area (TBSA), against physician consensus (±2 pp); and binary Fitzpatrick skin tone (FST; light I-III versus dark IV-VI). The first of these was the primary endpoint. RESULTS: The primary endpoint was in the range of 32.5-70.2%, with TBSA at 69.5-77.5%. All models compressed the estimation range (slopes 0.58-0.76, intercepts +10.0 to +22.2 pp); one multiplicative constant per model brought errors differing more than twofold into a 1.4 pp range. Across repeated queries, the median within-image range was 5.0-25.0 pp; averaging the five answers reduced error by only 0.24-2.22 pp. FST accuracy was 83.8-91.9% against a majority-class baseline of 81.0%. CONCLUSIONS: Averaging repeated answers removes only the smaller, random component; the larger, systematic one persists and requires calibration against reference data before clinical use can be considered.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Burn Extent and Fitzpatrick Skin Tone Assessment from Clinical Photographs: Systematic and Random Error in Multimodal Large Language Models. — 科研速览 Science Skim