Z. Guo, B. Jin, Z. Han, T. Ichikawa, N. Ozawa, X. B. Ling
Clinical Visual Memory demonstrates that optical compression can successfully unify highly heterogeneous clinical data into a single, high-fidelity visual-token representation.
Importance. Current clinical artificial intelligence is bottlenecked by highly specialized, fragmented models that reduce complex, multimodal data into isolated scalar predictions. By reframing multimodal perception as a pure vision problem, optical compression into a universal visual operating system can resolve internal data interaction bottlenecks and shift the paradigm toward continuous, generative clinical forecasting. Objective. To develop and internally validate Clinical Visual Memory, a generalist visual foundation model utilizing high-fidelity optical compression to fuse five heterogeneous intensive care unit (ICU) data modalities into a unified visual-token space, acting as a generative navigation system for critical-care trajectories. Design, Setting, and Participants. Retrospective modeling study using MIMIC-IV adult ICU encounters. After excluding stays shorter than 24 hours, 54,551 patients, 68,546 hospital admissions, and 74,829 ICU stays comprised the eligible cohort. Strict patient-level partitioning prevented cross-split leakage. Methods. Structured electronic health record (EHR) data, vital signs, 10-second electrocardiogram (ECG) waveforms, chest radiographs, and clinical notes were rendered as 2D images and encoded by a single frozen DINOv2 vision transformer. Through modality-aware latent-query cross-attention, these highly heterogeneous sources were optically compressed into a shared 1024-dimensional Clinical Visual Memory. To establish a universal output interface, this memory conditioned an conditional instruction-tuned image generator to decode eight-domain deterioration trajectories across 3- to 48-hour horizons. Visual compression fidelity was evaluated by reducing retained source-pixel area to 1%, and critical transitions were mapped using event-specific projected-axis geometry. Results. Operating as a generalist foundation, Clinical Visual Memory achieved AUROCs of 0.852 (95% CI, 0.821-0.882) for 48-hour mortality, 0.699 (0.673-0.725) for incident acute kidney injury (AKI), 0.723 (0.688-0.757) for high Sequential Organ Failure Assessment (SOFA), and 0.742 (0.726-0.757) for alive ICU discharge. Crucially, resolving the token bottleneck via an 8-fold reduction in source-pixel area preserved 98.9% of the uncompressed 48-hour mortality AUROC. Generative image-out decoding maintained clinical meaning, and trajectory geometry functioned as a proactive navigation system, yielding median warning lead times of 34.8 hours (IQR, 17.5-43.3) before death, 14.6 hours (6.5-33.2) before AKI, 15.0 hours (6.0-27.0) before high SOFA, and 20.8 hours (10.2-36.0) before recovery, with low false-alert burdens (0.040-0.109 per patient-day). Conclusions and Relevance. Clinical Visual Memory demonstrates that optical compression can successfully unify highly heterogeneous clinical data into a single, high-fidelity visual-token representation. By functioning as a continuous clinical navigation system rather than a discrete alert generator, this framework lays the architectural foundation for a generalist, generative visual operating system in critical care medicine. External and prospective validation are required before clinical deployment.