Lukas Pasold, Luca Eisentraut, Felix Waigner, Ricardo Buettner
Integrating DL-based imaging analysis with ML-based clinical modelling enhances the discriminative performance for ObCAD detection with complementary interpretability. This ensemble framework demonstrates potential to support clinical decision-making by identifying high-risk patients using routine cardiac CT combined with patient-level clinical data. Future studies using external validation and coronary artery calcium scores may further improve risk prediction.
This article presents a multi-label Sentinel-2 image dataset addressing the limited availability of remote sensing benchmarks that reflect the real-world co-occurrence of energy, transport, and storage infrastructure. The collection contains 10,000 true-color RGB PNG image chips from across Germany. Each image is independently annotated for electrical substations, power plants, storage tanks, railway infrastructure, and major road infrastructure, allowing none, one, or multiple infrastructure types to occur within the same scene. Candidate locations and labels were derived from OpenStreetMap geometries. Samples were selected from retained linear and polygonal features, spatially balanced across Germany, and accepted only when their complete image footprints satisfied the defined intersection and non-overlap criteria. Contextually similar no-label samples were generated near infrastructure locations to provide challenging negative examples. Predefined spatial folds keep neighboring samples, associated no-label samples, and samples from the same local OpenStreetMap feature together, supporting reproducible evaluation while avoiding spatial leakage. The image chips were generated from Copernicus Sentinel-2 Level-2A bands B04, B03, and B02 using a maximum cloud-cover threshold of 10 percent and least-cloud-cover mosaicking. Each 224 × 224 pixel image is a resampled representation of a 512 m × 512 m footprint. The released files include the image collection and CSV metadata containing filenames, binary labels, and spatial fold assignments. Spatial five-fold baseline experiments demonstrate the dataset's suitability for supervised multi-label classification, reproducible model comparison, and infrastructure mapping.