Solène Debuysère, Nicolas Trouvé, Nathan Letheule, Olivier Lévêque, Elise Colin
We present a framework for adapting a large pretrained latent diffusion model to very high-resolution Synthetic Aperture Radar (SAR) image generation. The approach enables controllable synthesis of rare or unseen scenes beyond the training set. Rather than training a task-specific small model from scratch, we adapt an open-source text-to-image foundation model to the SAR modality, using its semantic prior to align prompts with SAR imaging physics (side-looking geometry, slant-range projection, and coherent speckle with heavy-tailed statistics). Using a 100k-image SAR dataset at 40 cm resolution, we compare full fine-tuning and parameter-efficient Low-Rank Adaptation (LoRA) across the UNet diffusion backbone, the Variational Autoencoder (VAE), and the Text Encoders. Evaluation combines (i) statistical distances to real SAR amplitude distributions, (ii) GLCM-based textural similarity, (iii) semantic alignment with a SAR-specialized CLIP model, and (iv) SAR-specific analyses, including layover–shadow geometry on ∼ 150 buildings and Point Spread Function (PSF) fidelity via spectral main-lobe widths. Our results show that a hybrid strategy — full U-Net tuning with LoRA on the text encoders and a learned token embedding — best preserves SAR geometry and texture while maintaining prompt fidelity reducing the KL divergence. The framework supports text-based control and multimodal conditioning (e.g., segmentation maps, TerraSAR-X, or optical guidance), opening new paths for large-scale SAR scene data augmentation and unseen scenario simulation in Earth observation.