Danielle A Morrison, Abhinita S Mohanty, Lucian Sulica, Pegah Khosravi, Anaïs Rameau
StyleGAN3 produces synthetic laryngeal images demonstrating high perceptual realism. The significant real-versus-synthetic accuracy gap highlights that while true frames remain distinct, synthetic generations achieve profound ambiguity. Perceptual realism plateaus early (10,120-20,120 kimg), suggesting moderate training durations can produce human-perceived realism while reducing computational and environmental demands.
OBJECTIVE: To evaluate the perceptual realism of synthetic videostroboscopic laryngeal images and determine the relationship between training duration and clinician classification accuracy.
METHODS: Synthetic images were generated using StyleGAN3 from two age-stratified datasets: Dataset A (≥ 65 years) and Dataset B (< 65 years). A total of 114 clinicians evaluated a randomized, balanced-controlled 36-image survey drawn from a 144-image pool containing real frames and five StyleGAN3 training intervals (5120-25,000 kimg). Accuracy was analyzed across image classes, experience levels, and viewing devices.
RESULTS: Clinicians demonstrated significantly higher mean accuracy identifying true clinical frames than synthetic generations (70.9% vs. 57.8%, p < 0.001). Overall survey accuracy was 59.9% ± 14.2%, with realism peaking at 20,120 kimg (44.8% accuracy). Accuracy dropped significantly between 5120 kimg (79.6%) and 10,120 kimg (59.1%, p < 0.001), with diminishing returns thereafter. No significant accuracy difference existed between Datasets A and B (p = 0.27). The effect of specialty reached borderline significance (p = 0.059), though limited by severe subgroup imbalances (e.g., n = 5 fellows, n = 2 residents). Computer users (63.9%) were significantly more accurate than phone users (56.2%, p = 0.014). A confounder analysis confirmed specialty and device choice were statistically independent (p = 0.164).
CONCLUSION: StyleGAN3 produces synthetic laryngeal images demonstrating high perceptual realism. The significant real-versus-synthetic accuracy gap highlights that while true frames remain distinct, synthetic generations achieve profound ambiguity. Perceptual realism plateaus early (10,120-20,120 kimg), suggesting moderate training durations can produce human-perceived realism while reducing computational and environmental demands.
LEVEL OF EVIDENCE: N/A.