Zhirui Wang, Xiaogen Yuan, Haobin Zhan, Xiaoqiao Wang, Jianping Guo
Deep learning has become essential for nanophotonic inverse design, yet non-uniqueness-distinct structures producing near-identical spectra-severely undermines model reliability. Existing studies offer only qualitative observations of this problem. Here, we establish a quantitative framework that integrates DBSCAN clustering with adaptive spectral similarity thresholds to systematically evaluate how non-unique data degrade deep learning performance. Validated across three structurally diverse photonic systems, our analysis reveals two critical benchmarks: model convergence begins deteriorating when inter-spectral MSE falls below 1e-4, and severe performance collapse ensues below 1e-5. Complex architectures prove disproportionately vulnerable to non-uniqueness compared to simpler ones. We demonstrate that strategically filtering training data using these thresholds markedly restores prediction accuracy and convergence stability. This work provides the first quantitative diagnostic criteria and practical preprocessing guidelines for addressing non-uniqueness, directly applicable to improving the reliability of data-driven spectral analysis and photonic design.