Oliger Veronica Mendoza Betancourt, Delgi Peraza
This study proposes a deep reinforcement learning (DRL)-based power allocation framework for underwater MIMO-NOMA visible light communication (UVLC) systems, rigorously validated under real-world turbulence profiles calibrated with empirical oceanographic data. The framework integrates depth-dependent turbulence modeling (Kolmogorov and Anderson spectra) to simulate dynamic underwater channels, addressing salinity gradients, temperature variations, and non-stationary turbulence. A multi-objective reward function, optimized via systematic grid search and Pareto analysis, balances spectral efficiency, BER minimization, and energy efficiency while penalizing forward error correction (FEC) of 3.8× 10-3. Hardware-aware simulations demonstrate that the proposed Q-learning DRL strategy achieves a 50% BER reduction (1.8× 10-3) vs. 3.6× 10-3), 7 dB SNR gain, and 30% throughput improvement over conventional water-filling and genetic algorithms, with <10 ms inference latency suitable for embedded deployment. Compared to state-of-the-art transformer and graph neural network (GNN) approaches, the framework reduces computational overhead by 75% while maintaining robustness across depths (20-100 m) and turbulence intensities. Experimental validation via FPGA prototyping confirms feasibility for autonomous underwater vehicle (AUV) swarms and sensor networks, addressing scalability limitations of prior methods. This work bridges simulation and real-world deployment, offering a physics-guided, adaptive solution for turbulent underwater optical channels.