Nevzat Olgun
Identifying individuals in non-line-of-sight (NLOS) environments is a challenging problem. This difficulty stems from the absence of direct signal propagation and the dominance of multipath components in measurements. A significant portion of current studies focuses on suppressing these components as noise, which can lead to the loss of potentially discriminative individual-specific biometric information. In this study, a controlled, multichannel acoustic experimental setup representing NLOS conditions was designed. Using the data obtained from the designed experimental setup, multipath components were analyzed as a source of discriminative information, and an identification framework based on this analysis was developed. For this purpose, data were collected from 10 different subjects through an acoustic experimental setup consisting of 8 speakers and 8 microphones. In this study, secondary echo is treated as a delayed multipath component due to subject environment interaction, observed after the early time region dominated by direct speakers to microphone leakage and early wall reflections. Therefore, the raw signal from each microphone channel was divided into non-overlapping 1024 sample windows, and the last window exceeding the adaptive energy threshold was selected as the secondary echo segment. While this process temporally filters out early components, the selected segment may include the effects of room reverberation as well as delayed subject-environment interactions. Two different feature representations were extracted from the secondary echo segments. First, a time-averaged spectral representation (STFT) based on the short-time Fourier transform was obtained. Then, topological features based on the Limited Penetrable Visibility Graph (LPVG) were extracted. Spectral and topological representations were combined at the feature level to create a hybrid feature space. These obtained features were classified using Random Forest (RF), Extra Trees, and SVM. Model performance was evaluated under varying environmental conditions within the same room geometry. For this purpose, an object-configuration-based Leave-One-Scene-Out validation protocol was applied. The highest performance was obtained with the RF classifier under the configuration combining single and simultaneous multi speaker excitations. Under these conditions, STFT and LPVG representations achieved mean accuracies of 94.17% ± 6.48% and 88.83% ± 9.18%, respectively. The average accuracy of the fusion representation was found to be 96.17 ± 4.07%. The findings showed that the combined use of spectral and topological representations generally provided higher accuracy than either individual representation. A lower inter-fold standard deviation was also observed for the fusion approach. Furthermore, the results indicate that secondary acoustic echo components can carry discriminative patterns between individuals in NLOS environments. It was observed that environmental object configurations can influence discriminability by altering the multipath structure, and under certain conditions, this effect can be enhanced.