Hyouneek Jeon, Youngjoo Kim, Si-on Park, Junghwa Bahng
This brief communication proposes Beyond Hearing, an artificial intelligence-based multisensory feedback framework designed to support self-directed auditory rehabilitation beyond the temporal and spatial limitations of conventional face-to-face intervention. The framework integrates Whisper (OpenAI, San Francisco, CA, USA) based automatic speech recognition, MediaPipe (Google LLC, Mountain View, CA, USA) Face Mesh-based visual articulatory tracking, and Librosa-based acoustic feature analysis. These components are intended to document user responses, quantify lip movement patterns, and extract temporal and spectral speech features relevant to Korean speech perception and production. An adaptive training structure is proposed to adjust task difficulty and signal-to-noise ratio according to user performance. The proposed framework provides a conceptual model for combining auditory, visual, and acoustic information during home-based auditory training. By organizing repeated response data and providing multisensory feedback, the system may help users recognize error patterns and may assist audiologists in monitoring training progress. Beyond Hearing is presented as a conceptual and technical framework, not as a clinically validated intervention. Further studies are needed to examine its technical accuracy, usability, agreement with expert judgment, and clinical effectiveness in people with hearing impairment.