C. Yang, A. Thwaites, C. Wingfield, C. Zhang, A. Woolgar
The brain builds meaning from speech in stages, transforming acoustic input into linguistic comprehension. Yet where comprehension separates from general acoustic processing has been difficult to localize, because the two are tightly entangled in continuous speech. Here we align the activity of ~145,000 individual artificial neurons of an audio large language model with high-resolution magnetoencephalography, using single-neuron interpretability to track, layer by layer, which computations the cortex follows during natural listening. Contrasting native listeners with listeners hearing an unfamiliar language, while holding the acoustics physically identical, we find that comprehension selectively sustains brain-model alignment through the model's deep layers. Without comprehension, alignment in the acoustic encoder is preserved but deep-layer tracking abruptly collapses, and the sparse alignments that survive correspond only to physical acoustic features. This pinpoints the stage at which comprehension departs from perception and yields an interpretable, non-invasive index of whether speech is understood.