声维重塑一多源神经信息融合的听皮层维度再生系统
听觉皮层多维拓扑、分层组织与表征几何
这些研究从功能梯度、皮层层级结构、层2/3几何组织及沿听觉通路的表征变换等角度,揭示听皮层如何组织频谱、时间包络、声源特征和噪声不变性。共同重点是建立听觉皮层多维拓扑、分层加工和表征几何的神经生物学基础,为听觉维度再生提供结构先验。
- Auditory Cortical Gradients Integrate Bottom-Up and Top-Down Structure During Natural Sound Categorisation(David Haydock, R. Leech, Magdalena Kachlicka, Frederic Dick, 2025, bioRxiv)
- Canonical cortical architecture supports the emergence of noise-invariant auditory representations(Tomas Suarez Omedas, Ross S. Williamson, 2025, bioRxiv)
- The Geometry of Layer 2/3 Cortical Sound Processing in Slow Wave Sleep(Allan Muller, A. Filipchuk, S. Bagur, Brice Bathellier, 2025, Advancement of science)
- The geometry of cortical sound processing in slow wave sleep(Allan Muller, A. Filipchuk, S. Bagur, Brice Bathellier, 2025, bioRxiv)
- Orthogonal spectral and temporal envelope representation during the onset phase in auditory cortex(Kuniyuki Takahashi, Tian-Rui Guo, Tatsuya Yamagishi, Shinsuke Ohshima, Hiroaki Tsukano, Arata Horii, 2025, iScience)
- Noise-invariant representations of sound emerge along the canonical cortical hierarchy(Tomas Suarez Omedas, Ross S. Williamson, 2026, PLoS Biology)
听觉群体编码、时间动态与声音特征表征
这些文献聚焦神经元群体如何编码动态声音、空间位置、音高、声源信号及持续性刺激表征,并比较时间编码、群体编码和状态依赖性编码机制。它们共同刻画听觉表征的内容维度、时间维度和群体解码特征。
- Population coding of time-varying sounds in the non-lemniscal Inferior Colliculus.(Kaiwen Shi, Gunnar L. Quass, Meike M. Rogalla, Alexander N. Ford, Jordyn E. Czarny, P. Apostolides, 2024, Journal of Neurophysiology)
- Population coding of auditory space in the dorsal inferior colliculus persists with altered binaural cues(Meike M. Rogalla, Gunnar L. Quass, Harry Yardley, Clara Martinez-Voigt, Alexander N. Ford, Gunseli Wallace, Deepak Dileepkumar, Gabriel Corfas, P. Apostolides, 2024, bioRxiv)
- Cortical Representation of Pitch Perception in Mice(Jason W. Putnam, Abhay Kumar, Nasiru K. Gill, Jonathan Dinh, Franshesca Orellana Castellanos, Sofia Leusch, Sarah Vauhgn, Nikolas A. Francis, 2025, bioRxiv)
- Auditory representation of vocal signals in a pallial cortical circuit(Tarciso A. F. Velho, Dan Iancu, Rêmullo Brenno Galvão de Miranda Costa, Patrick D. Roberts, Claudio V. Mello, 2026, Journal of Neuroscience)
- Increased reliance on temporal coding when target sound is softer than the background(Nima Alamatsaz, Merri J. Rosen, Antje Ihlefeld, 2024, Scientific Reports)
- Auditory network persistence of stimulus representation in awake and naturally sleeping mice(Barak Hadad, Noa Regev, Uddi Kimchy, Ohad Rechnitz, Shai Abramson, Arseny Finkelstein, D. Derdikman, Y. Nir, 2025, bioRxiv)
选择性听觉注意、任务调制与注意状态解码
这些研究主要考察选择性听觉注意、任务目标和空间线索对听觉皮层表征及频谱—时间调制编码的影响,同时发展基于EEG、神经网络和无监督评估的注意解码方法。共同构成从注意机制到神经导向助听和听觉脑机接口的技术链条。
- Selective Auditory Attention Decoding in Naturalistic Conversations Using EEG-Based Speech Envelope Tracking in Multi-Speaker Environments(Gabriel Ivucic, Saurav Pahuja, Dashanka Da Silva, Tanja Schultz, 2025, Interspeech)
- Attention modulates the cortical representation of speech sounds in the spatial release from informational masking(Benjamin H. Zobel, R. Freyman, Lisa D. Sanders, 2025, Auditory Perception & Cognition)
- Auditory attention decoding based on neural-network for binaural beamforming applications(Roy Gueta, E. Zion-Golumbic, Jacob Goldberger, Sharon Gannot, 2025, Frontiers in Signal Processing)
- Unsupervised Accuracy Estimation for Brain–Computer Interfaces Based on Selective Auditory Attention Decoding(M. Lopez-Gordo, Simon Geirnaert, Alexander Bertrand, 2025, IEEE Transactions on Biomedical Engineering)
- Auditory Attention Detection for Thai Speech Using EEG Decoding(Shalong Samretngan, P. Israsena, S. Pan-ngum, 2026, International Conference on Electrical Engineering/Electronics, Computer, Telecommunications and Information Technology)
- Task-dependent cortical encoding of spectrotemporal modulations in human auditory cortex(Moïra-Phoebé Huet, Mounya Elhilali, 2025, Journal of the Acoustical Society of America)
复杂声场中的听觉对象绑定与目标声源分离
这些文献关注复杂声场中的听觉对象形成、目标声源绑定、语音跟踪和掩蔽条件下的表征变化。研究重点是联合群体编码、时间同步、注意驱动的特征增强以及目标与干扰声源之间的神经分离,体现听觉维度在真实场景中的动态重加权。
- Joint Population Coding and Temporal Coherence Link an Attended Talker’s Voice and Location Features in Naturalistic Multi-talker Scenes(Kiki van der Heijden, Prachi Patel, Stephan Bickel, J. Herrero, A. Mehta, N. Mesgarani, 2025, Journal of Neuroscience)
- The shape of attention: How cognitive goals sculpt cortical representation of speech(Moïra-Phoebé Huet, Mounya Elhilali, 2025, bioRxiv)
- Cortical Representation of Auditory Selective Attention in a Dichotic Listening Task: A Functional Near-Infrared Spectroscopy Study(T. Yamaguchi, Ryu-ichiro Hashimoto, Hiroki Sato, 2025, Brain Topography)
- Neural speech tracking and auditory attention decoding in everyday life(Lisa Straetmans, Kamil Adiloğlu, Stefan Debener, 2024, Frontiers in Human Neuroscience)
- Cortical representation of masked speech: effects of age, attention, and linguistic information in maskers(Anoop Basavanahalli Jagadeesh, A. Uppunda, 2025, Aging, Neuropsychology, and Cognition)
经验依赖可塑性、神经调制与预测性听觉重塑
这些研究揭示熟悉度、听觉训练、神经调质和感觉历史如何改变听觉皮层的群体表征及知觉决策。它们共同说明听觉维度再生不仅依赖输入信号,还需要经验依赖可塑性、预测机制和神经调制参与。
- The effect of familiarity on neural tracking of music stimuli is modulated by mind wandering(Joan Belo, Maureen Clerc, D. Schön, 2023, AIMS Neuroscience)
- Sound-evoked adenosine release in cooperation with neuromodulatory circuits permits auditory cortical plasticity and perceptual learning(I. Bayazitov, B. J. Teubner, Feng Feng, Zhaofa Wu, Yulong Li, J. Blundon, Stanislav S Zakharenko, 2024, Cell Reports)
- Auditory Training Alters the Cortical Representation of Complex Sounds(Huriye Atilgan, Kerry M. M. Walker, Andrew J. King, J. Schnupp, J. Bizley, 2025, Journal of Neuroscience)
- Cortical population codes for embedding sensory inputs into the prior context(Iacopo Hachen, Sebastian Reinartz, Alisea Stroligo, Alejandro Pequeño-Zurro, M. Diamond, 2026, bioRxiv)
基于ITD的听觉空间侧化与对手通道模型
两项研究均围绕双耳时间差和听觉空间侧化展开,重点检验对手通道模型、适应释放范式及群体调谐曲线对空间维度的解释能力。该组为听觉空间表征建模和实验范式可靠性评估提供专门依据。
- Auditory cortical responses to abrupt lateralization shifts do not reflect the activity of hemifield-specific units involved in opponent coding of auditory space.(B. Ilhan, Saliha B. Kurt, Pekcan Ungan, 2023, Neuropsychologia)
- Lateralization‐specific adaptation in auditory cortical evoked potentials: Comparison with frequency‐specificity(B. Ilhan, Saliha B. Kurt, Yavuz Bolay, Pekcan Ungan, 2024, European Journal of Neuroscience)
神经刺激与行为状态驱动的听觉皮层动态重配置
这些研究从迷走神经刺激和行为状态变化两方面考察听觉皮层表征的动态重配置,涉及层板、频段、空间感受野及离感受野响应的改变。共同价值在于为再生系统的闭环神经调控、状态依赖校准和刺激时机选择提供机制基础。
- Vagus nerve stimulation modulates information representation of sustained activity in layer specific manner in the rat auditory cortex(T. Shiramatsu, Kenji Ibayashi, Kensuke Kawai, Hirokazu Takahashi, 2025, bioRxiv)
- Dynamic representation of sound locations during task engagement in marmoset auditory cortex(Chenggang Chen, Evan D Remington, Xiaoqin Wang, 2025, bioRxiv)
多模态脑信号数据、跨模态对齐与神经表征融合
这些文献聚焦多模态脑信号采集、数据集与基准建设,以及视觉—听觉—语言等跨模态信息的时间对齐、语义映射和表征相似性分析。其共同重点是解决多源神经信息的标准化、皮层定位、跨模态知识迁移和融合建模问题,是系统工程与算法训练的基础。
- A Multimodal Seq2Seq Transformer for Predicting Brain Responses to Naturalistic Stimuli(Qianyi He, Yuan Chang Leong, 2025, arXiv.org)
- Visual-Auditory Multimodal Decoding from Multi-Channel Human Intracranial EEG(Bao-Wen Cheng, Ke Chen, Hao-Yu Hua, Zehan Wu, Zhuo-Fan Zhao, Liang Chen, Ying Mao, Meng Li, 2025, 31th International Conference on Neural Information Processing (ICONIP) abstracts)
- An EEG Dataset for Multimodal Semantic Alignment and Neural Decoding during Reading and Listening(Si-Tong Chen, Beiqianyi Li, Cuilin He, Dong-Yang Li, Mingyang Wu, Xin-Ke Shen, Song Wang, Xuetao Wei, Xindi Wang, Haiyan Wu, Quanying Liu, 2025, Scientific Data)
- Delayed knowledge transfer: Cross-modal knowledge transfer from delayed stimulus to EEG for continuous attention detection based on spike-represented EEG signals(Pengfei Sun, Jorg De Winne, Malu Zhang, P. Devos, D. Botteldooren, 2024, Neural Networks)
- YOTO (You Only Think Once): A Human EEG Dataset for Multisensory Perception and Mental Imagery(Yan-Han Chang, Hsi-An Chen, Min-Jiun Tsai, Chun-Lung Tseng, Ching-Huei Lo, Kuan-Chih Huang, Chun-Shu Wei, 2025, bioRxiv)
- Cross-Modal Correspondence Improves ErrP Decoding in Multimodal BCIs(Yixin Liu, Kang Yin, Hye-Bin Shin, 2026, Balkan Conference in Informatics)
- A Human EEG Dataset for Multisensory Perception and Mental Imagery(Yan-Han Chang, Hsi-An Chen, Min-Jiun Tsai, Chun-Lung Tseng, Ching-Huei Lo, Kuan-Chih Huang, Chun-Shu Wei, 2025, Scientific Data)
- CineBrain: A Large-Scale Multi-Modal Brain Dataset During Naturalistic Audiovisual Narrative Processing(Jianxiong Gao, Yichang Liu, Baofeng Yang, Jianfeng Feng, Yanwei Fu, 2025, arXiv.org)
- Multi-Modal EEG Datasets and Benchmarks for EEG-Based Neural Decoding Research(Szu-Chi Chiu, Zhige Chen, Jibin Wu, Kay Chen Tan, 2025, International Conference on Computational Intelligence and Security)
- Cortical Surface Reconstruction from 2D MRI with Segmentation-Constrained Super-Resolution and Representation Learning(Wenxuan Wu, Ruowen Qu, Dongzi Shi, Tong Xiong, Xiangmin Xu, Xiaofen Xing, Xin Zhang, 2024, International Conference on Medical Image Computing and Computer-Assisted Intervention)
语音、音乐与自然声音的非侵入式脑解码重建
这些研究直接从fMRI、EEG、MEG等非侵入式脑信号中解码或重建语音、音乐、自然声音及其内部听觉表征,方法涵盖深度生成模型、对比学习、最优传输和时空神经解码。共同目标是提升连续听觉内容、语义信息和声学结构的恢复能力。
- An fMRI-based auditory decoding framework combined with convolutional neural network for predicting the semantics of real-life sounds from brain activity(Mingqian Zhao, Baolin Liu, 2024, Applied Intelligence)
- Speech Decoding from Non-invasive Brain Signals Using Generative Learning Framework(Seo-Hyun Lee, Soowon Kim, Jun-Young Kim, Seong-Whan Lee, 2026, Balkan Conference in Informatics)
- Optimal Transport and Contrastive Learning for Brain Decoding of Musical Perception(Matteo Ciferri, M. Ferrante, N. Toschi, 2025, Annual International Conference of the IEEE Engineering in Medicine and Biology Society)
- Decoding Auditory Neural Representation Based on SERF-MEG and EEG(Chang-Zeng Liu, Huanqi Wu, Xiao-Yu Liang, Yang Gao, Min Xiang, Yuyu Ma, Xiaolin Ning, 2025, International Conference on Bioinformatics Research and Applications)
- Decoding the Spatiotemporal Dynamics of Neural Response Similarity in Auditory Processing: A Multivariate Analysis Based on OPM‐MEG(Chang-Zeng Liu, Yuyu Ma, Xiao-Yu Liang, Min Xiang, Huanqi Wu, Xiaolin Ning, 2025, Human Brain Mapping)
- Natural sounds can be reconstructed from human neuroimaging data using deep neural network representation(Jong-Yun Park, Mitsuaki Tsukamoto, Misato Tanaka, Y. Kamitani, 2025, PLoS Biology)
听觉神经假体、听力恢复与空间听觉重建
这些文献覆盖高密度神经接口、人工耳蜗、听觉脑干植入、儿童听力损失干预、空间听觉与声场重建,并讨论去传入后的音乐性幻觉等异常可塑性现象。共同关注听觉通路和皮层表征的功能恢复、神经假体转化、空间听觉改善及安全性风险。
- Minimally invasive implantation of scalable high-density cortical microelectrode arrays for multimodal neural decoding and stimulation(Mark Hettick, Elton Ho, Adam J. Poole, M. Monge, Demetrios Papageorgiou, Kazutaka Takahashi, Morgan LaMarca, Daniel Trietsch, Kyle Reed, Mark Murphy, Stephanie Rider, Kate R Gelman, Yoon Woo Byun, Joshua S Miller, Timothy L. Hanson, Vanessa Tolosa, Sang-Ho Lee, Sanjay Bhatia, Peter E. Konrad, Michael Mager, Craig H. Mermel, Benjamin I. Rapoport, 2025, Nature Biomedical Engineering)
- Effects of childhood hearing loss on the subcortical and cortical representation of speech(Axelle Calcus, Stuart Rosen, 2025, bioRxiv)
- Cortical neuroprostheses improve auditory coding and perception compared to cochlear implants(J. Taylor, N. Jovanovic, C. A. Navntoft, A. Albon, S. Rezaei-Mazinani, B. Bathellier, T. R. Barkat, 2025, bioRxiv)
- [Sound localization strategies for patients with bilateral cochlear implants].(J.-Y. Li, N.-Y. Wang, 2025, Zhonghua er bi yan hou tou jing wai ke za zhi = Chinese journal of otorhinolaryngology head and neck surgery)
- Sound field reconstruction using improved ℓ1-norm and the Cauchy penalty method(Lin-Sen Huang, Wangzeng Hui, Zhi-Yu Yang, Lihong Xia, Hao Zhang, Wei Zhang, 2024, Optimization and Engineering)
- [Cochlear implantation in a patient with profound bilateral deafness receiving osimertinib for advanced lung adenocarcinoma: a case report].(J. Mao, H.-M. Yang, S. Han, K. Rong, J. Xue, D.-M. Zhu, S.-D. Yu, J.-F. Li, 2026, Zhonghua er bi yan hou tou jing wai ke za zhi = Chinese journal of otorhinolaryngology head and neck surgery)
- [Auditory brainstem implant: current states and future prospects].(W. Tang, S. Zong, Pei-Yu Du, H.-J. Xiao, 2024, Zhonghua er bi yan hou tou jing wai ke za zhi = Chinese journal of otorhinolaryngology head and neck surgery)
- Musical Hallucinosis: Auditory Illusions After Hearing Loss and Cochlear Implantation(B. Dobkin, 2025, Neurorehabilitation and Neural Repair)
生物启发式听觉编码、语义声音地图与基础模型
这些研究从皮层微柱启发的稀疏时序编码、语义声音地图和听觉基础研究进展出发,探索声音表征的压缩、关联记忆、语义组织与生成建模。其独特贡献是为听觉维度再生系统提供生物启发式计算架构和高层语义表征先验。
- A biologically inspired auto-associative network with sparse temporal population coding(Ya Zhang, Kexin Shi, Xiaoling Luo, Yi Chen, Yuchen Wang, Hong Qu, 2023, Neural Networks)
- Navigable Semantic Sound Maps for Auditory Displays(Mika Alexander Sieweke, Thomas Hermann, 2025, Audio Mostly Conference)
- [Advances in basic research on key central neural substrates of tinnitus].(J.-Q. Wu, J. Lin, Y. Huo, J.-N. Zhang, 2025, Zhonghua er bi yan hou tou jing wai ke za zhi = Chinese journal of otorhinolaryngology head and neck surgery)
合并后形成十一条相互并列的研究主线:首先以听觉皮层拓扑、分层结构和群体编码揭示可再生的神经表征维度;其次分别讨论选择性注意、复杂声场中的对象绑定、经验与神经调制驱动的可塑性,以及ITD空间侧化等关键机制;随后以神经刺激和行为状态动态重配置补充闭环调控基础。在技术层面,研究进一步覆盖多模态脑信号数据与跨模态对齐、非侵入式语音和自然声音解码重建;在转化层面,涵盖神经假体、听力恢复、空间听觉重建及异常可塑性风险;最后以生物启发式编码和语义声音地图提供模型化与系统设计先验。重复出现的dynamic_representation_of_sound_locations_during仅归入神经刺激与行为状态驱动的动态重配置组,以保证分组互不交叉。
总计 58 篇相关文献
No abstract available
No abstract available
No abstract available
No abstract available
The Algonauts 2025 Challenge called on the community to develop encoding models that predict whole-brain fMRI responses to naturalistic multimodal movies. In this submission, we propose a sequence-to-sequence Transformer that autoregressively predicts fMRI activity from visual, auditory, and language inputs. Stimulus features were extracted using pretrained models including VideoMAE, HuBERT, Qwen, and BridgeTower. The decoder integrates information from prior brain states and current stimuli via dual cross-attention mechanisms that attend to both perceptual information extracted from the stimulus as well as narrative information provided by high-level summaries of the content. One core innovation of our approach is the use of sequences of multimodal context to predict sequences of brain activity, enabling the model to capture long-range temporal structure in both stimuli and neural responses. Another is the combination of a shared encoder with partial subject-specific decoder, which leverages common representational structure across subjects while accounting for individual variability. Our model achieves strong performance on both in-distribution and out-of-distribution data, demonstrating the effectiveness of temporally-aware, multimodal sequence modeling for brain activity prediction. The code is available at https://github.com/Angelneer926/Algonauts_challenge.
Most research decoding brain signals into images, often using them as priors for generative models, has focused only on visual content. This overlooks the brain's natural ability to integrate auditory and visual information, for instance, sound strongly influences how we perceive visual scenes. To investigate this, we propose a new task of reconstructing continuous video stimuli from multimodal brain signals recorded during audiovisual stimulation. To enable this, we introduce CineBrain, the first large-scale dataset that synchronizes fMRI and EEG during audiovisual viewing, featuring six hours of \textit{The Big Bang Theory} episodes for cross-modal alignment. We also conduct the first systematic exploration of combining fMRI and EEG for video reconstruction and present CineSync, a framework for reconstructing dynamic video using a Multi-Modal Fusion Encoder and a Neural Latent Decoder. CineSync achieves state-of-the-art performance in dynamic reconstruction, leveraging the complementary strengths of fMRI and EEG to improve visual fidelity. Our analysis shows that auditory cortical activations enhance decoding accuracy, highlighting the role of auditory input in visual perception. Project Page: https://jianxgao.github.io/CineBrain.
Perception requires more than passive sensing—it involves prioritizing the features most relevant to ongoing cognitive goals, a process guided by selective attention. A central question is whether attention operates by enhancing all features of a selected target, or by optimizing neural encoding around the specific demands of the task—i.e., is selective attention fundamentally anchored around task targets or around task goals? Here, we recorded electroencephalography (EEG) while participants performed two speech tasks—comprehension and detection—on identical auditory stimuli. Task difficulty was manipulated by introducing controlled background noise that increased cognitive demands without reducing speech intelligibility. We developed a novel EEG-based method, the Modulation Response Function (MRF), which captures cortical sensitivity to spectro-temporal features via spectrogram reconstruction. Behaviorally, comprehension performance declined with increased difficulty, with greater reliance on semantic cues, while detection performance remained near ceiling. Neurally, both envelope tracking and MRF magnitude were higher during comprehension, reflecting greater cognitive engagement. Critically, spectro-temporal tuning differed across tasks: formant-related modulations were selectively enhanced during comprehension, whereas pitch-related modulations were emphasized during detection. These findings support a discriminative model of attention, where cortical encoding is flexibly reshaped according to cognitive goals, selectively amplifying the features most relevant for successful task performance.
Reconstruction of perceptual experiences from brain activity offers a unique window into how population neural responses represent sensory information. Although decoding visual content from functional MRI (fMRI) has seen significant success, reconstructing arbitrary sounds remains challenging due to the fine temporal structure of auditory signals and the coarse temporal resolution of fMRI. Drawing on the hierarchical auditory features of deep neural networks (DNNs) with progressively larger time windows and their neural activity correspondence, we introduce a method for sound reconstruction that integrates brain decoding of DNN features and an audio-generative model. DNN features decoded from auditory cortical activity outperformed spectrotemporal and modulation-based features, enabling perceptually plausible reconstructions across diverse sound categories. Behavioral evaluations and objective measures confirmed that these reconstructions preserved short-term spectral and perceptual properties, capturing the characteristic timbre of speech, animal calls, and musical instruments, while the reconstructed sounds did not reproduce longer temporal sequences with fidelity. Leave-category-out analyses indicated that the method generalizes across sound categories. Reconstructions at higher DNN layers and from early auditory regions revealed distinct contributions to decoding performance. Applying the model to a selective auditory attention (“cocktail party”) task further showed that reconstructions reflected the attended sound more strongly than the unattended one in some of the subjects. Despite its inability to reconstruct exact temporal sequences, which may reflect the limited temporal resolution of fMRI, our framework demonstrates the feasibility of mapping brain activity to auditory experiences—a step toward more comprehensive understanding and reconstruction of internal auditory representations.
Speech decoding from non-invasive brain signal is a critical step toward neural interfaces that restore or augment human communication. Despite recent progress in deep learning for auditory perception and speech synthesis from cortical signals, non-invasive modalities such as electroencephalogram remain limited by low spatial resolution and susceptibility to noise. This work proposes a generative framework that integrates an encoder for spatio-temporal feature extraction and a generator–discriminator pair for latent representation regularization, enabling word decoding from non-invasive electroencephalogram signals. Using data from a participant performing spoken speech, the framework captured fine-grained temporal structure and yielded robust word-level performance. Peak accuracy reached 70.8 % with a balanced F1 score of 69.5 %, while phoneme reconstruction exhibited low mean squared error and high cosine similarity over training. These results demonstrate that the proposed encoder–decoder framework effectively models the temporal organization of spoken EEG signals, suggesting its potential to generalize toward imagined speech decoding for more flexible and non-invasive brain-to-speech communication. The approach provides an interpretable path from non-invasive neural signals to continuous speech representations and lays groundwork for future brain-computer interface systems.
One way to investigate the cortical tracking of continuous auditory stimuli is to use the stimulus reconstruction approach. However, the cognitive and behavioral factors impacting this cortical representation remain largely overlooked. Two possible candidates are familiarity with the stimulus and the ability to resist internal distractions. To explore the possible impacts of these two factors on the cortical representation of natural music stimuli, forty-one participants listened to monodic natural music stimuli while we recorded their neural activity. Using the stimulus reconstruction approach and linear mixed models, we found that familiarity positively impacted the reconstruction accuracy of music stimuli and that this effect of familiarity was modulated by mind wandering.
Auditory learning is supported by long-term changes in the neural processing of sound. We examined these task-depend changes in the auditory cortex by mapping neural sensitivity to timbre, pitch, and location cues in cues in trained (n = 5) and untrained control female ferrets (n = 5). Trained animals either identified vowels in a two-alternative forced choice task (n = 3) or discriminated when a repeating vowel changed in identity or pitch (n = 2). Neural responses were recorded under anesthesia in two primary auditory cortical fields and two tonotopically organized nonprimary fields. In trained animals, the overall sensitivity to sound timbre was reduced across three cortical fields compared with control animals, but maintained in a nonprimary field (the posterior pseudosylvian field). While training did not increase sensitivity to timbre across the auditory cortex, it did change the way in which neurons integrated spectral information, with neural responses in trained animals increasing their sensitivity to first and second formant frequencies, whereas in control animals cortical sensitivity to spectral timbre depended mostly on the second formant. Animals trained on timbre identification were required to generalize across pitch when discriminating timbre, and their neurons became less modulated by fundamental frequency relative to control animals. Finally, both trained groups showed increased spatial sensitivity and an enhanced response to sound source locations close to the midline, where the loudspeaker was located in the training chamber. These results demonstrate that training elicited widespread alterations in the cortical representation of complex sounds.
To advance the application of functional near-infrared spectroscopy (fNIRS) in brain-computer interface (BCI) technology, we investigated cortical activation patterns associated with auditory selective attention. Using a dichotic listening paradigm, participants were presented with simultaneous music and reading sounds to the left or right ear. During fNIRS recordings, they were instructed to selectively attend to the sound attribute (music vs. reading) or the spatial location (left vs. right ear). Cortical activity differences related to attentional targets were analyzed using a two-way analysis of variance (ANOVA), with sound attribute and spatial information as factors. Our results revealed a significant main effect of the sound attribute factor across multiple measurement channels. Notably, the right parietal region exhibited consistently greater activation when attention was directed toward music compared to reading sounds. Conversely, bilateral dorsolateral prefrontal cortex (DLPFC) channels showed higher activation when participants attended to reading sounds than to music. These findings indicate that cortical activation patterns are modulated by auditory attentional states based on sound attributes. Furthermore, preliminary classification analyses achieved an accuracy of 73.7% in discriminating attentional targets (music vs. reading sounds), demonstrating the feasibility of fNIRS-based BCI applications.
No abstract available
Knowledge of how vocal communication signals are represented in the auditory system is crucial for understanding the perceptual basis of vocal communication. Using male and female zebra finches, we identified differentially expressed molecular markers that helped define distinct (caudal, rostral, dorsal, and ventral) domains within the caudomedial nidopallium (NCM), a high-order cortical auditory area known for its song-selective responses. Using expression analysis of the activity-inducible gene zenk, we found that the number of activated neurons is more stimulus dependent in NCM than in the auditory midbrain or the caudomedial mesopallium and that information on the density and spatial distribution of responsive neurons in NCM is sufficient to discriminate responses to conspecific song from other stimuli. We observed stronger activation of dorsal NCM, higher selectivity of caudal NCM toward conspecific song, and strong activation of the inhibitory network of rostral NCM by nonconspecific song stimuli. The spatial organization of responsive cells was particularly sensitive to both spectral and temporal components of song. We also obtained evidence of broadly distributed song-selective neuronal ensembles and that individual NCM neurons participate in the representation of different conspecific songs, implying independent activation and molecular induction responses. We conclude that some basic aspects of the cortical response to complex auditory stimuli are topographically organized, a finding that has been elusive in other systems. These findings advance our knowledge of the functional organization of a key song-processing cortical area, providing novel insights into the auditory representation of vocal communication signals.
ABSTRACT Age-related declines in speech recognition are consistently linked to several interrelated factors – auditory acuity (peripheral and/or central), cognition, linguistic processing, neurophysiological efficiency, etc. These declines become more apparent in adverse listening conditions, such as informational masking, the additional masking effects that occur due to linguistic and/or cognitive confusions between the target sound and masker. In this study, we evaluated the age-related changes in the cortical representation of speech under different masking (informational vs energetic) and attention (active vs passive) conditions. We measured high-density cortical auditory evoked potentials (CAEPs) in 60 participants (30 young adults and 30 older adults) with clinically normal hearing (pure tone average <15 dB HL). CAEPs were recorded in the presence of informational and energetic (no linguistic information) maskers, while the participants either attended (active attention) or ignored (passive attention) the stimuli. Results showed that informational maskers caused significantly greater N1 latency delays than energetic maskers, with older adults showing significantly greater delays than younger adults. Active attention resulted in significantly larger N1 amplitudes compared to passive attention, with minimal age-related differences, possibly due to the strict criteria of auditory acuity in the older adult group. Further, microstate segmentation analyses, in addition to confirming the age-related delays in cortical responses (similar to N1 latencies), revealed longer engagement of fronto-central cortical regions under informational maskers, regardless of the presence of lexical-semantic information in the maskers. These findings, therefore, highlight the systematic effects of aging, attention, and linguistic complexity on the cortical representation of masked speech.
Little is known about the effects of childhood mild-to-moderate sensorineural hearing loss (MM HL) on the function of the auditory pathway. We aimed to examine the effect of childhood MM HL and the benefit of frequency-specific amplification on both subcortical and cortical auditory processing, and to relate it to speech-perceptual abilities. We recorded subcortical and cortical responses to speech syllables in nineteen children with congenital MM HL (unamplified and amplified), and sixteen children with typical hearing (unamplified sounds only). Speech perception was measured behaviourally. Congenital HL led to smaller subcortical and cortical responses to unamplified speech sounds. There was a significant benefit of amplification on subcortical and early, but not late, cortical responses, with some effects differing across age. No relationship was found between the neural and behavioural measures. Childhood MM HL affects both subcortical and cortical processing of speech. Amplification mostly benefits subcortical processing of speech in younger children. Childhood HL leads to functional changes in the processing of sounds, with amplification differentially affecting subcortical and cortical levels of the auditory pathway.
Spatial separation of multiple talkers reduces listener confusion (i.e., informational masking), improving speech processing. This spatial release from informational masking is thought to reflect automatic processes of stream segregation and improvements in selective attention. However, the relative contributions of bottom-up and top-down processing and their potential interactions remain unclear. This study investigated the role of attention in the spatial release from informational masking using event-related potentials. Noise-vocoded target words were presented with two-talker noise-vocoded masking babble, and the precedence effect manipulated the spatial cue to isolate effects on informational masking. In separate conditions, participants attended to the sounds to detect target words and ignored the sounds to perform a challenging visual task. Benefits of spatial separation on cortical auditory evoked potentials elicited by target words (N1-P2) were evident regardless of attention condition. Benefits of attending to the sounds were evident at later stages (P2, P3), and exploratory analysis also revealed attentional effects in earlier time windows (P1, late N1). These results showed strong contributions from preattentive bottom-up processing in the spatial release from informational masking. Attention benefitted later stages associated with the cognitive processing of sounds and may modulate early perceptual processing under some listening conditions.
Pitch perception arises from temporal and spectral cues in sound. We hypothesized that mice rely on temporal cues because they have wide auditory filters, resembling humans with hearing loss or cochlear implants. Using computational modeling, behavioral assays, and widefield calcium imaging, we found that the structure of periodotopy in auditory cortex predicts how well mice recognize temporal pitch cues, establishing mice as a robust model for temporal pitch perception.
‘Opponent channels model’ (OCM) is the widely accepted model for cortical representation of sound lateralization. Stimulus‐specific ‘release from adaptation’ (RFA) in cortical responses has been used in previous studies to test the predictions of this model. However, these attempts were shown to be prone to confounds of spurious responses such as those to auditory motion and sound onset. The present study aims to determine whether a multiple‐adaptor RFA algorithm could be employed for relatively confound‐free quantification of the population response of lateralization‐specific auditory cortical neurons, and provide useful data for estimation of the OCM hemifield tuning curves. Two experiments were conducted on 12 volunteers with normal hearing. In Exp.1, quadruple tone pips of either low or high frequency were presented as adaptor, followed by a single tone pip of either frequency as probe. In Exp.2, tone pips were replaced with dichotic click train pips with left‐leading and right‐leading interaural time difference (ITD). Frequency‐ and ITD‐specific RFA in cortical responses N1 and P2 was quantified using global field magnitude difference between ERPs to mismatched and matched adaptor‐probe pairs.
SUMMARY Meaningful auditory memories are formed in adults when acoustic information is delivered to the auditory cortex during heightened states of attention, vigilance, or alertness, as mediated by neuromodulatory circuits. Here, we identify that, in awake mice, acoustic stimulation triggers auditory thalamocortical projections to release adenosine, which prevents cortical plasticity (i.e., selective expansion of neural representation of behaviorally relevant acoustic stimuli) and perceptual learning (i.e., experience-dependent improvement in frequency discrimination ability). This sound-evoked adenosine release (SEAR) becomes reduced within seconds when acoustic stimuli are tightly paired with the activation of neuromodulatory (cholinergic or dopaminergic) circuits or periods of attentive wakefulness. If thalamic adenosine production is enhanced, then SEAR elevates further, the neuromodulatory circuits are unable to sufficiently reduce SEAR, and associative cortical plasticity and perceptual learning are blocked. This suggests that transient low-adenosine periods triggered by neuromodulatory circuits permit associative cortical plasticity and auditory perceptual learning in adults to occur.
Sensitivity to spectrotemporal modulations is critical for speech perception, yet how cortical representations of these modulations adapt to cognitive goals remains unclear. Using EEG in human listeners, we examined how task demands shape spectrotemporal encoding during speech perception under controlled background noise. Participants engaged in either a comprehension or detection task using identical speech stimuli, allowing us to isolate top-down effects on auditory processing of speech modulations. We developed the Modulation Response Function (MRF), a novel analysis method that reconstructs cortical sensitivity to spectrotemporal features. While both tasks elicited robust neural encoding, spectrotemporal tuning profiles diverged: comprehension selectively enhanced cortical sensitivity to formant-related modulations, whereas detection amplified pitch-related modulations. These task-specific patterns emerged despite identical sensory input, highlighting a flexible, goal-driven reshaping of spectrotemporal representation in auditory cortex. These findings provide evidence of a flexible regime of encoding spectrotemporal modulations in service of difference cognitive tasks in order to support speech perception in challenging environments.
In auditory cortex, neural responses to stimuli inside receptive fields (RFs) can be further facilitated by behavioral demands, such as attending to a spatial location. It is less clear how off-RF stimuli modulate neural responses and contribute to behavioral tasks. Our recent study revealed a particular form of location-specific facilitation evoked by repeated stimulation from an off-RF location, suggesting behavioral modulation of spatial RFs. To further explore this question, we trained marmosets to attend to sound locations that were either inside or outside the RFs of auditory cortical neurons. The majority of neurons showed increased firing rates at target locations inside their RFs. Interestingly, this increase also occurred outside the RFs, sometimes exceeding the responses at the RF center during passive listening. These task-related off-RF facilitation were much more common in the caudal area than in the rostral area and the primary auditory cortex. A normalization model reproduced the off-RF facilitation using widespread suppression. The model’s prediction was confirmed by experimental observations of widespread reductions in firing rate and hyperpolarized membrane potentials for off-RF stimuli. These results suggest that behavioral task demands recruit a broader range of neurons than those that are responsive to a target sound in the passive state.
The brain of a living organism enables stable information processing in response to constantly changing external environments and internal states. As one of such cortical modulation, the present study focused on the effect of vagus nerve stimulation (VNS) therapy on information representation of the auditory cortex. By quantifying sound representation using machine learning, we investigated whether VNS alters cortical information representation in a layer-specific and frequency band-specific manner. A microelectrode array meticulously mapped the band-specific power and phase-locking value of sustained activities in every layer of the rat auditory cortex. Sparse logistic regression was used to decode the test frequency from these neural characteristics. The comparison of decoding accuracy before and after the application of VNS indicated that sound representation of the high-gamma band activity was impaired in the deeper layers, i.e., layers 5 and 6, while it was slightly improved in the superficial layers, i.e., layers 2, 3, and 4. Moreover, there was an improvement of sound representation in theta band activity in the deeper layers, demonstrating the layer-specific and frequency band-specific effect of VNS. Given that the cortical laminar structure and oscillatory activity in multiple frequency bands helps the auditory cortex to act as a hub for feed-forward and feed-back pathways in various information processing, the current findings support the possibility that VNS provide complex effects on brain function by altering the balance of cortical activity between layers and frequency bands.
Persistent neural activity often outlasts sensory stimulation, bridging perception and action. While commonly linked to working memory and decision making, its existence during passive states and sleep remains unclear. Using chronic high-density electrophysiology in freely behaving mice, we show that population spiking activity across the auditory cortical hierarchy enables decoding of past stimuli long after their offset, during both wakefulness and sleep. Time-resolved decoding revealed that in wakefulness, persistent representations decay uniformly across sensory and association cortices, whereas during sleep, persistence is prolonged in association cortex but remains brief in early auditory regions. Recurrent neural network modeling showed that higher internal noise during wakefulness reproduces this pattern, suggesting that reduced interference during sleep stabilizes sensory traces in associative areas. Our results demonstrate that persistent representation is a passive, state-dependent feature of sensory processing, supporting sensory maintenance even in the absence of active engagement.
Summary Speech perception relies on two fundamental acoustic components: spectral and temporal. While spectral information is known to be represented in the auditory cortex through tonotopy, how temporal features are organized has remained unclear. Here, by varying onset rise-ramp steepness and frequencies, we reveal that temporal envelope steepness—a critical cue for phoneme discrimination and sound source perception—is systematically mapped in the mouse auditory cortex. Using widefield calcium imaging, we discovered that the envelope steepness is represented orthogonally to the tonotopic axis, forming a two-dimensional cortical map that mirrors the dual structure of sounds. This organization was observed in primary-like areas but not in higher-order-like areas, indicating distinct auditory processing streams. These findings uncover a principle of cortical organization, suggesting that the auditory cortex encodes sound along two independent axes and thereby provides a neural basis for the parallel processing of complex sounds such as speech and natural acoustic environments.
This paper presents a method to use t-SNE for training low-dimensional timbre maps obtained from musical recordings in a way that similar sound spectra are represented at similar location within the map. To use such maps for generating novel sounds and timbre trajectories, we use Kernel Regression Mapping for the inverse transformation from map space to timbre space and Griffin-Lim algorithm for phase reconstruction. With this mix of methods, we achieve a way to both visually explore timbral patterns and compress control data. As an application for navigable timbre maps we introduce a novel method – related to and inspired from Wave Space Sonification (WSS) – for the auditory exploration of patterns in multivariate time-series data by using the semantic sound map as the connecting representation. We demonstrate our approach by sonifying ECG data relating to cardiac pathologies.
No abstract available
Objective: Selective auditory attention decoding (AAD) algorithms process brain data such as electroencephalography to decode to which of multiple competing sound sources a person attends. Example use cases are neuro-steered hearing aids or communication via brain-computer interfaces (BCI). Recently, it has been shown that it is possible to train such AAD decoders based on stimulus reconstruction in an unsupervised setting, where no ground truth is available regarding which sound source is attended. In many practical scenarios, such ground-truth labels are absent, making it, moreover, difficult to quantify the accuracy of the decoders. In this paper, we aim to develop a completely unsupervised algorithm to estimate the accuracy of correlation-based AAD algorithms during a competing talker listening task. Methods: We use principles of digital communications by modeling the AAD decision system as a binary phase-shift keying channel with additive white gaussian noise. Results: We show that the proposed unsupervised performance estimation technique can accurately determine the AAD accuracy in a transparent-for-the-user way, for different amounts of training and estimation data and decision window lengths. Furthermore, since different applications demand different targeted accuracies, our approach can estimate the minimal amount of training required for any given target accuracy. Conclusion: Our proposed estimation technique accurately predicts the performance of a correlation-based AAD algorithm without access to ground-truth labels. Significance: In neuro-steered hearing aids, the accuracy estimates provided by our approach could support time-adaptive decoding, dynamic gain control, and neurofeedback. In BCIs, it could support a robust communication paradigm with accuracy feedback for caregivers.
In daily life, we effortlessly switch attention between different sound sources while filtering out distractions. Prior research has shown that the target speaker can be decoded from EEG in such settings, however, how the switching of attention influences decoding has not been measured. We investigated EEG-based target speaker decoding with exogenous attention switches in a multi-talker environment, where twenty participants alternated listening to two target speakers taking turns while ignoring two background speakers. Speech envelope reconstruction from EEG showed higher accuracy for the attended speaker than distractors, with single-trial classification at 76%. Our findings showed a brief increase in target speaker reconstruction accuracy after attention switches, suggesting heightened alertness when shifting attention to a new speaker. These findings enhance the understanding of dynamic auditory attention and support the development of attention-based hearing applications.
Auditory attention detection (AAD) aims to identify the sound source to which a listener is attending based on neural signals. In this work, AAD is investigated using electroencepha-lography (EEG), an approach widely used in research toward neuro-steered hearing devices for complex acoustic environments. While most AAD studies focus on non-tonal languages, tonal languages remain less studied despite constituting more than half of human languages. Here, Thai, a tonal language with five distinct pitch tones, is used as speech material to investigate EEG-based AAD. EEG signals were recorded using an 8-channel configuration from 30 normal-hearing subjects while listening to Thai speech presented through two-channel over-ear headphones. A linear stimulus reconstruction (LSR) decoder utilizing both EEG and speech achieved above-chance accuracy across decision windows ranging from 1 to 60 s, while a convolutional neural network (CNN) relying solely on EEG achieved comparable or superior decoding accuracy at the extended 60 s decision window. This work provides EEG data for studying AAD in tonal languages.
Background Phantom auditory percepts are especially prevalent in healthy persons with hearing loss. No first-person description of the not uncommon illusion called musical hallucinosis (MH) has been published in relation to possible neural mechanisms for its occurrence. Objectives The author presents his personal experience following implantation of a unilateral cochlear neuroprosthesis to try to compensate for progressive sensorineural hearing loss. Results The MH included abrupt onset of persistent, robust singing of the Star-Spangled Banner, then other familiar songs and nursery rhymes by a men’s choir without accompanying instrumentals, followed months later by continuous nonsense lyrics sung to a simpler stereotyped tune. The onset was associated with deafness as a complication of electrode placement within the cochlea, the early sizzling, synthetic, monotonal auditory sounds heard using the cochlear implant, and a burst of cacophonous tinnitus following a higher volume adjustment to the device. Conclusions Several physiological alterations, including deafferentation-induced spontaneous auditory pathway activity that triggers higher auditory cortical areas to place the ambiguous inputs within the individual’s prior experience of sound patterns, may help explain the evolution of MH and its persistence as a type of maladaptive neuroplasticity.
Neural information is represented through dynamic population coding in the brain. Decoding these patterns is not only crucial for understanding brain function, but also holds broad application value. This study incorporated the emerging Spin-Exchange Relaxation-Free Magnetoencephalography (SERF-MEG) technique with multivariate pattern analysis (MVPA) to investigate the dynamic neural representation of natural auditory stimuli. The results demonstrated that SERF-MEG’s high spatiotemporal resolution, when combined with MVPA, could capture the dynamic time course of neural representational similarity in the brain. The comparison with electroencephalogram (EEG) further revealed the difference in sensitivity between SERF-MEG and EEG during multivariate decoding. The source localization results showed that complex sound processing involves distributed dynamic neural activities across primary auditory cortex, higher auditory cortex, and memory-related areas. This study not only provides a novel multimodal analytical framework for decoding dynamic brain representations under complex stimuli but also expands the application potential of SERF-MEG in cognitive neuroscience research.
The human brain's response to video stimuli involves the engagement of numerous sensory modalities, including visual and auditory systems. Nevertheless, the neural underpinnings governing the processing of video stimuli remain incompletely understood. One critical factor contributing to this knowledge gap is the scarcity of corresponding human electrophysiology datasets, particularly those collected from deep brain regions. This work presented an intracranial EEG (iEEG) dataset acquired from 66 electrodes in an epilepsy patient over the course of 300 minutes while they viewed customized video stimuli. The dataset encompassed a wide range of cortical and subcortical brain regions, with the goal of enhancing our comprehension of the perceptual processing mechanisms involved during video observation. Based on this dataset, we employed a classification experiment to assess the significance of brain regions involved in the processing of diverse videos, auditory stimuli, and visual scenes within the human brain.
The growing availability of publicly shared EEG datasets involving audio, video, and linguistic stimuli offers great potential for advancing multi-modal neural decoding research. However, significant heterogeneity in file formats, annotation structures, and temporal alignment across datasets has hindered their direct usability and cross-dataset benchmarking. In this study, we present a unified preprocessing pipeline applied to six diverse EEG datasets that cover auditory attention, language comprehension, and visual perception. The proposed pipeline standardizes EEG representations and multimodal annotations, allowing the resulting data to be aligned in time and directly usable for large-scale modeling tasks. We summarize the resulting unified database and highlight its consistency in structure and format. To demonstrate usability, we provide benchmark classification results on label-eligible datasets using both traditional methods (rLDA) and deep learning (EEGNet). The resulting resource enables consistent multimodal EEG analyses and provides a structured basis for advancing research on cross-modal alignment and neural decoding. The accompanying benchmark experiments serve to illustrate the usability of the standardized datasets; achieving competitive performance was not the aim of these evaluations.
Error-related potentials (ErrPs) are key neural signatures for monitoring performance in brain-computer interfaces (BCIs), yet their reliability is often constrained by weak signal strength and large inter-subject variability. To examine whether multisensory alignment can enhance ErrP detection, we designed a multimodal BCI paradigm in which participants passively observed an autonomous agent navigating a maze while receiving congruent or incongruent combinations of visual, auditory, and tactile feedback. Electroencephalogram analysis revealed canonical N200-P300 components, with congruent conditions eliciting higher classification accuracy. A cross subjects, congruent stimuli improved ErrP decoding accuracy by 3.89 % and reduced inter-subject variance relative to incongruent feedback. Feature visualization using t-SNE showed that congruent and incongruent responses formed clearly separable clusters, indicating distinct neural representations modulated by cross- modal correspondence. These findings suggest that aligning sensory modalities may enhance the discriminability and stability of ErrP patterns in passive BCI settings, and should be interpreted as a preliminary proof-of-concept.
Brain decoding aims to reconstruct external stimuli from brain activity, providing insights into the neural representation of cognitive experiences. Music decoding from functional magnetic resonance imaging (fMRI) is particularly challenging due to the complexity of auditory processing and the temporal limitations of fMRI signals. In this study, we introduce a novel decoding framework that improves the alignment between fMRI activity and latent musical representations extracted using a pre-trained multimodal model (CLAP). We propose a dual-loss approach combining Optimal Transport and Contrastive Learning to enhance feature mapping and retrieval accuracy. The first loss ensures structural consistency between brain-predicted and true musical embeddings, while the contrastive loss refines the embedding space by maximizing similarities between corresponding pairs and minimizing non-correspondences. Using fMRI data from five subjects listening to music tracks from the GTZAN dataset, our method achieves improved decoding performance, surpassing traditional regression-based approaches from 22.1% top-1 accuracy to 29.3%. These results highlight the potential of integrating Optimal Transport and Contrastive Learning to improve brain decoding performance, paving the way for extending the approach to different sensory domains and applications in Brain-Computer Interfaces (BCI).Clinical relevance— This study could have clinical implications for understanding auditory processing disorders and developing neurorehabilitation strategies. By elucidating how the brain encodes complex auditory stimuli, this approach may contribute to BCI applications for speech and music perception restoration in individuals with hearing impairments or neurological conditions affecting auditory cognition.
The YOTO (You Only Think Once) dataset presents a human electroencephalog- raphy (EEG) resource for exploring multisensory perception and mental imagery. The study enrolled 26 participants who performed tasks involving both unimodal and multimodal stimuli. Researchers collected high-resolution EEG signals at a 1000 Hz sampling rate to capture high-temporal-resolution neural activity related to internal mental representations. The protocol incorporated visual, auditory, and combined cues to investigate the integration of multiple sensory modalities, and participants provided self-reported vividness ratings that indicate subjec- tive perceptual strength. Technical validation involved event-related potentials (ERPs) and power spectral density (PSD) analyses, which demonstrated the reli- ability of the data and confirmed distinct neural responses across stimuli. This dataset aims to foster studies on neural decoding, perception, and cognitive mod- eling, and it is publicly accessible for researchers who seek to advance multimodal mental imagery research and related applications.
The YOTO (You Only Think Once) dataset presents a human electroencephalography (EEG) resource for exploring multisensory perception and mental imagery. The study enrolled 26 participants who performed tasks involving both unimodal and multimodal stimuli. Researchers collected high-resolution EEG signals at a 1000 Hz sampling rate to capture high-temporal-resolution neural activity related to internal mental representations. The protocol incorporated visual, auditory, and combined cues to investigate the integration of multiple sensory modalities, and participants provided self-reported vividness ratings that indicate subjective perceptual strength. Technical validation involved event-related potentials (ERPs) and power spectral density (PSD) analyses, which demonstrated the reliability of the data and confirmed distinct neural responses across stimuli. This dataset aims to foster studies on neural decoding, perception, and cognitive modeling, and it is publicly accessible for researchers who seek to advance multimodal mental imagery research and related applications.
Decoding visual and auditory stimuli from brain activities, such as electroencephalography (EEG), offers promising advancements for enhancing machine-to-human interaction. However, effectively representing EEG signals remains a significant challenge. In this paper, we introduce a novel Delayed Knowledge Transfer (DKT) framework that employs spiking neurons for attention detection, using our experimental EEG dataset. This framework extracts patterns from audiovisual stimuli to model brain responses in EEG signals, while accounting for inherent response delays. By aligning audiovisual features with EEG signals through a shared embedding space, our approach improves the performance of brain-computer interface (BCI) systems. We also present WithMeAttention, a multimodal dataset designed to facilitate research in continuously distinguishing between target and distractor responses. Our methodology demonstrates a 3% improvement in accuracy on the WithMeAttention dataset compared to a baseline model that decodes EEG signals from scratch. This significant performance increase highlights the effectiveness of our approach Comprehensive analysis across four distinct conditions shows that rhythmic enhancement of visual information can optimize multi-sensory information processing. Notably, the two conditions featuring rhythmic target presentation - with and without accompanying beeps - achieved significantly superior performance compared to other scenarios. Furthermore, the delay distribution observed under different conditions indicates that our delay layer effectively emulates the neural processing delays in response to stimuli.
High-bandwidth brain–computer interfaces rely on invasive surgical procedures or brain-penetrating electrodes. Here we describe a cortical 1,024-channel thin-film microelectrode array and we demonstrate its minimally invasive surgical delivery that avoids craniotomy in porcine models and cadavers. We show recording and stimulation from the same electrodes to large portions of the cortical surface, and the reversibility of delivering the implants to multiple functional regions of the brain without damaging the cortical surface. We evaluate the performance of the interface for high-density neural recording and visualizing cortical surface activity at spatial and temporal resolutions and total spatial extents. We demonstrate accurate neural decoding of somatosensory, visual and volitional walking activity, and achieve focal neuromodulation through cortical stimulation at sub-millimetre scales. We report the feasibility of intraoperative use of the device in a five-patient pilot clinical study with anaesthetized and awake neurosurgical patients, characterizing the spatial scales at which sensorimotor activity and speech are represented at the cortical surface. The presented neural interface demonstrates the highly scalable nature of micro-electrocorticography and its utility for next-generation brain–computer interfaces. A 1,024-channel microelectrode array is delivered to the brain cortex via a minimally invasive incision in the skull and dura, and allows recording, stimulation and neural decoding across large portions of the brain in porcine models and human neurosurgical patients.
EEG-based neural decoding requires large-scale benchmark datasets. Paired brain-language data across speaking, listening, and reading modalities are essential for aligning neural activity with the semantic representation of large language models (LLMs). However, such datasets are rare, especially for non-English languages. Here, we present ChineseEEG-2, a high-density EEG dataset designed for benchmarking neural decoding models under real-world language tasks. Building on our previous ChineseEEG dataset, which focused on silent reading, ChineseEEG-2 adds two active modalities: Reading Aloud (RA) and Passive Listening (PL), using the same Chinese corpus. EEG and audio were simultaneously recorded from four participants during ~10.8 hours of reading aloud. These recordings were then played to eight other participants, collecting ~21.6 hours of EEG during listening. This setup enables precise temporal and semantic alignment across the RA and PL modalities. ChineseEEG-2 includes EEG signals, speech audio, aligned semantic embeddings from pre-trained language models, and task labels. Together with ChineseEEG, this dataset supports joint semantic alignment learning across speaking, listening, and reading. It enables benchmarking of neural decoding algorithms and promotes brain-LLM alignment under multimodal language tasks, especially in Chinese. ChineseEEG-2 provides a benchmark dataset for next-generation neural semantic decoding.
The brain represents information through the encoding of neural populations, where the activity patterns of these neural groups constitute the content of this information. Understanding these activity patterns and their dynamic changes is of significant importance to cognitive neuroscience and related research areas. Current studies focus more on brain regions that show differential responses to stimuli, but they lack the ability to capture information about the representational or process‐level dynamics within these regions. In this study, we recorded neural data from 10 healthy participants during auditory experiments using optically pumped magnetometer magnetoencephalography (OPM‐MEG) and electroencephalography (EEG). We constructed representational similarity matrices (RSMs) to investigate the similarity of neural response patterns during auditory decoding. The results indicate that RSA can reveal the dynamic changes in pattern similarity during different stages of auditory processing through the neural activity patterns reflected by OPM‐MEG. Comparisons with EEG results showed that both techniques captured the same processes during the early stages of auditory decoding. However, differences in sensitivity at later stages highlighted both common and distinct aspects of neural representation between the two modalities. Further analysis indicated that this process involved widespread neural network activation, including the Heschl's gyrus, superior temporal gyrus, middle temporal gyrus, inferior temporal gyrus, parahippocampal gyrus, and orbitofrontal gyrus. This study demonstrates that the combination of OPM‐MEG and RSA is sufficiently sensitive to detect changes in pattern similarity during neural representation processes and to identify their anatomical origins, offering new insights and references for the future application of RSA and other multivariate pattern analysis methods in the MEG field.
Individuals have the remarkable ability to differentiate between speakers and focus on a particular speaker, even amidst complex acoustic environments with multiple speakers, background noise and reverberations. This selective auditory attention, often illustrated by the cocktail party problem, has been extensively researched. With a considerable portion of the population experiencing hearing impairment and requiring hearing aids, there arises a necessity to separate and decode auditory signals artificially. The linearly constrained minimum variance (LCMV) beamforming design criterion has proven effective in isolating the desired source by steering a beam toward the target speaker while creating a null toward the interfering source. Preserving the binaural cues, e.g., interaural time difference (ITFD) and interaural level difference (ILD), is a prerequisite for producing a beamformer output suitable for hearing aid applications. For that, the binaural linearly constrained minimum variance (BLCMV) beamformer generates two outputs that satisfy the standard LCMV criterion while preserving the binaural cues between the left-ear and right-ear outputs. Identifying the attended speaker from the separated speakers and distinguishing it from the unattended speaker poses a fundamental challenge in the beamformer design. Several studies showed the ability to encode essential features of the attended speech from the cortex neural response, as recorded by the electroencephalography (EEG) signals. This led to the development of several algorithms addressing the auditory attention decoder (AAD) task. This paper investigates two neural network architectures for the AAD task. The first architecture leverages transfer learning. It is evaluated using both same-trial and cross-trial experiments. The second architecture employs an attention mechanism between the speech signal represented in the short time Fourier transform (STFT) domain and a multi-band filtered EEG signal. With the goal of alleviating the problem of same-trial overfitting, this architecture employs a new data organization structure that presents the neural network (NN) with a single speaker’s speech and the corresponding EEG signal as inputs. Finally, posterior probability post-processing is applied to the outputs of the NN to improve detection accuracy. The experimental study validates the applicability of the proposed scheme as an AAD method. Strategies for incorporating the AAD into BLCMV beamformer are discussed.
Introduction In our complex world, the auditory system plays a crucial role in perceiving and processing our environment. Humans are able to segment and stream concurrent auditory objects, allowing them to focus on specific sounds, such as speech, and suppress irrelevant auditory objects. The attentional enhancement or suppression of sound processing is evident in neural data through a phenomenon called neural speech tracking. Previous studies have identified correlates of neural speech tracking in electroencephalography (EEG) data, but EEG measures are susceptible to motion artefacts, and the association between neural data and auditory objects is vulnerable to distraction. Methods The current study investigated EEG-based auditory attention decoding in realistic everyday scenarios. N=20 participants were exposed to the sound of a busy cafeteria or walked along busy and quiet streets while listening to one or two simultaneous speech streams. We also investigated the robustness of neural speech tracking estimates within subjects. Linear decoding models were used to determine the magnitude of neural speech tracking. Results The results confirmed that neural speech tracking was strongest in single speaker scenarios. In dual speaker conditions, there was significantly stronger neural speech tracking for the attended speaker compared to the ignored speaker, even in complex environments such as a busy cafeteria or outdoor settings. Discussion In conclusion, EEG-based attention decoding is feasible in highly complex and realistic everyday conditions while humans behave naturally.
No abstract available
Listeners effortlessly extract multi-dimensional auditory objects, such as a localized talker, from complex acoustic scenes. However, the neural mechanisms that enable simultaneous encoding and linking of distinct sound features—such as a talker’s voice and location—are not fully understood. Using invasive intracranial recordings in seven neurosurgical patients (four male, three female), we investigated how the human auditory cortex processes and integrates these features during naturalistic multi-talker scenes and how attentional mechanisms modulate such feature integration. We found that cortical sites exhibit a continuum of feature sensitivity, ranging from single-feature-sensitive sites (responsive primarily to voice spectral features or to location features) to dual-feature-sensitive sites (responsive to both features). At the population level, neural response patterns from both single- and dual-feature-sensitive sites jointly encoded the attended talker’s voice and location. Notably, single-feature-sensitive sites encoded their primary feature with greater precision but also represented coarse information about the secondary feature. Sites selectively tracking a single, attended speech stream concurrently encoded both voice and location features, demonstrating a link between selective attention and feature integration. Additionally, attention selectively enhanced temporal coherence between voice- and location-sensitive sites, suggesting that temporal synchronization serves as a mechanism for linking these features. Our findings highlight two complementary neural mechanisms—joint population coding and temporal coherence—that enable the integration of voice and location features in the auditory cortex. These results provide new insights into the distributed, multi-dimensional nature of auditory object formation during active listening in complex environments.
Neurons in the auditory system must represent behaviorally relevant sounds in the presence of background noise (BN) to support noise-invariant perception and behavior. Although the primary auditory cortex (ACtx) has been implicated in constructing noise-invariant representations, it remains unclear which excitatory subpopulations within ACtx carry out this transformation from noise-dependent to noise-invariant coding. To address this, we presented pure tones with and without continuous BN to head-fixed mice and used two-photon calcium imaging to record sound-evoked activity from three major excitatory subpopulations in ACtx: layer (L)2/3 intratelencephalic (IT) neurons, L5 IT neurons, and L5 extratelencephalic (ET) neurons. L2/3 IT neurons exhibited strong noise dependence at the level of single-neuron responses, pairwise interactions, and population representations. In contrast, deep-layer pathways showed greater noise invariance, with L5 IT neurons preserving stable representations most consistently and L5 ET neurons exhibiting more limited invariance at the population level. These findings reveal a functional division of labor in ACtx, in which superficial neurons remain noise-dependent and deep-layer broadcast pathways, particularly L5 IT, preferentially carry noise-invariant representations, suggesting that excitatory subpopulations contribute differentially to the construction and propagation of noise-invariant codes.
Recent studies show that the classical model based on axonal delay-lines may not explain interaural time difference (ITD) based spatial coding in humans. Instead, a population-code model called "opponent channels model" (OCM) has been suggested. This model comprises two competing channels respectively for the two auditory hemifields, each with a sigmoidal tuning curve. Event-related potentials (ERPs) to ITD-changes are used in some studies to test the predictions of this model by considering the sounds before and after the change as adaptor and probe stimuli, respectively. It is assumed in these studies that the former stimulus causes adaptation of the neurons selective to its side, and that the ERP N1-P2 response to the ITD-change is the specific response of the neurons with selectivity to the side of probe sound. However, these ERP components are known as a global, non-specific acoustic change complex of cortical origin evoked by any change in the auditory environment. It probably does not genuinely reflect the activity of some stimulus-specific neuronal units that have escaped the refractory effect of the preceding adaptor, which means a violation of the crucial assumption in an adaptor-probe paradigm. To assess this viewpoint, we conducted two experiments. In the first one, we recorded ERPs to abrupt lateralization shifts of click trains having various pre- and post-shift ITDs within the physiological range of -600μ s to +600μ s. Magnitudes of the ERP components P1, N1, and P2 to these ITD-shifts did not comply with the additive behavior of partial probe responses presumed for an adaptor-probe paradigm, casting doubt on the accuracy of testing sensory coding models by using ERPs to abrupt lateralization changes. Findings of the second experiment, involving ERPs to conjoint outwards/transverse shift stimuli also supported this conclusion.
No abstract available
The inferior colliculus (IC) of the midbrain is important for complex sound processing, such as discriminating conspecific vocalizations and human speech. The IC's non-lemniscal, dorsal "shell" region is likely important for this process, as neurons in these layers project to higher-order thalamic nuclei that subsequently funnel acoustic signals to the amygdala and non-primary auditory cortices; forebrain circuits important for vocalization coding in a variety of mammals, including humans. However, the extent to which shell IC neurons transmit acoustic features necessary to discern vocalizations is less clear, owing to the technical difficulty of recording from neurons in the IC's superficial layers via traditional approaches. Here we use 2-photon Ca2+ imaging in mice of either sex to test how shell IC neuron populations encode the rate and depth of amplitude modulation, important sound cues for speech perception. Most shell IC neurons were broadly tuned, with a low neurometric discrimination of amplitude modulation rate; only a subset were highly selective to specific modulation rates. Nevertheless, neural network classifier trained on fluorescence data from shell IC neuron populations accurately classified amplitude modulation rate, and decoding accuracy was only marginally reduced when highly tuned neurons were omitted from training data. Rather, classifier accuracy increased monotonically with the modulation depth of the training data, such that classifiers trained on full-depth modulated sounds had median decoding errors of ~0.2 octaves. Thus, shell IC neurons may transmit time-varying signals via a population code, with perhaps limited reliance on the discriminative capacity of any individual neuron.
Cochlear implants have transformed the treatment of hearing loss by enabling auditory perception through direct electrical stimulation of the auditory nerve. However, their effectiveness can be limited in noisy environments, for fine frequency resolution, and in patients lacking an intact auditory nerve. The auditory cortex offers a promising alternative target for neuroprosthetic stimulation, but it remains unclear whether direct cortical input can evoke percepts with the complexity and structure of natural sounds. Here we show that surface cortical implants in mice support robust and flexible auditory behaviour, surpassing cochlear implants in tasks requiring fine spectral, temporal, and noise-resistant discrimination. We also show that animals generalize seamlessly between cortical and acoustic stimuli without additional training, indicating that cortical stimulation can generate interpretable, sound-like percepts. Electrophysiological recordings revealed that cortical stimulation drives spatially and temporally structured neural activity in auditory cortex, resembling responses to natural sound. These findings establish that the auditory cortex can interpret spatially and temporally patterned electrical input in a behavioural relevant way. By linking neuroprosthetic stimulation to naturalistic neural representations and behavioural generalization, this work provides a mechanistic and functional foundation for cortical auditory neuroprostheses and points toward new strategies for restoring hearing through direct brain stimulation.
Sound localization is critical for real-world hearing, such as segregating overlapping sound streams. For optimal flexibility, central representations of auditory space must adapt to peripheral changes in binaural cue availability, such as following asymmetric hearing loss in adulthood. However, whether the mature auditory system can reliably encode spatial auditory representations upon abrupt changes in binaural input is unclear. Here we use 2-photon Ca2+ imaging in awake head-fixed mice to determine how the higher-order "shell" layers of the inferior colliculus (IC) encode sound source location in the frontal azimuth, under binaural conditions and after acute monaural hearing loss induced by an ear plug ipsilateral to the imaged hemisphere. Spatial receptive fields were typically broad and not exclusively contralateral: Neurons responded reliably to multiple positions in the contra- and ipsi-lateral hemifields, with preferred positions tiling the entire frontal azimuth. Ear plugging broadened receptive fields and reduced spatial selectivity in a subset of neurons, in agreement with an inhibitory influence of ipsilateral sounds. However ear plugging also enhanced spatial tuning and/or unmasked receptive fields in other neurons, shifting the distribution of preferred angles ipsilaterally with minimal impact on the neuronal population’s overall spatial resolution; these effects occurred within 2 hours of ear plugging. Consequently, linear classifiers trained on fluorescence data from control and ear-plugged conditions had similar classification accuracy when tested on held out data from within, but not across hearing conditions. Spatially informative neuronal population codes therefore arise rapidly following monaural hearing loss, in absence of overt experience.
During wake, sound‐evoked and spontaneous neural activity of the auditory cortex evolves in distinct subspaces, whereas anesthesia disrupts sound responses and merges these spaces. To evaluate if similar modifications of sound representation geometry explain sensory disconnection during sleep, large neural populations of the mouse auditory cortex are followed across slow‐wave sleep and wakefulness. It is observed that sleep dampens sound responses but preserves the geometry of sound representations such that they remain separate from spontaneous activity. Moreover, response dampening is strongly coordinated across neurons and varied throughout sleep, spanning from fully preserved response patterns to population response failures on a fraction of sound presentations. These failures rarely occurred in wakefulness and are more common during high spindle‐band activity. Therefore, in sleep, the auditory system preserves sound feature selectivity up to the cortex for detailed acoustic surveillance but concurrently implements an intermittent gating mechanism leading to local sensory disconnections.
During wake, sound-evoked and spontaneous neural activity of the auditory cortex evolve in distinct subspaces whereas anesthesia disrupts sound responses and merges these spaces. To evaluate if similar modifications of the sound representation geometry explain sensory disconnection during sleep, we followed large neural populations of the mouse auditory cortex across slow wave sleep and wakefulness. We observed that sleep dampens sound responses but preserves the geometry of sound representations which remain separate from spontaneous activity. Moreover, response dampening was strongly coordinated across neurons and varied throughout sleep spanning from fully preserved response patterns to population response failures on a fraction of sound presentations. These failures are more common during high spindle-band activity and more rarely observed in wakefulness. Therefore, in sleep, the auditory system preserves sound feature selectivity up to the cortex for detailed acoustic surveillance, but concurrently implements an intermittent gating mechanism leading to local sensory disconnections.
Understanding how the brain organises natural categories is a central challenge in neuroscience. While prior work has shown that categories can be decoded from distributed activity patterns in auditory cortex, it remains unclear how these categories are globally arranged relative to one another, and how low-level acoustic and higher-level semantic structure jointly shape this organisation. Here, we addressed these questions by deriving low-dimensional functional gradients from high-depth functional magnetic resonance imaging (fMRI) data (three participants, ∼4.7 hours each) acquired during a category-specific one-back task. These gradients captured the principal axes of population activity in auditory cortex. Gradient-based models of the auditory cortex explained category structure more accurately than region-of-interest or whole-brain approaches, revealing that category information is distributed across multiple continuous axes rather than aligned with any single organisational dimension. Projecting acoustic (gammatone filter-bank) and behavioural similarity spaces directly into a shared framework with the fMRI functional axes showed that both contribute to the brain’s category geometry, with acoustic structure exerting a somewhat stronger influence. However, representational relationships varied across category pairs: some reflected primarily acoustic similarity, others semantic distinctions, and many a combination of both. This pairwise heterogeneity shows how auditory cortex may integrate multiple representational dimensions that define higher-level categories. Significance Statement Categorising natural sounds requires the brain to transform diverse acoustic signals into meaningful concepts. How this is achieved remains unclear: prior work has shown distributed activation patterns in auditory cortex, but not how category relationships are organised within its functional architecture. We show that natural sound categories are embedded across continuous cortical gradients that integrate bottom-up acoustic structure with top-down semantic information. Different categories emerge as flexible combinations of these spectral and behavioural dimensions. These findings reveal that natural sound categorisation arises from the geometry of cortical organisation, moving beyond localist accounts and offering a new framework for understanding how perception and cognition are linked in the human auditory system.
According to theories of the brain as a predictive network, perceptual decisions result from integrating incoming sensory inputs with prior experience. A neuronal population implementing this form of computation must not only retain information about past events, but also combine this information with current sensory evidence. To examine how this integration occurs, we recorded extracellular activity from both primary sensory cortex and a target frontal region in rats performing stimulus categorization. Psychophysical analysis showed that judgments were history-dependent. Sensory cortex represented the current stimulus largely independently of prior stimuli and failed to account for the history-dependence seen in behavior. By contrast, frontal cortex embedded current input within a representation of prior sensory information through collinearity in coding dimensions. This mechanism – predominantly mediated by fast-spiking neurons – explained trial-to-trial variability in decisions. Our findings argue for distinct roles of cortical regions in predictive processing, and identify a frontal stage where current and prior sensory information converge to inform decisions.
Associative system has attracted increasing attention for it can store basic information and then infer details to match perception with an efficient self-organization algorithm. However, the implementation of the associative system with the application of real-world data is relatively difficult. To address this issue, we propose a novel biologically inspired auto-associative (BIAA) network to explore the structure, encoding and formation of associative memory as well as to extend the ability to real-world application. Our network is constructed by imitating the organization of the cortical minicolumns where each minicolumn contains plenty of parallel biological spiking neurons. To allow the network to learn and predict one symbol per theta cycle, we incorporate synaptic delay and theta oscillation into the neuron dynamic process. Subsequently, we design a sparse temporal population (STP) coding scheme that allows each input symbol to be represented as stable, unique, and easily recallable sparsely distributed representations. By combining associative learning dynamics with the STP coding, our network realizes efficient storage and inference in an ordered manner. Experimental results indicate that the proposed network successfully performs sequence retrieval from partial text and sequence recovery from distorted information. BIAA network provides new insight into introducing biologically inspired mechanisms into associative system and has enormous potential for hardware and software applications.
Everyday environments often contain multiple concurrent sound sources that fluctuate over time. Normally hearing listeners can benefit from high signal-to-noise ratios (SNRs) in energetic dips of temporally fluctuating background sound, a phenomenon called dip-listening. Specialized mechanisms of dip-listening exist across the entire auditory pathway. Both the instantaneous fluctuating and the long-term overall SNR shape dip-listening. An unresolved issue regarding cortical mechanisms of dip-listening is how target perception remains invariant to overall SNR, specifically, across different tone levels with an ongoing fluctuating masker. Equivalent target detection over both positive and negative overall SNRs (SNR invariance) is reliably achieved in highly-trained listeners. Dip-listening is correlated with the ability to resolve temporal fine structure, which involves temporally-varying spike patterns. Thus the current work tests the hypothesis that at negative SNRs, neuronal readout mechanisms need to increasingly rely on decoding strategies based on temporal spike patterns, as opposed to spike count. Recordings from chronically implanted electrode arrays in core auditory cortex of trained and awake Mongolian gerbils that are engaged in a tone detection task in 10 Hz amplitude-modulated background sound reveal that rate-based decoding is not SNR-invariant, whereas temporal coding is informative at both negative and positive SNRs.
合并后形成十一条相互并列的研究主线:首先以听觉皮层拓扑、分层结构和群体编码揭示可再生的神经表征维度;其次分别讨论选择性注意、复杂声场中的对象绑定、经验与神经调制驱动的可塑性,以及ITD空间侧化等关键机制;随后以神经刺激和行为状态动态重配置补充闭环调控基础。在技术层面,研究进一步覆盖多模态脑信号数据与跨模态对齐、非侵入式语音和自然声音解码重建;在转化层面,涵盖神经假体、听力恢复、空间听觉重建及异常可塑性风险;最后以生物启发式编码和语义声音地图提供模型化与系统设计先验。重复出现的dynamic_representation_of_sound_locations_during仅归入神经刺激与行为状态驱动的动态重配置组,以保证分组互不交叉。