Jing Chen 0019

dblp:27/4364-19 · DBLP profile ↗
← Back
19ranked-venue papers
1as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 4 since 2021
YearPublicationVenuePosition
2025 A Novel Multimodal Method for Decoding Speech Perception from Brain Activities
abstract
Decoding speech from neural recordings has critical importance in application and scientific research. However, this task is still challenging with non-invasive recordings. Previous research has shown significant improvement in speech perception decoding task by leveraging wav2vec vectors and gives the potential for applications. To further explore this problem, we proposed a novel multimodal method by using functional magnetic resonance imaging (fMRI) and magnetoencephalography (MEG). In our method, separate encoders for fMRI and MEG are considered, then features extracted from both modalities are integrated and aligned with wav2vec vectors that were extracted from the speech. The multimodal method reaches averaged performance of 72.6% in top-10 accuracy with a negative sample size of 128. Performance evaluated with various metrics achieves steady improvement across subjects, demonstrating the effectiveness of the proposed data fusion method. Interpretation of the performance increment was also investigated by testing the correlation between encoder hidden outputs and different level of features extracted from the speech. Results demonstrate that MEG encoder learns more low-level information and fMRI encoder learns more high-level information, which indicates both complementary characteristics lead to the improvement. The result of this work shows the potential of multimodal methods for speech decoding.
Aoke Zhang, Bo Wang 0110, Xihong Wu, Jing Chen 0019
ICASSP4
2025 Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
abstract
Auditory Attention Decoding (AAD) can help to determine the identity of the attended speaker during an auditory selective attention task, by analyzing and processing measurements of electroencephalography (EEG) data. Most studies on AAD are based on scalp-EEG signals in two-speaker scenarios, which are far from real application. Ear-EEG has recently gained significant attention due to its motion tolerance and invisibility during data acquisition, making it easy to incorporate with other devices for applications. In this work, participants selectively attended to one of the four spatially separated speakers’ speech in an anechoic room. The EEG data were concurrently collected from a scalp-EEG system and an ear-EEG system (cEEGrids). Temporal response functions (TRFs) and stimulus reconstruction (SR) were utilized using ear-EEG data. Results showed that the attended speech TRFs were stronger than each unattended speech and decoding accuracy was 41.3% in the 60s (chance level of 25%). To further investigate the impact of electrode placement and quantity, SR was utilized in both scalp-EEG and ear-EEG, revealing that while the number of electrodes had a minor effect, their positioning had a significant influence on the decoding accuracy. One kind of auditory spatial attention detection (ASAD) method, STAnet, was testified with this ear-EEG database, resulting in 93.1% in 1-second decoding window. The implementation code and database for our work are available on GitHub: https://github.com/zhl486/Ear_EEG_code.git and Zenodo: https://zenodo.org/records/10803261.
Haolin Zhu, Yujie Yan, Xiran Xu, Zhongshu Ge, Pei Tian, Xihong Wu, Jing Chen 0019
ICASSP7
2025 Overestimated performance of auditory attention decoding caused by experimental design in EEG recordings
Yujie Yan, Xiran Xu, Haolin Zhu, Songyi Li, Bo Wang 0110, Xihong Wu, Jing Chen 0019
INTERSPEECH7
2024 Semantic Reconstruction of Continuous Language from Meg Signals
abstract
Decoding language from neural signals holds considerable theoretical and practical importance. Previous research has indicated the feasibility of decoding text or speech from invasive neural signals. However, when using non-invasive neural signals, significant challenges are encountered due to their low quality. In this study, we proposed a data-driven approach for decoding semantic of language from Magnetoencephalography (MEG) signals recorded while subjects were listening to continuous speech. First, a multi-subject decoding model was trained using contrastive learning to reconstruct continuous word embeddings from MEG data. Subsequently, a beam search algorithm was adopted to generate text sequences based on the reconstructed word embeddings. Given a candidate sentence in the beam, a language model was used to predict the subsequent words. The word embeddings of the subsequent words were correlated with the reconstructed word embedding. These correlations were then used as a measure of the probability for the next word. The results showed that the proposed continuous word embedding model can effectively leverage both subject-specific and subject-shared information. Additionally, the decoded text exhibited significant similarity to the target text, with an average BERTScore of 0.816.
Bo Wang 0110, Xiran Xu, Longxiang Zhang, Boda Xiao, Xihong Wu, Jing Chen 0019
ICASSP6
2024 A DenseNet-Based Method for Decoding Auditory Spatial Attention with EEG
abstract
Auditory spatial attention detection (ASAD) aims to decode the attended spatial location with EEG in a multiple-speaker setting. ASAD methods are inspired by the brain lateralization of cortical neural responses during the processing of auditory spatial attention, and show promising performance for the task of auditory attention decoding (AAD) with neural recordings. In the previous ASAD methods, the spatial distribution of EEG electrodes is not fully exploited, which may limit the performance of these methods. In the present work, by transforming the original EEG channels into a two-dimensional (2D) spatial topological map, the EEG data is transformed into a three-dimensional (3D) arrangement containing spatial-temporal information. And then a 3D deep convolutional neural network (DenseNet-3D) is used to extract temporal and spatial features of the neural representation for the attended locations. The results show that the proposed method achieves higher decoding accuracy than the state-of-the-art (SOTA) method (94.3% compared to XANet's 90.6%) with 1-second decision window for the widely used KULeuven (KUL) dataset, and the code to implement our work is available on Github: https://github.com/xuxiran/ASAD_DenseNet
Xiran Xu, Bo Wang 0110, Yujie Yan, Xihong Wu, Jing Chen 0019
ICASSP5
2024 Auditory Attention Decoding in Four-Talker Environment with EEG
Yujie Yan, Xiran Xu, Haolin Zhu, Pei Tian, Zhongshu Ge, Xihong Wu, Jing Chen 0019
INTERSPEECH7
2023 PGSS: Pitch-Guided Speech Separation
abstract
Monaural speech separation aims to separate concurrent speakers from a single-microphone mixture recording. Inspired by the effect of pitch priming in auditory scene analysis (ASA) mechanisms, a novel pitch-guided speech separation framework is proposed in this work. The prominent advantage of this framework is that both the permutation problem and the unknown speaker number problem existing in general models can be avoided by using pitch contours as the primary means to guide the target speaker. In addition, adversarial training is applied, instead of a traditional time-frequency mask, to improve the perceptual quality of separated speech. Specifically, the proposed framework can be divided into two phases: pitch extraction and speech separation. The former aims to extract pitch contour candidates for each speaker from the mixture, modeling the bottom-up process in ASA mechanisms. Any pitch contour can be selected as the condition in the second phase to separate the corresponding speaker, where a conditional generative adversarial network (CGAN) is applied. The second phase models the effect of pitch priming in ASA. Experiments on the WSJ0-2mix corpus reveal that the proposed approaches can achieve higher pitch extraction accuracy and better separation performance, compared to the baseline models, and have the potential to be applied to SOTA architectures.
Xiang Li 0072, Yiwen Wang 0009, Xihong Wu, Jing Chen 0019
AAAI5
2023 A Model-Based Hearing Compensation Method Using a Self-Supervised Framework
abstract
Hearing aids can improve auditory perception for hearing-impaired (HI) listeners, but even state-of-art devices provide only limited benefits if not configured correctly for the listeners. The prescriptive fittings of hearing aids ignore the individual difference among HI listeners with identical hearing thresholds. This paper proposes a model-based hearing compensation method using a self-supervised framework with a given auditory model. The influence of outer/inner hair cells dysfunction was simulated in the auditory model. And then, a neural network was trained to compensate for the given hearing impairment. Both objective and subjective experiments were conducted to evaluate the present method, and the results showed that listeners are sensitive to the parameter controlling the contribution of outer hair cells dysfunction. Additionally, the result indicated that listeners significantly preferred the speech processed by the proposed method to the traditional perspective fitting.
Yadong Niu, Xihong Wu, Jing Chen 0019
ICASSP4
2021 Effect of Carrier Bandwidth on Understanding Mandarin Sentences in Simulated Electric-Acoustic Hearing
Jing Chen 0019, Fei Chen 0011
Interspeech2
2020 Single-Channel Speech Separation Integrating Pitch Information Based on a Multi Task Learning Framework
abstract
Pitch is a critical cue for speech separation in humans' auditory perception. Although the technology of tracking pitch in single-talker speech succeeds in many applications, it's still a challenging problem to extract pitch information from speech mixtures in machine perception. In this paper, we aimed to combine speech separation and pitch tracking together to let them benefit from each other. A multi-task learning framework was proposed, in which a unified objective that considered both speech separation and pitch tracking was used, based on the utterance-level permutation invariant training (uPIT) as well as deep clustering (DPCL). In such framework, two tasks were optimized simultaneously and could benefit from each other through the sharing layers in the networks. Experimental results indicated the proposed multi-task framework outperformed the corresponding single-task framework, in terms of both speech separation and pitch tracking. The improvement was more significant for challenging same-gender mixtures.
Xiang Li 0072, Xihong Wu, Jing Chen 0019
ICASSP5
2020 Congruent Audiovisual Speech Enhances Cortical Envelope Tracking During Auditory Selective Attention
Zhen Fu, Jing Chen 0019
INTERSPEECH2
2019 A Spectral-change-aware Loss Function for DNN-based Speech Separation
abstract
Speech separation can be treated as a mask estimation problem where supervised learning is employed to construct the mapping from acoustic features to a mask. Interference can be reduced by applying the estimated mask on a time-frequency (T-F) representation of noisy speech, resulting in improved speech intelligibility. Most of existing learning networks for speech separation aim to minimize the Mean Square Error (MSE) over the training set, where the loss from each T-F representation is equally weighted. In this paper, we proposed a spectral-change-aware loss function, where loss from the T-F units with large spectral changes over time were assigned higher weights compared to the T-F units with minor spectral changes. Such spectral-change-aware loss function was evaluated on speech separation performance in terms of mask estimation accuracy, short-time objective intelligibility (STOI) and SNR gain of unvoiced segments. The results indicated that the proposed loss function could further improve the speech intelligibility and increase SNR gain of unvoiced segments even in the cost of increased error rate of estimated mask.
Xiang Li 0072, Xihong Wu, Jing Chen 0019
ICASSP3
2019 Integrating Spectrotemporal Context into Features Based on Auditory Perception for Classification-based Speech Separation
abstract
Speech separation, which has been a challenging task for decades, especially at low signal-to-noise ratios (SNRs), can be cast as a classification problem. In such adverse acoustic environment, extracting robust features from noisy mixtures is crucial for successful classification. In the past studies, features representing temporal dynamics, known as delta features, have been widely used. Combining basic features with their deltas yields better speech separation results than using basic features alone. In this study, the commonly used delta feature was modified according to the characteristics of auditory perception, which included auditory processing on spectral change and spectral contrast. Therefore, we proposed a feature which integrated spectrotemporal context via replacing the commonly used delta feature by spectral change feature and spectral contrast feature. Experimental results showed that the proposed feature could produce better speech segregation performance than the common delta feature.
Xiang Li 0072, Xihong Wu, Jing Chen 0019
ICASSP3
2019 Effects of Spectral and Temporal Cues to Mandarin Concurrent-Vowels Identification for Normal-Hearing and Hearing-Impaired Listeners
Zhen Fu, Xihong Wu, Jing Chen 0019
INTERSPEECH3
2018 A Time-Weighted Method for Predicting the Intelligibility of Speech in the Presence of Interfering Sounds
abstract
The speech intelligibility index (SII) has been widely used as an objective method of predicting speech intelligibility, but its traditional form is most effective predicting speech intelligibility scores under stationary noise but not more challenging conditions (e.g., competing noise interference). To address this limitation, the present work extended the SII model to predict the intelligibility of speech in both steady speech-spectral noise (SSN) and dual-talker speech (DTS), by using a time-weighted function that accounted for the relative perceptual importance of vowels and consonants in speech intelligibility. The performance of the new time-weighted SII (TW-SII) was compared to the other two well-known methods, i.e., the time-averaged SII (TA-SII) and coherence SII (CSII). Experimental results showed the intelligibility prediction accuracy of the three methods was similar for speech in SSN, but the prediction by TW-SII was more accurate than those by TA-SII and CSII for speech in DTS. The possible applications and limitations of the present intelligibility model were analyzed and discussed.
Mingjie Song, Fei Chen 0011, Xihong Wu, Jing Chen 0019
ICASSP4
2018 Measuring the Band Importance Function for Mandarin Chinese with a Bayesian Adaptive Procedure
Yufan Du, Yi Shen 0008, Hongying Yang, Xihong Wu, Jing Chen 0019
INTERSPEECH5
2016 Frequency importance function of the speech intelligibility index for Mandarin Chinese
Jing Chen 0019, Xihong Wu
Speech Commun.1
2007 Effect of number of masking talkers on speech-on-speech masking in Chinese
Xihong Wu, Jing Chen 0019
INTERSPEECH2
2007 The effect of voice cuing on releasing Chinese speech from informational masking
Jing Chen 0019, Xihong Wu, Bruce A. Schneider
Speech Commun.2