VLDB 2026 Research / reviewers in the wild / expert
Feifei Xiong
dblp:120/6947
· DBLP profile ↗
18ranked-venue papers
13as first author
7since 2021 · last 2025
0000-0001-9783-2169ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 11 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploiting spatial information and target speaker phoneme loss for multichannel directional speech enhancement and recognition
Cong Pang, Ye Ni, Lin Zhou 0001, Li Zhao 0003, Feifei Xiong |
Comput. Speech Lang. | 5 |
| 2024 | AS-pVAD: A Frame-Wise Personalized Voice Activity Detection Network with Attentive Score LossabstractWe present a lightweight neural network with attentive score loss for frame-wise personalized voice activity detection (i.e., AS-pVAD). Instead of using an external speaker embedding extractor with a large number of parameters, AS-pVAD employs a lightweight internal model to extract the target speaker embedding. A novel attentive score loss constraint is proposed to better exploit such embedding clues for pVAD compared to conventional embedding concatenation. Through joint training with a regular VAD, AS-pVAD can be further improved to identify the target speaker in the enrollment cases while it is able to function as a regular VAD in the enrollment-less cases. Experimental results show that AS-pVAD achieves over 0.9 of AUCROC on average in two-speaker talking scenario under various noisy and reverberant environments. Our test set is also publicly released to the community to facilitate the research in this area. Fenting Liu, Feifei Xiong, Yiya Hao, Kechenying Zhou, Jinwei Feng |
ICASSP | 2 |
| 2024 | EMDSQA: A Neural Speech Quality Assessment Model With Speaker EmbeddingabstractWe present a neural speech quality assessment model with speaker embedding. This model, i.e., EMDSQA, can precisely predict the Mean Opinion Score (MOS) of speech quality during online communications. Intrusive speech quality assessment methods such as perceptual objective listening quality analysis (POLQA) are not practical for online communications because every piece of degraded speech requires a corresponding clean reference. Non-intrusive methods can assess the quality of online speech, but have not reached the accuracy and robustness required for real-world applications. EMDSQA extracts the speaker embedding using an independent pipeline and feeds it as a prior feature to a self-attention-based MOS prediction model. Since EMDSQA does not need the corresponding clean reference, it is practical for real-world communication applications. An open-source test corpus, featuring real-world data, was also developed. Experimental results show that EMDSQA achieves a 0.92 Pearson correlation coefficient with the MOS measured from humans, surpassing other state-of-the-art intrusive or non-intrusive methods. Yiya Hao, Feifei Xiong, Nai Ding, Jinwei Feng |
IEEE Signal Process. Lett. | 2 |
| 2023 | Deep Subband Network for Joint Suppression of Echo, Noise and Reverberation in Real-Time Fullband Speech CommunicationabstractThis paper presents a deep and lightweight subband neural network which jointly suppresses the common interference in real-time fullband speech communication: echo, noise and reverberation. Preserving the advantages of spectro-temporal subband network (STSubNet) that requires small amount of resources for good generalization within a lightweight model, the proposed framework incorporates an adaptive filter and a modified time-domain loss function designed to balance the suppression efficiency among three types of interference. Extensive experimental results show that the proposed loss function significantly improves the residual echo suppression during far-end single talk scenario and balances between distortion to the desired signal and suppression on the undesired signal. In addition, we find that STSubNet requires adaptive filter output (with a better convergence preferred) to be the primary input to achieve a better performance. Competitive performance as compared to state-of-the-art separate models is achieved on three public benchmark test sets from individual echo suppression, denoising and dereverberation area. Feifei Xiong, Minya Dong, Kechenying Zhou, Houwei Zhu, Jinwei Feng |
ICASSP | 1 |
| 2023 | Blind Estimation of Room Impulse Response from Monaural Reverberant Speech with Segmental Generative Neural Network
Zhiheng Liao, Feifei Xiong, Juan Luo, Minjie Cai, Chng Eng Siong, Jinwei Feng, Xionghu Zhong |
INTERSPEECH | 2 |
| 2022 | Spectro-Temporal SubNet for Real-Time Monaural Speech Denoising and Dereverberation
Feifei Xiong, Weiguang Chen, Pengyu Wang 0010, Jinwei Feng |
INTERSPEECH | 1 |
| 2022 | Joint Estimation of Direction-of-Arrival and Distance for Arrays with Directional Sensors based on Sparse Bayesian Learning
Feifei Xiong, Pengyu Wang 0010, Zhongfu Ye, Jinwei Feng |
INTERSPEECH | 1 |
| 2020 | Source Domain Data Selection for Improved Transfer Learning Targeting Dysarthric Speech RecognitionabstractThis paper presents an improved transfer learning framework applied to robust personalised speech recognition models for speakers with dysarthria. As the baseline of transfer learning, a state-of-the-art CNN-TDNN-F ASR acoustic model trained solely on source domain data is adapted onto the target domain via neural network weight adaptation with the limited available data from target dysarthric speakers. Results show that linear weights in neural layers play the most important role for an improved modelling of dysarthric speech evaluated using UASpeech corpus, achieving averaged 11.6% and 7.6% relative recognition improvement in comparison to the conventional speaker-dependent training and data combination, respectively. To further improve the transferability towards target domain, we propose an utterance-based data selection of the source domain data based on the entropy of posterior probability, which is analysed to statistically obey a Gaussian distribution. Compared to a speaker-based data selection via dysarthria similarity measure, this allows for a more accurate selection of the potentially beneficial source domain data for either increasing the target domain training pool or constructing an intermediate domain for incremental transfer learning, resulting in a further absolute recognition performance improvement of nearly 2% added to transfer learning baseline for speakers with moderate to severe dysarthria. Feifei Xiong, Jon Barker, Zhengjun Yue, Heidi Christensen |
ICASSP | 1 |
| 2020 | Exploring Appropriate Acoustic and Language Modelling Choices for Continuous Dysarthric Speech RecognitionabstractThere has been much recent interest in building continuous speech recognition systems for people with severe speech impairments, e.g., dysarthria. However, the datasets that are commonly used are typically designed for tasks other than ASR development, or they contain only isolated words. As such, they contain much overlap in the prompts read by the speakers. Previous ASR evaluations have often neglected this, using language models (LMs) trained on non-disjoint training and test data, potentially producing unrealistically optimistic results. In this paper, we investigate the impact of LM design using the widely used TORGO database. We combine state-of-the-art acoustic models with LMs trained with data originating from LibriSpeech. Using LMs with varying vocabulary size, we examine the trade-off between the out-of-vocabulary rate and recognition confusions for speakers with varying degrees of dysarthria. It is found that the optimal LM complexity is highly speaker dependent, highlighting the need to design speaker-dependent LMs alongside speaker-dependent acoustic models when considering atypical speech. Zhengjun Yue, Feifei Xiong, Heidi Christensen, Jon Barker |
ICASSP | 2 |
| 2019 | Phonetic Analysis of Dysarthric Speech Tempo and Applications to Robust Personalised Dysarthric Speech RecognitionabstractImproving the accuracy of personalised speech recognition for speakers with dysarthria is a challenging research field. In this paper, we explore an approach that non-linearly modifies speech tempo to reduce mismatch between typical and atypical speech. Speech tempo analysis at the phonetic level is accomplished using a forced-alignment process from traditional GMM-HMM in automatic speech recognition (ASR). Estimated tempo adjustments are applied directly to the acoustic features rather than to the time-domain signals. Two approaches are considered: i) adjusting dysarthric speech towards typical speech for input into ASR systems trained with typical speech, and ii) adjusting typical speech towards dysarthric speech for data augmentation in personalised dysarthric ASR training. Experimental results show that the latter strategy with data augmentation is more effective, resulting in a nearly 7% absolute improvement in comparison to baseline speaker-dependent trained system evaluated using UASpeech corpus. Consistent recognition performance improvements are observed across speakers, with greatest benefit in cases of moderate and severe dysarthria. Feifei Xiong, Jon Barker, Heidi Christensen |
ICASSP | 1 |
| 2019 | Joint Estimation of Reverberation Time and Early-To-Late Reverberation Ratio From Single-Channel Speech SignalsabstractThe reverberation time (RT) and the early-to-late reverberation ratio (ELR) are two key parameters commonly used to characterize acoustic room environments. In contrast to conventional blind estimation methods that process the two parameters separately, we propose a model for joint estimation to predict the RT and the ELR simultaneously from single-channel speech signals from either full-band or sub-band frequency data, which is referred to as joint room parameter estimator (jROPE). An artificial neural network is employed to learn the mapping from acoustic observations to the RT and the ELR classes. Auditory-inspired acoustic features obtained by temporal modulation filtering of the speech time-frequency representations are used as input for the neural network. Based on an in-depth analysis of the dependency between the RT and the ELR, a two-dimensional (RT, ELR) distribution with constrained boundaries is derived, which is then exploited to evaluate four different configurations for jROPE. Experimental results show that-in comparison to the single-task ROPE system which individually estimates the RT or the ELR-jROPE provides improved results for both tasks in various reverberant and (diffuse) noisy environments. Among the four proposed joint types, the one incorporating multi-task learning with shared input and hidden layers yields the best estimation accuracies on average. When encountering extreme reverberant conditions with RTs and ELRs lying beyond the derived (RT, ELR) distribution, the type considering RT and ELR as a joint parameter performs robustly, in particular. From state-of-the-art algorithms that were tested in the acoustic characterization of environments challenge, jROPE achieves comparable results among the best for all individual tasks (RT and ELR estimation from full-band and sub-band signals). Feifei Xiong, Stefan Goetze, Birger Kollmeier, Bernd T. Meyer |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Exploring Auditory-Inspired Acoustic Features for Room Acoustic Parameter Estimation From Monaural SpeechabstractRoom acoustic parameters that characterize acoustic environments can help to improve signal enhancement algorithms such as for dereverberation, or automatic speech recognition by adapting models to the current parameter set. The reverberation time (RT) and the early-to-late reverberation ratio (ELR) are two key parameters. In this paper, we propose a blind ROom Parameter Estimator (ROPE) based on an artificial neural network that learns the mapping to discrete ranges of the RT and the ELR from single-microphone speech signals. Auditory-inspired acoustic features are used as neural network input, which are generated by a temporal modulation filter bank applied to the speech time-frequency representation. ROPE performance is analyzed in various reverberant environments in both clean and noisy conditions for both fullband and subband RT and ELR estimations. The importance of specific temporal modulation frequencies is analyzed by evaluating the contribution of individual filters to the ROPE performance. Experimental results show that ROPE is robust against different variations caused by room impulse responses (measured versus simulated), mismatched noise levels, and speech variability reflected through different corpora. Compared to state-of-the-art algorithms that were tested in the acoustic characterisation of environments (ACE) challenge, the ROPE model is the only one that is among the best for all individual tasks (RT and ELR estimation from fullband and subband signals). Improved fullband estimations are even obtained by ROPE when integrating speech-related frequency subbands. Furthermore, the model requires the least computational resources with a real time factor that is at least two times faster than competing algorithms. Results are achieved with an average observation window of 3 s, which is important for real-time applications. Feifei Xiong, Stefan Goetze, Birger Kollmeier, Bernd T. Meyer |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Combination strategy based on relative performance monitoring for multi-stream reverberant speech recognitionabstractA multi-stream framework with deep neural network (DNN) classifiers is applied to improve automatic speech recognition (ASR) in environments with different reverberation characteristics. We propose a room parameter estimation model to establish a reliable combination strategy which performs on either DNN posterior probabilities or word lattices. The model is implemented by training a multilayer perceptron incorporating auditory-inspired features in order to distinguish between and generalize to various reverberant conditions, and the model output is shown to be highly correlated to ASR performances between multiple streams, i.e., relative performance monitoring, in contrast to conventional mean temporal distance based performance monitoring for a single stream. Compared to traditional multi-condition training, average relative word error rate improvements of 7.7% and 9.4% have been achieved by the proposed combination strategies performing on posteriors and lattices, respectively, when the multi-stream ASR is tested in known and unknown simulated reverberant environments as well as realistically recorded conditions taken from REVERB Challenge evaluation set. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 1 |
| 2017 | On DNN posterior probability combination in multi-stream speech recognition for reverberant environmentsabstractA multi-stream framework with deep neural network (DNN) classifiers has been applied in this paper to improve automatic speech recognition (ASR) performance in environments with different reverberation characteristics. We propose a room parameter estimation model to determine the stream weights for DNN posterior probability combination with the aim of obtaining reliable log-likelihoods for decoding. The model is implemented by training a multi-layer perceptron to distinguish between various reverberant environments. The method is tested in known and unknown environments against approaches based on inverse entropy and autoencoders, with average relative word error rate improvements of 46% and 29%, respectively, when performing multi-stream ASR in different reverberant situations. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 1 |
| 2015 | A study on joint beamforming and spectral enhancement for robust speech recognition in reverberant environmentsabstractThis work evaluates multi-microphone beamforming and single-microphone spectral enhancement strategies to alleviate the reverberation effect for robust automatic speech recognition (ASR) systems in different reverberant environments characterized by different reverberation times T60 and direct-to-reverberation ratios (DRRs). The systems consist of minimum variance distortionless response (MVDR) beamformers in combination with minimum mean square error (MMSE) estimators, and late reverberation spectral variance (LRSV) estimators, the latter employing a generalized model of the room impulse response (RIR). Various system architectures are analyzed with a focus on optimal speech recognition performance. The system combining an MVDR beamformer and a subsequent MMSE estimator was found to lead to the best results, with relative reductions of 27.7% compared to the baseline system. This is attributed to a more accurate LRSV estimate from spatial averaging and diffuse field refinement for the MMSE estimator. Feifei Xiong, Bernd T. Meyer, Stefan Goetze |
ICASSP | 1 |
| 2014 | Estimating room acoustic parameters for speech recognizer adaptation and combination in reverberant environmentsabstractThis work analyzes the influence of reverberation on automatic speech recognition (ASR) systems and how to compensate its influence, with special focus on the important acoustical parameters i.e. room reverberation time T60and clarity index C50. A multilayer perceptron (MLP) using features of a spectro-temporal filter bank as input is employed to identify the acoustic conditions spanning various reverberant scenarios. The posterior probabilities of the MLP are used to design a novel selection scheme for adaptation in a cluster-based manner and for system combination achieved by recognizer output voting error reduction (ROVER). A comparison of word error rates is performed considering different training modes, and an average relative improvement of 7.1% is obtained by the proposed system compared to conventional multistyle training. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 1 |
| 2013 | Blind estimation of reverberation time based on spectro-temporal modulation filteringabstractA novel method for blind estimation of the reverberation time (RT60) is proposed based on applying spectro-temporal modulation filters to time-frequency representations. 2D-Gabor filters arranged in a filterbank enable an analysis of the properties of temporal, spectral, and spectro-temporal filtering for this task. Features are used as input to a multi-layer perceptron (MLP) classifier combined with a simple decision rule that attributes a specific RT60 to a given utterance and allows to assess the reliability of the approach for different resolutions of RT60 classification. While the filter set including temporal, spectral, and spectro-temporal filters already outperforms an MFCC baseline, the error rates are further reduced when relying on diagonal spectro-temporal filters alone. The average error rate is 1.9% for the best feature set, which corresponds to a relative reduction of 58.3% compared to the MFCC baseline for RT60s in 0.1 s resolution. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 1 |
| 2012 | System identification for listening-room compensation by means of acoustic echo cancellation and acoustic echo suppression filtersabstractSubsystems for dereverberation and acoustic echo cancellation (AEC)/acoustic echo suppression (AES) are important components in high-quality hands-free telecommunication systems. This contribution describes and analyzes a combined system for dereverberation and AEC/AES. The system identification inherently achieved by the AEC/AES system is used for the design of the room impulse response (RIR) equalization filter, i.e. the listening-room compensation (LRC) system. We use complex RIR smoothing and decoupled filtered-X least-mean-squares (dFxLMS) gradient algorithm for LRC and a combined AEC/AES system for the system identification necessary for the LRC filter design. The performance of the combined system and the mutual influences of LRC and AEC/AES are analyzed. Feifei Xiong, Jens-E. Appell, Stefan Goetze |
ICASSP | 1 |