EDBT 2026 Demo / reviewers in the wild / expert
Jinwei Feng
dblp:277/3933
· DBLP profile ↗
12ranked-venue papers
0as first author
11since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Practical Online Multichannel Dereverberation Approach with Data-Reuse TechniqueabstractOne of the most effective online dereverberation algorithms is the weighted prediction error (WPE) method and its improved version, switching WPE (SwWPE). This paper proposes a practical online dereverberation approach to improve SwWPE by introducing a datareuse technique, DR-SwWPE; we then show analytically that DR-SwWPE is equivalent to a SwWPE followed by a post-filter to suppress the residual reverberation. Secondly, we compare this suppression technique and two explicit post-filtering schemes to SwWPE using both simulated data and real-world recordings. Experimental results show that DR-SwWPE out-performs SwWPE with little additional computational cost. Weilong Huang, Jinwei Feng, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2024 | AS-pVAD: A Frame-Wise Personalized Voice Activity Detection Network with Attentive Score LossabstractWe present a lightweight neural network with attentive score loss for frame-wise personalized voice activity detection (i.e., AS-pVAD). Instead of using an external speaker embedding extractor with a large number of parameters, AS-pVAD employs a lightweight internal model to extract the target speaker embedding. A novel attentive score loss constraint is proposed to better exploit such embedding clues for pVAD compared to conventional embedding concatenation. Through joint training with a regular VAD, AS-pVAD can be further improved to identify the target speaker in the enrollment cases while it is able to function as a regular VAD in the enrollment-less cases. Experimental results show that AS-pVAD achieves over 0.9 of AUCROC on average in two-speaker talking scenario under various noisy and reverberant environments. Our test set is also publicly released to the community to facilitate the research in this area. Fenting Liu, Feifei Xiong, Yiya Hao, Kechenying Zhou, Jinwei Feng |
ICASSP | 6 |
| 2024 | EMDSQA: A Neural Speech Quality Assessment Model With Speaker EmbeddingabstractWe present a neural speech quality assessment model with speaker embedding. This model, i.e., EMDSQA, can precisely predict the Mean Opinion Score (MOS) of speech quality during online communications. Intrusive speech quality assessment methods such as perceptual objective listening quality analysis (POLQA) are not practical for online communications because every piece of degraded speech requires a corresponding clean reference. Non-intrusive methods can assess the quality of online speech, but have not reached the accuracy and robustness required for real-world applications. EMDSQA extracts the speaker embedding using an independent pipeline and feeds it as a prior feature to a self-attention-based MOS prediction model. Since EMDSQA does not need the corresponding clean reference, it is practical for real-world communication applications. An open-source test corpus, featuring real-world data, was also developed. Experimental results show that EMDSQA achieves a 0.92 Pearson correlation coefficient with the MOS measured from humans, surpassing other state-of-the-art intrusive or non-intrusive methods. Yiya Hao, Feifei Xiong, Nai Ding, Jinwei Feng |
IEEE Signal Process. Lett. | 5 |
| 2023 | Deep Subband Network for Joint Suppression of Echo, Noise and Reverberation in Real-Time Fullband Speech CommunicationabstractThis paper presents a deep and lightweight subband neural network which jointly suppresses the common interference in real-time fullband speech communication: echo, noise and reverberation. Preserving the advantages of spectro-temporal subband network (STSubNet) that requires small amount of resources for good generalization within a lightweight model, the proposed framework incorporates an adaptive filter and a modified time-domain loss function designed to balance the suppression efficiency among three types of interference. Extensive experimental results show that the proposed loss function significantly improves the residual echo suppression during far-end single talk scenario and balances between distortion to the desired signal and suppression on the undesired signal. In addition, we find that STSubNet requires adaptive filter output (with a better convergence preferred) to be the primary input to achieve a better performance. Competitive performance as compared to state-of-the-art separate models is achieved on three public benchmark test sets from individual echo suppression, denoising and dereverberation area. Feifei Xiong, Minya Dong, Kechenying Zhou, Houwei Zhu, Jinwei Feng |
ICASSP | 5 |
| 2023 | Blind Estimation of Room Impulse Response from Monaural Reverberant Speech with Segmental Generative Neural Network
Zhiheng Liao, Feifei Xiong, Juan Luo, Minjie Cai, Chng Eng Siong, Jinwei Feng, Xionghu Zhong |
INTERSPEECH | 6 |
| 2022 | Spectro-Temporal SubNet for Real-Time Monaural Speech Denoising and Dereverberation
Feifei Xiong, Weiguang Chen, Pengyu Wang 0010, Jinwei Feng |
INTERSPEECH | 5 |
| 2022 | Joint Estimation of Direction-of-Arrival and Distance for Arrays with Directional Sensors based on Sparse Bayesian Learning
Feifei Xiong, Pengyu Wang 0010, Zhongfu Ye, Jinwei Feng |
INTERSPEECH | 4 |
| 2021 | A Real-Time Speaker Diarization System Based on Spatial SpectrumabstractIn this paper we describe a speaker diarization system that enables localization and identification of all speakers present in a conversation or meeting. We propose a novel systematic approach to tackle several long-standing challenges in speaker diarization tasks: (1) to segment and separate overlapping speech from two speakers; (2) to estimate the number of speakers when participants may enter or leave the conversation at any time; (3) to provide accurate speaker identification on short text-independent utterances; (4) to track down speakers movement during the conversation; (5) to detect speaker change incidence real-time. First, a differential directional microphone array-based approach is exploited to capture the target speakers’ voice in far-field adverse environment. Second, an online speaker-location joint clustering approach is proposed to keep track of speaker location. Third, an instant speaker number detector is developed to trigger the mechanism that separates overlapped speech. The results suggest that our system effectively incorporates spatial information and achieves significant gains. Weilong Huang, Xianliang Wang, Hongbin Suo, Jinwei Feng, Zhijie Yan |
ICASSP | 5 |
| 2021 | Minimum-Norm Differential Beamforming for Linear Array with Directional Microphones
Weilong Huang, Jinwei Feng |
Interspeech | 2 |
| 2021 | Real-Time Multi-Channel Speech Enhancement Based on Neural Network Masking with Attention Model
Weilong Huang, Weiguang Chen, Jinwei Feng |
Interspeech | 4 |
| 2021 | Investigation of Spatial-Acoustic Features for Overlapping Speech Detection in Multiparty Meetings
Shiliang Zhang, Weilong Huang, Hongbin Suo, Jinwei Feng, Zhijie Yan |
Interspeech | 6 |
| 2020 | Differential Beamforming for Uniform Circular Array with Directional Microphones
Weilong Huang, Jinwei Feng |
INTERSPEECH | 2 |