Yuxuan Ke

dblp:160/0583 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2022 Bifurcation and Reunion: A Loss-Guided Two-Stage Approach for Monaural Speech Dereverberation
Xiaoxue Luo, Chengshi Zheng, Andong Li, Yuxuan Ke, Xiaodong Li 0002
INTERSPEECH4
2022 Analysis of trade-offs between magnitude and phase estimation in loss functions for speech denoising and dereverberation
Xiaoxue Luo, Chengshi Zheng, Andong Li, Yuxuan Ke, Xiaodong Li 0002
Speech Commun.4
2022 DBT-Net: Dual-Branch Federative Magnitude and Phase Estimation With Attention-in-Attention Transformer for Monaural Speech Enhancement
abstract
The decoupling-style concept begins to ignite in the speech enhancement area, which decouples the original complex spectrum estimation task into multiple easier sub-tasks (i.e., the magnitude-only recovery and residual complex spectrum estimation), resulting in better performance and easier interpretability. In this paper, we propose a dual-branch federative magnitude and phase estimation framework, dubbed DBT-Net, for monaural speech enhancement, aiming at recovering the coarse- and fine-grained regions of the overall spectrum in parallel. From the complementary perspective, the magnitude estimation branch is designed to filter out dominant noise components in the magnitude domain, while the complex spectrum purification branch is elaborately designed to inpaint the missing spectral details and implicitly estimate the phase information in the complex-valued spectral domain. To facilitate the information flow between each branch, interaction modules are introduced to leverage features learned from one branch, so as to suppress the undesired parts and recover the missing components of the other branch. Instead of adopting the conventional RNNs and temporal convolutional networks for sequence modeling, we employ a novel attention-in-attention transformer-based network within each branch for better feature learning. More specially, it is composed of several adaptive spectro-temporal attention transformer-based modules and an adaptive hierarchical attention module, aiming to capture long-term time-frequency dependencies and further aggregate intermediate hierarchical contextual information. Comprehensive evaluations on the WSJ0-SI84 + DNS-Challenge and VoiceBank + DEMAND dataset demonstrate that the proposed approach consistently outperforms previous advanced systems and yields state-of-the-art performance in terms of speech quality and intelligibility.
Guochen Yu, Andong Li, Hui Wang 0070, Yuxuan Ke, Chengshi Zheng
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Know Your Enemy, Know Yourself: A Unified Two-Stage Framework for Speech Enhancement
Andong Li, Yuxuan Ke, Chengshi Zheng, Xiaodong Li 0002
Interspeech3
2021 Distributed node-specific block-diagonal LCMV beamforming in wireless acoustic sensor networks
Minmin Yuan, Yuxuan Ke, Chengshi Zheng, Xiaodong Li 0002
Signal Process.3
2021 Finite data performance analysis of one-bit MVDR and phase-only MVDR
Weixin Meng, Yuxuan Ke, Chengshi Zheng, Xiaodong Li 0002
Signal Process.2
2021 Corrigendum to 'Finite data performance analysis of one-bit MVDR and phase-only MVDR' [Signal Processing 183 (2021) Article 108018]
Weixin Meng, Yuxuan Ke, Chengshi Zheng, Xiaodong Li 0002
Signal Process.2