EDBT 2026 Demo / reviewers in the wild / expert
Xuyi Zhuang
dblp:294/9932
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0000-9146-9952ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OACodec: Audio attribute disentanglement via orthogonal disentanglement and mutual information minimization
Yukun Qian, Lianyu Zhou, Xuyi Zhuang, Mingjiang Wang |
Neural Networks | 5 |
| 2026 | MS-VBRVQ: Multi-scale variable bitrate speech residual vector quantization
Yukun Qian, Shiyun Xu, Xuyi Zhuang, Mingjiang Wang |
Speech Commun. | 3 |
| 2025 | Joint Training Framework for Accent and Speech Recognition Based on Conformer Low-Rank AdaptationabstractIn real-world scenarios, accent variations often reduce Automatic Speech Recognition (ASR) accuracy. Addressing this typically involves a multi-task ASR and Accent Recognition (ASR-AR) framework, but there is limited research on optimizing task-specific feature extraction and enhancing ASR with AR information. This study introduces the Conformer Low-rank Adaptation for Joint Accent and Speech Recognition (CLAnSR), employing LoRA to augment both ASR and AR capabilities using a shared pre-trained base encoder. This approach significantly reduces the model’s parameter and training resource demands while facilitating the extraction of task-specific features. Additionally, we have incorporated accent-aware multi-channel embedding layers, which through spatially independent embeddings, enhance the model’s capacity to accurately represent tokens across diverse dialectical contexts. Tested on the KeSpeech dataset, CLAnSR reaches state-of-the-art AR accuracy 80.41% and competitive ASR CER 8.39%, outperforming non-LLM systems and matching those with LLMs. It reduces parameters by 34.02% and enhances both ASR and AR performance, effectively handling speech dialect variations and advancing the field. Xuyi Zhuang, Yukun Qian, Shiyun Xu, Mingjiang Wang |
ICASSP | 1 |
| 2025 | Hypformer: A Fast Hypothesis-Driven Rescoring Speech Recognition FrameworkabstractRecently, the performance of non-autoregressive ASR models has made significant progress but still lags behind hybrid CTC/attention systems. This paper introduces Hypformer, a fast hypothesis-driven rescoring speech recognition framework. Multiple hypothetical prefixes are realized by fast prefix generation algorithm. With two different rescoring methods, nar-ar rescoring and nar$^{2}$rescoring, Hypformer can flexibly switch between autoregressive and non-autoregressive decoding modes to perform rescoring of hypothesis prefixes. Experiments on the standard Mandarin datasets AISHELL-1 and AISHELL-2 demonstrate that Hypformer outperforms the state-of-the-art Hybrid CTC/Attention systems in ASR performance while achieving a speedup of over six times. Experiments on the Mandarin sub-dialect dataset KeSpeech indicate that Hypformer achieves more accurate recognition by leveraging richer contextual information. Xuyi Zhuang, Yukun Qian, Mingjiang Wang |
IEEE Signal Process. Lett. | 1 |
| 2024 | Lightweight Dynamic Sparse Transformer for Monaural Speech Enhancement
Xuyi Zhuang, Yukun Qian, Mingjiang Wang |
INTERSPEECH | 2 |
| 2023 | Half-Temporal and Half-Frequency Attention U2Net for Speech Signal ImprovementabstractDuring communication, volume changes, noise, and reverberation can disturb speech signals, significantly affecting the quality and intelligibility of speech. In the context of the ICASSP 2023 Signal Processing Grand Challenge, the first Speech Signal Improvement Grand Challenge (SIG) is organized to improve the quality of speech signals during communication. This paper proposes half-temporal and half-frequency attention U2Net for improving full-band speech signal. Channel-spectrum attention is proposed for the skip connection between the encoder and decoder. The proposed model achieves 0.353, 1.289, 0.604, 0.625, and 0.924 improvements in signal, noise, overall, reverberation, and loudness, respectively, in the SIG subjective test. The proposed model achieved fourth place in the SIG real-time track, showing excellent denoising and de-reverberation performance. Shiyun Xu, Xuyi Zhuang, Yukun Qian, Lianyu Zhou, Mingjiang Wang |
ICASSP | 3 |
| 2023 | Two-Stage UNet with Multi-Axis Gated Multilayer Perceptron for Monaural Noisy-Reverberant Speech EnhancementabstractIn denoising and de-reverberation tasks, the dominant methods are complex spectral masking and complex spectral mapping. To combine advantages and improve speech enhancement performance, we propose a two-stage UNet (TSUNet) to estimate complex spectral masking and complex spectral mapping. We use a multi-axis gated multilayer perceptron to build global and local attention modules of linear complexity for extracting speech features. Furthermore, we use the residual channel attention block to further filter out important speech features. On the blind test dataset of the Deep Noise Suppression Challenge, our proposed TSUNet has a massive advantage over other state-of-the-art models. TSUNet performs significantly better than the most recent models at noisy-reverberant speech enhancement. Shiyun Xu, Xuyi Zhuang, Lianyu Zhou, Heng Li 0013, Mingjiang Wang |
ICASSP | 3 |
| 2023 | Automatic Speech Recognition Transformer with Global Contextual Information Decoder
Yukun Qian, Xuyi Zhuang, Mingjiang Wang |
INTERSPEECH | 2 |
| 2022 | FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional NetworkabstractIn recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time enhancement of full-band (48 kHz) speech signals is still very challenging. Because of the low energy of spectral information in the high-frequency part, it is more difficult to directly model and enhance the full-band spectrum using neural networks. To solve this problem, this paper proposes a two-stage real-time speech enhancement model with extraction-interpolation mechanism for a full-band signal. The 48 kHz full-band time-domain signal is divided into three sub-channels by extracting, and a two-stage processing scheme of ‘masking + compensation’ is proposed to enhance the signal in the complex domain. After the two-stage enhancement, the enhanced full-band speech signal is restored by interval interpolation. In the subjective listening and word accuracy test, our proposed model achieves superior performance and outperforms the baseline model overall by 0.59 MOS and 4.0% WAcc for the non-personalized speech denoising task. Lu Zhang 0055, Xuyi Zhuang, Yukun Qian, Heng Li 0013, Mingjiang Wang |
ICASSP | 3 |
| 2022 | Coarse-Grained Attention Fusion With Joint Training Framework for Complex Speech Enhancement and End-to-End Speech Recognition
Xuyi Zhuang, Yukun Qian, Mingjiang Wang |
INTERSPEECH | 1 |