VLDB 2026 Research / reviewers in the wild / expert
Chengdong Liang
dblp:289/0544
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep learning based stage-wise two-dimensional speaker localization with large ad-hoc microphone arrays
Shupei Liu, Linfeng Feng, Yijun Gong, Chengdong Liang, Xuelong Li 0001 |
Speech Commun. | 4 |
| 2024 | Advancing speaker embedding learning: Wespeaker toolkit for research and production
Shuai Wang 0016, Zhengyang Chen, Bing Han 0008, Chengdong Liang, Xu Xiang, Wen Ding 0005, Johan Rohdin, Anna Silnova, Yanmin Qian, Haizhou Li 0001 |
Speech Commun. | 5 |
| 2023 | Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention FramesabstractRecently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy and latency. In this paper, we present fast-U2++, an enhanced version of U2++ to further reduce partial latency. The core idea of fast-U2++ is to output partial results of the bottom layers in its encoder with a small chunk, while using a large chunk in the top layers of its encoder to compensate the performance degradation caused by the small chunk. More-over, we use knowledge distillation method to reduce the token emission latency. We present extensive experiments on Aishell-1 dataset. Experiments and ablation studies show that compared to U2++, fast-U2++ reduces model latency from 320ms to 80ms, and achieves a character error rate (CER) of 5.06% with a streaming setup. Chengdong Liang, Xiao-Lei Zhang 0001, Di Wu 0061, Shengqiang Li, Xingchen Song, Zhendong Peng, Fuping Pan |
ICASSP | 1 |
| 2023 | Wespeaker: A Research and Production Oriented Speaker Embedding Learning ToolkitabstractSpeaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research and production oriented speaker embedding learning toolkit, Wespeaker. Wespeaker contains the implementation of scalable data management, state-of-the-art speaker embedding models, loss functions, and scoring back-ends, with highly competitive results achieved by structured recipes which were adopted in the winning systems in several speaker verification challenges. The application to other downstream tasks such as speaker diarization is also exhibited in the related recipe. Moreover, CPU- and GPU-compatible deployment codes are integrated for production-oriented development. The toolkit is publicly available at https://github.com/wenet-e2e/wespeaker. Chengdong Liang, Shuai Wang 0016, Zhengyang Chen, Xu Xiang, Yanlei Deng, Yanmin Qian |
ICASSP | 2 |
| 2023 | Branch-ECAPA-TDNN: A Parallel Branch Architecture to Capture Local and Global Features for Speaker Verification
Jiadi Yao, Chengdong Liang, Zhendong Peng, Xiao-Lei Zhang 0001 |
INTERSPEECH | 2 |
| 2022 | Multi-Channel Far-Field Speaker Verification with Large-Scale Ad-hoc Microphone ArraysabstractSpeaker verification based on ad-hoc microphone arrays has the potential of reducing the error significantly in adverse acoustic environments.However, existing approaches extract utterancelevel speaker embeddings from each channel of an ad-hoc microphone array, which does not consider fully the spatialtemporal information across the devices.In this paper, we propose to aggregate the multichannel signals of the ad-hoc microphone array at the frame-level by exploring the cross-channel information deeply with two attention mechanisms.The first one is a self-attention method.It consists of a cross-frame selfattention layer and a cross-channel self-attention layer successively, both working at the frame level.The second one learns the cross-frame and cross-channel information via two graph attention layers.Experimental results demonstrate that the proposed methods reach the state-of-the-art performance.Moreover, the graph-attention method is better than the self-attention method in most cases. Chengdong Liang, Jiadi Yao, Xiao-Lei Zhang 0001 |
INTERSPEECH | 1 |
| 2022 | Multi-class AUC Optimization for Robust Small-footprint Keyword Spotting with Limited Training Data
Menglong Xu, Shengqiang Li, Chengdong Liang, Xiao-Lei Zhang 0001 |
INTERSPEECH | 3 |