Chengdong Liang

dblp:289/0544 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Deep learning based stage-wise two-dimensional speaker localization with large ad-hoc microphone arrays
Shupei Liu, Linfeng Feng, Yijun Gong, Chengdong Liang, Xuelong Li 0001
Speech Commun.4
2024 Advancing speaker embedding learning: Wespeaker toolkit for research and production
Shuai Wang 0016, Zhengyang Chen, Bing Han 0008, Chengdong Liang, Xu Xiang, Wen Ding 0005, Johan Rohdin, Anna Silnova, Yanmin Qian, Haizhou Li 0001
Speech Commun.5
2023 Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames
abstract
Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy and latency. In this paper, we present fast-U2++, an enhanced version of U2++ to further reduce partial latency. The core idea of fast-U2++ is to output partial results of the bottom layers in its encoder with a small chunk, while using a large chunk in the top layers of its encoder to compensate the performance degradation caused by the small chunk. More-over, we use knowledge distillation method to reduce the token emission latency. We present extensive experiments on Aishell-1 dataset. Experiments and ablation studies show that compared to U2++, fast-U2++ reduces model latency from 320ms to 80ms, and achieves a character error rate (CER) of 5.06% with a streaming setup.
Chengdong Liang, Xiao-Lei Zhang 0001, Di Wu 0061, Shengqiang Li, Xingchen Song, Zhendong Peng, Fuping Pan
ICASSP1
2023 Wespeaker: A Research and Production Oriented Speaker Embedding Learning Toolkit
abstract
Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research and production oriented speaker embedding learning toolkit, Wespeaker. Wespeaker contains the implementation of scalable data management, state-of-the-art speaker embedding models, loss functions, and scoring back-ends, with highly competitive results achieved by structured recipes which were adopted in the winning systems in several speaker verification challenges. The application to other downstream tasks such as speaker diarization is also exhibited in the related recipe. Moreover, CPU- and GPU-compatible deployment codes are integrated for production-oriented development. The toolkit is publicly available at https://github.com/wenet-e2e/wespeaker.
Chengdong Liang, Shuai Wang 0016, Zhengyang Chen, Xu Xiang, Yanlei Deng, Yanmin Qian
ICASSP2
2023 Branch-ECAPA-TDNN: A Parallel Branch Architecture to Capture Local and Global Features for Speaker Verification
Jiadi Yao, Chengdong Liang, Zhendong Peng, Xiao-Lei Zhang 0001
INTERSPEECH2
2022 Multi-Channel Far-Field Speaker Verification with Large-Scale Ad-hoc Microphone Arrays
abstract
Speaker verification based on ad-hoc microphone arrays has the potential of reducing the error significantly in adverse acoustic environments.However, existing approaches extract utterancelevel speaker embeddings from each channel of an ad-hoc microphone array, which does not consider fully the spatialtemporal information across the devices.In this paper, we propose to aggregate the multichannel signals of the ad-hoc microphone array at the frame-level by exploring the cross-channel information deeply with two attention mechanisms.The first one is a self-attention method.It consists of a cross-frame selfattention layer and a cross-channel self-attention layer successively, both working at the frame level.The second one learns the cross-frame and cross-channel information via two graph attention layers.Experimental results demonstrate that the proposed methods reach the state-of-the-art performance.Moreover, the graph-attention method is better than the self-attention method in most cases.
Chengdong Liang, Jiadi Yao, Xiao-Lei Zhang 0001
INTERSPEECH1
2022 Multi-class AUC Optimization for Robust Small-footprint Keyword Spotting with Limited Training Data
Menglong Xu, Shengqiang Li, Chengdong Liang, Xiao-Lei Zhang 0001
INTERSPEECH3