VLDB 2026 Research / reviewers in the wild / expert
Shidong Shang
dblp:287/8090
· DBLP profile ↗
15ranked-venue papers
0as first author
15since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 15 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Full-text Error Correction for Chinese Speech Recognition with Large Language ModelabstractLarge Language Models (LLMs) have demonstrated substantial potential for error correction in Automatic Speech Recognition (ASR). However, most research focuses on utterances from short-duration speech recordings, which are the predominant form of speech data for supervised ASR training. This paper investigates the effectiveness of LLMs for error correction in full-text generated by ASR systems from longer speech recordings, such as transcripts from podcasts, news broadcasts, and meetings. First, we develop a Chinese dataset for full-text error correction, named ChFT, utilizing a pipeline that involves text-to-speech synthesis, ASR, and error-correction pair extractor. This dataset enables us to correct errors across contexts, including both full-text and segment, and to address a broader range of error types, such as punctuation restoration and inverse text normalization, thus making the correction process comprehensive. Second, we fine-tune a pre-trained LLM on the constructed dataset using a diverse set of prompts and target formats, and evaluate its performance on full-text error correction. Specifically, we design prompts based on full-text and segment, considering various output formats, such as directly corrected text and JSON-based error-correction pairs. Through various test settings, including homogeneous, up-to-date, and hard test sets, we find that the finetuned LLMs perform well in the full-text setting with different prompts, each presenting its own strengths and weaknesses. This establishes a promising baseline for further research. The dataset is available on the website1. Zhiyuan Tang, Shen Huang, Shidong Shang |
ICASSP | 4 |
| 2025 | AVS3P10 Standard for Real-time Speech CodingabstractAs the tenth part of the third-generation AVS standard series for real-time speech coding, AVS3P10 is the recent standard completed in the Audio Video Coding Standards Workgroup of China (AVS). Combining the state-of-the-art deep generative networks and signal processing methods, AVS3P10 targets defining new generation neural speech codecs with high quality at low bitrates, enabling excellent experiences even when the bitrate is at 5.9 kbps with excellent error resilience. Moreover, it provides wideband and super wideband coding modes, and it supports the extension of stereo coding. Both subjective listening test and objective measurement prove the merit of AVS3P10. Especially, a lightweight model with only 880k parameters is incorporated to maintain the practicality of AVS3P10 in computational efficiency. Conclusively, AVS3P10 demonstrates the maturity of neural speech coding with broad application perspectives in real-time communication. Weibei Dou, Gaoxiong Yi, Jingxin Li, Shidong Shang |
ICASSP | 6 |
| 2024 | Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models
Zhiyuan Tang, Shen Huang, Shidong Shang |
INTERSPEECH | 4 |
| 2023 | Inter-Subnet: Speech Enhancement with Subband InteractionabstractSubband-based approaches process subbands in parallel through the model with shared parameters to learn the commonality of local spectrums for noise reduction. In this way, they have achieved remarkable results with fewer parameters. However, in some complex environments, the lack of global spectral information has a negative impact on the performance of these subband-based approaches. To this end, this paper introduces the subband interaction as a new way to complement the subband model with the global spectral information such as cross-band dependencies and global spectral patterns, and proposes a new lightweight single-channel speech enhancement framework called Interactive Subband Network (Inter-SubNet). Experimental results on DNS Challenge - Interspeech 2021 dataset show that the proposed Inter-SubNet yields a significant improvement over the subband model and outperforms other state-of-the-art speech enhancement approaches, which demonstrate the effectiveness of subband interaction. Jun Chen 0024, Wei Rao 0002, Zilin Wang 0002, Jiuxin Lin, Zhiyong Wu 0001, Yannan Wang, Shidong Shang, Helen M. Meng |
ICASSP | 7 |
| 2023 | Gesper: A Unified Framework for General Speech RestorationabstractThis paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by speech enhancement. We propose a complex spectral mapping-based generative adversarial network (CSM-GAN) as the speech restoration module for the first time. For noise suppression and dereverberation, the enhancement module is presented with fullband-wideband parallel processing. On the blind test set of ICASSP 2023 SSI Challenge, the proposed Gesper system, which satisfies the real-time condition, achieves 3.27 P.804 overall mean opinion score (MOS) and 3.35 P.835 overall MOS, ranked 1st in both track 1 and track 2. Jun Chen 0024, Yupeng Shi, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001, Shidong Shang, Chengshi Zheng |
ICASSP | 9 |
| 2023 | Speech Enhancement with Intelligent Neural Homomorphic SynthesisabstractMost neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement. Specifically, we use homomorphic signal processing and cepstral analysis to obtain noisy speech’s excitation and vocal tract. Unlike traditional signal processing, we use an attentive recurrent network (ARN) model predicted ratio mask to replace the liftering separation function. Then two convolutional attentive recurrent network (CARN) networks are used to predict the excitation and vocal tract of clean speech, respectively. The system’s output is synthesized from the estimated excitation and vocal. Experiments prove that our proposed method performs better, with SI-SNR improving by 1.363dB compared to FullSubNet. Shulin He, Wei Rao 0002, Jun Chen 0024, Yukai Jv, Xueliang Zhang 0001, Yannan Wang, Shidong Shang |
ICASSP | 8 |
| 2023 | TEA-PSE 3.0: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System For ICASSP 2023 Dns-ChallengeabstractThis paper introduces the Unbeatable Team’s submission to the ICASSP 2023 Deep Noise Suppression (DNS) Challenge. We expand our previous work, TEA-PSE, to its upgraded version – TEA-PSE 3.0. Specifically, TEA-PSE 3.0 incorporates a residual LSTM after squeezed temporal convolution network (S-TCN) to enhance sequence modeling capabilities. Additionally, the local-global representation (LGR) structure is introduced to boost speaker information extraction, and multi-STFT resolution loss is used to effectively capture the time-frequency characteristics of the speech signals. Moreover, retraining methods are employed based on the freeze training strategy to fine-tune the system. According to the official results, TEA-PSE 3.0 ranks 1st in both ICASSP 2023 DNS-Challenge track 1 and track 2. Yukai Jv, Jun Chen 0024, Shulin He, Wei Rao 0002, Weixin Zhu, Yannan Wang, Shidong Shang |
ICASSP | 9 |
| 2023 | Multi-mode Neural Speech Coding Based on Deep Generative Networks
Shan Yang 0001, Yupeng Shi, Yuyong Kang, Dan Su 0002, Shidong Shang, Dong Yu 0001 |
INTERSPEECH | 8 |
| 2022 | Internet Streaming Audio Based Speech Reception Threshold Measurement in Cochlear Implant UsersabstractTraditional face-to-face subjective listening test has become a challenge due to the COVID-19 pandemic. We developed a remote assessment system with Tencent Meeting, a video conferencing application, to address this issue. This paper presents our work on evaluating the reliability of the remote assessment system. Two speech reception threshold (SRT) experiments were conducted to study the effects of noise suppression and maxima selection number on cochlear implant (CI) hearing. Both experiments were conducted locally and remotely, the correlations between the respective results were analyzed. Results showed that remote tests replicated the differences among testing conditions observed in local tests, but the absolute SRT values for individual conditions varied significantly between the two modes. The variations could be attributed to multiple reasons, such as online data transmission issues, audio playback devices, environmental conditions, and the training of participants. In conclusion, the relative variation of SRTs for CIs can be measured reliably, but the absolute SRT values should be carefully compared and explained according to objective and subjective experimental conditions. Yefei Mo, Kang Ouyang, Mingyue Shi, Huali Zhou, Yupeng Shi, Shidong Shang, Nengheng Zheng |
ICASSP | 8 |
| 2022 | TEA-PSE: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System for ICASSP 2022 DNS ChallengeabstractThis paper describes Tencent Ethereal Audio Lab – Northwestern Polytechnical University personalized speech enhancement (TEA-PSE) system submitted to track 2 of the ICASSP 2022 Deep Noise Suppression (DNS) challenge. Our system specifically combines the dual-stage network which is a superior real-time speech enhancement framework with the ECAPA-TDNN speaker embedding network which achieves state-of-the-art performance in speaker verification. The dual-stage network aims to decouple the primal speech enhancement problem into multiple easier sub-problems. Specifically, in stage 1, only the magnitude of the target speech is estimated, which is incorporated with the noisy phase to obtain a coarse complex spectrum estimation. To facilitate the formal estimation, in stage 2, an auxiliary network serves as a post-processing module, where residual noise and interfering speech are further suppressed and the phase information is effectively modified. With the asymmetric loss function to penalize over-suppression, more target speech is preserved, which is helpful for both speech recognition performance and subjective sense of hearing. Our system reaches 3.97 in overall audio quality (OVRL) MOS and 0.69 in word accuracy (WAcc) on the blind test set of the challenge, which outperforms the DNS baseline by 0.57 OVRL and ranks 1st in track 2. Yukai Jv, Wei Rao 0002, Xiaopeng Yan, Yihui Fu, Shubo Lv, Luyao Cheng, Yannan Wang, Lei Xie 0001, Shidong Shang |
ICASSP | 9 |
| 2022 | Speech Enhancement with Fullband-Subband Cross-Attention Network
Jun Chen 0024, Wei Rao 0002, Zilin Wang 0002, Zhiyong Wu 0001, Yannan Wang, Shidong Shang, Helen M. Meng |
INTERSPEECH | 7 |
| 2022 | ConferencingSpeech 2022 Challenge: Non-intrusive Objective Speech Quality Assessment (NISQA) Challenge for Online Conferencing ApplicationsabstractWith the advances in speech communication systems such as online conferencing applications, we can seamlessly work with people regardless of where they are. However, during online meetings, speech quality can be significantly affected by background noise, reverberation, packet loss, network jitter, etc. Because of its nature, speech quality is traditionally assessed in subjective tests in laboratories and lately also in crowdsourcing following the international standards from ITU-T Rec. P.800 series. However, those approaches are costly and cannot be applied to customer data. Therefore, an effective objective assessment approach is needed to evaluate or monitor the speech quality of the ongoing conversation. The ConferencingSpeech 2022 challenge targets the non-intrusive deep neural network models for the speech quality assessment task. We open-sourced a training corpus with more than 86K speech clips in different languages, with a wide range of synthesized and live degradations and their corresponding subjective quality scores through crowdsourcing. 18 teams submitted their models for evaluation in this challenge. The blind test sets included about 4300 clips from wide ranges of degradations. This paper describes the challenge, the datasets, and the evaluation methods and reports the final results. Gaoxiong Yi, Babak Naderi, Sebastian Möller 0001, Wafaa Wardah, Gabriel Mittag, Ross Cutler, Zhuohuang Zhang, Donald S. Williamson, Fei Chen 0011, Shidong Shang |
INTERSPEECH | 13 |
| 2022 | TEA-PSE 2.0: Sub-Band Network for Real-Time Personalized Speech EnhancementabstractPersonalized speech enhancement (PSE) utilizes additional cues like speaker embeddings to remove background noise and interfering speech and extract the speech from target speaker. Previous work, the Tencent-Ethereal-Audio-Lab personalized speech enhancement (TEA-PSE) system, ranked 1st in the ICASSP 2022 deep noise suppression (DNS2022) challenge. In this paper, we expand TEA-PSE to its sub-band version - TEA-PSE 2.0, to reduce computational complexity as well as further improve performance. Specifically, we adopt finite impulse response filter banks and spectrum splitting to reduce computational complexity. We introduce a time frequency convolution module (TFCM) to the system for increasing the receptive field with small convolution kernels. Besides, we explore several training strategies to optimize the two-stage network and investigate various loss functions in the PSE task. TEA-PSE 2.0 significantly outperforms TEA-PSE in both speech enhancement performance and computation complexity. Experimental results on the DNS2022 blind test set show that TEA-PSE 2.0 brings 0.102 OVRL personalized DNSMOS improvement with only 21.9% multiply-accumulate operations compared with the previous TEA-PSE. Yukai Jv, Wei Rao 0002, Yannan Wang, Lei Xie 0001, Shidong Shang |
SLT | 7 |
| 2021 | Conferencingspeech Challenge: Towards Far-Field Multi-Channel Speech Enhancement for Video ConferencingabstractThe ConferencingSpeech 2021 challenge is proposed to stimulate research on far-field multi-channel speech enhancement for video conferencing. The challenge consists of two separate tasks: 1) Task 1 is multi-channel speech enhancement with single microphone array and focusing on practical application with real-time requirement and 2) Task 2 is multi-channel speech enhancement with multiple distributed micro-phone arrays, which is a non-real-time track and does not have any constraints so that participants could explore any algorithms to obtain high speech quality. Targeting the real video conferencing room application, the challenge database was recorded from real speakers and all recording facilities were located by following the real setup of conferencing room. In this challenge, we open-sourced the list of open source clean speech and noise datasets, simulation scripts, and a baseline system for participants to develop their own system. The final ranking of the challenge will be decided by the subjective evaluation which is performed using Absolute Category Ratings (ACR) to estimate Mean Opinion Score (MOS), speech MOS (S-MOS), and noise MOS (N-MOS). This paper describes the challenge, tasks, datasets, subjective evaluation, and challenge results. The baseline system which is a complex ratio mask based neural network and its experimental results are also presented. Wei Rao 0002, Yihui Fu, Yanxin Hu, Yvkai Jv, Jiangyu Han, Zhongjie Jiang, Lei Xie 0001, Yannan Wang, Shinji Watanabe 0001, Zheng-Hua Tan, Hui Bu, Shidong Shang |
ASRU | 14 |
| 2021 | A Partitioned-Block Frequency-Domain Adaptive Kalman Filter for Stereophonic Acoustic Echo Cancellation
Feiran Yang 0001, Yuepeng Li, Shidong Shang |
Interspeech | 4 |