Lei Sun 0010

dblp:02/2264-10 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-7680-6455ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
Yusheng Dai, Jun Du 0002, Lei Sun 0010, Jianqing Gao, Ruoyu Wang 0029, Jiefeng Ma
ICCV7
2025 AudioAtlas: A Comprehensive and Balanced Benchmark Towards Movie-Oriented Text-to-Audio Generation
abstract
Recent rapid progress in Text-to-Audio (T2A) models contrasts sharply with the stagnation observed in the evolution of corresponding evaluation benchmarks. Existing benchmarks, such as AudioCaps, suffer from limited diversity and quality, as well as biased category distributions, leading to increasingly questionable reliability in assessing advanced T2A models. This paper introduces AudioAtlas, a comprehensive and balanced evaluation benchmark specifically designed for evaluating T2A models aimed at movie production. Based on an object-centric audio category system, AudioAtlas provides high-quality reference samples characterized by categorical balance and diversity. It includes detailed overall and event-level captions with rich descriptors, plus fine-grained temporal annotations from human experts, enabling thorough evaluation of temporal alignment and semantic accuracy. To enable precise evaluation of temporally-aligned generation across universal categories, two novel metrics are proposed leveraging recent advancements in large-scale Audio Language Models (AudioLLMs) and contrastive learning models. By re-benchmarking six currently influential T2A models, AudioAtlas provides evaluations better aligned with aesthetic considerations, offering clearer optimization directions for movie-production-oriented T2A systems. Additionally, we conduct a comprehensive comparative analysis on temporally-controllable T2A methods with training-based, and promising training-free approaches inspired by region-controllable image generation, clarifying current limitations and pointing out directions for future research. Audio specifically refers to sound event excluding speech and music. Further details are available on the project page: https://audioatlas.github.io/AudioAtlas/
Yusheng Dai, Lei Sun 0010, Jun Du 0002, Jianqing Gao
ACM Multimedia3
2024 A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition
abstract
Deep learning (DL)-based speaker diarization methods have proven powerful performance comparing to traditional clustering-based methods for multi-talker speech diarization and recognition in farfield scenes. However, most DL-based approaches cannot utilize the spatial information well due to the poor robustness to unknown array topology and acoustic scenario. In this paper, a spatial long-term iterative mask estimation (SLT-IME) method is proposed to improve the performance of speaker diarization in various real-world acoustic scenarios. First, the complex angular central gaussian mixture model (cACGMM) with diarization results as initial values is used to estimate the presence probability of each speaker at each time-frequency bin, namely speaker masks, in a long-term chunk. Then, the speaker masks are converted to speaker activities according to the threshold, which deliver the diarization information of which speaker is active and when. Finally, the estimated speaker activity can also serve as the initial input for the diarization system, resulting in improved ASR performance. Experimental results on the CHiME-7 three datasets (CHiME-6, DiPCo, Mixer 6) show proposed method can improve diarization and recognition systems performance simultaneously. It also plays a key role in the ensemble system that achieves the best performance in the main track of CHiME-7 DASR Challenge.
Yanhui Tu, Maokui He, Ruoyu Wang 0029, Shutong Niu, Lei Sun 0010, Zhongfu Ye, Jun Du 0002, Chin-Hui Lee 0001
ICASSP6
2024 Implicit Enhancement of Target Speaker in Speaker-Adaptive ASR through Efficient Joint Optimization
abstract
In multi-speaker scenarios, automatic speech recognition (ASR) models rely on pre-processed audio after speaker separation. However, when the target speaker is not accurately separated, ASR models face limitations in reaching their peak performance. To address this issue, we propose a speaker-adaptive ASR framework that possesses more implicit target speaker enhancement capability by efficiently joint-optimized speaker recognition (SR) and ASR models. Our framework introduces sharing self-supervised learning representation, optimization transfer and hierarchy speaker-gated attention. In this manner, it can maximize effectiveness of embedding bias and emphasize target speaker corresponding to semantic units. In the CHiME-7 DASR sub-track, the proposed method achieves a 28.19% relative reduction in word error rate (WER) on the development sets when compared to the official baseline. Notably, this framework has also been employed in the champion system for the CHiME-7 DASR.
Haitao Tang 0001, Jiahuan Fan, Ruoyu Wang 0029, Hang Chen 0001, Yanyong Zhang, Jun Du 0002, Hengshun Zhou, Lei Sun 0010, Tian Gao 0005, Genshun Wan, Jianqing Gao
ICASSP9
2023 An Experimental Study on Sound Event Localization and Detection Under Realistic Testing Conditions
abstract
We study four data augmentation (DA) techniques and two model architectures on realistic data for sound event localization and detection (SELD). First, based on ResNet-Conformer (RC), we compare the four DA approaches on the realistic DCASE 2022 SELD test set which is often not easy to handle due to room reverberations and audio overlaps in spontaneous recordings. Experimental results show that, except for audio channel swapping (ACS), the other three data augmentation methods that work well on the simulated SELD data set are no longer effective due to mismatches between simulated and realistic conditions. Next, using ACS-based augmentation, the two improved ResNet-Conformer networks further enhance SELD performances in realistic conditions. By incorporating these two sets of techniques, our overall system ranked the first place in SELD task of the DCASE 2022 Challenge.
Shutong Niu, Jun Du 0002, Qing Wang 0008, Li Chai 0002, Huaxin Wu, Zhaoxu Nian, Lei Sun 0010, Chin-Hui Lee 0001
ICASSP7
2023 Reducing the GAP Between Streaming and Non-Streaming Transducer-Based ASR by Adaptive Two-Stage Knowledge Distillation
abstract
Transducer is one of the mainstream frameworks for streaming speech recognition. There is a performance gap between the streaming and non-streaming transducer models due to limited context. To reduce this gap, an effective way is to ensure that their hidden and output distributions are consistent, which can be achieved by hierarchical knowledge distillation. However, it is difficult to ensure the distribution consistency simultaneously because the learning of the output distribution depends on the hidden one. In this paper, we propose an adaptive two-stage knowledge distillation method consisting of hidden layer learning and output layer learning. In the former stage, we learn hidden representation with full context by applying mean square error loss function. In the latter stage, we design a power transformation based adaptive smoothness method to learn stable output distribution. It achieved 19% relative reduction in word error rate, and a faster response for the first token compared with the original streaming model in LibriSpeech corpus.
Haitao Tang 0001, Yu Fu 0008, Lei Sun 0010, Jiabin Xue, Genshun Wan, Ming'en Zhao
ICASSP3
2023 QDM-SSD: Quality-Aware Dynamic Masking for Separation-Based Speaker Diarization
abstract
We improve iterative separation-based speaker diarization (ISSD) with quality-aware dynamic masking (QDM). We call the proposed framework QDM-SSD. Compared with ISSD, QDM-SSD enhances the simulated data used for model adaptation through QDM to alleviate the influence of errors in speaker priors. In addition to data quality purification, QDM-SSD also makes the adaptation data sparse by automatically adjusting speaker overlap ratios according to data quality. Furthermore, using a sliding window over the adaptation data, clean regions in speech segments can be better localized. Experiments on the two-speaker conversational telephone speech (CTS) corpus show that the proposed QDM-SSD framework can reduce the diarization error rate (DER) by 18.56% relatively compared with ISSD. Moreover, QDM-SSD is shown to generalize to other two-speaker non-conversation telephone speech data sets where ISSD fails to work. Finally, we demonstrate that QDM-SSD can serve as a front-end to improve the performances of back-end automatic speech recognition.
Shutong Niu, Jun Du 0002, Lei Sun 0010, Yu Hu 0003, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Improving Separation-Based Speaker Diarization Via Iterative Model Refinement And Speaker Embedding Based Post-Processing
abstract
In this paper, we propose an iterative separation-based speaker diarization (ISSD) approach to cope with the realistic data conditions. In the proposed ISSD, we iteratively generate adaptation data ac-cording to speaker priors and fine-tune the separation model, which leads to a gradual performance improvement. To further reduce some unavoidable speaker detection errors due to some undesirable prior errors using simple ISSD, we utilize speaker embedding information and propose two post-processing techniques, namely, speaker filtering and speaker recovery. We evaluate the diarization performance on the two-speaker conversational telephone speech (CTS) data set from DIHARD-III Challenge. When compared to state-of-the-art clustering-based speaker diarization (CSD) system, the proposed ISSD approach combined with the two post-processing schemes yields a 47.72 % and 46.97 % relative diarization error rate reduction on the development and evaluation sets, respectively. ISSD is also one key contributing factor to the best-performing system in DIHARD-III Challenge.
Shutong Niu, Jun Du 0002, Lei Sun 0010, Chin-Hui Lee 0001
ICASSP3
2022 External Text Based Data Augmentation for Low-Resource Speech Recognition in the Constrained Condition of OpenASR21 Challenge
Guolong Zhong, Hongyu Song, Ruoyu Wang 0029, Lei Sun 0010, Diyuan Liu, Jun Du 0002, Jie Zhang 0042, Li-Rong Dai 0001
INTERSPEECH4
2021 Scenario-Dependent Speaker Diarization for DIHARD-III Challenge
Jun Du 0002, Maokui He, Shutong Niu, Lei Sun 0010, Chin-Hui Lee 0001
Interspeech5
2020 Progressive Multi-Target Network Based Speech Enhancement with Snr-Preselection for Robust Speaker Diarization
abstract
In this paper, we design a novel front-end processing system for speaker diarization under realistic conditions with challenging background noises. To cope with diversified environments, we first extend our perviously proposed progressive learning based speech enhancement model by adding multi-task learning in each intermediate layer. The corresponding progressive multi-target (PMT) in various layers includes both progressive ratio mask (PRM) and progressively enhanced log-power spectra (PELPS) with specified signal-to-noise ratios (SNRs). Speech distortions are commonly introduced during the front-end processing, which often deteriorate the back-end performance. However, the proposed speech enhancement model can be regarded as a bagging of models with multiple learning objectives, which provides flexibility for selecting the most appropriate output for robust speaker diarzation. In addition, a global SNR estimation is performed using the results of deep neural network (DNN) based speech activity detection (SAD) to decide whether the audio should be enhanced. We evaluate the speaker diarzation performance on the second DIHARD dataset which includes several different realistic conditions. Compared with the original data, experiments demonstrate that the enhanced data processed by our proposed method can effectively avoid the performance loss of every single domain, and achieve consistent improvements in most domains.
Lei Sun 0010, Jun Du 0002, Xueyang Zhang, Tian Gao 0005, Chin-Hui Lee 0001
ICASSP1
2020 A Study of Child Speech Extraction Using Joint Speech Enhancement and Separation in Realistic Conditions
abstract
In this paper, we design a novel joint framework of speech enhancement and speech separation for child speech extraction in realistic conditions, targeting the problem of extracting child speech from daily conversations in BabyTrain mega corpus. To the best of our knowledge, it is the first discussion of a feasible method for child speech extraction in realistic conditions. First, we make detailed analysis of the BabyTrain mega corpus, which is recorded in adverse environments. We observe problems of background noises, reverberations and child speech that is partially obscured by adult speech (for instance due to speaker overlap but also imitation by the adult). Motivated by this, we conduct a joint framework of speech enhancement and speech separation for child speech extraction. To measure the extraction results in realistic conditions, we propose several objective measurements to evaluate the performance of the our system, which is different from those commonly used for simulation data. Compared with the unprocessed approach and classification approach, our proposed approach can yield the best performance among all subsets of BabyTrain.
Xin Wang 0037, Jun Du 0002, Alejandrina Cristià, Lei Sun 0010, Chin-Hui Lee 0001
ICASSP4
2020 A Space-and-Speaker-Aware Iterative Mask Estimation Approach to Multi-Channel Speech Recognition in the CHiME-6 Challenge
Yanhui Tu, Jun Du 0002, Lei Sun 0010, Chin-Hui Lee 0001
INTERSPEECH3
2019 Channel Adversarial Training for Cross-channel Text-independent Speaker Recognition
abstract
The conventional speaker recognition frameworks (e.g., the i-vector and CNN-based approach) have been successfully applied to various tasks when the channel of the enrolment dataset is similar to that of the test dataset. However, in real-world applications, mismatch always exists between these two datasets, which may severely deteriorate the recognition performance. Previously, a few channel compensation algorithms have been proposed, such as Linear Discriminant Analysis (LDA) and Probabilistic LDA. However, these methods always require the collections of different channels from a specific speaker, which is unrealistic to be satisfied in real scenarios. Inspired by domain adaptation, we propose a novel deep-learning based speaker recognition framework to learn the channel-invariant and speaker-discriminative speech representations via channel adversarial training. Specifically, we first employ a gradient reversal layer to remove variations across different channels. Then, the compressed information is projected into the same subspace by adversarial training. Experiments on test datasets with 54,133 speakers demonstrate that the proposed method is not only effective at alleviating the channel mismatch problem, but also outperforms state-of-the-art speaker recognition methods. Compared with the i-vector-based method and the CNN-based method, our proposed method achieves significant relative improvement of 44.7% and 22.6% respectively in terms of the Top1 recall.
Liang Zou, Lei Sun 0010, Zhen-Hua Ling
ICASSP4
2019 A Two-stage Single-channel Speaker-dependent Speech Separation Approach for Chime-5 Challenge
abstract
In this paper, we design a two-stage single-channel speaker-dependent speech separation approach for the CHiME-5 Challenge, targeting the problem of far-field and multi-talker conversational speech recognition in dinner party scenarios involving background noises, reverberations and overlapping speech. First, we make detailed analysis of the CHiME-5 data and observe problems of inaccurate human annotations and low-resource useable data for target speakers. Motivated by this, we conduct a first-stage speaker-dependent speech separation with a learning target for aggressive segregation to generate more and purer target speech data. Then a second-stage speaker-dependent speech separation with a new learning target is performed to obtain the final speech masks, which can be directly fed to back-end acoustic model. Compared with the official baseline, our proposed approach can yield an absolute word error rate reduction of 5.3%, namely from 81.3% to 76.0% in development test set. To the best of our knowledge, it is the first time to discuss a feasible method of single-channel speaker-dependent speech separation for such a challenging task although we make an assumption of oracle speaker diarization following the challenge rules. By integrating this crucial technique, our submitted systems achieved the first place of all four tasks in the CHiME-5 challenge.
Lei Sun 0010, Jun Du 0002, Tian Gao 0005, Chin-Hui Lee 0001
ICASSP1
2019 An iterative mask estimation approach to deep learning based multi-channel speech recognition
Yanhui Tu, Jun Du 0002, Lei Sun 0010, Hai-Kun Wang, Jingdong Chen, Chin-Hui Lee 0001
Speech Commun.3
2018 Enhancement and Analysis of Conversational Speech: JSALT 2017
abstract
Automatic speech recognition is more and more widely and effectively used. Nevertheless, in some automatic speech analysis tasks the state of the art is surprisingly poor. One of these is “diarization”, the task of determining who spoke when. Diarization is key to processing meeting audio and clinical interviews, extended recordings such as police body cam or child language acquisition data, and any other speech data involving multiple speakers whose voices are not cleanly separated into individual channels. Overlapping speech, environmental noise and suboptimal recording techniques make the problem harder. During the JSALT Summer Workshop at CMU in 2017, an international team of researchers worked on several aspects of this problem, including calibration of the state of the art, detection of overlaps, enhancement of noisy recordings, and classification of shorter speech segments. This paper sketches the workshop's results, and announces plans for a “Diarization Challenge” to encourage further progress.
Neville Ryant, Elika Bergelson, Kenneth Church 0001, Alejandrina Cristià, Jun Du 0002, Sriram Ganapathy, Sanjeev Khudanpur, Diana Kowalski, Mahesh Krishnamoorthy, Rajat Kulshreshta, Mark Y. Liberman, Yu-Ding Lu, Matthew Maciejewski, Florian Metze, Ján Profant, Lei Sun 0010, Yu Tsao 0001
ICASSP16
2018 A Novel LSTM-Based Speech Preprocessor for Speaker Diarization in Realistic Mismatch Conditions
abstract
In this study, we investigate on the effects of deep learning based speech enhancement as a preprocessor to speaker diarization in quite challenging realistic environments involving the background noises, reverberations and overlapping speech. To improve the generalization capability, the advanced long short-term memory (LSTM) architecture with the novel design of hidden layers via densely connected progressive learning and output layer via multiple-target learning is proposed for preprocessing. We build the deep model using synthesized training data pairs generated from WSJO reading-style speech and more than 100 noise types. Surprisingly, this proposed preprocessor demonstrates a strong generalization capability to speaker di-arization with the realistic noisy speech in highly mismatched conditions, in terms of the speaking style, interferences, and the interaction between them. Tested on three challenging tasks, namely AMI, ADOS, and SeedLings, the state-of-the-art diarization system with the novel LSTM-based speech preprocessor can yield consistent and significant reductions of diarization error rate (DER) over the systems using unprocessed noisy speech and traditional enhancement methods.
Lei Sun 0010, Jun Du 0002, Tian Gao 0005, Yu-Ding Lu, Yu Tsao 0001, Chin-Hui Lee 0001, Neville Ryant
ICASSP1
2018 Speaker Diarization with Enhancing Speech for the First DIHARD Challenge
Lei Sun 0010, Jun Du 0002, Xueyang Zhang, Chin-Hui Lee 0001
INTERSPEECH1
2017 On Design of Robust Deep Models for CHiME-4 Multi-Channel Speech Recognition with Multiple Configurations of Array Microphones
Yanhui Tu, Jun Du 0002, Lei Sun 0010, Chin-Hui Lee 0001
INTERSPEECH3