Shulin He

dblp:41/9706 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
20since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Robust Target Speaker Direction of Arrival Estimation
abstract
In multi-speaker environments the direction of arrival (DOA) of a target speaker is key for improving speech clarity and extracting target speaker’s voice. However, traditional DOA estimation methods often struggle in the presence of noise, reverberation, and particularly when competing speakers are present. To address these challenges, we propose RTS-DOA, a robust real-time DOA estimation system. This system innovatively uses the registered speech of the target speaker as a reference and leverages full-band and sub-band spectral information from a microphone array to estimate the DOA of the target speaker’s voice. Specifically, the system comprises a speech enhancement module for initially improving speech quality, a spatial module for learning spatial information, and a speaker module for extracting voiceprint features. Experimental results on the LibriSpeech dataset demonstrate that our RTS-DOA system effectively tackles multi-speaker scenarios and established new optimal benchmarks, achieving an Accuracy Rate that is 0.31 higher than the baseline.
Shulin He
ICASSP2
2025 TF-SkiMNet: Speech Enhancement Based on Inplace Modeling and Skipping Memory in Time-Frequency Domain
Shulin He, Jinglin Bai
INTERSPEECH2
2025 HWB-Net: A Novel High-Performance and Efficient Hybrid Waveform Bandwidth Extension Method
Shulin He
INTERSPEECH2
2025 Room Impulse Response as a Prompt for Acoustic Echo Cancellation
Shulin He
INTERSPEECH2
2025 An Intelligent Framework for Cluster-Based Side-Channel Analysis on Public-Key Cryptosystems
abstract
Classical cluster-based side-channel analysis (SCA) uses clustering algorithms to analyze power traces and often, principal component analysis to reduce the dimension of data, resulting in that clustering may not deal well with high-dimensional traces, such as cryptographic algorithm implementations with countermeasures. In this article, we propose an intelligent framework for cluster-based SCA, which includes three steps of clustering, classification and correction, for processing large high-dimensional data. By combining unsupervised clustering and supervised deep learning techniques, the framework succeeds in mining the data for additional in-depth information. In addition, unlike traditional cluster-based SCA, our approach focuses on deep learning and deliberately avoids over-reliance on cluster labels during classification. And metrics for correction are adopted to achieve a high level of reliability in key recovery. Experiments on the RSA smart card based on Montgomery ladder implementation and FPGA-based ECC with random delay demonstrate that our framework can significantly improve the success rate with strong robustness.
Congming Wei, Shulin He, An Wang 0001, Shaofei Sun, Yaoling Ding, Liehuang Zhu
IEEE Internet Things J.2
2025 Enhancing target speaker extraction with Hierarchical Speaker Representation Learning
abstract
Target speaker extraction aims to obtain the speech of the specific speaker from a mixture of multiple voices. The conventional approach exploits the target speaker embeddings from a pre-recorded speech segment as auxiliary information, providing prior for extraction. However, the naive single-vector embedding may lack attention to the subtle acoustic features such as pitch and harmonic distribution in the auxiliary speech, leading to an unsatisfying performance. Furthermore, traditional speaker embeddings are trained by speaker verification system and do not leverage the semantics of the auxiliary speech which may facilitate the extraction. To address these challenges, we propose a simple yet effective Hierarchical Speaker Representation Learning (HSRL). The proposed method comprises three modules: a Local Speaker Feature Extractor (LSFE), a Global Speaker Feature Extractor (GSFE), and a Hierarchical Cascading Input Strategy (HCIS). Specifically, the LSFE utilizes the fine-grained acoustic information in the anchor speech. In GSFE, we utilize ECAPA-TDNN to obtain the speaker embeddings of the target speaker, enhancing extraction performance with this global speaker information. In additional, a novel HCIS is proposed to integrate the output of the LSFE module to the input of the GSFE, which enables the global speaker features to focus on the semantic content of the pre-recorded speech. Experimental results on the Libri-2talker dataset demonstrate that our HSRL has achieved significant performance improvements and established new optimal benchmarks.
Shulin He, Huaiwen Zhang
Neural Networks1
2024 3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications
abstract
Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-time applications, espetially on resource-constrained platforms. In this paper, we address the TSE task using microphone array and introduce a novel three-stage solution that systematically decouples the process: First, a neural network is trained to estimate the direction of the target speaker. Second, with the direction determined, the Generalized Sidelobe Canceller (GSC) is used to extract the target speech. Third, an Inplace Convolutional Recurrent Neural Network (ICRN) acts as a denoising post-processor, refining the GSC output to yield the final separated speech. Our approach delivers superior performance while drastically reducing computational load, setting a new standard for efficient real-time target speaker extraction.
Shulin He, Hao Li 0046, Yang Yang 0121, Fei Chen 0011, Xueliang Zhang 0001
ICASSP1
2024 Hierarchical Speaker Representation for Target Speaker Extraction
abstract
Target speaker extraction aims to isolate a specific speaker’s voice from a composite of multiple sound sources, guided by an enrollment utterance or called anchor. Current methods predominantly derive speaker embeddings from the anchor and integrate them into the separation network to separate the voice of the target speaker. However, the representation of the speaker embedding is too simplistic, often being merely a 1×1024 vector. This dense information makes it difficult for the separation network to harness effectively. To address this limitation, we introduce a pioneering methodology called Hierarchical Representation (HR) that seamlessly fuses anchor data across granular and overarching 5 layers of the separation network, enhancing the precision of target extraction. HR amplifies the efficacy of anchors to improve target speaker isolation. On the Libri-2talker dataset, HR substantially outperforms state-of-the-art time-frequency domain techniques. Further demonstrating HR’s capabilities, we achieved first place in the prestigious ICASSP 2023 Deep Noise Suppression Challenge. The proposed HR methodology shows great promise for advancing target speaker extraction through enhanced anchor utilization.
Shulin He, Huaiwen Zhang, Wei Rao 0002, Kanghao Zhang, Yukai Jv, Yang Yang 0121, Xueliang Zhang 0001
ICASSP1
2024 SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques
abstract
Speech enhancement aims to improve speech quality and intelligibility, especially in noisy environments where background noise degrades speech signals. Currently, deep learning methods achieve great success in speech enhancement, e.g. the representative convolutional recurrent neural network (CRN) and its variants. However, CRN typically employs consecutive downsampling and upsampling convolution for frequency modeling, which destroys the inherent structure of the signal over frequency. Additionally, convolutional layers lacks of temporal modelling abilities. To address these issues, we propose an innovative module combing a State space model and Inplace Convolution (SIC), and to replace the conventional convolution in CRN, called SICRN. Specifically, a dual-path multidimensional State space model captures the global frequencies dependency and long-term temporal dependencies. Meanwhile, the 2D-inplace convolution is used to capture the local structure, which abandons the downsampling and upsampling. Systematic evaluations on the public INTERSPEECH 2020 DNS challenge dataset demonstrate SICRN’s efficacy. Compared to strong baselines, SICRN achieves performance close to state-of-the-art while having advantages in model parameters, computations, and algorithmic delay. The proposed SICRN shows great promise for improved speech enhancement.
Changjiang Zhao, Shulin He
ICASSP2
2024 Unified Audio Visual Cues for Target Speaker Extraction
Tianci Wu, Shulin He, Zhijian Mo
INTERSPEECH2
2024 Deep Echo Path Modeling for Acoustic Echo Cancellation
abstract
Acoustic echo cancellation (AEC) is a key audio processing technology that removes echoes from microphone inputs to enable natural-sounding full-duplex communication.In recent years, deep learning has shown great potential for advancing AEC.However, deep learning methods face challenges in generalizing to complex environments, especially unseen conditions not represented in training.In this paper, we propose a deep learning-based method to predict the echo path in the time-frequency domain.Specifically, we first estimate the echo path under single-talk scenario without near-end signal and then utilize these predicted echo paths as auxiliary labels to train the model on double-talk scenario with near-end signal.Experimental results show that our method outperforms the strong baselines and exhibits good generalization capabilities for unseen acoustic scenarios.By estimating the echo path using deep learning, this work advances AEC performance in the presence of complex conditions.
Chenggang Zhang, Shulin He
INTERSPEECH3
2024 FlashSpeech: Efficient Zero-Shot Speech Synthesis
abstract
Recent progress in large-scale zero-shot speech synthesis has been significantly advanced by language models and diffusion models. However, the generation process of both methods is slow and computationally intensive. Efficient speech synthesis using a lower computing budget to achieve quality on par with previous work remains a significant challenge. In this paper, we present FlashSpeech, a large-scale zero-shot speech synthesis system with approximately 5% of the inference time compared with previous work. FlashSpeech is built on the latent consistency model and applies a novel adversarial consistency training approach that can train from scratch without the need for a pre-trained diffusion model as the teacher. Furthermore, a new prosody generator module enhances the diversity of prosody, making the rhythm of the speech sound more natural. The generation processes of FlashSpeech can be achieved efficiently with one or two sampling steps while maintaining high audio quality and high similarity to the audio prompt for zero-shot speech generation. Our experimental results demonstrate the superior performance of FlashSpeech. Notably, FlashSpeech can be about 20 times faster than other zero-shot speech synthesis systems while maintaining comparable performance in terms of voice quality and similarity. Furthermore, FlashSpeech demonstrates its versatility by efficiently performing tasks like voice conversion, speech editing, and diverse speech sampling. Audio samples can be found in https://flashspeech.github.io/
Zhen Ye 0006, Zeqian Ju, Haohe Liu, Xu Tan 0003, Jianyi Chen, Peiwen Sun, Weizhen Bian, Shulin He, Wei Xue 0002, Yike Guo
ACM Multimedia10
2023 Gesper: A Unified Framework for General Speech Restoration
abstract
This paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by speech enhancement. We propose a complex spectral mapping-based generative adversarial network (CSM-GAN) as the speech restoration module for the first time. For noise suppression and dereverberation, the enhancement module is presented with fullband-wideband parallel processing. On the blind test set of ICASSP 2023 SSI Challenge, the proposed Gesper system, which satisfies the real-time condition, achieves 3.27 P.804 overall mean opinion score (MOS) and 3.35 P.835 overall MOS, ranked 1st in both track 1 and track 2.
Jun Chen 0024, Yupeng Shi, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001, Shidong Shang, Chengshi Zheng
ICASSP5
2023 Speech Enhancement with Intelligent Neural Homomorphic Synthesis
abstract
Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement. Specifically, we use homomorphic signal processing and cepstral analysis to obtain noisy speech’s excitation and vocal tract. Unlike traditional signal processing, we use an attentive recurrent network (ARN) model predicted ratio mask to replace the liftering separation function. Then two convolutional attentive recurrent network (CARN) networks are used to predict the excitation and vocal tract of clean speech, respectively. The system’s output is synthesized from the estimated excitation and vocal. Experiments prove that our proposed method performs better, with SI-SNR improving by 1.363dB compared to FullSubNet.
Shulin He, Wei Rao 0002, Jun Chen 0024, Yukai Jv, Xueliang Zhang 0001, Yannan Wang, Shidong Shang
ICASSP1
2023 TEA-PSE 3.0: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System For ICASSP 2023 Dns-Challenge
abstract
This paper introduces the Unbeatable Team’s submission to the ICASSP 2023 Deep Noise Suppression (DNS) Challenge. We expand our previous work, TEA-PSE, to its upgraded version – TEA-PSE 3.0. Specifically, TEA-PSE 3.0 incorporates a residual LSTM after squeezed temporal convolution network (S-TCN) to enhance sequence modeling capabilities. Additionally, the local-global representation (LGR) structure is introduced to boost speaker information extraction, and multi-STFT resolution loss is used to effectively capture the time-frequency characteristics of the speech signals. Moreover, retraining methods are employed based on the freeze training strategy to fine-tune the system. According to the official results, TEA-PSE 3.0 ranks 1st in both ICASSP 2023 DNS-Challenge track 1 and track 2.
Yukai Jv, Jun Chen 0024, Shulin He, Wei Rao 0002, Weixin Zhu, Yannan Wang, Shidong Shang
ICASSP4
2023 MC-SpEx: Towards Effective Speaker Extraction with Multi-Scale Interfusion and Conditional Speaker Modulation
Jun Chen 0024, Wei Rao 0002, Zilin Wang 0002, Jiuxin Lin, Yukai Jv, Shulin He, Yannan Wang, Zhiyong Wu 0001
INTERSPEECH6
2023 Gesper: A Restoration-Enhancement Framework for General Speech Reconstruction
Yupeng Shi, Jun Chen 0024, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001
INTERSPEECH5
2022 A Robust Deep Audio Splicing Detection Method via Singularity Detection Feature
abstract
There are many methods for detecting forged audio produced by conversion and synthesis. However, as a simpler method of forgery, splicing has not attracted widespread attention. Based on the characteristic that the tampering operation will cause singularities at high-frequency components, we propose a high-frequency singularity detection feature obtained by wavelet transform. The proposed feature can explicitly show the location of the tampering operation on the waveform. Moreover, the long short-term memory (LSTM) is introduced to the CNN-architecture LCNN to ensure that the sequence information can be fully learned. The proposed feature is sent to the improved RNN-architecture LCNN together with the widely used linear frequency cepstral coefficients (LFCC) to learn forgery characteristics where the LFCC is used as a supplement. Systematic evaluation and comparison show that the proposed method has greatly improved the accuracy and generalization.
Kanghao Zhang, Shan Liang 0007, Shuai Nie 0001, Shulin He, Xueliang Zhang 0001, Haoxin Ma, Jiangyan Yi
ICASSP4
2022 Speaker recognition-assisted robust audio deepfake detection
Shuai Nie 0001, Hui Zhang 0031, Shulin He, Kanghao Zhang, Shan Liang 0007, Xueliang Zhang 0001, Jianhua Tao 0001
INTERSPEECH4
2021 DBNet: A Dual-Branch Network Architecture Processing on Spectrum and Waveform for Single-Channel Speech Enhancement
abstract
In real acoustic environment, speech enhancement is an arduous task to improve the quality and intelligibility of speech interfered by background noise and reverberation.Over the past years, deep learning has shown great potential on speech enhancement.In this paper, we propose a novel real-time framework called DBNet which is a dual-branch structure with alternate interconnection.Each branch incorporates an encoderdecoder architecture with skip connections.The two branches are responsible for spectrum and waveform modeling, respectively.A bridge layer is adopted to exchange information between the two branches.Systematic evaluation and comparison show that the proposed system substantially outperforms related algorithms under very challenging environments.And in INTERSPEECH 2021 Deep Noise Suppression (DNS) challenge, the proposed system ranks the top 8 in real-time track 1 in terms of the Mean Opinion Score (MOS) of the ITU-T P.835 framework.
Kanghao Zhang, Shulin He, Hao Li 0046, Xueliang Zhang 0001
Interspeech2
2020 Speakerfilter: Deep Learning-Based Target Speaker Extraction Using Anchor Speech
abstract
Speaker extraction aims to separate a target speaker from multiple voices which is useful for applications, e.g. teleconference. In many practical cases, it has an opportunity to get a piece voice of the target speaker in advance, which provides useful information for speaker extraction. This paper addresses the problem of extracting the target speaker from the mixture using a short piece of anchor speech. To effectively utilize anchor speech, we propose a multi-level feature extraction and seamlessly integrate the features into a speech separation model. Experiments are conducted on the two-speaker dataset (WSJ0-mix2) which is widely used for speaker extraction. The systematic evaluation shows that the proposed method significantly outperforms the previous methods and achieves a signal-to-distortion ratio (SDR) improvement of 11.3 dB on the unprocessed mixture.
Shulin He, Hao Li 0046, Xueliang Zhang 0001
ICASSP1
2010 An accessibility measure for the combined travel demand model
Shulin He
Sci. China Inf. Sci.2