VLDB 2026 Research / reviewers in the wild / expert
Shulin He
dblp:41/9706
· DBLP profile ↗
22ranked-venue papers
5as first author
20since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Target Speaker Direction of Arrival EstimationabstractIn multi-speaker environments the direction of arrival (DOA) of a target speaker is key for improving speech clarity and extracting target speaker’s voice. However, traditional DOA estimation methods often struggle in the presence of noise, reverberation, and particularly when competing speakers are present. To address these challenges, we propose RTS-DOA, a robust real-time DOA estimation system. This system innovatively uses the registered speech of the target speaker as a reference and leverages full-band and sub-band spectral information from a microphone array to estimate the DOA of the target speaker’s voice. Specifically, the system comprises a speech enhancement module for initially improving speech quality, a spatial module for learning spatial information, and a speaker module for extracting voiceprint features. Experimental results on the LibriSpeech dataset demonstrate that our RTS-DOA system effectively tackles multi-speaker scenarios and established new optimal benchmarks, achieving an Accuracy Rate that is 0.31 higher than the baseline. Shulin He |
ICASSP | 2 |
| 2025 | TF-SkiMNet: Speech Enhancement Based on Inplace Modeling and Skipping Memory in Time-Frequency Domain
Shulin He, Jinglin Bai |
INTERSPEECH | 2 |
| 2025 | HWB-Net: A Novel High-Performance and Efficient Hybrid Waveform Bandwidth Extension Method
Shulin He |
INTERSPEECH | 2 |
| 2025 | Room Impulse Response as a Prompt for Acoustic Echo Cancellation
Shulin He |
INTERSPEECH | 2 |
| 2025 | An Intelligent Framework for Cluster-Based Side-Channel Analysis on Public-Key CryptosystemsabstractClassical cluster-based side-channel analysis (SCA) uses clustering algorithms to analyze power traces and often, principal component analysis to reduce the dimension of data, resulting in that clustering may not deal well with high-dimensional traces, such as cryptographic algorithm implementations with countermeasures. In this article, we propose an intelligent framework for cluster-based SCA, which includes three steps of clustering, classification and correction, for processing large high-dimensional data. By combining unsupervised clustering and supervised deep learning techniques, the framework succeeds in mining the data for additional in-depth information. In addition, unlike traditional cluster-based SCA, our approach focuses on deep learning and deliberately avoids over-reliance on cluster labels during classification. And metrics for correction are adopted to achieve a high level of reliability in key recovery. Experiments on the RSA smart card based on Montgomery ladder implementation and FPGA-based ECC with random delay demonstrate that our framework can significantly improve the success rate with strong robustness. Congming Wei, Shulin He, An Wang 0001, Shaofei Sun, Yaoling Ding, Liehuang Zhu |
IEEE Internet Things J. | 2 |
| 2025 | Enhancing target speaker extraction with Hierarchical Speaker Representation LearningabstractTarget speaker extraction aims to obtain the speech of the specific speaker from a mixture of multiple voices. The conventional approach exploits the target speaker embeddings from a pre-recorded speech segment as auxiliary information, providing prior for extraction. However, the naive single-vector embedding may lack attention to the subtle acoustic features such as pitch and harmonic distribution in the auxiliary speech, leading to an unsatisfying performance. Furthermore, traditional speaker embeddings are trained by speaker verification system and do not leverage the semantics of the auxiliary speech which may facilitate the extraction. To address these challenges, we propose a simple yet effective Hierarchical Speaker Representation Learning (HSRL). The proposed method comprises three modules: a Local Speaker Feature Extractor (LSFE), a Global Speaker Feature Extractor (GSFE), and a Hierarchical Cascading Input Strategy (HCIS). Specifically, the LSFE utilizes the fine-grained acoustic information in the anchor speech. In GSFE, we utilize ECAPA-TDNN to obtain the speaker embeddings of the target speaker, enhancing extraction performance with this global speaker information. In additional, a novel HCIS is proposed to integrate the output of the LSFE module to the input of the GSFE, which enables the global speaker features to focus on the semantic content of the pre-recorded speech. Experimental results on the Libri-2talker dataset demonstrate that our HSRL has achieved significant performance improvements and established new optimal benchmarks. Shulin He, Huaiwen Zhang |
Neural Networks | 1 |
| 2024 | 3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource ApplicationsabstractTarget speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-time applications, espetially on resource-constrained platforms. In this paper, we address the TSE task using microphone array and introduce a novel three-stage solution that systematically decouples the process: First, a neural network is trained to estimate the direction of the target speaker. Second, with the direction determined, the Generalized Sidelobe Canceller (GSC) is used to extract the target speech. Third, an Inplace Convolutional Recurrent Neural Network (ICRN) acts as a denoising post-processor, refining the GSC output to yield the final separated speech. Our approach delivers superior performance while drastically reducing computational load, setting a new standard for efficient real-time target speaker extraction. Shulin He, Hao Li 0046, Yang Yang 0121, Fei Chen 0011, Xueliang Zhang 0001 |
ICASSP | 1 |
| 2024 | Hierarchical Speaker Representation for Target Speaker ExtractionabstractTarget speaker extraction aims to isolate a specific speaker’s voice from a composite of multiple sound sources, guided by an enrollment utterance or called anchor. Current methods predominantly derive speaker embeddings from the anchor and integrate them into the separation network to separate the voice of the target speaker. However, the representation of the speaker embedding is too simplistic, often being merely a 1×1024 vector. This dense information makes it difficult for the separation network to harness effectively. To address this limitation, we introduce a pioneering methodology called Hierarchical Representation (HR) that seamlessly fuses anchor data across granular and overarching 5 layers of the separation network, enhancing the precision of target extraction. HR amplifies the efficacy of anchors to improve target speaker isolation. On the Libri-2talker dataset, HR substantially outperforms state-of-the-art time-frequency domain techniques. Further demonstrating HR’s capabilities, we achieved first place in the prestigious ICASSP 2023 Deep Noise Suppression Challenge. The proposed HR methodology shows great promise for advancing target speaker extraction through enhanced anchor utilization. Shulin He, Huaiwen Zhang, Wei Rao 0002, Kanghao Zhang, Yukai Jv, Yang Yang 0121, Xueliang Zhang 0001 |
ICASSP | 1 |
| 2024 | SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution TechniquesabstractSpeech enhancement aims to improve speech quality and intelligibility, especially in noisy environments where background noise degrades speech signals. Currently, deep learning methods achieve great success in speech enhancement, e.g. the representative convolutional recurrent neural network (CRN) and its variants. However, CRN typically employs consecutive downsampling and upsampling convolution for frequency modeling, which destroys the inherent structure of the signal over frequency. Additionally, convolutional layers lacks of temporal modelling abilities. To address these issues, we propose an innovative module combing a State space model and Inplace Convolution (SIC), and to replace the conventional convolution in CRN, called SICRN. Specifically, a dual-path multidimensional State space model captures the global frequencies dependency and long-term temporal dependencies. Meanwhile, the 2D-inplace convolution is used to capture the local structure, which abandons the downsampling and upsampling. Systematic evaluations on the public INTERSPEECH 2020 DNS challenge dataset demonstrate SICRN’s efficacy. Compared to strong baselines, SICRN achieves performance close to state-of-the-art while having advantages in model parameters, computations, and algorithmic delay. The proposed SICRN shows great promise for improved speech enhancement. Changjiang Zhao, Shulin He |
ICASSP | 2 |
| 2024 | Unified Audio Visual Cues for Target Speaker Extraction
Tianci Wu, Shulin He, Zhijian Mo |
INTERSPEECH | 2 |
| 2024 | Deep Echo Path Modeling for Acoustic Echo CancellationabstractAcoustic echo cancellation (AEC) is a key audio processing technology that removes echoes from microphone inputs to enable natural-sounding full-duplex communication.In recent years, deep learning has shown great potential for advancing AEC.However, deep learning methods face challenges in generalizing to complex environments, especially unseen conditions not represented in training.In this paper, we propose a deep learning-based method to predict the echo path in the time-frequency domain.Specifically, we first estimate the echo path under single-talk scenario without near-end signal and then utilize these predicted echo paths as auxiliary labels to train the model on double-talk scenario with near-end signal.Experimental results show that our method outperforms the strong baselines and exhibits good generalization capabilities for unseen acoustic scenarios.By estimating the echo path using deep learning, this work advances AEC performance in the presence of complex conditions. Chenggang Zhang, Shulin He |
INTERSPEECH | 3 |
| 2024 | FlashSpeech: Efficient Zero-Shot Speech SynthesisabstractRecent progress in large-scale zero-shot speech synthesis has been significantly advanced by language models and diffusion models. However, the generation process of both methods is slow and computationally intensive. Efficient speech synthesis using a lower computing budget to achieve quality on par with previous work remains a significant challenge. In this paper, we present FlashSpeech, a large-scale zero-shot speech synthesis system with approximately 5% of the inference time compared with previous work. FlashSpeech is built on the latent consistency model and applies a novel adversarial consistency training approach that can train from scratch without the need for a pre-trained diffusion model as the teacher. Furthermore, a new prosody generator module enhances the diversity of prosody, making the rhythm of the speech sound more natural. The generation processes of FlashSpeech can be achieved efficiently with one or two sampling steps while maintaining high audio quality and high similarity to the audio prompt for zero-shot speech generation. Our experimental results demonstrate the superior performance of FlashSpeech. Notably, FlashSpeech can be about 20 times faster than other zero-shot speech synthesis systems while maintaining comparable performance in terms of voice quality and similarity. Furthermore, FlashSpeech demonstrates its versatility by efficiently performing tasks like voice conversion, speech editing, and diverse speech sampling. Audio samples can be found in https://flashspeech.github.io/ Zhen Ye 0006, Zeqian Ju, Haohe Liu, Xu Tan 0003, Jianyi Chen, Peiwen Sun, Weizhen Bian, Shulin He, Wei Xue 0002, Yike Guo |
ACM Multimedia | 10 |
| 2023 | Gesper: A Unified Framework for General Speech RestorationabstractThis paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by speech enhancement. We propose a complex spectral mapping-based generative adversarial network (CSM-GAN) as the speech restoration module for the first time. For noise suppression and dereverberation, the enhancement module is presented with fullband-wideband parallel processing. On the blind test set of ICASSP 2023 SSI Challenge, the proposed Gesper system, which satisfies the real-time condition, achieves 3.27 P.804 overall mean opinion score (MOS) and 3.35 P.835 overall MOS, ranked 1st in both track 1 and track 2. Jun Chen 0024, Yupeng Shi, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001, Shidong Shang, Chengshi Zheng |
ICASSP | 5 |
| 2023 | Speech Enhancement with Intelligent Neural Homomorphic SynthesisabstractMost neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement. Specifically, we use homomorphic signal processing and cepstral analysis to obtain noisy speech’s excitation and vocal tract. Unlike traditional signal processing, we use an attentive recurrent network (ARN) model predicted ratio mask to replace the liftering separation function. Then two convolutional attentive recurrent network (CARN) networks are used to predict the excitation and vocal tract of clean speech, respectively. The system’s output is synthesized from the estimated excitation and vocal. Experiments prove that our proposed method performs better, with SI-SNR improving by 1.363dB compared to FullSubNet. Shulin He, Wei Rao 0002, Jun Chen 0024, Yukai Jv, Xueliang Zhang 0001, Yannan Wang, Shidong Shang |
ICASSP | 1 |
| 2023 | TEA-PSE 3.0: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System For ICASSP 2023 Dns-ChallengeabstractThis paper introduces the Unbeatable Team’s submission to the ICASSP 2023 Deep Noise Suppression (DNS) Challenge. We expand our previous work, TEA-PSE, to its upgraded version – TEA-PSE 3.0. Specifically, TEA-PSE 3.0 incorporates a residual LSTM after squeezed temporal convolution network (S-TCN) to enhance sequence modeling capabilities. Additionally, the local-global representation (LGR) structure is introduced to boost speaker information extraction, and multi-STFT resolution loss is used to effectively capture the time-frequency characteristics of the speech signals. Moreover, retraining methods are employed based on the freeze training strategy to fine-tune the system. According to the official results, TEA-PSE 3.0 ranks 1st in both ICASSP 2023 DNS-Challenge track 1 and track 2. Yukai Jv, Jun Chen 0024, Shulin He, Wei Rao 0002, Weixin Zhu, Yannan Wang, Shidong Shang |
ICASSP | 4 |
| 2023 | MC-SpEx: Towards Effective Speaker Extraction with Multi-Scale Interfusion and Conditional Speaker Modulation
Jun Chen 0024, Wei Rao 0002, Zilin Wang 0002, Jiuxin Lin, Yukai Jv, Shulin He, Yannan Wang, Zhiyong Wu 0001 |
INTERSPEECH | 6 |
| 2023 | Gesper: A Restoration-Enhancement Framework for General Speech Reconstruction
Yupeng Shi, Jun Chen 0024, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001 |
INTERSPEECH | 5 |
| 2022 | A Robust Deep Audio Splicing Detection Method via Singularity Detection FeatureabstractThere are many methods for detecting forged audio produced by conversion and synthesis. However, as a simpler method of forgery, splicing has not attracted widespread attention. Based on the characteristic that the tampering operation will cause singularities at high-frequency components, we propose a high-frequency singularity detection feature obtained by wavelet transform. The proposed feature can explicitly show the location of the tampering operation on the waveform. Moreover, the long short-term memory (LSTM) is introduced to the CNN-architecture LCNN to ensure that the sequence information can be fully learned. The proposed feature is sent to the improved RNN-architecture LCNN together with the widely used linear frequency cepstral coefficients (LFCC) to learn forgery characteristics where the LFCC is used as a supplement. Systematic evaluation and comparison show that the proposed method has greatly improved the accuracy and generalization. Kanghao Zhang, Shan Liang 0007, Shuai Nie 0001, Shulin He, Xueliang Zhang 0001, Haoxin Ma, Jiangyan Yi |
ICASSP | 4 |
| 2022 | Speaker recognition-assisted robust audio deepfake detection
Shuai Nie 0001, Hui Zhang 0031, Shulin He, Kanghao Zhang, Shan Liang 0007, Xueliang Zhang 0001, Jianhua Tao 0001 |
INTERSPEECH | 4 |
| 2021 | DBNet: A Dual-Branch Network Architecture Processing on Spectrum and Waveform for Single-Channel Speech EnhancementabstractIn real acoustic environment, speech enhancement is an arduous task to improve the quality and intelligibility of speech interfered by background noise and reverberation.Over the past years, deep learning has shown great potential on speech enhancement.In this paper, we propose a novel real-time framework called DBNet which is a dual-branch structure with alternate interconnection.Each branch incorporates an encoderdecoder architecture with skip connections.The two branches are responsible for spectrum and waveform modeling, respectively.A bridge layer is adopted to exchange information between the two branches.Systematic evaluation and comparison show that the proposed system substantially outperforms related algorithms under very challenging environments.And in INTERSPEECH 2021 Deep Noise Suppression (DNS) challenge, the proposed system ranks the top 8 in real-time track 1 in terms of the Mean Opinion Score (MOS) of the ITU-T P.835 framework. Kanghao Zhang, Shulin He, Hao Li 0046, Xueliang Zhang 0001 |
Interspeech | 2 |
| 2020 | Speakerfilter: Deep Learning-Based Target Speaker Extraction Using Anchor SpeechabstractSpeaker extraction aims to separate a target speaker from multiple voices which is useful for applications, e.g. teleconference. In many practical cases, it has an opportunity to get a piece voice of the target speaker in advance, which provides useful information for speaker extraction. This paper addresses the problem of extracting the target speaker from the mixture using a short piece of anchor speech. To effectively utilize anchor speech, we propose a multi-level feature extraction and seamlessly integrate the features into a speech separation model. Experiments are conducted on the two-speaker dataset (WSJ0-mix2) which is widely used for speaker extraction. The systematic evaluation shows that the proposed method significantly outperforms the previous methods and achieves a signal-to-distortion ratio (SDR) improvement of 11.3 dB on the unprocessed mixture. Shulin He, Hao Li 0046, Xueliang Zhang 0001 |
ICASSP | 1 |
| 2010 | An accessibility measure for the combined travel demand model
Shulin He |
Sci. China Inf. Sci. | 2 |