Yinghan Cao

dblp:394/5166 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021
YearPublicationVenuePosition
2026 ADS-BiMamba: Attentive Dynamic-Split Bidirectional Mamba for Multi-Channel Speech Enhancement
Shiyun Xu, Yinghan Cao, Changjun He, Mingjiang Wang
IEEE Signal Process. Lett.2
2026 CARVE: Content-Adaptive Rate-Variable Encoding for Neural Speech Codecs
abstract
Neural speech codecs that convert continuous wave forms into discrete tokens are an essential component for building compact speech representations in downstream applications. However, most existing codecs require both fixed and high frame rates to maintain reconstruction quality, which reduces the efficiency of downstream applications. To address this issue, we propose CARVE, a Content-Adaptive Rate-Variable Encoding strategy that performs compression at an externally specified frame rate within the latent space of a neural speech codec. CARVE comprises three modules: (i) Unsupervised Density Aware Scoring Module that combines inter-frame similarity with redundancy-aware cues to quantify the relative importance of each frame; (ii) Sparse Frame Segmentation Module that allocates finer segmentation to high-density information regions and coarser segmentation to low-density information regions based on these scores; and (iii) Context-Aware Frame Reconstruction Module that maps segment-level context back to frame-level latent representations. Integrating CARVE into the neural speech codec, extensive experiments show that it can perform compression at an externally specified frame rate while maintaining high reconstruction quality, thereby validating its effectiveness and practicality in low-frame-rate neural speech coding scenarios. Audio samples are available at: https://ethuil.github.io/CARVEdemo/
Yukun Qian, Yinghan Cao, Changjun He, Shiyun Xu, Mingjiang Wang
IEEE Signal Process. Lett.3
2025 Hybrid Feature Global Attention Network for Noisy-reverberant Speech Enhancement
abstract
Deep neural network-based speech enhancement methods have become widespread, with one of its fundamental aspects being the effective extraction and application of features in the time-frequency domain. This paper proposes a hybrid feature global attention network (HFGANet) designed to efficiently extract time-frequency domain features. HFGANet incorporates a hybrid gated multilayer perceptron (HgMLP) that effectively captures local, global, and inter-window features in the time-frequency domain to create hybrid representations. In contrast to traditional convolutional recurrent neural network architectures, this paper innovatively proposes a global attention structure to leverage these hybrid features. The proposed global attention block enhances the integration of local and global features. Additionally, we introduce Temporal Mamba and Frequency Mamba to further improve the model's ability to capture contextual information in both time and frequency dimensions. On the 1st Deep Noise Suppression Challenge blind test set with reverberation, HFGANet achieves 3.51 WB-PESQ, 95.03% STOI, and 17.72 SI-SDR, while maintaining a lower parameter count compared to state-of-the-art models. In the task of noisy-reverberant speech enhancement, our model achieved an improvement of 1.22 in PESQ, 17.4% in STOI, and 1.63 in DNSMOS.
Shiyun Xu, Yinghan Cao, Changjun He
ICASSP3
2025 PriorSinger: Singing Voice Synthesis Model with Prior Condition Cross Attention
abstract
The singing voice synthesis system is designed to generate realistic and expressive singing based on a given musical score. Generative Adversarial Networks (GANs) or diffusion models generate acoustic features, such as Mel-spectrograms, which are subsequently reconstructed into waveforms by a vocoder. In this work, the musical score is encoded as a prior condition to guide the diffusion denoiser through a novel prior cross-attention Transformer during the denoising process. Moreover, we introduce attention mechanisms in both the time and frequency domains within the diffusion denoiser to enhance the resolution of the generated acoustic features. Additionally, incorporating rotary positional encoding allows the model to better handle temporal and frequency positional information. Our model is capable of synthesizing singing with both higher quality and more vivid expressiveness. In subjective evaluations of the Opencpop dataset, our model outperforms state-of-the-art methods.
Bosong Yan, Yinghan Cao, Mingjiang Wang
ICASSP3
2025 SF-AN: A lightweight shuffle Fourier attention network for multi-channel speech enhancement
Shiyun Xu, Yinghan Cao, Yukun Qian, Changjun He, Mingjiang Wang
Speech Commun.2
2025 Two-stage UNet with channel and temporal-frequency attention for multi-channel speech enhancement
Shiyun Xu, Yinghan Cao, Mingjiang Wang
Speech Commun.2
2025 FSTF-AN: Fused Sparse Temporal-Frequency Attentive Network for Multi-Channel Speech Enhancement
abstract
The Transformer has achieved impressive performance in the multi-channel speech enhancement field; however, it struggles to capture local features, which leads to the loss of speech details. To enhance the extraction of local features in the network, we propose a fused sparse temporal-frequency attentive network (FSTF-AN), which aims to fully capture features across the temporal-frequency, frequency, and temporal dimensions. We propose top-$k$fused sparse self-attention, which employs a fusion strategy to adaptively retain the most crucial attention scores when computing self-attention maps, thereby eliminating irrelevant information interference and better aggregating features. Furthermore, we propose a multi-scale fused feed-forward network, which effectively captures multi-scale features, further enhancing the network's ability to capture local features. The experimental results demonstrate that FSTF-AN exhibits significant advantages over other SOTA models, effectively enhancing speech quality and intelligibility.
Shiyun Xu, Yinghan Cao, Mingjiang Wang
IEEE Signal Process. Lett.2