Guochen Yu

dblp:277/3520 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-7179-1044ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021
YearPublicationVenuePosition
2026 SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis
abstract
Although lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as mel-spectrograms or discrete self-supervised learning (SSL) tokens. The potential of latent diffusion models (LDMs) in this task remains largely unexplored. In this paper, we introduce SLD-L2S, a novel L2S framework built upon a hierarchical subspace latent diffusion model. Our method aims to directly map visual lip movements to the continuous latent space of a pre-trained neural audio codec, thereby avoiding the information loss inherent in traditional intermediate representations. The core of our method is a hierarchical architecture that processes visual representations through multiple parallel subspaces, initiated by a subspace decomposition module. To efficiently enhance interactions within and between these subspaces, we design the diffusion convolution block (DiCB) as our network backbone. Furthermore, we employ a reparameterized flow matching technique to directly generate the target latent vectors. This enables a principled inclusion of speech language model (SLM) and semantic losses during training, moving beyond conventional flow matching objectives and improving synthesized speech quality. Our experiments show that SLD-L2S achieves state-of-the-art generation quality on multiple benchmark datasets, surpassing existing methods in both objective and subjective evaluations.
Yifan Liang, Andong Li, Guochen Yu, Fangkun Liu, Lingling Dai, Xiaodong Li 0002, Chengshi Zheng
AAAI4
2024 BAE-Net: a Low Complexity and High Fidelity Bandwidth-Adaptive Neural Network for Speech Super-Resolution
abstract
Speech bandwidth extension (BWE) has demonstrated promising performance in enhancing the perceptual speech quality in real communication systems. Most existing BWE researches primarily focus on fixed upsampling ratios, disregarding the fact that the effective bandwidth of captured audio may fluctuate frequently due to various capturing devices and transmission conditions. In this paper, we propose a streaming adaptive bandwidth extension solution dubbed BAE-Net, which is suitable to handle the low-resolution speech with unknown and varying effective bandwidth. To address the challenges of recovering both the high-frequency magnitude and phase components of the speech content blindly, we devise a dual-stream architecture that incorporates the magnitude inpainting and phase refinement. For potential applications on edge devices, this paper also introduces BAE-NET-lite, which is a lightweight, streaming and efficient framework. Quantitative results demonstrate the superiority of BAE-Net in terms of performance and computational efficiency when compared with existing state-of-the-art BWE methods.
Guochen Yu, Xiguang Zheng, Runqiang Han, Chengshi Zheng
ICASSP1
2023 TaylorBeamixer: Learning Taylor-Inspired All-Neural Multi-Channel Speech Enhancement from Beam-Space Dictionary Perspective
Andong Li, Weixin Meng, Guochen Yu, Xiaodong Li 0002, Chengshi Zheng
INTERSPEECH3
2023 Hybrid TOA/AOA Indoor Positioning Based on Sparse Reconstruction and Map Matching
abstract
Indoor positioning technology, as a crucial foundation of location-based services, is experiencing a growing need for high precision driven by the Internet of Things (IoT). However, traditional positioning algorithms suffer from low sample utilization and susceptibility to noise. Moreover, the presence of indoor obstacles significantly affects positioning accuracy and leads to the issue of wall-penetrating positioning. To address these problems, this paper proposes a hybrid time-of-arrival/angle-of-arrival (TOA/AOA) indoor positioning algorithm based on sparse reconstruction and particle filtering-based map matching. Specifically, sparse reconstruction is employed to improve the utilization of samples, and iterative updating of the position estimation is performed during the multi-sample joint estimation process to enhance accuracy. Furthermore, to tackle the problem of wall-penetrating positioning, a particle filtering-based map matching algorithm is proposed to detect and eliminate the wall-penetrating particles using the map information matrix, which optimizes the positioning results obtained from sparse reconstruction. Simulation results demonstrate the effectiveness of the proposed algorithm in satisfying the demand for high-precision indoor positioning.
Chaoyang Du, Yang Liu 0063, Guochen Yu, Tianshuang Qiu
VTC Fall5
2023 A General Unfolding Speech Enhancement Method Motivated by Taylor's Theorem
abstract
While deep neural networks have facilitated significant advancements in the field of speech enhancement, most existing methods are developed following either empirical or relatively blind criteria, lacking adequate guidelines in pipeline design. Inspired by Taylor's theorem, we propose a general unfolding framework for both single- and multi-channel speech enhancement tasks. Concretely, we formulate the complex spectrum recovery into the spectral magnitude mapping in the neighborhood space of the noisy mixture, in which an unknown sparse term is introduced and applied for phase modification in advance. Based on that, the mapping function is decomposed into the superimposition of the 0th-order and high-order polynomials in Taylor's series, where the former coarsely removes the interference in the magnitude domain and the latter progressively complements the remaining spectral detail in the complex spectrum domain. In addition, we study the relation between adjacent order terms and reveal that each high-order term can be recursively estimated with its lower-order term, and each high-order term is then proposed to evaluate using a surrogate function with trainable weights, so that the whole system can be trained in an end-to-end manner. Given that the proposed framework is devised with the motivation of Taylor's theorem, it possesses improved internal flexibility. Extensive experiments are conducted on WSJ0-SI84, DNS-Challenge, Voicebank+Demand, spatialized Librispeech, and L3DAS22 multi-channel speech enhancement challenge datasets. Quantitative results show that the proposed approach yields competitive performance over existing top-performing approaches in terms of multiple objective metrics.
Andong Li, Guochen Yu, Chengshi Zheng, Xiaodong Li 0002
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Joint Magnitude Estimation and Phase Recovery Using Cycle-In-Cycle GAN for Non-Parallel Speech Enhancement
abstract
For the lack of adequate paired noisy-clean speech corpus in many real scenarios, non-parallel training is a promising task for DNN-based speech enhancement methods. However, because of the severe mismatch between input and target speeches, many previous studies only focus on the magnitude spectrum estimation and remain the phase unaltered, resulting in the degraded speech quality under low signal-to-noise ratio conditions. To tackle this problem, we decouple the difficult target w.r.t. original spectrum optimization into spectral magnitude and phase, and a novel Cycle-in-Cycle generative adversarial network (dubbed CinCGAN) is proposed to jointly estimate the spectral magnitude and phase information stage by stage under unpaired data. In the first stage, we pretrain a magnitude Cycle-GAN to coarsely estimate the spectral magnitude of clean speech. In the second stage, we incorporate the pretrained CycleGAN with a complex-valued CycleGAN as a cycle-in-cycle structure to simultaneously recover phase information and refine the overall spectrum. Experimental results demonstrate that the proposed approach significantly outperforms previous baselines under non-parallel training. The evaluation on training the models with standard paired data also shows that CinCGAN achieves remarkable performance especially in reducing background noise and speech distortion.
Guochen Yu, Andong Li, Yinuo Guo, Hui Wang 0070, Chengshi Zheng
ICASSP1
2022 Dual-Branch Attention-In-Attention Transformer for Single-Channel Speech Enhancement
abstract
Curriculum learning begins to thrive in the speech enhancement area, which decouples the original spectrum estimation task into multiple easier sub-tasks to achieve better performance. Motivated by that, we propose a dual-branch attention-in-attention transformer dubbed DB-AIAT to handle both coarse- and fine-grained regions of the spectrum in parallel. From a complementary perspective, a magnitude masking branch is proposed to coarsely estimate the overall magnitude spectrum, and simultaneously a complex refining branch is elaborately designed to compensate for the missing spectral details and implicitly derive phase information. Within each branch, we propose a novel attention-in-attention transformer-based module to replace the conventional RNNs and temporal convolutional networks for temporal sequence modeling. Specifically, the proposed attention-in-attention transformer consists of adaptive temporal-frequency attention transformer blocks and an adaptive hierarchical attention module, aiming to capture long-term temporal-frequency dependencies and further aggregate global hierarchical contextual information. Experimental results on Voice Bank + DEMAND demonstrate that DB-AIAT yields state-of-the-art performance (e.g., 3.31 PESQ, 95.6% STOI and 10.79dB SSNR) over previous advanced systems with a relatively small model size (2.81M).
Guochen Yu, Andong Li, Chengshi Zheng, Yinuo Guo, Hui Wang 0070
ICASSP1
2022 Taylor, Can You Hear Me Now? A Taylor-Unfolding Framework for Monaural Speech Enhancement
abstract
While the deep learning techniques promote the rapid development of the speech enhancement (SE) community, most schemes only pursue the performance in a black-box manner and lack adequate model interpretability. Inspired by Taylor's approximation theory, we propose an interpretable decoupling-style SE framework, which disentangles the complex spectrum recovery into two separate optimization problems i.e., magnitude and complex residual estimation. Specifically, serving as the 0th-order term in Taylor's series, a filter network is delicately devised to suppress the noise component only in the magnitude domain and obtain a coarse spectrum. To refine the phase distribution, we estimate the sparse complex residual, which is defined as the difference between target and coarse spectra, and measures the phase gap. In this study, we formulate the residual component as the combination of various high-order Taylor terms and propose a lightweight trainable module to replace the complicated derivative operator between adjacent terms. Finally, following Taylor's formula, we can reconstruct the target spectrum by the superimposition between 0th-order and high-order terms. Experimental results on two benchmark datasets show that our framework achieves state-of-the-art performance over previous competing baselines in various evaluation metrics. The source code is available at https://github.com/Andong-Li-speech/TaylorSENet.
Andong Li, Shan You, Guochen Yu, Chengshi Zheng, Xiaodong Li 0002
IJCAI3
2022 TMGAN-PLC: Audio Packet Loss Concealment using Temporal Memory Generative Adversarial Network
abstract
Real-time communications in packet-switched networks have become widely used in daily communication, while they inevitably suffer from network delays and data losses in constrained real-time conditions.To solve these problems, audio packet loss concealment (PLC) algorithms have been developed to mitigate voice transmission failures by reconstructing the lost information.Limited by the transmission latency and device memory, it is still intractable for PLC to accomplish high-quality voice reconstruction using a relatively small packet buffer.In this paper, we propose a temporal memory generative adversarial network for audio PLC, dubbed TMGAN-PLC, which is comprised of a novel nested-UNet generator and the time-domain/frequency-domain discriminators.Specifically, a combination of the nested-UNet and temporal featurewise linear modulation is elaborately devised in the generator to finely adjust the intra-frame information and establish inter-frame temporal dependencies.To complement the missing speech content caused by longer loss bursts, we employ multistage gated vector quantizers to capture the correct content and reconstruct the near-real smooth audio.Extensive experiments on the PLC Challenge dataset demonstrate that the proposed method yields promising performance in terms of speech quality, intelligibility, and PLCMOS.
Yuansheng Guan, Guochen Yu, Andong Li, Chengshi Zheng
INTERSPEECH2
2022 TaylorBeamformer: Learning All-Neural Beamformer for Multi-Channel Speech Enhancement from Taylor's Approximation Theory
abstract
While existing end-to-end beamformers achieve impressive performance in various front-end speech processing tasks, they usually encapsulate the whole process into a black box and thus lack adequate interpretability. As an attempt to fill the blank, we propose a novel neural beamformer inspired by Taylor's approximation theory called TaylorBeamformer for multi-channel speech enhancement. The core idea is that the recovery process can be formulated as the spatial filtering in the neighborhood of the input mixture. Based on that, we decompose it into the superimposition of the 0th-order non-derivative and high-order derivative terms, where the former serves as the spatial filter and the latter is viewed as the residual noise canceller to further improve the speech quality. To enable end-to-end training, we replace the derivative operations with trainable networks and thus can learn from training data. Extensive experiments are conducted on the synthesized dataset based on LibriSpeech and results show that the proposed approach performs favorably against the previous advanced baselines.
Andong Li, Guochen Yu, Chengshi Zheng, Xiaodong Li 0002
INTERSPEECH2
2022 Filtering and Refining: A Collaborative-Style Framework for Single-Channel Speech Enhancement
abstract
In low signal-to-noise ratio (SNR) acoustic scenarios, it remains fairly challenging to extract the target speech from its noisy mixture. In this paper, we propose a collaborative-style framework, namely, filtering and refining network (FRNet) for single-channel speech enhancement, recovering the complex spectrum of the target speech from coarse and fine-grained perspectives. Specifically, we devise a two-branch structure dubbed filtering-refining module (FRM). In the filtering block, the phase impact is ignored, and we only focus on coarse filtering in the magnitude domain. In the refining block, instead of predicting the irregular phase distribution directly, we estimate the complex residual for phase modification and spectrum rehabilitation, which takes the harmonic structure but with rather sparse energy distribution. By cascading FRMs repeatedly, we can reconstruct the target spectrum progressively. Furthermore, we propose a two-stream feature encoder to extract the feature representation of magnitude and phase individually, and the utilization of feature recalibration layers can preserve the prominent information from multiple scales. Extensive experiments are conducted on the WSJ0-SI84, Voicebank+Demand, and DNS-Challenge corpora. Evaluation results show that the proposed system performs favorably against previous advanced systems and achieves overall state-of-the-art performance in PESQ, ESTOI, SDR, and DNSMOS metrics.
Andong Li, Chengshi Zheng, Guochen Yu, Juanjuan Cai, Xiaodong Li 0002
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 DBT-Net: Dual-Branch Federative Magnitude and Phase Estimation With Attention-in-Attention Transformer for Monaural Speech Enhancement
abstract
The decoupling-style concept begins to ignite in the speech enhancement area, which decouples the original complex spectrum estimation task into multiple easier sub-tasks (i.e., the magnitude-only recovery and residual complex spectrum estimation), resulting in better performance and easier interpretability. In this paper, we propose a dual-branch federative magnitude and phase estimation framework, dubbed DBT-Net, for monaural speech enhancement, aiming at recovering the coarse- and fine-grained regions of the overall spectrum in parallel. From the complementary perspective, the magnitude estimation branch is designed to filter out dominant noise components in the magnitude domain, while the complex spectrum purification branch is elaborately designed to inpaint the missing spectral details and implicitly estimate the phase information in the complex-valued spectral domain. To facilitate the information flow between each branch, interaction modules are introduced to leverage features learned from one branch, so as to suppress the undesired parts and recover the missing components of the other branch. Instead of adopting the conventional RNNs and temporal convolutional networks for sequence modeling, we employ a novel attention-in-attention transformer-based network within each branch for better feature learning. More specially, it is composed of several adaptive spectro-temporal attention transformer-based modules and an adaptive hierarchical attention module, aiming to capture long-term time-frequency dependencies and further aggregate intermediate hierarchical contextual information. Comprehensive evaluations on the WSJ0-SI84 + DNS-Challenge and VoiceBank + DEMAND dataset demonstrate that the proposed approach consistently outperforms previous advanced systems and yields state-of-the-art performance in terms of speech quality and intelligibility.
Guochen Yu, Andong Li, Hui Wang 0070, Yuxuan Ke, Chengshi Zheng
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 A Simultaneous Denoising and Dereverberation Framework with Target Decoupling
abstract
Background noise and room reverberation are regarded as two major factors to degrade the subjective speech quality.In this paper, we propose an integrated framework to address simultaneous denoising and dereverberation under complicated scenario environments.It adopts a chain optimization strategy and designs four sub-stages accordingly.In the first two stages, we decouple the multi-task learning w.r.t.complex spectrum into magnitude and phase, and only implement noise and reverberation removal in the magnitude domain.Based on the estimated priors above, we further polish the spectrum in the third stage, where both magnitude and phase information are explicitly repaired with the residual learning.Due to the data mismatch and nonlinear effect of DNNs, the residual noise often exists in the DNN-processed spectrum.To resolve the problem, we adopt a light-weight algorithm as the post-processing module to capture and suppress the residual noise in the non-active regions.In the Interspeech 2021 Deep Noise Suppression (DNS) Challenge, our submitted system ranked top-1 for the real-time track in terms of Mean Opinion Score (MOS) with ITU-T P.835 framework.
Andong Li, Xiaoxue Luo, Guochen Yu, Chengshi Zheng, Xiaodong Li 0002
Interspeech4
2021 A two-stage complex network using cycle-consistent generative adversarial networks for speech enhancement
Guochen Yu, Hui Wang 0070, Qin Zhang 0009, Chengshi Zheng
Speech Commun.1