Wenhao Guan

dblp:220/0759 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
17since 2021 · last 2025
0009-0002-0015-2250ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Dynamic Language Group-based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical Routing
abstract
The Mixture of Experts (MoE) model is a promising approach for handling code-switching speech recognition (CS-ASR) tasks. However, the existing CS-ASR work on MoE has yet to leverage the advantages of MoE’s parameter scaling ability fully. This work proposes DLG-MoE, a Dynamic Language Group-based MoE, which can effectively handle the CS-ASR task and leverage the advantages of parameter scaling. DLG-MoE operates based on a hierarchical routing mechanism. First, the language router explicitly models the language attribute and dispatches the representations to the corresponding language expert groups. Subsequently, the unsupervised router within each language group implicitly models attributes beyond language and coordinates expert routing and collaboration. DLG-MoE outperforms the existing MoE methods on CS-ASR tasks while demonstrating great flexibility. It supports different top-k inference and streaming capabilities and can also prune the model parameters flexibly to obtain a monolingual sub-model.
Hukai Huang, Shenghui Lu, Yahui Shan, He Qu, Fengrun Zhang, Wenhao Guan, Qingyang Hong
ICASSP6
2025 SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
abstract
Recently, flow matching based speech synthesis has significantly enhanced the quality of synthesized speech while reducing the number of inference steps. In this paper, we introduce SlimSpeech, a lightweight and efficient speech synthesis system based on rectified flow. We have built upon the existing speech synthesis method utilizing the rectified flow model, modifying its structure to reduce parameters and serve as a teacher model. By refining the reflow operation, we directly derive a smaller model with a more straight sampling trajectory from the larger model, while utilizing distillation techniques to further enhance the model performance. Experimental results demonstrate that our proposed method, with significantly reduced model parameters, achieves comparable performance to larger models through one-step sampling.
Kaidi Wang 0001, Wenhao Guan, Shenghui Lu, Jianglong Yao, Qingyang Hong
ICASSP2
2025 InvoxSVC: Any-to-any Zero-shot Singing Voice Conversion with In-Context Learning in Latent Flow Matching
abstract
Recent advancements in singing voice conversion (SVC) have focused on achieving zero-shot, any-to-any voice transformation capabilities. Many approaches attempt to modify voice characteristics by incorporating global timbre variables into acoustic models. However, these methods often depend heavily on the capabilities of timbre extractors and lack an understanding of temporal local information. This limitation poses challenges, particularly in replicating specific voice qualities such as those of children. To address this issue, we introduce InvoxSVC, a latent flow matching model (LFM) designed for rapid and precise singing voice conversion with a particular emphasis on capturing temporal local features. While reducing the residual timbral information in the source singing encoding through singer-guidance, InvoxSVC enhances the model’s ability to capture temporal nuances by integrating in-context learning during inference. Additionally, the model employs a pre-trained high-fidelity variational autoencoder (VAE) to improve waveform generation. In comparative evaluations, InvoxSVC outperforms the open-source project So-VITS-SVC in both objective and subjective assessments.
Wangjin Zhou, Tianjiao Du, Wenhao Guan, Chenglin Xu, Yi Zhao 0006, Tatsuya Kawahara
ICME3
2025 DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
Peijie Chen, Wenhao Guan, Weijie Wu, Hukai Huang, Qingyang Hong
INTERSPEECH2
2025 ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
Wenhao Guan, Peijie Chen, Qingyang Hong
INTERSPEECH2
2025 Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
Kaidi Wang 0001, Wenhao Guan, Ziyue Jiang 0001, Hukai Huang, Peijie Chen, Weijie Wu, Qingyang Hong
INTERSPEECH2
2024 MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis
abstract
The style transfer task in Text-to-Speech (TTS) refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However, most existing style transfer approaches are either based on fixed emotional labels or reference speech clips, which cannot achieve flexible style transfer. Recently, some methods have adopted text descriptions to guide style transfer. In this paper, we propose a more flexible multi-modal and style controllable TTS framework named MM-TTS. It can utilize any modality as the prompt in unified multi-modal prompt space, including reference speech, emotional facial images, and text descriptions, to control the style of the generated speech in a system. The challenges of modeling such a multi-modal style controllable TTS mainly lie in two aspects: 1) aligning the multi-modal information into a unified style space to enable the input of arbitrary modality as the style prompt in a single system, and 2) efficiently transferring the unified style representation into the given text content, thereby empowering the ability to generate prompt style-related voice. To address these problems, we propose an aligned multi-modal prompt encoder that embeds different modalities into a unified style space, supporting style transfer for different modalities. Additionally, we present a new adaptive style transfer method named Style Adaptive Convolutions (SAConv) to achieve a better style representation. Furthermore, we design a Rectified Flow based Refiner to solve the problem of over-smoothing Mel-spectrogram and generate audio of higher fidelity. Since there is no public dataset for multi-modal TTS, we construct a dataset named MEAD-TTS, which is related to the field of expressive talking head. Our experiments on the MEAD-TTS dataset and out-of-domain datasets demonstrate that MM-TTS can achieve satisfactory results based on multi-modal prompts. The audio samples and constructed dataset are available at https://multimodal-tts.github.io.
Wenhao Guan, Yishuang Li, Hukai Huang, Jiayan Lin, Lingyan Huang, Qingyang Hong
AAAI1
2024 Reflow-TTS: A Rectified Flow Model for High-Fidelity Text-to-Speech
abstract
The diffusion models including Denoising Diffusion Probabilistic Models (DDPM) and score-based generative models have demonstrated excellent performance in speech synthesis tasks. However, its effectiveness comes at the cost of numerous sampling steps, resulting in prolonged sampling time required to synthesize high-quality speech. This drawback hinders its practical applicability in real-world scenarios. In this paper, we introduce ReFlow-TTS, a novel rectified flow based method for speech synthesis with high-fidelity. Specifically, our ReFlow-TTS is simply an Ordinary Differential Equation (ODE) model that transports Gaussian distribution to the ground-truth Mel-spectrogram distribution by straight line paths as much as possible. Furthermore, our proposed approach enables high-quality speech synthesis with a single sampling step and eliminates the need for training a teacher model. Our experiments on LJSpeech Dataset show that our ReFlow-TTS method achieves the best performance compared with other diffusion based models. And the ReFlow-TTS with one step sampling achieves competitive performance compared with existing one-step TTS models.
Wenhao Guan, Haodong Zhou, Shiyu Miao, Xingjia Xie, Qingyang Hong
ICASSP1
2024 Multivariate Fourier Distribution Perturbation: Domain Shifts with Uncertainty in Frequency Domain
abstract
Diversifying training data techniques have achieved tremendous success in Domain Generalization (DG) tasks. The key to diversifying domain data is by increasing the types of domain styles. After investigating this issue from the perspective of the Fourier transform, the domain cue is found to be implicitly encoded in the amplitude component of Fourier features, which is more indicative of domain-specific information than statistics (means and standard deviations). However, Fourier-based methods tend to augment amplitude components via linear interpolation between two samples, which limits the diversity. To break this limitation, we aim to augment novel amplitude components from a perturbation perspective, which is termed Multivariate Fourier Distribution Perturbation. Specially, we design channel-wise and pixel-wise random perturbations for in-sample and cross-sample distribution to expand the distribution scope of probabilistic feature amplitude components.
Weijie Chen 0006, Shicai Yang, Yishuang Li, Wenhao Guan
ICASSP5
2024 SR-HuBERT : An Efficient Pre-Trained Model for Speaker Verification
abstract
Recently, pre-trained models (PTMs) have been extensively applied in speaker verification (SV) and greatly boosted system performance. However, mainstream PTMs currently concentrate on using frame-level universal representations. In this paper, we propose a novel pre-training framework that jointly models speaker information — Speaker Related HuBERT, abbreviated as SR-HuBERT. This framework aims to further explore speaker-related information inherent in speech universal representations. The proposed SR-HuBERT utilizes an unsupervised clustering algorithm based on graph structures to generate speaker pseudo-labels and promotes the learning of segment-level speaker-related representations through a multi-task pre-training framework. Experimental results on VoxCeleb1 test set demonstrate the effectiveness of the proposed SR-HuBERT. Even in the scenarios of limited fine-tuning data, SR-HuBERT outperforms the other existing PTMs on SV tasks. Additionally, SR-HuBERT also performs well on speaker-related tasks of SUPERB benchmark.
Yishuang Li, Hukai Huang, Zhicong Chen, Wenhao Guan, Jiayan Lin, Qingyang Hong
ICASSP4
2024 Improving Multi-Speaker ASR With Overlap-Aware Encoding And Monotonic Attention
abstract
End-to-end (E2E) multi-speaker speech recognition with the serialized output training (SOT) strategy demonstrates good performance in modeling diverse speaker scenarios. However, the E2E architecture doesn’t explicitly address the modeling of overlapping speech areas, potentially limiting the model’s ability to generalize. To tackle this issue, we introduce two approaches: overlap-aware encoding method and monotonic attention loss. The former enables the model to acquire knowledge about overlapping speech through multitask learning, while the latter encourages the model to learn specific attention patterns associated with overlap by constraining the attention of adjacent text time steps. Our experimental results on the AliMeeting dataset show that the combination of these two methods effectively enhances the model’s performance.
Wenhao Guan, Lingyan Huang, Qingyang Hong
ICASSP3
2024 FastOcc: Accelerating 3D Occupancy Prediction by Fusing the 2D Bird's-Eye View and Perspective View
abstract
In autonomous driving, 3D occupancy prediction outputs voxel-wise status and semantic labels for more comprehensive understandings of 3D scenes compared with traditional perception tasks, such as 3D object detection and bird’s-eye view (BEV) semantic segmentation. Recent researchers have extensively explored various aspects of this task, including view transformation techniques, ground-truth label generation, and elaborate network design, aiming to achieve superior performance. However, the inference speed, crucial for running on an autonomous vehicle, is neglected. To this end, a new method, dubbed FastOcc, is proposed. By carefully analyzing the network effect and latency from four parts, including the input image resolution, image backbone, view transformation, and occupancy prediction head, it is found that the occupancy prediction head holds considerable potential for accelerating the model while keeping its accuracy. Targeted at improving this component, the time-consuming 3D convolution network is replaced with a novel residual-like architecture, where features are mainly digested by a lightweight 2D BEV convolution network and compensated by integrating the 3D voxel features interpolated from the original image features. Experiments on the Occ3D-nuScenes benchmark demonstrate that our FastOcc achieves state-of-the-art results with a fast inference speed.
Wenhao Guan, Di Feng, Yuheng Du, Xiangyang Xue 0001, Jian Pu
ICRA3
2024 LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
Wenhao Guan, Wangjin Zhou, Feng Deng, Qingyang Hong
INTERSPEECH1
2024 Efficient Integrated Features Based on Pre-trained Models for Speaker Verification
Yishuang Li, Wenhao Guan, Hukai Huang, Shiyu Miao, Qingyang Hong
INTERSPEECH2
2024 MinSpeech: A Corpus of Southern Min Dialect for Automatic Speech Recognition
Jiayan Lin, Shenghui Lu, Hukai Huang, Wenhao Guan, Hui Bu, Qingyang Hong
INTERSPEECH4
2024 Enhancing Code-Switching Speech Recognition With LID-Based Collaborative Mixture of Experts Model
abstract
Due to the inherent difficulty in modeling phonetic similarities across different languages, code-switching speech recognition presents a formidable challenge. This study proposes a Collaborative-MoE, a Mixture of Experts (MoE) model that leverages a collaborative mechanism among expert groups. Initially, a preceding routing network explicitly learns Language Identification (LID) tasks and selects experts based on acquired LID weights. This process ensures robust routing information to the MoE layer, mitigating interference from diverse language domains on expert network parameter updates. The LID weights are also employed to facilitate inter-group collaboration, enabling the integration of language-specific representations. Furthermore, within each language expert group, a gating network operates unsupervised to foster collaboration on attributes beyond language. Extensive experiments demonstrate the efficacy of our approach, achieving significant performance enhancements compared to alternative methods. Importantly, our method preserves the efficient inference capabilities characteristic of MoE models without necessitating additional pre-training.
Hukai Huang, Jiayan Lin, Yishuang Li, Wenhao Guan, Qingyang Hong
SLT5
2023 Interpretable Style Transfer for Text-to-Speech with ControlVAE and Diffusion Bridge
Wenhao Guan, Yishuang Li, Hukai Huang, Qingyang Hong
INTERSPEECH1