EDBT 2026 Demo / reviewers in the wild / expert
Yuehai Wang
dblp:60/4052
· DBLP profile ↗
18ranked-venue papers
1as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 since 2021Artificial intelligence and machine learning · 11 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language ModelsabstractText-to-speech (TTS) has seen significant advancements in high-quality, expressive speech synthesis. However, achieving diverse and natural prosody in synthesized speech remains challenging. In this paper, we propose ProsodyFlow, an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively. Our approach involves using a speech LLM to extract acoustic features, mapping these features into a prosody latent space, and then employing conditional flow matching to generate prosodic vectors conditioned on the input text. Experiments on the LJSpeech dataset show that ProsodyFlow improves synthesis quality and efficiency compared to existing models, achieving more prosodic and expressive speech synthesizing. Sizhe Shan, Yinlin Guo, Yuehai Wang |
COLING | 4 |
| 2025 | Do Less and Achieve More: Free Condition Video Outpainting with Diffusion ModelabstractVideo outpainting aims to extend the content of a video beyond its original spatial boundaries. Existing methods tend to condition the generation process on a single frame or caption, failing to address the challenge in long videos with multiple video clips. To address this, we extend the diffusion-based image inpainting to a 3D diffusion model with a motion module for video outpainting. Then, we design four conditioning schemes to seek the appropriate condition strategy: Frame-wise context, Frame-wise Adapter, Text, and Free Condition. Additionally, to enhance the continuity in long video generation, we propose a novel Motion Momentum Update (MMU) method based on a training-free technique. Experimental results demonstrate that our proposed Free Condition strategy, termed Free-Outpainter, achieves state-of-the-art performance in video outpainting tasks. Ablation studies also show that the condition-free training paradigm is the most effective in avoiding inaccurate external information and fully utilizing inherent video data. Benefiting from the Free Condition strategy, our pipeline reduces the inference time to 5 seconds for 512x512x16 frames with only 10 GB of GPU memory. More results can be found on our demo page https://lilyn3125.github.io/FreeOutpainting/. Haofan Huang, Yinlin Guo, Yening Lv, Sizhe Shan, Yuehai Wang |
ICASSP | 6 |
| 2024 | Audio Deepfake Detection With Self-Supervised Wavlm And Multi-Fusion Attentive ClassifierabstractWith the rapid development of speech synthesis and voice conversion technologies, Audio Deepfake has become a serious threat to the Automatic Speaker Verification (ASV) system. Numerous countermeasures are proposed to detect this type of attack. In this paper, we report our efforts to combine the self-supervised WavLM model and Multi-Fusion Attentive classifier for audio deepfake detection. Our method exploits the WavLM model to extract features that are more conducive to spoofing detection for the first time. Then, we propose a novel Multi-Fusion Attentive (MFA) classifier based on the Attentive Statistics Pooling (ASP) layer. The MFA captures the complementary information of audio features at both time and layer levels. Experiments demonstrate that our methods achieve state-of-the-art results on the ASVspoof 2021 DF set and provide competitive results on the ASVspoof 2019 and 2021 LA set. Yinlin Guo, Haofan Huang, Xi Chen 0025, Yuehai Wang |
ICASSP | 5 |
| 2024 | Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex RecordingsabstractTarget speaker extraction (TSE) aims to extract the target speaker’s voice from the input mixture. Previous studies have concentrated on high-overlapping scenarios. However, real-world applications usually meet more complex scenarios like variable speaker overlapping and target speaker absence. In this paper, we introduces a framework to perform continuous TSE (C-TSE), comprising a target speaker voice activation detection (TSVAD) and a TSE model. This framework significantly improves TSE performance on similar speakers and enhances personalization, which is lacking in traditional diarization methods. In detail, unlike conventional TSVAD deployed to refine the diarization results, the proposed Attention-target speaker voice activation detection (A-TSVAD) to directly generate timestamps of the target speaker. We also explore some different integration methods of A-TSVAD and TSE by comparing the cascaded and parallel methods. The framework’s effectiveness is assessed using a range of metrics, including diarization and enhancement metrics. Our experiments demonstrate that A-TSVAD outperforms conventional methods in reducing diarization errors, when integrating A-TSVAD and TSE in a sequential cascaded manner further enhances extraction accuracy. Audio demos are available on our demo page1. Hangting Chen, Jianwei Yu 0001, Yuehai Wang |
IJCNN | 4 |
| 2024 | FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
Yinlin Guo, Yening Lv, Jinqiao Dou, Yuehai Wang |
INTERSPEECH | 5 |
| 2023 | Specialty may be better: A decoupling multi-modal fusion network for Audio-visual event localizationabstractAudio and visual signals usually coexist in realistic scenes, and human brains can learn this multi-modal perception easily. So, it is crucial for the computer to learn how human brains work for solving multi-modal tasks. The Audio-visual event localization (AVEL) task involves two sub-tasks: find video segments that contain Audio-visual events, and determine the category of the events. However, the AVEL task remains challenging due to the severe background noise. Additionally, processing information from both modalities simultaneously is also a tough issue. The current approaches have two main problems. One is that the network tends to be influenced by noise and predicts unreasonable events for consecutive segments within the same video clip. The other is that the model will oscillate between the local and global targets due to the multi-objective learning. To address these problems, we propose a decoupling multi-modal fusion network, which not only suppresses the complex noise but also learns the local and global information exclusively. The proposal consists of two sub-networks: the Which-event sub-network for predicting the event category and the Is-event sub-network for determining the time boundary for the event. We evaluate our method on the standard AVE Dataset in both fully and weakly supervised settings, and the results verify the effectiveness of our method. Jinqiao Dou, Xi Chen 0025, Yuehai Wang |
IJCNN | 3 |
| 2022 | Masked Acoustic Unit for Mispronunciation Detection and CorrectionabstractComputer-Assisted Pronunciation Training (CAPT) plays an important role in language learning. Conventional ASR-based CAPT methods require expensive annotation of the ground truth pronunciation for the supervised training. Mean-while, certain undefined non-native phonemes cannot be correctly classified into standard phonemes, making the annotation process challenging and subjective. On the other hand, ASR-based CAPT methods only give the learner text-based feedback about the mispronunciation, but cannot teach the learner how to pronounce the sentence correctly. To solve these limitations, we propose to use the acoustic unit (AU) as the intermediary feature for both mispronunciation detection and correction. The proposed method uses the masked AU sequence and the target phonemes to detect the error AU and then corrects it. This method can give the learner speech-based self-imitating feedback, making our CAPT powerful for education. Zhan Zhang 0010, Yuehai Wang, Jianyi Yang 0003 |
ICASSP | 2 |
| 2022 | End-to-end Mispronunciation Detection with Simulated Error Distance
Zhan Zhang 0010, Yuehai Wang, Jianyi Yang 0003 |
INTERSPEECH | 2 |
| 2022 | BiCAPT: Bidirectional Computer-Assisted Pronunciation Training with Normalizing Flows
Zhan Zhang 0010, Yuehai Wang, Jianyi Yang 0003 |
INTERSPEECH | 2 |
| 2022 | Improve Speech Enhancement using Perception-High-Related Time-Frequency Loss
Ding Zhao, Yuehai Wang |
INTERSPEECH | 4 |
| 2021 | Text-conditioned Transformer for automatic pronunciation error detection
Zhan Zhang 0010, Yuehai Wang, Jianyi Yang 0003 |
Speech Commun. | 2 |
| 2020 | Collaborative Distillation for Ultra-Resolution Universal Style TransferabstractUniversal style transfer methods typically leverage rich representations from deep Convolutional Neural Network (CNN) models (e.g., VGG-19) pre-trained on large collections of images. Despite the effectiveness, its application is heavily constrained by the large model size to handle ultra-resolution images given limited memory. In this work, we present a new knowledge distillation method (named Collaborative Distillation) for encoder-decoder based neural style transfer to reduce the convolutional filters. The main idea is underpinned by a finding that the encoder-decoder pairs construct an exclusive collaborative relationship, which is regarded as a new kind of knowledge for style transfer models. Moreover, to overcome the feature size mismatch when applying collaborative distillation, a linear embedding loss is introduced to drive the student network to learn a linear embedding of the teacher’s features. Extensive experiments show the effectiveness of our method when applied to different universal style transfer approaches (WCT and AdaIN), even if the model size is reduced by 15.5 times. Especially, on WCT with the compressed models, we achieve ultra-resolution (over 40 megapixels) universal style transfer on a 12GB GPU for the first time. Further experiments on optimization-based stylization scheme show the generality of our algorithm on different stylization paradigms. Our code and trained models are available at https://github.com/mingsun-tse/collaborative-distillation. Huan Wang 0014, Yijun Li 0001, Yuehai Wang, Haoji Hu, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2020 | Deep quantised portrait mattingabstractPortrait matting is of vital importance for many applications such as portrait editing, background replacement, ecommerce demonstration, and augmented reality. The portrait matt can be accessed by predicting the α value of the original picture. Previous deep matting methods usually adopt a segmentation network to tackle portrait matting tasks. However, these traditional methods will introduce unpleasant blemishes in the matting results sometimes. The authors find that the key factor behind this phenomenon is how they model the matting problem. On the one hand, α value predicting can be modelled as a regression task. On the other hand, it can be viewed as a classification task of predicting background or foreground. To solve this problem, they explore different methods to model the nature of the α matting problem and propose a novel quantisation‐based adaption. Their method comes up with an α quantisation loss to achieve multi‐threshold filtering. Furthermore, they apply an α merging block to improve conventional regression methods. With their method, the gradient loss is reduced by 7.53% relatively, with mean square error and sum of absolute difference decreased by 14.7% relatively, leading to a more visually pleasant α matt in several segmentation backbones. Zhan Zhang 0010, Yuehai Wang, Jianyi Yang 0003 |
IET Comput. Vis. | 2 |
| 2019 | Structured Pruning for Efficient ConvNets via Incremental RegularizationabstractParameter pruning is a promising approach for CNN compression and acceleration by eliminating redundant model parameters with tolerable performance degrade. Despite its effectiveness, existing regularization-based parameter pruning methods usually drive weights towards zero with large and constant regularization factors, which neglects the fragility of the expressiveness of CNNs, and thus calls for a more gentle regularization scheme so that the networks can adapt during pruning. To achieve this, we propose a new and novel regularization-based pruning method, named IncReg, to incrementally assign different regularization factors to different weights based on their relative importance. Empirical analysis on CIFAR-10 dataset verifies the merits of IncReg. Further extensive experiments with popular CNNs on CIFAR-10 and ImageNet datasets show that IncReg achieves comparable to even better results compared with state-of-the-arts. Our source codes and trained models are available here: https://github.com/mingsun-tse/caffe_increg. Huan Wang 0014, Qiming Zhang 0001, Yuehai Wang, Lu Yu 0003, Haoji Hu |
IJCNN | 3 |
| 2018 | Structured Probabilistic Pruning for Convolutional Neural Network Acceleration
Huan Wang 0014, Qiming Zhang 0001, Yuehai Wang, Haoji Hu |
BMVC | 3 |
| 2017 | Robot as a Service in Information Science & Electronic Engineering EducationabstractThis paper reports the newly designed programming course developed and taught at Zhejiang University for a freshman class majoring in information science and electronic engineering, and disseminated to many other universities, to address the enrollment crisis as reported by ACM Curriculum Committee Review Task Force. The course teaches the basic information science and electronic engineering concepts and gives the students practical programming experience through robotics programming. The initial curriculum and the ongoing improvement of the course are presented. The course started with its experiment environment using ASU Visual IoT/Robotics Programming Environment (VIPLE). The environment supports a variety of platforms, including open architecture-based robots, such as Intel robots, PCDuino robots, vendor-specific robots, such as Lego EV3, as well as simulated robots, such as Web simulated robots in HTML5 and Unity-based 3D robots. Yuehai Wang, Yinong Chen 0004, Xiaofan Tong, Yubo Lee, Jianyi Yang 0003 |
ISADS | 1 |
| 2016 | Multiple scattering effects on the localization of two point scatterersabstractMultiple scattering effects are commonly ignored in the detection and estimation of scatterers in signal processing research, because the energy of the first-order scattering is much larger than that of higher-order components. Although multiple scattering can significantly increase the estimation precision of point scatterers, it does not always lead to an improvement. Identifying conditions under which multiple scattering is beneficial or detrimental to estimation in a general setup is still an open problem. In this paper, we consider the effects of multiple scattering on the localization of two point scatterers. By comparing the Fisher information matrix on location parameters when multiple scattering exists and does not exist, we show analytically that information on ranges can benefit estimating directions of arrival via multiple scattering when the two scatterers are in far-field and well resolved. Arye Nehorai, Hongwei Liu 0001, Bo Chen 0001, Yuehai Wang |
ICASSP | 5 |
| 2014 | A threshold-adaptive film mode detection method in video de-interlacing
Tao Nie, Zhipao Tu, Xiaohong Chen 0001, Yuehai Wang |
Multim. Tools Appl. | 5 |