Yewei Gu

dblp:285/1608 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 RLVC: Robust and Lightweight Voice Conversion Using Cross-Adaptive Instance Normalization
abstract
Voice conversion refers to transforming the speaker of a voice into a target speaker while keeping the content unchanged. Current solutions either rely on feature representations from large-scale pre-trained models or require complex model designs and intensive training, lacking exploration of intrinsic speech features. This restricts the exploration of lightweight and robust methods. In this study, we remove pre-trained models and depart from complex mutual information minimization for feature decoupling. Instead, we revisit decoupling methods based on instance normalization. To address it, we introduce a novel feature coupling module named cross-adaptive instance normalization (CAIN), which extends the adaptive instance normalization (AdaIN). Beyond offering style injection capabilities, CAIN is designed to maintain content consistency by reconstructing frame-level statistics in mel-spectrograms. The results indicate that CAIN, serving as a lightweight plugin, significantly improves conventional instance normalization-driven approaches. Building upon this, we propose RLVC, which achieves robust performance with a mere 5.29M parameters.
Yewei Gu, Xianfeng Zhao, Xiaowei Yi
ICME1
2023 Robust Feature Decoupling in Voice Conversion by Using Locality-Based Instance Normalization
Yewei Gu, Xianfeng Zhao, Xiaowei Yi
INTERSPEECH1
2022 Voice Conversion Using Learnable Similarity-Guided Masked Autoencoder
Yewei Gu, Xianfeng Zhao, Xiaowei Yi, Junchao Xiao
IWDW1
2021 FMFCC-A: A Challenging Mandarin Dataset for Synthetic Speech Detection
Yewei Gu, Xiaowei Yi, Xianfeng Zhao
IWDW2
2020 Deepfake Video Detection Using Audio-Visual Consistency
Yewei Gu, Xianfeng Zhao, Xiaowei Yi
IWDW1