EDBT 2026 Demo / reviewers in the wild / expert
Wangjin Zhou
dblp:318/1431
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
0009-0007-0693-5316ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KyotoMOS2: MOS Prediction for Speech Across Multiple Sampling RatesabstractWe propose KyotoMOS2, an automatic MOS prediction system capable of evaluating speech across varying sampling rates. We design a sampling-rate-aware SSL-MOS subsystem and evaluate 32 variants on both original and resampled audio. Based on system-level SRCC and MSE, 13 subsystems are selected and fused for the final prediction. A system-level SRCC-based early stopping strategy is used to train the fusion model. Our system (T13) ranked 2nd on SRCC in the AudioMOS Track 3. Wangjin Zhou, Keisuke Imoto, Tatsuya Kawahara |
ASRU | 1 |
| 2025 | InvoxSVC: Any-to-any Zero-shot Singing Voice Conversion with In-Context Learning in Latent Flow MatchingabstractRecent advancements in singing voice conversion (SVC) have focused on achieving zero-shot, any-to-any voice transformation capabilities. Many approaches attempt to modify voice characteristics by incorporating global timbre variables into acoustic models. However, these methods often depend heavily on the capabilities of timbre extractors and lack an understanding of temporal local information. This limitation poses challenges, particularly in replicating specific voice qualities such as those of children. To address this issue, we introduce InvoxSVC, a latent flow matching model (LFM) designed for rapid and precise singing voice conversion with a particular emphasis on capturing temporal local features. While reducing the residual timbral information in the source singing encoding through singer-guidance, InvoxSVC enhances the model’s ability to capture temporal nuances by integrating in-context learning during inference. Additionally, the model employs a pre-trained high-fidelity variational autoencoder (VAE) to improve waveform generation. In comparative evaluations, InvoxSVC outperforms the open-source project So-VITS-SVC in both objective and subjective assessments. Wangjin Zhou, Tianjiao Du, Wenhao Guan, Chenglin Xu, Yi Zhao 0006, Tatsuya Kawahara |
ICME | 1 |
| 2025 | Simple and Effective Content Encoder for Singing Voice Conversion via SSL-Embedding Dimension Reduction
Wangjin Zhou, Tianjiao Du, Chenglin Xu, Sheng Li 0010, Yi Zhao 0006, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2024 | Enhancing Realism in 3D Facial Animation Using Conformer-Based Generation and Automated Post-ProcessingabstractRecent progress has propelled the development of realistic talking-face videos for avatars. Yet, animating 3D cartoon avatars remains intricate due to the imprecise nature of facial-driven data. This often manifests as inconsistent mouth configurations and rigid facial expressions, curbing the animation’s realism. Addressing these issues, we introduce a conformer-based framework that derives expression coefficients directly from phonemes, thereby elevating prediction precision and minimizing manual oversight. Furthermore, by harnessing a pre-trained emotion blending module coupled with the keyframe of the target emotional character, we employ a zero-shot adaptation technique. This serves to amplify emotional expressions and bolster the authenticity of lip dynamics. Our methodology adeptly registers nuanced expression shifts in avatars, leading to remarkably lifelike animations, as substantiated by our experimental findings. Yi Zhao 0006, Chunyu Qiang, Hao Li 0078, Yulan Hu, Wangjin Zhou, Sheng Li 0010 |
ICASSP | 5 |
| 2024 | MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score PredictionabstractIEEE Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD) as we expect that MOS can be used to assess how close synthesized speech is to the natural human voice. We propose MOS-FAD, where MOS can be leveraged at two key points in FAD: training data selection and model fusion. In training data selection, we demonstrate that MOS enables effective filtering of samples from unbalanced datasets. In the model fusion, our results demonstrate that incorporating MOS as a gating mechanism in FAD model fusion enhances overall performance. Wangjin Zhou, Zhengdong Yang, Chenhui Chu, Sheng Li 0010, Raj Dabre, Yi Zhao 0006, Tatsuya Kawahara |
ICASSP | 1 |
| 2024 | LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
Wenhao Guan, Wangjin Zhou, Feng Deng, Qingyang Hong |
INTERSPEECH | 3 |
| 2024 | Disentangling Age and Identity with a Mutual Information Minimization for Cross-Age Speaker Verification
Fengrun Zhang, Wangjin Zhou, Wang Geng, Yahui Shan |
INTERSPEECH | 2 |
| 2023 | LE-SSL-MOS: Self-Supervised Learning MOS Prediction with Listener EnhancementabstractRecently, researchers have shown an increasing interest in automatically predicting the subjective evaluation for speech synthesis systems. This prediction is a challenging task, especially on the out-of-domain test set. In this paper, we proposed a novel fusion model for MOS prediction that combines both supervised and unsupervised approaches. In the supervised aspect, we developed a SSL-based predictor called LE-SSL-MOS. The LE-SSL-MOS utilizes pre-trained self-supervised learning models and further improves prediction accuracy by utilizing the opinion scores of each utterance in the listener enhancement branch. In the unsupervised aspect, two steps are contained: one is that we fine-tuned unit language model (ULM) using highly-intelligible domain data to improve the correlation of an unsupervised metric SpeechLMScore. Another is that we utilized ASR confidence as a new metric with the help of ensemble learning. To the best of our knowledge, this is the first architecture that fuses supervised and unsupervised methods for MOS prediction.With these approaches, our experimental results on the VoiceMOS Challenge 2023 show that LE-SSL-MOS performs better than the baseline. Our fusion system achieved an absolute improvement of 13 % over LE-SSL-MOS on the noisy and enhanced speech track. And our system ranked 1st and 2 nd respectively in the French speech synthesis track and the noisy and enhanced speech track of the challenge. Zili Qi, Xinhui Hu, Wangjin Zhou, Sheng Li 0010, Xinkang Xu |
ASRU | 3 |
| 2022 | Investigating Effective Domain Adaptation Method for Speaker Verification Task
Guangxing Li, Wangjin Zhou, Sheng Li 0010, Yi Zhao 0006, Hao Huang 0009 |
ICONIP (6) | 2 |
| 2022 | Fusion of Self-supervised Learned Models for MOS PredictionabstractWe participated in the mean opinion score (MOS) prediction challenge, 2022.This challenge aims to predict MOS scores of synthetic speech on two tracks, the main track and a more challenging sub-track: out-of-domain (OOD).To improve the accuracy of the predicted scores, we have explored several model fusion-related strategies and proposed a fused framework in which seven pretrained self-supervised learned (SSL) models have been engaged.These pretrained SSL models are derived from three ASR frameworks, including Wav2Vec, Hubert, and WavLM.For the OOD track, we followed the 7 SSL models selected on the main track and adopted a semi-supervised learning method to exploit the unlabeled data.According to the official analysis results, our system has achieved 1 st rank in 6 out of 16 metrics and is one of the top 3 systems for 13 out of 16 metrics.Specifically, we have achieved the highest LCC, SRCC, and KTAU scores at the system level on main track, as well as the best performance on the LCC, SRCC, and KTAU evaluation metrics at the utterance level on OOD track.Compared with the basic SSL models, the prediction accuracy of the fused system has been largely improved, especially on OOD sub-track. Zhengdong Yang, Wangjin Zhou, Chenhui Chu, Sheng Li 0010, Raj Dabre, Raphaël Rubino, Yi Zhao 0006 |
INTERSPEECH | 2 |