VLDB 2026 Research / reviewers in the wild / expert
Yuxun Tang
dblp:358/9271
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2025
0009-0002-4538-7440ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VERSA-v2: A Modular and Scalable Toolkit for Speech and Audio Evaluation with Expanded Metrics, Visualization, and LLM IntegrationabstractWe present VERSA-v2, a major upgrade of the Versatile Evaluation of Speech and Audio (VERSA) toolkit for standardized and scalable evaluation across speech, audio, and music tasks. It features a modular, object-oriented architecture that simplifies metric integration and now supports over 100 metrics, organized into curated task-specific packs. VERSA-v2 also introduces interactive visualizations, per-metric profiling, and prompt-based evaluation using both text- and audio-based large language models (LLMs). These advancements make VERSA-v2 a robust, extensible, and LLM-enabled platform for comprehensive and interpretable speech and audio evaluation. Jiatong Shi, Bo-Hao Su, Shikhar Bharadwaj, Shih-Heng Wang, Jionghao Hang, Wei Wang 0010, Wenhao Feng, Yuxun Tang, Nezih Topaloglu, Siddhant Arora, Jinchuan Tian, Hye-Jin Shim, Wangyou Zhang, Wen-Chin Huang, Shinji Watanabe 0001 |
ASRU | 10 |
| 2025 | Robust Training of Singing Voice Synthesis Using Prior and Posterior UncertaintyabstractSinging voice synthesis (SVS) has seen remarkable advancements in recent years. However, compared to speech and general audio data, publicly available singing datasets remain limited. In practice, this data scarcity often leads to performance degradation in long-tail scenarios, such as imbalanced pitch distributions or rare singing styles. To mitigate these challenges, we propose uncertainty-based optimization to improve the training process of end-to-end SVS models. First, we introduce differentiable data augmentation in the adversarial training, which operates in a sample-wise manner to increase the prior uncertainty. Second, we incorporate a frame-level uncertainty prediction module that estimates the posterior uncertainty, enabling the model to allocate more learning capacity to low-confidence segments. Empirical results on the Opencpop and Ofuton-P, across Chinese and Japanese, demonstrate that our approach improves performance in various perspectives. Jiatong Shi, Yuxun Tang, Shinji Watanabe 0001 |
ASRU | 3 |
| 2024 | The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
Xuankai Chang, Jiatong Shi, Jinchuan Tian, Yuning Wu 0001, Yuxun Tang, Yihan Wu 0008, Shinji Watanabe 0001, Yossi Adi, Xie Chen 0001, Qin Jin |
INTERSPEECH | 5 |
| 2024 | Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing
Jiatong Shi, Yueqian Lin, Xinyi Bai, Keyi Zhang, Yuning Wu 0001, Yuxun Tang, Qin Jin, Shinji Watanabe 0001 |
INTERSPEECH | 6 |
| 2024 | SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models
Yuxun Tang, Yuning Wu 0001, Jiatong Shi, Qin Jin |
INTERSPEECH | 1 |
| 2024 | TokSing: Singing Voice Synthesis based on Discrete Tokens
Yuning Wu 0001, Jiatong Shi, Yuxun Tang, Qin Jin |
INTERSPEECH | 4 |
| 2024 | CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
Yongyi Zang, Jiatong Shi, You Zhang 0001, Ryuichi Yamamoto, Jionghao Han, Yuxun Tang, Wen-Xiao Zhao, Tomoki Toda, Zhiyao Duan |
INTERSPEECH | 6 |
| 2024 | Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New ParadigmabstractThis research presents Muskits-ESPnet, a versatile toolkit that introduces new paradigms to Singing Voice Synthesis (SVS) through the application of pretrained audio models in both continuous and discrete approaches. Specifically, we explore discrete representations derived from SSL models and audio codecs and offer significant advantages in versatility and intelligence, supporting multi-format inputs and adaptable data processing workflows for various SVS models. The toolkit features automatic music score error detection and correction, as well as a perception auto-evaluation module to imitate human subjective evaluating scores. Muskits-ESPnet is available at https://github.com/espnet/espnet. Yuning Wu 0001, Jiatong Shi, Yuxun Tang, Yueqian Lin, Jionghao Han, Xinyi Bai, Shinji Watanabe 0001, Qin Jin |
ACM Multimedia | 4 |
| 2024 | ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs For Audio, Music, and SpeechabstractNeural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance downstream training efficiency and compatibility with autoregressive language models. However, as extensive downstream applications are investigated, challenges have arisen in ensuring fair comparisons across diverse applications. To address these issues, we present a new open-source platform ESPnet-Codec, which is built on ESPnet and focuses on neural codec training and evaluation. ESPnet-Codec offers various recipes in audio, music, and speech for training and evaluation using several widely adopted codec models. Together with ESPnet-Codec, we present VERSA, a standalone evaluation toolkit, which provides a comprehensive evaluation of codec performance over 20 audio evaluation metrics. Notably, we demonstrate that ESPnet-Codec can be integrated into six ESPnet tasks, supporting diverse applications. Jiatong Shi, Jinchuan Tian, Yihan Wu 0008, Jee-Weon Jung, Jia Qi Yip, Yoshiki Masuyama, Yuning Wu 0001, Yuxun Tang, Massa Baali, Dareen Alharthi, Ruifan Deng, Tejes Srivastava, Alexander H. Liu, Bhiksha Raj, Qin Jin, Ruihua Song, Shinji Watanabe 0001 |
SLT | 9 |
| 2024 | Visinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning RepresentationabstractSinging Voice Synthesis (SVS) has witnessed significant advancements with the advent of deep learning techniques. However, a significant challenge in SVS is the scarcity of labeled singing voice data, which limits the effectiveness of supervised learning methods. In response to this challenge, this paper introduces a novel approach to enhance the quality of SVS by leveraging unlabeled data from pre-trained self-supervised learning models. Building upon the existing VISinger2 framework, this study integrates additional spectral feature information into the system to enhance its performance. The integration aims to harness the rich acoustic features from the pre-trained models, thereby enriching the synthesis and yielding a more natural and expressive singing voice. Experimental results in various corpora demonstrate the efficacy of this approach in improving the overall quality of synthesized singing voices in both objective and subjective metrics. Jiatong Shi, Yuning Wu 0001, Yuxun Tang, Shinji Watanabe 0001 |
SLT | 4 |
| 2023 | Findings of the 2023 ML-Superb Challenge: Pre-Training And Evaluation Over More Languages And BeyondabstractThe 2023 Multilingual Speech Universal Performance Benchmark (ML-SUPERB) Challenge expands upon the acclaimed SUPERB framework, emphasizing self-supervised models in multilingual speech recognition and language identification. The challenge comprises a research track focused on applying ML-SUPERB to specific multilingual subjects, a Challenge Track for model submissions, and a New Language Track where language resource researchers can contribute and evaluate their low-resource language data in the context of the latest progress in multilingual speech recognition. The challenge garnered 12 model submissions and 54 language corpora, resulting in a comprehensive benchmark encompassing 154 languages. The findings indicate that merely scaling models is not the definitive solution for multilingual speech tasks, and a variety of speech/voice types present significant challenges in multilingual speech processing. Jiatong Shi, Dan Berrebbi, Hsiu-Hsuan Wang, Wei-Ping Huang, En-Pei Hu, Ho-Lam Chuang, Xuankai Chang, Yuxun Tang, Shang-Wen Li 0001, Abdel-rahman Mohamed, Hung-yi Lee, Shinji Watanabe 0001 |
ASRU | 9 |