VLDB 2026 Research / reviewers in the wild / expert
Yiquan Zhou
dblp:358/7001
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0002-1453-3820ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechabstractExisting autoregressive large-scale text-to-speech (TTS) models have advantages in speech naturalness, but their token-by-token generation mechanism makes it difficult to precisely control the duration of synthesized speech. This becomes a significant limitation in applications requiring strict audio-visual synchronization, such as video dubbing. This paper introduces IndexTTS2, which proposes a novel, general, and autoregressive model-friendly method for speech duration control. The method supports two generation modes: one explicitly specifies the number of generated tokens to precisely control speech duration; the other freely generates speech in an autoregressive manner without specifying the number of tokens, while faithfully reproducing the prosodic features of the input prompt. Furthermore, IndexTTS2 achieves disentanglement between emotional expression and speaker identity, enabling independent control over timbre and emotion. In the zero-shot setting, the model can accurately reconstruct the target timbre (from the timbre prompt) while perfectly reproducing the specified emotional tone (from the style prompt). To enhance speech clarity in highly emotional expressions, we incorporate GPT latent representations and design a novel three-stage training paradigm to improve the stability of the generated speech. Additionally, to lower the barrier for emotional control, we designed a soft instruction mechanism based on text descriptions by fine-tuning Qwen3, effectively guiding the generation of speech with the desired emotional orientation. Finally, experimental results on multiple datasets show that IndexTTS2 outperforms state-of-the-art zero-shot TTS models in terms of word error rate, speaker similarity, and emotional fidelity. Yiquan Zhou, Jinchao Wang, Jingchen Shu |
AAAI | 2 |
| 2026 | Character-Level Singing Technique Detection and Evaluation: A Two-Stage Approach
Yiquan Zhou, Haijun Duan |
IEEE Signal Process. Lett. | 3 |
| 2025 | SYKI-SVC: Advancing Singing Voice Conversion with Post-Processing Innovations and an Open-Source Professional TestsetabstractSinging voice conversion aims to transform a source singing voice into that of a target singer while preserving the original lyrics, melody, and various vocal techniques. In this paper, we propose a high-fidelity singing voice conversion system. Our system builds upon the SVCC T02 framework and consists of three key components: a feature extractor, a voice converter, and a post-processor. The feature extractor utilizes the ContentVec and Whisper models to derive F0 contours and extract speaker-independent linguistic features from the input singing voice. The voice converter then integrates the extracted timbre, F0, and linguistic content to synthesize the target speaker’s waveform. The post-processor augments high-frequency information directly from the source through simple and effective signal processing to enhance audio quality. Due to the lack of a standardized professional dataset for evaluating expressive singing conversion systems, we have created and made publicly available a specialized test set. Comparative evaluations demonstrate that our system achieves a remarkably high level of naturalness, and further analysis confirms the efficacy of our proposed system design. Yiquan Zhou, Hongwu Ding, Jiacheng Xu 0007, Jihua Zhu |
ICASSP | 1 |
| 2025 | AVENet: Disentangling Features by Approximating Average Features for Voice ConversionabstractVoice conversion (VC) has made progress in feature disentanglement, but it is still difficult to balance timbre and content information. This paper evaluates the pre-trained model features commonly used in voice conversion, and proposes an innovative method for disentangling speech feature representations. Specifically, we first propose an ideal content feature, referred to as the average feature, which is calculated by averaging the features within frame-level aligned parallel speech (FAPS) data. For generating FAPS data, we utilize a technique that involves freezing the duration predictor in a Text-to-Speech system and manipulating speaker embedding. To fit the average feature on traditional VC datasets, we then design the AVENet to take features as input and generate closely matching average features. Experiments are conducted on the performance of AVENet-extracted features within a VC system. The experimental results demonstrate its superiority over multiple current speech feature disentangling methods. These findings affirm the effectiveness of our disentanglement approach. Yiquan Zhou, Jihua Zhu, Hongwu Ding, Jiacheng Xu 0007 |
ICME | 2 |
| 2025 | FabasedVC: Enhancing Voice Conversion with Text Modality Fusion and Phoneme-Level SSL FeaturesabstractIn voice conversion (VC), it is crucial to preserve complete semantic information while accurately modeling the target speaker’s timbre and prosody. This paper proposes FabasedVC to achieve VC with enhanced similarity in timbre, prosody, and duration to the target speaker, as well as improved content integrity. It is an end-to-end VITS-based VC system that integrates relevant textual modality information, phoneme-level self-supervised learning (SSL) features, and a duration predictor. Specifically, we employ a text feature encoder to encode attributes such as text, phonemes, tones and BERT features. We then process the frame-level SSL features into phoneme-level features using two methods: average pooling and attention mechanism based on each phoneme’s duration. Moreover, a duration predictor is incorporated to better align the speech rate and prosody of the target speaker. Experimental results demonstrate that our method outperforms competing systems in terms of naturalness, similarity, and content integrity. We strongly recommend that readers listen to our samples.1 Zhetao Hu, Yiquan Zhou, Jiacheng Xu 0007, Zhiyu Wu |
MMAsia | 3 |
| 2024 | Switching Kalman Filter for State-of-Charge Estimation of Li-ion Battery Balancing SystemsabstractState of charge (SOC) estimation of lithium-ion batteries has been extensively studied, and most of the existing research focuses on SOC estimation of individual lithium-ion battery. In practical applications, however, lithium-ion batteries are connected to a balancing circuit to eliminate imbalances between batteries. When a balancing circuit is activated, the state space equation of its equivalent circuit will change. In this paper, we propose a switched extended Kalman filter method for SOC estimation of lithium-ion battery balance systems. The switching system model is established by combining the Li-ion battery equivalent circuit model and the switching resistance balance circuit. A switching extended Kalman filter is designed to estimate the SOC of a switching system. Heng Li 0005, Yiquan Zhou, Ren Zhu, Xiaoyang Chen 0003 |
HPCC | 2 |
| 2024 | State-of-Charge Estimation of Reconfigurable Lithium-ion Batteries Based on Nonlinear Switched SystemabstractIn the field of energy storage, precise estimation of the state-of-charge (SOC) of lithium-ion batteries is crucial for maximizing their efficiency and extending their lifespan. Existing research has predominantly focused on SOC estimation for individual battery cells. However, in practical applications, a balancing circuit is typically integrated into the battery management system (BMS) to mitigate cell imbalance. When the balancing circuit is activated, the battery cell transitions into a new operational mode, rendering conventional SOC estimation techniques ineffective. Reconfigurable circuits, recognized for their adaptability across diverse application environments, present a novel approach for SOC estimation in lithium-ion batteries, leveraging their dynamic reconfiguration capabilities. This paper introduces an innovative method for SOC estimation of reconfigurable lithium-ion batteries, employing an extended Kalman filter (EKF) within the context of reconfigurable circuits. The proposed methodology begins with the design of a switching system for lithium-ion batteries, facilitating equivalent circuit modeling. Subsequently, a nonlinear observer, based on the extended Kalman filter, is developed to estimate the SOC. An experimental platform was also constructed to validate the feasibility and efficiency of the proposed method. Experimental results demonstrate that this approach significantly enhances the accuracy and robustness of SOC estimation compared to traditional methods. Yingze Yang, Ren Zhu, Yiquan Zhou, Yunsheng Fan, Heng Li 0005 |
HPCC | 4 |
| 2024 | FGCL: Fine-Grained Contrastive Learning For Mandarin Stuttering Event DetectionabstractThis paper presents the T031 team’s approach to the StutteringSpeech Challenge in SLT2024. Mandarin Stuttering Event Detection (MSED) aims to detect instances of stuttering events in Mandarin speech. We propose a detailed acoustic analysis method to improve the accuracy of stutter detection by capturing subtle nuances that previous Stuttering Event Detection (SED) techniques have overlooked. To this end, we introduce the Fine-Grained Contrastive Learning (FGCL) framework for MSED. Specifically, we model the frame-level probabilities of stuttering events and introduce a mining algorithm to identify both easy and confusing frames. Then, we propose a stutter contrast loss to enhance the distinction between stuttered and fluent speech frames, thereby improving the discriminative capability of stuttered feature embeddings. Extensive evaluations on English and Mandarin datasets demonstrate the effectiveness of FGCL, achieving a significant increase of over 5.0% in F1 score on Mandarin data1.1FGCL won 3rd place in Mandarin stuttering event detection and automatic speech recognition in SLT2024 Han Jiang 0012, Yiquan Zhou, Hongwu Ding, Jiacheng Xu 0007, Jihua Zhu |
SLT | 3 |
| 2023 | VITS-Based Singing Voice Conversion System with DSPGAN Post-Processing for SVCC2023abstractThis paper presents the T02 team’s system for the Singing Voice Conversion Challenge 2023 (SVCC2023). Our system entails a VITS-based SVC model, incorporating three modules: a feature extractor, a voice converter, and a postprocessor. Specifically, the feature extractor provides F0 contours and extracts speaker-independent linguistic content from the input singing voice by leveraging a HuBERT model. The voice converter is employed to recompose the speaker timbre, F0, and linguistic content to generate the waveform of the target speaker. Besides, to further improve the audio quality, a fine-tuned DSPGAN vocoder is introduced to resynthesise the waveform. Given the limited target speaker data, we utilize a two-stage training strategy to adapt the base model to the target speaker. During model adaptation, several tricks, such as data augmentation and joint training with auxiliary singer data, are involved. Official challenge results show that our system achieves superior performance, especially in the cross-domain task, ranking 1st and 2nd in naturalness and similarity, respectively. Further ablation justifies the effectiveness of our system design. Yiquan Zhou, Jihua Zhu, Weifeng Zhao |
ASRU | 1 |