EDBT 2026 Demo / reviewers in the wild / expert
Jianhao Ye
dblp:314/6416
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Speech recognition and synthesis · 77% Generative modeling · 23% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
voice conversion |
1.7 | 2 | 2025 | Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling · ACL (1) 2025 StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching · AAAI 2025 |
Machine learning › Generative modeling
flow matching |
0.9 | 1 | 2025 | StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching · AAAI 2025 |
Natural language and speech › Speech recognition and synthesis › voice conversion
zero-shot voice conversion |
0.9 | 1 | 2025 | Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling · ACL (1) 2025 |
Processor architecture and microarchitecture
instruction set architecture |
0.9 | 1 | 2025 | Titan-I: An Open-Source, High Performance RISC-V Vector Core · MICRO 2025 |
Processor architecture and microarchitecture
out-of-order execution |
0.9 | 1 | 2025 | Titan-I: An Open-Source, High Performance RISC-V Vector Core · MICRO 2025 |
Processor architecture and microarchitecture › instruction set architecture › vector extension
RISC-V vector extension |
0.9 | 1 | 2025 | Titan-I: An Open-Source, High Performance RISC-V Vector Core · MICRO 2025 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.3 | 1 | 2025 | Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling · ACL (1) 2025 |
Processor architecture and microarchitecture
vector processing |
0.3 | 1 | 2025 | Titan-I: An Open-Source, High Performance RISC-V Vector Core · MICRO 2025 |
Methods — techniques the papers use, named apart from their topics
hybrid content encoding · 0.9dual attention mechanism · 0.9conditional flow matching · 0.9adaptive encoding · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow MatchingabstractZero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content. Despite recent advancements in zero-shot VC using language model-based or diffusion-based approaches, several challenges remain: 1) current approaches primarily focus on adapting timbre from unseen speakers and are unable to transfer style and timbre to different unseen speakers independently; 2) these approaches often suffer from slower inference speeds due to the autoregressive modeling methods or the need for numerous sampling steps; 3) the quality and similarity of the converted samples are still not fully satisfactory. To address these challenges, we propose a Style controllable zero-shot VC approach named StableVC, which aims to transfer timbre and style from source speech to different unseen target speakers. Specifically, we decompose speech into linguistic content, timbre, and style, and then employ a conditional flow matching module to reconstruct the high-quality mel-spectrogram based on these decomposed features. To effectively capture timbre and style in a zero-shot manner, we introduce a novel dual attention mechanism with an adaptive gate, rather than using conventional feature concatenation. With this non-autoregressive design, StableVC can efficiently capture the intricate timbre and style from different unseen speakers and generate high-quality speech significantly faster than real-time. Experiments demonstrate that our proposed StableVC outperforms state-of-the-art baseline systems in zero-shot VC and achieves flexible control over timbre and style from different unseen speakers. Moreover, StableVC offers approximately 25x and 1.65x faster sampling compared to autoregressive and diffusion-based baselines. Jixun Yao, Yuguang Yang 0005, Yu Pan 0008, Ziqian Ning, Jianhao Ye, Hongbin Zhou, Lei Xie 0001 |
AAAI | 5 |
| 2025 | Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre ModelingabstractYang Yuguang, Yu Pan, Jixun Yao, Xiang Zhang, Jianhao Ye, Hongbin Zhou, Lei Xie, Lei Ma, Jianjun Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yuguang Yang 0005, Yu Pan 0008, Jixun Yao, Jianhao Ye, Hongbin Zhou, Lei Xie 0001, Lei Ma 0003, Jianjun Zhao 0001 |
ACL (1) | 5 |
| 2025 | ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
Yu Pan 0008, Yanni Hu, Yuguang Yang 0005, Jixun Yao, Jianhao Ye, Hongbin Zhou, Lei Ma 0003, Jianjun Zhao 0001 |
INTERSPEECH | 5 |
| 2025 | Titan-I: An Open-Source, High Performance RISC-V Vector CoreabstractVector processing has evolved from early systems like the CDC STAR-100 and Cray-1 to modern ISAs like ARM's Scalable Vector Extension (SVE) and RISC-V Vector (RVV) extensions.However, scaling vector processing for contemporary workloads presents challenges due to overheads in traditional architectures.We introduce Titan-I (T1), an out-of-order (OoO) RVV architecture designed Jiuyang Liu, Qinjun Li, Yunqian Luo, Jiongjia Lu, Shupei Fan, Jianhao Ye, Yanqi Yang, Zewen Ye, Yuhang Zeng, Wei Cong, Xuecheng Zou, Mingyu Gao 0001 |
MICRO | 7 |
| 2022 | Improving Cross-Lingual Speech Synthesis with Triplet Training SchemeabstractRecent advances in cross-lingual text-to-speech (TTS) made it possible to synthesize speech in a language foreign to a monolingual speaker. However, there is still a large gap between the pronunciation of generated cross-lingual speech and that of native speakers in terms of naturalness and intelligibility. In this paper, a triplet training scheme is proposed to enhance the cross-lingual pronunciation by allowing previously unseen content and speaker combinations to be seen during training. Proposed method introduces an extra fine-tune stage with triplet loss during training, which efficiently draws the pronunciation of the synthesized foreign speech closer to those from the native anchor speaker, while preserving the non-native speaker’s timbre. Experiments are conducted based on a state-of-the-art baseline cross-lingual TTS system and its enhanced variants. All the objective and subjective evaluations show the proposed method brings significant improvement in both intelligibility and naturalness of the synthesized cross-lingual speech. Jianhao Ye, Hongbin Zhou, Zhiba Su, Wendi He, Kaimeng Ren, Heng Lu 0004 |
ICASSP | 1 |