EDBT 2026 Demo / reviewers in the wild / expert
Bohan Li 0003
dblp:123/2549-3
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AHAMask: Reliable Task Specification for Large Audio Language Models Without InstructionsabstractAlthough current large audio language models (LALMs) extend text large language models (LLMs) with generic acoustic understanding abilities, they usually suffer from prompt sensitivity, where different instructions of the same intention can yield drastically different outcomes. In this work, we propose AHAMask, where we simply mask some of the attention heads in the decoder-only LLM backbone of LALMs, to trigger specific acoustic task functionalities without instructions. These masks are efficiently obtained by training on an LALM, with the number of trainable parameters equal to the attention head count in its LLM backbone. We show by experiments that applying such selective attention head masks achieves comparable or even better performance than using instructions, either on single or composite tasks. Besides achieving reliable acoustic task specification for LALMs, this also reveals that LALMs exhibit certain ``functional pathways'' in their attention heads. Bohan Li 0003, Hankun Wang, Shuai Wang 0016, Xie Chen 0001, Kai Yu 0004 |
AAAI | 2 |
| 2026 | Recent Advances in Discrete Speech Tokens: A ReviewabstractThe rapid advancement of speech generation technologies in the era of large language models (LLMs) has established discrete speech tokens as a foundational paradigm for speech representation. These tokens, characterized by their discrete, compact, and concise nature, are not only advantageous for efficient transmission and storage, but also inherently compatible with the language modeling framework, enabling seamless integration of speech into text-dominated LLM architectures. Current research categorizes discrete speech tokens into two principal classes: acoustic tokens and semantic tokens, each of which has evolved into a rich research domain characterized by unique design philosophies and methodological approaches. This survey systematically synthesizes the existing taxonomy and recent innovations in discrete speech tokenization, conducts a critical examination of the strengths and limitations of each paradigm, and presents systematic experimental comparisons across token types. Furthermore, we identify persistent challenges in the field and propose potential research directions, aiming to offer actionable insights to inspire future advancements in the development and application of discrete speech tokens. Hankun Wang, Bohan Li 0003, Chongtian Shao, Hanglei Zhang, Chenpeng Du, Xie Chen 0001, Shujie Liu 0001, Kai Yu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task ConsistencyabstractNeural speech codecs excel in reconstructing clean speech signals; however, their efficacy in complex acoustic environments and downstream signal processing tasks remains underexplored. In this study, we introduce a novel benchmark named Environment-Resilient Speech Codec Benchmark (ERSB) to systematically evaluate whether neural speech codecs are environment-resilient. Specifically, we assess two key capabilities: (1) robust reconstruction, which measures the preservation of both speech and non-speech acoustic details, and (2) downstream task consistency, which ensures minimal deviation in downstream signal processing tasks when using reconstructed speech instead of the original. Our comprehensive experiments reveal that complex acoustic environments significantly degrade signal reconstruction and downstream task consistency. This work highlights the limitations of current speech codecs and raises a future direction that improves them for greater environmental resilience. Bohan Li 0003, Hankun Wang, Xie Chen 0001, Kai Yu 0004 |
ASRU | 3 |
| 2025 | Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative DecodingabstractThe auto-regressive (AR) architecture, exemplified by models such as GPT, is extensively utilized in modern Text-to-Speech (TTS) systems. However, it often leads to considerable inference delays, primarily due to the challenges associated with next-token prediction in long speech sequences. In this work, we introduce VADUSA, one of the first approaches to accelerate AR-based TTS through speculative decoding. Our findings demonstrate that VADUSA not only delivers a significant reduction in inference time but also enhances TTS quality by employing draft heads to predict future speech tokens in an auto-regressive manner. Additionally, the incorporation of a tolerance mechanism during the sampling process further boosts performance, yielding approximately a 3× speedup in AR TTS. Moreover, our approach exhibits strong generalization across diverse datasets and various speech token types. Bohan Li 0003, Hankun Wang, Situo Zhang, Kai Yu 0004 |
ICASSP | 1 |
| 2025 | Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict MechanismabstractDiffusion models have emerged as a powerful class of generative models across various modalities, including image, video, and audio synthesis. However, their deployment is often limited by significant inference latency, primarily due to the inherently sequential nature of the denoising process. While existing parallelization strategies attempt to accelerate inference by distributing computation across multiple devices, they typically incur high communication overhead, hindering deployment on commercial hardware. To address this challenge, we propose $\textbf{ParaStep}$, a novel parallelization method based on a reuse-then-predict mechanism that parallelizes diffusion inference by exploiting similarity between adjacent denoising steps. Unlike prior approaches that rely on layer-wise or stage-wise communication, ParaStep employs lightweight, step-wise communication, substantially reducing overhead. ParaStep achieves end-to-end speedups of up to $\textbf{3.88}$$\times$ on SVD, $\textbf{2.43}$$\times$ on CogVideoX-2b, and $\textbf{6.56}$$\times$ on AudioLDM2-large, while maintaining generation quality. These results highlight ParaStep as a scalable and communication-efficient solution for accelerating diffusion inference, particularly in bandwidth-constrained environments. Kunyun Wang, Bohan Li 0003, Kai Yu 0004, Minyi Guo, Jieru Zhao |
NeurIPS | 2 |
| 2024 | On the Effectiveness of Acoustic BPE in Decoder-Only TTS
Bohan Li 0003, Feiyu Shen, Shuai Wang 0016, Xie Chen 0001, Kai Yu 0004 |
INTERSPEECH | 1 |
| 2023 | VideoDubber: Machine Translation with Speech-Aware Length Control for Video DubbingabstractVideo dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation and speech synthesis. To ensure the translated speech to be well aligned with the corresponding video, the length/duration of the translated speech should be as close as possible to that of the original speech, which requires strict length control. Previous works usually control the number of words or characters generated by the machine translation model to be similar to the source sentence, without considering the isochronicity of speech as the speech duration of words/characters in different languages varies. In this paper, we propose VideoDubber, a machine translation system tailored for the task of video dubbing, which directly considers the speech duration of each token in translation, to match the length of source and target speech. Specifically, we control the speech length of generated sentence by guiding the prediction of each word with the duration information, including the speech duration of itself as well as how much duration is left for the remaining words. We design experiments on four language directions (German -> English, Spanish -> English, Chinese English), and the results show that VideoDubber achieves better length control ability on the generated speech than baseline methods. To make up the lack of real-world datasets, we also construct a real-world test set collected from films to provide comprehensive evaluations on the video dubbing task. Yihan Wu 0008, Junliang Guo, Xu Tan 0003, Chen Zhang 0020, Bohan Li 0003, Ruihua Song, Lei He 0005, Sheng Zhao 0002, Arul Menezes, Jiang Bian 0002 |
AAAI | 5 |
| 2022 | AdaSpeech 4: Adaptive Text to Speech in Zero-Shot ScenariosabstractAdaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new speakers. Considering seen and unseen speakers have diverse characteristics, zero-shot adaptive TTS requires strong generalization ability on speaker characteristics, which brings modeling challenges. In this paper, we develop AdaSpeech 4, a zero-shot adaptive TTS system for high-quality speech synthesis. We model the speaker characteristics systematically to improve the generalization on new speakers. Generally, the modeling of speaker characteristics can be categorized into three steps: extracting speaker representation, taking this speaker representation as condition, and synthesizing speech/mel-spectrogram given this speaker representation. Accordingly, we improve the modeling in three steps: 1) To extract speaker representation with better generalization, we factorize the speaker characteristics into basis vectors and extract speaker representation by weighted combining of these basis vectors through attention. 2) We leverage conditional layer normalization to integrate the extracted speaker representation to TTS model. 3) We propose a novel supervision loss based on the distribution of basis vectors to maintain the corresponding speaker characteristics in generated mel-spectrograms. Without any fine-tuning, AdaSpeech 4 achieves better voice quality and similarity than baselines in multiple datasets. Yihan Wu 0008, Xu Tan 0003, Bohan Li 0003, Lei He 0005, Sheng Zhao 0002, Ruihua Song, Tao Qin 0001, Tie-Yan Liu |
INTERSPEECH | 3 |
| 2021 | Adaspeech 2: Adaptive Text to Speech with Untranscribed DataabstractText to speech (TTS) is widely used to synthesize personal voice for a target speaker, where a well-trained source TTS model is fine-tuned with few paired adaptation data (speech and its transcripts) on this target speaker. However, in many scenarios, only untranscribed speech data is available for adaptation, which brings challenges to the previous TTS adaptation pipelines (e.g., AdaSpeech). In this paper, we develop AdaSpeech 2, an adaptive TTS system that only leverages untranscribed speech data for adaptation. Specifically, we introduce a mel-spectrogram encoder to a well-trained TTS model to conduct speech reconstruction, and at the same time constrain the output sequence of the mel-spectrogram encoder to be close to that of the original phoneme encoder. In adaptation, we use untranscribed speech data for speech reconstruction and only fine-tune the TTS decoder. AdaSpeech 2 has two advantages: 1) Pluggable: our system can be easily applied to existing trained TTS models without re-training. 2) Effective: our system achieves on-par voice quality with the transcribed TTS adaptation (e.g., AdaSpeech) with the same amount of untranscribed data, and achieves better voice quality than previous untranscribed adaptation methods1. Yuzi Yan, Xu Tan 0003, Bohan Li 0003, Tao Qin 0001, Sheng Zhao 0002, Yuan Shen 0001, Tie-Yan Liu |
ICASSP | 3 |
| 2021 | AdaSpeech: Adaptive Text to Speech for Custom Voice
Xu Tan 0003, Bohan Li 0003, Tao Qin 0001, Sheng Zhao 0002, Tie-Yan Liu |
ICLR | 3 |
| 2021 | Adaptive Text to Speech for Spontaneous Style
Yuzi Yan, Xu Tan 0003, Bohan Li 0003, Guangyan Zhang, Tao Qin 0001, Sheng Zhao 0002, Yuan Shen 0001, Tie-Yan Liu |
Interspeech | 3 |