EDBT 2026 Demo / reviewers in the wild / expert
Hao Li 0078
dblp:17/5705-78
· DBLP profile ↗
14ranked-venue papers
5as first author
7since 2021 · last 2024
0009-0005-4319-9026ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic CodingabstractRecently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations and using two sequence-to-sequence tasks to decouple TTS. However, existing methods suffer from three problems: the high-frequency waveform distortion of discrete speech representations, the prosodic averaging problem caused by the duration prediction model in non-autoregressive frameworks, and difficulty in prediction due to the information redundancy and dimension explosion of existing semantic coding methods. To address these problems, three progressive methods are proposed. First, we propose Diff-LM-Speech, an autoregressive structure consisting of a language model and diffusion models, which models the semantic embedding into the mel-spectrogram based on a diffusion model to achieve higher audio quality. We also introduce a prompt encoder structure based on a variational autoencoder and a prosody bottleneck to improve prompt representation ability. Second, we propose Tetra-Diff-Speech, a non-autoregressive structure consisting of four diffusion model-based modules that design a duration diffusion model to achieve diverse prosodic expressions. Finally, we propose Tri-Diff-Speech, a non-autoregressive structure consisting of three diffusion model-based modules that verify the non-necessity of existing semantic coding models and achieve the best results. Experimental results show that our proposed methods outperform baseline methods. We provide a website with audio samples.1 Chunyu Qiang, Hao Li 0078, He Qu, Ruibo Fu, Tao Wang 0074, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 2 |
| 2024 | Learning Speech Representation from Contrastive Token-Acoustic PretrainingabstractFor fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate representations extracted from speech should serve as a "bridge" between text and acoustic information, containing information from both modalities. The semantic content is emphasized, while the paralinguistic information such as speaker identity and acoustic details should be de-emphasized. However, existing methods for extracting fine-grained intermediate representations from speech suffer from issues of excessive redundancy and dimension explosion. Contrastive learning is a good method for modeling intermediate representations from two modalities. However, existing contrastive learning methods in the audio field focus on extracting global descriptive information for downstream audio classification tasks, making them unsuitable for TTS, VC, and ASR tasks. To address these issues, we propose a method named "Contrastive Token-Acoustic Pretraining (CTAP)", which uses two encoders to bring phoneme and speech into a joint multimodal space, learning how to connect phoneme and speech at the frame level. The CTAP model is trained on 210k speech and phoneme pairs, achieving minimally-supervised TTS, VC, and ASR. The proposed CTAP method offers a promising solution for fine-grained generation and recognition downstream tasks in speech processing. We provide a website with audio samples.1 Chunyu Qiang, Hao Li 0078, Yixin Tian, Ruibo Fu, Tao Wang 0074, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 2 |
| 2024 | High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion ModelsabstractText-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of discrete speech representations(semantic & acoustic) and using two sequence-to-sequence tasks to enable training with minimal supervision. However, existing methods suffer from information redundancy and dimension explosion in semantic representation, and high-frequency waveform distortion in discrete acoustic representation. Autoregressive frameworks exhibit typical instability and uncontrollability issues. And non-autoregressive frameworks suffer from prosodic averaging caused by duration prediction models. To address these issues, we propose a minimally-supervised high-fidelity speech synthesis method, where all modules are constructed based on the diffusion models. The non-autoregressive framework enhances controllability, and the duration diffusion model enables diversified prosodic expression. Contrastive Token-Acoustic Pretraining (CTAP) is used as an intermediate semantic representation to solve the problems of information redundancy and dimension explosion in existing semantic coding methods. Mel-spectrogram is used as the acoustic representation. Both semantic and acoustic representations are predicted by continuous variable regression tasks to solve the problem of high-frequency fine-grained waveform distortion. Experimental results show that our proposed method outperforms the baseline method. We provide audio samples on our website.1 Chunyu Qiang, Hao Li 0078, Yixin Tian, Yi Zhao 0006, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 2 |
| 2024 | Enhancing Realism in 3D Facial Animation Using Conformer-Based Generation and Automated Post-ProcessingabstractRecent progress has propelled the development of realistic talking-face videos for avatars. Yet, animating 3D cartoon avatars remains intricate due to the imprecise nature of facial-driven data. This often manifests as inconsistent mouth configurations and rigid facial expressions, curbing the animation’s realism. Addressing these issues, we introduce a conformer-based framework that derives expression coefficients directly from phonemes, thereby elevating prediction precision and minimizing manual oversight. Furthermore, by harnessing a pre-trained emotion blending module coupled with the keyframe of the target emotional character, we employ a zero-shot adaptation technique. This serves to amplify emotional expressions and bolster the authenticity of lip dynamics. Our methodology adeptly registers nuanced expression shifts in avatars, leading to remarkably lifelike animations, as substantiated by our experimental findings. Yi Zhao 0006, Chunyu Qiang, Hao Li 0078, Yulan Hu, Wangjin Zhou, Sheng Li 0010 |
ICASSP | 3 |
| 2024 | Jointly Recognizing Speech and Singing Voices Based on Multi-Task Audio Source SeparationabstractIn short video and live broadcasts, speech, singing voice, and background music often overlap and obscure each other. This complexity creates difficulties in structuring and recognizing the audio content, which may impair subsequent ASR and music understanding applications. This paper proposes a multi-task audio source separation (MTASS) based ASR model called JRSV, which Jointly Recognizes Speech and singing Voices. Specifically, the MTASS module separates the mixed audio into distinct speech and singing voice tracks while removing background music. The CTC/attention hybrid recognition module recognizes both tracks. Online distillation is proposed to improve the robustness of recognition further. To evaluate the proposed methods, a benchmark dataset is constructed and released. Experimental results demonstrate that JRSV can significantly improve recognition accuracy on each track of the mixed audio. Ye Bai 0001, Chenxing Li, Hao Li 0078 |
ICME | 3 |
| 2023 | HoloSinger: Semantics and Music Driven Motion Generation with Octahedral Holographic ProjectionabstractLyrics and music are both significant for a singer to perform a song. Therefore, it is important in singer's motion generation to model both semantic and acoustic correlation with motions at the same time. In this paper, we propose HoloSinger, a novel comprehensive system that synthesizes singing motions according to the given song. Additionally, we present singing avatar with octahedral holographic projection. For singing motion generation, we introduce a Transformer-VAE generative model to decompose lyrics and music, then fuse their impacts to synthesize singer's motions. Extensive experiments and user studies show that our method automatically generates realistic motions that adhere to musical choreography and reflect the lyric semantics appropriately. Furthermore, we design a desktop-level holographic projection device with an octahedral structure. It achieves high-definition holographic projection effects with smaller volume, larger imaging area ratio, and the ability of real-time AI interaction. Zeyu Jin, Zixuan Wang 0026, Qixin Wang 0002, Jia Jia 0001, Ye Bai 0001, Yi Zhao 0006, Hao Li 0078 |
ACM Multimedia | 7 |
| 2022 | Improving Spoken Language Understanding with Cross-Modal Contrastive Learning
Jingjing Dong, Jiayi Fu, Hao Li 0078 |
INTERSPEECH | 4 |
| 2018 | EMPHASIS: An Emotional Phoneme-based Acoustic Model for Speech Synthesis SystemabstractWe present EMPHASIS, an emotional phoneme-based acoustic model for speech synthesis system. EMPHASIS includes a phoneme duration prediction model and an acoustic parameter prediction model. It uses a CBHG-based regression network to model the dependencies between linguistic features and acoustic features. We modify the input and output layer structures of the network to improve the performance. For the linguistic features, we apply a feature grouping strategy to enhance emotional and prosodic features. The acoustic parameters are designed to be suitable for the regression task and waveform reconstruction. EMPHASIS can synthesize speech in real-time and generate expressive interrogative and exclamatory speech with high audio quality. EMPHASIS is designed to be a multi-lingual model and can synthesize Mandarin-English speech for now. In the experiment of emotional speech synthesis, it achieves better subjective results than other real-time speech synthesis systems. Hao Li 0078, Yongguo Kang |
INTERSPEECH | 1 |
| 2016 | Emotional head motion predicting from prosodic and linguistic features
Jinlin Jiang, Jianhua Tao 0001, Kaihui Mu, Hao Li 0078 |
Multim. Tools Appl. | 5 |
| 2015 | Evaluation of linear regression for speaker adaptation in HMM-based articulatory movements estimationabstractAcoustic-to-articulatory inversion problem is usually studied in speaker-specific manner because both articulatory data and acoustic features contain speaker-specific components. This paper presents our work on speaker-adaptation training for this problem. We implement speaker adaptation in HMM-based acoustic-to-articulatory inversion mapping, and evaluate different combinatorial structures of the articulatory data and acoustic features. The HMM-based inversion mapping models are built with single-stream and multistream, independent clustering and shared clustering structures. The speaker adaptation is implemented in stream-independent structure and shared adaptation structure. The constrained maximum likelihood linear regression method is used for the speaker-adaptive transformation. The experimental results show that the sharing of the speaker-adaptive transformation of the articulatory feature stream and acoustic feature stream can improve the estimation accuracy in inversion mapping. The multi-stream system with shared clustering and shared adaptive transformation has the best result among all the tested structures. Hao Li 0078, Jianhua Tao 0001 |
ICASSP | 1 |
| 2015 | Estimate articulatory MRI series from acoustic signal using deep architectureabstractThis paper presents our work on acoustic-to-articulatory inversion mapping, in which, the articulatory data is the MRI series for articulators on mid-sagittal plan. Deep architectures based on restricted Boltzmann machine (RBM) and linear regression are employed to construct the audio-visual mapping. We test two architectures to initialize the neural network: the bottom-up stacked RBM with top regression layer architecture and the one with extra Gaussian-Bernoulli RBM on the top of the former architecture. GMM-based mapping is used as baseline method. The MRI data from USC-TIMIT database is used for the training. The experimental results show that the deep regression network is an effective model to construct the mapping from acoustic speech signal to articulatory MRI series, and also indicate that it is a better strategy to initial the top layer as Gaussian-Bernoulli RBM to compress the MRI data before the liner regression. Hao Li 0078, Jianhua Tao 0001, Bin Liu 0041 |
ICASSP | 1 |
| 2015 | User behavior fusion in dialog management with multi-modal history cues
Jianhua Tao 0001, Linlin Chao, Hao Li 0078, Dawei Zhang 0001, Hao Che, Tingli Gao, Bin Liu 0041 |
Multim. Tools Appl. | 4 |
| 2014 | Tongue shape conversion with non-parallel training dataabstractArticulatory data is an indispensable resource for speech production research. It will facilitate this study if we can convert one speaker's articulatory data to adapt a given target speaker. In this paper, we propose a tongue shape conversion method for nonparallel training data. The method combines thin-plate spline approximation (TPSA) algorithm with codebook mapping. The TPSA is a spatial morph method with landmarks extracted from articulatory data with phonetic segmentations. The landmarks' degree of certainty is evaluated and be considered in the TPSA morph. The proposed method has the advantages of the spatial morph and the codebook mapping by considering both the spatial configuration and the acoustic parameters. The results of our experiments with electromagnetic articulography (EMA) data indicate that the proposed method yields better results than the spatial morph method and the codebook mapping regardless the amount of training data. Hao Li 0078, Jianhua Tao 0001 |
ICASSP | 1 |
| 2013 | Speaker-independent lips and tongue visualization of vowelsabstractThis paper proposes a scheme of speech-driven lips and tongue animation synthesis in a speaker-independent manner. Directional relative displacement (DRD) features are proposed based on the Electromagnetic Articulograph (EMA) data to describe human's lips and tongue movements, which are more stable across different speakers than the raw EMA data. Multi speakers' acoustic-articulatory data of vowels are used to learn the acoustic-toarticulatory inversion mapping. We build 2D geometric models of lips and tongue for visualization. With the trained mapping and the geometric models, visualization of lips and tongue movements from acoustic signal of vowels uttered by arbitrary speaker is realized. The experimental results demonstrate that the animations we synthesized are effective aids in helping people identifying vowels. Hao Li 0078, Jianhua Tao 0001 |
ICASSP | 1 |