VLDB 2026 Research / reviewers in the wild / expert
Xiang Li 0105
dblp:40/1491-105
· DBLP profile ↗
10ranked-venue papers
3as first author
7since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | SnakeGAN: A Universal Vocoder Leveraging DDSP Prior Knowledge and Periodic Inductive BiasabstractGenerative adversarial network (GAN)-based neural vocoders have been widely used in audio synthesis tasks due to their high generation quality, efficient inference, and small computation footprint. However, it is still challenging to train a universal vocoder which can generalize well to out-of-domain (OOD) scenarios, such as unseen speaking styles, non-speech vocalization, singing, and musical pieces. In this work, we propose SnakeGAN, a GAN-based universal vocoder, which can synthesize high-fidelity audio in various OOD scenarios. SnakeGAN takes a coarse-grained signal generated by a differentiable digital signal processing (DDSP) model as prior knowledge, aiming at recovering high-fidelity waveform from a Mel-spectrogram. We introduce periodic nonlinearities through the Snake activation function and anti-aliased representation into the generator, which further brings desired inductive bias for audio synthesis and significantly improves the extrapolation capacity for universal vocoding in unseen scenarios. To validate the effectiveness of our proposed method, we train SnakeGAN with only speech data and evaluate its performance for various OOD distributions with both subjective and objective metrics. Experimental results show that SnakeGAN significantly outperforms the compared approaches and can generate high-fidelity audio samples including unseen speakers with unseen styles, singing voices, instrumental pieces, and nonverbal vocalization. Sipan Li, Songxiang Liu, Xiang Li 0105, Yanyao Bian, Chao Weng, Zhiyong Wu 0001, Helen M. Meng |
ICME | 4 |
| 2023 | Diverse and Expressive Speech Prosody Prediction with Denoising Diffusion Probabilistic Model
Xiang Li 0105, Songxiang Liu, Max W. Y. Lam, Zhiyong Wu 0001, Chao Weng, Helen M. Meng |
INTERSPEECH | 1 |
| 2022 | Towards Cross-speaker Reading Style Transfer on Audiobook DatasetabstractCross-speaker style transfer aims to extract the speech style of the given reference speech, which can be reproduced in the timbre of arbitrary target speakers.Existing methods on this topic have explored utilizing utterance-level style labels to perform style transfer via either global or local scale style representations.However, audiobook datasets are typically characterized by both the local prosody and global genre, and are rarely accompanied by utterance-level style labels.Thus, properly transferring the reading style across different speakers remains a challenging task.This paper aims to introduce a chunk-wise multi-scale cross-speaker style model to capture both the global genre and the local prosody in audiobook speeches.Moreover, by disentangling speaker timbre and style with the proposed switchable adversarial classifiers, the extracted reading style is made adaptable to the timbre of different speakers.Experiment results confirm that the model manages to transfer a given reading style to new target speakers.With the support of local prosody and global genre type predictor, the potentiality of the proposed method in multi-speaker audiobook generation is further revealed. Xiang Li 0105, Changhe Song, Xianhao Wei, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng |
INTERSPEECH | 1 |
| 2022 | CALM: Constrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis
Xiang Li 0105, Zhiyong Wu 0001, Tingtian Li, Zixun Sun, Xinyu Xiao, Chi Sun, Hui Zhan, Helen M. Meng |
INTERSPEECH | 2 |
| 2022 | Content-Dependent Fine-Grained Speaker Embedding for Zero-Shot Speaker Adaptation in Text-to-Speech SynthesisabstractZero-shot speaker adaptation aims to clone an unseen speaker's voice without any adaptation time and parameters.Previous researches usually use a speaker encoder to extract a global fixed speaker embedding from reference speech, and several attempts have tried variable-length speaker embedding.However, they neglect to transfer the personal pronunciation characteristics related to phoneme content, leading to poor speaker similarity in terms of detailed speaking styles and pronunciation habits.To improve the ability of the speaker encoder to model personal pronunciation characteristics, we propose content-dependent fine-grained speaker embedding for zero-shot speaker adaptation.The corresponding local content embeddings and speaker embeddings are extracted from a reference speech, respectively.Instead of modeling the temporal relations, a reference attention module is introduced to model the content relevance between the reference speech and the input text, and to generate the finegrained speaker embedding for each phoneme encoder output.The experimental results show that our proposed method can improve speaker similarity of synthesized speeches, especially for unseen speakers. Yixuan Zhou 0002, Changhe Song, Xiang Li 0105, Zhiyong Wu 0001, Yanyao Bian, Dan Su 0002, Helen M. Meng |
INTERSPEECH | 3 |
| 2021 | Emotion Controllable Speech Synthesis Using Emotion-Unlabeled Dataset with the Assistance of Cross-Domain Speech Emotion RecognitionabstractNeural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach for emotional TTS synthesis on a TTS dataset without emotion labels. Specifically, our proposed method consists of a cross-domain speech emotion recognition (SER) model and an emotional TTS model. Firstly, we train the cross-domain SER model on both SER and TTS datasets. Then, we use emotion labels on the TTS dataset predicted by the trained SER model to build an auxiliary SER task and jointly train it with the TTS model. Experimental results show that our proposed method can generate speech with the specified emotional expressiveness and nearly no hindering on the speech quality. Xiong Cai, Dongyang Dai, Zhiyong Wu 0001, Xiang Li 0105, Jingbei Li, Helen M. Meng |
ICASSP | 4 |
| 2021 | Towards Multi-Scale Style Control for Expressive Speech SynthesisabstractThis paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis.The proposed method employs a multi-scale reference encoder to extract both the global-scale utterance-level and the local-scale quasi-phoneme-level style features of the target speech, which are then fed into the speech synthesis model as an extension to the input phoneme sequence.During training time, the multiscale style model could be jointly trained with the speech synthesis model in an end-to-end fashion.By applying the proposed method to style transfer task, experimental results indicate that the controllability of the multi-scale speech style model and the expressiveness of the synthesized speech are greatly improved.Moreover, by assigning different reference speeches to extraction of style on each scale, the flexibility of the proposed method is further revealed. Xiang Li 0105, Changhe Song, Jingbei Li, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng |
Interspeech | 1 |
| 2019 | Understanding the Teaching Styles by an Attention based Multi-task Cross-media Dimensional ModelingabstractTeaching style plays an influential role in helping students to achieve academic success. In this paper, we explore a new problem of effectively understanding teachers' teaching styles. Specifically, we study 1) how to quantitatively characterize various teachers' teaching styles for various teachers and 2) how to model the subtle relationship between cross-media teaching related data (speech, facial expressions and body motions, content et al.) and teaching styles. Using the adjectives selected from more than 10,000 feedback questionnaires provided by an educational enterprise, a novel concept called Teaching Style Semantic Space (TSSS) is developed based on the pleasure-arousal dimensional theory to describe teaching styles quantitatively and comprehensively. Then a multi-task deep learning based model, Attention-based Multi-path Multi-task Deep Neural Network (AMMDNN), is proposed to accurately and robustly capture the internal correlations between cross-media features and TSSS. Based on the benchmark dataset, we further develop a comprehensive data set including 4,541 full-annotated cross-modality teaching classes. Our experimental results demonstrate that the proposed AMMDNN outperforms (+0.0842% in terms of the concordance correlation coefficient (CCC) on average) baseline methods. To further demonstrate the advantages of the proposed TSSS and our model, several interesting case studies are carried out, such as teaching styles comparison among different teachers and courses, and leveraging the proposed method for teaching quality analysis. Suping Zhou, Jia Jia 0001, Yufeng Yin 0002, Xiang Li 0105, Zeyang Ye, Kehua Lei, Jialie Shen 0001 |
ACM Multimedia | 4 |
| 2018 | Few-Shot Charge Prediction with Discriminative Legal AttributesabstractAutomatic charge prediction aims to predict the final charges according to the fact descriptions in criminal cases and plays a crucial role in legal assistant systems. Existing works on charge prediction perform adequately on those high-frequency charges but are not yet capable of predicting few-shot charges with limited cases. Moreover, these exist many confusing charge pairs, whose fact descriptions are fairly similar to each other. To address these issues, we introduce several discriminative attributes of charges as the internal mapping between fact descriptions and charges. These attributes provide additional information for few-shot charges, as well as effective signals for distinguishing confusing charges. More specifically, we propose an attribute-attentive charge prediction model to infer the attributes and charges simultaneously. Experimental results on real-work datasets demonstrate that our proposed model achieves significant and consistent improvements than other state-of-the-art baselines. Specifically, our model outperforms other baselines by more than 50% in the few-shot scenario. Our codes and datasets can be obtained from https://github.com/thunlp/attribute_charge. Zikun Hu, Xiang Li 0105, Cunchao Tu, Zhiyuan Liu 0001, Maosong Sun 0001 |
COLING | 2 |
| 2018 | IcooBook: When the Picture Book for Children Encounters Aesthetics of InteractionabstractIn this work, we propose a novel PCA (Perception & Cognition & Affection) model from the prospective of aesthetics in interaction. Based on PCA, we establish a new electronic interactive picture book for children, named IcooBook. At the first level of perception, the proposed IcooBook provides interfaces of multi-sensory interaction; at the second level of cognition, IcooBook builds immersive interactive scenes; at the third level of affection, IcooBook creates high-level interaction modes based on automatic emotion recognition. The research on user study had proved the effectiveness of IcooBook in helping children being focusing on reading, getting better understanding about the context, and further encouraging children to appreciate the beauty of deep affective interaction. Yaohua Bu, Jia Jia 0001, Xiang Li 0105, Suping Zhou, Xiaobo Lu |
ACM Multimedia | 3 |