EDBT 2026 Demo / reviewers in the wild / expert
Zhiyong Wu 0001
dblp:24/968-1
· DBLP profile ↗
205ranked-venue papers
5as first author
135since 2021 · last 2026
0000-0001-8533-0524ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 182 · 4 first-author · 116 since 2021Artificial intelligence and machine learning · 94 · 3 first-author · 63 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human-Centric Video Generation via Collaborative Multi-Modal ConditioningabstractHuman-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, images, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of modality-complete data and the difficulty of jointly modeling triplet conditions without performance degradation. In this work, we present HuMo, a unified HCVG framework for collaborative multimodal control. For the first challenge, we construct an incomplete-yet-complementary dataset for improved data utilization efficiency and training scalability. For the second challenge, we propose a two-stage progressive multimodal training paradigm with task-specific strategies at each stage. In the first stage, to balance the text-following and subject-preservation abilities, we adopt the minimal-invasive image injection strategy. In the second stage, to enhance audio-visual sync, we propose a focus-by-predicting strategy that implicitly guides the model to associate audio with facial regions. For joint learning of controllabilities across multi-modal inputs, we progressively incorporate the audio-visual sync task, building on previously acquired capabilities. During inference, for flexible and fine-grained multimodal control, we design a stage-adaptive Classifier-Free Guidance strategy that dynamically adjusts guidance weights across denoising steps. Extensive experimental results demonstrate that HuMo surpasses specialized state-of-the-art methods in sub-tasks, establishing a unified framework for collaborative multimodal-conditioned HCVG. Liyang Chen, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Lijie Liu, Zhiyong Wu 0001 |
AAAI | 10 |
| 2026 | DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language ModelsabstractExtending pre-trained text Large Language Models (LLMs)’s speech understanding or generation abilities by introducing various effective speech tokens has attracted great attention in the speech research community. However, building a unified speech understanding and generation model still faces the following challenges: (1) Due to the huge modality gap between speech and text tokens, extending text LLMs to unified speech LLMs relies on large-scale paired data for fine-tuning, and (2) Generation and understanding tasks prefer information at different levels, e.g., generation benefits from detailed acoustic features, while understanding favors high-level semantics. This divergence leads to difficult performance optimization in one unified model. To solve these challenges, in this paper, we present two key insights in speech tokenization and speech language modeling. Specifically, we first propose an Understanding-driven Speech Tokenizer (USTokenizer), which extracts high-level semantic information essential for accomplishing understanding tasks using text LLMs. In this way, USToken enjoys better modality commonality with text, which reduces the difficulty of modality alignment in adapting text LLMs to speech LLMs. Secondly, we present DualSpeechLM, a dual-token modeling framework that concurrently models USToken as input and acoustic token as output within a unified, end-to-end framework, seamlessly integrating speech understanding and generation capabilities. Furthermore, we propose a novel semantic supervision loss and a Chain-of-Condition (CoC) strategy to stabilize model training and enhance speech generation performance. Experimental results demonstrate that our proposed approach effectively fosters a complementary relationship between understanding and generation tasks, highlighting the promising strategy of mutually enhancing both tasks in one unified model. Dongchao Yang, Yiwen Shao, Hangting Chen, Jiankun Zhao, Zhiyong Wu 0001, Helen M. Meng, Xixin Wu |
AAAI | 6 |
| 2026 | Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical FallaciesabstractWhile Large Language Models (LLMs) exhibit strong semantic capabilities, their resilience to manipulative linguistic patterns like logical fallacies remains an underexplored area.Prior work has focused on the ability of LLMs to identify or classify fallacies, but their robustness against these fallacies in persuasive contexts remains largely unexplored.To address this gap, we introduce LoFa (Logical Fallacy), a comprehensive benchmark to evaluate LLM robustness against fallacies.We first construct the LoFa dataset via a multi-agent pipeline, pairing factual questions with fallacious arguments.Then, we develop a multi-round debate framework to assess model resilience under sustained attacks.Furthermore, to disentangle robustness from a model's inherent knowledge limitations, we propose a new metric, LFR@k (Logical Fallacy Resistance), to quantify performance.Our experiments reveal that different LLMs exhibit varied robustness to distinct types of fallacies, highlighting unique vulnerability profiles across models....The ground beneath your feet was likely covered in sand, right?Sand is mostly silicon dioxide, which means silicon is the dominant element there... Question: What is the abundant element in earth's crust?, Source:https://en.wikipedia.org//w/index.php... Xin Wu 0003, Yi Cai 0001, Zhiyong Wu 0001 |
ACL (1) | 6 |
| 2026 | UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained AssessmentabstractYuanyuan Wang, Dongchao Yang, Yayue Deng, Zhiyong Wu, Steven Y. Guo, Helen M. Meng, Xixin Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dongchao Yang, Yayue Deng, Zhiyong Wu 0001, Steven Y. Guo, Helen M. Meng, Xixin Wu |
ACL (1) | 4 |
| 2025 | MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative RefinementabstractExisting works in single-image human reconstruction suffer from weak generalizability due to insufficient training data or 3D inconsistencies for a lack of comprehensive multi-view knowledge. In this paper, we introduce MagicMan, a human-specific multi-view diffusion model to generate high-quality novel views from a single reference image. As its core, we leverage a pre-trained 2D diffusion model as the generative prior for generalizability, with the parametric SMPL-X model as the 3D body prior to promote 3D awareness. To maintain consistency while generating denser views for improved 3D human reconstruction, we introduce hybrid multi-view attention to facilitate efficient and thorough information interchange across views. Besides, we present a geometry-aware dual branch to perform concurrent generation in both RGB and normal domains, further enhancing consistency via geometry cues. Last but not least, to address ill-shaped issues arising from inaccurate SMPL-X estimation, we propose a novel iterative refinement strategy, which progressively optimizes SMPL-X accuracy while enhancing the quality and consistency of the generated multi-views. Extensive experimental results demonstrate that our method significantly outperforms existing approaches in both novel view synthesis and subsequent 3D human reconstruction tasks. Zhiyong Wu 0001, Xiaoyu Li 0002, Chaopeng Zhang, Jiangnan Ye 0004, Liyang Chen, Xiangjun Gao, Haolin Zhuang |
AAAI | 2 |
| 2025 | Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio SynthesisabstractOur research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal synchronization adapter, our method significantly improves semantic integrity and the precision of beat point synchronization, particularly in fast-paced action sequences. Utilizing a contrastive audio-visual pre-trained encoder, our model is trained with video and high-quality audio data, improving the quality of the generated audio. This dual-adapter approach empowers users with enhanced control over audio semantics and beat effects, allowing the adjustment of the controller to achieve better results. Extensive experiments substantiate the effectiveness of our framework in achieving seamless audio-visual alignment. Zhiqi Huang 0005, Huan Liao, Zhiyong Wu 0001 |
ICASSP | 6 |
| 2025 | Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody FeaturesabstractMelody preservation is crucial in singing voice conversion (SVC). However, in many scenarios, audio is often accompanied with background music (BGM), which can cause audio distortion and interfere with the extraction of melody and other key features, significantly degrading SVC performance. Previous methods have attempted to address this by using more robust neural network-based melody extractors, but their performance drops sharply in the presence of complex accompaniment. Other approaches involve performing source separation before conversion, but this often introduces noticeable artifacts, leading to a significant drop in conversion quality and increasing the user’s operational costs. To address these issues, we introduce a novel SVC method that uses self-supervised representation-based melody features to improve melody modeling accuracy in the presence of BGM. In our experiments, we compare the effectiveness of different self-supervised learning (SSL) models for melody extraction and explore for the first time how SSL benefits the task of melody extraction. The experimental results demonstrate that our proposed SVC model significantly outperforms existing baseline methods in terms of melody accuracy and shows higher similarity and naturalness in both subjective and objective evaluations across noisy and clean audio environments. Wei Chen 0071, Binzhu Sha, Zhiyong Wu 0001 |
ICASSP | 6 |
| 2025 | Binary Representation Learning for Discriminative Acoustic Unit DiscoveryabstractAcoustic Unit Discovery (AUD) aims to obtain phoneme-like units that preserve linguistically significant information while removing paralinguistic details. Although Contrastive Predictive Coding (CPC) has emerged as a leading self-supervised representation learning method for this task, CPC-based methods still suffer from the limitation that the learned representations are susceptible to paralinguistic information and less discriminative. Inspired by the theory of distinctive features, we propose a new approach that builds a binary discriminative representation space and employs binary contrastive learning based on CPC to tackle the aforementioned issues. Experimental results show that our method achieves better results in AUD task and produces discriminative binary representations. Rui Niu, Changhe Song, Weihao Wu 0001, Zhiyong Wu 0001 |
ICASSP | 6 |
| 2025 | AudioComposer: Towards Fine-grained Audio Generation with Natural Language DescriptionsabstractCurrent Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by incorporating additional frame-level conditions or control networks. However, this usually leads to complex system design and difficulties due to the requirement for reference frame-level conditions. To address these challenges, we propose AudioComposer, a novel TTA generation framework that relies solely on natural language descriptions (NLDs) to provide both content specification and style control information. To further enhance audio generative modeling, we employ flow-based diffusion transformers with the cross-attention mechanism to incorporate text descriptions effectively into audio generation processes, which can not only simultaneously consider the content and style information in the text inputs, but also accelerate generation compared to other architectures. Furthermore, we propose a novel and comprehensive automatic data simulation pipeline to construct data with fine-grained text descriptions, which significantly alleviates the problem of data scarcity in the area. Experiments demonstrate the effectiveness of our framework using solely NLDs as inputs for content specification and style control. The generation quality and controllability surpass stateof-the-art TTA models, even with a smaller model size.1 Hangting Chen, Dongchao Yang, Zhiyong Wu 0001, Xixin Wu |
ICASSP | 4 |
| 2025 | DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion ModelsabstractConversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, existing CSS systems are limited to deterministic prediction, overlooking the diversity of potential responses. Moreover, they rarely employ language model (LM)-based TTS backbones, limiting the naturalness and quality of synthesized speech. To address these issues, in this paper, we propose DiffCSS, an innovative CSS framework that leverages diffusion models and an LM-based TTS backbone to generate diverse, expressive, and contextually coherent speech. A diffusion-based context-aware prosody predictor is proposed to sample diverse prosody embeddings conditioned on multimodal conversational context. Then a prosody-controllable LM-based TTS backbone is developed to synthesize high-quality speech with sampled prosody embeddings. Experimental results demonstrate that the synthesized speech from DiffCSS is more diverse, contextually coherent, and expressive than existing CSS systems1. Weihao Wu 0001, Yixuan Zhou 0002, Jingbei Li, Rui Niu, Songjun Cao, Zhiyong Wu 0001 |
ICASSP | 9 |
| 2025 | Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering DataabstractEmpathetic dialogue is crucial for natural human-computer interaction, allowing the dialogue system to respond in a more personalized and emotionally aware manner, improving user satisfaction and engagement. The emergence of large language models (LLMs) has revolutionized dialogue generation by harnessing their powerful capabilities and shown its potential in multimodal domains. Many studies have integrated speech with text-based LLMs to take speech question as input and output text response. However, the lack of spoken question-answering datasets that include speech style information to supervised fine-tuning (SFT) limits the performance of these systems. As a result, while these systems excel at understanding speech content, they often struggle to generate empathetic responses. In response, we propose a novel approach that circumvents the need for question-answering data, called Listen, Perceive, and Express (LPE). Our method employs a two-stage training process, initially guiding the LLM to listen the content and perceive the emotional aspects of speech. Subsequently, we utilize Chain-of-Thought (CoT) prompting to unlock the model’s potential for expressing empathetic responses based on listened spoken content and perceived emotional cues. We employ experiments to prove the effectiveness of proposed method. To our knowledge, this is the first attempt to leverage CoT for speech-based dialogue. Jingran Xie, Shun Lei, Xixin Wu, Zhiyong Wu 0001 |
ICASSP | 7 |
| 2025 | Identity-Preserving Audio-Driven Holistic Human Motion Video GenerationabstractGenerating realistic human motion videos is a pivotal challenge in advancing human-computer interaction. While existing approaches often focus on generating either head or gesture movements from audio, they lack unified control over full-body motion, frequently producing low-resolution and blurred outputs. Additionally, these methods struggle to maintain character identity throughout the generated content. In this paper, we introduce a novel framework that generates photorealistic, personalized human motion videos from audio by decoupling identity features. We integrate both visual features and voice timbre to enhance the preservation of character identity. Our approach follows a four-stage paradigm: (1) frame generation, (2) identity feature customization, (3) audio-motion modeling, and (4) motion-video rendering. Through the collaborative modeling of audio-motion and motion-video stages, our approach effectively maintains the consistency of character identity and background throughout the video, enhancing the realism and coherence of the generated video. Experimental results demonstrate that our framework delivers high-resolution videos with superior fidelity, establishing a new and effective baseline for holistic human motion video generation. Haiwei Xue, Zhensong Zhang, Minglei Li 0001, Zonghong Dai, Zhiyong Wu 0001 |
ICASSP | 5 |
| 2025 | RFWave: Multi-band Rectified Flow for Audio Waveform ReconstructionabstractRecent advancements in generative modeling have significantly enhanced the reconstruction of audio waveforms from various representations. While diffusion models are adept at this task, they are hindered by latency issues due to their operation at the individual sample point level and the need for numerous sampling steps. In this study, we introduce RFWave, a cutting-edge multi-band Rectified Flow approach designed to reconstruct high-fidelity audio waveforms from Mel-spectrograms or discrete acoustic tokens. RFWave uniquely generates complex spectrograms and operates at the frame level, processing all subbands simultaneously to boost efficiency. Leveraging Rectified Flow, which targets a straight transport trajectory, RFWave achieves reconstruction with just 10 sampling steps. Our empirical evaluations show that RFWave not only provides outstanding reconstruction quality but also offers vastly superior computational efficiency, enabling audio generation at speeds up to 160 times faster than real-time on a GPU. Both an online demonstration and the source code are accessible. Dongyang Dai, Zhiyong Wu 0001 |
ICLR | 3 |
| 2025 | AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech SynthesisabstractWith the advancement of speech synthesis technology, users have higher expectations for the naturalness and expressiveness of synthesized speech. But previous research ignores the importance of prompt selection. This study proposes a text-to-speech (TTS) framework based on Retrieval-Augmented Generation (RAG) technology, which can dynamically adjust the speech style according to the text content to achieve more natural and vivid communication effects. We have constructed a speech style knowledge database containing high-quality speech samples in various contexts and developed a style matching scheme. This scheme uses embeddings, extracted by Llama, PER-LLM-Embedder, and Moka, to match with samples in the knowledge database, selecting the most appropriate speech style for synthesis. Furthermore, our empirical research validates the effectiveness of the proposed method. Our demo can be viewed at: https://thuhcsi.github.io/icme2025-AutoStyle-TTS Chengyuan Ma, Wei Chen 0071, Zhiyong Wu 0001 |
ICME | 6 |
| 2025 | A Multi-Stage Framework for Multimodal Controllable Speech SynthesisabstractControllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with robustness and generalization due to data quality constraints, while text prompt methods offer limited diversity and fine-grained control. Although multimodal approaches aim to integrate various modalities, their reliance on fully matched training data significantly constrains their performance and applicability. This paper proposes a 3-stage multimodal controllable speech synthesis framework to address these challenges. For face encoder, we use supervised learning and knowledge distillation to tackle generalization issues. Furthermore, the text encoder is trained on both text-face and text-speech data to enhance the diversity of the generated speech. Experimental results demonstrate that this method outperforms single-modal baseline methods in both face based and text prompt based speech synthesis, highlighting its effectiveness in generating high-quality speech. Rui Niu, Weihao Wu 0001, Zhiyong Wu 0001 |
ICME | 5 |
| 2025 | UniSep: Universal Target Audio Separation with Language Models at ScaleabstractWe propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep is performed on unlimited source domains and unlimited source numbers. We formulate the separation task as a sequence-to-sequence problem, and a large language model (LLM) is used to model the audio sequence in the discrete latent space, leveraging the power of LLM in handling complex mixture audios with large-scale data. Moreover, a novel pre-training strategy is proposed to utilize audio-only data, which reduces the efforts of large-scale data simulation and enhances the ability of LLMs to understand the consistency and correlation of information within audio sequences. We also demonstrate the effectiveness of scaling datasets in an audio separation task: we use large-scale data (36.5k hours), including speech, music, and sound, to train a universal target audio separation model that is not limited to a specific domain. Experiments show that UniSep achieves competitive subjective and objective evaluation results compared with single-task models. Hangting Chen, Dongchao Yang, Guangzhi Li, Shan Yang 0001, Zhiyong Wu 0001, Helen M. Meng, Xixin Wu |
ICME | 8 |
| 2025 | VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweeningabstractWe propose VideoHumanMIB, a novel framework for Video Human Motion In-betweening that enables seamless transitions between different motion video clips, facilitating the generation of longer and more natural digital human videos. While existing video frame interpolation methods work well for similar motions in adjacent frames, they often struggle with complex human movements, resulting in artifacts and unrealistic transitions. To address these challenges, we introduce a two-stage approach: First, we design an Appearance Reconstruction AutoEncoder to decouple appearance and motion information, extracting robust appearance-invariant features. Second, we develop an enhanced diffusion pretrained network that leverages both motion optical flow and human pose as guidance conditions, enabling the model to learn comprehensive latent distributions of possible motions. Rather than operating directly in pixel space, our model works in a learned latent space, allowing it to better capture the underlying motion dynamics. The framework is optimized with a dual-frame constraint loss and a motion flow loss to ensure temporal consistency and natural movement transitions. Extensive experiments demonstrate that our approach generates highly realistic transition sequences that significantly outperform existing methods, particularly in challenging scenarios with large motion variations. The proposed VideoHumanMIB establishes a new baseline for human motion synthesis and enables more natural and controllable digital human animation. Haiwei Xue, Zhensong Zhang, Minglei Li 0001, Zonghong Dai, F. Richard Yu, Fei Ma 0006, Zhiyong Wu 0001 |
IJCAI | 7 |
| 2025 | DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
Wei Chen 0071, Binzhu Sha, Zhiyong Wu 0001 |
INTERSPEECH | 7 |
| 2025 | DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
Xueyuan Chen, Dongchao Yang, Minglin Wu, Xixin Wu, Zhiyong Wu 0001, Helen M. Meng |
INTERSPEECH | 7 |
| 2025 | StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
Fengjin Li, Yadong Niu, Jian Luan 0001, Zhiyong Wu 0001 |
INTERSPEECH | 7 |
| 2025 | Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
Jingran Xie, Xixin Wu, Zhiyong Wu 0001 |
INTERSPEECH | 7 |
| 2025 | WAKE: Watermarking Audio with Key Enrichment
Yaoxun Xu, Jianwei Yu 0001, Hangting Chen, Zhiyong Wu 0001, Xixin Wu, Dong Yu 0001, Rongzhi Gu, Yi Luo 0004 |
INTERSPEECH | 4 |
| 2025 | MuCodec: Ultra Low-Bitrate Music Codec for Music GenerationabstractMusic generation is pivotal in multimedia, aiding creation and lowering the creative threshold. It focuses on generating music with clear vocals and harmonious accompaniment based on lyrics, combining high artistic creativity with technical challenges. The music codec is an important bridging component in large language model-based music generation, connecting language models with the generated music. However, existing neural codecs typically require token rates exceeding 50 Hz to achieve acceptable music quality, resulting in a context length that surpasses 12,000 tokens for a 4-minute song-a scale that is computationally demanding. This highlights the need for high-compression, high-fidelity music codecs that can reconstruct both vocals and accompaniment with high quality at low frame rates and bitrates, thereby better assisting music generation. To address this, we introduce MuCodec, designed for high-quality music reconstruction at ultra-low bitrates, facilitating more efficient music generation. MuCodec employs a two-stage training method, enabling its encoder, MuEncoder, to extract semantic and acoustic features in a unified representation. These features are discretized using residual vector quantization and converted into Mel-VAE features through flow matching, with reconstruction quality improved by representation alignment during training. The Mel-VAE features are then reconstructed into music using a pretrained Mel-VAE decoder and HiFi-GAN. To the best of our knowledge, MuCodec is the first codec capable of reconstructing 48kHz stereo music at an ultra-low bitrate of 0.35 kbps (25 Hz), achieving state-of-the-art performance in both subjective and objective evaluations, and can more effectively support music generation. Code and Demo: https://mucodec.github.io/Mucodec/. Yaoxun Xu, Hangting Chen, Jianwei Yu 0001, Wei Tan 0011, Shun Lei, Rongzhi Gu, Zhiyong Wu 0001 |
ACM Multimedia | 8 |
| 2025 | HarmoniVox: Painting Voices to Match the Avatar's SoulabstractImagine James Bond speaking like Mr. Bean---such a mismatch would create a jarring dissonance and break the viewer's immersion. Current research on virtual avatar animation has focused on modeling 3D geometry, appearance, motion generation, however, neglecting the harmony between speech prosody and the avatar's visual presentation and contextual environment. In this paper, we seek to bridge this gap by firstly identifying and defining the key elements necessary for achieving audiovisual harmony, such as appearance, expression, body posture, backgrounds and colors. Subsequently, we propose a method that jointly models semantic consistency in avatar animation, named HarmoniVox, specifically on crafting prosodic speech consistent with the avatar's essence from given visual image. To achieve this, we implement a technical framework with a mutual modal contrastive learning strategy, enhancing multimodal alignment in a coarse-to-fine fashion. To support this method, we establish a experimental dataset HarAvaSpeech comprising 28,929 image-audio pairs, designed to encompass expressive speech prosody and rich avatar visual presentations across a wide range of contexts. Leveraging this dataset, our experiments demonstrate that the proposed method outperforms the baselines in manipulating the nuanced tone and harmonious rhythm of speech with the avatar visual presentations, and reveal generalizability on out-of-domain cases. Demo would be provided in https://harmonivox.github.io/harmonivox/. Songtao Zhou, Xiaoyu Qin 0001, Yixuan Zhou 0002, Qixin Wang 0002, Zeyu Jin, Zixuan Wang 0026, Zhiyong Wu 0001, Jia Jia 0001 |
ACM Multimedia | 7 |
| 2025 | LeVo: High-Quality Song Generation with Multi-Preference AlignmentabstractRecent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation.
However, existing approaches still struggle with the complex composition of songs and the scarcity of high-quality data, leading to limitations in audio quality, musicality, instruction following, and vocal-instrument harmony.
To address these challenges, we introduce LeVo, a language model based framework consisting of LeLM and Music Codec.
LeLM is capable of parallel modeling of two types of tokens: mixed tokens, which represent the combined audio of vocals and accompaniment to achieve better vocal-instrument harmony, and dual-track tokens, which separately encode vocals and accompaniment for high-quality song generation.
It employs two decoder-only transformers and a modular extension training strategy to prevent interference between different token types.
To further enhance musicality and instruction following ability, we introduce a multi-preference alignment method based on Direct Preference Optimization (DPO).
This method handles diverse human preferences through a semi-automatic data construction process and post-training.
Experimental results demonstrate that LeVo significantly outperforms existing open-source methods in both objective and subjective metrics, while performing competitively with industry systems.
Ablation studies further justify the effectiveness of our designs.
Audio examples and source code are available at https://levo-demo.github.io and https://github.com/tencent-ailab/songgeneration. Shun Lei, Yaoxun Xu, Huaicheng Zhang, Wei Tan 0011, Hangting Chen, Yixuan Zhang 0005, Haina Zhu, Shuai Wang 0016, Zhiyong Wu 0001, Dong Yu 0001 |
NeurIPS | 11 |
| 2025 | Echo: Enhancing Conversational Behavior Generation via Hierarchical Semantic Comprehension with Large Language ModelsabstractConversational behavior generation, being a crucial capability of embodied agents, is a significant factor influencing human-computer interaction. Generating high-quality conversational motions requires not only appropriate audio-motion mapping but also interactive responses to interlocutor behaviors and comprehensive understanding of conversational semantics. Existing methods primarily rely on audio signals and interlocutor motions for main agent motion generation, lacking high-level semantic understanding of the conversational content, leading to moderate quality motions that are not appropriate for the dialogue. To address these limitations, we leverage the powerful semantic understanding capabilities of large language models, to comprehend complex conversational contexts. Inspired by human conversation processes that conversational motions are highly related to both global and local semantic factors, including the conversational context, and the intentions, emotions, and passive or active states of the participants, we propose an agentic system named Echo that analyzes such information. To achieve comprehensive conversational understanding, Echo leverages multiple prompts and test-time recipes to guide large language models in decomposing conversational structures and extracting fine-grained semantic information. Furthermore, we design a hierarchical feature fusion network that systematically integrates from frame-level audio-motion features to sentence-level semantic understanding and finally to conversation-level contextual comprehension, organically combining fine-grained semantic features from large language models with audio and motion characteristics. Experimental results demonstrate that our framework can be effectively integrated with several state-of-the-art motion generation models to enhance their performance in generating high-quality conversational behaviors. Haiwei Xue, Yanbo Fan, Xuan Wang 0009, Zhiyong Wu 0001 |
SIGGRAPH Asia | 4 |
| 2025 | Human Motion Video Generation: A SurveyabstractHuman motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys in this field focus on individual methods, lacking a comprehensive overview of the entire generative process. This paper addresses this gap by providing an in-depth survey of human motion video generation, encompassing over ten sub-tasks, and detailing the five key phases of the generation process: input, motion planning, motion video generation, refinement, and output. Notably, this is the first survey that discusses the potential of large language models in enhancing human motion video generation. Our survey reviews the latest developments and technological trends in human motion video generation across three primary modalities: vision, text, and audio. By covering over two hundred papers, we offer a thorough overview of the field and highlight milestone works that have driven significant technological breakthroughs. Our goal for this survey is to unveil the prospects of human motion video generation and serve as a valuable resource for advancing the comprehensive applications of digital humans. Haiwei Xue, Xiangyang Luo 0002, Zhanghao Hu, Xin Zhang 0169, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li 0001, Jian Yang 0003, Fei Ma 0006, Zhiyong Wu 0001, Changpeng Yang, Zonghong Dai, F. Richard Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2025 | AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial AnimationabstractSpeech-driven 3D facial animation aims at generating facial movements that are synchronized with the driving speech, which has been widely explored recently. Existing works mostly neglect the person-specific talking style in generation, including facial expression and head pose styles. Several works intend to capture the personalities by fine-tuning modules. However, limited training data leads to the lack of vividness. In this work, we proposeAdaMesh, a novel adaptive speech-driven facial animation approach, which learns the personalized talking style from a reference video of about 10 seconds and generates vivid facial expressions and head poses. Specifically, we propose mixture-of-low-rank adaptation (MoLoRA) to fine-tune the expression adapter, which efficiently captures the facial expression style. For the personalized pose style, we propose a pose adapter by building a discrete pose prior and retrieving the appropriate style embedding with a semantic-aware pose style matrix without fine-tuning. Extensive experimental results show that our approach outperforms state-of-the-art methods, preserves the talking style in the reference video, and generates vivid facial animation. Liyang Chen, Weihong Bao, Shun Lei, Boshi Tang, Zhiyong Wu 0001, Shiyin Kang, Hao-Zhi Huang 0001, Helen M. Meng |
IEEE Trans. Multim. | 5 |
| 2024 | SimCalib: Graph Neural Network Calibration Based on Similarity between NodesabstractGraph neural networks (GNNs) have exhibited impressive performance in modeling graph data as exemplified in various applications. Recently, the GNN calibration problem has attracted increasing attention, especially in cost-sensitive scenarios. Previous work has gained empirical insights on the issue, and devised effective approaches for it, but theoretical supports still fall short. In this work, we shed light on the relationship between GNN calibration and nodewise similarity via theoretical analysis. A novel calibration framework, named SimCalib, is accordingly proposed to consider similarity between nodes at global and local levels. At the global level, the Mahalanobis distance between the current node and class prototypes is integrated to implicitly consider similarity between the current node and all nodes in the same class. At the local level, the similarity of node representation movement dynamics, quantified by nodewise homophily and relative degree, is considered. Informed about the application of nodewise movement patterns in analyzing nodewise behavior on the over-smoothing problem, we empirically present a possible relationship between over-smoothing and GNN calibration problem. Experimentally, we discover a correlation between nodewise similarity and model calibration improvement, in alignment with our theoretical results. Additionally, we conduct extensive experiments investigating different design factors and demonstrate the effectiveness of our proposed SimCalib framework for GNN calibration by achieving state-of-the-art performance on 14 out of 16 benchmarks. Boshi Tang, Zhiyong Wu 0001, Xixin Wu, Qiaochu Huang, Jun Chen 0024, Shun Lei, Helen M. Meng |
AAAI | 2 |
| 2024 | Explore 3D Dance Generation via Reward Model from Automatically-Ranked DemonstrationsabstractThis paper presents an Exploratory 3D Dance generation framework, E3D2, designed to address the exploration capability deficiency in existing music-conditioned 3D dance generation models. Current models often generate monotonous and simplistic dance sequences that misalign with human preferences because they lack exploration capabilities.The E3D2 framework involves a reward model trained from automatically-ranked dance demonstrations, which then guides the reinforcement learning process. This approach encourages the agent to explore and generate high quality and diverse dance movement sequences. The soundness of the reward model is both theoretically and experimentally validated. Empirical experiments demonstrate the effectiveness of E3D2 on the AIST++ dataset. Zilin Wang 0002, Haolin Zhuang, Yinmin Zhang, Junjie Zhong, Jun Chen 0024, Yu Yang 0016, Boshi Tang, Zhiyong Wu 0001 |
AAAI | 9 |
| 2024 | SECap: Speech Emotion Captioning with Large Language ModelabstractSpeech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human speech are often complex, and categorizing them into predefined groups can be insufficient to adequately represent speech emotions. On the contrary, describing speech emotions directly by means of natural language may be a more effective approach. Regrettably, there are not many studies available that have focused on this direction. Therefore, this paper proposes a speech emotion captioning framework named SECap, aiming at effectively describing speech emotions using natural language. Owing to the impressive capabilities of large language models in language comprehension and text generation, SECap employs LLaMA as the text decoder to allow the production of coherent speech emotion captions. In addition, SECap leverages HuBERT as the audio encoder to extract general speech features and Q-Former as the Bridge-Net to provide LLaMA with emotion-related speech features. To accomplish this, Q-Former utilizes mutual information learning to disentangle emotion-related speech features and speech contents, while implementing contrastive learning to extract more emotion-related speech features. The results of objective and subjective evaluations demonstrate that: 1) the SECap framework outperforms the HTSAT-BART baseline in all objective evaluations; 2) SECap can generate high-quality speech emotion captions that attain performance on par with human annotators in subjective mean opinion score tests. Yaoxun Xu, Hangting Chen, Jianwei Yu 0001, Qiaochu Huang, Zhiyong Wu 0001, Shixiong Zhang 0001, Guangzhi Li, Yi Luo 0004, Rongzhi Gu |
AAAI | 5 |
| 2024 | Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion ModelabstractCo-speech gestures, if presented in the lively form of videos, can achieve superior visual effects in human-machine interaction. While previous works mostly gener-ate structural human skeletons, resulting in the omission of appearance information, we focus on the direct gener-ation of audio-driven co-speech gesture videos in this work. There are two main challenges: 1) A suitable motion feature is needed to describe complex human movements with crucial appearance information. 2) Gestures and speech exhibit inherent dependencies and should be temporally aligned even of arbitrary length. To solve these problems, we present a novel motion-decoupled framework to gener-ate co-speech gesture videos. Specifically, we first intro-duce a well-designed nonlinear TPS transformation to ob-tain latent motion features preserving essential appearance information. Then a transformer-based diffusion model is proposed to learn the temporal correlation between gestures and speech, and performs generation in the latent motion space, followed by an optimal motion selection mod-ule to produce long-term coherent and consistent gesture videos. For better visual perception, we further design a refinement network focusing on missing details of cer-tain areas. Extensive experimental results show that our proposed framework significantly outperforms existing approaches in both motion and video-related evaluations. Our code, demos, and more resources are available at https://github.com/thuhcsi/S2G-MDDiffusion. Qiaochu Huang, Zhensong Zhang, Zhiyong Wu 0001, Minglei Li 0001, Songcen Xu |
CVPR | 5 |
| 2024 | Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech ReconstructionabstractDysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in complex, noisy acoustic environments. To address these challenges, we propose a novel multi-modal framework to utilize visual information, e.g., lip movements, in DSR as extra clues for reconstructing the highly abnormal pronunciations. The multi-modal framework consists of: (i) a multi-modal encoder to extract robust phoneme embeddings from dysarthric speech with auxiliary visual features; (ii) a variance adaptor to infer the normal phoneme duration and pitch contour from the extracted phoneme embeddings; (iii) a speaker encoder to encode the speaker’s voice characteristics; and (iv) a mel-decoder to generate the reconstructed mel-spectrogram based on the extracted phoneme embeddings, prosodic features and speaker embeddings. Both objective and subjective evaluations conducted on the commonly used UASpeech corpus show that our proposed approach can achieve significant improvements over baseline systems in terms of speech intelligibility and naturalness, especially for the speakers with more severe symptoms. Compared with original dysarthric speech, the reconstructed speech achieves 42.1% absolute word error rate reduction for patients with more severe dysarthria levels.1 Xueyuan Chen, Yuejiao Wang, Xixin Wu, Disong Wang, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 5 |
| 2024 | Stylespeech: Self-Supervised Style Enhancing with VQ-VAE-Based Pre-Training for Expressive Audiobook Speech SynthesisabstractThe expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these issues, in this paper, we propose a self-supervised style enhancing method with VQ-VAE-based pre-training for expressive audiobook speech synthesis. Firstly, a text style encoder is pre-trained with a large amount of unlabeled text-only data. Secondly, a spectrogram style extractor based on VQ-VAE is pre-trained in a self-supervised manner, with plenty of audio data that covers complex style variations. Then a novel architecture with two encoder-decoder paths is specially designed to model the pronunciation and high-level style expressiveness respectively, with the guidance of the style extractor. Both objective and subjective evaluations demonstrate that our proposed method can effectively improve the naturalness and expressiveness of the synthesized speech in audiobook synthesis especially for the role and out-of-domain scenarios.1 Xueyuan Chen, Xi Wang 0016, Shaofei Zhang, Lei He 0005, Zhiyong Wu 0001, Xixin Wu, Helen M. Meng |
ICASSP | 5 |
| 2024 | Enhancing Expressiveness in Dance Generation Via Integrating Frequency and Music Style InformationabstractDance generation, as a branch of human motion generation, has attracted increasing attention. Recently, a few works attempt to enhance dance expressiveness, which includes genre matching, beat alignment, and dance dynamics, from certain aspects. However, the enhancement is quite limited as they lack comprehensive consideration of the aforementioned three factors. In this paper, we propose ExpressiveBailando, a novel dance generation method designed to generate expressive dances, concurrently taking all three factors into account. Specifically, we mitigate the issue of speed homogenization by incorporating frequency information into VQ-VAE, thus improving dance dynamics. Additionally, we integrate music style information by extracting genre- and beat-related features with a pre-trained music model, hence achieving improvements in the other two factors. Extensive experimental results demonstrate that our proposed method can generate dances with high expressiveness and outperforms existing methods both qualitatively and quantitatively1. Qiaochu Huang, Boshi Tang, Haolin Zhuang, Liyang Chen, Shuochen Gao, Zhiyong Wu 0001, Haozhi Huang 0004, Helen M. Meng |
ICASSP | 7 |
| 2024 | Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic PromptsabstractZero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker’s voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language model-based TTS models show zero-shot speaker adaptation capabilities with only a 3-second acoustic prompt of an unseen speaker. However, they are limited by the length of the acoustic prompt, which makes it difficult to clone personal speaking style. In this paper, we propose a novel zero-shot TTS model with the multi-scale acoustic prompts based on a language model. A speaker-aware text encoder is proposed to learn the personal speaking style at the phoneme-level from the style prompt consisting of multiple sentences. Following that, a VALL-E based acoustic decoder is utilized to model the timbre from the timbre prompt at the frame-level and generate speech. The experimental results show that our proposed method outperforms baselines in terms of naturalness and speaker similarity, and can achieve better performance by scaling out to a longer style prompt1. Shun Lei, Yixuan Zhou 0002, Liyang Chen, Zhiyong Wu 0001, Xixin Wu, Shiyin Kang, Yahui Zhou, Yuxing Han 0001, Helen M. Meng |
ICASSP | 5 |
| 2024 | Generating Stereophonic Music with Single-Stage Language ModelsabstractThe recent success of audio language models (LMs) has revolutionized the field of neural music generation. Among all audio LM approaches, MusicGen has demonstrated the success of a single-stage LMs based music generation framework, without needing to train multiple LMs. Despite its promising performance in generating monophonic (mono) music, directly generating stereophonic (stereo) music following the previous framework has resulted in perceptible quality degradation. In this paper, we first discuss the difficulty of directly encoding stereo music with neural codec, and then provide a stable and practical solution based on a dual encoding approach. To utilize the dually encoded tokens in single-stage LMs, we also propose two forms of token sequence patterns. An extensive evaluation has been conducted using various aspects of stereo music audios to examine the performance of stereo neural codec approaches and the generation quality of single-stage LMs. Finally, our experimental results suggest that (i) our proposed dual encoding approach for neural codec is significantly better than the typical joint encoding approach in terms of reconstruction quality, and (ii) the stereo single-stage LMs trained with our proposed token sequence patterns substantially improved the perceptual quality of the state-of-the-art music generation model (i.e. MusicGen) in subjective tests. Xingda Li, Fan Zhuo, Jun Chen 0024, Shiyin Kang, Zhiyong Wu 0001, Yahui Zhou |
ICASSP | 6 |
| 2024 | Multi-View Midivae: Fusing Track- and Bar-View Representations for Long Multi-Track Symbolic Music GenerationabstractVariational Autoencoders (VAEs) constitute a crucial component of neural symbolic music generation, among which some works have yielded outstanding results and attracted considerable attention. Nevertheless, previous VAEs still encounter issues with overly long feature sequences and generated results lack contextual coherence, thus the challenge of modeling long multi-track symbolic music still remains unaddressed. To this end, we propose Multi-view MidiVAE, as one of the pioneers in VAE methods that effectively model and generate long multi-track symbolic music. The Multi-view MidiVAE utilizes the two-dimensional (2-D) representation, OctupleMIDI, to capture relationships among notes while reducing the feature sequences length. Moreover, we focus on instrumental characteristics and harmony as well as global and local information about the musical composition by employing a hybrid variational encoding-decoding strategy to integrate both Track- and Bar-view MidiVAE features. Objective and subjective experimental results on the CocoChorales dataset demonstrate that, compared to the baseline, Multi-view MidiVAE exhibits significant improvements in terms of modeling long multi-track symbolic music. Jun Chen 0024, Boshi Tang, Binzhu Sha, Yaolong Ju, Shiyin Kang, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 9 |
| 2024 | Unifying One-Shot Voice Conversion and Cloning with Disentangled Speech RepresentationsabstractWe propose unifying one-shot voice conversion and cloning into a single model that can be end-to-end optimized. To achieve this, we introduce a novel extension to a speech variational auto-encoder (VAE) that disentangles speech into content and speaker representations. Instead of using a fixed Gaussian prior as in the vanilla VAE, we incorporate a learnable text-aware prior as an informative guide for learning the content representation. This results in a content representation with reduced speaker information and more accurate linguistic information. The proposed model can sample the content representation using either the posterior conditioned on speech or the text-aware prior with textual input, enabling one-shot voice conversion and cloning, respectively. Experiments show that the proposed method achieves better or comparable overall performance for one-shot voice conversion and cloning compared to state-of-the-art voice conversion and cloning methods. Xixin Wu, Haohan Guo, Songxiang Liu, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 5 |
| 2024 | Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice ConversionabstractAny-to-any singing voice conversion (SVC) is confronted with the challenge of "timbre leakage" issue caused by inadequate disentanglement between the content and the speaker timbre. To address this issue, this study introduces NeuCoSVC, a novel neural concatenative SVC framework. It consists of a self-supervised learning (SSL) representation extractor, a neural harmonic signal generator, and a waveform synthesizer. The SSL extractor condenses audio into fixed-dimensional SSL features, while the harmonic signal generator leverages linear time-varying filters to produce both raw and filtered harmonic signals for pitch information. The synthesizer reconstructs waveforms using SSL features, harmonic signals, and loudness information. During inference, voice conversion is performed by substituting source SSL features with their nearest counterparts from a matching pool which comprises SSL features extracted from the reference audio, while preserving raw harmonic signals and loudness from the source audio. By directly utilizing SSL features from the reference audio, the proposed framework effectively resolves the "timbre leakage" issue caused by previous disentanglement-based approaches. Experimental results demonstrate that the proposed NeuCoSVC system outperforms the disentanglement-based speaker embedding approach in one-shot SVC across intra-language, cross-language, and cross-domain evaluations. Binzhu Sha, Xu Li 0015, Zhiyong Wu 0001, Ying Shan, Helen M. Meng |
ICASSP | 3 |
| 2024 | SCNet: Sparse Compression Network for Music Source SeparationabstractDeep learning-based methods have made significant achievements in music source separation. However, obtaining good results while maintaining a low model complexity remains challenging in super wide-band music source separation. Previous works either overlook the differences in subbands or inadequately address the problem of information loss when generating subband features. In this paper, we propose SCNet, a novel frequency-domain network to explicitly split the spectrogram of the mixture into several subbands and introduce a sparsity-based encoder to model different frequency bands. We use a higher compression ratio on subbands with less information to improve the information density and focus on modeling subbands with more information. In this way, the separation performance can be significantly improved using lower computational consumption. Experiment results show that the proposed model achieves a signal to distortion ratio (SDR) of 9.0 dB on the MUSDB18-HQ dataset without using extra data, which outperforms state-of-the-art methods. Specifically, SCNet’s CPU inference time is only 48% of HT Demucs, one of the previous state-of-the-art models. Weinan Tong, Jiaxu Zhu, Jun Chen 0024, Shiyin Kang, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 7 |
| 2024 | Consistent and Relevant: Rethink the Query Embedding in General Sound SeparationabstractThe query-based audio separation usually employs specific queries to extract target sources from a mixture of audio signals. Currently, most query-based separation models need additional networks to obtain query embedding. In this way, separation model is optimized to be adapted to the distribution of query embedding. However, query embedding may exhibit mismatches with separation models due to inconsistent structures and independent information. In this paper, we present CaRE-SEP, a consistent and relevant embedding network for general sound separation to encourage a comprehensive reconsideration of query usage in audio separation. CaRE-SEP alleviates the potential mismatch between queries and separation in two aspects, including sharing network structure and sharing feature information. First, a Swin-Unet model with a shared encoder is conducted to unify query encoding and sound separation into one model, eliminating the network architecture difference and generating consistent distribution of query and separation features. Second, by initializing CaRE-SEP with a pretrained classification network and allowing gradient backpropagation, the query embedding is optimized to be relevant to the separation feature, further alleviating the feature mismatch problem. Experimental results indicate the proposed CaRE-SEP model substantially improves the performance of separation tasks. Moreover, visualizations validate the potential mismatch and how CaRE-SEP solves it. Hangting Chen, Dongchao Yang, Jianwei Yu 0001, Chao Weng, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 6 |
| 2024 | Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion ModelsabstractAudio-driven co-speech human gesture generation has made remarkable advancements recently. However, most previous works only focus on single person audio-driven gesture generation. We aim at solving the problem of conversational co-speech gesture generation that considers multiple participants in a conversation, which is a novel and challenging task due to the difficulty of simultaneously incorporating semantic information and other relevant features from both the primary speaker and the interlocutor. To this end, we propose CoDiffuseGesture, a diffusion model-based approach for speech-driven interaction gesture generation via modeling bilateral conversational intention, emotion, and semantic context. Our method synthesizes appropriate interactive, speech-matched, high-quality gestures for conversational motions through the intention perception module and emotion reasoning module at the sentence level by a pretrained language model. Experimental results demonstrate the promising performance of the proposed method. Haiwei Xue, Zhensong Zhang, Zhiyong Wu 0001, Minglei Li 0001, Zonghong Dai, Helen M. Meng |
ICASSP | 4 |
| 2024 | FreeTalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker NaturalnessabstractCurrent talking avatars mostly generate co-speech gestures based on audio and text of the utterance, without considering the non-speaking motion of the speaker. Furthermore, previous works on co-speech gesture generation have designed network structures based on individual gesture datasets, which results in limited data volume, compromised generalizability, and restricted speaker movements. To tackle these issues, we introduce FreeTalker, which, to the best of our knowledge, is the first framework for the generation of both spontaneous (e.g., co-speech gesture) and non-spontaneous (e.g., moving around the podium) speaker motions. Specifically, we train a diffusion-based model for speaker motion generation that employs unified representations of both speech-driven gestures and text-driven motions, utilizing heterogeneous data sourced from various motion datasets. During inference, we utilize classifier-free guidance to highly control the style in the clips. Additionally, to create smooth transitions between clips, we utilize DoubleTake, a method that leverages a generative prior and ensures seamless motion blending. Extensive experiments show that our method generates natural and controllable speaker movements. Our code, model, and demo are are available at https://youngseng.github.io/FreeTalker/. Zunnan Xu, Haiwei Xue, Yongkang Cheng, Shaoli Huang, Mingming Gong, Zhiyong Wu 0001 |
ICASSP | 7 |
| 2024 | Hydraformer: One Encoder for All Subsampling RatesabstractIn automatic speech recognition, subsampling is essential for tackling diverse scenarios. However, the inadequacy of a single subsampling rate to address various real-world situations often necessitates training and deploying multiple models, consequently increasing associated costs. To address this issue, we propose HydraFormer, comprising HydraSub, a Conformer-based encoder, and a BiTransformer-based decoder. HydraSub encompasses multiple branches, each representing a distinct subsampling rate, allowing for the flexible selection of any branch during inference based on the specific use case. HydraFormer can efficiently manage different subsampling rates, significantly reducing training and deployment expenses. Experiments on AISHELL-1 and LibriSpeech datasets reveal that HydraFormer effectively adapts to various subsampling rates and languages while maintaining high recognition performance. Additionally, HydraFormer showcases exceptional stability, sustaining consistent performance under various initialization conditions, and exhibits robust transferability by learning from pretrained single subsampling rate automatic speech recognition models1. Yaoxun Xu, Xingchen Song, Zhiyong Wu 0001, Di Wu 0061, Zhendong Peng |
ICME | 3 |
| 2024 | NRAdapt: Noise-Robust Adaptive Text to Speech Using Untranscribed DataabstractVoice cloning, also known as personalized voice synthesis, is a significant branch of speech synthesis. It involves synthesizing speech with the same vocal characteristics as the target speaker. The field of voice cloning mainly faces two challenges: 1. Typically, only a small amount of voice data from the target speaker is available, necessitating the cloning of the speaker’s timbre using limited available data; 2. The voice data of the target speaker is often recorded with non-professional equipment in noisy environments, posing a significant challenge to the robustness of voice cloning systems. In this paper, we propose a noise-robust voice adaptation method based on speaker-independent bottleneck features.The method contains: 1. Using a disentanglement module based on autoencoder architecture to disentangle the timbre information from the speech data, resulting in the extraction of corresponding speaker-independent bottleneck features with environmental information; 2. With the help of the above module, the Text2BN (Text to Bottleneck) module is trained with high-quality voice data to establish a mapping from text to clean speaker-independent bottleneck features; 3. The decoder is fine-tuned using noisy target speaker speech data to adapt to the target speaker and is cascaded with the Text2BN module to synthesize clean audio. The disentanglement module does not require text transcriptions and not mandate the use of artificially paired clean/noisy datasets, enabling large-scale pre-training with massive real-world, untranscribed, noisy datasets to further enhance the model’s noise robustness and capabilities of decoupling timbre. During the cloning phase, there are no requirements for the recording conditions of the target speaker’s speech data or for text transcriptions, aligning more closely with real application scenarios. Since our model does not directly work on the noise level, it effectively avoids issues of insufficient robustness to out-of-distribution noise. Experiments demonstrate that the method can effectively utilize noisy target speaker speech data for voice cloning, achieving the preponderance of both speech quality and similarity in the synthesized speech. Shun Lei, Dongyang Dai, Zhiyong Wu 0001, Dading Chong |
IJCNN | 4 |
| 2024 | Representation Space Maintenance: Against Forgetting in Continual LearningabstractThe main purpose of Continual Learning is to prevent the catastrophic forgetting. A number of the recent methods utilize a replay buffer to keep the knowledge of previous tasks. However, the data from previous tasks may be limited. In this setting, the model face a challenge to strike a balance between plasticity and stability. If the model becomes overfit to the current task data, its performance will suffer when faced with new tasks. Conversely, if the model fails to retain the knowledge it has already learned, the new features will cover the old feature, resulting in catastrophic forgetting. To address this issue, we propose a novel method for continual learning, called RSM, which aims to maintain the features of old tasks and allocate new features while learning new tasks. RSM utilizes a CVAE to mimic the previous representation space. To maintain previous features while learning new tasks, we incorporate both real data and pseudo-features generated by CVAE in the representation learning and distillation process. Experiments conducted on benchmark datasets show that our method has demonstrated superior performance compared to replay-free and replay-based methods. Rui Niu, Zhiyong Wu 0001, Changhe Song |
IJCNN | 2 |
| 2024 | LoRA-MER: Low-Rank Adaptation of Pre-Trained Speech Models for Multimodal Emotion Recognition Using Mutual Information
Yunrui Cai, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng |
INTERSPEECH | 2 |
| 2024 | CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
Xueyuan Chen, Dongchao Yang, Dingdong Wang, Xixin Wu, Zhiyong Wu 0001, Helen M. Meng |
INTERSPEECH | 5 |
| 2024 | Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
Peiji Yang, Yicheng Zhong, Yixuan Zhou 0002, Zhisheng Wang 0001, Zhiyong Wu 0001, Xixin Wu, Helen M. Meng |
INTERSPEECH | 6 |
| 2024 | Speaker Change Detection with Weighted-sum Knowledge Distillation based on Self-supervised Pre-trained Models
Yuxiang Kong, Lichun Fan, Peng Gao 0013, Zhiyong Wu 0001 |
INTERSPEECH | 6 |
| 2024 | Comparing Discrete and Continuous Space LLMs for Speech Recognition
Yaoxun Xu, Shixiong Zhang 0001, Jianwei Yu 0001, Zhiyong Wu 0001, Dong Yu 0001 |
INTERSPEECH | 4 |
| 2024 | VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language ModellingabstractRecent AIGC systems possess the capability to generate digital multimedia content based on human language instructions, such as text, image and video. However, when it comes to speech, existing methods related to human instruction-to-speech generation exhibit two limitations. Firstly, they require the division of inputs into content prompt (transcript) and description prompt (style and speaker), instead of directly supporting human instruction. This division is less natural in form and does not align with other AIGC models. Secondly, the practice of utilizing an independent description prompt to model speech style, without considering the transcript content, restricts the ability to control speech at a fine-grained level. To address these limitations, we propose VoxInstruct, a novel unified multilingual codec language modeling framework that extends traditional text-to-speech tasks into a general human instruction-to-speech task. Our approach enhances the expressiveness of human instruction-guided speech generation and aligns the speech generation paradigm with other modalities. To enable the model to automatically extract the content of synthesized speech from raw text instructions, we introduce speech semantic tokens as an intermediate representation for instruction-to-content guidance. We also incorporate multiple Classifier-Free Guidance (CFG) strategies into our codec language model, which strengthens the generated speech following human instructions. Furthermore, our model architecture and training strategies allow for the simultaneous support of combining speech prompt and descriptive human instruction for expressive speech synthesis, which is a first-of-its-kind attempt. Codes, models and demos are at: https://github.com/thuhcsi/VoxInstruct. Yixuan Zhou 0002, Xiaoyu Qin 0001, Zeyu Jin, Shuoyi Zhou, Shun Lei, Songtao Zhou, Zhiyong Wu 0001, Jia Jia 0001 |
ACM Multimedia | 7 |
| 2024 | SpeechCraft: A Fine-Grained Expressive Speech Dataset with Natural Language DescriptionabstractSpeech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elaborate comprehension of speech style is urgently needed to facilitate insightful interplay between speech audio and natural language. However, constructing such datasets presents a major trade-off between large-scale data collection and high-quality annotation. To tackle this challenge, we propose an automatic speech annotation system for expressiveness interpretation that annotates in-the-wild speech clips with expressive and vivid human language descriptions. Initially, speech audios are processed by a series of expert classifiers and captioning models to capture diverse speech characteristics, followed by a fine-tuned LLaMA for customized annotation generation. Unlike previous tag/templet-based annotation frameworks with limited information and diversity, our system provides in-depth understandings of speech style through tailored natural language descriptions, thereby enabling accurate and voluminous data generation for large model training. With this system, we create SpeechCraft, a fine-grained bilingual expressive speech dataset. It is distinguished by highly descriptive natural language style prompts, containing approximately 2,000 hours of audio data and encompassing over two million speech clips. Extensive experiments demonstrate that the proposed dataset significantly boosts speech-language task performance in stylist speech synthesis and speech style understanding. Zeyu Jin, Jia Jia 0001, Qixin Wang 0002, Kehan Li 0007, Shuoyi Zhou, Songtao Zhou, Xiaoyu Qin 0001, Zhiyong Wu 0001 |
ACM Multimedia | 8 |
| 2024 | SongCreator: Lyrics-based Universal Song GenerationabstractMusic is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as singing voice, vocal composition and instrumental arrangement, etc., generating songs with both vocals and accompaniment given lyrics remains a significant challenge, hindering the application of music generation models in the real world. In this light, we propose SongCreator, a song-generation system designed to tackle this challenge. The model features two novel designs: a meticulously designed dual-sequence language model (DSLM) to capture the information of vocals and accompaniment for song generation, and a series of attention mask strategies for DSLM, which allows our model to understand, generate and edit songs, making it suitable for various songrelated generation tasks by utilizing specific attention masks. Extensive experiments demonstrate the effectiveness of SongCreator by achieving state-of-the-art or competitive performances on all eight tasks. Notably, it surpasses previous works by a large margin in lyrics-to-song and lyrics-to-vocals. Additionally, it is able to independently control the acoustic conditions of the vocals and accompaniment in the generated song through different audio prompts, exhibiting its potential applicability. Our samples are available at https://thuhcsi.github.io/SongCreator/. Shun Lei, Yixuan Zhou 0002, Boshi Tang, Max W. Y. Lam, Jingcheng Wu, Shiyin Kang, Zhiyong Wu 0001, Helen M. Meng |
NeurIPS | 9 |
| 2024 | Joint Multiscale Cross-Lingual Speaking Style Transfer With Bidirectional Attention Mechanism for Automatic DubbingabstractAutomatic dubbing, which generates a corresponding version of the input speech in another language, can be widely utilized in many real-world scenarios, such as video and game localization. In addition to synthesizing the translated scripts, automatic dubbing further transfers the speaking style in the original language to the dubbed speeches to give audiences the impression that the characters are speaking in their native tongue. However, state-of-the-art automatic dubbing systems only model the transfer on the duration and speaking rate, disregarding the other aspects of speaking style, such as emotion, intonation and emphasis, which are also crucial to fully understand the characters and speech. In this paper, we propose a joint multiscale cross-lingual speaking style transfer framework to simultaneously model the bidirectional speaking style transfer between two languages at both the global scale (i.e., utterance level) and local scale (i.e., word level). The global and local speaking styles in each language are extracted and utilized to predict the global and local speaking styles in the other language with an encoder-decoder framework for each direction and a shared bidirectional attention mechanism for both directions. A multiscale speaking style-enhanced FastSpeech 2 is then utilized to synthesize the desired speech with the predicted global and local speaking styles for each language. The experimental results demonstrate the effectiveness of our proposed framework, which outperforms a baseline with only duration transfer in objective and subjective evaluations. Jingbei Li, Sipan Li, Zhiyong Wu 0001, Helen M. Meng, Qiao Tian 0001, Yuping Wang 0005, Yuxuan Wang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | What Does Your Face Sound Like? 3D Face Shape towards VoiceabstractFace-based speech synthesis provides a practical solution to generate voices from human faces. However, directly using 2D face images leads to the problems of uninterpretability and entanglement. In this paper, to address the issues, we introduce 3D face shape which (1) has an anatomical relationship between voice characteristics, partaking in the "bone conduction" of human timbre production, and (2) is naturally independent of irrelevant factors by excluding the blending process. We devise a three-stage framework to generate speech from 3D face shapes. Fully considering timbre production in anatomical and acquired terms, our framework incorporates three additional relevant attributes including face texture, facial features, and demographics. Experiments and subjective tests demonstrate our method can generate utterances matching faces well, with good audio quality and voice diversity. We also explore and visualize how the voice changes with the face. Case studies show that our method upgrades the face-voice inference to personalized custom-made voice creating, revealing a promising prospect in virtual human and dubbing applications. Zhiyong Wu 0001, Ying Shan, Jia Jia 0001 |
AAAI | 2 |
| 2023 | QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture GenerationabstractSpeech-driven gesture generation is highly challenging due to the random jitters of human motion. In addition, there is an inherent asynchronous relationship between human speech and gestures. To tackle these challenges, we introduce a novel quantization-based and phase-guided motion matching framework. Specifically, we first present a gesture VQ-VAE module to learn a codebook to summarize meaningful gesture units. With each code representing a unique gesture, random jittering problems are alleviated effectively. We then use Levenshtein distance to align diverse gestures with different speech. Levenshtein distance based on audio quantization as a similarity metric of corresponding speech of gestures helps match more appropriate gestures with speech, and solves the alignment problem of speech and gestures well. Moreover, we introduce phase to guide the optimal gesture matching based on the semantics of context or rhythm of audio. Phase guides when text-based or speech-based gestures should be performed to make the generated gestures more natural. Extensive experiments show that our method outperforms recent approaches on speech-driven gesture generation. Our code, database, pre-trained models and demos are available at https://github.com/YoungSeng/QPGesture. Zhiyong Wu 0001, Minglei Li 0001, Zhensong Zhang, Weihong Bao, Haolin Zhuang |
CVPR | 2 |
| 2023 | Wavsyncswap: End-To-End Portrait-Customized Audio-Driven Talking Face GenerationabstractAudio-driven talking face with portrait customization enhances the flexibility of avatar applications for different scenarios, such as on-line meetings, mixed reality, and data generation. Among the existing methods, audio-driven talking face and face swapping are typically viewed as separate tasks that are cascaded to achieve the objective. Using state-of-the-art methods Wav2Lip and SimSwap for this purpose, we meet some issues: affected mouth synchronization, lost texture information, and slow inference speed. To resolve these issues, we propose an end-to-end model that combines the advantages of both approaches. Our approach generates highly-synchronized mouth with the aid of a pre-trained lip-sync discriminator. And identity information is provided by ArcFace and the ID injection module in the model because of its strong correlation with facial texture. Experimental results demonstrate that our method achieves lip-sync accuracy comparable to real synced videos, preserves more texture details than cascade methods, and alleviates the blurring of Wav2Lip. Also, our approach improves the inference speed.1 Weihong Bao, Liyang Chen, Chaoyong Zhou, Zhiyong Wu 0001 |
ICASSP | 5 |
| 2023 | Inter-Subnet: Speech Enhancement with Subband InteractionabstractSubband-based approaches process subbands in parallel through the model with shared parameters to learn the commonality of local spectrums for noise reduction. In this way, they have achieved remarkable results with fewer parameters. However, in some complex environments, the lack of global spectral information has a negative impact on the performance of these subband-based approaches. To this end, this paper introduces the subband interaction as a new way to complement the subband model with the global spectral information such as cross-band dependencies and global spectral patterns, and proposes a new lightweight single-channel speech enhancement framework called Interactive Subband Network (Inter-SubNet). Experimental results on DNS Challenge - Interspeech 2021 dataset show that the proposed Inter-SubNet yields a significant improvement over the subband model and outperforms other state-of-the-art speech enhancement approaches, which demonstrate the effectiveness of subband interaction. Jun Chen 0024, Wei Rao 0002, Zilin Wang 0002, Jiuxin Lin, Zhiyong Wu 0001, Yannan Wang, Shidong Shang, Helen M. Meng |
ICASSP | 5 |
| 2023 | Gesper: A Unified Framework for General Speech RestorationabstractThis paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by speech enhancement. We propose a complex spectral mapping-based generative adversarial network (CSM-GAN) as the speech restoration module for the first time. For noise suppression and dereverberation, the enhancement module is presented with fullband-wideband parallel processing. On the blind test set of ICASSP 2023 SSI Challenge, the proposed Gesper system, which satisfies the real-time condition, achieves 3.27 P.804 overall mean opinion score (MOS) and 3.35 P.835 overall MOS, ranked 1st in both track 1 and track 2. Jun Chen 0024, Yupeng Shi, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001, Shidong Shang, Chengshi Zheng |
ICASSP | 8 |
| 2023 | LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-SpeechabstractRecent advances in neural text-to-speech (TTS) models bring thousands of TTS applications into daily life, where models are deployed in cloud to provide services for customs. Among these models are diffusion probabilistic models (DPMs), which can be stably trained and are more parameter-efficient compared with other generative models. As transmitting data between customs and the cloud introduces high latency and the risk of exposing private data, deploying TTS models on edge devices is preferred. When implementing DPMs onto edge devices, there are two practical problems. First, current DPMs are not lightweight enough for resource-constrained devices. Second, DPMs require many denoising steps in inference, which increases latency. In this work, we present LightGrad, a lightweight DPM for TTS. LightGrad is equipped with a lightweight U-Net diffusion decoder and a training-free fast sampling technique, reducing both model parameters and inference latency. Streaming inference is also implemented in LightGrad to reduce latency further. Compared with Grad-TTS, LightGrad achieves 62.2% reduction in paramters, 65.7% reduction in latency, while preserving comparable speech quality on both Chinese Mandarin and English in 4 denoising steps1. Xingchen Song, Zhendong Peng, Fuping Pan, Zhiyong Wu 0001 |
ICASSP | 6 |
| 2023 | Context-Aware Coherent Speaking Style Prediction with Hierarchical Transformers for Audiobook Speech SynthesisabstractRecent advances in text-to-speech have significantly improved the expressiveness of synthesized speech. However, it is still challenging to generate speech with contextually appropriate and coherent speaking style for multi-sentence text in audiobooks. In this paper, we propose a context-aware coherent speaking style prediction method for audiobook speech synthesis. To predict the style embedding of the current utterance, a hierarchical transformer-based context-aware style predictor with a mixture attention mask is designed, considering both text-side context information and speech- side style information of previous speeches. Based on this, we can generate long-form speech with coherent style and prosody sentence by sentence. Objective and subjective evaluations on a Mandarin audiobook dataset demonstrate that our proposed model can generate speech with more expressive and coherent speaking style than baselines, for both single-sentence and multi-sentence test1. Shun Lei, Yixuan Zhou 0002, Liyang Chen, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng |
ICASSP | 4 |
| 2023 | Av-Sepformer: Cross-Attention Sepformer for Audio-Visual Target Speaker ExtractionabstractVisual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross- and self-attention to fuse and model features from audio and visual. AV-SepFormer splits the audio feature into a number of chunks, equivalent to the length of the visual feature. Then self- and cross-attention are employed to model the multi-modal features. Furthermore, we use a novel 2D positional encoding, that introduces the positional information between and within chunks and provides significant gains over the traditional positional encoding. Our model has two key advantages: the time granularity of audio chunked feature is synchronized to the visual feature, which alleviates the harm caused by the inconsistency of audio and video sampling rate; by combining self- and cross-attention, feature fusion and speech extraction processes are unified within an attention paradigm. The experimental results show that AV-SepFormer significantly outperforms other existing methods. Jiuxin Lin, Xinyu Cai, Heinrich Dinkel, Jun Chen 0024, Zhiyong Yan, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 8 |
| 2023 | TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length PenaltyabstractIn this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply length penalty (i.e., by trimming trailing frames, see Fig. 1-(b)) directly on the spectrogram of input utterances, which does not require any alignment. We demonstrate that TrimTail is computationally cheap and can be applied online and optimized with any training loss or any model architecture on any dataset without any extra effort by applying it on various end-to-end streaming ASR networks either trained with CTC loss [1] or Transducer loss [2]. We achieve 100 ~ 200ms latency reduction with equal or even better accuracy on both Aishell-1 and Librispeech. Moreover, by using TrimTail, we can achieve a 400ms algorithmic improvement of User Sensitive Delay (USD) with an accuracy loss of less than 0.2. Xingchen Song, Di Wu 0061, Zhiyong Wu 0001, Yuekai Zhang, Zhendong Peng, Wenpeng Li, Fuping Pan, Changbao Zhu |
ICASSP | 3 |
| 2023 | TFCnet: Time-Frequency Domain Corrector for Speech SeparationabstractDeep learning-based methods have made significant achievements in speech separation. Especially the time-domain separation methods have achieved the best performance in recent years. However, time-domain methods are unstable for waveform transformation, which is prone to amplitude and phase errors. Considering the robustness of time-frequency (T-F) domain methods, we propose an innovative network architecture called Time-Frequency Domain Corrector Network (TFCNet), which consists of a time-domain separator and a specially-designed T-F domain corrector. The corrector module is added after the time-domain separation step to correct the real and imaginary parts information in the T-F domain. The proposed model achieves state-of-the-art performance with an SI-SDRi of 22.2dB on the WSJ0-2mix dataset and an SI-SDRi of 19.4dB on the Libri-2mix dataset. Weinan Tong, Jiaxu Zhu, Jun Chen 0024, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng |
ICASSP | 4 |
| 2023 | A Synthetic Corpus Generation Method for Neural Vocoder TrainingabstractNowadays, neural vocoders are preferred for their ability to synthesize high-fidelity audio. However, training a neural vocoder requires a massive corpus of high-quality real audio, and the audio recording process is often labor-intensive. In this work, we propose a synthetic corpus generation method for neural vocoder training, which can easily generate synthetic audio with an unlimited number at nearly no cost. We explicitly model the prior characteristics of audio from multiple target domains simultaneously (e.g., speeches, singing voices, and instrumental pieces) to equip the generated audio data with these characteristics. And we show that our synthetic corpus allows the neural vocoder to achieve competitive results without any real audio in the training process. To validate the effectiveness of our proposed method, we performed empirical experiments on both speech and music utterances in subjective and objective metrics. The experimental results show that the neural vocoder trained with the synthetic corpus produced by our method can generalize to multiple target scenarios and has excellent singing voice (MOS: 4.20) and instrumental piece (MOS: 4.00) synthesis results. Zilin Wang 0002, Jun Chen 0024, Sipan Li, Jinfeng Bai, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 7 |
| 2023 | DASA: Difficulty-Aware Semantic Augmentation for Speaker VerificationabstractData augmentation is vital to the generalization ability and robustness of deep neural networks (DNNs) models. Existing augmentation methods for speaker verification manipulate the raw signal, which are time-consuming and the augmented samples lack diversity. In this paper, we present a novel difficulty-aware semantic augmentation (DASA) approach for speaker verification, which can generate diversified training samples in speaker embedding space with negligible extra computing cost. Firstly, we augment training samples by perturbing speaker embeddings along semantic directions, which are obtained from speaker-wise covariance matrices. Secondly, accurate covariance matrices are estimated from robust speaker embeddings during training, so we introduce difficulty-aware additive margin softmax (DAAM-Softmax) to obtain optimal speaker embeddings. Finally, we assume the number of augmented samples goes to infinity and derive a closed-form upper bound of the expected loss with DASA, which achieves compatibility and efficiency. Extensive experiments demonstrate the proposed approach can achieve a remarkable performance improvement. The best result achieves a 14.6% relative reduction in EER metric on CN-Celeb evaluation set. Yang Zhang 0025, Zhiyong Wu 0001, Tao Wei 0003, Helen M. Meng |
ICASSP | 3 |
| 2023 | CB-Conformer: Contextual Biasing Conformer for Biased Word RecognitionabstractDue to the mismatch between the source and target domains, how to better utilize the biased word information to improve the performance of the automatic speech recognition model in the target domain becomes a hot research topic. Previous approaches either decode with a fixed external language model or introduce a sizeable biasing module, which leads to poor adaptability and slow inference. In this work, we propose CB-Conformer to improve biased word recognition by introducing the Contextual Biasing Module and the Self-Adaptive Language Model to vanilla Conformer. The Contextual Biasing Module combines audio fragments and contextual information, with only 0.2% model parameters of the original Conformer. The Self-Adaptive Language Model modifies the internal weights of biased words based on their recall and precision, resulting in a greater focus on biased words and more successful integration with the automatic speech recognition model than the standard fixed language model. In addition, we construct and release an open-source Mandarin biased-word dataset based on WenetSpeech. Experiments indicate that our proposed method brings a 15.34% character error rate reduction, a 14.13% biased word recall increase, and a 6.80% biased word F1-score increase compared with the base Conformer. Yaoxun Xu, Baiji Liu, Qiaochu Huang, Xingchen Song, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng |
ICASSP | 5 |
| 2023 | Keyword-Specific Acoustic Model Pruning for Open-Vocabulary Keyword SpottingabstractThe open-vocabulary KWS system allows users to customize wake words, but its application is limited by the model size. In this paper, we design a dynamic acoustic model with input-dependent parameters. We find that acoustic frames with similar pronunciation generate similar subnetworks, and different parameters contribute to recognizing different phonemes. Based on this observation, we further constrain the structural similarity among the subnetworks with the same phoneme pseudo-label, thus independent subnetworks to recognize different phonemes could be pruned out. When used in the end-to-end KWS system, the subnetworks recognizing phonemes in the keyword would be combined as a keyword-specific acoustic model, and the parameters that do not contribute to recognizing the keyword are pruned off. Experiments demonstrate that the proposed method can prune more than 80% of the parameters without performance loss. Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 3 |
| 2023 | Enhancing the Vocal Range of Single-Speaker Singing Voice Synthesis with Melody-Unsupervised Pre-TrainingabstractThe single-speaker singing voice synthesis (SVS) usually underperforms at pitch values that are out of the singer's vocal range or associated with limited training samples. Based on our previous work, this work proposes a melody-unsupervised multi-speaker pretraining method conducted on a multi-singer dataset to enhance the vocal range of the single-speaker, while not degrading the timbre similarity. This pre-training method can be deployed to a large-scale multi-singer dataset, which only contains audio-and-lyrics pairs without phonemic timing information and pitch annotation. Specifically, in the pre-training step, we design a phoneme predictor to produce the frame-level phoneme probability vectors as the phonemic timing information and a speaker encoder to model the timbre variations of different singers, and directly estimate the frame-level f0 values from the audio to provide the pitch information. These pre-trained model parameters are delivered into the fine-tuning step as prior knowledge to enhance the single speaker's vocal range. Moreover, this work also contributes to improving the sound quality and rhythm naturalness of the synthesized singing voices. It is the first to introduce a differentiable duration regulator to improve the rhythm naturalness of the synthesized voice, and a bi-directional flow model to improve the sound quality. Experimental results verify that the proposed SVS system outperforms the baseline on both sound quality and naturalness. Shaohuan Zhou, Xu Li 0015, Zhiyong Wu 0001, Ying Shan, Helen M. Meng |
ICASSP | 3 |
| 2023 | GTN-Bailando: Genre Consistent long-Term 3D Dance Generation Based on Pre-Trained Genre Token NetworkabstractMusic-driven 3D dance generation has become an intensive research topic in recent years with great potential for real-world applications. Most existing methods lack the consideration of genre, which results in genre inconsistency in the generated dance movements. In addition, the correlation between the dance genre and the music has not been investigated. To address these issues, we propose a genre-consistent dance generation framework, GTN-Bailando. First, we propose the Genre Token Network (GTN), which infers the genre from music to enhance the genre consistency of long-term dance generation. Second, to improve the generalization capability of the model, the strategy of pre-training and fine-tuning is adopted. Experimental results on the AIST++ dataset show that the proposed dance generation framework outperforms state-of-the-art methods in terms of motion quality and genre consistency1. Haolin Zhuang, Shun Lei, Long Xiao, Liyang Chen, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng |
ICASSP | 7 |
| 2023 | SnakeGAN: A Universal Vocoder Leveraging DDSP Prior Knowledge and Periodic Inductive BiasabstractGenerative adversarial network (GAN)-based neural vocoders have been widely used in audio synthesis tasks due to their high generation quality, efficient inference, and small computation footprint. However, it is still challenging to train a universal vocoder which can generalize well to out-of-domain (OOD) scenarios, such as unseen speaking styles, non-speech vocalization, singing, and musical pieces. In this work, we propose SnakeGAN, a GAN-based universal vocoder, which can synthesize high-fidelity audio in various OOD scenarios. SnakeGAN takes a coarse-grained signal generated by a differentiable digital signal processing (DDSP) model as prior knowledge, aiming at recovering high-fidelity waveform from a Mel-spectrogram. We introduce periodic nonlinearities through the Snake activation function and anti-aliased representation into the generator, which further brings desired inductive bias for audio synthesis and significantly improves the extrapolation capacity for universal vocoding in unseen scenarios. To validate the effectiveness of our proposed method, we train SnakeGAN with only speech data and evaluate its performance for various OOD distributions with both subjective and objective metrics. Experimental results show that SnakeGAN significantly outperforms the compared approaches and can generate high-fidelity audio samples including unseen speakers with unseen styles, singing voices, instrumental pieces, and nonverbal vocalization. Sipan Li, Songxiang Liu, Xiang Li 0105, Yanyao Bian, Chao Weng, Zhiyong Wu 0001, Helen M. Meng |
ICME | 7 |
| 2023 | Adversarial Speaker Disentanglement Using Unannotated External Data for Self-supervised Representation-based Voice ConversionabstractNowadays, recognition-synthesis-based methods have been quite popular with voice conversion (VC). By introducing linguistics features with good disentangling characters extracted from an automatic speech recognition (ASR) model, the VC performance achieved considerable breakthroughs. Recently, self-supervised learning (SSL) methods trained with a large-scale unannotated speech corpus have been applied to downstream tasks focusing on the content information, which is suitable for VC tasks. However, a huge amount of speaker information in SSL representations degrades timbre similarity and the quality of converted speech significantly. To address this problem, we proposed a high-similarity any-to-one voice conversion method with the input of SSL representations. We incorporated adversarial training mechanisms in the synthesis module using external unannotated corpora. Two auxiliary discriminators were trained to distinguish whether a sequence of mel-spectrograms has been converted by the acoustic model and whether a sequence of content embeddings contains speaker information from external corpora. Experimental results show that our proposed method achieves comparable similarity and higher naturalness than the supervised method, which needs a huge amount of annotated corpora for training and is applicable to improve similarity for VC methods with other SSL representations as input. Xintao Zhao, Shuai Wang 0016, Yang Chao, Zhiyong Wu 0001, Helen M. Meng |
ICME | 4 |
| 2023 | DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion ModelsabstractThe art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the rhythm and semantics of the gesture to the corresponding speech. To address these problems, we present DiffuseStyleGesture, a diffusion model based speech-driven gesture generation approach. It generates high-quality, speech-matched, stylized, and diverse co-speech gestures based on given speeches of arbitrary length. Specifically, we introduce cross-local attention and self-attention to the gesture diffusion pipeline to generate better speech matched and realistic gestures. We then train our model with classifier-free guidance to control the gesture style by interpolation or extrapolation. Additionally, we improve the diversity of generated gestures with different initial gestures and noise. Extensive experiments show that our method outperforms recent approaches on speech-driven gesture generation. Our code, pre-trained models, and demos are available at https://github.com/YoungSeng/DiffuseStyleGesture. Zhiyong Wu 0001, Minglei Li 0001, Zhensong Zhang, Weihong Bao, Long Xiao |
IJCAI | 2 |
| 2023 | MC-SpEx: Towards Effective Speaker Extraction with Multi-Scale Interfusion and Conditional Speaker Modulation
Jun Chen 0024, Wei Rao 0002, Zilin Wang 0002, Jiuxin Lin, Yukai Jv, Shulin He, Yannan Wang, Zhiyong Wu 0001 |
INTERSPEECH | 8 |
| 2023 | Diverse and Expressive Speech Prosody Prediction with Denoising Diffusion Probabilistic Model
Xiang Li 0105, Songxiang Liu, Max W. Y. Lam, Zhiyong Wu 0001, Chao Weng, Helen M. Meng |
INTERSPEECH | 4 |
| 2023 | Towards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis
Shun Lei, Qiaochu Huang, Yixuan Zhou 0002, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng |
INTERSPEECH | 5 |
| 2023 | Focus on the Sound around You: Monaural Target Speaker Extraction via Distance and Speaker Information
Jiuxin Lin, Heinrich Dinkel, Jun Chen 0024, Zhiyong Wu 0001, Zhiyong Yan |
INTERSPEECH | 5 |
| 2023 | Gesper: A Restoration-Enhancement Framework for General Speech Reconstruction
Yupeng Shi, Jun Chen 0024, Wei Rao 0002, Shulin He, Andong Li, Yannan Wang, Zhiyong Wu 0001 |
INTERSPEECH | 8 |
| 2023 | ZeroPrompt: Streaming Acoustic Encoders are Zero-Shot Masked LMsabstractIn this paper, we present ZeroPrompt (Figure 1-(a)) and the corresponding Prompt-and-Refine strategy (Figure 3), two simple but effective training-free methods to decrease the Token Display Time (TDT) of streaming ASR models without any accuracy loss.The core idea of ZeroPrompt is to append zeroed content to each chunk during inference, which acts like a prompt to encourage the model to predict future tokens even before they were spoken.We argue that streaming acoustic encoders naturally have the modeling ability of Masked Language Models and our experiments demonstrate that ZeroPrompt is engineering cheap and can be applied to streaming acoustic encoders on any dataset without any accuracy loss.Specifically, compared with our baseline models, we achieve 350 ∼ 700ms reduction on First Token Display Time (TDT-F) and 100 ∼ 400ms reduction on Last Token Display Time (TDT-L), with theoretically and experimentally equal WER on both Aishell-1 and Librispeech datasets. Xingchen Song, Di Wu 0061, Zhendong Peng, Bo Dang 0004, Fuping Pan, Zhiyong Wu 0001 |
INTERSPEECH | 7 |
| 2023 | Prosody Modeling with 3D Visual Information for Expressive Video Dubbing
Shansong Liu, Xu Li 0015, Haozhe Wu, Zhiyong Wu 0001, Ying Shan, Jia Jia 0001 |
INTERSPEECH | 5 |
| 2023 | SememeASR: Boosting Performance of End-to-End Speech Recognition against Domain and Long-Tailed Data Shift with Sememe Semantic KnowledgeabstractRecently, excellent progress has been made in speech recognition. However, pure data-driven approaches have struggled to solve the problem in domain-mismatch and long-tailed data. Considering that knowledge-driven approaches can help data-driven approaches alleviate their flaws, we introduce sememe-based semantic knowledge information to speech recognition (SememeASR). Sememe, according to the linguistic definition, is the minimum semantic unit in a language and is able to represent the implicit semantic information behind each word very well. Our experiments show that the introduction of sememe information can improve the effectiveness of speech recognition. In addition, our further experiments show that sememe knowledge can improve the model's recognition of long-tailed data and enhance the model's domain generalization ability. Jiaxu Zhu, Changhe Song, Zhiyong Wu 0001, Helen M. Meng |
INTERSPEECH | 3 |
| 2023 | Text-Only Domain Adaptation for End-to-End Speech Recognition through Down-Sampling Acoustic RepresentationabstractMapping two modalities, speech and text, into a shared representation space, is a research topic of using text-only data to improve end-to-end automatic speech recognition (ASR) performance in new domains. However, the length of speech representation and text representation is inconsistent. Although the previous method up-samples the text representation to align with acoustic modality, it may not match the expected actual duration. In this paper, we proposed novel representations match strategy through down-sampling acoustic representation to align with text modality. By introducing a continuous integrate-and-fire (CIF) module generating acoustic representations consistent with token length, our ASR model can learn unified representations from both modalities better, allowing for domain adaptation using text-only data of the target domain. Experiment results of new domain data demonstrate the effectiveness of the proposed method. Jiaxu Zhu, Weinan Tong, Yaoxun Xu, Changhe Song, Zhiyong Wu 0001, Zhao You, Dan Su 0002, Dong Yu 0001, Helen M. Meng |
INTERSPEECH | 5 |
| 2023 | SpeechTripleNet: End-to-End Disentangled Speech Representation Learning for Content, Timbre and ProsodyabstractDisentangled speech representation learning aims to separate different factors of variation from speech into disjoint representations. This paper focuses on disentangling speech into representations for three factors: spoken content, speaker timbre, and speech prosody. Many previous methods for speech disentanglement have focused on separating spoken content and speaker timbre. However, the lack of explicit modeling of prosodic information leads to degraded speech generation performance and uncontrollable prosody leakage into content and/or speaker representations. While some recent methods have utilized explicit speaker labels or pre-trained models to facilitate triple-factor disentanglement, there are no end-to-end methods to simultaneously disentangle three factors using only unsupervised or self-supervised learning objectives. This paper introduces SpeechTripleNet, an end-to-end method to disentangle speech into representations for content, timbre, and prosody. Based on VAE, SpeechTripleNet restricts the structures of the latent variables and the amount of information captured in them to induce disentanglement. It is a pure unsupervised/self-supervised learning method that only requires speech data and no additional labels. Our qualitative and quantitative results demonstrate that SpeechTripleNet is effective in achieving triple-factor speech disentanglement, as well as controllable speech editing concerning different factors. Xixin Wu, Zhiyong Wu 0001, Helen M. Meng |
ACM Multimedia | 3 |
| 2023 | UnifiedGesture: A Unified Gesture Synthesis Model for Multiple SkeletonsabstractThe automatic co-speech gesture generation draws much attention in computer animation. Previous works designed network structures on individual datasets, which resulted in a lack of data volume and generalizability across different motion capture standards. In addition, it is a challenging task due to the weak correlation between speech and gestures. To address these problems, we present UnifiedGesture, a novel diffusion model-based speech-driven gesture synthesis approach, trained on multiple gesture datasets with different skeletons. Specifically, we first present a retargeting network to learn latent homeomorphic graphs for different motion capture standards, unifying the representations of various gestures while extending the dataset. We then capture the correlation between speech and gestures based on a diffusion model architecture using cross-local attention and self-attention to generate better speech-matched and realistic gestures. To further align speech and gesture and increase diversity, we incorporate reinforcement learning on the discrete gesture units with a learned reward function. Extensive experiments show that UnifiedGesture outperforms recent approaches on speech-driven gesture generation in terms of CCA, FGD, and human-likeness. Zilin Wang 0002, Zhiyong Wu 0001, Minglei Li 0001, Zhensong Zhang, Qiaochu Huang, Songcen Xu, Changpeng Yang, Zonghong Dai |
ACM Multimedia | 3 |
| 2023 | Lite-RTSE: Exploring a Cost-Effective Lite DNN Model for Real-Time Speech Enhancement in RTC ScenariosabstractThe noise reduction performance of DNN-based monaural speech enhancement (SE) methods has been significantly improved in recent years, while the complexity of the model has also been increased several times. Therefore, it is highly desirable to explore more ‘cost-effective’ speech enhancement methods for a wider range of hardware platforms. In this letter, we investigate low-cost model design strategies and propose a lite real-time speech enhancement (Lite-RTSE) model. This real-time SE model achieves efficient speech enhancement by leveraging the low-dimensional long short-term memory (LSTM) units and a novel multi-order convolution block. A two-stage complex spectrum reconstruction scheme of ‘masking + residual’ contributes to better quality and intelligibility of enhanced speech. Experimental results show that Lite-RTSE model is able to achieve competitive speech denoising performance compared with state-of-the-art SE models, while only containing 1.56 M parameters at 0.55 G multiply-accumulate operations per second (MAC/S). Xingwei Liang, Lu Zhang 0055, Zhiyong Wu 0001, Ruifeng Xu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | MSStyleTTS: Multi-Scale Style Modeling With Hierarchical Context Information for Expressive Speech SynthesisabstractExpressive speech synthesis is crucial for many human-computer interaction scenarios, such as audiobooks, podcasts, and voice assistants. Previous works focus on predicting the style embeddings at one single scale from the information within the current sentence. Whereas, context information in neighboring sentences and multi-scale nature of style in human speech are neglected, making it challenging to convert multi-sentence text into natural and expressive speech. In this paper, we propose MSStyleTTS, a style modeling method for expressive speech synthesis, to capture and predict styles at different levels from a wider range of context rather than a sentence. Two sub-modules, including multi-scale style extractor and multi-scale style predictor, are trained together with a FastSpeech 2 based acoustic model. The predictor is designed to explore the hierarchical context information by considering structural relationships in context and predict style embeddings at global-level, sentence-level and subword-level. The extractor extracts multi-scale style embedding from the ground-truth speech and explicitly guides the style prediction. Evaluations on both in-domain and out-of-domain audiobook datasets demonstrate that the proposed method significantly outperforms the three baselines. In addition, we conduct the analysis of the context information and multi-scale style representations that have never been discussed before. Shun Lei, Yixuan Zhou 0002, Liyang Chen, Zhiyong Wu 0001, Xixin Wu, Shiyin Kang, Helen M. Meng |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Hiformer: Sequence Modeling Networks With Hierarchical Attention MechanismsabstractThe attention-based encoder-decoder structure, such as the Transformer, has achieved state-of-the-art performance on various sequence modeling tasks,e.g., machine translation (MT) and automatic speech recognition (ASR), benefited from the superior capability of layer-wise self-attention mechanism in the encoder/decoder to access long-distance contextual information. Recently, analysis on the Transformer layers has shown that different levels of information,e.g., phoneme level, word level and semantic level, are represented at different layers. Effectively integrating information from various levels is important for structured prediction. However, the self-attention in the conventional Transformer structure only focuses on intra-layer integration, and does not explicitly model inter-layer information relationships. Also, attention across the encoder and decoder (cross-coder) only focuses on the top encoder layer but ignores the intermediate layers. In this paper, we propose a sequence modeling structure equipped with a hierarchical attention mechanism, named Hiformer, that can consider the inter-layer and cross-coder hierarchical information to improve structured prediction performance. Extensive experiments conducted on both MT and ASR tasks demonstrate the effectiveness of the proposed Hiformer model. Xixin Wu, Kun Li 0003, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech SynthesisabstractNaturalness and expressiveness are crucial for audiobook speech synthesis, but now are limited by the averaged global-scale speaking style representation. In this paper, we propose an unsupervised multi-scale context-sensitive text-to-speech model for audiobooks. A multi-scale hierarchical context encoder is specially designed to predict both global-scale context style embedding and local-scale context style embedding from a wider context of input text in a hierarchical manner. Likewise, a multi-scale reference encoder is introduced to extract reference style embeddings at both global and local scales from the reference speech, which is used to guide the prediction of speaking styles. On top of these, a bi-reference attention mechanism is used to align both local-scale reference style embedding sequence and local-scale context style embedding sequence with corresponding phoneme embedding sequence. Both objective and subjective experiment results on a real-world multi-speaker Mandarin novel audio dataset demonstrate the excellent performance of our proposed method over all baselines in terms of naturalness and expressiveness of the synthesized speech. Xueyuan Chen, Shun Lei, Zhiyong Wu 0001, Weifeng Zhao, Helen M. Meng |
COLING | 3 |
| 2022 | A Character-Level Span-Based Model for Mandarin Prosodic Structure PredictionabstractThe accuracy of prosodic structure prediction is crucial to the naturalness of synthesized speech in Mandarin text-to-speech system, but now is limited by widely-used sequence-to-sequence framework and error accumulation from previous word segmentation results. In this paper, we propose a span-based Mandarin prosodic structure prediction model to obtain an optimal prosodic structure tree, which can be converted to corresponding prosodic label sequence. Instead of the prerequisite for word segmentation, rich linguistic features are provided by Chinese character-level BERT and sent to encoder with self-attention architecture. On top of this, span representation and label scoring are used to describe all possible prosodic structure trees, of which each tree has its corresponding score. To find the optimal tree with the highest score for a given sentence, a bottom-up CKYstyle algorithm is further used. The proposed method can predict prosodic labels of different levels at the same time and accomplish the process directly from Chinese characters in an end-to-end manner. Experiment results on two real-world datasets demonstrate the excellent performance of our span-based method over all sequence-to-sequence baseline approaches. Xueyuan Chen, Changhe Song, Yixuan Zhou 0002, Zhiyong Wu 0001, Changbin Chen, Zhongqin Wu, Helen M. Meng |
ICASSP | 4 |
| 2022 | Transformer-S2A: Robust and Efficient Speech-to-AnimationabstractWe propose a novel robust and efficient Speech-to-Animation (S2A) approach for synchronized facial animation generation in human-computer interaction. Compared with conventional approaches, the proposed approach utilizes phonetic posteriorgrams (PPGs) of spoken phonemes as input to ensure the cross-language and cross-speaker ability, and introduces corresponding prosody features (i.e. pitch and energy) to further enhance the expression of generated animation. Mixture-of-experts (MOE)-based Transformer is employed to better model contextual information while provide significant optimization on computation efficiency. Experiments demonstrate the effectiveness of the proposed approach on both objective and subjective evaluation with 17× inference speedup compared with the state-of-the-art approach. Liyang Chen, Zhiyong Wu 0001, Jun Ling, Runnan Li, Xu Tan 0003, Sheng Zhao 0002 |
ICASSP | 2 |
| 2022 | FullSubNet+: Channel Attention Fullsubnet with Complex Spectrograms for Speech EnhancementabstractPreviously proposed FullSubNet has achieved outstanding performance in Deep Noise Suppression (DNS) Challenge and attracted much attention. However, it still encounters issues such as input-output mismatch and coarse processing for frequency bands. In this paper, we propose an extended single-channel real-time speech enhancement framework called FullSubNet+ with following significant improvements. First, we design a lightweight multi-scale time sensitive channel attention (MulCA) module which adopts multi-scale convolution and channel attention mechanism to help the network focus on more discriminative frequency bands for noise reduction. Then, to make full use of the phase information in noisy speech, our model takes all the magnitude, real and imaginary spectrograms as inputs. Moreover, by replacing the long short-term memory (LSTM) layers in original full-band model with stacked temporal convolutional network (TCN) blocks, we design a more efficient full-band module called full-band extractor. The experimental results in DNS Challenge dataset show the superior performance of our FullSubNet+, which reaches the state-of-the-art (SOTA) performance and outperforms other existing speech enhancement approaches. Jun Chen 0024, Zilin Wang 0002, Deyi Tuo, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng |
ICASSP | 4 |
| 2022 | An End-to-End Chinese Text Normalization Model Based on Rule-Guided Flat-Lattice TransformerabstractText normalization, defined as a procedure transforming nonstandard words to spoken-form words, is crucial to the intelligibility of synthesized speech in text-to-speech system. Rule-based methods without considering context can not eliminate ambiguation, whereas sequence-to-sequence neural network based methods suffer from the unexpected and uninterpretable errors problem. Recently proposed hybrid system treats rule-based model and neural model as two cascaded sub-modules, where limited interaction capability makes neural network model cannot fully utilize expert knowledge contained in the rules. Inspired by Flat-LAttice Transformer (FLAT), we propose an end-to-end Chinese text normalization model, which accepts Chinese characters as direct input and integrates expert knowledge contained in rules into the neural network, both contribute to the superior performance of proposed model for the text normalization task. We also release a first publicly accessible large-scale dataset for Chinese text normalization. Our proposed model has achieved excellent results on this dataset. Wenlin Dai, Changhe Song, Xiang Li 0067, Zhiyong Wu 0001, Huashan Pan, Xiulin Li, Helen M. Meng |
ICASSP | 4 |
| 2022 | Towards Expressive Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech SynthesisabstractPrevious works on expressive speech synthesis mainly focus on current sentence. The context in adjacent sentences is neglected, resulting in inflexible speaking style for the same text, which lacks speech variations. In this paper, we propose a hierarchical framework to model speaking style from context. A hierarchical context encoder is proposed to explore a wider range of contextual information considering structural relationship in context, including inter-phrase and inter-sentence relations. Moreover, to encourage this encoder to learn style representation better, we introduce a novel training strategy with knowledge distillation, which provides the target for encoder training. Both objective and subjective evaluations on a Mandarin lecture dataset demonstrate that the proposed method can significantly improve the naturalness and expressiveness of the synthesized speech1. Shun Lei, Yixuan Zhou 0002, Liyang Chen, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng |
ICASSP | 4 |
| 2022 | Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-Based Multi-Modal Context ModelingabstractComparing with traditional text-to-speech (TTS) systems, conversational TTS systems are required to synthesize speeches with proper speaking style confirming to the conversational context. However, state-of-the-art context modeling methods in conversational TTS only model the textual information in context with a recurrent neural network (RNN). Such methods have limited ability in modeling the inter-speaker influence in conversations, and also neglect the speaking styles and the intra-speaker inertia inside each speaker. Inspired by DialogueGCN and its superiority in modeling such conversational influences than RNN based approaches, we propose a graph-based multi-modal context modeling method and adopt it to conversational TTS to enhance the speaking styles of synthesized speeches. Both the textual and speaking style information in the context are extracted and processed by DialogueGCN to model the inter- and intra-speaker influence in conversations. The outputs of DialogueGCN are then summarized by attention mechanism, and converted to the enhanced speaking style for current utterance. An English conversation corpus is collected and annotated for our research and released to public. Experiment results on this corpus demonstrate the effectiveness of our proposed approach, which outperforms the state-of-the-art context modeling method in conversational TTS in both MOS and ABX preference rate. Jingbei Li, Zhiyong Wu 0001, Helen M. Meng, Chao Weng, Dan Su 0002 |
ICASSP | 4 |
| 2022 | Neufa: Neural Network Based End-to-End Forced Alignment with Bidirectional Attention MechanismabstractAlthough deep learning and end-to-end models have been widely used and shown their superiority in automatic speech recognition (ASR) and text-to-speech (TTS) synthesis, state-of-the-art forced alignment (FA) models are still based on hidden Markov model (HMM). HMM has limited view of contextual information and is developed with long pipelines, leading to error accumulation and unsatisfactory performance. Inspired by the capability of attention mechanism in capturing long term contextual information and learning alignments in ASR and TTS, we propose a neural network based end-to-end forced aligner called NeuFA, in which a novel bidirectional attention mechanism plays an essential role. NeuFA integrates the alignment learning of both ASR and TTS tasks in a unified framework by learning bidirectional alignment information from a shared attention matrix in the proposed bidirectional attention mechanism. Alignments are extracted from the learnt attention weights and optimized by the ASR, TTS and FA tasks in a multi-task learning manner. Experimental results demonstrate the effectiveness of our proposed model, with mean absolute error (MAE) on test set drops from 25.8 ms to 23.7 ms at word level, and from 18.0 ms to 15.7 ms at phoneme level compared with state-of-the-art HMM based model. Jingbei Li, Zhiyong Wu 0001, Helen M. Meng, Qiao Tian 0001, Yuping Wang 0005, Yuxuan Wang 0002 |
ICASSP | 3 |
| 2022 | Adversarial Sample Detection for Speaker Verification by Neural VocodersabstractAutomatic speaker verification (ASV), one of the most important technology for biometric identification, has been widely adopted in security-critical applications. However, ASV is seriously vulnerable to recently emerged adversarial attacks, yet effective counter-measures against them are limited. In this paper, we adopt neural vocoders to spot adversarial samples for ASV. We use the neural vocoder to re-synthesize audio and find that the difference between the ASV scores for the original and re-synthesized audio is a good indicator for discrimination between genuine and adversarial samples. This effort is, to the best of our knowledge, among the first to pursue such a technical direction for detecting time-domain adversarial samples for ASV, and hence there is a lack of established baselines for comparison. Consequently, we implement the Griffin-Lim algorithm as the detection baseline. The proposed approach achieves effective detection performance that outperforms the baselines in all the settings. We also show that the neural vocoder adopted in the detection framework is dataset-independent. Our codes will be made open-source for future works to do fair comparison1. Po-Chun Hsu, Ji Gao, Shen Huang, Jian Kang 0006, Zhiyong Wu 0001, Helen M. Meng, Hung-yi Lee |
ICASSP | 7 |
| 2022 | Neural Architecture Search for Speech Emotion RecognitionabstractDeep neural networks have brought significant advancements to speech emotion recognition (SER). However, the architecture design in SER is mainly based on expert knowledge and empirical (trial-and-error) evaluations, which is time-consuming and resource intensive. In this paper, we propose to apply neural architecture search (NAS) techniques to automatically configure the SER models. To accelerate the candidate architecture optimization, we propose a uniform path dropout strategy to encourage all candidate architecture operations to be equally optimized. Experimental results of two different neural structures on IEMOCAP show that NAS can improve SER performance (54.89% to 56.28%) while maintaining model parameter sizes. The proposed dropout strategy also shows superiority over the previous approaches. Xixin Wu, Shoukang Hu, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 3 |
| 2022 | An Approach to Mispronunciation Detection and Diagnosis with Acoustic, Phonetic and Linguistic (APL) EmbeddingsabstractMany mispronunciation detection and diagnosis (MD&D) research approaches try to exploit both the acoustic and linguistic features as input. Yet the improvement of the performance is limited, partially due to the shortage of large amount annotated training data at the phoneme level. Phonetic embeddings, extracted from ASR models trained with huge amount of word level annotations, can serve as a good representation of the content of input speech, in a noise-robust and speaker-independent manner. These embeddings, when used as implicit phonetic supplementary information, can alleviate the data shortage of explicit phoneme annotations. We propose to utilize Acoustic, Phonetic and Linguistic (APL) embedding features jointly for building a more powerful MD&D system. Experimental results obtained on the L2-ARCTIC database show the proposed approach outperforms the baseline by 9.93%, 10.13% and 6.17% on the detection accuracy, diagnosis error rate and the F-measure, respectively. Wenxuan Ye, Shaoguang Mao, Frank K. Soong, Wenshan Wu, Yan Xia 0005, Jonathan Tien, Zhiyong Wu 0001 |
ICASSP | 7 |
| 2022 | Disentangling Content and Fine-Grained Prosody Information Via Hybrid ASR Bottleneck Features for Voice ConversionabstractNon-parallel data voice conversion (VC) have achieved considerable breakthroughs recently through introducing bottleneck features (BNFs) extracted by the automatic speech recognition(ASR) model. However, selection of BNFs have a significant impact on VC result. For example, when extracting BNFs from ASR trained with Cross Entropy loss (CE-BNFs) and feeding into neural network to train a VC system, the timbre similarity of converted speech is significantly degraded. If BNFs are extracted from ASR trained using Connectionist Temporal Classification loss (CTC-BNFs), the naturalness of the converted speech may decrease. This phenomenon is caused by the difference of information contained in BNFs. In this paper, we proposed an any-to-one VC method using hybrid bottleneck features extracted from CTC-BNFs and CE-BNFs to complement each other advantages. Gradient reversal layer and instance normalization were used to extract prosody information from CE-BNFs and content information from CTC-BNFs. Auto-regressive decoder and Hifi-GAN vocoder were used to generate high-quality waveform. Experimental results show that our proposed method achieves higher similarity, naturalness, quality than baseline method and reveals the differences between the information contained in CE-BNFs and CTC-BNFs as well as the influence they have on the converted speech. Xintao Zhao, Changhe Song, Zhiyong Wu 0001, Shiyin Kang, Deyi Tuo, Helen M. Meng |
ICASSP | 4 |
| 2022 | Learning from Designers: Fashion Compatibility Analysis Via Dataset DistillationabstractLearning fashion compatibility is of great significance to both academic research and industry, which serves as a key technique for many real applications like online shopping recommendation and clothing generation. In previous studies, user-generated data (e.g. outfits from social media platform) are usually used for learning item embeddings and further modeling the compatibility. However, due to the noisy and messy nature of such data, one can hardly learn a representation that can clearly characterize the fashion-related attributes (e.g. color, material). In this paper, we propose an Attention-based Dataset Distillation Graph Neural Network (ADD-GNN) to leverage the designer-generated data as a guidance on modeling the outfit compatibility. Specifically, we jointly optimize two components which distill knowledge from fashion designers for feature representation learning and model the overall compatibility through attention-based graph neural network. Experimental results on real world fashion datasets clearly demonstrate the superiority of our proposed ADD-GNN against several competitive baselines in outfit compatibility tasks, which proves the effectiveness of distilling knowledge from designers. Yulan Chen, Zhiyong Wu 0001, Zheyan Shen, Jia Jia 0001 |
ICIP | 2 |
| 2022 | The ReprGesture entry to the GENEA Challenge 2022abstractThis paper describes the ReprGesture entry to the Generation and Evaluation of Non-verbal Behaviour for Embodied Agents (GENEA) challenge 2022. The GENEA challenge provides the processed datasets and performs crowdsourced evaluations to compare the performance of different gesture generation systems. In this paper, we explore an automatic gesture generation system based on multimodal representation learning. We use WavLM features for audio, FastText features for text and position and rotation matrix features for gesture. Each modality is projected to two distinct subspaces: modality-invariant and modality-specific. To learn inter-modality-invariant commonalities and capture the characters of modality-specific representations, gradient reversal layer based adversarial classifier and modality reconstruction decoders are used during training. The gesture decoder generates proper gestures using all representations and features related to the rhythm in the audio. Our code, pre-trained models and demo are available at https://github.com/YoungSeng/ReprGesture. Zhiyong Wu 0001, Minglei Li 0001, Mengchen Zhao, Jiuxin Lin, Liyang Chen, Weihong Bao |
ICMI | 2 |
| 2022 | Speaker Characteristics Guided Speech SynthesisabstractTalking head techniques are widely researched. Most of the previous works focus on the association among tones, prosody, and visual cues, such as head motion, lip movement, and gestures. However, it is widely believed the timbre, matching the voice with the speaker's identity, shall be considered, since people obtain speaker-specific information from both the auditory and visual modalities. This paper aims to generate proper voice characteristics in line with the speaker characteristics we select. We first select six speaker characteristics related to the voice qualities: gender, age, race, body mass index, face shape, and personality. We then train a Conditional Variational AutoEncoder with attention (attentionCVAE) model to infer speaker embeddings from speaker characteristics and employ a multi-speaker text-to-speech system to generate utterances of nonexistent speakers we set. Subjective tests indicate the proposed method successfully reconstructs real-world speaker embedding and generates realistic embedding from speaker characteristics. The further analysis uncovers how and to what extent the speaker characteristics influence the voice qualities of speakers. Zhiyong Wu 0001, Jia Jia 0001 |
IJCNN | 2 |
| 2022 | Speech Enhancement with Fullband-Subband Cross-Attention Network
Jun Chen 0024, Wei Rao 0002, Zilin Wang 0002, Zhiyong Wu 0001, Yannan Wang, Shidong Shang, Helen M. Meng |
INTERSPEECH | 4 |
| 2022 | Improving Mandarin Prosodic Structure Prediction with Multi-level Contextual InformationabstractFor text-to-speech (TTS) synthesis, prosodic structure prediction (PSP) plays an important role in producing natural and intelligible speech. Although inter-utterance linguistic information can influence the speech interpretation of the target utterance, previous works on PSP mainly focus on utilizing intrautterance linguistic information of the current utterance only. This work proposes to use inter-utterance linguistic information to improve the performance of PSP. Multi-level contextual information, which includes both inter-utterance and intrautterance linguistic information, is extracted by a hierarchical encoder from character level, utterance level and discourse level of the input text. Then a multi-task learning (MTL) decoder predicts prosodic boundaries from multi-level contextual information. Objective evaluation results on two datasets show that our method achieves better F1 scores in predicting prosodic word (PW), prosodic phrase (PPH) and intonational phrase (IPH). It demonstrates the effectiveness of using multi-level contextual information for PSP. Subjective preference tests also indicate the naturalness of synthesized speeches are improved. Changhe Song, Deyi Tuo, Xixin Wu, Shiyin Kang, Zhiyong Wu 0001, Helen M. Meng |
INTERSPEECH | 6 |
| 2022 | Towards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech SynthesisabstractPrevious works on expressive speech synthesis focus on modelling the mono-scale style embedding from the current sentence or context, but the multi-scale nature of speaking style in human speech is neglected.In this paper, we propose a multiscale speaking style modelling method to capture and predict multi-scale speaking style for improving the naturalness and expressiveness of synthetic speech.A multi-scale extractor is proposed to extract speaking style embeddings at three different levels from the ground-truth speech, and explicitly guide the training of a multi-scale style predictor based on hierarchical context information.Both objective and subjective evaluations on a Mandarin audiobooks dataset demonstrate that our proposed method can significantly improve the naturalness and expressiveness of the synthesized speech 1 . Shun Lei, Yixuan Zhou 0002, Liyang Chen, Jiankun Hu, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng |
INTERSPEECH | 5 |
| 2022 | Towards Cross-speaker Reading Style Transfer on Audiobook DatasetabstractCross-speaker style transfer aims to extract the speech style of the given reference speech, which can be reproduced in the timbre of arbitrary target speakers.Existing methods on this topic have explored utilizing utterance-level style labels to perform style transfer via either global or local scale style representations.However, audiobook datasets are typically characterized by both the local prosody and global genre, and are rarely accompanied by utterance-level style labels.Thus, properly transferring the reading style across different speakers remains a challenging task.This paper aims to introduce a chunk-wise multi-scale cross-speaker style model to capture both the global genre and the local prosody in audiobook speeches.Moreover, by disentangling speaker timbre and style with the proposed switchable adversarial classifiers, the extracted reading style is made adaptable to the timbre of different speakers.Experiment results confirm that the model manages to transfer a given reading style to new target speakers.With the support of local prosody and global genre type predictor, the potentiality of the proposed method in multi-speaker audiobook generation is further revealed. Xiang Li 0105, Changhe Song, Xianhao Wei, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng |
INTERSPEECH | 4 |
| 2022 | CALM: Constrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis
Xiang Li 0105, Zhiyong Wu 0001, Tingtian Li, Zixun Sun, Xinyu Xiao, Chi Sun, Hui Zhan, Helen M. Meng |
INTERSPEECH | 3 |
| 2022 | Speech Representation Disentanglement with Adversarial Mutual Information Learning for One-shot Voice Conversion
Methawee Tantrawenith, Haolin Zhuang, Zhiyong Wu 0001, Aolan Sun, Jianzong Wang, Ning Cheng 0001, Huaizhen Tang, Xintao Zhao, Helen M. Meng |
INTERSPEECH | 4 |
| 2022 | MFA-Conformer: Multi-scale Feature Aggregation Conformer for Automatic Speaker VerificationabstractIn this paper, we present Multi-scale Feature Aggregation Conformer (MFA-Conformer), an easy-to-implement, simple but effective backbone for automatic speaker verification based on the Convolution-augmented Transformer (Conformer).The architecture of the MFA-Conformer is inspired by recent stateof-the-art models in speech recognition and speaker verification.Firstly, we introduce a convolution subsampling layer to decrease the computational cost of the model.Secondly, we adopt Conformer blocks which combine Transformers and convolution neural networks (CNNs) to capture global and local features effectively.Finally, the output feature maps from all Conformer blocks are concatenated to aggregate multi-scale representations before final pooling.We evaluate the MFA-Conformer on the widely used benchmarks.The best system obtains 0.64%, 1.29% and 1.63% EER on VoxCeleb1-O, SITW.Dev, and SITW.Eval set, respectively.MFA-Conformer significantly outperforms the popular ECAPA-TDNN systems in both recognition performance and inference speed.Last but not the least, the ablation studies clearly demonstrate that the combination of global and local feature learning can lead to robust and accurate speaker embedding extraction.We have also released the code 1 for future comparison. Yang Zhang 0025, Zhiqiang Lv, Pengfei Hu 0004, Zhiyong Wu 0001, Hung-yi Lee, Helen M. Meng |
INTERSPEECH | 6 |
| 2022 | Towards Improving the Expressiveness of Singing Voice Synthesis with BERT Derived Semantic InformationabstractThis paper presents an end-to-end high-quality singing voice synthesis (SVS) system that uses bidirectional encoder representation from Transformers (BERT) derived semantic embeddings to improve the expressiveness of the synthesized singing voice.Based on the main architecture of recently proposed VISinger, we put forward several specific designs for expressive singing voice synthesis.First, different from the previous SVS models, we use text representation of lyrics extracted from pre-trained BERT as additional input to the model.The representation contains information about semantics of the lyrics, which could help SVS system produce more expressive and natural voice.Second, we further introduce an energy predictor to stabilize the synthesized voice and model the wider range of energy variations that also contribute to the expressiveness of singing voice.Last but not the least, to attenuate the off-key issues, the pitch predictor is re-designed to predict the real to note pitch ratio.Both objective and subjective experimental results indicate that the proposed SVS system can produce singing voice with higher-quality outperforming VISinger 1 . Shaohuan Zhou, Shun Lei, Weiya You, Deyi Tuo, Yuren You, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng |
INTERSPEECH | 6 |
| 2022 | Enhancing Word-Level Semantic Representation via Dependency Structure for Expressive Text-to-Speech SynthesisabstractExploiting rich linguistic information in raw text is crucial for expressive text-to-speech (TTS).As large scale pre-trained text representation develops, bidirectional encoder representations from Transformers (BERT) has been proven to embody semantic information and employed to TTS recently.However, original or simply fine-tuned BERT embeddings still cannot provide sufficient semantic knowledge that expressive TTS models should take into account.In this paper, we propose a wordlevel semantic representation enhancing method based on dependency structure and pre-trained BERT embedding.The BERT embedding of each word is reprocessed considering its specific dependencies and related words in the sentence, to generate more effective semantic representation for TTS.To better utilize the dependency structure, relational gated graph network (RGGN) is introduced to make semantic information flow and aggregate through the dependency structure.The experimental results show that the proposed method can further improve the naturalness and expressiveness of synthesized speeches on both Mandarin and English datasets 1 . Yixuan Zhou 0002, Changhe Song, Jingbei Li, Zhiyong Wu 0001, Yanyao Bian, Dan Su 0002, Helen M. Meng |
INTERSPEECH | 4 |
| 2022 | Content-Dependent Fine-Grained Speaker Embedding for Zero-Shot Speaker Adaptation in Text-to-Speech SynthesisabstractZero-shot speaker adaptation aims to clone an unseen speaker's voice without any adaptation time and parameters.Previous researches usually use a speaker encoder to extract a global fixed speaker embedding from reference speech, and several attempts have tried variable-length speaker embedding.However, they neglect to transfer the personal pronunciation characteristics related to phoneme content, leading to poor speaker similarity in terms of detailed speaking styles and pronunciation habits.To improve the ability of the speaker encoder to model personal pronunciation characteristics, we propose content-dependent fine-grained speaker embedding for zero-shot speaker adaptation.The corresponding local content embeddings and speaker embeddings are extracted from a reference speech, respectively.Instead of modeling the temporal relations, a reference attention module is introduced to model the content relevance between the reference speech and the input text, and to generate the finegrained speaker embedding for each phoneme encoder output.The experimental results show that our proposed method can improve speaker similarity of synthesized speeches, especially for unseen speakers. Yixuan Zhou 0002, Changhe Song, Xiang Li 0105, Zhiyong Wu 0001, Yanyao Bian, Dan Su 0002, Helen M. Meng |
INTERSPEECH | 5 |
| 2022 | Inferring Speaking Styles from Multi-modal Conversational Context by Multi-scale Relational Graph Convolutional NetworksabstractTo support applications of speech-driven interactive systems in various conversational scenarios, text-to-speech (TTS) synthesis needs to understand the conversational context and determine appropriate speaking styles in its synthesized speeches. These speaking styles are influenced by the dependencies between the multi-modal information in the context at both global scale (i.e. utterance level) and local scale (i.e. word level). However, the dependency modeling and speaking style inference at the local scale are largely missing in state-of-the-art TTS systems, resulting in the synthesis of incorrect or improper speaking styles. In this paper, to learn the dependencies in conversations at both global and local scales and to improve the synthesis of speaking styles, we propose a context modeling method which models the dependencies among the multi-modal information in context with multi-scale relational graph convolutional network (MSRGCN). The learnt multi-modal context information at multiple scales is then utilized to infer the global and local speaking styles of the current utterance for speech synthesis. Experiments demonstrate the effectiveness of the proposed approach, and ablation studies reflect the contributions from modeling multi-modal information and multi-scale dependencies. Jingbei Li, Xixin Wu, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng, Qiao Tian 0001, Yuping Wang 0005, Yuxuan Wang 0002 |
ACM Multimedia | 4 |
| 2022 | Disentangled Speech Representation Learning for One-Shot Cross-Lingual Voice Conversion Using ß-VAEabstractWe propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot cross-lingual voice conversion task to demonstrate the effectiveness of the disentanglement. Inspired by ß- VAE, we introduce a learning objective that balances between the information captured by the content and speaker representations. In addition, the inductive biases from the architectural design and the training dataset further encourage the desired disentanglement. Both objective and subjective evaluations show the effectiveness of the proposed method in speech disentanglement and in one-shot cross-lingual voice conversion. Disong Wang, Xixin Wu, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
SLT | 4 |
| 2022 | Improving the Adversarial Robustness for Speaker Verification by Self-Supervised LearningabstractPrevious works have shown that automatic speaker verification (ASV) is seriously vulnerable to malicious spoofing attacks, such as replay, synthetic speech, and recently emerged adversarial attacks. Great efforts have been dedicated to defending ASV against replay and synthetic speech; however, only a few approaches have been explored to deal with adversarial attacks. All the existing approaches to tackle adversarial attacks for ASV require the knowledge for adversarial samples generation, but it is impractical for defenders to know the exact attack algorithms that are applied by the in-the-wild attackers. This work is among the first to perform adversarial defense for ASV without knowing the specific attack algorithms. Inspired by self-supervised learning models (SSLMs) that possess the merits of alleviating the superficial noise in the inputs and reconstructing clean samples from the interrupted ones, this work regards adversarial perturbations as one kind of noise and conducts adversarial defense for ASV by SSLMs. Specifically, we propose to perform adversarial defense from two perspectives: 1) adversarial perturbation purification and 2) adversarial perturbation detection. The purification module aims at alleviating the adversarial perturbations in the samples and pulling the contaminated adversarial inputs back towards the decision boundary. Experimental results show that our proposed purification module effectively counters adversarial attacks and outperforms traditional filters from both alleviating the adversarial noise and maintaining the performance of genuine samples. The detection module aims at detecting adversarial samples from genuine ones based on the statistical properties of ASV scores derived by a unique ASV integrating with different number of SSLMs. Experimental results show that our detection module helps shield the ASV by detecting adversarial samples. Both purification and detection methods are helpful for defending against different kinds of attack algorithms. Moreover, since there is no common metric for evaluating the ASV performance under adversarial attacks, this work also formalizes evaluation metrics for adversarial defense considering both purification and detection based approaches into account. We sincerely encourage future works to benchmark their approaches based on the proposed evaluation framework. Xu Li 0015, Andy T. Liu, Zhiyong Wu 0001, Helen M. Meng, Hung-yi Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Inferring Emotion from Large-scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning ApproachabstractEffective emotion inference from user queries helps to give a more personified response for Voice Dialogue Applications(VDAs). The tremendous amounts of VDA users bring in diverse emotion expressions. How to achieve a high emotion inferring performance from large-scale Internet Voice Data in VDAs? Traditionally, researches on speech emotion recognition are based on acted voice datasets, which have limited speakers but strong and clear emotion expressions. Inspired by this, in this paper, we propose a novel approach to leverage acted voice data with strong emotion expressions to enhance large-scale unlabeled internet voice data with diverse emotion expressions for emotion inferring. Specifically, we propose a novel semi-supervised multi-modal curriculum augmentation deep learning framework. First, to learn more general emotion cues, we adopt a curriculum learning based epoch-wise training strategy, which trains our model guided by strong and balanced emotion samples from acted voice data and sub-sequently leverages weak and unbalanced emotion samples from internet voice data.Second, to employ more diverse emotion expressions, we design a Multi-path Mix-match Multimodal Deep Neural Network(MMMD), which effectively learns feature representations for multiple modalities and trains labeled and unlabeled data in hybrid semi-supervised methods for superior generalization and robustness. Experiments on an internet voice dataset with 500,000 utterances show our method outperforms (+10.09% in terms of F1) several alternative baselines, while an acted corpus with 2,397 utterances contributes 4.35%. To further compare our method with state-of-the-art techniques in traditionally acted voice datasets, we also conduct experiments on public dataset IEMOCAP. The results reveal the effectiveness of the proposed approach. Suping Zhou, Jia Jia 0001, Zhiyong Wu 0001, Wei Chen 0071, Shuo Huang 0005, Jialie Shen 0001 |
AAAI | 3 |
| 2021 | Reconstructing Dual Learning for Neural Voice Conversion Using Relatively Few SamplesabstractThis paper introduces a dual learning system for neural voice conversion (DualVC) using relatively few samples based on the symmetry of the speech conversion task. The system contains a pair of sequence-to-sequence neural networks that have the same structure but are trained in opposite directions. The objective function of the dual model training is the sum of paired conversion loss and reconstruction loss during the dual training circle. The models in the two directions are trained alternately to guide each other by the corresponding reconstruction loss. Furthermore, curriculum learning techniques are used to load models in existing fields into the current task to accelerate the rapid iteration and convergence of the model. The experiment on the voice conversion task with the proposed DualVC and curriculum learning strategy obtained a comparable naturalness and similarity with only a 30% dataset than the BaseVC model trained on the full dataset. Aolan Sun, Jianzong Wang, Ning Cheng 0001, Methawee Tantrawenith, Zhiyong Wu 0001, Helen M. Meng, Edward Xiao, Jing Xiao 0006 |
ASRU | 5 |
| 2021 | PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective FeedbackabstractSecond language (L2) English learners often find it difficult to improve their pronunciations due to the lack of expressive and personalized corrective feedback. In this paper, we present Pronunciation Teacher (PTeacher), a Computer-Aided Pronunciation Training (CAPT) system that provides personalized exaggerated audio-visual corrective feedback for mispronunciations. Though the effectiveness of exaggerated feedback has been demonstrated, it is still unclear how to define the appropriate degrees of exaggeration when interacting with individual learners. To fill in this gap, we interview 100 L2 English learners and 22 professional native teachers to understand their needs and experiences. Three critical metrics are proposed for both learners and teachers to identify the best exaggeration levels in both audio and visual modalities. Additionally, we incorporate the personalized dynamic feedback mechanism given the English proficiency of learners. Based on the obtained insights, a comprehensive interactive pronunciation training course is designed to help L2 learners rectify mispronunciations in a more perceptible, understandable, and discriminative manner. Extensive user studies demonstrate that our system significantly promotes the learners’ learning efficiency. Yaohua Bu, Hang Zhou 0009, Jia Jia 0001, Shengqi Chen 0001, Dachuan Shi, Haozhe Wu, Kun Li 0003, Zhiyong Wu 0001, Yuanchun Shi, Xiaobo Lu, Ziwei Liu 0002 |
CHI | 12 |
| 2021 | Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language UnderstandingabstractLack of training data presents a grand challenge to scaling out spoken language understanding (SLU) to low-resource languages.Although various data augmentation approaches have been proposed to synthesize training data in low-resource target languages, the augmented data sets are often noisy, and thus impede the performance of SLU models.In this paper we focus on mitigating noise in augmented data.We develop a denoising training approach.Multiple models are trained with data produced by various augmented methods.Those models provide supervision signals to each other.The experimental results show that our method outperforms the existing state of the art by 3.05 and 4.24 percentage points on two benchmark datasets, respectively.The code will be made open sourced on github. Yingmei Guo, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Mingxing Xu, Zhiyong Wu 0001, Daxin Jiang |
EMNLP (1) | 6 |
| 2021 | Emotion Controllable Speech Synthesis Using Emotion-Unlabeled Dataset with the Assistance of Cross-Domain Speech Emotion RecognitionabstractNeural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach for emotional TTS synthesis on a TTS dataset without emotion labels. Specifically, our proposed method consists of a cross-domain speech emotion recognition (SER) model and an emotional TTS model. Firstly, we train the cross-domain SER model on both SER and TTS datasets. Then, we use emotion labels on the TTS dataset predicted by the trained SER model to build an auxiliary SER task and jointly train it with the TTS model. Experimental results show that our proposed method can generate speech with the specified emotional expressiveness and nearly no hindering on the speech quality. Xiong Cai, Dongyang Dai, Zhiyong Wu 0001, Xiang Li 0105, Jingbei Li, Helen M. Meng |
ICASSP | 3 |
| 2021 | Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder InputabstractNon-autoregressive (NAR) transformer models have achieved significantly inference speedup but at the cost of inferior accuracy compared to autoregressive (AR) models in automatic speech recognition (ASR). Most of the NAR transformers take a fixed-length sequence filled with MASK tokens or a redundant sequence copied from encoder states as decoder input, they cannot provide efficient target-side information thus leading to accuracy degradation. To address this problem, we propose a CTC-enhanced NAR transformer, which generates target sequence by refining predictions of the CTC module. Experimental results show that our method outperforms all previous NAR counterparts and achieves 50x faster decoding speed than a strong AR baseline with only 0.0∼ 0.3 absolute CER degradation on Aishell-1 and Aishell-2 datasets. Xingchen Song, Zhiyong Wu 0001, Chao Weng, Dan Su 0002, Helen M. Meng |
ICASSP | 2 |
| 2021 | Syntactic Representation Learning For Neural Network Based TTS with Syntactic Parse Tree TraversalabstractSyntactic structure of a sentence text is correlated with the prosodic structure of the speech that is crucial for improving the prosody and naturalness of a text-to-speech (TTS) system. Nowadays TTS systems usually try to incorporate syntactic structure information with manually designed features based on expert knowledge. In this paper, we propose a syntactic representation learning method based on syntactic parse tree traversal to automatically utilize the syntactic structure information. Two constituent label sequences are linearized through left-first and right-first traversals from constituent parse tree. Syntactic representations are then extracted at word level from each constituent label sequence by a corresponding uni-directional gated recurrent unit (GRU) network. Meanwhile, nuclear-norm maximization loss is introduced to enhance the discriminability and diversity of the embeddings of constituent labels. Upsampled syntactic representations and phoneme embeddings are concatenated to serve as the encoder input of Tacotron2. Experimental results demonstrate the effectiveness of our proposed approach, with mean opinion score (MOS) increasing from 3.70 to 3.82 and ABX preference exceeding by 17% compared with the baseline. In addition, for sentences with multiple syntactic parse trees, prosodic differences can be clearly perceived from the synthesized speeches. Changhe Song, Jingbei Li, Yixuan Zhou 0002, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 4 |
| 2021 | Improving Pronunciation Assessment Via Ordinal Regression with Anchored Reference SamplesabstractSentence level pronunciation assessment is important for Computer Assisted Language Learning (CALL). Traditional speech pronunciation assessment, based on the Goodness of Pronunciation (GOP) algorithm, has some weakness in assessing a speech utterance: 1) Phoneme GOP scores cannot be easily translated into a sentence score with a simple average for effective assessment; 2) The rank ordering information has not been well exploited in GOP scoring for delivering a robust assessment and correlate well with a human rater’s evaluations. In this paper, we propose two new statistical features, average GOP (aGOP) and confusion GOP (cGOP) and use them to train a binary classifier in Ordinal Regression with Anchored Reference Samples (ORARS). When the proposed approach is tested on Microsoft mTutor ESL Dataset, a relative improvement of Pearson correlation coefficient of 26.9% is obtained over the conventional GOP-based one. The performance is at a human-parity level or better than human raters. Shaoguang Mao, Frank K. Soong, Yan Xia 0005, Jonathan Tien, Zhiyong Wu 0001 |
ICASSP | 6 |
| 2021 | The Huya Multi-Speaker and Multi-Style Speech Synthesis System for M2voc Challenge 2020abstractText-to-speech systems now can generate speech that is hard to distinguish from human speech. In this paper, we propose the Huya multi-speaker and multi-style speech synthesis system which is based on DurIAN and HiFi-GAN to generate high-fidelity speech even under low-resource condition. We use the fine-grained linguistic representation which leverages the similarity in pronunciation between different languages and promotes the speech quality of code-switch speech synthesis. Our TTS system uses the HiFi-GAN as the neural vocoder which has higher synthesis stability for unseen speakers and can generate higher quality speech with noisy training data than WaveRNN in the challenge tasks. The model is trained on the datasets released by the organizer as well as CMU-ARCTIC, AIShell-1 and THCHS-30 as the external datasets and the results were evaluated by the organizer. We participated in all four tracks and three of them entered high score lists. The evaluation results show that our system outperforms the majority of all participating teams. Yuren You, Deyi Tuo, Shiyin Kang, Zhiyong Wu 0001, Helen M. Meng |
ICASSP | 6 |
| 2021 | Adversarial Defense for Automatic Speaker Verification by Cascaded Self-Supervised Learning ModelsabstractAutomatic speaker verification (ASV) is one of the core technologies in biometric identification. With the ubiquitous usage of ASV systems in safety-critical applications, more and more malicious attackers attempt to launch adversarial attacks at ASV systems. In the midst of the arms race between attack and defense in ASV, how to effectively improve the robustness of ASV against adversarial attacks remains an open question. We note that the self-supervised learning models possess the ability to mitigate superficial perturbations in the input after pretraining. Hence, with the goal of effective defense in ASV against adversarial attacks, we propose a standard and attack-agnostic method based on cascaded self-supervised learning models to purify the adversarial perturbations. Experimental results demonstrate that the proposed method achieves effective defense performance and can successfully counter adversarial attacks in scenarios where attackers may either be aware or unaware of the self-supervised learning models. Xu Li 0015, Andy T. Liu, Zhiyong Wu 0001, Helen M. Meng, Hung-yi Lee |
ICASSP | 4 |
| 2021 | The Multi-Speaker Multi-Style Voice Cloning Challenge 2021abstractThe Multi-speaker Multi-style Voice Cloning Challenge (M2VoC) aims to provide a common sizable dataset as well as a fair testbed for the benchmarking of the popular voice cloning task. Specifically, we formulate the challenge to adapt an average TTS model to the stylistic target voice with limited data from target speaker, evaluated by speaker identity and style similarity. The challenge consists of two tracks, namely few-shot track and one-shot track, where the participants are required to clone multiple target voices with 100 and 5 samples respectively. There are also two sub-tracks in each track. For sub-track a, to fairly compare different strategies, the participants are allowed to use only the training data provided by the organizer strictly. For sub-track b, the participants are allowed to use any data publicly available. In this paper, we present a detailed explanation on the tasks and data used in the challenge, followed by a summary of submitted systems and evaluation results. Qicong Xie, Xiaohai Tian, Guanghou Liu, Lei Xie 0001, Zhiyong Wu 0001, Haizhou Li 0001, Fen Hong, Hui Bu |
ICASSP | 6 |
| 2021 | Towards Multi-Scale Style Control for Expressive Speech SynthesisabstractThis paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis.The proposed method employs a multi-scale reference encoder to extract both the global-scale utterance-level and the local-scale quasi-phoneme-level style features of the target speech, which are then fed into the speech synthesis model as an extension to the input phoneme sequence.During training time, the multiscale style model could be jointly trained with the speech synthesis model in an end-to-end fashion.By applying the proposed method to style transfer task, experimental results indicate that the controllability of the multi-scale speech style model and the expressiveness of the synthesized speech are greatly improved.Moreover, by assigning different reference speeches to extraction of style on each scale, the flexibility of the proposed method is further revealed. Xiang Li 0105, Changhe Song, Jingbei Li, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng |
Interspeech | 4 |
| 2021 | VAENAR-TTS: Variational Auto-Encoder Based Non-AutoRegressive Text-to-Speech SynthesisabstractThis paper describes a variational auto-encoder based nonautoregressive text-to-speech (VAENAR-TTS) model.The autoregressive TTS (AR-TTS) models based on the sequenceto-sequence architecture can generate high-quality speech, but their sequential decoding process can be time-consuming.Recently, non-autoregressive TTS (NAR-TTS) models have been shown to be more efficient with the parallel decoding process.However, these NAR-TTS models rely on phoneme-level durations to generate a hard alignment between the text and the spectrogram.Obtaining duration labels, either through forced alignment or knowledge distillation, is cumbersome.Furthermore, hard alignment based on phoneme expansion can degrade the naturalness of the synthesized speech.In contrast, the proposed model of VAENAR-TTS is an end-to-end approach that does not require phoneme-level durations.The VAENAR-TTS model does not contain recurrent structures and is completely non-autoregressive in both the training and inference phases.Based on the VAE architecture, the alignment information is encoded in the latent variable, and attention-based soft alignment between the text and the latent variable is used in the decoder to reconstruct the spectrogram.Experiments show that VAENAR-TTS achieves state-of-the-art synthesis quality, while the synthesis speed is comparable with other NAR-TTS models. Zhiyong Wu 0001, Xixin Wu, Xu Li 0015, Shiyin Kang, Xunying Liu, Helen M. Meng |
Interspeech | 2 |
| 2021 | Adversarially Learning Disentangled Speech Representations for Robust Multi-Factor Voice ConversionabstractFactorizing speech as disentangled speech representations is vital to achieve highly controllable style transfer in voice conversion (VC).Conventional speech representation learning methods in VC only factorize speech as speaker and content, lacking controllability on other prosody-related factors.State-ofthe-art speech representation learning methods for more speech factors are using primary disentangle algorithms such as random resampling and ad-hoc bottleneck layer size adjustment, which however is hard to ensure robust speech representation disentanglement.To increase the robustness of highly controllable style transfer on multiple factors in VC, we propose a disentangled speech representation learning framework based on adversarial learning.Four speech representations characterizing content, timbre, rhythm and pitch are extracted, and further disentangled by an adversarial Mask-And-Predict (MAP) network inspired by BERT.The adversarial network is used to minimize the correlations between the speech representations, by randomly masking and predicting one of the representations from the others.Experimental results show that the proposed framework significantly improves the robustness of VC on multiple factors by increasing the speech quality MOS from 2.79 to 3.30 and decreasing the MCD from 3.89 to 3.58. Jingbei Li, Xintao Zhao, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng |
Interspeech | 4 |
| 2021 | Voting for the Right Answer: Adversarial Defense for Speaker VerificationabstractAutomatic speaker verification (ASV) is a well developed technology for biometric identification, and has been ubiquitous implemented in security-critic applications, such as banking and access control.However, previous works have shown that ASV is under the radar of adversarial attacks, which are very similar to their original counterparts from human's perception, yet will manipulate the ASV render wrong prediction.Due to the very late emergence of adversarial attacks for ASV, effective countermeasures against them are limited.Given that the security of ASV is of high priority, in this work, we propose the idea of "voting for the right answer" to prevent risky decisions of ASV in blind spot areas, by employing random sampling and voting.Experimental results show that our proposed method improves the robustness against both the limited-knowledge attackers by pulling the adversarial samples out of the blind spots, and the sufficient-knowledge attackers by introducing randomness and increasing the attackers' budgets. Yang Zhang 0025, Zhiyong Wu 0001, Hung-yi Lee |
Interspeech | 3 |
| 2021 | Controllable Emphatic Speech Synthesis based on Forward Attention for Expressive Speech SynthesisabstractIn speech interaction scenarios, speech emphasis is essential for expressing the underlying intention and attitude. Recently, end-to-end emphatic speech synthesis greatly improves the naturalness of synthetic speech, but also brings new problems: 1) lack of interpretability for how emphatic codes affect the model; 2) no separate control of emphasis on duration and on intonation and energy. We propose a novel way to build an interpretable and controllable emphatic speech synthesis framework based on forward attention. Firstly, we explicitly model the local variation of speaking rate for emphasized words and neutral words with modified forward attention to manifest emphasized words in terms of duration. The 2-layers LSTM in decoder is further divided into attention-RNN and decoder-RNN to disentangle the influence of emphasis on duration and on intonation and energy. The emphasis information is injected into decoder-RNN for highlighting emphasized words in the aspects of intonation and energy. Experimental results have shown that our model can not only provide separate control of emphasis on duration and on intonation and energy, but also generate more robust and prominent emphatic speech with high quality and naturalness. Liangqi Liu, Jiankun Hu, Zhiyong Wu 0001, Songfan Yang, Jia Jia 0001, Helen M. Meng |
SLT | 3 |
| 2021 | Exemplar-Based Emotive Speech SynthesisabstractExpressive text-to-speech (E-TTS) synthesis is important for enhancing user experience in communication with machines using the speech modality. However, one of the challenges in E-TTS is the lack of a precise description of emotions. Previous categorical specifications may be insufficient for describing complex emotions. The dimensional specifications face the difficulty of ambiguity in annotation. This work advocates a new approach of describing emotive speech acoustics using spoken exemplars. We investigate methods to extract emotion descriptions from the input exemplar of emotive speech. The measures are combined to form two descriptors, based on capsule network (CapNet) and residual error network (RENet). The first is designed to consider the spatial information in the input exemplary spectrogram, and the latter is to capture the contrastive information between emotive acoustic expressions. Two different approaches are applied for conversion from the variable-length feature sequence to fixed-size description vector: (1) dynamic routing groups similar capsules to the output description; and (2) recurrent neural network's hidden states store the temporal information for the description. The two descriptors are integrated to a state-of-the-art sequence-to-sequence architecture to obtain an end-to-end architecture that is optimized as a whole towards the same goal of generating correct emotive speech. Experimental results on a public audiobook dataset demonstrate that the two exemplar-based approaches achieve significant performance improvement over the baseline system in both emotion similarity and speech quality. Xixin Wu, Yuewen Cao, Songxiang Liu, Shiyin Kang, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2021 | Speech Emotion Recognition Using Sequential Capsule NetworksabstractSpeech emotion recognition (SER) is an indispensable part of fluid human-machine interaction and attracts lots of research attentions. Recent work on SER has successfully applied convolutional neural networks (CNNs) to learn feature representations from speech spectrograms. However, the fundamental problem of CNNs is that the spatial information in spectrograms is lost, which includes positional and relationship information of low-level features, such as pitch and formant frequencies. We propose a novel architecture of sequential capsule networks (CapNets) by leveraging the advantange of CapNets that spatial information can be preserved in capsules and passed to upper capsule layers via dynamic routing. Also, the dynamic routing algorithm provides an effective alternative to pooling or storing recurrent hidden states for obtaining utterance-level features from the sequential capsule outputs. To further improve the model's ability to capture contextual information, we introduce a recurrent connection to the sequential structure. The experimental comparison of the proposed systems and previously published systems using CNNs and recurrent neural networks (RNNs) based on the IEMOCAP corpus demonstrates the effectiveness of the proposed sequential CapNets. Xixin Wu, Yuewen Cao, Songxiang Liu, Disong Wang, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | Code-Switched Speech Synthesis Using Bilingual Phonetic Posteriorgram with Only Monolingual CorporaabstractSynthesizing fluent code-switched (CS) speech with consistent voice using only monolingual corpora is still a challenging task, since language alternation seldom occurs during training and the speaker identity is directly correlated with language. In this paper, we present a bilingual phonetic posteriorgram (PPG) based CS speech synthesizer using only monolingual corpora. The bilingual PPG is used to bridge across speakers and languages, which is formed by stacking two monolingual PPGs extracted from two monolingual speaker-independent speech recognition systems. It is assumed that bilingual PPG can represent the articulation of speech sounds speaker-independently and captures accurate phonetic information of both languages in the same feature space. The proposed model first extracts bilingual PPGs from training data. Then an encoder- decoder based model is used to learn the relationship between input text and bilingual PPGs, and the bilingual PPGs are mapped to acoustic features using bidirectional long-short term memory based model conditioned on speaker embedding to control speaker identity. Experiments validate the effectiveness of the proposed model in terms of speech intelligibility, audio fidelity and speaker consistency of the generated code-switched speech. Yuewen Cao, Songxiang Liu, Xixin Wu, Shiyin Kang, Zhiyong Wu 0001, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng |
ICASSP | 6 |
| 2020 | End-To-End Accent Conversion Without Using Native UtterancesabstractTechniques for accent conversion (AC) aim to convert non-native to native accented speech. Conventional AC methods try to convert only the speaker identity of a native speaker's voice to that of the non-native accented target speaker, leaving the underlying content and pronunciations unchanged. This hinders their practical use in real-world applications, because native-accented utterances are required at conversion stage. In this paper, we present an end-to-end framework, which is able to conduct AC from non-native-accented utterances without using any native-accented utterances during online conversion. We achieve this by independently extracting linguistic and speaker representations from non-native accented speech and condition a speech synthesis model on these representations to generate native-accented speech. Experiments on open-source data corpora show that the proposed system can convert Hindi-accented English speech into native American English speech with high naturalness, which is indistinguishable from native-accented recordings in terms of accent. Songxiang Liu, Disong Wang, Yuewen Cao, Lifa Sun, Xixin Wu, Shiyin Kang, Zhiyong Wu 0001, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng |
ICASSP | 7 |
| 2020 | Channel-Wise Dense Connection Graph Convolutional Network for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition task has drawn much attention for many years. Graph Convolutional Network (GCN) has proved its effectiveness in this task. However, how to improve the model's robustness to different human actions and how to make effective use of features produced by the network are main topics needed to be further explored. Human actions are time series sequence, meaning that temporal information is a key factor to model the representation of data. The ranges of body parts involved in small actions (e.g. raise a glass or shake head) and big actions (e.g. walking or jumping) are diverse. It's crucial for the model to generate and utilize more features that can be adaptive to a wider range of actions. Furthermore, feature channels are specific with the action class, the model needs to weigh their importance and pay attention to more related ones. To address these problems, in this work, we propose a two-stream channel-wise dense connection GCN (2s-CDGCN). Specifically, the skeleton data was extracted and processed into spatial and temporal information for better feature representation. A channel-wise attention module was used to select and emphasize the more useful features generated by the network. Moreover, to ensure maximum information flow, dense connection was introduced to the network structure, which enables the network to reuse the skeleton features and generate more information adaptive and related to different human actions. Our model has shown its ability to improve the accuracy of human action recognition task on two large datasets, NTU-RGB+D and Kinetics. Extensive evaluations were conducted to prove the effectiveness of our model. Michael Lao BanTeng, Zhiyong Wu 0001 |
ICPR | 2 |
| 2020 | Enhancing Monotonicity for Robust Autoregressive Transformer TTS
Xiangyu Liang, Zhiyong Wu 0001, Runnan Li, Sheng Zhao 0002, Helen M. Meng |
INTERSPEECH | 2 |
| 2020 | SpecSwap: A Simple Data Augmentation Method for End-to-End Speech Recognition
Xingcheng Song, Zhiyong Wu 0001, Dan Su 0002, Helen M. Meng |
INTERSPEECH | 2 |
| 2020 | Speech-XLNet: Unsupervised Acoustic Model Pretraining for Self-Attention NetworksabstractSelf-attention network (SAN) can benefit significantly from the bi-directional representation learning through unsupervised pretraining paradigms such as BERT and XLNet.In this paper, we present an XLNet-like pretraining scheme "Speech-XLNet" to learn speech representations with self-attention networks (SANs).Firstly, we find that by shuffling the speech frame orders, Speech-XLNet serves as a strong regularizer which encourages the SAN network to make inferences by focusing on global structures through its attention weights.Secondly, Speech-XLNet also allows the model to explore bi-directional context information while maintaining the autoregressive training manner.Visualization results show that our approach can generalize better with more flattened and widely distributed optimas compared to the conventional approach.Experimental results on TIMIT demonstrate that Speech-XLNet greatly improves hybrid SAN/HMM in terms of both convergence speed and recognition accuracy.Our best systems achieve a relative improvement of 15.2% on the TIMIT task.Besides, we also apply our pretrained model to an End-to-End SAN with WSJ dataset and WER is reduced by up to 68% when only a few hours of transcribed data is used. Xingchen Song, Guangsen Wang, Zhiyong Wu 0001, Dan Su 0002, Helen M. Meng |
INTERSPEECH | 4 |
| 2020 | Re-Weighted Interval Loss for Handling Data Imbalance Problem of End-to-End Keyword Spotting
Zhiyong Wu 0001, Daode Yuan, Jian Luan 0001, Jia Jia 0001, Helen M. Meng, Binheng Song |
INTERSPEECH | 2 |
| 2019 | End-to-end Code-switched TTS with Mix of Monolingual RecordingsabstractState-of-the-art text-to-speech (TTS) synthesis models can produce monolingual speech with high intelligibility and naturalness. However, when the models are applied to synthesize code-switched (CS) speech, the performance declines seriously. Conventionally, developing a CS TTS system requires multilingual data to incorporate language-specific and cross-lingual knowledge. Recently, end-to-end (E2E) architecture has achieved satisfactory results in monolingual TTS. The architecture enables the training from one end of alphabetic text input to the other end of acoustic feature output. In this paper, we explore the use of E2E framework for CS TTS, using a combination of Mandarin and English monolingual speech corpus uttered by two female speakers. To handle alphabetic input from different languages, we explore two kinds of encoders: (1) shared multilingual encoder with explicit language embedding (LDE); (2) separated monolingual encoder (SPE) for each language. The two systems use identical decoder architecture, where a discriminative code is incorporated to enable the model to generate speech in one speaker's voice consistently. Experiments confirm the effectiveness of the proposed modifications on the E2E TTS framework in terms of quality and speaker similarity of the generated speech. Moreover, our proposed systems can generate controllable foreign-accented speech at character-level using only mixture of monolingual training data. Yuewen Cao, Xixin Wu, Songxiang Liu, Jianwei Yu 0001, Xu Li 0015, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 6 |
| 2019 | Learning Discriminative Features from Spectrograms Using Center Loss for Speech Emotion RecognitionabstractIdentifying the emotional state from speech is essential for the natural interaction of the machine with the speaker. However, extracting effective features for emotion recognition is difficult, as emotions are ambiguous. We propose a novel approach to learn discriminative features from variable length spectrograms for emotion recognition by cooperating soft-max cross-entropy loss and center loss together. The soft-max cross-entropy loss enables features from different emotion categories separable, and center loss efficiently pulls the features belonging to the same emotion category to their center. By combining the two losses together, the discriminative power will be highly enhanced, which leads to network learning more effective features for emotion recognition. As demonstrated by the experimental results, after introducing center loss, both the unweighted accuracy and weighted accuracy are improved by over 3% on Mel-spectrogram input, and more than 4% on Short Time Fourier Transform spectrogram input. Dongyang Dai, Zhiyong Wu 0001, Runnan Li, Xixin Wu, Jia Jia 0001, Helen M. Meng |
ICASSP | 2 |
| 2019 | Dilated Residual Network with Multi-head Self-attention for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) plays an important role in intelligent speech interaction. One vital challenge in SER is to extract emotion-relevant features from speech signals. In state-of-the-art SER techniques, deep learning methods, e.g, Convolutional Neural Networks (CNNs), are widely employed for feature learning and have achieved significant performance. However, in the CNN-oriented methods, two performance limitations have raised: 1) the loss of temporal structure of speech in the progressive resolution reduction; 2) the ignoring of relative dependencies between elements in suprasegmental feature sequence. In this paper, we proposed the combining use of Dilated Residual Network (DRN) and Multi-head Self-attention to alleviate the above limitations. By employing DRN, the network can retain high resolution of temporal structure in feature learning, with similar size of receptive field to CNN based approach. By employing Multi-head Self-attention, the network can model the inner dependencies between elements with different positions in the learned suprasegmental feature sequence, which enhances the importing of emotion-salient information. Experiments on emotional benchmarking dataset IEMOCAP have demonstrated the effectiveness of the proposed framework, with 11.7% to 18.6% relative improvement to state-of-the-art approaches. Runnan Li, Zhiyong Wu 0001, Jia Jia 0001, Sheng Zhao 0002, Helen M. Meng |
ICASSP | 2 |
| 2019 | A Compact Framework for Voice Conversion Using Wavenet Conditioned on Phonetic PosteriorgramsabstractVoice conversion can benefit from WaveNet vocoder with improvement in converted speech's naturalness and quality. However, nowadays approaches segregate the training of conversion module and WaveNet vocoder towards different optimization objectives, which might lead to the difficulty in model tuning and coordination. In this paper, we propose a compact framework to unify the conversion and the vocoder parts. Multi-head self-attention structure and bidirectional long short-term memory (BLSTM) recurrent neural network (RNN) are employed to encode speaker independent phonetic posteriorgrams (PPGs) into an intermediate representation which is used as the condition input of WaveNet to generate target speaker's waveform. In this way, we unify the conversion and vocoder parts into a compact system in which all parameters can be tuned simultaneously for global optimization. We compared the proposed method with the baseline system that consists of separately trained conversion module and WaveNet vocoder. Subjective evaluations show that the proposed method can achieve better results in both naturalness and speaker similarity. Zhiyong Wu 0001, Runnan Li, Shiyin Kang, Jia Jia 0001, Helen M. Meng |
ICASSP | 2 |
| 2019 | NN-based Ordinal Regression for Assessing Fluency of ESL SpeechabstractAutomatic assessment of a language learner's speech fluency is highly desirable for language education, e.g. for English as a Second Language (ESL) learning. In this paper, we formulate the fluency assessment as a problem of Ordinal Regression with Anchored Reference Samples (ORARS), where the fluency of a speech utterance is predicted by an ordinal regression neural network (NN) trained with anchored reference samples. The ORARS is trained and tested by: picking human expert labeled samples in each mean opinion score (MOS) bucket as the anchored reference samples and pairing them with input speech samples as training couplets; training an NN-based binary classifier to determine which sample in a pair is better in fluency; predicting the rank (MOS) of a test sample based upon the posteriors of all binary comparisons between the test sample and all anchored reference samples. Experimentally, our proposed approach outperforms the traditional NN-based methods and reaches a performance of "human parity", i.e. as comparable as human experts, in its fluency assessment of collected ESL speech. To the best of our knowledge, this is the first attempt to assess speech fluency with an ordinal regression framework where a test input is paired with bucketed and anchored reference samples. Shaoguang Mao, Zhiyong Wu 0001, Jingshuai Jiang, Peiyun Liu, Frank K. Soong |
ICASSP | 2 |
| 2019 | Quasi-fully Convolutional Neural Network with Variational Inference for Speech SynthesisabstractRecurrent neural networks, such as gated recurrent units (GRUs) and long short-term memory (LSTM), are widely used on acoustic modeling for speech synthesis. However, such sequential generating processes are not friendly to today’s massively parallel computing devices. We introduce a fully convolutional neural network (CNN) model, which can effiently run on parallel processers, for speech synthesis. To improve the quality of the generated acoustic features, we strengthen our model with variational inference. We also use quasi-recurrent neural networks (QRNNs) to smoothen the generated acoustic features. Finally, a high-quality parallel WaveNet model is used to generate audio samples. Our contributions are twofold. First, we show that CNNs with variational inference can generate highly natural speech on a par with end-to-end models; the use of QRNNs further improves the synthetic quality by reducing trembling of generated acoustic features and introduces very little run-time overheads. Second, we show some techniques to further speed up the sampling process of the parallel WaveNet model. Xixin Wu, Zhiyong Wu 0001, Shiyin Kang, Deyi Tuo, Guangzhi Li, Dan Su 0002, Dong Yu 0001, Helen M. Meng |
ICASSP | 3 |
| 2019 | Speech Emotion Recognition Using Capsule NetworksabstractSpeech emotion recognition (SER) is a fundamental step towards fluent human-machine interaction. One challenging problem in SER is obtaining utterance-level feature representation for classification. Recent works on SER have made significant progress by using spectrogram features and introducing neural network methods, e.g., convolutional neural networks (CNNs). However the fundamental problem of CNNs is that the spatial information in spectrograms is not captured, which are basically position and relationship information of low-level features like pitch and formant frequencies. This paper presents a novel architecture based on the capsule networks (CapsNets) for SER. The proposed system can take into account the spatial relationship of speech features in spectrograms, and provide an effective pooling method for obtaining utterance global features. We also introduce a recurrent connection to CapsNets to improve the model's time sensitivity. We compare the proposed model to previous published results based on combined CNN-long short-term memory (CNN-LSTM) models on the benchmark corpus IEMOCAP over four emotions, i.e., neutral, angry, happy and sad. Experimental results show that our model achieves better results than the baseline system on weighted accuracy (WA) (72.73% vs. 68.8%) and un-weighted accuracy (UA) (59.71% vs. 59.4%), which demonstrates the effectiveness of CapsNets for SER. Xixin Wu, Songxiang Liu, Yuewen Cao, Xu Li 0015, Jianwei Yu 0001, Dongyang Dai, Xi Ma, Shoukang Hu, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 9 |
| 2019 | Modeling Emotion Influence Using Attention-based Graph Convolutional Recurrent NetworkabstractUser emotion modeling is a vital problem of social media analysis. In previous studies, content and topology information of social networks have been considered in emotion modeling tasks, but the inflence of current emotion states of other users was not considered. We define emotion influence as the emotional impact from user’s friends in social networks, which is determined by both network structure and node attributes (the features of friends). In this paper, we try to model the emotion influence to help analyze user’s emotion. The key challenges to this problem are: 1) how to combine content features and network structures together to model emotion influence; 2) how to selectively focus on the major social network information related to emotion influence. To tackle these challenges, we propose an attention-based graph convolutional recurrent network to bring in emotion influence and content data. Firstly, we use an attention-based graph convolutional network to selectively aggregate the features of the user’s friends with specific attention. Then an LSTM model is used to learn user’s own content features and emotion influence. The model we proposed is more capable of quantifying the emotion influence in social networks as well as combining them together to analyze the user emotion status. We conduct emotion classification experiments to evaluate the effectiveness of our model on a real world dataset called Sina Weibo1. Results show that our model outperforms several state-of-the-art methods. Yulan Chen, Jia Jia 0001, Zhiyong Wu 0001 |
ICMI | 3 |
| 2019 | Towards Discriminative Representation Learning for Speech Emotion RecognitionabstractIn intelligent speech interaction, automatic speech emotion recognition (SER) plays an important role in understanding user intention. While sentimental speech has different speaker characteristics but similar acoustic attributes, one vital challenge in SER is how to learn robust and discriminative representations for emotion inferring. In this paper, inspired by human emotion perception, we propose a novel representation learning component (RLC) for SER system, which is constructed with Multi-head Self-attention and Global Context-aware Attention Long Short-Term Memory Recurrent Neutral Network (GCA-LSTM). With the ability of Multi-head Self-attention mechanism in modeling the element-wise correlative dependencies, RLC can exploit the common patterns of sentimental speech features to enhance emotion-salient information importing in representation learning. By employing GCA-LSTM, RLC can selectively focus on emotion-salient factors with the consideration of entire utterance context, and gradually produce discriminative representation for emotion inferring. Experiments on public emotional benchmark database IEMOCAP and a tremendous realistic interaction database demonstrate the outperformance of the proposed SER framework, with 6.6% to 26.7% relative improvement on unweighted accuracy compared to state-of-the-art techniques. Runnan Li, Zhiyong Wu 0001, Jia Jia 0001, Yaohua Bu, Sheng Zhao 0002, Helen M. Meng |
IJCAI | 2 |
| 2019 | Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-Trained BERT
Dongyang Dai, Zhiyong Wu 0001, Shiyin Kang, Xixin Wu, Jia Jia 0001, Dan Su 0002, Dong Yu 0001, Helen M. Meng |
INTERSPEECH | 2 |
| 2019 | Knowledge-Based Linguistic Encoding for End-to-End Mandarin Text-to-Speech Synthesis
Jingbei Li, Zhiyong Wu 0001, Runnan Li, Pengpeng Zhi, Helen M. Meng |
INTERSPEECH | 2 |
| 2019 | One-Shot Voice Conversion with Global Speaker Embeddings
Zhiyong Wu 0001, Dongyang Dai, Runnan Li, Shiyin Kang, Jia Jia 0001, Helen M. Meng |
INTERSPEECH | 2 |
| 2018 | Emphatic Speech Generation with Conditioned Input Layer and Bidirectional LSTMS for Expressive Speech SynthesisabstractBy highlighting the focus of an utterance to draw attention, emphasis in speech interaction plays an important role for speaker intention expressing and understanding. Therefore, emphatic speech synthesis draws increasing interest in the text-to-speech (TTS) area. For emphatic speech synthesis, three problems still exist: 1) sparseness of emphatic speech data; 2) flexibility of trained model; 3) modelling shortage for secondary emphasis. Recently, recurrent neural networks (RNNs) and their bidirectional long short term memory (BLSTM) variants based statistical parametric speech synthesis (SPSS) systems have shown their adaptability and controllability in acoustic modelling thus can solve aforementioned problems. In this paper, we propose a novel conditional input layer for conventional BLSTM-RNN based approach combining using emphasis-specific vectors and linguistic features as input to produce emphatic speech trajectories. Experimental results from objective and subjective evaluations demonstrate the proposed approach can produce emphatic speech trajectories with high quality and naturalness only requiring an additional small-scale emphatic speech corpus. Runnan Li, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng, Lianhong Cai |
ICASSP | 2 |
| 2018 | Unsupervised Discovery of an Extended Phoneme Set in L2 English Speech for Mispronunciation Detection and DiagnosisabstractSecond language (L2) speech is often labelled with the native, phoneme categories. Hence, we often observe segments for which it is difficult, if not impossible, to decide on a categorical phoneme label. We refer to these segments as “non-categorical” phoneme units. Existing approaches to mispronunciation detection and diagnosis (MDD) mostly focus on categorical phoneme errors, where one native phoneme is substituted for another. However, noncategorical errors are not considered. To better represent L2 speech for improved MDD, this work aims to discover an Extended Phoneme Set in L2 speech (L2-EPS) which includes not only the categorical phonemes based on the native set, but also non-categorical phoneme units. We apply an optimized k-means algorithm to cluster phoneme-based phonemic posterior-grams (PPGs), which are generated through an acoustic-phonemic model (APM). Then we find the L2-EPS based on analysis of the clusters obtained. We verified experimentally that the non-categorical phonemes in L2-EPS can extend the native phoneme categories to better describe L2 speech. Hence L2-EPS can enrich the existing approaches to MDD for better performance. Shaoguang Mao, Xu Li 0015, Kun Li 0003, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 4 |
| 2018 | Applying Multitask Learning to Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English SpeechabstractFor mispronunciation detection and diagnosis (MDD), nowadays approaches generally treat the phonemes in correct and mispronunciations as the same despite the fact they may actually carry different characteristics. Furthermore, serious data imbalance issue between correct and mispronunciation in dataset further influences the performances. To address these problems, this paper investigates the use of multi-task (MT) learning technique to enhance the acoustic-phonemic model (APM) for MDD. The phonemes in correct and mispronunciations are processed separately but in multi-task manner considering both correct and mispronunciation recognition tasks. A feature representation module is further proposed to improve performance. Compared with baseline APM, the proposed MT-APM, R-MT-APM achieve better performance not only in Precision, Recall and F-Measure, but also in mispronunciation detection and diagnosis accuracies. With feature representation module, R-MT-APM achieves the highest mispronunciation detection accuracy. Shaoguang Mao, Zhiyong Wu 0001, Runnan Li, Xu Li 0015, Helen M. Meng, Lianhong Cai |
ICASSP | 2 |
| 2018 | Feature Based Adaptation for Speaking Style SynthesisabstractSpeaking style plays an important role in the expressivity of speech for communication. Hence speaking style is very important for synthetic speech as well. Speaking style adaptation faces the difficulty that the data of specific styles may be limited and difficult to obtain in large amounts. A possible solution is to leverage data from speaking styles that are more available, to train the speech synthesizer and then adapt it to the target style for which the data is scarce. Conventional DNN adaptation approaches directly update the top layers of a well-trained, style-dependent model towards the target style. The detailed local context-level mismatch between the original and the target styles is not considered. In order to address this issue, two frame-level input feature-based style adaptation techniques are investigated in this paper. We will use style features extracted from (1) a target-style data trained bottleneck DNN, and (2) a novel cross-style residual feature regression DNN. These features are used for top-layer adaptation of a well-trained style-dependent synthesis network. Experimental results on adapting the declarative sty le to the interrogative sty le demonstrate the effectiveness of our proposed style features in improving the expressiveness of synthesizing speech for the interrogative style, while maintaining speech quality. Xixin Wu, Lifa Sun, Shiyin Kang, Songxiang Liu, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 5 |
| 2018 | Integrating Articulatory Features into Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English SpeechabstractThis paper proposes novel approaches to mispronunciation detection and diagnosis (MDD) on second-language (L2) learners' speech with articulatory features. Here, articulatory features are the positions of articulators when pronouncing phonemes and reflect the pronunciation mechanisms of each phoneme. The use of articulatory features in MDD is helpful in distinguishing phonemes. Three models with articulatory features are proposed based on acoustic-phonemic model (APM): 1) articulatory-acoustic-phonemic model (AAPM) that embeds articulatory features directly into input features; 2) AAPM with feature representation (R-AAPM) to represent original input features with articulatory features; and 3) articulatory multi-task acoustic-phonemic model (A-MT-APM) where phoneme recognizer and articulatory feature classifiers are trained simultaneously in multi-task manner. Compared with baseline phoneme-based APM, proposed approaches perform better in mispronunciation detection and diagnosis measured with Precision, Recall and F1-Measure metrics. Specifically, the A-MT-APM approach gains 5.6% and 7.0% improvement in F1-Measure and diagnostic accuracy respectively. The contributions include: 1) introducing the articulatory features to MDD in deep learning framework; 2) investigating several model architectures for better exploiting articulatory features. Shaoguang Mao, Zhiyong Wu 0001, Xu Li 0015, Runnan Li, Xixin Wu, Helen M. Meng |
ICME | 2 |
| 2018 | Emotion Recognition from Variable-Length Speech Segments Using Deep Learning on Spectrograms
Xi Ma, Zhiyong Wu 0001, Jia Jia 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai |
INTERSPEECH | 2 |
| 2018 | Rapid Style Adaptation Using Residual Error Embedding for Expressive Speech Synthesis
Xixin Wu, Yuewen Cao, Songxiang Liu, Shiyin Kang, Zhiyong Wu 0001, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng |
INTERSPEECH | 6 |
| 2018 | Detection of Glottal Closure Instants from Speech Signals: A Convolutional Neural Network Based Method
Zhiyong Wu 0001, Binbin Shen, Helen M. Meng |
INTERSPEECH | 2 |
| 2018 | Siamese Recurrent Auto-Encoder Representation for Query-by-Example Spoken Term Detection
Ziwei Zhu 0003, Zhiyong Wu 0001, Runnan Li, Helen M. Meng, Lianhong Cai |
INTERSPEECH | 2 |
| 2018 | Inferring User Emotive State Changes in Realistic Human-Computer Conversational DialogsabstractHuman-computer conversational interactions are increasingly pervasive in real-world applications, such as chatbots and virtual assistants. The user experience can be enhanced through affective design of such conversational dialogs, especially in enabling the computer to understand the emotive state in the user's input, and to generate an appropriate system response within the dialog turn. Such a system response may further influence the user's emotive state in the subsequent dialog turn. In this paper, we focus on the change in the user's emotive states in adjacent dialog turns, to which we refer as user emotive state change. We propose a multi-modal, multi-task deep learning framework to infer the user's emotive states and emotive state changes simultaneously. Multi-task learning convolution fusion auto-encoder is applied to fuse the acoustic and textual features to generate a robust representation of the user's input. Long-short term memory recurrent auto-encoder is employed to extract features of system responses at the sentence-level to better capture factors affecting user emotive states. Multi-task learned structured output layer is adopted to model the dependency of user emotive state change, conditioned upon the user input's emotive states and system response in current dialog turn. Experimental results demonstrate the effectiveness of the proposed method. Runnan Li, Zhiyong Wu 0001, Jia Jia 0001, Jingbei Li, Wei Chen 0071, Helen M. Meng |
ACM Multimedia | 2 |
| 2018 | Automatic lexical stress and pitch accent detection for L2 English speech using multi-distribution deep neural networks
Kun Li 0003, Shaoguang Mao, Xu Li 0015, Zhiyong Wu 0001, Helen M. Meng |
Speech Commun. | 4 |
| 2017 | Multi-Task Deep Learning for User Intention Understanding in Speech Interaction SystemsabstractSpeech interaction systems have been gaining popularity in recent years. The main purpose of these systems is to generate more satisfactory responses according to users' speech utterances, in which the most critical problem is to analyze user intention. Researches show that user intention conveyed through speech is not only expressed by content, but also closely related with users' speaking manners (e.g. with or without acoustic emphasis). How to incorporate these heterogeneous attributes to infer user intention remains an open problem. In this paper, we define Intention Prominence (IP) as the semantic combination of focus by text and emphasis by speech, and propose a multi-task deep learning framework to predict IP. Specifically, we first use long short-term memory (LSTM) which is capable of modeling long short-term contextual dependencies to detect focus and emphasis, and incorporate the tasks for focus and emphasis detection with multi-task learning (MTL) to reinforce the performance of each other. We then employ Bayesian network (BN) to incorporate multimodal features (focus, emphasis, and location reflecting users' dialect conventions) to predict IP based on feature correlations. Experiments on a data set of 135,566 utterances collected from real-world Sogou Voice Assistant illustrate that our method can outperform the comparison methods over 6.9-24.5% in terms of F1-measure. Moreover, a real practice in the Sogou Voice Assistant indicates that our method can improve the performance on user intention understanding by 7%. Yishuang Ning, Jia Jia 0001, Zhiyong Wu 0001, Runnan Li, Yongsheng An, Helen M. Meng |
AAAI | 3 |
| 2017 | Multi-task learning of structured output layer bidirectional LSTMS for speech synthesisabstractRecurrent neural networks (RNNs) and their bidirectional long short term memory (BLSTM) variants are powerful sequence modelling approaches. Their inherently strong ability in capturing long range temporal dependencies allow BLSTM-RNN speech synthesis systems to produce higher quality and smoother speech trajectories than conventional deep neural networks (DNNs). In this paper, we improve the conventional BLSTM-RNN based approach by introducing a multi-task learned structured output layer where spectral parameter targets are conditioned upon pitch parameters prediction. Both objective and subjective experimental results demonstrated the effectiveness of the proposed technique. Runnan Li, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng, Lianhong Cai |
ICASSP | 2 |
| 2017 | Learning cross-lingual knowledge with multilingual BLSTM for emphasis detection with limited training dataabstractBidirectional long short-term memory (BLSTM) recurrent neural network (RNN) has achieved state-of-the-art performance in many sequence processing problems given its capability in capturing contextual information. However, for languages with limited amount of training data, it is still difficult to obtain a high quality BLSTM model for emphasis detection, the aim of which is to recognize the emphasized speech segments from natural speech. To address this problem, in this paper, we propose a multilingual BLSTM (MTL-BLSTM) model where the hidden layers are shared across different languages while the softmax output layer is language-dependent. The MTL-BLSTM can learn cross-lingual knowledge and transfer this knowledge to both languages to improve the emphasis detection performance. Experimental results demonstrate our method can outperform the comparison methods over 2-15.6% and 2.9-15.4% on the English corpus and Mandarin corpus in terms of relative F1-measure, respectively. Yishuang Ning, Zhiyong Wu 0001, Runnan Li, Jia Jia 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai |
ICASSP | 2 |
| 2017 | Multi-Task Learning for Prosodic Structure Generation Using BLSTM RNN with Structured Output Layer
Zhiyong Wu 0001, Runnan Li, Helen M. Meng, Lianhong Cai |
INTERSPEECH | 2 |
| 2017 | Spectro-Temporal Modelling with Time-Frequency LSTM and Structured Output Layer for Voice Conversion
Runnan Li, Zhiyong Wu 0001, Yishuang Ning, Lifa Sun, Helen M. Meng, Lianhong Cai |
INTERSPEECH | 2 |
| 2017 | Speech Emotion Recognition with Emotion-Pair Based Framework Considering Emotion Distribution Information in Dimensional Emotion Space
Xi Ma, Zhiyong Wu 0001, Jia Jia 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai |
INTERSPEECH | 2 |
| 2017 | Movie Recommendation via BLSTM
Song Tang 0001, Zhiyong Wu 0001, Kang Chen 0001 |
MMM (2) | 2 |
| 2016 | Low level descriptors based DBLSTM bottleneck feature for speech driven talking avatarabstractSpeech is bimodal in nature. There are close correlations between the acoustic speech signals and the visual gestures such as lip movements, facial expressions and head motions. For speech driven talking avatar, how to derive more representative acoustic features from which to predict more accurate and realistic visual gestures still remains the research problem. Inspired by the promising performance of low level descriptors (LLD) in speech emotion recognition, in this work, we investigate the usage of LLD feature for the task of speech driven talking avatar. Furthermore, visual gestures also demonstrate correlations with not only context information of past or future acoustic features (e.g. anticipatory co-articulation phenomena) but also textual information (e.g. textual hints for lip movement). To incorporate such information, we also propose to use deep bidirectional long short-term memory (DBLSTM) as the bottleneck feature extractor, which can combine LLD feature with contextual information. Experimental results indicate that the proposed LLD based DBLSTM bottleneck feature outperforms the conventional spectrum related features for the task of speech driven talking avatar, and more sophisticated contextual information can further improve the performance. Xinyu Lan, Xu Li 0015, Yishuang Ning, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Lianhong Cai |
ICASSP | 4 |
| 2016 | Question detection from acoustic features using recurrent neural network with gated recurrent unitabstractQuestion detection is of importance for many speech applications. Only parts of the speech utterances can provide useful clues for question detection. Previous work of question detection using acoustic features in Mandarin conversation is weak in capturing such proper time context information, which could be modeled essentially in recurrent neural network (RNN) structure. In this paper, we conduct an investigation on recurrent approaches to cope with this problem. Based on gated recurrent unit (GRU), we build different RNN and bidirectional RNN (BRNN) models to extract efficient features at segment and utterance level. The particular advantage of GRU is it can determine a proper time scale to extract high-level contextual features. Experimental results show that the features extracted within proper time scale make the classifier perform better than the baseline method with pre-designed lexical and acoustic feature set. Yaodong Tang, Zhiyong Wu 0001, Helen M. Meng, Mingxing Xu, Lianhong Cai |
ICASSP | 3 |
| 2016 | Learning cross-lingual information with multilingual BLSTM for speech synthesis of low-resource languagesabstractBidirectional long short-term memory (BLSTM) based speech synthesis has shown great potential in improving the quality of the synthetic speech. However, for low-resource languages, it is difficult to obtain a high quality BLSTM model. BLSTM based speech synthesis can be viewed as a transformation between the input features and the output features. We assume that the input and output layers of BLSTM are language-dependent while the hidden layers can be language-independent if trained properly. We investigate whether sufficient training data of another language (auxiliary) can benefit the BLSTM training of a new language (target) that has only limited training data. In this paper, we propose 1) a multilingual BLSTM that shares hidden layers across different languages and 2) a specific training approach that can best utilize the training data from both the auxiliary and target languages. Experimental results demonstrate the effectiveness of the proposed approach. The multilingual BLSTM can learn the cross-lingual information, and can predict more accurate acoustic features for speech synthesis of the target language than the monolingual BLSTM that is trained with only the data from the target language. Subjective test also indicates that multilingual BLSTM outperforms the monolingual BLSTM in generating higher quality synthetic speech. Quanjie Yu, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng, Lianhong Cai |
ICASSP | 3 |
| 2016 | Heterogeneity-entropy based unsupervised feature learning for personality prediction with cross-media dataabstractPersonality prediction has broad prospects of application in real life. It can be accomplished by analyzing massive and variant data in social networks, which conveys one's personal traits through user generated contents, user's social relationships and behaviors. However, it is difficult to design an effective feature representation from such complex data to predict user's personality as well as high-level and abstract psychological concepts. In this paper, we propose a novel unsupervised cross-modal feature learning algorithm, named Heterogeneity Entropy Neural Network (HENN), to extract the common information between modalities and map it to the user's personality. HENN is constructed hierarchically on Deep Belief Networks (DBNs) and Auto-encoder (AE) with a modified loss function, in which an additional term named Heterogeneity Entropy (HE) is added to measure common information among different modalities. Experiments on a cross-media dataset collected from two famous Chinese social network platforms, i.e., Renren and SinaMicroblog, demonstrate the superiority of our method over several existing algorithms. Haishu Xianyu, Mingxing Xu, Zhiyong Wu 0001, Lianhong Cai |
ICME | 3 |
| 2016 | Phoneme Embedding and its Application to Speech Driven Talking Avatar Synthesis
Xu Li 0015, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Xiaoyan Lou, Lianhong Cai |
INTERSPEECH | 2 |
| 2016 | Expressive Speech Driven Talking Avatar Synthesis with DBLSTM Using Limited Amount of Emotional Bimodal Data
Xu Li 0015, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Xiaoyan Lou, Lianhong Cai |
INTERSPEECH | 2 |
| 2016 | Combining CNN and BLSTM to Extract Textual and Acoustic Features for Recognizing Stances in Mandarin Ideological Debate Competition
Linchuan Li, Zhiyong Wu 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai |
INTERSPEECH | 2 |
| 2016 | Analysis on Gated Recurrent Unit Based Question Detection Approach
Yaodong Tang, Zhiyong Wu 0001, Helen M. Meng, Mingxing Xu, Lianhong Cai |
INTERSPEECH | 2 |
| 2015 | Understanding speaking styles of internet speech data with LSTM and low-resource trainingabstractSpeech are widely used to express one's emotion, intention, desire, etc. in social network communication, deriving abundant of internet speech data with different speaking styles. Such data provides a good resource for social multimedia research. However, regarding different styles are mixed together in the internet speech data, how to classify such data remains a challenging problem. In previous work, utterance-level statistics of acoustic features are utilized as features in classifying speaking styles, ignoring the local context information. Long short-term memory (LSTM) recurrent neural network (RNN) has achieved exciting success in lots of research areas, such as speech recognition. It is able to retrieve context information for long time duration, which is important in characterizing speaking styles. To train LSTM, huge number of labeled training data is required. While for the scenario of internet speech data classification, it is quite difficult to get such large scale labeled data. On the other hand, we can get some publicly available data for other tasks (such as speech emotion recognition), which offers us a new possibility to exploit LSTM in the low-resource task. We adopt retraining strategy to train LSTM to recognize speaking styles in speech data by training the network on emotion and speaking style datasets sequentially without reset the weights of the network. Experimental results demonstrate that retraining improves the training speed and the accuracy of network in speaking style classification. Xixin Wu, Zhiyong Wu 0001, Yishuang Ning, Jia Jia 0001, Lianhong Cai, Helen M. Meng |
ACII | 2 |
| 2015 | A deep recurrent approach for acoustic-to-articulatory inversionabstractTo solve the acoustic-to-articulatory inversion problem, this paper proposes a deep bidirectional long short term memory recurrent neural network and a deep recurrent mixture density network. The articulatory parameters of the current frame may have correlations with the acoustic features many frames before or after. The traditional pre-designed fixed-length context window may be either insufficient or redundant to cover such correlation information. The advantage of recurrent neural network is that it can learn proper context information on its own without the requirement of externally specifying a context window. Experimental results indicate that recurrent model can produce more accurate predictions for acoustic-to-articulatory inversion than deep neural network having fixed-length context window. Furthermore, the predicted articulatory trajectory curve of recurrent neural network is smooth. Average root mean square error of 0.816 mm on the MNGU0 test set is achieved without any post-filtering, which is state-of-the-art inversion accuracy. Quanjie Yu, Zhiyong Wu 0001, Shiyin Kang, Helen M. Meng, Lianhong Cai |
ICASSP | 3 |
| 2015 | HMM-based emphatic speech synthesis for corrective feedback in computer-aided pronunciation trainingabstractThis paper investigates the incorporation of hidden Markov model (HMM) based emphatic speech synthesis for audio exaggeration into an audio-visual speech synthesis framework for the corrective feedback in computer-aided pronunciation training (CAPT). To improve the voice quality of the synthetic emphatic speech, this paper proposes a new method for HMM training. In this method, the contextual questions for decision tree building are extended by considering the emphasis-related information. HMMs are then trained using a small scale emphatic corpus together with a large scale neutral corpus. The emphatic corpus is used to ensure the quality of the emphatic speech segments whereas the neutral corpus is to further improve the quality of both the non-emphatic speech segments and the emphatic ones. Finally, emphatic speech synthesis is achieved by extending the Flite+hts_engine. Experimental results show that our method can synthesize emphatic speech with high quality and make the feedback more discriminatively perceptible. Yishuang Ning, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng, Lianhong Cai |
ICASSP | 2 |
| 2015 | Modelling High-Dimensional Sequences with LSTM-RTRBM: Application to Polyphonic Music Generation
Qi Lyu, Zhiyong Wu 0001, Jun Zhu 0001, Helen M. Meng |
IJCAI | 2 |
| 2015 | Using tilt for automatic emphasis detection with Bayesian networks
Yishuang Ning, Zhiyong Wu 0001, Xiaoyan Lou, Helen M. Meng, Jia Jia 0001, Lianhong Cai |
INTERSPEECH | 2 |
| 2015 | Polyphonic Music Modelling with LSTM-RTRBMabstractRecent interest in music information retrieval and related technologies is exploding. However, very few of the existing techniques take advantage of the recent advancements in neural networks. The challenges of developing effective browsing, searching and organization techniques for the growing bodies of music collections call for more powerful statistical models. In this paper, we present LSTM-RTRBM, a new neural network model for the problem of creating accurate yet flexible models of polyphonic music. Our model integrates the ability of Long Short-Term Memory (LSTM) in memorizing and retrieving useful history information, together with the advantage of Restricted Boltzmann Machine (RBM) in high dimensional data modelling. Our approach greatly improves the performance of polyphonic music sequence modelling, achieving the state-of-the-art results on multiple datasets. Qi Lyu, Zhiyong Wu 0001, Jun Zhu 0001 |
ACM Multimedia | 2 |
| 2015 | Generating emphatic speech with hidden Markov model for expressive speech synthesis
Zhiyong Wu 0001, Yishuang Ning, Xiao Zang, Jia Jia 0001, Helen M. Meng, Lianhong Cai |
Multim. Tools Appl. | 1 |
| 2015 | Acoustic to articulatory mapping with deep neural network
Zhiyong Wu 0001, Xixin Wu, Xinyu Lan, Helen M. Meng |
Multim. Tools Appl. | 1 |
| 2014 | Learning dynamic features with neural networks for phoneme recognitionabstractDynamic features such as delta and delta-delta of basic acoustic features have long been used in various speech applications and give satisfactory performance. The explicit physical meaning and simplicity of dynamic features clearly compound their prevalence. In this paper, we propose a new framework with neural network to learn the alternatives of traditional delta and higher order differences. Instead of embracing the interpretability and simplicity, our framework is able to learn a new transformation that simulates what differences do but is more relevant to a specific task such as phoneme recognition. We determine the best way to learn such a new transformation among several most probable alternatives. Our experiments indicate that dynamic features obtained with transformation learned this way are better than traditional differences in both frame classification and phoneme recognition. The improvement of performance is even clearer when higher-order of differences are applied. Zhiyong Wu 0001, Helen M. Meng, Lianhong Cai |
ICASSP | 2 |
| 2014 | Contrastive auto-encoder for phoneme recognitionabstractSpeech data typically contains task irrelevant information lying within features. Specifically, phonetic information, speaker characteristic information, emotional information and noise are always mixed together and tend to impair one another for certain task. We propose a new type of auto-encoder for feature learning called contrastive auto-encoder. Unlike other variants of auto-encoders, contrastive auto-encoder is able to leverage class labels in constructing its representation layer. We achieve this by modeling two autoencoders together and making their differences contribute to the total loss function. The transformation built with contrastive auto-encoder can be seen as a task-specific and invariant feature learner. Our experiments on TIMIT clearly show the superiority of the feature extracted from contrastive auto-encoder over original acoustic feature, feature extracted from deep auto-encoder, and feature extracted from a model that contrastive auto-encoder originates from. Zhiyong Wu 0001, Helen M. Meng, Lianhong Cai |
ICASSP | 2 |
| 2014 | Using conditional random fields to predict focus word pair in spontaneous spoken EnglishabstractThis paper addresses the problem of automatically labeling focus word pairs in spontaneous spoken English, where a focus word pair refers to salient part of text or speech and the word motivating it.The prediction of focus word pairs is important for speech applications such as expressive text-tospeech (TTS) synthesis and speech recognition.It can also help in better textual and intention understanding for spoken dialog systems.Traditional approaches such as support vector machines (SVMs) prediction neglect the dependency between words and meet the obstacle of the imbalanced distribution of positive and negative samples of dataset.This paper introduces conditional random fields (CRFs) to the task of automatically predicting focus word pair from lexical, syntactic and semantic features.Furthermore, several new features related to syntactic and semantic information are proposed to achieve better performance.Experiments on the publicly available Switchboard corpus demonstrate that CRF model outperforms the baseline and SVM model for focus word pair prediction, and newly proposed features can further improve performance for CRF based predictor.Specifically, compared to the low recall rate of 11.31% achieved by the SVM model, the proposed CRF based predictor can yield a high recall rate of 70.88% with little impact on precision. Xiao Zang, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Lianhong Cai |
INTERSPEECH | 2 |
| 2014 | Multi-channel speech enhancement using sparse coding on local time-frequency structures
Zhaogui Ding, Weifeng Li 0001, Zhiyong Wu 0001, Longbiao Wang, Qingmin Liao |
INTERSPEECH | 4 |
| 2014 | Head and facial gestures synthesis using PAD model for an expressive talking avatar
Jia Jia 0001, Zhiyong Wu 0001, Helen M. Meng, Lianhong Cai |
Multim. Tools Appl. | 2 |
| 2014 | Synthesizing English emphatic speech for multimodal corrective feedback in computer-aided pronunciation training
Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng, Lianhong Cai |
Multim. Tools Appl. | 2 |
| 2013 | Investigation of tandem deep belief network approach for phoneme recognitionabstractThis paper proposes using tandem DBN approach - a hierarchical architecture that consists of two or more deep belief networks (DBNs) in tandem manner - for phoneme recognition task on TIMIT. First we describe the standard DBN approach applied in phoneme recognition and discuss the motivation of combining it with tandem classifier approach. We then perform series of experiments to find out the best configuration for the DBN in the second level and discover the full potential of this method. The experiments show that for the DBN in the second level, (a) 2048 units in each hidden layer is better than 1024 and 512 units, (b) for sufficient length of temporal context, two hidden layers are better, (c) the one gives best performance on development set shows 4% relative improvement on coretest set. Zhiyong Wu 0001, Binbin Shen, Helen M. Meng, Lianhong Cai |
ICASSP | 2 |
| 2012 | Hierarchical English Emphatic Speech Synthesis Based on HMM with Limited Training DataabstractEmphasis is an important form of expressiveness in speech. Hidden Markov model (HMM) based speech synthesis has shown great flexibility in generating expressive speech. This paper proposes a hierarchical model based on HMMs aiming at synthesizing emphatic speech of both high emphasis quality and high naturalness with the limited amount of data. Decision trees (DTs) are constructed with non-emphasis-related questions using both neutral and emphasis corpora. The data in each leaf node of the DTs are classified into 6 emphasis categories according to the emphasis-related questions. The data in the same emphasis category are grouped into one sub-node and are used to train one HMM. As there might be no data of some specific emphasis categories in the leaf nodes of the DTs, a method based on cost calculation is proposed to select a suitable HMM in the same leaf node for predicting parameters. Further a compensation model is proposed to adjust the predicted parameters. Experiments show that the proposed hierarchical model can synthesize emphatic speech with high quality for both naturalness and emphasis, using limited amount of training data. Index Terms: emphatic speech synthesis, hidden Markov model (HMM), hierarchy, compensation model 1. Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Lianhong Cai |
INTERSPEECH | 2 |
| 2012 | Comparison of adaptation methods for GMM-SVM based speech emotion recognitionabstractThe required length of the utterance is one of the key factors affecting the performance of automatic emotion recognition. To gain the accuracy rate of emotion distinction, adaptation algorithms that can be manipulated on short utterances are highly essential. Regarding this, this paper compares two classical model adaptation methods, maximum a posteriori (MAP) and maximum likelihood linear regression (MLLR), in GMM-SVM based emotion recognition, and tries to find which method can perform better on different length of the enrollment of the utterances. Experiment results show that MLLR adaptation performs better for very short enrollment utterances (with the length shorter than 2s) while MAP adaptation is more effective for longer utterances. Jianbo Jiang, Zhiyong Wu 0001, Mingxing Xu, Jia Jia 0001, Lianhong Cai |
SLT | 2 |
| 2011 | Combining Active and Semi-Supervised Learning for Homograph Disambiguation in Mandarin Text-to-Speech SynthesisabstractGrapheme-to-phoneme conversion (G2P) is a crucial step for Mandarin text-to-speech (TTS) synthesis, where homograph disambiguation is the core issue. Several machine learning algorithms have been proposed to solve the issue by building models from well annotated training corpus. However, the preparation of such well annotated corpus is very laboring and time-consuming which requires lots of manual hand-label work to validate the proper pronunciations of the homographs. This work tries to cover this problem by introducing the active learning (AL) and semi-supervised learning (SSL) algorithms for the homograph disambiguation task using unlabeled data. Experiments show that the proposed framework can greatly reduce the cost of manual hand-label work while preserving the performance of the trained model. Index Terms: text-to-speech (TTS) synthesis, homograph disambiguation, active learning (AL), Yarowsky algorithm, semi-supervised learning (SSL) 1. Binbin Shen, Zhiyong Wu 0001, Lianhong Cai |
INTERSPEECH | 2 |
| 2010 | Comparison of Syllable/Phone HMM Based Mandarin TTSabstractThe performance of HMM-based text to speech (TTS) system is affected by the basic modeling units and the size of training data. This paper compares two HMM based Mandarin TTS systems using syllable and phone as basic units respectively with 1000, 3000 and 5000 sentences' training data. Two female speakers' corpora are used as training data for evaluation. For both corpora, the system using syllable as basic unit outperforms the system using phone as basic unit with 3000 and 5000 sentences' training data. Quansheng Duan, Shiyin Kang, Zhiyong Wu 0001, Lianhong Cai, Zhiwei Shuang, Yong Qin 0001 |
ICPR | 3 |
| 2009 | Modeling the Expressivity of Input Text Semantics for Chinese Text-to-Speech Synthesis in a Spoken Dialog SystemabstractThis work focuses on the development of expressive text-to-speech synthesis techniques for a Chinese spoken dialog system, where the expressivity is driven by the message content. We adapt the three-dimensional pleasure-displeasure, arousal-nonarousal and dominance-submissiveness (PAD) model for describing expressivity in input text semantics. The context of our study is based on response messages generated by a spoken dialog system in the tourist information domain. We use theP(pleasure) andA(arousal) dimensions to describe expressivity at the prosodic word level based onlexicalsemantics. TheD(dominance) dimension is used to describe expressivity at the utterance level based ondialogacts. We analyze contrastive (neutral versus expressive) speech recordings to develop a nonlinear perturbation model that incorporates the PAD values of a response message to transform neutral speech into expressive speech. Two levels of perturbations are implemented-local perturbation at the prosodic word level, as well as global perturbation at the utterance level. Perceptual experiments involving 14 subjects indicate that the proposed approach can significantly enhance expressivity in response generation for a spoken dialog system. Zhiyong Wu 0001, Helen M. Meng, Hongwu Yang, Lianhong Cai |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Facial Expression Synthesis Using PAD Emotional Parameters for a Chinese Expressive Avatar
Zhiyong Wu 0001, Helen M. Meng, Lianhong Cai |
ACII | 2 |
| 2007 | Head Movement Synthesis Based on Semantic and Prosodic Features for a Chinese Expressive AvatarabstractThis paper proposes an approach for text-to-visual speech synthesis, where the synthetic head movements are rendered with an expressive talking avatar speaking Cantonese Chinese. The input text consists of descriptive information sourced from the Hong Kong tourism domain. The text is segmented into prosodic words (PW) and we adopt the PAD model to describe the expressivity of a prosodic word based on its semantics. Within the PW, we consider two prosodic features relevant to head movement synthesis, namely, the stress and tone of the Chinese syllable. We designed and recorded an audiovisual speech corpus and analyzed the data to derive statistical correspondences between different (P,A) values for a Chinese prosodic word and head movement coordinates. These statistics help parameter selection in a sinusoidal movement model. Corpus analyses also enable us to locate "peak points" of head movements that are synchronized with prosodic features within a prosodic word. These help the design of three heuristics that control head movements within a prosodic word. Perceptual evaluation based on the expressive talking avatar shows that head movement synthesis can raise the MOS by 1.04 points on average, when compared to the baseline which only shows lip articulations without head movements. Zhiyong Wu 0001, Helen M. Meng, Lianhong Cai |
ICASSP (4) | 2 |
| 2006 | Real-time synthesis of Chinese visual speech and facial expressions using MPEG-4 FAP features in a three-dimensional avatarabstractThis paper describes our initial work in developing a real-time audio-visual Chinese speech synthesizer with a 3D expressive avatar. The avatar model is parameterized according to the MPEG-4 facial animation standard [1]. This standard offers a compact set of facial animation parameters (FAPs) and feature points (FPs) to enable realization of 20 Chinese visemes and 7 facial expressions (i.e. 27 target facial configurations). The Xface [2] open source toolkit enables us to define the influence zone for each FP and the deformation function that relates them. Hence we can easily animate a large number of coordinates in the 3D model by specifying values for a small set of FAPs and their FPs. FAP values for 27 target facial configurations were estimated from available corpora. We extended the dominance blending approach to effect animations for coarticulated visemes superposed with expression changes. We selected six sentiment-carrying text messages and synthesized expressive visual speech (for all expressions, in randomized order) with neutral audio speech. A perceptual experiment involving 11 subjects shows that they can identify the facial expression that matches the text message’s sentiment 85% of the time. Zhiyong Wu 0001, Lianhong Cai, Helen M. Meng |
INTERSPEECH | 1 |
| 2006 | Modelling the Global acoustic Correlates of Expressivity for Chinese Text-to-speech SynthesisabstractThis paper proposed a novel approach for describing the expressive elements in dialog response messages for expressive text-to-speech synthesis. We adopt the three-dimensional PAD emotional model in describing expressivity based on response message content and its dialog state. In particular, we use the P (pleasure) and A (arousal) descriptors to describe expressivity at the local, prosodic-word level based on its semantics. We also use the D (dominance) descriptor to describe expressivity at the global, utterance level based on its dialog act. Our context of study is based on response messages of a spoken dialog system in the Hong Kong tourism domain. We also prepared contrastive (neutral versus expressive) recordings to aid identification of the acoustic correlates of expressivity at both local and global levels. We utilized the acoustic analysis of these contrastive recordings to establish a nonlinear model that can be used to modulate input neutral speech at both local and global levels to generate output expressive speech. This work focuses on the nonlinear relationship between the D (dominance) values and their acoustic correlates. Perceptual evaluation indicates that local modulation of input neutral speech produces over 73% utterances carry appropriate expressivity. The combined uses of both local and global modulations produce nearly 84% expressive utterances. Hongwu Yang, Helen M. Meng, Zhiyong Wu 0001, Lianhong Cai |
SLT | 3 |
| 2000 | Research on dynamic characters of Chinese pitch contours
Zhiyong Wu 0001, Lianhong Cai, Tongchun Zhou |
INTERSPEECH | 1 |