Shujie Liu 0001

dblp:54/2695 · DBLP profile ↗
← Back
127ranked-venue papers
14as first author
67since 2021 · last 2026
0009-0008-2599-6752ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 88 · 6 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 64 · 6 first-author · 43 since 2021Computer networks · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Shanks: Simultaneous Hearing and Thinking for Spoken Language Models
abstract
Cheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin, Kevin Lin, Shujie Liu, Zhendong Wang, Zhengyuan Yang, Hung-yi Lee, Lijuan Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Cheng-Han Chiang, Chung-Ching Lin, Shujie Liu 0001, Zhengyuan Yang, Hung-yi Lee
ACL (1)6
2026 AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
abstract
Hong Ting Tsang, Jiaxin Bai, Haoyu Huang, Qiao Xiao, Tianshi Zheng, Baixuan Xu, Shujie Liu, Yangqiu Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hong Ting Tsang, Jiaxin Bai, Qiao Xiao, Tianshi Zheng, Baixuan Xu, Shujie Liu 0001, Yangqiu Song
ACL (1)7
2026 Closing the Modality Reasoning Gap for Speech Large Language Models
abstract
Although Speech Large Language Models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on speech inputs is markedly weaker than on text.This gap could be associated with representational drift across Transformer layers and behavior deviations in long-chain reasoning.To address this issue, we introduce TARS, a reinforcement-learning framework that aligns text-conditioned and speech-conditioned trajectories through an asymmetric reward design.The framework employs two dense and complementary signals: representation alignment, which measures layer-wise hidden-state similarity between speech-and text-conditioned trajectories, and behavior alignment, which evaluates semantic consistency between generated outputs and reference text completions.Experiments on challenging reasoning benchmarks, including MMSU and OBQA, show that our approach significantly narrows the modality reasoning gap and achieves state-of-the-art performance among 7B-scale Speech LLMs.
Chaoren Wang, Xueyao Zhang, Shujie Liu 0001, Yan Lu 0001, Jinyu Li 0001, Zhizheng Wu 0001
ACL (1)4
2026 SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
abstract
Hui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu, Junyang Chen, Yanzhe Zhang, Shiwan Zhao, Jinyu Li, Jiaming Zhou, Haoqin Sun, Yan Lu, Yong Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hui Wang 0075, Jinghua Zhao 0004, Yifan Yang 0005, Shujie Liu 0001, Shiwan Zhao, Jinyu Li 0001, Jiaming Zhou 0001, Haoqin Sun, Yan Lu 0001
ACL (1)4
2026 Recent Advances in Discrete Speech Tokens: A Review
abstract
The rapid advancement of speech generation technologies in the era of large language models (LLMs) has established discrete speech tokens as a foundational paradigm for speech representation. These tokens, characterized by their discrete, compact, and concise nature, are not only advantageous for efficient transmission and storage, but also inherently compatible with the language modeling framework, enabling seamless integration of speech into text-dominated LLM architectures. Current research categorizes discrete speech tokens into two principal classes: acoustic tokens and semantic tokens, each of which has evolved into a rich research domain characterized by unique design philosophies and methodological approaches. This survey systematically synthesizes the existing taxonomy and recent innovations in discrete speech tokenization, conducts a critical examination of the strengths and limitations of each paradigm, and presents systematic experimental comparisons across token types. Furthermore, we identify persistent challenges in the field and propose potential research directions, aiming to offer actionable insights to inspire future advancements in the development and application of discrete speech tokens.
Hankun Wang, Bohan Li 0003, Chongtian Shao, Hanglei Zhang, Chenpeng Du, Xie Chen 0001, Shujie Liu 0001, Kai Yu 0004
IEEE Trans. Pattern Anal. Mach. Intell.9
2025 Autoregressive Speech Synthesis without Vector Quantization
abstract
Lingwei Meng, Long Zhou, Shujie Liu, Sanyuan Chen, Bing Han, Shujie Hu, Yanqing Liu, Jinyu Li, Sheng Zhao, Xixin Wu, Helen M. Meng, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Lingwei Meng, Shujie Liu 0001, Sanyuan Chen, Bing Han 0008, Shujie Hu, Jinyu Li 0001, Sheng Zhao 0002, Xixin Wu, Helen M. Meng, Furu Wei
ACL (1)3
2025 TrInk: Ink Generation with Transformer Network
abstract
Zezhong Jin, Shubhang Desai, Xu Chen, Biyi Fang, Zhuoyi Huang, Zhe Li, Chong-Xin Gan, Xiao Tu, Man-Wai Mak, Yan Lu, Shujie Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zezhong Jin, Shubhang Desai, Biyi Fang, Zhuoyi Huang, Zhe Li 0030, Chong-Xin Gan, Xiao Tu, Man-Wai Mak, Yan Lu 0001, Shujie Liu 0001
EMNLP11
2025 V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
abstract
In this paper, we introduce V2SFlow, a novel Video-to-Speech (V2S) framework designed to generate natural and intelligible speech directly from silent talking face videos. While recent V2S systems have shown promising results on constrained datasets with limited speakers and vocabularies, their performance often degrades on real-world, unconstrained datasets due to the inherent variability and complexity of speech signals. To address these challenges, we decompose the speech signal into manageable subspaces (content, pitch, and speaker information), each representing distinct speech attributes, and predict them directly from the visual input. To generate coherent and realistic speech from these predicted attributes, we employ a rectified flow matching decoder built on a Transformer architecture, which models efficient probabilistic pathways from random noise to the target speech distribution. Extensive experiments demonstrate that V2SFlow significantly outperforms state-of-the-art methods, even surpassing the naturalness of ground truth utterances.
Jeongsoo Choi, Jinyu Li 0001, Joon Son Chung, Shujie Liu 0001
ICASSP5
2025 Boosting Large Language Model for Speech Synthesis: An Empirical Study
abstract
Large language models (LLMs) have made significant advancements in natural language processing and are concurrently extending the language ability to other modalities, such as speech and vision. Nevertheless, most of the previous work focuses on prompting LLMs with perception abilities like auditory comprehension, and the effective approach for augmenting LLMs with speech synthesis capabilities remains ambiguous. In this paper, we conduct a comprehensive empirical exploration of boosting LLMs with the ability to generate speech, by combining pre-trained LLM LLaMA/OPT and text-to-speech synthesis model VALL-E. We compare three integration methods between LLMs and speech synthesis models, including directly fine-tuned LLMs, superposed layers of LLMs and VALL-E, and coupled LLMs and VALL-E using LLMs as a powerful text encoder. Experimental results show that, using LoRA method to fine-tune LLMs directly to boost the speech synthesis capability does not work well, and superposed LLMs and VALL-E can improve the quality of generated speech both in speaker similarity and word error rate (WER). Among these three methods, coupled methods leveraging LLMs as the text encoder can achieve the best performance, making it outperform original speech synthesis models with a consistently better speaker similarity and a significant (10.9%) WER reduction.
Hongkun Hao, Shujie Liu 0001, Jinyu Li 0001, Shujie Hu, Rui Wang 0015, Furu Wei
ICASSP3
2025 ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation
abstract
Text-to-video (T2V) models have recently undergone rapid and substantial advancements. Nevertheless, due to limitations in data and computational resources, achieving efficient generation of long videos with rich motion dynamics remains a significant challenge. To generate high-quality, dynamic, and temporally consistent long videos, this paper presents ARLON, a novel framework that boosts diffusion Transformers with autoregressive (\textbf{AR}) models for long (\textbf{LON}) video generation, by integrating the coarse spatial and long-range temporal information provided by the AR model to guide the DiT model effectively. Specifically, ARLON incorporates several key innovations: 1) A latent Vector Quantized Variational Autoencoder (VQ-VAE) compresses the input latent space of the DiT model into compact and highly quantized visual tokens, bridging the AR and DiT models and balancing the learning complexity and information density; 2) An adaptive norm-based semantic injection module integrates the coarse discrete visual units from the AR model into the DiT model, ensuring effective guidance during video generation; 3) To enhance the tolerance capability of noise introduced from the AR inference, the DiT model is trained with coarser visual latent tokens incorporated with an uncertainty sampling module. Experimental results demonstrate that ARLON significantly outperforms the baseline OpenSora-V1.2 on eight out of eleven metrics selected from VBench, with notable improvements in dynamic degree and aesthetic quality, while delivering competitive results on the remaining three and simultaneously accelerating the generation process. In addition, ARLON achieves state-of-the-art performance in long video generation, outperforming other open-source models in this domain. Detailed analyses of the improvements in inference efficiency are presented, alongside a practical application that demonstrates the generation of long videos using progressive text prompts. Project page: \url{http://aka.ms/arlon}.
Zongyi Li, Shujie Hu, Shujie Liu 0001, Jeongsoo Choi, Lingwei Meng, Jinyu Li 0001, Furu Wei
ICLR3
2025 Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
abstract
Recent zero-shot text-to-speech (TTS) systems face a common dilemma: autoregressive (AR) models suffer from slow generation and lack duration controllability, while non-autoregressive (NAR) models lack temporal modeling and typically require complex designs. In this paper, we introduce a novel pseudo-autoregressive (PAR) codec language modeling approach that unifies AR and NAR modeling. Combining explicit temporal modeling from AR with parallel generation from NAR, PAR generates dynamic-length spans at fixed time steps. Building on PAR, we propose PALLE, a two-stage TTS system that leverages PAR for initial generation followed by NAR refinement. In the first stage, PAR progressively generates speech tokens along the time dimension, with each step predicting all positions in parallel but only retaining the left-most span. In the second stage, low-confidence tokens are iteratively refined in parallel, leveraging the global contextual information. Experiments demonstrate that PALLE, trained on LibriTTS, outperforms state-of-the-art systems trained on large-scale data, including F5-TTS, E2-TTS, and MaskGCT, on the LibriSpeech test-clean set in terms of speech quality, speaker similarity, and intelligibility, while achieving up to ten times faster inference speed. Audio samples are available at https://microsoft.com/research/project/vall-e-x/palle.
Yifan Yang 0005, Shujie Liu 0001, Jinyu Li 0001, Yuxuan Hu 0003, Hui Wang 0075, Jianwei Yu 0001, Lingwei Meng, Haiyang Sun 0004, Yan Lu 0001, Kai Yu 0004, Xie Chen 0001
ACM Multimedia2
2025 FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
abstract
To advance continuous token modeling and temporal-coherence enforcement, we propose FELLE, an autoregressive model that integrates language modeling with token-wise flow matching. By leveraging the autoregressive nature of language models and the generative efficacy of flow matching, FELLE effectively predicts continuous-valued tokens (mel-spectrograms). For each continuous-valued token, FELLE modifies the general prior distribution in flow matching by incorporating information from the previous step, improving coherence and stability. Furthermore, to enhance synthesis quality, FELLE introduces a coarse-to-fine flow-matching mechanism, generating continuous-valued tokens hierarchically, conditioned on the language model's output. Experimental results demonstrate the potential of incorporating flow-matching techniques in autoregressive mel-spectrogram modeling, leading to significant improvements in TTS generation quality, as shown in https://aka.ms/felle.
Hui Wang 0075, Shujie Liu 0001, Lingwei Meng, Jinyu Li 0001, Yifan Yang 0005, Shiwan Zhao, Haiyang Sun 0004, Haoqin Sun, Jiaming Zhou 0001, Yan Lu 0001
ACM Multimedia2
2025 StreamMel: Real-Time Zero-Shot Text-to-Speech Via Interleaved Continuous Autoregressive Modeling
abstract
Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming TTS paradigms often rely on multi-stage pipelines and discrete representations, leading to increased computational cost and suboptimal system performance. In this work, we propose StreamMel, a pioneering single-stage streaming TTS framework that models continuous mel-spectrograms. By interleaving text tokens with acoustic frames, StreamMel enables low-latency, autoregressive synthesis while preserving high speaker similarity and naturalness. Experiments on LibriSpeech demonstrate that StreamMel outperforms existing streaming TTS baselines in both quality and latency. It even achieves performance comparable to offline systems while supporting efficient real-time generation, showcasing broad prospects for integration with real-time speech large language models. Audio samples are available at:https://aka.ms/StreamMel.
Hui Wang 0075, Yifan Yang 0005, Shujie Liu 0001, Jinyu Li 0001, Lingwei Meng, Tie-Yan Liu, Jiaming Zhou 0001, Haoqin Sun, Yan Lu 0001
IEEE Signal Process. Lett.3
2024 COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
Yashesh Gaur, Sunit Sivasankaran, Zhuo Chen 0006, Shujie Liu 0001, Jinyu Li 0001
INTERSPEECH6
2024 TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
abstract
There is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation. However, most end-to-end models struggle to outperform cascade models, i.e., a pipeline framework by concatenating speech recognition, machine translation and text-to-speech models. The primary challenges stem from the inherent complexities involved in direct translation tasks and the scarcity of data. In this study, we introduce a novel model framework TransVIP that leverages diverse datasets in a cascade fashion yet facilitates end-to-end inference through joint probability. Furthermore, we propose two separated encoders to preserve the speaker’s voice characteristics and isochrony from the source speech during the translation process, making it highly suitable for scenarios such as video dubbing. Our experiments on the French-English language pair demonstrate that our model outperforms the current state-of-the-art speech-to-speech translation model.
Chenyang Le, Yao Qian, Dongmei Wang, Shujie Liu 0001, Xiaofei Wang 0009, Midia Yousefi, Yanmin Qian, Jinyu Li 0001, Sheng Zhao 0002, Michael Zeng 0001
NeurIPS5
2024 CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations
abstract
Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to be a challenge. In this paper, we introduce CoVoMix: Conversational Voice Mixture Generation, a novel model for zero-shot, human-like, multi-speaker, multi-round dialogue speech generation. CoVoMix first converts dialogue text into multiple streams of discrete tokens, with each token stream representing semantic information for individual talkers. These token streams are then fed into a flow-matching based acoustic model to generate mixed mel-spectrograms. Finally, the speech waveforms are produced using a HiFi-GAN model. Furthermore, we devise a comprehensive set of metrics for measuring the effectiveness of dialogue modeling and generation. Our experimental results show that CoVoMix can generate dialogues that are not only human-like in their naturalness and coherence but also involve multiple talkers engaging in multiple rounds of conversation. This is exemplified by instances generated in a single channel where one speaker's utterance is seamlessly mixed with another's interjections or laughter, indicating the latter's role as an attentive listener. Audio samples are enclosed in the supplementary.
Leying Zhang, Yao Qian, Shujie Liu 0001, Dongmei Wang, Xiaofei Wang 0009, Midia Yousefi, Yanmin Qian, Jinyu Li 0001, Lei He 0005, Sheng Zhao 0002, Michael Zeng 0001
NeurIPS4
2024 Investigating Neural Audio Codecs For Speech Language Model-Based Speech Generation
abstract
Neural audio codec tokens serve as the fundamental building blocks for speech language model (SLM)-based speech generation. However, there is no systematic understanding on how the codec system affects the speech generation performance of the SLM. In this work, we examine codec tokens within SLM framework for speech generation to provide insights for effective codec design. We retrain existing high-performing neural codec models on the same data set and loss functions to compare their performance in a uniform setting. We integrate codec tokens into two SLM systems: masked-based parallel speech generation system and an auto-regressive (AR) plus non-auto-regressive (NAR) model-based system. Our findings indicate that better speech reconstruction in codec systems does not guarantee improved speech generation in SLM. A high-quality codec decoder is crucial for natural speech production in SLM, while speech intelligibility depends more on quantization mechanism.
Jiaqi Li 0030, Dongmei Wang, Xiaofei Wang 0009, Yao Qian, Shujie Liu 0001, Midia Yousefi, Canrun Li, Chung-Hsien Tsai, Jun-Kun Chen, Sheng Zhao 0002, Jinyu Li 0001, Zhizheng Wu 0001, Michael Zeng 0001
SLT6
2024 NDVQ: Robust Neural Audio Codec With Normal Distribution-Based Vector Quantization
abstract
Built upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and autoregressive audio generation. However, existing models face substantial challenges in perceptual quality and signal distortion, especially when operating in extremely low bandwidth, rooted in the sensitivity of the VQ codebook to noise. This degradation poses significant challenges for several downstream tasks, such as codec-based speech synthesis. To address this issue, we propose a novel VQ method, Normal Distribution-based Vector Quantization (NDVQ), by introducing an explicit margin between the VQ codes via learning a variance. Specifically, our approach involves mapping the waveform to a latent space and quantizing it by selecting the most likely normal distribution, with each codebook entry representing a unique normal distribution defined by its mean and variance. Using these distribution-based VQ codec codes, a decoder reconstructs the input waveform. NDVQ is trained with additional distribution-related losses, alongside reconstruction and discrimination losses. Experiments demonstrate that NDVQ outperforms existing audio compression baselines, such as EnCodec, in terms of audio quality and zero-shot TTS, particularly in very low bandwidth scenarios.
Zhikang Niu, Sanyuan Chen, Ziyang Ma 0001, Xie Chen 0001, Shujie Liu 0001
SLT6
2024 DDTSE: Discriminative Diffusion Model for Target Speech Extraction
abstract
Diffusion models have gained attention in speech enhancement tasks, providing an alternative to conventional discriminative methods. However, research on target speech extraction under multispeaker noisy conditions remains relatively unexplored. Moreover, the superior quality of diffusion methods typically comes at the cost of slower inference speed. In this paper, we introduce the Discriminative Diffusion model for Target Speech Extraction (DDTSE). We apply the same forward process as diffusion models and utilize the reconstruction loss similar to discriminative methods. Furthermore, we devise a two-stage training strategy to emulate the inference process during model training. DDTSE not only works as a standalone system, but also can further improve the performance of discriminative models without additional retraining. Experimental results demonstrate that DDTSE not only achieves higher perceptual quality but also accelerates the inference process by 3 times compared to the conventional diffusion model.
Leying Zhang, Yao Qian, Linfeng Yu, Heming Wang, Hemin Yang, Shujie Liu 0001, Yanmin Qian
SLT6
2024 Advanced Long-Content Speech Recognition With Factorized Neural Transducer
abstract
Long-form automatic speech recognition (ASR) has obtained increasing interest in recent years, as it captures the relationship among consecutive historical sentences while decoding the current sentence. In this paper, we propose two novel approaches, which integrate long-form information into the factorized neural transducer (FNT) based architecture in both non-streaming (referred to asLongFNT) and streaming (referred to asSLongFNT) scenarios. We first investigate whether long-form transcriptions can improve the vanilla conformer transducer (C-T) models. Our experiments indicate that the vanilla C-T models do not exhibit improved performance when utilizing long-form transcriptions, possibly due to the predictor network of C-T models not functioning as a pure language model. Instead, FNT shows its potential in utilizing long-form information, where we propose theLongFNTmodel and explore the impact of long-form information in both text (LongFNT-Text) and speech (LongFNT-Speech). The proposed LongFNT-Text and LongFNT-Speech models further complement each other to achieve better performance, with transcription history proving more valuable to the model. The effectiveness of our LongFNT approach is evaluated on LibriSpeech and GigaSpeech corpora, and obtains relative 19% and 12% word error rate reduction, respectively. Furthermore, we extend the LongFNT model to the streaming scenario, which is namedSLongFNT, consisting of SLongFNT-Text and SLongFNT-Speech approaches to utilize long-form text and speech information. Experiments show that the proposed SLongFNT model achieves relative 26% and 17% WER reduction on LibriSpeech and GigaSpeech respectively while keeping a good latency, compared to the FNT baseline. Overall, our proposedLongFNTandSLongFNThighlight the significance of considering long-form speech and transcription knowledge for improving both non-streaming and streaming speech recognition systems.
Xun Gong 0005, Yu Wu 0012, Jinyu Li 0001, Shujie Liu 0001, Rui Zhao 0017, Xie Chen 0001, Yanmin Qian
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
abstract
Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text speech generation tasks involving transforming input speech and processing audio captured in adverse acoustic conditions. This paper introduces SpeechX, a versatile speech generation model capable of zero-shot TTS and various speech transformation tasks, dealing with both clean and noisy signals. SpeechX combines neural codec language modeling with multi-task learning using task-dependent prompting, enabling unified and extensible modeling and providing a consistent way for leveraging textual input in speech enhancement and transformation tasks. Experimental results show SpeechX's efficacy in various tasks, including zero-shot TTS, noise suppression, target speaker extraction, speech removal, and speech editing with or without background noise, achieving comparable or superior performance to specialized models across tasks.
Xiaofei Wang 0007, Manthan Thakker, Zhuo Chen 0006, Naoyuki Kanda, Sefik Emre Eskimez, Sanyuan Chen, Shujie Liu 0001, Jinyu Li 0001, Takuya Yoshioka
IEEE ACM Trans. Audio Speech Lang. Process.8
2024 VioLA: Conditional Language Models for Speech Recognition, Synthesis, and Translation
abstract
Recent research shows a big convergence in model architecture, training objectives, and inference methods across various tasks for different modalities. In this paper, we proposeVioLA, a single auto-regressive Transformer decoder-only network that unifies various cross-modal tasks involving speech and text, such as speech-to-text, text-to-text, text-to-speech, and speech-to-speech tasks, as a conditional language model task via multi-task learning framework. To accomplish this, we first convert the speech utterances to discrete tokens (similar to the textual data) using an offline neural codec encoder. In such a way, all these tasks are converted to token-based sequence prediction problems, which can be naturally handled with one conditional language model. We further integrate task IDs (TID), language IDs (LID), and LSTM-based acoustic embedding into the proposed model to enhance the modeling capability of handling different languages and tasks. Experimental results demonstrate that the proposedVioLAmodel can support both single-modal and cross-modal tasks well, and the decoder-only model achieves a comparable and even better performance than the strong baselines.
Tianrui Wang, Yu Wu 0012, Shujie Liu 0001, Yashesh Gaur, Zhuo Chen 0006, Jinyu Li 0001, Furu Wei
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 SpeechLM: Enhanced Speech Pre-Training With Unpaired Textual Data
abstract
How to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics. In this paper, we propose a cross-modalSpeechandLanguageModel (SpeechLM) to explicitly align speech and text pre-training with a pre-defined unified discrete representation. Specifically, we introduce two alternative discrete tokenizers to bridge the speech and text modalities, including phoneme-unit and hidden-unit tokenizers, which can be trained using unpaired speech or a small amount of paired speech-text data. Based on the trained tokenizers, we convert the unlabeled speech and text data into tokens of phoneme units or hidden units. The pre-training objective is designed to unify the speech and the text into the same discrete semantic space with a unified Transformer network. We evaluate SpeechLM on various spoken language processing tasks including speech recognition, speech translation, and universal representation evaluation framework SUPERB, demonstrating significant improvements on content-related tasks. Code and models are available athttps://aka.ms/SpeechLM.
Sanyuan Chen, Yu Wu 0012, Shuo Ren 0002, Shujie Liu 0001, Zhuoyuan Yao, Xun Gong 0005, Li-Rong Dai 0001, Jinyu Li 0001, Furu Wei
IEEE ACM Trans. Audio Speech Lang. Process.6
2024 VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation Learning
abstract
Although speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework to integrate different modal information and leverage different resources (e.g., visual-audio pairs, audio-text pairs, unlabeled speech, and unlabeled text) to facilitate speech representation learning was not well explored. In this paper, we propose a unified cross-modal representation learning frameworkVatLM(Visual-Audio-Text Language Model). The proposedVatLMemploys a unified backbone network to model the modality-independent information and utilizes three simple modality-dependent modules to preprocess visual, speech, and text inputs. In order to integrate these three modalities into one shared semantic space,VatLMis optimized with a masked prediction task of unified tokens, given by our proposed unified tokenizer. We evaluate the pre-trainedVatLMon audio-visual related downstream tasks, including audio-visual speech recognition (AVSR), and visual speech recognition (VSR) tasks. Results show that the proposedVatLMoutperforms previous state-of-the-art models, such as the audio-visual pre-trained AV-HuBERT model, and analysis also demonstrates thatVatLMis capable of aligning different modalities into the same space. To facilitate future research, we release the code and pre-trained models athttps://aka.ms/vatlm.
Qiushi Zhu, Shujie Liu 0001, Binxing Jiao, Jie Zhang 0042, Li-Rong Dai 0001, Daxin Jiang, Jinyu Li 0001, Furu Wei
IEEE Trans. Multim.4
2023 Prompting Large Language Models for Zero-Shot Domain Adaptation in Speech Recognition
abstract
The integration of Language Models (LMs) has proven to be an effective way to address domain shifts in speech recognition. However, these approaches usually require a significant amount of target domain text data for the training of LMs. Different from these methods, in this work, with only a domain-specific text prompt, we propose two zero-shot ASR domain adaptation methods using LLaMA, a 7-billionparameter large language model (LLM). LLM is used in two ways: 1) second-pass rescoring: reranking N-best hypotheses of a given ASR system with LLaMA; 2) deep LLM-fusion: incorporating LLM into the decoder of an encoder-decoder based ASR system. Experiments show that, with only one domain prompt, both methods can effectively reduce word error rates (WER) on out-of-domain TedLium-2 and SPGISpeech datasets. Especially, the deep LLM-fusion has the advantage of better recall of entity and out-of-vocabulary words.
Yuang Li, Yu Wu 0012, Jinyu Li 0001, Shujie Liu 0001
ASRU4
2023 Building High-Accuracy Multilingual ASR With Gated Language Experts and Curriculum Training
abstract
We propose gated language experts and curriculum training to enhance multilingual transformer transducer models without requiring user input for language identification (LID) during inference. Our method incorporates a gating mechanism and LID loss to enable transformer experts to learn language-specific information. Linear experts are applied on joint network to stabilize training. The curriculum training scheme leverages LID to guide gated experts in improving their respective language-specific performance. Experimental results on an English and Spanish bilingual task show significant average relative word error reductions of 12.5 % and 7.3 % compared to the baseline bilingual and monolingual models, respectively. Our models even perform similarly to upper-bound models with oracle LID. Extending our approach to trilingual, quadrilingual, and pentalingual models reveals similar advantages to those seen in the bilingual models, highlighting its ease of extension to multiple languages.
Eric Sun, Jinyu Li 0001, Yuxuan Hu 0003, Yimeng Zhu, Linquan Liu, Shujie Liu 0001, Edward Lin, Yifan Gong 0001
ASRU9
2023 On Decoder-Only Architecture For Speech-to-Text and Large Language Model Integration
abstract
Large language models (LLMs) have achieved remarkable success in the field of natural language processing, enabling better human-computer interaction using natural language. However, the seamless integration of speech signals into LLMs has not been explored well. The “decoder-only“ architecture has also not been well studied for speech processing tasks. In this research, we introduce Speech-LLaMA, a novel approach that effectively incorporates acoustic information into text-based large language models. Our method leverages Connectionist Temporal Classification and a simple audio encoder to map the compressed acoustic features to the continuous semantic space of the LLM. In addition, we further probe the decoder-only architecture for speech-to-text tasks by training a smaller scale randomly initialized speech-LLaMA model from speech-text paired data alone. We conduct experiments on multilingual speech-to-text translation tasks and demonstrate a significant improvement over strong baselines, highlighting the potential advantages of decoder-only models for speech-to-text conversion.
Jian Wu 0027, Yashesh Gaur, Zhuo Chen 0006, Yimeng Zhu, Tianrui Wang, Jinyu Li 0001, Shujie Liu 0001, Linquan Liu, Yu Wu 0012
ASRU8
2023 Prosody-Aware Speecht5 for Expressive Neural TTS
abstract
SpeechT5, a multimodal learning framework which explores encoder-decoder pre-training by leveraging both unlabeled speech and text, has been proven to be effective on a wide variety of speech processing tasks. In this paper, we enhance SpeechT5 by adding a new sub-task on prosody modeling (prosody-aware SpeechT5) for neural text-to-speech (TTS), which can improve the model capability to learn richer contextual representations through multi-task learning. In the prosody-aware SpeechT5 training framework, most modules in neural TTS can be pre-trained with large-scale unlabeled speech and text corpus, including encoder, decoder, and variance adaptor. Experimental results show that the proposed prosody-aware SpeechT5 is effective at improving the expressiveness of neural TTS1: 1) the CMOS (comparison mean opinion score) gain is 0.154 for texts from news domain and 0.114 for texts from audiobook domain; 2) the prosody related issues in synthetic speech are reduced by 19.02% in subjective evaluation.
Yuanhao Yi, Shujie Liu 0001, Lei He 0005
ICASSP4
2023 LongFNT: Long-Form Speech Recognition with Factorized Neural Transducer
abstract
Traditional automatic speech recognition (ASR) systems usually focus on individual utterances, without considering long-form speech with useful historical information, which is more practical in real scenarios. Simply attending longer transcription history for a vanilla neural transducer model shows no much gain in our preliminary experiments, since the prediction network is not a pure language model. This motivates us to leverage the factorized neural transducer structure, containing a real language model, the vocabulary predictor. We propose the LongFNT-Text architecture, which fuses the sentence-level long-form features directly with the output of the vocabulary predictor and then embeds token-level long-form features inside the vocabulary predictor, with a pre-trained contextual encoder RoBERTa to further boost the performance. Moreover, we propose the LongFNT architecture by extending the long-form speech to the original speech input and achieve the best performance. The effectiveness of our LongFNT approach is validated on LibriSpeech and GigaSpeech corpora with 19% and 12% relative word error rate (WER) reduction, respectively.
Xun Gong 0005, Yu Wu 0012, Jinyu Li 0001, Shujie Liu 0001, Rui Zhao 0017, Xie Chen 0001, Yanmin Qian
ICASSP4
2023 Target Sound Extraction with Variable Cross-Modality Clues
abstract
Automatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture of sources. It often uses a model conditioned on a fixed form of target sound clues, such as a sound class label, which limits the ways in which users can interact with the model to specify the target sounds. To leverage variable number of clues cross modalities available in the inference phase, including a video, a sound event class, and a text caption, we propose a unified transformer-based TSE model architecture, where a multi-clue attention module integrates all the clues across the modalities. Since there is no off-the-shelf benchmark to evaluate our proposed approach, we build a dataset1based on public corpora, Audioset and AudioCaps. Experimental results for seen and unseen target-sound evaluation sets show that our proposed TSE model can effectively deal with a varying number of clues which improves the TSE performance and robustness against partially compromised clues.
Chenda Li, Yao Qian, Zhuo Chen 0006, Dongmei Wang, Takuya Yoshioka, Shujie Liu 0001, Yanmin Qian, Michael Zeng 0001
ICASSP6
2023 DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation Tasks
abstract
Self-supervised learning has been successfully applied to various speech recognition and understanding tasks. However, for generative tasks such as speech enhancement and speech separation, most self-supervised speech representations did not show substantial improvements. To deal with this problem, in this paper, we propose data2vec-SG (Speech Generation), which is a teacher-student learning framework that addresses speech generation tasks. Our data2vec-SG introduces a reconstruction module into data2vec [1] and enforces the representations to contain not only the semantic information but also the acoustic knowledge to generate clean speech waveforms. Experimental results demonstrate that the proposed framework boosts the performance of various speech generation tasks including speech enhancement, speech separation, and packet loss concealment. Meanwhile, the learned representation is also capable of helping other downstream tasks, which is demonstrated by the good performance in the speech recognition task in both clean and noisy conditions.
Heming Wang, Yao Qian, Hemin Yang, Naoyuki Kanda, Takuya Yoshioka, Xiaofei Wang 0009, Shujie Liu 0001, Zhuo Chen 0006, DeLiang Wang, Michael Zeng 0001
ICASSP9
2023 Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation
abstract
Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity problem because the corpora from the speech of the source language to the speech of the target language are very rare. To address this issue, we propose in this paper a Speech2S model, which is jointly pre-trained with unpaired speech and bilingual text data for direct speech-to-speech translation tasks. By effectively leveraging the paired text data, Speech2S is capable of modeling the cross-lingual speech conversion from source to target language. We verify the performance of the proposed Speech2S on Europarl-ST and VoxPopuli datasets. Experimental results demonstrate that Speech2S gets an improvement of about 5 BLEU scores compared to encoder-only pre-training models, and achieves a competitive or even better performance than existing state-of-the-art models1.
Shujie Liu 0001, Lei He 0005, Jinyu Li 0001, Furu Wei
ICASSP5
2023 Code-Switching Text Generation and Injection in Mandarin-English ASR
abstract
Code-switching speech refers to a means of expression by mixing two or more languages within a single utterance. Automatic Speech Recognition (ASR) with End-to-End (E2E) modeling for such speech can be a challenging task due to the lack of data. In this study, we investigate text generation and injection for improving the performance of an industry commonly-used streaming model, Transformer-Transducer (T-T), in Mandarin-English code-switching speech recognition. We first propose a strategy to generate codeswitching text data and then investigate injecting generated text into T-T model explicitly by Text-To-Speech (TTS) conversion or implicitly by tying speech and text latent spaces. Experimental results on the T-T model trained with a dataset containing 1,800 hours of real Mandarin-English code-switched speech show that our approaches to inject generated code-switching text significantly boost the performance of T-T models, i.e., 16% relative Token-based Error Rate (TER) reduction averaged on three evaluation sets, and the approach of tying speech and text latent spaces is superior to that of TTS conversion on the evaluation set which contains more homogeneous data with the training set.
Yuxuan Hu 0003, Yao Qian, Ma Jin, Linquan Liu, Shujie Liu 0001, Yu Shi 0001, Yanmin Qian, Edward Lin, Michael Zeng 0001
ICASSP6
2023 Robust Data2VEC: Noise-Robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning
abstract
Self-supervised pre-training methods based on contrastive learning or regression tasks can utilize more unlabeled data to improve the performance of automatic speech recognition (ASR). However, the robustness impact of combining the two pre-training tasks and constructing different negative samples for contrastive learning still remains unclear. In this paper, we propose a noise-robust data2vec for self-supervised speech representation learning by jointly optimizing the contrastive learning and regression tasks in the pre-training stage. Furthermore, we present two improved methods to facilitate contrastive learning. More specifically, we first propose to construct patch-based non-semantic negative samples to boost the noise robustness of the pre-training model, which is achieved by dividing the features into patches at different sizes (i.e., so-called negative samples). Second, by analyzing the distribution of positive and negative samples, we propose to remove the easily distinguishable negative samples to improve the discriminative capacity for pre-training models. Experimental results on the CHiME-4 dataset show that our method is able to improve the performance of the pre-trained model in noisy scenarios. We find that joint training of the contrastive learning and regression tasks can avoid the model collapse to some extent compared to only training the regression task.
Qiushi Zhu, Jie Zhang 0042, Shujie Liu 0001, Yu-Chen Hu, Li-Rong Dai 0001
ICASSP4
2023 BEATs: Audio Pre-Training with Acoustic Tokenizers
abstract
We introduce a self-supervised learning (SSL) framework BEATs for general audio representation pre-training, where we optimize an acoustic tokenizer and an audio SSL model by iterations. Unlike the previous audio SSL models that employ reconstruction loss for pre-training, our audio SSL model is trained with the discrete label prediction task, where the labels are generated by a semantic-rich acoustic tokenizer. We propose an iterative pipeline to jointly optimize the tokenizer and the pre-trained model, aiming to abstract high-level semantics and discard the redundant details for audio. The experimental results demonstrate our acoustic tokenizers can generate discrete labels with rich audio semantics and our audio SSL models achieve state-of-the-art (SOTA) results across various audio classification benchmarks, even outperforming previous models that use more training data and model parameters significantly. Specifically, we set a new SOTA mAP 50.6% on AudioSet-2M without using any external data, and 98.1% accuracy on ESC-50. The code and pre-trained models are available at https://aka.ms/beats.
Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Shujie Liu 0001, Daniel Tompkins, Zhuo Chen 0006, Wanxiang Che, Xiangzhan Yu, Furu Wei
ICML4
2023 GRAVO: Learning to Generate Relevant Audio from Visual Features with Noisy Online Videos
Youngdo Ahn, Chengyi Wang 0002, Yu Wu 0012, Jong Won Shin, Shujie Liu 0001
INTERSPEECH5
2023 Accelerating Transducers through Adjacent Token Merging
Yuang Li, Yu Wu 0012, Jinyu Li 0001, Shujie Liu 0001
INTERSPEECH4
2023 LAMASSU: A Streaming Language-Agnostic Multilingual Speech Recognition and Translation Model Using Neural Transducers
Eric Sun, Yu Wu 0012, Yashesh Gaur, Shujie Liu 0001, Jinyu Li 0001
INTERSPEECH7
2023 ComSL: A Composite Speech-Language Model for End-to-End Speech-to-Text Translation
abstract
Joint speech-language training is challenging due to the large demand for training data and GPU consumption, as well as the modality gap between speech and language. We present ComSL, a speech-language model built atop a composite architecture of public pre-trained speech-only and language-only models and optimized data-efficiently for spoken language tasks. Particularly, we propose to incorporate cross-modality learning into transfer learning and conduct them simultaneously for downstream tasks in a multi-task learning manner. Our approach has demonstrated effectiveness in end-to-end speech-to-text translation tasks, achieving a new state-of-the-art average BLEU score of 31.5 on the multilingual speech to English text translation task for 21 languages, as measured on the public CoVoST2 evaluation set.
Chenyang Le, Yao Qian, Shujie Liu 0001, Yanmin Qian, Michael Zeng 0001, Xuedong Huang 0001
NeurIPS4
2022 SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
abstract
Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Junyi Ao, Rui Wang 0073, Chengyi Wang 0002, Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Tom Ko, Qing Li 0001, Yu Zhang 0006, Zhihua Wei 0001, Yao Qian, Jinyu Li 0001, Furu Wei
ACL (1)7
2022 SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training
abstract
The rapid development of single-modal pretraining has prompted researchers to pay more attention to cross-modal pre-training methods.In this paper, we propose a unified-modal speech-unit-text pre-training model, SpeechUT, to connect the representations of a speech encoder and a text decoder with a shared unit encoder.Leveraging hidden-unit as an interface to align speech and text, we can decompose the speech-to-text model into a speech-to-unit model and a unit-to-text model, which can be jointly pre-trained with unpaired speech and text data respectively.Our proposed SpeechUT is fine-tuned and evaluated on automatic speech recognition (ASR) and speech translation (ST) tasks.Experimental results show that SpeechUT gets substantial improvements over strong baselines, and achieves state-of-the-art performance on both the Lib-riSpeech ASR and MuST-C ST tasks.To better understand the proposed SpeechUT, detailed analyses are conducted.The code and pretrained models are available at https://aka. ms/SpeechUT.
Junyi Ao, Shujie Liu 0001, Li-Rong Dai 0001, Jinyu Li 0001, Furu Wei
EMNLP4
2022 Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
abstract
The speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks. In this paper, we explore the limits of speech representations learned by different self-supervised objectives and datasets for automatic speaker verification (ASV), especially with a well-recognized SOTA ASV model, ECAPA-TDNN [1], as a downstream model. The representations from all hidden layers of the pre-trained model are firstly averaged with learnable weights and then fed into the ECAPA-TDNN as input features. The experimental results on Voxceleb dataset show that the weighted average representation is significantly superior to FBank, a conventional handcrafted feature for ASV. Our best single system achieves 0.537%, 0.569%, and 1.180% equal error rate (EER) on the three official trials of VoxCeleb1, separately. Accordingly, the ensemble system with three pre-trained models can further improve the EER to 0.479%, 0.536% and 1.023%. Among the three evaluation trials, our best system outperforms the winner system [2] of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC2021) on the VoxCeleb1-E trial.
Zhengyang Chen, Sanyuan Chen, Yu Wu 0012, Yao Qian, Chengyi Wang 0002, Shujie Liu 0001, Yanmin Qian, Michael Zeng 0001
ICASSP6
2022 Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-Training
abstract
Self-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years have witnessed great successes in applying self-supervised learning in speech recognition, while limited exploration was attempted in applying SSL for modeling speaker characteristics. In this paper, we aim to improve the existing SSL framework for speaker representation learning. Two methods are introduced for enhancing the unsupervised speaker information extraction. First, we apply multi-task learning to the current SSL framework, where we integrate utterance-wise contrastive loss with the SSL objective function. Second, for better speaker discrimination, we propose an utterance mixing strategy for data augmentation, where additional overlapped utterances are created unsupervisely and incorporated during training. We integrate the proposed methods into the HuBERT framework. Experiment results on the SUPERB benchmark show that the proposed system achieves state-of-the-art performance in universal representation learning, especially for speaker identification oriented tasks. An ablation study is performed verifying the efficacy of each proposed method. Finally, we scale up the training dataset to 94 thousand hours of public audio data and achieve further performance improvement in all SUPERB tasks.
Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Zhengyang Chen, Zhuo Chen 0006, Shujie Liu 0001, Jian Wu 0027, Yao Qian, Furu Wei, Jinyu Li 0001, Xiangzhan Yu
ICASSP6
2022 Multi-View Self-Attention Based Transformer for Speaker Recognition
abstract
Initially developed for natural language processing (NLP), Transformer model is now widely used for speech processing tasks such as speaker recognition, due to its powerful sequence modeling capabilities. However, conventional self-attention mechanisms are originally designed for modeling textual sequence without considering the characteristics of speech and speaker modeling. Besides, different Transformer variants for speaker recognition have not been well studied. In this work, we propose a novel multi-view self-attention mechanism and present an empirical study of different Transformer variants with or without the proposed attention mechanism for speaker recognition. Specifically, to balance the capabilities of capturing global dependencies and modeling the locality, we propose a multi-view self-attention mechanism for speaker Transformer, in which different attention heads can attend to different ranges of the receptive field. Furthermore, we introduce and compare five Transformer variants with different network architectures, embedding locations, and pooling methods to learn speaker embeddings. Experimental results on the VoxCeleb1 and VoxCeleb2 datasets show that the proposed multi-view self-attention mechanism achieves improvement in the performance of speaker recognition, and the proposed speaker Transformer network attains excellent results compared with state-of-the-art models.
Rui Wang 0073, Junyi Ao, Shujie Liu 0001, Zhihua Wei 0001, Tom Ko, Qing Li 0001, Yu Zhang 0006
ICASSP4
2022 Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction
abstract
Noise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a preprocessing module that conducts speech enhancement, and then feed the enhanced speech to an ASR backend. In this work, instead of suppressing background noise with a conventional cascaded pipeline, we employ a noise-robust representation learned by a refined self-supervised framework for noisy speech recognition. We propose to combine a reconstruction module with contrastive learning and perform multi-task continual pre-training on noisy data. The reconstruction module is used for auxiliary learning to improve the noise robustness of the learned representation and thus is not required during inference. Experiments demonstrate the effectiveness of our proposed method. Our model substantially reduces the word error rate (WER) for the synthesized noisy LibriSpeech test sets, and yields around 4.1/7.5% WER reduction on noisy clean/other test sets compared to data augmentation. For the real-world noisy speech from the CHiME-4 challenge (1-channel track), we have obtained the state of the art ASR performance without any denoising front-end. Moreover, we achieve comparable performance to the best supervised approach reported with only 16% of labeled data.
Heming Wang, Yao Qian, Xiaofei Wang 0009, Chengyi Wang 0002, Shujie Liu 0001, Takuya Yoshioka, Jinyu Li 0001, DeLiang Wang
ICASSP6
2022 Optimizing Alignment of Speech and Language Latent Spaces for End-To-End Speech Recognition and Understanding
abstract
The advances in attention-based encoder-decoder (AED) networks have brought great progress to end-to-end (E2E) automatic speech recognition (ASR). One way to further improve the performance of AED-based E2E ASR is to introduce an extra text encoder for leveraging extensive text data and thus capture more context-aware linguistic information. However, this approach brings a mismatch problem between the speech encoder and the text encoder due to the different units used for modeling. In this paper, we propose an embedding aligner and modality switch training to better align the speech and text latent spaces. The embedding aligner is a shared linear projection between text encoder and speech encoder trained by masked language modeling (MLM) loss and connectionist temporal classification (CTC), respectively. The modality switch training randomly swaps speech and text embeddings based on the forced alignment result to learn a joint representation space. Experimental results show that our proposed approach achieves a relative 14% to 19% word error rate (WER) reduction on Librispeech ASR task. We further verify its effectiveness on spoken language understanding (SLU), i.e., an absolute 2.5% to 2.8% F1 score improvement on SNIPS slot filling task.
Wei Wang 0010, Shuo Ren 0002, Yao Qian, Shujie Liu 0001, Yu Shi 0001, Yanmin Qian, Michael Zeng 0001
ICASSP4
2022 Improving Self-Supervised Learning for Speech Recognition with Intermediate Layer Supervision
abstract
Recently, pioneer work finds that self-supervised pre-training methods can improve multiple downstream speech tasks, because the model utilizes bottom layers to learn speaker-related information and top layers to encode content-related information. Since the network capacity is limited, we believe the speech recognition performance could be further improved if the model is dedicated to audio content information learning. To this end, we propose Intermediate Layer Supervision for Self-Supervised Learning (ILS-SSL), which forces the model to concentrate on content information as much as possible by adding an additional SSL loss on the intermediate layers. Experiments on LibriSpeech test-other set show that our method outperforms HuBERT significantly, which achieves a 23.5%/11.6% relative word error rate reduction in the w/o language model setting for Base/Large models. Detailed analysis shows the bottom layers of our model have a better correlation with phonetic units, which is consistent with our intuition and explains the success of our method for ASR. We will release our code and model at https://github.com/microsoft/UniSpeech.
Chengyi Wang 0002, Yu Wu 0012, Sanyuan Chen, Shujie Liu 0001, Jinyu Li 0001, Yao Qian, Zhenglu Yang
ICASSP4
2022 A Configurable Multilingual Model is All You Need to Recognize All Languages
abstract
Multilingual automatic speech recognition models have shown great promise in recent years because of the simple model training and deployment process. Conventional methods either train a universal multilingual model without taking any language information or with a 1-hot language ID (LID) vector to guide the recognition of the target language. In practice, a multilingual user can be prompted to preselect several languages he/she can speak. The multilingual model without LID cannot well utilize the language information set by the user while the multilingual model with 1-hot LID can only handle one pre-selected language. In this paper, we propose a novel configurable multilingual model (CMM) which is trained only once but can be configured as different models based on users’ choices by extracting language-specific modules together with a universal module from the trained CMM. Particularly, a single CMM can be deployed to any user scenario where the users can pre-select any combination of languages. Trained with 75K hours of transcribed anonymized Microsoft multilingual data and evaluated with 10-language test sets, the proposed CMM improves from the universal multilingual model by 26.0%, 16.9%, and 10.4% relative word error reduction when the user selects 1, 2, or 3 languages, respectively.
Jinyu Li 0001, Eric Sun, Shujie Liu 0001
ICASSP4
2022 Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data
abstract
This paper studies a novel pre-training technique with unpaired speech data, Speech2C, for encoder-decoder based automatic speech recognition (ASR).Within a multi-task learning framework, we introduce two pre-training tasks for the encoderdecoder network using acoustic units, i.e., pseudo codes, derived from an offline clustering model.One is to predict the pseudo codes via masked language modeling in encoder output, like HuBERT model, while the other lets the decoder learn to reconstruct pseudo codes autoregressively instead of generating textual scripts.In this way, the decoder learns to reconstruct original speech information with codes before learning to generate correct text.Comprehensive experiments on the LibriSpeech corpus show that the proposed Speech2C can relatively reduce the word error rate (WER) by 19.2% over the method without decoder pre-training, and also outperforms significantly the state-of-the-art wav2vec 2.0 and Hu-BERT on fine-tuning subsets of 10h and 100h.We release our code and model at https://github.com/microsoft/SpeechT5/tree/main/Speech2C.
Junyi Ao, Shujie Liu 0001, Haizhou Li 0001, Tom Ko, Li-Rong Dai 0001, Jinyu Li 0001, Yao Qian, Furu Wei
INTERSPEECH4
2022 Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?
abstract
Recently, self-supervised learning (SSL) has demonstrated strong performance in speaker recognition, even if the pretraining objective is designed for speech recognition.In this paper, we study which factor leads to the success of selfsupervised learning on speaker-related tasks, e.g.speaker verification (SV), through a series of carefully designed experiments.Our empirical results on the Voxceleb-1 dataset suggest that the benefit of SSL to SV task is from a combination of mask speech prediction loss, data scale, and model size, while the SSL quantizer has a minor impact.We further employ the integrated gradients attribution method and loss landscape visualization to understand the effectiveness of self-supervised learning for speaker recognition performance.
Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Shujie Liu 0001, Zhuo Chen 0006, Gang Liu 0001, Jinyu Li 0001, Jian Wu 0027, Xiangzhan Yu, Furu Wei
INTERSPEECH4
2022 Speech Pre-training with Acoustic Piece
abstract
Previous speech pre-training methods, such as wav2vec2.0 and HuBERT, pre-train a Transformer encoder to learn deep representations from audio data, with objectives predicting either elements from latent vector quantized space or pre-generated labels (known as target codes) with offline clustering. However, those training signals (quantized elements or codes) are independent across different tokens without considering their relations. According to our observation and analysis, the target codes share obvious patterns aligned with phonemized text data. Based on that, we propose to leverage those patterns to better pre-train the model considering the relations among the codes. The patterns we extracted, called "acoustic piece"s, are from the sentence piece result of HuBERT codes. With the acoustic piece as the training signal, we can implicitly bridge the input audio and natural language, which benefits audio-to-text tasks, such as automatic speech recognition (ASR). Simple but effective, our method "HuBERT-AP" significantly outperforms strong baselines on the LibriSpeech ASR task.
Shuo Ren 0002, Shujie Liu 0001, Yu Wu 0012, Furu Wei
INTERSPEECH2
2022 Supervision-Guided Codebooks for Masked Prediction in Speech Pre-training
abstract
Recently, masked prediction pre-training has seen remarkable progress in self-supervised learning (SSL) for speech recognition.It usually requires a codebook obtained in an unsupervised way, making it less accurate and difficult to interpret.We propose two supervision-guided codebook generation approaches to improve automatic speech recognition (ASR) performance and also the pre-training efficiency, either through decoding with a hybrid ASR system to generate phoneme-level alignments (named PBERT), or performing clustering on the supervised speech features extracted from an end-to-end CTC model (named CTC clustering).Both the hybrid and CTC models are trained on the same small amount of labeled speech as used in fine-tuning.Experiments demonstrate significant superiority of our methods to various SSL and self-training baselines, with up to 17.0% relative WER reduction.Our pre-trained models also show good transferability in a non-ASR speech task.
Chengyi Wang 0002, Yu Wu 0012, Sanyuan Chen, Jinyu Li 0001, Shujie Liu 0001, Furu Wei
INTERSPEECH6
2022 Separating Long-Form Speech with Group-wise Permutation Invariant Training
abstract
Multi-talker conversational speech processing has drawn many interests for various applications such as meeting transcription.Speech separation is often required to handle overlapped speech that is commonly observed in conversation.Although the original utterancelevel permutation invariant training-based continuous speech separation approach has proven to be effective in various conditions, it lacks the ability to leverage the long-span relationship of utterances and is computationally inefficient due to the highly overlapped sliding windows.To overcome these drawbacks, we propose a novel training scheme named Group-PIT, which allows direct training of the speech separation models on the long-form speech with a low computational cost for label assignment.Two different speech separation approaches with Group-PIT are explored, including direct long-span speech separation and short-span speech separation with long-span tracking.The experiments on the simulated meeting-style data demonstrate the effectiveness of our proposed approaches, especially in dealing with a very long speech input.
Wangyou Zhang, Zhuo Chen 0006, Naoyuki Kanda, Shujie Liu 0001, Jinyu Li 0001, Sefik Emre Eskimez, Takuya Yoshioka, Zhong Meng, Yanmin Qian, Furu Wei
INTERSPEECH4
2022 Two-Stream Network for Sign Language Recognition and Translation
abstract
Sign languages are visual languages using manual articulations and non-manual elements to convey information. For sign language recognition and translation, the majority of existing approaches directly encode RGB videos into hidden representations. RGB videos, however, are raw signals with substantial visual redundancy, leading the encoder to overlook the key information for sign language understanding. To mitigate this problem and better incorporate domain knowledge, such as handshape and body movement, we introduce a dual visual encoder containing two separate streams to model both the raw videos and the keypoint sequences generated by an off-the-shelf keypoint estimator. To make the two streams interact with each other, we explore a variety of techniques, including bidirectional lateral connection, sign pyramid network with auxiliary supervision, and frame-level self-distillation. The resulting model is called TwoStream-SLR, which is competent for sign language recognition (SLR). TwoStream-SLR is extended to a sign language translation (SLT) model, TwoStream-SLT, by simply attaching an extra translation network. Experimentally, our TwoStream-SLR and TwoStream-SLT achieve state-of-the-art performance on SLR and SLT tasks across a series of datasets including Phoenix-2014, Phoenix-2014T, and CSL-Daily.
Ronglai Zuo, Fangyun Wei, Yu Wu 0012, Shujie Liu 0001, Brian Kan-Wing Mak
NeurIPS5
2022 Exploring WavLM on Speech Enhancement
abstract
There is a surge in interest in self-supervised learning approaches for end-to-end speech encoding in recent years as they have achieved great success. Especially, WavLM showed state-of-the-art performance on various speech processing tasks. To better understand the efficacy of self-supervised learning models for speech enhancement, in this work, we design and conduct a series of experiments with three resource conditions by combining WavLM and two high-quality speech enhancement systems. Also, We propose a regression-based WavLM training objective and a noise-mixing data configuration to further boost the downstream enhancement performance. The experiments on the DNS challenge dataset and a simulation dataset show that the WavLM benefits the speech enhancement task in terms of both speech quality and speech recognition accuracy, especially for low fine-tuning resources. For the high fine-tuning resource condition, only the word error rate is substantially improved.
Hyungchan Song, Sanyuan Chen, Zhuo Chen 0006, Yu Wu 0012, Takuya Yoshioka, Jong Won Shin, Shujie Liu 0001
SLT8
2021 SemFace: Pre-training Encoder and Decoder with a Semantic Interface for Neural Machine Translation
abstract
Shuo Ren, Long Zhou, Shujie Liu, Furu Wei, Ming Zhou, Shuai Ma. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Shuo Ren 0002, Shujie Liu 0001, Furu Wei, Ming Zhou 0001, Shuai Ma 0001
ACL/IJCNLP (1)3
2021 Jointly Learning to Repair Code and Generate Commit Message
abstract
We propose a novel task of jointly repairing program codes and generating commit messages.Code repair and commit message generation are two essential and related tasks for software development.However, existing work usually performs the two tasks independently.We construct a multilingual triple dataset including buggy code, fixed code, and commit messages for this novel task.We provide the cascaded models as baseline, which are enhanced with different training approaches, including the teacher-student method, the multi-task method, and the backtranslation method.To deal with the error propagation problem of the cascaded method, the joint model is proposed that can both repair the code and generate the commit message in a unified framework.Experimental results show that the enhanced cascaded model with teacher-student method and multitask-learning method achieves the best score on different metrics of automated code repair, and the joint model behaves better than the cascaded model on commit message generation.
Jiaqi Bai 0001, Ambrosio Blanco, Shujie Liu 0001, Furu Wei, Ming Zhou 0001, Zhoujun Li 0001
EMNLP (1)4
2021 Knowledge Enhanced Fine-Tuning for Better Handling Unseen Entities in Dialogue Generation
abstract
Although pre-training models have achieved great success in dialogue generation, their performance drops dramatically when the input contains an entity that does not appear in pretraining and fine-tuning datasets (unseen entity).To address this issue, existing methods leverage an external knowledge base to generate appropriate responses.In real-world scenario, the entity may not be included by the knowledge base or suffer from the precision of knowledge retrieval.To deal with this problem, instead of introducing knowledge base as the input, we force the model to learn a better semantic representation by predicting the information in the knowledge base, only based on the input context.Specifically, with the help of a knowledge base, we introduce two auxiliary training objectives: 1) Interpret Masked Word, which conjectures the meaning of the masked entity given the context; 2) Hypernym Generation, which predicts the hypernym of the entity based on the context.Experiment results on two dialogue corpus verify the effectiveness of our methods under both knowledge available and unavailable settings.
Leyang Cui, Yu Wu 0012, Shujie Liu 0001, Yue Zhang 0004
EMNLP (1)3
2021 Don't Shoot Butterfly with Rifles: Multi-Channel Continuous Speech Separation with Early Exit Transformer
abstract
With its strong modeling capacity that comes from a multi-head and multi-layer structure, Transformer is a very powerful model for learning a sequential representation and has been successfully applied to speech separation recently. However, multi-channel speech separation sometimes does not necessarily need such a heavy structure for all time frames especially when the cross-talker challenge happens only occasionally. For example, in conversation scenarios, most regions contain only a single active speaker, where the separation task downgrades to a single speaker enhancement problem. It turns out that using a very deep network structure for dealing with signals with a low overlap ratio not only negatively affects the inference efficiency but also hurts the separation performance. To deal with this problem, we propose an early exit mechanism, which enables the Transformer model to handle different cases with adaptive depth. Experimental results indicate that not only does the early exit mechanism accelerate the inference, but it also improves the accuracy.
Sanyuan Chen, Yu Wu 0012, Zhuo Chen 0006, Takuya Yoshioka, Shujie Liu 0001, Jinyu Li 0001, Xiangzhan Yu
ICASSP5
2021 Developing Real-Time Streaming Transformer Transducer for Speech Recognition on Large-Scale Dataset
abstract
Recently, Transformer based end-to-end models have achieved great success in many areas including speech recognition. However, compared to LSTM models, the heavy computational cost of the Transformer during inference is a key issue to prevent their applications. In this work, we explored the potential of Transformer Transducer (T-T) models for the fist pass decoding with low latency and fast speed on a large-scale dataset. We combine the idea of Transformer- XL and chunk-wise streaming processing to design a streamable Transformer Transducer model. We demonstrate that T-T outperforms the hybrid model, RNN Transducer (RNN-T), and streamable Transformer attention-based encoder-decoder model in the streaming scenario. Furthermore, the runtime cost and latency can be optimized with a relatively small look-ahead.
Xie Chen 0001, Yu Wu 0012, Shujie Liu 0001, Jinyu Li 0001
ICASSP4
2021 Continuous Speech Separation with Conformer
abstract
Continuous speech separation was recently proposed to deal with the overlapped speech in natural conversations. While it was shown to significantly improve the speech recognition performance for multichannel conversation transcription, its effectiveness has yet to be proven for a single-channel recording scenario. This paper examines the use of Conformer architecture in lieu of recurrent neural networks for the separation model. Conformer allows the separation model to efficiently capture both local and global context information, which is helpful for speech separation. Experimental results using the LibriCSS dataset show that the Conformer separation model achieves the state of the art results for both single-channel and multi-channel settings. Results for real meeting recordings are also presented, showing significant performance gains in both word error rate (WER) and speaker-attributed WER.
Sanyuan Chen, Yu Wu 0012, Zhuo Chen 0006, Jian Wu 0027, Jinyu Li 0001, Takuya Yoshioka, Chengyi Wang 0002, Shujie Liu 0001, Ming Zhou 0001
ICASSP8
2021 Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 2020
abstract
This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2020. We will first explain our system design to address issues in handling real multi-talker recordings. We then present the details of the components, which include Res2Net-based speaker embedding extractor, conformer-based continuous speech separation with leakage filtering, and a modified DOVER (short for Diarization Output Voting Error Reduction) method for system fusion. We evaluate the systems with the data set provided by VoxSRC challenge 2020, which contains real-life multi-talker audio collected from YouTube. Our best system achieves 3.71% and 6.23% of the diarization error rate (DER) on development set and evaluation set, respectively, being ranked the 1st at the diarization track of the challenge.
Naoyuki Kanda, Zhuo Chen 0006, Tianyan Zhou, Takuya Yoshioka, Sanyuan Chen, Yong Zhao 0008, Gang Liu 0001, Yu Wu 0012, Jian Wu 0027, Shujie Liu 0001, Jinyu Li 0001, Yifan Gong 0001
ICASSP11
2021 GraphCodeBERT: Pre-training Code Representations with Data Flow
Daya Guo, Shuo Ren 0002, Zhangyin Feng, Duyu Tang, Shujie Liu 0001, Nan Duan 0001, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin 0001, Daxin Jiang, Ming Zhou 0001
ICLR6
2021 UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data
abstract
In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both labeled and unlabeled data, in which supervised phonetic CTC learning and phonetically-aware contrastive self-supervised learning are conducted in a multi-task learning manner. The resultant representations can capture information more correlated with phonetic structures and improve the generalization across languages and domains. We evaluate the effectiveness of UniSpeech for cross-lingual representation learning on public CommonVoice corpus. The results show that UniSpeech outperforms self-supervised pretraining and supervised transfer learning for speech recognition by a maximum of 13.4% and 26.9% relative phone error rate reductions respectively (averaged over all testing languages). The transferability of UniSpeech is also verified on a domain-shift speech recognition task, i.e., a relative word error rate reduction of 6% against the previous approach.
Chengyi Wang 0002, Yu Wu 0012, Yao Qian, Ken'ichi Kumatani, Shujie Liu 0001, Furu Wei, Michael Zeng 0001, Xuedong Huang 0001
ICML5
2021 Ultra Fast Speech Separation Model with Teacher Student Learning
abstract
Transformer has been successfully applied to speech separation recently with its strong long-dependency modeling capacity using a self-attention mechanism. However, Transformer tends to have heavy run-time costs due to the deep encoder layers, which hinders its deployment on edge devices. A small Transformer model with fewer encoder layers is preferred for computational efficiency, but it is prone to performance degradation. In this paper, an ultra fast speech separation Transformer model is proposed to achieve both better performance and efficiency with teacher student learning (T-S learning). We introduce layer-wise T-S learning and objective shifting mechanisms to guide the small student model to learn intermediate representations from the large teacher model. Compared with the small Transformer model trained from scratch, the proposed T-S learning method reduces the word error rate (WER) by more than 5% for both multi-channel and single-channel speech separation on LibriCSS dataset. Utilizing more unlabeled speech data, our ultra fast speech separation models achieve more than 10% relative WER reduction.
Sanyuan Chen, Yu Wu 0012, Zhuo Chen 0006, Jian Wu 0027, Takuya Yoshioka, Shujie Liu 0001, Jinyu Li 0001, Xiangzhan Yu
Interspeech6
2021 Improving Multilingual Transformer Transducer Models by Reducing Language Confusions
Eric Sun, Jinyu Li 0001, Zhong Meng, Yu Wu 0012, Shujie Liu 0001, Yifan Gong 0001
Interspeech6
2021 Investigation of Practical Aspects of Single Channel Speech Separation for ASR
abstract
Speech separation has been successfully applied as a frontend processing module of conversation transcription systems thanks to its ability to handle overlapped speech and its flexibility to combine with downstream tasks such as automatic speech recognition (ASR). However, a speech separation model often introduces target speech distortion, resulting in a sub-optimum word error rate (WER). In this paper, we describe our efforts to improve the performance of a single channel speech separation system. Specifically, we investigate a two-stage training scheme that firstly applies a feature level optimization criterion for pretraining, followed by an ASR-oriented optimization criterion using an end-to-end (E2E) speech recognition model. Meanwhile, to keep the model light-weight, we introduce a modified teacher-student learning technique for model compression. By combining those approaches, we achieve a absolute average WER improvement of 2.70% and 0.77% using models with less than 10M parameters compared with the previous state-of-the-art results on the LibriCSS dataset for utterance-wise evaluation and continuous evaluation, respectively
Jian Wu 0027, Zhuo Chen 0006, Sanyuan Chen, Yu Wu 0012, Takuya Yoshioka, Naoyuki Kanda, Shujie Liu 0001, Jinyu Li 0001
Interspeech7
2020 RobuTrans: A Robust Transformer-Based Text-to-Speech Model
abstract
Recently, neural network based speech synthesis has achieved outstanding results, by which the synthesized audios are of excellent quality and naturalness. However, current neural TTS models suffer from the robustness issue, which results in abnormal audios (bad cases) especially for unusual text (unseen context). To build a neural model which can synthesize both natural and stable audios, in this paper, we make a deep analysis of why the previous neural TTS models are not robust, based on which we propose RobuTrans (Robust Transformer), a robust neural TTS model based on Transformer. Comparing to TransformerTTS, our model first converts input texts to linguistic features, including phonemic features and prosodic features, then feed them to the encoder. In the decoder, the encoder-decoder attention is replaced with a duration-based hard attention mechanism, and the causal self-attention is replaced with a "pseudo non-causal attention" mechanism to model the holistic information of the input. Besides, the position embedding is replaced with a 1-D CNN, since it constrains the maximum length of synthesized audio. With these modifications, our model not only fix the robustness problem, but also achieves on parity MOS (4.36) with TransformerTTS (4.37) and Tacotron2 (4.37) on our general set.
Naihan Li, Yu Wu 0012, Shujie Liu 0001, Sheng Zhao 0002
AAAI4
2020 Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech Translation
abstract
End-to-end speech translation, a hot topic in recent years, aims to translate a segment of audio into a specific language with an end-to-end model. Conventional approaches employ multi-task learning and pre-training methods for this task, but they suffer from the huge gap between pre-training and fine-tuning. To address these issues, we propose a Tandem Connectionist Encoding Network (TCEN) which bridges the gap by reusing all subnets in fine-tuning, keeping the roles of subnets consistent, and pre-training the attention module. Furthermore, we propose two simple but effective methods to guarantee the speech encoder outputs and the MT encoder inputs are consistent in terms of semantic representation and sequence length. Experimental results show that our model leads to significant improvements in En-De and En-Fr translation irrespective of the backbones.
Chengyi Wang 0002, Yu Wu 0012, Shujie Liu 0001, Zhenglu Yang, Ming Zhou 0001
AAAI3
2020 A Dataset for Low-Resource Stylized Sequence-to-Sequence Generation
abstract
Low-resource stylized sequence-to-sequence (S2S) generation is in high demand. However, its development is hindered by the datasets which have limitations on scale and automatic evaluation methods. We construct two large-scale, multiple-reference datasets for low-resource stylized S2S, the Machine Translation Formality Corpus (MTFC) that is easy to evaluate and the Twitter Conversation Formality Corpus (TCFC) that tackles an important problem in chatbots. These datasets contain context to source style parallel data, source style to target parallel data, and non-parallel sentences in the target style to enable the semi-supervised learning. We provide three baselines, the pivot-based method, the teacher-student method, and the back-translation method. We find that the pivot-based method is the worst, and the other two methods achieve the best score on different metrics.
Yu Wu 0012, Yunli Wang, Shujie Liu 0001
AAAI3
2020 MuTual: A Dataset for Multi-Turn Dialogue Reasoning
abstract
Non-task oriented dialogue systems have achieved great success in recent years due to largely accessible conversation data and the development of deep learning techniques.Given a context, current systems are able to yield a relevant and fluent response, but sometimes make logical mistakes because of weak reasoning capabilities.To facilitate the conversation reasoning research, we introduce Mu-Tual, a novel dataset for Multi-Turn dialogue Reasoning, consisting of 8,860 manually annotated dialogues based on Chinese student English listening comprehension exams.Compared to previous benchmarks for non-task oriented dialogue systems, MuTual is much more challenging since it requires a model that can handle various reasoning problems.Empirical results show that state-of-the-art methods only reach 71%, which is far behind the human performance of 94%, indicating that there is ample room for improving reasoning ability.MuTual is available at https://github. com/Nealcly/MuTual. * Contribution during internship at MSRA.M: Ma'am
Leyang Cui, Yu Wu 0012, Shujie Liu 0001, Yue Zhang 0004, Ming Zhou 0001
ACL3
2020 A Graph-based Coarse-to-fine Method for Unsupervised Bilingual Lexicon Induction
abstract
Unsupervised bilingual lexicon induction is the task of inducing word translations from monolingual corpora of two languages.Recent methods are mostly based on unsupervised cross-lingual word embeddings, the key to which is to find initial solutions of word translations, followed by the learning and refinement of mappings between the embedding spaces of two languages.However, previous methods find initial solutions just based on word-level information, which may be (1) limited and inaccurate, and (2) prone to contain some noise introduced by the insufficiently pre-trained embeddings of some words.To deal with those issues, in this paper, we propose a novel graph-based paradigm to induce bilingual lexicons in a coarse-to-fine way.We first build a graph for each language with its vertices representing different words.Then we extract word cliques from the graphs and map the cliques of two languages.Based on that, we induce the initial word translation solution with the central words of the aligned cliques.This coarse-to-fine approach not only leverages clique-level information, which is richer and more accurate, but also effectively reduces the bad effect of the noise in the pre-trained embeddings.Finally, we take the initial solution as the seed to learn cross-lingual embeddings, from which we induce bilingual lexicons.Experiments show that our approach improves the performance of bilingual lexicon induction compared with previous methods.
Shuo Ren 0002, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001
ACL2
2020 A Retrieve-and-Rewrite Initialization Method for Unsupervised Machine Translation
abstract
The commonly used framework for unsupervised machine translation builds initial translation models of both translation directions, and then performs iterative back-translation to jointly boost their translation performance.The initialization stage is very important since bad initialization may wrongly squeeze the search space, and too much noise introduced in this stage may hurt the final performance.In this paper, we propose a novel retrieval and rewriting based method to better initialize unsupervised translation models.We first retrieve semantically comparable sentences from monolingual corpora of two languages and then rewrite the target side to minimize the semantic gap between the source and retrieved targets with a designed rewriting model.The rewritten sentence pairs are used to initialize SMT models which are used to generate pseudo data for two NMT models, followed by the iterative back-translation.Experiments show that our method can build better initial unsupervised translation models and improve the final translation performance by over 4 BLEU scores.
Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001
ACL3
2020 Curriculum Pre-training for End-to-End Speech Translation
abstract
End-to-end speech translation poses a heavy burden on the encoder because it has to transcribe, understand, and learn cross-lingual semantics simultaneously.To obtain a powerful encoder, traditional methods pre-train it on ASR data to capture speech features.However, we argue that pre-training the encoder only through simple speech recognition is not enough, and high-level linguistic knowledge should be considered.Inspired by this, we propose a curriculum pre-training method that includes an elementary course for transcription learning and two advanced courses for understanding the utterance and mapping words in two languages.The difficulty of these courses is gradually increasing.Experiments show that our curriculum pre-training method leads to significant improvements on En-De and En-Fr speech translation benchmarks.
Chengyi Wang 0002, Yu Wu 0012, Shujie Liu 0001, Ming Zhou 0001, Zhenglu Yang
ACL3
2020 Semantic Mask for Transformer Based End-to-End Speech Recognition
abstract
Attention-based encoder-decoder model has achieved impressive results for both automatic speech recognition (ASR) and text-to-speech (TTS) tasks.This approach takes advantage of the memorization capacity of neural networks to learn the mapping from the input sequence to the output sequence from scratch, without the assumption of prior knowledge such as the alignments.However, this model is prone to overfitting, especially when the amount of training data is limited.Inspired by SpecAugment and BERT, in this paper, we propose a semantic mask based regularization for training such kind of end-toend (E2E) model.The idea is to mask the input features corresponding to a particular output token, e.g., a word or a wordpiece, in order to encourage the model to fill the token based on the contextual information.While this approach is applicable to the encoder-decoder framework with any type of neural network architecture, we study the transformer-based model for ASR in this work.We perform experiments on Librispeech 960h and TedLium2 data sets, and achieve the state-of-the-art performance on the test set in the scope of E2E models.
Chengyi Wang 0002, Yu Wu 0012, Yujiao Du, Jinyu Li 0001, Shujie Liu 0001, Liang Lu 0001, Shuo Ren 0002, Guoli Ye, Sheng Zhao 0002, Ming Zhou 0001
INTERSPEECH5
2020 Low Latency End-to-End Streaming Speech Recognition with a Scout Network
abstract
The attention-based Transformer model has achieved promising results for speech recognition (SR) in the offline mode.However, in the streaming mode, the Transformer model usually incurs significant latency to maintain its recognition accuracy when applying a fixed-length look-ahead window in each encoder layer.In this paper, we propose a novel low-latency streaming approach for Transformer models, which consists of a scout network and a recognition network.The scout network detects the whole word boundary without seeing any future frames, while the recognition network predicts the next subword by utilizing the information from all the frames before the predicted boundary.Our model achieves the best performance (2.7/6.4WER) with only 639 ms latency on the test-clean and test-other data sets of Librispeech.
Chengyi Wang 0002, Yu Wu 0012, Liang Lu 0001, Shujie Liu 0001, Jinyu Li 0001, Guoli Ye, Ming Zhou 0001
INTERSPEECH4
2020 On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
abstract
Recently, there has been a strong push to transition from hybrid models to end-to-end (E2E) models for automatic speech recognition.Currently, there are three promising E2E methods: recurrent neural network transducer (RNN-T), RNN attentionbased encoder-decoder (AED), and Transformer-AED.In this study, we conduct an empirical comparison of RNN-T, RNN-AED, and Transformer-AED models, in both non-streaming and streaming modes.We use 65 thousand hours of Microsoft anonymized training data to train these models.As E2E models are more data hungry, it is better to compare their effectiveness with large amount of training data.To the best of our knowledge, no such comprehensive study has been conducted yet.We show that although AED models are stronger than RNN-T in the non-streaming mode, RNN-T is very competitive in streaming mode if its encoder can be properly initialized.Among all three E2E models, transformer-AED achieved the best accuracy in both streaming and non-streaming mode.We show that both streaming RNN-T and transformer-AED models can obtain better accuracy than a highly-optimized hybrid model.
Jinyu Li 0001, Yu Wu 0012, Yashesh Gaur, Chengyi Wang 0002, Rui Zhao 0017, Shujie Liu 0001
INTERSPEECH6
2020 MoBoAligner: A Neural Alignment Model for Non-Autoregressive TTS with Monotonic Boundary Search
abstract
To speed up the inference of neural speech synthesis, nonautoregressive models receive increasing attention recently.In non-autoregressive models, additional durations of text tokens are required to make a hard alignment between the encoder and the decoder.The duration-based alignment plays a crucial role since it controls the correspondence between text tokens and spectrum frames and determines the rhythm and speed of synthesized audio.To get better duration-based alignment and improve the quality of non-autoregressive speech synthesis, in this paper, we propose a novel neural alignment model named MoBoAligner.Given the pairs of the text and mel spectrum, MoBoAligner tries to identify the boundaries of text tokens in the given mel spectrum frames based on the tokenframe similarity in the neural semantic space with an end-toend framework.With these boundaries, durations can be extracted and used in the training of non-autoregressive TTS models.Compared with the duration extracted by TransformerTTS, MoBoAligner brings improvement for the non-autoregressive TTS model on MOS (our 3.74 comparing to baseline's 3.44).Besides, MoBoAligner is task-specified and lightweight, which reduces the parameter number by 45% and the training time consuming by 30%.
Naihan Li, Shujie Liu 0001, Sheng Zhao 0002, Ming Zhou 0001
INTERSPEECH2
2020 A Hierarchical Clustering Approach to Fuzzy Semantic Representation of Rare Words in Neural Machine Translation
abstract
Rare words are usually replaced with a singletoken in the current encoder-decoder style of neural machine translation, challenging the translation modeling by an obscured context. In this article, we propose to build a fuzzy semantic representation (FSR) method for rare words through a hierarchical clustering method to group rare words together, and integrate it into the encoder-decoder framework. This hierarchical structure can compensate for the semantic information in both source and target sides, and providing fuzzy context information to capture the semantic of rare words. The introduced FSR can also alleviate the data sparseness, which is the bottleneck in dealing with rare words in neural machine translation. In particular, our method is easily extended to the transformer-based neural machine translation model and learns the FSRs of all in-vocabulary words to enhance the sentence representations in addition to rare words. Our experiments on Chinese-to-English translation tasks confirm a significant improvement in the translation quality brought by the proposed method.
Muyun Yang, Shujie Liu 0001, Kehai Chen, Enbo Zhao, Tiejun Zhao
IEEE Trans. Fuzzy Syst.2
2020 Joint Learning of Question Answering and Question Generation
abstract
Question answering (QA) and question generation (QG) are closely related tasks that could improve each other; however, the connection of these two tasks is not well explored in the literature. In this paper, we present two training algorithms for learning better QA and QG models through leveraging one another. The first algorithm extends Generative Adversarial Network (GAN), which selectively incorporates artificially generated instances as additional QA training data. The second algorithm is an extension of dual learning, which incorporates the probabilistic correlation of QA and QG as additional regularization in training objectives. To test the scalability of our algorithms, we conduct experiments on both document based and table based question answering tasks. Results show that both algorithms improve a QA model in terms of accuracy and QG model in terms of BLEU score. Moreover, we find that the performance of a QG model could be easily improved by a QA model via policy gradient, however, directly applying GAN that regards all the generated questions as negative instances could not improve the accuracy of the QA model. Our algorithm that selectively assigns labels to generated questions would bring a performance boost.
Duyu Tang, Nan Duan 0001, Tao Qin 0001, Shujie Liu 0001, Ming Zhou 0001, Yuanhua Lv, Wenpeng Yin 0001, Bing Qin 0001, Ting Liu 0001
IEEE Trans. Knowl. Data Eng.5
2019 Neural Speech Synthesis with Transformer Network
abstract
Although end-to-end neural text-to-speech (TTS) methods (such as Tacotron2) are proposed and achieve state-of-theart performance, they still suffer from two problems: 1) low efficiency during training and inference; 2) hard to model long dependency using current recurrent neural networks (RNNs). Inspired by the success of Transformer network in neural machine translation (NMT), in this paper, we introduce and adapt the multi-head attention mechanism to replace the RNN structures and also the original attention mechanism in Tacotron2. With the help of multi-head self-attention, the hidden states in the encoder and decoder are constructed in parallel, which improves training efficiency. Meanwhile, any two inputs at different times are connected directly by a self-attention mechanism, which solves the long range dependency problem effectively. Using phoneme sequences as input, our Transformer TTS network generates mel spectrograms, followed by a WaveNet vocoder to output the final audio results. Experiments are conducted to test the efficiency and performance of our new network. For the efficiency, our Transformer TTS network can speed up the training about 4.25 times faster compared with Tacotron2. For the performance, rigorous human tests show that our proposed model achieves state-of-the-art performance (outperforms Tacotron2 with a gap of 0.048) and is very close to human quality (4.39 vs 4.44 in MOS).
Naihan Li, Shujie Liu 0001, Sheng Zhao 0002
AAAI2
2019 Unsupervised Neural Machine Translation with SMT as Posterior Regularization
abstract
Without real bilingual corpus available, unsupervised Neural Machine Translation (NMT) typically requires pseudo parallel data generated with the back-translation method for the model training. However, due to weak supervision, the pseudo data inevitably contain noises and errors that will be accumulated and reinforced in the subsequent training process, leading to bad translation performance. To address this issue, we introduce phrase based Statistic Machine Translation (SMT) models which are robust to noisy data, as posterior regularizations to guide the training of unsupervised NMT models in the iterative back-translation process. Our method starts from SMT models built with pre-trained language models and word-level translation tables inferred from cross-lingual embeddings. Then SMT and NMT models are optimized jointly and boost each other incrementally in a unified EM framework. In this way, (1) the negative effect caused by errors in the iterative back-translation process can be alleviated timely by SMT filtering noises from its phrase tables; meanwhile, (2) NMT can compensate for the deficiency of fluency inherent in SMT. Experiments conducted on en-fr and en-de translation tasks show that our method outperforms the strong baseline and achieves new state-of-the-art unsupervised machine translation performance.
Shuo Ren 0002, Zhirui Zhang, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001
AAAI3
2019 Regularizing Neural Machine Translation by Target-Bidirectional Agreement
abstract
Although Neural Machine Translation (NMT) has achieved remarkable progress in the past several years, most NMT systems still suffer from a fundamental shortcoming as in other sequence generation tasks: errors made early in generation process are fed as inputs to the model and can be quickly amplified, harming subsequent sequence generation. To address this issue, we propose a novel model regularization method for NMT training, which aims to improve the agreement between translations generated by left-to-right (L2R) and right-to-left (R2L) NMT decoders. This goal is achieved by introducing two Kullback-Leibler divergence regularization terms into the NMT training objective to reduce the mismatch between output probabilities of L2R and R2L models. In addition, we also employ a joint training strategy to allow L2R and R2L models to improve each other in an interactive update process. Experimental results show that our proposed method significantly outperforms state-of-the-art baselines on Chinese-English and English-German translation tasks.
Zhirui Zhang, Shuangzhi Wu, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Tong Xu 0001
AAAI3
2019 Explicit Cross-lingual Pre-training for Unsupervised Machine Translation
abstract
Shuo Ren, Yu Wu, Shujie Liu, Ming Zhou, Shuai Ma. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Ming Zhou 0001, Shuai Ma 0001
EMNLP/IJCNLP (1)3
2019 Unsupervised Context Rewriting for Open Domain Conversation
abstract
Kun Zhou, Kai Zhang, Yu Wu, Shujie Liu, Jingsong Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yu Wu 0012, Shujie Liu 0001, Jingsong Yu
EMNLP/IJCNLP (1)4
2018 Assertion-Based QA With Question-Aware Open Information Extraction
abstract
We present assertion based question answering (ABQA), an open domain question answering task that takes a question and a passage as inputs, and outputs a semi-structured assertion consisting of a subject, a predicate and a list of arguments. An assertion conveys more evidences than a short answer span in reading comprehension, and it is more concise than a tedious passage in passage-based QA. These advantages make ABQA more suitable for human-computer interaction scenarios such as voice-controlled speakers. Further progress towards improving ABQA requires richer supervised dataset and powerful models of text understanding. To remedy this, we introduce a new dataset called WebAssertions, which includes hand-annotated QA labels for 358,427 assertions in 55,960 web passages. To address ABQA, we develop both generative and extractive approaches. The backbone of our generative approach is sequence to sequence learning. In order to capture the structure of the output assertion, we introduce a hierarchical decoder that first generates the structure of the assertion and then generates the words of each field. The extractive approach is based on learning to rank. Features at different levels of granularity are designed to measure the semantic relevance between a question and an assertion. Experimental results show that our approaches have the ability to infer question-aware assertions from a passage. We further evaluate our approaches by incorporating the ABQA results as additional features in passage-based QA. Results on two datasets show that ABQA features significantly improve the accuracy on passage-based QA.
Duyu Tang, Nan Duan 0001, Shujie Liu 0001, Daxin Jiang, Ming Zhou 0001, Zhoujun Li 0001
AAAI4
2018 Joint Training for Neural Machine Translation Models with Monolingual Data
abstract
Monolingual data have been demonstrated to be helpful in improving translation quality of both statistical machine translation (SMT) systems and neural machine translation (NMT) systems, especially in resource-poor or domain adaptation tasks where parallel data are not rich enough. In this paper, we propose a novel approach to better leveraging monolingual data for neural machine translation by jointly learning source-to-target and target-to-source NMT models for a language pair with a joint EM optimization method. The training process starts with two initial NMT models pre-trained on parallel data for each direction, and these two models are iteratively updated by incrementally decreasing translation losses on training data.In each iteration step, both NMT models are first used to translate monolingual data from one language to the other, forming pseudo-training data of the other NMT model. Then two new NMT models are learnt from parallel data together with the pseudo training data. Both NMT models are expected to be improved and better pseudo-training data can be generated in next step. Experiment results on Chinese-English and English-German translation tasks show that our approach can simultaneously improve translation quality of source-to-target and target-to-source models, significantly outperforming strong baseline systems which are enhanced with monolingual data for model training including back-translation.
Zhirui Zhang, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Enhong Chen
AAAI2
2018 Triangular Architecture for Rare Language Translation
abstract
Neural Machine Translation (NMT) performs poor on the low-resource language pair (X, Z), especially when Z is a rare language.By introducing another rich language Y , we propose a novel triangular training architecture (TA-NMT) to leverage bilingual data (Y, Z) (may be small) and (X, Y ) (can be rich) to improve the translation performance of lowresource pairs.In this triangular architecture, Z is taken as the intermediate latent variable, and translation models of Z are jointly optimized with a unified bidirectional EM algorithm under the goal of maximizing the translation likelihood of (X, Y ).Empirical results demonstrate that our method significantly improves the translation quality of rare languages on MultiUN and IWSLT2012 datasets, and achieves even better performance combining back-translation methods.
Shuo Ren 0002, Wenhu Chen, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Shuai Ma 0001
ACL (1)3
2018 Bidirectional Generative Adversarial Networks for Neural Machine Translation
abstract
Generative Adversarial Network (GAN) has been proposed to tackle the exposure bias problem of Neural Machine Translation (NMT).However, the discriminator typically results in the instability of the GAN training due to the inadequate training problem: the search space is so huge that sampled translations are not sufficient for discriminator training.To address this issue and stabilize the GAN training, in this paper, we propose a novel Bidirectional Generative Adversarial Network for Neural Machine Translation (BGAN-NMT), which aims to introduce a generator model to act as the discriminator, whereby the discriminator naturally considers the entire translation space so that the inadequate training problem can be alleviated.To satisfy this property, generator and discriminator are both designed to model the joint probability of sentence pairs, with the difference that, the generator decomposes the joint probability with a source language model and a source-to-target translation model, while the discriminator is formulated as a target language model and a target-to-source translation model.To further leverage the symmetry of them, an auxiliary GAN is introduced and adopts generator and discriminator models of original one as its own discriminator and generator respectively.Two GANs are alternately trained to update the parameters.Experiment results on German-English and Chinese-English translation tasks demonstrate that our method not only stabilizes GAN training but also achieves significant improvements over baseline systems.
Zhirui Zhang, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Enhong Chen
CoNLL2
2018 Generative Bridging Network for Neural Sequence Prediction
abstract
Wenhu Chen, Guanlin Li, Shuo Ren, Shujie Liu, Zhirui Zhang, Mu Li, Ming Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Wenhu Chen, Shuo Ren 0002, Shujie Liu 0001, Zhirui Zhang, Mu Li 0001, Ming Zhou 0001
NAACL-HLT4
2018 Learning to Collaborate for Question Answering and Asking
abstract
Duyu Tang, Nan Duan, Zhao Yan, Zhirui Zhang, Yibo Sun, Shujie Liu, Yuanhua Lv, Ming Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Duyu Tang, Nan Duan 0001, Zhirui Zhang, Shujie Liu 0001, Yuanhua Lv, Ming Zhou 0001
NAACL-HLT6
2018 Coarse-To-Fine Learning for Neural Machine Translation
Zhirui Zhang, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Enhong Chen
NLPCC (1)2
2017 Chunk-based Decoder for Neural Machine Translation
abstract
Shonosuke Ishiwatari, Jingtao Yao, Shujie Liu, Mu Li, Ming Zhou, Naoki Yoshinaga, Masaru Kitsuregawa, Weijia Jia. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Shonosuke Ishiwatari, JingTao Yao 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Naoki Yoshinaga 0001, Masaru Kitsuregawa, Weijia Jia 0001
ACL (1)3
2017 Stack-based Multi-layer Attention for Transition-based Dependency Parsing
abstract
Although sequence-to-sequence (seq2seq) network has achieved significant success in many NLP tasks such as machine translation and text summarization, simply applying this approach to transition-based dependency parsing cannot yield a comparable performance gain as in other stateof-the-art methods, such as stack-LSTM and head selection.In this paper, we propose a stack-based multi-layer attention model for seq2seq learning to better leverage structural linguistics information.In our method, two binary vectors are used to track the decoding stack in transition-based parsing, and multi-layer attention is introduced to capture multiple word dependencies in partial trees.We conduct experiments on PTB and CTB datasets, and the results show that our proposed model achieves state-of-the-art accuracy and significant improvement in labeled precision with respect to the baseline seq2seq model.
Zhirui Zhang, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Enhong Chen
EMNLP2
2017 Modeling Indicative Context for Statistical Machine Translation
Shuangzhi Wu, Dongdong Zhang 0001, Shujie Liu 0001, Ming Zhou 0001
NLPCC3
2016 Knowledge-Based Semantic Embedding for Machine Translation
abstract
In this paper, with the help of knowledge base, we build and formulate a semantic space to connect the source and target languages, and apply it to the sequence-to-sequence framework to propose a Knowledge-Based Semantic Embedding (KBSE) method.In our KB-SE method, the source sentence is firstly mapped into a knowledge based semantic space, and the target sentence is generated using a recurrent neural network with the internal meaning preserved.Experiments are conducted on two translation tasks, the electric business data and movie data, and the results show that our proposed method can achieve outstanding performance, compared with both the traditional SMT methods and the existing encoder-decoder models.
Shujie Liu 0001, Shuo Ren 0002, Mu Li 0001, Ming Zhou 0001, Xu Sun 0001, Houfeng Wang
ACL (1)2
2016 Improving Attention Modeling with Implicit Distortion and Fertility for Machine Translation
abstract
In neural machine translation, the attention mechanism facilitates the translation process by producing a soft alignment between the source sentence and the target sentence. However, without dedicated distortion and fertility models seen in traditional SMT systems, the learned alignment may not be accurate, which can lead to low translation quality. In this paper, we propose two novel models to improve attention-based neural machine translation. We propose a recurrent attention mechanism as an implicit distortion model, and a fertility conditioned decoder as an implicit fertility model. We conduct experiments on large-scale Chinese–English translation tasks. The results show that our models significantly improve both the alignment and translation quality compared to the original attention mechanism and several other variations.
Shujie Liu 0001, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001, Kenny Q. Zhu
COLING2
2015 Hierarchical Recurrent Neural Network for Document Modeling
abstract
This paper proposes a novel hierarchical recurrent neural network language model (HRNNLM) for document modeling.After establishing a RNN to capture the coherence between sentences in a document, HRNNLM integrates it as the sentence history information into the word level RNN to predict the word sequence with cross-sentence contextual information.A two-step training approach is designed, in which sentence-level and word-level language models are approximated for the convergence in a pipeline style.Examined by the standard sentence reordering scenario, HRNNLM is proved for its better accuracy in modeling the sentence coherence.And at the word level, experimental results also indicate a significant lower model perplexity, followed by a practical better translation result when applied to a Chinese-English document translation reranking task.
Shujie Liu 0001, Muyun Yang, Mu Li 0001, Ming Zhou 0001, Sheng Li 0003
EMNLP2
2015 Entity Translation with Collective Inference in Knowledge Graph
Qinglin Li, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001
NLPCC2
2015 A Maximum Entropy Approach to Discourse Coherence Modeling
abstract
This paper introduces a maximum entropy method to Discourse Coherence Modeling (DCM). Different from the state-of-art supervised entity-grid model and unsupervised cohesion-driven model, the model we proposed only takes as input lexicon features, which increases the training speed and decoding speed significantly. We conduct an evaluation on two publicly available benchmark data sets via sentence ordering tasks, and the results confirm the effectiveness of our maximum entropy based approach in DCM.
Muyun Yang, Shujie Liu 0001, Sheng Li 0003, Tiejun Zhao
NLPCC3
2015 A Statistical Parsing Framework for Sentiment Classification
abstract
We present a statistical parsing framework for sentence-level sentiment classification in this article. Unlike previous works that use syntactic parsing results for sentiment analysis, we develop a statistical parser to directly analyze the sentiment structure of a sentence. We show that complicated phenomena in sentiment analysis (e.g., negation, intensification, and contrast) can be handled the same way as simple and straightforward sentiment expressions in a unified and probabilistic way. We formulate the sentiment grammar upon Context-Free Grammars (CFGs), and provide a formal description of the sentiment parsing framework. We develop the parsing model to obtain possible sentiment parse trees for a sentence, from which the polarity model is proposed to derive the sentiment strength and polarity, and the ranking model is dedicated to selecting the best sentiment tree. We train the parser directly from examples of sentences annotated only with sentiment polarity labels but without any syntactic annotations or polarity annotations of constituents within sentences. Therefore we can obtain training data easily. In particular, we train a sentiment parser, s.parser, from a large amount of review sentences with users' ratings as rough sentiment polarity labels. Extensive experiments on existing benchmark data sets show significant improvements over baseline sentiment classification approaches.
Li Dong 0004, Furu Wei, Shujie Liu 0001, Ming Zhou 0001, Ke Xu 0001
Comput. Linguistics3
2015 Towards Machine Translation in Semantic Vector Space
abstract
Measuring the quality of the translation rules and their composition is an essential issue in the conventional statistical machine translation (SMT) framework. To express the translation quality, the previous lexical and phrasal probabilities are calculated only according to the co-occurrence statistics in the bilingual corpus and may be not reliable due to the data sparseness problem. To address this issue, we propose measuring the quality of the translation rules and their composition in the semantic vector embedding space (VES). We present a recursive neural network (RNN)-based translation framework, which includes two submodels. One is the bilingually-constrained recursive auto-encoder, which is proposed to convert the lexical translation rules into compact real-valued vectors in the semantic VES. The other is a type-dependent recursive neural network, which is proposed to perform the decoding process by minimizing the semantic gap (meaning distance) between the source language string and its translation candidates at each state in a bottom-up structure. The RNN-based translation model is trained using a max-margin objective function that maximizes the margin between the reference translation and the n-best translations in forced decoding. In the experiments, we first show that the proposed vector representations for the translation rules are very reliable for application in translation modeling. We further show that the proposed type-dependent, RNN-based model can significantly improve the translation quality in the large-scale, end-to-end Chinese-to-English translation evaluation.
Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2014 Mind the Gap: Machine Translation by Minimizing the Semantic Gap in Embedding Space
abstract
The conventional statistical machine translation (SMT) methods perform the decoding process by compositing a set of the translation rules which are associated with high probabilities. However, the probabilities of the translation rules are calculated only according to the cooccurrence statistics in the bilingual corpus rather than the semantic meaning similarity. In this paper, we propose a Recursive Neural Network (RNN) based model that converts each translation rule into a compact real-valued vector in the semantic embedding space and performs the decoding process by minimizing the semantic gap between the source language string and its translation candidates at each state in a bottom-up structure. The RNN-based translation model is trained using a max-margin objective function. Extensive experiments on Chinese-to-English translation show that our RNN-based model can significantly improve the translation quality by up to 1.68 BLEU score.
Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong
AAAI2
2014 Learning Topic Representation for SMT with Neural Networks
abstract
Statistical Machine Translation (SMT) usually utilizes contextual information to disambiguate translation candidates. However, it is often limited to contexts within sentence boundaries, hence broader topical information cannot be leveraged. In this paper, we propose a novel approach to learning topic representation for paral-lel data using a neural network architec-ture, where abundant topical contexts are embedded via topic relevant monolingual data. By associating each translation rule with the topic representation, topic rele-vant rules are selected according to the dis-tributional similarity with the source text during SMT decoding. Experimental re-sults show that our method significantly improves translation accuracy in the NIST Chinese-to-English translation task com-pared to a state-of-the-art baseline. 1
Lei Cui 0001, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Muyun Yang
ACL (1)3
2014 A Recursive Recurrent Neural Network for Statistical Machine Translation
abstract
In this paper, we propose a novel recursive recurrent neural network (R 2 NN) to model the end-to-end decoding process for statistical machine translation.R 2 NN is a combination of recursive neural network and recurrent neural network, and in turn integrates their respective capabilities: (1) new information can be used to generate the next hidden state, like recurrent neural networks, so that language model and translation model can be integrated naturally; (2) a tree structure can be built, as recursive neural networks, so as to generate the translation candidates in a bottom up manner.A semi-supervised training approach is proposed to train the parameters, and the phrase pair embedding is explored to model translation confidence directly.Experiments on a Chinese to English translation task show that our proposed R 2 NN can outperform the stateof-the-art baseline by about 1.5 points in BLEU.
Shujie Liu 0001, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001
ACL (1)1
2014 Bilingually-constrained Phrase Embeddings for Machine Translation
abstract
We propose Bilingually-constrained Recursive Auto-encoders (BRAE) to learn semantic phrase embeddings (compact vector representations for phrases), which can distinguish the phrases with different semantic meanings.The BRAE is trained in a way that minimizes the semantic distance of translation equivalents and maximizes the semantic distance of nontranslation pairs simultaneously.After training, the model learns how to embed each phrase semantically in two languages and also learns how to transform semantic embedding space in one language to the other.We evaluate our proposed method on two end-to-end SMT tasks (phrase table pruning and decoding with phrasal semantic similarities) which need to measure semantic similarity between a source phrase and its translation candidates.Extensive experiments show that the BRAE is remarkably effective in these two tasks.
Jiajun Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Chengqing Zong
ACL (1)2
2013 Word Alignment Modeling with Context Dependent Deep Neural Network
Nan Yang 0002, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Nenghai Yu
ACL (1)2
2013 Multi-Domain Adaptation for SMT Using Multi-Task Learning
abstract
Domain adaptation for SMT usually adapts models to an individual specific domain.However, it often lacks some correlation among different domains where common knowledge could be shared to improve the overall translation quality.In this paper, we propose a novel multi-domain adaptation approach for SMT using Multi-Task Learning (MTL), with in-domain models tailored for each specific domain and a general-domain model shared by different domains.The parameters of these models are tuned jointly via MTL so that they can learn general knowledge more accurately and exploit domain knowledge better.Our experiments on a largescale English-to-Chinese translation task validate that the MTL-based adaptation approach significantly and consistently improves the translation quality compared to a non-adapted baseline.Furthermore, it also outperforms the individual adaptation of each specific domain.
Lei Cui 0001, Xilun Chen 0002, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001
EMNLP4
2013 Efficient Collective Entity Linking with Stacking
abstract
Entity disambiguation works by linking ambiguous mentions in text to their corresponding real-world entities in knowledge base.Recent collective disambiguation methods enforce coherence among contextual decisions at the cost of non-trivial inference processes.We propose a fast collective disambiguation approach based on stacking.First, we train a local predictor g 0 with learning to rank as base learner, to generate initial ranking list of candidates.Second, top k candidates of related instances are searched for constructing expressive global coherence features.A global predictor g 1 is trained in the augmented feature space and stacking is employed to tackle the train/test mismatch problem.The proposed method is fast and easy to implement.Experiments show its effectiveness over various algorithms on several public datasets.By learning a rich semantic relatedness measure between entity categories and context document, performance is further improved.
Zhengyan He, Shujie Liu 0001, Yang Song 0021, Mu Li 0001, Ming Zhou 0001, Houfeng Wang
EMNLP2
2013 Collective Corpus Weighting and Phrase Scoring for SMT Using Graph-Based Random Walk
Lei Cui 0001, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001
NLPCC3
2012 Learning Translation Consensus with Structured Label Propagation
Shujie Liu 0001, Chi-Ho Li, Mu Li 0001, Ming Zhou 0001
ACL (1)1
2012 Re-training Monolingual Parser Bilingually for Syntactic SMT
Shujie Liu 0001, Chi-Ho Li, Mu Li 0001, Ming Zhou 0001
EMNLP-CoNLL1
2012 Virtual View Reconstruction Using Temporal Information
abstract
The most significant problem in generating virtual views from a limited number of video camera views is handling areas that have become dis-occluded by shifting the virtual view away from the camera view. We propose using temporal information to address this problem, based on the notion that dis-occluded areas may have been seen by some camera in some previous frames. We formulate the problem as one of estimating the underlying state of the object in a stochastic dynamical system, given a sequence of observations. We apply the formulation to improving the visual quality of virtual views generated from a single “color plus depth” camera, and show that our algorithm achieves better results than depth image based rendering using standard inpainting.
Shujie Liu 0001, Philip A. Chou, Cha Zhang, Zhengyou Zhang, Chang Wen Chen
ICME1
2012 A novel 3D video transcoding scheme for adaptive 3D video transmission to heterogeneous terminals
abstract
Three-dimensional video (3DV) is attracting many interests with its enhanced viewing experience and more user driven features. 3DV has several unique characteristics different from 2D video: (1) It has a much larger amount of data captured and compressed, and corresponding video compression techniques can be much more complicated in order to explore data redundancy. This will lead to more constraints on users' network access and computational capability, (2) Most users only need part of the 3DV data at any given time, while the users' requirements exhibit large diversity, (3) Only a limited number of views are captured and transmitted for 3DV. View rendering is thus necessary to generate virtual views based on the received 3DV data. However, many terminal devices do not have the functionality to generate virtual views. To enable 3DV experience for the majority of users with limited capabilities, adaptive 3DV transmission is necessary to extract/generate the required data content and represent it with supported formats and bitrates for heterogeneous terminal devices. 3DV transcoding is an emerging and effective technique to achieve desired adaptive 3DV transmission. In this article, we propose the first efficient 3DV transcoding scheme that can obtain any desired view, either an encoded one or a virtual one, and compress it with more universal H.264/AVC. The key idea of the proposed scheme is to appropriately utilize motion information contained in the bitstream to generate candidate motion information. Original information of both the desired view and reference views are used to obtain this candidate information and a proper motion refinement process is carried out for certain blocks. Simulation results show that, compared to the straightforward cascade algorithm, the proposed scheme is able to output compressed bitstream of the required view with significantly reduced complexity while incurring negligible performance loss. Such a 3DV transcoding can be applied to most gateways that usually have constraints on computational complexity and time delay.
Shujie Liu 0001, Chang Wen Chen
ACM Trans. Multim. Comput. Commun. Appl.1
2011 Transductive Minimum Error Rate Training for Statistical Machine Translation
Yinggong Zhao, Shujie Liu 0001, Yangsheng Ji, Jiajun Chen 0001, Guodong Zhou 0001
IJCNLP2
2011 Statistic Machine Translation Boosted with Spurious Word Deletion
Shujie Liu 0001, Chi-Ho Li, Ming Zhou 0001
MTSummit1
2011 A Unified SMT Framework Combining MIRA and MERT
Shujie Liu 0001, Chi-Ho Li, Ming Zhou 0001
MTSummit1
2011 Scalable video transmission: packet loss induced distortion modeling and estimation
abstract
To provide enhanced multimedia services for heterogeneous networks and terminal devices, Scalable Video Coding (SVC) has been developed to embed different quality of video in a single bitstream. Similar to classical compressed video transmission, different packets of a video bitstream have different impacts on received video quality. Therefore, distortion modeling and estimation are necessary in designing a robust video transmission strategy under various network conditions. In the paper, we present the first scheme of packet loss induced distortion modeling and estimation in SVC transmission. The proposed scheme is applicable to numerous video communication and networking scenarios in which accurate distortion information can be utilized to enhance the performance of video transmission. One major challenge in scalable video distortion estimation is due to the adoption of more complicated prediction structure in SVC, which makes the tracking of error propagation much more difficult than the non-scalable encoded video. In this research, we tackle such challenge by systematically tracking the propagation of errors under various prediction trajectories. Supplemental information about the compressed video is embedded into data packets to substantially simplify the modeling and estimation. Moreover, with supplemental data of inter prediction information, distortion estimation can be processed without parsing video bitstream which results in much lower computation and memory cost. With negligible effects on the data size, experimental results show that the proposed scheme is able to track and estimate the distortion with very high accuracy. This first ever scalable video transmission distortion modeling and estimation scheme can be deployed at either gateways or receivers because of its low computation and memory cost.
Shujie Liu 0001, Chang Wen Chen
NOSSDAV1
2010 Discriminative Pruning for Discriminative ITG Alignment
Shujie Liu 0001, Chi-Ho Li, Ming Zhou 0001
ACL1
2010 Sparse dyadic mode for depth map compression
abstract
In order to enable new video applications such as 3DTV and free-viewpoint video, new data formats including both 2D video sequences and corresponding depth map sequences have been proposed. One major characteristic making the depth maps different from video frames is that they typically consist of homogeneous areas separated by sharp edges representing depth discontinuities. Another characteristic of depth map sequences is that the edges exhibit quite similar boundary behaviors as the edges in the corresponding video frames. In this paper, we propose a novel sparse dyadic mode in the design of an efficient depth map compression algorithm through appropriately exploiting these characteristics. With sparse representations of depth blocks and effective reference of edge information from the corresponding video frames, sparse dyadic mode can achieve up to 1.5 dB gain on rendering quality as compared to depth sequences coded using MVC at the same bitrate.
Shujie Liu 0001, PoLin Lai, Dong Tian, Cristina Gomila, Chang Wen Chen
ICIP1
2010 A novel prioritized spatial multiplexing for MIMO wireless system with application to H.264 SVC video
abstract
Spatial multiplexing MIMO wireless system is an excellent choice for next generation broadband multimedia communications because of its parallel transmission capability. A major challenge for such MIMO system is that it usually requires appropriate decomposition of both wireless channels and multimedia data bitstreams. Existing approaches to MIMO channel decomposition have been passively based on simple channel feedback. To facilitate proper match between decomposed channels with compressed video that exhibit variable priority for different layer of bitstreams, it is necessary to proactively prioritize the decomposed MIMO channels. We present in this paper a novel pre-coding scheme capable of integrating both channel and source characteristics in order to achieve the desired prioritized spatial multiplexing. This pre-coding scheme is applied to the transmission of video data compressed with H.264 Scalable Video Coding (SVC) standard. Based on the desired bit-error-rate (BER) and signal-to-noise-ratio (SNR) for each SVC video layer, the base layer video is proactively assigned the highest priority sub-channel to guarantee minimum required quality while enhancement layers are mapped to other lower priority sub-channels. Comparing with existing schemes, the proposed pre-coder has low computational complexity and is suitable for rapid hardware implementation. The experimental results demonstrate that the proposed prioritized spatial multiplexing approach can efficiently enhance the quality of SVC video streaming over MIMO systems.
Qian Liu 0001, Shujie Liu 0001, Chang Wen Chen
ICME2
2010 3D video transcoding for virtual views
abstract
Recent emerging development of three dimensional video (3DV) has been vigorously driving the Multiview Video Coding (MVC) standard developed by Joint Video Team as an amendment to H.264/AVC and the new 3DV standard developed by MPEG. It is expected that 3DV contents will soon be available to various media consumers. Because of the heterogeneous networks and terminal devices, a variety of end users shall not have (1) the 3DV decoder and view synthesize software installed on their devices, and (2) adequate network bandwidth for 3DV content delivery. 3DV transcoder is therefore necessary for users without 3DV decoder to enjoy 3DV services, as well as for service providers to better control virtual view quality of all 3DV customers. In this paper, we report the first 3DV transcoding scheme for virtual view, which is able to generate bitstream of one single view that can be decoded by H.264/AVC. The key idea of the proposed transcoding is to appropriately generate candidate motion and modes based on motion information from other views in the original bitstream making best use of inter-view correlation. These candidate motion and modes are then properly used to encode a virtual view with significantly reduced complexity for decoding by H.264/AVC. Simulation results show that, compared to the straightforward cascade algorithm, the proposed transcoding experiences minor performance loss with a significantly reduced computation cost. Such low complexity transcoding can be deployed at various media gateways to deliver 3DV content to non-3DV devices.
Shujie Liu 0001, Chang Wen Chen
ACM Multimedia1
2010 Joint trilateral filtering for depth map compression
abstract
New data formats including 2D video and the corresponding depth maps enable new video applications in which virtual views can be rendered, such as 3DTV and free-viewpoint video (FVV). Different from video frames, depth maps typically consist of homogeneous areas (with no textures) separated by sharp edges representing depth value changes such as between foreground and background. Conventional video coding techniques with transforms followed by quantization typically result in large artifacts along such sharp edges. To suppress these coding artifacts while preserving edges, we propose in this paper a novel filtering method for depth coding, joint trilateral filter. The main contribution in the proposed filter design is the utilization of edge information in the collocated video frame as well as in the depth map. The filtering weights are determined by the following three factors: a domain (spatial) filter which measures the proximity of pixel positions, and two range filters. One range filter takes into account the similarity among depth samples and the other one considers the similarity among the collocated pixels in the video frame. By replacing the deblocking filter in H.264/AVC with the proposed trilateral filter, simulation results demonstrate up to 0.8 dB gain in rendering quality at given bitrate for depth signal.
Shujie Liu 0001, PoLin Lai, Dong Tian, Cristina Gomila, Chang Wen Chen
VCIP1
2009 Multiview Video transcoding: From multiple views to single view
abstract
As multiview video is gaining more and more attentions, Multiview Video Coding (MVC) standard has been under development by the Joint Video Team as an extension to H.264/AVC. There will be increasingly more multiview video sources for both high end and low end consumers. Since a variety of end user devices shall not have multiview video decoders installed, multiview video transcoder is therefore needed in order for these users to enjoy the videos encoded by MVC standard. In this paper, we propose an algorithm for transcoding any viewpoint in bitstreams encoded by MVC to bitstreams that can be decoded by H.264/AVC decoders, regardless of the MVC prediction structure. The key component of this algorithm is the design of motion information reuse to implement simple motion refinement for inter-view coded macroblocks. Such reduced complexity implementation is especially applicable to the low power mobile devices. Experimental results show that, comparing with cascaded algorithm, the proposed scheme, with much reduced complexity, only suffers moderate performance loss.
Shujie Liu 0001, Chang Wen Chen
PCS1
2008 Diagnostic Evaluation of Machine Translation Systems Using Automatically Constructed Linguistic Check-Points
Ming Zhou 0001, Shujie Liu 0001, Mu Li 0001, Dongdong Zhang 0001, Tiejun Zhao
COLING3
2008 Low-complexity asymmetric multiview video coding
abstract
Multiview video coding (MVC) is currently under development by the Joint Video Team (JVT) as an extension to Advanced Video Coding (H264/AVC). Based on the suppression theory in binocular vision, the fidelity of one of the two views of a stereoscopic display can be reduced without noticeable degradation of subjective quality. Thus, in MVC, a subset of views can be coded with lower spatial resolution at negligible cost to subjective quality. Due to different resolutions, a downsampling process is required in an MVC decoder in order to enable motion compensation (MC) between views. In this paper, a low-complexity MC algorithm is proposed for MVC to enable inter-view prediction between pictures with different resolutions. It requires lower memory consumption and lower computational complexity compared with the conventional downsampled inter-view prediction, while providing comparable efficiency, as shown by the simulation results.
Ying Chen 0011, Shujie Liu 0001, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj
ICME2
2008 Frame loss error concealment for multiview video coding
abstract
The Multiview Video Coding (MVC) standard is currently under development by the Joint Video Team as an extension of the Advanced Video Coding (H.264/AVC) standard. An MVC encoder compresses more than one viewpoint of a scene captured by different cameras. Redundancies between views can be used for inter-view prediction in encoding as well as error concealment in decoding. In this paper, a new algorithm utilizing motion information of pictures from other views to conceal a lost picture is proposed. The algorithm first derives motion information for a lost picture based on motion fields of pictures in adjacent views. Then, traditional motion compensation is invoked within the view containing the lost picture to derive a concealed frame. Experimental results show that the proposed algorithm can improve video quality with a negligible computational complexity overhead compared to simple temporal error concealment algorithms.
Shujie Liu 0001, Ying Chen 0011, Ye-Kui Wang, Moncef Gabbouj, Miska M. Hannuksela, Houqiang Li
ISCAS1