VLDB 2026 Research / reviewers in the wild / expert
Yang Feng 0004
dblp:07/6095-4
· DBLP profile ↗
85ranked-venue papers
8as first author
57since 2021 · last 2026
0000-0003-2579-1366ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 82 · 8 first-author · 55 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible MultilingualityabstractLarge language models (LLMs) exhibit strong general intelligence, yet their multilingual performance remains highly imbalanced.Although LLMs encode substantial cross-lingual knowledge in a unified semantic space, they often struggle to reliably interface this knowledge with low-resource or unseen languages.Fortunately, pretrained encoder-decoder translation models already possess balanced multilingual capability, suggesting a natural complement to LLMs.In this work, we propose XBridge, a compositional encoder-LLM-decoder architecture that offloads multilingual understanding and generation to external pretrained translation models, while preserving the LLM as an English-centric core for general knowledge processing.To address the resulting representation misalignment across models, we introduce lightweight cross-model mapping layers and an optimal transport-based alignment objective, enabling fine-grained semantic consistency for multilingual generation.Experiments on four LLMs across multilingual understanding, reasoning, summarization, and generation indicate that XBridge outperforms strong baselines, especially on low-resource and previously unseen languages, without retraining the LLM. Mengyu Bu, Yang Feng 0004 |
ACL (1) | 2 |
| 2025 | Large Language Models Are Read/Write Policy-Makers for Simultaneous GenerationabstractSimultaneous generation models write generation results while reading streaming inputs, necessitating a policy-maker to determine the appropriate output timing. Existing simultaneous generation methods generally adopt the traditional encoder-decoder architecture and learn the generation and policy-making capabilities through complex dynamic programming techniques. Although LLMs excel at text generation, they face challenges in taking on the role of policy-makers through traditional training methods, limiting their exploration in simultaneous generation. To overcome these limitations, we propose a novel LLM-driven Simultaneous Generation (LSG) framework, which allows the off-the-shelf LLM to decide the generation timing and produce output concurrently. Specifically, LSG selects the generation policy that minimizes latency as the baseline policy. Referring to the baseline policy, LSG enables the LLM to devise an improved generation policy that better balances latency and generation quality, and writes generation results accordingly. Experiments on simultaneous translation and streaming automatic speech recognition tasks show that our method can achieve state-of-the-art performance utilizing the open-source LLMs and demonstrate practicality in real-world scenarios. Shoutao Guo, Shaolei Zhang 0001, Zhengrui Ma, Yang Feng 0004 |
AAAI | 4 |
| 2025 | LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech SynthesisabstractReal-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction.Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language models (LLMs).In this paper, we introduce LLaMA-Omni 2, a series of speech language models (SpeechLMs) ranging from 0.5B to 14B parameters, capable of achieving highquality real-time speech interaction.LLaMA-Omni 2 is built upon the Qwen2.5 series models, integrating a speech encoder and an autoregressive streaming speech decoder.Despite being trained on only 200K multi-turn speech dialogue samples, LLaMA-Omni 2 demonstrates strong performance on several spoken question answering and speech instruction following benchmarks, surpassing previous state-of-theart SpeechLMs like GLM-4-Voice, which was trained on millions of hours of speech data. 1 Qingkai Fang, Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 5 |
| 2025 | AlignX: Advancing Multilingual Large Language Models with Multilingual Representation AlignmentabstractMultilingual large language models (LLMs) possess impressive multilingual understanding and generation capabilities.However, their performance and cross-lingual alignment often lag for non-dominant languages.A common solution is to fine-tune LLMs on largescale and more balanced multilingual corpora, but such approaches often lead to imprecise alignment and suboptimal knowledge transfer, struggling with limited improvements across languages.In this paper, we propose AlignX to bridge the multilingual performance gap, which is a two-stage representationlevel framework for enhancing multilingual performance of pre-trained LLMs.In the first stage, we align multilingual representations with multilingual semantic alignment and language feature integration.In the second stage, we stimulate the multilingual capability of LLMs via multilingual instruction fine-tuning.Experimental results on several pre-trained LLMs demonstrate that our approach enhances LLMs' multilingual general and cross-lingual generation capability.Further analysis indicates that AlignX brings the multilingual representations closer and improves the cross-lingual alignment.1 Mengyu Bu, Shaolei Zhang 0001, Zhongjun He, Hua Wu 0003, Yang Feng 0004 |
EMNLP | 5 |
| 2025 | IG-Pruning: Input-Guided Block Pruning for Large Language ModelsabstractWith the growing computational demands of large language models (LLMs), efficient inference has become increasingly critical for practical deployment. Depth pruning has emerged as a promising approach for reducing the computational costs of large language models by removing transformer layers. However, existing methods typically rely on fixed block masks, which can lead to suboptimal performance across different tasks and inputs. In this paper, we propose IG-Pruning, a novel input-aware block-wise pruning method that dynamically selects layer masks at inference time. Our approach consists of two stages: (1) Discovering diverse mask candidates through semantic clustering and L0 optimization, and (2) Implementing efficient dynamic pruning without the need for extensive training. Experimental results demonstrate that our method consistently outperforms state-of-the-art static depth pruning methods, making it particularly suitable for resource-constrained deployment scenarios. Kangyu Qiao, Shaolei Zhang 0001, Yang Feng 0004 |
EMNLP | 3 |
| 2025 | LLaMA-Omni: Seamless Speech Interaction with Large Language ModelsabstractModels like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how to build speech interaction models based on open-source LLMs. To address this, we propose LLaMA-Omni, a novel model architecture designed for low-latency and high-quality speech interaction with LLMs. LLaMA-Omni integrates a pretrained speech encoder, a speech adaptor, an LLM, and a streaming speech decoder. It eliminates the need for speech transcription, and can simultaneously generate text and speech responses directly from speech instructions with extremely low latency. We build our model based on the latest Llama-3.1-8B-Instruct model. To align the model with speech interaction scenarios, we construct a dataset named InstructS2S-200K, which includes 200K speech instructions and corresponding speech responses. Experimental results show that compared to previous speech-language models, LLaMA-Omni provides better responses in both content and style, with a response latency as low as 226ms. Additionally, training LLaMA-Omni takes less than 3 days on just 4 GPUs, paving the way for the efficient development of speech-language models in the future. Qingkai Fang, Shoutao Guo, Zhengrui Ma, Shaolei Zhang 0001, Yang Feng 0004 |
ICLR | 6 |
| 2025 | LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision TokenabstractThe advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision tokens (continuous representations) and integrate them and textual instructions into the context of large language models (LLMs), where large-scale parameters and numerous context tokens (predominantly vision tokens) result in substantial computational overhead. Previous efforts towards efficient LMMs always focus on replacing the LLM backbone with smaller models, while neglecting the crucial issue of token quantity. In this paper, we introduce LLaVA-Mini, an efficient LMM with minimal vision tokens. To achieve a high compression ratio of vision tokens while preserving visual information, we first analyze how LMMs understand vision tokens and find that most vision tokens only play a crucial role in the early layers of LLM backbone, where they mainly fuse visual information into text tokens. Building on this finding, LLaVA-Mini introduces modality pre-fusion to fuse visual information into text tokens in advance, thereby facilitating the extreme compression of vision tokens fed to LLM backbone into one token. LLaVA-Mini is a unified large multimodal model that can support the understanding of images, high-resolution images, and videos in an efficient manner. Experiments across 11 image-based and 7 video-based benchmarks demonstrate that LLaVA-Mini outperforms LLaVA-v1.5 with just 1 vision token instead of 576. Efficiency analyses reveal that LLaVA-Mini can reduce FLOPs by 77%, deliver low-latency responses within 40 milliseconds, and process over 10,000 frames of video on the GPU hardware with 24GB of memory. Shaolei Zhang 0001, Qingkai Fang, Yang Feng 0004 |
ICLR | 4 |
| 2025 | Overcoming Non-monotonicity in Transducer-based Streaming GenerationabstractStreaming generation models are utilized across fields, with the Transducer architecture being popular in industrial applications. However, its input-synchronous decoding mechanism presents challenges in tasks requiring non-monotonic alignments, such as simultaneous translation. In this research, we address this issue by integrating Transducer’s decoding with the history of input stream via a learnable monotonic attention. Our approach leverages the forward-backward algorithm to infer the posterior probability of alignments between the predictor states and input timestamps, which is then used to estimate the monotonic context representations, thereby avoiding the need to enumerate the exponentially large alignment space during training. Extensive experiments show that our MonoAttn-Transducer effectively handles non-monotonic alignments in streaming scenarios, offering a robust solution for complex generation tasks. Code is available at https://github.com/ictnlp/MonoAttn-Transducer. Zhengrui Ma, Yang Feng 0004, Min Zhang 0005 |
ICML | 2 |
| 2025 | MoCE: Adaptive Mixture of Contextualization Experts for Byte-based Neural Machine TranslationabstractLanglin Huang, Mengyu Bu, Yang Feng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Langlin Huang, Mengyu Bu, Yang Feng 0004 |
NAACL (Long Papers) | 3 |
| 2025 | FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech ProcessingabstractThe rapid advancement of Large Language Models (LLMs) has spurred significant progress in Large Speech-Language Models (LSLMs), enhancing their capabilities in both speech understanding and generation. While existing LSLMs often concentrate on augmenting speech generation or tackling a diverse array of short-speech tasks, the efficient processing of long-form speech remains a critical yet underexplored challenge. This gap is primarily attributed to the scarcity of long-speech training datasets and the high computational costs associated with long sequences. To address these limitations, we introduce FastLongSpeech, a novel framework designed to extend LSLM capabilities for efficient long-speech processing without necessitating dedicated long-speech training data. FastLongSpeech incorporates an iterative fusion strategy that can compress excessively long-speech sequences into manageable lengths. To adapt LSLMs for long-speech inputs, it introduces a dynamic compression training approach, which exposes the model to short-speech sequences at varying compression ratios, thereby transferring the capabilities of LSLMs to long-speech tasks. To assess the long-speech capabilities of LSLMs, we develop a long-speech understanding benchmark called LongSpeech-Eval. Experiments show that our method exhibits strong performance in both long-speech and short-speech tasks, while greatly improving inference efficiency. Shoutao Guo, Shaolei Zhang 0001, Qingkai Fang, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
NeurIPS | 6 |
| 2025 | Efficient Speech Language Modeling via Energy Distance in Continuous Latent SpaceabstractWe introduce \emph{SLED}, an alternative approach to speech language modeling by encoding speech waveforms into sequences of continuous latent representations and modeling them autoregressively using an energy distance objective. The energy distance offers an analytical measure of the distributional gap by contrasting simulated and target samples, enabling efficient training to capture the underlying continuous autoregressive distribution. By bypassing reliance on residual vector quantization, SLED avoids discretization errors and eliminates the need for the complicated hierarchical architectures common in existing speech language models. It simplifies the overall modeling pipeline while preserving the richness of speech information and maintaining inference efficiency. Empirical results demonstrate that SLED achieves strong performance in both zero-shot and streaming speech synthesis, showing its potential for broader applications in general-purpose speech language models. Demos and code are available at \url{https://github.com/ictnlp/SLED-TTS}. Zhengrui Ma, Yang Feng 0004, Chenze Shao, Fandong Meng, Jie Zhou 0016, Min Zhang 0005 |
NeurIPS | 2 |
| 2024 | Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?abstractTwo-pass direct speech-to-speech translation (S2ST) models have shown promising results which decompose S2ST into speech-to-text translation (S2TT) and text-to-speech (TTS), yet conduct end-to-end training by sharing the target text representation between S2TT and TTS models.However, the training of these models still requires large-scale parallel speech data comprising triplets, which is extremely challenging to collect.On the other hand, S2TT and TTS have accumulated a large amount of data and numerous pretrained models, which can be used to reduce the reliance on parallel speech data.To this end, we propose a composite S2ST model named ComSpeech, which connects pretrained S2TT and TTS models by introducing a vocabulary adaptor based on connectionist temporal classification (CTC).The vocabulary adaptor is employed to adapt the output text sequence of S2TT to the input text sequence of TTS, which are different due to the use of different vocabularies.In this way, ComSpeech can still be trained end-to-end and only needs a small amount of parallel speech data to finetune.We further propose a novel training method ComSpeech-ZS to eliminate the reliance on parallel speech data by aligning the text representation space of S2TT and TTS.Experimental results on the CVSS dataset show that when the parallel speech data is available, ComSpeech surpasses previous two-pass models like UnitY and Translatotron 2 in both translation quality and decoding speed.When there is no parallel speech data, ComSpeech-ZS lags behind ComSpeech by only 0.7 ASR-BLEU and outperforms the cascaded models. 1 Qingkai Fang, Shaolei Zhang 0001, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
ACL (1) | 5 |
| 2024 | Decoder-only Streaming Transformer for Simultaneous TranslationabstractSimultaneous Machine Translation (SiMT) generates translation while reading source tokens, essentially producing the target prefix based on the source prefix.To achieve good performance, it leverages the relationship between source and target prefixes to exact a policy to guide the generation of translations.Although existing SiMT methods primarily focus on the Encoder-Decoder architecture, we explore the potential of Decoder-only architecture, owing to its superior performance in various tasks and its inherent compatibility with SiMT.However, directly applying the Decoder-only architecture to SiMT poses challenges in terms of training and inference.To alleviate the above problems, we propose the first Decoder-only SiMT model, named Decoder-only Streaming Transformer (DST).Specifically, DST separately encodes the positions of the source and target prefixes, ensuring that the position of the target prefix remains unaffected by the expansion of the source prefix.Furthermore, we propose a Streaming Self-Attention (SSA) mechanism tailored for the Decoder-only architecture.It is capable of obtaining translation policy by assessing the sufficiency of input source information and integrating with the soft-attention mechanism to generate translations.Experiments demonstrate that our approach achieves state-of-theart performance on three translation tasks 1 . Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 3 |
| 2024 | A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any TranslationabstractSimultaneous translation models play a crucial role in facilitating communication.However, existing research primarily focuses on text-totext or speech-to-text models, necessitating additional cascade components to achieve speechto-speech translation.These pipeline methods suffer from error propagation and accumulate delays in each cascade component, resulting in reduced synchronization between the speaker and listener.To overcome these challenges, we propose a novel non-autoregressive generation framework for simultaneous speech translation (NAST-S2x 1 ), which integrates speechto-text and speech-to-speech tasks into a unified end-to-end framework.We develop a nonautoregressive decoder capable of concurrently generating multiple text or acoustic unit tokens upon receiving fixed-length speech chunks.The decoder can generate blank or repeated tokens and employ CTC decoding to dynamically adjust its latency.Experimental results show that NAST-S2x outperforms state-of-theart models in both speech-to-text and speechto-speech tasks.It achieves high-quality simultaneous interpretation within a delay of less than 3 seconds and provides a 28× decoding speedup in offline generation. 2 Zhengrui Ma, Qingkai Fang, Shaolei Zhang 0001, Shoutao Guo, Yang Feng 0004, Min Zhang 0005 |
ACL (1) | 5 |
| 2024 | StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task LearningabstractSimultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication.Beyond accomplishing translation between speech, Simul-S2ST requires a policy to control the model to generate corresponding target speech at the opportune moment within speech inputs, thereby posing a double challenge of translation and policy.In this paper, we propose StreamSpeech, a direct Simul-S2ST model that jointly learns translation and simultaneous policy in a unified framework of multi-task learning.Adhering to a multi-task learning approach, StreamSpeech can perform offline and simultaneous speech recognition, speech translation and speech synthesis via an "All-in-One" seamless model.Experiments on CVSS benchmark demonstrate that StreamSpeech achieves state-of-the-art performance in both offline S2ST and Simul-S2ST tasks.Besides, StreamSpeech is able to present high-quality intermediate results (i.e., ASR or translation results) during simultaneous translation process, offering a more comprehensive real-time communication experience 1 . Shaolei Zhang 0001, Qingkai Fang, Shoutao Guo, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
ACL (1) | 6 |
| 2024 | TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceabstractLarge Language Models (LLMs) sometimes suffer from producing hallucinations, especially LLMs may generate untruthful responses despite knowing the correct knowledge.Activating the truthfulness within LLM is the key to fully unlocking LLM's knowledge potential.In this paper, we propose TruthX, an inference-time intervention method to activate the truthfulness of LLM by identifying and editing the features within LLM's internal representations that govern the truthfulness.TruthX employs an auto-encoder to map LLM's representations into semantic and truthful latent spaces respectively, and applies contrastive learning to identify a truthful editing direction within the truthful space.During inference, by editing LLM's internal representations in truthful space, TruthX effectively enhances the truthfulness of LLM.Experiments show that TruthX improves the truthfulness of 13 advanced LLMs by an average of 20% on Truth-fulQA benchmark.Further analyses suggest that TruthX can control LLM to produce truthful or hallucinatory responses via editing only one vector in LLM's internal representations 1 . Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 3 |
| 2024 | Glancing Future for Simultaneous Machine TranslationabstractSimultaneous machine translation (SiMT) outputs translation while reading the source sentence. Unlike conventional sequence-to-sequence (seq2seq) training, existing SiMT methods adopt the prefix-to-prefix (prefix2prefix) training, where the model predicts target tokens based on partial source tokens. However, the prefix2prefix training diminishes the ability of the model to capture global information and introduces forced predictions due to the absence of essential source information. Consequently, it is crucial to bridge the gap between the prefix2prefix training and seq2seq training to enhance the translation capability of the SiMT model. In this paper, we propose a novel method that glances future in curriculum learning to achieve the transition from the seq2seq training to prefix2prefix training. Specifically, we gradually reduce the available source information from the whole sentence to the prefix corresponding to that latency. Our method is applicable to a wide range of SiMT methods and experiments demonstrate that our method outperforms strong baselines1. Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ICASSP | 3 |
| 2024 | Speculative Decoding with CTC-based Draft Model for LLM Inference AccelerationabstractInference acceleration of large language models (LLMs) has been put forward in many application scenarios and speculative decoding has shown its advantage in addressing inference acceleration. Speculative decoding usually introduces a draft model to assist the base LLM where the draft model produces drafts and the base LLM verifies the draft for acceptance or rejection. In this framework, the final inference speed is decided by the decoding speed of the draft model and the acceptance rate of the draft provided by the draft model. Currently the widely used draft models usually generate draft tokens for the next several positions in a non-autoregressive way without considering the correlations between draft tokens. Therefore, it has a high decoding speed but an unsatisfactory acceptance rate. In this paper, we focus on how to improve the performance of the draft model and aim to accelerate inference via a high acceptance rate. To this end, we propose a CTC-based draft model which strengthens the correlations between draft tokens during the draft phase, thereby generating higher-quality draft candidate sequences. Experiment results show that compared to strong baselines, the proposed method can achieve a higher acceptance rate and hence a faster inference speed. Zhuofan Wen 0003, Shangtong Gui, Yang Feng 0004 |
NeurIPS | 3 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 11 |
| 2023 | Rephrasing the Reference for Non-autoregressive Machine TranslationabstractNon-autoregressive neural machine translation (NAT) models suffer from the multi-modality problem that there may exist multiple possible translations of a source sentence, so the reference sentence may be inappropriate for the training when the NAT output is closer to other translations. In response to this problem, we introduce a rephraser to provide a better training target for NAT by rephrasing the reference sentence according to the NAT output. As we train NAT based on the rephraser output rather than the reference sentence, the rephraser output should fit well with the NAT output and not deviate too far from the reference, which can be quantified as reward functions and optimized by reinforcement learning. Experiments on major WMT benchmarks and NAT baselines show that our approach consistently improves the translation quality of NAT. Specifically, our best variant achieves comparable performance to the autoregressive Transformer, while being 14.7 times more efficient in inference. Chenze Shao, Jinchao Zhang 0001, Jie Zhou 0016, Yang Feng 0004 |
AAAI | 4 |
| 2023 | Back Translation for Speech-to-text Translation Without TranscriptsabstractThe success of end-to-end speech-to-text translation (ST) is often achieved by utilizing source transcripts, e.g., by pre-training with automatic speech recognition (ASR) and machine translation (MT) tasks, or by introducing additional ASR and MT data.Unfortunately, transcripts are only sometimes available since numerous unwritten languages exist worldwide.In this paper, we aim to utilize large amounts of targetside monolingual data to enhance ST without transcripts.Motivated by the remarkable success of back translation in MT, we develop a back translation algorithm for ST (BT4ST) to synthesize pseudo ST data from monolingual target data.To ease the challenges posed by short-to-long generation and one-to-many mapping, we introduce self-supervised discrete units and achieve back translation by cascading a target-to-unit model and a unit-to-speech model.With our synthetic ST data, we achieve an average boost of 2.3 BLEU on MuST-C En→De, En→Fr, and En→Es datasets.More experiments show that our method is especially effective in low-resource scenarios.12 Qingkai Fang, Yang Feng 0004 |
ACL (1) | 2 |
| 2023 | Understanding and Bridging the Modality Gap for Speech TranslationabstractHow to achieve better end-to-end speech translation (ST) by leveraging (text) machine translation (MT) data?Among various existing techniques, multi-task learning is one of the effective ways to share knowledge between ST and MT in which additional MT data can help to learn source-to-target mapping.However, due to the differences between speech and text, there is always a gap between ST and MT.In this paper, we first aim to understand this modality gap from the target-side representation differences, and link the modality gap to another well-known problem in neural machine translation: exposure bias.We find that the modality gap is relatively small during training except for some difficult cases, but keeps increasing during inference due to the cascading effect.To address these problems, we propose the Cross-modal Regularization with Scheduled Sampling (CRESS) method.Specifically, we regularize the output predictions of ST and MT, whose target-side contexts are derived by sampling between ground truth words and self-generated words with a varying probability.Furthermore, we introduce token-level adaptive training which assigns different training weights to target tokens to handle difficult cases with large modality gaps.Experiments and analysis show that our approach effectively bridges the modality gap, and achieves promising results in all eight directions of the MuST-C dataset. 1 Qingkai Fang, Yang Feng 0004 |
ACL (1) | 2 |
| 2023 | Learning Optimal Policy for Simultaneous Machine Translation via Binary SearchabstractSimultaneous machine translation (SiMT) starts to output translation while reading the source sentence and needs a precise policy to decide when to output the generated translation.Therefore, the policy determines the number of source tokens read during the translation of each target token.However, it is difficult to learn a precise translation policy to achieve good latency-quality trade-offs, because there is no golden policy corresponding to parallel sentences as explicit supervision.In this paper, we present a new method for constructing the optimal policy online via binary search.By employing explicit supervision, our approach enables the SiMT model to learn the optimal policy, which can guide the model in completing the translation during inference.Experiments on four translation tasks show that our method can exceed strong baselines across all latency scenarios 1 Shoutao Guo, Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 3 |
| 2023 | CMOT: Cross-modal Mixup via Optimal Transport for Speech TranslationabstractEnd-to-end speech translation (ST) is the task of translating speech signals in the source language into text in the target language.As a cross-modal task, end-to-end ST is difficult to train with limited data.Existing methods often try to transfer knowledge from machine translation (MT), but their performances are restricted by the modality gap between speech and text.In this paper, we propose Cross-modal Mixup via Optimal Transport (CMOT) to overcome the modality gap.We find the alignment between speech and text sequences via optimal transport and then mix up the sequences from different modalities at a token level using the alignment.Experiments on the MuST-C ST benchmark demonstrate that CMOT achieves an average BLEU of 30.0 in 8 translation directions, outperforming previous methods.Further analysis shows CMOT can adaptively find the alignment between modalities, which helps alleviate the modality gap between speech and text. Qingkai Fang, Yang Feng 0004 |
ACL (1) | 3 |
| 2023 | Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationabstractMultimodal machine translation (MMT) simultaneously takes the source sentence and a relevant image as input for translation.Since there is no paired image available for the input sentence in most cases, recent studies suggest utilizing powerful text-to-image generation models to provide image inputs.Nevertheless, synthetic images generated by these models often follow different distributions compared to authentic images.Consequently, using authentic images for training and synthetic images for inference can introduce a distribution shift, resulting in performance degradation during inference.To tackle this challenge, in this paper, we feed synthetic and authentic images to the MMT model, respectively.Then we minimize the gap between the synthetic and authentic images by drawing close the input image representations of the Transformer Encoder and the output distributions of the Transformer Decoder.Therefore, we mitigate the distribution disparity introduced by the synthetic images during inference, thereby freeing the authentic images from the inference process.Experimental results show that our approach achieves state-ofthe-art performance on the Multi30K En-De and En-Fr datasets, while remaining independent of authentic images during inference. Wenyu Guo, Qingkai Fang, Dong Yu 0003, Yang Feng 0004 |
EMNLP | 4 |
| 2023 | Non-autoregressive Streaming Transformer for Simultaneous TranslationabstractSimultaneous machine translation (SiMT) models are trained to strike a balance between latency and translation quality.However, training these models to achieve high quality while maintaining low latency often leads to a tendency for aggressive anticipation.We argue that such issue stems from the autoregressive architecture upon which most existing SiMT models are built.To address those issues, we propose non-autoregressive streaming Transformer (NAST) which comprises a unidirectional encoder and a non-autoregressive decoder with intra-chunk parallelism.We enable NAST to generate the blank token or repetitive tokens to adjust its READ/WRITE strategy flexibly, and train it to maximize the nonmonotonic latent alignment with an alignmentbased latency loss.Experiments on various SiMT benchmarks demonstrate that NAST outperforms previous strong autoregressive SiMT baselines. Zhengrui Ma, Shaolei Zhang 0001, Shoutao Guo, Chenze Shao, Min Zhang 0005, Yang Feng 0004 |
EMNLP | 6 |
| 2023 | Fuzzy Alignments in Directed Acyclic Graph for Non-Autoregressive Machine Translation
Zhengrui Ma, Chenze Shao, Shangtong Gui, Min Zhang 0005, Yang Feng 0004 |
ICLR | 5 |
| 2023 | Hidden Markov Transformer for Simultaneous Machine Translation
Shaolei Zhang 0001, Yang Feng 0004 |
ICLR | 2 |
| 2023 | DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech TranslationabstractDirect speech-to-speech translation (S2ST) translates speech from one language into another using a single model. However, due to the presence of linguistic and acoustic diversity, the target speech follows a complex multimodal distribution, posing challenges to achieving both high-quality translations and fast decoding speeds for S2ST models. In this paper, we propose DASpeech, a non-autoregressive direct S2ST model which realizes both fast and high-quality S2ST. To better capture the complex distribution of the target speech, DASpeech adopts the two-pass architecture to decompose the generation process into two steps, where a linguistic decoder first generates the target text, and an acoustic decoder then generates the target speech based on the hidden states of the linguistic decoder. Specifically, we use the decoder of DA-Transformer as the linguistic decoder, and use FastSpeech 2 as the acoustic decoder. DA-Transformer models translations with a directed acyclic graph (DAG). To consider all potential paths in the DAG during training, we calculate the expected hidden states for each target token via dynamic programming, and feed them into the acoustic decoder to predict the target mel-spectrogram. During inference, we select the most probable path and take hidden states on that path as input to the acoustic decoder. Experiments on the CVSS Fr$\rightarrow$En benchmark demonstrate that DASpeech can achieve comparable or even better performance than the state-of-the-art S2ST model Translatotron 2, while preserving up to 18.53$\times$ speedup compared to the autoregressive baseline. Compared with the previous non-autoregressive S2ST model, DASpeech does not rely on knowledge distillation and iterative decoding, achieving significant improvements in both translation quality and decoding speed. Furthermore, DASpeech shows the ability to preserve the speaker's voice of the source speech during translation. Qingkai Fang, Yang Feng 0004 |
NeurIPS | 3 |
| 2023 | Non-autoregressive Machine Translation with Probabilistic Context-free GrammarabstractNon-autoregressive Transformer(NAT) significantly accelerates the inference of neural machine translation. However, conventional NAT models suffer from limited expression power and performance degradation compared to autoregressive (AT) models due to the assumption of conditional independence among target tokens. To address these limitations, we propose a novel approach called PCFG-NAT, which leverages a specially designed Probabilistic Context-Free Grammar (PCFG) to enhance the ability of NAT models to capture complex dependencies among output tokens. Experimental results on major machine translation benchmarks demonstrate that PCFG-NAT further narrows the gap in translation quality between NAT and AT models. Moreover, PCFG-NAT facilitates a deeper understanding of the generated sentences, addressing the lack of satisfactory explainability in neural machine translation. Code is publicly available at https://github.com/ictnlp/PCFG-NAT. Shangtong Gui, Chenze Shao, Zhengrui Ma, Xishan Zhang, Yunji Chen, Yang Feng 0004 |
NeurIPS | 6 |
| 2023 | Beyond MLE: Convex Learning for Text GenerationabstractMaximum likelihood estimation (MLE) is a statistical method used to estimate the parameters of a probability distribution that best explain the observed data. In the context of text generation, MLE is often used to train generative language models, which can then be used to generate new text. However, we argue that MLE is not always necessary and optimal, especially for closed-ended text generation tasks like machine translation. In these tasks, the goal of model is to generate the most appropriate response, which does not necessarily require it to estimate the entire data distribution with MLE. To this end, we propose a novel class of training objectives based on convex functions, which enables text generation models to focus on highly probable outputs without having to estimate the entire data distribution. We investigate the theoretical properties of the optimal predicted distribution when applying convex functions to the loss, demonstrating that convex functions can sharpen the optimal distribution, thereby enabling the model to better capture outputs with high probabilities. Experiments on various text generation tasks and models show the effectiveness of our approach. It enables autoregressive models to bridge the gap between greedy and beam search, and facilitates the learning of non-autoregressive models with a maximum improvement of 9+ BLEU points. Moreover, our approach also exhibits significant impact on large language models (LLMs), substantially enhancing their generative capability on various tasks. Source code is available at \url{https://github.com/ictnlp/Convex-Learning}. Chenze Shao, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
NeurIPS | 4 |
| 2023 | Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationabstractSimultaneous sequence generation is a pivotal task for real-time scenarios, such as streaming speech recognition, simultaneous machine translation and simultaneous speech translation, where the target sequence is generated while receiving the source sequence. The crux of achieving high-quality generation with low latency lies in identifying the optimal moments for generating, accomplished by learning a mapping between the source and target sequences. However, existing methods often rely on task-specific heuristics for different sequence types, limiting the model’s capacity to adaptively learn the source-target mapping and hindering the exploration of multi-task learning for various simultaneous tasks. In this paper, we propose a unified segment-to-segment framework (Seg2Seg) for simultaneous sequence generation, which learns the mapping in an adaptive and unified manner. During the process of simultaneous generation, the model alternates between waiting for a source segment and generating a target segment, making the segment serve as the natural bridge between the source and target. To accomplish this, Seg2Seg introduces a latent segment as the pivot between source to target and explores all potential source-target mappings via the proposed expectation training, thereby learning the optimal moments for generating. Experiments on multiple simultaneous generation tasks demonstrate that Seg2Seg achieves state-of-the-art performance and exhibits better generality across various tasks. Shaolei Zhang 0001, Yang Feng 0004 |
NeurIPS | 2 |
| 2022 | Neural Machine Translation with Phrase-Level Universal Visual RepresentationsabstractMultimodal machine translation (MMT) aims to improve neural machine translation (NMT) with additional visual information, but most existing MMT methods require paired input of source sentence and image, which makes them suffer from shortage of sentence-image pairs.In this paper, we propose a phrase-level retrieval-based method for MMT to get visual information for the source input from existing sentence-image data sets so that MMT can break the limitation of paired sentence-image input.Our method performs retrieval at the phrase level and hence learns visual information from pairs of source phrase and grounded region, which can mitigate data sparsity.Furthermore, our method employs the conditional variational auto-encoder to learn visual representations which can filter redundant visual information and only retain visual information related to the phrase.Experiments show that the proposed method significantly outperforms strong baselines on multiple MMT datasets, especially when the textual context is limited. Qingkai Fang, Yang Feng 0004 |
ACL (1) | 2 |
| 2022 | STEMM: Self-learning with Speech-text Manifold Mixup for Speech TranslationabstractHow to learn a better speech representation for end-to-end speech-to-text translation (ST) with limited labeled data?Existing techniques often attempt to transfer powerful machine translation (MT) capabilities to ST, but neglect the representation discrepancy across modalities.In this paper, we propose the Speech-TExt Manifold Mixup (STEMM) method to calibrate such discrepancy.Specifically, we mix up the representation sequences of different modalities, and take both unimodal speech sequences and multimodal mixed sequences as input to the translation model in parallel, and regularize their output predictions with a selflearning framework.Experiments on MuST-C speech translation benchmark and further analysis show that our method effectively alleviates the cross-modal representation discrepancy, and achieves significant improvements over a strong baseline on eight translation directions.* indicates corresponding authors. Qingkai Fang, Rong Ye, Lei Li 0005, Yang Feng 0004, Mingxuan Wang |
ACL (1) | 4 |
| 2022 | Prediction Difference Regularization against Perturbation for Neural Machine TranslationabstractRegularization methods applying input perturbation have drawn considerable attention and have been frequently explored for NMT tasks in recent years.Despite their simplicity and effectiveness, we argue that these methods are limited by the under-fitting of training data.In this paper, we utilize prediction difference for ground-truth tokens to analyze the fitting of token-level samples and find that underfitting is almost as common as over-fitting.We introduce prediction difference regularization (PD-R), a simple and effective method that can reduce over-fitting and under-fitting at the same time.For all token-level samples, PD-R minimizes the prediction difference between the original pass and the input-perturbed pass, making the model less sensitive to small input changes, thus more robust to both perturbations and under-fitted training data.Experiments on three widely used WMT translation tasks show that our approach can significantly improve over existing perturbation regularization methods.On WMT16 En-De task, our model achieves 1.80 SacreBLEU improvement over vanilla transformer. Dengji Guo, Zhengrui Ma, Min Zhang 0005, Yang Feng 0004 |
ACL (1) | 4 |
| 2022 | Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine TranslationabstractNeural networks tend to gradually forget the previously learned knowledge when learning multiple tasks sequentially from dynamic data distributions.This problem is called catastrophic forgetting, which is a fundamental challenge in the continual learning of neural networks.In this work, we observe that catastrophic forgetting not only occurs in continual learning but also affects the traditional static training.Neural networks, especially neural machine translation models, suffer from catastrophic forgetting even if they learn from a static training set.To be specific, the final model pays imbalanced attention to training samples, where recently exposed samples attract more attention than earlier samples.The underlying cause is that training samples do not get balanced training in each model update, so we name this problem imbalanced training.To alleviate this problem, we propose Complementary Online Knowledge Distillation (COKD), which uses dynamically updated teacher models trained on specific data orders to iteratively provide complementary knowledge to the student model.Experimental results on multiple machine translation tasks show that our method successfully alleviates the problem of imbalanced training and achieves substantial improvements over strong baseline systems. 1 Chenze Shao, Yang Feng 0004 |
ACL (1) | 2 |
| 2022 | Modeling Dual Read/Write Paths for Simultaneous Machine TranslationabstractSimultaneous machine translation (SiMT) outputs translation while reading source sentence and hence requires a policy to decide whether to wait for the next source word (READ) or generate a target word (WRITE), the actions of which form a read/write path.Although the read/write path is essential to SiMT performance, no direct supervision is given to the path in the existing methods.In this paper, we propose a method of dual-path SiMT which introduces duality constraints to direct the read/write path.According to duality constraints, the read/write path in source-totarget and target-to-source SiMT models can be mapped to each other.As a result, the two SiMT models can be optimized jointly by forcing their read/write paths to satisfy the mapping.Experiments on En↔Vi and De↔En tasks show that our method can outperform strong baselines under all latency. Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 2 |
| 2022 | Reducing Position Bias in Simultaneous Machine Translation with Length-Aware FrameworkabstractSimultaneous machine translation (SiMT) starts translating while receiving the streaming source inputs, and hence the source sentence is always incomplete during translating.Different from the full-sentence MT using the conventional seq-to-seq architecture, SiMT often applies prefix-to-prefix architecture, which forces each target word to only align with a partial source prefix to adapt to the incomplete source in streaming inputs.However, the source words in the front positions are always illusoryly considered more important since they appear in more prefixes, resulting in position bias, which makes the model pay more attention on the front source positions in testing.In this paper, we first analyze the phenomenon of position bias in SiMT, and develop a Length-Aware Framework to reduce the position bias by bridging the structural gap between SiMT and fullsentence MT.Specifically, given the streaming inputs, we first predict the full-sentence length and then fill the future source position with positional encoding, thereby turning the streaming inputs into a pseudo full-sentence.The proposed framework can be integrated into most existing SiMT methods to further improve performance.Experiments on two representative SiMT methods, including the state-of-theart adaptive policy, show that our method successfully reduces the position bias and thereby achieves better SiMT performance. Shaolei Zhang 0001, Yang Feng 0004 |
ACL (1) | 2 |
| 2022 | Continual Learning of Neural Machine Translation within Low Forgetting Risk RegionsabstractThis paper considers continual learning of large-scale pretrained neural machine translation model without accessing the previous training data or introducing model separation.We argue that the widely used regularization-based methods, which perform multi-objective learning with an auxiliary loss, suffer from the misestimate problem and cannot always achieve a good balance between the previous and new tasks.To solve the problem, we propose a twostage training method based on the local features of the real loss.We first search low forgetting risk regions, where the model can retain the performance on the previous task as the parameters are updated, to avoid the catastrophic forgetting problem.Then we can continually train the model within this region only with the new training data to fit the new task.Specifically, we propose two methods to search the low forgetting risk regions, which are based on the curvature of loss and the impacts of the parameters on the model output, respectively.We conduct experiments on domain adaptation and more challenging language adaptation tasks, and the experimental results show that our method can achieve significant improvements compared with several strong baselines. Shuhao Gu, Bojie Hu, Yang Feng 0004 |
EMNLP | 3 |
| 2022 | Counterfactual Data Augmentation via Perspective Transition for Open-Domain DialoguesabstractThe construction of open-domain dialogue systems requires high-quality dialogue datasets.The dialogue data admits a wide variety of responses for a given dialogue history, especially responses with different semantics.However, collecting high-quality such a dataset in most scenarios is labor-intensive and timeconsuming.In this paper, we propose a data augmentation method to automatically augment high-quality responses with different semantics by counterfactual inference.Specifically, given an observed dialogue, our counterfactual generation model first infers semantically different responses by replacing the observed reply perspective with substituted ones.Furthermore, our data selection method filters out detrimental augmented responses.Experimental results show that our data augmentation method can augment high-quality responses with different semantics for a given dialogue history, and can outperform competitive baselines on multiple downstream tasks. Jiao Ou, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016 |
EMNLP | 3 |
| 2022 | Low-resource Neural Machine Translation with Cross-modal AlignmentabstractHow to achieve neural machine translation with limited parallel data?Existing techniques often rely on large-scale monolingual corpora, which is impractical for some low-resource languages.In this paper, we turn to connect several low-resource languages to a particular high-resource one by additional visual modality.Specifically, we propose a cross-modal contrastive learning method to learn a shared space for all languages, where both a coarsegrained sentence-level objective and a finegrained token-level one are introduced.Experimental results and further analysis show that our method can effectively learn the crossmodal and cross-lingual alignment with a small amount of image-text pairs and achieves significant improvements over the text-only baseline under both zero-shot and few-shot scenarios.Our code could be found at https: //github.com/ictnlp/LNMT-CA. Qingkai Fang, Yang Feng 0004 |
EMNLP | 3 |
| 2022 | Information-Transport-based Policy for Simultaneous TranslationabstractSimultaneous translation (ST) outputs translation while receiving the source inputs, and hence requires a policy to determine whether to translate a target token or wait for the next source token.The major challenge of ST is that each target token can only be translated based on the current received source tokens, where the received source information will directly affect the translation quality.So naturally, how much source information is received for the translation of the current target token is supposed to be the pivotal evidence for the ST policy to decide between translating and waiting.In this paper, we treat the translation as information transport from source to target and accordingly propose an Information-Transport-based Simultaneous Translation (ITST).ITST quantifies the transported information weight from each source token to the current target token, and then decides whether to translate the target token according to its accumulated received information.Experiments on both text-to-text ST and speech-to-text ST (a.k.a., streaming speech translation) tasks show that ITST outperforms strong baselines and achieves state-of-the-art performance 1 . Shaolei Zhang 0001, Yang Feng 0004 |
EMNLP | 2 |
| 2022 | One Reference Is Not Enough: Diverse Distillation with Reference Selection for Non-Autoregressive TranslationabstractNon-autoregressive neural machine translation (NAT) suffers from the multi-modality problem: the source sentence may have multiple correct translations, but the loss function is calculated only according to the reference sentence.Sequence-level knowledge distillation makes the target more deterministic by replacing the target with the output from an autoregressive model.However, the multi-modality problem in the distilled dataset is still nonnegligible.Furthermore, learning from a specific teacher limits the upper bound of the model capability, restricting the potential of NAT models.In this paper, we argue that one reference is not enough and propose diverse distillation with reference selection (DDRS) for NAT.Specifically, we first propose a method called SeedDiv for diverse machine translation, which enables us to generate a dataset containing multiple high-quality reference translations for each source sentence.During the training, we compare the NAT output with all references and select the one that best fits the NAT output to train the model.Experiments on widely-used machine translation benchmarks demonstrate the effectiveness of DDRS, which achieves 29.82 BLEU with only one decoding pass on WMT14 En-De, improving the state-of-the-art performance for NAT by over 1 BLEU. 1 Chenze Shao, Xuanfu Wu, Yang Feng 0004 |
NAACL-HLT | 3 |
| 2022 | Non-Monotonic Latent Alignments for CTC-Based Non-Autoregressive Machine TranslationabstractNon-autoregressive translation (NAT) models are typically trained with the cross-entropy loss, which forces the model outputs to be aligned verbatim with the target sentence and will highly penalize small shifts in word positions. Latent alignment models relax the explicit alignment by marginalizing out all monotonic latent alignments with the CTC loss. However, they cannot handle non-monotonic alignments, which is non-negligible as there is typically global word reordering in machine translation. In this work, we explore non-monotonic latent alignments for NAT. We extend the alignment space to non-monotonic alignments to allow for the global word reordering and further consider all alignments that overlap with the target sentence. We non-monotonically match the alignments to the target sentence and train the latent alignment model to maximize the F1 score of non-monotonic matching. Extensive experiments on major WMT benchmarks show that our method substantially improves the translation performance of CTC-based models. Our best model achieves 30.06 BLEU on WMT14 En-De with only one-iteration decoding, closing the gap between non-autoregressive and autoregressive models. Chenze Shao, Yang Feng 0004 |
NeurIPS | 2 |
| 2021 | Future-Guided Incremental Transformer for Simultaneous TranslationabstractSimultaneous translation (ST) starts translations synchronously while reading source sentences, and is used in many online scenarios. The previous wait-k policy is concise and achieved good results in ST. However, wait-k policy faces two weaknesses: low training speed caused by the recalculation of hidden states and lack of future source information to guide training. For the low training speed, we propose an incremental Transformer with an average embedding layer (AEL) to accelerate the speed of calculation of the hidden states during training. For future-guided training, we propose a conventional Transformer as the teacher of the incremental Transformer, and try to invisibly embed some future information in the model through knowledge distillation. We conducted experiments on Chinese-English and German-English simultaneous translation tasks and compared with the wait-k policy to evaluate the proposed method. Our method can effectively increase the training speed by about 28 times on average at different k and implicitly embed some predictive abilities in the model, achieving better translation quality than wait-k baseline. Shaolei Zhang 0001, Yang Feng 0004, Liangyou Li |
AAAI | 2 |
| 2021 | Guiding Teacher Forcing with Seer Forcing for Neural Machine TranslationabstractYang Feng, Shuhao Gu, Dengji Guo, Zhengxin Yang, Chenze Shao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yang Feng 0004, Shuhao Gu, Dengji Guo, Zhengxin Yang, Chenze Shao |
ACL/IJCNLP (1) | 1 |
| 2021 | Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue UtterancesabstractZekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zekang Li, Jinchao Zhang 0001, Zhengcong Fei, Yang Feng 0004, Jie Zhou 0016 |
ACL/IJCNLP (1) | 4 |
| 2021 | GTM: A Generative Triple-wise Model for Conversational Question GenerationabstractLei Shen, Fandong Meng, Jinchao Zhang, Yang Feng, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Lei Shen 0001, Fandong Meng, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016 |
ACL/IJCNLP (1) | 4 |
| 2021 | Importance-based Neuron Allocation for Multilingual Neural Machine TranslationabstractWanying Xie, Yang Feng, Shuhao Gu, Dong Yu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wanying Xie, Yang Feng 0004, Shuhao Gu, Dong Yu 0003 |
ACL/IJCNLP (1) | 2 |
| 2021 | Universal Simultaneous Machine Translation with Mixture-of-Experts Wait-k PolicyabstractSimultaneous machine translation (SiMT) generates translation before reading the entire source sentence and hence it has to trade off between translation quality and latency.To fulfill the requirements of different translation quality and latency in practical applications, the previous methods usually need to train multiple SiMT models for different latency levels, resulting in large computational costs.In this paper, we propose a universal SiMT model with Mixture-of-Experts Wait-k Policy to achieve the best translation quality under arbitrary latency with only one trained model.Specifically, our method employs multi-head attention to accomplish the mixture of experts where each head is treated as a wait-k expert with its own waiting words number, and given a test latency and source inputs, the weights of the experts are accordingly adjusted to produce the best translation.Experiments on three datasets show that our method outperforms all the strong baselines under different latency, including the state-of-the-art adaptive policy. Shaolei Zhang 0001, Yang Feng 0004 |
EMNLP (1) | 2 |
| 2021 | Learning to Select Context in a Hierarchical and Global Perspective for Open-Domain Dialogue GenerationabstractOpen-domain multi-turn conversations mainly have three features, which are hierarchical semantic structure, redundant information, and long-term dependency. Grounded on these, selecting relevant context becomes a challenge step for multiturn dialogue generation. However, existing methods cannot differentiate both useful words and utterances in long distances from a response. Besides, previous work just performs context selection based on a state in the decoder, which lacks a global guidance and could lead some focuses on irrelevant or unnecessary information. In this paper, we propose a novel model with hierarchical self-attention mechanism and distant supervision to not only detect relevant words and utterances in short and long distances, but also discern related information globally when decoding. Experimental results on two public datasets of both automatic and human evaluations show that our model significantly outperforms other baselines in terms of fluency, coherence, and informativeness. Lei Shen 0001, Haolan Zhan, Yang Feng 0004 |
ICASSP | 4 |
| 2021 | SE-DAE: Style-Enhanced Denoising Auto-Encoder for Unsupervised Text Style TransferabstractText style transfer aims to change the style of sentences while preserving the semantic meanings. Due to the lack of parallel data, the Denoising Auto-Encoder (DAE) is widely used in this task to model distributions of different sentence styles. However, because of the conflict between the target of the conventional denoising procedure and the target of style transfer task, the vanilla DAE can not produce satisfying enough results. To improve the transferability of the model, most of the existing works combine DAE with various complicated unsupervised networks, which makes the whole system become over-complex. In this work, we design a novel DAE model named Style-Enhanced DAE (SE-DAE), which is specifically designed for the text style transfer task. Compared with previous complicated style-transfer models, our model do not consist of any complicated unsupervised networks, but only relies on the high-quality pseudo-parallel data generated by a novel data refinement mechanism. Moreover, to alleviate the conflict between the targets of the conventional denoising procedure and the style transfer task, we propose another novel style denoising mechanism, which is more compatible with the target of the style transfer task. We validate the effectiveness of our model on two style benchmark datasets. Both automatic evaluation and human evaluation show that our proposed model is highly competitive compared with previous strong the state of the art (SOTA) approaches and greatly outperforms the vanilla DAE. Yang Feng 0004, Jiao Ou |
IJCNN | 2 |
| 2021 | Better Learning and Fusing Multi-Granularity Context Representations for Relevant Response GenerationabstractIn open-domain multi-turn dialogue, capturing related information from conversational history is key to generating relevant responses. In this paper, we propose a method to extract and fuse related context information at different levels of granularity with a word-level encoder, an utterance-level encoder, and a fusion decoder. The word-level encoder aims to obtain a context-aware representation for the last utterance, called a post, by attending the post to the context utterances at the word level. The utterance-level encoder is used to model the correlation among context utterances with the self-matching attention mechanism. Finally, the multi-granularity context representations from the two encoders are fused via the fusion decoder to improve the relevance of generated responses. Experimental results on the Ubuntu Dialog Corpus and Cornell Movie Dialog Corpus show that our methods significantly outperform all the strong baselines. Jiao Ou, Yang Feng 0004 |
IJCNN | 2 |
| 2021 | Modeling Coverage for Non-Autoregressive Neural Machine TranslationabstractNon-Autoregressive Neural Machine Translation (NAT) has achieved significant inference speedup by generating all tokens simultaneously. Despite its high efficiency, NAT usually suffers from two kinds of translation errors: over-translation (e.g. repeated tokens) and under-translation (e.g. missing translations), which eventually limits the translation quality. In this paper, we argue that these issues of NAT can be addressed through coverage modeling, which has been proved to be useful in autoregressive decoding. We propose a novel Coverage-NAT to model the coverage information directly by a token-level coverage iterative refinement mechanism and a sentence-level coverage agreement, which can remind the model if a source token has been translated or not and improve the semantics consistency between the translation and the source, respectively. Experimental results on WMT14 En↔De and WMT16 En↔Ro translation tasks show that our method can alleviate those errors and achieve strong improvements over the baseline system. Yong Shan, Yang Feng 0004, Chenze Shao |
IJCNN | 2 |
| 2021 | Pruning-then-Expanding Model for Domain Adaptation of Neural Machine TranslationabstractDomain Adaptation is widely used in practical applications of neural machine translation, which aims to achieve good performance on both general domain and in-domain data.However, the existing methods for domain adaptation usually suffer from catastrophic forgetting, large domain divergence, and model explosion.To address these three problems, we propose a method of "divide and conquer" which is based on the importance of neurons or parameters for the translation model.In this method, we first prune the model and only keep the important neurons or parameters, making them responsible for both generaldomain and in-domain translation.Then we further train the pruned model supervised by the original whole model with knowledge distillation.Last we expand the model to the original size and fine-tune the added parameters for the in-domain translation.We conducted experiments on different language pairs and domains and the results show that our method can achieve significant improvements compared with several strong baselines. Shuhao Gu, Yang Feng 0004, Wanying Xie |
NAACL-HLT | 2 |
| 2021 | Sequence-Level Training for Non-Autoregressive Neural Machine TranslationabstractAbstract In recent years, Neural Machine Translation (NMT) has achieved notable results in various translation tasks. However, the word-by-word generation manner determined by the autoregressive mechanism leads to high translation latency of the NMT and restricts its low-latency applications. Non-Autoregressive Neural Machine Translation (NAT) removes the autoregressive mechanism and achieves significant decoding speedup by generating target words independently and simultaneously. Nevertheless, NAT still takes the word-level cross-entropy loss as the training objective, which is not optimal because the output of NAT cannot be properly evaluated due to the multimodality problem. In this article, we propose using sequence-level training objectives to train NAT models, which evaluate the NAT outputs as a whole and correlates well with the real translation quality. First, we propose training NAT models to optimize sequence-level evaluation metrics (e.g., BLEU) based on several novel reinforcement algorithms customized for NAT, which outperform the conventional method by reducing the variance of gradient estimation. Second, we introduce a novel training objective for NAT models, which aims to minimize the Bag-of-N-grams (BoN) difference between the model output and the reference sentence. The BoN training objective is differentiable and can be calculated efficiently without doing any approximations. Finally, we apply a three-stage training strategy to combine these two methods to train the NAT model. We validate our approach on four translation tasks (WMT14 En↔De, WMT16 En↔Ro), which shows that our approach largely outperforms NAT baselines and achieves remarkable performance on all translation tasks. The source code is available at https://github.com/ictnlp/Seq-NAT. Chenze Shao, Yang Feng 0004, Jinchao Zhang 0001, Fandong Meng, Jie Zhou 0016 |
Comput. Linguistics | 2 |
| 2021 | Bridging Text and Video: A Universal Multimodal Transformer for Audio-Visual Scene-Aware DialogabstractAudio-Visual Scene-Aware Dialog (AVSD) is a task to generate responses when chatting about a given video, which is organized as a track of the 8th Dialog System Technology Challenge (DSTC8). There are two challenges in this task: 1) making effective interaction among different modalities; 2) better understanding dialogues and generating informative responses. To tackle the challenges, we propose a universal multimodal transformer and introduce the multi-task learning method to learn joint representations among different modalities as well as generate informative and fluent responses by leveraging the pre-trained language model. Our method extends the natural language generation pre-trained model to multimodal dialogue generation task, which allows fine-tuning language models to capture information across both visual and textual modalities. Our system achieves the best performance in the objective evaluation in both DSTC7-AVSD and DSTC8-AVSD dataset and achieves an impressive 98.4% of the human performance based on human ratings in the DSTC8-AVSD challenge. Zekang Li, Zongjia Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Modeling Fluency and Faithfulness for Diverse Neural Machine TranslationabstractNeural machine translation models usually adopt the teacher forcing strategy for training which requires the predicted sequence matches ground truth word by word and forces the probability of each prediction to approach a 0-1 distribution. However, the strategy casts all the portion of the distribution to the ground truth word and ignores other words in the target vocabulary even when the ground truth word cannot dominate the distribution. To address the problem of teacher forcing, we propose a method to introduce an evaluation module to guide the distribution of the prediction. The evaluation module accesses each prediction from the perspectives of fluency and faithfulness to encourage the model to generate the word which has a fluent connection with its past and future translation and meanwhile tends to form a translation equivalent in meaning to the source. The experiments on multiple translation tasks show that our method can achieve significant improvements over strong baselines. Yang Feng 0004, Wanying Xie, Shuhao Gu, Chenze Shao, Wen Zhang 0009, Zhengxin Yang, Dong Yu 0003 |
AAAI | 1 |
| 2020 | Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine TranslationabstractNon-Autoregressive Neural Machine Translation (NAT) achieves significant decoding speedup through generating target words independently and simultaneously. However, in the context of non-autoregressive translation, the word-level cross-entropy loss cannot model the target-side sequential dependency properly, leading to its weak correlation with the translation quality. As a result, NAT tends to generate influent translations with over-translation and under-translation errors. In this paper, we propose to train NAT to minimize the Bag-of-Ngrams (BoN) difference between the model output and the reference sentence. The bag-of-ngrams training objective is differentiable and can be efficiently calculated, which encourages NAT to capture the target-side sequential dependency and correlates well with the translation quality. We validate our approach on three translation tasks and show that our approach largely outperforms the NAT baseline by about 5.0 BLEU scores on WMT14 En↔De and about 2.5 BLEU scores on WMT16 En↔Ro. Chenze Shao, Jinchao Zhang 0001, Yang Feng 0004, Fandong Meng, Jie Zhou 0016 |
AAAI | 3 |
| 2020 | A Contextual Hierarchical Attention Network with Adaptive Objective for Dialogue State TrackingabstractRecent studies in dialogue state tracking (DST) leverage historical information to determine states which are generally represented as slot-value pairs. However, most of them have limitations to efficiently exploit relevant context due to the lack of a powerful mechanism for modeling interactions between the slot and the dialogue history. Besides, existing methods usually ignore the slot imbalance problem and treat all slots indiscriminately, which limits the learning of hard slots and eventually hurts overall performance. In this paper, we propose to enhance the DST through employing a contextual hierarchical attention network to not only discern relevant information at both word level and turn level but also learn contextual representations. We further propose an adaptive objective to alleviate the slot imbalance problem by dynamically adjust weights of different slots during training. Experimental results show that our approach reaches 52.68% and 58.55% joint accuracy on MultiWOZ 2.0 and MultiWOZ 2.1 datasets respectively and achieves new state-of-the-art performance with considerable improvements (+1.24% and +5.98%). Yong Shan, Zekang Li, Jinchao Zhang 0001, Fandong Meng, Yang Feng 0004, Cheng Niu, Jie Zhou 0016 |
ACL | 5 |
| 2020 | CDL: Curriculum Dual Learning for Emotion-Controllable Response GenerationabstractEmotion-controllable response generation is an attractive and valuable task that aims to make open-domain conversations more empathetic and engaging. Existing methods mainly enhance the emotion expression by adding regularization terms to standard cross-entropy loss and thus influence the training process. However, due to the lack of further consideration of content consistency, the common problem of response generation tasks, safe response, is intensified. Besides, query emotions that can help model the relationship between query and response are simply ignored in previous models, which would further hurt the coherence. To alleviate these problems, we propose a novel framework named Curriculum Dual Learning (CDL) which extends the emotion-controllable response generation to a dual task to generate emotional responses and emotional queries alternatively. CDL utilizes two rewards focusing on emotion and content to improve the duality. Additionally, it applies curriculum learning to gradually generate high-quality responses based on the difficulties of expressing various emotions. Experimental results show that CDL significantly outperforms the baselines in terms of coherence, diversity, and relation to emotion factors. Lei Shen 0001, Yang Feng 0004 |
ACL | 2 |
| 2020 | Investigating Catastrophic Forgetting During Continual Training for Neural Machine TranslationabstractNeural machine translation (NMT) models usually suffer from catastrophic forgetting during continual training where the models tend to gradually forget previously learned knowledge and swing to fit the newly added data which may have a different distribution, e.g. a different domain. Although many methods have been proposed to solve this problem, we cannot get to know what causes this phenomenon yet. Under the background of domain adaptation, we investigate the cause of catastrophic forgetting from the perspectives of modules and parameters (neurons). The investigation on the modules of the NMT model shows that some modules have tight relation with the general-domain knowledge while some other modules are more essential in the domain adaptation. And the investigation on the parameters shows that some parameters are important for both the general-domain and in-domain translation and the great change of them during continual training brings about the performance decline in general-domain. We conducted experiments across different language pairs and domains to ensure the validity and reliability of our findings. Shuhao Gu, Yang Feng 0004 |
COLING | 2 |
| 2020 | Token-level Adaptive Training for Neural Machine TranslationabstractThere exists a token imbalance phenomenon in natural language as different tokens appear with different frequencies, which leads to different learning difficulties for tokens in Neural Machine Translation (NMT).The vanilla NMT model usually adopts trivial equal-weighted objectives for target tokens with different frequencies and tends to generate more high-frequency tokens and less lowfrequency tokens compared with the golden token distribution.However, low-frequency tokens may carry critical semantic information that will affect the translation quality once they are neglected.In this paper, we explored target token-level adaptive objectives based on token frequencies to assign appropriate weights for each target token during training.We aimed that those meaningful but relatively low-frequency words could be assigned with larger weights in objectives to encourage the model to pay more attention to these tokens.Our method yields consistent improvements in translation quality on ZH-EN, EN-RO, and EN-DE translation tasks, especially on sentences that contain more low-frequency tokens where we can get 1.68, 1.02, and 0.52 BLEU increases compared with baseline, respectively.Further analyses show that our method can also improve the lexical diversity of translation. Shuhao Gu, Jinchao Zhang 0001, Fandong Meng, Yang Feng 0004, Wanying Xie, Jie Zhou 0016, Dong Yu 0003 |
EMNLP (1) | 4 |
| 2020 | Generating Diverse Translation from Model Distribution with DropoutabstractDespite the improvement of translation quality, neural machine translation (NMT) often suffers from the lack of diversity in its generation.In this paper, we propose to generate diverse translations by deriving a large number of possible models with Bayesian modelling and sampling models from them for inference.The possible models are obtained by applying concrete dropout to the NMT model and each of them has specific confidence for its prediction, which corresponds to a posterior model distribution under specific training data in the principle of Bayesian modeling.With variational inference, the posterior model distribution can be approximated with a variational distribution, from which the final models for inference are sampled.We conducted experiments on Chinese-English and English-German translation tasks and the results shows that our method makes a better trade-off between diversity and accuracy. Xuanfu Wu, Yang Feng 0004, Chenze Shao |
EMNLP (1) | 2 |
| 2020 | Bridging the Gap between Training and Inference for Neural Machine Translation (Extended Abstract)abstractNeural Machine Translation (NMT) generates target words sequentially in the way of predicting the next word conditioned on the context words. At training time, it predicts with the ground truth words as context while at inference it has to generate the entire sequence from scratch. This discrepancy of the fed context leads to error accumulation among the translation. Furthermore, word-level training requires strict matching between the generated sequence and the ground truth sequence which leads to overcorrection over different but reasonable translations. In this paper, we address these issues by sampling context words not only from the ground truth sequence but also from the predicted sequence during training. Experimental results on NIST Chinese->English and WMT2014 English->German translation tasks demonstrate that our method can achieve significant improvements on multiple data sets compared to strong baselines. Wen Zhang 0009, Yang Feng 0004, Qun Liu 0001 |
IJCAI | 2 |
| 2019 | Incremental Transformer with Deliberation Decoder for Document Grounded ConversationsabstractDocument Grounded Conversations is a task to generate dialogue responses when chatting about the content of a given document. Obviously, document knowledge plays a critical role in Document Grounded Conversations, while existing dialogue models do not exploit this kind of knowledge effectively enough. In this paper, we propose a novel Transformer-based architecture for multi-turn document grounded conversations. In particular, we devise an Incremental Transformer to encode multi-turn utterances along with knowledge in related documents. Motivated by the human cognitive process, we design a two-pass decoder (Deliberation Decoder) to improve context coherence and knowledge correctness. Our empirical study on a real-world Document Grounded Dataset proves that responses generated by our model significantly outperform competitive baselines on both context coherence and knowledge relevance. Zekang Li, Cheng Niu, Fandong Meng, Yang Feng 0004, Jie Zhou 0016 |
ACL (1) | 4 |
| 2019 | Retrieving Sequential Information for Non-Autoregressive Neural Machine TranslationabstractNon-Autoregressive Transformer (NAT) aims to accelerate the Transformer model through discarding the autoregressive mechanism and generating target words independently, which fails to exploit the target sequential information.Over-translation and under-translation errors often occur for the above reason, especially in the long sentence translation scenario.In this paper, we propose two approaches to retrieve the target sequential information for NAT to enhance its translation ability while preserving the fast-decoding property.Firstly, we propose a sequence-level training method based on a novel reinforcement algorithm for NAT (Reinforce-NAT) to reduce the variance and stabilize the training procedure.Secondly, we propose an innovative Transformer decoder named FS-decoder to fuse the target sequential information into the top layer of the decoder.Experimental results on three translation tasks show that the Reinforce-NAT surpasses the baseline NAT system by a significant margin on BLEU without decelerating the decoding speed and the FS-decoder achieves comparable translation performance to the autoregressive Transformer with considerable speedup. Chenze Shao, Yang Feng 0004, Jinchao Zhang 0001, Fandong Meng, Xilin Chen 0001, Jie Zhou 0016 |
ACL (1) | 2 |
| 2019 | Bridging the Gap between Training and Inference for Neural Machine TranslationabstractNeural Machine Translation (NMT) generates target words sequentially in the way of predicting the next word conditioned on the context words.At training time, it predicts with the ground truth words as context while at inference it has to generate the entire sequence from scratch.This discrepancy of the fed context leads to error accumulation among the way.Furthermore, word-level training requires strict matching between the generated sequence and the ground truth sequence which leads to overcorrection over different but reasonable translations.In this paper, we address these issues by sampling context words not only from the ground truth sequence but also from the predicted sequence by the model during training, where the predicted sequence is selected with a sentence-level optimum.Experiment results on Chinese→English and WMT'14 English→German translation tasks demonstrate that our approach can achieve significant improvements on multiple datasets. Wen Zhang 0009, Yang Feng 0004, Fandong Meng, Di You, Qun Liu 0001 |
ACL (1) | 2 |
| 2019 | Enhancing Context Modeling with a Query-Guided Capsule Network for Document-level TranslationabstractZhengxin Yang, Jinchao Zhang, Fandong Meng, Shuhao Gu, Yang Feng, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhengxin Yang, Jinchao Zhang 0001, Fandong Meng, Shuhao Gu, Yang Feng 0004, Jie Zhou 0016 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Improving Multi-head Attention with Capsule Networks
Shuhao Gu, Yang Feng 0004 |
NLPCC (1) | 2 |
| 2019 | Neural Machine Translation with Bilingual History Involved Attention
Haiyang Xue, Yang Feng 0004, Di You, Wen Zhang 0009 |
NLPCC (2) | 2 |
| 2019 | A New Algorithm for Component Decomposition and Type Recognition of Tibetan Syllable
Shugen Wang, Yanru Wu, Meijing Guan, Yang Feng 0004 |
NLPCC (2) | 5 |
| 2018 | Knowledge Diffusion for Neural Dialogue GenerationabstractEnd-to-end neural dialogue generation has shown promising results recently, but it does not employ knowledge to guide the generation and hence tends to generate short, general, and meaningless responses.In this paper, we propose a neural knowledge diffusion (NKD) model to introduce knowledge into dialogue generation.This method can not only match the relevant facts for the input utterance but diffuse them to similar entities.With the help of facts matching and entity diffusion, the neural dialogue generation is augmented with the ability of convergent and divergent thinking over the knowledge base.Our empirical study on a real-world dataset proves that our model is capable of generating meaningful, diverse and natural responses for both factoid-questions and knowledge grounded chi-chats.The experiment results also show that our model outperforms competitive baseline models significantly. Shuman Liu, Hongshen Chen, Zhaochun Ren, Yang Feng 0004, Qun Liu 0001, Dawei Yin 0001 |
ACL (1) | 4 |
| 2018 | Refining Source Representations with Relation Networks for Neural Machine TranslationabstractAlthough neural machine translation with the encoder-decoder framework has achieved great success recently, it still suffers drawbacks of forgetting distant information, which is an inherent disadvantage of recurrent neural network structure, and disregarding relationship between source words during encoding step. Whereas in practice, the former information and relationship are often useful in current step. We target on solving these problems and thus introduce relation networks to learn better representations of the source. The relation networks are able to facilitate memorization capability of recurrent neural network via associating source words with each other, this would also help retain their relationships. Then the source representations and all the relations are fed into the attention component together while decoding, with the main encoder-decoder framework unchanged. Experiments on several datasets show that our method can improve the translation performance significantly over the conventional encoder-decoder model and even outperform the approach involving supervised syntactic knowledge. Wen Zhang 0009, Yang Feng 0004, Qun Liu 0001 |
COLING | 3 |
| 2018 | Greedy Search with Probabilistic N-gram Matching for Neural Machine TranslationabstractNeural machine translation (NMT) models are usually trained with the word-level loss using the teacher forcing algorithm, which not only evaluates the translation improperly but also suffers from exposure bias.Sequence-level training under the reinforcement framework can mitigate the problems of the word-level loss, but its performance is unstable due to the high variance of the gradient estimation.On these grounds, we present a method with a differentiable sequence-level training objective based on probabilistic n-gram matching which can avoid the reinforcement framework.In addition, this method performs greedy search in the training which uses the predicted words as context just as at inference to alleviate the problem of exposure bias.Experiment results on the NIST Chinese-to-English translation tasks show that our method significantly outperforms the reinforcement-based algorithms and achieves an improvement of 1.5 BLEU points on average over a strong baseline system. Chenze Shao, Xilin Chen 0001, Yang Feng 0004 |
EMNLP | 3 |
| 2018 | Speeding Up Neural Machine Translation Decoding by Cube PruningabstractAlthough neural machine translation has achieved promising results, it suffers from slow translation speed.The direct consequence is that a trade-off has to be made between translation quality and speed, thus its performance can not come into full play.We apply cube pruning, a popular technique to speed up dynamic programming, into neural machine translation to speed up the translation.To construct the equivalence class, similar target hidden states are combined, leading to less RNN expansion operations on the target side and less softmax operations over the large target vocabulary.The experiments show that, at the same or even better translation quality, our method can translate faster compared with naive beam search by 3.3× on GPUs and 3.5× on CPUs. Wen Zhang 0009, Liang Huang 0001, Yang Feng 0004, Lei Shen 0001, Qun Liu 0001 |
EMNLP | 3 |
| 2017 | Flexible and Creative Chinese Poetry Generation Using Neural MemoryabstractJiyuan Zhang, Yang Feng, Dong Wang, Yang Wang, Andrew Abel, Shiyue Zhang, Andi Zhang. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Jiyuan Zhang 0001, Yang Feng 0004, Dong Wang 0013, Andrew Abel, Shiyue Zhang 0001, Andi Zhang 0002 |
ACL (1) | 2 |
| 2017 | Memory-augmented Neural Machine TranslationabstractNeural machine translation (NMT) has achieved notable success in recent times, however it is also widely recognized that this approach has limitations with handling infrequent words and word pairs.This paper presents a novel memoryaugmented NMT (M-NMT) architecture, which stores knowledge about how words (usually infrequently encountered ones) should be translated in a memory and then utilizes them to assist the neural model.We use this memory mechanism to combine the knowledge learned from a conventional statistical machine translation system and the rules learned by an NMT system, and also propose a solution for out-of-vocabulary (OOV) words based on this framework.Our experiments on two Chinese-English translation tasks demonstrated that the M-NMT architecture outperformed the NMT baseline by 9.0 and 2.7 BLEU points on the two tasks, respectively.Additionally, we found this architecture resulted in a much more effective OOV treatment compared to competitive methods. Yang Feng 0004, Shiyue Zhang 0001, Andi Zhang 0002, Dong Wang 0013, Andrew Abel |
EMNLP | 1 |
| 2017 | Memory visualization for gated recurrent neural networks in speech recognitionabstractRecurrent neural networks (RNNs) have shown clear superiority in sequence modeling, particularly the ones with gated units, such as long short-term memory (LSTM) and gated recurrent unit (GRU). However, the dynamic properties behind the remarkable performance remain unclear in many applications, e.g., automatic speech recognition (ASR). This paper employs visualization techniques to study the behavior of LSTM and GRU when performing speech recognition tasks. Our experiments show some interesting patterns in the gated memory, and some of them have inspired simple yet effective modifications on the network structure. We report two of such modifications: (1) lazy cell update in LSTM, and (2) shortcut connections for residual learning. Both modifications lead to more comprehensible and powerful networks. Zhiyuan Tang, Ying Shi 0001, Dong Wang 0013, Yang Feng 0004, Shiyue Zhang 0001 |
ICASSP | 4 |
| 2014 | Factored Markov Translation with Robust ModelingabstractPhrase-based translation models usually memorize local translation literally and make independent assumption between phrases which makes it neither generalize well on unseen data nor model sentence-level effects between phrases. In this pa-per we present a new method to model correlations between phrases as a Markov model and meanwhile employ a robust smoothing strategy to provide better gen-eralization. This method defines a re-cursive estimation process and backs off in parallel paths to infer richer structures. Our evaluation shows an 1.1–3.2 % BLEU improvement over competitive baselines for Chinese-English and Arabic-English translation. 1 Yang Feng 0004, Trevor Cohn, Xinkai Du |
CoNLL | 1 |
| 2013 | A Markov Model of Machine Translation using Non-parametric Bayesian Inference
Yang Feng 0004, Trevor Cohn |
ACL (1) | 1 |
| 2012 | Hierarchical Chunk-to-String Translation
Yang Feng 0004, Dongdong Zhang 0001, Mu Li 0001, Qun Liu 0001 |
ACL (1) | 1 |
| 2012 | Left-to-Right Tree-to-String Decoding with Prediction
Yang Feng 0004, Yang Liu 0005, Qun Liu 0001, Trevor Cohn |
EMNLP-CoNLL | 1 |
| 2009 | Joint Decoding with Multiple Translation Models
Yang Liu 0005, Haitao Mi, Yang Feng 0004, Qun Liu 0001 |
ACL/IJCNLP | 3 |
| 2009 | Lattice-based System Combination for Statistical Machine Translation
Yang Feng 0004, Yang Liu 0005, Haitao Mi, Qun Liu 0001, Yajuan Lü |
EMNLP | 1 |