Lemao Liu

dblp:41/10887 · DBLP profile ↗
← Back
68ranked-venue papers
12as first author
40since 2021 · last 2025
0000-0003-3804-5768ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 63 · 12 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task
abstract
Junjie Wu, Mo Yu, Lemao Liu, Dit-Yan Yeung, Jie Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Junjie Wu 0007, Mo Yu, Lemao Liu, Dit-Yan Yeung, Jie Zhou 0016
NAACL (Long Papers)3
2025 The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding
abstract
Mo Yu, Lemao Liu, Junjie Wu, Tsz Ting Chung, Shunchi Zhang, Jiangnan Li, Dit-Yan Yeung, Jie Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Mo Yu, Lemao Liu, Junjie Wu 0007, Tsz Ting Chung, Shunchi Zhang, Dit-Yan Yeung, Jie Zhou 0016
NAACL (Long Papers)2
2025 TF-Attack: Transferable and fast adversarial attacks on large language models
Kehai Chen, Lemao Liu, Xuefeng Bai 0001, Yang Xiang 0003, Min Zhang 0005
Knowl. Based Syst.3
2024 Advancement in Graph Understanding: A Multimodal Benchmark and Fine-Tuning of Vision-Language Models
abstract
Qihang Ai, Jiafan Li, Jincheng Dai, Jianwu Zhou, Lemao Liu, Haiyun Jiang, Shuming Shi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Qihang Ai, Jiafan Li, Jincheng Dai, Jianwu Zhou, Lemao Liu, Haiyun Jiang, Shuming Shi 0001
ACL (1)5
2024 Context Consistency between Training and Inference in Simultaneous Machine Translation
abstract
Simultaneous Machine Translation (SiMT) aims to yield a real-time partial translation with a monotonically growing source-side context.However, there is a counterintuitive phenomenon about the context usage between training and inference: e.g., in wait-k inference, model consistently trained with wait-k is much worse than that model inconsistently trained with wait-k ′ (k ′ ̸ = k) in terms of translation quality.To this end, we first investigate the underlying reasons behind this phenomenon and uncover the following two factors: 1) the limited correlation between translation quality and training loss; 2) exposure bias between training and inference.Based on both reasons, we then propose an effective training approach called context consistency training accordingly, which encourages consistent context usage between training and inference by optimizing translation quality and latency as bi-objectives and exposing the predictions to the model during the training.The experiments on three language pairs demonstrate that our SiMT system encouraging context consistency outperforms existing SiMT systems with context inconsistency for the first time.1
Meizhi Zhong, Lemao Liu, Kehai Chen, Min Zhang 0005
ACL (1)2
2024 Rethinking the Evaluation of In-Context Learning for LLMs
abstract
In-context learning (ICL) has demonstrated excellent performance across various downstream NLP tasks, especially when synergized with powerful large language models (LLMs).Existing studies evaluate ICL methods primarily based on downstream task performance.This evaluation protocol overlooks the significant cost associated with the demonstration configuration process, i.e., tuning the demonstration as the ICL prompt.However, in this work, we point out that the evaluation protocol leads to unfair comparisons and potentially biased evaluation, because we surprisingly find the correlation between the configuration costs and task performance.Then we call for a twodimensional evaluation paradigm that considers both of these aspects, facilitating a fairer comparison.Finally, based on our empirical finding that the optimized demonstration on one language model generalizes across language models of different sizes, we introduce a simple yet efficient strategy that can be applied to any ICL method as a plugin, yielding a better trade-off between the two dimensions according to the proposed evaluation paradigm.
Guoxin Yu, Lemao Liu, Mo Yu, Xiang Ao 0001
EMNLP2
2024 Rethinking Targeted Adversarial Attacks for Neural Machine Translation
abstract
Targeted adversarial attacks are widely used to evaluate the robustness of neural machine translation systems. Unfortunately, this paper first identifies a critical issue in the existing settings of NMT targeted adversarial attacks, where their attacking results are largely overestimated. To this end, this paper presents a new setting for NMT targeted adversarial attacks that could lead to reliable attacking results. Under the new setting, it then proposes a Targeted Word Gradient adversarial Attack (TWGA) method to craft adversarial examples. Experimental results demonstrate that our proposed setting could provide faithful attacking results for targeted adversarial attacks on NMT systems, and the proposed TWGA method can effectively attack such victim NMT systems. In-depth analyses on a large-scale dataset further illustrate some valuable findings.1Our code and data are available at https://github.com/wujunjie1998/TWGA.
Junjie Wu 0007, Lemao Liu, Wei Bi, Dit-Yan Yeung
ICASSP2
2024 Hint-Enhanced In-Context Learning Wakes Large Language Models Up For Knowledge-Intensive Tasks
abstract
In-context learning (ICL) ability has emerged with the increasing scale of large language models (LLMs), enabling them to learn input-label mappings from demonstrations and perform well on downstream tasks. However, under the standard ICL setting, LLMs may sometimes neglect query-related information in demonstrations, leading to incorrect predictions. To address this limitation, we propose a new paradigm called Hint-enhanced In-Context Learning (HICL) to explore the power of ICL in open-domain question answering, an important form in knowledge-intensive tasks. HICL leverages LLMs’ reasoning ability to extract query-related knowledge from demonstrations, then concatenates the knowledge to prompt LLMs in a more explicit way. Furthermore, we track the source of this knowledge to identify specific examples, and introduce a Hint-related Example Retriever (HER) to select informative examples for enhanced demonstrations. We evaluate HICL with HER on 3 open-domain QA benchmarks, and observe average performance gains of 2.89 EM score and 2.52 F1 score on gpt-3.5-turbo, 7.62 EM score and 7.27 F1 score on LLaMA-2-Chat-7B compared with standard setting.
Qingyan Guo, Xinzhe Ni, Chufan Shi, Lemao Liu, Haiyun Jiang, Yujiu Yang 0001
ICASSP5
2024 The Reasonableness Behind Unreasonable Translation Capability of Large Language Model
abstract
Multilingual large language models trained on non-parallel data yield impressive translation capabilities. Existing studies demonstrate that incidental sentence-level bilingualism within pre-training data contributes to the LLM's translation abilities. However, it has also been observed that LLM's translation capabilities persist even when incidental sentence-level bilingualism are excluded from the training corpus. In this study, we comprehensively investigate the unreasonable effectiveness and the underlying mechanism for LLM's translation abilities, specifically addressing the question why large language models learn to translate without parallel data, using the BLOOM model series as a representative example. Through extensive experiments, our findings suggest the existence of unintentional bilingualism in the pre-training corpus, especially word alignment data significantly contributes to the large language model's acquisition of translation ability. Moreover, the translation signal derived from word alignment data is comparable to that from sentence-level bilingualism. Additionally, we study the effects of monolingual data and parameter-sharing in assisting large language model to learn to translate. Together, these findings present another piece of the broader puzzle of trying to understand how large language models acquire translation capability.
Tingchen Fu, Lemao Liu, Deng Cai 0002, Guoping Huang, Shuming Shi 0001, Rui Yan 0001
ICLR2
2024 An Energy-based Model for Word-level AutoCompletion in Computer-aided Translation
abstract
Abstract Word-level AutoCompletion (WLAC) is a rewarding yet challenging task in Computer-aided Translation. Existing work addresses this task through a classification model based on a neural network that maps the hidden vector of the input context into its corresponding label (i.e., the candidate target word is treated as a label). Since the context hidden vector itself does not take the label into account and it is projected to the label through a linear classifier, the model cannot sufficiently leverage valuable information from the source sentence as verified in our experiments, which eventually hinders its overall performance. To alleviate this issue, this work proposes an energy-based model for WLAC, which enables the context hidden vector to capture crucial information from the source sentence. Unfortunately, training and inference suffer from efficiency and effectiveness challenges, therefore we employ three simple yet effective strategies to put our model into practice. Experiments on four standard benchmarks demonstrate that our reranking-based approach achieves substantial improvements (about 6.07%) over the previous state-of-the-art model. Further analyses show that each strategy of our approach contributes to the final performance.1
Cheng Yang 0007, Guoping Huang, Mo Yu, Zhirui Zhang, Siheng Li, Shuming Shi 0001, Yujiu Yang 0001, Lemao Liu
Trans. Assoc. Comput. Linguistics9
2024 Datastore Distillation for Nearest Neighbor Machine Translation
abstract
Nearest neighbor machine translation (i.e.,$k$NN-MT) is a promising approach to enhance translation quality by equipping pre-trained neural machine translation (NMT) models with the nearest neighbor retrieval. Despite its great success,$k$NN-MT typically requires ample space to store its token-level datastore, causing$k$NN-MT to be less practical in edge devices or online scenarios. In this paper, inspired by the concept of knowledge distillation, we provide a new perspective to ease the storage overhead by datastore distillation, which is formalized as a constrained optimization problem. We further design a novel model-agnostic iterative nearest neighbor merging method for the datastore distillation problem to obtain an effective and efficient solution. Experiments on three benchmark datasets indicate that our approach not only reduces the volume of the datastore by up to 50% without significant performance degradation, but also outperforms other baselines by a large margin at the same compression rate. Another experiment conducted on WikiText-103 further demonstrates the effectiveness of our method in the language model task.
Yuhan Dai, Zhirui Zhang, Yichao Du, Shengcai Liu, Lemao Liu, Tong Xu 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 On the Compositional Generalization in Versatile Open-domain Dialogue
abstract
Previous research has demonstrated the potential of multi-task learning to foster a conversational agent's ability to acquire a variety of skills.However, these approaches either suffer from interference among different datasets (also known as negative transfer), or fail to effectively reuse knowledge and skills learned from other datasets.In contrast to previous works, we develop a sparsely activated modular network: (1) We propose a wellrounded set of operators and instantiate each operator with an independent module; (2) We formulate dialogue generation as the execution of a generated programme which recursively composes and assembles modules.Extensive experiments on 9 datasets verify the efficacy of our methods through automatic evaluation and human evaluation.Notably, our model outperforms state-of-the-art supervised approaches on 4 datasets with only 10% training data thanks to the modular architecture and multi-task learning.1 † Tingchen Fu and Xueliang Zhao contribute equally to this work.
Tingchen Fu, Xueliang Zhao, Lemao Liu, Rui Yan 0001
ACL (1)3
2023 Rethinking Word-Level Auto-Completion in Computer-Aided Translation
abstract
Word-Level Auto-Completion (WLAC) plays a crucial role in Computer-Assisted Translation.It aims at providing word-level autocompletion suggestions for human translators.While previous studies have primarily focused on designing complex model architectures, this paper takes a different perspective by rethinking the fundamental question: what kind of words are good auto-completions?We introduce a measurable criterion to answer this question and discover that existing WLAC models often fail to meet this criterion.Building upon this observation, we propose an effective approach to enhance WLAC performance by promoting adherence to the criterion.Notably, the proposed approach is general and can be applied to various encoder-based architectures.Through extensive experiments, we demonstrate that our approach outperforms the top-performing system submitted to the WLAC shared tasks in WMT2022, while utilizing significantly smaller model sizes ¶ .
Lemao Liu, Guoping Huang, Zhirui Zhang, Shuming Shi 0001, Rui Wang 0015
EMNLP2
2023 Nearest Neighbor Machine Translation is Meta-Optimizer on Output Projection Layer
abstract
Nearest Neighbor Machine Translation (kNN-MT) has achieved great success in domain adaptation tasks by integrating pre-trained Neural Machine Translation (NMT) models with domain-specific token-level retrieval.However, the reasons underlying its success have not been thoroughly investigated.In this paper, we comprehensively analyze kNN-MT through theoretical and empirical studies.Initially, we provide new insights into the working mechanism of kNN-MT as an efficient technique to implicitly execute gradient descent on the output projection layer of NMT, indicating that it is a specific case of model fine-tuning.Subsequently, we conduct multi-domain experiments and word-level analysis to examine the differences in performance between kNN-MT and entire-model fine-tuning.Our findings suggest that: (i) Incorporating kNN-MT with adapters yields comparable translation performance to fine-tuning on in-domain test sets, while achieving better performance on out-of-domain test sets; (ii) Fine-tuning significantly outperforms kNN-MT on the recall of in-domain low-frequency words, but this gap could be bridged by optimizing the context representations with additional adapter layers. 1
Zhirui Zhang, Yichao Du, Lemao Liu, Rui Wang 0015
EMNLP4
2023 IMTLab: An Open-Source Platform for Building, Evaluating, and Diagnosing Interactive Machine Translation Systems
abstract
Xu Huang, Zhirui Zhang, Ruize Gao, Yichao Du, Lemao Liu, Guoping Huang, Shuming Shi, Jiajun Chen, Shujian Huang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Zhirui Zhang, Yichao Du, Lemao Liu, Guoping Huang, Shuming Shi 0001, Jiajun Chen 0001, Shujian Huang
EMNLP5
2023 SimCSE++: Improving Contrastive Learning for Sentence Embeddings from Two Perspectives
abstract
This paper improves contrastive learning for sentence embeddings from two perspectives: handling dropout noise and addressing feature corruption.Specifically, for the first perspective, we identify that the dropout noise from negative pairs affects the model's performance.Therefore, we propose a simple yet effective method to deal with such type of noise.Secondly, we pinpoint the rank bottleneck of current solutions to feature corruption and propose a dimension-wise contrastive learning objective to address this issue.Both proposed methods are generic and can be applied to any contrastive learning based models for sentence embeddings.Experimental results on standard benchmarks demonstrate that combining both proposed methods leads to a gain of 1.8 points compared to the strong baseline SimCSE configured with BERT base.Furthermore, applying the proposed method to DiffCSE, another strong contrastive learning based baseline, results in a gain of 1.4 points.* The source code is available at https://github.com/ Jiahao004/SimCSE-plus-plus.
Jiahao Xu 0001, Wei Shao 0009, Lihui Chen 0001, Lemao Liu
EMNLP4
2023 A Simple Yet Effective Approach to Structured Knowledge Distillation
abstract
Structured prediction models aim at solving tasks where the output is a complex structure, rather than a single variable. Performing knowledge distillation for such problems is non- trivial due to their exponentially large output space. Previous works address this problem by developing particular distillation strategies (e.g., dynamic programming) that are both complicated and of low run-time efficiency. In this work, we propose an approach that is much simpler in its formulation, far more efficient for training than existing methods, and even performs better than our baselines. Specifically, we transfer the knowledge from a teacher model to its student by locally matching their computations on all internal structures rather than the final outputs. In this manner, we avoid time-consuming techniques like Monte Carlo Sampling for decoding output structures, permitting parallel computation and efficient training. Besides, we show that it encourages the student model to better mimic the internal behavior of the teacher model. Experiments on two structured prediction tasks demonstrate that our approach not only halves the time cost, but also outperforms previous methods on two widely adopted benchmark datasets.1 2
Wenye Lin, Yangming Li, Lemao Liu, Shuming Shi 0001, Hai-Tao Zheng 0002
ICASSP3
2023 Federated Nearest Neighbor Machine Translation
Yichao Du, Zhirui Zhang, Bingzhe Wu, Lemao Liu, Tong Xu 0001, Enhong Chen
ICLR4
2023 Lift Yourself Up: Retrieval-augmented Text Generation with Self-Memory
abstract
With direct access to human-written reference as memory, retrieval-augmented generation has achieved much progress in a wide range of text generation tasks. Since better memory would typically prompt better generation (we define this as primal problem). The traditional approach for memory retrieval involves selecting memory that exhibits the highest similarity to the input. However, this method is constrained by the quality of the fixed corpus from which memory is retrieved. In this paper, by exploring the duality of the primal problem: better generation also prompts better memory, we propose a novel framework, selfmem, which addresses this limitation by iteratively employing a retrieval-augmented generator to create an unbounded memory pool and using a memory selector to choose one output as memory for the subsequent generation round. This enables the model to leverage its own output, referred to as self-memory, for improved generation. We evaluate the effectiveness of selfmem on three distinct text generation tasks: neural machine translation, abstractive text summarization, and dialogue generation, under two generation paradigms: fine-tuned small model and few-shot LLM. Our approach achieves state-of-the-art results in four directions in JRC-Acquis translation dataset, 50.3 ROUGE-1 in XSum, and 62.9 ROUGE-1 in BigPatent, demonstrating the potential of self-memory in enhancing retrieval-augmented generation models. Furthermore, we conduct thorough analyses of each component in the selfmem framework to identify current system bottlenecks and provide insights for future research.
Xin Cheng 0002, Xiuying Chen, Lemao Liu, Dongyan Zhao 0001, Rui Yan 0001
NeurIPS4
2023 Repetition In Repetition Out: Towards Understanding Neural Text Degeneration from the Data Perspective
abstract
There are a number of diverging hypotheses about the neural text degeneration problem, i.e., generating repetitive and dull loops, which makes this problem both interesting and confusing. In this work, we aim to advance our understanding by presenting a straightforward and fundamental explanation from the data perspective. Our preliminary investigation reveals a strong correlation between the degeneration issue and the presence of repetitions in training data. Subsequent experiments also demonstrate that by selectively dropping out the attention to repetitive words in training data, degeneration can be significantly minimized. Furthermore, our empirical analysis illustrates that prior works addressing the degeneration issue from various standpoints, such as the high-inflow words, the likelihood objective, and the self-reinforcement phenomenon, can be interpreted by one simple explanation. That is, penalizing the repetitions in training data is a common and fundamental factor for their effectiveness. Moreover, our experiments reveal that penalizing the repetitions in training data remains critical even when considering larger model sizes and instruction tuning.
Tian Lan 0003, Deng Cai 0002, Lemao Liu, Nigel Collier, Taro Watanabe, Yixuan Su
NeurIPS5
2023 Fairness-guided Few-shot Prompting for Large Language Models
abstract
Large language models have demonstrated surprising ability to perform in-context learning, i.e., these models can be directly applied to solve numerous downstream tasks by conditioning on a prompt constructed by a few input-output examples. However, prior research has shown that in-context learning can suffer from high instability due to variations in training examples, example order, and prompt formats. Therefore, the construction of an appropriate prompt is essential for improving the performance of in-context learning. In this paper, we revisit this problem from the view of predictive bias. Specifically, we introduce a metric to evaluate the predictive bias of a fixed prompt against labels or a given attributes. Then we empirically show that prompts with higher bias always lead to unsatisfactory predictive quality. Based on this observation, we propose a novel search strategy based on the greedy search to identify the near-optimal prompt for improving the performance of in-context learning. We perform comprehensive experiments with state-of-the-art mainstream models such as GPT-3 on various downstream tasks. Our results indicate that our method can enhance the model's in-context learning performance in an effective and interpretable manner.
Huan Ma 0006, Changqing Zhang 0002, Yatao Bian, Lemao Liu, Zhirui Zhang, Peilin Zhao, Shu Zhang 0013, Huazhu Fu, Qinghua Hu, Bingzhe Wu
NeurIPS4
2023 Discourse-Aware Graph Networks for Textual Logical Reasoning
abstract
Textual logical reasoning, especially question-answering (QA) tasks with logical reasoning, requires awareness of particular logical structures. The passage-level logical relations represent entailment or contradiction between propositional units (e.g., a concluding sentence). However, such structures are unexplored as current QA systems focus on entity-based relations. In this work, we propose logic structural-constraint modeling to solve the logical reasoning QA and introduce discourse-aware graph networks (DAGNs). The networks first construct logic graphs leveraging in-line discourse connectives and generic logic theories, then learn logic representations by end-to-end evolving the logic relations with an edge-reasoning mechanism and updating the graph features. This pipeline is applied to a general encoder, whose fundamental features are joined with the high-level logic features for answer prediction. Experiments on three textual logical reasoning datasets demonstrate the reasonability of the logical structures built in DAGNs and the effectiveness of the learned logic features. Moreover, zero-shot transfer results show the features' generality to unseen logical texts.
Yinya Huang, Lemao Liu, Kun Xu 0005, Liang Lin 0004, Xiaodan Liang
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Learning from Sibling Mentions with Scalable Graph Inference in Fine-Grained Entity Typing
abstract
Yi Chen, Jiayang Cheng, Haiyun Jiang, Lemao Liu, Haisong Zhang, Shuming Shi, Ruifeng Xu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yi Chen 0019, Cheng Jiayang, Haiyun Jiang, Lemao Liu, Haisong Zhang, Shuming Shi 0001, Ruifeng Xu 0001
ACL (1)4
2022 Rethinking Negative Sampling for Handling Missing Entity Annotations
abstract
Negative sampling is highly effective in handling missing annotations for named entity recognition (NER).One of our contributions is an analysis on how it makes sense through introducing two insightful concepts: missampling and uncertainty.Empirical studies show low missampling rate and high uncertainty are both essential for achieving promising performances with negative sampling.Based on the sparsity of named entities, we also theoretically derive a lower bound for the probability of zero missampling rate, which is only relevant to sentence length.The other contribution is an adaptive and weighted sampling distribution that further improves negative sampling via our former analysis.Experiments on synthetic datasets and well-annotated datasets (e.g., CoNLL-2003) show that our proposed approach benefits negative sampling in terms of F1 score and loss convergence.Besides, models with improved negative sampling have achieved new state-of-the-art results on realworld datasets (e.g., EC).
Yangming Li, Lemao Liu, Shuming Shi 0001
ACL (1)2
2022 BiTIIMT: A Bilingual Text-infilling Method for Interactive Machine Translation
abstract
Yanling Xiao, Lemao Liu, Guoping Huang, Qu Cui, Shujian Huang, Shuming Shi, Jiajun Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yanling Xiao, Lemao Liu, Guoping Huang, Qu Cui, Shujian Huang, Shuming Shi 0001, Jiajun Chen 0001
ACL (1)2
2022 Neural Machine Translation with Contrastive Translation Memories
abstract
Retrieval-augmented Neural Machine Translation models have been successful in many translation scenarios.Different from previous works that make use of mutually similar but redundant translation memories (TMs), we propose a new retrieval-augmented NMT to model contrastively retrieved translation memories that are holistically similar to the source sentence while individually contrastive to each other providing maximal information gains in three phases.First, in TM retrieval phase, we adopt a contrastive retrieval algorithm to avoid redundancy and uninformativeness of similar translation pieces.Second, in memory encoding stage, given a set of TMs we propose a novel Hierarchical Group Attention module to gather both local context of each TM and global context of the whole TM set.Finally, in training phase, a Multi-TM contrastive learning objective is introduced to learn salient feature of each TM with respect to target sentence.Experimental results show that our framework obtains improvements over strong baselines on the benchmark datasets.
Xin Cheng 0002, Shen Gao, Lemao Liu, Dongyan Zhao 0001, Rui Yan 0001
EMNLP3
2022 On the Evaluation Metrics for Paraphrase Generation
abstract
In this paper we revisit automatic metrics for paraphrase evaluation and obtain two findings that disobey conventional wisdom:(1) Reference-free metrics achieve better performance than their reference-based counterparts.(2) Most commonly used metrics do not align well with human annotation.Underlying reasons behind the above findings are explored through additional experiments and in-depth analyses.Based on the experiments and analyses, we propose ParaScore, a new evaluation metric for paraphrase generation.It possesses the merits of referencebased and reference-free metrics and explicitly models lexical divergence.Based on our analysis and improvements, our proposed reference-based outperforms than referencefree metrics.Experimental results demonstrate that ParaScore significantly outperforms existing metrics.Our codes and toolkit are released in https://github.com/ shadowkiller33/ParaScore.
Lingfeng Shen, Lemao Liu, Haiyun Jiang, Shuming Shi 0001
EMNLP2
2022 Towards Efficient Dialogue Pre-training with Transferable and Interpretable Latent Structure
abstract
With the availability of massive generaldomain dialogue data, pre-trained dialogue generation appears to be super appealing to transfer knowledge from the general domain to downstream applications.In most existing work, such transferable ability is mainly obtained by fitting a large model with hundreds of millions of parameters on massive data in an exhaustive way, leading to inefficient running and poor interpretability.This paper proposes a novel dialogue generation model with a latent structure that is easily transferable from the general domain to downstream tasks in a lightweight and transparent way.Experiments on two benchmarks validate the effectiveness of the proposed model.Thanks to the transferable latent structure, our model is able to yield better dialogue responses than four strong baselines in terms of both automatic and human evaluations, and our model with about 22% parameters particularly delivers a 5x speedup in running time compared with the strongest baseline.Moreover, the proposed model is explainable by interpreting the discrete latent variables.
Xueliang Zhao, Lemao Liu, Tingchen Fu, Shuming Shi 0001, Dongyan Zhao 0001, Rui Yan 0001
EMNLP2
2022 On Synthetic Data for Back Translation
abstract
Jiahao Xu, Yubin Ruan, Wei Bi, Guoping Huang, Shuming Shi, Lihui Chen, Lemao Liu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jiahao Xu 0001, Yubin Ruan, Wei Bi, Guoping Huang, Shuming Shi 0001, Lihui Chen 0001, Lemao Liu
NAACL-HLT7
2022 Recent Advances in Retrieval-Augmented Text Generation
abstract
Recently retrieval-augmented text generation has achieved state-of-the-art performance in many NLP tasks and has attracted increasing attention of the NLP and IR community, this tutorial thereby aims to present recent advances in retrieval-augmented text generation comprehensively and comparatively. It firstly highlights the generic paradigm of retrieval-augmented text generation, then reviews notable works for different text generation tasks including dialogue generation, machine translation, and other generation tasks, and finally points out some limitations and shortcomings to facilitate future research.
Deng Cai 0002, Yan Wang 0060, Lemao Liu, Shuming Shi 0001
SIGIR3
2021 Neural Machine Translation with Monolingual Translation Memory
abstract
Deng Cai, Yan Wang, Huayang Li, Wai Lam, Lemao Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Deng Cai 0002, Yan Wang 0060, Wai Lam, Lemao Liu
ACL/IJCNLP (1)5
2021 Fast and Accurate Neural Machine Translation with Translation Memory
abstract
Qiuxiang He, Guoping Huang, Qu Cui, Li Li, Lemao Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Qiuxiang He, Guoping Huang, Qu Cui, Li Li 0006, Lemao Liu
ACL/IJCNLP (1)5
2021 GWLAN: General Word-Level AutocompletioN for Computer-Aided Translation
abstract
Huayang Li, Lemao Liu, Guoping Huang, Shuming Shi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Lemao Liu, Guoping Huang, Shuming Shi 0001
ACL/IJCNLP (1)2
2021 Engage the Public: Poll Question Generation for Social Media Posts
abstract
Zexin Lu, Keyang Ding, Yuji Zhang, Jing Li, Baolin Peng, Lemao Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Keyang Ding, Yuji Zhang 0002, Jing Li 0049, Baolin Peng, Lemao Liu
ACL/IJCNLP (1)6
2021 An Empirical Study on Multiple Information Sources for Zero-Shot Fine-Grained Entity Typing
abstract
Auxiliary information from multiple sources has been demonstrated to be effective in zeroshot fine-grained entity typing (ZFET).However, there lacks a comprehensive understanding about how to make better use of the existing information sources and how they affect the performance of ZFET.In this paper, we empirically study three kinds of auxiliary information: context consistency, type hierarchy and background knowledge (e.g., prototypes and descriptions) of types, and propose a multi-source fusion model (MSF) targeting these sources.The performance obtains up to 11.42% and 22.84% absolute gains over stateof-the-art baselines on BBN and Wiki respectively with regard to macro F1 scores.More importantly, we further discuss the characteristics, merits and demerits of each information source and provide an intuitive understanding of the complementarity among them.
Yi Chen 0019, Haiyun Jiang, Lemao Liu, Shuming Shi 0001, Chuang Fan, Min Yang 0007, Ruifeng Xu 0001
EMNLP (1)3
2021 Fine-grained Entity Typing without Knowledge Base
abstract
Existing work on Fine-grained Entity Typing (FET) typically trains automatic models on the datasets obtained by using Knowledge Bases (KB) as distant supervision.However, the reliance on KB means this training setting can be hampered by the lack of or the incompleteness of the KB.To alleviate this limitation, we propose a novel setting for training FET models: FET without accessing any knowledge base.Under this setting, we propose a two-step framework to train FET models.In the first step, we automatically create pseudo data with fine-grained labels from a large unlabeled dataset.Then a neural network model is trained based on the pseudo data, either in an unsupervised way or using self-training under the weak guidance from a coarse-grained Named Entity Recognition (NER) model.Experimental results show that our method achieves competitive performance with respect to the models trained on the original KB-supervised datasets.* The first two authors (Jing and Yibin) contributed equally to this work during the internships at Tencent AI Lab.
Lemao Liu, Yangming Li, Haiyun Jiang, Haisong Zhang, Shuming Shi 0001
EMNLP (1)3
2021 Empirical Analysis of Unlabeled Entity Problem in Named Entity Recognition
Yangming Li, Lemao Liu, Shuming Shi 0001
ICLR2
2021 Neural Sequence Segmentation as Determining the Leftmost Segments
abstract
Prior methods to text segmentation are mostly at token level.Despite the adequacy, this nature limits their full potential to capture the long-term dependencies among segments.In this work, we propose a novel framework that incrementally segments natural language sentences at segment level.For every step in segmentation, it recognizes the leftmost segment of the remaining sequence.Implementations involve LSTM-minus technique to construct the phrase representations and recurrent neural networks (RNN) to model the iterations of determining the leftmost segments.We have conducted extensive experiments on syntactic chunking and Chinese part-of-speech (POS) tagging across 3 datasets, demonstrating that our methods have significantly outperformed previous all baselines and achieved new stateof-the-art results.Moreover, qualitative analysis and the study on segmenting long-length sentences verify its effectiveness in modeling long-term dependencies.
Yangming Li, Lemao Liu, Kaisheng Yao
NAACL-HLT2
2021 Attending From Foresight: A Novel Attention Mechanism for Neural Machine Translation
abstract
Machines translation (MT) is an essential task in natural language processing or even in artificial intelligence. Statistical machine translation has been the dominant approach to MT for decades, but recently neural machine translation achieves increasing interest because of its appealing model architecture and impressive translation performance. In neural machine translation, an attention model is used to identify the aligned source words for the next target word, i.e., target foresight word, to select translation context. However, it does not make use of any information about this target foresight word at all. Previous work proposed an approach to improve the attention model by explicitly accessing this target foresight word and demonstrating substantial alignment tasks. However, this approach cannot be applied in machine translation tasks where the target foresight word is unavailable. This paper proposes several novel enhanced attention models by introducing hidden information (such as part-of-speech) of the target foresight word for the translation task. We incorporate the novel enhanced attention employing hidden information about the target foresight word into both recurrent and self-attention-based neural translation models and theoretically justify that such hidden information can make translation prediction easier. Empirical experiments on four datasets further verify that the proposed attention models deliver significant improvements in translation quality.
Lemao Liu, Zhaopeng Tu, Shuming Shi 0001, Max Q.-H. Meng
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Detecting Source Contextual Barriers for Understanding Neural Machine Translation
abstract
In machine translation evaluation, the traditional wisdom measures model's generalization ability in an average sense, for example by using corpus BLEU. However, the statistics of corpus BLEU cannot provide comprehensive understanding and fine-grained analysis on model's generalization ability. As a remedy, this paper attempts to understand NMT at fine-grained level, by detecting contextual barriers within an unseen input sentence that \textit{cause} the degradation in model's translation quality. It proposes a principled definition of source contextual barriers as well as its modified version which is tractable in computation and operates at word-level. Based on the modified one, three simple methods are proposed for barrier detection by search-aware risk estimation through counterfactual generation. Extensive analyses are conducted on those detected contextual barrier words on both Zh$\Leftrightarrow$En NIST benchmarks. Potential usages motivated from barrier words are also discussed.
Lemao Liu, Conghui Zhu, Rui Wang 0015, Tiejun Zhao, Shuming Shi 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Balancing Quality and Human Involvement: An Effective Approach to Interactive Neural Machine Translation
abstract
Conventional interactive machine translation typically requires a human translator to validate every generated target word, even though most of them are correct in the advanced neural machine translation (NMT) scenario. Previous studies have exploited confidence approaches to address the intensive human involvement issue, which request human guidance only for a few number of words with low confidences. However, such approaches do not take the history of human involvement into account, and optimize the models only for the translation quality while ignoring the cost of human involvement. In response to these pitfalls, we propose a novel interactive NMT model, which explicitly accounts the history of human involvements and particularly is optimized towards two objectives corresponding to the translation quality and the cost of human involvement, respectively. Specifically, the model jointly predicts a target word and a decision on whether to request human guidance, which is based on both the partial translation and the history of human involvements. Since there is no explicit signals on the decisions of requesting human guidance in the bilingual corpus, we optimize the model with the reinforcement learning technique which enables our model to accurately predict when to request human guidance. Simulated and real experiments show that the proposed model can achieve higher translation quality with similar or less human involvement over the confidence-based baseline.
Tianxiang Zhao 0001, Lemao Liu, Guoping Huang, Yingling Liu, Guiquan Liu, Shuming Shi 0001
AAAI2
2020 Evaluating Explanation Methods for Neural Machine Translation
abstract
Recently many efforts have been devoted to interpreting the black-box NMT models, but little progress has been made on metrics to evaluate explanation methods.Word Alignment Error Rate can be used as such a metric that matches human understanding, however, it can not measure explanation methods on those target words that are not aligned to any source word.This paper thereby makes an initial attempt to evaluate explanation methods from an alternative viewpoint.To this end, it proposes a principled metric based on fidelity in regard to the predictive behavior of the NMT model.As the exact computation for this metric is intractable, we employ an efficient approach as its approximation.On six standard translation tasks, we quantitatively evaluate several explanation methods in terms of the proposed metric and we reveal some valuable findings for these explanation methods in our experiments.
Jierui Li, Lemao Liu, Guoping Huang, Shuming Shi 0001
ACL2
2020 Regularized Context Gates on Transformer for Machine Translation
abstract
Context gates are effective to control the contributions from the source and target contexts in the recurrent neural network (RNN) based neural machine translation (NMT).However, it is challenging to extend them into the advanced Transformer architecture, which is more complicated than RNN.This paper first provides a method to identify source and target contexts and then introduce a gate mechanism to control the source and target contributions in Transformer.In addition, to further reduce the bias problem in the gate mechanism, this paper proposes a regularization method to guide the learning of the gates with supervision automatically generated using pointwise mutual information.Extensive experiments on 4 translation datasets demonstrate that the proposed model obtains an averaged gain of 1.0 BLEU score over a strong Transformer baseline.
Lemao Liu, Rui Wang 0015, Guoping Huang, Max Meng
ACL2
2020 Agreement on Target-Bidirectional Recurrent Neural Networks for Sequence-to-Sequence Learning
abstract
Recurrent neural networks are extremely appealing for sequence-to-sequence learning tasks. Despite their great success, they typically suffer from a shortcoming: they are prone to generate unbalanced targets with good prefixes but bad suffixes, and thus performance suffers when dealing with long sequences. We propose a simple yet effective approach to overcome this shortcoming. Our approach relies on the agreement between a pair of target-directional RNNs, which generates more balanced targets. In addition, we develop two efficient approximate search methods for agreement that are empirically shown to be almost optimal in terms of either sequence level or non-sequence level metrics. Extensive experiments were performed on three standard sequence-to-sequence transduction tasks: machine transliteration, grapheme-to-phoneme transformation and machine translation. The results show that the proposed approach achieves consistent and substantial improvements, compared to many state-of-the-art systems.
Lemao Liu, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita
J. Artif. Intell. Res.1
2020 Neural Machine Translation With Noisy Lexical Constraints
abstract
In neural machine translation, lexically constrained decoding generates translation outputs strictly including the constraints predefined by users, and it is beneficial to improve translation quality at the cost of more decoding overheads if the constraints are perfect. Unfortunately, those constraints may contain mistakes in real-world situations and incorrect constraints will undermine lexically constrained decoding. In this article, we propose a novel framework that is capable of improving the translation quality even if the constraints are noisy. The key to our framework is to treat the lexical constraints as external memories. More concretely, it encodes the constraints by a memory encoder and then leverages the memories by a memory integrator. Experiments demonstrate that our framework can not only deliver substantial BLEU gains in handling noisy constraints, but also achieve speedup in decoding. These results motivate us to apply our models to a new scenario where the constraints are generated without the help of users. Experiments show that our models can indeed improve the translation quality with the automatically generated constraints.
Guoping Huang, Deng Cai 0002, Lemao Liu
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Graph Based Translation Memory for Neural Machine Translation
abstract
A translation memory (TM) is proved to be helpful to improve neural machine translation (NMT). Existing approaches either pursue the decoding efficiency by merely accessing local information in a TM or encode the global information in a TM yet sacrificing efficiency due to redundancy. We propose an efficient approach to making use of the global information in a TM. The key idea is to pack a redundant TM into a compact graph and perform additional attention mechanisms over the packed graph for integrating the TM representation into the decoding network. We implement the model by extending the state-of-the-art NMT, Transformer. Extensive experiments on three language pairs show that the proposed approach is efficient in terms of running time and space occupation, and particularly it outperforms multiple strong baselines in terms of BLEU scores.
Mengzhou Xia, Guoping Huang, Lemao Liu, Shuming Shi 0001
AAAI3
2019 On the Word Alignment from Neural Machine Translation
abstract
Prior researches suggest that neural machine translation (NMT) captures word alignment through its attention mechanism, however, this paper finds attention may almost fail to capture word alignment for some NMT models.This paper thereby proposes two methods to induce word alignment which are general and agnostic to specific NMT models.Experiments show that both methods induce much better word alignment than attention.This paper further visualizes the translation through the word alignment induced by NMT.In particular, it analyzes the effect of alignment errors on translation errors at word level and its quantitative analysis over many testing examples consistently demonstrate that alignment errors are likely to lead to translation errors measured by different metrics.
Lemao Liu, Max Meng, Shuming Shi 0001
ACL (1)3
2019 Understanding Data Augmentation in Neural Machine Translation: Two Perspectives towards Generalization
abstract
Guanlin Li, Lemao Liu, Guoping Huang, Conghui Zhu, Tiejun Zhao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Lemao Liu, Guoping Huang, Conghui Zhu, Tiejun Zhao
EMNLP/IJCNLP (1)2
2019 Word Position Aware Translation Memory for Neural Machine Translation
Qiuxiang He, Guoping Huang, Lemao Liu, Li Li 0006
NLPCC (1)3
2018 Improving Sequence-to-Sequence Constituency Parsing
abstract
Sequence-to-sequence constituency parsing casts the tree structured prediction problem as a general sequential problem by top-down tree linearization,and thus it is very easy to train in parallel with distributed facilities. Despite its success, it relies on a probabilistic attention mechanism for a general purpose, which can not guarantee the selected context to be informative in the specific parsing scenario. Previous work introduced a deterministic attention to select the informative context for sequence-to-sequence parsing, but it is based on the bottom-up linearization even if it was observed that top-down linearization is better than bottom-up linearization for standard sequence-to-sequence constituency parsing. In this paper, we thereby extend the deterministic attention to directly conduct on the top-down tree linearization. Intensive experiments show that our parser delivers substantial improvements over the bottom-up linearization in accuracy, and it achieves 92.3 Fscore on the Penn English Treebank section 23 and 85.4 Fscore on the Penn Chinese Treebank test dataset, without reranking or semi-supervised training.
Lemao Liu, Muhua Zhu, Shuming Shi 0001
AAAI1
2018 Target Foresight Based Attention for Neural Machine Translation
abstract
Xintong Li, Lemao Liu, Zhaopeng Tu, Shuming Shi, Max Meng. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Lemao Liu, Zhaopeng Tu, Shuming Shi 0001, Max Meng
NAACL-HLT2
2018 A Neural Approach to Source Dependence Based Context Model for Statistical Machine Translation
abstract
In statistical machine translation, translation prediction considers not only the aligned source word itself but also its source contextual information. Learning context representation is a promising method for improving translation results, particularly through neural networks. Most of the existing methods process context words sequentially and neglect source long-distance dependencies. In this paper, we propose a novel neural approach to source dependence-based context representation for translation prediction. The proposed model is capable of not only encoding source long-distance dependencies but also capturing functional similarities to better predict translations (i.e., word form translations and ambiguous word translations). To verify our method, the proposed mode is incorporated into phrase-based and hierarchical phrase-based translation models, respectively. Experiments on large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves significant improvement over the baseline systems and outperforms several existing context-enhanced methods.
Kehai Chen, Tiejun Zhao, Muyun Yang, Lemao Liu, Akihiro Tamura, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Sentence Selection and Weighting for Neural Machine Translation Domain Adaptation
abstract
Neural machine translation (NMT) has been prominent in many machine translation tasks. However, in some domain-specific tasks, only the corpora from similar domains can improve translation performance. If out-of-domain corpora are directly added into the in-domain corpus, the translation performance may even degrade. Therefore, domain adaptation techniques are essential to solve the NMT domain problem. Most existing methods for domain adaptation are designed for the conventional phrase-based machine translation. For NMT domain adaptation, there have been only a few studies on topics such as fine tuning, domain tags, and domain features. In this paper, we have four goals for sentence level NMT domain adaptation. First, the NMT's internal sentence embedding is exploited and the sentence embedding similarity is used to select out-of-domain sentences that are close to the in-domain corpus. Second, we propose three sentence weighting methods, i.e., sentence weighting, domain weighting, and batch weighting, to balance the data distribution during NMT training. Third, in addition, we propose dynamic training methods to adjust the sentence selection and weighting during NMT training. Fourth, to solve the multidomain problem in a real-world NMT scenario where the domain distributions of training and testing data often mismatch, we proposed a multidomain sentence weighting method to balance the domain distributions of training data and match the domain distributions of training and testing data. The proposed methods are evaluated in international workshop on spoken language translation (IWSLT) English-to-French/German tasks and a multidomain English-to-French task. Empirical results show that the sentence selection and weighting methods can significantly improve the NMT performance, outperforming the existing baselines.
Rui Wang 0015, Masao Utiyama, Andrew M. Finch, Lemao Liu, Kehai Chen, Eiichiro Sumita
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Translation Prediction with Source Dependency-Based Context Representation
abstract
Learning context representations is very promising to improve translation results, particularly through neural networks. Previous efforts process the context words sequentially and neglect their internal syntactic structure. In this paper, we propose a novel neural network based on bi-convolutional architecture to represent the source dependency-based context for translation prediction. The proposed model is able to not only encode the long-distance dependencies but also capture the functional similarities for better translation prediction (i.e., ambiguous words translation and word forms translation). Examined by a large-scale Chinese-English translation task, the proposed approach achieves a significant improvement (of up to +1.9 BLEU points) over the baseline system, and meanwhile outperforms a number of context-enhanced comparison system.
Kehai Chen, Tiejun Zhao, Muyun Yang, Lemao Liu
AAAI4
2017 Deterministic Attention for Sequence-to-Sequence Constituent Parsing
abstract
The sequence-to-sequence model is proven to be extremely successful in constituent parsing. It relies on one key technique, the probabilistic attention mechanism, to automatically select the context for prediction. Despite its successes, the probabilistic attention model does not always select the most important context. For example, the headword and boundary words of a subtree have been shown to be critical when predicting the constituent label of the subtree, but this contextual information becomes increasingly difficult to learn as the length of the sequence increases. In this study, we proposed a deterministic attention mechanism that deterministically selects the important context and is not affected by the sequence length. We implemented two different instances of this framework. When combined with a novel bottom-up linearization method, our parser demonstrated better performance than that achieved by the sequence-to-sequence parser with probabilistic attention mechanism.
Chunpeng Ma, Lemao Liu, Akihiro Tamura, Tiejun Zhao, Eiichiro Sumita
AAAI2
2017 Neural Machine Translation with Source Dependency Representation
abstract
Source dependency information has been successfully introduced into statistical machine translation.However, there are only a few preliminary attempts for Neural Machine Translation (NMT), such as concatenating representations of source word and its dependency label together.In this paper, we propose a novel attentional NMT with source dependency representation to improve translation performance of NMT, especially on long sentences.Empirical results on NIST Chinese-to-English translation task show that our method achieves 1.6 BLEU improvements on average over a strong NMT system.
Kehai Chen, Rui Wang 0015, Masao Utiyama, Lemao Liu, Akihiro Tamura, Eiichiro Sumita, Tiejun Zhao
EMNLP4
2017 Instance Weighting for Neural Machine Translation Domain Adaptation
abstract
Instance weighting has been widely applied to phrase-based machine translation domain adaptation.However, it is challenging to be applied to Neural Machine Translation (NMT) directly, because NMT is not a linear model.In this paper, two instance weighting technologies, i.e., sentence weighting and domain weighting with a dynamic weight learning strategy, are proposed for NMT domain adaptation.Empirical results on the IWSLT English-German/French tasks show that the proposed methods can substantially improve NMT performance by up to 2.7-6.7 BLEU points, outperforming the existing baselines by up to 1.6-3.6BLEU points.
Rui Wang 0015, Masao Utiyama, Lemao Liu, Kehai Chen, Eiichiro Sumita
EMNLP3
2017 Translation Quality Estimation Using Only Bilingual Corpora
abstract
In computer-aided translation scenarios, quality estimation of machine translation hypotheses plays a critical role. Existing methods for word-level translation quality estimation (TQE) rely on the availability of manually annotated TQE training data obtained via direct annotation or postediting. However, due to the cost of human labor, such data are either limited in size or is only available for few tasks in practice. To avoid the reliance on such annotated TQE data, this paper proposes an approach to train word-level TQE models using bilingual corpora, which are typically used in machine translation training and is relatively easier to access. We formalize the training of our proposed method under the framework of maximum marginal likelihood estimation. To avoid degenerated solutions, we propose a novel regularized training objective whose optimization is achieved by an efficient approximation. Extensive experiments on both written and spoken language datasets empirically show that our approach yields comparable performance to the standard training on annotated data.
Lemao Liu, Atsushi Fujita, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Agreement on Target-Bidirectional LSTMs for Sequence-to-Sequence Learning
abstract
Recurrent neural networks, particularly the long short- term memory networks, are extremely appealing for sequence-to-sequence learning tasks. Despite their great success, they typically suffer from a fundamental short- coming: they are prone to generate unbalanced targets with good prefixes but bad suffixes, and thus perfor- mance suffers when dealing with long sequences. We propose a simple yet effective approach to overcome this shortcoming. Our approach relies on the agreement between a pair of target-directional LSTMs, which generates more balanced targets. In addition, we develop two efficient approximate search methods for agreement that are empirically shown to be almost optimal in terms of sequence-level losses. Extensive experiments were performed on two standard sequence-to-sequence trans- duction tasks: machine transliteration and grapheme-to- phoneme transformation. The results show that the proposed approach achieves consistent and substantial im- provements, compared to six state-of-the-art systems. In particular, our approach outperforms the best reported error rates by a margin (up to 9% relative gains) on the grapheme-to-phoneme task.
Lemao Liu, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita
AAAI1
2016 Neural Machine Translation with Supervised Attention
abstract
The attention mechanism is appealing for neural machine translation, since it is able to dynamically encode a source sentence by generating a alignment between a target word and source words. Unfortunately, it has been proved to be worse than conventional alignment models in alignment accuracy. In this paper, we analyze and explain this issue from the point view of reordering, and propose a supervised attention which is learned with guidance from conventional alignment models. Experiments on two Chinese-to-English translation tasks show that the supervised attention mechanism yields better alignments leading to substantial gains over the standard attention based NMT.
Lemao Liu, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita
COLING1
2016 Local fisher discriminant analysis for spoken language identification
abstract
I-vector is a state-of-the-art technique widely used in spoken language identification systems. Since i-vectors include total variability factors, discriminant analysis methods have been introduced to find the most discriminative features while removing the undesired variables for language identification, for example, linear discriminant analysis (LDA) and nonparametric discriminant analysis (NDA). However, these methods either do not consider or use weak local structures of the data. In this study, we introduce a local Fisher discriminant analysis (LFDA) as a post-processing discriminant analysis method to extract the discriminative features from i-vectors. LFDA is a full-rank method which takes the local structure of the data into account for non-Gaussian distribution data, i.e., multimodal. Compared with LDA and NDA, LFDA is a pair-wise local method which enhances the centralization of the distribution of samples in the same class to obtain larger amounts of discriminative features. Experimental results indicate that LFDA is more effective than LDA and NDA for the i-vector-based language identification task.
Xugang Lu, Lemao Liu, Hisashi Kawai
ICASSP3
2016 Agreement on Target-bidirectional Neural Machine Translation
abstract
Lemao Liu, Masao Utiyama, Andrew Finch, Eiichiro Sumita. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Lemao Liu, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita
HLT-NAACL1
2014 Search-Aware Tuning for Machine Translation
abstract
Parameter tuning is an important problem in statistical machine translation, but surprisingly, most existing methods such as MERT, MIRA and PRO are agnostic about search, while search errors could severely degrade translation quality.We propose a searchaware framework to promote promising partial translations, preventing them from being pruned.To do so we develop two metrics to evaluate partial derivations.Our technique can be applied to all of the three above-mentioned tuning methods, and extensive experiments on Chinese-to-English and English-to-Chinese translation show up to +2.6 BLEU gains over search-agnostic baselines.
Lemao Liu, Liang Huang 0001
EMNLP1
2014 Discriminative Training for Log-Linear Based SMT: Global or Local Methods
abstract
In statistical machine translation, the standard methods such as MERT tune a single weight with regard to a given development data. However, these methods suffer from two problems due to the diversity and uneven distribution of source sentences. First, their performance is highly dependent on the choice of a development set, which may lead to an unstable performance for testing. Second, the sentence level translation quality is not assured since tuning is performed on the document level rather than on sentence level. In contrast with the standard global training in which a single weight is learned, we propose novel local training methods to address these two problems. We perform training and testing in one step by locally learning the sentence-wise weight for each input sentence. Since the time of each tuning step is unnegligible and learning sentence-wise weights for the entire test set means many passes of tuning, it is a great challenge for the efficiency of local training. We propose an efficient two-phase method to put the local training into practice by employing the ultraconservative update. On NIST Chinese-to-English translation tasks with both medium and large scales of training data, our local training methods significantly outperform standard methods with the maximal improvements up to 2.0 BLEU points, meanwhile their efficiency is comparable to that of the standard methods.
Lemao Liu, Tiejun Zhao, Taro Watanabe, Hailong Cao, Conghui Zhu
ACM Trans. Asian Lang. Inf. Process.1
2013 Additive Neural Networks for Statistical Machine Translation
Lemao Liu, Taro Watanabe, Eiichiro Sumita, Tiejun Zhao
ACL (1)1
2013 Tuning SMT with a Large Number of Features via Online Feature Grouping
Lemao Liu, Tiejun Zhao, Taro Watanabe, Eiichiro Sumita
IJCNLP1
2012 Locally Training the Log-Linear Model for SMT
Lemao Liu, Hailong Cao, Taro Watanabe, Tiejun Zhao, Mo Yu, Conghui Zhu
EMNLP-CoNLL1
2011 A Unified and Discriminative Soft Syntactic Constraint Model for Hierarchical Phrase-based Translation
Lemao Liu, Tiejun Zhao, Hailong Cao
MTSummit1