Xuebo Liu 0002

dblp:166/0029-2 · DBLP profile ↗
← Back
50ranked-venue papers
5as first author
47since 2021 · last 2026
0000-0001-8524-2006ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 5 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021
YearPublicationVenuePosition
2026 Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation
abstract
REtrieval-Augmented LLM-based Machine Translation (REAL-MT) shows promise for knowledge-intensive tasks like idiomatic translation, but its reliability under noisy retrieval, a common challenge in real-world deployment, remains poorly understood. To address this gap, we propose a noise synthesis framework and new metrics to systematically evaluate REAL-MT’s reliability across high-, medium-, and low-resource language pairs. Using both open- and closed-sourced models, including standard LLMs and large reasoning models (LRMs), we find that models heavily rely on retrieved context, and this dependence is significantly more detrimental in low-resource language pairs, producing nonsensical translations. Although LRMs possess enhanced reasoning capabilities, they show no improvement in error correction and are even more susceptible to noise, tending to rationalize incorrect contexts. Attention analysis reveals a shift from the source idiom to noisy content, while confidence increases despite declining accuracy, indicating poor self-monitoring. To mitigate these issues, we investigate training-free and fine-tuning strategies, which improve robustness at the cost of performance in clean contexts, revealing a fundamental trade-off. Our findings highlight the limitations of current approaches, underscoring the need for self-verifying integration mechanisms.
Runzhe Zhan, Chi Seng Cheang, Xuebo Liu 0002, Yuyao Niu, Fengying Ye, Kaixin Lan, Lidia S. Chao, Derek F. Wong
AAAI5
2026 DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate Reasoning
abstract
Rongqing Jiang, Xuebo Liu, Shengxin Liu, Yutong Wang, Min Zhang, Shimin Tao, Daimeng Wei, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Rongqing Jiang, Xuebo Liu 0002, Shengxin Liu, Min Zhang 0005, Shimin Tao, Daimeng Wei
ACL (1)2
2026 Exploring and enhancing the transfer of distribution in knowledge distillation for autoregressive language models
Jun Rao, Xuebo Liu 0002, Zepeng Lin, Liang Ding 0006, Jing Li 0034, Min Zhang 0005
Knowl. Based Syst.2
2026 Orchestrating Prompt Expertise: Enhancing Knowledge Distillation via Expert-Guided Tuning
abstract
Multi-teacher knowledge distillation transfers knowledge from multiple large teacher models to a small student model and has performed well on many downstream tasks. However, when distilling knowledge from multiple teachers, it always suffers from the severe problems of being time-consuming and storage-extensive for multiple teacher models training and inference. We present MoE-KD, a simple but effective framework that produces supervision for training the student model from one single teacher model, which fixes the above problems and improves effectiveness. In the proposed MoE-KD, multiple trainable prompts are used to extract different views of samples from a single pre-trained language model and only a few parameters (prompts) need to be trained and stored. To guarantee the generated supervision signals with increased robustness and correctness, we introduce an uncertainty-based mechanism and a selector module, which routes the input instance to its corresponding teacher. We have also extended MoE KD to lifelong learning scenarios, proposing a lightweight solution for catastrophic forgetting. We conduct experiments on traditional KD scenarios and lifelong learning scenarios. MoE-KD yields improvements up to 1.1% and 140% in accuracy and efficiency in knowledge distillation and 2.8% improvements on average in lifelong learning, compared with the strong baseline methods.
Jun Rao, Shuhan Qi, Xuebo Liu 0002, Lei Wang 0203, Min Zhang 0005, Xuan Wang 0002
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2026 Domain Adaptive Machine Translation with Synthetic Feedback for Large Language Models
abstract
Domain-specific machine translation (MT) significantly benefits from large language models (LLMs) due to their strong instruction-following abilities and in-context learning (ICL) capabilities. Appropriate demonstration samples and feedback are essential for helping LLMs refine their translation outputs in real-world applications. However, the scarcity of in-domain samples and professional feedback creates practical limitations. Furthermore, the current ICL paradigm does not offer the fine-grained domain features in addition to parallel translation pairs. To address these challenges, we propose a pipeline that collects in-domain translations from LLMs and generates synthetic, human-like feedback for revising these translations. The translations and their corresponding feedback are stored together to build a demonstration database, with each instance paired with the original in-domain translation and its revision. During online translation, similar in-domain translations can be retrieved as revision demonstrations. This process guides LLMs in iteratively refining their outputs by learning from demonstrations. We evaluate the proposed pipeline using open-source models like Llama3-8B-Instruct and Mistral-7B-Instruct-v0.3, on five domain-specific benchmarks for English-centric, Chinese-centric, and Portuguese-centric translation. The results demonstrate the effectiveness of the pipeline in tailoring in-domain translations and improving translation performance compared to direct translation instructions. Additionally, we discuss the experimental results from the following perspectives: (1) the effectiveness of different in-context retrieval methods; (2) the observed differences across selected domains and language; (3) the quantitative analysis of sentence-level and word-level statistics; and (4) the effect of ICL retrieval database size and decoding parameters.
Xinyi Yang 0008, Runzhe Zhan, Junchao Wu, Yue Zhang 0004, Xuebo Liu 0002, Lidia S. Chao, Derek F. Wong
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2026 Exploiting Multimodal Knowledge Graph for Multimodal Machine Translation
abstract
A neural Multimodal Machine Translation (MMT) system utilizes multimodal information, particularly images, to enhance traditional text-only models and achieve superior performance. However, the effectiveness of MMT heavily depends on the availability of extensive collections of bilingual parallel sentence pairs and manually annotated images, which poses a challenge due to the scarcity of such pairs. To address this issue, we propose incorporating the Multimodal Knowledge Graph (MMKG) for data augmentation in MMT. By utilizing MMKG as an additional source of knowledge, we can overcome the limitations of existing sentence-image pairings. This allows us to expand the original parallel corpus and generate corresponding images, creating new synthetic data pairs that facilitate effective data augmentation. Experiments conducted on two translation datasets, Multi30k and IKEA, demonstrate that the proposed MMKG enhancement method significantly improves performance across multiple baseline methods, ultimately outperforming all baseline approaches. Additionally, experiments under low-resource conditions reveal that our method achieves exceptional enhancement effects in low-resource corpora, surpassing other data augmentation baseline methods. These results indicate the efficacy and potential of the proposed method for enhancing the performance of multimodal models across diverse datasets.
Tianjiao Xu, Xuebo Liu 0002, Derek F. Wong, Yue Zhang 0004, Lidia S. Chao, Min Zhang 0005, Tian Gan 0002
IEEE Trans. Multim.2
2025 Knowledge Editing with Dynamic Knowledge Graphs for Multi-Hop Question Answering
abstract
Multi-hop question answering (MHQA) poses a significant challenge for large language models (LLMs) due to the extensive knowledge demands involved. Knowledge editing, which aims to precisely modify the LLMs to incorporate specific knowledge without negatively impacting other unrelated knowledge, offers a potential solution for addressing MHQA challenges with LLMs. However, current solutions struggle to effectively resolve issues of knowledge conflicts. Most parameter-preserving editing methods are hindered by inaccurate retrieval and overlook secondary editing issues, which can introduce noise into the reasoning process of LLMs. In this paper, we introduce KEDKG, a novel knowledge editing method that leverages a dynamic knowledge graph for MHQA, designed to ensure the reliability of answers. KEDKG involves two primary steps: dynamic knowledge graph construction and knowledge graph augmented generation. Initially, KEDKG autonomously constructs a dynamic knowledge graph to store revised information while resolving potential knowledge conflicts. Subsequently, it employs a fine-grained retrieval strategy coupled with an entity and relation detector to enhance the accuracy of graph retrieval for LLM generation. Experimental results on benchmarks show that KEDKG surpasses previous state-of-the-art models, delivering more accurate and reliable answers in environments with dynamic information.
Yigeng Zhou, Jing Li 0034, Yequan Wang, Xuebo Liu 0002, Daojing He, Fangming Liu, Min Zhang 0005
AAAI5
2025 SGIC: A Self-Guided Iterative Calibration Framework for RAG
abstract
Recent research in retrieval-augmented generation (RAG) has concentrated on retrieving useful information from candidate documents.However, numerous methodologies frequently neglect the calibration capabilities of large language models (LLMs), which capitalize on their robust in-context reasoning prowess.This work illustrates that providing LLMs with specific cues substantially improves their calibration efficacy, especially in multi-round calibrations.We present a new SGIC: Self-Guided Iterative Calibration Framework that employs uncertainty scores as a tool.Initially, this framework calculates uncertainty scores to determine both the relevance of each document to the query and the confidence level in the responses produced by the LLMs.Subsequently, it reevaluates these scores iteratively, amalgamating them with prior responses to refine calibration.Furthermore, we introduce an innovative approach for constructing an iterative self-calibration training set, which optimizes LLMs to efficiently harness uncertainty scores for capturing critical information and enhancing response accuracy.Our proposed framework significantly improves performance on both closed-source and open-weight LLMs.
Guanhua Chen 0006, Yutong Yao, Lidia S. Chao, Xuebo Liu 0002, Derek F. Wong
ACL (1)4
2025 DRPruning: Efficient Large Language Model Pruning through Distributionally Robust Optimization
abstract
Large language models (LLMs) deliver impressive results but face challenges from increasing model sizes and computational costs.Structured pruning reduces model size and speeds up inference but often causes uneven degradation across domains, leading to biased performance.To address this, we propose DR-Pruning, a method that dynamically adjusts the data distribution during training to restore balanced performance across heterogeneous and multi-tasking data.Experiments in monolingual and multilingual settings show that DR-Pruning surpasses similarly sized models in both pruning and continued pretraining over perplexity, downstream tasks, and instruction tuning.Further analysis demonstrates the robustness of DRPruning towards various domains and distribution shifts.Furthermore, DRPruning can determine optimal reference losses and data ratios automatically, suggesting potential for broader applications.
Hexuan Deng, Wenxiang Jiao, Xuebo Liu 0002, Min Zhang 0005, Zhaopeng Tu
ACL (1)3
2025 AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration
abstract
Multi-agent systems (MAS) based on large language models (LLMs) have demonstrated significant potential in collaborative problemsolving.However, they still face substantial challenges of low communication efficiency and suboptimal task performance, making the careful design of the agents' communication topologies particularly important.Inspired by the management theory that roles in an efficient team are often dynamically adjusted, we propose AgentDropout, which identifies redundant agents and communication across different communication rounds by optimizing the adjacency matrices of the communication graphs and eliminates them to enhance both token efficiency and task performance.Compared to state-of-the-art methods, AgentDropout achieves an average reduction of 21.6% in prompt token consumption and 18.4% in completion token consumption, along with a performance improvement of 1.14 on the tasks.Furthermore, the extended experiments demonstrate that AgentDropout achieves notable domain transferability and structure robustness, revealing its reliability and effectiveness.We release our code at https://github. com/wangzx1219/AgentDropout.
Zhexuan Wang, Xuebo Liu 0002, Liang Ding 0006, Miao Zhang 0037, Jie Liu 0001, Min Zhang 0005
ACL (1)3
2025 Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
abstract
The efficacy of detectors for texts generated by large language models (LLMs) substantially depends on the availability of large-scale training data. However, white-box zero-shot detectors, which require no such data, are limited by the accessibility of the source model of the LLM-generated text. In this paper, we propose a simple yet effective black-box zero-shot detection approach based on the observation that, from the perspective of LLMs, human-written texts typically contain more grammatical errors than LLM-generated texts. This approach involves calculating the Grammar Error Correction Score (GECScore) for the given text to differentiate between human-written and LLM-generated text. Experimental results show that our method outperforms current state-of-the-art (SOTA) zero-shot and supervised methods, achieving an average AUROC of 98.62% across XSum and Writing Prompts dataset. Additionally, our approach demonstrates strong reliability in the wild, exhibiting robust generalization and resistance to paraphrasing attacks. Data and code are available at: https://github.com/NLP2CT/GECScore.
Junchao Wu, Runzhe Zhan, Derek F. Wong, Shu Yang 0010, Xuebo Liu 0002, Lidia S. Chao, Min Zhang 0005
COLING5
2025 AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs
abstract
Despite the impressive performance of large language models (LLMs) in general domains, they often underperform in specialized domains.Existing approaches typically rely on data synthesis methods and yield promising results by using unlabeled data to capture domain-specific features.However, these methods either incur high computational costs or suffer from performance limitations, while also demonstrating insufficient generalization across different tasks.To address these challenges, we propose AQuilt, a framework for constructing instruction-tuning data for any specialized domains from corresponding unlabeled data, including Answer, Question, Unlabeled data, Inspection, Logic, and Task type.By incorporating logic and inspection, we encourage reasoning processes and self-inspection to enhance model performance.Moreover, customizable task instructions enable high-quality data generation for any task.As a result, we construct a dataset of 703k examples to train a powerful data synthesis model.Experiments show that AQuilt is comparable to DeepSeek-V3 while utilizing just 17% of the production cost.Further analysis demonstrates that our generated data exhibits higher relevance to downstream tasks.
Xiaopeng Ke, Hexuan Deng, Xuebo Liu 0002, Jun Rao, Zhenxi Song, Jun Yu 0002, Min Zhang 0005
EMNLP3
2025 Weight-Aware Activation Sparsity with Constrained Bayesian Optimization Scheduling for Large Language Models
abstract
Activation sparsity provides a dynamic, inputdependent alternative to weight pruning for accelerating inference in large language models (LLMs), effectively reducing unnecessary computations and memory accesses during the forward pass.Despite its promise, existing activation sparsification methods suffer from two major limitations: (1) solely relying on activation magnitude for sparsification, ignoring the coupling influence with the corresponding weights, (2) applying uniform sparsity rates across all blocks without considering block-wise sparsity sensitivity.To address these issues, this paper proposes a novel training-free weightaware activation sparsity framework, called WAS.Firstly, with analyzing the coupling relationship between weight and activation, we introduce a weight-aware scoring method to measure the activation importance in sparsification.Then, a novel constrained Bayesian optimization algorithm is further devised to set a suitable sparsity ratio for all blocks based on the sparsity sensitivity.Finally, we implement a custom GPU sparsity kernel to support the resulting sparsity patterns for wallclock decoding speed-ups.Our WAS achieves competitive performance at 60% model-level sparsity and significantly outperforms prior methods at higher sparsity levels, achieving up to 1.68× inference speed-up-at no retraining or weight update.Codes are available at https://github.com/HITSZ-Miao-Group/WAS.
Xuebo Liu 0002, Liqiang Nie
EMNLP3
2025 DelTA: An Online Document-Level Translation Agent Based on Multi-Level Memory
abstract
Large language models (LLMs) have achieved reasonable quality improvements in machine translation (MT). However, most current research on MT-LLMs still faces significant challenges in maintaining translation consistency and accuracy when processing entire documents. In this paper, we introduce DelTA, a Document-levEL Translation Agent designed to overcome these limitations. DelTA features a multi-level memory structure that stores information across various granularities and spans, including Proper Noun Records, Bilingual Summary, Long-Term Memory, and Short-Term Memory, which are continuously retrieved and updated by auxiliary LLM-based components. Experimental results indicate that DelTA significantly outperforms strong baselines in terms of translation consistency and quality across four open/closed-source LLMs and two representative document translation datasets, achieving an increase in consistency scores by up to 4.58 percentage points and in COMET scores by up to 3.16 points on average. DelTA employs a sentence-by-sentence translation strategy, ensuring no sentence omissions and offering a memory-efficient solution compared to the mainstream method. Furthermore, DelTA improves pronoun and context-dependent translation accuracy, and the summary component of the agent also shows promise as a tool for query-based summarization tasks. The code and data of our approach are released at https://github.com/YutongWang1216/DocMTAgent.
Jiali Zeng, Xuebo Liu 0002, Derek F. Wong, Fandong Meng, Jie Zhou 0016, Min Zhang 0005
ICLR3
2024 Revisiting Demonstration Selection Strategies in In-Context Learning
abstract
Keqin Peng, Liang Ding, Yancheng Yuan, Xuebo Liu, Min Zhang, Yuanxin Ouyang, Dacheng Tao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Keqin Peng, Liang Ding 0006, Yancheng Yuan, Xuebo Liu 0002, Min Zhang 0005, Yuanxin Ouyang, Dacheng Tao
ACL (1)4
2024 TasTe: Teaching Large Language Models to Translate through Self-Reflection
abstract
Large language models (LLMs) have exhibited remarkable performance in various natural language processing tasks.Techniques like instruction tuning have effectively enhanced the proficiency of LLMs in the downstream task of machine translation.However, the existing approaches fail to yield satisfactory translation outputs that match the quality of supervised neural machine translation (NMT) systems.One plausible explanation for this discrepancy is that the straightforward prompts employed in these methodologies are unable to fully exploit the acquired instruction-following capabilities.To this end, we propose the TASTE framework, which stands for translating through selfreflection.The self-reflection process includes two stages of inference.In the first stage, LLMs are instructed to generate preliminary translations and conduct self-assessments on these translations simultaneously.In the second stage, LLMs are tasked to refine these preliminary translations according to the evaluation results.The evaluation results in four language directions on the WMT22 benchmark reveal the effectiveness of our approach compared to existing methods.Our work presents a promising approach to unleash the potential of LLMs and enhance their capabilities in MT.The codes and datasets are open-sourced at https://github. com/YutongWang1216/ReflectionLLMMT.
Jiali Zeng, Xuebo Liu 0002, Fandong Meng, Jie Zhou 0016, Min Zhang 0005
ACL (1)3
2024 Speech Sense Disambiguation: Tackling Homophone Ambiguity in End-to-End Speech Translation
abstract
End-to-end speech translation (ST) presents notable disambiguation challenges as it necessitates simultaneous cross-modal and crosslingual transformations.While word sense disambiguation is an extensively investigated topic in textual machine translation, the exploration of disambiguation strategies for ST models remains limited.Addressing this gap, this paper introduces the concept of speech sense disambiguation (SSD), specifically emphasizing homophones -words pronounced identically but with different meanings.To facilitate this, we first create a comprehensive homophone dictionary and an annotated dataset rich with homophone information established based on speech-text alignment.Building on this unique dictionary, we introduce AmbigST, an innovative homophone-aware contrastive learning approach that integrates a homophone-aware masking strategy.Our experiments on different MuST-C and CoVoST ST benchmarks demonstrate that AmbigST sets new performance standards.Specifically, it achieves SOTA results on BLEU scores for English to German, Spanish, and French ST tasks, underlining its effectiveness in reducing speech sense ambiguity.
Tengfei Yu, Xuebo Liu 0002, Liang Ding 0006, Kehai Chen, Dacheng Tao, Min Zhang 0005
ACL (1)2
2024 LRQuant: Learnable and Robust Post-Training Quantization for Large Language Models
abstract
Post-training quantization (PTQ) for large language models (LLMs) significantly accelerates model inference and relieves memory constraints, without incurring model training.A "smoothing paradigm" is commonly used in LLM quantization, which transfers the quantization difficulty of activation to weight quantization using mathematically equivalent transformations.However, existing methods face two issues: 1) Most smoothing parameters are hand-crafted defined which leads to suboptimal results; 2) There are significant performance degradations when tested on unseen datasets.To address these challenges, this paper introduces a robust learnable smooth-based PTQ framework, called LRQuant.Firstly, we consider a learnable paradigm to find optimal smoothing parameters which are initialized by logarithmic activation equivalent.In addition, we empirically found that only relying on MSE loss could hardly lead to optimal quantization results, and we then propose a novel loss function based on the negative logarithm of cosine similarity (NLC loss) between outputs of full-precision and quantized block.At last, we pioneeringly introduce Testtime adaptation (TTA) into LLM quantization, which allows for rapid model adaptation during testing to improve generalization performance.More surprisingly, we find that by using our TTA method, we can achieve better results on test sets than directly using test sets for calibration in some cases while avoiding catastrophic forgetting.Codes are available at https://github.com/zjq0455/RLQ.
Miao Zhang 0022, Xuebo Liu 0002, Liqiang Nie
ACL (1)5
2024 3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
abstract
Multimodal machine translation (MMT) is a challenging task that seeks to improve translation quality by incorporating visual information. However, recent studies have indicated that the visual information provided by existing MMT datasets is insufficient, causing models to disregard it and overestimate their capabilities. This issue presents a significant obstacle to the development of MMT research. This paper presents a novel solution to this issue by introducing 3AM, an ambiguity-aware MMT dataset comprising 26,000 parallel sentence pairs in English and Chinese, each with corresponding images. Our dataset is specifically designed to include more ambiguity and a greater variety of both captions and images than other MMT datasets. We utilize a word sense disambiguation model to select ambiguous data from vision-and-language datasets, resulting in a more challenging dataset. We further benchmark several state-of-the-art MMT models on our proposed dataset. Experimental results show that MMT models trained on our dataset exhibit a greater ability to exploit visual information than those trained on other MMT datasets. Our work provides a valuable resource for researchers in the field of multimodal learning and encourages further exploration in this area. The data, code and scripts are freely available at https://github.com/MaxyLee/3AM.
Xuebo Liu 0002, Derek F. Wong, Jun Rao, Liang Ding 0006, Lidia S. Chao, Dacheng Tao, Min Zhang 0005
LREC/COLING2
2024 Pluggable Neural Machine Translation Models via Memory-augmented Adapters
abstract
Although neural machine translation (NMT) models perform well in the general domain, it remains rather challenging to control their generation behavior to satisfy the requirement of different users. Given the expensive training cost and the data scarcity challenge of learning a new model from scratch for each user requirement, we propose a memory-augmented adapter to steer pretrained NMT models in a pluggable manner. Specifically, we construct a multi-granular memory based on the user-provided text samples and propose a new adapter architecture to combine the model representations and the retrieved results. We also propose a training strategy using memory dropout to reduce spurious dependencies between the NMT model and the memory. We validate our approach on both style- and domain-specific experiments and the results indicate that our method can outperform several representative pluggable baselines.
Yuzhuang Xu, Shuo Wang 0013, Peng Li 0030, Xuebo Liu 0002, Xiaolong Wang 0014, Yang Liu 0005
LREC/COLING4
2024 EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
abstract
The vision and language generative models have been overgrown in recent years. For video generation, various open-sourced models and public-available services have been developed to generate high-quality videos. However, these methods often use a few metrics, e.g., FVD [56] or IS [45], to evaluate the performance. We argue that it is hard to judge the large conditional generative models from the simple metrics since these models are often trained on very large datasets with multi-aspect abilities. Thus, we propose a novel framework and pipeline for exhaustively evaluating the performance of the generated videos. Our approach involves generating a diverse and comprehensive list of 700 prompts for text-to-video generation, which is based on an analysis of real-world user data and generated with the assistance of a large language model. Then, we evaluate the state-of-the-art video generative models on our carefully designed benchmark, in terms of visual qualities, content qualities, motion qualities, and text-video alignment with 17 well-selected objective metrics. To obtain the finalleaderboard of the models, we further fit a series of coefficients to align the objective metrics to the users' opinions. Based on the proposed human alignment method, our final score shows a higher correlation than simply averaging the metrics, showing the effectiveness of the proposed evaluation method.
Yaofang Liu, Xiaodong Cun, Xuebo Liu 0002, Xintao Wang 0002, Yong Zhang 0034, Haoxin Chen, Yang Liu 0005, Tieyong Zeng, Raymond Chan 0001, Ying Shan
CVPR3
2024 Can LLMs Learn Uncertainty on Their Own? Expressing Uncertainty Effectively in A Self-Training Manner
abstract
Large language models (LLMs) often exhibit excessive, random, and uninformative uncertainty, rendering them unsuitable for decisionmaking in human-computer interactions.In this paper, we aim to instigate a heightened awareness of self-uncertainty in LLMs, enabling them to express uncertainty more effectively.To accomplish this, we propose an uncertainty-aware instruction tuning (UaIT) method, aligning LLMs' perception with the probabilistic uncertainty of the generation.We conducted experiments using LLaMA2 and Mistral on multiple free-form QA tasks.Experimental results revealed a surprising 45.2% improvement in the effectiveness of uncertainty expression by LLMs, accompanied by reasonably good out-of-domain generalization capabilities.Moreover, this uncertainty expression can serve as a valuable real-time basis for human decision-making, e.g., retrieving external documents and incorporating stronger LLMs 1 .
Shudong Liu 0004, Zhaocong Li, Xuebo Liu 0002, Runzhe Zhan, Derek F. Wong, Lidia S. Chao, Min Zhang 0005
EMNLP3
2024 Curriculum Consistency Learning for Conditional Sentence Generation
abstract
Consistency learning (CL) has proven to be a valuable technique for improving the robustness of models in conditional sentence generation (CSG) tasks by ensuring stable predictions across various input data forms.However, models augmented with CL often face challenges in optimizing consistency features, which can detract from their efficiency and effectiveness.To address these challenges, we introduce Curriculum Consistency Learning (CCL), a novel strategy that guides models to learn consistency in alignment with their current capacity to differentiate between features.CCL is designed around the inherent aspects of CL-related losses, promoting task independence and simplifying implementation.Implemented across four representative CSG tasks, including instruction tuning (IT) for large language models and machine translation (MT) in three modalities (text, speech, and vision), CCL demonstrates marked improvements.Specifically, it delivers +2.0 average accuracy point improvement compared with vanilla IT and an average increase of +0.7 in COMET scores over traditional CL methods in MT tasks.Our comprehensive analysis further indicates that models utilizing CCL are particularly adept at managing complex instances, showcasing the effectiveness and efficiency of CCL in improving CSG models.Code and scripts are available at https://github.com/xinxinxing/ Curriculum-Consistency-Learning.
Liangxin Liu, Xuebo Liu 0002, Lian Lian, Shengjun Cheng, Jun Rao, Tengfei Yu, Hexuan Deng, Min Zhang 0005
EMNLP2
2024 CommonIT: Commonality-Aware Instruction Tuning for Large Language Models via Data Partitions
abstract
With instruction tuning, Large Language Models (LLMs) can enhance their ability to adhere to commands.Diverging from most works focusing on data mixing, our study concentrates on enhancing the model's capabilities from the perspective of data sampling during training.Drawing inspiration from the human learning process, where it is generally easier to master solutions to similar topics through focused practice on a single type of topic, we introduce a novel instruction tuning strategy termed Com-monIT: Commonality-aware Instruction Tuning.Specifically, we cluster instruction datasets into distinct groups with three proposed metrics (TASK, EMBEDDING and LENGTH).We ensure each training mini-batch, or "partition", consists solely of data from a single group, which brings about both data randomness across mini-batches and intra-batch data similarity.Rigorous testing on LLaMa models demonstrates CommonIT's effectiveness in enhancing the instruction-following capabilities of LLMs through IT datasets (FLAN, CoT, and Alpaca) and models (LLaMa2-7B, Qwen2-7B, LLaMa 13B, and BLOOM 7B).CommonIT consistently boosts an average improvement of 2.1% on the general domain (i.e., the average score of Knowledge, Reasoning, Multilinguality and Coding) with the LENGTH metric, and 5.2% on the special domain (i.e., GSM, Openfunctions and Code) with the TASK metric, and 3.8% on the specific tasks (i.e., MMLU) with the EMBEDDING metric.Code is available at https://github.com/raojay7/CommonIT.
Jun Rao, Xuebo Liu 0002, Lian Lian, Shengjun Cheng, Yunjie Liao, Min Zhang 0005
EMNLP2
2024 Self-Powered LLM Modality Expansion for Large Speech-Text Models
abstract
Large language models (LLMs) exhibit remarkable performance across diverse tasks, indicating their potential for expansion into large speech-text models (LSMs) by integrating speech capabilities.Although unified speech-text pre-training and multimodal data instruction-tuning offer considerable benefits, these methods generally entail significant resource demands and tend to overfit specific tasks.This study aims to refine the use of speech datasets for LSM training by addressing the limitations of vanilla instruction tuning.We explore the instruction-following dynamics within LSMs, identifying a critical issue termed speech anchor bias-a tendency for LSMs to over-rely on speech inputs, mistakenly interpreting the entire speech modality as directives, thereby neglecting textual instructions.To counteract this bias, we introduce a self-powered LSM that leverages augmented automatic speech recognition data generated by the model itself for more effective instruction tuning.Our experiments across a range of speech-based tasks demonstrate that selfpowered LSM mitigates speech anchor bias and improves the fusion of speech and text modalities in LSMs.
Tengfei Yu, Xuebo Liu 0002, Zhiyi Hou, Liang Ding 0006, Dacheng Tao, Min Zhang 0005
EMNLP2
2024 NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates
abstract
Despite their remarkable abilities in various tasks, large language models (LLMs) still struggle with real-time information (e.g., new facts and terms) due to the knowledge cutoff in their development process. However, existing benchmarks focus on outdated content and limited fields, facing difficulties in real-time updating and leaving new terms unexplored. To address this problem, we propose an adaptive benchmark, NewTerm, for real-time evaluation of new terms. We design a highly automated construction method to ensure high-quality benchmark construction with minimal human effort, allowing flexible updates for real-time information. Empirical results on various LLMs demonstrate over 20% performance reduction caused by new terms. Additionally, while updates to the knowledge cutoff of LLMs can cover some of the new terms, they are unable to generalize to more distant new terms. We also analyze which types of terms are more challenging and why LLMs struggle with new terms, paving the way for future research. Finally, we construct NewTerm 2022 and 2023 to evaluate the new terms updated each year and will continue updating annually. The benchmark and codes can be found at https://anonymous.4open.science/r/NewTerms.
Hexuan Deng, Wenxiang Jiao, Xuebo Liu 0002, Min Zhang 0005, Zhaopeng Tu
NeurIPS3
2024 SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
abstract
Instruction tuning (IT) is crucial to tailoring large language models (LLMs) towards human-centric interactions. Recent advancements have shown that the careful selection of a small, high-quality subset of IT data can significantly enhance the performance of LLMs. Despite this, common approaches often rely on additional models or data, which increases costs and limits widespread adoption. In this work, we propose a novel approach, termed $\textit{SelectIT}$, that capitalizes on the foundational capabilities of the LLM itself. Specifically, we exploit the intrinsic uncertainty present in LLMs to more effectively select high-quality IT data, without the need for extra resources. Furthermore, we introduce a curated IT dataset, the $\textit{Selective Alpaca}$, created by applying SelectIT to the Alpaca-GPT4 dataset. Empirical results demonstrate that IT using Selective Alpaca leads to substantial model ability enhancement. The robustness of SelectIT has also been corroborated in various foundation models and domain-specific tasks. Our findings suggest that longer and more computationally intensive IT data may serve as superior sources of IT, offering valuable insights for future research in this area. Data, code, and scripts are freely available at https://github.com/Blue-Raincoat/SelectIT.
Liangxin Liu, Xuebo Liu 0002, Derek F. Wong, Dongfang Li 0002, Baotian Hu, Min Zhang 0005
NeurIPS2
2024 Understanding and Improving Low-Resource Neural Machine Translation with Shallow Features
Xuebo Liu 0002, Derek F. Wong, Yuchu Lin, Runzhe Zhan, Lidia S. Chao, Min Zhang 0005
NLPCC (3)2
2024 Parameter-Efficient and Student-Friendly Knowledge Distillation
abstract
Pre-trained models are frequently employed in multimodal learning. However, these models have too many parameters and need too much effort to fine-tune the downstream tasks. Knowledge distillation (KD) is a method to transfer knowledge using the soft label from this pre-trained teacher model to a smaller student, where the parameters of the teacher are fixed (or partially) during training. Recent studies show that this mode may cause difficulties in knowledge transfer due to the mismatched model capacities. To alleviate the mismatch problem, adjustment of temperature parameters, label smoothing and teacher-student joint training methods (online distillation) to smooth the soft label of a teacher network, have been proposed. But those methods rarely explain the effect of smoothed soft labels to enhance the KD performance. The main contributions of our work are the discovery, analysis, and validation of the effect of the smoothed soft label and a less time-consuming and adaptive transfer of the pre-trained teacher's knowledge method, namely PESF-KD by adaptive tuning soft labels of the teacher network. Technically, we first mathematically formulate the mismatch as the sharpness gap between teacher's and student's predictive distributions, where we show such a gap can be narrowed with the appropriate smoothness of the soft label. Then, we introduce an adapter module for the teacher and only update the adapter to obtain soft labels with appropriate smoothness. Experiments on various benchmarks including CV and NLP show that PESF-KD can significantly reduce the training cost while obtaining competitive results compared to advanced online distillation methods.
Jun Rao, Xv Meng, Liang Ding 0006, Shuhan Qi, Xuebo Liu 0002, Min Zhang 0005, Dacheng Tao
IEEE Trans. Multim.5
2023 Improving Simultaneous Machine Translation with Monolingual Data
abstract
Simultaneous machine translation (SiMT) is usually done via sequence-level knowledge distillation (Seq-KD) from a full-sentence neural machine translation (NMT) model. However, there is still a significant performance gap between NMT and SiMT. In this work, we propose to leverage monolingual data to improve SiMT, which trains a SiMT student on the combination of bilingual data and external monolingual data distilled by Seq-KD. Preliminary experiments on En-Zh and En-Ja news domain corpora demonstrate that monolingual data can significantly improve translation quality (e.g., +3.15 BLEU on En-Zh). Inspired by the behavior of human simultaneous interpreters, we propose a novel monolingual sampling strategy for SiMT, considering both chunk length and monotonicity. Experimental results show that our sampling strategy consistently outperforms the random sampling strategy (and other conventional typical NMT monolingual sampling strategies) by avoiding the key problem of SiMT -- hallucination, and has better scalability. We achieve +0.72 BLEU improvements on average against random sampling on En-Zh and En-Ja. Data and codes can be found at https://github.com/hexuandeng/Mono4SiMT.
Hexuan Deng, Liang Ding 0006, Xuebo Liu 0002, Meishan Zhang, Dacheng Tao, Min Zhang 0005
AAAI3
2023 Revisiting Commonsense Reasoning in Machine Translation: Training, Evaluation and Challenge
abstract
The ability of commonsense reasoning (CR) decides whether a neural machine translation (NMT) model can move beyond pattern recognition.Despite the rapid advancement of NMT and the use of pretraining to enhance NMT models, research on CR in NMT is still in its infancy, leaving much to be explored in terms of effectively training NMT models with high CR abilities and devising accurate automatic evaluation metrics.This paper presents a comprehensive study aimed at expanding the understanding of CR in NMT.For the training, we confirm the effectiveness of incorporating pretrained knowledge into NMT models and subsequently utilizing these models as robust testbeds for investigating CR in NMT.For the evaluation, we propose a novel entity-aware evaluation method that takes into account both the NMT candidate and important entities in the candidate, which is more aligned with human judgement.Based on the strong testbed and evaluation methods, we identify challenges in training NMT models with high CR abilities and suggest directions for further unlabeled data utilization and model design.We hope that our methods and findings will contribute to advancing the research of CR in NMT.
Xuebo Liu 0002, Derek F. Wong, Runzhe Zhan, Liangxuan Yu, Min Zhang 0005
ACL (1)1
2023 TemplateGEC: Improving Grammatical Error Correction with Detection Template
abstract
Yinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong, Derek F. Wong, Yang Gao, Heyan Huang, Min Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xuebo Liu 0002, Shuo Wang 0013, Peiyuan Gong, Derek F. Wong, Yang Gao 0016, Heyan Huang, Min Zhang 0005
ACL (1)2
2023 kNN-TL: k-Nearest-Neighbor Transfer Learning for Low-Resource Neural Machine Translation
abstract
Shudong Liu, Xuebo Liu, Derek F. Wong, Zhaocong Li, Wenxiang Jiao, Lidia S. Chao, Min Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shudong Liu 0004, Xuebo Liu 0002, Derek F. Wong, Zhaocong Li, Wenxiang Jiao, Lidia S. Chao, Min Zhang 0005
ACL (1)2
2023 Test-time Adaptation for Machine Translation Evaluation by Uncertainty Minimization
abstract
The neural metrics recently received considerable attention from the research community in the automatic evaluation of machine translation.Unlike text-based metrics that have interpretable and consistent evaluation mechanisms for various data sources, the reliability of neural metrics in assessing out-of-distribution data remains a concern due to the disparity between training data and real-world data.This paper aims to address the inference bias of neural metrics through uncertainty minimization during test time, without requiring additional data.Our proposed method comprises three steps: uncertainty estimation, test-time adaptation, and inference.Specifically, the model employs the prediction uncertainty of the current data as a signal to update a small fraction of parameters during test time and subsequently refine the prediction through optimization.To validate our approach, we apply the proposed method to three representative models and conduct experiments on the WMT21 benchmarks.The results obtained from both in-domain and out-of-distribution evaluations consistently demonstrate improvements in correlation performance across different models.Furthermore, we provide evidence that the proposed method effectively reduces model uncertainty.The code is publicly available at https://github.com/NLP2CT/TaU.
Runzhe Zhan, Xuebo Liu 0002, Derek F. Wong, Cuilian Zhang, Lidia S. Chao, Min Zhang 0005
ACL (1)2
2023 Revisiting Token Dropping Strategy in Efficient BERT Pretraining
abstract
Token dropping is a recently-proposed strategy to speed up the pretraining of masked language models, such as BERT, by skipping the computation of a subset of the input tokens at several middle layers.It can effectively reduce the training time without degrading much performance on downstream tasks.However, we empirically find that token dropping is prone to a semantic loss problem and falls short in handling semantic-intense tasks ( §2).Motivated by this, we propose a simple yet effective semantic-consistent learning method (SCTD) to improve the token dropping.SCTD aims to encourage the model to learn how to preserve the semantic information in the representation space.Extensive experiments on 12 tasks show that, with the help of our SCTD, token dropping can achieve consistent and significant performance gains across all task types and model sizes.More encouragingly, SCTD saves up to 57% of pretraining time and brings up to +1.56% average improvement over the vanilla token dropping.
Qihuang Zhong, Liang Ding 0006, Juhua Liu, Xuebo Liu 0002, Min Zhang 0005, Bo Du 0001, Dacheng Tao
ACL (1)4
2023 Can LMs Generalize to Future Data? An Empirical Analysis on Text Summarization
abstract
Recent pre-trained language models (PLMs) achieve promising results in existing abstractive summarization datasets.However, existing summarization benchmarks overlap in time with the standard pre-training corpora and finetuning datasets.Hence, the strong performance of PLMs may rely on the parametric knowledge that is memorized during pre-training and fine-tuning.Moreover, the knowledge memorized by PLMs may quickly become outdated, which affects the generalization performance of PLMs on future data.In this work, we propose TEMPOSUM, a novel benchmark that contains data samples from 2010 to 2022, to understand the temporal generalization ability of abstractive summarization models.Through extensive human evaluation, we show that parametric knowledge stored in summarization models significantly affects the faithfulness of the generated summaries on future data.Moreover, existing faithfulness enhancement methods cannot reliably improve the faithfulness of summarization models on future data.Finally, we discuss several recommendations to the research community on how to evaluate and improve the temporal generalization capability of text summarization models. 1
Chi Seng Cheang, Hou Pong Chan, Derek F. Wong, Xuebo Liu 0002, Zhaocong Li, Shudong Liu 0004, Lidia S. Chao
EMNLP4
2023 Clustering Pseudo Language Family in Multilingual Translation Models with Fisher Information Matrix
abstract
In multilingual translation research, the comprehension and utilization of language families are of paramount importance.Nevertheless, clustering languages based solely on their ancestral families can yield suboptimal results due to variations in the datasets employed during the model's training phase.To mitigate this challenge, we introduce an innovative method that leverages the fisher information matrix (FIM) to cluster language families, anchored on the multilingual translation model's characteristics.We hypothesize that language pairs with similar effects on model parameters exhibit a considerable degree of linguistic congruence and should thus be grouped cohesively.This concept has led us to define pseudo language families.We provide an in-depth discussion regarding the inception and application of these pseudo language families.Empirical evaluations reveal that employing these pseudo language families enhances performance over conventional language families in adapting a multilingual translation model to unfamiliar language pairs.The proposed methodology may also be extended to scenarios requiring language similarity measurements.The source code and associated scripts can be accessed at https://github.com/ecoli-hit/PseudoFamily.
Xuebo Liu 0002, Min Zhang 0005
EMNLP2
2023 PromptST: Abstract Prompt Learning for End-to-End Speech Translation
abstract
An end-to-end speech-to-text (S2T) translation model is usually initialized from a pretrained speech recognition encoder and a pretrained text-to-text (T2T) translation decoder.Although this straightforward setting has been shown empirically successful, there do not exist clear answers to the research questions: 1) how are speech and text modalities fused in S2T model and 2) how to better fuse the two modalities?In this paper, we take the first step toward understanding the fusion of speech and text features in S2T model.We first design and release a 10GB linguistic probing benchmark, namely Speech-Senteval, to investigate the acoustic and linguistic behaviors of S2T models.Preliminary analysis reveals that the uppermost encoder layers of the S2T model can not learn linguistic knowledge efficiently, which is crucial for accurate translation.Based on the finding, we further propose a simple plug-in prompt-learning strategy on the uppermost encoder layers to broaden the abstract representation power of the encoder of S2T models.We call such a promptenhanced S2T model PromptST.Experimental results on four widely-used S2T datasets show that PromptST can deliver significant improvements over a strong baseline by capturing richer linguistic knowledge.Benchmarks,
Tengfei Yu, Liang Ding 0006, Xuebo Liu 0002, Kehai Chen, Meishan Zhang, Dacheng Tao, Min Zhang 0005
EMNLP3
2022 ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation
abstract
Bei Li, Quan Du, Tao Zhou, Yi Jing, Shuhan Zhou, Xin Zeng, Tong Xiao, JingBo Zhu, Xuebo Liu, Min Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Quan Du, Yi Jing, Shuhan Zhou, Tong Xiao 0001, Xuebo Liu 0002, Min Zhang 0005
ACL (1)9
2022 Revisiting Grammatical Error Correction Evaluation and Beyond
abstract
Pretraining-based (PT-based) automatic evaluation metrics (e.g., BERTScore and BARTScore) have been widely used in several sentence generation tasks (e.g., machine translation and text summarization) due to their better correlation with human judgments over traditional overlapbased methods.Although PT-based methods have become the de facto standard for training grammatical error correction (GEC) systems, GEC evaluation still does not benefit from pretrained knowledge.This paper takes the first step towards understanding and improving GEC evaluation with pretraining.We first find that arbitrarily applying PT-based metrics to GEC evaluation brings unsatisfactory correlation results because of the excessive attention to inessential systems outputs (e.g., unchanged parts).To alleviate the limitation, we propose a novel GEC evaluation metric to achieve the best of both worlds, namely PT-M 2 , which only uses PT-based metrics to score those corrected parts.Experimental results on the CoNLL14 evaluation task show that PT-M 2 significantly outperforms existing methods, achieving a new state-of-the-art result of 0.949 Pearson correlation.Further analysis reveals that PT-M 2 is robust to evaluate competitive GEC systems.
Peiyuan Gong, Xuebo Liu 0002, Heyan Huang, Min Zhang 0005
EMNLP2
2022 ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine Translation
abstract
Transfer learning is a simple and powerful method that can be used to boost model performance of low-resource neural machine translation (NMT).Existing transfer learning methods for NMT are static, which simply transfer knowledge from a parent model to a child model once via parameter initialization.In this paper, we propose a novel transfer learning method for NMT, namely ConsistTL, which can continuously transfer knowledge from the parent model during the training of the child model.Specifically, for each training instance of the child model, ConsistTL constructs the semantically-equivalent instance for the parent model and encourages prediction consistency between the parent and child for this instance, which is equivalent to the child model learning each instance under the guidance of the parent model.Experimental results on five low-resource NMT tasks demonstrate that ConsistTL results in significant improvements over strong transfer learning baselines, with a gain up to 1.7 BLEU over the existing backtranslation model on the widely-used WMT17 Turkish-English benchmark.Further analysis reveals that ConsistTL can improve the inference calibration of the child model.Code and scripts are freely available at https://github. com/NLP2CT/ConsistTL.
Zhaocong Li, Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao, Min Zhang 0005
EMNLP2
2022 Breaking the Representation Bottleneck of Chinese Characters: Neural Machine Translation with Stroke Sequence Modeling
abstract
Existing research generally treats Chinese character as a minimum unit for representation.However, such Chinese character representation will suffer two bottlenecks: 1) Learning bottleneck, the learning cannot benefit from its rich internal features (e.g., radicals and strokes); and 2) Parameter bottleneck, each individual character has to be represented by a unique vector.In this paper, we introduce a novel representation method for Chinese characters to break the bottlenecks, namely StrokeNet, which represents a Chinese character by a Latinized stroke sequence (e.g., "凹(concave)" to "ajaie" and "凸(convex)" to "aeaqe").Specifically, StrokeNet maps each stroke to a specific Latin character, thus allowing similar Chinese characters to have similar Latin representations.With the introduction of StrokeNet to neural machine translation (NMT), many powerful but not applicable techniques to non-Latin languages (e.g., shared subword vocabulary learning and ciphertextbased data augmentation) can now be perfectly implemented.Experiments on the widelyused NIST Chinese-English, WMT17 Chinese-English and IWSLT17 Japanese-English NMT tasks show that StrokeNet can provide a significant performance boost over the strong baselines with fewer model parameters, achieving 26.5 BLEU on the WMT17 Chinese-English task which is better than any previously reported results without using monolingual data.
Xuebo Liu 0002, Min Zhang 0005
EMNLP2
2021 Meta-Curriculum Learning for Domain Adaptation in Neural Machine Translation
abstract
Meta-learning has been sufficiently validated to be beneficial for low-resource neural machine translation (NMT). However, we find that meta-trained NMT fails to improve the translation performance of the domain unseen at the meta-training stage. In this paper, we aim to alleviate this issue by proposing a novel meta-curriculum learning for domain adaptation in NMT. During meta-training, the NMT first learns the similar curricula from each domain to avoid falling into a bad local optimum early, and finally learns the curricula of individualities to improve the model robustness for learning domain-specific knowledge. Experimental results on 10 different low-resource domains show that meta-curriculum learning can improve the translation performance of both familiar and unfamiliar domains. All the codes and data are freely available at https://github.com/NLP2CT/Meta-Curriculum.
Runzhe Zhan, Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao
AAAI2
2021 Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive Translation
abstract
Liang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong, Dacheng Tao, Zhaopeng Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Liang Ding 0006, Longyue Wang, Xuebo Liu 0002, Derek F. Wong, Dacheng Tao, Zhaopeng Tu
ACL/IJCNLP (1)3
2021 Understanding and Improving Encoder Layer Fusion in Sequence-to-Sequence Learning
Xuebo Liu 0002, Longyue Wang, Derek F. Wong, Liang Ding 0006, Lidia S. Chao, Zhaopeng Tu
ICLR1
2021 Understanding and Improving Lexical Choice in Non-Autoregressive Translation
Liang Ding 0006, Longyue Wang, Xuebo Liu 0002, Derek F. Wong, Dacheng Tao, Zhaopeng Tu
ICLR3
2021 Exploiting Translation Model for Parallel Corpus Mining
abstract
Parallel corpus mining (PCM) is beneficial for many corpus-based natural language processing tasks, e.g., machine translation and bilingual dictionary induction, especially in low-resource languages and domains. It relies heavily on cross-lingual representations to model the interdependencies between different languages and determine whether sentences are parallel or not. In this paper, we take the first step towards exploiting the multilingual Transformer translation model to produce expressive sentence representations for PCM. Since the traditional Transformer lacks an immediate sentence representation, we pool the output representation of the encoder as the sentence representation, which is further optimized as a part of the training flow of the translation model. Experiments conducted on the BUCC PCM task show that the proposed method improves mining performance over the existing methods with the assistance of the pre-trained multilingual BERT. To further test the usability of the proposed method, we mine parallel sentences from public resources and find that the mined sentences can indeed enhance low-resource machine translation.
Chongman Leong, Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Norm-Based Curriculum Learning for Neural Machine Translation
abstract
A neural machine translation (NMT) system is expensive to train, especially with highresource settings.As the NMT architectures become deeper and wider, this issue gets worse and worse.In this paper, we aim to improve the efficiency of training an NMT by introducing a novel norm-based curriculum learning method.We use the norm (aka length or module) of a word embedding as a measure of 1) the difficulty of the sentence, 2) the competence of the model, and 3) the weight of the sentence.The normbased sentence difficulty takes the advantages of both linguistically motivated and modelbased sentence difficulties.It is easy to determine and contains learning-dependent features.The norm-based model competence makes NMT learn the curriculum in a fully automated way, while the norm-based sentence weight further enhances the learning of the vector representation of the NMT.Experimental results for the WMT'14 English-German and WMT'17 Chinese-English translation tasks demonstrate that the proposed method outperforms strong baselines in terms of BLEU score (+1.17/+1.56)and training speedup (2.22x/3.33x).
Xuebo Liu 0002, Houtim Lai, Derek F. Wong, Lidia S. Chao
ACL1
2019 Shared-Private Bilingual Word Embeddings for Neural Machine Translation
abstract
Word embedding is central to neural machine translation (NMT), which has attracted intensive research interest in recent years.In NMT, the source embedding plays the role of the entrance while the target embedding acts as the terminal.These layers occupy most of the model parameters for representation learning.Furthermore, they indirectly interface via a soft-attention mechanism, which makes them comparatively isolated.In this paper, we propose shared-private bilingual word embeddings, which give a closer relationship between the source and target embeddings, and which also reduce the number of model parameters.For similar source and target words, their embeddings tend to share a part of the features and they cooperatively learn these common representation units.Experiments on 5 language pairs belonging to 6 different language families and written in 5 different alphabets demonstrate that the proposed model provides a significant performance boost over the strong baselines with dramatically fewer model parameters.
Xuebo Liu 0002, Derek F. Wong, Yang Liu 0005, Lidia S. Chao, Tong Xiao 0001
ACL (1)1
2019 Latent Attribute Based Hierarchical Decoder for Neural Machine Translation
abstract
Neural machine translation (NMT) has achieved state-of-the-art performance in many translation tasks. However, because the computational cost increases with the size of the search space for predicting the target words, the translation quality of NMT is constrained by the limited vocabulary. To alleviate this problem, we propose a novel dynamic hierarchical decoder for NMT to utilize all of the target words in the training and decoding process. In the proposed model, a target word is represented by two latent attribute vectors rather than a word vector. The model is trained to dynamically put together those words that share similar linguistic attributes. The prediction of a target word is, therefore, turned into the prediction of attribute vectors, where the $\mathrm{softmax}$ functions are performed at the attribute level. This greatly reduces the model size and the decoding time. Our experimental results demonstrate that the proposed model significantly outperforms the NMT baselines in both Chinese-English and English-German translation tasks.
Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.1