Fandong Meng

dblp:117/4056 · DBLP profile ↗
← Back
122ranked-venue papers
7as first author
93since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 118 · 7 first-author · 90 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GRAM-R²: Self-Training Generative Foundation Reward Models for Reward Reasoning
abstract
Major progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs to generalist reward models. Despite this trend, developing effective reward models remains a fundamental challenge: the heavy reliance on large-scale labeled preference data. Pre-training on abundant unlabeled data offers a promising direction, but existing approaches fall short in instilling explicit reasoning capabilities into reward models. To bridge this gap, we propose a self-training approach that can leverage unlabeled data to scale up reward reasoning in reward models. Based on this approach, we develop GRAM-R² a generative reward model trained to produce not only preference labels but also accompanying reward rationales. GRAM-R² can serve as a foundation model for reward reasoning and can be applied to a wide range of tasks with minimal or no additional fine-tuning. It can support downstream applications such as policy optimization and task-specific reward tuning. Experiments on response ranking, task adaptation, and reinforcement learning from human feedback demonstrate that GRAM-R² consistently delivers strong performance, outperforming several strong discriminative and generative baselines.
Chenglong Wang 0002, Yongyu Mu, Yifu Huo, Jiali Zeng, Murun Yang, Xiaoyang Hao, Chunliang Zhang, Fandong Meng, Tong Xiao 0001
AAAI11
2026 Figure It Out: Improve the Frontier of Reasoning with Executable Visual States
abstract
Complex reasoning problems often involve implicit spatial and geometric relationships that are not explicitly encoded in text. While recent reasoning models perform well across many domains, purely text-based reasoning struggles to capture structural constraints in complex settings. In this paper, we introduce FIGR, which integrates executable visual construction into multi-turn reasoning via end-to-end reinforcement learning. Rather than relying solely on textual chains of thought, FIGR externalizes intermediate hypotheses by generating executable code that constructs diagrams within the reasoning loop. An adaptive reward mechanism selectively regulates when visual construction is invoked, enabling more consistent reasoning over latent global properties that are difficult to infer from text alone. Experiments on seven challenging mathematical benchmarks demonstrate that FIGR outperforms strong text-only chain-of-thought baselines, improving the base model by 13.12% on AIME 2025 and 11.00% on BeyondAIME. These results highlight the effectiveness of precise, controllable figure construction of FIGR in enhancing complex reasoning ability.
Meiqi Chen 0001, Fandong Meng, Jie Zhou 0016
ACL (1)2
2026 APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
abstract
Yuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao, Weilin Zhao, Ao Sun, Ziqi Yuan, Hao Zhou, Fandong Meng, Zhiyuan Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Hao Zhou 0012, Fandong Meng, Zhiyuan Liu 0001
ACL (1)9
2026 Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
abstract
Zhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lean Wang, Yuanxin Liu, Lei Li 0009, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
ACL (1)6
2026 Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning
abstract
Xue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang, Kaiyu Huang, Yufeng Chen, Xu Jinan, Jie Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yunlong Liang, Fandong Meng, Songming Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
ACL (1)3
2026 Cross-layer Attention Sharing for Pre-trained Large Language Models
abstract
Abstract To enhance the efficiency of the attention mechanism within large language models (LLMs), previous works primarily compress the Key-Value cache or group attention heads, while largely overlooking redundancy between layers. Our comprehensive analyses across various LLMs show that highly similar attention patterns persist within most layers. It’s intuitive to reduce the redundancy by sharing attention weights across layers. However, further analysis reveals two challenges: (1) Directly sharing the weight matrix without carefully rearranging the attention heads proves to be ineffective; (2) Shallow layers are vulnerable to small deviations in attention weights. Driven by these insights, we introduce LiSA, a lightweight substitute for self-attention in well-trained LLMs. LiSA employs tiny feed-forward networks to align attention heads between adjacent layers and low-rank matrices to approximate differences in layer-wise attention weights. Evaluations encompassing 13 typical benchmarks demonstrate that LiSA maintains high response quality in terms of accuracy and perplexity while reducing redundant attention calculations within 53% −84% of the total layers. Our implementations of LiSA achieve a 6 × compression of Q and K matrices within the attention mechanism, with maximum throughput improvements 19.5%, 32.3%, and 40.1% for LLaMA3-8B, LLaMA2-7B, and LLaMA2-13B, respectively. Our code is available at https://github.com/takagi97/lisa.
Yongyu Mu, Yuzhang Wu, Yuchun Fan, Chenglong Wang 0002, Jiali Zeng, Qiaozhi He, Murun Yang, Fandong Meng, Jie Zhou 0016, Tong Xiao 0001
Trans. Assoc. Comput. Linguistics9
2026 DeepTrans: Deep Reasoning Translation via Reinforcement Learning
abstract
Abstract Recently, deep reasoning LLMs (e.g., OpenAI o1 and DeepSeek-R1) have shown promising performance in various downstream tasks. Free translation is an important and interesting task in the multilingual world, which requires going beyond word-for-word translation. However, the task is still under-explored in deep reasoning LLMs. In this paper, we introduce DeepTrans, a deep reasoning translation model that learns free translation via reinforcement learning (RL). Specifically, we carefully build a reward model with pre-defined scoring criteria on both the translation results and the thought processes. The reward model teaches DeepTrans how to think and free-translate the given sentences during RL. Besides, our RL training does not need any labeled translations, avoiding the human-intensive annotation or resource-intensive data synthesis. Experimental results show the effectiveness of DeepTrans. Using Qwen2.5-7B as the backbone, DeepTrans improves performance by 16.3% in literature translation, and outperforms strong deep reasoning LLMs. Moreover, we summarize the failures and interesting findings during our RL exploration. We hope this work could inspire other researchers in free translation.1
Jiaan Wang, Fandong Meng, Jie Zhou 0016
Trans. Assoc. Comput. Linguistics2
2025 THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation
abstract
The sparse Mixture-of-Experts (MoE) has achieved significant progress for neural machine translation (NMT). However, there exist two limitations in current MoE solutions which may lead to sub-optimal performance: 1) they directly use the task knowledge of NMT into MoE (e.g., domain/linguistics-specific knowledge), which are generally unavailable at practical application and neglect the naturally grouped domain/linguistic properties; 2) the expert selection only depends on the localized token representation without considering the context, which fully grasps the state of each token in a global view. To address the above limitations, we propose THOR-MoE via arming the MoE with hierarchical task-guided and context-responsive routing policies. Specifically, it 1) firstly predicts the domain/language label and then extracts mixed domain/language representation to allocate task-level experts in a hierarchical manner; 2) injects the context information to enhance the token routing from the pre-selected task-level experts set, which can help each token to be accurately routed to more specialized and suitable experts. Extensive experiments on multi-domain translation and multilingual translation benchmarks with different architectures consistently demonstrate the superior performance of THOR-MoE. Additionally, the THOR-MoE operates as a plug-and-play module compatible with existing Top-(CITATION) or Top-(CITATION) routing schemes, ensuring broad applicability across diverse MoE architectures. For instance, compared with vanilla Top- (CITATION) routing, the context-aware manner can achieve an average improvement of 0.75 BLEU with less than 22% activated parameters on multi-domain translation tasks.
Yunlong Liang, Fandong Meng, Jie Zhou 0016
ACL (1)2
2025 PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
abstract
Kun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kun Ouyang, Yuanxin Liu, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
ACL (1)6
2025 An Empirical Study of Many-to-Many Summarization with Large Language Models
abstract
Jiaan Wang, Fandong Meng, Zengkui Sun, Yunlong Liang, Yuxuan Cao, Jiarong Xu, Haoxiang Shi, Jie Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jiaan Wang, Fandong Meng, Zengkui Sun, Yunlong Liang, Jiarong Xu, Haoxiang Shi, Jie Zhou 0016
ACL (1)2
2025 Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts
abstract
Continually expanding new languages for existing large language models (LLMs) is a promising yet challenging approach to building powerful multilingual LLMs.The biggest challenge is to make the model continuously learn new languages while preserving the proficient ability of old languages.To achieve this, recent work utilizes the Mixture-of-Experts (MoE) architecture to expand new languages by adding new experts and avoid catastrophic forgetting of old languages by routing corresponding tokens to the original model backbone (old experts).Although intuitive, this kind of method is parameter-costly when expanding new languages and still inevitably impacts the performance of old languages.To address these limitations, we analyze the language characteristics of different layers in LLMs and propose a layer-wise expert allocation algorithm (LayerMoE) to determine the appropriate number of new experts for each layer.Specifically, we find different layers in LLMs exhibit different representation similarities between languages and then utilize the similarity as the indicator to allocate experts for each layer, i.e., the higher similarity, the fewer experts.Additionally, to further mitigate the forgetting of old languages, we add a classifier in front of the router network on the layers with higher similarity to guide the routing of old language tokens.Experimental results show that our method outperforms the previous state-of-the-art baseline with 60% fewer experts in the single-expansion setting and with 33.3% fewer experts in the lifelong-expansion setting, demonstrating the effectiveness of our method.
Yunlong Liang, Fandong Meng, Songming Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
ACL (1)3
2025 Advancing SMoE for Continuous Domain Adaptation of MLLMs: Adaptive Router and Domain-Specific Loss
abstract
Recent studies have explored Continual Instruction Tuning (CIT) in Multimodal Large Language Models (MLLMs), with a primary focus on Task-incremental CIT, where MLLMs are required to continuously acquire new tasks.However, the more practical and challenging Domain-incremental CIT, focused on the continual adaptation of MLLMs to new domains, remains underexplored.In this paper, we propose a new Sparse Mixture of Expert (SMoE) based method for domain-incremental CIT in MLLMs.During training, we learn a domainspecific SMoE module for each new domain in every FFN sub-layer of MLLMs, preventing catastrophic forgetting caused by inter-domain conflicts.Moreover, we equip the SMoE module with a domain-specific autoregressive loss (DSAL), which is used to identify the most suitable SMoE module for processing each test instruction during inference.To further enhance the SMoE module's ability to learn domain knowledge, we design an adaptive thresholdbased router (AT-Router) that allocates computing resources (experts) to instruction tokens based on their importance.Finally, we establish a new benchmark to evaluate the efficacy of our method and advance future research.Extensive experiments show that our method consistently outperforms all competitive baselines.
Ziyao Lu, Fandong Meng, Jie Zhou 0016, Jinsong Su
ACL (1)3
2025 A Self-Denoising Model for Robust Few-Shot Relation Extraction
abstract
The few-shot relation extraction (FSRE) aims at enhancing the model's generalization to new relations with a few labeled instances.Existing studies usually adopt prototype networks (Pro-toNets) for FSRE and assume that the support set, adapting the model to new relations, only contains accurately labeled support instances.However, this assumption is often unrealistic, as even carefully-annotated datasets commonly contain mislabeled instances.In this paper, we first conduct a preliminary study that reveals the high sensitivity of ProtoNets to noisy labels in the support set.Moreover, we find that fully leveraging mislabeled support instances is crucial for enhancing the FSRE model robustness.Thus, we propose a self-denoising model for FSRE, designed to improve model robustness by automatically correcting mislabeled support instances.It comprises two core components: 1) a label correction module (LCM), used to correct noisy labels of support instances based on the distances between them in the embedding space, and 2) a relation classification module (RCM), aimed to achieve more accurate predictions for new relations based on the corrected labels produced by LCM.Moreover, we propose a feedback-based training strategy, which focuses on training LCM and RCM to synergistically handle noisy labels in support set.Experimental results on two public datasets confirm the robustness of our model.Particularly, even in scenarios without noisy labels, our model significantly outperforms all baselines.
Ziyao Lu, Fandong Meng, Jie Zhou 0016, Jinsong Su
ACL (1)4
2025 Multilingual Knowledge Editing with Language-Agnostic Factual Neurons
abstract
Multilingual knowledge editing (MKE) aims to simultaneously update factual knowledge across multiple languages within large language models (LLMs). Previous research indicates that the same knowledge across different languages within LLMs exhibits a degree of shareability. However, most existing MKE methods overlook the connections of the same knowledge between different languages, resulting in knowledge conflicts and limited edit performance. To address this issue, we first investigate how LLMs process multilingual factual knowledge and discover that the same factual knowledge in different languages generally activates a shared set of neurons, which we call language-agnostic factual neurons (LAFNs). These neurons represent the same factual knowledge shared across languages and imply the semantic connections among multilingual knowledge. Inspired by this finding, we propose a new MKE method by Locating and Updating Language-Agnostic Factual Neurons (LU-LAFNs) to edit multilingual knowledge simultaneously, which avoids knowledge conflicts and thus improves edit performance. Experimental results on Bi-ZsRE and MzsRE benchmarks demonstrate that our method achieves the best edit performance, indicating the effectiveness and importance of modeling the semantic connections among multilingual knowledge.
Yunlong Liang, Fandong Meng, Songming Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
COLING3
2025 ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
abstract
Ziqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Guanbo Wang, Fandong Meng, Jie Zhou, Ju Ren, Yaoxue Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Ziqing Qiao, Yongheng Deng, Jiali Zeng, Guanbo Wang, Fandong Meng, Jie Zhou 0016, Ju Ren 0001, Yaoxue Zhang
EMNLP7
2025 MiniPLM: Knowledge Distillation for Pre-training Language Models
abstract
Knowledge distillation (KD) is widely used to train small, high-performing student language models (LMs) using large teacher LMs. While effective in fine-tuning, KD during pre-training faces efficiency, flexibility, and effectiveness issues. Existing methods either incur high computational costs due to online teacher inference, require tokenization matching between teacher and student LMs, or risk losing the difficulty and diversity of the teacher-generated training data. In this work, we propose **MiniPLM**, a KD framework for pre-training LMs by refining the training data distribution with the teacher LM's knowledge. For efficiency, MiniPLM performs offline teacher inference, allowing KD for multiple student LMs without adding training costs. For flexibility, MiniPLM operates solely on the training corpus, enabling KD across model families. For effectiveness, MiniPLM leverages the differences between large and small LMs to enhance the training data difficulty and diversity, helping student LMs acquire versatile and sophisticated knowledge. Extensive experiments demonstrate that MiniPLM boosts the student LMs' performance on 9 common downstream tasks, improves language modeling capabilities, and reduces pre-training computation. The benefit of MiniPLM extends to larger training scales, evidenced by the scaling curve extrapolation. Further analysis reveals that MiniPLM supports KD across model families and enhances the pre-training data utilization. Our code, data, and models can be found at https://github.com/thu-coai/MiniPLM.
Yuxian Gu, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Minlie Huang
ICLR3
2025 Beyond Next Token Prediction: Patch-Level Training for Large Language Models
abstract
The prohibitive training costs of Large Language Models (LLMs) have emerged as a significant bottleneck in the development of next-generation LLMs. In this paper, we show that it is possible to significantly reduce the training costs of LLMs without sacrificing their performance. Specifically, we introduce patch-level training for LLMs, in which multiple tokens are aggregated into a unit of higher information density, referred to as a `patch', to serve as the fundamental text unit for training LLMs. During patch-level training, we feed the language model shorter sequences of patches and train it to predict the next patch, thereby processing the majority of the training data at a significantly reduced cost. Following this, the model continues token-level training on the remaining training data to align with the inference mode. Experiments on a diverse range of models (370M-2.7B parameters) demonstrate that patch-level training can reduce the overall training costs to 0.5$\times$, without compromising the model performance compared to token-level training. Source code: \url{https://github.com/shaochenze/PatchTrain}.
Chenze Shao, Fandong Meng, Jie Zhou 0016
ICLR2
2025 DelTA: An Online Document-Level Translation Agent Based on Multi-Level Memory
abstract
Large language models (LLMs) have achieved reasonable quality improvements in machine translation (MT). However, most current research on MT-LLMs still faces significant challenges in maintaining translation consistency and accuracy when processing entire documents. In this paper, we introduce DelTA, a Document-levEL Translation Agent designed to overcome these limitations. DelTA features a multi-level memory structure that stores information across various granularities and spans, including Proper Noun Records, Bilingual Summary, Long-Term Memory, and Short-Term Memory, which are continuously retrieved and updated by auxiliary LLM-based components. Experimental results indicate that DelTA significantly outperforms strong baselines in terms of translation consistency and quality across four open/closed-source LLMs and two representative document translation datasets, achieving an increase in consistency scores by up to 4.58 percentage points and in COMET scores by up to 3.16 points on average. DelTA employs a sentence-by-sentence translation strategy, ensuring no sentence omissions and offering a memory-efficient solution compared to the mainstream method. Furthermore, DelTA improves pronoun and context-dependent translation accuracy, and the summary component of the agent also shows promise as a tool for query-based summarization tasks. The code and data of our approach are released at https://github.com/YutongWang1216/DocMTAgent.
Jiali Zeng, Xuebo Liu 0002, Derek F. Wong, Fandong Meng, Jie Zhou 0016, Min Zhang 0005
ICLR5
2025 Continuous Visual Autoregressive Generation via Score Maximization
abstract
Conventional wisdom suggests that autoregressive models are used to process discrete data. When applied to continuous modalities such as visual data, Visual AutoRegressive modeling (VAR) typically resorts to quantization-based approaches to cast the data into a discrete space, which can introduce significant information loss. To tackle this issue, we introduce a Continuous VAR framework that enables direct visual autoregressive generation without vector quantization. The underlying theoretical foundation is strictly proper scoring rules, which provide powerful statistical tools capable of evaluating how well a generative model approximates the true distribution. Within this framework, all we need is to select a strictly proper score and set it as the training objective to optimize. We primarily explore a class of training objectives based on the energy score, which is likelihood-free and thus overcomes the difficulty of making probabilistic predictions in the continuous space. Previous efforts on continuous autoregressive generation, such as GIVT and diffusion loss, can also be derived from our framework using other strictly proper scores. Source code: \url{https://github.com/shaochenze/EAR}.
Chenze Shao, Fandong Meng, Jie Zhou 0016
ICML2
2025 Personalized Language Model Learning on Text Data Without User Identifiers
abstract
In many practical natural language applications, user data are highly sensitive, requiring anonymous uploads of text data from mobile devices to the cloud without user identifiers. However, the absence of user identifiers restricts the ability of cloud-based language models to provide personalized services, which are essential for catering to diverse user needs. The trivial method of replacing an explicit user identifier with a static user embedding as model input still compromises data anonymization. In this work, we propose to let each mobile device maintain a user-specific distribution to dynamically generate user embeddings, thereby breaking the one-to-one mapping between an embedding and a specific user. We further theoretically demonstrate that to prevent the cloud from tracking users via uploaded embeddings, the local distributions of different users should either be derived from a linearly dependent space to avoid identifiability or be close to each other to prevent accurate attribution. Evaluation on both public and industrial datasets using different language models reveals a remarkable improvement in accuracy from incorporating anonymous user embeddings, while preserving real-time inference requirement.
Yangwenjian Tan, Chaoyue Niu, Fandong Meng, Jie Zhou 0016, Fan Wu 0006, Guihai Chen
KDD (1)5
2025 Efficient Speech Language Modeling via Energy Distance in Continuous Latent Space
abstract
We introduce \emph{SLED}, an alternative approach to speech language modeling by encoding speech waveforms into sequences of continuous latent representations and modeling them autoregressively using an energy distance objective. The energy distance offers an analytical measure of the distributional gap by contrasting simulated and target samples, enabling efficient training to capture the underlying continuous autoregressive distribution. By bypassing reliance on residual vector quantization, SLED avoids discretization errors and eliminates the need for the complicated hierarchical architectures common in existing speech language models. It simplifies the overall modeling pipeline while preserving the richness of speech information and maintaining inference efficiency. Empirical results demonstrate that SLED achieves strong performance in both zero-shot and streaming speech synthesis, showing its potential for broader applications in general-purpose speech language models. Demos and code are available at \url{https://github.com/ictnlp/SLED-TTS}.
Zhengrui Ma, Yang Feng 0004, Chenze Shao, Fandong Meng, Jie Zhou 0016, Min Zhang 0005
NeurIPS4
2025 Bridging the reality gap: A benchmark for physical reasoning in general world models with various physical phenomena beyond mechanics
Jin An Xu, Huiqi Hu, Xiuwen Xu, Zijian Jin, Fandong Meng, Jie Zhou 0016, Wenjuan Han
Expert Syst. Appl.8
2024 Generative Multi-Modal Knowledge Retrieval with Large Language Models
abstract
Knowledge retrieval with multi-modal queries plays a crucial role in supporting knowledge-intensive multi-modal applications. However, existing methods face challenges in terms of their effectiveness and training efficiency, especially when it comes to training and integrating multiple retrievers to handle multi-modal queries. In this paper, we propose an innovative end-to-end generative framework for multi-modal knowledge retrieval. Our framework takes advantage of the fact that large language models (LLMs) can effectively serve as virtual knowledge bases, even when trained with limited data. We retrieve knowledge via a two-step process: 1) generating knowledge clues related to the queries, and 2) obtaining the relevant document by searching databases using the knowledge clue. In particular, we first introduce an object-aware prefix-tuning technique to guide multi-grained visual learning. Then, we align multi-grained visual features into the textual feature space of the LLM, employing the LLM to capture cross-modal interactions. Subsequently, we construct instruction data with a unified format for model training. Finally, we propose the knowledge-guided generation strategy to impose prior constraints in the decoding steps, thereby promoting the generation of distinctive knowledge clues. Through experiments conducted on three benchmarks, we demonstrate significant improvements ranging from 3.0% to 14.6% across all evaluation metrics when compared to strong baselines.
Xinwei Long, Jiali Zeng, Fandong Meng, Zhiyuan Ma 0005, Bowen Zhou 0002, Jie Zhou 0016
AAAI3
2024 Teaching Large Language Models to Translate with Comparison
abstract
Open-sourced large language models (LLMs) have demonstrated remarkable efficacy in various tasks with instruction tuning. However, these models can sometimes struggle with tasks that require more specialized knowledge such as translation. One possible reason for such deficiency is that instruction tuning aims to generate fluent and coherent text that continues from a given instruction without being constrained by any task-specific requirements. Moreover, it can be more challenging to tune smaller LLMs with lower-quality training data. To address this issue, we propose a novel framework using examples in comparison to teach LLMs to learn translation. Our approach involves output comparison and preference comparison, presenting the model with carefully designed examples of correct and incorrect translations and an additional preference loss for better regularization. Empirical evaluation on four language directions of WMT2022 and FLORES-200 benchmarks shows the superiority of our proposed method over existing methods. Our findings offer a new perspective on fine-tuning LLMs for translation tasks and provide a promising solution for generating high-quality translations. Please refer to Github for more details: https://github.com/lemon0830/TIM.
Jiali Zeng, Fandong Meng, Yongjing Yin, Jie Zhou 0016
AAAI2
2024 Tree-of-Reasoning Question Decomposition for Complex Question Answering with Large Language Models
abstract
Large language models (LLMs) have recently demonstrated remarkable performance across various Natual Language Processing tasks. In the field of multi-hop reasoning, the Chain-of-thought (CoT) prompt method has emerged as a paradigm, using curated stepwise reasoning demonstrations to enhance LLM's ability to reason and produce coherent rational pathways. To ensure the accuracy, reliability, and traceability of the generated answers, many studies have incorporated information retrieval (IR) to provide LLMs with external knowledge. However, existing CoT with IR methods decomposes questions into sub-questions based on a single compositionality type, which limits their effectiveness for questions involving multiple compositionality types. Additionally, these methods suffer from inefficient retrieval, as complex questions often contain abundant information, leading to the retrieval of irrelevant information inconsistent with the query's intent. In this work, we propose a novel question decomposition framework called TRQA for multi-hop question answering, which addresses these limitations. Our framework introduces a reasoning tree (RT) to represent the structure of complex questions. It consists of four components: the Reasoning Tree Constructor (RTC), the Question Generator (QG), the Retrieval and LLM Interaction Module (RAIL), and the Answer Aggregation Module (AAM). Specifically, the RTC predicts diverse sub-question structures to construct the reasoning tree, allowing a more comprehensive representation of complex questions. The QG generates sub-questions for leaf-node in the reasoning tree, and we explore two methods for QG: prompt-based and T5-based approaches. The IR module retrieves documents aligned with sub-questions, while the LLM formulates answers based on the retrieved information. Finally, the AAM aggregates answers along the reason tree, producing a definitive response from bottom to top.
Kun Zhang 0041, Jiali Zeng, Fandong Meng, Yuanzhuo Wang, Shiqi Sun 0003, Long Bai 0002, Huawei Shen, Jie Zhou 0016
AAAI3
2024 CSCD-NS: a Chinese Spelling Check Dataset for Native Speakers
abstract
In this paper, we present CSCD-NS, the first Chinese spelling check (CSC) dataset designed for native speakers, containing 40,000 samples from a Chinese social platform.Compared with existing CSC datasets aimed at Chinese learners, CSCD-NS is ten times larger in scale and exhibits a distinct error distribution, with a significantly higher proportion of word-level errors.To further enhance the data resource, we propose a novel method that simulates the input process through an input method, generating large-scale and high-quality pseudo data that closely resembles the actual error distribution and outperforms existing methods.Moreover, we investigate the performance of various models in this scenario, including large language models (LLMs), such as ChatGPT.The result indicates that generative models underperform BERT-like classification models due to strict length and pronunciation constraints.The high prevalence of word-level errors also makes CSC for native speakers challenging enough, leaving substantial room for improvement.
Fandong Meng, Jie Zhou 0016
ACL (1)2
2024 Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine Translation
abstract
Incrementally expanding the capability of an existing translation model to solve new domain tasks over time is a fundamental and practical problem, which usually suffers from catastrophic forgetting.Generally, multi-domain learning can be seen as a good solution.However, there are two drawbacks: 1) it requires having the training data for all domains available at the same time, which may be unrealistic due to storage or privacy concerns; 2) it requires re-training the model on the data of all domains from scratch when adding a new domain and this is time-consuming and computationally expensive.To address these issues, we present a semi-supervised contrastive distillation framework for incremental neural machine translation.Specifically, to avoid catastrophic forgetting, we propose to exploit unlabeled data from the same distributions of the older domains through knowledge distillation.Further, to ensure the distinct domain characteristics in the model as the number of domains increases, we devise a cross-domain contrastive objective to enhance the distilled knowledge.Extensive experiments on domain translation benchmarks show that our approach, without accessing any previous training data or re-training on all domains from scratch, can significantly prevent the model from forgetting previously learned knowledge while obtaining good performance on the incrementally added domains.
Yunlong Liang, Fandong Meng, Jiaan Wang, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)2
2024 Understanding and Addressing the Under-Translation Problem from the Perspective of Decoding Objective
abstract
Neural Machine Translation (NMT) has made remarkable progress over the past years.However, under-translation and over-translation remain two challenging problems in state-of-theart NMT systems.In this work, we conduct an in-depth analysis on the underlying cause of under-translation in NMT, providing an explanation from the perspective of decoding objective.To optimize the beam search objective, the model tends to overlook words it is less confident about, leading to the under-translation phenomenon.Correspondingly, the model's confidence in predicting the End Of Sentence (EOS) diminishes when under-translation occurs, serving as a mild penalty for under-translated candidates.Building upon this analysis, we propose employing the confidence of predicting EOS as a detector for under-translation, and strengthening the confidence-based penalty to penalize candidates with a high risk of under-translation.Experiments on both synthetic and real-world data show that our method can accurately detect and rectify under-translated outputs, with minor impact on other correct translations.
Chenze Shao, Fandong Meng, Jiali Zeng, Jie Zhou 0016
ACL (1)2
2024 Cross-Lingual Knowledge Editing in Large Language Models
abstract
Knowledge editing aims to change language models' performance on several special cases (i.e., editing scope) by infusing the corresponding expected knowledge into them.With the recent advancements in large language models (LLMs), knowledge editing has been shown as a promising technique to adapt LLMs to new knowledge without retraining from scratch.However, most of the previous studies neglect the multi-lingual nature of some main-stream LLMs (e.g., LLaMA, ChatGPT and GPT-4), and typically focus on monolingual scenarios, where LLMs are edited and evaluated in the same language.As a result, it is still unknown the effect of source language editing on a different target language.In this paper, we aim to figure out this cross-lingual effect in knowledge editing.Specifically, we first collect a largescale cross-lingual synthetic dataset by translating ZsRE from English to Chinese.Then, we conduct English editing on various knowledge editing methods covering different paradigms, and evaluate their performance in Chinese, and vice versa.To give deeper analyses of the crosslingual effect, the evaluation includes four aspects, i.e., reliability, generality, locality and portability.Furthermore, we analyze the inconsistent behaviors of the edited models and discuss their specific challenges. 1
Jiaan Wang, Yunlong Liang, Zengkui Sun, Jiarong Xu, Fandong Meng
ACL (1)6
2024 TasTe: Teaching Large Language Models to Translate through Self-Reflection
abstract
Large language models (LLMs) have exhibited remarkable performance in various natural language processing tasks.Techniques like instruction tuning have effectively enhanced the proficiency of LLMs in the downstream task of machine translation.However, the existing approaches fail to yield satisfactory translation outputs that match the quality of supervised neural machine translation (NMT) systems.One plausible explanation for this discrepancy is that the straightforward prompts employed in these methodologies are unable to fully exploit the acquired instruction-following capabilities.To this end, we propose the TASTE framework, which stands for translating through selfreflection.The self-reflection process includes two stages of inference.In the first stage, LLMs are instructed to generate preliminary translations and conduct self-assessments on these translations simultaneously.In the second stage, LLMs are tasked to refine these preliminary translations according to the evaluation results.The evaluation results in four language directions on the WMT22 benchmark reveal the effectiveness of our approach compared to existing methods.Our work presents a promising approach to unleash the potential of LLMs and enhance their capabilities in MT.The codes and datasets are open-sourced at https://github. com/YutongWang1216/ReflectionLLMMT.
Jiali Zeng, Xuebo Liu 0002, Fandong Meng, Jie Zhou 0016, Min Zhang 0005
ACL (1)4
2024 Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation
abstract
Retrieval-augmented generation (RAG) enhances large language models (LLMs) by incorporating additional information from retrieval.However, studies have shown that LLMs still face challenges in effectively using the retrieved information, even ignoring it or being misled by it.The key reason is that the training of LLMs does not clearly make LLMs learn how to utilize input retrieved texts with varied quality.In this paper, we propose a novel perspective that considers the role of LLMs in RAG as "Information Refiner", which means that regardless of correctness, completeness, or usefulness of retrieved texts, LLMs can consistently integrate knowledge within the retrieved texts and model parameters to generate the texts that are more concise, accurate, and complete than the retrieved texts.To this end, we propose an information refinement training method named INFO-RAG that optimizes LLMs for RAG in an unsupervised manner.INFO-RAG is low-cost and general across various tasks.Extensive experiments on zero-shot prediction of 11 datasets in diverse tasks including Question Answering, Slot-Filling, Language Modeling, Dialogue, and Code Generation show that INFO-RAG improves the performance of LLaMA2 by an average of 9.39% relative points.INFO-RAG also shows advantages in in-context learning and robustness of RAG.LLMs
Liang Pang 0001, Mo Yu, Fandong Meng, Huawei Shen, Xueqi Cheng 0001, Jie Zhou 0016
ACL (1)4
2024 UMTIT: Unifying Recognition, Translation, and Generation for Multimodal Text Image Translation
abstract
Prior research in Image Machine Translation (IMT) has focused on either translating the source image solely into the target language text or exclusively into the target image. As a result, the former approach lacked the capacity to generate target images, while the latter was insufficient in producing target text. In this paper, we present a Unified Multimodal Text Image Translation (UMTIT) model that not only translates text images into the target language but also generates consistent target images. The UMTIT model consists of two image-text modality conversion steps: the first step converts images to text to recognize the source text and generate translations, while the second step transforms text to images to create target images based on the translations. Due to the limited availability of public datasets, we have constructed two multimodal image translation datasets. Experimental results show that our UMTIT model is versatile enough to handle tasks across multiple modalities and outperforms previous methods. Notably, UMTIT surpasses the state-of-the-art TrOCR in text recognition tasks, achieving a lower Character Error Rate (CER); it also outperforms cascading methods in text translation tasks, obtaining a higher BLEU score; and, most importantly, UMTIT can generate high-quality target text images.
Liqiang Niu, Fandong Meng, Jie Zhou 0016
LREC/COLING2
2024 DC-MBR: Distributional Cooling for Minimum Bayesian Risk Decoding
abstract
Minimum Bayesian Risk Decoding (MBR) emerges as a promising decoding algorithm in Neural Machine Translation. However, MBR performs poorly with label smoothing, which is surprising as label smoothing provides decent improvement with beam search and improves generality in various tasks. In this work, we show that the issue arises from the inconsistency of label smoothing on the token-level and sequence-level distributions. We demonstrate that even though label smoothing only causes a slight change in the token level, the sequence-level distribution is highly skewed. We coin the issue autoregressive over-smoothness. To address this issue, we propose a simple and effective method, Distributional Cooling MBR (DC-MBR), which manipulates the entropy of output distributions by tuning down the Softmax temperature. We theoretically prove the equivalence between the pre-tuning label smoothing factor and distributional cooling. Extensive experiments on NMT benchmarks validate that distributional cooling improves MBR in various settings.
Jianhao Yan, Fandong Meng, Jie Zhou 0016, Yue Zhang 0004
LREC/COLING3
2024 C-LLM: Learn to Check Chinese Spelling Errors Character by Character
abstract
Chinese Spell Checking (CSC) aims to detect and correct spelling errors in sentences.Despite Large Language Models (LLMs) exhibit robust capabilities and are widely applied in various tasks, their performance on CSC is often unsatisfactory.We find that LLMs fail to meet the Chinese character-level constraints of the CSC task, namely equal length and phonetic similarity, leading to a performance bottleneck.Further analysis reveals that this issue stems from the granularity of tokenization, as current mixed character-word tokenization struggles to satisfy these characterlevel constraints.To address this issue, we propose C-LLM, a Large Language Modelbased Chinese Spell Checking method that learns to check errors Character by Character.Character-level tokenization enables the model to learn character-level alignment, effectively mitigating issues related to character-level constraints.Furthermore, CSC is simplified to replication-dominated and substitutionsupplemented tasks.Experiments on two CSC benchmarks demonstrate that C-LLM achieves an average improvement of 10% over existing methods.Specifically, it shows a 2.1% improvement in general scenarios and a significant 12% improvement in vertical domain scenarios, establishing state-of-the-art performance.The source code can be accessed at https://github.com/ktlKTL/C-LLM.
Kunting Li, Liang He 0003, Fandong Meng, Jie Zhou 0016
EMNLP4
2024 Multi-Level Cross-Modal Alignment for Speech Relation Extraction
abstract
Liang Zhang, Zhen Yang, Biao Fu, Ziyao Lu, Liangying Shao, Shiyu Liu, Fandong Meng, Jie Zhou, Xiaoli Wang, Jinsong Su. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Biao Fu, Ziyao Lu, Liangying Shao, Fandong Meng, Jie Zhou 0016, Xiaoli Wang 0002, Jinsong Su
EMNLP7
2024 Towards Codable Watermarking for Injecting Multi-Bits Information to LLMs
abstract
As large language models (LLMs) generate texts with increasing fluency and realism, there is a growing need to identify the source of texts to prevent the abuse of LLMs. Text watermarking techniques have proven reliable in distinguishing whether a text is generated by LLMs by injecting hidden patterns. However, we argue that existing LLM watermarking methods are encoding-inefficient and cannot flexibly meet the diverse information encoding needs (such as encoding model version, generation time, user id, etc.). In this work, we conduct the first systematic study on the topic of **Codable Text Watermarking for LLMs** (CTWL) that allows text watermarks to carry multi-bit customizable information. First of all, we study the taxonomy of LLM watermarking technologies and give a mathematical formulation for CTWL. Additionally, we provide a comprehensive evaluation system for CTWL: (1) watermarking success rate, (2) robustness against various corruptions, (3) coding rate of payload information, (4) encoding and decoding efficiency, (5) impacts on the quality of the generated text. To meet the requirements of these non-Pareto-improving metrics, we follow the most prominent vocabulary partition-based watermarking direction, and devise an advanced CTWL method named **Balance-Marking**. The core idea of our method is to use a proxy language model to split the vocabulary into probability-balanced parts, thereby effectively maintaining the quality of the watermarked text. Our code is available at https://github.com/lancopku/codable-watermarking-for-llm.
Lean Wang, Wenkai Yang, Deli Chen, Hao Zhou 0012, Yankai Lin 0001, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
ICLR6
2024 Large Language Models Are Not Robust Multiple Choice Selectors
abstract
Multiple choice questions (MCQs) serve as a common yet important task format in the evaluation of large language models (LLMs). This work shows that modern LLMs are vulnerable to option position changes in MCQs due to their inherent “selection bias”, namely, they prefer to select specific option IDs as answers (like “Option A”). Through extensive empirical analyses with 20 LLMs on three benchmarks, we pinpoint that this behavioral bias primarily stems from LLMs’ token bias, where the model a priori assigns more probabilistic mass to specific option ID tokens (e.g., A/B/C/D) when predicting answers from the option IDs. To mitigate selection bias, we propose a label-free, inference-time debiasing method, called PriDe, which separates the model’s prior bias for option IDs from the overall prediction distribution. PriDe first estimates the prior by permutating option contents on a small number of test samples, and then applies the estimated prior to debias the remaining samples. We demonstrate that it achieves interpretable and transferable debiasing with high computational efficiency. We hope this work can draw broader research attention to the bias and robustness of modern LLMs.
Chujie Zheng, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Minlie Huang
ICLR3
2024 Language Generation with Strictly Proper Scoring Rules
abstract
Language generation based on maximum likelihood estimation (MLE) has become the fundamental approach for text generation. Maximum likelihood estimation is typically performed by minimizing the log-likelihood loss, also known as the logarithmic score in statistical decision theory. The logarithmic score is strictly proper in the sense that it encourages honest forecasts, where the expected score is maximized only when the model reports true probabilities. Although many strictly proper scoring rules exist, the logarithmic score is the only local scoring rule among them that depends exclusively on the probability of the observed sample, making it capable of handling the exponentially large sample space of natural text. In this work, we propose a straightforward strategy for adapting scoring rules to language generation, allowing for language modeling with any non-local scoring rules. Leveraging this strategy, we train language generation models using two classic strictly proper scoring rules, the Brier score and the Spherical score, as alternatives to the logarithmic score. Experimental results indicate that simply substituting the loss function, without adjusting other hyperparameters, can yield substantial improvements in model’s generation capabilities. Moreover, these improvements can scale up to large language models (LLMs) such as LLaMA-7B and LLaMA-13B. Source code: https://github.com/shaochenze/ScoringRulesLM.
Chenze Shao, Fandong Meng, Yijin Liu, Jie Zhou 0016
ICML2
2024 On Prompt-Driven Safeguarding for Large Language Models
abstract
Prepending model inputs with safety prompts is a common practice for safeguarding large language models (LLMs) against queries with harmful intents. However, the underlying working mechanisms of safety prompts have not been unraveled yet, restricting the possibility of automatically optimizing them to improve LLM safety. In this work, we investigate how LLMs’ behavior (i.e., complying with or refusing user queries) is affected by safety prompts from the perspective of model representation. We find that in the representation space, the input queries are typically moved by safety prompts in a "higher-refusal" direction, in which models become more prone to refusing to provide assistance, even when the queries are harmless. On the other hand, LLMs are naturally capable of distinguishing harmful and harmless queries without safety prompts. Inspired by these findings, we propose a method for safety prompt optimization, namely DRO (Directed Representation Optimization). Treating a safety prompt as continuous, trainable embeddings, DRO learns to move the queries’ representations along or opposite the refusal direction, depending on their harmfulness. Experiments with eight LLMs on out-of-domain and jailbreak benchmarks demonstrate that DRO remarkably improves the safeguarding performance of human-crafted safety prompts, without compromising the models’ general performance.
Chujie Zheng, Fan Yin, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Kai-Wei Chang 0001, Minlie Huang, Nanyun Peng 0001
ICML4
2024 How Effectively Do Code Language Models Understand Poor-Readability Code?
abstract
Code language models such as CodeT5 and CodeLlama have demonstrated substantial achievement in code comprehension. While the majority of research efforts have focused on improving model architectures and training processes, we find that the current benchmarks used for evaluating code comprehension models are confined to high-readability code, regardless of the popularity of low-readability code in reality. As such, they are inadequate to demonstrate the full spectrum of the model's ability, particularly the robustness to varying readability degrees. In this paper, we analyze the robustness of code summarization models to code with varying readability, including seven obfuscated datasets derived from existing benchmarks. Our findings indicate that current code summarization models are vulnerable to code with poor readability. In particular, their performance predominantly depends on semantic cues within the code, often neglecting the syntactic aspects. Existing benchmarks are biased toward evaluating semantic features, thereby overlooking the models' ability to understand nonsensitive syntactic features. Based on the findings, we present Poor-CodeSumEval, a new evaluation benchmark on code summarization tasks. PoorCodeSumEval innovatively introduces readability into the testing process, considering semantic, syntactic, and their cross-obfuscation, thereby providing a more comprehensive and rigorous evaluation of code summarization models. Our studies also provide more insightful suggestions for future research, such as constructing multi-readability benchmarks to evaluate the robustness of models on poor-readability code, proposing readability-awareness metrics, and automatic methods for code data cleaning and normalization.
Yitian Chai, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xiaodong Gu 0002
ASE4
2024 On Large Language Models' Hallucination with Regard to Known Facts
abstract
Che Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Yang Cheng, Fandong Meng, Mo Yu, Bowen Zhou, Jie Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Che Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Fandong Meng, Mo Yu, Bowen Zhou 0002, Jie Zhou 0016
NAACL-HLT6
2024 XAL: EXplainable Active Learning Makes Classifiers Better Low-resource Learners
abstract
Yun Luo, Zhen Yang, Fandong Meng, Yingjie Li, Fang Guo, Qinglin Qi, Jie Zhou, Yue Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Fandong Meng, Yingjie Li 0008, Qinglin Qi, Jie Zhou 0016, Yue Zhang 0004
NAACL-HLT3
2024 Complex Question Enhanced Transfer Learning for Zero-Shot Joint Information Extraction
abstract
Zero-shot information extraction (IE) tasks have attracted great attention recently. However, how to jointly model multiple IE tasks in the zero-shot scenario is still an open question. In this article, we focus on zero-shot joint IE tasks and highlight how to transfer the knowledge of cross-task relations from the source domain to the target domain. To solve this problem, we first unify all IE tasks with a machine reading comprehension (MRC) framework, which can make the most of training data and enhance its ability on span extraction. Then, we generatecomplex questionsto explicitly model cross-task relations with natural language descriptions, thereby providing prior knowledge for pre-defined types and building more general linkages among different entities and triggers as well. Specifically, we define three operations for generating templates for complex questions, i.e.,intersecting,connecting, andcomposing. Besides, we design an efficient training strategy to exploit the synthetic data with complex questions. We evaluate our approach on four datasets from different domains for various IE tasks. Experimental results show the effectiveness of our approach in improving the performance of zero-shot joint IE tasks in multiple domains.
Ying Zhang 0084, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Summary-Oriented Vision Modeling for Multimodal Abstractive Summarization
abstract
Multimodal abstractive summarization (MAS) aims to produce a concise summary given the multimodal data (text and vision).Existing studies mainly focus on how to effectively use the visual features from the perspective of an article, having achieved impressive success on the high-resource English dataset.However, less attention has been paid to the visual features from the perspective of the summary, which may limit the model performance, especially in the low-and zero-resource scenarios.In this paper, we propose to improve the summary quality through summary-oriented visual features.To this end, we devise two auxiliary tasks including vision to summary task and masked image modeling task.Together with the main summarization task, we optimize the MAS model via the training objectives of all these tasks.By these means, the MAS model can be enhanced by capturing the summaryoriented visual features, thereby yielding more accurate summaries.Experiments on 44 languages, covering mid-high-, low-, and zeroresource scenarios, verify the effectiveness and superiority of the proposed approach, which achieves state-of-the-art performance under all scenarios.Additionally, we will contribute a large-scale multilingual multimodal abstractive summarization (MM-Sum) dataset. 1
Yunlong Liang, Fandong Meng, Jin An Xu, Jiaan Wang, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)2
2023 Towards Unifying Multi-Lingual and Cross-Lingual Summarization
abstract
Jiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang, Zhixu Li, Jianfeng Qu, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang, Zhixu Li, Jianfeng Qu, Jie Zhou 0016
ACL (1)2
2023 Consistency Regularization Training for Compositional Generalization
abstract
Existing neural models have difficulty generalizing to unseen combinations of seen components.To achieve compositional generalization, models are required to consistently interpret (sub)expressions across contexts.Without modifying model architectures, we improve the capability of Transformer on compositional generalization through consistency regularization training, which promotes representation consistency across samples and prediction consistency for a single sample.Experimental results on semantic parsing and machine translation benchmarks empirically demonstrate the effectiveness and generality of our method.In addition, we find that the prediction consistency scores on in-distribution validation sets can be an alternative for evaluating models during training, when commonly-used metrics are not informative.
Yongjing Yin, Jiali Zeng, Yafu Li, Fandong Meng, Jie Zhou 0016, Yue Zhang 0004
ACL (1)4
2023 Personality Understanding of Fictional Characters during Book Reading
abstract
Mo Yu, Jiangnan Li, Shunyu Yao, Wenjie Pang, Xiaochen Zhou, Zhou Xiao, Fandong Meng, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Mo Yu, Wenjie Pang, Xiaochen Zhou, Xiao Zhou 0004, Fandong Meng, Jie Zhou 0016
ACL (1)7
2023 Soft Language Clustering for Multilingual Model Pre-training
abstract
Jiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou 0016
ACL (1)5
2023 Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning
abstract
In-context learning (ICL) emerges as a promising capability of large language models (LLMs) by providing them with demonstration examples to perform diverse tasks.However, the underlying mechanism of how LLMs learn from the provided context remains under-explored.In this paper, we investigate the working mechanism of ICL through an information flow lens.Our findings reveal that label words in the demonstration examples function as anchors:(1) semantic information aggregates into label word representations during the shallow computation layers' processing; (2) the consolidated information in label words serves as a reference for LLMs' final predictions.Based on these insights, we introduce an anchor re-weighting method to improve ICL performance, a demonstration compression technique to expedite inference, and an analysis framework for diagnosing ICL errors in GPT2-XL.The promising applications of our findings again validate the uncovered ICL working mechanism and pave the way for future studies. 1
Lean Wang, Lei Li 0039, Damai Dai, Deli Chen, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
EMNLP6
2023 HyperNetwork-based Decoupling to Improve Model Generalization for Few-Shot Relation Extraction
abstract
Few-shot relation extraction (FSRE) aims to train a model that can deal with new relations using only a few labeled examples.Most existing studies employ Prototypical Networks for FSRE, which usually overfits the relation classes in the training set and cannot generalize well to unseen relations.By investigating the class separation of an FSRE model, we find that model upper layers are prone to learn relation-specific knowledge.Therefore, in this paper, we propose a HyperNetworkbased Decoupling approach to improve the generalization of FSRE models.Specifically, our model consists of an encoder, a network generator (for producing relation classifiers) and the generated-then-finetuned classifiers for every N -way-K-shot episode.Meanwhile, we design a two-step training strategy along with a class-agnostic aligner, by which the generated classifiers focus on acquiring relation-specific knowledge and the encoder is encouraged to learn more general relation knowledge.In this way, the roles of upper and lower layers in our FSRE model are explicitly decoupled, thus enhancing its generalizing capability during testing.Experiments on two public datasets demonstrate the effectiveness of our method.Our source code is available at https: //github.com/DeepLearnXMU/FSRE-HDN.
Chulun Zhou, Fandong Meng, Jinsong Su, Yidong Chen 0001, Jie Zhou 0016
EMNLP3
2023 Towards Better Multi-modal Keyphrase Generation via Visual Entity Enhancement and Multi-granularity Image Noise Filtering
abstract
Multi-modal keyphrase generation aims to produce a set of keyphrases that represent the core points of the input text-image pair. In this regard, dominant methods mainly focus on multi-modal fusion for keyphrase generation. Nevertheless, there are still two main drawbacks: 1) only a limited number of sources, such as image captions, can be utilized to provide auxiliary information. However, they may not be sufficient for the subsequent keyphrase generation. 2) the input text and image are often not perfectly matched, and thus the image may introduce noise into the model. To address these limitations, in this paper, we propose a novel multi-modal keyphrase generation model, which not only enriches the model input with external knowledge, but also effectively filters image noise. First, we introduce external visual entities of the image as the supplementary input to the model, which benefits the cross-modal semantic alignment for keyphrase generation. Second, we simultaneously calculate an image-text matching score and image region-text correlation scores to perform multi-granularity image noise filtering. Particularly, we introduce the correlation scores between image regions and ground-truth keyphrases to refine the calculation of the previously-mentioned correlation scores. To demonstrate the effectiveness of our model, we conduct several groups of experiments on the benchmark dataset. Experimental results and in-depth analyses show that our model achieves the state-of-the-art performance. Our code is available on https://github.com/DeepLearnXMU/MM-MKP.
Suhang Wu, Fandong Meng, Jie Zhou 0016, Xiaoli Wang 0002, Jinsong Su
ACM Multimedia3
2023 Fed-FA: Theoretically Modeling Client Data Divergence for Federated Language Backdoor Defense
abstract
Federated learning algorithms enable neural network models to be trained across multiple decentralized edge devices without sharing private data. However, they are susceptible to backdoor attacks launched by malicious clients. Existing robust federated aggregation algorithms heuristically detect and exclude suspicious clients based on their parameter distances, but they are ineffective on Natural Language Processing (NLP) tasks. The main reason is that, although text backdoor patterns are obvious at the underlying dataset level, they are usually hidden at the parameter level, since injecting backdoors into texts with discrete feature space has less impact on the statistics of the model parameters. To settle this issue, we propose to identify backdoor clients by explicitly modeling the data divergence among clients in federated NLP systems. Through theoretical analysis, we derive the f-divergence indicator to estimate the client data divergence with aggregation updates and Hessians. Furthermore, we devise a dataset synthesization method with a Hessian reassignment mechanism guided by the diffusion theory to address the key challenge of inaccessible datasets in calculating clients' data Hessians. We then present the novel Federated F-Divergence-Based Aggregation~(\textbf{Fed-FA}) algorithm, which leverages the f-divergence indicator to detect and discard suspicious clients. Extensive empirical results show that Fed-FA outperforms all the parameter distance-based methods in defending against backdoor attacks among various natural language backdoor attack scenarios.
Zhiyuan Zhang 0001, Deli Chen, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
NeurIPS4
2023 Multi-modal graph contrastive encoding for neural machine translation
Yongjing Yin, Jiali Zeng, Jinsong Su, Chulun Zhou, Fandong Meng, Jie Zhou 0016, Degen Huang, Jiebo Luo 0001
Artif. Intell.5
2023 A Multi-Task Multi-Stage Transitional Training Framework for Neural Chat Translation
abstract
Neural chat translation (NCT) aims to translate a cross-lingual chat between speakers of different languages. Existing context-aware NMT models cannot achieve satisfactory performances due to the following inherent problems: 1) limited resources of annotated bilingual dialogues; 2) the neglect of modelling conversational properties; 3) training discrepancy between different stages. To address these issues, in this paper, we propose a multi-task multi-stage transitional (MMT) training framework, where an NCT model is trained using the bilingual chat translation dataset and additional monolingual dialogues. We elaborately design two auxiliary tasks, namely utterance discrimination and speaker discrimination, to introduce the modelling of dialogue coherence and speaker characteristic into the NCT model. The training process consists of three stages: 1) sentence-level pre-training on large-scale parallel corpus; 2) intermediate training with auxiliary tasks using additional monolingual dialogues; 3) context-aware fine-tuning with gradual transition. Particularly, the second stage serves as an intermediate phase that alleviates the training discrepancy between the pre-training and fine-tuning stages. Moreover, to make the stage transition smoother, we train the NCT model using a gradual transition strategy, i.e., gradually transiting from using monolingual to bilingual dialogues. Extensive experiments on two language pairs demonstrate the effectiveness and superiority of our proposed training framework.
Chulun Zhou, Yunlong Liang, Fandong Meng, Jie Zhou 0016, Jin An Xu, Min Zhang 0005, Jinsong Su
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Scheduled Multi-task Learning for Neural Chat Translation
abstract
Neural Chat Translation (NCT) aims to translate conversational text into different languages.Existing methods mainly focus on modeling the bilingual dialogue characteristics (e.g., coherence) to improve chat translation via multi-task learning on small-scale chat translation data.Although the NCT models have achieved impressive success, it is still far from satisfactory due to insufficient chat translation data and simple joint training manners.To address the above issues, we propose a scheduled multi-task learning framework for NCT.Specifically, we devise a three-stage training framework to incorporate the large-scale in-domain chat translation data into training by adding a second pre-training stage between the original pre-training and fine-tuning stages.Further, we investigate where and how to schedule the dialogue-related auxiliary tasks in multiple training stages to effectively enhance the main chat translation task.Extensive experiments on four language directions (English↔Chinese and English↔German) verify the effectiveness and superiority of the proposed approach.Additionally, we will make the large-scale indomain paired bilingual dialogue dataset publicly available for the research community.1
Yunlong Liang, Fandong Meng, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)2
2022 MSCTD: A Multimodal Sentiment Chat Translation Dataset
abstract
Multimodal machine translation and textual chat translation have received considerable attention in recent years.Although the conversation in its natural form is usually multimodal, there still lacks work on multimodal machine translation in conversations.In this work, we introduce a new task named Multimodal Chat Translation (MCT), aiming to generate more accurate translations with the help of the associated dialogue history and visual context.To this end, we firstly construct a Multimodal Sentiment Chat Translation Dataset (MSCTD) containing 142,871 English-Chinese utterance pairs in 14,762 bilingual dialogues and 30,370 English-German utterance pairs in 3,079 bilingual dialogues.Each utterance pair, corresponding to the visual context that reflects the current conversational scene, is annotated with a sentiment label.Then, we benchmark the task by establishing multiple baseline systems that incorporate multimodal and sentiment features for MCT.Preliminary experiments on four language directions (English↔Chinese and English↔German) verify the potential of contextual and multimodal information fusion and the positive impact of sentiment on the MCT task.Additionally, as a by-product of the MSCTD, it also provides two new benchmarks on multimodal dialogue sentiment analysis.Our work can facilitate research on both multimodal chat translation and multimodal dialogue sentiment analysis.1
Yunlong Liang, Fandong Meng, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)2
2022 A Variational Hierarchical Model for Neural Cross-Lingual Summarization
abstract
Yunlong Liang, Fandong Meng, Chulun Zhou, Jinan Xu, Yufeng Chen, Jinsong Su, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yunlong Liang, Fandong Meng, Chulun Zhou, Jin An Xu, Yufeng Chen 0005, Jinsong Su, Jie Zhou 0016
ACL (1)2
2022 EAG: Extract and Generate Multi-way Aligned Corpus for Complete Multi-lingual Neural Machine Translation
abstract
Complete Multi-lingual Neural Machine Translation (C-MNMT) achieves superior performance against the conventional MNMT by constructing multi-way aligned corpus, i.e., aligning bilingual training examples from different language pairs when either their source or target sides are identical.However, since exactly identical sentences from different language pairs are scarce, the power of the multi-way aligned corpus is limited by its scale.To handle this problem, this paper proposes "Extract and Generate" (EAG), a two-step approach to construct large-scale and high-quality multi-way aligned corpus from bilingual data.Specifically, we first extract candidate aligned examples by pairing the bilingual examples from different language pairs with highly similar source or target sentences; and then generate the final aligned examples from the candidates with a welltrained generation model.With this two-step pipeline, EAG can construct a large-scale and multi-way aligned corpus whose diversity is almost identical to the original bilingual corpus.Experiments on two publicly available datasets i.e., WMT-5 and OPUS-100, show that the proposed method achieves significant improvements over strong baselines, with +1.1 and +1.4 BLEU points improvements on the two datasets respectively.
Fandong Meng, Jie Zhou 0016
ACL (1)3
2022 Conditional Bilingual Mutual Information Based Adaptive Training for Neural Machine Translation
abstract
Songming Zhang, Yijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu, Jian Liu, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Songming Zhang 0001, Yijin Liu, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jian Liu 0032, Jie Zhou 0016
ACL (1)3
2022 Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine Translation
abstract
Most dominant neural machine translation (NMT) models are restricted to make predictions only according to the local context of preceding words in a left-to-right manner.Although many previous studies try to incorporate global information into NMT models, there still exist limitations on how to effectively exploit bidirectional global context.In this paper, we propose a Confidence Based Bidirectional Global Context Aware (CB-BGCA) training framework for NMT, where the NMT model is jointly trained with an auxiliary conditional masked language model (CMLM).The training consists of two stages:(1) multi-task joint training; (2) confidence based knowledge distillation.At the first stage, by sharing encoder parameters, the NMT model is additionally supervised by the signal from the CMLM decoder that contains bidirectional global contexts.Moreover, at the second stage, using the CMLM as teacher, we further pertinently incorporate bidirectional global context to the NMT model on its unconfidently-predicted target words via knowledge distillation.Experimental results show that our proposed CB-BGCA training framework significantly improves the NMT model by +1.02, +1.30 and +0.57BLEU scores on three large-scale translation datasets, namely WMT'14 Englishto-German, WMT'19 Chinese-to-English and WMT'14 English-to-French, respectively.
Chulun Zhou, Fandong Meng, Jie Zhou 0016, Min Zhang 0005, Jinsong Su
ACL (1)2
2022 TAKE: Topic-shift Aware Knowledge sElection for Dialogue Generation
abstract
Knowledge-grounded dialogue generation consists of two subtasks: knowledge selection and response generation. The knowledge selector generally constructs a query based on the dialogue context and selects the most appropriate knowledge to help response generation. Recent work finds that realizing who (the user or the agent) holds the initiative and utilizing the role-initiative information to instruct the query construction can help select knowledge. It depends on whether the knowledge connection between two adjacent rounds is smooth to assign the role. However, whereby the user takes the initiative only when there is a strong semantic transition between two rounds, probably leading to initiative misjudgment. Therefore, it is necessary to seek a more sensitive reason beyond the initiative role for knowledge selection. To address the above problem, we propose a Topic-shift Aware Knowledge sElector(TAKE). Specifically, we first annotate the topic shift and topic inheritance labels in multi-round dialogues with distant supervision. Then, we alleviate the noise problem in pseudo labels through curriculum learning and knowledge distillation. Extensive experiments on WoW show that TAKE performs better than strong baselines.
Chenxu Yang, Zheng Lin 0001, Fandong Meng, Weiping Wang 0005, Lanrui Wang, Jie Zhou 0016
COLING4
2022 Categorizing Semantic Representations for Neural Machine Translation
abstract
Modern neural machine translation (NMT) models have achieved competitive performance in standard benchmarks. However, they have recently been shown to suffer limitation in compositional generalization, failing to effectively learn the translation of atoms (e.g., words) and their semantic composition (e.g., modification) from seen compounds (e.g., phrases), and thus suffering from significantly weakened translation performance on unseen compounds during inference. We address this issue by introducing categorization to the source contextualized representations. The main idea is to enhance generalization by reducing sparsity and overfitting, which is achieved by finding prototypes of token representations over the training set and integrating their embeddings into the source encoding. Experiments on a dedicated MT dataset (i.e., CoGnition) show that our method reduces compositional generalization error rates by 24% error reduction. In addition, our conceptually simple method gives consistently better results than the Transformer baseline on a range of general MT datasets.
Yongjing Yin, Yafu Li, Fandong Meng, Jie Zhou 0016, Yue Zhang 0004
COLING3
2022 TSAM: A Two-Stream Attention Model for Causal Emotion Entailment
abstract
Causal Emotion Entailment (CEE) aims to discover the potential causes behind an emotion in a conversational utterance. Previous works formalize CEE as independent utterance pair classification problems, with emotion and speaker information neglected. From a new perspective, this paper considers CEE in a joint framework. We classify multiple utterances synchronously to capture the correlations between utterances in a global view and propose a Two-Stream Attention Model (TSAM) to effectively model the speaker’s emotional influences in the conversational history. Specifically, the TSAM comprises three modules: Emotion Attention Network (EAN), Speaker Attention Network (SAN), and interaction module. The EAN and SAN incorporate emotion and speaker information in parallel, and the subsequent interaction module effectively interchanges relevant information between the EAN and SAN via a mutual BiAffine transformation. Extensive experimental results demonstrate that our model achieves new State-Of-The-Art (SOTA) performance and outperforms baselines remarkably.
Duzhen Zhang, Fandong Meng, Xiuyi Chen, Jie Zhou 0016
COLING3
2022 Towards Robust k-Nearest-Neighbor Machine Translation
abstract
k-Nearest-Neighbor Machine Translation (kNN-MT) becomes an important research direction of NMT in recent years.Its main idea is to retrieve useful key-value pairs from an additional datastore to modify translations without updating the NMT model.However, the underlying retrieved noisy pairs will dramatically deteriorate the model performance.In this paper, we conduct a preliminary study and find that this problem results from not fully exploiting the prediction of the NMT model.To alleviate the impact of noise, we propose a confidence-enhanced kNN-MT model with robust training.Concretely, we introduce the NMT confidence to refine the modeling of two important components of kNN-MT: kNN distribution and the interpolation weight.Meanwhile we inject two types of perturbations into the retrieved pairs for robust training.Experimental results on four benchmark datasets demonstrate that our model not only achieves significant improvements over current kNN-MT models, but also exhibits better robustness.Our code is available at https://github.com/ DeepLearnXMU/Robust-knn-mt.
Ziyao Lu, Fandong Meng, Chulun Zhou, Jie Zhou 0016, Degen Huang, Jinsong Su
EMNLP3
2022 Cross-Align: Modeling Deep Cross-lingual Interactions for Word Alignment
abstract
Word alignment which aims to extract lexicon translation equivalents between source and target sentences, serves as a fundamental tool for natural language processing.Recent studies in this area have yielded substantial improvements by generating alignments from contextualized embeddings of the pre-trained multilingual language models.However, we find that the existing approaches capture few interactions between the input sentence pairs, which degrades the word alignment quality severely, especially for the ambiguous words in the monolingual context.To remedy this problem, we propose Cross-Align to model deep interactions between the input sentence pairs, in which the source and target sentences are encoded separately with the shared self-attention modules in the shallow layers, while cross-lingual interactions are explicitly constructed by the crossattention modules in the upper layers.Besides, to train our model effectively, we propose a two-stage training framework, where the model is trained with a simple Translation Language Modeling (TLM) objective in the first stage and then finetuned with a self-supervised alignment objective in the second stage.Experiments show that the proposed Cross-Align achieves the state-of-the-art (SOTA) performance on four out of five language pairs. 1
Siyu Lai, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
EMNLP3
2022 ClidSum: A Benchmark Dataset for Cross-Lingual Dialogue Summarization
abstract
We present CLIDSUM, a benchmark dataset towards building cross-lingual summarization systems on dialogue documents.It consists of 67k+ dialogue documents and 112k+ annotated summaries in different target languages.Based on the proposed CLIDSUM, we introduce two benchmark settings for supervised and semi-supervised scenarios, respectively.We then build various baseline systems in different paradigms (pipeline and end-to-end) and conduct extensive experiments on CLIDSUM to provide deeper analyses.Furthermore, we propose mDIALBART which extends mBART via further pre-training, where the multiple objectives help the pre-trained model capture the structural characteristics as well as key content in dialogues and the transformation from source to the target language.Experimental results show the superiority of mDIALBART, as an end-to-end model, outperforms strong pipeline models on CLIDSUM.Finally, we discuss specific challenges that current approaches faced with this task and give multiple promising directions for future research.
Jiaan Wang, Fandong Meng, Ziyao Lu, Duo Zheng, Zhixu Li, Jianfeng Qu, Jie Zhou 0016
EMNLP2
2022 Digging Errors in NMT: Evaluating and Understanding Model Errors from Partial Hypothesis Space
abstract
Solid evaluation of neural machine translation (NMT) is key to its understanding and improvement.Current evaluation of an NMT system is usually built upon a heuristic decoding algorithm (e.g., beam search) and an evaluation metric assessing similarity between the translation and golden reference.However, this system-* Equal contribution.
Jianhao Yan, Chenming Wu, Fandong Meng, Jie Zhou 0016
EMNLP3
2022 WeTS: A Benchmark for Translation Suggestion
abstract
Translation suggestion (TS), which provides alternatives for specific words or phrases given the entire documents generated by machine translation (MT), has been proven to play a significant role in post-editing (PE).There are two main pitfalls for existing researches in this line.First, most conventional works only focus on the overall performance of PE but ignore the exact performance of TS, which makes the progress of PE sluggish and less explainable; Second, as no publicly available golden dataset exists to support in-depth research for TS, almost all of the previous works conduct experiments on their in-house datasets or the noisy datasets built automatically, which makes their experiments hard to be reproduced and compared.To break these limitations mentioned above and spur the research in TS, we create a benchmark dataset, called WeTS, which is a golden corpus annotated by expert translators on four translation directions.Apart from the golden corpus, we also propose several methods to generate synthetic corpora which can be used to improve the performance substantially through pre-training.As for the model, we propose the segment-aware self-attention based Transformer for TS.Experimental results show that our approach achieves the best results on all four directions, including Englishto-German, German-to-English, Chinese-to-English, and English-to-Chinese.Codes and corpus can be found at https://github.com/ ZhenYangIACAS/WeTS.git.
Fandong Meng, Yingxue Zhang 0003, Ernan Li, Jie Zhou 0016
EMNLP2
2022 Neutral Utterances are Also Causes: Enhancing Conversational Causal Emotion Entailment with Social Commonsense Knowledge
abstract
Conversational Causal Emotion Entailment aims to detect causal utterances for a non-neutral targeted utterance from a conversation. In this work, we build conversations as graphs to overcome implicit contextual modelling of the original entailment style. Following the previous work, we further introduce the emotion information into graphs. Emotion information can markedly promote the detection of causal utterances whose emotion is the same as the targeted utterance. However, it is still hard to detect causal utterances with different emotions, especially neutral ones. The reason is that models are limited in reasoning causal clues and passing them between utterances. To alleviate this problem, we introduce social commonsense knowledge (CSK) and propose a Knowledge Enhanced Conversation graph (KEC). KEC propagates the CSK between two utterances. As not all CSK is emotionally suitable for utterances, we therefore propose a sentiment-realized knowledge selecting strategy to filter CSK. To process KEC, we further construct the Knowledge Enhanced Directed Acyclic Graph networks. Experimental results show that our method outperforms baselines and infers more causes with different emotions from the targeted utterance.
Fandong Meng, Zheng Lin 0001, Rui Liu 0032, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005, Jie Zhou 0016
IJCAI2
2022 Visual Dialog for Spotting the Differences between Pairs of Similar Images
abstract
Visual dialog has witnessed great progress after introducing various vision-oriented goals into the conversation. Much of previous work focuses on tasks where only one image can be accessed by two interlocutors, such as VisDial and GuessWhat. The work on situations where two interlocutors access different images has received less attention. Those situations are common in real world and bring some different challenges compared with one-image tasks. The lack of such types of dialog tasks and corresponding large-scale datasets makes it impossible to carry out in-depth research. This paper therefore first proposes a new visual dialog task named Dial-the-Diff, where two interlocutors accessing two similar images respectively try to spot the difference between the images through conversing in natural language. The task raises new challenges to the dialog strategy and the ability of categorizing objects. We then build a large-scale multi-modal dataset for the task, named DialDiff, which contains 87k Virtual Reality images and 78k dialogs. Some details of the data are given and analyzed to highlight the challenges behind the task. Finally, we propose benchmark models for this task, and conduct extensive experiments to evaluate their performance as well as its problems remained.
Duo Zheng, Fandong Meng, Qingyi Si, Hairun Fan, Zipeng Xu, Jie Zhou 0016, Fangxiang Feng, Xiaojie Wang 0006
ACM Multimedia2
2022 Generating Authentic Adversarial Examples beyond Meaning-preserving with Doubly Round-trip Translation
abstract
Siyu Lai, Zhen Yang, Fandong Meng, Xue Zhang, Yufeng Chen, Jinan Xu, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Siyu Lai, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
NAACL-HLT3
2022 Learning to Win Lottery Tickets in BERT Transfer via Task-agnostic Mask Training
abstract
Yuanxin Liu, Fandong Meng, Zheng Lin, Peng Fu, Yanan Cao, Weiping Wang, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yuanxin Liu, Fandong Meng, Zheng Lin 0001, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005, Jie Zhou 0016
NAACL-HLT2
2022 A Win-win Deal: Towards Sparse and Robust Pre-trained Language Models
abstract
Despite the remarkable success of pre-trained language models (PLMs), they still face two challenges: First, large-scale PLMs are inefficient in terms of memory footprint and computation. Second, on the downstream tasks, PLMs tend to rely on the dataset bias and struggle to generalize to out-of-distribution (OOD) data. In response to the efficiency problem, recent studies show that dense PLMs can be replaced with sparse subnetworks without hurting the performance. Such subnetworks can be found in three scenarios: 1) the fine-tuned PLMs, 2) the raw PLMs and then fine-tuned in isolation, and even inside 3) PLMs without any parameter fine-tuning. However, these results are only obtained in the in-distribution (ID) setting. In this paper, we extend the study on PLMs subnetworks to the OOD setting, investigating whether sparsity and robustness to dataset bias can be achieved simultaneously. To this end, we conduct extensive experiments with the pre-trained BERT model on three natural language understanding (NLU) tasks. Our results demonstrate that \textbf{sparse and robust subnetworks (SRNets) can consistently be found in BERT}, across the aforementioned three scenarios, using different training and compression methods. Furthermore, we explore the upper bound of SRNets using the OOD information and show that \textbf{there exist sparse and almost unbiased BERT subnetworks}. Finally, we present 1) an analytical study that provides insights on how to promote the efficiency of SRNets searching process and 2) a solution to improve subnetworks' performance at high sparsity. The code is available at \url{https://github.com/llyx97/sparse-and-robust-PLM}.
Yuanxin Liu, Fandong Meng, Zheng Lin 0001, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005, Jie Zhou 0016
NeurIPS2
2022 Emotional conversation generation with heterogeneous graph neural network
Yunlong Liang, Fandong Meng, Ying Zhang 0084, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
Artif. Intell.2
2022 A Survey on Cross-Lingual Summarization
abstract
Abstract Cross-lingual summarization is the task of generating a summary in one language (e.g., English) for the given document(s) in a different language (e.g., Chinese). Under the globalization background, this task has attracted increasing attention of the computational linguistics community. Nevertheless, there still remains a lack of comprehensive review for this task. Therefore, we present the first systematic critical review on the datasets, approaches, and challenges in this field. Specifically, we carefully organize existing datasets and approaches according to different construction methods and solution paradigms, respectively. For each type of dataset or approach, we thoroughly introduce and summarize previous efforts and further compare them with each other to provide deeper analyses. In the end, we also discuss promising directions and offer our thoughts to facilitate future research. This survey is for both beginners and experts in cross-lingual summarization, and we hope it will serve as a starting point as well as a source of new ideas for researchers and engineers interested in this area.
Jiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang, Zhixu Li, Jianfeng Qu, Jie Zhou 0016
Trans. Assoc. Comput. Linguistics2
2022 An AST Structure Enhanced Decoder for Code Generation
abstract
Currently, the most dominant neural code generation modelsare often equipped with a tree-structured LSTM decoder, which outputs a sequence of actions to construct an Abstract Syntax Tree (AST) via pre-order traversal. However, such a decoder has two obvious drawbacks. First, except for the parent action, other faraway and important history actions rarely contribute to the current decision. Second, it also neglects future actions, which may be crucial for the prediction of the current action. To deal with these issues, in this paper, we propose a novel AST structure enhanced decoder for code generation, which significantly extends the decoder with respect to the above two aspects. First, we introduce an AST information enhanced attention mechanism to fully exploit history actions, of which impacts are further distinguished according to their syntactic distances, action types and relative positions; Second, we jointly model the predictions of current action and its important future action via multi-task learning, where the learned hidden state of the latter can be further leveraged to improve the former. Experimental results on commonly-used datasets demonstrate the effectiveness of our proposed decoder.1
Linfeng Song, Yubin Ge, Fandong Meng, Junfeng Yao, Jinsong Su
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Infusing Multi-Source Knowledge with Heterogeneous Graph Neural Network for Emotional Conversation Generation
abstract
The success of emotional conversation systems depends on sufficient perception and appropriate expression of emotions. In a real-world conversation, we firstly instinctively perceive emotions from multi-source information, including the emotion flow of dialogue history, facial expressions, and personalities of speakers, and then express suitable emotions according to our personalities, but these multiple types of information are insufficiently exploited in emotional conversation fields. To address this issue, we propose a heterogeneous graph-based model for emotional conversation generation. Specifically, we design a Heterogeneous Graph-Based Encoder to represent the conversation content (i.e., the dialogue history, its emotion flow, facial expressions, and speakers' personalities) with a heterogeneous graph neural network, and then predict suitable emotions for feedback. After that, we employ an Emotion-Personality-Aware Decoder to generate a response not only relevant to the conversation context but also with appropriate emotions, by taking the encoded graph representations, the predicted emotions from the encoder and the personality of the current speaker as inputs. Experimental results show that our model can effectively perceive emotions from multi-source knowledge and generate a satisfactory response, which significantly outperforms previous state-of-the-art models.
Yunlong Liang, Fandong Meng, Ying Zhang 0084, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
AAAI2
2021 Faster Depth-Adaptive Transformers
abstract
Depth-adaptive neural networks can dynamically adjust depths according to the hardness of input words, and thus improve efficiency. The main challenge is how to measure such hardness and decide the required depths (i.e., layers) to conduct. Previous works generally build a halting unit to decide whether the computation should continue or stop at each layer. As there is no specific supervision of depth selection, the halting unit may be under-optimized and inaccurate, which results in suboptimal and unstable performance when modeling sentences. In this paper, we get rid of the halting unit and estimate the required depths in advance, which yields a faster depth-adaptive model. Specifically, two approaches are proposed to explicitly measure the hardness of input words and estimate corresponding adaptive depth, namely 1) mutual information (MI) based estimation and 2) reconstruction loss based estimation. We conduct experiments on the text classification task with 24 datasets in various sizes and domains. Results confirm that our approaches can speed up the vanilla Transformer (up to 7x) while preserving high accuracy. Moreover, efficiency and robustness are significantly improved when compared with other depth-adaptive approaches.
Yijin Liu, Fandong Meng, Jie Zhou 0016, Yufeng Chen 0005, Jin An Xu
AAAI2
2021 Selective Knowledge Distillation for Neural Machine Translation
abstract
Fusheng Wang, Jianhao Yan, Fandong Meng, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Fusheng Wang 0008, Jianhao Yan, Fandong Meng, Jie Zhou 0016
ACL/IJCNLP (1)3
2021 Exploring Dynamic Selection of Branch Expansion Orders for Code Generation
abstract
Hui Jiang, Chulun Zhou, Fandong Meng, Biao Zhang, Jie Zhou, Degen Huang, Qingqiang Wu, Jinsong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chulun Zhou, Fandong Meng, Biao Zhang 0002, Jie Zhou 0016, Degen Huang, Qingqiang Wu 0001, Jinsong Su
ACL/IJCNLP (1)3
2021 Modeling Bilingual Conversational Characteristics for Neural Chat Translation
abstract
Yunlong Liang, Fandong Meng, Yufeng Chen, Jinan Xu, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yunlong Liang, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
ACL/IJCNLP (1)2
2021 Marginal Utility Diminishes: Exploring the Minimum Knowledge for BERT Knowledge Distillation
abstract
Yuanxin Liu, Fandong Meng, Zheng Lin, Weiping Wang, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yuanxin Liu, Fandong Meng, Zheng Lin 0001, Weiping Wang 0005, Jie Zhou 0016
ACL/IJCNLP (1)2
2021 Prevent the Language Model from being Overconfident in Neural Machine Translation
abstract
Mengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Mengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou, Jie Zhou 0016
ACL/IJCNLP (1)2
2021 GTM: A Generative Triple-wise Model for Conversational Question Generation
abstract
Lei Shen, Fandong Meng, Jinchao Zhang, Yang Feng, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Lei Shen 0001, Fandong Meng, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016
ACL/IJCNLP (1)2
2021 Modeling Explicit Concerning States for Reinforcement Learning in Visual Dialogue
Zipeng Xu, Fandong Meng, Xiaojie Wang 0006, Duo Zheng, Chenxu Lv, Jie Zhou 0016
BMVC2
2021 Improving Graph-based Sentence Ordering with Iteratively Predicted Pairwise Orderings
abstract
Shaopeng Lai, Ante Wang, Fandong Meng, Jie Zhou, Yubin Ge, Jiali Zeng, Junfeng Yao, Degen Huang, Jinsong Su. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Shaopeng Lai, Ante Wang, Fandong Meng, Jie Zhou 0016, Yubin Ge, Jiali Zeng, Junfeng Yao, Degen Huang, Jinsong Su
EMNLP (1)3
2021 Towards Making the Most of Dialogue Characteristics for Neural Chat Translation
abstract
Neural Chat Translation (NCT) aims to translate conversational text between speakers of different languages.Despite the promising performance of sentence-level and context-aware neural machine translation models, there still remain limitations in current NCT models because the inherent dialogue characteristics of chat, such as dialogue coherence and speaker personality, are neglected.In this paper, we propose to promote the chat translation by introducing the modeling of dialogue characteristics into the NCT model.To this end, we design four auxiliary tasks including monolingual response generation, cross-lingual response generation, next utterance discrimination, and speaker identification.Together with the main chat translation task, we optimize the NCT model through the training objectives of all these tasks.By this means, the NCT model can be enhanced by capturing the inherent dialogue characteristics, thus generating more coherent and speaker-relevant translations.Comprehensive experiments on four language directions (English⇔German and English⇔Chinese) verify the effectiveness and superiority of the proposed approach.
Yunlong Liang, Chulun Zhou, Fandong Meng, Jin An Xu, Yufeng Chen 0005, Jinsong Su, Jie Zhou 0016
EMNLP (1)3
2021 Scheduled Sampling Based on Decoding Steps for Neural Machine Translation
abstract
Scheduled sampling is widely used to mitigate the exposure bias problem for neural machine translation.Its core motivation is to simulate the inference scene during training by replacing ground-truth tokens with predicted tokens, thus bridging the gap between training and inference.However, vanilla scheduled sampling is merely based on training steps and equally treats all decoding steps.Namely, it simulates an inference scene with uniform error rates, which disobeys the real inference scene, where larger decoding steps usually have higher error rates due to error accumulations.To alleviate the above discrepancy, we propose scheduled sampling methods based on decoding steps, increasing the selection chance of predicted tokens with the growth of decoding steps.Consequently, we can more realistically simulate the inference scene during training, thus better bridging the gap between training and inference.Moreover, we investigate scheduled sampling based on both training steps and decoding steps for further improvements.Experimentally, our approaches significantly outperform the Transformer baseline and vanilla scheduled sampling on three large-scale WMT tasks.Additionally, our approaches also generalize well to the text summarization task on two popular benchmarks.
Yijin Liu, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
EMNLP (1)2
2021 Context Tracking Network: Graph-based Context Modeling for Implicit Discourse Relation Recognition
abstract
Yingxue Zhang, Fandong Meng, Peng Li, Ping Jian, Jie Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yingxue Zhang 0003, Fandong Meng, Peng Li 0030, Ping Jian, Jie Zhou 0016
NAACL-HLT2
2021 Sequence-Level Training for Non-Autoregressive Neural Machine Translation
abstract
Abstract In recent years, Neural Machine Translation (NMT) has achieved notable results in various translation tasks. However, the word-by-word generation manner determined by the autoregressive mechanism leads to high translation latency of the NMT and restricts its low-latency applications. Non-Autoregressive Neural Machine Translation (NAT) removes the autoregressive mechanism and achieves significant decoding speedup by generating target words independently and simultaneously. Nevertheless, NAT still takes the word-level cross-entropy loss as the training objective, which is not optimal because the output of NAT cannot be properly evaluated due to the multimodality problem. In this article, we propose using sequence-level training objectives to train NAT models, which evaluate the NAT outputs as a whole and correlates well with the real translation quality. First, we propose training NAT models to optimize sequence-level evaluation metrics (e.g., BLEU) based on several novel reinforcement algorithms customized for NAT, which outperform the conventional method by reducing the variance of gradient estimation. Second, we introduce a novel training objective for NAT models, which aims to minimize the Bag-of-N-grams (BoN) difference between the model output and the reference sentence. The BoN training objective is differentiable and can be calculated efficiently without doing any approximations. Finally, we apply a three-stage training strategy to combine these two methods to train the NAT model. We validate our approach on four translation tasks (WMT14 En↔De, WMT16 En↔Ro), which shows that our approach largely outperforms NAT baselines and achieves remarkable performance on all translation tasks. The source code is available at https://github.com/ictnlp/Seq-NAT.
Chenze Shao, Yang Feng 0004, Jinchao Zhang 0001, Fandong Meng, Jie Zhou 0016
Comput. Linguistics4
2021 A dependency syntactic knowledge augmented interactive architecture for end-to-end aspect-based sentiment analysis
Yunlong Liang, Fandong Meng, Jinchao Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
Neurocomputing2
2021 Simulated annealing for optimization of graphs and sequences
Xianggen Liu, Pengyong Li, Fandong Meng, Hao Zhou 0012, Huasong Zhong, Jie Zhou 0016, Lili Mou, Sen Song
Neurocomputing3
2021 MS-Ranker: Accumulating evidence from potentially correct candidates via reinforcement learning for answer selection
Yingxue Zhang 0003, Fandong Meng, Peng Li 0030, Ping Jian, Jie Zhou 0016
Neurocomputing2
2020 DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual Dialog
abstract
Visual Dialog is a vision-language task that requires an AI agent to engage in a conversation with humans grounded in an image. It remains a challenging task since it requires the agent to fully understand a given question before making an appropriate response not only from the textual dialog history, but also from the visually-grounded information. While previous models typically leverage single-hop reasoning or single-channel reasoning to deal with this complex multimodal reasoning task, which is intuitively insufficient. In this paper, we thus propose a novel and more powerful Dual-channel Multi-hop Reasoning Model for Visual Dialog, named DMRM. DMRM synchronously captures information from the dialog history and the image to enrich the semantic representation of the question by exploiting dual-channel reasoning. Specifically, DMRM maintains a dual channel to obtain the question- and history-aware image features and the question- and image-aware dialog history features by a mulit-hop reasoning process in each channel. Additionally, we also design an effective multimodal attention to further enhance the decoder to generate more accurate responses. Experimental results on the VisDial v0.9 and v1.0 datasets demonstrate that the proposed model is effective and outperforms compared models by a significant margin.
Fandong Meng, Jiaming Xu 0001, Peng Li 0030, Bo Xu 0002, Jie Zhou 0016
AAAI2
2020 Multi-Zone Unit for Recurrent Neural Networks
Fandong Meng, Jinchao Zhang 0001, Yang Liu 0005, Jie Zhou 0016
AAAI1
2020 Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine Translation
abstract
Non-Autoregressive Neural Machine Translation (NAT) achieves significant decoding speedup through generating target words independently and simultaneously. However, in the context of non-autoregressive translation, the word-level cross-entropy loss cannot model the target-side sequential dependency properly, leading to its weak correlation with the translation quality. As a result, NAT tends to generate influent translations with over-translation and under-translation errors. In this paper, we propose to train NAT to minimize the Bag-of-Ngrams (BoN) difference between the model output and the reference sentence. The bag-of-ngrams training objective is differentiable and can be efficiently calculated, which encourages NAT to capture the target-side sequential dependency and correlates well with the translation quality. We validate our approach on three translation tasks and show that our approach largely outperforms the NAT baseline by about 5.0 BLEU scores on WMT14 En↔De and about 2.5 BLEU scores on WMT16 En↔Ro.
Chenze Shao, Jinchao Zhang 0001, Yang Feng 0004, Fandong Meng, Jie Zhou 0016
AAAI4
2020 Enhancing Pointer Network for Sentence Ordering with Pairwise Ordering Predictions
abstract
Dominant sentence ordering models use a pointer network decoder to generate ordering sequences in a left-to-right fashion. However, such a decoder only exploits the noisy left-side encoded context, which is insufficient to ensure correct sentence ordering. To address this deficiency, we propose to enhance the pointer network decoder by using two pairwise ordering prediction modules: The FUTURE module predicts the relative orientations of other unordered sentences with respect to the candidate sentence, and the HISTORY module measures the local coherence between several (e.g., 2) previously ordered sentences and the candidate sentence, without the influence of noisy left-side context. Using the pointer mechanism, we then incorporate this dynamically generated information into the decoder as a supplement to the left-side context for better predictions. On several commonly-used datasets, our model significantly outperforms other baselines, achieving the state-of-the-art performance. Further analyses verify that pairwise ordering predictions indeed provide extra useful context as expected, leading to better sentence ordering. We also evaluate our sentence ordering models on a downstream task, multi-document summarization, and the summaries reordered by our model achieve the best coherence scores. Our code is available at https://github.com/DeepLearnXMU/Pairwise.git.
Yongjing Yin, Fandong Meng, Jinsong Su, Yubin Ge, Linfeng Song, Jie Zhou 0016, Jiebo Luo 0001
AAAI2
2020 Unsupervised Paraphrasing by Simulated Annealing
abstract
We propose UPSA, a novel approach that accomplishes Unsupervised Paraphrasing by Simulated Annealing.We model paraphrase generation as an optimization problem and propose a sophisticated objective function, involving semantic similarity, expression diversity, and language fluency of paraphrases.UPSA searches the sentence space towards this objective by performing a sequence of local edits.We evaluate our approach on various datasets, namely, Quora, Wikianswers, MSCOCO, and Twitter.Extensive results show that UPSA achieves the state-of-the-art performance compared with previous unsupervised methods in terms of both automatic and human evaluations.Further, our approach outperforms most existing domain-adapted supervised models, showing the generalizability of UPSA. 1
Xianggen Liu, Lili Mou, Fandong Meng, Hao Zhou 0012, Jie Zhou 0016, Sen Song
ACL3
2020 A Contextual Hierarchical Attention Network with Adaptive Objective for Dialogue State Tracking
abstract
Recent studies in dialogue state tracking (DST) leverage historical information to determine states which are generally represented as slot-value pairs. However, most of them have limitations to efficiently exploit relevant context due to the lack of a powerful mechanism for modeling interactions between the slot and the dialogue history. Besides, existing methods usually ignore the slot imbalance problem and treat all slots indiscriminately, which limits the learning of hard slots and eventually hurts overall performance. In this paper, we propose to enhance the DST through employing a contextual hierarchical attention network to not only discern relevant information at both word level and turn level but also learn contextual representations. We further propose an adaptive objective to alleviate the slot imbalance problem by dynamically adjust weights of different slots during training. Experimental results show that our approach reaches 52.68% and 58.55% joint accuracy on MultiWOZ 2.0 and MultiWOZ 2.1 datasets respectively and achieves new state-of-the-art performance with considerable improvements (+1.24% and +5.98%).
Yong Shan, Zekang Li, Jinchao Zhang 0001, Fandong Meng, Yang Feng 0004, Cheng Niu, Jie Zhou 0016
ACL4
2020 A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine Translation
abstract
Multi-modal neural machine translation (NMT) aims to translate source sentences into a target language paired with images. However, dominant multi-modal NMT models do not fully exploit fine-grained semantic correspondences between semantic units of different modalities, which have potential to refine multi-modal representation learning. To deal with this issue, in this paper, we propose a novel graph-based multi-modal fusion encoder for NMT. Specifically, we first represent the input sentence and image using a unified multi-modal graph, which captures various semantic relationships between multi-modal semantic units (words and visual objects). We then stack multiple graph-based multi-modal fusion layers that iteratively perform semantic interactions to learn node representations. Finally, these representations provide an attention-based context vector for the decoder. We evaluate our proposed encoder on the Multi30K datasets. Experimental results and in-depth analysis show the superiority of our multi-modal NMT model.
Yongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou, Zhengyuan Yang, Jie Zhou 0016, Jiebo Luo 0001
ACL2
2020 Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue Generation
abstract
Knowledge selection plays an important role in knowledge-grounded dialogue, which is a challenging task to generate more informative responses by leveraging external knowledge.Recently, latent variable models have been proposed to deal with the diversity of knowledge selection by using both prior and posterior distributions over knowledge and achieve promising performance.However, these models suffer from a huge gap between prior and posterior knowledge selection.Firstly, the prior selection module may not learn to select knowledge properly because of lacking the necessary posterior information.Secondly, latent variable models suffer from the exposure bias that dialogue generation is based on the knowledge selected from the posterior distribution at training but from the prior distribution at inference.Here, we deal with these issues on two aspects: (1) We enhance the prior selection module with the necessary posterior information obtained from the specially designed Posterior Information Prediction Module (PIPM); (2) We propose a Knowledge Distillation Based Training Strategy (KDBTS) to train the decoder with the knowledge selected from the prior distribution, removing the exposure bias of knowledge selection.Experimental results on two knowledge-grounded dialogue datasets show that both PIPM and KDBTS achieve performance improvement over the state-of-theart latent variable model and their combination shows further improvement.
Xiuyi Chen, Fandong Meng, Peng Li 0030, Bo Xu 0002, Jie Zhou 0016
EMNLP (1)2
2020 Token-level Adaptive Training for Neural Machine Translation
abstract
There exists a token imbalance phenomenon in natural language as different tokens appear with different frequencies, which leads to different learning difficulties for tokens in Neural Machine Translation (NMT).The vanilla NMT model usually adopts trivial equal-weighted objectives for target tokens with different frequencies and tends to generate more high-frequency tokens and less lowfrequency tokens compared with the golden token distribution.However, low-frequency tokens may carry critical semantic information that will affect the translation quality once they are neglected.In this paper, we explored target token-level adaptive objectives based on token frequencies to assign appropriate weights for each target token during training.We aimed that those meaningful but relatively low-frequency words could be assigned with larger weights in objectives to encourage the model to pay more attention to these tokens.Our method yields consistent improvements in translation quality on ZH-EN, EN-RO, and EN-DE translation tasks, especially on sentences that contain more low-frequency tokens where we can get 1.68, 1.02, and 0.52 BLEU increases compared with baseline, respectively.Further analyses show that our method can also improve the lexical diversity of translation.
Shuhao Gu, Jinchao Zhang 0001, Fandong Meng, Yang Feng 0004, Wanying Xie, Jie Zhou 0016, Dong Yu 0003
EMNLP (1)3
2020 Multi-Unit Transformers for Neural Machine Translation
abstract
Transformer models (Vaswani et al., 2017) achieve remarkable success in Neural Machine Translation.Many efforts have been devoted to deepening the Transformer by stacking several units (i.e., a combination of Multihead Attentions and FFN) in a cascade, while the investigation over multiple parallel units draws little attention.In this paper, we propose the Multi-Unit TransformErs (MUTE), which aim to promote the expressiveness of the Transformer by introducing diverse and complementary units.Specifically, we use several parallel units and show that modeling with multiple units improves model performance and introduces diversity.Further, to better leverage the advantage of the multi-unit setting, we design biased module and sequential dependency that guide and encourage complementariness among different units.Experimental results on three machine translation tasks, the NIST Chinese-to-English, WMT'14 English-to-German and WMT'18 Chinese-to-English, show that the MUTE models significantly outperform the Transformer-Base, by up to +1.52, +1.90 and +1.10 BLEU points, with only a mild drop in inference speed (about 3.1%).In addition, our methods also surpass the Transformer-Big model, with only 54% of its parameters.These results demonstrate the effectiveness of the MUTE, as well as its efficiency in both the inference process and parameter usage. 1
Jianhao Yan, Fandong Meng, Jie Zhou 0016
EMNLP (1)2
2020 Dynamic Context-guided Capsule Network for Multimodal Machine Translation
abstract
Multimodal machine translation (MMT), which mainly focuses on enhancing text-only translation with visual features, has attracted considerable attention from both computer vision and natural language processing communities. Most current MMT models resort to attention mechanism, global context modeling or multimodal joint representation learning to utilize visual features. However, the attention mechanism lacks sufficient semantic interactions between modalities while the other two provide fixed visual context, which is unsuitable for modeling the observed variability when generating translation. To address the above issues, in this paper, we propose a novel Dynamic Context-guided Capsule Network (DCCN) for MMT. Specifically, at each timestep of decoding, we first employ the conventional source-target attention to produce a timestep-specific source-side context vector. Next, DCCN takes this vector as input and uses it to guide the iterative extraction of related visual features via a context-guided dynamic routing mechanism. Particularly, we represent the input image with global and regional visual features, we introduce two parallel DCCNs to model multimodal context vectors with visual features at different granularities. Finally, we obtain two multimodal context vectors, which are fused and incorporated into the decoder for the prediction of the target word. Experimental results on the Multi30K dataset of English-to-German and English-to-French translation demonstrate the superiority of DCCN. Our code is available on https://github.com/DeepLearnXMU/MM-DCCN.
Fandong Meng, Jinsong Su, Yongjing Yin, Zhengyuan Yang, Yubin Ge, Jie Zhou 0016, Jiebo Luo 0001
ACM Multimedia2
2019 DTMT: A Novel Deep Transition Architecture for Neural Machine Translation
abstract
Past years have witnessed rapid developments in Neural Machine Translation (NMT). Most recently, with advanced modeling and training techniques, the RNN-based NMT (RNMT) has shown its potential strength, even compared with the well-known Transformer (self-attentional) model. Although the RNMT model can possess very deep architectures through stacking layers, the transition depth between consecutive hidden states along the sequential axis is still shallow. In this paper, we further enhance the RNN-based NMT through increasing the transition depth between consecutive hidden states and build a novel Deep Transition RNN-based Architecture for Neural Machine Translation, named DTMT. This model enhances the hidden-to-hidden transition with multiple non-linear transformations, as well as maintains a linear transformation path throughout this deep transition by the well-designed linear transformation mechanism to alleviate the gradient vanishing problem. Experiments show that with the specially designed deep transition modules, our DTMT can achieve remarkable improvements on translation quality. Experimental results on Chinese⇒English translation task show that DTMT can outperform the Transformer model by +2.09 BLEU points and achieve the best results ever reported in the same dataset. On WMT14 English⇒German and English⇒French translation tasks, DTMT shows superior quality to the state-of-the-art NMT systems, including the Transformer and the RNMT+.
Fandong Meng, Jinchao Zhang 0001
AAAI1
2019 Incremental Transformer with Deliberation Decoder for Document Grounded Conversations
abstract
Document Grounded Conversations is a task to generate dialogue responses when chatting about the content of a given document. Obviously, document knowledge plays a critical role in Document Grounded Conversations, while existing dialogue models do not exploit this kind of knowledge effectively enough. In this paper, we propose a novel Transformer-based architecture for multi-turn document grounded conversations. In particular, we devise an Incremental Transformer to encode multi-turn utterances along with knowledge in related documents. Motivated by the human cognitive process, we design a two-pass decoder (Deliberation Decoder) to improve context coherence and knowledge correctness. Our empirical study on a real-world Document Grounded Dataset proves that responses generated by our model significantly outperform competitive baselines on both context coherence and knowledge relevance.
Zekang Li, Cheng Niu, Fandong Meng, Yang Feng 0004, Jie Zhou 0016
ACL (1)3
2019 GCDT: A Global Context Enhanced Deep Transition Architecture for Sequence Labeling
abstract
Current state-of-the-art systems for the sequence labeling tasks are typically based on the family of Recurrent Neural Networks (RNNs).However, the shallow connections between consecutive hidden states of RNNs and insufficient modeling of global information restrict the potential performance of those models.In this paper, we try to address these issues, and thus propose a Global Context enhanced Deep Transition architecture for sequence labeling named GCDT.We deepen the state transition path at each position in a sentence, and further assign every token with a global representation learned from the entire sentence.Experiments on two standard sequence labeling tasks show that, given only training data and the ubiquitous word embeddings (Glove), our GCDT achieves 91.96 F 1 on the CoNLL03 NER task and 95.43 F 1 on the CoNLL2000 Chunking task, which outperforms the best reported results under the same settings.Furthermore, by leveraging BERT as an additional resource, we establish new stateof-the-art results with 93.47 F 1 on NER and 97.30F 1 on Chunking 1 .
Yijin Liu, Fandong Meng, Jinchao Zhang 0001, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)2
2019 Retrieving Sequential Information for Non-Autoregressive Neural Machine Translation
abstract
Non-Autoregressive Transformer (NAT) aims to accelerate the Transformer model through discarding the autoregressive mechanism and generating target words independently, which fails to exploit the target sequential information.Over-translation and under-translation errors often occur for the above reason, especially in the long sentence translation scenario.In this paper, we propose two approaches to retrieve the target sequential information for NAT to enhance its translation ability while preserving the fast-decoding property.Firstly, we propose a sequence-level training method based on a novel reinforcement algorithm for NAT (Reinforce-NAT) to reduce the variance and stabilize the training procedure.Secondly, we propose an innovative Transformer decoder named FS-decoder to fuse the target sequential information into the top layer of the decoder.Experimental results on three translation tasks show that the Reinforce-NAT surpasses the baseline NAT system by a significant margin on BLEU without decelerating the decoding speed and the FS-decoder achieves comparable translation performance to the autoregressive Transformer with considerable speedup.
Chenze Shao, Yang Feng 0004, Jinchao Zhang 0001, Fandong Meng, Xilin Chen 0001, Jie Zhou 0016
ACL (1)4
2019 Bridging the Gap between Training and Inference for Neural Machine Translation
abstract
Neural Machine Translation (NMT) generates target words sequentially in the way of predicting the next word conditioned on the context words.At training time, it predicts with the ground truth words as context while at inference it has to generate the entire sequence from scratch.This discrepancy of the fed context leads to error accumulation among the way.Furthermore, word-level training requires strict matching between the generated sequence and the ground truth sequence which leads to overcorrection over different but reasonable translations.In this paper, we address these issues by sampling context words not only from the ground truth sequence but also from the predicted sequence by the model during training, where the predicted sequence is selected with a sentence-level optimum.Experiment results on Chinese→English and WMT'14 English→German translation tasks demonstrate that our approach can achieve significant improvements on multiple datasets.
Wen Zhang 0009, Yang Feng 0004, Fandong Meng, Di You, Qun Liu 0001
ACL (1)3
2019 A Novel Aspect-Guided Deep Transition Model for Aspect Based Sentiment Analysis
abstract
Yunlong Liang, Fandong Meng, Jinchao Zhang, Jinan Xu, Yufeng Chen, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yunlong Liang, Fandong Meng, Jinchao Zhang 0001, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
EMNLP/IJCNLP (1)2
2019 CM-Net: A Novel Collaborative Memory Network for Spoken Language Understanding
abstract
Yijin Liu, Fandong Meng, Jinchao Zhang, Jie Zhou, Yufeng Chen, Jinan Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yijin Liu, Fandong Meng, Jinchao Zhang 0001, Jie Zhou 0016, Yufeng Chen 0005, Jin An Xu
EMNLP/IJCNLP (1)2
2019 Enhancing Context Modeling with a Query-Guided Capsule Network for Document-level Translation
abstract
Zhengxin Yang, Jinchao Zhang, Fandong Meng, Shuhao Gu, Yang Feng, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zhengxin Yang, Jinchao Zhang 0001, Fandong Meng, Shuhao Gu, Yang Feng 0004, Jie Zhou 0016
EMNLP/IJCNLP (1)3
2018 Towards Robust Neural Machine Translation
abstract
Small perturbations in the input can severely distort intermediate representations and thus impact translation quality of neural machine translation (NMT) models.In this paper, we propose to improve the robustness of NMT models with adversarial stability training.The basic idea is to make both the encoder and decoder in NMT models robust against input perturbations by enabling them to behave similarly for the original input and its perturbed counterpart.Experimental results on Chinese-English, English-German and English-French translation tasks show that our approaches can not only achieve significant improvements over strong NMT systems but also improve the robustness of NMT models.
Yong Cheng 0003, Zhaopeng Tu, Fandong Meng, Junjie Zhai, Yang Liu 0005
ACL (1)3
2018 Modeling Localness for Self-Attention Networks
abstract
Self-attention networks have proven to be of profound value for its strength of capturing global dependencies.In this work, we propose to model localness for self-attention networks, which enhances the ability of capturing useful local context.We cast localness modeling as a learnable Gaussian bias, which indicates the central and scope of the local region to be paid more attention.The bias is then incorporated into the original attention distribution to form a revised distribution.To maintain the strength of capturing long distance dependencies and enhance the ability of capturing shortrange dependencies, we only apply localness modeling to lower layers of self-attention networks.Quantitative and qualitative analyses on Chinese⇒English and English⇒German translation tasks demonstrate the effectiveness and universality of the proposed approach.
Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, Tong Zhang 0001
EMNLP4
2018 Neural Machine Translation with Key-Value Memory-Augmented Attention
abstract
Although attention-based Neural Machine Translation (NMT) has achieved remarkable progress in recent years, it still suffers from issues of repeating and dropping translations. To alleviate these issues, we propose a novel key-value memory-augmented attention model for NMT, called KVMEMATT. Specifically, we maintain a timely updated keymemory to keep track of attention history and a fixed value-memory to store the representation of source sentence throughout the whole translation process. Via nontrivial transformations and iterative interactions between the two memories, the decoder focuses on more appropriate source word(s) for predicting the next target word at each decoding step, therefore can improve the adequacy of translations. Experimental results on Chinese)English and WMT17 German,English translation tasks demonstrate the superiority of the proposed model.
Fandong Meng, Zhaopeng Tu, Yong Cheng 0003, Junjie Zhai, Yuekui Yang
IJCAI1
2016 Interactive Attention for Neural Machine Translation
abstract
Conventional attention-based Neural Machine Translation (NMT) conducts dynamic alignment in generating the target sentence. By repeatedly reading the representation of source sentence, which keeps fixed after generated by the encoder (Bahdanau et al., 2015), the attention mechanism has greatly enhanced state-of-the-art NMT. In this paper, we propose a new attention mechanism, called INTERACTIVE ATTENTION, which models the interaction between the decoder and the representation of source sentence during translation by both reading and writing operations. INTERACTIVE ATTENTION can keep track of the interaction history and therefore improve the translation performance. Experiments on NIST Chinese-English translation task show that INTERACTIVE ATTENTION can achieve significant improvements over both the previous attention-based NMT baseline and some state-of-the-art variants of attention-based NMT (i.e., coverage models (Tu et al., 2016)). And neural machine translator with our INTERACTIVE ATTENTION can outperform the open source attention-based NMT system Groundhog by 4.22 BLEU points and the open source phrase-based system Moses by 3.94 BLEU points averagely on multiple test sets.
Fandong Meng, Zhengdong Lu, Hang Li 0001, Qun Liu 0001
COLING1
2016 Topic-based term translation models for statistical machine translation
Deyi Xiong, Fandong Meng, Qun Liu 0001
Artif. Intell.2
2015 Encoding Source Language with Convolutional Neural Network for Machine Translation
abstract
Fandong Meng, Zhengdong Lu, Mingxuan Wang, Hang Li, Wenbin Jiang, Qun Liu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Fandong Meng, Zhengdong Lu, Mingxuan Wang, Hang Li 0001, Wenbin Jiang 0002, Qun Liu 0001
ACL (1)1
2014 A Dependency Edge-based Transfer Model for Statistical Machine Translation
Hongshen Chen, Fandong Meng, Wenbin Jiang 0002, Qun Liu 0001
COLING3
2014 Modeling Term Translation for Document-informed Machine Translation
abstract
Term translation is of great importance for statistical machine translation (SMT), especially document-informed SMT.In this paper, we investigate three issues of term translation in the context of documentinformed SMT and propose three corresponding models: (a) a term translation disambiguation model which selects desirable translations for terms in the source language with domain information, (b) a term translation consistency model that encourages consistent translations for terms with a high strength of translation consistency throughout a document, and (c) a term bracketing model that rewards translation hypotheses where bracketable source terms are translated as a whole unit.We integrate the three models into hierarchical phrase-based SMT and evaluate their effectiveness on NIST Chinese-English translation tasks with large-scale training data.Experiment results show that all three models can achieve significant improvements over the baseline.Additionally, we can obtain a further improvement when combining the three models.
Fandong Meng, Deyi Xiong, Wenbin Jiang 0002, Qun Liu 0001
EMNLP1
2013 Translation with Source Constituency and Dependency Trees
abstract
We present a novel translation model, which simultaneously exploits the constituency and dependency trees on the source side, to combine the advantages of two types of trees.We take head-dependents relations of dependency trees as backbone and incorporate phrasal nodes of constituency trees as the source side of our translation rules, and the target side as strings.Our rules hold the property of long distance reorderings and the compatibility with phrases.Large-scale experimental results show that our model achieves significantly improvements over the constituency-to-string (+2.45 BLEU on average) and dependencyto-string (+0.91 BLEU on average) models, which only employ single type of trees, and significantly outperforms the state-of-theart hierarchical phrase-based model (+1.12BLEU on average), on three Chinese-English NIST test sets.
Fandong Meng, Linfeng Song, Yajuan Lü, Qun Liu 0001
EMNLP1
2012 Iterative Annotation Transformation with Predict-Self Reestimation for Chinese Word Segmentation
Wenbin Jiang 0002, Fandong Meng, Qun Liu 0001, Yajuan Lü
EMNLP-CoNLL2