Yufeng Chen 0005

dblp:64/5715-5 · DBLP profile ↗
← Back
65ranked-venue papers
0as first author
54since 2021 · last 2026
0000-0003-0437-6788ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 63 · 52 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning
abstract
Xue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang, Kaiyu Huang, Yufeng Chen, Xu Jinan, Jie Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yunlong Liang, Fandong Meng, Songming Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
ACL (1)6
2026 CroSearch-R1: Better Leveraging Cross-lingual Knowledge for Retrieval-Augmented Generation
abstract
A multilingual collection may contain useful knowledge in other languages to supplement and correct the facts in the original language for Retrieval-Augmented Generation (RAG). However, the vanilla approach that simply concatenates multiple pieces of knowledge from different languages into the context may fail to improve effectiveness due to the potential disparities across languages. To better leverage multilingual knowledge, we propose CroSearch-R1, a search-augmented reinforcement learning framework to integrate multilingual knowledge into the Group Relative Policy Optimization (GRPO) process. In particular, the approach adopts a multi-turn retrieval strategy with cross-lingual knowledge integration to dynamically align the knowledge from other languages as supplementary evidence into a unified representation space. Furthermore, we introduce a multilingual rollout mechanism to optimize reasoning transferability across languages. Experimental results demonstrate that our framework effectively leverages cross-lingual complementarity and improves the effectiveness of RAG with multilingual collections.
Fengran Mo, Sijin Lu, Yufeng Chen 0005, Jian-Yun Nie
SIGIR4
2026 DKF: Domain knowledge fusion in progressive incremental learning for multi-domain machine translation
Zhibo Man, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu
Expert Syst. Appl.4
2025 AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
abstract
In modern large language models (LLMs), LLM alignment is of crucial importance and is typically achieved through methods such as reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO).However, in most existing methods for LLM alignment, all tokens in the response are optimized using a sparse, response-level reward or preference annotation.The ignorance of token-level rewards may erroneously punish high-quality tokens or encourage lowquality tokens, resulting in suboptimal performance and slow convergence speed.To address this issue, we propose AlignDistil, an RLHFequivalent distillation method for token-level reward optimization.Specifically, we introduce the reward learned by DPO into the RLHF objective and theoretically prove the equivalence between this objective and a token-level distillation process, where the teacher distribution linearly combines the logits from the DPO model and a reference model.On this basis, we further bridge the accuracy gap between the reward from the DPO model and the pure reward model, by building a contrastive DPO reward with a normal and a reverse DPO model.Moreover, to avoid under-and over-optimization on different tokens, we design a token adaptive logit extrapolation mechanism to construct an appropriate teacher distribution for each token.Experimental results demonstrate the superiority of our AlignDistil over existing methods and showcase fast convergence due to its tokenlevel distributional reward optimization.
Songming Zhang 0001, Bojie Hu, Yufeng Chen 0005, Jin An Xu
ACL (1)5
2025 Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts
abstract
Continually expanding new languages for existing large language models (LLMs) is a promising yet challenging approach to building powerful multilingual LLMs.The biggest challenge is to make the model continuously learn new languages while preserving the proficient ability of old languages.To achieve this, recent work utilizes the Mixture-of-Experts (MoE) architecture to expand new languages by adding new experts and avoid catastrophic forgetting of old languages by routing corresponding tokens to the original model backbone (old experts).Although intuitive, this kind of method is parameter-costly when expanding new languages and still inevitably impacts the performance of old languages.To address these limitations, we analyze the language characteristics of different layers in LLMs and propose a layer-wise expert allocation algorithm (LayerMoE) to determine the appropriate number of new experts for each layer.Specifically, we find different layers in LLMs exhibit different representation similarities between languages and then utilize the similarity as the indicator to allocate experts for each layer, i.e., the higher similarity, the fewer experts.Additionally, to further mitigate the forgetting of old languages, we add a classifier in front of the router network on the layers with higher similarity to guide the routing of old language tokens.Experimental results show that our method outperforms the previous state-of-the-art baseline with 60% fewer experts in the single-expansion setting and with 33.3% fewer experts in the lifelong-expansion setting, demonstrating the effectiveness of our method.
Yunlong Liang, Fandong Meng, Songming Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
ACL (1)5
2025 Multilingual Knowledge Editing with Language-Agnostic Factual Neurons
abstract
Multilingual knowledge editing (MKE) aims to simultaneously update factual knowledge across multiple languages within large language models (LLMs). Previous research indicates that the same knowledge across different languages within LLMs exhibits a degree of shareability. However, most existing MKE methods overlook the connections of the same knowledge between different languages, resulting in knowledge conflicts and limited edit performance. To address this issue, we first investigate how LLMs process multilingual factual knowledge and discover that the same factual knowledge in different languages generally activates a shared set of neurons, which we call language-agnostic factual neurons (LAFNs). These neurons represent the same factual knowledge shared across languages and imply the semantic connections among multilingual knowledge. Inspired by this finding, we propose a new MKE method by Locating and Updating Language-Agnostic Factual Neurons (LU-LAFNs) to edit multilingual knowledge simultaneously, which avoids knowledge conflicts and thus improves edit performance. Experimental results on Bi-ZsRE and MzsRE benchmarks demonstrate that our method achieves the best edit performance, indicating the effectiveness and importance of modeling the semantic connections among multilingual knowledge.
Yunlong Liang, Fandong Meng, Songming Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
COLING5
2025 Boosting Data Utilization for Multilingual Dense Retrieval
abstract
Multilingual dense retrieval aims to retrieve relevant documents across different languages based on a unified retriever model.The challenge lies in aligning representations of different languages in a shared vector space.The common practice is to fine-tune the dense retriever via contrastive learning, whose effectiveness highly relies on the quality of the negative samples and the efficacy of mini-batch data.Different from the existing studies that focus on developing sophisticated model architecture, we propose a method to boost data utilization for multilingual dense retrieval by obtaining high-quality hard negative samples and effective mini-batch data.The extensive experimental results on a multilingual retrieval benchmark, MIRACL, with 16 languages demonstrate the effectiveness of our method by outperforming several existing strong baselines.
Fengran Mo, Yufeng Chen 0005, Changhao Guan, Zhenrui Yue, Xinyu Wang 0061, Jin An Xu
EMNLP3
2025 TriG-RAG: Triple-Granularity Fusion for Retrieval-Augmented Generation with Adaptive Context-Relation Balance
Jingrui Zhang, Yufeng Chen 0005, Jin An Xu
NLPCC (3)2
2024 Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine Translation
abstract
Incrementally expanding the capability of an existing translation model to solve new domain tasks over time is a fundamental and practical problem, which usually suffers from catastrophic forgetting.Generally, multi-domain learning can be seen as a good solution.However, there are two drawbacks: 1) it requires having the training data for all domains available at the same time, which may be unrealistic due to storage or privacy concerns; 2) it requires re-training the model on the data of all domains from scratch when adding a new domain and this is time-consuming and computationally expensive.To address these issues, we present a semi-supervised contrastive distillation framework for incremental neural machine translation.Specifically, to avoid catastrophic forgetting, we propose to exploit unlabeled data from the same distributions of the older domains through knowledge distillation.Further, to ensure the distinct domain characteristics in the model as the number of domains increases, we devise a cross-domain contrastive objective to enhance the distilled knowledge.Extensive experiments on domain translation benchmarks show that our approach, without accessing any previous training data or re-training on all domains from scratch, can significantly prevent the model from forgetting previously learned knowledge while obtaining good performance on the incrementally added domains.
Yunlong Liang, Fandong Meng, Jiaan Wang, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)5
2024 CollabKG: A Learnable Human-Machine-Cooperative Information Extraction Toolkit for (Event) Knowledge Graph Construction
abstract
In order to construct or extend entity-centric and event-centric knowledge graphs (KG and EKG), the information extraction (IE) annotation toolkit is essential. However, existing IE toolkits have several non-trivial problems, such as not supporting multi-tasks, and not supporting automatic updates. In this work, we present CollabKG, a learnable human-machine-cooperative IE toolkit for KG and EKG construction. Specifically, for the multi-task issue, CollabKG unifies different IE subtasks, including named entity recognition (NER), entity-relation triple extraction (RE), and event extraction (EE), and supports both KG and EKG. Then, combining advanced prompting-based IE technology, the human-machine-cooperation mechanism with Large Language Models (LLMs) as the assistant machine is presented which can provide a lower cost as well as a higher performance. Lastly, owing to the two-way interaction between the human and machine, CollabKG with learning ability allows self-renewal. Besides, CollabKG has several appealing features (e.g., customization, training-free, and label propagation) that make the system powerful and high-productivity. We holistically compare our toolkit with other existing tools on these features. Human evaluation quantitatively illustrates that CollabKG significantly improves annotation quality, efficiency, and stability simultaneously.
Yufeng Chen 0005, Xingyu Cui, Jin An Xu, Wenjuan Han
LREC/COLING2
2024 A Reinforcement Learning Approach to Improve Low-Resource Machine Translation Leveraging Domain Monolingual Data
abstract
Due to the lack of parallel data, the mainstream fine-tuning-based domain adaptation methods have the overfitting problem in the translation of low-resource domains, and it is difficult for the model to learn the in-domain generalization knowledge. To address the above issue, in this work, we propose a novel Reinforcement Learning Domain Adaptation method for Neural Machine Translation (RLDA-NMT) in the low-resource domain. RLDA-NMT utilizes in-domain source monolingual data to make up for the lack of parallel data, and reinforces domain features learning to make the translation model learn the domain-specific knowledge more fully. Specifically, we first train a ranking-based model with a small-scale in-domain parallel corpus, and then adopt it as the reward model to select higher-quality generated translations for reinforcement when fine-tuning pre-trained NMT model using in-domain source monolingual data. We conduct experiments on Education, Laws, Thesis, and Patent domains of Chinese⇔English translation tasks. Experimental results demonstrate that RLDA-NMT can alleviate overfitting and reinforce the NMT model to learn domain-specific knowledge. Additionally, the results also show that RLDA-NMT and back-translation (BT) are nicely complementary to each other, where combining RLDA-NMT with BT can further improve translation quality.
Hongxiao Zhang, Mingtong Liu, Chunyou Li, Yufeng Chen 0005, Jin An Xu
LREC/COLING4
2024 Dual-Space Knowledge Distillation for Large Language Models
abstract
Knowledge distillation (KD) is known as a promising solution to compress large language models (LLMs) via transferring their knowledge to smaller models.During this process, white-box KD methods usually minimize the distance between the output distributions of the two models so that more knowledge can be transferred.However, in the current whitebox KD framework, the output distributions are from the respective output spaces of the two models, using their own prediction heads.We argue that the space discrepancy will lead to low similarity between the teacher model and the student model on both representation and distribution levels.Furthermore, this discrepancy also hinders the KD process between models with different vocabularies, which is common for current LLMs.To address these issues, we propose a dual-space knowledge distillation (DSKD) framework that unifies the output spaces of the two models for KD.On the basis of DSKD, we further develop a cross-model attention mechanism, which can automatically align the representations of the two models with different vocabularies.Thus, our framework is not only compatible with various distance functions for KD (e.g., KL divergence) like the current framework, but also supports KD between any two LLMs regardless of their vocabularies.Experiments on task-agnostic instructionfollowing benchmarks show that DSKD significantly outperforms the current white-box KD framework with various distance functions, and also surpasses existing KD methods for LLMs with different vocabularies 1 .* Yufeng Chen is the corresponding author.vocabulary, which, however, is hardly satisfied for various LLMs in this era ( §2.2.2).Towards these limitations, we then propose a new framework for white-box KD, named dualspace knowledge distillation (DSKD), which is as simple as the current white-box KD framework but addresses the issues due to the space discrepancy.Specifically, DSKD unifies the output spaces of the two models by projecting the output hidden states 2 of the teacher/student to the representation spaces of the student/teacher, where we can use the shared prediction heads to produce the two distributions in the same output spaces.In particular, for models with different vocabularies, we further develop a cross-model attention (CMA) mechanism to automatically align the tokens in two differently tokenized sequences.Like the current framework, DSKD is also compatible with existing distance functions for distributions, including KL divergence, JS divergence, and so on.Meanwhile, with CMA, we can transform distributions of the two LLMs into the same shape, which makes our framework more general and can be applied to any two LLMs regardless of their vocabularies.We evaluate our framework on instructionfollowing benchmarks under both settings that the two LLMs have the same/different vocabularies.Experimental results showcase that for LLMs with the same vocabulary, our DSKD framework significantly outperforms the current white-box KD framework on various distance functions.Moreover, DSKD with CMA surpasses all existing KD methods for LLMs with different vocabularies.To sum up, the contributions are as follows:• We empirically reveal that the current whitebox KD framework limits the similarity between the student and the teacher due to their different output spaces.• As a solution, we propose a new framework for white-box KD, named dual-space knowledge distillation (DSKD), which unifies the output spaces of the distributions from the teacher and the student for more effective KD.• Based on DSKD, we further develop a crossmodel attention mechanism to support KD between LLMs with different vocabularies.
Songming Zhang 0001, Zengkui Sun, Yufeng Chen 0005, Jin An Xu
EMNLP4
2024 Model-Agnostic Knowledge Distillation Between Heterogeneous Models
Jiaxin Shen, Yanyao Liu, Yong Jiang 0005, Yufeng Chen 0005, Wenjuan Han
NLPCC (1)4
2024 An Ensemble Strategy with Gradient Conflict for Multi-Domain Neural Machine Translation
abstract
Multi-domain neural machine translation aims to construct a unified neural machine translation model to translate sentences across various domains. Nevertheless, previous studies have one limitation is the incapacity to acquire both domain-general and domain-specific representations concurrently. To this end, we propose an ensemble strategy with gradient conflict for multi-domain neural machine translation that automatically learns model parameters by identifying both domain-shared and domain-specific features. Specifically, our approach consists of (1) a parameter-sharing framework, where the parameters of all the layers are originally shared and equivalent to each domain, and (2) ensemble strategy, in which we design an Extra Ensemble strategy via a piecewise condition function to learn direction and distance-based gradient conflict. In addition, we give a detailed theoretical analysis of the gradient conflict to further validate the effectiveness of our approach. Experimental results on two multi-domain datasets show the superior performance of our proposed model compared to previous work.
Zhibo Man, Yu Li 0025, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2024 WDSRL: Multi-Domain Neural Machine Translation With Word-Level Domain-Sensitive Representation Learning
abstract
Due to the strong reliance on domain-specific knowledge, the joint learning manner of domain discrimination and translation has been widely considered in the Multi-Domain Neural Machine Translation (MDNMT) task. However, the word ambiguity problem still inevitably exists in MDNMT, especially when mixed multi-domain data is brought into the model training phase. Although word-level MDNMT can mitigate this problem to some extent, poor domain discrimination yet remains and severely hinders performance. Based on the above limitation, we observed that coarser granularity strings may provide more specific semantics, which is more conducive to domain discrimination. Thus, we propose a Word-level Domain-Sensitive Representation Learning (WDSRL) method. Specifically, we focus on two aspects of our approach: domain representation and domain discrimination. To extend the scope of domain representation, we adopt Convolution Neural Networks (CNN) to encode Local Domain Representation at different granularities, and then integrate Topic Knowledge Representation into each word. By doing so, context features related to the domain could be comprehensively enriched. Regarding domain discrimination, we design a Domain-Sensitive Discriminator, which could not only generate domain features for each word but also enhance domain representation learning. Experimental results demonstrate our substantial improvements over several representative baselines on multiple language pairs. Furthermore, the extensive analysis also indicates the superiority of our proposed domain-sensitive feature encoding strategy and domain-sensitive discriminator for word-level representation learning.
Zhibo Man, Zengcheng Huang, Yu Li 0025, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu
IEEE ACM Trans. Audio Speech Lang. Process.6
2024 Complex Question Enhanced Transfer Learning for Zero-Shot Joint Information Extraction
abstract
Zero-shot information extraction (IE) tasks have attracted great attention recently. However, how to jointly model multiple IE tasks in the zero-shot scenario is still an open question. In this article, we focus on zero-shot joint IE tasks and highlight how to transfer the knowledge of cross-task relations from the source domain to the target domain. To solve this problem, we first unify all IE tasks with a machine reading comprehension (MRC) framework, which can make the most of training data and enhance its ability on span extraction. Then, we generatecomplex questionsto explicitly model cross-task relations with natural language descriptions, thereby providing prior knowledge for pre-defined types and building more general linkages among different entities and triggers as well. Specifically, we define three operations for generating templates for complex questions, i.e.,intersecting,connecting, andcomposing. Besides, we design an efficient training strategy to exploit the synthetic data with complex questions. We evaluate our approach on four datasets from different domains for various IE tasks. Experimental results show the effectiveness of our approach in improving the performance of zero-shot joint IE tasks in multiple domains.
Ying Zhang 0084, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Summary-Oriented Vision Modeling for Multimodal Abstractive Summarization
abstract
Multimodal abstractive summarization (MAS) aims to produce a concise summary given the multimodal data (text and vision).Existing studies mainly focus on how to effectively use the visual features from the perspective of an article, having achieved impressive success on the high-resource English dataset.However, less attention has been paid to the visual features from the perspective of the summary, which may limit the model performance, especially in the low-and zero-resource scenarios.In this paper, we propose to improve the summary quality through summary-oriented visual features.To this end, we devise two auxiliary tasks including vision to summary task and masked image modeling task.Together with the main summarization task, we optimize the MAS model via the training objectives of all these tasks.By these means, the MAS model can be enhanced by capturing the summaryoriented visual features, thereby yielding more accurate summaries.Experiments on 44 languages, covering mid-high-, low-, and zeroresource scenarios, verify the effectiveness and superiority of the proposed approach, which achieves state-of-the-art performance under all scenarios.Additionally, we will contribute a large-scale multilingual multimodal abstractive summarization (MM-Sum) dataset. 1
Yunlong Liang, Fandong Meng, Jin An Xu, Jiaan Wang, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)5
2023 Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation
abstract
Songming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen, Wenjuan Han, Jian Liu, Jinan Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Songming Zhang 0001, Yunlong Liang, Shuaibo Wang, Yufeng Chen 0005, Wenjuan Han, Jian Liu 0032, Jin An Xu
ACL (1)4
2023 MT2: Towards a Multi-Task Machine Translation Model with Translation-Specific In-Context Learning
abstract
Sentence-level translation, document-level translation, translation memory, and terminology constrained translation play an important role in machine translation.Most of the previous work uses separate models or methods to solve these tasks, which is not conducive to knowledge transfer of different tasks and increases the complexity of system construction.In this work, we explore the potential of pre-trained language model in machine translation tasks and propose a Multi-Task Machine Translation (MT2) model to integrate these translation tasks.We design a novel translationspecific In-Context Learning (ICL) paradigm for model training, in which all of the translation tasks can be modeled as context-learning tasks that integrate contextual information for performance improvement.Specifically, we propose a retrieval and alignment method to obtain a large scale context-enhancement training data, then we train the model in an in-context learning manner.Furthermore, we adopt two context-dependent training strategies to encourage the model to better understand and utilize contextual information for translation.Extensive experiments on translation memory, terminology constrained translation, document-level translation, and few-shot domain-adaptation tasks demonstrate the superior performance of our model, verifying the effectiveness of our proposed approach.
Chunyou Li, Mingtong Liu, Hongxiao Zhang, Yufeng Chen 0005, Jin An Xu
EMNLP4
2023 Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFs
abstract
Real-world named entity recognition (NER) datasets are notorious for their noisy nature, attributed to annotation errors, inconsistencies, and subjective interpretations.Such noises present a substantial challenge for traditional supervised learning methods.In this paper, we present a new and unified approach to tackle annotation noises for NER.Our method considers NER as a constituency tree parsing problem, utilizing a tree-structured Conditional Random Fields (CRFs) with uncertainty evaluation for integration.Through extensive experiments conducted on four realworld datasets, we demonstrate the effectiveness of our model in addressing both partial and incorrect annotation errors.Remarkably, our model exhibits superb performance even in extreme scenarios with 90% annotation noise.
Jian Liu 0032, Weichang Liu, Yufeng Chen 0005, Jin An Xu, Zhe Zhao 0006
EMNLP3
2023 A Quality-based Syntactic Template Retriever for Syntactically-Controlled Paraphrase Generation
abstract
Existing syntactically-controlled paraphrase generation (SPG) models perform promisingly with human-annotated or well-chosen syntactic templates.However, the difficulty of obtaining such templates actually hinders the practical application of SPG models.For one thing, the prohibitive cost makes it unfeasible to manually design decent templates for every source sentence.For another, the templates automatically retrieved by current heuristic methods are usually unreliable for SPG models to generate qualified paraphrases.To escape this dilemma, we propose a novel Quality-based Syntactic Template Retriever (QSTR) to retrieve templates based on the quality of the to-be-generated paraphrases.Furthermore, for situations requiring multiple paraphrases for each source sentence, we design a Diverse Templates Search (DTS) algorithm, which can enhance the diversity between paraphrases without sacrificing quality.Experiments demonstrate that QSTR can significantly surpass existing retrieval methods in generating high-quality paraphrases and even perform comparably with human-annotated templates in terms of reference-free metrics.Additionally, human evaluation and the performance on downstream tasks using our generated paraphrases for data augmentation showcase the potential of our QSTR and DTS algorithm in practical scenarios.
Songming Zhang 0001, Yunlong Liang, Yufeng Chen 0005, Jian Liu 0032, Wenjuan Han, Jin An Xu
EMNLP4
2023 Exploring Domain-shared and Domain-specific Knowledge in Multi-Domain Neural Machine Translation
abstract
Currently, multi-domain neural machine translation (NMT) has become a significant research topic in domain adaptation machine translation, which trains a single model by mixing data from multiple domains. Multi-domain NMT aims to improve the performance of the low-resources domain through data augmentation. However, mixed domain data brings more translation ambiguity. Previous work focused on domain-general or domain-context knowledge learning, respectively. Therefore, there is a challenge for acquiring domain-general or domain-context knowledge simultaneously. To this end, we propose a unified framework for learning simultaneously domain-general and domain-specific knowledge, we are the first to apply parameter differentiation in multi-domain NMT. Specifically, we design the differentiation criterion and differentiation granularity to obtain domain-specific parameters. Experimental results on multi-domain UM-corpus English-to-Chinese and OPUS German-to-English datasets show that the average BLEU scores of the proposed method exceed the strong baseline by 1.22 and 1.87, respectively. In addition, we investigate the case study to illustrate the effectiveness of the proposed method in acquiring domain knowledge.
Zhibo Man, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu
MTSummit (1)4
2023 A Neighborhood Re-Ranking Model With Relation Constraint for Knowledge Graph Completion
abstract
Knowledge graph completion (KGC) aims to predict missing links based on observed triples. However, current KGC models are still limited by the following two aspects. (1) the entity semantics is implicitly learned by neural network and merely depends on existing facts, which mostly suffers from less additional specific knowledge. Although previous studies have noticed that entity type information can effectively improve KGC task, most of them rely on labeled type-specific data. (2) the recent graph-based models mainly concentrate on Graph Neural Network (GNN) to update source entity representation, regardless of the separate role that neighborhood information plays and may mix noisy neighbor features for target prediction. To address the above two issues, we propose a neighborhood re-ranking model with relation constraint for KGC task. We suggest that both relation constraint and structured information located in triples can boost the model performance. More importantly, we automatically generate explicit constraints as additional type feature to enrich entity representation instead of depending on human annotated labels. Meanwhile, we construct a neighborhood completion module to re-rank candidate entities for full use of the neighbor structure rather than traditional GNN updating manner. Extensive experiments on seven benchmarks demonstrate that our model achieves the competitive results in comparison to the recent advanced baselines.
Yu Li 0025, Bojie Hu, Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Scheduled Multi-task Learning for Neural Chat Translation
abstract
Neural Chat Translation (NCT) aims to translate conversational text into different languages.Existing methods mainly focus on modeling the bilingual dialogue characteristics (e.g., coherence) to improve chat translation via multi-task learning on small-scale chat translation data.Although the NCT models have achieved impressive success, it is still far from satisfactory due to insufficient chat translation data and simple joint training manners.To address the above issues, we propose a scheduled multi-task learning framework for NCT.Specifically, we devise a three-stage training framework to incorporate the large-scale in-domain chat translation data into training by adding a second pre-training stage between the original pre-training and fine-tuning stages.Further, we investigate where and how to schedule the dialogue-related auxiliary tasks in multiple training stages to effectively enhance the main chat translation task.Extensive experiments on four language directions (English↔Chinese and English↔German) verify the effectiveness and superiority of the proposed approach.Additionally, we will make the large-scale indomain paired bilingual dialogue dataset publicly available for the research community.1
Yunlong Liang, Fandong Meng, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)4
2022 MSCTD: A Multimodal Sentiment Chat Translation Dataset
abstract
Multimodal machine translation and textual chat translation have received considerable attention in recent years.Although the conversation in its natural form is usually multimodal, there still lacks work on multimodal machine translation in conversations.In this work, we introduce a new task named Multimodal Chat Translation (MCT), aiming to generate more accurate translations with the help of the associated dialogue history and visual context.To this end, we firstly construct a Multimodal Sentiment Chat Translation Dataset (MSCTD) containing 142,871 English-Chinese utterance pairs in 14,762 bilingual dialogues and 30,370 English-German utterance pairs in 3,079 bilingual dialogues.Each utterance pair, corresponding to the visual context that reflects the current conversational scene, is annotated with a sentiment label.Then, we benchmark the task by establishing multiple baseline systems that incorporate multimodal and sentiment features for MCT.Preliminary experiments on four language directions (English↔Chinese and English↔German) verify the potential of contextual and multimodal information fusion and the positive impact of sentiment on the MCT task.Additionally, as a by-product of the MSCTD, it also provides two new benchmarks on multimodal dialogue sentiment analysis.Our work can facilitate research on both multimodal chat translation and multimodal dialogue sentiment analysis.1
Yunlong Liang, Fandong Meng, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)4
2022 A Variational Hierarchical Model for Neural Cross-Lingual Summarization
abstract
Yunlong Liang, Fandong Meng, Chulun Zhou, Jinan Xu, Yufeng Chen, Jinsong Su, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yunlong Liang, Fandong Meng, Chulun Zhou, Jin An Xu, Yufeng Chen 0005, Jinsong Su, Jie Zhou 0016
ACL (1)5
2022 Saliency as Evidence: Event Detection with Trigger Saliency Attribution
abstract
Event detection (ED) is a critical subtask of event extraction that seeks to identify event triggers of certain types in texts.Despite significant advances in ED, existing methods typically follow a "one model fits all types" approach, which sees no differences between event types and often results in a quite skewed performance.Finding the causes of skewed performance is crucial for the robustness of an ED model, but to date there has been little exploration of this problem.This research examines the issue in depth and presents a new concept termed trigger salience attribution, which can explicitly quantify the underlying patterns of events.On this foundation, we develop a new training mechanism for ED, which can distinguish between triggerdependent and context-dependent types and achieve promising performance on two benchmarks.Finally, by highlighting many distinct characteristics of trigger-dependent and context-dependent types, our work may promote more research into this problem.
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
ACL (1)2
2022 Conditional Bilingual Mutual Information Based Adaptive Training for Neural Machine Translation
abstract
Songming Zhang, Yijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu, Jian Liu, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Songming Zhang 0001, Yijin Liu, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jian Liu 0032, Jie Zhou 0016
ACL (1)4
2022 Iterative Constrained Back-Translation for Unsupervised Domain Adaptation of Machine Translation
abstract
Back-translation has been proven to be effective in unsupervised domain adaptation of neural machine translation (NMT). However, the existing back-translation methods mainly improve domain adaptability by generating in-domain pseudo-parallel data that contains sentence-structural knowledge, paying less attention to the in-domain lexical knowledge, which may lead to poor translation of unseen in-domain words. In this paper, we propose an Iterative Constrained Back-Translation (ICBT) method to incorporate in-domain lexical knowledge on the basis of BT for unsupervised domain adaptation of NMT. Specifically, we apply lexical constraints into back-translation to generate pseudo-parallel data with in-domain lexical knowledge, and then perform round-trip iterations to incorporate more lexical knowledge. Based on this, we further explore sampling strategies of constrained words in ICBT to introduce more targeted lexical knowledge, via domain specificity and confidence estimation. Experimental results on four domains show that our approach achieves state-of-the-art results, improving the BLEU score by up to 3.08 compared to the strongest baseline, which demonstrates the effectiveness of our approach.
Hongxiao Zhang, Yufeng Chen 0005, Jin An Xu, Jian Liu 0032
COLING4
2022 Cross-Align: Modeling Deep Cross-lingual Interactions for Word Alignment
abstract
Word alignment which aims to extract lexicon translation equivalents between source and target sentences, serves as a fundamental tool for natural language processing.Recent studies in this area have yielded substantial improvements by generating alignments from contextualized embeddings of the pre-trained multilingual language models.However, we find that the existing approaches capture few interactions between the input sentence pairs, which degrades the word alignment quality severely, especially for the ambiguous words in the monolingual context.To remedy this problem, we propose Cross-Align to model deep interactions between the input sentence pairs, in which the source and target sentences are encoded separately with the shared self-attention modules in the shallow layers, while cross-lingual interactions are explicitly constructed by the crossattention modules in the upper layers.Besides, to train our model effectively, we propose a two-stage training framework, where the model is trained with a simple Translation Language Modeling (TLM) objective in the first stage and then finetuned with a self-supervised alignment objective in the second stage.Experiments show that the proposed Cross-Align achieves the state-of-the-art (SOTA) performance on four out of five language pairs. 1
Siyu Lai, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
EMNLP4
2022 Long Text Generation with Topic-aware Discrete Latent Variable Model
abstract
Generating coherent long texts is an important yet challenging task, particularly for the openended generation task.Prior work based on discrete latent codes focuses on the modeling of discourse relation, resulting in discrete codes only learning shallow semantics (Ji and Huang, 2021).A natural text always revolves around several related topics and the transition across them is natural and smooth.In this work, we investigate whether discrete latent codes can learn information of topics.To this end, we build a topic-aware latent code-guided text generation model.To encourage discrete codes to model information about topics, we propose a span-level bag-of-words training objective for the model.Automatic and manual evaluation experiments show that our method can generate more topic-relevant and coherent texts.
Erguang Yang, Mingtong Liu, Deyi Xiong, Yufeng Chen 0005, Jin An Xu
EMNLP5
2022 Low-Resource NER by Data Augmentation With Prompting
abstract
Named entity recognition (NER) is a fundamental information extraction task that seeks to identify entity mentions of certain types in text. Despite numerous advances, the existing NER methods rely on extensive supervision for model training, which struggle in a low-resource scenario with limited training data. In this paper, we propose a new data augmentation method for low-resource NER, by eliciting knowledge from BERT with prompting strategies. Particularly, we devise a label-conditioned word replacement strategy that can produce more label-consistent examples by capturing the underlying word-label dependencies, and a prompting with question answering method to generate new training data from unlabeled texts. The experimental results have widely confirmed the effectiveness of our approach. Particularly, in a low-resource scenario with only 150 training sentences, our approach outperforms previous methods without data augmentation by over 40% in F1 and prior best data augmentation methods by over 2.0% in F1. Furthermore, our approach also fits with a zero-shot scenario, yielding promising results without using any human-labeled data for the task.
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
IJCAI2
2022 Multimedia Event Extraction From News With a Unified Contrastive Learning Framework
abstract
Extracting events from news have seen many benefits in downstream applications. Today's event extraction (EE) systems, however, usually focus on a single modality --- either for text or image, and such methods suffer from incomplete information because a news document is typically presented in a multimedia format. In this paper, we propose a new method for multimedia EE by bridging the textual and visual modalities with a unified contrastive learning framework. Our central idea is to create a shared space for texts and images in order to improve their similar representation. This is accomplished by training on text-image pairs in general, and we demonstrate that it is possible to use this framework to boost learning for one modality by investigating the complementary of the other modality. On the benchmark dataset, our approach establishes a new state-of-the-art performance and shows a 3 percent improvement in F1. Furthermore, we demonstrate that it can achieve cutting-edge performance for visual EE even in a zero-shot scenario with no annotated data in the visual modality.
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
ACM Multimedia2
2022 Generating Authentic Adversarial Examples beyond Meaning-preserving with Doubly Round-trip Translation
abstract
Siyu Lai, Zhen Yang, Fandong Meng, Xue Zhang, Yufeng Chen, Jinan Xu, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Siyu Lai, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
NAACL-HLT5
2022 Emotional conversation generation with heterogeneous graph neural network
Yunlong Liang, Fandong Meng, Ying Zhang 0084, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
Artif. Intell.4
2022 Document-level event argument linking as machine reading comprehension
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
Neurocomputing2
2022 Improving generation diversity via syntax-controlled paraphrasing
Erguang Yang, Mingtong Liu, Deyi Xiong, Jin An Xu, Yufeng Chen 0005
Neurocomputing7
2022 Exploiting Morpheme and Cross-lingual Knowledge to Enhance Mongolian Named Entity Recognition
abstract
Mongolian named entity recognition (NER) is not only one of the most crucial and fundamental tasks in Mongolian natural language processing, but also an important step to improve the performance of downstream tasks such as information retrieval, machine translation, and dialog system. However, traditional Mongolian NER models heavily rely on the feature engineering. Even worse, the complex morphological structure of Mongolian words makes the data sparser. To alleviate the feature engineering and data sparsity in Mongolian named entity recognition, we propose a novel NER framework with Multi-Knowledge Enhancement (MKE-NER) . Specifically, we introduce both linguistic knowledge through Mongolian morpheme representation and cross-lingual knowledge from Mongolian-Chinese parallel corpus. Furthermore, we design two methods to exploit cross-lingual knowledge sufficiently, i.e., cross-lingual representation and cross-lingual annotation projection. Experimental results demonstrate the effectiveness of our MKE-NER model, which outperforms strong baselines and achieves the best performance (94.04% F1 score) on the traditional Mongolian benchmark. Particularly, extensive experiments with different data scales highlight the superiority of our method in low-resource scenarios.
Songming Zhang 0001, Ying Zhang 0084, Yufeng Chen 0005, Du Wu, Jin An Xu, Jian Liu 0032
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2022 MRCAug: Data Augmentation via Machine Reading Comprehension for Document-Level Event Argument Extraction
abstract
Document-level event argument extraction (EAE) is a critical event semantic understanding task that requires a model to identify an event's global arguments beyond the sentence level. Existing approaches to this problem are based on supervised learning, which require a large amount of labeled data for model training. However, due to the complicated structure of an event, human annotation for this task is costly, and the issue of inadequacy of training data has long hampered the study. In this study, we propose a novel approach to mitigating the data sparsity problem faced by document-level EAE, by linking the task with machine reading comprehension (MRC). Particularly, we devise two data augmentation regimes via MRC, including an implicit knowledge transfer method, which enables knowledge transfer from other tasks to the document-level EAE task, and an explicit data generation method, which can explicitly generate new training examples by treating a pre-trained MRC model as an annotator. Furthermore, we propose a self-training based noise reduction strategy that can effectively addresses the out-of-domain noise introduced by the data augmentation methods. The extensive assessments on three benchmarks have validated the effectiveness of our approach — it not only achieves state-of-the-art performance but also demonstrates superior results in the data-low scenario.
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Cross-Domain Slot Filling as Machine Reading Comprehension: A New Perspective
abstract
With intelligent dialogue systems becoming more and more important in our daily lives, slot filling, one of the most important components of an intelligent dialogue system, has gotten a lot of attention from academia and industry. Despite many advancements in the single-domain learning paradigm for slot filling, leveraging resources from different domains to boost learning for a target domain remains a challenge. In contrast to prior methods that supplemented a sequence labeling model with slot meta-information, we address cross-domain slot filling as a machine reading comprehension (MRC) problem for the first time, where the extraction of slot values is viewed as a question answering process. In the framework above, we present both static and dynamic question generating mechanisms, which have complimentary effects in diverse cross-domain contexts. Furthermore, we devise a dynamic question generation approach that can generate numerous values for a slot at the same time. Finally, we construct a pre-training and fine-tuning training approach that enables us to improve learning by utilizing MRC’s resources. We conducted extensive experiments on four datasets to evaluate our approach, and the experimental results clearly justified the advantages of our approach in various cross-domain settings.
Jian Liu 0032, Mengshi Yu, Yufeng Chen 0005, Jin An Xu
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Infusing Multi-Source Knowledge with Heterogeneous Graph Neural Network for Emotional Conversation Generation
abstract
The success of emotional conversation systems depends on sufficient perception and appropriate expression of emotions. In a real-world conversation, we firstly instinctively perceive emotions from multi-source information, including the emotion flow of dialogue history, facial expressions, and personalities of speakers, and then express suitable emotions according to our personalities, but these multiple types of information are insufficiently exploited in emotional conversation fields. To address this issue, we propose a heterogeneous graph-based model for emotional conversation generation. Specifically, we design a Heterogeneous Graph-Based Encoder to represent the conversation content (i.e., the dialogue history, its emotion flow, facial expressions, and speakers' personalities) with a heterogeneous graph neural network, and then predict suitable emotions for feedback. After that, we employ an Emotion-Personality-Aware Decoder to generate a response not only relevant to the conversation context but also with appropriate emotions, by taking the encoded graph representations, the predicted emotions from the encoder and the personality of the current speaker as inputs. Experimental results show that our model can effectively perceive emotions from multi-source knowledge and generate a satisfactory response, which significantly outperforms previous state-of-the-art models.
Yunlong Liang, Fandong Meng, Ying Zhang 0084, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
AAAI4
2021 Faster Depth-Adaptive Transformers
abstract
Depth-adaptive neural networks can dynamically adjust depths according to the hardness of input words, and thus improve efficiency. The main challenge is how to measure such hardness and decide the required depths (i.e., layers) to conduct. Previous works generally build a halting unit to decide whether the computation should continue or stop at each layer. As there is no specific supervision of depth selection, the halting unit may be under-optimized and inaccurate, which results in suboptimal and unstable performance when modeling sentences. In this paper, we get rid of the halting unit and estimate the required depths in advance, which yields a faster depth-adaptive model. Specifically, two approaches are proposed to explicitly measure the hardness of input words and estimate corresponding adaptive depth, namely 1) mutual information (MI) based estimation and 2) reconstruction loss based estimation. We conduct experiments on the text classification task with 24 datasets in various sizes and domains. Results confirm that our approaches can speed up the vanilla Transformer (up to 7x) while preserving high accuracy. Moreover, efficiency and robustness are significantly improved when compared with other depth-adaptive approaches.
Yijin Liu, Fandong Meng, Jie Zhou 0016, Yufeng Chen 0005, Jin An Xu
AAAI4
2021 Modeling Bilingual Conversational Characteristics for Neural Chat Translation
abstract
Yunlong Liang, Fandong Meng, Yufeng Chen, Jinan Xu, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yunlong Liang, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
ACL/IJCNLP (1)3
2021 Towards Making the Most of Dialogue Characteristics for Neural Chat Translation
abstract
Neural Chat Translation (NCT) aims to translate conversational text between speakers of different languages.Despite the promising performance of sentence-level and context-aware neural machine translation models, there still remain limitations in current NCT models because the inherent dialogue characteristics of chat, such as dialogue coherence and speaker personality, are neglected.In this paper, we propose to promote the chat translation by introducing the modeling of dialogue characteristics into the NCT model.To this end, we design four auxiliary tasks including monolingual response generation, cross-lingual response generation, next utterance discrimination, and speaker identification.Together with the main chat translation task, we optimize the NCT model through the training objectives of all these tasks.By this means, the NCT model can be enhanced by capturing the inherent dialogue characteristics, thus generating more coherent and speaker-relevant translations.Comprehensive experiments on four language directions (English⇔German and English⇔Chinese) verify the effectiveness and superiority of the proposed approach.
Yunlong Liang, Chulun Zhou, Fandong Meng, Jin An Xu, Yufeng Chen 0005, Jinsong Su, Jie Zhou 0016
EMNLP (1)5
2021 Machine Reading Comprehension as Data Augmentation: A Case Study on Implicit Event Argument Extraction
abstract
Implicit event argument extraction (EAE) is a crucial document-level information extraction task that aims to identify event arguments beyond the sentence level.Despite many efforts for this task, the lack of enough training data has long impeded the study.In this paper, we take a new perspective to address the data sparsity issue faced by implicit EAE, by bridging the task with machine reading comprehension (MRC).Particularly, we devise two data augmentation regimes via MRC, including: 1) implicit knowledge transfer, which enables knowledge transfer from other tasks, by building a unified training framework in the MRC formulation, and 2) explicit data augmentation, which can explicitly generate new training examples, by treating MRC models as an annotator.The extensive experiments have justified the effectiveness of our approach -it not only obtains state-of-the-art performance on two benchmarks, but also demonstrates superior results in a data-low scenario.
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
EMNLP (1)2
2021 Scheduled Sampling Based on Decoding Steps for Neural Machine Translation
abstract
Scheduled sampling is widely used to mitigate the exposure bias problem for neural machine translation.Its core motivation is to simulate the inference scene during training by replacing ground-truth tokens with predicted tokens, thus bridging the gap between training and inference.However, vanilla scheduled sampling is merely based on training steps and equally treats all decoding steps.Namely, it simulates an inference scene with uniform error rates, which disobeys the real inference scene, where larger decoding steps usually have higher error rates due to error accumulations.To alleviate the above discrepancy, we propose scheduled sampling methods based on decoding steps, increasing the selection chance of predicted tokens with the growth of decoding steps.Consequently, we can more realistically simulate the inference scene during training, thus better bridging the gap between training and inference.Moreover, we investigate scheduled sampling based on both training steps and decoding steps for further improvements.Experimentally, our approaches significantly outperform the Transformer baseline and vanilla scheduled sampling on three large-scale WMT tasks.Additionally, our approaches also generalize well to the text summarization task on two popular benchmarks.
Yijin Liu, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
EMNLP (1)3
2021 Syntactically-Informed Unsupervised Paraphrasing with Non-Parallel Data
abstract
Previous works on syntactically controlled paraphrase generation heavily rely on largescale parallel paraphrase data that are not easily available for many languages and domains.In this paper, we take this research direction to the extreme and investigate whether it is possible to learn syntactically controlled paraphrase generation with non-parallel data.We propose a syntactically-informed unsupervised paraphrasing model based on conditional variational auto-encoder (VAE) which can generate texts in a specified syntactic structure.Particularly, we design a two-stage learning method to effectively train the model using non-parallel data.The conditional VAE is trained to reconstruct the input sentence according to the given input and its syntactic structure.Furthermore, to improve the syntactic controllability and semantic consistency of the pre-trained conditional VAE, we finetune it using syntax controlling and cycle reconstruction learning objectives, and employ Gumbel-Softmax to combine these new learning objectives.Experiment results demonstrate that the proposed model trained only on non-parallel data is capable of generating diverse paraphrases with specified structures.Additionally, we further validate the effectiveness of our method for generating syntactically adversarial examples on a sentiment analysis task.
Erguang Yang, Mingtong Liu, Deyi Xiong, Changjian Hu, Jin An Xu, Yufeng Chen 0005
EMNLP (1)8
2021 Discourse-Level Event Temporal Ordering with Uncertainty-Guided Graph Completion
abstract
Learning to order events at discourse-level is a crucial text understanding task. Despite many efforts for this task, the current state-of-the-art methods rely heavily on manually designed features, which are costly to produce and are often specific to tasks/domains/datasets. In this paper, we propose a new graph perspective on the task, which does not require complex feature engineering but can assimilate global features and learn inter-dependencies effectively. Specifically, in our approach, each document is considered as a temporal graph, in which the nodes and edges represent events and event-event relations respectively. In this sense, the temporal ordering task corresponds to constructing edges for an empty graph. To train our model, we design a graph mask pre-training mechanism, which can learn inter-dependencies of temporal relations by learning to recover a masked edge following graph topology. In the testing stage, we design an certain-first strategy based on model uncertainty, which can decide the prediction orders and reduce the risk of error propagation. The experimental results demonstrate that our approach outperforms previous methods consistently and can meanwhile maintain good global consistency.
Jian Liu 0032, Jin An Xu, Yufeng Chen 0005
IJCAI3
2021 Improving Stylized Neural Machine Translation with Iterative Dual Knowledge Transfer
abstract
Stylized neural machine translation (NMT) aims to translate sentences of one style into sentences of another style, which is essential for the application of machine translation in a real-world scenario. However, a major challenge in this task is the scarcity of high-quality parallel data which is stylized paired. To address this problem, we propose an iterative dual knowledge transfer framework that utilizes informal training data of machine translation and formality style transfer data to create large-scale stylized paired data, for the training of stylized machine translation model. Specifically, we perform bidirectional knowledge transfer between translation model and text style transfer model iteratively through knowledge distillation. Then, we further propose a data-refinement module to process the noisy synthetic parallel data generated during knowledge transfer. Experiment results demonstrate the effectiveness of our method, achieving an improvement over the existing best model by 5 BLEU points on MTFC dataset. Meanwhile, extensive analyses illustrate our method can also improve the accuracy of formality style transfer.
Xuanxuan Wu, Jian Liu 0032, Xinjie Li 0003, Jin An Xu, Yufeng Chen 0005
IJCAI5
2021 Cross-Domain Slot Filling as Machine Reading Comprehension
abstract
With task-oriented dialogue systems being widely applied in everyday life, slot filling, the essential component of task-oriented dialogue systems, is required to be quickly adapted to new domains that contain domain-specific slots with few or no training data. Previous methods for slot filling usually adopt sequence labeling framework, which, however, often has limited ability when dealing with the domain-specific slots. In this paper, we take a new perspective on cross-domain slot filling by framing it as a machine reading comprehension (MRC) problem. Our approach firstly transforms slot names into well-designed queries, which contain rich informative prior knowledge and are very helpful for the detection of domain-specific slots. In addition, we utilize the large-scale MRC dataset for pre-training, which further alleviates the data scarcity problem. Experimental results on SNIPS and ATIS datasets show that our approach consistently outperforms the existing state-of-the-art methods by a large margin.
Mengshi Yu, Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
IJCAI3
2021 Contrastive Learning for Machine Translation Quality Estimation
Hui Di, Jian Liu 0032, Yufeng Chen 0005, Kazushige Ouchi, Jin An Xu
NLPCC (1)4
2021 Explore Coarse-Grained Structures for Syntactically Controllable Paraphrase Generation
Erguang Yang, Mingtong Liu, Deyi Xiong, Changjian Hu, Jin An Xu, Yufeng Chen 0005
NLPCC (1)8
2021 Deep bi-directional interaction network for sentence matching
Mingtong Liu, Jin An Xu, Yufeng Chen 0005
Appl. Intell.4
2021 A dependency syntactic knowledge augmented interactive architecture for end-to-end aspect-based sentiment analysis
Yunlong Liang, Fandong Meng, Jinchao Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
Neurocomputing4
2020 A Learning-Exploring Method to Generate Diverse Paraphrases with Multi-Objective Deep Reinforcement Learning
abstract
Paraphrase generation (PG) is of great importance to many downstream tasks in natural language processing.Diversity is an essential nature to PG for enhancing generalization capability and robustness of downstream applications.Recently, neural sequence-to-sequence (Seq2Seq) models have shown promising results in PG.However, traditional model training for PG focuses on optimizing model prediction against single reference and employs cross-entropy loss, which objective is unable to encourage model to generate diverse paraphrases.In this work, we present a novel approach with multi-objective learning to PG.We propose a learning-exploring method to generate sentences as learning objectives from the learned data distribution, and employ reinforcement learning to combine these new learning objectives for model training.We first design a sample-based algorithm to explore diverse sentences.Then we introduce several reward functions to evaluate the sampled sentences as learning signals in terms of expressive diversity and semantic fidelity, aiming to generate diverse and high-quality paraphrases.To effectively optimize model performance satisfying different evaluating aspects, we use a GradNorm-based algorithm that automatically balances these training objectives.Experiments and analyses on Quora and Twitter datasets demonstrate that our proposed method not only gains a significant increase in diversity but also improves generation quality over several state-of-the-art baselines.
Mingtong Liu, Erguang Yang, Deyi Xiong, Changjian Hu, Jin An Xu, Yufeng Chen 0005
COLING8
2020 Exploring Bilingual Parallel Corpora for Syntactically Controllable Paraphrase Generation
abstract
Paraphrase generation is of great importance to many downstream tasks in natural language processing. Recent efforts have focused on generating paraphrases in specific syntactic forms, which, generally, heavily relies on manually annotated paraphrase data that is not easily available for many languages and domains. In this paper, we propose a novel end-to-end framework to leverage existing large-scale bilingual parallel corpora to generate paraphrases under the control of syntactic exemplars. In order to train one model over the two languages of parallel corpora, we embed sentences of them into the same content and style spaces with shared content and style encoders using cross-lingual word embeddings. We propose an adversarial discriminator to disentangle the content and style space, and employ a latent variable to model the syntactic style of a given exemplar in order to guide the two decoders for generation. Additionally, we introduce cycle and masking learning schemes to efficiently train the model. Experiments and analyses demonstrate that the proposed model trained only on bilingual parallel data is capable of generating diverse paraphrases with desirable syntactic styles. Fine-tuning the trained model on a small paraphrase corpus makes it substantially outperform state-of-the-art paraphrase generation models trained on a larger paraphrase dataset.
Mingtong Liu, Erguang Yang, Deyi Xiong, Chen Sheng, Changjian Hu, Jin An Xu, Yufeng Chen 0005
IJCAI8
2020 Ensemble Distilling Pretrained Language Models for Machine Translation Quality Estimation
Hui Di, Jin An Xu, Kazushige Ouchi, Yufeng Chen 0005
NLPCC (2)5
2019 GCDT: A Global Context Enhanced Deep Transition Architecture for Sequence Labeling
abstract
Current state-of-the-art systems for the sequence labeling tasks are typically based on the family of Recurrent Neural Networks (RNNs).However, the shallow connections between consecutive hidden states of RNNs and insufficient modeling of global information restrict the potential performance of those models.In this paper, we try to address these issues, and thus propose a Global Context enhanced Deep Transition architecture for sequence labeling named GCDT.We deepen the state transition path at each position in a sentence, and further assign every token with a global representation learned from the entire sentence.Experiments on two standard sequence labeling tasks show that, given only training data and the ubiquitous word embeddings (Glove), our GCDT achieves 91.96 F 1 on the CoNLL03 NER task and 95.43 F 1 on the CoNLL2000 Chunking task, which outperforms the best reported results under the same settings.Furthermore, by leveraging BERT as an additional resource, we establish new stateof-the-art results with 93.47 F 1 on NER and 97.30F 1 on Chunking 1 .
Yijin Liu, Fandong Meng, Jinchao Zhang 0001, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)5
2019 A Novel Aspect-Guided Deep Transition Model for Aspect Based Sentiment Analysis
abstract
Yunlong Liang, Fandong Meng, Jinchao Zhang, Jinan Xu, Yufeng Chen, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yunlong Liang, Fandong Meng, Jinchao Zhang 0001, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
EMNLP/IJCNLP (1)5
2019 CM-Net: A Novel Collaborative Memory Network for Spoken Language Understanding
abstract
Yijin Liu, Fandong Meng, Jinchao Zhang, Jie Zhou, Yufeng Chen, Jinan Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yijin Liu, Fandong Meng, Jinchao Zhang 0001, Jie Zhou 0016, Yufeng Chen 0005, Jin An Xu
EMNLP/IJCNLP (1)5
2019 Original Semantics-Oriented Attention and Deep Fusion Network for Sentence Matching
abstract
Mingtong Liu, Yujie Zhang, Jinan Xu, Yufeng Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mingtong Liu, Jin An Xu, Yufeng Chen 0005
EMNLP/IJCNLP (1)4
2019 A Sequence-to-Action Architecture for Character-Based Chinese Dependency Parsing with Status History
Hang Liu 0005, Meng Chen 0006, Jin An Xu, Yufeng Chen 0005
NLPCC (2)5
2019 Improved Quality Estimation of Machine Translation with Pre-trained Language Representation
Guoyi Miao, Hui Di, Jin An Xu, Zhongcheng Yang, Yufeng Chen 0005, Kazushige Ouchi
NLPCC (1)5
2018 Improved Character-Based Chinese Dependency Parsing by Using Stack-Tree LSTM
Hang Liu 0005, Mingtong Liu, Jin An Xu, Yufeng Chen 0005
NLPCC (2)5
2017 A Semantic Concept Based Unknown Words Processing Method in Neural Machine Translation
Shaotong Li, Jin An Xu, Guoyi Miao, Yufeng Chen 0005
NLPCC5