VLDB 2026 Research / reviewers in the wild / expert
Jin An Xu
dblp:67/3124 · also Jinan Xu
· DBLP profile ↗
89ranked-venue papers
1as first author
71since 2021 · last 2026
0000-0003-0170-626XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 83 · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STOLA: Self-Adaptive Touch-Language Framework for Tactile Commonsense Reasoning in Open-Ended ScenariosabstractThis paper explores the challenges of integrating tactile sensing into intelligent systems for multimodal reasoning, particularly in enabling commonsense reasoning about the open-ended physical world. We identify two key challenges: modality discrepancy, where existing touch-language models often treat touch as a mere sub-modality of language without further addressing the semantic differences, and open-ended tactile data scarcity, where current datasets lack the diversity, open-endedness, and complexity needed for reasoning. To overcome these challenges, we introduce SToLa, a Self-Adaptive Touch-Language framework. SToLa utilizes Mixture of Experts (MoE) to dynamically process, unify, and manage tactile and language modalities, capturing their unique characteristics. Crucially, we also present a comprehensive tactile commonsense reasoning dataset and benchmark featuring free-form questions and responses, 8 physical properties, 4 interactive characteristics, and diverse commonsense knowledge. Experiments show SToLa exhibits competitive performance compared to existing models on the PHYSICLEAR benchmark and self-constructed datasets, proving the effectiveness of the Mixture of Experts architecture in multimodal management and the performance advantages for open-scenario tactile commonsense reasoning tasks. Jin An Xu, Jialing Chen, Bin Fang 0003, Wenjuan Han |
AAAI | 2 |
| 2026 | Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement LearningabstractXue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang, Kaiyu Huang, Yufeng Chen, Xu Jinan, Jie Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yunlong Liang, Fandong Meng, Songming Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
ACL (1) | 7 |
| 2026 | TRTF: A Two-Stage Robust Training Framework for Visual Question Answering
Yu Li 0025, Jin An Xu |
ICPR (10) | 2 |
| 2026 | DKF: Domain knowledge fusion in progressive incremental learning for multi-domain machine translation
Zhibo Man, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu |
Expert Syst. Appl. | 5 |
| 2026 | GLINT: Global-local fusion with attention intervention for MLLM hallucination mitigation caused by insufficient visual resolution
Jin An Xu, Songming Zhang 0001, Chengkai Wang, Wenjuan Han |
Inf. Process. Manag. | 2 |
| 2025 | AlignDistil: Token-Level Language Model Alignment as Adaptive Policy DistillationabstractIn modern large language models (LLMs), LLM alignment is of crucial importance and is typically achieved through methods such as reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO).However, in most existing methods for LLM alignment, all tokens in the response are optimized using a sparse, response-level reward or preference annotation.The ignorance of token-level rewards may erroneously punish high-quality tokens or encourage lowquality tokens, resulting in suboptimal performance and slow convergence speed.To address this issue, we propose AlignDistil, an RLHFequivalent distillation method for token-level reward optimization.Specifically, we introduce the reward learned by DPO into the RLHF objective and theoretically prove the equivalence between this objective and a token-level distillation process, where the teacher distribution linearly combines the logits from the DPO model and a reference model.On this basis, we further bridge the accuracy gap between the reward from the DPO model and the pure reward model, by building a contrastive DPO reward with a normal and a reverse DPO model.Moreover, to avoid under-and over-optimization on different tokens, we design a token adaptive logit extrapolation mechanism to construct an appropriate teacher distribution for each token.Experimental results demonstrate the superiority of our AlignDistil over existing methods and showcase fast convergence due to its tokenlevel distributional reward optimization. Songming Zhang 0001, Bojie Hu, Yufeng Chen 0005, Jin An Xu |
ACL (1) | 6 |
| 2025 | Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-ExpertsabstractContinually expanding new languages for existing large language models (LLMs) is a promising yet challenging approach to building powerful multilingual LLMs.The biggest challenge is to make the model continuously learn new languages while preserving the proficient ability of old languages.To achieve this, recent work utilizes the Mixture-of-Experts (MoE) architecture to expand new languages by adding new experts and avoid catastrophic forgetting of old languages by routing corresponding tokens to the original model backbone (old experts).Although intuitive, this kind of method is parameter-costly when expanding new languages and still inevitably impacts the performance of old languages.To address these limitations, we analyze the language characteristics of different layers in LLMs and propose a layer-wise expert allocation algorithm (LayerMoE) to determine the appropriate number of new experts for each layer.Specifically, we find different layers in LLMs exhibit different representation similarities between languages and then utilize the similarity as the indicator to allocate experts for each layer, i.e., the higher similarity, the fewer experts.Additionally, to further mitigate the forgetting of old languages, we add a classifier in front of the router network on the layers with higher similarity to guide the routing of old language tokens.Experimental results show that our method outperforms the previous state-of-the-art baseline with 60% fewer experts in the single-expansion setting and with 33.3% fewer experts in the lifelong-expansion setting, demonstrating the effectiveness of our method. Yunlong Liang, Fandong Meng, Songming Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
ACL (1) | 6 |
| 2025 | Multilingual Knowledge Editing with Language-Agnostic Factual NeuronsabstractMultilingual knowledge editing (MKE) aims to simultaneously update factual knowledge across multiple languages within large language models (LLMs). Previous research indicates that the same knowledge across different languages within LLMs exhibits a degree of shareability. However, most existing MKE methods overlook the connections of the same knowledge between different languages, resulting in knowledge conflicts and limited edit performance. To address this issue, we first investigate how LLMs process multilingual factual knowledge and discover that the same factual knowledge in different languages generally activates a shared set of neurons, which we call language-agnostic factual neurons (LAFNs). These neurons represent the same factual knowledge shared across languages and imply the semantic connections among multilingual knowledge. Inspired by this finding, we propose a new MKE method by Locating and Updating Language-Agnostic Factual Neurons (LU-LAFNs) to edit multilingual knowledge simultaneously, which avoids knowledge conflicts and thus improves edit performance. Experimental results on Bi-ZsRE and MzsRE benchmarks demonstrate that our method achieves the best edit performance, indicating the effectiveness and importance of modeling the semantic connections among multilingual knowledge. Yunlong Liang, Fandong Meng, Songming Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
COLING | 6 |
| 2025 | Boosting Data Utilization for Multilingual Dense RetrievalabstractMultilingual dense retrieval aims to retrieve relevant documents across different languages based on a unified retriever model.The challenge lies in aligning representations of different languages in a shared vector space.The common practice is to fine-tune the dense retriever via contrastive learning, whose effectiveness highly relies on the quality of the negative samples and the efficacy of mini-batch data.Different from the existing studies that focus on developing sophisticated model architecture, we propose a method to boost data utilization for multilingual dense retrieval by obtaining high-quality hard negative samples and effective mini-batch data.The extensive experimental results on a multilingual retrieval benchmark, MIRACL, with 16 languages demonstrate the effectiveness of our method by outperforming several existing strong baselines. Fengran Mo, Yufeng Chen 0005, Changhao Guan, Zhenrui Yue, Xinyu Wang 0061, Jin An Xu |
EMNLP | 7 |
| 2025 | Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning ModelabstractThe rapid development of Multimodal Large Reasoning Models (MLRMs) has demonstrated broad application potential, yet their safety and reliability remain critical concerns that require systematic exploration.To address this gap, we conduct a comprehensive and systematic safety evaluation of 13 MLRMs across 5 benchmarks and unveil prevalent safety degradation phenomena in most advanced models.Moreover, our analysis reveals distinct safety patterns across different benchmarks: significant safety degradation is observed across jailbreak robustness benchmarks, whereas safetyawareness benchmarks demonstrate less pronounced degradation.In particular, the long thought process in some scenarios even enhances safety performance.Therefore, it is a potential approach to address safety issues in MLRMs by leveraging the intrinsic reasoning capabilities of the model to detect unsafe intent.To operationalize this insight, we construct a multimodal tuning dataset that incorporates a safety-oriented thought process.Experimental results from fine-tuning existing MLRMs with this dataset effectively enhance the safety on both jailbreak robustness and safety-awareness benchmarks.This study provides a new perspective for developing safe MLRMs. 1 Warning: this paper contains example data that may be offensive or harmful. Xinyue Lou, You Li 0010, Jin An Xu, Chi Chen 0005 |
EMNLP | 3 |
| 2025 | DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain TranslationabstractCurrently, Large Language Models (LLMs) have achieved remarkable results in machine translation.However, their performance in multi-domain translation (MDT) is less satisfactory, the meanings of words can vary across different domains, highlighting the significant ambiguity inherent in MDT.Therefore, evaluating the disambiguation ability of LLMs in MDT remains an open problem.To this end, we present an evaluation and analysis of LLMs on disambiguation in multi-domain translation (DMDTEval), our systematic evaluation framework consisting of three aspects: (1) we construct a translation test set with multi-domain ambiguous word annotation, (2) we curate a diverse set of disambiguation prompt strategies, and (3) we design precise disambiguation metrics, and study the efficacy of various prompt strategies on multiple state-of-the-art LLMs.We conduct comprehensive experiments across 4 language pairs and 13 domains, our extensive experiments reveal a number of crucial findings that we believe will pave the way and also facilitate further research in the critical area of improving the disambiguation of LLMs. Zhibo Man, Yuanmeng Chen, Jin An Xu |
EMNLP | 4 |
| 2025 | WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language ModelsabstractAlthough Large Language Models (LLMs) excel in NLP tasks, they still need external tools to extend their ability. Current research on tool learning with LLMs often assumes mandatory tool use, which does not always align with real-world situations, where the necessity for tools is uncertain, and incorrect or unnecessary use of tools can damage the general abilities of LLMs. Therefore, we propose to explore whether LLMs can discern their ability boundaries and use tools flexibly. We then introduce the Whether-or-not tool usage Evaluation benchmark (WTU-Eval) to assess LLMs with eleven datasets, where six of them are tool-usage datasets, and five are general datasets. LLMs are prompted to use tools according to their needs. The results of eight LLMs on WTU-Eval reveal that LLMs frequently struggle to determine tool use in general datasets, and LLMs’ performance in tool-usage datasets improves when their ability is similar to ChatGPT. Jian Liu 0032, Kangyun Ning, Yisong Su, Wenjuan Han, Jin An Xu, Yuanzhe Zhang |
ICASSP | 5 |
| 2025 | TriG-RAG: Triple-Granularity Fusion for Retrieval-Augmented Generation with Adaptive Context-Relation Balance
Jingrui Zhang, Yufeng Chen 0005, Jin An Xu |
NLPCC (3) | 3 |
| 2025 | Bridging the reality gap: A benchmark for physical reasoning in general world models with various physical phenomena beyond mechanics
Jin An Xu, Huiqi Hu, Xiuwen Xu, Zijian Jin, Fandong Meng, Jie Zhou 0016, Wenjuan Han |
Expert Syst. Appl. | 2 |
| 2024 | Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine TranslationabstractIncrementally expanding the capability of an existing translation model to solve new domain tasks over time is a fundamental and practical problem, which usually suffers from catastrophic forgetting.Generally, multi-domain learning can be seen as a good solution.However, there are two drawbacks: 1) it requires having the training data for all domains available at the same time, which may be unrealistic due to storage or privacy concerns; 2) it requires re-training the model on the data of all domains from scratch when adding a new domain and this is time-consuming and computationally expensive.To address these issues, we present a semi-supervised contrastive distillation framework for incremental neural machine translation.Specifically, to avoid catastrophic forgetting, we propose to exploit unlabeled data from the same distributions of the older domains through knowledge distillation.Further, to ensure the distinct domain characteristics in the model as the number of domains increases, we devise a cross-domain contrastive objective to enhance the distilled knowledge.Extensive experiments on domain translation benchmarks show that our approach, without accessing any previous training data or re-training on all domains from scratch, can significantly prevent the model from forgetting previously learned knowledge while obtaining good performance on the incrementally added domains. Yunlong Liang, Fandong Meng, Jiaan Wang, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016 |
ACL (1) | 4 |
| 2024 | DoRA: Enhancing Parameter-Efficient Fine-Tuning with Dynamic Rank DistributionabstractFine-tuning large-scale pre-trained models is inherently a resource-intensive task.While it can enhance the capabilities of the model, it also incurs substantial computational costs, posing challenges to the practical application of downstream tasks.Existing parameter-efficient fine-tuning (PEFT) methods such as Low-Rank Adaptation (LoRA) rely on a bypass framework that ignores the differential parameter budget requirements across weight matrices, which may lead to suboptimal fine-tuning outcomes.To address this issue, we introduce the Dynamic Low-Rank Adaptation (DoRA) method.DoRA decomposes high-rank LoRA layers into structured single-rank components, allowing for dynamic pruning of parameter budget based on their importance to specific tasks during training, which makes the most of the limited parameter budget.Experimental results demonstrate that DoRA can achieve competitive performance compared with LoRA and full model fine-tuning, and outperform various strong baselines with the same storage parameter budget.Our code is available at https: //github.com/MIkumikumi0116/DoRA Yulong Mao, Changhao Guan, Ganglin Bao, Fengran Mo, Jin An Xu |
ACL (1) | 6 |
| 2024 | PersonalityScanner: Exploring the Validity of Personality Assessment Based on Multimodal Signals in Virtual Reality
Huiqi Hu, Xianhao Yu, Jin An Xu, Yujia Peng, Wenjuan Han |
CogSci | 6 |
| 2024 | CollabKG: A Learnable Human-Machine-Cooperative Information Extraction Toolkit for (Event) Knowledge Graph ConstructionabstractIn order to construct or extend entity-centric and event-centric knowledge graphs (KG and EKG), the information extraction (IE) annotation toolkit is essential. However, existing IE toolkits have several non-trivial problems, such as not supporting multi-tasks, and not supporting automatic updates. In this work, we present CollabKG, a learnable human-machine-cooperative IE toolkit for KG and EKG construction. Specifically, for the multi-task issue, CollabKG unifies different IE subtasks, including named entity recognition (NER), entity-relation triple extraction (RE), and event extraction (EE), and supports both KG and EKG. Then, combining advanced prompting-based IE technology, the human-machine-cooperation mechanism with Large Language Models (LLMs) as the assistant machine is presented which can provide a lower cost as well as a higher performance. Lastly, owing to the two-way interaction between the human and machine, CollabKG with learning ability allows self-renewal. Besides, CollabKG has several appealing features (e.g., customization, training-free, and label propagation) that make the system powerful and high-productivity. We holistically compare our toolkit with other existing tools on these features. Human evaluation quantitatively illustrates that CollabKG significantly improves annotation quality, efficiency, and stability simultaneously. Yufeng Chen 0005, Xingyu Cui, Jin An Xu, Wenjuan Han |
LREC/COLING | 5 |
| 2024 | A Reinforcement Learning Approach to Improve Low-Resource Machine Translation Leveraging Domain Monolingual DataabstractDue to the lack of parallel data, the mainstream fine-tuning-based domain adaptation methods have the overfitting problem in the translation of low-resource domains, and it is difficult for the model to learn the in-domain generalization knowledge. To address the above issue, in this work, we propose a novel Reinforcement Learning Domain Adaptation method for Neural Machine Translation (RLDA-NMT) in the low-resource domain. RLDA-NMT utilizes in-domain source monolingual data to make up for the lack of parallel data, and reinforces domain features learning to make the translation model learn the domain-specific knowledge more fully. Specifically, we first train a ranking-based model with a small-scale in-domain parallel corpus, and then adopt it as the reward model to select higher-quality generated translations for reinforcement when fine-tuning pre-trained NMT model using in-domain source monolingual data. We conduct experiments on Education, Laws, Thesis, and Patent domains of Chinese⇔English translation tasks. Experimental results demonstrate that RLDA-NMT can alleviate overfitting and reinforce the NMT model to learn domain-specific knowledge. Additionally, the results also show that RLDA-NMT and back-translation (BT) are nicely complementary to each other, where combining RLDA-NMT with BT can further improve translation quality. Hongxiao Zhang, Mingtong Liu, Chunyou Li, Yufeng Chen 0005, Jin An Xu |
LREC/COLING | 5 |
| 2024 | Dual-Space Knowledge Distillation for Large Language ModelsabstractKnowledge distillation (KD) is known as a promising solution to compress large language models (LLMs) via transferring their knowledge to smaller models.During this process, white-box KD methods usually minimize the distance between the output distributions of the two models so that more knowledge can be transferred.However, in the current whitebox KD framework, the output distributions are from the respective output spaces of the two models, using their own prediction heads.We argue that the space discrepancy will lead to low similarity between the teacher model and the student model on both representation and distribution levels.Furthermore, this discrepancy also hinders the KD process between models with different vocabularies, which is common for current LLMs.To address these issues, we propose a dual-space knowledge distillation (DSKD) framework that unifies the output spaces of the two models for KD.On the basis of DSKD, we further develop a cross-model attention mechanism, which can automatically align the representations of the two models with different vocabularies.Thus, our framework is not only compatible with various distance functions for KD (e.g., KL divergence) like the current framework, but also supports KD between any two LLMs regardless of their vocabularies.Experiments on task-agnostic instructionfollowing benchmarks show that DSKD significantly outperforms the current white-box KD framework with various distance functions, and also surpasses existing KD methods for LLMs with different vocabularies 1 .* Yufeng Chen is the corresponding author.vocabulary, which, however, is hardly satisfied for various LLMs in this era ( §2.2.2).Towards these limitations, we then propose a new framework for white-box KD, named dualspace knowledge distillation (DSKD), which is as simple as the current white-box KD framework but addresses the issues due to the space discrepancy.Specifically, DSKD unifies the output spaces of the two models by projecting the output hidden states 2 of the teacher/student to the representation spaces of the student/teacher, where we can use the shared prediction heads to produce the two distributions in the same output spaces.In particular, for models with different vocabularies, we further develop a cross-model attention (CMA) mechanism to automatically align the tokens in two differently tokenized sequences.Like the current framework, DSKD is also compatible with existing distance functions for distributions, including KL divergence, JS divergence, and so on.Meanwhile, with CMA, we can transform distributions of the two LLMs into the same shape, which makes our framework more general and can be applied to any two LLMs regardless of their vocabularies.We evaluate our framework on instructionfollowing benchmarks under both settings that the two LLMs have the same/different vocabularies.Experimental results showcase that for LLMs with the same vocabulary, our DSKD framework significantly outperforms the current white-box KD framework on various distance functions.Moreover, DSKD with CMA surpasses all existing KD methods for LLMs with different vocabularies.To sum up, the contributions are as follows:• We empirically reveal that the current whitebox KD framework limits the similarity between the student and the teacher due to their different output spaces.• As a solution, we propose a new framework for white-box KD, named dual-space knowledge distillation (DSKD), which unifies the output spaces of the distributions from the teacher and the student for more effective KD.• Based on DSKD, we further develop a crossmodel attention mechanism to support KD between LLMs with different vocabularies. Songming Zhang 0001, Zengkui Sun, Yufeng Chen 0005, Jin An Xu |
EMNLP | 5 |
| 2024 | Empowering Vision-Language Models for Reasoning Ability through Large Language ModelsabstractVision-language models (VLM) have shown excellent performance in vision-language tasks. However, they sometimes lack sufficient reasoning ability. In contrast, large language models (LLMs) have emerged with powerful reasoning capabilities. Therefore, we propose a framework called TReE, which transfers the reasoning ability of the LLM to the VLM in learning-free settings. TReE is a three-stage framework: observation, thinking, and re-thinking. The observation stage requires the VLM to obtain overall visual information about the image. Then, the thinking stage combines the visual information and task description as the prompt for the LLM, allowing it to present the thinking process (namely, rationale). Lastly, the re-thinking stage learns useful information from the rationale and then predicts the final result using the VLM. We are the first to explore enhancing the VLM’s reasoning ability without any training, finetuning, or access to the LLM’s parameters, which we refer to as a plug-in mode, leading to the model-agnostic feature. Experiments show that TReE performed well on general visual questionanswering (VQA) tasks and outperformed KOSMOS-1 on the challenging Raven IQ test dataset by 6%. Furthermore, with additional lightweight finetuning using a smaller amount of parameters, TReE achieved a high accuracy of 81.7% on GQA and 67.3% on VQAv2. Jin An Xu, Wenjuan Han |
ICASSP | 3 |
| 2024 | Low-Resource Event Causality Identification With Global Consistency Constraints
Kangyun Ning, Jian Liu 0032, Jin An Xu |
NLPCC (1) | 3 |
| 2024 | Curriculum pre-training for stylized neural machine translation
Aixiao Zou, Xuanxuan Wu, Xinjie Li 0003, Fuwei Cui, Jin An Xu |
Appl. Intell. | 6 |
| 2024 | LegalAsst: Human-centered and AI-empowered machine to enhance court productivity and legal assistance
Wenjuan Han, Jiaxin Shen, Yanyao Liu, Jin An Xu, Fangxu Hu, Xueli Yu, Huaqing Wang, Zhijing Liu, Yajie Yang, Tianshui Shi, Mengyao Ge |
Inf. Sci. | 5 |
| 2024 | An Ensemble Strategy with Gradient Conflict for Multi-Domain Neural Machine TranslationabstractMulti-domain neural machine translation aims to construct a unified neural machine translation model to translate sentences across various domains. Nevertheless, previous studies have one limitation is the incapacity to acquire both domain-general and domain-specific representations concurrently. To this end, we propose an ensemble strategy with gradient conflict for multi-domain neural machine translation that automatically learns model parameters by identifying both domain-shared and domain-specific features. Specifically, our approach consists of (1) a parameter-sharing framework, where the parameters of all the layers are originally shared and equivalent to each domain, and (2) ensemble strategy, in which we design an Extra Ensemble strategy via a piecewise condition function to learn direction and distance-based gradient conflict. In addition, we give a detailed theoretical analysis of the gradient conflict to further validate the effectiveness of our approach. Experimental results on two multi-domain datasets show the superior performance of our proposed model compared to previous work. Zhibo Man, Yu Li 0025, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2024 | WDSRL: Multi-Domain Neural Machine Translation With Word-Level Domain-Sensitive Representation LearningabstractDue to the strong reliance on domain-specific knowledge, the joint learning manner of domain discrimination and translation has been widely considered in the Multi-Domain Neural Machine Translation (MDNMT) task. However, the word ambiguity problem still inevitably exists in MDNMT, especially when mixed multi-domain data is brought into the model training phase. Although word-level MDNMT can mitigate this problem to some extent, poor domain discrimination yet remains and severely hinders performance. Based on the above limitation, we observed that coarser granularity strings may provide more specific semantics, which is more conducive to domain discrimination. Thus, we propose a Word-level Domain-Sensitive Representation Learning (WDSRL) method. Specifically, we focus on two aspects of our approach: domain representation and domain discrimination. To extend the scope of domain representation, we adopt Convolution Neural Networks (CNN) to encode Local Domain Representation at different granularities, and then integrate Topic Knowledge Representation into each word. By doing so, context features related to the domain could be comprehensively enriched. Regarding domain discrimination, we design a Domain-Sensitive Discriminator, which could not only generate domain features for each word but also enhance domain representation learning. Experimental results demonstrate our substantial improvements over several representative baselines on multiple language pairs. Furthermore, the extensive analysis also indicates the superiority of our proposed domain-sensitive feature encoding strategy and domain-sensitive discriminator for word-level representation learning. Zhibo Man, Zengcheng Huang, Yu Li 0025, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2024 | Complex Question Enhanced Transfer Learning for Zero-Shot Joint Information ExtractionabstractZero-shot information extraction (IE) tasks have attracted great attention recently. However, how to jointly model multiple IE tasks in the zero-shot scenario is still an open question. In this article, we focus on zero-shot joint IE tasks and highlight how to transfer the knowledge of cross-task relations from the source domain to the target domain. To solve this problem, we first unify all IE tasks with a machine reading comprehension (MRC) framework, which can make the most of training data and enhance its ability on span extraction. Then, we generatecomplex questionsto explicitly model cross-task relations with natural language descriptions, thereby providing prior knowledge for pre-defined types and building more general linkages among different entities and triggers as well. Specifically, we define three operations for generating templates for complex questions, i.e.,intersecting,connecting, andcomposing. Besides, we design an efficient training strategy to exploit the synthetic data with complex questions. We evaluate our approach on four datasets from different domains for various IE tasks. Experimental results show the effectiveness of our approach in improving the performance of zero-shot joint IE tasks in multiple domains. Ying Zhang 0084, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Summary-Oriented Vision Modeling for Multimodal Abstractive SummarizationabstractMultimodal abstractive summarization (MAS) aims to produce a concise summary given the multimodal data (text and vision).Existing studies mainly focus on how to effectively use the visual features from the perspective of an article, having achieved impressive success on the high-resource English dataset.However, less attention has been paid to the visual features from the perspective of the summary, which may limit the model performance, especially in the low-and zero-resource scenarios.In this paper, we propose to improve the summary quality through summary-oriented visual features.To this end, we devise two auxiliary tasks including vision to summary task and masked image modeling task.Together with the main summarization task, we optimize the MAS model via the training objectives of all these tasks.By these means, the MAS model can be enhanced by capturing the summaryoriented visual features, thereby yielding more accurate summaries.Experiments on 44 languages, covering mid-high-, low-, and zeroresource scenarios, verify the effectiveness and superiority of the proposed approach, which achieves state-of-the-art performance under all scenarios.Additionally, we will contribute a large-scale multilingual multimodal abstractive summarization (MM-Sum) dataset. 1 Yunlong Liang, Fandong Meng, Jin An Xu, Jiaan Wang, Yufeng Chen 0005, Jie Zhou 0016 |
ACL (1) | 3 |
| 2023 | Document-Level Event Argument Extraction With a Chain Reasoning ParadigmabstractDocument-level event argument extraction aims to identify event arguments beyond sentence level, where a significant challenge is to model long-range dependencies.Focusing on this challenge, we present a new chain reasoning paradigm for the task, which can generate decomposable first-order logic rules for reasoning.This paradigm naturally captures long-range interdependence due to the chains' compositional nature, which also improves interpretability by explicitly modeling the reasoning process.We introduce T-norm fuzzy logic for optimization, which permits end-toend learning and shows promise for integrating the expressiveness of logical reasoning with the generalization of neural networks.In experiments, we show that our approach outperforms previous methods by a significant margin on two standard benchmarks (over 6 points in F1).Moreover, it is data-efficient in lowresource scenarios and robust enough to defend against adversarial attacks. Jian Liu 0032, Jin An Xu, Haoyan Liu 0001, Zhe Zhao 0006 |
ACL (1) | 3 |
| 2023 | Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationabstractSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen, Wenjuan Han, Jian Liu, Jinan Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Songming Zhang 0001, Yunlong Liang, Shuaibo Wang, Yufeng Chen 0005, Wenjuan Han, Jian Liu 0032, Jin An Xu |
ACL (1) | 7 |
| 2023 | MT2: Towards a Multi-Task Machine Translation Model with Translation-Specific In-Context LearningabstractSentence-level translation, document-level translation, translation memory, and terminology constrained translation play an important role in machine translation.Most of the previous work uses separate models or methods to solve these tasks, which is not conducive to knowledge transfer of different tasks and increases the complexity of system construction.In this work, we explore the potential of pre-trained language model in machine translation tasks and propose a Multi-Task Machine Translation (MT2) model to integrate these translation tasks.We design a novel translationspecific In-Context Learning (ICL) paradigm for model training, in which all of the translation tasks can be modeled as context-learning tasks that integrate contextual information for performance improvement.Specifically, we propose a retrieval and alignment method to obtain a large scale context-enhancement training data, then we train the model in an in-context learning manner.Furthermore, we adopt two context-dependent training strategies to encourage the model to better understand and utilize contextual information for translation.Extensive experiments on translation memory, terminology constrained translation, document-level translation, and few-shot domain-adaptation tasks demonstrate the superior performance of our model, verifying the effectiveness of our proposed approach. Chunyou Li, Mingtong Liu, Hongxiao Zhang, Yufeng Chen 0005, Jin An Xu |
EMNLP | 5 |
| 2023 | Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFsabstractReal-world named entity recognition (NER) datasets are notorious for their noisy nature, attributed to annotation errors, inconsistencies, and subjective interpretations.Such noises present a substantial challenge for traditional supervised learning methods.In this paper, we present a new and unified approach to tackle annotation noises for NER.Our method considers NER as a constituency tree parsing problem, utilizing a tree-structured Conditional Random Fields (CRFs) with uncertainty evaluation for integration.Through extensive experiments conducted on four realworld datasets, we demonstrate the effectiveness of our model in addressing both partial and incorrect annotation errors.Remarkably, our model exhibits superb performance even in extreme scenarios with 90% annotation noise. Jian Liu 0032, Weichang Liu, Yufeng Chen 0005, Jin An Xu, Zhe Zhao 0006 |
EMNLP | 4 |
| 2023 | A Quality-based Syntactic Template Retriever for Syntactically-Controlled Paraphrase GenerationabstractExisting syntactically-controlled paraphrase generation (SPG) models perform promisingly with human-annotated or well-chosen syntactic templates.However, the difficulty of obtaining such templates actually hinders the practical application of SPG models.For one thing, the prohibitive cost makes it unfeasible to manually design decent templates for every source sentence.For another, the templates automatically retrieved by current heuristic methods are usually unreliable for SPG models to generate qualified paraphrases.To escape this dilemma, we propose a novel Quality-based Syntactic Template Retriever (QSTR) to retrieve templates based on the quality of the to-be-generated paraphrases.Furthermore, for situations requiring multiple paraphrases for each source sentence, we design a Diverse Templates Search (DTS) algorithm, which can enhance the diversity between paraphrases without sacrificing quality.Experiments demonstrate that QSTR can significantly surpass existing retrieval methods in generating high-quality paraphrases and even perform comparably with human-annotated templates in terms of reference-free metrics.Additionally, human evaluation and the performance on downstream tasks using our generated paraphrases for data augmentation showcase the potential of our QSTR and DTS algorithm in practical scenarios. Songming Zhang 0001, Yunlong Liang, Yufeng Chen 0005, Jian Liu 0032, Wenjuan Han, Jin An Xu |
EMNLP | 7 |
| 2023 | Exploring Domain-shared and Domain-specific Knowledge in Multi-Domain Neural Machine TranslationabstractCurrently, multi-domain neural machine translation (NMT) has become a significant research topic in domain adaptation machine translation, which trains a single model by mixing data from multiple domains. Multi-domain NMT aims to improve the performance of the low-resources domain through data augmentation. However, mixed domain data brings more translation ambiguity. Previous work focused on domain-general or domain-context knowledge learning, respectively. Therefore, there is a challenge for acquiring domain-general or domain-context knowledge simultaneously. To this end, we propose a unified framework for learning simultaneously domain-general and domain-specific knowledge, we are the first to apply parameter differentiation in multi-domain NMT. Specifically, we design the differentiation criterion and differentiation granularity to obtain domain-specific parameters. Experimental results on multi-domain UM-corpus English-to-Chinese and OPUS German-to-English datasets show that the average BLEU scores of the proposed method exceed the strong baseline by 1.22 and 1.87, respectively. In addition, we investigate the case study to illustrate the effectiveness of the proposed method in acquiring domain knowledge. Zhibo Man, Yuanmeng Chen, Yufeng Chen 0005, Jin An Xu |
MTSummit (1) | 5 |
| 2023 | Multi-source inverse-curriculum-based training for low-resource dialogue generation
Fuwei Cui, Hui Di, Hongjie Ren, Kazushige Ouchi, Jin An Xu |
Appl. Intell. | 7 |
| 2023 | A Multi-Task Multi-Stage Transitional Training Framework for Neural Chat TranslationabstractNeural chat translation (NCT) aims to translate a cross-lingual chat between speakers of different languages. Existing context-aware NMT models cannot achieve satisfactory performances due to the following inherent problems: 1) limited resources of annotated bilingual dialogues; 2) the neglect of modelling conversational properties; 3) training discrepancy between different stages. To address these issues, in this paper, we propose a multi-task multi-stage transitional (MMT) training framework, where an NCT model is trained using the bilingual chat translation dataset and additional monolingual dialogues. We elaborately design two auxiliary tasks, namely utterance discrimination and speaker discrimination, to introduce the modelling of dialogue coherence and speaker characteristic into the NCT model. The training process consists of three stages: 1) sentence-level pre-training on large-scale parallel corpus; 2) intermediate training with auxiliary tasks using additional monolingual dialogues; 3) context-aware fine-tuning with gradual transition. Particularly, the second stage serves as an intermediate phase that alleviates the training discrepancy between the pre-training and fine-tuning stages. Moreover, to make the stage transition smoother, we train the NCT model using a gradual transition strategy, i.e., gradually transiting from using monolingual to bilingual dialogues. Extensive experiments on two language pairs demonstrate the effectiveness and superiority of our proposed training framework. Chulun Zhou, Yunlong Liang, Fandong Meng, Jie Zhou 0016, Jin An Xu, Min Zhang 0005, Jinsong Su |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Combination of Loss-based Active Learning and Semi-supervised Learning for Recognizing Entities in Chinese Electronic Medical RecordsabstractThe recognition of entities in an electronic medical record (EMR) is especially important to downstream tasks, such as clinical entity normalization and medical dialogue understanding. However, in the medical professional field, training a high-quality named entity recognition system always requires large-scale annotated datasets, which are highly expensive to obtain. In this article, to lower the cost of data annotation and maximizing the use of unlabeled data, we propose a hybrid approach to recognizing the entities in Chinese electronic medical record, which is in combination of loss-based active learning and semi-supervised learning. Specifically, we adopted a dynamic balance strategy to dynamically balance the minimum loss predicted by a named entity recognition decoder and a loss prediction module at different stages in the process. Experimental results demonstrated our proposed framework’s effectiveness and efficiency, achieving higher performances than existing approaches on Chinese EMR entity recognition datasets under limited labeling resources. Jinghui Yan, Chengqing Zong, Jin An Xu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | A Neighborhood Re-Ranking Model With Relation Constraint for Knowledge Graph CompletionabstractKnowledge graph completion (KGC) aims to predict missing links based on observed triples. However, current KGC models are still limited by the following two aspects. (1) the entity semantics is implicitly learned by neural network and merely depends on existing facts, which mostly suffers from less additional specific knowledge. Although previous studies have noticed that entity type information can effectively improve KGC task, most of them rely on labeled type-specific data. (2) the recent graph-based models mainly concentrate on Graph Neural Network (GNN) to update source entity representation, regardless of the separate role that neighborhood information plays and may mix noisy neighbor features for target prediction. To address the above two issues, we propose a neighborhood re-ranking model with relation constraint for KGC task. We suggest that both relation constraint and structured information located in triples can boost the model performance. More importantly, we automatically generate explicit constraints as additional type feature to enrich entity representation instead of depending on human annotated labels. Meanwhile, we construct a neighborhood completion module to re-rank candidate entities for full use of the neighbor structure rather than traditional GNN updating manner. Extensive experiments on seven benchmarks demonstrate that our model achieves the competitive results in comparison to the recent advanced baselines. Yu Li 0025, Bojie Hu, Jian Liu 0032, Yufeng Chen 0005, Jin An Xu |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Scheduled Multi-task Learning for Neural Chat TranslationabstractNeural Chat Translation (NCT) aims to translate conversational text into different languages.Existing methods mainly focus on modeling the bilingual dialogue characteristics (e.g., coherence) to improve chat translation via multi-task learning on small-scale chat translation data.Although the NCT models have achieved impressive success, it is still far from satisfactory due to insufficient chat translation data and simple joint training manners.To address the above issues, we propose a scheduled multi-task learning framework for NCT.Specifically, we devise a three-stage training framework to incorporate the large-scale in-domain chat translation data into training by adding a second pre-training stage between the original pre-training and fine-tuning stages.Further, we investigate where and how to schedule the dialogue-related auxiliary tasks in multiple training stages to effectively enhance the main chat translation task.Extensive experiments on four language directions (English↔Chinese and English↔German) verify the effectiveness and superiority of the proposed approach.Additionally, we will make the large-scale indomain paired bilingual dialogue dataset publicly available for the research community.1 Yunlong Liang, Fandong Meng, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016 |
ACL (1) | 3 |
| 2022 | MSCTD: A Multimodal Sentiment Chat Translation DatasetabstractMultimodal machine translation and textual chat translation have received considerable attention in recent years.Although the conversation in its natural form is usually multimodal, there still lacks work on multimodal machine translation in conversations.In this work, we introduce a new task named Multimodal Chat Translation (MCT), aiming to generate more accurate translations with the help of the associated dialogue history and visual context.To this end, we firstly construct a Multimodal Sentiment Chat Translation Dataset (MSCTD) containing 142,871 English-Chinese utterance pairs in 14,762 bilingual dialogues and 30,370 English-German utterance pairs in 3,079 bilingual dialogues.Each utterance pair, corresponding to the visual context that reflects the current conversational scene, is annotated with a sentiment label.Then, we benchmark the task by establishing multiple baseline systems that incorporate multimodal and sentiment features for MCT.Preliminary experiments on four language directions (English↔Chinese and English↔German) verify the potential of contextual and multimodal information fusion and the positive impact of sentiment on the MCT task.Additionally, as a by-product of the MSCTD, it also provides two new benchmarks on multimodal dialogue sentiment analysis.Our work can facilitate research on both multimodal chat translation and multimodal dialogue sentiment analysis.1 Yunlong Liang, Fandong Meng, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016 |
ACL (1) | 3 |
| 2022 | A Variational Hierarchical Model for Neural Cross-Lingual SummarizationabstractYunlong Liang, Fandong Meng, Chulun Zhou, Jinan Xu, Yufeng Chen, Jinsong Su, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yunlong Liang, Fandong Meng, Chulun Zhou, Jin An Xu, Yufeng Chen 0005, Jinsong Su, Jie Zhou 0016 |
ACL (1) | 4 |
| 2022 | Saliency as Evidence: Event Detection with Trigger Saliency AttributionabstractEvent detection (ED) is a critical subtask of event extraction that seeks to identify event triggers of certain types in texts.Despite significant advances in ED, existing methods typically follow a "one model fits all types" approach, which sees no differences between event types and often results in a quite skewed performance.Finding the causes of skewed performance is crucial for the robustness of an ED model, but to date there has been little exploration of this problem.This research examines the issue in depth and presents a new concept termed trigger salience attribution, which can explicitly quantify the underlying patterns of events.On this foundation, we develop a new training mechanism for ED, which can distinguish between triggerdependent and context-dependent types and achieve promising performance on two benchmarks.Finally, by highlighting many distinct characteristics of trigger-dependent and context-dependent types, our work may promote more research into this problem. Jian Liu 0032, Yufeng Chen 0005, Jin An Xu |
ACL (1) | 3 |
| 2022 | Conditional Bilingual Mutual Information Based Adaptive Training for Neural Machine TranslationabstractSongming Zhang, Yijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu, Jian Liu, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Songming Zhang 0001, Yijin Liu, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jian Liu 0032, Jie Zhou 0016 |
ACL (1) | 5 |
| 2022 | Iterative Constrained Back-Translation for Unsupervised Domain Adaptation of Machine TranslationabstractBack-translation has been proven to be effective in unsupervised domain adaptation of neural machine translation (NMT). However, the existing back-translation methods mainly improve domain adaptability by generating in-domain pseudo-parallel data that contains sentence-structural knowledge, paying less attention to the in-domain lexical knowledge, which may lead to poor translation of unseen in-domain words. In this paper, we propose an Iterative Constrained Back-Translation (ICBT) method to incorporate in-domain lexical knowledge on the basis of BT for unsupervised domain adaptation of NMT. Specifically, we apply lexical constraints into back-translation to generate pseudo-parallel data with in-domain lexical knowledge, and then perform round-trip iterations to incorporate more lexical knowledge. Based on this, we further explore sampling strategies of constrained words in ICBT to introduce more targeted lexical knowledge, via domain specificity and confidence estimation. Experimental results on four domains show that our approach achieves state-of-the-art results, improving the BLEU score by up to 3.08 compared to the strongest baseline, which demonstrates the effectiveness of our approach. Hongxiao Zhang, Yufeng Chen 0005, Jin An Xu, Jian Liu 0032 |
COLING | 5 |
| 2022 | Cross-Align: Modeling Deep Cross-lingual Interactions for Word AlignmentabstractWord alignment which aims to extract lexicon translation equivalents between source and target sentences, serves as a fundamental tool for natural language processing.Recent studies in this area have yielded substantial improvements by generating alignments from contextualized embeddings of the pre-trained multilingual language models.However, we find that the existing approaches capture few interactions between the input sentence pairs, which degrades the word alignment quality severely, especially for the ambiguous words in the monolingual context.To remedy this problem, we propose Cross-Align to model deep interactions between the input sentence pairs, in which the source and target sentences are encoded separately with the shared self-attention modules in the shallow layers, while cross-lingual interactions are explicitly constructed by the crossattention modules in the upper layers.Besides, to train our model effectively, we propose a two-stage training framework, where the model is trained with a simple Translation Language Modeling (TLM) objective in the first stage and then finetuned with a self-supervised alignment objective in the second stage.Experiments show that the proposed Cross-Align achieves the state-of-the-art (SOTA) performance on four out of five language pairs. 1 Siyu Lai, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
EMNLP | 5 |
| 2022 | Long Text Generation with Topic-aware Discrete Latent Variable ModelabstractGenerating coherent long texts is an important yet challenging task, particularly for the openended generation task.Prior work based on discrete latent codes focuses on the modeling of discourse relation, resulting in discrete codes only learning shallow semantics (Ji and Huang, 2021).A natural text always revolves around several related topics and the transition across them is natural and smooth.In this work, we investigate whether discrete latent codes can learn information of topics.To this end, we build a topic-aware latent code-guided text generation model.To encourage discrete codes to model information about topics, we propose a span-level bag-of-words training objective for the model.Automatic and manual evaluation experiments show that our method can generate more topic-relevant and coherent texts. Erguang Yang, Mingtong Liu, Deyi Xiong, Yufeng Chen 0005, Jin An Xu |
EMNLP | 6 |
| 2022 | Low-Resource NER by Data Augmentation With PromptingabstractNamed entity recognition (NER) is a fundamental information extraction task that seeks to identify entity mentions of certain types in text. Despite numerous advances, the existing NER methods rely on extensive supervision for model training, which struggle in a low-resource scenario with limited training data. In this paper, we propose a new data augmentation method for low-resource NER, by eliciting knowledge from BERT with prompting strategies. Particularly, we devise a label-conditioned word replacement strategy that can produce more label-consistent examples by capturing the underlying word-label dependencies, and a prompting with question answering method to generate new training data from unlabeled texts. The experimental results have widely confirmed the effectiveness of our approach. Particularly, in a low-resource scenario with only 150 training sentences, our approach outperforms previous methods without data augmentation by over 40% in F1 and prior best data augmentation methods by over 2.0% in F1. Furthermore, our approach also fits with a zero-shot scenario, yielding promising results without using any human-labeled data for the task. Jian Liu 0032, Yufeng Chen 0005, Jin An Xu |
IJCAI | 3 |
| 2022 | Multimedia Event Extraction From News With a Unified Contrastive Learning FrameworkabstractExtracting events from news have seen many benefits in downstream applications. Today's event extraction (EE) systems, however, usually focus on a single modality --- either for text or image, and such methods suffer from incomplete information because a news document is typically presented in a multimedia format. In this paper, we propose a new method for multimedia EE by bridging the textual and visual modalities with a unified contrastive learning framework. Our central idea is to create a shared space for texts and images in order to improve their similar representation. This is accomplished by training on text-image pairs in general, and we demonstrate that it is possible to use this framework to boost learning for one modality by investigating the complementary of the other modality. On the benchmark dataset, our approach establishes a new state-of-the-art performance and shows a 3 percent improvement in F1. Furthermore, we demonstrate that it can achieve cutting-edge performance for visual EE even in a zero-shot scenario with no annotated data in the visual modality. Jian Liu 0032, Yufeng Chen 0005, Jin An Xu |
ACM Multimedia | 3 |
| 2022 | Generating Authentic Adversarial Examples beyond Meaning-preserving with Doubly Round-trip TranslationabstractSiyu Lai, Zhen Yang, Fandong Meng, Xue Zhang, Yufeng Chen, Jinan Xu, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Siyu Lai, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
NAACL-HLT | 6 |
| 2022 | Emotional conversation generation with heterogeneous graph neural network
Yunlong Liang, Fandong Meng, Ying Zhang 0084, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
Artif. Intell. | 5 |
| 2022 | Modeling semantic and emotional relationship in multi-turn emotional conversations using multi-task learning
Fuwei Cui, Hui Di, Lei Shen 0001, Kazushige Ouchi, Jin An Xu |
Appl. Intell. | 6 |
| 2022 | Document-level event argument linking as machine reading comprehension
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu |
Neurocomputing | 3 |
| 2022 | Improving generation diversity via syntax-controlled paraphrasing
Erguang Yang, Mingtong Liu, Deyi Xiong, Jin An Xu, Yufeng Chen 0005 |
Neurocomputing | 6 |
| 2022 | Document-level event argument extraction with self-augmentation and a cross-domain joint training mechanism
Jian Liu 0032, Jin An Xu |
Knowl. Based Syst. | 3 |
| 2022 | Exploiting Morpheme and Cross-lingual Knowledge to Enhance Mongolian Named Entity RecognitionabstractMongolian named entity recognition (NER) is not only one of the most crucial and fundamental tasks in Mongolian natural language processing, but also an important step to improve the performance of downstream tasks such as information retrieval, machine translation, and dialog system. However, traditional Mongolian NER models heavily rely on the feature engineering. Even worse, the complex morphological structure of Mongolian words makes the data sparser. To alleviate the feature engineering and data sparsity in Mongolian named entity recognition, we propose a novel NER framework with Multi-Knowledge Enhancement (MKE-NER) . Specifically, we introduce both linguistic knowledge through Mongolian morpheme representation and cross-lingual knowledge from Mongolian-Chinese parallel corpus. Furthermore, we design two methods to exploit cross-lingual knowledge sufficiently, i.e., cross-lingual representation and cross-lingual annotation projection. Experimental results demonstrate the effectiveness of our MKE-NER model, which outperforms strong baselines and achieves the best performance (94.04% F1 score) on the traditional Mongolian benchmark. Particularly, extensive experiments with different data scales highlight the superiority of our method in low-resource scenarios. Songming Zhang 0001, Ying Zhang 0084, Yufeng Chen 0005, Du Wu, Jin An Xu, Jian Liu 0032 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2022 | MRCAug: Data Augmentation via Machine Reading Comprehension for Document-Level Event Argument ExtractionabstractDocument-level event argument extraction (EAE) is a critical event semantic understanding task that requires a model to identify an event's global arguments beyond the sentence level. Existing approaches to this problem are based on supervised learning, which require a large amount of labeled data for model training. However, due to the complicated structure of an event, human annotation for this task is costly, and the issue of inadequacy of training data has long hampered the study. In this study, we propose a novel approach to mitigating the data sparsity problem faced by document-level EAE, by linking the task with machine reading comprehension (MRC). Particularly, we devise two data augmentation regimes via MRC, including an implicit knowledge transfer method, which enables knowledge transfer from other tasks to the document-level EAE task, and an explicit data generation method, which can explicitly generate new training examples by treating a pre-trained MRC model as an annotator. Furthermore, we propose a self-training based noise reduction strategy that can effectively addresses the out-of-domain noise introduced by the data augmentation methods. The extensive assessments on three benchmarks have validated the effectiveness of our approach — it not only achieves state-of-the-art performance but also demonstrates superior results in the data-low scenario. Jian Liu 0032, Yufeng Chen 0005, Jin An Xu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Cross-Domain Slot Filling as Machine Reading Comprehension: A New PerspectiveabstractWith intelligent dialogue systems becoming more and more important in our daily lives, slot filling, one of the most important components of an intelligent dialogue system, has gotten a lot of attention from academia and industry. Despite many advancements in the single-domain learning paradigm for slot filling, leveraging resources from different domains to boost learning for a target domain remains a challenge. In contrast to prior methods that supplemented a sequence labeling model with slot meta-information, we address cross-domain slot filling as a machine reading comprehension (MRC) problem for the first time, where the extraction of slot values is viewed as a question answering process. In the framework above, we present both static and dynamic question generating mechanisms, which have complimentary effects in diverse cross-domain contexts. Furthermore, we devise a dynamic question generation approach that can generate numerous values for a slot at the same time. Finally, we construct a pre-training and fine-tuning training approach that enables us to improve learning by utilizing MRC’s resources. We conducted extensive experiments on four datasets to evaluate our approach, and the experimental results clearly justified the advantages of our approach in various cross-domain settings. Jian Liu 0032, Mengshi Yu, Yufeng Chen 0005, Jin An Xu |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Infusing Multi-Source Knowledge with Heterogeneous Graph Neural Network for Emotional Conversation GenerationabstractThe success of emotional conversation systems depends on sufficient perception and appropriate expression of emotions. In a real-world conversation, we firstly instinctively perceive emotions from multi-source information, including the emotion flow of dialogue history, facial expressions, and personalities of speakers, and then express suitable emotions according to our personalities, but these multiple types of information are insufficiently exploited in emotional conversation fields. To address this issue, we propose a heterogeneous graph-based model for emotional conversation generation. Specifically, we design a Heterogeneous Graph-Based Encoder to represent the conversation content (i.e., the dialogue history, its emotion flow, facial expressions, and speakers' personalities) with a heterogeneous graph neural network, and then predict suitable emotions for feedback. After that, we employ an Emotion-Personality-Aware Decoder to generate a response not only relevant to the conversation context but also with appropriate emotions, by taking the encoded graph representations, the predicted emotions from the encoder and the personality of the current speaker as inputs. Experimental results show that our model can effectively perceive emotions from multi-source knowledge and generate a satisfactory response, which significantly outperforms previous state-of-the-art models. Yunlong Liang, Fandong Meng, Ying Zhang 0084, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
AAAI | 5 |
| 2021 | Faster Depth-Adaptive TransformersabstractDepth-adaptive neural networks can dynamically adjust depths according to the hardness of input words, and thus improve efficiency. The main challenge is how to measure such hardness and decide the required depths (i.e., layers) to conduct. Previous works generally build a halting unit to decide whether the computation should continue or stop at each layer. As there is no specific supervision of depth selection, the halting unit may be under-optimized and inaccurate, which results in suboptimal and unstable performance when modeling sentences. In this paper, we get rid of the halting unit and estimate the required depths in advance, which yields a faster depth-adaptive model. Specifically, two approaches are proposed to explicitly measure the hardness of input words and estimate corresponding adaptive depth, namely 1) mutual information (MI) based estimation and 2) reconstruction loss based estimation. We conduct experiments on the text classification task with 24 datasets in various sizes and domains. Results confirm that our approaches can speed up the vanilla Transformer (up to 7x) while preserving high accuracy. Moreover, efficiency and robustness are significantly improved when compared with other depth-adaptive approaches. Yijin Liu, Fandong Meng, Jie Zhou 0016, Yufeng Chen 0005, Jin An Xu |
AAAI | 5 |
| 2021 | Modeling Bilingual Conversational Characteristics for Neural Chat TranslationabstractYunlong Liang, Fandong Meng, Yufeng Chen, Jinan Xu, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yunlong Liang, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
ACL/IJCNLP (1) | 4 |
| 2021 | Towards Making the Most of Dialogue Characteristics for Neural Chat TranslationabstractNeural Chat Translation (NCT) aims to translate conversational text between speakers of different languages.Despite the promising performance of sentence-level and context-aware neural machine translation models, there still remain limitations in current NCT models because the inherent dialogue characteristics of chat, such as dialogue coherence and speaker personality, are neglected.In this paper, we propose to promote the chat translation by introducing the modeling of dialogue characteristics into the NCT model.To this end, we design four auxiliary tasks including monolingual response generation, cross-lingual response generation, next utterance discrimination, and speaker identification.Together with the main chat translation task, we optimize the NCT model through the training objectives of all these tasks.By this means, the NCT model can be enhanced by capturing the inherent dialogue characteristics, thus generating more coherent and speaker-relevant translations.Comprehensive experiments on four language directions (English⇔German and English⇔Chinese) verify the effectiveness and superiority of the proposed approach. Yunlong Liang, Chulun Zhou, Fandong Meng, Jin An Xu, Yufeng Chen 0005, Jinsong Su, Jie Zhou 0016 |
EMNLP (1) | 4 |
| 2021 | Machine Reading Comprehension as Data Augmentation: A Case Study on Implicit Event Argument ExtractionabstractImplicit event argument extraction (EAE) is a crucial document-level information extraction task that aims to identify event arguments beyond the sentence level.Despite many efforts for this task, the lack of enough training data has long impeded the study.In this paper, we take a new perspective to address the data sparsity issue faced by implicit EAE, by bridging the task with machine reading comprehension (MRC).Particularly, we devise two data augmentation regimes via MRC, including: 1) implicit knowledge transfer, which enables knowledge transfer from other tasks, by building a unified training framework in the MRC formulation, and 2) explicit data augmentation, which can explicitly generate new training examples, by treating MRC models as an annotator.The extensive experiments have justified the effectiveness of our approach -it not only obtains state-of-the-art performance on two benchmarks, but also demonstrates superior results in a data-low scenario. Jian Liu 0032, Yufeng Chen 0005, Jin An Xu |
EMNLP (1) | 3 |
| 2021 | Scheduled Sampling Based on Decoding Steps for Neural Machine TranslationabstractScheduled sampling is widely used to mitigate the exposure bias problem for neural machine translation.Its core motivation is to simulate the inference scene during training by replacing ground-truth tokens with predicted tokens, thus bridging the gap between training and inference.However, vanilla scheduled sampling is merely based on training steps and equally treats all decoding steps.Namely, it simulates an inference scene with uniform error rates, which disobeys the real inference scene, where larger decoding steps usually have higher error rates due to error accumulations.To alleviate the above discrepancy, we propose scheduled sampling methods based on decoding steps, increasing the selection chance of predicted tokens with the growth of decoding steps.Consequently, we can more realistically simulate the inference scene during training, thus better bridging the gap between training and inference.Moreover, we investigate scheduled sampling based on both training steps and decoding steps for further improvements.Experimentally, our approaches significantly outperform the Transformer baseline and vanilla scheduled sampling on three large-scale WMT tasks.Additionally, our approaches also generalize well to the text summarization task on two popular benchmarks. Yijin Liu, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
EMNLP (1) | 4 |
| 2021 | Syntactically-Informed Unsupervised Paraphrasing with Non-Parallel DataabstractPrevious works on syntactically controlled paraphrase generation heavily rely on largescale parallel paraphrase data that are not easily available for many languages and domains.In this paper, we take this research direction to the extreme and investigate whether it is possible to learn syntactically controlled paraphrase generation with non-parallel data.We propose a syntactically-informed unsupervised paraphrasing model based on conditional variational auto-encoder (VAE) which can generate texts in a specified syntactic structure.Particularly, we design a two-stage learning method to effectively train the model using non-parallel data.The conditional VAE is trained to reconstruct the input sentence according to the given input and its syntactic structure.Furthermore, to improve the syntactic controllability and semantic consistency of the pre-trained conditional VAE, we finetune it using syntax controlling and cycle reconstruction learning objectives, and employ Gumbel-Softmax to combine these new learning objectives.Experiment results demonstrate that the proposed model trained only on non-parallel data is capable of generating diverse paraphrases with specified structures.Additionally, we further validate the effectiveness of our method for generating syntactically adversarial examples on a sentiment analysis task. Erguang Yang, Mingtong Liu, Deyi Xiong, Changjian Hu, Jin An Xu, Yufeng Chen 0005 |
EMNLP (1) | 7 |
| 2021 | Discourse-Level Event Temporal Ordering with Uncertainty-Guided Graph CompletionabstractLearning to order events at discourse-level is a crucial text understanding task. Despite many efforts for this task, the current state-of-the-art methods rely heavily on manually designed features, which are costly to produce and are often specific to tasks/domains/datasets. In this paper, we propose a new graph perspective on the task, which does not require complex feature engineering but can assimilate global features and learn inter-dependencies effectively. Specifically, in our approach, each document is considered as a temporal graph, in which the nodes and edges represent events and event-event relations respectively. In this sense, the temporal ordering task corresponds to constructing edges for an empty graph. To train our model, we design a graph mask pre-training mechanism, which can learn inter-dependencies of temporal relations by learning to recover a masked edge following graph topology. In the testing stage, we design an certain-first strategy based on model uncertainty, which can decide the prediction orders and reduce the risk of error propagation. The experimental results demonstrate that our approach outperforms previous methods consistently and can meanwhile maintain good global consistency. Jian Liu 0032, Jin An Xu, Yufeng Chen 0005 |
IJCAI | 2 |
| 2021 | Improving Stylized Neural Machine Translation with Iterative Dual Knowledge TransferabstractStylized neural machine translation (NMT) aims to translate sentences of one style into sentences of another style, which is essential for the application of machine translation in a real-world scenario. However, a major challenge in this task is the scarcity of high-quality parallel data which is stylized paired. To address this problem, we propose an iterative dual knowledge transfer framework that utilizes informal training data of machine translation and formality style transfer data to create large-scale stylized paired data, for the training of stylized machine translation model. Specifically, we perform bidirectional knowledge transfer between translation model and text style transfer model iteratively through knowledge distillation. Then, we further propose a data-refinement module to process the noisy synthetic parallel data generated during knowledge transfer. Experiment results demonstrate the effectiveness of our method, achieving an improvement over the existing best model by 5 BLEU points on MTFC dataset. Meanwhile, extensive analyses illustrate our method can also improve the accuracy of formality style transfer. Xuanxuan Wu, Jian Liu 0032, Xinjie Li 0003, Jin An Xu, Yufeng Chen 0005 |
IJCAI | 4 |
| 2021 | Cross-Domain Slot Filling as Machine Reading ComprehensionabstractWith task-oriented dialogue systems being widely applied in everyday life, slot filling, the essential component of task-oriented dialogue systems, is required to be quickly adapted to new domains that contain domain-specific slots with few or no training data. Previous methods for slot filling usually adopt sequence labeling framework, which, however, often has limited ability when dealing with the domain-specific slots. In this paper, we take a new perspective on cross-domain slot filling by framing it as a machine reading comprehension (MRC) problem. Our approach firstly transforms slot names into well-designed queries, which contain rich informative prior knowledge and are very helpful for the detection of domain-specific slots. In addition, we utilize the large-scale MRC dataset for pre-training, which further alleviates the data scarcity problem. Experimental results on SNIPS and ATIS datasets show that our approach consistently outperforms the existing state-of-the-art methods by a large margin. Mengshi Yu, Jian Liu 0032, Yufeng Chen 0005, Jin An Xu |
IJCAI | 4 |
| 2021 | Contrastive Learning for Machine Translation Quality Estimation
Hui Di, Jian Liu 0032, Yufeng Chen 0005, Kazushige Ouchi, Jin An Xu |
NLPCC (1) | 6 |
| 2021 | Explore Coarse-Grained Structures for Syntactically Controllable Paraphrase Generation
Erguang Yang, Mingtong Liu, Deyi Xiong, Changjian Hu, Jin An Xu, Yufeng Chen 0005 |
NLPCC (1) | 7 |
| 2021 | Deep bi-directional interaction network for sentence matching
Mingtong Liu, Jin An Xu, Yufeng Chen 0005 |
Appl. Intell. | 3 |
| 2021 | A dependency syntactic knowledge augmented interactive architecture for end-to-end aspect-based sentiment analysis
Yunlong Liang, Fandong Meng, Jinchao Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016 |
Neurocomputing | 5 |
| 2020 | A Learning-Exploring Method to Generate Diverse Paraphrases with Multi-Objective Deep Reinforcement LearningabstractParaphrase generation (PG) is of great importance to many downstream tasks in natural language processing.Diversity is an essential nature to PG for enhancing generalization capability and robustness of downstream applications.Recently, neural sequence-to-sequence (Seq2Seq) models have shown promising results in PG.However, traditional model training for PG focuses on optimizing model prediction against single reference and employs cross-entropy loss, which objective is unable to encourage model to generate diverse paraphrases.In this work, we present a novel approach with multi-objective learning to PG.We propose a learning-exploring method to generate sentences as learning objectives from the learned data distribution, and employ reinforcement learning to combine these new learning objectives for model training.We first design a sample-based algorithm to explore diverse sentences.Then we introduce several reward functions to evaluate the sampled sentences as learning signals in terms of expressive diversity and semantic fidelity, aiming to generate diverse and high-quality paraphrases.To effectively optimize model performance satisfying different evaluating aspects, we use a GradNorm-based algorithm that automatically balances these training objectives.Experiments and analyses on Quora and Twitter datasets demonstrate that our proposed method not only gains a significant increase in diversity but also improves generation quality over several state-of-the-art baselines. Mingtong Liu, Erguang Yang, Deyi Xiong, Changjian Hu, Jin An Xu, Yufeng Chen 0005 |
COLING | 7 |
| 2020 | Exploring Bilingual Parallel Corpora for Syntactically Controllable Paraphrase GenerationabstractParaphrase generation is of great importance to many downstream tasks in natural language processing. Recent efforts have focused on generating paraphrases in specific syntactic forms, which, generally, heavily relies on manually annotated paraphrase data that is not easily available for many languages and domains. In this paper, we propose a novel end-to-end framework to leverage existing large-scale bilingual parallel corpora to generate paraphrases under the control of syntactic exemplars. In order to train one model over the two languages of parallel corpora, we embed sentences of them into the same content and style spaces with shared content and style encoders using cross-lingual word embeddings. We propose an adversarial discriminator to disentangle the content and style space, and employ a latent variable to model the syntactic style of a given exemplar in order to guide the two decoders for generation. Additionally, we introduce cycle and masking learning schemes to efficiently train the model. Experiments and analyses demonstrate that the proposed model trained only on bilingual parallel data is capable of generating diverse paraphrases with desirable syntactic styles. Fine-tuning the trained model on a small paraphrase corpus makes it substantially outperform state-of-the-art paraphrase generation models trained on a larger paraphrase dataset. Mingtong Liu, Erguang Yang, Deyi Xiong, Chen Sheng, Changjian Hu, Jin An Xu, Yufeng Chen 0005 |
IJCAI | 7 |
| 2020 | Ensemble Distilling Pretrained Language Models for Machine Translation Quality Estimation
Hui Di, Jin An Xu, Kazushige Ouchi, Yufeng Chen 0005 |
NLPCC (2) | 3 |
| 2019 | GCDT: A Global Context Enhanced Deep Transition Architecture for Sequence LabelingabstractCurrent state-of-the-art systems for the sequence labeling tasks are typically based on the family of Recurrent Neural Networks (RNNs).However, the shallow connections between consecutive hidden states of RNNs and insufficient modeling of global information restrict the potential performance of those models.In this paper, we try to address these issues, and thus propose a Global Context enhanced Deep Transition architecture for sequence labeling named GCDT.We deepen the state transition path at each position in a sentence, and further assign every token with a global representation learned from the entire sentence.Experiments on two standard sequence labeling tasks show that, given only training data and the ubiquitous word embeddings (Glove), our GCDT achieves 91.96 F 1 on the CoNLL03 NER task and 95.43 F 1 on the CoNLL2000 Chunking task, which outperforms the best reported results under the same settings.Furthermore, by leveraging BERT as an additional resource, we establish new stateof-the-art results with 93.47 F 1 on NER and 97.30F 1 on Chunking 1 . Yijin Liu, Fandong Meng, Jinchao Zhang 0001, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016 |
ACL (1) | 4 |
| 2019 | A Novel Aspect-Guided Deep Transition Model for Aspect Based Sentiment AnalysisabstractYunlong Liang, Fandong Meng, Jinchao Zhang, Jinan Xu, Yufeng Chen, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yunlong Liang, Fandong Meng, Jinchao Zhang 0001, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | CM-Net: A Novel Collaborative Memory Network for Spoken Language UnderstandingabstractYijin Liu, Fandong Meng, Jinchao Zhang, Jie Zhou, Yufeng Chen, Jinan Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yijin Liu, Fandong Meng, Jinchao Zhang 0001, Jie Zhou 0016, Yufeng Chen 0005, Jin An Xu |
EMNLP/IJCNLP (1) | 6 |
| 2019 | Original Semantics-Oriented Attention and Deep Fusion Network for Sentence MatchingabstractMingtong Liu, Yujie Zhang, Jinan Xu, Yufeng Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingtong Liu, Jin An Xu, Yufeng Chen 0005 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Sequence-to-Action Architecture for Character-Based Chinese Dependency Parsing with Status History
Hang Liu 0005, Meng Chen 0006, Jin An Xu, Yufeng Chen 0005 |
NLPCC (2) | 4 |
| 2019 | Improved Quality Estimation of Machine Translation with Pre-trained Language Representation
Guoyi Miao, Hui Di, Jin An Xu, Zhongcheng Yang, Yufeng Chen 0005, Kazushige Ouchi |
NLPCC (1) | 3 |
| 2018 | Improved Character-Based Chinese Dependency Parsing by Using Stack-Tree LSTM
Hang Liu 0005, Mingtong Liu, Jin An Xu, Yufeng Chen 0005 |
NLPCC (2) | 4 |
| 2017 | A Semantic Concept Based Unknown Words Processing Method in Neural Machine Translation
Shaotong Li, Jin An Xu, Guoyi Miao, Yufeng Chen 0005 |
NLPCC | 2 |
| 2016 | Automatic Cross-Lingual Similarization of Dependency Grammars for Tree-based Machine TranslationabstractStructural isomorphism between languages benefits the performance of cross-lingual applications.We propose an automatic algorithm for cross-lingual similarization of dependency grammars, which automatically learns grammars with high cross-lingual similarity.The algorithm similarizes the annotation styles of the dependency grammars for two languages in the level of classification decisions, and gradually improves the cross-lingual similarity without losing linguistic knowledge resorting to iterative crosslingual cooperative learning.The dependency grammars given by cross-lingual similarization have much higher cross-lingual similarity while maintaining non-triviality.As applications, the cross-lingually similarized grammars significantly improve the performance of dependency tree-based machine translation. Wenbin Jiang 0002, Jin An Xu, Rangjia Cai |
EMNLP | 3 |
| 2014 | Augment Dependency-to-String Translation with Fixed and Floating Structures
Jin An Xu, Qun Liu 0001 |
COLING | 2 |
| 2014 | Case Frame Constraints for Hierarchical Phrase-Based Translation: Japanese-Chinese as an Example
Jiangming Liu, Jin An Xu |
NLPCC | 2 |
| 2013 | An Approach of Hybrid Hierarchical Structure for Word Similarity Computing by HowNet
Jiangming Liu, Jin An Xu |
IJCNLP | 2 |
| 2013 | A Method to Construct Chinese-Japanese Named Entity Translation Equivalents Using Monolingual Corpora
Kuang Ru, Jin An Xu, Peihao Wu |
NLPCC | 2 |
| 2013 | Exploring Multiple Chinese Word Segmentation Results Based on Linear Model
Jin An Xu |
NLPCC | 4 |
| 2006 | A SVM-based personal recommendation system for TV programsabstractThis paper presents a SVM-based prediction approach for constructing personal recommendation system for TV programs. We have applied support vector machine (SVM) to personal prediction of online Internet electronic program guide (IEPG). Our basic idea is to combine SVM and feedback processing into our system, using user-watched histories as retraining data, to realize personal predictions. We evaluate the precision by experiments with open data. The results show that the proposed polynomial kernel SVM system offers a statistically significant increase in performance compared to other method, and this system demonstrates good dynamically adaptive capability. Jin An Xu, Kenji Araki |
MMM | 1 |