Jian Liu 0032

dblp:35/295-32 · DBLP profile ↗
← Back
31ranked-venue papers
19as first author
25since 2021 · last 2025
0000-0002-4728-5678ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 17 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 8 since 2021
YearPublicationVenuePosition
2025 WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models
abstract
Although Large Language Models (LLMs) excel in NLP tasks, they still need external tools to extend their ability. Current research on tool learning with LLMs often assumes mandatory tool use, which does not always align with real-world situations, where the necessity for tools is uncertain, and incorrect or unnecessary use of tools can damage the general abilities of LLMs. Therefore, we propose to explore whether LLMs can discern their ability boundaries and use tools flexibly. We then introduce the Whether-or-not tool usage Evaluation benchmark (WTU-Eval) to assess LLMs with eleven datasets, where six of them are tool-usage datasets, and five are general datasets. LLMs are prompted to use tools according to their needs. The results of eight LLMs on WTU-Eval reveal that LLMs frequently struggle to determine tool use in general datasets, and LLMs’ performance in tool-usage datasets improves when their ability is similar to ChatGPT.
Jian Liu 0032, Kangyun Ning, Yisong Su, Wenjuan Han, Jin An Xu, Yuanzhe Zhang
ICASSP1
2025 Flexible Optimal Transport With Contrastive Graphical Modeling for Multimodal Hate Detection
abstract
Multimodal hate detection plays a crucial role in maintaining harmonious online environments by identifying harmful content, such as hateful memes. Although previous research has made significant progress in detecting explicit hate speech, there remains a critical gap in analyzing implicit hate, which is particularly challenging due to the absence of explicit harmful text claims or demographic visual cues. Despite the promising results based on cross-modal attention, previous methods may suffer from the distributional modality gap caused by the non-literal associations between multimodal elements, which lacks apparent alignment in implicit hateful contents. In this work, we propose a novel framework: Flexible Optimal Transport (FLOT) to capture the non-literal cross-modal alignment for multimodal hate in the context of memes. FLOT formulates the problem of cross-modal alignment as finding optimal transportation plans, which leverages a kernel method to capture complementary information from multiple modalities. The kernel embeddings reproduce a kernel Hilbert space (RKHS) to serve as a non-linear transformation of alignment, which effectively reduces the distributional modality gap with more interpretability. Moreover, we established topological structures with contrastive modeling for the aligned representations, which are optimized to achieve comprehensive alignment between different modalities, and facilitate local reasoning based on multimodal elements. Experimental results have demonstrated that our FLOT achieved state-of-the-art performance on three publicly available benchmark datasets. Furthermore, extensive qualitative analysis confirms the superior ability of FLOT in capturing implicit cross-modal alignment.
Linhao Zhang, Li Jin 0001, Xiaoyu Li 0004, Xian Sun 0001, Xin Wang 0117, Zequn Zhang, Jian Liu 0032, Zhicong Lu, Guangluan Xu
IEEE Trans. Multim.7
2024 Video Event Extraction with Multi-View Interaction Knowledge Distillation
abstract
Video event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, which reflects the relation between objects; (2) inter-modality interaction, which aligns the features from text and video modality. In this paper, we propose a Multi-view Interaction with knowledge Distillation (MID) framework to solve the above problems with the Knowledge Distillation (KD) mechanism. Specifically, we propose the self-Relational KD (self-RKD) to enhance the inter-object interaction, where the relation between objects is measured by distance metric, and the high-level relational knowledge from the deeper layer is taken as the guidance for boosting the shallow layer in the video encoder. Meanwhile, to improve the inter-modality interaction, the Layer-to-layer KD (LKD) is proposed, which integrates additional cross-modal supervisions (i.e., the results of cross-attention) with the textual supervising signal for training each transformer decoder layer. Extensive experiments show that without any additional parameters, MID achieves the state-of-the-art performance compared to other strong methods in VEE.
Kaiwen Wei, Runyan Du, Li Jin 0001, Jian Liu 0032, Jianhua Yin 0001, Linhao Zhang, Nayu Liu, Zhi Guo
AAAI4
2024 Low-Resource Event Causality Identification With Global Consistency Constraints
Kangyun Ning, Jian Liu 0032, Jin An Xu
NLPCC (1)2
2023 Document-Level Event Argument Extraction With a Chain Reasoning Paradigm
abstract
Document-level event argument extraction aims to identify event arguments beyond sentence level, where a significant challenge is to model long-range dependencies.Focusing on this challenge, we present a new chain reasoning paradigm for the task, which can generate decomposable first-order logic rules for reasoning.This paradigm naturally captures long-range interdependence due to the chains' compositional nature, which also improves interpretability by explicitly modeling the reasoning process.We introduce T-norm fuzzy logic for optimization, which permits end-toend learning and shows promise for integrating the expressiveness of logical reasoning with the generalization of neural networks.In experiments, we show that our approach outperforms previous methods by a significant margin on two standard benchmarks (over 6 points in F1).Moreover, it is data-efficient in lowresource scenarios and robust enough to defend against adversarial attacks.
Jian Liu 0032, Jin An Xu, Haoyan Liu 0001, Zhe Zhao 0006
ACL (1)1
2023 Learning with Partial Annotations for Event Detection
abstract
Event detection (ED) seeks to discover and classify event instances in plain texts.Previous methods for ED typically adopt supervised learning, requiring fully labeled and high-quality training data.However, in a realworld application, we may not obtain clean training data but only partially labeled one, which could substantially impede the learning process.In this work, we conduct a seminal study for learning with partial annotations for ED.We propose a new trigger localization formulation using contrastive learning to distinguish ground-truth triggers from contexts, showing a decent robustness for addressing partial annotation noise.Impressively, in an extreme scenario where more than 90% of events are unlabeled, our approach achieves an F1 score of over 60%.In addition, we reannotate and make available two fully annotated subsets of ACE 2005 to serve as an unbiased benchmark for event detection.We hope our approach and data will inspire future studies on this vital yet understudied problem.
Jian Liu 0032, Dianbo Sui, Kang Liu 0001, Haoyan Liu 0001, Zhe Zhao 0006
ACL (1)1
2023 Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation
abstract
Songming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen, Wenjuan Han, Jian Liu, Jinan Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Songming Zhang 0001, Yunlong Liang, Shuaibo Wang, Yufeng Chen 0005, Wenjuan Han, Jian Liu 0032, Jin An Xu
ACL (1)6
2023 Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFs
abstract
Real-world named entity recognition (NER) datasets are notorious for their noisy nature, attributed to annotation errors, inconsistencies, and subjective interpretations.Such noises present a substantial challenge for traditional supervised learning methods.In this paper, we present a new and unified approach to tackle annotation noises for NER.Our method considers NER as a constituency tree parsing problem, utilizing a tree-structured Conditional Random Fields (CRFs) with uncertainty evaluation for integration.Through extensive experiments conducted on four realworld datasets, we demonstrate the effectiveness of our model in addressing both partial and incorrect annotation errors.Remarkably, our model exhibits superb performance even in extreme scenarios with 90% annotation noise.
Jian Liu 0032, Weichang Liu, Yufeng Chen 0005, Jin An Xu, Zhe Zhao 0006
EMNLP1
2023 A Quality-based Syntactic Template Retriever for Syntactically-Controlled Paraphrase Generation
abstract
Existing syntactically-controlled paraphrase generation (SPG) models perform promisingly with human-annotated or well-chosen syntactic templates.However, the difficulty of obtaining such templates actually hinders the practical application of SPG models.For one thing, the prohibitive cost makes it unfeasible to manually design decent templates for every source sentence.For another, the templates automatically retrieved by current heuristic methods are usually unreliable for SPG models to generate qualified paraphrases.To escape this dilemma, we propose a novel Quality-based Syntactic Template Retriever (QSTR) to retrieve templates based on the quality of the to-be-generated paraphrases.Furthermore, for situations requiring multiple paraphrases for each source sentence, we design a Diverse Templates Search (DTS) algorithm, which can enhance the diversity between paraphrases without sacrificing quality.Experiments demonstrate that QSTR can significantly surpass existing retrieval methods in generating high-quality paraphrases and even perform comparably with human-annotated templates in terms of reference-free metrics.Additionally, human evaluation and the performance on downstream tasks using our generated paraphrases for data augmentation showcase the potential of our QSTR and DTS algorithm in practical scenarios.
Songming Zhang 0001, Yunlong Liang, Yufeng Chen 0005, Jian Liu 0032, Wenjuan Han, Jin An Xu
EMNLP5
2023 A Neighborhood Re-Ranking Model With Relation Constraint for Knowledge Graph Completion
abstract
Knowledge graph completion (KGC) aims to predict missing links based on observed triples. However, current KGC models are still limited by the following two aspects. (1) the entity semantics is implicitly learned by neural network and merely depends on existing facts, which mostly suffers from less additional specific knowledge. Although previous studies have noticed that entity type information can effectively improve KGC task, most of them rely on labeled type-specific data. (2) the recent graph-based models mainly concentrate on Graph Neural Network (GNN) to update source entity representation, regardless of the separate role that neighborhood information plays and may mix noisy neighbor features for target prediction. To address the above two issues, we propose a neighborhood re-ranking model with relation constraint for KGC task. We suggest that both relation constraint and structured information located in triples can boost the model performance. More importantly, we automatically generate explicit constraints as additional type feature to enrich entity representation instead of depending on human annotated labels. Meanwhile, we construct a neighborhood completion module to re-rank candidate entities for full use of the neighbor structure rather than traditional GNN updating manner. Extensive experiments on seven benchmarks demonstrate that our model achieves the competitive results in comparison to the recent advanced baselines.
Yu Li 0025, Bojie Hu, Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Saliency as Evidence: Event Detection with Trigger Saliency Attribution
abstract
Event detection (ED) is a critical subtask of event extraction that seeks to identify event triggers of certain types in texts.Despite significant advances in ED, existing methods typically follow a "one model fits all types" approach, which sees no differences between event types and often results in a quite skewed performance.Finding the causes of skewed performance is crucial for the robustness of an ED model, but to date there has been little exploration of this problem.This research examines the issue in depth and presents a new concept termed trigger salience attribution, which can explicitly quantify the underlying patterns of events.On this foundation, we develop a new training mechanism for ED, which can distinguish between triggerdependent and context-dependent types and achieve promising performance on two benchmarks.Finally, by highlighting many distinct characteristics of trigger-dependent and context-dependent types, our work may promote more research into this problem.
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
ACL (1)1
2022 Conditional Bilingual Mutual Information Based Adaptive Training for Neural Machine Translation
abstract
Songming Zhang, Yijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu, Jian Liu, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Songming Zhang 0001, Yijin Liu, Fandong Meng, Yufeng Chen 0005, Jin An Xu, Jian Liu 0032, Jie Zhou 0016
ACL (1)6
2022 Iterative Constrained Back-Translation for Unsupervised Domain Adaptation of Machine Translation
abstract
Back-translation has been proven to be effective in unsupervised domain adaptation of neural machine translation (NMT). However, the existing back-translation methods mainly improve domain adaptability by generating in-domain pseudo-parallel data that contains sentence-structural knowledge, paying less attention to the in-domain lexical knowledge, which may lead to poor translation of unseen in-domain words. In this paper, we propose an Iterative Constrained Back-Translation (ICBT) method to incorporate in-domain lexical knowledge on the basis of BT for unsupervised domain adaptation of NMT. Specifically, we apply lexical constraints into back-translation to generate pseudo-parallel data with in-domain lexical knowledge, and then perform round-trip iterations to incorporate more lexical knowledge. Based on this, we further explore sampling strategies of constrained words in ICBT to introduce more targeted lexical knowledge, via domain specificity and confidence estimation. Experimental results on four domains show that our approach achieves state-of-the-art results, improving the BLEU score by up to 3.08 compared to the strongest baseline, which demonstrates the effectiveness of our approach.
Hongxiao Zhang, Yufeng Chen 0005, Jin An Xu, Jian Liu 0032
COLING6
2022 Low-Resource NER by Data Augmentation With Prompting
abstract
Named entity recognition (NER) is a fundamental information extraction task that seeks to identify entity mentions of certain types in text. Despite numerous advances, the existing NER methods rely on extensive supervision for model training, which struggle in a low-resource scenario with limited training data. In this paper, we propose a new data augmentation method for low-resource NER, by eliciting knowledge from BERT with prompting strategies. Particularly, we devise a label-conditioned word replacement strategy that can produce more label-consistent examples by capturing the underlying word-label dependencies, and a prompting with question answering method to generate new training data from unlabeled texts. The experimental results have widely confirmed the effectiveness of our approach. Particularly, in a low-resource scenario with only 150 training sentences, our approach outperforms previous methods without data augmentation by over 40% in F1 and prior best data augmentation methods by over 2.0% in F1. Furthermore, our approach also fits with a zero-shot scenario, yielding promising results without using any human-labeled data for the task.
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
IJCAI1
2022 Multimedia Event Extraction From News With a Unified Contrastive Learning Framework
abstract
Extracting events from news have seen many benefits in downstream applications. Today's event extraction (EE) systems, however, usually focus on a single modality --- either for text or image, and such methods suffer from incomplete information because a news document is typically presented in a multimedia format. In this paper, we propose a new method for multimedia EE by bridging the textual and visual modalities with a unified contrastive learning framework. Our central idea is to create a shared space for texts and images in order to improve their similar representation. This is accomplished by training on text-image pairs in general, and we demonstrate that it is possible to use this framework to boost learning for one modality by investigating the complementary of the other modality. On the benchmark dataset, our approach establishes a new state-of-the-art performance and shows a 3 percent improvement in F1. Furthermore, we demonstrate that it can achieve cutting-edge performance for visual EE even in a zero-shot scenario with no annotated data in the visual modality.
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
ACM Multimedia1
2022 Document-level event argument linking as machine reading comprehension
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
Neurocomputing1
2022 Document-level event argument extraction with self-augmentation and a cross-domain joint training mechanism
Jian Liu 0032, Jin An Xu
Knowl. Based Syst.1
2022 Exploiting Morpheme and Cross-lingual Knowledge to Enhance Mongolian Named Entity Recognition
abstract
Mongolian named entity recognition (NER) is not only one of the most crucial and fundamental tasks in Mongolian natural language processing, but also an important step to improve the performance of downstream tasks such as information retrieval, machine translation, and dialog system. However, traditional Mongolian NER models heavily rely on the feature engineering. Even worse, the complex morphological structure of Mongolian words makes the data sparser. To alleviate the feature engineering and data sparsity in Mongolian named entity recognition, we propose a novel NER framework with Multi-Knowledge Enhancement (MKE-NER) . Specifically, we introduce both linguistic knowledge through Mongolian morpheme representation and cross-lingual knowledge from Mongolian-Chinese parallel corpus. Furthermore, we design two methods to exploit cross-lingual knowledge sufficiently, i.e., cross-lingual representation and cross-lingual annotation projection. Experimental results demonstrate the effectiveness of our MKE-NER model, which outperforms strong baselines and achieves the best performance (94.04% F1 score) on the traditional Mongolian benchmark. Particularly, extensive experiments with different data scales highlight the superiority of our method in low-resource scenarios.
Songming Zhang 0001, Ying Zhang 0084, Yufeng Chen 0005, Du Wu, Jin An Xu, Jian Liu 0032
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2022 MRCAug: Data Augmentation via Machine Reading Comprehension for Document-Level Event Argument Extraction
abstract
Document-level event argument extraction (EAE) is a critical event semantic understanding task that requires a model to identify an event's global arguments beyond the sentence level. Existing approaches to this problem are based on supervised learning, which require a large amount of labeled data for model training. However, due to the complicated structure of an event, human annotation for this task is costly, and the issue of inadequacy of training data has long hampered the study. In this study, we propose a novel approach to mitigating the data sparsity problem faced by document-level EAE, by linking the task with machine reading comprehension (MRC). Particularly, we devise two data augmentation regimes via MRC, including an implicit knowledge transfer method, which enables knowledge transfer from other tasks to the document-level EAE task, and an explicit data generation method, which can explicitly generate new training examples by treating a pre-trained MRC model as an annotator. Furthermore, we propose a self-training based noise reduction strategy that can effectively addresses the out-of-domain noise introduced by the data augmentation methods. The extensive assessments on three benchmarks have validated the effectiveness of our approach — it not only achieves state-of-the-art performance but also demonstrates superior results in the data-low scenario.
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Cross-Domain Slot Filling as Machine Reading Comprehension: A New Perspective
abstract
With intelligent dialogue systems becoming more and more important in our daily lives, slot filling, one of the most important components of an intelligent dialogue system, has gotten a lot of attention from academia and industry. Despite many advancements in the single-domain learning paradigm for slot filling, leveraging resources from different domains to boost learning for a target domain remains a challenge. In contrast to prior methods that supplemented a sequence labeling model with slot meta-information, we address cross-domain slot filling as a machine reading comprehension (MRC) problem for the first time, where the extraction of slot values is viewed as a question answering process. In the framework above, we present both static and dynamic question generating mechanisms, which have complimentary effects in diverse cross-domain contexts. Furthermore, we devise a dynamic question generation approach that can generate numerous values for a slot at the same time. Finally, we construct a pre-training and fine-tuning training approach that enables us to improve learning by utilizing MRC’s resources. We conducted extensive experiments on four datasets to evaluate our approach, and the experimental results clearly justified the advantages of our approach in various cross-domain settings.
Jian Liu 0032, Mengshi Yu, Yufeng Chen 0005, Jin An Xu
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Machine Reading Comprehension as Data Augmentation: A Case Study on Implicit Event Argument Extraction
abstract
Implicit event argument extraction (EAE) is a crucial document-level information extraction task that aims to identify event arguments beyond the sentence level.Despite many efforts for this task, the lack of enough training data has long impeded the study.In this paper, we take a new perspective to address the data sparsity issue faced by implicit EAE, by bridging the task with machine reading comprehension (MRC).Particularly, we devise two data augmentation regimes via MRC, including: 1) implicit knowledge transfer, which enables knowledge transfer from other tasks, by building a unified training framework in the MRC formulation, and 2) explicit data augmentation, which can explicitly generate new training examples, by treating MRC models as an annotator.The extensive experiments have justified the effectiveness of our approach -it not only obtains state-of-the-art performance on two benchmarks, but also demonstrates superior results in a data-low scenario.
Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
EMNLP (1)1
2021 Discourse-Level Event Temporal Ordering with Uncertainty-Guided Graph Completion
abstract
Learning to order events at discourse-level is a crucial text understanding task. Despite many efforts for this task, the current state-of-the-art methods rely heavily on manually designed features, which are costly to produce and are often specific to tasks/domains/datasets. In this paper, we propose a new graph perspective on the task, which does not require complex feature engineering but can assimilate global features and learn inter-dependencies effectively. Specifically, in our approach, each document is considered as a temporal graph, in which the nodes and edges represent events and event-event relations respectively. In this sense, the temporal ordering task corresponds to constructing edges for an empty graph. To train our model, we design a graph mask pre-training mechanism, which can learn inter-dependencies of temporal relations by learning to recover a masked edge following graph topology. In the testing stage, we design an certain-first strategy based on model uncertainty, which can decide the prediction orders and reduce the risk of error propagation. The experimental results demonstrate that our approach outperforms previous methods consistently and can meanwhile maintain good global consistency.
Jian Liu 0032, Jin An Xu, Yufeng Chen 0005
IJCAI1
2021 Improving Stylized Neural Machine Translation with Iterative Dual Knowledge Transfer
abstract
Stylized neural machine translation (NMT) aims to translate sentences of one style into sentences of another style, which is essential for the application of machine translation in a real-world scenario. However, a major challenge in this task is the scarcity of high-quality parallel data which is stylized paired. To address this problem, we propose an iterative dual knowledge transfer framework that utilizes informal training data of machine translation and formality style transfer data to create large-scale stylized paired data, for the training of stylized machine translation model. Specifically, we perform bidirectional knowledge transfer between translation model and text style transfer model iteratively through knowledge distillation. Then, we further propose a data-refinement module to process the noisy synthetic parallel data generated during knowledge transfer. Experiment results demonstrate the effectiveness of our method, achieving an improvement over the existing best model by 5 BLEU points on MTFC dataset. Meanwhile, extensive analyses illustrate our method can also improve the accuracy of formality style transfer.
Xuanxuan Wu, Jian Liu 0032, Xinjie Li 0003, Jin An Xu, Yufeng Chen 0005
IJCAI2
2021 Cross-Domain Slot Filling as Machine Reading Comprehension
abstract
With task-oriented dialogue systems being widely applied in everyday life, slot filling, the essential component of task-oriented dialogue systems, is required to be quickly adapted to new domains that contain domain-specific slots with few or no training data. Previous methods for slot filling usually adopt sequence labeling framework, which, however, often has limited ability when dealing with the domain-specific slots. In this paper, we take a new perspective on cross-domain slot filling by framing it as a machine reading comprehension (MRC) problem. Our approach firstly transforms slot names into well-designed queries, which contain rich informative prior knowledge and are very helpful for the detection of domain-specific slots. In addition, we utilize the large-scale MRC dataset for pre-training, which further alleviates the data scarcity problem. Experimental results on SNIPS and ATIS datasets show that our approach consistently outperforms the existing state-of-the-art methods by a large margin.
Mengshi Yu, Jian Liu 0032, Yufeng Chen 0005, Jin An Xu
IJCAI2
2021 Contrastive Learning for Machine Translation Quality Estimation
Hui Di, Jian Liu 0032, Yufeng Chen 0005, Kazushige Ouchi, Jin An Xu
NLPCC (1)3
2020 Graph-Based Knowledge Integration for Question Answering over Dialogue
abstract
Question answering over dialogue, a specialized machine reading comprehension task, aims to comprehend a dialogue and to answer specific questions.Despite many advances, existing approaches for this task did not consider dialogue structure and background knowledge (e.g., relationships between speakers).In this paper, we introduce a new approach for the task, featured by its novelty in structuring dialogue and integrating background knowledge for reasoning.Specifically, different from previous "structure-less" approaches, our method organizes a dialogue as a "relational graph", using edges to represent relationships between entities.To encode this relational graph, we devise a relational graph convolutional network (R-GCN), which can traverse the graph's topological structure and effectively encode multi-relational knowledge for reasoning.The extensive experiments have justified the effectiveness of our approach over competitive baselines.Moreover, a deeper analysis shows that our model is better at tackling complex questions requiring relational reasoning and defending adversarial attacks with distracting sentences.
Jian Liu 0032, Dianbo Sui, Kang Liu 0001, Jun Zhao 0001
COLING1
2020 Event Extraction as Machine Reading Comprehension
abstract
Event extraction (EE) is a crucial information extraction task that aims to extract event information in texts.Previous methods for EE typically model it as a classification task, which are data-hungry and suffer from the data scarcity problem.In this paper, we propose a new learning paradigm of EE, by explicitly casting it as a machine reading comprehension problem (MRC).Our approach includes an unsupervised question generation process, which can transfer event schema into a set of natural questions, followed by a BERTbased question-answering process to retrieve answers as EE results.This learning paradigm enables us to strengthen the reasoning process of EE, by introducing sophisticated models in MRC, and relieve the data scarcity problem, by introducing the large-scale datasets in MRC.The empirical results show that: i) our approach attains state-of-the-art performance by considerable margins over previous methods.ii) Our model is excelled in the data-scarce scenario, for example, obtaining 49.8% in F1 for event argument extraction with only 1% data, compared with 2.2% of the previous method.iii) Our model also fits with zero-shot scenarios, achieving 37.0% and 16% in F1 on two datasets without using any EE training data.
Jian Liu 0032, Yubo Chen 0001, Kang Liu 0001, Wei Bi, Xiaojiang Liu
EMNLP (1)1
2020 Knowledge Enhanced Event Causality Identification with Mention Masking Generalizations
abstract
Identifying causal relations of events is a crucial language understanding task. Despite many efforts for this task, existing methods lack the ability to adopt background knowledge, and they typically generalize poorly to new, previously unseen data. In this paper, we present a new method for event causality identification, aiming to address limitations of previous methods. On the one hand, our model can leverage external knowledge for reasoning, which can greatly enrich the representation of events; On the other hand, our model can mine event-agnostic, context-specific patterns, via a mechanism called event mention masking generalization, which can greatly enhance the ability of our model to handle new, previously unseen cases. In experiments, we evaluate our model on three benchmark datasets and show our model outperforms previous methods by a significant margin. Moreover, we perform 1) cross-topic adaptation, 2) exploiting unseen predicates, and 3) cross-task adaptation to evaluate the generalization ability of our model. Experimental results show that our model demonstrates a definite advantage over previous methods.
Jian Liu 0032, Yubo Chen 0001, Jun Zhao 0001
IJCAI1
2019 Exploiting the Ground-Truth: An Adversarial Imitation Based Knowledge Distillation Approach for Event Detection
abstract
The ambiguity in language expressions poses a great challenge for event detection. To disambiguate event types, current approaches rely on external NLP toolkits to build knowledge representations. Unfortunately, these approaches work in a pipeline paradigm and suffer from error propagation problem. In this paper, we propose an adversarial imitation based knowledge distillation approach, for the first time, to tackle the challenge of acquiring knowledge from rawsentences for event detection. In our approach, a teacher module is first devised to learn the knowledge representations from the ground-truth annotations. Then, we set up a student module that only takes the raw-sentences as the input. The student module is taught to imitate the behavior of the teacher under the guidance of an adversarial discriminator. By this way, the process of knowledge distillation from rawsentence has been implicitly integrated into the feature encoding stage of the student module. To the end, the enhanced student is used for event detection, which processes raw texts and requires no extra toolkits, naturally eliminating the error propagation problem faced by pipeline approaches. We conduct extensive experiments on the ACE 2005 datasets, and the experimental results justify the effectiveness of our approach.
Jian Liu 0032, Yubo Chen 0001, Kang Liu 0001
AAAI1
2019 Neural Cross-Lingual Event Detection with Minimal Parallel Resources
abstract
Jian Liu, Yubo Chen, Kang Liu, Jun Zhao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jian Liu 0032, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001
EMNLP/IJCNLP (1)1
2018 Event Detection via Gated Multilingual Attention Mechanism
abstract
Identifying event instance in text plays a critical role in building NLP applications such as Information Extraction (IE) system. However, most existing methods for this task focus only on monolingual clues of a specific language and ignore the massive information provided by other languages. Data scarcity and monolingual ambiguity hinder the performance of these monolingual approaches. In this paper, we propose a novel multilingual approach---dubbed as Gated Multilingual Attention (GMLATT) framework---to address the two issues simultaneously. In specific, to alleviate data scarcity problem, we exploit the consistent information in multilingual data via context attention mechanism. Which takes advantage of the consistent evidence in multilingual data other than learning only from monolingual data. To deal with monolingual ambiguity problem, we propose gated cross-lingual attention to exploit the complement information conveyed by multilingual data, which is helpful for the disambiguation. The cross-lingual attention gate serves as a sentinel modelling the confidence of the clues provided by other languages and controls the information integration of various languages. We have conducted extensive experiments on the ACE 2005 benchmark. Experimental results show that our approach significantly outperforms state-of-the-art methods.
Jian Liu 0032, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001
AAAI1