Yajuan Lyu

dblp:190/7920 · DBLP profile ↗
← Back
29ranked-venue papers
0as first author
17since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Inferential Knowledge-Enhanced Integrated Reasoning for Video Question Answering
abstract
Recently, video question answering has attracted growing attention. It involves answering a question based on a fine-grained understanding of video multi-modal information. Most existing methods have successfully explored the deep understanding of visual modality. We argue that a deep understanding of linguistic modality is also essential for answer reasoning, especially for videos that contain character dialogues. To this end, we propose an Inferential Knowledge-Enhanced Integrated Reasoning method. Our method consists of two main components: 1) an Inferential Knowledge Reasoner to generate inferential knowledge for linguistic modality inputs that reveals deeper semantics, including the implicit causes, effects, mental states, etc. 2) an Integrated Reasoning Mechanism to enhance video content understanding and answer reasoning by leveraging the generated inferential knowledge. Experimental results show that our method achieves significant improvement on two mainstream datasets. The ablation study further demonstrates the effectiveness of each component of our approach.
Jianguo Mao, Wenbin Jiang 0002, Hong Liu 0007, Yajuan Lyu
AAAI5
2023 WeCheck: Strong Factual Consistency Checker via Weakly Supervised Learning
abstract
A crucial issue of current text generation models is that they often uncontrollably generate text that is factually inconsistent with inputs.Due to lack of annotated data, existing factual consistency metrics usually train evaluation models on synthetic texts or directly transfer from other related tasks, such as question answering (QA) and natural language inference (NLI).Bias in synthetic text or upstream tasks makes them perform poorly on text actually generated by language models, especially for general evaluation for various tasks.To alleviate this problem, we propose a weakly supervised framework named WeCheck that is directly trained on actual generated samples from language models with weakly annotated labels.WeCheck first utilizes a generative model to infer the factual labels of generated samples by aggregating weak labels from multiple resources.Next, we train a simple noise-aware classification model as the target metric using the inferred weakly supervised information.Comprehensive experiments on various tasks demonstrate the strong performance of WeCheck, achieving an average absolute improvement of 3.3% on the TRUE benchmark over 11B state-of-the-art methods using only 435M parameters.Furthermore, it is up to 30× faster than previous evaluation methods, greatly improving the accuracy and efficiency of factual consistency evaluation. 1
Wei Li 0176, Xinyan Xiao, Sujian Li, Yajuan Lyu
ACL (1)6
2023 S2ynRE: Two-stage Self-training with Synthetic data for Low-resource Relation Extraction
abstract
Current relation extraction methods suffer from the inadequacy of large-scale annotated data.While distant supervision alleviates the problem of data quantities, there still exists domain disparity in data qualities due to its reliance on domain-restrained knowledge bases. In this work, we propose S2ynRE, a framework of two-stage Self-training with Synthetic data for Relation Extraction.We first leverage the capability of large language models to adapt to the target domain and automatically synthesize large quantities of coherent, realistic training data.We then propose an accompanied two-stage self-training algorithm that iteratively and alternately learns from synthetic and golden data together.We conduct comprehensive experiments and detailed ablations on popular relation extraction datasets to demonstrate the effectiveness of the proposed framework.
Benfeng Xu, Quan Wang 0002, Yajuan Lyu, Dai Dai, Yongdong Zhang 0001, Zhendong Mao 0001
ACL (1)3
2023 IM-TQA: A Chinese Table Question Answering Dataset with Implicit and Multi-type Table Structures
abstract
Mingyu Zheng, Yang Hao, Wenbin Jiang, Zheng Lin, Yajuan Lyu, QiaoQiao She, Weiping Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Mingyu Zheng, Wenbin Jiang 0002, Zheng Lin 0001, Yajuan Lyu, Qiaoqiao She, Weiping Wang 0005
ACL (1)5
2023 Semantic-Driven Instance Generation for Table Question Answering
Wenbin Jiang 0002, Xiang Ao 0001, Xinwei Feng, Yajuan Lyu, Qiaoqiao She, Qing He 0003
DASFAA (1)6
2023 $k$NN Prompting: Beyond-Context Learning with Calibration-Free Nearest Neighbor Inference
Benfeng Xu, Quan Wang 0002, Zhendong Mao 0001, Yajuan Lyu, Qiaoqiao She, Yongdong Zhang 0001
ICLR4
2023 Neural Knowledge Bank for Pretrained Transformers
Damai Dai, Wenbin Jiang 0002, Qingxiu Dong, Yajuan Lyu, Zhifang Sui
NLPCC (2)4
2023 Mixture-of-Experts for Biomedical Question Answering
Damai Dai, Wenbin Jiang 0002, Yajuan Lyu, Zhifang Sui, Baobao Chang
NLPCC (1)4
2023 FactGen: Faithful Text Generation by Factuality-aware Pre-training and Contrastive Ranking Fine-tuning
abstract
Conditional text generation is supposed to generate a fluent and coherent target text that is faithful to the source text. Although pre-trained models have achieved promising results, they still suffer from the crucial factuality problem. To deal with this issue, we propose a factuality-aware pretraining-finetuning framework named FactGen, which fully considers factuality during two training stages. Specifically, at the pre-training stage, we utilize a natural language inference model to construct target texts that are entailed by the source texts, resulting in a more factually consistent pre-training objective. Then, during the fine-tuning stage, we further introduce a contrastive ranking loss to encourage the model to generate factually consistent text with higher probability. Extensive experiments on three conditional text generation tasks demonstrate the effectiveness and generality of our training framework.
Zhibin Lan, Wei Li 0176, Jinsong Su, Xinyan Xiao, Yajuan Lyu
J. Artif. Intell. Res.7
2022 Hierarchical Representation-based Dynamic Reasoning Network for Biomedical Question Answering
abstract
Recently, Biomedical Question Answering (BQA) has attracted growing attention due to its application value and technical challenges. Most existing works treat it as a semantic matching task that predicts answers by computing confidence among questions, options and evidence sentences, which is insufficient for scenarios that require complex reasoning based on a deep understanding of biomedical evidences. We propose a novel model termed Hierarchical Representation-based Dynamic Reasoning Network (HDRN) to tackle this problem. It first constructs the hierarchical representations for biomedical evidences to learn semantics within and among evidences. It then performs dynamic reasoning based on the hierarchical representations of evidences to solve complex biomedical problems. Against the existing state-of-the-art model, the proposed model significantly improves more than 4.5%, 3% and 1.3% on three mainstream BQA datasets, PubMedQA, MedQA-USMLE and NLPEC. The ablation study demonstrates the superiority of each improvement of our model. The code will be released after the paper is published.
Jianguo Mao, Zengfeng Zeng, Weihua Peng, Wenbin Jiang 0002, Hong Liu 0007, Yajuan Lyu
COLING8
2022 A Transition-based Method for Complex Question Understanding
abstract
Complex Question Understanding (CQU) parses complex questions to Question Decomposition Meaning Representation (QDMR) which is a sequence of atomic operators. Existing works are based on end-to-end neural models which do not explicitly model the intermediate states and lack interpretability for the parsing process. Besides, they predict QDMR in a mismatched granularity and do not model the step-wise information which is an essential characteristic of QDMR. To alleviate the issues, we treat QDMR as a computational graph and propose a transition-based method where a decider predicts a sequence of actions to build the graph node-by-node. In this way, the partial graph at each step enables better representation of the intermediate states and better interpretability. At each step, the decider encodes the intermediate state with specially designed encoders and predicts several candidates of the next action and its confidence. For inference, a searcher seeks the optimal graph based on the predictions of the decider to alleviate the error propagation. Experimental results demonstrate the parsing accuracy of our method against several strong baselines. Moreover, our method has transparent and human-readable intermediate results, showing improved interpretability.
Wenbin Jiang 0002, Yajuan Lyu, Sujian Li
COLING3
2022 Explainable Question Answering based on Semantic Graph by Global Differentiable Learning and Dynamic Adaptive Reasoning
abstract
Multi-hop Question Answering is an agent task for testing the reasoning ability.With the development of pre-trained models, the implicit reasoning ability has been surprisingly improved and can even surpass human performance.However, the nature of the black box hinders the construction of explainable intelligent systems.Several researchers have explored explainable neural-symbolic reasoning methods based on question decomposition techniques.The undifferentiable symbolic operations and the error propagation in the reasoning process lead to poor performance.To alleviate it, we propose a simple yet effective Global Differentiable Learning strategy to explore optimal reasoning paths from the latent probability space so that the model learns to solve intermediate reasoning processes without expert annotations.We further design a Dynamic Adaptive Reasoner to enhance the generalization of unseen questions.Our method achieves 17% improvements in F1-score against BreakRC and shows better interpretability.We take a step forward in building interpretable reasoning methods.
Jianguo Mao, Wenbin Jiang 0002, Hong Liu 0007, Yajuan Lyu, Qiaoqiao She
EMNLP6
2022 Precisely the Point: Adversarial Augmentations for Faithful and Informative Text Generation
abstract
Though model robustness has been extensively studied in language understanding, the robustness of Seq2Seq generation remains understudied.In this paper, we conduct the first quantitative analysis on the robustness of pre-trained Seq2Seq models.We find that even current SOTA pre-trained Seq2Seq model (BART) is still vulnerable, which leads to significant degeneration in faithfulness and informativeness for text generation tasks.This motivated us to further propose a novel adversarial augmentation framework, namely AdvSeq, for generally improving faithfulness and informativeness of Seq2Seq models via enhancing their robustness.AdvSeq automatically constructs two types of adversarial augmentations during training, including implicit adversarial samples by perturbing word representations and explicit adversarial samples by word swapping, both of which effectively improve Seq2Seq robustness.Extensive experiments on three popular text generation tasks demonstrate that AdvSeq significantly improves both the faithfulness and informativeness of Seq2Seq generation under both automatic and human evaluation settings.
Wei Li 0176, Xinyan Xiao, Sujian Li, Yajuan Lyu
EMNLP6
2022 CLOP: Video-and-Language Pre-Training with Knowledge Regularizations
abstract
Video-and-language pre-training has shown promising results for learning generalizable representations. Most existing approaches usually model video and text in an implicit manner, without considering explicit structural representations of the multi-modal content. We denote such form of representations as structural knowledge, which express rich semantics of multiple granularities. There are related works that propose object-aware approaches to inject similar knowledge as inputs. However, the existing methods usually fail to effectively utilize such knowledge as regularizations to shape a superior cross-modal representation space. To this end, we propose a Cross-modaL knOwledge-enhanced Pre-training (CLOP) method with Knowledge Regularizations. There are two key designs of ours: 1) a simple yet effective Structural Knowledge Prediction (SKP) task to pull together the latent representations of similar videos; and 2) a novel Knowledge-guided sampling approach for Contrastive Learning (KCL) to push apart cross-modal hard negative samples. We evaluate our method on four text-video retrieval tasks and one multi-choice QA task. The experiments show clear improvements, outperforming prior works by a substantial margin. Besides, we provide ablations and insights of how our methods affect the latent representation space, demonstrating the value of incorporating knowledge regularizations into video-and-language pre-training.
Guohao Li 0002, Zhifan Feng, Yajuan Lyu, Hua Wu 0003, Haifeng Wang 0001
ACM Multimedia5
2022 Dynamic Multistep Reasoning based on Video Scene Graph for Video Question Answering
abstract
Jianguo Mao, Wenbin Jiang, Xiangdong Wang, Zhifan Feng, Yajuan Lyu, Hong Liu, Yong Zhu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jianguo Mao, Wenbin Jiang 0002, Zhifan Feng, Yajuan Lyu, Hong Liu 0007, Yong Zhu 0004
NAACL-HLT5
2022 EmRel: Joint Representation of Entities and Embedded Relations for Multi-triple Extraction
abstract
Benfeng Xu, Quan Wang, Yajuan Lyu, Yabing Shi, Yong Zhu, Jie Gao, Zhendong Mao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Benfeng Xu, Quan Wang 0002, Yajuan Lyu, Yabing Shi, Yong Zhu 0004, Zhendong Mao 0001
NAACL-HLT3
2021 Entity Structure Within and Throughout: Modeling Mention Dependencies for Document-Level Relation Extraction
abstract
Entities, as the essential elements in relation extraction tasks, exhibit certain structure. In this work, we formulate such entity structure as distinctive dependencies between mention pairs. We then propose SSAN, which incorporates these structural dependencies within the standard self-attention mechanism and throughout the overall encoding stage. Specifically, we design two alternative transformation modules inside each self-attention building block to produce attentive biases so as to adaptively regularize its attention flow. Our experiments demonstrate the usefulness of the proposed entity structure and the effectiveness of SSAN. It significantly outperforms competitive baselines, achieving new state-of-the-art results on three popular document-level relation extraction datasets. We further provide ablation and visualization to show how the entity structure guides the model for better relation extraction. Our code is publicly available.
Benfeng Xu, Quan Wang 0002, Yajuan Lyu, Yong Zhu 0004, Zhendong Mao 0001
AAAI3
2020 Capturing Sentence Relations for Answer Sentence Selection with Multi-Perspective Graph Encoding
abstract
This paper focuses on the answer sentence selection task. Unlike previous work, which only models the relation between the question and each candidate sentence, we propose Multi-Perspective Graph Encoder (MPGE) to take the relations among the candidate sentences into account and capture the relations from multiple perspectives. By utilizing MPGE as a module, we construct two answer sentence selection models which are based on traditional representation and pre-trained representation, respectively. We conduct extensive experiments on two datasets, WikiQA and SQuAD. The results show that the proposed MPGE is effective for both types of representation. Moreover, the overall performance of our proposed model surpasses the state-of-the-art on both datasets. Additionally, we further validate the robustness of our method by the adversarial examples of AddSent and AddOneSent.
Zhixing Tian, Yuanzhe Zhang, Xinwei Feng, Wenbin Jiang 0002, Yajuan Lyu, Kang Liu 0001, Jun Zhao 0001
AAAI5
2020 DuEE: A Large-Scale Dataset for Chinese Event Extraction in Real-World Scenarios
Fayuan Li, Yuguang Chen, Weihua Peng, Quan Wang 0002, Yajuan Lyu, Yong Zhu 0004
NLPCC (2)7
2019 Joint Extraction of Entities and Overlapping Relations Using Position-Attentive Sequence Labeling
abstract
Joint entity and relation extraction is to detect entity and relation using a single model. In this paper, we present a novel unified joint extraction model which directly tags entity and relation labels according to a query word position p, i.e., detecting entity at p, and identifying entities at other positions that have relationship with the former. To this end, we first design a tagging scheme to generate n tag sequences for an n-word sentence. Then a position-attention mechanism is introduced to produce different sentence representations for every query position to model these n tag sequences. In this way, our method can simultaneously extract all entities and their type, as well as all overlapping relations. Experiment results show that our framework performances significantly better on extracting overlapping relations as well as detecting long-range relation, and thus we achieve state-of-the-art performance on two public datasets.
Dai Dai, Xinyan Xiao, Yajuan Lyu, Shan Dou, Qiaoqiao She, Haifeng Wang 0001
AAAI3
2019 Enhancing Pre-Trained Language Representations with Rich Knowledge for Machine Reading Comprehension
abstract
Machine reading comprehension (MRC) is a crucial and challenging task in NLP.Recently, pre-trained language models (LMs), especially BERT, have achieved remarkable success, presenting new state-of-the-art results in MRC.In this work, we investigate the potential of leveraging external knowledge bases (KBs) to further improve BERT for MRC.We introduce KT-NET, which employs an attention mechanism to adaptively select desired knowledge from KBs, and then fuses selected knowledge with BERT to enable context-and knowledgeaware predictions.We believe this would combine the merits of both deep LMs and curated KBs towards better MRC.Experimental results indicate that KT-NET offers significant and consistent improvements over BERT, outperforming competitive baselines on ReCoRD and SQuAD1.1 benchmarks.Notably, it ranks the 1st place on the ReCoRD leaderboard, and is also the best single model on the SQuAD1.1 leaderboard at the time of submission (March 4th, 2019). 1
An Yang, Quan Wang 0002, Jing Liu 0022, Kai Liu 0023, Yajuan Lyu, Hua Wu 0003, Qiaoqiao She, Sujian Li
ACL (1)5
2019 Machine Reading Comprehension Using Structural Knowledge Graph-aware Network
abstract
Delai Qiu, Yuanzhe Zhang, Xinwei Feng, Xiangwen Liao, Wenbin Jiang, Yajuan Lyu, Kang Liu, Jun Zhao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Delai Qiu, Yuanzhe Zhang, Xinwei Feng, Xiangwen Liao, Wenbin Jiang 0002, Yajuan Lyu, Kang Liu 0001, Jun Zhao 0001
EMNLP/IJCNLP (1)6
2019 DuIE: A Large-Scale Chinese Dataset for Information Extraction
Shuangjie Li, Yabing Shi, Wenbin Jiang 0002, Haijin Liang, Yajuan Lyu, Yong Zhu 0004
NLPCC (2)8
2019 An Overview of the 2019 Language and Intelligence Challenge
Quan Wang 0002, Wenquan Wu, Yabing Shi, Wei He 0014, Ying Chen 0011, Yajuan Lyu, Hua Wu 0003
NLPCC (2)9
2018 Multi-Passage Machine Reading Comprehension with Cross-Passage Answer Verification
abstract
Yizhong Wang, Kai Liu, Jing Liu, Wei He, Yajuan Lyu, Hua Wu, Sujian Li, Haifeng Wang. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Yizhong Wang, Kai Liu 0023, Jing Liu 0022, Wei He 0014, Yajuan Lyu, Hua Wu 0003, Sujian Li, Haifeng Wang 0001
ACL (1)5
2018 Joint Training of Candidate Extraction and Answer Selection for Reading Comprehension
abstract
While sophisticated neural-based techniques have been developed in reading comprehension, most approaches model the answer in an independent manner, ignoring its relations with other answer candidates.This problem can be even worse in open-domain scenarios, where candidates from multiple passages should be combined to answer a single question.In this paper, we formulate reading comprehension as an extract-then-select twostage procedure.We first extract answer candidates from passages, then select the final answer by combining information from all the candidates.Furthermore, we regard candidate extraction as a latent variable and train the two-stage process jointly with reinforcement learning.As a result, our approach has improved the state-ofthe-art performance significantly on two challenging open-domain reading comprehension datasets.Further analysis demonstrates the effectiveness of our model components, especially the information fusion of all the candidates and the joint training of the extract-then-select procedure.
Xinyan Xiao, Yajuan Lyu
ACL (1)4
2018 Improving Neural Abstractive Document Summarization with Explicit Information Selection Modeling
abstract
Information selection is the most important component in document summarization task.In this paper, we propose to extend the basic neural encoding-decoding framework with an information selection layer to explicitly model and optimize the information selection process in abstractive document summarization.Specifically, our information selection layer consists of two parts: gated global information filtering and local sentence selection.Unnecessary information in the original document is first globally filtered, then salient sentences are selected locally while generating each summary sentence sequentially.To optimize the information selection process directly, distantly-supervised training guided by the golden summary is also imported.Experimental results demonstrate that the explicit modeling and optimizing of the information selection process improves document summarization performance significantly, which enables our model to generate more informative and concise summaries, and thus significantly outperform state-of-the-art neural abstractive methods.
Wei Li 0176, Xinyan Xiao, Yajuan Lyu, Yuanzhuo Wang
EMNLP3
2018 Improving Neural Abstractive Document Summarization with Structural Regularization
abstract
Recent neural sequence-to-sequence models have shown significant progress on short text summarization.However, for document summarization, they fail to capture the longterm structure of both documents and multisentence summaries, resulting in information loss and repetitions.In this paper, we propose to leverage the structural information of both documents and multi-sentence summaries to improve the document summarization performance.Specifically, we import both structural-compression and structuralcoverage regularization into the summarization process in order to capture the information compression and information coverage properties, which are the two most important structural properties of document summarization.Experimental results demonstrate that the structural regularization improves the document summarization performance significantly, which enables our model to generate more informative and concise summaries, and thus significantly outperforms state-of-the-art neural abstractive methods.
Wei Li 0176, Xinyan Xiao, Yajuan Lyu, Yuanzhuo Wang
EMNLP3
2018 Answer-focused and Position-aware Neural Question Generation
abstract
In this paper, we focus on the problem of question generation (QG).Recent neural networkbased approaches employ the sequence-tosequence model which takes an answer and its context as input and generates a relevant question as output.However, we observe two major issues with these approaches: (1) The generated interrogative words (or question words) do not match the answer type.(2) The model copies the context words that are far from and irrelevant to the answer, instead of the words that are close and relevant to the answer.To address these two issues, we propose an answer-focused and position-aware neural question generation model.(1) By answerfocused, we mean that we explicitly model question word generation by incorporating the answer embedding, which can help generate an interrogative word matching the answer type.(2) By position-aware, we mean that we model the relative distance between the context words and the answer.Hence the model can be aware of the position of the context words when copying them to generate a question.We conduct extensive experiments to examine the effectiveness of our model.The experimental results show that our model significantly improves the baseline and outperforms the state-of-the-art system.
Xingwu Sun, Jing Liu 0022, Yajuan Lyu, Wei He 0014, Yanjun Ma
EMNLP3