Wenlin Yao

dblp:203/8711 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0002-4502-0350ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 5 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMs
abstract
Large language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation framework to assess LLMs’ proficiency in following instructions on diverse real-world tasks. We construct a hierarchical task tree encompassing seven major areas covering over 200 categories and over 800 tasks, which covers diverse capabilities such as question answering, reasoning, multi-turn dialogue, and text generation, to evaluate LLMs in a comprehensive and in-depth manner. We also design detailed evaluation standards and processes to facilitate consistent, unbiased judgments from human evaluators. A test set of over 3,000 instances is released, spanning different difficulty levels and knowledge domains. Our work provides a standardized methodology to evaluate human alignment in LLMs for both English and Chinese. We also analyze the feasibility of automating parts of evaluation with a strong LLM (GPT-4). Our framework supports a thorough assessment of LLMs as they are integrated into real-world applications. We have made publicly available the task tree, TencentLLMEval dataset, and evaluation methodology which have been demonstrated as effective in assessing the performance of Tencent Hunyuan LLMs. By doing so, we aim to facilitate the benchmarking of advances in the development of safe and human-aligned LLMs.
Shuyi Xie, Wenlin Yao, Yong Dai 0001, Zishan Xu, Fan Lin, Donglin Zhou, Lifeng Jin, Xinhua Feng, Pengzhi Wei, Zhichao Hu, Dong Yu 0001, Zhengyou Zhang
ACM Trans. Intell. Syst. Technol.2
2026 Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
abstract
Explainable AI (XAI) refers to techniques that provide human-understandable insights into the workings of AI models. Recently, the focus of XAI has been extended toward explaining Large Language Models (LLMs). This extension calls for a significant transformation in the XAI methodologies for two reasons. First, many existing XAI methods cannot be directly applied to LLMs due to their complexity and advanced capabilities. Second, as LLMs are increasingly deployed in diverse applications, the role of XAI shifts from merely opening the “black box” to actively enhancing the productivity and applicability of LLMs in real-world settings. Meanwhile, the conversation and generation abilities of LLMs can reciprocally enhance XAI. Therefore, in this article, we introduce Usable XAI in the context of LLMs by analyzing (1) how XAI can explain and improve LLM-based AI systems and (2) how XAI techniques can be improved by using LLMs. We introduce 10 strategies, introducing the key techniques for each and discussing their associated challenges. We also provide case studies to demonstrate how to obtain and leverage explanations.
Xuansheng Wu, Haiyan Zhao 0003, Yaochen Zhu, Fan Yang 0023, Lijie Hu, Tianming Liu 0001, Xiaoming Zhai, Wenlin Yao, Jundong Li, Mengnan Du, Ninghao Liu 0001
ACM Trans. Knowl. Discov. Data9
2025 OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
abstract
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Hongming Zhang, Tianqing Fang, Zhenzhong Lan, Dong Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Tianqing Fang, Zhen-Zhong Lan, Dong Yu 0001
ACL (1)2
2025 WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
abstract
Zhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang, Qin Lu, Liang Qiu, Changlong Yu, Puyang Xu, Chao Zhang, Bing Yin, Hyokun Yun, Lihong Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zhepei Wei, Wenlin Yao, Changlong Yu, Puyang Xu, Chao Zhang 0014, Hyokun Yun, Lihong Li 0001
EMNLP2
2025 DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
abstract
Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) have demonstrated impressive language/vision reasoning abilities, igniting the recent trend of building agents for targeted applications such as shopping assistants or AI software engineers. Recently, many data science benchmarks have been proposed to investigate their performance in the data science domain. However, existing data science benchmarks still fall short when compared to real-world data science applications due to their simplified settings. To bridge this gap, we introduce DSBench, a comprehensive benchmark designed to evaluate data science agents with realistic tasks. This benchmark includes 466 data analysis tasks and 74 data modeling tasks, sourced from Eloquence and Kaggle competitions. DSBench offers a realistic setting by encompassing long contexts, multimodal task backgrounds, reasoning with large data files and multi-table structures, and performing end-to-end data modeling tasks. Our evaluation of state-of-the-art LLMs, LVLMs, and agents shows that they struggle with most tasks, with the best agent solving only 34.12% of data analysis tasks and achieving a 34.74% Relative Performance Gap (RPG). These findings underscore the need for further advancements in developing more practical, intelligent, and autonomous data science agents.
Liqiang Jing, Zhehui Huang, Xiaoyang Wang 0001, Wenlin Yao, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Xinya Du, Dong Yu 0001
ICLR4
2025 DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
abstract
Enhancing the capability of large language models (LLMs) in reasoning has gained significant attention in recent years. Previous studies have demonstrated the effectiveness of various prompting strategies in aiding LLMs in reasoning (called "reasoning actions"), such as step-by-step thinking, reflecting before answering, solving with programs, and their combinations. However, these approaches often applied static, predefined reasoning actions uniformly to all questions, without considering the specific characteristics of each question or the capability of the task-solving LLM. In this paper, we propose DOTS, an approach enabling LLMs to reason Dynamically via Optimal reasoning Trajectories Search, tailored to the specific characteristics of each question and the inherent capability of the task-solving LLM. Our approach involves three key steps: i) defining atomic reasoning action modules that can be composed into various reasoning action trajectories; ii) searching for the optimal action trajectory for each training question through iterative exploration and evaluation for the specific task-solving LLM; and iii) using the collected optimal trajectories to train an LLM to plan for the reasoning trajectories of unseen questions. In particular, we propose two learning paradigms, i.e., fine-tuning an external LLM as a planner to guide the task-solving LLM, or directly fine-tuning the task-solving LLM with an internalized capability for reasoning actions planning. Our experiments across eight reasoning tasks show that our method consistently outperforms static reasoning techniques and the vanilla instruction tuning approach. Further analysis reveals that our method enables LLMs to adjust their computation based on problem complexity, allocating deeper thinking and reasoning to harder problems.
Murong Yue, Wenlin Yao, Haitao Mi, Dian Yu 0001, Ziyu Yao 0002, Dong Yu 0001
ICLR2
2025 Ask a Strong LLM Judge when Your Reward Model is Uncertain
abstract
Reward model (RM) plays a pivotal role in reinforcement learning with human feedback (RLHF) for aligning large language models (LLMs). However, classical RMs trained on human preferences are vulnerable to reward hacking and generalize poorly to out-of-distribution (OOD) inputs. By contrast, strong LLM judges equipped with reasoning capabilities demonstrate superior generalization, even without additional training, but incur significantly higher inference costs, limiting their applicability in online RLHF. In this work, we propose an uncertainty-based routing framework that efficiently complements a fast RM with a strong but costly LLM judge. Our approach formulates advantage estimation in policy gradient (PG) methods as pairwise preference classification, enabling principled uncertainty quantification to guide routing. Uncertain pairs are forwarded to the LLM judge, while confident ones are evaluated by the RM. Experiments on RM benchmarks demonstrate that our uncertainty-based routing strategy significantly outperforms random judge calling at the same cost, and downstream alignment results showcase its effectiveness in improving online RLHF.
Zhenghao Xu, Qingru Zhang, Ilgee Hong, Changlong Yu, Wenlin Yao, Haoming Jiang, Lihong Li 0001, Hyokun Yun, Tuo Zhao
NeurIPS7
2024 WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
abstract
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, Dong Yu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Yong Dai 0001, Hongming Zhang 0009, Zhen-Zhong Lan, Dong Yu 0001
ACL (1)2
2024 MinT: Boosting Generalization in Mathematical Reasoning via Multi-view Fine-tuning
abstract
Reasoning in mathematical domains remains a significant challenge for relatively small language models (LMs). Many current methods focus on specializing LMs in mathematical reasoning and rely heavily on distilling knowledge from powerful yet inefficient large LMs (LLMs). In this work, we explore a new direction that avoids over-reliance on LLM teachers, introducing a multi-view fine-tuning method that efficiently exploits existing mathematical problem datasets with diverse annotation styles. Our approach uniquely considers the various annotation formats as different “views” that may help each other and leverage them in training the model. By postpending distinct instructions to input questions, models can learn to generate solutions in diverse formats in a flexible manner. Experimental results show that our strategy enables relatively small LMs to outperform prior approaches that heavily rely on knowledge distillation, as well as carefully established baselines. Additionally, the proposed method grants the models promising generalization ability across various views and datasets, and the capability to learn from inaccurate or incomplete noisy data. We hope our multi-view training paradigm could inspire future studies in other machine reasoning domains.
Zhenwen Liang, Dian Yu 0001, Xiaoman Pan, Wenlin Yao, Xiangliang Zhang 0001, Dong Yu 0001
LREC/COLING4
2024 When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives
abstract
Reasoning is most powerful when an LLM accurately aggregates relevant information.We examine the critical role of information aggregation in reasoning by requiring the LLM to analyze sports narratives.To succeed at this task, an LLM must infer points from actions, identify related entities, attribute points accurately to players and teams, and compile key statistics to draw conclusions.We conduct comprehensive experiments with real NBA basketball data and present SPORTSGEN, a new method to synthesize game narratives.By synthesizing data, we can rigorously evaluate LLMs' reasoning capabilities under complex scenarios with varying narrative lengths and density of information.Our findings show that most models, including GPT-4o, often fail to accurately aggregate basketball scores due to frequent scoring patterns.Open-source models like Llama-3 further suffer from significant score hallucinations.Finally, the effectiveness of reasoning is influenced by narrative complexity, information density, and domain-specific terms, highlighting the challenges in analytical reasoning tasks. 1 * Work done during Yebowen Hu's internship; Kaiqiang Song and Sangwoo Cho were full-time researchers at Tencent AI Lab, Seattle, USA at the time of this work. https://github.com/YebowenHu/SportsGenAnalyze the team-player affiliations and play-by-play descriptions below to determine the total points scored by each team (player).Please explain your reasoning step by step and provide the final results in JSON format.Start with: {New York Knicks: 0, Denver Nuggets: 0} ({Andrea Bargnani: 0, Timofey Mozgov: 0, ...
Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Wenlin Yao, Hassan Foroosh, Dong Yu 0001, Fei Liu 0004
EMNLP5
2024 MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
abstract
Fuxiao Liu, Xiaoyang Wang, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Fuxiao Liu, Xiaoyang Wang 0001, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, Dong Yu 0001
NAACL-HLT3
2024 From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
abstract
Xuansheng Wu, Wenlin Yao, Jianshu Chen, Xiaoman Pan, Xiaoyang Wang, Ninghao Liu, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xuansheng Wu, Wenlin Yao, Jianshu Chen, Xiaoman Pan, Xiaoyang Wang 0001, Ninghao Liu 0001, Dong Yu 0001
NAACL-HLT2
2024 IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation
abstract
As Large Language Models (LLMs) become more capable of handling increasingly complex tasks, the evaluation set must keep pace with these advancements to ensure it remains sufficiently discriminative. Item Discrimination (ID) theory, which is widely used in educational assessment, measures the ability of individual test items to differentiate between high and low performers. Inspired by this theory, we propose an ID-induced prompt synthesis framework for evaluating LLMs so that the evaluation set continually updates and refines according to model abilities. Our data synthesis framework prioritizes both breadth and specificity. It can generate prompts that comprehensively evaluate the capabilities of LLMs while revealing meaningful performance differences between models, allowing for effective discrimination of their relative strengths and weaknesses across various tasks and domains. To produce high-quality data, we incorporate a self-correct mechanism into our generalization framework and develop two models to predict prompt discrimination and difficulty score to facilitate our data synthesis framework, contributing valuable tools to evaluation data synthesis research. We apply our generated data to evaluate five SOTA models. Our data achieves an average score of 51.92, accompanied by a variance of 10.06. By contrast, previous works (i.e., SELF-INSTRUCT and WizardLM) obtain an average score exceeding 67, with a variance below 3.2. The results demonstrate that the data generated by our framework is more challenging and discriminative compared to previous works. We will release a dataset of over 3,000 carefully crafted prompts to facilitate evaluation research of LLMs.
Fan Lin, Shuyi Xie, Yong Dai 0001, Wenlin Yao, Tianjiao Lang, Yu Zhang 0004
NeurIPS4
2024 Could Small Language Models Serve as Recommenders? Towards Data-centric Cold-start Recommendation
Xuansheng Wu, Huachi Zhou, Wenlin Yao, Xiao Huang 0001, Ninghao Liu 0001
WWW4
2024 DIRECT: Dual Interpretable Recommendation with Multi-aspect Word Attribution
abstract
Recommending products to users with intuitive explanations helps improve the system in transparency, persuasiveness, and satisfaction. Existing interpretation techniques include post hoc methods and interpretable modeling. The former category could quantitatively analyze input contribution to model prediction but has limited interpretation faithfulness, while the latter could explain model internal mechanisms but may not directly attribute model predictions to input features. In this study, we propose a novel Dual Interpretable Recommendation model called DIRECT, which integrates ideas of the two interpretation categories to inherit their advantages and avoid limitations. Specifically, DIRECT makes use of item descriptions as explainable evidence for recommendation. First, similar to the post hoc interpretation, DIRECT could attribute the prediction of a user preference score to textual words of the item descriptions. The attribution of each word is related to its sentiment polarity and word importance, where a word is important if it corresponds to an item aspect that the user is interested in. Second, to improve the interpretability of embedding space, we propose to extract high-level concepts from embeddings, where each concept corresponds to an item aspect. To learn discriminative concepts, we employ a concept bottleneck layer and maximize the coding rate reduction on word-aspect embeddings by leveraging a word–word affinity graph extracted from a pre-trained language model. In this way, DIRECT simultaneously achieves faithful attribution and usable interpretation of embedding space. We also show that DIRECT achieves linear inference time complexity regarding the length of item reviews. We conduct experiments including ablation studies on five real-world datasets. Quantitative analysis, visualizations, and case studies verify the interpretability of DIRECT. Our code is available at: https://github.com/JacksonWuxs/DIRECT .
Xuansheng Wu, Hanqin Wan, Qiaoyu Tan, Wenlin Yao, Ninghao Liu 0001
ACM Trans. Intell. Syst. Technol.4
2023 How do Words Contribute to Sentence Semantics? Revisiting Sentence Embeddings with a Perturbation Method
abstract
Wenlin Yao, Lifeng Jin, Hongming Zhang, Xiaoman Pan, Kaiqiang Song, Dian Yu, Dong Yu, Jianshu Chen. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Wenlin Yao, Lifeng Jin, Hongming Zhang 0009, Xiaoman Pan, Kaiqiang Song, Dian Yu 0001, Dong Yu 0001, Jianshu Chen
EACL1
2023 Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional Operations
abstract
Traditional sentence embedding models encode sentences into vector representations to capture useful properties such as the semantic similarity between sentences.However, in addition to similarity, sentence semantics can also be interpreted via compositional operations such as sentence fusion or difference.It is unclear whether the compositional semantics of sentences can be directly reflected as compositional operations in the embedding space.To more effectively bridge the continuous embedding and discrete text spaces, we explore the plausibility of incorporating various compositional properties into the sentence embedding space that allows us to interpret embedding transformations as compositional sentence operations.We propose INTERSENT, an end-toend framework for learning interpretable sentence embeddings that supports compositional sentence operations in the embedding space.Our method optimizes operator networks and a bottleneck encoder-decoder model to produce meaningful and interpretable sentence embeddings.Experimental results demonstrate that our method significantly improves the interpretability of sentence embeddings on four textual generation tasks over existing approaches while maintaining strong performance on traditional semantic similarity tasks. 1 .
James Y. Huang, Wenlin Yao, Kaiqiang Song, Hongming Zhang 0009, Muhao Chen 0001, Dong Yu 0001
EMNLP2
2023 Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models
Xiaoman Pan, Wenlin Yao, Hongming Zhang 0009, Dian Yu 0001, Dong Yu 0001, Jianshu Chen
ICLR2
2023 Thrust: Adaptively Propels Large Language Models with External Knowledge
abstract
Although large-scale pre-trained language models (PTLMs) are shown to encode rich knowledge in their model parameters, the inherent knowledge in PTLMs can be opaque or static, making external knowledge necessary. However, the existing information retrieval techniques could be costly and may even introduce noisy and sometimes misleading knowledge. To address these challenges, we propose the instance-level adaptive propulsion of external knowledge (IAPEK), where we only conduct the retrieval when necessary. To achieve this goal, we propose to model whether a PTLM contains enough knowledge to solve an instance with a novel metric, Thrust, which leverages the representation distribution of a small amount of seen instances. Extensive experiments demonstrate that Thrust is a good measurement of models' instance-level knowledgeability. Moreover, we can achieve higher cost-efficiency with the Thrust score as the retrieval indicator than the naive usage of external knowledge on 88% of the evaluated tasks with 26% average performance improvement. Such findings shed light on the real-world practice of knowledge-enhanced LMs with a limited budget for knowledge seeking due to computation latency or costs.
Hongming Zhang 0009, Xiaoman Pan, Wenlin Yao, Dong Yu 0001, Jianshu Chen
NeurIPS4
2022 Salience Allocation as Guidance for Abstractive Summarization
abstract
Fei Wang, Kaiqiang Song, Hongming Zhang, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang, Muhao Chen, Dong Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Fei Wang 0060, Kaiqiang Song, Hongming Zhang 0009, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang 0001, Muhao Chen 0001, Dong Yu 0001
EMNLP6
2022 Z-LaVI: Zero-Shot Language Solver Fueled by Visual Imagination
abstract
Large-scale pretrained language models have made significant advances in solving downstream language understanding tasks.However, they generally suffer from reporting bias, the phenomenon describing the lack of explicit commonsense knowledge in written text, e.g., "an orange is orange".To overcome this limitation, we develop a novel approach, Z-LaVI, to endow language models with visual imagination capabilities.Specifically, we leverage two complementary types of "imaginations": (i) recalling existing images through retrieval and (ii) synthesizing nonexistent images via text-toimage generation.Jointly exploiting the language inputs and the imagination, a pretrained vision-language model (e.g., CLIP) eventually composes a zero-shot solution to the original language tasks.Notably, fueling language models with imagination can effectively leverage visual knowledge to solve plain language tasks.In consequence, Z-LaVI consistently improves the zero-shot performance of existing language models across a diverse set of language tasks. 1 * Work was done during the internship at Tencent AI Lab.(a) Word Sense Disambiguation (b) Science Question Answering (c) Topic Classification sense1: bank(institute) sense2: bank(geography) Input: The species prefers the {bank} of pond.
Wenlin Yao, Hongming Zhang 0009, Xiaoyang Wang 0001, Dong Yu 0001, Jianshu Chen
EMNLP2
2021 Connect-the-Dots: Bridging Semantics between Words and Definitions via Aligning Word Sense Inventories
abstract
Word Sense Disambiguation (WSD) aims to automatically identify the exact meaning of one word according to its context.Existing supervised models struggle to make correct predictions on rare word senses due to limited training data and can only select the best definition sentence from one predefined word sense inventory (e.g., WordNet).To address the data sparsity problem and generalize the model to be independent of one predefined inventory, we propose a gloss alignment algorithm that can align definition sentences (glosses) with the same meaning from different sense inventories to collect rich lexical knowledge.We then train a model to identify semantic equivalence between a target word in context and one of its glosses using these aligned inventories, which exhibits strong transfer capability to many WSD tasks 1 .Experiments on benchmark datasets show that the proposed method improves predictions on both frequent and rare word senses, outperforming prior work by 1.2% on the All-Words WSD Task and 4.3% on the Low-Shot WSD Task.Evaluation on WiC Task also indicates that our method can better capture word meanings in context.
Wenlin Yao, Xiaoman Pan, Lifeng Jin, Jianshu Chen, Dian Yu 0001, Dong Yu 0001
EMNLP (1)1
2020 Weakly-Supervised Fine-Grained Event Recognition on Social Media Texts for Disaster Management
abstract
People increasingly use social media to report emergencies, seek help or share information during disasters, which makes social networks an important tool for disaster management. To meet these time-critical needs, we present a weakly supervised approach for rapidly building high-quality classifiers that label each individual Twitter message with fine-grained event categories. Most importantly, we propose a novel method to create high-quality labeled data in a timely manner that automatically clusters tweets containing an event keyword and asks a domain expert to disambiguate event word senses and label clusters quickly. In addition, to process extremely noisy and often rather short user-generated messages, we enrich tweet representations using preceding context tweets and reply tweets in building event recognition classifiers. The evaluation on two hurricanes, Harvey and Florence, shows that using only 1-2 person-hours of human supervision, the rapidly trained weakly supervised classifiers outperform supervised classifiers trained using more than ten thousand annotated tweets created in over 50 person-hours.
Wenlin Yao, Cheng Zhang 0006, Shiva Saravanan, Ruihong Huang, Ali Mostafavi
AAAI1
2020 Weakly Supervised Subevent Knowledge Acquisition
abstract
Subevents elaborate an event and widely exist in event descriptions.Subevent knowledge is useful for discourse analysis and event-centric applications.Acknowledging the scarcity of subevent knowledge, we propose a weakly supervised approach to extract subevent relation tuples from text and build the first large scale subevent knowledge base.We first obtain the initial set of event pairs that are likely to have the subevent relation, by exploiting two observations that 1) subevents are temporally contained by the parent event, and 2) the definitions of the parent event can be used to further guide the identification of subevents.Then, we collect rich weak supervision using the initial seed subevent pairs to train a contextual classifier using BERT and apply the classifier to identify new subevent pairs.The evaluation showed that the acquired subevent tuples (239K) are of high quality (90.1% accuracy) and cover a wide range of event types.The acquired subevent knowledge has been shown useful for discourse analysis and identifying a range of event-event relations 1 .
Wenlin Yao, Zeyu Dai 0001, Maitreyi Ramaswamy, Bonan Min, Ruihong Huang
EMNLP (1)1
2018 Temporal Event Knowledge Acquisition via Identifying Narratives
abstract
Inspired by the double temporality characteristic of narrative texts, we propose a novel approach for acquiring rich temporal "before/after" event knowledge across sentences in narrative stories.The double temporality states that a narrative story often describes a sequence of events following the chronological order and therefore, the temporal order of events matches with their textual order.We explored narratology principles and built a weakly supervised approach that identifies 287k narrative paragraphs from three large text corpora.We then extracted rich temporal event knowledge from these narrative paragraphs.Such event knowledge is shown useful to improve temporal relation classification and outperform several recent neural network models on the narrative cloze task.
Wenlin Yao, Ruihong Huang
ACL (1)1