Yufang Hou 0001

dblp:85/10131-1 · DBLP profile ↗
← Back
46ranked-venue papers
11as first author
31since 2021 · last 2026
0000-0003-2897-6075ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 11 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models
abstract
Javier Carnerero-Cano, Massimiliano Pronesti, Radu Marinescu, Tigran T. Tchrakian, James Barry, Jasmina Gajcin, Yufang Hou, Alessandra Pascale, Elizabeth M. Daly. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Javier Carnerero-Cano, Massimiliano Pronesti, Radu Marinescu 0002, Tigran T. Tchrakian, James Barry, Jasmina Gajcin, Yufang Hou 0001, Alessandra Pascale, Elizabeth Daly
ACL (1)7
2025 The Nature of NLP: Analyzing Contributions in NLP Papers
abstract
Natural Language Processing (NLP) is an established and dynamic field.Despite this, what constitutes NLP research remains debated.In this work, we address the question by quantitatively examining NLP research papers.We propose a taxonomy of research contributions and introduce NLPContributions, a dataset of nearly 2k NLP research paper abstracts, carefully annotated to identify scientific contributions and classify their types according to this taxonomy.We also introduce a novel task of automatically identifying contribution statements and classifying their types from research papers.We present experimental results for this task and apply our model to ∼29k NLP research papers to analyze their contributions, aiding in the understanding of the nature of NLP research.We show that NLP research has taken a winding path -with the focus on language and human-centric studies being prominent in the 1970s and 80s, tapering off in the 1990s and 2000s, and starting to rise again since the late 2010s.Alongside this revival, we observe a steady rise in dataset and methodological contributions since the 1990s, such that today, on average, individual NLP papers contribute in more ways than ever before.Our dataset and analyses offer a powerful lens for tracing research trends and offer potential for generating informed, datadriven literature surveys. 1
Aniket Pramanick, Yufang Hou 0001, Saif M. Mohammad, Iryna Gurevych
ACL (1)2
2025 Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies
abstract
Massimiliano Pronesti, Joao H Bettencourt-Silva, Paul Flanagan, Alessandra Pascale, Oisín Redmond, Anya Belz, Yufang Hou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Massimiliano Pronesti, Joao H. Bettencourt-Silva, Paul Flanagan, Alessandra Pascale, Oisin Redmond, Anya Belz, Yufang Hou 0001
ACL (1)7
2025 Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoning
abstract
Systematic reviews in medicine play a critical role in evidence-based decision-making by aggregating findings from multiple studies.A central bottleneck in automating this process is extracting numeric evidence and determining study-level conclusions for specific outcomes and comparisons.Prior work has framed this problem as a textual inference task by retrieving relevant content fragments and inferring conclusions from them.However, such approaches often rely on shallow textual cues and fail to capture the underlying numeric reasoning behind expert assessments.In this work, we conceptualise the problem as one of quantitative reasoning.Rather than inferring conclusions from surface text, we extract structured numerical evidence (e.g., event counts or standard deviations) and apply domain knowledge informed logic to derive outcome-specific conclusions.We develop a numeric reasoning system composed of a numeric data extraction model and an effect estimate component, enabling more accurate and interpretable inference aligned with the domain expert principles.We train the numeric data extraction model using different strategies, including supervised fine-tuning (SFT), and reinforcement learning (RL) with a new value reward model.When evaluated on the COCHRANEFOREST benchmark, our best-performing approach -using RL to train a small-scale number extraction modelyields up to a 21% absolute improvement in F1 score over retrieval-based systems and outperforms general-purpose LLMs of over 400B parameters by up to 9%.Our results demonstrate the promise of reasoning-driven approaches for automating systematic evidence synthesis.
Massimiliano Pronesti, Michela Lorandi, Paul Flanagan, Oisin Redmond, Anya Belz, Yufang Hou 0001
EMNLP6
2025 A Position Paper on the Automatic Generation of Machine Learning Leaderboards
abstract
An important task in machine learning (ML) research is comparing prior work, which is often performed via ML leaderboards: a tabular overview of experiments with comparable conditions (e.g., same task, dataset, and metric).However, the growing volume of literature creates challenges in creating and maintaining these leaderboards.To ease this burden, researchers have developed methods to extract leaderboard entries from research papers for automated leaderboard curation.Yet, prior work varies in problem framing, complicating comparisons and limiting real-world applicability.In this position paper, we present the first overview of Automatic Leaderboard Generation (ALG) research, identifying fundamental differences in assumptions, scope, and output formats.We propose an ALG unified conceptual framework to standardise how the ALG task is defined.We offer ALG benchmarking guidelines, including recommendations for datasets and metrics that promote fair, reproducible evaluation.Lastly, we outline challenges and new directions for ALG, such as, advocating for broader coverage by including all reported results and richer metadata.
Roelien C. Timmer, Yufang Hou 0001, Stephen Wan 0001
EMNLP2
2025 Grounding Fallacies Misrepresenting Scientific Publications in Evidence
abstract
Max Glockner, Yufang Hou, Preslav Nakov, Iryna Gurevych. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Max Glockner, Yufang Hou 0001, Preslav Nakov, Iryna Gurevych
NAACL (Long Papers)2
2024 Missci: Reconstructing Fallacies in Misrepresented Science
abstract
Health-related misinformation on social networks can lead to poor decision-making and real-world dangers.Such misinformation often misrepresents scientific publications and cites them as "proof" to gain perceived credibility.To effectively counter such claims automatically, a system must explain how the claim was falsely derived from the cited publication.Current methods for automated fact-checking or fallacy detection neglect to assess the (mis)used evidence in relation to misinformation claims, which is required to detect the mismatch between them.To address this gap, we introduce MISSCI, a novel argumentation theoretical model for fallacious reasoning together with a new dataset for real-world misinformation detection that misrepresents biomedical publications.Unlike previous fallacy detection datasets, MISSCI (i) focuses on implicit fallacies between the relevant content of the cited publication and the inaccurate claim, and (ii) requires models to verbalize the fallacious reasoning in addition to classifying it.We present MISSCI as a dataset to test the critical reasoning abilities of large language models (LLMs), which are required to reconstruct real-world fallacious arguments, in a zero-shot setting.We evaluate two representative LLMs and the impact of providing different levels of detail about the fallacy classes to the LLMs via prompts.Our experiments and human evaluation show promising results for GPT 4, while also demonstrating the difficulty of this task. 1 1 Code and data are available at: https://github. com/UKPLab/acl2024-missci.Claim: Hydroxychloroquine is a cure for COVID-19. Accurate premise ( ): Chloroquine reduced infection of the coronavirus. Fallacy of CompositionFallacious premise ( ) SARS-CoV-1 and SARS-CoV-2 are both coronaviruses.Therefore, they can be treated the same way.
Max Glockner, Yufang Hou 0001, Preslav Nakov, Iryna Gurevych
ACL (1)2
2024 Systematic Task Exploration with LLMs: A Study in Citation Text Generation
abstract
Large language models (LLMs) bring unprecedented flexibility in defining and executing complex, creative natural language generation (NLG) tasks.Yet, this flexibility brings new challenges, as it introduces new degrees of freedom in formulating the task inputs and instructions and in evaluating model performance.To facilitate the exploration of creative NLG tasks, we propose a three-component research framework that consists of systematic input manipulation, reference data, and output measurement.We use this framework to explore citation text generation -a popular scholarly NLP task that lacks consensus on the task definition and evaluation metric and has not yet been tackled within the LLM paradigm.Our results highlight the importance of systematically investigating both task instruction and input configuration when prompting LLMs, and reveal non-trivial relationships between different evaluation metrics used for citation text generation.Additional human generation and human evaluation experiments provide new qualitative insights into the task to guide future research in citation text generation.We make our code 1 and data 2 publicly available.
Furkan Sahinuç, Ilia Kuznetsov, Yufang Hou 0001, Iryna Gurevych
ACL (1)3
2024 How to Handle Different Types of Out-of-Distribution Scenarios in Computational Argumentation? A Comprehensive and Fine-Grained Field Study
abstract
The advent of pre-trained Language Models (LMs) has markedly advanced natural language processing, but their efficacy in out-of-distribution (OOD) scenarios remains a significant challenge. Computational argumentation (CA), modeling human argumentation processes, is a field notably impacted by these challenges because complex annotation schemes and high annotation costs naturally lead to resources barely covering the multiplicity of available text sources and topics. Due to this data scarcity, generalization to data from uncovered covariant distributions is a common challenge for CA tasks like stance detection or argument classification. This work systematically assesses LMs’ capabilities for such OOD scenarios. While previous work targets specific OOD types like topic shifts or OOD uniformly, we address three prevalent OOD scenarios in CA: topic shift, domain shift, and language shift. Our findings challenge the previously asserted general superiority of in-context learning (ICL) for OOD. We find that the efficacy of such learning paradigms varies with the type of OOD. Specifically, while ICL excels for domain shifts, prompt-based fine-tuning surpasses for topic shifts. To sum up, we navigate the heterogeneity of OOD scenarios in CA and empirically underscore the potential of base-sized LMs in overcoming these challenges.
Andreas Waldis, Yufang Hou 0001, Iryna Gurevych
ACL (1)2
2024 Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
abstract
Scientific leaderboards are standardized ranking systems that facilitate evaluating and comparing competitive methods.Typically, a leaderboard is defined by a task, dataset, and evaluation metric (TDM) triple, allowing objective performance assessment and fostering innovation through benchmarking.However, the exponential increase in publications has made it infeasible to construct and maintain these leaderboards manually.Automatic leaderboard construction has emerged as a solution to reduce manual labor.Existing datasets for this task are based on the community-contributed leaderboards without additional curation.Our analysis shows that a large portion of these leaderboards are incomplete, and some of them contain incorrect information.In this work, we present SCILEAD, a manually-curated Scientific Leaderboard dataset that overcomes the aforementioned problems.Building on this dataset, we propose three experimental settings that simulate real-world scenarios where TDM triples are fully defined, partially defined, or undefined during leaderboard construction.While previous research has only explored the first setting, the latter two are more representative of real-world applications.To address these diverse settings, we develop a comprehensive LLM-based framework for constructing leaderboards.Our experiments and analysis reveal that various LLMs often correctly identify TDM triples while struggling to extract result values from publications.We make our code 1 and data 2 publicly available.
Furkan Sahinuç, Thy Thy Tran, Yulia Grishina, Yufang Hou 0001, Iryna Gurevych
EMNLP4
2024 WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
abstract
Retrieval-augmented generation (RAG) has emerged as a promising solution to mitigate the limitations of large language models (LLMs), such as hallucinations and outdated information. However, it remains unclear how LLMs handle knowledge conflicts arising from different augmented retrieved passages, especially when these passages originate from the same source and have equal trustworthiness. In this work, we conduct a comprehensive evaluation of LLM-generated answers to questions that have varying answers based on contradictory passages from Wikipedia, a dataset widely regarded as a high-quality pre-training resource for most LLMs. Specifically, we introduce WikiContradict, a benchmark consisting of 253 high-quality, human-annotated instances designed to assess the performance of LLMs in providing a complete perspective on conflicts from the retrieved documents, rather than choosing one answer over another, when augmented with retrieved passages containing real-world knowledge conflicts. We benchmark a diverse range of both closed and open-source LLMs under different QA scenarios, including RAG with a single passage, and RAG with 2 contradictory passages. Through rigorous human evaluations on a subset of WikiContradict instances involving 5 LLMs and over 3,500 judgements, we shed light on the behaviour and limitations of these models. For instance, when provided with two passages containing contradictory facts, all models struggle to generate answers that accurately reflect the conflicting nature of the context, especially for implicit conflicts requiring reasoning. Since human evaluation is costly, wealso introduce an automated model that estimates LLM performance using a strong open-source language model, achieving an F-score of 0.8. Using this automated metric, we evaluate more than 1,500 answers from seven LLMs across all WikiContradict instances.
Yufang Hou 0001, Alessandra Pascale, Javier Carnerero-Cano, Tigran T. Tchrakian, Radu Marinescu 0002, Elizabeth Daly, Inkit Padhi, Prasanna Sattigeri
NeurIPS1
2024 Holmes ⌕ A Benchmark to Assess the Linguistic Competence of Language Models
abstract
Abstract We introduce Holmes, a new benchmark designed to assess language models’ (LMs’) linguistic competence—their unconscious understanding of linguistic phenomena. Specifically, we use classifier-based probing to examine LMs’ internal representations regarding distinct linguistic phenomena (e.g., part-of-speech tagging). As a result, we meet recent calls to disentangle LMs’ linguistic competence from other cognitive abilities, such as following instructions in prompting-based evaluations. Composing Holmes, we review over 270 probing studies and include more than 200 datasets to assess syntax, morphology, semantics, reasoning, and discourse phenomena. Analyzing over 50 LMs reveals that, aligned with known trends, their linguistic competence correlates with model size. However, surprisingly, model architecture and instruction tuning also significantly influence performance, particularly in morphology and syntax. Finally, we propose FlashHolmes, a streamlined version that reduces the computation load while maintaining high-ranking precision.
Andreas Waldis, Yotam Perlitz, Leshem Choshen, Yufang Hou 0001, Iryna Gurevych
Trans. Assoc. Comput. Linguistics4
2023 Matching Pairs: Attributing Fine-Tuned Models to their Pre-Trained Large Language Models
abstract
Myles Foley, Ambrish Rawat, Taesung Lee, Yufang Hou, Gabriele Picco, Giulio Zizzo. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Myles Foley, Ambrish Rawat, Taesung Lee, Yufang Hou 0001, Gabriele Picco, Giulio Zizzo
ACL (1)4
2023 Are Fairy Tales Fair? Analyzing Gender Bias in Temporal Narrative Event Chains of Children's Fairy Tales
abstract
Paulina Toro Isaza, Guangxuan Xu, Toye Oloko, Yufang Hou, Nanyun Peng, Dakuo Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Paulina Toro Isaza, Guangxuan Xu, Toye Oloko, Yufang Hou 0001, Nanyun Peng 0001, Dakuo Wang
ACL (1)4
2023 PairSpanBERT: An Enhanced Language Model for Bridging Resolution
abstract
We present PAIRSPANBERT, a SPANBERTbased pre-trained model specialized for bridging resolution.PAIRSPANBERT is pre-trained with a novel objective that aims to learn the contexts in which two mentions are implicitly linked to each other from a large amount of data automatically generated either heuristically or via distance supervision with a knowledge graph.Despite the noise inherent in the automatically generated data, we achieve the best results reported to date on three evaluation datasets for bridging resolution when replacing SPANBERT with PAIRSPANBERT in a stateof-the-art resolver that jointly performs entity coreference resolution and bridging resolution.
Hideo Kobayashi, Yufang Hou 0001, Vincent Ng 0001
ACL (1)2
2023 A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for Summarization
abstract
Lining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, João Sedoc. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Lining Zhang, Simon Mille, Yufang Hou 0001, Daniel Deutsch, Elizabeth Clark, Yixin Liu 0003, Saad Mahamood, Sebastian Gehrmann, Miruna-Adriana Clinciu, Khyathi Raghavi Chandu, João Sedoc
ACL (1)3
2023 'Don't Get Too Technical with Me': A Discourse Structure-Based Framework for Automatic Science Journalism
abstract
Science journalism refers to the task of reporting technical findings of a scientific paper as a less technical news article to the general public audience.We aim to design an automated system to support this real-world task (i.e., automatic science journalism) by 1) introducing a newly-constructed and real-world dataset (SCITECHNEWS), with tuples of a publiclyavailable scientific paper, its corresponding news article, and an expert-written short summary snippet; 2) proposing a novel technical framework that integrates a paper's discourse structure with its metadata to guide generation; and, 3) demonstrating with extensive automatic and human experiments that our framework outperforms other baseline methods (e.g.Alpaca and ChatGPT) in elaborating a content plan meaningful for the target audience, simplifying the information selected, and producing a coherent final report in a layman's style.
Ronald Cardenas, Bingsheng Yao, Dakuo Wang, Yufang Hou 0001
EMNLP4
2023 CiteBench: A Benchmark for Scientific Citation Text Generation
abstract
Science progresses by building upon the prior body of knowledge documented in scientific publications.The acceleration of research makes it hard to stay up-to-date with the recent developments and to summarize the evergrowing body of prior work.To address this, the task of citation text generation aims to produce accurate textual summaries given a set of papers-to-cite and the citing paper context.Due to otherwise rare explicit anchoring of cited documents in the citing paper, citation text generation provides an excellent opportunity to study how humans aggregate and synthesize textual knowledge from sources.Yet, existing studies are based upon widely diverging task definitions, which makes it hard to study this task systematically.To address this challenge, we propose CITEBENCH: a benchmark for citation text generation that unifies multiple diverse datasets and enables standardized evaluation of citation text generation models across task designs and domains.Using the new benchmark, we investigate the performance of multiple strong baselines, test their transferability between the datasets, and deliver new insights into the task definition and evaluation to guide future research in citation text generation.We make the code for CITEBENCH publicly available at https://github.com/ UKPLab/citebench.
Martin Funkquist, Ilia Kuznetsov, Yufang Hou 0001, Iryna Gurevych
EMNLP3
2023 A Diachronic Analysis of Paradigm Shifts in NLP Research: When, How, and Why?
abstract
Understanding the fundamental concepts and trends in a scientific field is crucial for keeping abreast of its continuous advancement.In this study, we propose a systematic framework for analyzing the evolution of research topics in a scientific field using causal discovery and inference techniques.We define three variables to encompass diverse facets of the evolution of research topics within NLP and utilize a causal discovery algorithm to unveil the causal connections among these variables using observational data.Subsequently, we leverage this structure to measure the intensity of these relationships.By conducting extensive experiments on the ACL Anthology corpus, we demonstrate that our framework effectively uncovers evolutionary trends and the underlying causes for a wide range of NLP research topics.Specifically, we show that tasks and methods are primary drivers of research in NLP, with datasets following, while metrics have minimal impact. 1
Aniket Pramanick, Yufang Hou 0001, Saif M. Mohammad, Iryna Gurevych
EMNLP2
2022 Constrained Multi-Task Learning for Bridging Resolution
abstract
We examine the extent to which supervised bridging resolvers can be improved without employing additional labeled bridging data by proposing a novel constrained multi-task learning framework for bridging resolution, within which we (1) design cross-task consistency constraints to guide the learning process; (2) pretrain the entity coreference model in the multitask framework on the large amount of publicly available coreference data; and (3) integrate prior knowledge encoded in rule-based resolvers.Our approach achieves state-of-theart results on three standard evaluation corpora.
Hideo Kobayashi, Yufang Hou 0001, Vincent Ng 0001
ACL (1)2
2022 Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative Comprehension
abstract
Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Li, Nora Bradford, Branda Sun, Tran Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, Mark Warschauer. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Dakuo Wang, Mo Yu, Daniel Ritchie 0002, Bingsheng Yao, Sherry Tongshuang Wu, Zheng Zhang 0043, Toby Jia-Jun Li, Nora Bradford, Branda Sun, Tran Bao Hoang, Yisi Sang, Yufang Hou 0001, Xiaojuan Ma, Diyi Yang, Nanyun Peng 0001, Zhou Yu 0005, Mark Warschauer
ACL (1)13
2022 Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric Summarization
abstract
Generating educational questions of fairytales or storybooks is vital for improving children's literacy ability.However, it is challenging to generate questions that capture the interesting aspects of a fairytale story with educational meaningfulness.In this paper, we propose a novel question generation method that first learns the question type distribution of an input story paragraph, and then summarizes salient events which can be used to generate high-cognitive-demand questions.To train the event-centric summarizer, we finetune a pre-trained transformer-based sequenceto-sequence model using silver samples composed by educational question-answer pairs.On a newly proposed educational questionanswering dataset FairytaleQA, we show good performance of our method on both automatic and human evaluation metrics.Our work indicates the necessity of decomposing question type distribution learning and event-centric summary generation for educational question generation.
Zhenjie Zhao, Yufang Hou 0001, Dakuo Wang, Mo Yu, Chengzhong Liu, Xiaojuan Ma
ACL (1)2
2022 End-to-End Neural Bridging Resolution
abstract
The state of bridging resolution research is rather unsatisfactory: not only are state-of-the-art resolvers evaluated in unrealistic settings, but the neural models underlying these resolvers are weaker than those used for entity coreference resolution. In light of these problems, we evaluate bridging resolvers in an end-to-end setting, strengthen them with better encoders, and attempt to gain a better understanding of them via perturbation experiments and a manual analysis of their outputs.
Hideo Kobayashi, Yufang Hou 0001, Vincent Ng 0001
COLING2
2022 Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation
abstract
Misinformation emerges in times of uncertainty when credible information is limited.This is challenging for NLP-based fact-checking as it relies on counter-evidence, which may not yet be available.Despite increasing interest in automatic fact-checking, it is still unclear if automated approaches can realistically refute harmful real-world misinformation.Here, we contrast and compare NLP fact-checking with how professional fact-checkers combat misinformation in the absence of counter-evidence.In our analysis, we show that, by design, existing NLP task definitions for fact-checking cannot refute misinformation as professional fact-checkers do for the majority of claims.We then define two requirements that the evidence in datasets must fulfill for realistic factchecking: It must be (1) sufficient to refute the claim and (2) not leaked from existing fact-checking articles.We survey existing factchecking datasets and find that all of them fail to satisfy both criteria.Finally, we perform experiments to demonstrate that models trained on a large-scale fact-checking dataset rely on leaked evidence, which makes them unsuitable in real-world scenarios.Taken together, we show that current NLP fact-checking cannot realistically combat real-world misinformation because it depends on unrealistic assumptions about counter-evidence in the data 1 .
Max Glockner, Yufang Hou 0001, Iryna Gurevych
EMNLP2
2022 Privacy-aware supervised classification: An informative subspace based multi-objective approach
Chandan Biswas, Debasis Ganguly, Partha Sarathi Mukherjee, Ujjwal Bhattacharya, Yufang Hou 0001
Pattern Recognit.5
2021 Employing Argumentation Knowledge Graphs for Neural Argument Generation
abstract
Khalid Al Khatib, Lukas Trautner, Henning Wachsmuth, Yufang Hou, Benno Stein. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Khalid Al-Khatib, Lukas Trautner, Henning Wachsmuth, Yufang Hou 0001, Benno Stein 0001
ACL/IJCNLP (1)4
2021 Outcome Prediction from Behaviour Change Intervention Evaluations using a Combination of Node and Word Embedding
Debasis Ganguly, Martin Gleize, Yufang Hou 0001, Charles Jochim, Francesca Bonin, Alessandra Pascale, Pierpaolo Tommasi, Pol Mac Aonghusa, Marie Johnston, Mike Kelly, Susan Michie
AMIA3
2021 TDMSci: A Specialized Corpus for Scientific Literature Entity Tagging of Tasks Datasets and Metrics
abstract
Yufang Hou, Charles Jochim, Martin Gleize, Francesca Bonin, Debasis Ganguly. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Yufang Hou 0001, Charles Jochim, Martin Gleize, Francesca Bonin, Debasis Ganguly
EACL1
2021 Probing for Bridging Inference in Transformer Language Models
abstract
We probe pre-trained transformer language models for bridging inference.We first investigate individual attention heads in BERT and observe that attention heads at higher layers prominently focus on bridging relations incomparison with the lower and middle layers, also, few specific attention heads concentrate consistently on bridging.More importantly, we consider language models as a whole in our second approach where bridging anaphora resolution is formulated as a masked token prediction task (Of-Cloze test).Our formulation produces optimistic results without any finetuning, which indicates that pre-trained language models substantially capture bridging inference.Our further investigation shows that the distance between anaphor-antecedent and the context provided to language models play an important role in the inference.
Onkar Pandit, Yufang Hou 0001
NAACL-HLT2
2021 D2S: Document-to-Slide Generation Via Query-Based Text Summarization
abstract
Edward Sun, Yufang Hou, Dakuo Wang, Yunfeng Zhang, Nancy X. R. Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Edward Sun, Yufang Hou 0001, Dakuo Wang, Nancy Xin Ru Wang
NAACL-HLT2
2021 Ensembling Graph Predictions for AMR Parsing
abstract
In many machine learning tasks, models are trained to predict structure data such as graphs. For example, in natural language processing, it is very common to parse texts into dependency trees or abstract meaning representation (AMR) graphs. On the other hand, ensemble methods combine predictions from multiple models to create a new one that is more robust and accurate than individual predictions. In the literature, there are many ensembling techniques proposed for classification or regression problems, however, ensemble graph prediction has not been studied thoroughly. In this work, we formalize this problem as mining the largest graph that is the most supported by a collection of graph predictions. As the problem is NP-Hard, we propose an efficient heuristic algorithm to approximate the optimal solution. To validate our approach, we carried out experiments in AMR parsing problems. The experimental results demonstrate that the proposed approach can combine the strength of state-of-the-art AMR parsers to create new predictions that are more accurate than any individual models in five standard benchmark datasets.
Hoang Thanh Lam, Gabriele Picco, Yufang Hou 0001, Young-Suk Lee 0001, Lam M. Nguyen, Dzung T. Phan, Vanessa López, Ramón Fernandez Astudillo
NeurIPS3
2020 Corpus Wide Argument Mining - A Working Solution
abstract
One of the main tasks in argument mining is the retrieval of argumentative content pertaining to a given topic. Most previous work addressed this task by retrieving a relatively small number of relevant documents as the initial source for such content. This line of research yielded moderate success, which is of limited use in a real-world system. Furthermore, for such a system to yield a comprehensive set of relevant arguments, over a wide range of topics, it requires leveraging a large and diverse corpus in an appropriate manner. Here we present a first end-to-end high-precision, corpus-wide argument mining system. This is made possible by combining sentence-level queries over an appropriate indexing of a very large corpus of newspaper articles, with an iterative annotation scheme. This scheme addresses the inherent label bias in the data and pinpoints the regions of the sample space whose manual labeling is required to obtain high-precision among top-ranked candidates.
Liat Ein-Dor, Eyal Shnarch, Lena Dankin, Alon Halfon, Benjamin Sznajder, Ariel Gera, Carlos Alzate, Martin Gleize, Leshem Choshen, Yufang Hou 0001, Yonatan Bilu, Ranit Aharonov, Noam Slonim
AAAI10
2020 End-to-End Argumentation Knowledge Graph Construction
abstract
This paper studies the end-to-end construction of an argumentation knowledge graph that is intended to support argument synthesis, argumentative question answering, or fake news detection, among others. The study is motivated by the proven effectiveness of knowledge graphs for interpretable and controllable text generation and exploratory search. Original in our work is that we propose a model of the knowledge encapsulated in arguments. Based on this model, we build a new corpus that comprises about 16k manual annotations of 4740 claims with instances of the model's elements, and we develop an end-to-end framework that automatically identifies all modeled types of instances. The results of experiments show the potential of the framework for building a web-based argumentation graph that is of high quality and large scale.
Khalid Al-Khatib, Yufang Hou 0001, Henning Wachsmuth, Charles Jochim, Francesca Bonin, Benno Stein 0001
AAAI2
2020 Bridging Anaphora Resolution as Question Answering
abstract
Most previous studies on bridging anaphora resolution (Poesio et al., 2004; Hou et al., 2013b; Hou, 2018a) use the pairwise model to tackle the problem and assume that the gold mention information is given. In this paper, we cast bridging anaphora resolution as question answering based on context. This allows us to find the antecedent for a given anaphor without knowing any gold mention information (except the anaphor itself). We present a question answering framework (BARQA) for this task, which leverages the power of transfer learning. Furthermore, we propose a novel method to generate a large amount of "quasi-bridging" training data. We show that our model pre-trained on this dataset and fine-tuned on a small amount of in-domain dataset achieves new state-of-the-art results for bridging anaphora resolution on two bridging corpora (ISNotes (Markert et al., 2012) and BASHI (Ro ̈siger, 2018)).
Yufang Hou 0001
ACL1
2020 Knowledge Extraction and Prediction from Behavior Science Randomized Controlled Trials: A Case Study in Smoking Cessation
Francesca Bonin, Martin Gleize, Yufang Hou 0001, Debasis Ganguly, Ailbhe Finnerty, Charles Jochim, Alessandra Pascale, Pierpaolo Tommasi, Pol Mac Aonghusa, Susan Michie
AMIA3
2020 Fine-grained Information Status Classification Using Discourse Context-Aware BERT
abstract
Previous work on bridging anaphora recognition (Hou et al., 2013a) casts the problem as a subtask of learning fine-grained information status (IS).However, these systems heavily depend on many hand-crafted linguistic features.In this paper, we propose a simple discourse context-aware BERT model for fine-grained IS classification.On the ISNotes corpus (Markert et al., 2012), our model achieves new state-of-the-art performance on fine-grained IS classification, obtaining a 4.8 absolute overall accuracy improvement compared to Hou et al. (2013a).More importantly, we also show an improvement of 10.5 F1 points for bridging anaphora recognition without using any complex hand-crafted semantic features designed for capturing the bridging phenomenon.We further analyze the trained model and find that the most attended signals for each IS category correspond well to linguistic notions of information status.
Yufang Hou 0001
COLING1
2020 HBCP Corpus: A New Resource for the Analysis of Behavioural Change Intervention Reports
abstract
Due to the fast pace at which research reports in behaviour change are published, researchers, consultants and policymakers would benefit from more automatic ways to process these reports. Automatic extraction of the reports’ intervention content, population, settings and their results etc. are essential in synthesising and summarising the literature. However, to the best of our knowledge, no unique resource exists at the moment to facilitate this synthesis. In this paper, we describe the construction of a corpus of published behaviour change intervention evaluation reports aimed at smoking cessation. We also describe and release the annotation of 57 entities, that can be used as an off-the-shelf data resource for tasks such as entity recognition, etc. Both the corpus and the annotation dataset are being made available to the community.
Francesca Bonin, Martin Gleize, Ailbhe Finnerty, Candice Moore, Charles Jochim, Emma Norris, Yufang Hou 0001, Alison J. Wright, Debasis Ganguly, Emily Hayes, Silje Zink, Alessandra Pascale, Pol Mac Aonghusa, Susan Michie
LREC7
2019 Identification of Tasks, Datasets, Evaluation Metrics, and Numeric Scores for Scientific Leaderboards Construction
abstract
While the fast-paced inception of novel tasks and new datasets helps foster active research in a community towards interesting directions, keeping track of the abundance of research activity in different areas on different datasets is likely to become increasingly difficult.The community could greatly benefit from an automatic system able to summarize scientific results, e.g., in the form of a leaderboard.In this paper we build two datasets and develop a framework (TDMS-IE) aimed at automatically extracting task, dataset, metric and score from NLP papers, towards the automatic construction of leaderboards.Experiments show that our model outperforms several baselines by a large margin.Our model is a first step towards automatic leaderboard construction, e.g., in the NLP domain.
Yufang Hou 0001, Charles Jochim, Martin Gleize, Francesca Bonin, Debasis Ganguly
ACL (1)1
2018 A Deterministic Algorithm for Bridging Anaphora Resolution
abstract
Previous work on bridging anaphora resolution (Poesio et al., 2004; Hou et al., 2013b) use syntactic preposition patterns to calculate word relatedness.However, such patterns only consider NPs' head nouns and hence do not fully capture the semantics of NPs.Recently, Hou (2018) created word embeddings (embeddings PP) to capture associative similarity (i.e., relatedness) between nouns by exploring the syntactic structure of noun phrases.But embeddings PP only contains word representations for nouns.In this paper, we create new word vectors by combining embeddings PP with GloVe.This new word embeddings (embeddings bridging) are a more general lexical knowledge resource for bridging and allow us to represent the meaning of an NP beyond its head easily.We therefore develop a deterministic approach for bridging anaphora resolution, which represents the semantics of an NP based on its head noun and modifications.We show that this simple approach achieves the competitive results compared to the best system in Hou et al. (2013b) which explores Markov Logic Networks to model the problem.Additionally, we further improve the results for bridging anaphora resolution reported in Hou (2018) by combining our simple deterministic approach with Hou et al. (2013b)'s best system MLN II.
Yufang Hou 0001
EMNLP1
2018 Unrestricted Bridging Resolution
abstract
In contrast to identity anaphors, which indicate coreference between a noun phrase and its antecedent, bridging anaphors link to their antecedent(s) via lexico-semantic, frame, or encyclopedic relations. Bridging resolution involves recognizing bridging anaphors and finding links to antecedents. In contrast to most prior work, we tackle both problems. Our work also follows a more wide-ranging definition of bridging than most previous work and does not impose any restrictions on the type of bridging anaphora or relations between anaphor and antecedent. We create a corpus (ISNotes) annotated for information status (IS), bridging being one of the IS subcategories. The annotations reach high reliability for all categories and marginal reliability for the bridging subcategory. We use a two-stage statistical global inference method for bridging resolution. Given all mentions in a document, the first stage, bridging anaphora recognition, recognizes bridging anaphors as a subtask of learning fine-grained IS. We use a cascading collective classification method where (i) collective classification allows us to investigate relations among several mentions and autocorrelation among IS classes and (ii) cascaded classification allows us to tackle class imbalance, important for minority classes such as bridging. We show that our method outperforms current methods both for IS recognition overall as well as for bridging, specifically. The second stage, bridging antecedent selection, finds the antecedents for all predicted bridging anaphors. We investigate the phenomenon of semantically or syntactically related bridging anaphors that share the same antecedent, a phenomenon we call sibling anaphors. We show that taking sibling anaphors into account in a joint inference model improves antecedent selection performance. In addition, we develop semantic and salience features for antecedent selection and suggest a novel method to build the candidate antecedent list for an anaphor, using the discourse scope of the anaphor. Our model outperforms previous work significantly.
Yufang Hou 0001, Katja Markert, Michael Strube 0001
Comput. Linguistics1
2017 Computational Argumentation Quality Assessment in Natural Language
abstract
Henning Wachsmuth, Nona Naderi, Yufang Hou, Yonatan Bilu, Vinodkumar Prabhakaran, Tim Alberdingk Thijm, Graeme Hirst, Benno Stein. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Henning Wachsmuth, Nona Naderi, Yufang Hou 0001, Yonatan Bilu, Vinodkumar Prabhakaran, Tim Alberdingk Thijm, Graeme Hirst, Benno Stein 0001
EACL (1)3
2016 Incremental Fine-grained Information Status Classification Using Attention-based LSTMs
abstract
Information status plays an important role in discourse processing. According to the hearer’s common sense knowledge and his comprehension of the preceding text, a discourse entity could be old, mediated or new. In this paper, we propose an attention-based LSTM model to address the problem of fine-grained information status classification in an incremental manner. Our approach resembles how human beings process the task, i.e., decide the information status of the current discourse entity based on its preceding context. Experimental results on the ISNotes corpus (Markert et al., 2012) reveal that (1) despite its moderate result, our model with only word embedding features captures the necessary semantic knowledge needed for the task by a large extent; and (2) when incorporating with additional several simple features, our model achieves the competitive results compared to the state-of-the-art approach (Hou et al., 2013) which heavily depends on lots of hand-crafted semantic features.
Yufang Hou 0001
COLING1
2014 A Rule-Based System for Unrestricted Bridging Resolution: Recognizing Bridging Anaphora and Finding Links to Antecedents
abstract
Bridging resolution plays an important role in establishing (local) entity coherence.This paper proposes a rule-based approach for the challenging task of unrestricted bridging resolution, where bridging anaphors are not limited to definite NPs and semantic relations between anaphors and their antecedents are not restricted to meronymic relations.The system consists of eight rules which target different relations based on linguistic insights.Our rule-based system significantly outperforms a reimplementation of a previous rule-based system (Vieira and Poesio, 2000).Furthermore, it performs better than a learning-based approach which has access to the same knowledge resources as the rule-based system.Additionally, incorporating the rules and more features into the learning-based system yields a minor improvement over the rule-based system.
Yufang Hou 0001, Katja Markert, Michael Strube 0001
EMNLP1
2013 Cascading Collective Classification for Bridging Anaphora Recognition using a Rich Linguistic Feature Set
abstract
Recognizing bridging anaphora is difficult due to the wide variation within the phenomenon, the resulting lack of easily identifiable surface markers and their relative rarity.We develop linguistically motivated discourse structure, lexico-semantic and genericity detection features and integrate these into a cascaded minority preference algorithm that models bridging recognition as a subtask of learning finegrained information status (IS).We substantially improve bridging recognition without impairing performance on other IS classes.
Yufang Hou 0001, Katja Markert, Michael Strube 0001
EMNLP1
2013 Global Inference for Bridging Anaphora Resolution
Yufang Hou 0001, Katja Markert, Michael Strube 0001
HLT-NAACL1
2012 Collective Classification for Fine-grained Information Status
Katja Markert, Yufang Hou 0001, Michael Strube 0001
ACL (1)2