VLDB 2026 Research / reviewers in the wild / expert
Andreas Vlachos 0001
dblp:18/1071-1
· DBLP profile ↗
66ranked-venue papers
5as first author
39since 2021 · last 2026
0000-0003-2123-5071ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 4 first-author · 38 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form GenerationabstractHallucination remains a major challenge for the safe and trustworthy deployment of large language models (LLMs) in factual content generation.Prior work has explored confidence estimation as an effective approach to hallucination detection, but often relies on post-hoc self-consistency methods that require computationally expensive sampling.Verbalized confidence offers a more efficient alternative, but existing approaches are largely limited to shortform question answering (QA) tasks and do not generalize well to open-ended generation.In this paper, we propose LOVEC (Long-form Verbalized Confidence), a novel reinforcement learning (RL)-based method that trains LLMs to append an on-the-fly numerical confidence score to each generated statement during longform generation.The confidence score serves as a direct and interpretable signal of the factuality of generation.We introduce two evaluation settings, free-form tagging and iterative tagging, to assess different verbalized confidence estimation methods.Experiments on three long-form QA datasets show that our RLtrained models achieve better calibration and generalize robustly across domains.Also, our method is highly efficient, being 20× faster than traditional self-consistency methods while achieving better calibration. Caiqi Zhang, Chengzu Li, Nigel Collier, Andreas Vlachos 0001 |
ACL (1) | 5 |
| 2026 | Ev2R: Evaluating Evidence Retrieval in Automated Fact-CheckingabstractAbstract Current automated fact-checking (AFC) approaches typically evaluate evidence either implicitly via the predicted verdicts or through exact matches with predefined closed knowledge sources, such as Wikipedia. However, these methods are limited due to their reliance on evaluation metrics originally designed for other purposes and constraints from closed knowledge sources. In this work, we introduce Ev2R which combines the strengths of reference-based evaluation and verdict-level proxy scoring. Ev2R jointly assesses how well the evidence aligns with the gold references and how reliably it supports the verdict, addressing the shortcomings of prior methods. We evaluate Ev2R against three types of evidence evaluation approaches: reference-based, proxy-reference, and reference-less baselines. Assessments against human ratings and adversarial tests demonstrate that Ev2R consistently outperforms existing scoring approaches in accuracy and robustness. It achieves stronger correlation with human judgments and greater robustness to adversarial perturbations, establishing it as a reliable metric for evidence evaluation in AFC.1 Mubashara Akhtar, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Causal Estimation of Tokenisation BiasabstractModern language models are typically trained over subword sequences, but ultimately define probabilities over character-strings.Ideally, the choice of the tokeniser-which maps characterstrings to subwords-should not affect the probability assigned to the underlying characterstring; in practice, it does.We define this mismatch as tokenisation bias.In this work, we quantify one particular type of tokenisation bias: the effect of including or not a subword (e.g., ⟨hello⟩) in a tokeniser's vocabulary on the probability a trained model assigns to the corresponding characters (i.e., "hello").Estimating this effect is challenging because each model is trained with only one tokeniser.We address this by framing tokenisation bias as a causal effect and estimating it using the regression discontinuity design.Specifically, we exploit the fact that tokenisation algorithms rank subwords and add the first K to a tokeniser's vocabulary, where K is an arbitrary cutoff point.As such, we can estimate a causal effect by comparing similar subwords around this cutoff.Experimentally, we find that tokenisation consistently affects models' outputs across scales, vocabularies, and tokenisers.Notably, a subword's presence in a small model's vocabulary may increase its characters' probability by up to 17 times, highlighting tokenisation as a key design choice in language modelling. Pietro Lesci, Clara Meister, Thomas Hofmann 0001, Andreas Vlachos 0001, Tiago Pimentel |
ACL (1) | 4 |
| 2025 | Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language ModelsabstractDiffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text.Token-level diffusion doesn't model wordorder dependencies explicitly and operates on short, fixed output windows, while passagelevel diffusion struggles with learning robust representations for long-form text.To address these challenges, we propose Segment-Level Diffusion (SLD), a framework that enhances diffusion-based text generation through text segmentation, robust representation training with adversarial and contrastive learning, and improved latent-space guidance.By segmenting long-form outputs into multiple latent representations and decoding them with an autoregressive decoder, SLD simplifies diffusion predictions and improves scalability.Experiments on four datasets demonstrate that, when compared to other diffusion and autoregressive baselines SLD achieves competitive or superior fluency, coherence, and contextual compatibility in automatic and human evaluations.1 1 Our code is available at: https://github.com/ SpaceHunterInf/Segment_Level_Diffusion Encoder 𝑖 !𝑖 " Georgi Karadzhov, Chenxi Whitehouse, Andreas Vlachos 0001 |
ACL (1) | 4 |
| 2025 | Conformity in Large Language ModelsabstractThe conformity effect describes the tendency of individuals to align their responses with the majority.Studying this bias in large language models (LLMs) is crucial, as LLMs are increasingly used in various information-seeking and decision-making tasks as conversation partners to improve productivity.Thus, conformity to incorrect responses can compromise their effectiveness.In this paper, we adapt psychological experiments to examine the extent of conformity in popular LLMs.Our findings reveal that all tested models exhibit varying levels of conformity toward the majority, regardless of their initial choice or correctness, across different knowledge domains.Notably, we are the first to show that LLMs are more likely to conform when they are more uncertain in their own prediction.We further explore factors that influence conformity, such as training paradigms and input characteristics, finding that instruction-tuned models are less susceptible to conformity, while increasing the naturalness of majority tones amplifies conformity.Finally, we propose two interventions, Devil's Advocate and Question Distillation, to mitigate conformity, providing insights into building more robust language models. What is the oldest college in Cambridge?It is Peterhouse College. What is the oldest college in Cambridge?King's. Caiqi Zhang, Tom Stafford 0002, Nigel Collier, Andreas Vlachos 0001 |
ACL (1) | 5 |
| 2025 | Social Good or Scientific Curiosity? Uncovering the Research Framing Behind NLP ArtefactsabstractClarifying the research framing of NLP artefacts (e.g., models, datasets, etc.) is crucial to aligning research with practical applications when researchers claim that their findings have real-world impact.Recent studies manually analyzed NLP research across domains, showing that few papers explicitly identify key stakeholders, intended uses, or appropriate contexts.In this work, we propose to automate this analysis, developing a three-component system that infers research framings by first extracting key elements (means, ends, stakeholders), then linking them through interpretable rules and contextual reasoning.We evaluate our approach on two domains: automated factchecking using an existing dataset, and hate speech detection for which we annotate a new dataset 1 -achieving consistent improvements over strong LLM baselines.Finally, we apply our system to recent automated fact-checking papers and uncover three notable trends: a rise in underspecified research goals, increased emphasis on scientific exploration over application, and a shift toward supporting human factcheckers rather than pursuing full automation.General Framing Description AFC HS Automated deployment System replaces a human task with minimal intervention.Automated external fact-checking Automated content moderation Assistive deployment System supports human decision-making.Assisted internal/external fact-checking Assisted content moderation Knowledge access and curation Organizes/synthesizes knowledge for future use.Assisted knowledge curation Assisted knowledge curation Knowledge exploration Explores models or data without specific application goals. Eric Chamoun, Nedjma Ousidhoum, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
EMNLP | 4 |
| 2025 | Improving Zero-shot Sentence Decontextualisation with Content Selection and PlanningabstractExtracting individual sentences from a document as evidence or reasoning steps is commonly done in many NLP tasks.However, extracted sentences often lack context necessary to make them understood, e.g., coreference and background information.To this end, we propose a content selection and planning framework for zero-shot decontextualisation, which determines what content should be mentioned and in what order for a sentence to be understood out of context.Specifically, given a potentially ambiguous sentence and its context, we first segment it into basic semanticallyindependent units.We then identify potentially ambiguous units from the given sentence, and extract relevant units from the context based on their discourse relations.Finally, we generate a content plan to rewrite the sentence by enriching each ambiguous unit with its relevant units.Experimental results demonstrate that our approach is competitive for sentence decontextualisation, producing sentences that exhibit better semantic integrity and discourse coherence, outperforming existing methods. Zhenyun Deng, Yulong Chen 0001, Andreas Vlachos 0001 |
EMNLP | 3 |
| 2025 | TCP: a Benchmark for Temporal Constraint-Based PlanningabstractTemporal reasoning and planning are essential capabilities for large language models (LLMs), yet most existing benchmarks evaluate them in isolation and under limited forms of complexity.To address this gap, we introduce the Temporal Constraint-based Planning (TCP) benchmark, that jointly assesses both capabilities.Each instance in TCP features a naturalistic dialogue around a collaborative project, where diverse and interdependent temporal constraints are explicitly or implicitly expressed, and models must infer an optimal schedule that satisfies all constraints.To construct TCP, we generate abstract problem prototypes that are then paired with realistic scenarios from various domains and enriched into dialogues using an LLM.A human quality check is performed on a sampled subset to confirm the reliability of our benchmark.We evaluate state-of-the-art LLMs and find that even the strongest models may struggle with TCP, highlighting its difficulty and revealing limitations in LLMs' temporal constraint-based planning abilities.We analyze underlying failure cases, open source our benchmark 1 , and hope our findings can inspire future research. Zifeng Ding, Sikuan Yan, Moy Yuan, Xianglong Hu, Fangru Lin, Andreas Vlachos 0001 |
EMNLP | 6 |
| 2025 | TSVer: A Benchmark for Fact Verification Against Time-Series EvidenceabstractReasoning over temporal and numerical data, such as time series, is a crucial aspect of factchecking.While many systems have recently been developed to handle this form of evidence, their evaluation remains limited by existing datasets, which often lack structured evidence, provide insufficient justifications for verdicts, or rely on synthetic claims.In this paper, we introduce TSVER, a new benchmark dataset for fact verification focusing on temporal and numerical reasoning with time-series evidence.TSVER contains 287 real-world claims sourced from 38 fact-checking organizations and a curated database of 400 time series covering diverse domains.Each claim is annotated with time frames across all pertinent time series, along with a verdict and justifications reflecting how the evidence is used to reach the verdict.Using an LLM-assisted multi-step annotation process, we improve the quality of our annotations and achieve an inter-annotator agreement of κ = 0.745 on verdicts.We also develop a baseline for verifying claims against timeseries evidence and show that even the state-ofthe-art reasoning models like Gemini-2.5-Proare challenged by time series, achieving a 63.37 accuracy score on verdicts and an Ev 2 R score of 48.63 on verdict justifications. Marek Strong, Andreas Vlachos 0001 |
EMNLP | 2 |
| 2025 | A Bayesian Optimization Approach to Machine Translation RerankingabstractJulius Cheng, Maike Züfle, Vilém Zouhar, Andreas Vlachos. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Julius Cheng, Maike Züfle, Vilém Zouhar, Andreas Vlachos 0001 |
NAACL (Long Papers) | 4 |
| 2025 | AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the WebabstractTextual claims are often accompanied by images to enhance their credibility and spread on social media, but this also raises concerns about the spread of misinformation.Existing datasets for automated verification of image-text claims remain limited, as they often consist of synthetic claims and lack evidence annotations to capture the reasoning behind the verdict.In this work, we introduce AVerImaTeC, a dataset consisting of 1,297 real-world image-text claims. Each claim is annotated with question-answer (QA) pairs containing evidence from the web, reflecting a decomposed reasoning regarding the verdict.We mitigate common challenges in fact-checking datasets such as contextual dependence, temporal leakage, and evidence insufficiency, via claim normalization, temporally constrained evidence annotation, and a two-stage sufficiency check. We assess the consistency of the annotation in AVerImaTeC via inter-annotator studies, achieving a $\kappa=0.742$ on verdicts and $74.7\%$ consistency on QA pairs. We also propose a novel evaluation method for evidence retrieval and conduct extensive experiments to establish baselines for verifying image-text claims using open-web evidence. Zifeng Ding, Zhijiang Guo, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
NeurIPS | 5 |
| 2025 | Introducing FOReCAst: The Future Outcome Reasoning and Confidence Assessment BenchmarkabstractForecasting is an important task in many domains. However, existing forecasting benchmarks lack comprehensive confidence assessment, focusing on limited question types, and often consist of artificial questions that do not reflect real-world needs. To address these gaps, we introduce FOReCAst (Future Outcome Reasoning and Confidence Assessment), a benchmark that evaluates models' ability to make predictions and their confidence in them. FOReCAst spans diverse forecasting scenarios involving Boolean questions, timeframe prediction, and quantity estimation, enabling a comprehensive evaluation of both prediction accuracy and confidence calibration for real-world applications. Zhangdie Yuan, Zifeng Ding, Andreas Vlachos 0001 |
NeurIPS | 3 |
| 2024 | Document-level Claim Extraction and Decontextualisation for Fact-CheckingabstractSelecting which claims to check is a timeconsuming task for human fact-checkers, especially from documents consisting of multiple sentences and containing multiple claims.However, existing claim extraction approaches focus more on identifying and extracting claims from individual sentences, e.g., identifying whether a sentence contains a claim or the exact boundaries of the claim within a sentence.In this paper, we propose a method for documentlevel claim extraction for fact-checking, which aims to extract check-worthy claims from documents and decontextualise them so that they can be understood out of context.Specifically, we first recast claim extraction as extractive summarization in order to identify central sentences from documents, then rewrite them to include necessary context from the originating document through sentence decontextualisation.Evaluation with both automatic metrics and a fact-checking professional shows that our method is able to extract check-worthy claims from documents more accurately than previous work, while also improving evidence retrieval. Zhenyun Deng, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
ACL (1) | 3 |
| 2024 | Causal Estimation of Memorisation ProfilesabstractUnderstanding memorisation in language models has practical and societal implications, e.g., studying models' training dynamics or preventing copyright infringements.Prior work defines memorisation as the causal effect of training with an instance on the model's ability to predict that instance.This definition relies on a counterfactual: the ability to observe what would have happened had the model not seen that instance.Existing methods struggle to provide computationally efficient and accurate estimates of this counterfactual.Further, they often estimate memorisation for a model architecture rather than for a specific model instance.This paper fills an important gap in the literature, proposing a new, principled, and efficient method to estimate memorisation based on the difference-in-differences design from econometrics.Using this method, we characterise a model's memorisation profile-its memorisation trends across training-by only observing its behaviour on a small set of instances throughout training.In experiments with the Pythia model suite, we find that memorisation (i) is stronger and more persistent in larger models, (ii) is determined by data order and learning rate, and (iii) has stable trends across model sizes, thus making memorisation in larger models predictable from smaller ones. pietrolesci/memorisation-profiles Pietro Lesci, Clara Meister, Thomas Hofmann 0001, Andreas Vlachos 0001, Tiago Pimentel |
ACL (1) | 4 |
| 2024 | The effect of diversity on group decision-making
Georgi Karadzhov, Andreas Vlachos 0001, Tom Stafford 0002 |
CogSci | 2 |
| 2024 | Measuring Uncertainty in Neural Machine Translation with Similarity-Sensitive EntropyabstractUncertainty estimation is an important diagnostic tool for statistical models, and is often used to assess the confidence of model predictions.Previous work shows that neural machine translation (NMT) is an intrinsically uncertain task where there are often multiple correct and semantically equivalent translations, and that well-trained NMT models produce good translations despite spreading probability mass among many semantically similar translations.These findings suggest that popular measures of uncertainty based on token-and sequencelevel entropies which measure surface form diversity may not be good proxies of the more useful quantity of interest, semantic diversity.We propose to adapt similarity-sensitive Shannon entropy (S3E), a concept borrowed from theoretical ecology, for NMT.By demonstrating significantly improved correlation between S3E and task performance on quality estimation and named entity recall, we show that S3E is a useful framework for measuring uncertainty in NMT. Julius Cheng, Andreas Vlachos 0001 |
EACL (1) | 2 |
| 2024 | Do We Need Language-Specific Fact-Checking Models? The Case of ChineseabstractThis paper investigates the potential benefits of language-specific fact-checking models, focusing on the case of Chinese using CHEF dataset.To better reflect real-world fact-checking, we first develop a novel Chinese document-level evidence retriever, achieving state-of-the-art performance.We then demonstrate the limitations of translation-based methods and multilingual language models, highlighting the need for language-specific systems.To better analyze token-level biases in different systems, we construct an adversarial dataset based on the CHEF dataset, where each instance has a large word overlap with the original one but holds the opposite veracity label.Experimental results on the CHEF dataset and our adversarial dataset show that our proposed method outperforms translation-based methods and multilingual language models and is more robust toward biases, emphasizing the importance of language-specific fact-checking systems. 1 Verifiers Retrievers Caiqi Zhang, Zhijiang Guo, Andreas Vlachos 0001 |
EMNLP | 3 |
| 2024 | An LLM Feature-based Framework for Dialogue Constructiveness AssessmentabstractResearch on dialogue constructiveness assessment focuses on (i) analysing conversational factors that influence individuals to take specific actions, win debates, change their perspectives or broaden their open-mindedness and (ii) predicting constructiveness outcomes following dialogues for such use cases.These objectives can be achieved by training either interpretable feature-based models (which often involve costly human annotations) or neural models such as pre-trained language models (which have empirically shown higher task accuracy but lack interpretability).In this paper we propose an LLM feature-based framework for dialogue constructiveness assessment that combines the strengths of feature-based and neural approaches, while mitigating their downsides.The framework first defines a set of dataset-independent and interpretable linguistic features, which can be extracted by both prompting an LLM and simple heuristics.Such features are then used to train LLM featurebased models.We apply this framework to three datasets of dialogue constructiveness and find that our LLM feature-based models outperform or performs at least as well as standard feature-based models and neural models.We also find that the LLM feature-based model learns more robust prediction rules instead of relying on superficial shortcuts, which often trouble neural models. 1 Lexin Zhou, Youmna Farag, Andreas Vlachos 0001 |
EMNLP | 3 |
| 2024 | AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced DatasetsabstractPietro Lesci, Andreas Vlachos. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Pietro Lesci, Andreas Vlachos 0001 |
NAACL-HLT | 2 |
| 2024 | TabVer: Tabular Fact Verification with Natural LogicabstractAbstract Fact verification on tabular evidence incentivizes the use of symbolic reasoning models where a logical form is constructed (e.g., a LISP-style program), providing greater verifiability than fully neural approaches. However, these logical forms typically rely on well-formed tables, restricting their use in many scenarios. An emerging symbolic reasoning paradigm for textual evidence focuses on natural logic inference, which constructs proofs by modeling set-theoretic relations between a claim and its evidence in natural language. This approach provides flexibility and transparency but is less compatible with tabular evidence since the relations do not extend to arithmetic functions. We propose a set-theoretic interpretation of numerals and arithmetic functions in the context of natural logic, enabling the integration of arithmetic expressions in deterministic proofs. We leverage large language models to generate arithmetic expressions by generating questions about salient parts of a claim which are answered by executing appropriate functions on tables. In a few-shot setting on FEVEROUS, we achieve an accuracy of 71.4, outperforming both fully neural and symbolic reasoning models by 3.4 points. When evaluated on TabFact without any further training, our method remains competitive with an accuracy lead of 0.5 points. Rami Aly, Andreas Vlachos 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2024 | AmbiFC: Fact-Checking Ambiguous Claims with EvidenceabstractAbstract Automated fact-checking systems verify claims against evidence to predict their veracity. In real-world scenarios, the retrieved evidence may not unambiguously support or refute the claim and yield conflicting but valid interpretations. Existing fact-checking datasets assume that the models developed with them predict a single veracity label for each claim, thus discouraging the handling of such ambiguity. To address this issue we present AmbiFC,1 a fact-checking dataset with 10k claims derived from real-world information needs. It contains fine-grained evidence annotations of 50k passages from 5k Wikipedia pages. We analyze the disagreements arising from ambiguity when comparing claims against evidence in AmbiFC, observing a strong correlation of annotator disagreement with linguistic phenomena such as underspecification and probabilistic reasoning. We develop models for predicting veracity handling this ambiguity via soft labels, and find that a pipeline that learns the label distribution for sentence-level evidence selection and veracity prediction yields the best performance. We compare models trained on different subsets of AmbiFC and show that models trained on the ambiguous instances perform better when faced with the identified linguistic phenomena. Max Glockner, Ieva Staliunaite, James Thorne, Gisela Vallejo, Andreas Vlachos 0001, Iryna Gurevych |
Trans. Assoc. Comput. Linguistics | 5 |
| 2023 | QA-NatVer: Question Answering for Natural Logic-based Fact VerificationabstractFact verification systems assess a claim's veracity based on evidence.An important consideration in designing them is faithfulness, i.e. generating explanations that accurately reflect the reasoning of the model.Recent works have focused on natural logic, which operates directly on natural language by capturing the semantic relation of spans between an aligned claim with its evidence via set-theoretic operators.However, these approaches rely on substantial resources for training, which are only available for high-resource languages.To this end, we propose to use question answering to predict natural logic operators, taking advantage of the generalization capabilities of instruction-tuned language models.Thus, we obviate the need for annotated training data while still relying on a deterministic inference system.In a few-shot setting on FEVER, our approach outperforms the best baseline by 4.3 accuracy points, including a state-of-the-art pre-trained seq2seq natural logic system, as well as a state-of-the-art prompt-based classifier.Our system demonstrates its robustness and portability, achieving competitive performance on a counterfactual dataset and surpassing all approaches without further annotation on a Danish verification dataset.A human evaluation indicates that our approach produces more plausible proofs with fewer erroneous natural logic operators than previous natural logic-based systems. Rami Aly, Marek Strong, Andreas Vlachos 0001 |
EMNLP | 3 |
| 2023 | Automated Fact-Checking in Dialogue: Are Specialized Models Needed?abstractPrior research has shown that typical factchecking models for stand-alone claims struggle with claims made in dialogues.As a solution, fine-tuning these models on labelled dialogue data has been proposed.However, creating separate models for each use case is impractical, and we show that fine-tuning models for dialogue results in poor performance on typical fact-checking.To overcome this challenge, we present techniques that allow us to use the same models for both dialogue and typical fact-checking.These mainly focus on retrieval adaptation and transforming conversational inputs so that they can be accurately predicted by models trained on stand-alone claims.We demonstrate that a typical fact-checking model incorporating these techniques is competitive with state-of-the-art models fine-tuned for dialogue, while maintaining its accuracy on stand-alone claims. Eric Chamoun, Marzieh Saeidi, Andreas Vlachos 0001 |
EMNLP | 3 |
| 2023 | Faster Minimum Bayes Risk Decoding with Confidence-based PruningabstractMinimum Bayes risk (MBR) decoding outputs the hypothesis with the highest expected utility over the model distribution for some utility function.It has been shown to improve accuracy over beam search in conditional language generation problems and especially neural machine translation, in both human and automatic evaluations.However, the standard samplingbased algorithm for MBR is substantially more computationally expensive than beam search, requiring a large number of samples as well as a quadratic number of calls to the utility function, limiting its applicability.We describe an algorithm for MBR which gradually grows the number of samples used to estimate the utility while pruning hypotheses that are unlikely to have the highest utility according to confidence estimates obtained with bootstrap sampling.Our method requires fewer samples and drastically reduces the number of calls to the utility function compared to standard MBR while being statistically indistinguishable in terms of accuracy.We demonstrate the effectiveness of our approach in experiments on three language pairs, using chrF++ and COMET as utility/evaluation metrics. Julius Cheng, Andreas Vlachos 0001 |
EMNLP | 2 |
| 2023 | AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the WebabstractExisting datasets for automated fact-checking have substantial limitations, such as relying on artificial claims, lacking annotations for evidence and intermediate reasoning, or including evidence published after the claim. In this paper we introduce AVeriTeC, a new dataset of 4,568 real-world claims covering fact-checks by 50 different organizations. Each claim is annotated with question-answer pairs supported by evidence available online, as well as textual justifications explaining how the evidence combines to produce a verdict. Through a multi-round annotation process, we avoid common pitfalls including context dependence, evidence insufficiency, and temporal leakage, and reach a substantial inter-annotator agreement of $\kappa=0.619$ on verdicts. We develop a baseline as well as an evaluation scheme for verifying claims through question-answering against the open web. Michael Sejr Schlichtkrull, Zhijiang Guo, Andreas Vlachos 0001 |
NeurIPS | 3 |
| 2023 | DeliData: A Dataset for Deliberation in Multi-party Problem SolvingabstractGroup deliberation enables people to collaborate and solve problems, however, it is understudied due to a lack of resources. To this end, we introduce the first publicly available dataset containing collaborative conversations on solving a well-established cognitive task, consisting of 500 group dialogues and 14k utterances. In 64% of these conversations, the group members are able to find a better solution than they had identified individually, and in 43.8% of the groups who had a correct answer as their final solution, none of the participants had solved the task correctly by themselves. Furthermore, we propose a novel annotation schema that captures deliberation cues and release all 14k utterances annotated with it. Finally, we use the proposed dataset to develop and evaluate two methods for generating deliberation utterances. The data collection platform, dataset and annotated corpus are publicly available at https://delibot.xyz. Georgi Karadzhov, Tom Stafford 0002, Andreas Vlachos 0001 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2022 | Leveraging Wikipedia article evolution for promotional tone detectionabstractDetecting biased language is useful for a variety of applications, such as identifying hyperpartisan news sources or flagging onesided rhetoric.In this work we introduce WikiEvolve, a dataset for document-level promotional tone detection in English.Unlike previously proposed datasets, it contains seven versions of the same article from Wikipedia, from different points in its revision history; one with promotional tone, and six without it.We adapt the gradient reversal layer framework to encode two article versions simultaneously, and thus leverage the training signal present in the multiple versions.In our experiments, our proposed adaptation of gradient reversal improves the accuracy of four different architectures on both in-domain and outof-domain evaluation. Christine de Kock, Andreas Vlachos 0001 |
ACL (1) | 2 |
| 2022 | Natural Logic-guided Autoregressive Multi-hop Document Retrieval for Fact VerificationabstractA key component of fact verification is the evidence retrieval, often from multiple documents.Recent approaches use dense representations and condition the retrieval of each document on the previously retrieved ones.The latter step is performed over all the documents in the collection, requiring storing their dense representations in an index, thus incurring a high memory footprint.An alternative paradigm is retrieve-and-rerank, where documents are retrieved using methods such as BM25, their sentences are reranked, and further documents are retrieved conditioned on these sentences, reducing the memory requirements.However, such approaches can be brittle as they rely on heuristics and assume hyperlinks between documents.We propose a novel retrieve-and-rerank method for multi-hop retrieval, that consists of a retriever that jointly scores documents in the knowledge source and sentences from previously retrieved documents using an autoregressive formulation and is guided by a proof system based on natural logic that dynamically terminates the retrieval process if the evidence is deemed sufficient.This method is competitive with current state-of-the-art methods on FEVER, HoVer and FEVEROUS-S, while using 5 to 10 times less memory than competing systems.Evaluation on an adversarial dataset indicates improved stability of our approach compared to commonly deployed thresholdbased methods.Finally, the proof system helps humans predict model decisions correctly more often than using the evidence alone. Rami Aly, Andreas Vlachos 0001 |
EMNLP | 2 |
| 2022 | How to disagree well: Investigating the dispute tactics used on WikipediaabstractDisagreements are frequently studied from the perspective of either detecting toxicity or analysing argument structure.We propose a framework of dispute tactics which unifies these two perspectives, as well as other dialogue acts which play a role in resolving disputes, such as asking questions and providing clarification.This framework includes a preferential ordering among rebuttal-type tactics, ranging from ad hominem attacks to refuting the central argument.Using this framework, we annotate 213 disagreements (3,865 utterances) from Wikipedia Talk pages.This allows us to investigate research questions around the tactics used in disagreements; for instance, we provide empirical validation of the approach to disagreement recommended by Wikipedia.We develop models for multilabel prediction of dispute tactics in an utterance, achieving the best performance with a transformer-based label powerset model.Adding an auxiliary task to incorporate the ordering of rebuttal tactics further yields a statistically significant increase.Finally, we show that these annotations can be used to provide useful additional signals to improve performance on the task of predicting escalation. Christine de Kock, Andreas Vlachos 0001 |
EMNLP | 2 |
| 2022 | Varifocal Question Generation for Fact-checkingabstractFact-checking requires retrieving evidence related to a claim under investigation.The task can be formulated as question generation based on a claim, followed by question answering.However, recent question generation approaches assume that the answer is known and typically contained in a passage given as input, whereas such passages are what is being sought when verifying a claim.In this paper, we present Varifocal, a method that generates questions based on different focal points within a given claim, i.e. different spans of the claim and its metadata, such as its source and date.Our method outperforms previous work on a fact-checking question generation dataset on a wide range of automatic evaluation metrics.These results are corroborated by our manual evaluation, which indicates that our method generates more relevant and informative questions.We further demonstrate the potential of focal points in generating sets of clarification questions for product descriptions. Nedjma Ousidhoum, Zhangdie Yuan, Andreas Vlachos 0001 |
EMNLP | 3 |
| 2022 | What makes you change your mind? An empirical investigation in online group decision-making conversationsabstractPeople leverage group discussions to collaborate in order to solve complex tasks, e.g. in project meetings or hiring panels.By doing so, they engage in a variety of conversational strategies where they try to convince each other of the best approach and ultimately reach a decision.In this work, we investigate methods for detecting what makes someone change their mind.To this end, we leverage a recently introduced dataset containing group discussions of people collaborating to solve a task.To find out what makes someone change their mind, we incorporate various techniques such as neural text classification and language-agnostic change point detection.Evaluation of these methods shows that while the task is not trivial, the best way to approach it is using a languageaware model with learning-to-rank training.Finally, we examine the cues that the models develop as indicative of the cause of a change of mind. Georgi Karadzhov, Tom Stafford 0002, Andreas Vlachos 0001 |
SIGDIAL | 3 |
| 2022 | A Survey on Automated Fact-CheckingabstractAbstract Fact-checking has become increasingly important due to the speed with which both information and misinformation can spread in the modern media ecosystem. Therefore, researchers have been exploring how fact-checking can be automated, using techniques based on natural language processing, machine learning, knowledge representation, and databases to automatically predict the veracity of claims. In this paper, we survey automated fact-checking stemming from natural language processing, and discuss its connections to related tasks and disciplines. In this process, we present an overview of existing datasets and models, aiming to unify the various definitions given and identify common concepts. Finally, we highlight challenges for future research. Zhijiang Guo, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | ProoFVer: Natural Logic Theorem Proving for Fact VerificationabstractAbstract Fact verification systems typically rely on neural network classifiers for veracity prediction, which lack explainability. This paper proposes ProoFVer, which uses a seq2seq model to generate natural logic-based inferences as proofs. These proofs consist of lexical mutations between spans in the claim and the evidence retrieved, each marked with a natural logic operator. Claim veracity is determined solely based on the sequence of these operators. Hence, these proofs are faithful explanations, and this makes ProoFVer faithful by construction. Currently, ProoFVer has the highest label accuracy and the second best score in the FEVER leaderboard. Furthermore, it improves by 13.21% points over the next best model on a dataset with counterfactual instances, demonstrating its robustness. As explanations, the proofs show better overlap with human rationales than attention-based highlights and the proofs help humans predict model decisions correctly more often than using the evidence directly.1 Amrith Krishna, Sebastian Riedel 0001, Andreas Vlachos 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Leveraging Type Descriptions for Zero-shot Named Entity Recognition and ClassificationabstractRami Aly, Andreas Vlachos, Ryan McDonald. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Rami Aly, Andreas Vlachos 0001, Ryan McDonald |
ACL/IJCNLP (1) | 2 |
| 2021 | Evidence-based Factual Error CorrectionabstractJames Thorne, Andreas Vlachos. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. James Thorne, Andreas Vlachos 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Incremental Beam Manipulation for Natural Language GenerationabstractThe performance of natural language generation systems has improved substantially with modern neural networks.At test time they typically employ beam search to avoid locally optimal but globally suboptimal predictions.However, due to model errors, a larger beam size can lead to deteriorating performance according to the evaluation metric.For this reason, it is common to rerank the output of beam search, but this relies on beam search to produce a good set of hypotheses, which limits the potential gains.Other alternatives to beam search require changes to the training of the model, which restricts their applicability compared to beam search.This paper proposes incremental beam manipulation, i.e. reranking the hypotheses in the beam during decoding instead of only at the end.This way, hypotheses that are unlikely to lead to a good final output are discarded, and in their place hypotheses that would have been ignored will be considered instead.Applying incremental beam manipulation leads to an improvement of 1.93 and 5.82 BLEU points over vanilla beam search for the test sets of the E2E and WebNLG challenges respectively.The proposed method also outperformed a strong reranker by 1.04 BLEU points on the E2E challenge, while being on par with it on the WebNLG dataset. James Hargreaves, Andreas Vlachos 0001, Guy Emerson |
EACL | 2 |
| 2021 | I Beg to Differ: A study of constructive disagreement in online conversationsabstractDisagreements are pervasive in human communication.In this paper we investigate what makes disagreement constructive.To this end, we construct WikiDisputes, a corpus of 7 425 Wikipedia Talk page conversations that contain content disputes, and define the task of predicting whether disagreements will be escalated to mediation by a moderator.We evaluate feature-based models with linguistic markers from previous work, and demonstrate that their performance is improved by using features that capture changes in linguistic markers throughout the conversations, as opposed to averaged values.We develop a variety of neural models and show that taking into account the structure of the conversation improves predictive accuracy, exceeding that of feature-based models.We assess our best neural model in terms of both predictive accuracy and uncertainty by evaluating its behaviour when it is only exposed to the beginning of the conversation, finding that model accuracy improves and uncertainty reduces as models are exposed to more information. Christine de Kock, Andreas Vlachos 0001 |
EACL | 2 |
| 2021 | Elastic weight consolidation for better bias inoculationabstractThe biases present in training datasets have been shown to affect models for sentence pair classification tasks such as natural language inference (NLI) and fact verification.While fine-tuning models on additional data has been used to mitigate them, a common issue is that of catastrophic forgetting of the original training dataset.In this paper, we show that elastic weight consolidation (EWC) allows finetuning of models to mitigate biases while being less susceptible to catastrophic forgetting.In our evaluation on fact verification and NLI stress tests, we show that fine-tuning with EWC dominates standard fine-tuning, yielding models with lower levels of forgetting on the original (biased) dataset for equivalent gains in accuracy on the fine-tuning (unbiased) dataset. James Thorne, Andreas Vlachos 0001 |
EACL | 2 |
| 2021 | Cross-Policy Compliance Detection via Question AnsweringabstractPolicy compliance detection is the task of ensuring that a scenario conforms to a policy (e.g. a claim is valid according to government rules or a post in an online platform conforms to community guidelines).This task has been previously instantiated as a form of textual entailment, which results in poor accuracy due to the complexity of the policies.In this paper we propose to address policy compliance detection via decomposing it into question answering, where questions check whether the conditions stated in the policy apply to the scenario, and an expression tree combines the answers to obtain the label.Despite the initial upfront annotation cost, we demonstrate that this approach results in better accuracy, especially in the cross-policy setup where the policies during testing are unseen in training.In addition, it allows us to use existing question answering models pre-trained on existing large datasets.Finally, it explicitly identifies the information missing from a scenario in case policy compliance cannot be determined.We conduct our experiments using a recent dataset consisting of government policies, which we augment with expert annotations and find that the cost of annotating question answering decomposition is largely offset by improved interannotator agreement and speed. Marzieh Saeidi, Majid Yazdani, Andreas Vlachos 0001 |
EMNLP (1) | 3 |
| 2020 | Generating Fact Checking BriefsabstractAngela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos, Antoine Bordes, Sebastian Riedel. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Angela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos 0001, Antoine Bordes, Sebastian Riedel 0001 |
EMNLP (1) | 6 |
| 2019 | Merge and Label: A Novel Neural Network Architecture for Nested NERabstractNamed entity recognition (NER) is one of the best studied tasks in natural language processing.However, most approaches are not capable of handling nested structures which are common in many applications.In this paper we introduce a novel neural network architecture that first merges tokens and/or entities into entities forming nested structures, and then labels each of them independently.Unlike previous work, our merge and label approach predicts real-valued instead of discrete segmentation structures, which allow it to combine word and nested entity embeddings while maintaining differentiability.We evaluate our approach using the ACE 2005 Corpus, where it achieves state-of-the-art F1 of 74.6, further improved with contextual embeddings (BERT) to 82.4, an overall improvement of close to 8 F1 points over previous approaches trained on the same data.Additionally we compare it against BiLSTM-CRFs, the dominant approach for flat NER structures, demonstrating that its ability to predict nested structures does not impact performance in simpler cases. 1 Joseph Fisher, Andreas Vlachos 0001 |
ACL (1) | 2 |
| 2019 | HighRES: Highlight-based Reference-less Evaluation of SummarizationabstractThere has been substantial progress in summarization research enabled by the availability of novel, often large-scale, datasets and recent advances on neural network-based approaches.However, manual evaluation of the system generated summaries is inconsistent due to the difficulty the task poses to human non-expert readers.To address this issue, we propose a novel approach for manual evaluation, HIGHlight-based Reference-less Evaluation of Summarization (HIGHRES), in which summaries are assessed by multiple annotators against the source document via manually highlighted salient content in the latter.Thus summary assessment on the source document by human judges is facilitated, while the highlights can be used for evaluating multiple systems.To validate our approach we employ crowd-workers to augment with highlights a recently proposed dataset and compare two state-of-the-art systems.We demonstrate that HIGHRES improves inter-annotator agreement in comparison to using the source document directly, while they help emphasize differences among systems that would be ignored under other evaluation approaches. 1 Hardy, Shashi Narayan, Andreas Vlachos 0001 |
ACL (1) | 3 |
| 2019 | Model-Agnostic Meta-Learning for Relation Classification with Limited SupervisionabstractIn this paper we frame the task of supervised relation classification as an instance of metalearning.We propose a model-agnostic metalearning protocol for training relation classifiers to achieve enhanced predictive performance in limited supervision settings.During training, we aim to not only learn good parameters for classifying relations with sufficient supervision, but also learn model parameters that can be fine-tuned to enhance predictive performance for relations with limited supervision.In experiments conducted on two relation classification datasets, we demonstrate that the proposed meta-learning approach improves the predictive performance of two state-of-the-art supervised relation classification models. Abiola Obamuyide, Andreas Vlachos 0001 |
ACL (1) | 2 |
| 2019 | Incorporating Label Dependencies in Multilabel Stance DetectionabstractWilliam Ferreira, Andreas Vlachos. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. William Ferreira 0001, Andreas Vlachos 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Neural Generative Rhetorical Structure ParsingabstractAmandla Mabona, Laura Rimell, Stephen Clark, Andreas Vlachos. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Amandla Mabona, Laura Rimell, Stephen Clark, Andreas Vlachos 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Evaluating adversarial attacks against multiple fact verification systemsabstractJames Thorne, Andreas Vlachos, Christos Christodoulopoulos, Arpit Mittal. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. James Thorne, Andreas Vlachos 0001, Christos Christodoulopoulos 0001, Arpit Mittal |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Automated Fact Checking in the News RoomabstractFact checking is an essential task in journalism; its importance has been highlighted due to recently increased concerns and efforts in combating misinformation. In this paper, we present an automated fact checking platform which given a claim, it retrieves relevant textual evidence from a document collection, predicts whether each piece of evidence supports or refutes the claim, and returns a final verdict. We describe the architecture of the system and the user interface, focusing on the choices made to improve its user friendliness and transparency. We conduct a user study of the fact-checking platform in a journalistic setting: we integrated it with a collection of news articles and provide an evaluation of the platform using feedback from journalists in their workflow. We found that the predictions of our platform were correct 58% of the time, and 59% of the returned evidence was relevant. Sebastião Miranda, David Nogueira, Afonso Mendes, Andreas Vlachos 0001, Andrew Secker, Rebecca Garrett, Jeff Mitchell 0001, Zita Marinho |
WWW | 4 |
| 2018 | Topic or Style? Exploring the Most Useful Features for Authorship AttributionabstractApproaches to authorship attribution, the task of identifying the author of a document, are based on analysis of individuals’ writing style and/or preferred topics. Although the problem has been widely explored, no previous studies have analysed the relationship between dataset characteristics and effectiveness of different types of features. This study carries out an analysis of four widely used datasets to explore how different types of features affect authorship attribution accuracy under varying conditions. The results of the analysis are applied to authorship attribution models based on both discrete and continuous representations. We apply the conclusions from our analysis to an extension of an existing approach to authorship attribution and outperform the prior state-of-the-art on two out of the four datasets used. Yunita Sari, Andreas Vlachos 0001 |
COLING | 3 |
| 2018 | Automated Fact Checking: Task Formulations, Methods and Future DirectionsabstractThe recently increased focus on misinformation has stimulated research in fact checking, the task of assessing the truthfulness of a claim. Research in automating this task has been conducted in a variety of disciplines including natural language processing, machine learning, knowledge representation, databases, and journalism. While there has been substantial progress, relevant papers and articles have been published in research communities that are often unaware of each other and use inconsistent terminology, thus impeding understanding and further progress. In this paper we survey automated fact checking research stemming from natural language processing and related disciplines, unifying the task formulations and methodologies across papers and authors. Furthermore, we highlight the use of evidence as an important distinguishing factor among them cutting across task formulations and methods. We conclude with proposing avenues for future NLP research on automated fact checking. James Thorne, Andreas Vlachos 0001 |
COLING | 2 |
| 2018 | Guided Neural Language Generation for Abstractive Summarization using Abstract Meaning RepresentationabstractRecent work on abstractive summarization has made progress with neural encoder-decoder architectures.However, such models are often challenged due to their lack of explicit semantic modeling of the source document and its summary.In this paper, we extend previous work on abstractive summarization using Abstract Meaning Representation (AMR) with a neural language generation stage which we guide using the source document.We demonstrate that this guidance improves summarization results by 7.4 and 10.5 points in ROUGE-2 using gold standard AMR parses and parses obtained from an off-the-shelf parser respectively.We also find that the summarization performance using the latter is 2 ROUGE-2 points higher than that of a well-established neural encoderdecoder approach trained on a larger dataset. Hardy, Andreas Vlachos 0001 |
EMNLP | 2 |
| 2018 | FEVER: a Large-scale Dataset for Fact Extraction and VERificationabstractJames Thorne, Andreas Vlachos, Christos Christodoulopoulos, Arpit Mittal. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. James Thorne, Andreas Vlachos 0001, Christos Christodoulopoulos 0001, Arpit Mittal |
NAACL-HLT | 2 |
| 2016 | Noise reduction and targeted exploration in imitation learning for Abstract Meaning Representation parsingabstractSemantic parsers map natural language statements into meaning representations, and must abstract over syntactic phenomena, resolve anaphora, and identify word senses to eliminate ambiguous interpretations.Abstract meaning representation (AMR) is a recent example of one such semantic formalism which, similar to a dependency parse, utilizes a graph to represent relationships between concepts (Banarescu et al., 2013).As with dependency parsing, transition-based approaches are a common approach to this problem.However, when trained in the traditional manner these systems are susceptible to the accumulation of errors when they find undesirable states during greedy decoding.Imitation learning algorithms have been shown to help these systems recover from such errors.To effectively use these methods for AMR parsing we find it highly beneficial to introduce two novel extensions: noise reduction and targeted exploration.The former mitigates the noise in the feature representation, a result of the complexity of the task.The latter targets the exploration steps of imitation learning towards areas which are likely to provide the most information in the context of a large action-space.We achieve state-ofthe art results, and improve upon standard transition-based parsing by 4.7 F 1 points. Andreas Vlachos 0001, Jason Naradowsky |
ACL (1) | 2 |
| 2016 | Imitation learning for language generation from unaligned dataabstractNatural language generation (NLG) is the task of generating natural language from a meaning representation. Current rule-based approaches require domain-specific and manually constructed linguistic resources, while most machine-learning based approaches rely on aligned training data and/or phrase templates. The latter are needed to restrict the search space for the structured prediction task defined by the unaligned datasets. In this work we propose the use of imitation learning for structured prediction which learns an incremental model that handles the large search space by avoiding explicit enumeration of the outputs. We focus on the Locally Optimal Learning to Search framework which allows us to train against non-decomposable loss functions such as the BLEU or ROUGE scores while not assuming gold standard alignments. We evaluate our approach on three datasets using both automatic measures and human judgements and achieve results comparable to the state-of-the-art approaches developed for each of them. Gerasimos Lampouras, Andreas Vlachos 0001 |
COLING | 2 |
| 2016 | Stance Detection with Bidirectional Conditional EncodingabstractStance detection is the task of classifying the attitude expressed in a text towards a target such as Hillary Clinton to be "positive", negative" or "neutral". Previous work has assumed that either the target is mentioned in the text or that training data for every target is given. This paper considers the more challenging version of this task, where targets are not always mentioned and no training data is available for the test targets. We experiment with conditional LSTM encoding, which builds a representation of the tweet that is dependent on the target, and demonstrate that it outperforms encoding the tweet and the target independently. Performance is improved further when the conditional model is augmented with bidirectional encoding. We evaluate our approach on the SemEval 2016 Task 6 Twitter Stance Detection corpus achieving performance second best only to a system trained on semi-automatically labelled tweets for the test target. When such weak supervision is added, our approach achieves state-of-the-art results. Isabelle Augenstein, Tim Rocktäschel, Andreas Vlachos 0001, Kalina Bontcheva |
EMNLP | 3 |
| 2016 | Timeline extraction using distant supervision and joint inferenceabstractIn timeline extraction the goal is to order all the events in which a target entity is involved in a timeline.Due to the lack of explicitly annotated data, previous work is primarily rule-based and uses pre-trained temporal linking systems.In this work, we propose a distantly supervised approach by heuristically aligning timelines with documents.The noisy training data created allows us to learn models that anchor events to temporal expressions and entities; during testing, the predictions of these models are combined to produce the timeline.Furthermore, we show how to improve performance using joint inference.In experiments in the SemEval-2015 TimeLine task we show that our distantly supervised approach matches the state-of-the-art performance while joint inference further improves on it by 3.2 F-score points. Savelie Cornegruta, Andreas Vlachos 0001 |
EMNLP | 2 |
| 2016 | Emergent: a novel data-set for stance classificationabstractWe present Emergent, a novel data-set derived from a digital journalism project for rumour debunking.The data-set contains 300 rumoured claims and 2,595 associated news articles, collected and labelled by journalists with an estimation of their veracity (true, false or unverified).Each associated article is summarized into a headline and labelled to indicate whether its stance is for, against, or observing the claim, where observing indicates that the article merely repeats the claim.Thus, Emergent provides a real-world data source for a variety of natural language processing tasks in the context of fact-checking.Further to presenting the dataset, we address the task of determining the article headline stance with respect to the claim.For this purpose we use a logistic regression classifier and develop features that examine the headline and its agreement with the claim.The accuracy achieved was 73% which is 26% higher than the one achieved by the Excitement Open Platform (Magnini et al., 2014). William Ferreira 0001, Andreas Vlachos 0001 |
HLT-NAACL | 2 |
| 2015 | Extracting Relations between Non-Standard Entities using Distant Supervision and Imitation LearningabstractDistantly supervised approaches have become popular in recent years as they allow training relation extractors without textbound annotation, using instead known relations from a knowledge base and a large textual corpus from an appropriate domain.While state of the art distant supervision approaches use off-theshelf named entity recognition and classification (NERC) systems to identify relation arguments, discrepancies in domain or genre between the data used for NERC training and the intended domain for the relation extractor can lead to low performance.This is particularly problematic for "non-standard" named entities such as album which would fall into the MISC category.We propose to ameliorate this issue by jointly training the named entity classifier and the relation extractor using imitation learning which reduces structured prediction learning to classification learning.We further experiment with Web features different features and compare against using two off-the-shelf supervised NERC systems, Stanford NER and FIGER, for named entity classification.Our experiments show that imitation learning improves average precision by 4 points over an one-stage classification model, while removing Web features results in a 6 points reduction.Compared to using FIGER and Stanford NER, average precision is 10 points and 19 points higher with our imitation learning approach. Isabelle Augenstein, Andreas Vlachos 0001, Diana Maynard |
EMNLP | 2 |
| 2015 | A Strong Lexical Matching Method for the Machine Comprehension TestabstractMachine comprehension of text is the overarching goal of a great deal of research in natural language processing.The Machine Comprehension Test (Richardson et al., 2013) was recently proposed to assess methods on an open-domain, extensible, and easy-to-evaluate task consisting of two datasets.In this paper we develop a lexical matching method that takes into account multiple context windows, question types and coreference resolution.We show that the proposed method outperforms the baseline of Richardson et al. (2013), and despite its relative simplicity, is comparable to recent work using machine learning.We hope that our approach will inform future work on this task.Furthermore, we argue that MC500 is harder than MC160 due to the way question answer pairs were created. Ellery Smith, Nicola Greco, Matko Bosnjak, Andreas Vlachos 0001 |
EMNLP | 4 |
| 2015 | Identification and Verification of Simple Claims about Statistical PropertiesabstractIn this paper we study the identification and verification of simple claims about statistical properties, e.g.claims about the population or the inflation rate of a country.We show that this problem is similar to extracting numerical information from text and following recent work, instead of annotating data for each property of interest in order to learn supervised models, we develop a distantly supervised baseline approach using a knowledge base and raw text.In experiments on 16 statistical properties about countries from Freebase we show that our approach identifies simple statistical claims about properties with 60% precision, while it is able to verify these claims without requiring any explicit supervision for either tasks.Furthermore, we evaluate our approach as a statistical property extractor and we show it achieves 0.11 mean absolute percentage error. Andreas Vlachos 0001, Sebastian Riedel 0001 |
EMNLP | 1 |
| 2014 | A New Corpus and Imitation Learning Framework for Context-Dependent Semantic ParsingabstractSemantic parsing is the task of translating natural language utterances into a machine-interpretable meaning representation. Most approaches to this task have been evaluated on a small number of existing corpora which assume that all utterances must be interpreted according to a database and typically ignore context. In this paper we present a new, publicly available corpus for context-dependent semantic parsing. The MRL used for the annotation was designed to support a portable, interactive tourist information system. We develop a semantic parser for this corpus by adapting the imitation learning algorithm DAgger without requiring alignment information during training. DAgger improves upon independently trained classifiers by 9.0 and 4.8 points in F-score on the development and test sets respectively. Andreas Vlachos 0001, Stephen Clark |
Trans. Assoc. Comput. Linguistics | 1 |
| 2013 | Dependency Language Models for Sentence CompletionabstractSentence completion is a challenging semantic modeling task in which models must choose the most appropriate word from a given set to complete a sentence.Although a variety of language models have been applied to this task in previous work, none of the existing approaches incorporate syntactic information.In this paper we propose to tackle this task using a pair of simple language models in which the probability of a sentence is estimated as the probability of the lexicalisation of a given syntactic dependency tree.We apply our approach to the Microsoft Research Sentence Completion Challenge and show that it improves on n-gram language models by 8.7 percentage points, achieving the highest accuracy reported to date apart from neural language models that are more complex and expensive to train. Joseph Gubbins, Andreas Vlachos 0001 |
EMNLP | 2 |
| 2012 | Biomedical event extraction from abstracts and full papers using search-based structured predictionabstractBACKGROUND: Biomedical event extraction has attracted substantial attention as it can assist researchers in understanding the plethora of interactions among genes that are described in publications in molecular biology. While most recent work has focused on abstracts, the BioNLP 2011 shared task evaluated the submitted systems on both abstracts and full papers. In this article, we describe our submission to the shared task which decomposes event extraction into a set of classification tasks that can be learned either independently or jointly using the search-based structured prediction framework. Our intention is to explore how these two learning paradigms compare in the context of the shared task. RESULTS: We report that models learned using search-based structured prediction exceed the accuracy of independently learned classifiers by 8.3 points in F-score, with the gains being more pronounced on the more complex Regulation events (13.23 points). Furthermore, we show how the trade-off between recall and precision can be adjusted in both learning paradigms and that search-based structured prediction achieves better recall at all precision points. Finally, we report on experiments with a simple domain-adaptation method, resulting in the second-best performance achieved by a single system. CONCLUSIONS: We demonstrate that joint inference using the search-based structured prediction framework can achieve better performance than independently learned classifiers, thus demonstrating the potential of this learning paradigm for event extraction and other similarly complex information-extraction tasks. Andreas Vlachos 0001, Mark Craven |
BMC Bioinform. | 1 |
| 2011 | Search-based Structured Prediction applied to Biomedical Event Extraction
Andreas Vlachos 0001, Mark Craven |
CoNLL | 1 |
| 2009 | The infinite HMM for unsupervised PoS tagging
Jurgen Van Gael, Andreas Vlachos 0001, Zoubin Ghahramani |
EMNLP | 2 |
| 2008 | Natural Language Processing in aid of FlyBase curatorsabstractBACKGROUND: Despite increasing interest in applying Natural Language Processing (NLP) to biomedical text, whether this technology can facilitate tasks such as database curation remains unclear. RESULTS: PaperBrowser is the first NLP-powered interface that was developed under a user-centered approach to improve the way in which FlyBase curators navigate an article. In this paper, we first discuss how observing curators at work informed the design and evaluation of PaperBrowser. Then, we present how we appraise PaperBrowser's navigational functionalities in a user-based study using a text highlighting task and evaluation criteria of Human-Computer Interaction. Our results show that PaperBrowser reduces the amount of interactions between two highlighting events and therefore improves navigational efficiency by about 58% compared to the navigational mechanism that was previously available to the curators. Moreover, PaperBrowser is shown to provide curators with enhanced navigational utility by over 74% irrespective of the different ways in which they highlight text in the article. CONCLUSION: We show that state-of-the-art performance in certain NLP tasks such as Named Entity Recognition and Anaphora Resolution can be combined with the navigational functionalities of PaperBrowser to support curation quite successfully. Nikiforos Karamanis, Ruth L. Seal, Ian Lewin, Peter McQuilton, Andreas Vlachos 0001, Caroline Gasperin, Rachel A. Drysdale, Ted Briscoe |
BMC Bioinform. | 5 |
| 2008 | A stopping criterion for active learning
Andreas Vlachos 0001 |
Comput. Speech Lang. | 1 |