VLDB 2026 Research / reviewers in the wild / expert
Michael Sejr Schlichtkrull
dblp:186/7091
· DBLP profile ↗
13ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-8666-0856ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ev2R: Evaluating Evidence Retrieval in Automated Fact-CheckingabstractAbstract Current automated fact-checking (AFC) approaches typically evaluate evidence either implicitly via the predicted verdicts or through exact matches with predefined closed knowledge sources, such as Wikipedia. However, these methods are limited due to their reliance on evaluation metrics originally designed for other purposes and constraints from closed knowledge sources. In this work, we introduce Ev2R which combines the strengths of reference-based evaluation and verdict-level proxy scoring. Ev2R jointly assesses how well the evidence aligns with the gold references and how reliably it supports the verdict, addressing the shortcomings of prior methods. We evaluate Ev2R against three types of evidence evaluation approaches: reference-based, proxy-reference, and reference-less baselines. Assessments against human ratings and adversarial tests demonstrate that Ev2R consistently outperforms existing scoring approaches in accuracy and robustness. It achieves stronger correlation with human judgments and greater robustness to adversarial perturbations, establishing it as a reliable metric for evidence evaluation in AFC.1 Mubashara Akhtar, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | Social Good or Scientific Curiosity? Uncovering the Research Framing Behind NLP ArtefactsabstractClarifying the research framing of NLP artefacts (e.g., models, datasets, etc.) is crucial to aligning research with practical applications when researchers claim that their findings have real-world impact.Recent studies manually analyzed NLP research across domains, showing that few papers explicitly identify key stakeholders, intended uses, or appropriate contexts.In this work, we propose to automate this analysis, developing a three-component system that infers research framings by first extracting key elements (means, ends, stakeholders), then linking them through interpretable rules and contextual reasoning.We evaluate our approach on two domains: automated factchecking using an existing dataset, and hate speech detection for which we annotate a new dataset 1 -achieving consistent improvements over strong LLM baselines.Finally, we apply our system to recent automated fact-checking papers and uncover three notable trends: a rise in underspecified research goals, increased emphasis on scientific exploration over application, and a shift toward supporting human factcheckers rather than pursuing full automation.General Framing Description AFC HS Automated deployment System replaces a human task with minimal intervention.Automated external fact-checking Automated content moderation Assistive deployment System supports human decision-making.Assisted internal/external fact-checking Assisted content moderation Knowledge access and curation Organizes/synthesizes knowledge for future use.Assisted knowledge curation Assisted knowledge curation Knowledge exploration Explores models or data without specific application goals. Eric Chamoun, Nedjma Ousidhoum, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
EMNLP | 3 |
| 2025 | Attacks by Content: Automated Fact-checking is an AI Security IssueabstractWhen AI agents retrieve and reason over external documents, adversaries can manipulate the data they receive to subvert their behaviour.Previous research has studied indirect prompt injection, where the attacker injects malicious instructions.We argue that injection of instructions is not necessary to manipulate agentsattackers could instead supply biased, misleading, or false information.We term this an attack by content.Existing defenses, which focus on detecting hidden commands, are ineffective against attacks by content.To defend themselves and their users, agents must critically evaluate retrieved information, corroborating claims with external evidence and evaluating source trustworthiness.We argue that this is analogous to an existing NLP task, automated fact-checking, which we propose to repurpose as a cognitive self-defense tool for agents. Michael Sejr Schlichtkrull |
EMNLP | 1 |
| 2025 | AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the WebabstractTextual claims are often accompanied by images to enhance their credibility and spread on social media, but this also raises concerns about the spread of misinformation.Existing datasets for automated verification of image-text claims remain limited, as they often consist of synthetic claims and lack evidence annotations to capture the reasoning behind the verdict.In this work, we introduce AVerImaTeC, a dataset consisting of 1,297 real-world image-text claims. Each claim is annotated with question-answer (QA) pairs containing evidence from the web, reflecting a decomposed reasoning regarding the verdict.We mitigate common challenges in fact-checking datasets such as contextual dependence, temporal leakage, and evidence insufficiency, via claim normalization, temporally constrained evidence annotation, and a two-stage sufficiency check. We assess the consistency of the annotation in AVerImaTeC via inter-annotator studies, achieving a $\kappa=0.742$ on verdicts and $74.7\%$ consistency on QA pairs. We also propose a novel evaluation method for evidence retrieval and conduct extensive experiments to establish baselines for verifying image-text claims using open-web evidence. Zifeng Ding, Zhijiang Guo, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
NeurIPS | 4 |
| 2024 | Document-level Claim Extraction and Decontextualisation for Fact-CheckingabstractSelecting which claims to check is a timeconsuming task for human fact-checkers, especially from documents consisting of multiple sentences and containing multiple claims.However, existing claim extraction approaches focus more on identifying and extracting claims from individual sentences, e.g., identifying whether a sentence contains a claim or the exact boundaries of the claim within a sentence.In this paper, we propose a method for documentlevel claim extraction for fact-checking, which aims to extract check-worthy claims from documents and decontextualise them so that they can be understood out of context.Specifically, we first recast claim extraction as extractive summarization in order to identify central sentences from documents, then rewrite them to include necessary context from the originating document through sentence decontextualisation.Evaluation with both automatic metrics and a fact-checking professional shows that our method is able to extract check-worthy claims from documents more accurately than previous work, while also improving evidence retrieval. Zhenyun Deng, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
ACL (1) | 2 |
| 2023 | Are Embedded Potatoes Still Vegetables? On the Limitations of WordNet Embeddings for Lexical SemanticsabstractKnowledge Base Embedding (KBE) models have been widely used to encode structured information from knowledge bases, including WordNet.However, the existing literature has predominantly focused on link prediction as the evaluation task, often neglecting exploration of the models' semantic capabilities.In this paper, we investigate the potential disconnect between the performance of KBE models of WordNet on link prediction and their ability to encode semantic information, highlighting the limitations of current evaluation protocols.Our findings reveal that some top-performing KBE models on the WN18RR benchmark exhibit subpar results on two semantic tasks and two downstream tasks.These results demonstrate the inadequacy of link prediction benchmarks for evaluating the semantic capabilities of KBE models, suggesting the need for a more targeted assessment approach. Xuyou Cheng, Michael Sejr Schlichtkrull, Guy Emerson |
EMNLP | 2 |
| 2023 | AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the WebabstractExisting datasets for automated fact-checking have substantial limitations, such as relying on artificial claims, lacking annotations for evidence and intermediate reasoning, or including evidence published after the claim. In this paper we introduce AVeriTeC, a new dataset of 4,568 real-world claims covering fact-checks by 50 different organizations. Each claim is annotated with question-answer pairs supported by evidence available online, as well as textual justifications explaining how the evidence combines to produce a verdict. Through a multi-round annotation process, we avoid common pitfalls including context dependence, evidence insufficiency, and temporal leakage, and reach a substantial inter-annotator agreement of $\kappa=0.619$ on verdicts. We develop a baseline as well as an evaluation scheme for verifying claims through question-answering against the open web. Michael Sejr Schlichtkrull, Zhijiang Guo, Andreas Vlachos 0001 |
NeurIPS | 1 |
| 2022 | A Survey on Automated Fact-CheckingabstractAbstract Fact-checking has become increasingly important due to the speed with which both information and misinformation can spread in the modern media ecosystem. Therefore, researchers have been exploring how fact-checking can be automated, using techniques based on natural language processing, machine learning, knowledge representation, and databases to automatically predict the veracity of claims. In this paper, we survey automated fact-checking stemming from natural language processing, and discuss its connections to related tasks and disciplines. In this process, we present an overview of existing datasets and models, aiming to unify the various definitions given and identify common concepts. Finally, we highlight challenges for future research. Zhijiang Guo, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | Joint Verification and Reranking for Open Fact Checking Over TablesabstractMichael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Wen-tau Yih, Sebastian Riedel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Michael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Scott Yih, Sebastian Riedel 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | Interpreting Graph Neural Networks for NLP With Differentiable Edge Masking
Michael Sejr Schlichtkrull, Nicola De Cao, Ivan Titov 0001 |
ICLR | 1 |
| 2020 | How do Decisions Emerge across Layers in Neural Models? Interpretation with Differentiable MaskingabstractAttribution methods assess the contribution of inputs to the model prediction.One way to do so is erasure: a subset of inputs is considered irrelevant if it can be removed without affecting the prediction.Though conceptually simple, erasure's objective is intractable and approximate search remains expensive with modern deep NLP models.Erasure is also susceptible to the hindsight bias: the fact that an input can be dropped does not mean that the model 'knows' it can be dropped.The resulting pruning is over-aggressive and does not reflect how the model arrives at the prediction.To deal with these challenges, we introduce Differentiable Masking.DIFFMASK learns to maskout subsets of the input while maintaining differentiability.The decision to include or disregard an input token is made with a simple model based on intermediate hidden layers of the analyzed model.First, this makes the approach efficient because we predict rather than search.Second, as with probing classifiers, this reveals what the network 'knows' at the corresponding layers.This lets us not only plot attribution heatmaps but also analyze how decisions are formed across network layers.We use DIFFMASK to study BERT models on sentiment classification and question answering.1Question: Where did the Broncos practice for the Super Bowl ? Nicola De Cao, Michael Sejr Schlichtkrull, Wilker Aziz, Ivan Titov 0001 |
EMNLP (1) | 2 |
| 2018 | Modeling Relational Data with Graph Convolutional Networks
Michael Sejr Schlichtkrull, Thomas Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov 0001, Max Welling |
ESWC | 1 |
| 2017 | Cross-Lingual Dependency Parsing with Late Decoding for Truly Low-Resource LanguagesabstractIn cross-lingual dependency annotation projection, information is often lost during transfer because of early decoding.We present an end-to-end graph-based neural network dependency parser that can be trained to reproduce matrices of edge scores, which can be directly projected across word alignments.We show that our approach to cross-lingual dependency parsing is not only simpler, but also achieves an absolute improvement of 2.25% averaged across 10 languages compared to the previous state of the art. Michael Sejr Schlichtkrull, Anders Søgaard |
EACL (1) | 1 |