EDBT 2026 Demo / reviewers in the wild / expert
Konstantin Todorov
dblp:36/6753
· DBLP profile ↗
20ranked-venue papers in the field
2as first author
8since 2021 · last 2026
0000-0002-9116-6692ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 11Information Retrieval & Web Search · 7Other / Interdisciplinary · 2 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The CLEF-2026 CheckThat! Lab: Advancing Multilingual Fact-Checking
Julia Maria Struß, Sebastian Schellhammer, Stefan Dietze, Venktesh V., Vinay Setty, Tanmoy Chakraborty 0002, Preslav Nakov, Avishek Anand, Primakov Chungkham, Salim Hafid, Dhruv Sahnan, Konstantin Todorov |
ECIR (4) | 12 |
| 2026 | Linguistic Signatures for Enhanced Emotion DetectionabstractEmotion detection is a central problem in NLP, with recent progress driven by transformer-based models trained on established datasets. However, little is known about the linguistic regularities that characterize how emotions are expressed across different corpora and labels. This study examines whether linguistic features can serve as reliable interpretable signals for emotion recognition in text. We extract emotion-specific linguistic signatures from 13 English datasets and evaluate how incorporating these features into transformer models impacts performance. Our RoBERTa-based models enriched with high level linguistic features achieve consistent performance gains of up to +2.4 macro F1 on the GoEmotions benchmark, showing that explicit lexical cues can complement neural representations and improve robustness in predicting emotion categories. Florian Lecourt, Madalina Croitoru, Konstantin Todorov |
WWW | 3 |
| 2025 | The CLEF-2025 CheckThat! Lab: Subjectivity, Fact-Checking, Claim Normalization, and Retrieval
Firoj Alam, Julia Maria Struß, Tanmoy Chakraborty 0002, Stefan Dietze, Salim Hafid, Katerina Korre, Arianna Muti, Preslav Nakov, Federico Ruggeri, Sebastian Schellhammer, Vinay Setty, Megha Sundriyal, Konstantin Todorov, Venktesh V |
ECIR (5) | 13 |
| 2025 | Graph Embeddings Meet Link Keys Discovery for Entity MatchingabstractEntity Matching (EM) automates the discovery of identity links between entities within different Knowledge Graphs (KGs). Link keys are crucial for EM, serving as rules allowing to identify identity links across different KGs, possibly described using different ontologies. However, the approach for extracting link keys struggles to scale on large KGs. While embedding-based EM methods efficiently handle large KGs they lack explainability. This paper proposes a novel hybrid EM approach to guarantee the scalability link key extraction approach and improve the explainability of embedding-based EM methods. First, embedding-based EM approaches are used to sample the KGs based on the identity links they generate, thereby reducing the search space to relevant sub-graphs for link key extraction. Second, rules (in the form of link keys) are extracted to explain the generation of identity links by the embedding-based methods. Experimental results demonstrate that the proposed approach allows link key extraction to scale on large KGs, preserving the quality of the extracted link keys. Additionally, it shows that link keys can improve the explainability of the identity links generated by embedding-methods, allowing for the regeneration of 77% of the identity links produced for a specific EM task, thereby providing an approximation of the reasons behind their generation. Chloé Khadija Jradeh, Ensiyeh Raoufi, Jérôme David, Pierre Larmande, François Scharffe, Konstantin Todorov, Cássia Trojahn dos Santos |
WWW | 6 |
| 2025 | An In-depth Analysis of the Linguistic Characteristics of Science Claims on the Web and their Impact on Fact-checkingabstractWeb claims, seen as assertions shared on the web and eligible for fact-checking, are at the heart of online discourse. They have been studied extensively on a variety of downstream tasks such as fact-checking, claim retrieval, bias detection, argument mining, or viewpoint discovery. On the other hand, claims originating from scientific publications have also been the subject of several downstream NLP tasks. However, research carried out so far has yet to focus on scientific web claims, which are scientific claims made on the web (e.g., on social media and news articles). The process of detecting and fact-checking a claim from the web can be very different depending on whether the claim is scientific or not, thus making it crucial for the developed datasets, methods, and models to make a distinction between the two. With this work, we aim at understanding what makes this distinction necessary, by understanding the linguistic differences between scientific and non-scientific claims on the web, and the impact those differences have on existing downstream tasks. To do so, we manually annotate 1,524 web claims from established benchmarks for fact-checking-related tasks, and we run statistical tests to analyze and compare the linguistic features of each group. We find that scientific claims on the web use more analytical speech, but also use more sentiment-related speech, more expressions of physical motion, and have distinct parts of speech (PoS) and punctuation styles. We also conduct experiments showing that BERT-based language models perform worse on scientific web claims by up to 17 F1 points for several downstream tasks. To understand why, we develop a novel methodology to map predictive tokens of language models to explainable linguistic features and find that language models fail to detect a specific subset of predictive features of scientific web claims. We conclude by stating that language models aimed at studying scientific web claims ought to be trained on scientific web discourse, as opposed to being trained only on generic web discourse or only on scientific text from scientific publications. Salim Hafid, Sebastian Schellhammer, Yavuz Selim Kartal, Thomas Papastergiou, Stefan Dietze, Sandra Bringay, Konstantin Todorov |
ACM Trans. Web | 7 |
| 2022 | SciTweets - A Dataset and Annotation Framework for Detecting Scientific Online DiscourseabstractScientific topics, claims and resources are increasingly debated as part of online discourse, where prominent examples include discourse related to COVID-19 or climate change. This has led to both significant societal impact and increased interest in scientific online discourse from various disciplines. For instance, communication studies aim at a deeper understanding of biases, quality or spreading patterns of scientific information, whereas computational methods have been proposed to extract, classify or verify scientific claims using NLP and IR techniques. However, research across disciplines currently suffers from both a lack of robust definitions of the various forms of science-relatedness as well as appropriate ground truth data for distinguishing them. In this work, we contribute (a) an annotation framework and corresponding definitions for different forms of scientific relatedness of online discourse in tweets, (b) an expert-annotated dataset of 1261 tweets obtained through our labeling framework reaching an average Fleiss Kappa κ of 0.63, (c) a multi-label classifier trained on our data able to detect science- relatedness with 89% F1 and also able to detect distinct forms of scientific knowledge (claims, references). With this work, we aim to lay the foundation for developing and evaluating robust methods for analysing science as part of large-scale online discourse. Salim Hafid, Sebastian Schellhammer, Sandra Bringay, Konstantin Todorov, Stefan Dietze |
CIKM | 4 |
| 2022 | IMGT-KG: A Knowledge Graph for Immunogenetics
Gaoussou Sanou, Véronique Giudicelli, Nika Abdollahi, Sophia Kossida, Konstantin Todorov, Patrice Duroux |
ISWC | 5 |
| 2021 | AgroLD: A Knowledge Graph for the Plant Sciences
Pierre Larmande, Konstantin Todorov |
ISWC | 2 |
| 2019 | ClaimsKG: A Knowledge Graph of Fact-Checked ClaimsabstractVarious research areas at the intersection of computer and social sciences require a ground truth of contextualized claims labelled with their truth values in order to facilitate supervision, validation or reproducibility of approaches dealing, for example, with fact-checking or analysis of societal debates. So far, no reasonably large, up-to-date and queryable corpus of structured information about claims and related metadata is publicly available. In an attempt to fill this gap, we introduce ClaimsKG, a knowledge graph of fact-checked claims, which facilitates structured queries about their truth values, authors, dates, journalistic reviews and other kinds of metadata. ClaimsKG is generated through a semi-automated pipeline, which harvests data from popular fact-checking websites on a regular basis, annotates claims with related entities from DBpedia, and lifts the data to RDF using an RDF/S model that makes use of established vocabularies. In order to harmonise data originating from diverse fact-checking sites, we introduce normalised ratings as well as a simple claims coreference resolution strategy. The current knowledge graph, extensible to new information, consists of 28,383 claims published since 1996, amounting to 6,606,032 triples. Andon Tchechmedjiev, Pavlos Fafalios, Katarina Boland, Malo Gasquet, Matthäus Zloch, Benjamin Zapilko, Stefan Dietze, Konstantin Todorov |
ISWC (2) | 8 |
| 2019 | Linking and disambiguating entities across heterogeneous RDF graphs
Manel Achichi, Zohra Bellahsene, Mohamed Ben Ellefi, Konstantin Todorov |
J. Web Semant. | 4 |
| 2018 | DOREMUS: A Graph of Linked Musical WorksabstractThree major French cultural institutions—the French National Library (BnF), Radio France and the Philharmonie de Paris—have come together in order to develop shared methods to describe semantically their catalogs of music works and events. This process comprises the construction of knowledge graphs representing the data contained in these catalogs following a novel agreed upon ontology that extends CIDOC-CRM and FRBRoo, the linking of these graphs and their open publication on the web. A number of specialized tools that allow for the reproduction of this process are developed, as well as web applications for easy access and navigation through the data. The paper presents one of the main outcomes of this project—the DOREMUS knowledge graph, consisting of three linked datasets describing classical music works and their associated events (e.g., performances in concerts). This resource fills an important gap between library content description and music metadata. We present the DOREMUS pipeline for lifting and linking the data, the tools developed for these purposes, as well as a search application allowing to explore the data. Manel Achichi, Pasquale Lisena, Konstantin Todorov, Raphaël Troncy, Jean Delahousse |
ISWC (2) | 3 |
| 2017 | KeyRanker: Automatic RDF Key Ranking for Data LinkingabstractAutomatic approaches to key discovery on RDF datasets generate sets of discriminative properties that can be used to configure data linking systems relying on link specifications. These keys often come in large numbers, generated independently for two datasets to be linked, lacking an assessment of their usefulness for the linking task. We propose a novel generic algorithm for selecting keys, valid in two datasets, and ranking them with respect to their individual likelihood to generate identity links. In addition, we explore the combined use of several complementary keys improving their individual performance. We evaluate our approach on diverse synthetic and real-world benchmark data, showing its robustness with respect to different linking tools and domains. Houssameddine Farah, Danai Symeonidou, Konstantin Todorov |
K-CAP | 3 |
| 2017 | Ontolex JeuxDeMots and Its Alignment to the Linguistic Linked Open Data Cloud
Andon Tchechmedjiev, Théophile Mandon, Mathieu Lafourcade, Anne Laurent, Konstantin Todorov |
ISWC (1) | 5 |
| 2016 | Automatic Key Selection for Data Linking
Manel Achichi, Mohamed Ben Ellefi, Danai Symeonidou, Konstantin Todorov |
EKAW | 4 |
| 2016 | Selecting Optimal Background Knowledge Sources for the Ontology Matching Task
Abdel Nasser Tigrine, Zohra Bellahsene, Konstantin Todorov |
EKAW | 3 |
| 2016 | Dataset Recommendation for Data Linking: An Intensional Approach
Mohamed Ben Ellefi, Zohra Bellahsene, Stefan Dietze, Konstantin Todorov |
ESWC | 4 |
| 2016 | Beyond Established Knowledge Graphs-Recommending Web Datasets for Data Linking
Mohamed Ben Ellefi, Zohra Bellahsene, Stefan Dietze, Konstantin Todorov |
ICWE | 4 |
| 2013 | Opening the Black Box of Ontology Matching
DuyHoa Ngo, Zohra Bellahsene, Konstantin Todorov |
ESWC | 3 |
| 2009 | Detecting Ontology Mappings via Descriptive Statistical MethodsabstractInstance-based ontology mapping comprises a collection of theoretical approaches and applications for identifying the implicit semantic similarities between two ontologies on the basis of the instances that populate their concepts. The current paper situates this general problem in the realm of finding mappings between the nodes of two different Web directories populated with text documents (the Web pages that they intend to organize). We propose a novel approach to detect potential concept mappings based on principle component analysis and discriminant analysis and introduce a resulting concept similarity measure. The procedure can be used as an independent concept mapping technique, or as a support to a concept similarity measure of other nature. Konstantin Todorov |
ICIW | 1 |
| 2008 | Combining Structural and Instance-Based Ontology Similarities for Mapping Web DirectoriesabstractOntologies are knowledge bodies describing the semantics of data and are broadly applied in supporting knowledge exchange between different parties. However, the adoption of the same ontology over a certain domain of interest by different people or organizations is unlikely, mainly due to the decentralized nature of ontology development. The paper presents a procedure for mapping hierarchical ontologies, populated with instances taken from properly classified text documents. It combines a structural and instance-based approaches in order to yield concept-to- concept mapping assertions between two input ontologies. It can be successfully applied to finding correspondences between the elements of two directories. Konstantin Todorov |
ICIW | 1 |