EDBT 2026 Demo / reviewers in the wild / expert
Stefan Dietze
dblp:25/5167
· DBLP profile ↗
41ranked-venue papers in the field
3as first author
16since 2021 · last 2026
0009-0001-4364-9243ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 21Knowledge Engineering, Semantic Web & Information Systems · 16 (3 first)Database Systems & Data Management · 3Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Overview of Touché 2026: Argumentation Systems - Extended Abstract
Johannes Kiesel, Marc Feger, Tim Hagen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Katarina Boland, Wilhelm Pertsch, Julia Romberg, Ines Zelch, Stefan Dietze, Matthias Hagen, Martin Potthast, Benno Stein 0001 |
ECIR (4) | 11 |
| 2026 | The CLEF-2026 CheckThat! Lab: Advancing Multilingual Fact-Checking
Julia Maria Struß, Sebastian Schellhammer, Stefan Dietze, Venktesh V., Vinay Setty, Tanmoy Chakraborty 0002, Preslav Nakov, Avishek Anand, Primakov Chungkham, Salim Hafid, Dhruv Sahnan, Konstantin Todorov |
ECIR (4) | 3 |
| 2025 | The CLEF-2025 CheckThat! Lab: Subjectivity, Fact-Checking, Claim Normalization, and Retrieval
Firoj Alam, Julia Maria Struß, Tanmoy Chakraborty 0002, Stefan Dietze, Salim Hafid, Katerina Korre, Arianna Muti, Preslav Nakov, Federico Ruggeri, Sebastian Schellhammer, Vinay Setty, Megha Sundriyal, Konstantin Todorov, Venktesh V |
ECIR (5) | 4 |
| 2025 | Research Knowledge Graphs: The Shifting Paradigm of Scholarly Information Representation
Matthäus Zloch, Danilo Dessì, Jennifer D'Souza 0001, Leyla Jael Castro, Benjamin Zapilko, Saurav Karmakar, Brigitte Mathiak, Markus Stocker, Wolfgang Otto 0002, Sören Auer, Stefan Dietze |
ESWC (2) | 11 |
| 2025 | TeleScope A Longitudinal Dataset for Investigating Online Discourse and Information Interaction on TelegramabstractTelegram is a globally popular instant messaging platform known for its strong emphasis on security, privacy, and unique social networking features. It has recently emerged as the host for various cross-domain analysis and research works, such as social media influence, propaganda studies, and extremism. This paper introduces TeleScope, an extensive dataset suite that, to our knowledge, is the largest of its kind. It comprises metadata for about 500K Telegram channels and downloaded message metadata for about 71K public channels, accounting for around 120M crawled messages. We also release channel connections and user interaction data built using Telegram’s message-forwarding feature to study multiple use cases, such as information spread and message-forwarding patterns. In addition, we provide data enrichments, such as language detection, active message posting periods for each channel, and Telegram entities extracted from messages, that enable online discourse analysis beyond what is possible with the original data alone. The dataset is designed for diverse applications, independent of specific research objectives, and sufficiently versatile to facilitate the replication of social media studies comparable to those conducted on platforms like X (formerly Twitter). Susmita Gangopadhyay, Danilo Dessì, Dimitar Dimitrov 0002, Stefan Dietze |
ICWSM | 4 |
| 2025 | An In-depth Analysis of the Linguistic Characteristics of Science Claims on the Web and their Impact on Fact-checkingabstractWeb claims, seen as assertions shared on the web and eligible for fact-checking, are at the heart of online discourse. They have been studied extensively on a variety of downstream tasks such as fact-checking, claim retrieval, bias detection, argument mining, or viewpoint discovery. On the other hand, claims originating from scientific publications have also been the subject of several downstream NLP tasks. However, research carried out so far has yet to focus on scientific web claims, which are scientific claims made on the web (e.g., on social media and news articles). The process of detecting and fact-checking a claim from the web can be very different depending on whether the claim is scientific or not, thus making it crucial for the developed datasets, methods, and models to make a distinction between the two. With this work, we aim at understanding what makes this distinction necessary, by understanding the linguistic differences between scientific and non-scientific claims on the web, and the impact those differences have on existing downstream tasks. To do so, we manually annotate 1,524 web claims from established benchmarks for fact-checking-related tasks, and we run statistical tests to analyze and compare the linguistic features of each group. We find that scientific claims on the web use more analytical speech, but also use more sentiment-related speech, more expressions of physical motion, and have distinct parts of speech (PoS) and punctuation styles. We also conduct experiments showing that BERT-based language models perform worse on scientific web claims by up to 17 F1 points for several downstream tasks. To understand why, we develop a novel methodology to map predictive tokens of language models to explainable linguistic features and find that language models fail to detect a specific subset of predictive features of scientific web claims. We conclude by stating that language models aimed at studying scientific web claims ought to be trained on scientific web discourse, as opposed to being trained only on generic web discourse or only on scientific text from scientific publications. Salim Hafid, Sebastian Schellhammer, Yavuz Selim Kartal, Thomas Papastergiou, Stefan Dietze, Sandra Bringay, Konstantin Todorov |
ACM Trans. Web | 5 |
| 2024 | Understanding the Impact of Entity Linking on the Topology of Entity Co-occurrence Networks for Social Media Analysis
James Nevin, Dimitar Dimitrov 0002, Michael Lees, Paul Groth, Stefan Dietze |
EKAW | 6 |
| 2024 | A mutually enhanced multi-scale relation-aware graph convolutional network for argument pair extraction
Xiaofei Zhu, Yidan Liu, Jiafeng Guo, Stefan Dietze |
J. Intell. Inf. Syst. | 6 |
| 2023 | Dynamic global structure enhanced multi-channel graph neural network for session-based recommendation
Xiaofei Zhu, Gu Tang, Pengfei Wang 0009, Chenliang Li 0005, Jiafeng Guo, Stefan Dietze |
Inf. Sci. | 6 |
| 2023 | Exploring rich structure information for aspect-based sentiment classification
Xiaofei Zhu, Jiafeng Guo, Stefan Dietze |
J. Intell. Inf. Syst. | 4 |
| 2022 | SaL-Lightning Dataset: Search and Eye Gaze Behavior, Resource Interactions and Knowledge Gain during Web SearchabstractThe emerging research field Search as Learning (SAL) investigates how the Web facilitates learning through modern information retrieval systems. SAL research requires significant amounts of data that capture both search behavior of users and their acquired knowledge in order to obtain conclusive insights or train supervised machine learning models. However, the creation of such datasets is costly and requires interdisciplinary efforts in order to design studies and capture a wide range of features. In this paper, we address this issue and introduce an extensive dataset based on a user study, in which 114 participants were asked to learn about the formation of lightning and thunder. Participants’ knowledge states were measured before and after Web search through multiple-choice questionnaires and essay-based free recall tasks. To enable future research in SAL-related tasks we recorded a plethora of features and person-related attributes. Besides the screen recordings, visited Web pages, and detailed browsing histories, a large number of behavioral features and resource features were monitored. We underline the usefulness of the dataset by describing three, already published, use cases. Christian Otto, Markus Rokicki, Georg Pardi, Wolfgang Gritz, Daniel Hienert, Ran Yu 0001, Johannes von Hoyer, Anett Hoppe, Stefan Dietze, Peter Holtz, Yvonne Kammerer, Ralph Ewerth |
CHIIR | 9 |
| 2022 | SciTweets - A Dataset and Annotation Framework for Detecting Scientific Online DiscourseabstractScientific topics, claims and resources are increasingly debated as part of online discourse, where prominent examples include discourse related to COVID-19 or climate change. This has led to both significant societal impact and increased interest in scientific online discourse from various disciplines. For instance, communication studies aim at a deeper understanding of biases, quality or spreading patterns of scientific information, whereas computational methods have been proposed to extract, classify or verify scientific claims using NLP and IR techniques. However, research across disciplines currently suffers from both a lack of robust definitions of the various forms of science-relatedness as well as appropriate ground truth data for distinguishing them. In this work, we contribute (a) an annotation framework and corresponding definitions for different forms of scientific relatedness of online discourse in tweets, (b) an expert-annotated dataset of 1261 tweets obtained through our labeling framework reaching an average Fleiss Kappa κ of 0.63, (c) a multi-label classifier trained on our data able to detect science- relatedness with 89% F1 and also able to detect distinct forms of scientific knowledge (claims, references). With this work, we aim to lay the foundation for developing and evaluating robust methods for analysing science as part of large-scale online discourse. Salim Hafid, Sebastian Schellhammer, Sandra Bringay, Konstantin Todorov, Stefan Dietze |
CIKM | 5 |
| 2022 | Lattice-based progressive author disambiguation
Tobias Backes, Stefan Dietze |
Inf. Syst. | 2 |
| 2022 | Exploiting stance hierarchies for cost-sensitive stance detection of Web documents
Arjun Roy 0001, Pavlos Fafalios, Asif Ekbal, Xiaofei Zhu, Stefan Dietze |
J. Intell. Inf. Syst. | 5 |
| 2021 | SoMeSci- A 5 Star Open Data Gold Standard Knowledge Graph of Software Mentions in Scientific ArticlesabstractKnowledge about software used in scientific investigations is important for several reasons, for instance, to enable an understanding of provenance and methods involved in data handling. However, software is usually not formally cited, but rather mentioned informally within the scholarly description of the investigation, raising the need for automatic information extraction and disambiguation. Given the lack of reliable ground truth data, we present SoMeSci-Software Mentions in Science-a gold standard knowledge graph of software mentions in scientific articles. It contains high quality annotations (IRR: K=.82) of 3756 software mentions in 1367 PubMed Central articles. Besides the plain mention of the software, we also provide relation labels for additional information, such as the version, the developer, a URL or citations. Moreover, we distinguish between different types, such as application, plugin or programming environment, as well as different types of mentions, such as usage or creation. To the best of our knowledge, SoMeSci is the most comprehensive corpus about software mentions in scientific articles, providing training samples for Named Entity Recognition, Relation Extraction, Entity Disambiguation, and Entity Linking. Finally, we sketch potential use cases and provide baseline results. David Schindler, Felix Bensmann, Stefan Dietze, Frank Krüger 0001 |
CIKM | 3 |
| 2021 | Topic-independent modeling of user knowledge in informational search sessionsabstractAbstract Web search is among the most frequent online activities. In this context, widespread informational queries entail user intentions to obtain knowledge with respect to a particular topic or domain. To serve learning needs better, recent research in the field of interactive information retrieval has advocated the importance of moving beyond relevance ranking of search results and considering a user’s knowledge state within learning oriented search sessions. Prior work has investigated the use of supervised models to predict a user’s knowledge gain and knowledge state from user interactions during a search session. However, the characteristics of the resources that a user interacts with have neither been sufficiently explored, nor exploited in this task. In this work, we introduce a novel set of resource-centric features and demonstrate their capacity to significantly improve supervised models for the task of predicting knowledge gain and knowledge state of users in Web search sessions. We make important contributions, given that reliable training data for such tasks is sparse and costly to obtain. We introduce various feature selection strategies geared towards selecting a limited subset of effective and generalizable features. Ran Yu 0001, Markus Rokicki, Ujwal Gadiraju, Stefan Dietze |
Inf. Retr. J. | 5 |
| 2020 | The Role of Word-Eye-Fixations for Query Term PredictionabstractThroughout the search process, the user's gaze on inspected SERPs and websites can reveal his or her search interests. Gaze behavior can be captured with eye tracking and described with word-eye-fixations. Word-eye-fixations contain the user's accumulated gaze fixation duration on each individual word of a web page. In this work, we analyze the role of word-eye-fixations for predicting query terms. We investigate the relationship between a range of in-session features, in particular, gaze data, with the query terms and train models for predicting query terms. We use a dataset of 50 search sessions obtained through a lab study in the social sciences domain. Using established machine learning models, we can predict query terms with comparably high accuracy, even with only little training data. Feature analysis shows that the categories Fixation, Query Relevance and Session Topic contain the most effective features for our task. Masoud Davari, Daniel Hienert, Dagmar Kern, Stefan Dietze |
CHIIR | 4 |
| 2020 | TweetsCOV19 - A Knowledge Base of Semantically Annotated Tweets about the COVID-19 PandemicabstractPublicly available social media archives facilitate research in the social sciences and provide corpora for training and testing a wide range of machine learning and natural language processing methods. With respect to the recent outbreak of the Coronavirus disease 2019 (COVID-19), online discourse on Twitter reflects public opinion and perception related to the pandemic itself as well as mitigating measures and their societal impact. Understanding such discourse, its evolution, and interdependencies with real-world events or (mis)information can foster valuable insights. On the other hand, such corpora are crucial facilitators for computational methods addressing tasks such as sentiment analysis, event detection, or entity recognition. However, obtaining, archiving, and semantically annotating large amounts of tweets is costly. In this paper, we describe TweetsCOV19, a publicly available knowledge base of currently more than 8 million tweets, spanning October 2019 - April 2020. Metadata about the tweets as well as extracted entities, hashtags, user mentions, sentiments, and URLs are exposed using established RDF/S vocabularies, providing an unprecedented knowledge base for a range of knowledge discovery tasks. Next to a description of the dataset and its extraction and annotation process, we present an initial analysis and use cases of the corpus. Dimitar Dimitrov 0002, Erdal Baran, Pavlos Fafalios, Ran Yu 0001, Xiaofei Zhu, Matthäus Zloch, Stefan Dietze |
CIKM | 7 |
| 2020 | Crosstown traffic - supervised prediction of impact of planned special events on urban traffic
Nicolas Tempelmeier, Stefan Dietze, Elena Demidova |
GeoInformatica | 2 |
| 2019 | A Software Framework and Datasets for the Analysis of Graph Measures on RDF GraphsabstractAs the availability and the inter-connectivity of RDF datasets grow, so does the necessity to understand the structure of the data. Understanding the topology of RDF graphs can guide and inform the development of, e.g. synthetic dataset generators, sampling methods, index structures, or query optimizers. In this work, we propose two resources: (i) a software framework (Resource URL of the framework: https://doi.org/10.5281/zenodo.2109469 ) able to acquire, prepare, and perform a graph-based analysis on the topology of large RDF graphs, and (ii) results on a graph-based analysis of 280 datasets (Resource URL of the datasets: https://doi.org/10.5281/zenodo.1214433 ) from the LOD Cloud with values for 28 graph measures computed with the framework. We present a preliminary analysis based on the proposed resources and point out implications for synthetic dataset generators. Finally, we identify a set of measures, that can be used to characterize graphs in the Semantic Web. Matthäus Zloch, Maribel Acosta, Daniel Hienert, Stefan Dietze, Stefan Conrad 0001 |
ESWC | 4 |
| 2019 | ClaimsKG: A Knowledge Graph of Fact-Checked ClaimsabstractVarious research areas at the intersection of computer and social sciences require a ground truth of contextualized claims labelled with their truth values in order to facilitate supervision, validation or reproducibility of approaches dealing, for example, with fact-checking or analysis of societal debates. So far, no reasonably large, up-to-date and queryable corpus of structured information about claims and related metadata is publicly available. In an attempt to fill this gap, we introduce ClaimsKG, a knowledge graph of fact-checked claims, which facilitates structured queries about their truth values, authors, dates, journalistic reviews and other kinds of metadata. ClaimsKG is generated through a semi-automated pipeline, which harvests data from popular fact-checking websites on a regular basis, annotates claims with related entities from DBpedia, and lifts the data to RDF using an RDF/S model that makes use of established vocabularies. In order to harmonise data originating from diverse fact-checking sites, we introduce normalised ratings as well as a simple claims coreference resolution strategy. The current knowledge graph, extensible to new information, consists of 28,383 claims published since 1996, amounting to 6,606,032 triples. Andon Tchechmedjiev, Pavlos Fafalios, Katarina Boland, Malo Gasquet, Matthäus Zloch, Benjamin Zapilko, Stefan Dietze, Konstantin Todorov |
ISWC (2) | 7 |
| 2018 | Analyzing Knowledge Gain of Users in Informational Search Sessions on the WebabstractWeb search is frequently used by people to acquire new knowledge and to satisfy learning-related objectives, but little is known about how a user»s knowledge evolves through the course of a search session. We present a study addressing the knowledge gain of users in informational search sessions. Using crowdsourcing, we recruited 500 distinct users and orchestrated real-world search sessions spanning 10 different topics and information needs. By using scientifically formulated knowledge tests we calibrated the knowledge of users before and after their search sessions, quantifying their knowledge gain. We investigated the impact of information needs on the search behavior and knowledge gain of users, revealing a significant effect of information need on user queries and navigational patterns, but no direct effect on the knowledge gain. Users on average exhibited a higher knowledge gain through search sessions pertaining to topics they were less familiar with. Ujwal Gadiraju, Ran Yu 0001, Stefan Dietze, Peter Holtz |
CHIIR | 3 |
| 2018 | TweetsKB: A Public and Large-Scale RDF Corpus of Annotated Tweets
Pavlos Fafalios, Vasileios Iosifidis, Eirini Ntoutsi, Stefan Dietze |
ESWC | 4 |
| 2018 | Predicting User Knowledge Gain in Informational Search SessionsabstractWeb search is frequently used by people to acquire new knowledge and to satisfy learning-related objectives. In this context, informational search missions with an intention to obtain knowledge pertaining to a topic are prominent. The importance of learning as an outcome of web search has been recognized. Yet, there is a lack of understanding of the impact of web search on a user's knowledge state. Predicting the knowledge gain of users can be an important step forward if web search engines that are currently optimized for relevance can be molded to serve learning outcomes. In this paper, we introduce a supervised model to predict a user's knowledge state and knowledge gain from features captured during the search sessions. To measure and predict the knowledge gain of users in informational search sessions, we recruited 468 distinct users using crowdsourcing and orchestrated real-world search sessions spanning 11 different topics and information needs. By using scientifically formulated knowledge tests, we calibrated the knowledge of users before and after their search sessions, quantifying their knowledge gain. Our supervised models utilise and derive a comprehensive set of features from the current state of the art and compare performance of a range of feature sets and feature selection strategies. Through our results, we demonstrate the ability to predict and classify the knowledge state and gain using features obtained during search sessions, exhibiting superior performance to an existing baseline in the knowledge state prediction task. Ran Yu 0001, Ujwal Gadiraju, Peter Holtz, Markus Rokicki, Philipp Kemkes, Stefan Dietze |
SIGIR | 6 |
| 2018 | Inferring Missing Categorical Information in Noisy and Sparse Web MarkupabstractEmbedded markup of Web pages has seen widespread adoption throughout the past years driven by standards such as RDFa and Microdata and initiatives such as schema.org, where recent studies show an adoption by 39% of all Web pages already in 2016. While this constitutes an important information source for tasks such as Web search, Web page classification or knowledge graph augmentation, individual markup nodes are usually sparsely described and often lack essential information. For instance, from 26 million nodes describing events within the Common Crawl in 2016, 59% of nodes provide less than six statements and only 257,000 nodes (0.96%) are typed with more specific event subtypes. Nevertheless, given the scale and diversity of Web markup data, nodes that provide missing information can be obtained from the Web in large quantities, in particular for categorical properties. Such data constitutes potential training data for inferring missing information to significantly augment sparsely described nodes. In this work, we introduce a supervised approach for inferring missing categorical properties in Web markup. Our experiments, conducted on properties of events and movies, show a performance of 79% and 83% F1 score correspondingly, significantly outperforming existing baselines. Nicolas Tempelmeier, Elena Demidova, Stefan Dietze |
WWW | 3 |
| 2017 | FuseM: Query-Centric Data Fusion on Structured Web MarkupabstractEmbedded markup based on Microdata, RDFa, and Microformats have become prevalent on the Web and constitute an unprecedented source of data. However, RDF statements extracted from markup are fundamentally different from traditional RDF graphs: entity descriptions are flat, facts are highly redundant, and despite very frequent co-references explicit links are missing. Therefore, carrying out typical entity-centric tasks such as retrieval and summarisation cannot be tackled sufficiently with state-of-the-art methods and require preliminary data fusion. Given the scale and dynamics of Web markup, the applicability of general data fusion approaches is limited. We present a novel query-centric data fusion approach which overcomes such issues through a combination of entity retrieval and fusion techniques geared towards the specific challenges associated with embedded markup. To ensure precise and diverse entity descriptions, we follow a supervised learning approach and train a classifier for data fusion of a pool of candidate facts relevant to a given query and obtained through a preliminary entity retrieval step. We perform a thorough evaluation on a subset of the Web Data Commons dataset and show significant improvement over existing baselines. In addition, an investigation into the coverage and complementarity of facts from the constructed entity descriptions compared to DBpedia, shows potential for aiding tasks such as knowledge base population. Ran Yu 0001, Ujwal Gadiraju, Besnik Fetahu, Stefan Dietze |
ICDE | 4 |
| 2016 | Dataset Recommendation for Data Linking: An Intensional Approach
Mohamed Ben Ellefi, Zohra Bellahsene, Stefan Dietze, Konstantin Todorov |
ESWC | 3 |
| 2016 | Beyond Established Knowledge Graphs-Recommending Web Datasets for Data Linking
Mohamed Ben Ellefi, Zohra Bellahsene, Stefan Dietze, Konstantin Todorov |
ICWE | 3 |
| 2015 | Improving Entity Retrieval on Structured Data
Besnik Fetahu, Ujwal Gadiraju, Stefan Dietze |
ISWC (1) | 3 |
| 2015 | Adaptive Focused Crawling of Linked Data
Ran Yu 0001, Ujwal Gadiraju, Besnik Fetahu, Stefan Dietze |
WISE (1) | 4 |
| 2014 | A Scalable Approach for Efficiently Generating Structured Dataset Topic Profiles
Besnik Fetahu, Stefan Dietze, Bernardo Pereira Nunes, Marco A. Casanova, Davide Taibi 0002, Wolfgang Nejdl |
ESWC | 2 |
| 2014 | Two Approaches to the Dataset Interlinking Recommendation Problem
Giseli Rabello Lopes, Luiz André P. Paes Leme, Bernardo Pereira Nunes, Marco A. Casanova, Stefan Dietze |
WISE (1) | 5 |
| 2013 | Complex Matching of RDF Datatype Properties
Bernardo Pereira Nunes, Alexander Arturo Mera Caraballo, Marco A. Casanova, Besnik Fetahu, Luiz André P. Paes Leme, Stefan Dietze |
DEXA (1) | 6 |
| 2013 | Combining a Co-occurrence-Based and a Semantic Measure for Entity Linking
Bernardo Pereira Nunes, Stefan Dietze, Marco A. Casanova, Ricardo Kawase, Besnik Fetahu, Wolfgang Nejdl |
ESWC | 2 |
| 2013 | Summaries on the Fly: Query-Based Extraction of Structured Knowledge from Web Documents
Besnik Fetahu, Bernardo Pereira Nunes, Stefan Dietze |
ICWE | 3 |
| 2013 | Identifying Candidate Datasets for Data Interlinking
Luiz André P. Paes Leme, Giseli Rabello Lopes, Bernardo Pereira Nunes, Marco A. Casanova, Stefan Dietze |
ICWE | 5 |
| 2013 | Recommending Tripleset Interlinking through a Social Network Approach
Giseli Rabello Lopes, Luiz André P. Paes Leme, Bernardo Pereira Nunes, Marco A. Casanova, Stefan Dietze |
WISE (1) | 5 |
| 2012 | Exploiting the Social and Semantic Web for Guided Web Archiving
Thomas Risse 0001, Stefan Dietze, Wim Peters, Katerina Doka, Yannis Stavrakas, Pierre Senellart |
TPDL | 2 |
| 2011 | SmartLink: A Web-Based Editor and Search Environment for Linked Services
Stefan Dietze, HongQing Yu, Carlos Pedrinaci, Dong Liu 0015, John Domingue |
ESWC (2) | 1 |
| 2008 | Conceptual Situation Spaces for Semantic Situation-Driven Processes
Stefan Dietze, Alessio Gugliotta, John Domingue |
ESWC | 1 |
| 2007 | A Semantic Web Service Oriented Framework for Adaptive Learning Environments
Stefan Dietze, Alessio Gugliotta, John Domingue |
ESWC | 1 |