VLDB 2026 Research / reviewers in the wild / expert
Jörg Schlötterer
dblp:160/1725 · also Joerg Schloetterer
· DBLP profile ↗
22ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-3678-0390ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weakly Supervised Shortcut Learning Mitigation Using Sparse AutoencodersabstractReliance on spurious features that coincidentally correlate with task labels (i.e., shortcut learning) remains a major barrier to the reliable deployment of machine learning models, particularly in high-stakes domains like medical diagnostics.Moreover, in such settings retraining models or collecting and labeling additional data is often impractical, limiting the applicability of many existing shortcut mitigation methods.In this paper we propose a lightweight framework that leverages sparse autoencoders to disentangle spurious from core features to mitigate shortcut learning.Our approach requires no model retraining and works even when group annotations are scarce or unavailable for certain classes.Results on standard benchmarks demonstrate that, even with as few as 50 labeled examples, reliance on spurious features can be significantly reduced. Sari Sadiya, Despina Tawadros, Phuong Quynh Le, Jörg Schlötterer, Christin Seifert, Gemma Roig |
ESANN | 5 |
| 2025 | Efficient Unsupervised Shortcut Learning Detection and Mitigation in TransformersabstractShortcut learning, i.e., a model's reliance on undesired features not directly relevant to the task, is a major challenge that severely limits the applications of machine learning algorithms, particularly when deploying them to assist in making sensitive decisions, such as in medical diagnostics. In this work, we leverage recent advancements in machine learning to create an unsupervised framework that is capable of both detecting and mitigating shortcut learning in transformers. We validate our method on multiple datasets. Results demonstrate that our framework significantly improves both worst-group accuracy (samples misclassified due to shortcuts) and average accuracy, while minimizing human annotation effort. Moreover, we demonstrate that the detected shortcuts are meaningful and informative to human experts, and that our framework is computationally efficient, allowing it to be run on consumer hardware. Lukas Kuhn, Sari Sadiya, Jörg Schlötterer, Florian Buettner 0001, Christin Seifert, Gemma Roig |
ICCV | 3 |
| 2025 | Has this Fact been Edited? Detecting Knowledge Edits in Language ModelsabstractPaul Youssef, Zhixue Zhao, Christin Seifert, Jörg Schlötterer. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Paul Youssef, Zhixue Zhao, Christin Seifert, Jörg Schlötterer |
NAACL (Long Papers) | 4 |
| 2025 | How to Make LLMs Forget: On Reversing In-Context Knowledge EditsabstractPaul Youssef, Zhixue Zhao, Jörg Schlötterer, Christin Seifert. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Paul Youssef, Zhixue Zhao, Jörg Schlötterer, Christin Seifert |
NAACL (Long Papers) | 3 |
| 2024 | InfoLossQA: Characterizing and Recovering Information Loss in Text SimplificationabstractJan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu, Byron Wallace, Junyi Jessy Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu 0004, Byron C. Wallace, Junyi Jessy Li |
ACL (1) | 3 |
| 2024 | Comprehensive Study on German Language Models for Clinical and Biomedical Text UnderstandingabstractRecent advances in natural language processing (NLP) can be largely attributed to the advent of pre-trained language models such as BERT and RoBERTa. While these models demonstrate remarkable performance on general datasets, they can struggle in specialized domains such as medicine, where unique domain-specific terminologies, domain-specific abbreviations, and varying document structures are common. This paper explores strategies for adapting these models to domain-specific requirements, primarily through continuous pre-training on domain-specific data. We pre-trained several German medical language models on 2.4B tokens derived from translated public English medical data and 3B tokens of German clinical data. The resulting models were evaluated on various German downstream tasks, including named entity recognition (NER), multi-label classification, and extractive question answering. Our results suggest that models augmented by clinical and translation-based pre-training typically outperform general domain models in medical contexts. We conclude that continuous pre-training has demonstrated the ability to match or even exceed the performance of clinical models trained from scratch. Furthermore, pre-training on clinical data or leveraging translated texts have proven to be reliable methods for domain adaptation in medical NLP tasks. Ahmad Idrissi-Yaghir, Amin Dada, Henning Schäfer, Kamyar Arzideh, Giulia Baldini 0001, Jan Trienes, Max Hasin, Jeanette Bewersdorff, Cynthia Sabrina Schmidt, Marie Bauer, Kaleb E. Smith, Jiang Bian 0001, Yonghui Wu 0001, Jörg Schlötterer, Torsten Zesch, Peter A. Horn, Christin Seifert, Felix Nensa, Jens Kleesiek, Christoph M. Friedrich |
LREC/COLING | 14 |
| 2024 | A Second Look on BASS - Boosting Abstractive Summarization with Unified Semantic Graphs - A Replication Study
Osman Alperen Koras, Jörg Schlötterer, Christin Seifert |
ECIR (4) | 2 |
| 2024 | CEval: A Benchmark for Evaluating Counterfactual Text GenerationabstractCounterfactual text generation aims to minimally change a text, such that it is classified differently.Assessing progress in method development for counterfactual text generation is hindered by a non-uniform usage of data sets and metrics in related work.We propose CEval, a benchmark for comparing counterfactual text generation methods.CEval unifies counterfactual and text quality metrics, includes common counterfactual datasets with human annotations, standard baselines (MICE, GDBA, CREST) and the open-source language model LLAMA-2.Our experiments found no perfect method for generating counterfactual text.Methods that excel at counterfactual metrics often produce lower-quality text while LLMs with simple prompts generate high-quality text but struggle with counterfactual criteria.By making CEval available as an open-source Python library, we encourage the community to contribute additional methods and maintain consistent evaluation in future work. 1 Van Bach Nguyen, Christin Seifert, Jörg Schlötterer |
INLG | 3 |
| 2024 | Corpus Considerations for Annotator Modeling and ScalingabstractOlufunke O. Sarumi, Béla Neuendorf, Joan Plepi, Lucie Flek, Jörg Schlötterer, Charles Welch. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Olufunke Oluyemi Sarumi, Béla Neuendorf, Joan Plepi, Lucie Flek, Jörg Schlötterer, Charles Welch |
NAACL-HLT | 5 |
| 2023 | PIP-Net: Patch-Based Intuitive Prototypes for Interpretable Image ClassificationabstractInterpretable methods based on prototypical patches recognize various components in an image in order to explain their reasoning to humans. However, existing prototype-based methods can learn prototypes that are not in line with human visual perception, i.e., the same prototype can refer to different concepts in the real world, making interpretation not intuitive. Driven by the principle of explainability-by-design, we introduce PIP-Net (Patch-based Intuitive Prototypes Network): an interpretable image classification model that learns prototypical parts in a self-supervised fashion which correlate better with human vision. PIP-Net can be interpreted as a sparse scoring sheet where the presence of a prototypical part in an image adds evidence for a class. The model can also abstain from a decision for out-of-distribution data by saying “I haven't seen this before”. We only use image-level labels and do not rely on any part annotations. PIP-Net is globally interpretable since the set of learned prototypes shows the entire reasoning of the model. A smaller local explanation locates the relevant prototypes in one image. We show that our prototypes correlate with ground-truth object parts, indicating that PIP-Net closes the “semantic gap” between latent space and pixel space. Hence, our PIP-Net with interpretable prototypes enables users to interpret the decision making process in an intuitive, faithful and semantically meaningful way. Code is available at https://github.com/M-Nauta/PIPNet. Meike Nauta, Jörg Schlötterer, Maurice van Keulen, Christin Seifert |
CVPR | 2 |
| 2023 | Benchmarking eXplainable AI - A Survey on Available Toolkits and Open ChallengesabstractThe goal of Explainable AI (XAI) is to make the reasoning of a machine learning model accessible to humans, such that users of an AI system can evaluate and judge the underlying model. Due to the blackbox nature of XAI methods it is, however, hard to disentangle the contribution of a model and the explanation method to the final output. It might be unclear on whether an unexpected output is caused by the model or the explanation method. Explanation models, therefore, need to be evaluated in technical (e.g. fidelity to the model) and user-facing (correspondence to domain knowledge) terms. A recent survey has identified 29 different automated approaches to quantitatively evaluate explanations. In this work, we take an additional perspective and analyse which toolkits and data sets are available. We investigate which evaluation metrics are implemented in the toolkits and whether they produce the same results. We find that only a few aspects of explanation quality are currently covered, data sets are rare and evaluation results are not comparable across different toolkits. Our survey can serve as a guide for the XAI community for identifying future directions of research, and most notably, standardisation of evaluation. Phuong Quynh Le, Meike Nauta, Van Bach Nguyen, Shreyasi Pathak, Jörg Schlötterer, Christin Seifert |
IJCAI | 5 |
| 2023 | Guidance in Radiology Report Summarization: An Empirical Evaluation and Error AnalysisabstractAutomatically summarizing radiology reports into a concise impression can reduce the manual burden of clinicians and improve the consistency of reporting.Previous work aimed to enhance content selection and factuality through guided abstractive summarization.However, two key issues persist.First, current methods heavily rely on domain-specific resources to extract the guidance signal, limiting their transferability to domains and languages where those resources are unavailable.Second, while automatic metrics like ROUGE show progress, we lack a good understanding of the errors and failure modes in this task.To bridge these gaps, we first propose a domain-agnostic guidance signal in form of variable-length extractive summaries.Our empirical results on two English benchmarks demonstrate that this guidance signal improves upon unguided summarization while being competitive with domain-specific methods.Additionally, we run an expert evaluation of four systems according to a taxonomy of 11 fine-grained errors.We find that the most pressing differences between automatic summaries and those of radiologists relate to content selection including omissions (up to 52%) and additions (up to 57%).We hypothesize that latent reporting factors and corpus-level inconsistencies may limit models to reliably learn content selection from the available data, presenting promising directions for future work.Incorrect location of a finding?Incorrect severity of a finding?Any other error?Please describe... Jan Trienes, Paul Youssef, Jörg Schlötterer, Christin Seifert |
INLG | 3 |
| 2022 | Evaluating Simulated User Interaction and Search Behaviour
Saber Zerhoudi, Michael Granitzer, Christin Seifert, Jörg Schlötterer |
ECIR (2) | 4 |
| 2020 | Influence of Random Walk Parametrization on Graph Embeddings
Fabian Schliski, Jörg Schlötterer, Michael Granitzer |
ECIR (2) | 2 |
| 2018 | Shortest Path Distance Approximation Using Deep Learning TechniquesabstractComputing shortest path distances between nodes lies at the heart of many graph algorithms and applications. Traditional exact methods such as breadth-first-search (BFS) do not scale up to contemporary, rapidly evolving today's massive networks. Therefore, it is required to find approximation methods to enable scalable graph processing with a significant speedup. In this paper, we utilize vector embeddings learnt by deep learning techniques to approximate the shortest paths distances in large graphs. We show that a feedforward neural network fed with embeddings can approximate distances with relatively low distortion error. The suggested method is evaluated on the Facebook, BlogCatalog, Youtube and Flickr social networks. Fatemeh Salehi Rizi, Jörg Schlötterer, Michael Granitzer |
ASONAM | 2 |
| 2018 | QueryCrumbs for Experts: A Compact Visual Query Support System to Facilitate Insights into Search Engine InternalsabstractSearch experts use advanced query language and search tactics to formulate their queries. However, the effectiveness of those advanced techniques depends on the search engine internals. We propose QueryCrumbs for Experts, a compact visualization, which facilitates insights to the search engine internals and therefore allows the searcher to determine effective search strategies. Treating the search engine as a black box, QueryCrumbs can be seamlessly integrated into existing search interfaces, guiding the user's exploration and assessment of results. QueryCrumbs for Experts visualize the recent search history alongside with a simple and also a qualitative comparison of the result sets, from which conclusions about the search engine internals can be drawn. The evaluation shows that, by identifying specific patterns in the visualization, expert users can gain valuable insights into search engine internals, empowering them to adapt their search accordingly. Jörg Schlötterer, Christin Seifert, Michael Granitzer |
IV | 1 |
| 2017 | On Joint Representation Learning of Network Structure and Document Content
Jörg Schlötterer, Christin Seifert, Michael Granitzer |
CD-MAKE | 1 |
| 2017 | Focus Paragraph Detection for Online Zero-Effort Queries: Lessons learned from Eye-Tracking DataabstractIn order to realize zero-effort retrieval in a web-context, it is crucial to identify the part of the web page the user is focusing on. In this paper, we investigate the identification of focus paragraphs in web pages. Starting from a naive baseline for paragraph and focus paragraph detection, we conducted an eye-tracking study to evaluate the most promising features. We found that single features (mouse position, paragraph position, mouse activity) are less predictive for gaze which confirms findings from other studies. The results indicate that an algorithm for focus paragraph detection needs to incorporate a weighted combination of those features as well as additional features, e.g. semantic context derived from the user's web history. Christin Seifert, Annett Mitschick, Jörg Schlötterer, Raimund Dachselt |
CHIIR | 3 |
| 2017 | QueryCrumbs: A Compact Visualization for Navigating the Search Query HistoryabstractModels of human information seeking reveal that search, in particular ad-hoc retrieval, is non-linear and iterative. Despite these findings, todays search user interfaces do not support non-linear navigation, like for example backtracking in time. In this work, we propose QueryCrumbs, a compact and easy-to-understand visualization for navigating the search query history supporting iterative query refinement. We apply a multi-layered interface design to support novices and firsttime users as well as intermediate users. The formative evaluation with first-time and intermediate users showed that the interactions can be easily performed, and the visual encodings were well understood without instructions. Results indicate that QueryCrumbs can support users when searching for information in an iterative manner. Christin Seifert, Jörg Schlötterer, Michael Granitzer |
IV | 2 |
| 2016 | Supporting Web Surfers in Finding Related Material in Digital Library Repositories
Jörg Schlötterer, Christin Seifert, Michael Granitzer |
TPDL | 1 |
| 2015 | From Context-Aware to Context-Based: Mobile Just-In-Time Retrieval of Cultural Heritage Objects
Jörg Schlötterer, Christin Seifert, Michael Granitzer |
ECIR | 1 |
| 2015 | User Interface Considerations for Browser-Based Just-in-Time-RetrievalabstractWith the availability of free online enrichment services injection of additional, external resources in existing Web content becomes more and more widespread. For the specific area of just-in-time retrieval of digital resources based on web page content, there are no specific guidelines of how to design and integrate the additional user interface components. In this paper, we conceptualise related user interface issues, investigating the central questions: (i) how can a user be visually notified that additional results are available, and (ii) with which user interface elements should the results be presented. Concretely, we identified four different notification styles and six different result presentation styles. In a survey-based study with 75 participants we elicited the users' preferences, revealing a clear preference for the representation style (split pane) and a strong preference for three notification styles (notification bubble, icon appearance and change of icon's appearance). The latter preferences are related to the preferred browser. The results can serve as guideline for designing web-based user interfaces for just-in-time retrieval. Christin Seifert, Jörg Schlötterer, Michael Granitzer |
IV | 2 |