VLDB 2026 Research / reviewers in the wild / expert
Michael Vössing
dblp:198/5318
· DBLP profile ↗
15ranked-venue papers
0as first author
15since 2021 · last 2025
0000-0002-7722-6142ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 9 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Complementarity in human-AI collaboration: concept, sources, and evidenceabstractArtificial intelligence (AI) has the potential to significantly enhance human performance across various domains. Ideally, collaboration between humans and AI should result in complementary team performance (CTP)—a level of performance that neither of them can attain individually. So far, however, CTP has rarely been observed, suggesting an insufficient understanding of the principle and the application of complementarity. Therefore, we develop a general concept of complementarity and formalize its theoretical potential as well as the actual realized effect in decision-making situations. Moreover, we identify information and capability asymmetry as the two key sources of complementarity. Finally, we illustrate the impact of each source on complementarity potential and effect in two empirical studies. Our work provides researchers with a comprehensive theoretical foundation of human-AI complementarity in decision-making and demonstrates that leveraging these sources constitutes a viable pathway towards designing effective human-AI collaboration, i.e., the realization of CTP. Patrick Hemmer, Max Schemmer, Niklas Kühl 0001, Michael Vössing, Gerhard Satzger |
Eur. J. Inf. Syst. | 4 |
| 2025 | Improving Label Error Detection and Elimination with Uncertainty QuantificationabstractIdentifying and handling label errors can significantly enhance the accuracy of supervised machine learning models. Recent approaches for identifying label errors demonstrate that a low self-confidence of models with respect to a certain label represents a good indicator of an erroneous label. However, latest work has built on softmax probabilities to measure selfconfidence. In this paper, we argue that—as softmax probabilities do not reflect a model’s predictive uncertainty accurately— label error detection requires more sophisticated measures of model uncertainty. Therefore, we develop a range of novel, model-agnostic algorithms for Uncertainty Quantification-Based Label Error Detection (UQ-LED), which combine the techniques of confident learning (CL), Monte Carlo Dropout (MCD), model uncertainty measures (e.g., entropy), and ensemble learning to enhance label error detection. We comprehensively evaluate our algorithms on four image classification benchmark datasets in two stages. In the first stage, we demonstrate that our UQ-LED algorithms outperform state-of-the-art confident learning in identifying label errors. In the second stage, we show that removing all identified errors from the training data based on our approach results in higher accuracies than training on all available labeled data. Importantly, besides our contributions to the detection of label errors, we particularly propose a novel approach to generate realistic, class-dependent label errors synthetically. Overall, our study demonstrates that selectively cleaning datasets with UQ-LED algorithms leads to more accurate classifications than using larger, noisier datasets. Johannes Jakubik, Michael Vössing, Manil Maskey, Christopher Wölfle, Gerhard Satzger |
J. Artif. Intell. Res. | 2 |
| 2025 | AI Reliance and Decision Quality: Fundamentals, Interdependence, and the Effects of InterventionsabstractIn AI-assisted decision-making, a central promise of having a human-in-the-loop is that they should be able to complement the AI system by overriding its wrong recommendations. In practice, however, we often see that humans cannot assess the correctness of AI recommendations and, as a result, adhere to wrong or override correct advice. Different ways of relying on AI recommendations have immediate, yet distinct, implications for decision quality. Unfortunately, reliance and decision quality are often inappropriately conflated in the current literature on AI-assisted decision-making. In this work, we disentangle and formalize the relationship between reliance and decision quality, and we characterize the conditions under which human-AI complementarity is achievable. To illustrate how reliance and decision quality relate to one another, we propose a visual framework and demonstrate its usefulness for interpreting empirical findings, including the effects of interventions like explanations. Overall, our research highlights the importance of distinguishing between reliance behavior and decision quality in AI-assisted decision-making. Jakob Schöffer, Johannes Jakubik, Michael Vössing, Niklas Kühl 0001, Gerhard Satzger |
J. Artif. Intell. Res. | 3 |
| 2025 | Human Delegation Behavior in Human-AI Collaboration: The Effect of Contextual InformationabstractThe integration of artificial intelligence (AI) into human decision-making processes at the workplace presents both opportunities and challenges. One promising approach to leverage existing complementary capabilities is allowing humans to delegate individual instances of decision tasks to AI. However, enabling humans to delegate instances effectively requires them to assess several factors. One key factor is the analysis of both their own capabilities and those of the AI in the context of the given task. In this work, we conduct a behavioral study to explore the effects of providing contextual information to support this delegation decision. Specifically, we investigate how contextual information about the AI and the task domain influence humans' delegation decisions to an AI and their impact on the human-AI team performance. Our findings reveal that access to contextual information significantly improves human-AI team performance in delegation settings. Finally, we show that the delegation behavior changes with the different types of contextual information. Overall, this research advances the understanding of computer-supported, collaborative work and provides actionable insights for designing more effective collaborative systems. Philipp Spitzer, Joshua Holstein, Patrick Hemmer, Michael Vössing, Niklas Kühl 0001, Dominik Martin, Gerhard Satzger |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2025 | Towards Understanding AI Delegation: The Role of Self-Efficacy and Visual Processing AbilityabstractRecent work has proposed AI models that can learn to decide whether to make a prediction for a task instance or to delegate it to a human by considering both parties’ capabilities. In simulations with synthetically generated or context-independent human predictions, delegation can help improve the performance of human-AI teams—compared to humans or the AI model completing the task alone. However, so far, it remains unclear how humans perform and how they perceive the task when individual instances of a task are delegated to them by an AI model. In an experimental study with 196 participants, we show that task performance and task satisfaction improve for the instances delegated by the AI model, regardless of whether humans are aware of the delegation. Additionally, we identify humans’ increased levels of self-efficacy as the underlying mechanism for these improvements in performance and satisfaction, and one dimension of cognitive ability as a moderator to this effect. In particular, AI delegation can buffer potential negative effects on task performance and task satisfaction for humans with low visual processing ability. Our findings provide initial evidence that allowing AI models to take over more management responsibilities can be an effective form of human-AI collaboration in workplaces. Monika Westphal, Patrick Hemmer, Michael Vössing, Max Schemmer, Sebastian Vetter, Gerhard Satzger |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2025 | Balancing the Unknown: Exploring Human Reliance on AI Advice under Aleatoric and Epistemic UncertaintyabstractArtificial intelligence (AI) systems increasingly support decision-making across a broad range of domains. The complexity of real-world tasks, however, introduces uncertainty into the prediction capabilities of these systems. This uncertainty can manifest as aleatoric uncertainty arising from inherent variability in outcomes or epistemic uncertainty stemming from limitations in the AI system’s knowledge. While prior research has investigated uncertainty as a monolithic concept, the distinct effects of communicating aleatoric or epistemic uncertainty on humans and their reliance behavior remain unexplored. In this work, we present two behavioral experiments that systematically examine how participants rely on AI advice when faced with different types of uncertainty. While the first experiment manipulates the source of uncertainty, specifying it as either aleatoric or epistemic, the second decomposes uncertainty into its individual components, presenting aleatoric and epistemic uncertainty simultaneously. This work contributes to a deeper understanding of the multifaceted impact of different uncertainty types on human–AI interaction. Joshua Holstein, Lars Boecking, Philipp Spitzer, Niklas Kühl 0001, Michael Vössing, Gerhard Satzger |
ACM Trans. Comput. Hum. Interact. | 5 |
| 2024 | Redefining the Laparoscopic Spatial Sense: AI-Based Intra- and Postoperative Measurement from StereoimagesabstractA significant challenge in image-guided surgery is the accurate measurement task of relevant structures such as vessel segments, resection margins, or bowel lengths. While this task is an essential component of many surgeries, it involves substantial human effort and is prone to inaccuracies. In this paper, we develop a novel human-AI-based method for laparoscopic measurements utilizing stereo vision that has been guided by practicing surgeons. Based on a holistic qualitative requirements analysis, this work proposes a comprehensive measurement method, which comprises state-of-the-art machine learning architectures, such as RAFT-Stereo and YOLOv8. The developed method is assessed in various realistic experimental evaluation environments. Our results outline the potential of our method achieving high accuracies in distance measurements with errors below 1 mm. Furthermore, on-surface measurements demonstrate robustness when applied in challenging environments with textureless regions. Overall, by addressing the inherent challenges of image-guided surgery, we lay the foundation for a more robust and accurate solution for intra- and postoperative measurements, enabling more precise, safe, and efficient surgical procedures. Leopold Müller, Patrick Hemmer, Moritz Queisner, Igor M. Sauer, Simeon Allmendinger, Johannes Jakubik, Michael Vössing, Niklas Kühl 0001 |
AAAI | 7 |
| 2023 | Learning to Defer with Limited Expert PredictionsabstractRecent research suggests that combining AI models with a human expert can exceed the performance of either alone. The combination of their capabilities is often realized by learning to defer algorithms that enable the AI to learn to decide whether to make a prediction for a particular instance or defer it to the human expert. However, to accurately learn which instances should be deferred to the human expert, a large number of expert predictions that accurately reflect the expert's capabilities are required—in addition to the ground truth labels needed to train the AI. This requirement shared by many learning to defer algorithms hinders their adoption in scenarios where the responsible expert regularly changes or where acquiring a sufficient number of expert predictions is costly. In this paper, we propose a three-step approach to reduce the number of expert predictions required to train learning to defer algorithms. It encompasses (1) the training of an embedding model with ground truth labels to generate feature representations that serve as a basis for (2) the training of an expertise predictor model to approximate the expert's capabilities. (3) The expertise predictor generates artificial expert predictions for instances not yet labeled by the expert, which are required by the learning to defer algorithms. We evaluate our approach on two public datasets. One with "synthetically" generated human experts and another from the medical domain containing real-world radiologists' predictions. Our experiments show that the approach allows the training of various learning to defer algorithms with a minimal number of human expert predictions. Furthermore, we demonstrate that even a small number of expert predictions per class is sufficient for these algorithms to exceed the performance the AI and the human expert can achieve individually. Patrick Hemmer, Lukas Thede, Michael Vössing, Johannes Jakubik, Niklas Kühl 0001 |
AAAI | 3 |
| 2023 | Online Emotions during the Storming of the U.S. Capitol: Evidence from the Social Media Network ParlerabstractThe storming of the U.S. Capitol on January 6, 2021 has led to the killing of 5 people and is widely regarded as an attack on democracy. The storming was largely coordinated through social media networks such as Twitter and "Parler". Yet little is known regarding how users interacted on Parler during the storming of the Capitol. In this work, we examine the emotion dynamics on Parler during the storming with regard to heterogeneity across time and users. For this, we segment the user base into different groups (e.g., Trump supporters and QAnon supporters). We use affective computing to infer the emotions in content, thereby allowing us to provide a comprehensive assessment of online emotions. Our evaluation is based on a large-scale dataset from Parler, comprising of 717,300 posts from 144,003 users. We find that the user base responded to the storming of the Capitol with an overall negative sentiment. Akin to this, Trump supporters also expressed a negative sentiment and high levels of unbelief. In contrast to that, QAnon supporters did not express a more negative sentiment during the storming. We further provide a cross-platform analysis and compare the emotion dynamics on Parler and Twitter. Our findings point at a comparatively less negative response to the incidents on Parler compared to Twitter accompanied by higher levels of disapproval and outrage. Our contribution to research is three-fold: (1) We identify online emotions that were characteristic of the storming; (2) we assess emotion dynamics across different user groups on Parler; (3) we compare the emotion dynamics on Parler and Twitter. Thereby, our work offers important implications for actively managing online emotions to prevent similar incidents in the future. Johannes Jakubik, Michael Vössing, Nicolas Pröllochs, Dominik Bär, Stefan Feuerriegel |
ICWSM | 2 |
| 2023 | Toward Foundation Models for Earth Monitoring: Generalizable Deep Learning Models for Natural Hazard SegmentationabstractClimate change results in an increased probability of extreme weather events that put societies and businesses at risk on a global scale. Therefore, near real-time mapping of natural hazards is an emerging priority for the support of natural disaster relief, risk management, and informed governmental policy decisions. Current remote sensing based approaches to near real-time natural hazard mapping increasingly leverage advantage of deep learning (DL). Nevertheless, DL-based approaches are mainly designed for one specific task in a single geographic region based on specific frequency bands of satellite data. For that reason, DL models used to map specific natural hazards struggle with their generalization to other types of natural hazards in unseen regions. In this work, we propose a methodology to significantly improve the generalizability of DL natural hazards mappers based on pre-training on a suitable pre-task. Without access to any data from the target domain, we demonstrate that this methodology improved generalizability across four U-Net architectures for the segmentation of unseen natural hazards, such as flood events, landslides, and massive glacier collapses. Importantly, our method is strongly invariant to geographic differences and the type of input frequency bands of satellite data. That is confirmed by obtaining a balanced accuracy of up to 0.74 in comparison with performance of reference baselines. By leveraging characteristics of unlabeled images from the target domain that are publicly available, our approach is able to further improve the generalization behavior of DL models without fine-tuning. That is reflected in performance metrics. Thereby, our approach is one of first attempts to support the development of foundation models for earth monitoring with the objective of directly segmenting unseen natural hazards across novel geographic regions from different sources of satellite imagery. Johannes Jakubik, Michal Muszynski, Michael Vössing, Niklas Kühl 0001, Thomas Brunschwiler |
IGARSS | 3 |
| 2023 | Human-AI Collaboration: The Effect of AI Delegation on Human Task Performance and Task SatisfactionabstractRecent work has proposed artificial intelligence (AI) models that can learn to decide whether to make a prediction for an instance of a task or to delegate it to a human by considering both parties’ capabilities. In simulations with synthetically generated or context-independent human predictions, delegation can help improve the performance of human-AI teams—compared to humans or the AI model completing the task alone. However, so far, it remains unclear how humans perform and how they perceive the task when they are aware that an AI model delegated task instances to them. In an experimental study with 196 participants, we show that task performance and task satisfaction improve through AI delegation, regardless of whether humans are aware of the delegation. Additionally, we identify humans’ increased levels of self-efficacy as the underlying mechanism for these improvements in performance and satisfaction. Our findings provide initial evidence that allowing AI models to take over more management responsibilities can be an effective form of human-AI collaboration in workplaces. Patrick Hemmer, Monika Westphal, Max Schemmer, Sebastian Vetter, Michael Vössing, Gerhard Satzger |
IUI | 5 |
| 2023 | What a MESS: Multi-Domain Evaluation of Zero-Shot Semantic SegmentationabstractWhile semantic segmentation has seen tremendous improvements in the past, there are still significant labeling efforts necessary and the problem of limited generalization to classes that have not been present during training. To address this problem, zero-shot semantic segmentation makes use of large self-supervised vision-language models, allowing zero-shot transfer to unseen classes. In this work, we build a benchmark for Multi-domain Evaluation of Zero-Shot Semantic Segmentation (MESS), which allows a holistic analysis of performance across a wide range of domain-specific datasets such as medicine, engineering, earth monitoring, biology, and agriculture. To do this, we reviewed 120 datasets, developed a taxonomy, and classified the datasets according to the developed taxonomy. We select a representative subset consisting of 22 datasets and propose it as the MESS benchmark. We evaluate eight recently published models on the proposed MESS benchmark and analyze characteristics for the performance of zero-shot transfer models. The toolkit is available at https://github.com/blumenstiel/MESS. Benedikt Blumenstiel, Johannes Jakubik, Hilde Kühne, Michael Vössing |
NeurIPS | 4 |
| 2022 | Designing a Human-in-the-Loop System for Object Detection in Floor PlansabstractIn recent years, companies in the Architecture, Engineering, and Construction (AEC) industry have started exploring how artificial intelligence (AI) can reduce time-consuming and repetitive tasks. One use case that can benefit from the adoption of AI is the determination of quantities in floor plans. This information is required for several planning and construction steps. Currently, the task requires companies to invest a significant amount of manual effort. Either digital floor plans are not available for existing buildings, or the formats cannot be processed due to lack of standardization. In this paper, we therefore propose a human-in-the-loop approach for the detection and classification of symbols in floor plans. The developed system calculates a measure of uncertainty for each detected symbol which is used to acquire the knowledge of human experts for those symbols that are difficult to classify. We evaluate our approach with a real-world dataset provided by an industry partner and find that the selective acquisition of human expert knowledge enhances the model’s performance by up to 10.5%—resulting in an overall prediction accuracy of 92.1% on average. We further design a pipeline for the generation of synthetic training data that allows the systems to be adapted to new construction projects with minimal manual effort. Overall, our work supports professionals in the AEC industry on their journey to the data-driven generation of business value. Johannes Jakubik, Patrick Hemmer, Michael Vössing, Benedikt Blumenstiel, Andrea Bartos, Kamilla Mohr |
AAAI | 3 |
| 2022 | A Meta-Analysis of the Utility of Explainable Artificial Intelligence in Human-AI Decision-MakingabstractResearch in artificial intelligence (AI)-assisted decision-making is experiencing tremendous growth with a constantly rising number of studies evaluating the effect of AI with and without techniques from the field of explainable AI (XAI) on human decision-making performance. However, as tasks and experimental setups vary due to different objectives, some studies report improved user decision-making performance through XAI, while others report only negligible effects. Therefore, in this article, we present an initial synthesis of existing research on XAI studies using a statistical meta-analysis to derive implications across existing research. We observe a statistically positive impact of XAI on users' performance. Additionally, the first results indicate that human-AI decision-making tends to yield better task performance on text data. However, we find no effect of explanations on users' performance compared to sole AI predictions. Our initial synthesis gives rise to future research investigating the underlying causes and contributes to further developing algorithms that effectively benefit human decision-makers by providing meaningful explanations. Max Schemmer, Patrick Hemmer, Maximilian Nitsche, Niklas Kühl 0001, Michael Vössing |
AIES | 5 |
| 2022 | Forming Effective Human-AI Teams: Building Machine Learning Models that Complement the Capabilities of Multiple ExpertsabstractMachine learning (ML) models are increasingly being used in application domains that often involve working together with human experts. In this context, it can be advantageous to defer certain instances to a single human expert when they are difficult to predict for the ML model. While previous work has focused on scenarios with one distinct human expert, in many real-world situations several human experts with varying capabilities may be available. In this work, we propose an approach that trains a classification model to complement the capabilities of multiple human experts. By jointly training the classifier together with an allocation system, the classifier learns to accurately predict those instances that are difficult for the human experts, while the allocation system learns to pass each instance to the most suitable team member—either the classifier or one of the human experts. We evaluate our proposed approach in multiple experiments on public datasets with “synthetic” experts and a real-world medical dataset annotated by multiple radiologists. Our approach outperforms prior work and is more accurate than the best human expert or a classifier. Furthermore, it is flexibly adaptable to teams of varying sizes and different levels of expert diversity. Patrick Hemmer, Sebastian Schellhammer, Michael Vössing, Johannes Jakubik, Gerhard Satzger |
IJCAI | 3 |