VLDB 2026 Research / reviewers in the wild / expert
Markus Langer
dblp:155/5791
· DBLP profile ↗
12ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-8165-1803ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design Considerations for Human Oversight of AI: Insights from Co-Design Workshops and Work Design TheoryabstractAs AI systems become increasingly capable and autonomous, domain experts’ roles are shifting from performing tasks themselves to overseeing AI-generated outputs. Such oversight is critical, as undetected errors can have serious consequences or undermine the benefits of AI. Effective oversight, however, depends not only on detecting and correcting AI errors but also on the motivation and engagement of the oversight personnel and the meaningfulness they see in their work. Yet little is known about how domain experts approach and experience the oversight task and what should be considered to design effective and motivational interfaces that support human oversight. To address these questions, we conducted four co-design workshops with domain experts from psychology and computer science. We asked them to first oversee an AI-based grading system, and then discuss their experiences and needs during oversight. Finally, they collaboratively prototyped interfaces that could support them in their oversight task. Our thematic analysis revealed four key user requirements: understanding tasks and responsibilities, gaining insight into the AI’s decision-making, contributing meaningfully to the process, and collaborating with peers and the AI. We integrated these empirical insights with the SMART model of work design to develop a framework of twelve design considerations with increased transferability compared to the identified user requirements. Our framework links interface characteristics and user requirements to the psychological processes underlying effective and satisfying work. Being grounded in work design theory and overlapping with existing guidelines for human–AI interaction, we expect these considerations to be applicable across domains and discuss how they go beyond existing guidelines for human-AI interaction to inform the design of engaging and meaningful interfaces that support human oversight of AI-based systems. Cedric Faas, Sophie Kerstan, Richard Uth, Markus Langer, Anna Maria Feit |
IUI | 4 |
| 2026 | Trust the Explanation or my Expectation? Effects of Output Accuracy and Explanations on Expectation Violations and Trust in AI-Supported Decisionsabstract• Inaccurate AI outputs led to expectation violations. • Expectation violations mediated the effects of AI output accuracy on trust. • Explanations did not moderate the link between accuracy and expectation violations. • For inaccurate AI outputs, explanations led to more trusting behavior. Systems based on Artificial Intelligence (AI) increasingly support decision-making, but their outputs may be inaccurate. Prior research has suggested that explanations might help detect inaccuracies, aiding successful human-AI interaction. This study investigates how the accuracy of system outputs influences users’ trust, trusting behavior, and trustworthiness perceptions, the role of expectation violations in this process, and how explanations for the system outputs influence these effects. In an online study with a 2(explanation vs. no explanation) × 2(accurate vs. inaccurate outputs) between-within design, 218 participants evaluated six job applicants. They received CVs and algorithmic evaluations of applicants’ suitability. For three applicants, outputs were accurate; for the other three, outputs reflected a 40% lower suitability than their true suitability. Half of the participants received explanations. Accurate outputs led to higher trustworthiness, trust, and trusting behavior than inaccurate outputs. Expectation violation fully mediated how accuracy affected trust and trustworthiness, and partially how accuracy influenced trusting behavior. Moreover, there was a significant interaction between explanations and output accuracy concerning trusting behavior: when outputs were accurate, explanations had little effect on trusting behavior; however, when outputs were inaccurate, explanations led to stronger trusting behavior, as participants less strongly deviated from the inaccurate outputs. We conclude that users are able to deviate from inaccurate outputs, and we highlight the importance of expectation violations in this regard. However, our findings also show possible detrimental effects of explanations as they can increase the decisional weight of inaccurate outputs instead of facilitating the detection of inaccuracies. Tim Hunsicker, Isabel Duhl, Pascal Haubert, Linda Onnasch, Markus Langer |
Int. J. Hum. Comput. Stud. | 5 |
| 2025 | Lay Perceptions of Algorithmic Discrimination in the Context of Systemic InjusticeabstractAlgorithmic fairness research often disregards concerns related to systemic injustice.We study how contextualizing algorithms within systemic injustice impacts lay perceptions of algorithmic discrimination.Using the hiring domain as a case-study, we conduct a 2x3 between-participants experiment (𝑁 =716), studying how people's views of algorithmic fairness are influenced by information about (i) systemic injustice in historical hiring decisions and (ii) algorithms' propensity to perpetuate biases learned from past human decisions.We find that shedding light on systemic injustice has heterogeneous effects: participants from historically advantaged groups became more negative about discriminatory algorithms, while those from disadvantaged groups reported more positive attitudes.Explaining that algorithms learn from past human decisions had null effects on people's views, adding nuances to calls for improving public understanding of algorithms.Our findings reveal that contextualizing algorithms in systemic injustice can have unintended consequences and show how different ways of framing existing inequalities influence perceptions of injustice. Gabriel Lima, Nina Grgic-Hlaca, Markus Langer, Yixin Zou |
CHI | 3 |
| 2025 | Software doping analysis for human oversightabstractAbstract This article introduces a framework that is meant to assist in mitigating societal risks that software can pose. Concretely, this encompasses facets of software doping as well as unfairness and discrimination in high-risk decision-making systems. The term software doping refers to software that contains surreptitiously added functionality that is against the interest of the user. A prominent example of software doping are the tampered emission cleaning systems that were found in millions of cars around the world when the diesel emissions scandal surfaced. The first part of this article combines the formal foundations of software doping analysis with established probabilistic falsification techniques to arrive at a black-box analysis technique for identifying undesired effects of software. We apply this technique to emission cleaning systems in diesel cars but also to high-risk systems that evaluate humans in a possibly unfair or discriminating way. We demonstrate how our approach can assist humans-in-the-loop to make better informed and more responsible decisions. This is to promote effective human oversight, which will be a central requirement enforced by the European Union’s upcoming AI Act. We complement our technical contribution with a juridically, philosophically, and psychologically informed perspective on the potential problems caused by such systems. Sebastian Biewer, Kevin Baum 0001, Sarah Sterz, Holger Hermanns, Sven Hetmank, Markus Langer, Anne Lauber-Rönsberg, Franz Lehr |
Formal Methods Syst. Des. | 6 |
| 2025 | RelEYEance: Gaze-based Assessment of Users' AI-reliance at Run-timeabstractIn time-critical detection tasks, such as drone monitoring, a key condition for users to effectively leverage AI assistance is to find an appropriate trade-off between making fast decisions and verifying AI suggestions, which we refer to as appropriate user reliance. However, assessing such reliance is often oversimplified by focusing solely on task outcomes, potentially overlooking whether users properly verify AI messages. We collected eye-tracking data from an AI-assisted monitoring task and developed a gaze-based reliance model: RelEYEance, to assess the extent of user reliance on AI-suggested alarms. We found that gaze patterns related to verification behaviors distinguish between appropriate reliance, over-reliance, and under-reliance, influencing task performance. We validated our model in a second user study, showing it can reliably detect users' over- and under-reliance at run-time, which could be used e.g. for issuing intervention messages. The results demonstrate the potential for real-time human-AI reliance assessment, facilitating adaptive reliance calibration. Zekun Wu 0001, Yao Wang 0018, Markus Langer, Anna Maria Feit |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | Taming the AI Monster: Monitoring of Individual Fairness for Effective Human Oversight
Kevin Baum 0001, Sebastian Biewer, Holger Hermanns, Sven Hetmank, Markus Langer, Anne Lauber-Rönsberg, Sarah Sterz |
SPIN | 5 |
| 2022 | "Look! It's a Computer Program! It's an Algorithm! It's AI!": Does Terminology Affect Human Perceptions and Evaluations of Algorithmic Decision-Making Systems?abstractIn the media, in policy-making, but also in research articles, algorithmic decision-making (ADM) systems are referred to as algorithms, artificial intelligence, and computer programs, amongst other terms. We hypothesize that such terminological differences can affect people’s perceptions of properties of ADM systems, people’s evaluations of systems in application contexts, and the replicability of research as findings may be influenced by terminological differences. In two studies (N = 397, N = 622), we show that terminology does indeed affect laypeople’s perceptions of system properties (e.g., perceived complexity) and evaluations of systems (e.g., trust). Our findings highlight the need to be mindful when choosing terms to describe ADM systems, because terminology can have unintended consequences, and may impact the robustness and replicability of HCI research. Additionally, our findings indicate that terminology can be used strategically (e.g., in communication about ADM systems) to influence people’s perceptions and evaluations of these systems. Markus Langer, Tim Hunsicker, Tina Feldkamp, Cornelius J. König, Nina Grgic-Hlaca |
CHI | 1 |
| 2021 | What do we want from Explainable Artificial Intelligence (XAI)? - A stakeholder perspective on XAI and a conceptual model guiding interdisciplinary XAI research
Markus Langer, Daniel Oster, Timo Speith, Holger Hermanns, Lena Kästner, Eva Schmidt, Andreas Sesing-Wagenpfeil, Kevin Baum 0001 |
Artif. Intell. | 1 |
| 2019 | Can Social Agents elicit Shame as Humans do?abstractThis paper presents a study that examines whether social agents can elicit the social emotion shame as humans do. For that, we use job interviews, which are highly evaluative situations per se. We vary the interview style (shame-eliciting vs. neutral) and the job interviewer (human vs. social agent). Our dependent variables include observational data regarding the social signals of shame and shame regulation as well as self-assessment questionnaires regarding the felt uneasiness and discomfort in the situation. Our results indicate that social agents can elicit shame to the same amount as humans. This gives insights about the impact of social agents on users and the emotional connection between them. Tanja Schneeberger, Mirella Scholtes, Bernhard Hilpert, Markus Langer, Patrick Gebhard |
ACII | 4 |
| 2019 | Explainability as a Non-Functional RequirementabstractRecent research efforts strive to aid in designing explainable systems. Nevertheless, a systematic and overarching approach to ensure explainability by design is still missing. Often it is not even clear what precisely is meant when demanding explainability. To address this challenge, we investigate the elicitation, specification, and verification of explainablity as a Non-Functional Requirement (NFR) with the long-term vision of establishing a standardized certification process for the explainability of software-driven systems in tandem with appropriate development techniques. In this work, we carve out different notions of explainability and high-level requirements people have in mind when demanding explainability, and sketch how explainability concerns may be approached in a hypothetical hiring scenario. We provide a conceptual analysis which unifies the different notions of explainability and the corresponding explainability demands. Maximilian A. Köhl, Kevin Baum 0001, Markus Langer, Daniel Oster, Timo Speith, Dimitri Bohlender |
RE | 3 |
| 2019 | Serious Games for Training Social Skills in Job InterviewsabstractIn this paper, we focus on experience-based role play with virtual agents to provide young adults at the risk of exclusion with social skill training. We present a scenario-based serious game simulation platform. It comes with a social signal interpretation component, a scripted and autonomous agent dialog and social interaction behavior model, and an engine for 3-D rendering of lifelike virtual social agents in a virtual environment. We show how two training systems developed on the basis of this simulation platform can be used to educate people in showing appropriate socioemotive reactions in job interviews. Furthermore, we give an overview of four conducted studies investigating the effect of the agents' portrayed personality and the appearance of the environment on the players' perception of the characters and the learning experience. Patrick Gebhard, Tanja Schneeberger, Elisabeth André, Tobias Baur 0001, Ionut Damian, Gregor Mehlmann, Cornelius J. König, Markus Langer |
IEEE Trans. Games | 8 |
| 2012 | Deeply Coupled GPS/INS integration in pedestrian navigation systems in weak signal conditionsabstractThis paper describes non-coherent Deeply Coupled GPS/INS integration in a pedestrian navigation system to improve position accuracy and availability in weak signal conditions. A pedestrian navigation system consists of several sensors to calculate a position of a person to guide for example rescue missions. The system presented in this paper consists of a torso mounted IMU and is used for step detection and step length and heading estimation. Additionally a barometer, magnetometer and a GPS sensor for absolute positioning are used. Since pedestrian navigation systems often are used in challenging environments like urban canyons or indoors, the use of GPS signals is often restricted. We will show that by using a Deeply Coupled GPS/INS integration system, tracking of GPS signals under weak signal conditions is possible and a seamless transition between Indoor and outdoor situations is achieved. By applying the information of a position displacement between two steps from the step length and heading estimation GPS tracking and position accuracy can be increased. For an optimal performance the system uses a deeply acquisition and re-acquisition routine. Therefore additional satellites can be used which could not have been acquired before, due to low signal to noise ratios. By carefully weighting the GPS measurements accordingly to their C / No and having a larger set of satellites available, position accuracy is increased compared to a non-vector tracking approach. The sensor fusion itself is realized in an error state space kalman filter and the step length update is performed using a state cloning technique preserving realistic position uncertainties in the filter. With this approach tracking and acquisition of GPS signals inside buildings with C / No below 20dBHz is possible. In this paper we will show that using deep integration in GPS signal tracking including step length estimations increases position accuracy of a pedestrian navigation system and availability of GPS position updates. Markus Langer, Stefan Kiesel, Christian Ascher, Gert F. Trommer |
IPIN | 1 |