VLDB 2026 Research / reviewers in the wild / expert
Ulrike Kuhl
dblp:254/7626
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-9405-918XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Performance: Comprehensive Evaluation Strategies for Impactful Machine LearningabstractEvaluation is an integral part of developing machine learning and AI-based systems for real-world applications.Given the transformative changes induced by ML/AI systems like large language models, evaluation needs to go beyond performance and include robustness, fairness, user perception, and legal compliance to ensure responsible usage.Further, evaluation practices need to also consider the full range of applications beyond the classic batch setting of machine learning, i.e., data streams, recommender systems, reinforcement learning, and foundation models.This paper provides an analysis of current evaluation practices and gaps across settings and dimensions, and argues for holistic, reproducible evaluation beyond benchmark performance.Recently, machine learning (ML) and artificial intelligence (AI) research has been grappling with an evaluation paradox: while ML systems, especially large language models (LLMs), appear to perform ever better in benchmarks, outperforming humans in many cases, this performance does not translate to success of deployed ML/AI systems in the real world, with 95% of AI projects in industry failing to provide meaningful return on investment [1,2,3].This highlights the need to rethink evaluation to keep pace with current developments.First, the appearance of foundation models, including LLMs, poses new challenges as foundation models are intended for, and hence need to be evaluated on, many different tasks at the same time [1,3].Being trained on huge amounts of data from the internet, data leakage between test benchmarks and training data becomes a considerable risk [4].Second, we observe an increase in real-world applications beyond batch machine learning, inducing additional challenges, like noisy data and dynamic environments, which need to be reflected in the evaluation process [5].Third, when applying ML/AI systems in settings that affect human users, there is a need for evaluation beyond performance: Considering robustness and safety, fairness, explainability, privacy, and legal aspects in addition to performance is paramount to ensure performant and accountable application of ML/AI systems [6].In this tutorial paper, we will analyze the state of evaluation along two axes (see Table 1), namely evaluation dimensions and settings.In terms of * VV, UK and BP gratefully acknowledge funding for the project KI-Akademie OWL, financed by the Federal Ministry of Research, Technology and Space (BMFTR) and supported by the VDI Valerie Vaquet, Ulrike Kuhl, Sasa Brdnik, Benjamin Paaßen |
ESANN | 2 |
| 2025 | The role of user feedback in enhancing understanding and trust in counterfactual explanations for explainable AI
Muhammad Suffian Nizami, Ulrike Kuhl, Alessandro Bogliolo, Jose Maria Alonso-Moral |
Int. J. Hum. Comput. Stud. | 2 |
| 2024 | Generation Gap or Diffusion Trap? How Age Affects the Detection of Personalized AI-Generated Images
René Lüdemann, Alexander Schulz 0001, Ulrike Kuhl |
CHIRA (2) | 3 |
| 2024 | Automatic Matchmaking in Two-Versus-Two Sports
Sören Rüttgers, Ulrike Kuhl, Benjamin Paaßen |
EDM | 2 |
| 2022 | Intuitiveness in Active TeachingabstractWhile machine learning (ML) gives rise to astonishing results in automated systems, it is usually at the cost of large data requirements. This makes many successful algorithms from ML unsuitable for human-machine interaction, where the machine must learn from a small number of training samples that can be provided by a user within a reasonable time frame. Fortunately, the user can tailor the training data they create to be as useful as possible, severely limiting its necessary size—as long as they know about the machine’s requirements and limitations. Of course, acquiring this knowledge can in turn be cumbersome and costly. This raises the question of how easy ML algorithms are to interact with. In this work, we address this issue by analyzing the intuitiveness of certain algorithms when they are actively taught by users. After developing a theoretical framework of intuitiveness as a property of algorithms, we introduce an active teaching paradigm involving a prototypical two-dimensional spatial learning task as a method to judge the efficacy of human-machine interactions. Finally, we present and discuss the results of a large-scale user study into the performance and teaching strategies of 800 users interacting with two prominent ML algorithms in our system, providing first evidence for the role of intuition as an important factor impacting human-machine interaction. Jan Philip Göpfert, Ulrike Kuhl, Lukas Hindemith, Heiko Wersing, Barbara Hammer |
IEEE Trans. Hum. Mach. Syst. | 2 |