VLDB 2026 Research / reviewers in the wild / expert
Merlin Knaeble
dblp:266/9034 · also Merlin Knäble
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-5108-4609ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Paintings, Not Noise - The Role of Presentation Sequence in LabelingabstractAbstract Labeling is critical in creating training datasets for supervised machine learning, and is a common form of crowd work heteromation. It typically requires manual labor, is badly compensated and not infrequently bores the workers involved. Although task variety is known to drive human autonomy and intrinsic motivation, there is little research in this regard in the labeling context. Against this backdrop, we manipulate the presentation sequence of a labeling task in an online experiment and use the theoretical lens of self-determination theory to explain psychological work outcomes and work performance. We rely on 176 crowd workers contributing with group comparisons between three presentation sequences (by label, by image, random) and a mediation path analysis along the phenomena studied. Surprising among our key findings is that the task variety when sorting by label is perceived higher than when sorting by image and the random group. Naturally, one would assume that the random group would be perceived as most varied. We choose a visual metaphor to explain this phenomenon, whereas paintings offer a structured presentation of coloured pixels, as opposed to random noise. Merlin Knaeble, Mario Nadj, Alexander Maedche |
Interact. Comput. | 1 |
| 2025 | StoryPoint: GenAI-supported domain-specific data story authoring for enterprisesabstractIn today’s data-driven world, enterprises face the dual challenge of deriving value from vast datasets and effectively communicating insights. While data storytelling is integral for conveying insights, existing authoring tools often fail to fully leverage domain expertise or support the entire storytelling process. This paper introduces StoryPoint, an open-source data story authoring tool that combines domain expertise with Generative AI to dynamically enhance data story creation. We found that enabling domain experts to intuitively visualize data via natural language inputs, supported by automatically generated charts and narratives, helps narrow the gap between visualization and interpretation. To design StoryPoint, we followed a literature-grounded and user-centered design approach. A formative evaluation with eight domain experts reveals StoryPoint’s efficacy in rapid data story prototyping. A summative evaluation, including (i) a small-scale experiment comparing StoryPoint with a benchmark, (ii) a large-scale online experiment with 104 crowdworkers, and (iii) three real-world industry cases, underscores its utility. Our findings highlight that StoryPoint reduces data story creation time by more than 50% compared to the benchmark, receives significantly higher usability ratings, and supports the creation of data stories rating higher in readability, fluency, clarity, and trustworthiness. • StoryPoint - first open-source data story authoring tool for enterprises. • Combining LLMs with domain expertise enhances enterprise data storytelling workflows. • Empirical evaluations with domain experts and real-world use cases demonstrate StoryPoint’s efficiency, usability, and ability to create high-quality data stories. • Evaluation scales for assessing narrative elements and data story quality in enterprise contexts. Jonas Gunklach, Elias Müller, Merlin Knaeble, Alexander Maedche |
Int. J. Hum. Comput. Stud. | 3 |
| 2024 | HILL: A Hallucination Identifier for Large Language ModelsabstractLarge language models (LLMs) are prone to hallucinations, i.e., nonsensical, unfaithful, and undesirable text. Users tend to overrely on LLMs and corresponding hallucinations which can lead to misinterpretations and errors. To tackle the problem of overreliance, we propose HILL, the "Hallucination Identifier for Large Language Models". First, we identified design features for HILL with a Wizard of Oz approach with nine participants. Subsequently, we implemented HILL based on the identified design features and evaluated HILL’s interface design by surveying 17 participants. Further, we investigated HILL’s functionality to identify hallucinations based on an existing question-answering dataset and five user interviews. We find that HILL can correctly identify and highlight hallucinations in LLM responses which enables users to handle LLM responses with more caution. With that, we propose an easy-to-implement adaptation to existing LLMs and demonstrate the relevance of user-centered designs of AI artifacts. Florian Leiser, Sven Eckhardt, Valentin Leuthe, Merlin Knaeble, Alexander Maedche, Gerhard Schwabe, Ali Sunyaev |
CHI | 4 |
| 2023 | "Garbage In, Garbage Out": Mitigating Human Biases in Data Entry by Means of Artificial Intelligence
Sven Eckhardt, Merlin Knaeble, Andreas Bucher, Dario Staehelin, Mateusz Dolata, Doris Agotai, Gerhard Schwabe |
INTERACT (3) | 2 |
| 2022 | Towards Automatic Parsing of Structured Visual Content through the Use of Synthetic DataabstractStructured Visual Content (SVC) such as graphs, flow charts, or the like are used by authors to illustrate various concepts. While such depictions allow the average reader to better understand the contents, images containing SVCs are typically not machine-readable. This, in turn, not only hinders automated knowledge aggregation, but also the perception of displayed information for visually impaired people. In this work, we propose a synthetic dataset, containing SVCs in the form of images as well as ground truths. We show the usage of this dataset by an application that automatically extracts a graph representation from an SVC image. This is done by training a model via common supervised learning methods. As there currently exist no large-scale public datasets for the detailed analysis of SVC, we propose the Synthetic SVC (SSVC) dataset comprising 12,000 images with respective bounding box annotations and detailed graph representations. Our dataset enables the development of strong models for the interpretation of SVCs while skipping the time-consuming dense data annotation.We evaluate our model on both synthetic and manually annotated data and show the transferability of synthetic to real via various metrics, given the presented application. Here, we evaluate that this proof of concept is possible to some extend and lay down a solid baseline for this task. We discuss the limitations of our approach for further improvements. Our utilized metrics can be used as a tool for future comparisons in this domain. To enable further research on this task, the dataset is publicly available at https://bit.ly/3jN1pJJ. Lukas Schölch, Jonas Steinhäuser, Maximilian Beichter, Constantin Seibold, Kailun Yang 0001, Merlin Knaeble, Thorsten Schwarz, Alexander Maedche, Rainer Stiefelhagen |
ICPR | 6 |