EDBT 2026 Demo / reviewers in the wild / expert
Johannes Kirmayr
dblp:304/3211
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2026
0009-0004-6661-5005ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 50% Reinforcement learning · 25% Knowledge representation and reasoning · 25% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 81% Interaction techniques and input · 19% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
agent evaluation |
1.0 | 1 | 2026 | CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty · ACL (1) 2026 |
Natural language and speech › Language models and text generation
LLM agents |
1.0 | 1 | 2026 | CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty · ACL (1) 2026 |
Natural language and speech › Language models and text generation › LLM agents
tool use |
1.0 | 1 | 2026 | CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty · ACL (1) 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
uncertainty reasoning |
1.0 | 1 | 2026 | CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty · ACL (1) 2026 |
Interaction techniques and input
in-vehicle interaction |
0.3 | 1 | 2026 | "What Are You Doing?": Effects of Intermediate Feedback from Agentic LLM In-Car Assistants During Multi-Step Processing · CHI 2026 |
Human-AI interaction
voice assistants |
0.3 | 1 | 2026 | "What Are You Doing?": Effects of Intermediate Feedback from Agentic LLM In-Car Assistants During Multi-Step Processing · CHI 2026 |
Methods — techniques the papers use, named apart from their topics
mixed-methods study · 1.0interview study · 1.0dual-task paradigm · 1.0benchmark · 1.0LLM-simulated user · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World UncertaintyabstractExisting benchmarks for Large Language Model (LLM) agents focus on task completion under idealistic settings but overlook reliability in real-world, user-facing applications.In domains, such as in-car voice assistants, users often issue incomplete or ambiguous requests, creating intrinsic uncertainty that agents must manage through dialogue, tool use, and policy adherence.We introduce CAR-bench, a benchmark for evaluating consistency, uncertainty handling, and capability awareness in multi-turn, tool-using LLM agents in an in-car assistant domain.The environment features an LLM-simulated user, domain policies, and 58 interconnected tools spanning navigation, productivity, charging, and vehicle control.Beyond standard task completion, CAR-bench introduces Hallucination tasks that test agents' limit-awareness under missing tools or information, and Disambiguation tasks that require resolving uncertainty through clarification or internal information gathering.Baseline results reveal large gaps between occasional and consistent success on all task types.Even frontier reasoning LLMs achieve less than 50% consistent pass rate on Disambiguation tasks due to premature actions, and frequently violate policies or fabricate information to satisfy user requests in Hallucination tasks, underscoring the need for more reliable and self-aware LLM agents in real-world settings.1 Johannes Kirmayr, Lukas Stappen, Elisabeth André |
ACL (1) | 1 |
| 2026 | "What Are You Doing?": Effects of Intermediate Feedback from Agentic LLM In-Car Assistants During Multi-Step ProcessingabstractAgentic AI assistants that autonomously perform multi-step tasks raise open questions for user experience: how should such systems communicate progress and reasoning during extended operations, especially in attention-critical contexts such as driving? We investigate feedback timing and verbosity from agentic LLM-based in-car assistants through a controlled, mixed-methods study (N=45) comparing planned steps and intermediate results feedback against silent operation with final-only response. Using a dual-task paradigm with an in-car voice assistant, we found that intermediate feedback significantly improved perceived speed, trust, and user experience while reducing task load - effects that held across varying task complexities and interaction contexts. Interviews further revealed user preferences for an adaptive approach: high initial transparency to establish trust, followed by progressively reducing verbosity as systems prove reliable, with adjustments based on task stakes and situational context. We translate our empirical findings into design implications for feedback timing and verbosity in agentic in-car assistants, balancing transparency and efficiency. Johannes Kirmayr, Raphael Wennmacher, Khanh Huynh, Lukas Stappen, Elisabeth André, Florian Alt |
CHI | 1 |