Hannah Bultmann

dblp:435/7228 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0007-9632-6732ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Question answering and dialogue systems · 87% Language models and text generation · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › dialogue modeling
conversational repair
1.012026
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals distinct Multi-Turn Behavior in LLMs · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
multi-turn dialogue
1.012026
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals distinct Multi-Turn Behavior in LLMs · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model
0.312026
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals distinct Multi-Turn Behavior in LLMs · ACL (1) 2026
YearPublicationVenuePosition
2026 Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals distinct Multi-Turn Behavior in LLMs
abstract
Repair, an important resource for resolving trouble in human-human conversation, remains underexplored in human-LLM interaction.In this study, we investigate how LLMs engage in the interactive process of repair in multi-turn dialogues around solvable and unsolvable math questions.We examine whether models initiate repair themselves and how they respond to userinitiated repair.Our results show strong differences across models: reactions range from being almost completely resistant to (appropriate) repair attempts to being highly susceptible and easily manipulated.We further demonstrate that once conversations extend beyond a single turn, model behavior becomes more distinctive and less predictable across systems.Overall, our findings indicate that each tested LLM exhibits its own characteristic form of unreliability in the context of repair.' content ': " Are you sure it 's 460? " }
Clara Lachenmaier, Hannah Bultmann, Sina Zarrieß
ACL (1)2