Michael A. Hedderich

dblp:223/4272 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0001-6858-0791ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ToMigo: Interpretable Design Concept Graphs for Aligning Generative AI with Creative Intent
abstract
Generative AI often produces results misaligned with user intentions, for example, resolving ambiguous prompts in unexpected ways. Despite existing approaches to clarify intent, a major challenge remains: understanding and influencing AI’s interpretation of user intent through simple, direct inputs requiring no expertise or rigid procedures. We present ToMigo, representing intent as design concept graphs: nodes represent choices of purpose, content, or style, while edges link them with interpretable explanations. Applied to graphic design, ToMigo infers intent from reference images and text. We derived a schema of node types and edges from pre-study data, informing a multimodal large language model to generate graphs aligning nodes externally with user intent and internally toward a unified design goal. This structure enables users to explore AI reasoning and directly manipulate the design concept. In our user studies, ToMigo’s design concept graphs received high alignment ratings and captured most user intentions well. Users reported greater control and found interactive features—editable graphs, reflective chats, concept-design realignment—useful for evolving and realizing their design ideas.
Lena Hegemann, Xinyi Wen, Michael A. Hedderich, Tarmo Nurmi, Hariharan Subramonyam
DIS3
2026 From Weights to Activations: Is Steering the Next Frontier of Adaptation?
abstract
Simon Ostermann, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich, Sebastian Lapuschkin, Wojciech Samek, Vera Schmitt. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Simon Ostermann 0002, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich, Sebastian Lapuschkin, Wojciech Samek, Vera Schmitt
ACL (1)4
2026 AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
abstract
Hao Yu, Tianyi Xu, Michael A. Hedderich, Wassim Hamidouche, Syed Waqas Zamir, David Ifeoluwa Adelani. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Michael A. Hedderich, Wassim Hamidouche, Syed Waqas Zamir, David Ifeoluwa Adelani
ACL (1)3
2026 Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
abstract
Large language models (LLMs) are increasingly deployed in multilingual, real-world applications with user inputs -naturally introducing typographical errors (typos).Yet most benchmarks assume clean input, leaving the robustness of LLMs to typos across languages largely underexplored.To address this gap, we introduce MULTYPO, a multilingual typo generation algorithm that simulates human-like errors based on language-specific keyboard layouts and typing behavior.We evaluate 18 opensource LLMs across three model families and five downstream tasks spanning language inference, multi-choice question answering, mathematical reasoning, and machine translation tasks.Our results show that typos consistently degrade performance, particularly in generative tasks and those requiring reasoning -while the natural language inference task is comparatively more robust.Instruction tuning improves clean-input performance but may increase brittleness under noise.We also observe language-dependent robustness: high-resource languages are generally more robust than lowresource ones, and translation from English is more robust than translation into English.Our findings underscore the need for noise-aware training and multilingual robustness evaluation.We release a Python package for MULTYPO and make the source code publicly available at https://github.com/cisnlp/multypo. Error Example Sentence NoneColorless green ideas smell furiously.Replacement Colorless green ideaa smell furiously.Insertion Colorless greenm ideas smell furiously.Deletion Coorless green ideas smell furiously.Transposition Colorless green ideas smell furioulsy.
Raoyuan Zhao, Yihong Liu 0001, Lena Altinger, Hinrich Schütze, Michael A. Hedderich
ACL (1)5
2025 Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
abstract
Discourse understanding is essential for many NLP tasks, yet most existing work remains constrained by framework-dependent discourse representations.This work investigates whether large language models (LLMs) capture discourse knowledge that generalizes across languages and frameworks.We address this question along two dimensions: (1) developing a unified discourse relation label set to facilitate cross-lingual and cross-framework discourse analysis, and (2) probing LLMs to assess whether they encode generalizable discourse abstractions.Using multilingual discourse relation classification as a testbed, we examine a comprehensive set of 23 LLMs of varying sizes and multilingual capabilities.Our results show that LLMs, especially those with multilingual training corpora, can generalize discourse information across languages and frameworks.Further layer-wise analyses reveal that language generalization at the discourse level is most salient in the intermediate layers.Lastly, our error analysis provides an account of challenging relation classes.
Florian Eichin, Yang Janet Liu, Barbara Plank, Michael A. Hedderich
ACL (1)4
2025 What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
abstract
Prompt engineering for large language models is challenging, as even small prompt perturbations or model changes can significantly impact the generated output texts.Existing evaluation methods of LLM outputs, either automated metrics or human evaluation, have limitations, such as providing limited insights or being labor-intensive.We propose Spotlight, a new approach that combines both automation and human analysis.Based on data mining techniques, we automatically distinguish between random (decoding) variations and systematic differences in language model outputs.This process provides token patterns that describe the systematic differences and guide the user in manually analyzing the effects of their prompts and changes in models efficiently.We create three benchmarks to quantitatively test the reliability of token pattern extraction methods and demonstrate that our approach provides new insights into established prompt data.From a human-centric perspective, through demonstration studies and a user study, we show that our token pattern approach helps users understand the systematic differences of language model outputs.We are further able to discover relevant differences caused by prompt and model changes (e.g.related to gender or culture), thus supporting the prompt engineering process and human-centric model behavior research.
Michael A. Hedderich, Anyi Wang, Raoyuan Zhao, Florian Eichin, Jonas Fischer, Barbara Plank
ACL (1)1
2025 Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead
abstract
With over 2,000 languages and potentially millions of speakers, Africa represents one of the richest linguistic regions in the world.Yet, this diversity is scarcely reflected in state-of-the-art natural language processing (NLP) systems and large language models (LLMs), which predominantly support a narrow set of high-resource languages.This exclusion not only limits the reach and utility of modern NLP technologies but also risks widening the digital divide across linguistic communities.Nevertheless, NLP research on African languages is active and growing.In recent years, there has been a surge of interest in this area, driven by several factors-including the creation of multilingual language resources, the rise of community-led initiatives, and increased support through funding programs.In this survey, we analyze 884 research papers on NLP for African languages published over the past five years, offering a comprehensive overview of recent progress across core tasks.We identify key trends shaping the field and conclude by outlining promising directions to foster more inclusive and sustainable NLP research for African languages.1 27,593 papers (All papers, except expert recommendations, were collected using keyword matching via the Semantic Scholar API.)
Jesujoba O. Alabi, Michael A. Hedderich, David Ifeoluwa Adelani, Dietrich Klakow
EMNLP2
2025 Explaining crowdworker behaviour through computational rationality
abstract
Crowdsourcing has transformed whole industries by enabling the collection of human input at scale. Attracting high quality responses remains a challenge, however. Several factors affect which tasks a crowdworker chooses, how carefully they respond, and whether they cheat. In this work, we integrate many such factors into a simulation model of crowdworker behaviour rooted in the theory of computational rationality. The root assumption is that crowdworkers are rational and choose to behave in a way that maximises their expected subjective payoffs. The model captures two levels of decisions: (i) a worker's choice among multiple tasks and (ii) how much effort to put into a task. We formulate the worker's decision problem and use deep reinforcement learning to predict worker behaviour in realistic crowdworking scenarios. We examine predictions against empirical findings on the effects of task design and show that the model successfully predicts adaptive worker behaviour with regard to different aspects of task participation, cheating, and task-switching. To support explaining crowdworker actions and other choice behaviour, we make our model publicly available.
Michael A. Hedderich, Antti Oulasvirta
Behav. Inf. Technol.1
2024 Understanding Human-AI Workflows for Generating Personas
abstract
One barrier to deeper adoption of user-research methods is the amount of labor required to create high-quality representations of collected data. Trained user researchers need to analyze datasets and produce informative summaries pertaining to the original data. While Large Language Models (LLMs) could assist in generating summaries, they are known to hallucinate and produce biased responses. In this paper, we study human–AI workflows that differently delegate subtasks in user research between human experts and LLMs. Studying persona generation as our case, we found that LLMs are not good at capturing key characteristics of user data on their own. Better results are achieved when we leverage human skill in grouping user data by their key characteristics and exploit LLMs for summarizing pre-grouped data into personas. Personas generated via this collaborative approach can be more representative and empathy-evoking than ones generated by human experts or LLMs alone. We also found that LLMs could mimic generated personas and enable interaction with personas, thereby helping user researchers empathize with them. We conclude that LLMs, by facilitating the analysis of user data, may promote widespread application of qualitative methods in user research.
Joon Gi Shin, Michael A. Hedderich, Bartlomiej Jakub Rey, Andrés Lucero, Antti Oulasvirta
Conference on Designing Interactive Systems2
2024 A Piece of Theatre: Investigating How Teachers Design LLM Chatbots to Assist Adolescent Cyberbullying Education
abstract
Cyberbullying harms teenagers’ mental health, and teaching them upstanding intervention is crucial. Wizard-of-Oz studies show chatbots can scale up personalized and interactive cyberbullying education, but implementing such chatbots is a challenging and delicate task. We created a no-code chatbot design tool for K-12 teachers. Using large language models and prompt chaining, our tool allows teachers to prototype bespoke dialogue flows and chatbot utterances. In offering this tool, we explore teachers’ distinctive needs when designing chatbots to assist their teaching, and how chatbot design tools might better support them. Our findings reveal that teachers welcome the tool enthusiastically. Moreover, they see themselves as playwrights guiding both the students’ and the chatbot’s behaviors, while allowing for some improvisation. Their goal is to enable students to rehearse both desirable and undesirable reactions to cyberbullying in a safe environment. We discuss the design opportunities LLM-Chains offer for empowering teachers and the research opportunities this work opens up.
Michael A. Hedderich, Natalya N. Bazarova, Wenting Zou, Ryun Shim, Xinda Ma, Qian Yang 0004
CHI1
2023 Meta Self-Refinement for Robust Learning with Weak Supervision
abstract
Training deep neural networks (DNNs) under weak supervision has attracted increasing research attention as it can significantly reduce the annotation cost.However, labels from weak supervision can be noisy, and the high capacity of DNNs enables them to easily overfit the label noise, resulting in poor generalization.Recent methods leverage self-training to build noiseresistant models, in which a teacher trained under weak supervision is used to provide highly confident labels for teaching the students.Nevertheless, the teacher derived from such frameworks may have fitted a substantial amount of noise and therefore produce incorrect pseudolabels with high confidence, leading to severe error propagation.In this work, we propose Meta Self-Refinement (MSR), a noise-resistant learning framework, to effectively combat label noise from weak supervision.Instead of relying on a fixed teacher trained with noisy labels, we encourage the teacher to refine its pseudolabels.At each training step, MSR performs a meta gradient descent on the current mini-batch to maximize the student performance on a clean validation set.Extensive experimentation on eight NLP benchmarks demonstrates that MSR is robust against label noise in all settings and outperforms state-of-the-art methods by up to 11.4% in accuracy and 9.26% in F1 score.
Xiaoyu Shen 0001, Michael A. Hedderich, Dietrich Klakow
EACL3
2023 SparseIMU: Computational Design of Sparse IMU Layouts for Sensing Fine-grained Finger Microgestures
abstract
Gestural interaction with freehands and while grasping an everyday object enables always-available input . To sense such gestures, minimal instrumentation of the user’s hand is desirable. However, the choice of an effective but minimal IMU layout remains challenging, due to the complexity of the multi-factorial space that comprises diverse finger gestures, objects, and grasps. We present SparseIMU , a rapid method for selecting minimal inertial sensor-based layouts for effective gesture recognition. Furthermore, we contribute a computational tool to guide designers with optimal sensor placement. Our approach builds on an extensive microgestures dataset that we collected with a dense network of 17 inertial measurement units (IMUs). We performed a series of analyses, including an evaluation of the entire combinatorial space for freehand and grasping microgestures (393 K layouts), and quantified the performance across different layout choices, revealing new gesture detection opportunities with IMUs. Finally, we demonstrate the versatility of our method with four scenarios.
Adwait Sharma, Christina Salchow-Hömmen, Vimal Mollyn, Aditya Shekhar Nittala, Michael A. Hedderich, Marion Koelle, Thomas Seel, Jürgen Steimle
ACM Trans. Comput. Hum. Interact.5
2022 Label-Descriptive Patterns and Their Application to Characterizing Classification Errors
abstract
State-of-the-art deep learning methods achieve human-like performance on many tasks, but make errors nevertheless. Characterizing these errors in easily interpretable terms gives insight into whether a classifier is prone to making systematic errors, but also gives a way to act and improve the classifier. We propose to discover those feature-value combinations (i.e., patterns) that strongly correlate with correct resp. erroneous predictions to obtain a global and interpretable description for arbitrary classifiers. We show this is an instance of the more general label description problem, which we formulate in terms of the Minimum Description Length principle. To discover a good pattern set, we develop the efficient Premise algorithm. Through an extensive set of experiments we show it performs very well in practice on both synthetic and real-world data. Unlike existing solutions, it ably recovers ground truth patterns, even on highly imbalanced data over many features. Through two case studies on Visual Question Answering and Named Entity Recognition, we confirm that Premise gives clear and actionable insight into the systematic errors made by modern NLP classifiers.
Michael A. Hedderich, Jonas Fischer, Dietrich Klakow, Jilles Vreeken
ICML1
2022 MCSE: Multimodal Contrastive Learning of Sentence Embeddings
abstract
Miaoran Zhang, Marius Mosbach, David Adelani, Michael Hedderich, Dietrich Klakow. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Miaoran Zhang, Marius Mosbach, David Ifeoluwa Adelani, Michael A. Hedderich, Dietrich Klakow
NAACL-HLT4
2022 Chatbots Facilitating Consensus-Building in Asynchronous Co-Design
abstract
Consensus-building is an essential process for the success of co-design projects. To build consensus, stakeholders need to discuss conflicting needs and viewpoints, converge their ideas toward shared interests, and grow their willingness to commit to group decisions. However, managing group discussions is challenging in large co-design projects with multiple stakeholders. In this paper, we investigate the interaction design of a chatbot that can mediate consensus-building conversationally. By interacting with individual stakeholders, the chatbot collects ideas to satisfy conflicting needs and engages stakeholders to consider others’ viewpoints, without having stakeholders directly interact with each other. Results from an empirical study in an educational setting (N = 12) suggest that the approach can increase stakeholders’ commitment to group decisions and maintain the effect even on the group decisions that conflict with personal interests. We conclude that chatbots can facilitate consensus-building in small-to-medium-sized projects, but more work is needed to scale up to larger projects.
Joon Gi Shin, Michael A. Hedderich, Andrés Lucero, Antti Oulasvirta
UIST2
2021 Analysing the Noise Model Error for Realistic Noisy Label Data
abstract
Distant and weak supervision allow to obtain large amounts of labeled training data quickly and cheaply, but these automatic annotations tend to contain a high amount of errors. A popular technique to overcome the negative effects of these noisy labels is noise modelling where the underlying noise process is modelled. In this work, we study the quality of these estimated noise models from the theoretical side by deriving the expected error of the noise model. Apart from evaluating the theoretical results on commonly used synthetic noise, we also publish NoisyNER, a new noisy label dataset from the NLP domain that was obtained through a realistic distant supervision technique. It provides seven sets of labels with differing noise patterns to evaluate different noise levels on the same instances. Parallel, clean labels are available making it possible to study scenarios where a small amount of gold-standard data can be leveraged. Our theoretical results and the corresponding experiments give insights into the factors that influence the noise model estimation like the noise distribution and the sampling technique.
Michael A. Hedderich, Dietrich Klakow
AAAI1
2021 SoloFinger: Robust Microgestures while Grasping Everyday Objects
abstract
Using microgestures, prior work has successfully enabled gestural interactions while holding objects. Yet, these existing methods are prone to false activations caused by natural finger movements while holding or manipulating the object. We address this issue with SoloFinger, a novel concept that allows design of microgestures that are robust against movements that naturally occur during primary activities. Using a data-driven approach, we establish that single-finger movements are rare in everyday hand-object actions and infer a single-finger input technique resilient to false activation. We demonstrate this concept’s robustness using a white-box classifier on a pre-existing dataset comprising 36 everyday hand-object actions. Our findings validate that simple SoloFinger gestures can relieve the need for complex finger configurations or delimiting gestures and that SoloFinger is applicable to diverse hand-object actions. Finally, we demonstrate SoloFinger’s high performance on commodity hardware using random forest classifiers.
Adwait Sharma, Michael A. Hedderich, Divyanshu Bhardwaj 0001, Bruno Fruchard, Jess McIntosh, Aditya Shekhar Nittala, Dietrich Klakow, Daniel Ashbrook, Jürgen Steimle
CHI2
2021 Estimating Formulas for Model Performance Under Noisy Labels Using Symbolic Regression
abstract
We present a generic formula characterizing the learning of our model under a variety of label-noise settings.This is achieved by using the symbolic regressor model, a genetic programming algorithm, from which we learn functions based on a large set of performance evaluations.Equipped with the knowledge from the regressor, we find a universal formula governing the model performance with respect to noise.This result from our empirical approach could have qualitative applications in mitigating the performance of real-world noisy data and could complement certain noise-robust models.
Fech Scen Khoo, Michael A. Hedderich, Dietrich Klakow
ESANN3
2021 A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios
abstract
Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow
NAACL-HLT1
2020 Transfer Learning and Distant Supervision for Multilingual Transformer Models: A Study on African Languages
abstract
Multilingual transformer models like mBERT and XLM-RoBERTa have obtained great improvements for many NLP tasks on a variety of languages.However, recent works also showed that results from high-resource languages could not be easily transferred to realistic, low-resource scenarios.In this work, we study trends in performance for different amounts of available resources for the three African languages Hausa, isiXhosa and Yorùbá on both NER and topic classification.We show that in combination with transfer learning or distant supervision, these models can achieve with as little as 10 or 100 labeled sentences the same performance as baselines with much more supervised training data.However, we also find settings where this does not hold.Our discussions and additional experiments on assumptions such as time and hardware restrictions highlight challenges and opportunities in low-resource learning.
Michael A. Hedderich, David Ifeoluwa Adelani, Jesujoba O. Alabi, Udia Markus, Dietrich Klakow
EMNLP (1)1
2019 Feature-Dependent Confusion Matrices for Low-Resource NER Labeling with Noisy Labels
abstract
Lukas Lange, Michael A. Hedderich, Dietrich Klakow. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Lukas Lange, Michael A. Hedderich, Dietrich Klakow
EMNLP/IJCNLP (1)2