VLDB 2026 Research / reviewers in the wild / expert
Jan Alexandersson
dblp:40/581
· DBLP profile ↗
23ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-6676-3145ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MultiMediate '25: Cross-cultural Multi-domain Engagement EstimationabstractEstimating momentary conversational engagement is central to assistive, socially aware AI systems, yet models are typically trained and evaluated within a single domain, limiting real-world robustness. The MultiMediate '25 challenge advances engagement estimation to more challenging, cross-cultural, and multi-domain settings. Building on prior challenge editions, we expand beyond NOXI as the sole training source by introducing NOXI-J, a new multilingual corpus covering Japanese and Chinese interactions, enabling both training and evaluation in diverse linguistic contexts. Although NOXI-J conceptually extends NOXI, we treat it as a distinct domain because linguistic, cultural, capture, and annotation differences induce measurable distribution shifts. In this paper, we present new annotations, precomputed multi-modal features (visual, vocal, and verbal), baseline evaluations, and an analysis of the best performing challenge solutions. Beyond accuracy, we quantify fairness using Conditional Demographic Disparity for gender and language. Our baselines confirm strong in-domain performance (e.g., paralinguistic eGeMAPS and video-transformer features) and reveal notable cross-domain drops, underscoring the challenge of cultural, linguistic, and interactional shifts. Fairness analyses indicate generally small discrepancies for our baselines. We observe the largest disparities for the proposed challenge solutions on the Chinese language test set. All annotations, features, code, and leaderboards are made publicly available to foster sustained progress on robust and fair engagement estimation. Daksitha Withanage, Marius Funk, Michal Balazia, Huajian Qiu, Shogo Okada, François Brémond, Jan Alexandersson, Andreas Bulling, Elisabeth André, Philipp Müller 0001 |
ACM Multimedia | 7 |
| 2024 | Recognizing Emotion Regulation Strategies from Human Behavior with Large Language ModelsabstractHuman emotions are often not expressed directly, but regulated according to internal processes and social display rules. For affective computing systems, an understanding of how users regulate their emotions can be highly useful, for example to provide feedback in job interview training, or in psychotherapeutic scenarios. However, at present no method to automatically classify different emotion regulation strategies in a cross-user scenario exists. At the same time, recent studies showed that instruction-tuned Large Language Models (LLMs) can reach impressive performance across a variety of affect recognition tasks such as categorical emotion recognition or sentiment analysis. While these results are promising, it remains unclear to what extent the representational power of LLMs can be utilized in the more subtle task of classifying users' internal emotion regulation strategy. To close this gap, we make use of the recently introduced Deep corpus for modeling the social display of the emotion shame, where each point in time is annotated with one of seven different emotion regulation classes. We fine-tune Llama2-7B as well as the recently introduced Gemma model using Low-rank Optimization on prompts generated from different sources of information on the Deep corpus. These include verbal and nonverbal behavior, person factors, as well as the results of an indepth interview after the interaction. Our results show, that a fine-tuned Llama2-7B LLM is able to classify the utilized emotion regulation strategy with high accuracy (0.84) without needing access to data from post-interaction interviews. This represents a significant improvement over previous approaches based on Bayesian Networks and highlights the importance of modeling verbal behavior in emotion regulation. Philipp Müller 0001, Alexander Heimerl, Sayed Muddashir Hossain, Lea Siegel, Jan Alexandersson, Patrick Gebhard, Elisabeth André, Tanja Schneeberger |
ACII | 5 |
| 2024 | M3TCM: Multi-modal Multi-task Context Model for Utterance Classification in Motivational InterviewsabstractAccurate utterance classification in motivational interviews is crucial to automatically understand the quality and dynamics of client-therapist interaction, and it can serve as a key input for systems mediating such interactions. Motivational interviews exhibit three important characteristics. First, there are two distinct roles, namely client and therapist. Second, they are often highly emotionally charged, which can be expressed both in text and in prosody. Finally, context is of central importance to classify any given utterance. Previous works did not adequately incorporate all of these characteristics into utterance classification approaches for mental health dialogues. In contrast, we present M3TCM, a Multi-modal, Multi-task Context Model for utterance classification. Our approach for the first time employs multi-task learning to effectively model both joint and individual components of therapist and client behaviour. Furthermore, M3TCM integrates information from the text and speech modality as well as the conversation context. With our novel approach, we outperform the state of the art for utterance classification on the recently introduced AnnoMI dataset with a relative improvement of 20% for the client- and by 15% for therapist utterance classification. In extensive ablation studies, we quantify the improvement resulting from each contribution. Sayed Muddashir Hossain, Jan Alexandersson, Philipp Müller 0001 |
LREC/COLING | 2 |
| 2024 | MultiMediate'24: Multi-Domain Engagement EstimationabstractEstimating the momentary level of participant's engagement is an important prerequisite for assistive systems that support human interactions. Previous work has addressed this task in within-domain evaluation scenarios, i.e. training and testing on the same dataset. This is in contrast to real-life scenarios where domain shifts between training and testing data frequently occur. With MultiMediate'24, we present the first challenge addressing multi-domain engagement estimation. As training data, we utilise the NOXI database of dyadic novice-expert interactions. In addition to within-domain test data, we add two new test domains. First, we introduce recordings following the NOXI protocol but covering languages that are not present in the NOXI training data. Second, we collected novel engagement annotations on the MPIIGroupInteraction dataset which consists of group discussions between three to four people. In this way, MultiMediate'24 evaluates the ability of approaches to generalise across factors such as language and cultural background, group size, task, and screen-mediated vs. face-to-face interaction. This paper describes the MultiMediate'24 challenge and presents baseline results. In addition, we discuss selected challenge solutions. Philipp Müller 0001, Michal Balazia, Tobias Baur 0001, Michael Dietz, Alexander Heimerl, Anna Penzkofer, Dominik Schiller, François Brémond, Jan Alexandersson, Elisabeth André, Andreas Bulling |
ACM Multimedia | 9 |
| 2023 | MultiMediate '23: Engagement Estimation and Bodily Behaviour Recognition in Social InteractionsabstractAutomatic analysis of human behaviour is a fundamental prerequisite for the creation of machines that can effectively interact with- and support humans in social interactions. In MultiMediate'23, we address two key human social behaviour analysis tasks for the first time in a controlled challenge: engagement estimation and bodily behaviour recognition in social interactions. This paper describes the MultiMediate'23 challenge and presents novel sets of annotations for both tasks. For engagement estimation we collected novel annotations on the NOvice eXpert Interaction (NOXI) database. For bodily behaviour recognition, we annotated test recordings of the MPIIGroupInteraction corpus with the BBSI annotation scheme. In addition, we present baseline results for both challenge tasks. Philipp Müller 0001, Michal Balazia, Tobias Baur 0001, Michael Dietz, Alexander Heimerl, Dominik Schiller, Mohammed Guermal, Dominike Thomas, François Brémond, Jan Alexandersson, Elisabeth André, Andreas Bulling |
ACM Multimedia | 10 |
| 2022 | Towards Improving EEG-Based Intent Recognition in Visual Search Tasks
Maurice Rekrut, Jan Alexandersson, Antonio Krüger |
ICONIP (3) | 3 |
| 2022 | Too old for technology? Stereotype threat and technology use by older adultsabstractOlder adults are often stereotyped as having less technological ability than younger age groups. As a result, older individuals may avoid using technology due to stereotype threat, the fear of confirming negative stereotypes about their social group. The present research examined the role of stereotype threat within the Technology Acceptance Model (TAM). Across two studies, experiencing stereotype threat in the technological domain was indirectly associated with lower levels of technology use among older adults. This was found for subjective (Study 1) and objective measures (Study 2) of use behaviour, and for technology use in general (Study 1) and computer use in particular (Study 2). In line with the predictions of the Technology Acceptance Model, this relationship was mediated by anxiety, perceived ease of use, perceived usefulness, and behavioural intention. Specifically, stereotype threat was negatively associated with perceived ease of use (Studies 1 and 2) and anxiety mediated this relationship (Study 2). These findings suggest that older adults underuse technology due to the threat of confirming ageist stereotypes targeting their age group. Stereotype threat may thus be an important barrier to technology acceptance and usage in late adulthood. João Mariano, Sibila Marques, Miguel R. Ramos, Filomena Gerardo, Cátia Lage da Cunha, Andrey Girenko, Jan Alexandersson, Bernard Stree, Michele Lamanna, Maurizio Lorenzatto, Louise Pierrel Mikkelsen, Uffe Bundgaard-Joergensen, Sílvia Rêgo, Hein de Vries |
Behav. Inf. Technol. | 7 |
| 2021 | Spinning Icons: Introducing a Novel SSVEP-BCI Paradigm Based on RotationabstractSteady-State-Visually-Evoked-Potential (SSVEP) Brain-Computer Interfaces (BCIs) make use of flickering stimuli to determine the target a user is looking at and select commands accordingly. Those types of BCI can be operated with little to no training, achieve high classification accuracies and are robust in application. A drawback of this approach is the reduced user comfort due to the constant flickering of the stimuli which can be annoying and tiring to look at. Existing studies addressing this issue try to make use of motion to disguise the oscillating patterns. However, this makes them look abstract and restricts the design of those applications as those patterns do not blend in to conventional user interfaces. In this work we introduce the concept of spinning icons to evoke SSVEPs. The icons are rotating in a certain frequency around their vertical axis and are supposed to appear more natural and be less stressing for the human eye. Furthermore this concept is not bound to any kind of abstract motion based pattern but rather supposed to work with any type of icon or image. The newly designed stimuli were evaluated in an application-oriented scenario and compared to standard and state-of-the-art movement-based SSVEP stimuli regarding the classification accuracy and experienced visual fatigue. The results show that the newly created stimuli performed equally well and partially even better in terms of classification accuracy and were rated throughout better concerning visual fatigue by the study participants. This work therefore lays the foundation for more comfortable SSVEP-BCIs which can be used with basically every icon or UI element spinning around their vertical axis. Maurice Rekrut, Tobias Jungbluth, Jan Alexandersson, Antonio Krüger |
IUI | 3 |
| 2018 | The Metalogue Debate Trainee Corpus: Data Collection and Annotations
Volha Petukhova, Andrei Malchanau, Youssef Oualil, Dietrich Klakow, Saturnino Luz, Fasih Haider, Nick Campbell 0001, Dimitris Koryzis, Dimitris Spiliotopoulos, Pierre Albert, Nicklas Linz, Jan Alexandersson |
LREC | 12 |
| 2017 | DiDiER - digitized services in dietary counselling for people with increased health risks related to malnutrition and food allergiesabstractThe goal of the DiDiER project is to verifiably improve services in the field of dietary counselling. This will be achieved by digitising information to increase counselling intensity and to improve workflows for the service provider. The project will develop an IT-based support system for dietary counselling, covering two use cases, facilitation and support of the work of nutritionists in ambulatory allergological nutrition counselling and of nutritionists involved in the care of geriatric patients, especially of those with frailty. One of the project's significant features is that the user's sensitive data remain under his or her personal control at all times. Patrick Elfert, Marco Eichelberg, Johannes Tröger, Jochen Britz, Jan Alexandersson, Daniel Bieber, Jürgen M. Bauer, Susanne Teichmann, Ludwig Kuhn, Martin Thielen, Janina Sauer, Alexander Münzberg, Norbert Rösch, Julia Woizischke, Rebecca Diekmann, Andreas Hein 0001 |
ISCC | 5 |
| 2014 | Metalogue: A Multiperspective Multimodal Dialogue System with Metacognitive Abilities for Highly Adaptive and Flexible Dialogue ManagementabstractThis poster paper presents a high-level description of the Metalogue project that is developing a multi-modal dialogue system that is able to implement interactive behaviors that seem natural to users and is flexible enough to exploit the full potential of multimodal interaction. We provide an outline of the initial work undertaken to define a an open architecture for the integrated Metalogue system. This system includes components that are necessary for the implementation of the processing stages for a variety of application domains: initialization, training, information gathering, orchestration, multimodality, dialogue management, speech recognition, speech synthesis and user modelling. Jan Alexandersson, Maria Aretoulaki, Nick Campbell 0001, Michael Gardner, Andrey Girenko, Dietrich Klakow, Dimitris Koryzis, Volha Petukhova, Marcus Specht, Dimitris Spiliotopoulos, Alexander Stricker, Niels Taatgen |
Intelligent Environments | 1 |
| 2014 | Bridging the Gap between Smart Home and AgentsabstractNoticeable, within the area of Smart Homes, Intelligent Environment and/or Ambient Assisted Living the research deviates into two main branches. On the one hand, there are the agent-based distributed approaches providing intelligent services with focus on data analysis. On the other hand there are the centralized middleware platforms with the main focus on interfacing appliances and services and providing accessible user personalized and interfaces. To bridge this gap, we propose to integrate FIPA Specifications compliant MAS approaches with a standardized middleware platform, the ISO/IEC 24752 Universal Remote Console (URC). By doing so, we resolve two problems that have been neglected so far. Firstly, we provide for MASs a standardized way of interfacing arbitrary appliances and services. Secondly, we allow to deploy personalized and accessible user interfaces for MAS. Additionally, from the point of view of the URC technology, we provide a uniform way to interact with agents. We illustrate the capabilities of the proposed framework with a centralized, decentralized, and hybrid setup scenario. Jochen Britz, Jochen Frey, Jan Alexandersson |
Intelligent Environments | 3 |
| 2012 | ISO 24617-2: A semantically-based standard for dialogue annotation
Harry Bunt, Jan Alexandersson, Jae-Woong Choe, Alex Chengyu Fang, Kôiti Hasida, Volha Petukhova, Andrei Popescu-Belis, David R. Traum |
LREC | 2 |
| 2011 | Designing with and for the Visually Impaired: Vocabulary, Spelling and the Screen Reader
Verena Stein, Robert Neßelrath, Jan Alexandersson, Johannes Tröger |
CSEDU (2) | 3 |
| 2011 | SmartCase: A Smart Home Environment in a SuitcaseabstractTypically, intelligent environments include a heterogeneous set of devices and services offering complex functionalities. To establish a real assistive intelligent environment the available functionalities, their interdependencies and possible side effects must be comprehensible for the end-users and for the diverse set of developers, architects and craftsmen involved in the development process. In this paper, we present a detailed small-scale model of an instrumented two-rooms apartment and show how the model can be utilized for designing, evaluating and controlling intelligent environments. Furthermore, we present an architectural foundation based on Universal Remote Console technology and discuss how to synchronize the model with the real environments. Jochen Frey, Simon Bergweiler, Jan Alexandersson, Ehsan Gholamsaghaee, Norbert Reithinger, Christoph Stahl |
Intelligent Environments | 3 |
| 2010 | Improving Spelling Skills for Blind Language Learners - Orthographic Feedback in an Auditory Vocabulary Trainer
Verena Stein, Robert Neßelrath, Jan Alexandersson |
CSEDU (2) | 3 |
| 2010 | Towards an ISO Standard for Dialogue Act Annotation
Harry Bunt, Jan Alexandersson, Jean Carletta, Jae-Woong Choe, Alex Chengyu Fang, Kôiti Hasida, Kiyong Lee, Volha Petukhova, Andrei Popescu-Belis, Laurent Romary, Claudia Soria, David R. Traum |
LREC | 2 |
| 2003 | SmartKom: adaptive and flexible multimodal access to multiple applicationsabstractThe development of an intelligent user interface that supports multimodal access to multiple applications is a challenging task. In this paper we present a generic multimodal interface system where the user interacts with an anthropomorphic personalized interface agent using speech and natural gestures. The knowledge-based and uniform approach of SmartKom enables us to realize a comprehensive system that understands imprecise, ambiguous, or incomplete multimodal input and generates coordinated, cohesive, and coherent multimodal presentations for three scenarios, currently addressing more than 50 different functionalities of 14 applications. We demonstrate the main ideas in a walk through the main processing steps from modality fusion to modality fission. Norbert Reithinger, Jan Alexandersson, Tilman Becker, Anselm Blocher, Ralf Engel, Markus Löckelt, Jochen Müller 0001, Norbert Pfleger, Peter Poller, Michael Streit, Valentin Tschernomas |
ICMI | 2 |
| 2000 | Summarizing Multilingual Spoken Negotiation DialoguesabstractWe present the multilingual summarization functionality for VERB-MOBIL, a speech translation system. We reuse resources of the system to create a summary. After content extraction, we interpret the results in the dialog context. A summary generator provides the input to generation. A first evaluation indicates the feasibility of the approach. Norbert Reithinger, Michael Kipp, Ralf Engel, Jan Alexandersson |
ACL | 4 |
| 2000 | Multilingual Summary Generation in a Speech-To-Speech Translation System for Multilingual DialoguesabstractThis paper describes a novel functionality of the VERBMOBIL system, a large scale translation system designed for spontaneously spoken multilingual negotiation dialogues. The task is the on-demand generation of dialogue scripts and result summaries of dialogues. We focus on summary generation and show how the relevant data are selected from the dialogue memory and how they are packed into an appropriate abstract representation. Finally, we demonstrate how the existing generation module of VERBMOBIL was extended to produce multilingual and result summaries from these representations. Jan Alexandersson, Peter Poller, Michael Kipp, Ralf Engel |
INLG | 1 |
| 1998 | Towards Multilingual Protocol Generation For Spontaneous Speech Dialogues
Jan Alexandersson, Peter Poller |
INLG | 1 |
| 1997 | Learning dialogue structures from a corpusabstractThis paper demonstrates some aspects of a plan processor which is a subcomponent of the dialogue module of verbmobil. We describe how we transfer results from the research area of grammar extraction for the semi-automatic acquisition of plan operators for turn classes. We exploit statistical knowledge acquired during learning the grammar and incorporate top down predictions to enhance the correct analysis of turn classes described. A first evaluation shows a relative recognition rate of around 70% on unseen data. Jan Alexandersson, Norbert Reithinger |
EUROSPEECH | 1 |
| 1995 | A Robust and Efficient Three-Layered Dialogue Component for a Speech-to-Speech Translation System
Jan Alexandersson, Elisabeth Maier, Norbert Reithinger |
EACL | 1 |