VLDB 2026 Research / reviewers in the wild / expert
Rafael Wampfler
dblp:01/2930
· DBLP profile ↗
12ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-0158-1305ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Joint Personality-Emotion Framework for Personality-Consistent Conversational AgentsabstractArousal Valence Figure 1: Conceptual overview of the proposed framework.Left: Personality descriptors projected into the valence-arousal space using the EMoLon lexicon [9].Center: Kernel Density Estimation (KDE) applied to the projected descriptors, illustrating the density distribution of personality-related adjectives.Right: Warped emotion topology derived from the KDE. Nikola Kovacevic, Christian Holz 0001, Markus Gross 0001, Rafael Wampfler |
IVA | 4 |
| 2025 | BEE: Belief-Value-Aligned, Explainable, and Extensible Cognitive Framework for Conversational AgentsabstractRecent advances in large language models have enabled virtual agents to exhibit increasingly believable social behaviors.However, creating social agents that remain consistent with a defined profile and explain their reasoning remains challenging.We introduce a cognitive framework designed to address these gaps.Our framework features a graph-based memory module (concept pool) and a decision-making process inspired by human cognition.The concept pool contains an agent's beliefs, values, and background stories for context-dependent retrieval.The decision-making process uses the concept pool to produce belief-value aligned responses of the virtual agent and intuitive, human-readable explanations of the reasoning.To evaluate the effectiveness of our framework, we created two virtual agents based on historical figures and compared them to baseline agents.Our evaluation combined quantitative assessments of belief-value alignment with a user study (n=48) examining explainable agency.Results show that our framework exhibits model-agnostic improved belief-value alignment and produces more detailed, relevant, and understandable explanations.By grounding virtual agent behavior in structured memories and cognitive principles, our framework offers a compelling step toward more coherent and socially intelligent virtual agents. Markus Gross 0001, Rafael Wampfler |
IVA | 3 |
| 2025 | PhonemeNet: A Transformer Pipeline for Text-Driven Facial AnimationabstractWe present a fully text-driven framework for 3D facial animation that eliminates the need for audio input or explicit prosodic cues. Our architecture extracts rich phoneme embeddings from text using a pre-trained TTS encoder, aligns them with quantized motion embeddings via a transformer decoder, and decodes the result into mesh deformations through a pre-trained transformer decoder. We explore two scenarios of our pipeline: (1) In the single-subject setting, we find that phoneme embeddings alone can yield accurate lip motion. (2) In a multi-subject setting, where speaker articulation varies widely, we introduce stochastic latent modulation to model residual variability conditioned on both phoneme context and speaker identity. We evaluate our approach quantitatively and qualitatively: We demonstrate accurate lip sync in the single-subject case, and compare against audio-driven baselines on a large multi-subject dataset. Our results show that PhonemeNet not only achieves competitive lip sync and motion quality, but also offers flexibility, modularity, and scalability as an alternative to audio-driven facial animation. Philine Witzig, Barbara Solenthaler, Markus Gross 0001, Rafael Wampfler |
MIG | 4 |
| 2025 | egoEMOTION: Egocentric Vision and Physiological Signals for Emotion and Personality Recognition in Real-world TasksabstractUnderstanding affect is central to anticipating human behavior, yet current egocentric vision benchmarks largely ignore the person’s emotional states that shape their decisions and actions. Existing tasks in egocentric perception focus on physical activities, hand-object interactions, and attention modeling—assuming neutral affect and uniform personality. This limits the ability of vision systems to capture key internal drivers of behavior. In this paper, we present egoEMOTION, the first dataset that couples egocentric visual and physiological signals with dense self-reports of emotion and personality across controlled and real-world scenarios. Our dataset includes over 50 hours of recordings from 43 participants, captured using Meta’s Project Aria glasses. Each session provides synchronized eye-tracking video, head-mounted photoplethysmography, inertial motion data, and physiological baselines for reference. Participants completed emotion-elicitation tasks and naturalistic activities while self-reporting their affective state using the Circumplex Model and Mikels’ Wheel as well as their personality via the Big Five model. We define three benchmark tasks: (1) continuous affect classification (valence, arousal, dominance); (2) discrete emotion classification; and (3) trait-level personality inference. We show that a classical learning-based method, as a simple baseline in real-world affect prediction, produces better estimates from signals captured on egocentric vision systems than processing physiological signals. Our dataset establishes emotion and personality as core dimensions in egocentric perception and opens new directions in affect-driven modeling of behavior, intent, and interaction. Matthias Jammot, Björn Braun, Paul Streli, Rafael Wampfler, Christian Holz 0001 |
NeurIPS | 4 |
| 2024 | On Multimodal Emotion Recognition for Human-Chatbot Interaction in the WildabstractThe field of natural language generation is swiftly evolving, giving rise to powerful conversational characters for use in different applications such as entertainment, education, and healthcare. A central aspect of these applications is providing personalized interactions, driven by the ability of the characters to recognize and adapt to user emotions. Current emotion recognition models primarily rely on datasets collected from actors or in controlled laboratory settings focusing on human-human interactions, which hinders their adaptability to real-world applications for conversational agents. In this work, we unveil the complexity of human-chatbot emotion recognition in the wild. We collected a multimodal dataset consisting of text, audio, and video recordings from 99 participants while they conversed with a GPT-3-based chatbot over three weeks. Using different transformer-based multimodal emotion recognition networks, we provide evidence for a strong domain gap between human-human interaction and human-chatbot interaction that is attributed to the subjective nature of self-reported emotion labels, the reduced activation and expressivity of the face, and the inherent subtlety of emotions in such settings, emphasizing the challenges of recognizing user emotions in real-world contexts. We show how personalizing our model to the user increases the model performance by up to 38% (user emotions) and up to 41% (perceived chatbot emotions), highlighting the potential of personalization for overcoming the observed domain gap. Nikola Kovacevic, Christian Holz 0001, Markus Gross 0001, Rafael Wampfler |
ICMI | 4 |
| 2024 | EmoSpaceTime: Decoupling Emotion and Content through Contrastive Learning for Expressive 3D Speech AnimationabstractEquipping stylized conversational characters with facial animations tailored to specific emotions enhances coherence and authenticity. Many data-driven speech animation methods lack dynamic facial expressions since they rely on explicit semantic control signals, leading to static emotional expressions. We present a Transformer-AE for disentangling emotion and content within the facial motion latent space. Our method processes animation control parameters in the frequency domain, enabling a more fine-grained separation of emotion and content based on frequencies. Through contrastive learning, the model is encouraged to learn similar representations for similar emotional states and the same linguistic content. Capturing the full dynamics of an emotional episode spatially and temporally, this approach enables emotion swapping, enhances expressiveness, and gives artists fine control over emotion, e.g., through emotion interpolation. Our analyses show that the Transformer-AE effectively separates emotion from content, enabling more nuanced and realistic facial animation for conversational characters. Philine Witzig, Barbara Solenthaler, Markus Gross 0001, Rafael Wampfler |
MIG | 4 |
| 2023 | Personality Trait Recognition Based on Smartphone Typing Characteristics in the WildabstractAs governed by personality trait theory, humans tackle problems differently depending on their long-term behavioral characteristics. Computational awareness of personality traits fuels affective computing research, which investigates how to reliably recognize and utilize personality traits. Applications are diverse, including therapy monitoring, learning assistance, and recommender systems. Data-driven approaches are a promising path forward towards personality-aware human-computer interactions. Thereby, central challenges are the non-disruptive data acquisition, the time frame over which data must be collected before predictions become accurate, and the feature-centered data reduction to train reliable and lightweight machine learning models. In this work, we address these challenges by presenting a fully-automatic feature extraction and machine learning pipeline that makes accurate personality trait predictions for the widely-used Five Factor Model from passively-collected, short-term smartphone typing data collected from 76 participants (68 university students) in the wild. Our model allows for personality trait assessments after one day of data collection, demonstrating that, despite being a long-term behavioral trend, personality traits can be inferred accurately from shorter time periods. We demonstrate that our system can accurately predict personality traits on two levels (low and high) with up to 74.5% accuracy and 0.72 AUC for a single day, and up to 84.5% accuracy and 0.79 AUC after subsequent refinement over 10 weeks. Nikola Kovacevic, Christian Holz 0001, Tobias Günther, Markus Gross 0001, Rafael Wampfler |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | Affective State Prediction from Smartphone Touch and Sensor Data in the WildabstractKnowledge of users’ affective states can improve their interaction with smartphones by providing more personalized experiences (e.g., search results and news articles). We present an affective state classification model based on data gathered on smartphones in real-world environments. From touch events during keystrokes and the signals from the inertial sensors, we extracted two-dimensional heat maps as input into a convolutional neural network to predict the affective states of smartphone users. For evaluation, we conducted a data collection in the wild with 82 participants over 10 weeks. Our model accurately predicts three levels (low, medium, high) of valence (AUC up to 0.83), arousal (AUC up to 0.85), and dominance (AUC up to 0.84). We also show that using the inertial sensor data alone, our model achieves a similar performance (AUC up to 0.83), making our approach less privacy-invasive. By personalizing our model to the user, we show that performance increases by an additional 0.07 AUC. Rafael Wampfler, Severin Klingler, Barbara Solenthaler, Victor R. Schinazi, Markus Gross 0001, Christian Holz 0001 |
CHI | 1 |
| 2020 | Affective State Prediction Based on Semi-Supervised Learning from Smartphone Touch DataabstractGaining awareness of the user's affective states enables smartphones to support enriched interactions that are sensitive to the user's context. To accomplish this on smartphones, we propose a system that analyzes the user's text typing behavior using a semi-supervised deep learning pipeline for predicting affective states measured by valence, arousal, and dominance. Using a data collection study with 70 participants on text conversations designed to trigger different affective responses, we developed a variational auto-encoder to learn efficient feature embeddings of two-dimensional heat maps generated from touch data while participants engaged in these conversations. Using the learned embedding in a cross-validated analysis, our system predicted three levels (low, medium, high) of valence (AUC up to 0.84), arousal (AUC up to 0.82), and dominance (AUC up to 0.82). These results demonstrate the feasibility of our approach to accurately predict affective states based only on touch data. Rafael Wampfler, Severin Klingler, Barbara Solenthaler, Victor R. Schinazi, Markus Gross 0001 |
CHI | 1 |
| 2020 | Image Reconstruction of Tablet Front Camera Recordings in Educational Settings
Rafael Wampfler, Andreas Emch, Barbara Solenthaler, Markus Gross 0001 |
EDM | 1 |
| 2019 | Affective State Prediction in a Mobile Setting using Wearable Biometric Sensors and Stylus
Rafael Wampfler, Severin Klingler, Barbara Solenthaler, Victor R. Schinazi, Markus Gross 0001 |
EDM | 1 |
| 2017 | Efficient Feature Embeddings for Student Classification with Variational Auto-encoders
Severin Klingler, Rafael Wampfler, Tanja Käser, Barbara Solenthaler, Markus Gross 0001 |
EDM | 2 |