EDBT 2026 Demo / reviewers in the wild / expert
Wim T. J. L. Pouw
dblp:183/8833
· DBLP profile ↗
11ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0003-2729-6502ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DIMS Dashboard for Exploring Dynamic Interactions and Multimodal Signals
Grace Qiyuan Miao, James P. Trujillo, Landry S. Bulls, Mark A. Thornton, Rick Dale, Wim T. J. L. Pouw |
CogSci | 6 |
| 2025 | Multimodal Quantitative Measures for Multiparty Behavior EvaluationabstractDigital humans are emerging as autonomous agents in multiparty interactions, yet existing evaluation metrics largely ignore contextual coordination dynamics. We introduce a unified, intervention-driven framework for objective assessment of multiparty social behaviour in skeletal motion data, spanning three complementary dimensions: (1) synchrony via Cross-Recurrence Quantification Analysis, (2) temporal alignment via Multiscale Empirical Mode Decompositionbased Beat Consistency, and (3) structural similarity via Soft Dynamic Time Warping. We validate metric sensitivity through three theory-driven perturbations -- gesture kinematic dampening, uniform speech-gesture delays, and prosodic pitch-variance reduction-applied to $\approx 145$ 30-second thin slices of group interactions from the DnD dataset. Mixed-effects analyses reveal predictable, joint-independent shifts: dampening increases CRQA determinism and reduces beat consistency, delays weaken cross-participant coupling, and pitch flattening elevates F0 Soft-DTW costs. A complementary perception study ($N=27$) compares judgments of full-video and skeleton-only renderings to quantify representation effects. Our three measures deliver orthogonal insights into spatial structure, timing alignment, and behavioural variability. Thereby forming a robust toolkit for evaluating and refining socially intelligent agents. Code available on \href{https://github.com/tapri-lab/gig-interveners}{GitHub}. Ojas Shirekar, Wim T. J. L. Pouw, Chenxu Hao, Vrushank Phadnis, Thabo Beeler, Chirag Raman |
ICMI | 2 |
| 2024 | Analysing Cross-Speaker Convergence in Face-to-Face Dialogue through the Lens of Automatically Detected Shared Linguistic Constructions
Esam Ghaleb, Marlou Rasenberg, Wim T. J. L. Pouw, Ivan Toni, Judith Holler, Asli Özyürek, Raquel Fernández |
CogSci | 3 |
| 2024 | Children's multimodal coordination during collaborative problem solving
Lisette De Jonge-Hoekstra, Wim T. J. L. Pouw, Steffie Van der Steen, Ralf F. A. Cox, James A. Dixon |
CogSci | 2 |
| 2024 | What do we mean when we say gestures are more expressive than vocalizations? An experimental and simulation study
Sárka Kadavá, Aleksandra Cwiek, Susanne Fuchs, Wim T. J. L. Pouw |
CogSci | 4 |
| 2024 | Towards a movement science of communication
Sárka Kadavá, Lara Pearson, James P. Trujillo, Wim T. J. L. Pouw |
CogSci | 4 |
| 2024 | Automated Recognition of Grooming Behavior in Wild Chimpanzees
Yana van de Sande, Wim T. J. L. Pouw, Lara M. Southern |
CogSci | 2 |
| 2024 | Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic EvaluationabstractIn face-to-face dialogues, the form-meaning relationship of co-speech gestures varies depending on contextual factors such as what the gestures refer to and the individual characteristics of speakers. These factors make co-speech gesture representation learning challenging. How can we learn meaningful gestures representations considering gestures’ variability and relationship with speech? This paper tackles this challenge by employing self-supervised contrastive learning techniques to learn gesture representations from skeletal and speech information. We propose an approach that includes both unimodal and multimodal pre-training to ground gesture representations in co-occurring speech. For training, we utilize a face-to-face dialogue dataset rich with representational iconic gestures. We conduct thorough intrinsic evaluations of the learned representations through comparison with human-annotated pairwise gesture similarity. Moreover, we perform a diagnostic probing analysis to assess the possibility of recovering interpretable gesture features from the learned representations. Our results show a significant positive correlation with human-annotated gesture similarity and reveal that the similarity between the learned representations is consistent with well-motivated patterns related to the dynamics of dialogue interaction. Moreover, our findings demonstrate that several features concerning the form of gestures can be recovered from the latent representations. Overall, this study shows that multimodal contrastive learning is a promising approach for learning gesture representations, which opens the door to using such representations in larger-scale gesture analysis studies. Esam Ghaleb, Bulat Khaertdinov, Wim T. J. L. Pouw, Marlou Rasenberg, Judith Holler, Asli Özyürek, Raquel Fernández |
ICMI | 3 |
| 2024 | Co-Speech Gesture Detection through Multi-Phase Sequence LabelingabstractGestures are integral components of face-to-face communication. They unfold over time, often following predictable movement phases of preparation, stroke, and retraction. Yet, the prevalent approach to automatic gesture detection treats the problem as binary classification, classifying a segment as either containing a gesture or not, thus failing to capture its inherently sequential and contextual nature. To address this, we introduce a novel framework that reframes the task as a multi-phase sequence labeling problem rather than binary classification. Our model processes sequences of skeletal movements over time windows, uses Transformer encoders to learn contextual embeddings, and leverages Conditional Random Fields to perform sequence labeling. We evaluate our proposal on a large dataset of diverse co-speech gestures in task-oriented face-to-face dialogues. The results consistently demonstrate that our method significantly outperforms strong baseline models in detecting gesture strokes. Furthermore, applying Transformer encoders to learn contextual embeddings from movement sequences substantially improves gesture unit detection. These results highlight our framework’s capacity to capture the fine-grained dynamics of co-speech gesture phases, paving the way for more nuanced and accurate gesture detection and analysis. Esam Ghaleb, Ilya Burenko, Marlou Rasenberg, Wim T. J. L. Pouw, Peter Uhrig, Judith Holler, Ivan Toni, Asli Özyürek, Raquel Fernández |
WACV | 4 |
| 2024 | A toolkit for the dynamic study of air sacs in siamang and other elastic circular structuresabstractBiological structures are defined by rigid elements, such as bones, and elastic elements, like muscles and membranes. Computer vision advances have enabled automatic tracking of moving animal skeletal poses. Such developments provide insights into complex time-varying dynamics of biological motion. Conversely, the elastic soft-tissues of organisms, like the nose of elephant seals, or the buccal sac of frogs, are poorly studied and no computer vision methods have been proposed. This leaves major gaps in different areas of biology. In primatology, most critically, the function of air sacs is widely debated; many open questions on the role of air sacs in the evolution of animal communication, including human speech, remain unanswered. To support the dynamic study of soft-tissue structures, we present a toolkit for the automated tracking of semi-circular elastic structures in biological video data. The toolkit contains unsupervised computer vision tools (using Hough transform) and supervised deep learning (by adapting DeepLabCut) methodology to track inflation of laryngeal air sacs or other biological spherical objects (e.g., gular cavities). Confirming the value of elastic kinematic analysis, we show that air sac inflation correlates with acoustic markers that likely inform about body size. Finally, we present a pre-processed audiovisual-kinematic dataset of 7+ hours of closeup audiovisual recordings of siamang (Symphalangus syndactylus) singing. This toolkit (https://github.com/WimPouw/AirSacTracker) aims to revitalize the study of non-skeletal morphological structures across multiple species. Lara S. Burchardt, Yana van de Sande, Mounia Kehy, Marco Gamba, Andrea Ravignani, Wim T. J. L. Pouw |
PLoS Comput. Biol. | 6 |
| 2017 | Is ambiguity detection in haptic imagery possible? Evidence for Enactive imaginings
Wim T. J. L. Pouw, Asimina Aslanidou, Kevin Kamermans, Fred Paas |
CogSci | 1 |