Ronald Poppe

dblp:63/4070 · also Ronald W. Poppe · DBLP profile ↗
← Back
53ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-0843-7878ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 9 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 11 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Explicitly modeling trajectories and correlations for video analysis
abstract
Video analysis tasks such as action recognition and sign language recognition require that both the appearance and dynamics of objects and people of interest are captured. Extraction of distinctive local appearance changes of regions under motion, such as finger movement while gesturing, is often essential to make correct classifications. Without disentangling gross and fine spatio-temporal patterns, both will be confounded, leading to reduced accuracy. We introduce two innovations to tackle this issue. First, we explicitly model motion of image patches by temporally aligning visual tokens using optical flow. In addition to obtaining a coarse motion representation, for a query token, we apply self-attention along the trajectory. This essentially cancels out gross movement and makes possible the extraction of distinctive local patterns of regions under motion. Our second innovation is a dynamic attention mechanism that filters out irrelevant frame regions. It assigns dynamic key–value tokens from correlated regions to each query to focus on coordinated appearance changes such as joint hand and mouth movements while gesturing. We combine these innovations in a novel Trajectory and Correlation (TC) block, a hybrid network that effectively models spatio-temporal information from trajectories and correlated regions. We experiment on four sign language (PHOENIX14, PHOENIX14-T, CSL, and CSL-Daily) and two action recognition (Kinetics-400 and Something-Something V2) datasets. Using TC blocks in different backbones consistently achieves improved performance with a modest increase in parameters, and 30% additional computational cost.
Albert Ali Salah, Ronald Poppe
Image Vis. Comput.3
2026 Improving the generalization of ViTs for action understanding with VLM pre-training
abstract
Owing to their ability to extract powerful video embeddings, Vision Transformers (ViTs) are currently the best performing models in video action understanding. However, when these models are frozen and applied to downstream tasks, their performance drops significantly, revealing limited generalization. In this paper, we describe the Four-Tiered Prompts (FTP) framework that introduces feature processors to transform the ViT’s output. In a pre-training stage, each feature processor is trained using contrastive learning to align the ViT’s visual embeddings with a vision language model’s (VLM) textual embeddings. We use four feature processors, each linked to the output of a VLM prompt that reflects the fundamental aspects of human action: category, components, description, and context. With the FTP framework, we increase the ViT’s generalization ability by forcing the visual encoder to incorporate relevant, semantic information. Importantly, we only employ the VLM during training. Subsequently, inference incurs a limited computation cost. For video action recognition and detection, employing the FTP framework consistently yields state-of-the-art performance after fine-tuning. Extensive experiments demonstrate how different choices contribute to the overall increase in performance. 1
Albert Ali Salah, Ronald Poppe
Pattern Recognit.3
2025 Snakes and Ladders: Two Steps Up for VideoMamba
Albert Ali Salah, Ronald Poppe
ICCV3
2025 About Time: Advances, Challenges, and Outlooks of Action Understanding
abstract
Abstract We have witnessed impressive advances in video action understanding. Increased dataset sizes, variability, and computation availability have enabled leaps in performance and task diversification. Current systems can provide coarse- and fine-grained descriptions of video scenes, extract segments corresponding to queries, synthesize unobserved parts of videos, and predict context across multiple modalities. This survey comprehensively reviews advances in uni- and multi-modal action understanding across a range of tasks. We focus on prevalent challenges, overview widely adopted datasets, and survey seminal works with an emphasis on recent advances. We broadly distinguish between three temporal scopes: (1) recognition tasks of actions observed in full, (2) prediction tasks for ongoing partially observed actions, and (3) forecasting tasks for subsequent unobserved action(s). This division allows us to identify specific action modeling and video representation challenges. Finally, we outline future directions to address current shortcomings.
Alexandros Stergiou, Ronald Poppe
Int. J. Comput. Vis.2
2024 TCNet: Continuous Sign Language Recognition from Trajectories and Correlated Regions
abstract
A key challenge in continuous sign language recognition (CSLR) is to efficiently capture long-range spatial interactions over time from the video input. To address this challenge, we propose TCNet, a hybrid network that effectively models spatio-temporal information from Trajectories and Correlated regions. TCNet's trajectory module transforms frames into aligned trajectories composed of continuous visual tokens. This facilitates extracting region trajectory patterns. In addition, for a query token, self-attention is learned along the trajectory. As such, our network can also focus on fine-grained spatio-temporal patterns, such as finger movement, of a region in motion. TCNet's correlation module utilizes a novel dynamic attention mechanism that filters out irrelevant frame regions. Additionally, it assigns dynamic key-value tokens from correlated regions to each query. Both innovations significantly reduce the computation cost and memory. We perform experiments on four large-scale datasets: PHOENIX14, PHOENIX14-T, CSL, and CSL-Daily. Our results demonstrate that TCNet consistently achieves state-of-the-art performance. For example, we improve over the previous state-of-the-art by 1.5\% and 1.0\% word error rate on PHOENIX14 and PHOENIX14-T, respectively. Code is available at https://github.com/hotfinda/TCNet
Albert Ali Salah, Ronald Poppe
AAAI3
2024 Compensation Sampling for Improved Convergence in Diffusion Models
Albert Ali Salah, Ronald Poppe
ECCV (61)3
2024 Survey of Automated Methods for Nonverbal Behavior Analysis in Parent-Child Interactions
abstract
Social interactions are fundamental for human beings, motivating the abundance of studies into the behavioral correlates of constructs such as personality and relationship. The primary drivers of this research are video-taped recordings of interactions. Recent advancements in automatic behavior analysis provide a cost-effective and more objective alternative to manual coding by trained experts. Still, the use of automated analysis is far from trivial. In this literature survey, we discuss the current state-of-the-art in automated parent-child interaction analysis, and critically assess opportunities and limitations. We focus on parent-child interactions as they reflect various aspects of a child's development, and provide distinct challenges for the automated measurement and interpretation of the interactive behavior. We briefly discuss single-person and dyadic nonverbal measurements, and identify measurement challenges. We then provide an overview of various developmental constructs that can be measured through the classification of extracted cues. Finally, we outline persistent limitations of the current state-of-the-art, and we highlight promising directions to bridge the gap between manual and automated measurements.
Berfu Karaca, Albert Ali Salah, Jaap Denissen, Ronald Poppe, Sonja M. C. de Zwarte
FG4
2024 Decoding Contact: Automatic Estimation of Contact Signatures in Parent-Infant Free Play Interactions
abstract
In parent-child interactions (PCIs), there is frequent physical contact between the two actors. Quantifying this contact provides valuable input to assess the nature of the interaction or the relation between parent and child. Here, we explore the application of vision-based techniques to automatically detect contact signatures at each frame of video recordings of playful parent-infant interactions. We employ two separate models: (i) a multimodal convolutional neural network (CNN) that integrates 2D pose and body part information, and (ii) a unimodal graph convolutional neural network (GCN) that utilizes only 2D pose. We showcase the potential and limitations of automatic contact signature estimation through quantitative and qualitative assessments using a parent-infant free play interaction dataset consisting of 100 parent-child dyadic interactions, covering 20 hours. Additionally, our experiments provide insights into various design choices through systematic experimentation. By releasing our annotations and code, we aim to enable further research in the automatic contact signature estimation during free play interactions between parents and infants.
Metehan Doyran, Albert Ali Salah, Ronald Poppe
ICMI3
2023 SCFormer: Integrating hybrid Features in Vision Transformers
abstract
Hybrid modules that combine self-attention and convolution operations can benefit from the advantages of both, and consequently achieve higher performance than either operation alone. However, current hybrid modules do not capitalize directly on the intrinsic relation between self-attention and convolution, but rather introduce external mechanisms that come with increased computation cost. In this paper, we propose a new hybrid vision transformer called Shift and Concatenate Transformer(SCFormer), which benefits from the intrinsic relationship between convolution and self-attention. SCFormer roots in the Shift and Concatenate Attention (SCA) block, that integrates convolution and self-attention features. We propose a shifting mechanism and corresponding aggregation rules for the feature integration of SCA blocks such that generated features more closely approximate the optimal output features. Extensive experiments show that, with comparable computational complexity, SCFormer consistently achieves improved results over competitive baselines on image recognition and downstream tasks. Our code is available at: https://github.com/hotfinda/SCFormer.
Ronald Poppe, Albert Ali Salah
ICME2
2023 LA-layer: General local attention layer for full attention networks
abstract
Attention layers have contributed to state-of-the-art results on vision tasks. Still, they leave room for improvement because position information is used in a fixed manner, and the computation cost is typically high. To mitigate both issues, we propose a convolution-style local attention layer (LA-layer) as a replacement for traditional attention layers. LA-layers not only encode the position information of pixels in a convolutional manner, but also produce position offsets following a novel constrained rule so that keys will deform and result in larger receptive fields. Query and keys are processed by a novel aggregation function that outputs attention weights for the values. In our experiments with different types of ResNets, we replace convolutional layers with LA-layers and address image recognition, object detection and instance segmentation tasks. We consistently demonstrate performance gains, despite having fewer FLOPs and training parameters. Our code is available at: https://github.com/hotfinda/LA-layer.
Ronald Poppe, Albert Ali Salah
ICME2
2023 Multi-Dataset, Multitask Learning of Egocentric Vision Tasks
abstract
For egocentric vision tasks such as action recognition, there is a relative scarcity of labeled data. This increases the risk of overfitting during training. In this paper, we address this issue by introducing a multitask learning scheme that employs related tasks as well as related datasets in the training process. Related tasks are indicative of the performed action, such as the presence of objects and the position of the hands. By including related tasks as additional outputs to be optimized, action recognition performance typically increases because the network focuses on relevant aspects in the video. Still, the training data is limited to a single dataset because the set of action labels usually differs across datasets. To mitigate this issue, we extend the multitask paradigm to include datasets with different label sets. During training, we effectively mix batches with samples from multiple datasets. Our experiments on egocentric action recognition in the EPIC-Kitchens, EGTEA Gaze+, ADL and Charades-EGO datasets demonstrate the improvements of our approach over single-dataset baselines. On EGTEA we surpass the current state-of-the-art by 2.47 percent. We further illustrate the cross-dataset task correlations that emerge automatically with our novel training scheme.
Georgios Kapidis, Ronald Poppe, Remco C. Veltkamp
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 AdaPool: Exponential Adaptive Pooling for Information-Retaining Downsampling
abstract
Pooling layers are essential building blocks of convolutional neural networks (CNNs), to reduce computational overhead and increase the receptive fields of proceeding convolutional operations. Their goal is to produce downsampled volumes that closely resemble the input volume while, ideally, also being computationally and memory efficient. Meeting both these requirements remains a challenge. To this end, we propose an adaptive and exponentially weighted pooling method: adaPool. Our method learns a regional-specific fusion of two sets of pooling kernels that are based on the exponent of the Dice-Sørensen coefficient and the exponential maximum, respectively. AdaPool improves the preservation of detail on a range of tasks including image and video classification and object detection. A key property of adaPool is its bidirectional nature. In contrast to common pooling methods, the learned weights can also be used to upsample activation maps. We term this method adaUnPool. We evaluate adaUnPool on image and video super-resolution and frame interpolation. For benchmarking, we introduce Inter4K, a novel high-quality, high frame-rate video dataset. Our experiments demonstrate that adaPool systematically achieves better results across tasks and backbones, while introducing a minor additional computational and memory overhead.
Alexandros Stergiou, Ronald Poppe
IEEE Trans. Image Process.2
2021 Incremental Few-Shot Instance Segmentation
abstract
Few-shot instance segmentation methods are promising when labeled training data for novel classes is scarce. However, current approaches do not facilitate flexible addition of novel classes. They also require that examples of each class are provided at train and test time, which is memory intensive. In this paper, we address these limitations by presenting the first incremental approach to few-shot instance segmentation: iMTFA. We learn discriminative embeddings for object instances that are merged into class representatives. Storing embedding vectors rather than images effectively solves the memory overhead problem. We match these class embeddings at the RoI-level using cosine similarity. This allows us to add new classes without the need for further training or access to previous training data. In a series of experiments, we consistently outperform the current state-of-the-art. Moreover, the reduced memory requirements allow us to evaluate, for the first time, few-shot instance segmentation performance on all classes in COCO jointly1.
Dan Andrei Ganea, Bas Boom, Ronald Poppe
CVPR3
2021 Refining activation downsampling with SoftPool
abstract
Convolutional Neural Networks (CNNs) use pooling to decrease the size of activation maps. This process is crucial to increase the receptive fields and to reduce computational requirements of subsequent convolutions. An important feature of the pooling operation is the minimization of information loss, with respect to the initial activation maps, without a significant impact on the computation and memory overhead. To meet these requirements, we propose SoftPool: a fast and efficient method for exponentially weighted activation downsampling. Through experiments across a range of architectures and pooling methods, we demonstrate that SoftPool can retain more information in the reduced activation maps. This refined downsampling leads to improvements in a CNN’s classification accuracy. Experiments with pooling layer substitutions on ImageNet1K show an increase in accuracy over both original architectures and other pooling methods. We also test SoftPool on video datasets for action recognition. Again, through the direct replacement of pooling layers, we observe consistent performance improvements while computational loads and memory requirements remain limited1.
Alexandros Stergiou, Ronald Poppe, Grigorios Kalliatakis
ICCV2
2021 Multi-Temporal Convolutions for Human Action Recognition in Videos
abstract
Effective extraction of temporal patterns is crucial for the recognition of temporally varying actions in video. We argue that the fixed-sized spatio-temporal convolution kernels used in convolutional neural networks (CNNs) can be improved to extract informative motions that are executed at different time scales. To address this challenge, we present a novel convolution block that is capable of extracting spatio-temporal patterns at multiple temporal resolutions. Our proposed multi-temporal convolution (MTConv) blocks utilize two branches that focus on brief and prolonged spatio-temporal patterns, respectively. The extracted time-varying features are aligned in a third branch, with respect to global motion patterns through recurrent cells. The proposed blocks are lightweight and can be integrated into any 3D-CNN architecture. This introduces a substantial reduction in computational costs. Extensive experiments on Kinetics, Moments in Time and HACS action recognition benchmark datasets demonstrate competitive performance of MTConvs compared to the state-of-the-art with a significantly lower computational footprint11Our code is available at: https://git.io/JfuPi.
Alexandros Stergiou, Ronald Poppe
IJCNN2
2021 Learn to cycle: Time-consistent feature discovery for action recognition
abstract
Generalizing over temporal variations is a prerequisite for effective action recognition in videos. Despite significant advances in deep neural networks, it remains a challenge to focus on short-term discriminative motions in relation to the overall performance of an action. We address this challenge by allowing some flexibility in discovering relevant spatio-temporal features. We introduce Squeeze and Recursion Temporal Gates (SRTG), an approach that favors inputs with similar activations with potential temporal variations. We implement this idea with a novel CNN block that uses an LSTM to encapsulate feature dynamics, in conjunction with a temporal gate that is responsible for evaluating the consistency of the discovered dynamics and the modeled features. We show consistent improvement when using SRTG blocks, with only a minimal increase in the number of GFLOPs. On Kinetics-700, we perform on par with current state-of-the-art models, and outperform these on HACS, Moments in Time, UCF-101 and HMDB-51.1
Alexandros Stergiou, Ronald Poppe
Pattern Recognit. Lett.2
2020 Light Field Saliency Detection With Deep Convolutional Networks
abstract
Light field imaging presents an attractive alternative to RGB imaging because of the recording of the direction of the incoming light. The detection of salient regions in a light field image benefits from the additional modeling of angular patterns. For RGB imaging, methods using CNNs have achieved excellent results on a range of tasks, including saliency detection. However, it is not trivial to use CNN-based methods for saliency detection on light field images because these methods are not specifically designed for processing light field inputs. In addition, current light field datasets are not sufficiently large to train CNNs. To overcome these issues, we present a new Lytro Illum dataset, which contains 640 light fields and their corresponding ground-truth saliency maps. Compared to current publicly available light field saliency datasets [1], [2], our new dataset is larger, of higher quality, contains more variation and more types of light field inputs. This makes our dataset suitable for training deeper networks and benchmarking. Furthermore, we propose a novel end-to-end CNN-based framework for light field saliency detection. Specifically, we propose three novel MAC (Model Angular Changes) blocks to process light field micro-lens images. We systematically study the impact of different architecture variants and compare light field saliency with regular 2D saliency. Our extensive comparisons indicate that our novel network significantly outperforms state-of-the-art methods on the proposed dataset and has desired generalization abilities on other existing datasets.
Jun Zhang 0017, Yamei Liu, Shengping Zhang, Ronald Poppe, Meng Wang 0001
IEEE Trans. Image Process.4
2019 Saliency Tubes: Visual Explanations for Spatio-Temporal Convolutions
abstract
Deep learning approaches have been established as the main methodology for video classification and recognition. Recently, 3-dimensional convolutions have been used to achieve state-of-the-art performance in many challenging video datasets. Because of the high level of complexity of these methods, as the convolution operations are also extended to an additional dimension in order to extract features from it as well, providing a visualization for the signals that the network interpret as informative, is a challenging task. An effective notion of understanding the network's innerworkings would be to isolate the spatio-temporal regions on the video that the network finds most informative. We propose a method called Saliency Tubes which demonstrate the foremost points and regions in both frame level and over time that are found to be the main focus points of the network. We demonstrate our findings on widely used datasets for thirdperson and egocentric action classification and enhance the set of methods and visualizations that improve 3D Convolutional Neural Networks (CNNs) intelligibility. Our code1and a demo video2are also available.
Alexandros Stergiou, Georgios Kapidis, Grigorios Kalliatakis, Christos Chrysoulas, Remco C. Veltkamp, Ronald Poppe
ICIP6
2019 Spatio-Temporal FAST 3D Convolutions for Human Action Recognition
abstract
Effective processing of video input is essential for the recognition of temporally varying events such as human actions. Motivated by the often distinctive temporal characteristics of actions in either horizontal or vertical direction, we introduce a novel convolution block for CNN architectures with video input. Our proposed Fractioned Adjacent Spatial and Temporal (FAST) 3D convolutions are a natural decomposition of a regular 3D convolution. Each convolution block consist of three sequential convolution operations: a 2D spatial convolution followed by spatio-temporal convolutions in the horizontal and vertical direction, respectively. Additionally, we introduce a FAST variant that treats horizontal and vertical motion in parallel. Experiments on benchmark action recognition datasets UCF-101 and HMDB-51 with ResNet architectures demonstrate consistent increased performance of FAST 3D convolution blocks over traditional 3D convolutions. The lower validation loss indicates better generalization, especially for deeper networks. We also evaluate the performance of CNN architectures with similar memory requirements, based either on Two-stream networks or with 3D convolution blocks. DenseNet-121 with FAST 3D convolutions was shown to perform best, giving further evidence of the merits of the decoupled spatio-temporal convolutions.
Alexandros Stergiou, Ronald Poppe
ICMLA2
2019 Analyzing human-human interactions: A survey
Alexandros Stergiou, Ronald Poppe
Comput. Vis. Image Underst.2
2019 Automated and unobtrusive measurement of physical activity in an interactive playground
Alejandro Moreno, Ronald Poppe, Jenny L. Gibson, Dirk Heylen
Int. J. Hum. Comput. Stud.2
2019 Interactive rodent behavior annotation in video using active learning
abstract
Manual annotation of rodent behaviors in video is time-consuming. By learning a classifier, we can automate the labeling process. Still, this strategy requires a sufficient number of labeled examples. Moreover, we need to train new classifiers when there is a change in the set of behaviors that we consider or in the manifestation of these behaviors in video. Consequently, there is a need for an efficient way to annotate rodent behaviors. In this paper we introduce a framework for interactive behavior annotation in video based on active learning. By putting a human in the loop, we alternate between learning and labeling. We apply the framework to three rodent behavior datasets and show that we can train accurate behavior classifiers with a strongly reduced number of labeled samples. We confirm the efficacy of the tool in a user study demonstrating that interactive annotation facilitates efficient, high-quality behavior measurements in practice.
Malte Lorbach, Ronald Poppe, Remco C. Veltkamp
Multim. Tools Appl.2
2019 A survey of variational and CNN-based optical flow techniques
Zhigang Tu 0001, Wei Xie 0008, Dejun Zhang, Ronald Poppe, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001
Signal Process. Image Commun.4
2018 International Workshop on Multimodal Analyses Enabling Artificial Agents in Human-Machine Interaction (Workshop Summary)
abstract
In this paper a brief overview of the third workshop on Multimodal Analyses enabling Artificial Agents in Human-Machine Interaction. The paper is focussing on the main aspects intended to be discussed in the workshop reflecting the main scope of the papers presented during the meeting. The MA3HMI 2018 workshop is held in conjunction with the 18th ACM International Conference on Mulitmodal Interaction (ICMI 2018) taking place in Boulder, USA, in October 2018. This year, we have solicited papers concerning the different phases of the development of multimodal systems. Tools and systems that address real-time conversations with artificial agents and technical systems are also within the scope.
Ronald Böck, Francesca Bonin, Nick Campbell 0001, Ronald Poppe
ICMI4
2018 Multi-stream CNN: Learning representations based on human-related regions for action recognition
Zhigang Tu 0001, Wei Xie 0008, Qianqing Qin, Ronald Poppe, Remco C. Veltkamp, Baoxin Li, Junsong Yuan 0001
Pattern Recognit.4
2017 A Thing of Beauty: Steering Behavior in an Interactive Playground
abstract
Interactive playgrounds are spaces where players engage in collocated, playful activities, in which added digital technology can be designed to promote cognitive, social, and motor skills development. To promote such development, different strategies can be used to implement game mechanics that change player's in-game behavior. One of such strategies is enticing players to take action through incentives akin to game achievements. We explored if this strategy could be used to influence players' proxemic behavior in the Interactive Tag Playground, an installation that enhances the traditional game of tag. We placed the ITP in an art gallery, observed hundreds of play sessions, and refined the mechanics, which consisted in projecting collectible particles around the tagger that upon collection by runners resulted only in the embellishment of their circles. We implemented the refined mechanics in a study with 48 children. The playground automatically collected the players' positions, and analyses show that runners got closer to and moved more towards taggers when using our enticing strategy. This suggests an enticing strategy can be used to influence physical in-game behavior.
Robby van Delden, Alejandro Moreno, Ronald Poppe, Dennis Reidsma, Dirk Heylen
CHI3
2017 Lend Me a Hand: Auxiliary Image Data Helps Interaction Detection
abstract
In social settings, people interact in close proximity. When analyzing such encounters from video, we are typically interested in distinguishing between a large number of different interactions. Here, we address training deformable part models (DPMs) for the detection of such interactions from video, in both space and time. When we consider a large number of interaction classes, we face two challenges. First, we need to distinguish between interactions that are visually more similar. Second, it becomes more difficult to obtain sufficient specific training examples for each interaction class. In this paper, we address both challenges and focus on the latter. Specifically, we introduce a method to train body part detectors from nonspecific images with pose information. Such resources are widely available. We introduce a training scheme and an adapted DPM formulation to allow for the inclusion of this auxiliary data. We perform cross-dataset experiments to evaluate the generalization performance of our method. We demonstrate that our method can still achieve decent performance, from as few as five training examples.
Coert Van Gemeren, Ronald Poppe, Remco C. Veltkamp
FG2
2017 Variational method for joint optical flow estimation and edge-aware image restoration
Zhigang Tu 0001, Wei Xie 0008, Coert Van Gemeren, Ronald Poppe, Remco C. Veltkamp
Pattern Recognit.5
2016 International workshop on multimodal analyses enabling artificial agents in human- machine interaction (workshop summary)
abstract
In this paper a brief overview of the third workshop on Multimodal Analyses enabling Artificial Agents in Human-Machine Interaction. The paper is focussing on the main aspects intended to be discussed in the workshop reflecting the main scope of the papers presented during the meeting. The MA3HMI 2016 workshop is held in conjunction with the 18th ACM International Conference on Mulitmodal Interaction (ICMI 2016) taking place in Tokyo, Japan, in November 2016. This year, we have solicited papers concerning the different phases of the development of multimodal systems. Tools and systems that address real-time conversations with artificial agents and technical systems are also within the scope.
Ronald Böck, Francesca Bonin, Nick Campbell 0001, Ronald Poppe
ICMI4
2016 Locating human interactions with discriminatively trained deformable pose+motion parts
abstract
We model dyadic (two-person) interactions by discriminatively training a spatio-temporal deformable part model of fine-grained human interactions. All interactions involve at most two persons. Our models are capable of localizing human interactions in unsegmented videos, marking the interactions of interest in space and time. Our contributions are as follows: First, we create a model that localizes human interactions in space and time. Second, our models use multiple pose and motion features per part. Third, we experiment with different ways of training our models discriminatively. When testing on the target class our models achieve a mean average precision score of 0.86. Cross dataset tests show that our models generalize well to different environments.
Coert Van Gemeren, Ronald Poppe, Remco C. Veltkamp
ICPR2
2016 Weighted local intensity fusion method for variational optical flow estimation
Zhigang Tu 0001, Ronald Poppe, Remco C. Veltkamp
Pattern Recognit.2
2016 Adaptive guided image filter for warping in variational optical flow computation
Zhigang Tu 0001, Ronald Poppe, Remco C. Veltkamp
Signal Process.2
2014 Towards an Interactive Leisure Activity for People with PIMD
Robby van Delden, Dennis Reidsma, Wietske van Oorsouw, Ronald Poppe, Peter van der Vos, Andries Lohmeijer, Petri Embregts, Vanessa Evers, Dirk Heylen
ICCHP (1)4
2014 Touching the Void - Introducing CoST: Corpus of Social Touch
abstract
Touch behavior is of great importance during social interaction. To transfer the tactile modality from interpersonal interaction to other areas such as Human-Robot Interaction (HRI) and remote communication automatic recognition of social touch is necessary. This paper introduces CoST: Corpus of Social Touch, a collection containing 7805 instances of 14 different social touch gestures. The gestures were performed in three variations: gentle, normal and rough, on a sensor grid wrapped around a mannequin arm. Recognition of the rough variations of these 14 gesture classes using Bayesian classifiers and Support Vector Machines (SVMs) resulted in an overall accuracy of 54% and 53%, respectively. Furthermore, this paper provides more insight into the challenges of automatic recognition of social touch gestures, including which gestures can be recognized more easily and which are more difficult to recognize.
Merel M. Jung, Ronald Poppe, Mannes Poel, Dirk Heylen
ICMI2
2014 Twente Debate Corpus ― A Multimodal Corpus for Head Movement Analysis
Bayu Rahayudi, Ronald Poppe, Dirk Heylen
LREC2
2013 Perceptual evaluation of backchannel strategies for artificial listeners
Ronald Poppe, Khiet P. Truong, Dirk Heylen
Auton. Agents Multi Agent Syst.1
2012 Online Behavior Evaluation with the Switching Wizard of Oz
Ronald Poppe, Mark ter Maat, Dirk Heylen
IVA1
2012 Facing scalability: Naming faces in an online social network
Ronald Poppe
Pattern Recognit.1
2011 Facial and bodily expressions for control and adaptation of games (ECAG'11)
abstract
Novel sensors open up new avenues for facial and bodily interaction, either consciously and unconsciously. In the former, a user controls the interaction, whereas in the latter, the interaction is adapted based on observations of the user. In the second International Workshop on Facial and Bodily Expressions for Control and Adaptation of Games (ECAG'11), challenges in using body and face in interaction are addressed.
Anton Nijholt, Ronald Poppe
FG2
2011 Scalable face labeling in online social networks
abstract
Face labeling is the process of assigning names to faces. In this paper, we start from a weakly-supervised setting where names are linked to photos, not faces. We introduce two face labeling strategies that scale well to large data sets and allow for labeling parts thereof. This is a useful property especially for data sets where photos are frequently added or (re)labeled. We evaluate our and two related face labeling strategies on a novel corpus of 34,763 faces, gathered from an online social network for dance party visitors. We achieve a speed-up of an order of magnitude over the state-of-the-art approach while the labeling quality is almost unaffected. On a subset of the faces, the speed-up is even more apparent, reaching at least two orders of magnitude.
Ronald Poppe
FG1
2011 A Multimodal Analysis of Vocal and Visual Backchannels in Spontaneous Dialogs
abstract
Backchannels (BCs) are short vocal and visual listener responses that signal attention, interest, and understanding to the speaker. Previous studies have investigated BC prediction in telephone-style dialogs from prosodic cues. In contrast, we consider spontaneous face-to-face dialogs. The additional visual modality allows speaker and listener to monitor each other's attention continuously, and we hypothesize that this affects the BC-inviting cues. In this study, we investigate how gaze, in addition to prosody, can cue BCs. Moreover, we focus on the type of BC performed, with the aim to find out whether vocal and visual BCs are invited by similar cues. In contrast to telephone-style dialogs, we do not find rising/falling pitch to be a BC-inviting cue. However, in a face-to-face setting, gaze appears to cue BCs. In addition, we find that mutual gaze occurs significantly more often during visual BCs. Moreover, vocal BCs are more likely to be timed during pauses in the speaker's speech.
Khiet P. Truong, Ronald Poppe, Iwan de Kok, Dirk Heylen
INTERSPEECH2
2011 Backchannels: Quantity, Type and Timing Matters
Ronald Poppe, Khiet P. Truong, Dirk Heylen
IVA1
2010 Towards affective state modeling in narrative and conversational settings
abstract
We carry out two studies on affective state modeling for communication settings that involve unilateral intent on the part of one participant (the evoker) to shift the affective state of another participant (the experiencer). The first investigates viewer response in a narrative setting using a corpus of docu-mentaries annotated with viewer-reported narrative peaks. The second investigates affective triggers in a conversational set-ting using a corpus of recorded interactions, annotated with continuous affective ratings, between a human interlocutor and an emotionally colored agent. In each case, we build a “one-sided ” model using indicators derived from the speech of one participant. Our classification experiments confirm the viabil-ity of our models and provide insight into useful features. Index Terms: affect, speech recognition, audio analysis, natural language communication
Bart Jochems, Martha A. Larson, Roeland Ordelman, Ronald Poppe, Khiet P. Truong
INTERSPEECH4
2010 A rule-based backchannel prediction model using pitch and pause information
abstract
We manually designed rules for a backchannel (BC) prediction model based on pitch and pause information. In short, the model predicts a BC when there is a pause of a certain length that is preceded by a falling or rising pitch. This model was validated against the Dutch IFADV Corpus in a corpus-based evaluation method. The results showed that our model performs slightly better than another well-known rule-based BC prediction model that uses only pitch information. We observed that the length of a pause preceding a BC is one of the important features in this model, next to the duration of the pitch slope at the end of an utterance. Further, we discuss implications of a corpus-based approach to BC prediction evaluation.
Khiet P. Truong, Ronald Poppe, Dirk Heylen
INTERSPEECH2
2010 Backchannel Strategies for Artificial Listeners
Ronald Poppe, Khiet P. Truong, Dennis Reidsma, Dirk Heylen
IVA1
2010 A survey on vision-based human action recognition
Ronald Poppe
Image Vis. Comput.1
2010 Differences in head orientation behavior for speakers and listeners: An experiment in a virtual environment
abstract
An experiment was conducted to investigate whether human observers use knowledge of the differences in focus of attention in multiparty interaction to identify the speaker amongst the meeting participants. A virtual environment was used to have good stimulus control. Head orientations were displayed as the only cue for focus attention. The orientations were derived from a corpus of tracked head movements. We present some properties of the relation between head orientations and speaker--listener status, as found in the corpus. With respect to the experiment, it appears that people use knowledge of the patterns in focus of attention to distinguish the speaker from the listeners. However, the human speaker identification results were rather low. Head orientations (or focus of attention) alone do not provide a sufficient cue for reliable identification of the speaker in a multiparty setting.
Rutger Rienks, Ronald Poppe, Dirk Heylen
ACM Trans. Appl. Percept.2
2008 Facial and bodily expressions for control and adaptation of games (ECAG'08)
abstract
In this paper, the emphasis is on research on facial and bodily expressions for the control and adaptation of games. We distinguish between two forms of expressions, depending on whether the user has the initiative and consciously uses his or her movements and expressions to control the interface, or whether the application takes the initiative to adapt itself to the affective state of the user as it can be interpreted from the user's expressive behavior.
Anton Nijholt, Ronald Poppe
FG2
2008 Meeting behavior detection in smart environments: Nonverbal cues that help to obtain natural interaction
abstract
Unobtrusive and multiple-sensor interfaces enable observation of natural human behavior in smart environments. Being able to detect, analyze and interpret this activity allows for the implementation of various applications, including real-time surveillance and real-time support in smart home and office environments. We focus on smart meeting environments, where nonverbal behavioral cues sometimes tell more about issues such as discussion participation, involvement and contribution, than information obtained from verbal contributions. An important aspect of this behavior is the interaction between meeting participants. We regard the special case where some of the participants are at physically different locations, which hinders natural interaction. We discuss how we can exploit the ability of a sensor-equipped environment to detect nonverbal interaction cues and use these to allow and improve natural interaction between collaborating participants at distributed locations. We focus on research efforts to detect nonverbal interaction cues by looking at various modalities. Also, we discuss how information obtained from the fusion of these modalities allows us to generate and display behavioral cues that allow remote participants to take part in a distributed meeting in a natural way.
Mannes Poel, Ronald Poppe, Anton Nijholt
FG2
2008 Discriminative human action recognition using pairwise CSP classifiers
abstract
We present a discriminative approach to human action recognition. At the heart of our approach is the use of common spatial patterns (CSP), a spatial filter technique that transforms temporal feature data by using differences in variance between two classes. Such a transformation focusses on differences between classes, rather than on modelling each class individually. As a results, to distinguish between two classes, we can use simple distance metrics in the low-dimensional transformed space. The most likely class is found by pairwise evaluation of all discriminant functions. Our image representations are silhouette boundary gradients, spatially binned into cells. We achieve scores of approximately 96% on a standard action dataset, and show that reasonable results can be obtained when training on only a single subject. Future work is aimed at combining our approach with automatic human detection.
Ronald Poppe, Mannes Poel
FG1
2007 Person or Puppet? The Role of Stimulus Realism in Attributing Emotion to Static Body Postures
Marco Pasch, Ronald Poppe
ACII2
2007 Vision-based human motion analysis: An overview
abstract
Markerless vision-based human motion analysis has the potential to provide an inexpensive, non-obtrusive solution for the estimation of body poses. The significant research effort in this domain has been motivated by the fact that many application areas, including surveillance, Human–Computer Interaction and automatic annotation, will benefit from a robust solution. In this paper, we discuss the characteristics of human motion analysis. We divide the analysis into a modeling and an estimation phase. Modeling is the construction of the likelihood function, estimation is concerned with finding the most likely pose given the likelihood surface. We discuss model-free approaches separately. This taxonomy allows us to highlight trends in the domain and to point out limitations of the current state of the art.
Ronald Poppe
Comput. Vis. Image Underst.1
2006 Towards Bi-directional Dancing Interaction
Dennis Reidsma, Herwin van Welbergen, Ronald Poppe, Pieter Bos, Anton Nijholt
ICEC3