VLDB 2026 Research / reviewers in the wild / expert
Gwenn Englebienne
dblp:79/6571
· DBLP profile ↗
42ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-3130-2082ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 16 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 since 2021Systems, architecture and hardware · 5 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-Grained Cross-Modal Retrieval in Art via Region-Level Grounding of Symbolic NarrativesabstractRetrieving specific symbolic elements within paintings, such as a pomegranate representing fertility, requires fine-grained cross-modal understanding beyond whole-image-level matching. Existing art datasets lack region-level annotations that link localized objects to their iconographic descriptions, limiting semantic search to object labels. We present RichArt, a benchmark dataset of 7,087 paintings with 20,882 region-level annotations pairing visual elements with their symbolic narratives extracted from museum catalogs using a scalable, semi-automated LLM-based pipeline, followed by human validation. To support bidirectional retrieval (text to region and region to text), we introduce MARGE-GD (Multi-modal Alignment of RichArt Grounding Embeddings via Grounding DINO). Building on Grounding DINO’s architecture, MARGE-GD projects region and text representations into a shared embedding space via MLP heads trained with contrastive loss, enabling retrieval while preserving visual grounding capability. On RichArt, MARGE-GD achieves an MRR of 0.75 (text to region) and 0.61 (region to text), outperforming CLIP by 2.59 × and 2.77 × respectively, demonstrating significant improvements in fine-grained semantic retrieval. Mihai-Bogdan Bîndila, Shenghui Wang 0001, Gwenn Englebienne |
ICMR | 3 |
| 2025 | Collecting Object-level Affordance for RGBD DatasetsabstractAccurate interpretation of the environment is both essential to automated robots and highly beneficial for teleoperated robots. Going beyond obstacle recognition, interpreting the semantics of the environment and the actions it affords, enables robots to interact with environments made for humans in a human-like manner. This paper describes the collection of affordance labels at the object level for multiple indoors datasets, to train computer vision algorithms for detecting object affordances in indoor spaces. It is a first step towards determining high level "semantic" affordances, to allow reasoning about what to do with objects, rather than "functional" affordances, which allow reasoning about how to use the same. A baseline model is provided, which highlights the value of the affordance labels in a variety of robotics applications. Luc Schoot Uiterkamp, Gwenn Englebienne, Dirk Heylen |
RO-MAN | 2 |
| 2023 | A GNN-Based Architecture for Group Detection from Spatio-Temporal Trajectory Data
Maedeh Nasri, Zhizhou Fang, Mitra Baratchi, Gwenn Englebienne, Shenghui Wang 0001, Alexander Koutamanis, Carolien Rieffe |
IDA | 4 |
| 2023 | Toward Standard Guidelines to Design the Sense of Embodiment in Teleoperation Applications: A Review and ToolboxabstractWe present a literature review and a toolbox to help the reader find the best method to design for and assess Sense of Embodiment (SoE) in several application scenarios. The main examples are based on teleoperation applications, due the challenges that these applications present. The three embodiment components that we consider to describe SoE are sense of ownership, sense of agency, and sense of self-location. We relate each embodiment component to the most often used assessment measures, test tasks, and application scenarios. The toolbox is built to efficiently design, test, and assess an embodiment experience, following seven concrete steps. We provide four main contributions: 1) a literature review of the assessment measures and strategies used to measure SoE; 2) a systematic categorization of SoE measures; 3) A categorization of the main test tasks used in SoE assessment; and 4) a toolbox consisting of seven steps as guidance to design SoE. We included several examples and tables to guide the user step by step through the design of an embodiment experience. Sara Falcone, Gwenn Englebienne, Jan B. F. van Erp, Dirk Heylen |
Hum. Comput. Interact. | 2 |
| 2022 | Pupil Diameter as Implicit Measure to Estimate Sense of Embodiment
Sara Falcone, Anne-Marie Brouwer, Dirk Heylen, Jan B. F. van Erp, Saket Sachin Pradhan, Ivo V. Stuldreher, Ioana Cocu, Martijn Heuvel, Pieter S. de Vries, Kaj Gijsbertse, Gwenn Englebienne |
CogSci | 12 |
| 2022 | EMG-based Feedback Modulation for Increased Transparency in TeleoperationabstractIn interacting with stiff environments through teleoperated systems, time delays cause a mismatch between haptic feedback and the expected feedback by the operator. This mismatch causes artefacts in the feedback, which decrease transparency, but so does filtering these artefacts. Through modelling of operator stiffness and the expected feedback force with EMG, the artifacts can be selectively filtered without loss of transparency. We developed several feedback modulation techniques to bring the feedback force closer to the expected force: 1) the average between the modelled operator force and the feedback force, 2) a low pass filter and 3) a scaling modulation. To control for overdamping, a transparency check is included. We show that the averaging approach yields significantly better contacts than unmodulated feedback. None of the modulation algorithms differ significantly from the unmodulated feedback in transparency. Luc Schoot Uiterkamp, Francesco Porcini, Gwenn Englebienne, Antonio Frisoli, Douwe Dresscher |
IROS | 3 |
| 2022 | What Comes After Telepresence? Embodiment, Social Presence and Transporting One's Functional and Social SelfabstractAdvances in robotics and multisensory displays allow extending telepresence ambitions beyond only the feeling of being present at a remote location In this paper, we discuss what may lie beyond telepresence and how we can transport both the functional and social self of a user. We introduce the embodiment illusion and its potential contribution to task performance and list important cues to evoke this illusion, including synchronicity in multisensory information, a first-person visual perspective, and a human-like visual appearance and anatomy of the telepresence robot. We also introduce the concept of social presence and the important bidirectional social cues it needs, including eye contact, facial expression, posture, gestures, and social touch. For all these multisensory and social cues, we explain how they can be implemented in a telepresence system and describe our solution consisting of a closed control pod and a humanoid telepresence robot. Jan B. F. van Erp, Camille Sallaberry, Christiaan Brekelmans, Douwe Dresscher, Frank Bart ter Haar, Gwenn Englebienne, Jeanine van Bruggen, Joachim de Greeff, Leonor Fermoselle, Alexander Toet, Nirul Hoeba, Robin Lieftink, Sara Falcone, Tycho J. H. Brug |
SMC | 6 |
| 2022 | It's Complicated: The Relationship between User Trust, Model Accuracy and Explanations in AIabstractAutomated decision-making systems become increasingly powerful due to higher model complexity. While powerful in prediction accuracy, Deep Learning models are black boxes by nature, preventing users from making informed judgments about the correctness and fairness of such an automated system. Explanations have been proposed as a general remedy to the black box problem. However, it remains unclear if effects of explanations on user trust generalise over varying accuracy levels. In an online user study with 959 participants, we examined the practical consequences of adding explanations for user trust: We evaluated trust for three explanation types on three classifiers of varying accuracy. We find that the influence of our explanations on trust differs depending on the classifier’s accuracy. Thus, the interplay between trust and explanations is more complex than previously reported. Our findings also reveal discrepancies between self-reported and behavioural trust, showing that the choice of trust measure impacts the results. Andrea Papenmeier, Dagmar Kern, Gwenn Englebienne, Christin Seifert |
ACM Trans. Comput. Hum. Interact. | 3 |
| 2021 | Individual Action and Group Activity Recognition in Soccer Videos from a Static Panoramic CameraabstractData and statistics are key to soccer analytics and have important roles in player evaluation and fan engagement. Automatic recognition of soccer events - such as passes and corners - would ease the data gathering process, potentially opening up the market for soccer analytics at non professional clubs. Existing approaches extract events on group level only and rely on television broadcasts or recordings from multiple camera viewpoints. We propose a novel method for the recognition of individual actions and group activities in panoramic videos from a single viewpoint. Three key contributions in the proposed method are (1) player snippets as model input, (2) independent extraction of spatio-temporal features per player, and (3) feature contextuali-sation using zero-padding and feature suppression in graph attention networks. Our method classifies video samples in eight action and eleven activity types, and reaches accuracies above 75% for ten of these classes. Beerend G. A. Gerats, Henri Bouma, Wouter Uijens, Gwenn Englebienne, Luuk J. Spreeuwers |
ICPRAM | 4 |
| 2021 | Towards Analyzing and Predicting the Experience of Live Performances with Wearable SensingabstractWe present an approach to interpret the response of audiences to live performances by processing mobile sensor data. We apply our method on three different datasets obtained from three live performances, where each audience member wore a single tri-axial accelerometer and proximity sensor embedded inside a smart sensor pack. Using these sensor data, we developed a novel approach to predict audience members’ self-reported experience of the performances in terms of enjoyment, immersion, willingness to recommend the event to others, and change in mood. The proposed method uses an unsupervised method to identify informative intervals of the event, using the linkage of the audience members’ bodily movements, and uses data from these intervals only to estimate the audience members’ experience. We also analyze how the relative location of members of the audience can affect their experience and present an automatic way of recovering neighborhood information based on proximity sensors. We further show that the linkage of the audience members’ bodily movements is informative of memorable moments which were later reported by the audience. Ekin Gedik, Laura Cabrera Quiros, Claudio Martella, Gwenn Englebienne, Hayley Hung |
IEEE Trans. Affect. Comput. | 4 |
| 2019 | Non-Parametric Subject Prediction
Shenghui Wang 0001, Rob Koopman, Gwenn Englebienne |
TPDL | 3 |
| 2017 | Learning spectro-temporal features with 3D CNNs for speech emotion recognitionabstractIn this paper, we propose to use deep 3-dimensional convolutional networks (3D CNNs) in order to address the challenge of modelling spectro-temporal dynamics for speech emotion recognition (SER). Compared to a hybrid of Convolutional Neural Network and Long-Short-Term-Memory (CNN-LSTM), our proposed 3D CNNs simultaneously extract short-term and long-term spectral features with a moderate number of parameters. We evaluated our proposed and other state-of-the-art methods in a speaker-independent manner using aggregated corpora that give a large and diverse set of speakers. We found that 1) shallow temporal and moderately deep spectral kernels of a homogeneous architecture are optimal for the task; and 2) our 3D CNNs are more effective for spectro-temporal feature learning compared to other methods. Finally, we visualised the feature space obtained with our proposed method using t-distributed stochastic neighbour embedding (T-SNE) and could observe distinct clusters of emotions. Jaebok Kim, Khiet P. Truong, Gwenn Englebienne, Vanessa Evers |
ACII | 3 |
| 2017 | Towards Speech Emotion Recognition "in the Wild" Using Aggregated Corpora and Deep Multi-Task LearningabstractOne of the challenges in Speech Emotion Recognition (SER) "in the wild" is the large mismatch between training and test data (e.g.speakers and tasks).In order to improve the generalisation capabilities of the emotion models, we propose to use Multi-Task Learning (MTL) and use gender and naturalness as auxiliary tasks in deep neural networks.This method was evaluated in within-corpus and various cross-corpus classification experiments that simulate conditions "in the wild".In comparison to Single-Task Learning (STL) based state of the art methods, we found that our MTL method proposed improved performance significantly.Particularly, models using both gender and naturalness achieved more gains than those using either gender or naturalness separately.This benefit was also found in the high-level representations of the feature space, obtained from our method proposed, where discriminative emotional clusters could be observed. Jaebok Kim, Gwenn Englebienne, Khiet P. Truong, Vanessa Evers |
INTERSPEECH | 2 |
| 2017 | Deep Temporal Models using Identity Skip-Connections for Speech Emotion RecognitionabstractDeep architectures using identity skip-connections have demonstrated groundbreaking performance in the field of image classification. Recently, empirical studies suggested that identity skip-connections enable ensemble-like behaviour of shallow networks, and that depth is not a solo ingredient for their success. Therefore, we examine the potential of identity skip-connections for the task of Speech Emotion Recognition (SER) where moderately deep temporal architectures are often employed. To this end, we propose a novel architecture which regulates unimpeded feature flows and captures long-term dependencies via gate-based skip-connections and a memory mechanism. Our proposed architecture is compared to other state-of-the-art methods of SER and is evaluated on large aggregated corpora recorded in different contexts. Our proposed architecture outperforms the state-of-the-art methods by 9 - 15% and achieves an Unweighted Accuracy of 80.5% in an imbalanced class distribution. In addition, we examine a variant adopting simplified skip-connections of Residual Networks (ResNet) and show that gate-based skip-connections are more effective than simplified skip-connections. Jaebok Kim, Gwenn Englebienne, Khiet P. Truong, Vanessa Evers |
ACM Multimedia | 2 |
| 2017 | Blame my telepresence robot joint effect of proxemics and attribution on interpersonal attractionabstractWhen remote users share autonomy with a telepresence robot, questions arise as to how the behaviour of the robot is interpreted by local users. We investigated how a robot's violations of social norms under shared autonomy influence the local user's evaluation of the robot's remote users. Specifically, we examined how attribution of such violations to either the robot or the remote user influences social perception of the remote user. Using personal space invasion as a salient social norm violation, we conducted a within-subject experiment (n=20) to investigate these questions. Participants saw several people introducing themselves through a telepresence robot, personal space invasion and attribution were manipulated. We found a significant (p=0.007) joint effect of the manipulations on interpersonal attraction. After these first 20 participants our robot broke down, and we had to continue with another robot (n=20). We found a difference between the two robots, causing us to discard this data from our main analysis. Subsequent video annotation and comparison of the two robots suggests that accuracy of the followed trajectory modifies attribution. Our results offer insights into the mechanisms of attribution in interactions with a telepresence robot as a mediator. Josca van Houwelingen-Snippe, Jered Vroon, Gwenn Englebienne, Pim Haselager |
RO-MAN | 3 |
| 2017 | Learning to Recognize Human Activities Using Soft LabelsabstractHuman activity recognition system is of great importance in robot-care scenarios. Typically, training such a system requires activity labels to be both completely and accurately annotated. In this paper, we go beyond such restriction and propose a learning method that allow labels to be incomplete and uncertain. We introduce the idea of soft labels which allows annotators to assign multiple, and weighted labels to data segments. This is very useful in many situations, e.g., when the labels are uncertain, when part of the labels are missing, or when multiple annotators assign inconsistent labels. We formulate the activity recognition task as a sequential labeling problem. Latent variables are embedded in the model in order to exploit sub-level semantics for better estimation. We propose a max-margin framework which incorporate soft labels for learning the model parameters. The model is evaluated on two challenging datasets. To simulate the uncertainty in data annotation, we randomly change the labels for transition segments. The results show significant improvement over the state-of-the-art approach. Ninghang Hu, Gwenn Englebienne, Zhongyu Lou, Ben J. A. Kröse |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Unsupervised visit detection in smart homes
Ahmed Nait Aicha, Gwenn Englebienne, Ben J. A. Kröse |
Pervasive Mob. Comput. | 2 |
| 2017 | Delta Features From Ambient Sensor Data are Good Predictors of Change in Functional HealthabstractSensor systems can be deployed in the homes of older adults living alone for functional health assessments. Their information is very useful for health care specialists. The problem lies in developing person independent models while facing a large variability in behavior. We address this problem by, first, proposing a new feature extraction method for data from ambient motion sensors. The method uses functional similarities between houses and daily structure to extract meaningful features. Second, we propose a change-based approach for analyzing data, taking difference scores of both the sensor features and health metrics. To evaluate our approach, experiments on longitudinal data were conducted, where the relationship between sensor data and health measurements was modeled with linear regression and (nonlinear) regression forests. These experiments show that the change-based approach yields better results and that the resulting models can be used as a reliable metric for (functional) health. In addition, feature analysis can help health care specialists understand relevant aspects of behavior. Prediction of health metrics is possible even with simple sensors. With such sensors, it is possible to detect problems and health decline in an early stage. This will have great impact on clinical practice. Saskia Robben, Gwenn Englebienne, Ben J. A. Kröse |
IEEE J. Biomed. Health Informatics | 2 |
| 2016 | Human intent forecasting using intrinsic kinematic constraintsabstractThe performance of human-robot collaboration tasks can be improved by incorporating predictions of the human collaborator's movement intentions. These predictions allow a collaborative robot to both provide appropriate assistance and plan its own motion so it does not interfere with the human. In the specific case of human reach intent prediction, prior work has divided the task into two pieces: recognition of human activities and prediction of reach intent. In this work, we propose a joint model for simultaneous recognition of human activities and prediction of reach intent based on skeletal pose. Since future reach intent is tightly linked to the action a person is performing at present, we hypothesize that this joint model will produce better performance on the recognition and prediction tasks than past approaches. In addition, our approach incorporates a simple human kinematic model which allows us to generate features that compactly capture the reachability of objects in the environment and the motion cost to reach those objects, which we anticipate will improve performance. Experiments using the CAD-120 benchmark dataset show that both the joint modeling approach and the human kinematic features give improved F1 scores versus the previous state of the art. Ninghang Hu, Aaron M. Bestick, Gwenn Englebienne, Ruzena Bajcsy, Ben J. A. Kröse |
IROS | 3 |
| 2016 | Incorporating perception uncertainty in human-aware navigation: A comparative studyabstractIn this work, we present a novel approach to human-aware navigation by probabilistically modelling the uncertainty of perception for a social robotic system and investigating its effect on the overall social navigation performance. The model of the social costmap around a person has been extended to consider this new uncertainty factor, which has been widely neglected despite playing an important role in situations with noisy perception. A social path planner based on the fast marching method has been augmented to account for the uncertainty in the positions of people. The effectiveness of the proposed approach has been tested in extensive experiments carried out with real robots and in simulation. Real experiments have been conducted, given noisy perception, in the presence of single/multiple, static/dynamic humans. Results show how this approach has been able to achieve trajectories that are able to keep a more appropriate social distance to the people, compared to those of the basic navigation approach, and the human-aware navigation approach which relies solely on perfect perception, when the complexity of the environment increases. Accounting for uncertainty of perception is shown to result in smoother trajectories with lower jerk that are more natural from the point of view of humans. Zeynab Talebpour, Deepak Viswanathan, Rodrigo M. M. Ventura, Gwenn Englebienne, Alcherio Martinoli |
RO-MAN | 4 |
| 2016 | Mixture of Switching Linear Dynamics to Discover Behavior Patterns in Object TracksabstractWe present a novel non-parametric Bayesian model to jointly discover the dynamics of low-level actions and high-level behaviors of tracked objects. In our approach, actions capture both linear, low-level object dynamics, and an additional spatial distribution on where the dynamic occurs. Furthermore, behavior classes capture high-level temporal motion dependencies in Markov chains of actions, thus each learned behavior is a switching linear dynamical system. The number of actions and behaviors is discovered from the data itself using Dirichlet Processes. We are especially interested in cases where tracks can exhibit large kinematic and spatial variations, e.g. person tracks in open environments, as found in the visual surveillance and intelligent vehicle domains. The model handles real-valued features directly, so no information is lost by quantizing measurements into 'visual words', and variations in standing, walking and running can be discovered without discrete thresholds. We describe inference using Markov Chain Monte Carlo sampling and validate our approach on several artificial and real-world pedestrian track datasets from the surveillance and intelligent vehicle domain. We show that our model can distinguish between relevant behavior patterns that an existing state-of-the-art hierarchical model for clustering and simpler model variants cannot. The software and the artificial and surveillance datasets are made publicly available for benchmarking purposes. Julian F. P. Kooij, Gwenn Englebienne, Dariu Gavrila |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | A hierarchical representation for human activity recognition with noisy labelsabstractHuman activity recognition is an essential task for robots to effectively and efficiently interact with the end users. Many machine learning approaches for activity recognition systems have been proposed recently. Most of these methods are built upon a strong assumption that the labels in the training data are noise-free, which is often not realistic. In this paper, we incorporate the uncertainty of labels into a max-margin learning algorithm, and the algorithm allows the labels to deviate over iterations in order to find a better solution. This is incorporated with a hierarchical approach where we jointly estimate activities at two different levels of granularity. The model is tested on two datasets, i.e., the CAD-120 dataset and the Accompany dataset, and the proposed model shows outperforming results over the state-of-the-art methods. Ninghang Hu, Gwenn Englebienne, Zhongyu Lou, Ben J. A. Kröse |
IROS | 2 |
| 2015 | How Was It?: Exploiting Smartphone Sensing to Measure Implicit Audience Responses to Live PerformancesabstractIn this paper, we present an approach to understand the response of an audience to a live dance performance by the processing of mobile sensor data. We argue that exploiting sensing capabilities already available in smart phones enables a potentially large scale measurement of an audience's implicit response to a performance. In this work, we leverage both tri-axial accelerometers, worn by ordinary members of the public during a dance performance, to predict responses to a number of survey answers, comprising enjoyment, immersion, willingness to recommend the event to others, and change in mood. We also analyse how behaviour as a result of seeing a dance performance might be reflected in a people's subsequent social behaviour using proximity and acceleration sensing. To our knowledge, this is the first work where pervasive mobile sensing has been used to investigate spontaneous responses to predict the affective evaluation of a live performance. Using a single body worn accelerometer to monitor a set of audience members, we were able to predict whether they enjoyed the event with a balanced classification accuracy of 90\%. The collective coordination of the audience's bodily movements also highlighted memorable moments that were reported later by the audience. The effective use of body movements to measure affective responses in such a setting is particularly surprising given that traditionally, physiological signals such as skin conductance or brain-based signals are the more commonly accepted methods to measure implicit affective response. Our experiments open interesting new directions for research on both automated techniques and applications for the implicit tagging of real world events via spontaneous and implicit audience responses during as well as after a performance. Claudio Martella, Ekin Gedik, Laura Cabrera Quiros, Gwenn Englebienne, Hayley Hung |
ACM Multimedia | 4 |
| 2015 | Dynamics of social positioning patterns in group-robot interactionsabstractWhen a mobile robot interacts with a group of people, it has to consider its position and orientation. We introduce a novel study aimed at generating hypotheses on suitable behavior for such social positioning, explicitly focusing on interaction with small groups of users and allowing for the temporal and social dynamics inherent in most interactions. In particular, the interactions we look at are approach, converse and retreat. In this study, groups of three participants and a telepresence robot (controlled remotely by a fourth participant) solved a task together while we collected quantitative and qualitative data, including tracking of positioning/orientation and ratings of the behaviors used. In the data we observed a variety of patterns that can be extrapolated to hypotheses using inductive reasoning. One such pattern/hypothesis is that a (telepresence) robot could pass through a group when retreating, without this affecting how comfortable that retreat is for the group members. Another is that a group will rate the position/orientation of a (telepresence) robot as more comfortable when it is aimed more at the center of that group. Jered Vroon, Michiel Joosse, Manja Lohse, Jan Kolkmeier, Jaebok Kim, Khiet P. Truong, Gwenn Englebienne, Dirk Heylen, Vanessa Evers |
RO-MAN | 7 |
| 2015 | Identifying multiple objects from their appearance in inaccurate detections
Julian F. P. Kooij, Gwenn Englebienne, Dariu Gavrila |
Comput. Vis. Image Underst. | 2 |
| 2015 | RARE: people detection in crowded passages by range image reconstruction
Tim van Oosterhout, Gwenn Englebienne, Ben J. A. Kröse |
Mach. Vis. Appl. | 2 |
| 2015 | Latent Hierarchical Model for Activity RecognitionabstractWe present a novel hierarchical model for human activity recognition. In contrast with approaches that successively recognize actions and activities, our approach jointly models actions and activities in a unified framework, and their labels are simultaneously predicted. The model is embedded with a latent layer that is able to capture a richer class of contextual information in both state-state and observation-state pairs. Although loops are present in the model, the model has an overall linear-chain structure, where the exact inference is tractable. Therefore, the model is very efficient in both inference and learning. The parameters of the graphical model are learned with a structured support vector machine. A data-driven approach is used to initialize the latent variables; therefore, no manual labeling for the latent states is required. The experimental results from using two benchmark datasets show that our model outperforms the state-of-the-art approach, and our model is computationally more efficient. Ninghang Hu, Gwenn Englebienne, Zhongyu Lou, Ben J. A. Kröse |
IEEE Trans. Robotics | 2 |
| 2014 | Detecting conversing groups with a single worn accelerometerabstractIn this paper we propose the novel task of detecting groups of conversing people using only a single body-worn accelerometer per person. Our approach estimates each individual's social actions and uses the co-ordination of these social actions between pairs to identify group membership. The aim of such an approach is to be deployed in dense crowded environments. Our work differs significantly from previous approaches, which have tended to rely on audio and/or proximity sensing, often in much less crowded scenarios, for estimating whether people are talking together or who is speaking. Ultimately, we are interested in detecting who is speaking, who is conversing with whom, and from that, to infer socially relevant information about the interaction such as whether people are enjoying themselves, or the quality of their relationship in these extremely dense crowded scenarios. Striving towards this long-term goal, this paper presents a systematic study to understand how to detect groups of people who are conversing together in this setting, where we achieve a $64%$ classification accuracy using a fully automated system. Hayley Hung, Gwenn Englebienne, Laura Cabrera Quiros |
ICMI | 2 |
| 2014 | Learning latent structure for activity recognitionabstractWe present a novel latent discriminative model for human activity recognition. Unlike the approaches that require conditional independence assumptions, our model is very flexible in encoding the full connectivity among observations, latent states, and activity states. The model is able to capture richer class of contextual information in both state-state and observation-state pairs. Although loops are present in the model, we can consider the graphical model as a linear-chain structure, where the exact inference is tractable. Thereby the model is very efficient in both inference and learning. The parameters of the graphical model are learned with the Structured-Support Vector Machine (Structured-SVM). A data-driven approach is used to initialize the latent variables, thereby no hand labeling for the latent states is required. Experimental results on the CAD-120 benchmark dataset show that our model outperforms the state-of-the-art approach by over 5% in both precision and recall, while our model is more efficient in computation. Ninghang Hu, Gwenn Englebienne, Zhongyu Lou, Ben J. A. Kröse |
ICRA | 2 |
| 2014 | A two-layered approach to recognize high-level human activitiesabstractAutomated human activity recognition is an essential task for Human Robot Interaction (HRI). A successful activity recognition system enables an assistant robot to provide precise services. In this paper, we present a two-layered approach that can recognize sub-level activities and high-level activities successively. In the first layer, the low-level activities are recognized based on the RGB-D video. In the second layer, we use the recognized low-level activities as input features for estimating high-level activities. Our model is embedded with a latent node, so that it can capture a richer class of sub-level semantics compared with the traditional approach. Our model is evaluated on a challenging benchmark dataset. We show that the proposed approach outperforms the single-layered approach, suggesting that the hierarchical nature of the model is able to better explain the observed data. The results also show that our model outperforms the state-of-the-art approach in accuracy, precision and recall. Ninghang Hu, Gwenn Englebienne, Ben J. A. Kröse |
RO-MAN | 2 |
| 2014 | Behavior analysis of elderly using topic models
Kristin Rieping, Gwenn Englebienne, Ben J. A. Kröse |
Pervasive Mob. Comput. | 2 |
| 2013 | Solving Person Re-identification in Non-overlapping Camera using Efficient Gibbs SamplingabstractThis paper proposes a novel probabilistic approach for appearance-based person reidentification in non-overlapping camera networks.It accounts for varying illumination, varying camera gain and has low computational complexity.More specifically, we present a graphical model where we model the person's appearance in addition to camera illumination and gain.We analytically derive the solutions for the person's appearance and camera properties, and use a novel constant time Gibbs sampling scheme to estimate the identification labels.We validate our algorithm on two indoor datasets and perform a comparative analysis with existing algorithms.We demonstrate significantly increased re-identification accuracy in addition to significantly reducing the computational complexity on our datasets. Vijay John, Gwenn Englebienne, Ben J. A. Kröse |
BMVC | 2 |
| 2013 | Classifying social actions with a single accelerometerabstractIn this paper, we estimate different types of social actions from a single body-worn accelerometer in a crowded social setting. Accelerometers have many advantages in such settings: they are impervious to environmental noise, unobtrusive, cheap, low-powered, and their readings are specific to a single person. Our experiments show that they are surprisingly informative of different types of social actions. The social actions we address in this paper are whether a person is speaking, laughing, gesturing, drinking, or stepping. To our knowledge, this is the first work to carry out experiments on estimating social actions from conversational behavior using only a wearable accelerometer. The ability to estimate such actions using just the acceleration opens up the potential for analyzing more about social aspects of people's interactions without explicitly recording what they are saying. Hayley Hung, Gwenn Englebienne, Jeroen Kools |
UbiComp | 2 |
| 2013 | Person re-identification using height-based gait in colour depth cameraabstractWe address the problem of person re-identification in colour-depth camera using the height temporal information of people. Our proposed gait-based feature corresponds to the frequency response of the height temporal information. We demonstrate that the discriminative periodic motion associated with human gait is encoded within the height temporal information. Additionally, we also investigate the discriminative ability of a novel feature vector obtained by the integration of the height temporal information with a color and height-based appearance model. Given the proposed features we adopt a feature selection scheme for each person, based on KL-divergence, to identify the discriminative subset of frequency bins for the height-based gait feature that enhances the overall identification accuracy. To identify the test person, we formulate a probabilistic matching framework incorporating the selected frequency bins. We validate our algorithm on the publicly available TUM-GAID dataset and our studio datasets and report over 80% accuracy for 75 people with our combined feature, significantly better than standard colour-based features. Additionally, we also observe that height-based gait features reports over 90% for smaller population and are rotation invariant, robust to appearance noise and occlusion. Vijay John, Gwenn Englebienne, Ben J. A. Kröse |
ICIP | 2 |
| 2013 | Posture recognition with a top-view cameraabstractWe describe a system that recognizes human postures with heavy self-occlusion. In particular, we address posture recognition in a robot assisted-living scenario, where the environment is equipped with a top-view camera for monitoring human activities. This setup is very useful because top-view cameras lead to accurate localization and limited inter-occlusion between persons, but conversely they suffer from body parts being frequently self-occluded. The conventional way of posture recognition relies on good estimation of body part positions, which turns out to be unstable in the top-view due to occlusion and foreshortening. In our approach, we learn a posture descriptor for each specific posture category. The posture descriptor encodes how well the person in the image can be `explained' by the model. The postures are subsequently recognized from the matching scores returned by the posture descriptors. We select the state-of-the-art approach of pose estimation as our posture descriptor. The results show that our method is able to correctly classify 79.7% of the test sample, which outperforms the conventional approach by over 23%. Ninghang Hu, Gwenn Englebienne, Ben J. A. Kröse |
IROS | 2 |
| 2012 | A Non-parametric Hierarchical Model to Discover Behavior Dynamics from Tracks
Julian F. P. Kooij, Gwenn Englebienne, Dariu Gavrila |
ECCV (6) | 2 |
| 2012 | Multimodal Speaker DiarizationabstractWe present a novel probabilistic framework that fuses information coming from the audio and video modality to perform speaker diarization. The proposed framework is a Dynamic Bayesian Network (DBN) that is an extension of a factorial Hidden Markov Model (fHMM) and models the people appearing in an audiovisual recording as multimodal entities that generate observations in the audio stream, the video stream, and the joint audiovisual space. The framework is very robust to different contexts, makes no assumptions about the location of the recording equipment, and does not require labeled training data as it acquires the model parameters using the Expectation Maximization (EM) algorithm. We apply the proposed model to two meeting videos and a news broadcast video, all of which come from publicly available data sets. The results acquired in speaker diarization are in favor of the proposed multimodal framework, which outperforms the single modality analysis results and improves over the state-of-the-art audio-based speaker diarization. Athanasios K. Noulas, Gwenn Englebienne, Ben J. A. Kröse |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Move, and i will tell you who you are: detecting deceptive roles in low-quality dataabstractMotion, like speech, provides information about one's emotional state. This work introduces an automated non-verbal audio-visual approach for detecting deceptive roles in multi-party conversations using low resolution video. We show how using simple features extracted from motion and speech improves over speech-only for the detection of deceptive roles. Our results show that deceptive players were recognised with significantly higher precision when video features were used. We improve the classification performance with 22.6% compared to our baseline. Nimrod Raiman, Hayley Hung, Gwenn Englebienne |
ICMI | 3 |
| 2010 | An activity monitoring system for elderly care using generative and discriminative modelsabstractAn activity monitoring system allows many applications to assist in care giving for elderly in their homes. In this paper we present a wireless sensor network for unintrusive observations in the home and show the potential of generative and discriminative models for recognizing activities from such observations. Through a large number of experiments using four real world datasets we show the effectiveness of the generative hidden Markov model and the discriminative conditional random fields in activity recognition. Tim van Kasteren, Gwenn Englebienne, Ben J. A. Kröse |
Pers. Ubiquitous Comput. | 2 |
| 2008 | Accurate activity recognition in a home settingabstractA sensor system capable of automatically recognizing activities would allow many potential ubiquitous applications. In this paper, we present an easy to install sensor network and an accurate but inexpensive annotation method. A recorded dataset consisting of 28 days of sensor data and its annotation is described and made available to the community. Through a number of experiments we show how the hidden Markov model and conditional random fields perform in recognizing activities. We achieve a timeslice accuracy of 95.6% and a class accuracy of 79.4%. Tim van Kasteren, Athanasios K. Noulas, Gwenn Englebienne, Ben J. A. Kröse |
UbiComp | 3 |
| 2008 | Learning Concept Mappings from Instance Similarity
Shenghui Wang 0001, Gwenn Englebienne, Stefan Schlobach |
ISWC | 2 |
| 2007 | A probabilistic model for generating realistic lip movements from speechabstractThe present work aims to model the correspondence between facial motion and speech. The face and sound are modelled separately, with phonemes being the link between both. We propose a sequential model and evaluate its suitability for the generation of the facial animation from a sequence of phonemes, which we obtain from speech. We evaluate the results both by computing the error between generated sequences and real video, as well as with a rigorous double-blind test with human subjects. Experiments show that our model compares favourably to other existing methods and that the sequences generated are comparable to real video sequences. Gwenn Englebienne, Timothy F. Cootes, Magnus Rattray |
NIPS | 1 |