EDBT 2026 Demo / reviewers in the wild / expert
Jean-Marc Odobez
dblp:98/1208
· DBLP profile ↗
136ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0002-9537-9898ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 81 · 10 first-author · 11 since 2021Artificial intelligence and machine learning · 67 · 2 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 23 · 2 since 2021Systems, architecture and hardware · 6 · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing
Masoumeh Chapariniya, Jean-Marc Odobez, Volker Dellwo, Teodora Vukovic |
FG | 2 |
| 2026 | GeoSACS: Geometric Shared Autonomy via Canal SurfacesabstractShared autonomy (SA), which combines user inputs with autonomous capabilities, presents a significant opportunity for assistive robotics. A key challenge in SA is the dimensionality gap: the mismatch between low-dimensional user inputs from familiar interfaces (e.g., 2D joysticks) and the high-dimensional control required by robot manipulators. To enhance usability and acceptance, this mapping must be as simple and intuitive as possible. We introduce GeoSACS, a geometric framework for SA. GeoSACS uses canal surfaces to encode task structure with as few as two demonstrations. While the robot moves autonomously along the canal, users can then make corrections on the 2D planar circular cross-sections orthogonal to the robot motion. By leveraging geometric structure to partition the 6D control space between the robot and the user, GeoSACS allows the intuitive mapping of 2D user inputs to 6D end-effector control. We describe GeoSACS and evaluate its underlying assumptions in a user study against two baselines. Results from the study demonstrate reduced workload and improved performance, providing insights for the design of future SA systems. Shalutha Rajapakshe, Atharva Dastenavar, Michael Hagenow, Jean-Marc Odobez, Emmanuel Senft |
HRI | 4 |
| 2025 | Enhancing 3D Gaze Estimation in the Wild using Weak Supervision with Gaze Following Labelsabstractin-the-wild 3D gaze datasets. To address these challenges, we introduce a novel Self-Training Weakly-Supervised Gaze Estimation framework (ST-WSGE). This two-stage learning framework leverages diverse 2D gaze datasets, such as gaze-following data, which offer rich variations in appearances, natural scenes, and gaze distributions, and proposes an approach to generate 3D pseudo-labels and enhance model generalization. Furthermore, traditional modality-specific models, designed separately for images or videos, limit the effective use of available training data. To overcome this, we propose the Gaze Transformer (GaT), a modality-agnostic architecture capable of simultaneously learning static and dynamic gaze information from both image and video datasets. By combining 3D video datasets with 2D gaze target labels from gaze following tasks, our approach achieves the following key contributions: (i) Significant state-of-the-art improvements in within-domain and cross-domain generalization on unconstrained benchmarks like Gaze360 and GFIE, with notable cross-modal gains in video gaze estimation; (ii) Superior cross-domain performance on datasets such as MPIIFaceGaze and Gaze360 compared to frontal face methods. Code and pre-trained models will be released to the community. Pierre Vuillecard, Jean-Marc Odobez |
CVPR | 2 |
| 2025 | Giving Sense to Inputs: Toward an Accessible Control Framework for Shared AutonomyabstractWhile shared autonomy offers significant potential for assistive robotics, key questions remain about how to effectively map 2D control inputs to 6D robot motions. An intuitive framework should allow users to input commands effortlessly, with the robot responding as expected, without users needing to anticipate the impact of their inputs. In this article, we propose a dynamic input mapping framework that links joystick movements to motions on control frames defined along a trajectory encoded with canal surfaces. We evaluate our method in a user study with 20 participants, demonstrating that our input mapping framework reduces the workload and improves usability compared to a baseline mapping with similar motion encoding. To prepare for deployment in assistive scenarios, we built on the development from the accessible gaming community to select an accessible control interface. We then tested the system in an exploratory study, where three wheelchair users controlled the robot for both daily living activities and a creative painting task, demonstrating its feasibility for users closer to our target population. Shalutha Rajapakshe, Jean-Marc Odobez, Emmanuel Senft |
HRI | 2 |
| 2025 | Loose Social-Interaction Recognition in Real-World Therapy ScenariosabstractThe computer vision community has explored dyadic interactions for atomic actions such as pushing, carrying-object, etc. However, with the advancement in deep learning models, there is a need to explore more complex dyadic situations such as loose interactions. These are interactions where two people perform certain atomic activities to complete a global action irrespective of temporal synchronisation and physical engagement, like cooking-together for example. Analysing these types of dyadic-interactions has several useful applications in the medical domain for social-skills development and mental health diagnosis. To achieve this, we propose a novel dual-path architecture to capture the loose interaction between two individuals. Our model learns global abstract features from each stream via a CNNs backbone and fuses them using a new Global-Layer-Attention module based on a cross-attention strategy. We evaluate our model on real-world autism diagnoses such as our Loose-Interaction dataset, and the publicly available Autism dataset for loose interactions. Our network achieves baseline results on the Loose-Interaction and SOTA results on the Autism datasets. Moreover, we study different social interactions by experimenting on a publicly available dataset i.e. NTU-RGB+D (interactive classes from both NTU-60 and NTU-120). We have found that different interactions require different network designs. We also compare a slightly different version of our method (details in Section 3.6) by incorporating time information to address tight interactions achieving SOTA results. Abid Ali 0002, Rui Dai 0001, Ashish Marisetty, Guillaume Astruc, Monique Thonnat, Jean-Marc Odobez, Susanne Thümmler, François Brémond |
WACV | 6 |
| 2025 | From Forest to Zoo: Great Ape Behavior Recognition with ChimpBehaveabstractAbstract This paper addresses the significant challenge of recognizing behaviors in non-human primates, specifically focusing on chimpanzees. Automated behavior recognition is crucial for both conservation efforts and the advancement of behavioral research. However, it is often hindered by the labor-intensive process of manual video annotation. Despite the availability of large-scale animal behavior datasets, effectively applying machine learning models across varied environmental settings remains a critical challenge due to the variability in data collection contexts and the specificity of annotations. In this paper, we introduce ChimpBehave , a novel dataset comprising over 2 h and 20 min of video (approximately 215,000 frames) of zoo-housed chimpanzees, annotated with bounding boxes and fine-grained locomotive behavior labels. Uniquely, ChimpBehave aligns its behavior classes with those in PanAf, an existing dataset collected in distinct visual environments, enabling the study of cross-dataset generalization - where models are trained on one dataset and tested on another with differing data distributions. We benchmark ChimpBehave using state-of-the-art video-based and skeleton-based action recognition models, establishing performance baselines for both within-dataset and cross-dataset evaluations. Our results highlight the strengths and limitations of different model architectures, providing insights into the application of automated behavior recognition across diverse visual settings. The dataset, models, and code can be accessed at: https://github.com/MitchFuchs/ChimpBehave Michael Fuchs 0005, Emilie Genty, Adrian Bangerter, Klaus Zuberbühler, Jean-Marc Odobez, Paul Cotofrei |
Int. J. Comput. Vis. | 5 |
| 2024 | Weakly-Supervised Autism Severity Assessment in Long VideosabstractAutism Spectrum Disorder (ASD) is a diverse collection of neurobiological conditions marked by challenges in social communication and reciprocal interactions, as well as repetitive and stereotypical behaviors. Atypical behavior patterns in a long, untrimmed video can serve as biomarkers for children with ASD. In this paper, we propose a video-based weakly-supervised method that takes spatio-temporal features of long videos to learn typical and atypical behaviors for autism detection. On top of that, we propose a shallow TCN-MLP network, which is designed to further categorize the severity score. We evaluate our method on actual evaluation videos of children with autism collected and annotated (for severity score) by clinical professionals. Experimental results demonstrate the effectiveness of behaviors biomarkers that could help clinicians in autism spectrum analysis. Abid Ali 0002, Camilla Barbini, Séverine Dubuisson, Jean-Marc Odobez, François Brémond, Susanne Thümmler |
CBMI | 5 |
| 2024 | Sharingan: A Transformer Architecture for Multi-Person Gaze FollowingabstractGaze is a powerful form of non-verbal communication that humans develop from an early age. As such, modeling this behavior is an important task that can benefit a broad set of application domains ranging from robotics to sociology. In particular, the gaze following task in computer vision is defined as the prediction of the 2D pixel coordinates where a person in the image is looking. Previous attempts in this area have primarily centered on CNN-based architectures, but they have been constrained by the need to process one person at a time, which proves to be highly inefficient. In this paper, we introduce a novel and effective multi-person transformer-based architecture for gaze prediction. While there exist prior works using transformers for multi-person gaze prediction [38], [39], they use a fixed set of learnable embeddings to decode both the person and its gaze target, which requires a matching step afterward to link the predictions with the annotations. Thus, it is difficult to quantitatively evaluate these methods reliably with the available benchmarks, or integrate them into a larger human behavior understanding system. Instead, we are the first to propose a multi-person transformer-based architecture that maintains the original task formulation and ensures control over the people fed as input. Our main contribution lies in encoding the person-specific information into a single controlled token to be processed alongside image tokens and using its output for prediction based on a novel multiscale decoding mechanism. Our new architecture achieves state-of-the-art results on the GazeFollow, VideoAttentionTarget, and ChildPlay datasets and outperforms comparable multi-person architectures with a notable margin. Our code, checkpoints, and data extractions will be made publicly available soon. Samy Tafasca, Jean-Marc Odobez |
CVPR | 3 |
| 2024 | A Unified Model for Gaze Following and Social Gaze PredictionabstractHuman gaze plays a crucial role in communication and social interaction. Many recent studies have focused on predicting the 2D pixel location of a person's gaze target in an image. However, this approach has limitations when it comes to studying gaze for downstream applications that require analysis of higher-level social gaze behaviors. Previous works have post-processed the predicted 2D gaze target for social gaze prediction, however, we show that this approach is insufficient. Our proposed method jointly predicts the gaze target and social gaze behaviour, explicitly incorporating people interaction for state of the art results on three social gaze tasks - looking at heads, mutual gaze and shared attention. Additionally, we introduce evaluation protocols for these tasks, presenting a promising avenue for future research in gaze behavior analysis. Samy Tafasca, Naravich Chutisilp, Jean-Marc Odobez |
FG | 4 |
| 2024 | CCDb-HG: Novel Annotations and Gaze-Aware Representations for Head Gesture RecognitionabstractDespite remarkable progress in various human behavior perception tasks, head gesture recognition (HGR) has received limited attention in terms of datasets, benchmarks, and methods. In this work, we aim to address this gap and make two main contributions. First, we densely annotated the existing large-scale conversational dataset CCDb with diverse head gesture categories. This results in the CCDb-HG dataset, which can serve as a comprehensive benchmark for HGR research. Secondly, while previous gesture recognition methods have largely relied on head pose or facial landmarks as input, we propose to explore in addition the use of gaze to resolve ambiguous cases. This follows from the fact that head dynamics in interactions is driven by two main functions: communication (i.e. head gestures) and attention (i.e. gazing at other people or objects of interest). In fact, the head dynamics associated with attention activities can be confused for communication gestures, even though the gaze patterns are quite different in the two cases. In addition, we study several geometric and temporal data augmentation techniques to improve the generalization across novel viewpoints, as well as different model architectures to establish baseline performance on CCDb-HG. Our findings provide insights into various aspects of HGR and motivate further research in this field. To facilitate reproducibility, we will release the CCDb-HG annotations, code, and HGR models. Pierre Vuillecard, Arya Farkhondeh, Michael Villamizar, Jean-Marc Odobez |
FG | 4 |
| 2024 | MTGS: A Novel Framework for Multi-Person Temporal Gaze Following and Social Gaze PredictionabstractGaze following and social gaze prediction are fundamental tasks providing insights into human communication behaviors, intent, and social interactions. Most previous approaches addressed these tasks separately, either by designing highly specialized social gaze models that do not generalize to other social gaze tasks or by considering social gaze inference as an ad-hoc post-processing of the gaze following task. Furthermore, the vast majority of gaze following approaches have proposed models that can handle only one person at a time and are static, therefore failing to take advantage of social interactions and temporal dynamics. In this paper, we address these limitations and introduce a novel framework to jointly predict the gaze target and social gaze label for all people in the scene. It comprises (i) a temporal, transformer-based architecture that, in addition to frame tokens, handles person-specific tokens capturing the gaze information related to each individual; (ii) a new dataset, VSGaze, built from multiple gaze following and social gaze datasets by extending and validating head detections and tracks, and unifying annotation types. We demonstrate that our model can address and benefit from training on all tasks jointly, achieving state-of-the-art results for multi-person gaze following and social gaze prediction. Our annotations and code will be made publicly available. Samy Tafasca, Arya Farkhondeh, Pierre Vuillecard, Jean-Marc Odobez |
NeurIPS | 5 |
| 2024 | Toward Semantic Gaze Target DetectionabstractFrom the onset of infanthood, humans naturally develop the ability to closely observe and interpret the visual gaze of others. This skill, known as gaze following, holds significance in developmental theory as it enables us to grasp another person’s mental state, emotions, intentions, and more. In computer vision, gaze following is defined as the prediction of the pixel coordinates where a person in the image is focusing their attention. Existing methods in this research area have predominantly centered on pinpointing the gaze target by predicting a gaze heatmap or gaze point. However, a notable drawback of this approach is its limited practical value in gaze applications, as mere localization may not fully capture our primary interest — understanding the underlying semantics, such as the nature of the gaze target, rather than just its 2D pixel location. To address this gap, we extend the gaze following task, and introduce a novel architecture that simultaneously predicts the localization and semantic label of the gaze target. We devise a pseudo-annotation pipeline for the GazeFollow dataset, propose a new benchmark, develop an experimental protocol and design a suitable baseline for comparison. Our method sets a new state-of-the-art on the main GazeFollow benchmark for localization and achieves competitive results in the recognition task on both datasets compared to the baseline, with 40% fewer parameters Samy Tafasca, Victor Bros, Jean-Marc Odobez |
NeurIPS | 4 |
| 2023 | ChildPlay: A New Benchmark for Understanding Children's Gaze BehaviourabstractGaze behaviors such as eye-contact or shared attention are important markers for diagnosing developmental disorders in children. While previous studies have looked at some of these elements, the analysis is usually performed on private datasets and is restricted to lab settings. Furthermore, all publicly available gaze target prediction benchmarks mostly contain instances of adults, which makes models trained on them less applicable to scenarios with young children. In this paper, we propose the first study for predicting the gaze target of children and interacting adults. To this end, we introduce the ChildPlay dataset: a curated collection of short video clips featuring children playing and interacting with adults in uncontrolled environments (e.g. kindergarten, therapy centers, preschools etc.), which we annotate with rich gaze information. We further propose a new model for gaze target prediction that is geometrically grounded by explicitly identifying the scene parts in the 3D field of view (3DFoV) of the person, leveraging recent geometry preserving depth inference methods. Our model achieves state of the art results on benchmark datasets and ChildPlay. Furthermore, results show that looking at faces prediction performance on children is much worse than on adults, and can be significantly improved by fine-tuning models using child gaze annotations. Our dataset is available at https://www.idiap.ch/en/dataset/childplay-gaze. Code will be made available soon. Samy Tafasca, Jean-Marc Odobez |
ICCV | 3 |
| 2023 | A Multitask and Kernel Approach for Learning to Push Objects with a Target-Parameterized Deep Q-NetworkabstractPushing is an essential motor skill involved in several manipulation tasks, and has been an important research topic in robotics. Recent works have shown that Deep Q-Networks (DQNs) can learn pushing policies (when, where to push, and how) to solve manipulation tasks, potentially in synergy with other skills (e.g. grasping). Nevertheless, DQNs often assume a fixed setting and task, which may limit their deployment in practice. Furthermore, they suffer from sparse-gradient backpropagation when the action space is very large, a problem exacerbated by the fact that they are trained to predict state-action values based on a single reward function aggregating several facets of the task, rendering the model training challenging. To address these issues, we propose a multi-head target-parameterized DQN to learn robotic manipulation tasks, in particular pushing policies, and make the following contributions: i) we show that learning to predict different reward and task aspects can be beneficial compared to predicting a single value function where reward factors are not disentangled; ii) we study several alternatives to generalize a policy by encoding the target parameters either into the network layers or visually in the input; iii) we propose a kernelized version of the loss function, allowing to obtain better, faster and more stable training performance. Extensive experiments on simulations validate our design choices, and we show that our architecture learned on simulated data can achieve high performance in a real-robot setup involving a Franka Emika robot arm and unseen objects. Marco Ewerton, Michael Villamizar, Julius Jankowski, Sylvain Calinon, Jean-Marc Odobez |
IROS | 5 |
| 2022 | Robust Unsupervised Gaze Calibration Using Conversation and Manipulation Attention PriorsabstractGaze estimation is a difficult task, even for humans. However, as humans, we are good at understanding a situation and exploiting it to guess the expected visual focus of attention of people, and we usually use this information to retrieve people’s gaze. In this article, we propose to leverage such situation-based expectation about people’s visual focus of attention to collect weakly labeled gaze samples and perform person-specific calibration of gaze estimators in an unsupervised and online way. In this context, our contributions are the following: (i) we show how task contextual attention priors can be used to gather reference gaze samples, which is a cumbersome process otherwise; (ii) we propose a robust estimation framework to exploit these weak labels for the estimation of the calibration model parameters; and (iii) we demonstrate the applicability of this approach on two human-human and human-robot interaction settings, namely conversation and manipulation. Experiments on three datasets validate our approach, providing insights on the priors effectiveness and on the impact of different calibration models, particularly the usefulness of taking head pose into account. Rémy Siegfried, Jean-Marc Odobez |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Multi-Task Neural Network for Robust Multiple Speaker Embedding ExtractionabstractThis paper introduces a novel approach for extracting speaker embeddings from audio mixtures of multiple overlapping voices. This approach is based on a multi-task neural network. The network first extracts a latent feature for each direction. This feature is used for detecting sound sources as well as identifying speakers. In contrast to traditional approaches, the proposed method does not rely on explicit sound source separation. The neural network model learns from data to extract the most suitable features of the sounds at different directions. The experiments using audio recordings of overlapping sound sources show that the proposed approach outperforms a beamforming-based traditional method. Weipeng He, Petr Motlícek, Jean-Marc Odobez |
Interspeech | 3 |
| 2021 | An Efficient Image-to-Image Translation HourGlass-based Architecture for Object Pushing Policy LearningabstractHumans effortlessly solve pushing tasks in everyday life but unlocking these capabilities remains a challenge in robotics because physics models of these tasks are often inaccurate or unattainable. State-of-the-art data-driven approaches learn to compensate for these inaccuracies or replace the approximated physics models altogether. Nevertheless, approaches like Deep Q-Networks (DQNs) suffer from local optima in large state-action spaces. Furthermore, they rely on well-chosen deep learning architectures and learning paradigms. In this paper, we propose to frame the learning of pushing policies (where to push and how) by DQNs as an image-to-image translation problem and exploit an Hourglass-based architecture. We present an architecture combining a predictor of which pushes lead to changes in the environment with a state-action value predictor dedicated to the pushing task. Moreover, we investigate positional information encoding to learn position-dependent policy behaviors. We demonstrate in simulation experiments with a UR5 robot arm that our overall architecture helps the DQN learn faster and achieve higher performance in a pushing task involving objects with unknown dynamics. Marco Ewerton, Ángel Martínez-González, Jean-Marc Odobez |
IROS | 3 |
| 2021 | IEEE SLT 2021 Alpha-Mini Speech Challenge: Open Datasets, Tracks, Rules and BaselinesabstractThe IEEE Spoken Language Technology Workshop (SLT) 2021 Alpha-mini Speech Challenge (ASC) is intended to improve research on keyword spotting (KWS) and sound source location (SSL) on humanoid robots. Many publications report significant improvements in deep learning based KWS and SSL on open source datasets in recent years. For deep learning model training, it is necessary to expand the data coverage to improve the model robustness. Thus, simulating multi-channel noisy and reverberant data from single-channel speech, noise, echo and room impulsive response (RIR) is widely adopted. However, this approach may generate mismatch between simulated data and recorded data in real application scenarios, especially echo data. In this challenge, we open source a sizable speech, keyword, echo and noise corpus for promoting data-driven methods, particularly deep-learning approaches on KWS and SSL. We also choose Alpha-mini, a humanoid robot produced by UBTECH equipped with a built-in four-microphone array on its head, to record development and evaluation sets under the actual Alpha-mini robot application scenario, including environ-mental noise as well as echo and mechanical noise generated by the robot itself for model evaluation. Furthermore, we illustrate the rules, evaluation methods and baselines for re-searchers to quickly assess their achievements and optimize their models. Yihui Fu, Zhuoyuan Yao, Weipeng He, Jian Wu 0027, Zhanheng Yang, Lei Xie 0001, Dong-Yan Huang, Hui Bu, Petr Motlícek, Jean-Marc Odobez |
SLT | 12 |
| 2021 | A Differential Approach for Gaze EstimationabstractMost non-invasive gaze estimation methods regress gaze directions directly from a single face or eye image. However, due to important variabilities in eye shapes and inner eye structures amongst individuals, universal models obtain limited accuracies and their output usually exhibit high variance as well as subject dependent biases. Thus, increasing accuracy is usually done through calibration, allowing gaze predictions for a subject to be mapped to her actual gaze. In this article, we introduce a novel approach, which works by directly training a differential convolutional neural network to predict gaze differences between two eye input images of the same subject. Then, given a set of subject specific calibration images, we can use the inferred differences to predict the gaze direction of a novel eye sample. The assumption is that by comparing eye images of the same user, annoyance factors (alignment, eyelid closing, illumination perturbations) which usually plague single image prediction methods can be much reduced, allowing better prediction altogether. Furthermore, the differential network itself can be adapted via finetuning to make predictions consistent with the available user reference pairs. Experiments on 3 public datasets validate our approach which constantly outperforms state-of-the-art methods even when using only one calibration sample or those relying on subject specific gaze adaptation. Gang Liu 0013, Yu Yu 0003, Kenneth Alberto Funes Mora, Jean-Marc Odobez |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Neural Network Adaptation and Data Augmentation for Multi-Speaker Direction-of-Arrival EstimationabstractDeep neural networks have been successfully applied to sound direction-of-arrival estimation under challenging conditions. However, such a learning-based approach requires a large amount of labeled training data, which is difficult to acquire. To address this problem, we propose a novel approach for multi-speaker direction-of-arrival estimation with data augmentation and weakly-supervised domain adaptation. We generate source domain data with simulation, and collect real data annotated with the number of sound sources as the weak labels. The real data are further augmented by mixing single-source segments. Then, weakly-supervised domain adaptation is applied to models pre-trained on the simulated data. We define a loss function for the adaptation process which exploits the weak labels and the mixture component information in the augmented data. Experiments with real robot audio data show that our proposed approach achieves similar performance as if the fully-labeled real data are used. This paper suggests an effective development procedure for DOA estimation models applied to new types of microphone arrays with minimal data collection efforts. Weipeng He, Petr Motlícek, Jean-Marc Odobez |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Unsupervised Representation Learning for Gaze EstimationabstractAlthough automatic gaze estimation is very important to a large variety of application areas, it is difficult to train accurate and robust gaze models, in great part due to the difficulty in collecting large and diverse data (annotating 3D gaze is expensive and existing datasets use different setups). To address this issue, our main contribution in this paper is to propose an effective approach to learn a low dimensional gaze representation without gaze annotations, which to the best of our best knowledge, is the first work to do so. The main idea is to rely on a gaze redirection network and use the gaze representation difference of the input and target images (of the redirection network) as the redirection variable. A redirection loss in image domain allows the joint training of both the redirection network and the gaze representation network. In addition, we propose a warping field regularization which not only provides an explicit physical meaning to the gaze representations but also avoids redirection distortions. Promising results on few-shot gaze estimation (competitive results can be achieved with as few as ≤ 100 calibration samples), cross-dataset gaze estimation, gaze network pretraining, and another task (head pose estimation) demonstrate the validity of our framework. Yu Yu 0003, Jean-Marc Odobez |
CVPR | 2 |
| 2020 | Residual Pose: A Decoupled Approach for Depth-based 3D Human Pose EstimationabstractWe propose to leverage recent advances in reliable 2D pose estimation with Convolutional Neural Networks (CNN) to estimate the 3D pose of people from depth images in multi-person Human-Robot Interaction (HRI) scenarios. Our method is based on the observation that using the depth information to obtain 3D lifted points from 2D body landmark detections provides a rough estimate of the true 3D human pose, thus requiring only a refinement step. In that line our contributions are threefold. (i) we propose to perform 3D pose estimation from depth images by decoupling 2D pose estimation and 3D pose refinement; (ii) we propose a deep-learning approach that regresses the residual pose between the lifted 3D pose and the true 3D pose; (iii) we show that despite its simplicity, our approach achieves very competitive results both in accuracy and speed on two public datasets and is therefore appealing for multi-person HRI compared to recent state-of-the-art methods. Ángel Martínez-González, Michael Villamizar, Olivier Canévet, Jean-Marc Odobez |
IROS | 4 |
| 2020 | The MuMMER Data Set for Robot Perception in Multi-party HRI ScenariosabstractThis paper presents the MuMMER data set, a data set for human-robot interaction scenarios that is available for research purposes1. It comprises 1h 29 min of multimodal recordings of people interacting with the social robot Pepper in entertainment scenarios, such as quiz, chat, and route guidance. In the 33 clips (of 1 to 4 min long) recorded from the robot point of view, the participants are interacting with the robot in an unconstrained manner.The data set exhibits interesting features and difficulties, such as people leaving the field of view, robot moving (head rotation with embedded camera in the head), different illumination conditions. The data set contains color and depth videos from a Kinect v2, an Intel D435, and the video from Pepper.All the visual faces and the identities in the data set were manually annotated, making the identities consistent across time and clips. The goal of the data set is to evaluate perception algorithms in multi-party human/robot interaction, in particular the re-identification part when a track is lost, as this ability is crucial for keeping the dialog history. The data set can easily be extended with other types of annotations.We also present a benchmark on this data set that should serve as a baseline for future comparison. The baseline system, IHPER2(Idiap Human Perception system) is available for research and is evaluated on the MuMMER data set. We show that an identity precision and recall of ~80% and a MOTA score above 80% are obtained. Olivier Canévet, Weipeng He, Petr Motlícek, Jean-Marc Odobez |
RO-MAN | 4 |
| 2020 | WatchNet++: efficient and accurate depth-based network for detecting people attacks and intrusion
Michael Villamizar, Ángel Martínez-González, Olivier Canévet, Jean-Marc Odobez |
Mach. Vis. Appl. | 4 |
| 2020 | Multi-scale sequential network for semantic text segmentation and localization
Michael Villamizar, Olivier Canévet, Jean-Marc Odobez |
Pattern Recognit. Lett. | 3 |
| 2020 | Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose EstimationabstractAchieving robust multi-person 2D body landmark localization and pose estimation is essential for human behavior and interaction understanding as encountered for instance in HRI settings. Accurate methods have been proposed recently, but they usually rely on rather deep Convolutional Neural Network (CNN) architecture, thus requiring large computational and training resources. In this paper, we investigate different architectures and methodologies to address these issues and achieve fast and accurate multi-person 2D pose estimation. To foster speed, we propose to work with depth images, whose structure contains sufficient information about body landmarks while being simpler than textured color images and thus potentially requiring less complex CNNs for processing. In this context, we make the following contributions. i) we study several CNN architecture designs combining pose machines relying on the cascade of detectors concept with lightweight and efficient CNN structures; ii) to address the need for large training datasets with high variability, we rely on semi-synthetic data combining multi-person synthetic depth data with real sensor backgrounds; iii) we explore domain adaptation techniques to address the performance gap introduced by testing on real depth images; iv) to increase the accuracy of our fast lightweight CNN models, we investigate knowledge distillation at several architecture levels which effectively enhance performance. Experiments and results on synthetic and real data highlight the impact of our design choices, providing insights into methods addressing standard issues normally faced in practical applications, and resulting in architectures effectively matching our goal in both performance and speed. Ángel Martínez-González, Michael Villamizar, Olivier Canévet, Jean-Marc Odobez |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Improving Few-Shot User-Specific Gaze Adaptation via Gaze Redirection SynthesisabstractAs an indicator of human attention gaze is a subtle behavioral cue which can be exploited in many applications. However, inferring 3D gaze direction is challenging even for deep neural networks given the lack of large amount of data (groundtruthing gaze is expensive and existing datasets use different setups) and the inherent presence of gaze biases due to person-specific difference. In this work, we address the problem of person-specific gaze model adaptation from only a few reference training samples. The main and novel idea is to improve gaze adaptation by generating additional training samples through the synthesis of gaze-redirected eye images from existing reference samples. In doing so, our contributions are threefold:(i) we design our gaze redirection framework from synthetic data, allowing us to benefit from aligned training sample pairs to predict accurate inverse mapping fields; (ii) we proposed a self-supervised approach for domain adaptation; (iii) we exploit the gaze redirection to improve the performance of person-specific gaze estimation. Extensive experiments on two public datasets demonstrate the validity of our gaze retargeting and gaze estimation framework. Yu Yu 0003, Gang Liu 0013, Jean-Marc Odobez |
CVPR | 3 |
| 2019 | A deep learning approach for robust head pose independent eye movements recognition from videosabstractRecognizing eye movements is important for gaze behavior understanding like in human communication analysis (human-human or robot interactions) or for diagnosis (medical, reading impairments). In this paper, we address this task using remote RGB-D sensors to analyze people behaving in natural conditions. This is very challenging given that such sensors have a normal sampling rate of 30 Hz and provide low-resolution eye images (typically 36×60 pixels), and natural scenarios introduce many variabilities in illumination, shadows, head pose, and dynamics. Hence gaze signals one can extract in these conditions have lower precision compared to dedicated IR eye trackers, rendering previous methods less appropriate for the task. To tackle these challenges, we propose a deep learning method that directly processes the eye image video streams to classify them into fixation, saccade, and blink classes, and allows to distinguish irrelevant noise (illumination, low-resolution artifact, inaccurate eye alignment, difficult eye shapes) from true eye motion signals. Experiments on natural 4-party interactions demonstrate the benefit of our approach compared to previous methods, including deep learning models applied to gaze outputs. Rémy Siegfried, Yu Yu 0003, Jean-Marc Odobez |
ETRA | 3 |
| 2019 | Adaptation of Multiple Sound Source Localization Neural Networks with Weak Supervision and Domain-adversarial TrainingabstractDespite the recent success of deep neural network-based approaches in sound source localization, these approaches suffer the limitations that the required annotation process is costly, and the mismatch between the training and test conditions undermines the performance. This paper addresses the question of how models trained with simulation can be exploited for multiple sound source localization in real scenarios by domain adaptation. In particular, two domain adaptation methods are investigated: weak supervision and domain-adversarial training. Our experiments show that the weak supervision with the knowledge of the number of sources can significantly improve the performance of an unadapted model. However, the domain-adversarial training does not yield significant improvement for this particular problem. Weipeng He, Petr Motlícek, Jean-Marc Odobez |
ICASSP | 3 |
| 2019 | Improving speech embedding using crossmodal transfer learning with audio-visual dataabstractLearning a discriminative voice embedding allows speaker turns to be compared directly and efficiently, which is crucial for tasks such as diarization and verification. This paper investigates several transfer learning approaches to improve a voice embedding using knowledge transferred from a face representation. The main idea of our crossmodal approaches is to constrain the target voice embedding space to share latent attributes with the source face embedding space. The shared latent attributes can be formalized as geometric properties or distribution characterics between these embedding spaces. We propose four transfer learning approaches belonging to two categories: the first category relies on the structure of the source face embedding space to regularize at different granularities the speaker turn embedding space. The second category -a domain adaptation approach- improves the embedding space of speaker turns by applying a maximum mean discrepancy loss to minimize the disparity between the distributions of the embedded features. Experiments are conducted on TV news datasets, REPERE and ETAPE, to demonstrate our methods. Quantitative results in verification and clustering tasks show promising improvement, especially in cases where speaker turns are short or the training data size is limited. The analysis also gives insights the embedding spaces and shows their potential applications. Nam Le 0001, Jean-Marc Odobez |
Multim. Tools Appl. | 2 |
| 2018 | UNICITY: A depth maps database for people detection in security airlocksabstractWe introduce a new dataset, dubbed UNICITY1, for the task of detecting people in security airlocks in top view depth images. If security companies have been relying on computer systems and algorithms for a long time, very few are trusting artificial intelligence and more specifically machine learning approaches in production environments. We are confident that the recent advances in these domains, especially with the democratization of deep learning, will open new horizons for security systems. We release this dataset to encourage the development of such approaches in the scientific community.UNICITY consists of 58k images collected from 65 recorded sequences with one or two people performing different behaviors including attacks and trickeries (e.g. tailgating2). It also provides full annotation of people such as the location of head and shoulders. As as result, UNICITY is perfectly suited for training and adapting machine learning algorithms for video surveillance applications. This paper presents the data collection, an evaluation protocol, as well as two baseline methods for attack detection. Joël Dumoulin, Olivier Canévet, Michael Villamizar, Hugo Nunes, Omar Abou Khaled, Elena Mugellini, Fabrice Moscheni, Jean-Marc Odobez |
AVSS | 8 |
| 2018 | WatchNet: Efficient and Depth-based Network for People Detection in Video Surveillance SystemsabstractWe propose a deep-learning approach for people detection on depth imagery. The approach is designed to be deployed as an autonomous appliance for identifying people attacks and intrusion in video surveillance scenarios. To this end, we propose a fully-convolutional and sequential network, named WatchNet, that localizes people in depth images by predicting human body landmarks such as head and shoulders. We use a large synthetic dataset to train the network with abundant data and generate automatic annotations. Adaptation to real data is performed via fine tuning with real depth images.The proposed method is validated in a novel and challenging database with about 29k top view images collected from several sequences including different people assaults. A comparative evaluation is given between our approach and other standard methods, showing remarkable detection results and efficiency. The network runs in 10 and 28 FPS using CPU and GPU, respectively. Michael Villamizar, Ángel Martínez-González, Olivier Canévet, Jean-Marc Odobez |
AVSS | 4 |
| 2018 | A Differential Approach for Gaze Estimation with Calibration
Gang Liu 0013, Yu Yu 0003, Kenneth Alberto Funes Mora, Jean-Marc Odobez |
BMVC | 4 |
| 2018 | Deep Neural Networks for Multiple Speaker Detection and LocalizationabstractWe propose to use neural networks for simultaneous detection and localization of multiple sound sources in human-robot interaction. In contrast to conventional signal processing techniques, neural network-based sound source localization methods require fewer strong assumptions about the environment. Previous neural network-based methods have been focusing on localizing a single sound source, which do not extend to multiple sources in terms of detection and localization. In this paper, we thus propose a likelihood-based encoding of the network output, which naturally allows the detection of an arbitrary number of sources. In addition, we investigate the use of sub-band cross-correlation information as features for better localization in sound mixtures, as well as three different network architectures based on different motivations. Experiments on real data recorded from a robot show that our proposed methods significantly outperform the popular spatial spectrum-based approaches. Weipeng He, Petr Motlícek, Jean-Marc Odobez |
ICRA | 3 |
| 2018 | Joint Localization and Classification of Multiple Sound Sources Using a Multi-task Neural NetworkabstractWe propose a novel multi-task neural network-based approach for joint sound source localization and speech/non-speech classification in noisy environments. The network takes raw short time Fourier transform as input and outputs the likelihood values for the two tasks, which are used for the simultaneous detection, localization and classification of an unknown number of overlapping sound sources, Tested with real recorded data, our method achieves significantly better performance in terms of speech/non-speech classification and localization of speech sources, compared to method that performs localization and classification separately. In addition, we demonstrate that incorporating the temporal context can further improve the performance. Weipeng He, Petr Motlícek, Jean-Marc Odobez |
INTERSPEECH | 3 |
| 2018 | Robust and Discriminative Speaker Embedding via Intra-Class Distance Variance RegularizationabstractLearning a good speaker embedding is critical for many speech processing tasks, including recognition, verification, and diarization. To this end, we propose a complementary optimizing goal called intra-class loss to improve deep speaker embed dings learned with triplet loss. This loss function is formulated as a soft constraint on the averaged pair-wise distance between samples from the same class. Its goal is to prevent the scattering of these samples within the embedding space to increase the intra-class compactncss.When intra-class loss is jointly optimized with triplet loss, we can observe 2 major improvements: the deep embedding network can achieve a more robust and discriminative representation and the training process is more stable with a faster convergence rate. We conduct experiments on 2 large public benchmarking datasets for speaker verification, VoxCeleb and VoxForge. The results show that intra-class loss helps accelerating the convergence of deep network training and significantly improves the overall performance of the resulted embeddings. Nam Le 0001, Jean-Marc Odobez |
INTERSPEECH | 2 |
| 2018 | Leveraging Convolutional Pose Machines for Fast and Accurate Head Pose EstimationabstractWe propose a head pose estimation framework that leverages on a recent keypoint detection model. More specifically, we apply the convolutional pose machines (CPMs) to input images, extract different types of facial keypoint features capturing appearance information and keypoint relationships, and train multilayer perceptrons (MLPs) and convolutional neural networks (CNNs) for head pose estimation. The benefit of leveraging on the CPMs (which we apply anyway for other purposes like tracking) is that we can design highly efficient models for practical usage. We evaluate our approach on the Annotated Facial Landmarks in the Wild (AFLW) dataset and achieve competitive results with the state-of-the-art. Yuanzhouhan Cao, Olivier Canévet, Jean-Marc Odobez |
IROS | 3 |
| 2018 | Real-time Convolutional Networks for Depth-based Human Pose EstimationabstractWe propose to combine recent Convolutional Neural Networks (CNN) models with depth imaging to obtain a reliable and fast multi-person pose estimation algorithm applicable to Human Robot Interaction (HRI) scenarios. Our hypothesis is that depth images contain less structures and are easier to process than RGB images while keeping the required information for human detection and pose inference, thus allowing the use of simpler networks for the task. Our contributions are threefold. (i) we propose a fast and efficient network based on residual blocks (called RPM) for body landmark localization from depth images; (ii) we created a public dataset DIH comprising more than 170k synthetic images of human bodies with various shapes and viewpoints as well as real (annotated) data for evaluation; (iii) we show that our model trained on synthetic data from scratch can perform well on real data, obtaining similar results to larger models initialized with pre-trained networks. It thus provides a good trade-off between performance and computation. Experiments on real data demonstrate the validity of our approach. Ángel Martínez-González, Michael Villamizar, Olivier Canévet, Jean-Marc Odobez |
IROS | 4 |
| 2018 | Facing Employers and Customers: What Do Gaze and Expressions Tell About Soft Skills?abstractEye gaze and facial expressions are central to face-to-face social interactions. These behavioral cues and their connections to first impressions have been widely studied in psychology and computing literature, but limited to a single situation. Utilizing ubiquitous multimodal sensors coupled with advances in computer vision and machine learning, we investigate the connections between these behavioral cues and perceived soft skills in two diverse workplace situations (job interviews and reception desk). Pearson's correlation analysis shows a moderate connection between certain facial expressions, eye gaze cues and perceived soft skills in job interviews (r ϵ [-30,30]) and desk (r ϵ [20,36]) situations. Results of our computational framework to infer perceived soft skills indicates a low predictive power of eye gaze, facial expressions, and their combination in both interviews (R2 ϵ [0.02,0.21]) and desk (R2 ϵ [0.05,0.15]) situations. Our work has important implications for employee training and behavioral feedback systems. Skanda Muralidhar, Rémy Siegfried, Jean-Marc Odobez, Daniel Gatica-Perez |
MUM | 3 |
| 2018 | HeadFusion: 360° Head Pose Tracking Combining 3D Morphable Model and 3D ReconstructionabstractHead pose estimation is a fundamental task for face and social related research. Although 3D morphable model (3DMM) based methods relying on depth information usually achieve accurate results, they usually require frontal or mid-profile poses which preclude a large set of applications where such conditions can not be garanteed, like monitoring natural interactions from fixed sensors placed in the environment. A major reason is that 3DMM models usually only cover the face region. In this paper, we present a framework which combines the strengths of a 3DMM model fitted online with a prior-free reconstruction of a 3D full head model providing support for pose estimation from any viewpoint. In addition, we also proposes a symmetry regularizer for accurate 3DMM fitting under partial observations, and exploit visual tracking to address natural head dynamics with fast accelerations. Extensive experiments show that our method achieves state-of-the-art performance on the public BIWI dataset, as well as accurate and robust results on UbiPose, an annotated dataset of natural interactions that we make public and where adverse poses, occlusions or fast motions regularly occur. Yu Yu 0003, Kenneth Alberto Funes Mora, Jean-Marc Odobez |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Maya Codical Glyph Segmentation: A Crowdsourcing ApproachabstractThis paper focuses on the crowd-annotation of an ancient Maya glyph dataset derived from the three ancient codices that survived up to date. More precisely, nonexpert annotators are asked to segment glyph-blocks into their constituent glyph entities. As a means of supervision, available glyph variants are provided to the annotators during the crowdsourcing task. Compared to object recognition in natural images or handwriting transcription tasks, designing an engaging task and dealing with crowd behavior is challenging in our case. This challenge originates from the inherent complexity of Maya writing and an incomplete understanding of the signs and semantics in the existing catalogs. We elaborate on the evolution of the crowdsourcing task design, and discuss the choices for providing supervision during the task. We analyze the distributions of similarity and task difficulty scores, and the segmentation performance of the crowd. A unique dataset of over 9000 Maya glyphs from 291 categories individually segmented from the three codices was created and will be made publicly available thanks to this process. This dataset lends itself to automatic glyph classification tasks. We provide baseline methods for glyph classification using traditional shape descriptors and convolutional neural networks. Gulcan Can, Jean-Marc Odobez, Daniel Gatica-Perez |
IEEE Trans. Multim. | 2 |
| 2017 | Robust and Accurate 3D Head Pose Estimation through 3DMM and Online Head Model ReconstructionabstractAccurate and robust 3D head pose estimation is important for face related analysis. Though high accuracy has been achieved by previous works based on 3D morphable model (3DMM), their performance drops with extreme head poses because such models usually only represent the frontal face region. In this paper, we present a robust head pose estimation framework by complementing a 3DMM model with an online 3D reconstruction of the full head providing more support when handling extreme head poses. The approach includes a robust online 3DMM fitting step based on multi-view observation samples as well as smooth and face-neutral synthetic samples generated from the reconstructed 3D head model. Experiments show that our framework achieves state-of-the-art pose estimation accuracy on the BIWI dataset, and has robust performance for extreme head poses when tested on natural interaction sequences. Yu Yu 0003, Kenneth Alberto Funes Mora, Jean-Marc Odobez |
FG | 3 |
| 2017 | A domain adaptation approach to improve speaker turn embedding using face representationabstractThis paper proposes a novel approach to improve speaker modeling using knowledge transferred from face representation. In particular, we are interested in learning a discriminative metric which allows speaker turns to be compared directly, which is beneficial for tasks such as diarization and dialogue analysis. Our method improves the embedding space of speaker turns by applying maximum mean discrepancy loss to minimize the disparity between the distributions of facial and acoustic embedded features. This approach aims to discover the shared underlying structure of the two embedded spaces, thus enabling the transfer of knowledge from the richer face representation to the counterpart in speech. Experiments are conducted on broadcast TV news datasets, REPERE and ETAPE, to demonstrate the validity of our method. Quantitative results in verification and clustering tasks show promising improvement, especially in cases where speaker turns are short or the training data size is limited. Nam Le 0001, Jean-Marc Odobez |
ICMI | 2 |
| 2017 | Towards the use of social interaction conventions as prior for gaze model adaptationabstractGaze is an important non-verbal cue involved in many facets of social interactions like communication, attentiveness or attitudes. Nevertheless, extracting gaze directions visually and remotely usually suffers large errors because of low resolution images, inaccurate eye cropping, or large eye shape variations across the population, amongst others. This paper hypothesizes that these challenges can be addressed by exploiting multimodal social cues for gaze model adaptation on top of an head-pose independent 3D gaze estimation framework. First, a robust eye cropping refinement is achieved by combining a semantic face model with eye landmark detections. Investigations on whether temporal smoothing can overcome instantaneous refinement limitations is conducted. Secondly, to study whether social interaction convention could be used as priors for adaptation, we exploited the speaking status and head pose constraints to derive soft gaze labels and infer person-specific gaze bias using robust statistics. Experimental results on gaze coding in natural interactions from two different settings demonstrate that the two steps of our gaze adaptation method contribute to reduce gaze errors by a large margin over the baseline and can be generalized to several identities in challenging scenarios. Rémy Siegfried, Yu Yu 0003, Jean-Marc Odobez |
ICMI | 3 |
| 2017 | Active Online Anomaly Detection Using Dirichlet Process Mixture Model and Gaussian Process ClassificationabstractWe present a novel anomaly detection (AD) system for streaming videos. Different from prior methods that rely on unsupervised learning of clip representations, that are usually coarse in nature, and batch-mode learning, we propose the combination of two non-parametric models for our task: (i) Dirichlet process mixture models (DPMM) based modeling of object motion and directions in each cell, and (ii) Gaussian process based active learning paradigm involving labeling by a domain expert. Whereas conventional clip representation methods adopt quantizing only motion directions leading to a lossy, coarse representation that are inadequate, our clip representation approach results in fine grained clusters at each cell that model the scene activities (both direction and speed) more effectively. For active anomaly detection, we adapt a Gaussian Process framework to process incoming samples (video snippets) sequentially, seek labels for confusing or informative samples and and update the AD model online. Furthermore, the proposed video representation along with a novel query criterion to select informative samples for labeling that incorporates both exploration and exploitation criteria is proposed, and is found to outperform competing criteria on two challenging traffic scene datasets. Jagannadan Varadarajan, Subramanian Ramanathan, Narendra Ahuja, Pierre Moulin, Jean-Marc Odobez |
WACV | 5 |
| 2016 | Training on the job: behavioral analysis of job interviews in hospitalityabstractFirst impressions play a critical role in the hospitality industry and have been shown to be closely linked to the behavior of the person being judged.In this work, we implemented a behavioral training framework for hospitality students with the goal of improving the impressions that other people make about them. We outline the challenges associated with designing such a framework and embedding it in the everyday practice of a real hospitality school. We collected a dataset of 169 laboratory sessions where two role-plays were conducted, job interviews and reception desk scenarios, for a total of 338 interactions. For job interviews, we evaluated the relationship between automatically extracted nonverbal cues and various perceived social variables in a correlation analysis. Furthermore, our system automatically predicted first impressions from job interviews in a regression task, and was able to explain up to 32% of the variance, thus extending the results in existing literature, and showing gender differences, corroborating previous findings in psychology. This work constitutes a step towards applying social sensing technologies to the real world by designing and implementing a living lab for students of an international hospitality management school. Skanda Muralidhar, Laurent Son Nguyen, Denise Frauendorfer, Jean-Marc Odobez, Marianne Schmid Mast, Daniel Gatica-Perez |
ICMI | 4 |
| 2016 | Towards building an attentive artificial listener: on the perception of attentiveness in audio-visual feedback tokensabstractCurrent dialogue systems typically lack a variation of audio-visual feedback tokens. Either they do not encompass feedback tokens at all, or only support a limited set of stereotypical functions. However, this does not mirror the subtleties of spontaneous conversations. If we want to be able to build an artificial listener, as a first step towards building an empathetic artificial agent, we also need to be able to synthesize more subtle audio-visual feedback tokens. In this study, we devised an array of monomodal and multimodal binary comparison perception tests and experiments to understand how different realisations of verbal and visual feedback tokens influence third-party perception of the degree of attentiveness. This allowed us to investigate i) which features (amplitude, frequency, duration...) of the visual feedback influences attentiveness perception; ii) whether visual or verbal backchannels are perceived to be more attentive iii) whether the fusion of unimodal tokens with low perceived attentiveness increases the degree of perceived attentiveness compared to unimodal tokens with high perceived attentiveness taken alone; iv) the automatic ranking of audio-visual feedback token in terms of conveyed degree of attentiveness. Catharine Oertel, José Lopes 0001, Yu Yu 0003, Kenneth Alberto Funes Mora, Joakim Gustafson, Alan W. Black, Jean-Marc Odobez |
ICMI | 7 |
| 2016 | Temporally subsampled detection for accurate and efficient face tracking and diarizationabstractFace diarization, i.e. face tracking and clustering within video documents, is useful and important for video indexing and fast browsing but it is also a difficult and time consuming task. In this paper, we address the tracking aspect and propose a novel algorithm with two main contributions. First, we propose an approach that leverages state-of-the-art deformable part-based model (DPM) face detector with a multi-cue discriminant tracking-by-detection framework that relies on automatically learned long-term time-interval sensitive association costs specific to each document type. Secondly to improve performance, we propose an explicit false alarm removal step at the track level to efficiently filter out wrong detections (and resulting tracks). Altogether, the method is able to skip frames, i.e. process only 3 to 4 frames per second - thus cutting down computational cost - while performing better than state-of-the-art methods as evaluated on three public benchmarks from different context including a movie and broadcast data. Nam Le 0001, Alexandre Heili, Di Wu 0009, Jean-Marc Odobez |
ICPR | 4 |
| 2016 | Learning Multimodal Temporal Representation for Dubbing Detection in Broadcast MediaabstractPerson discovery in the absence of prior identity knowledge requires accurate association of visual and auditory cues. In broadcast data, multimodal analysis faces additional challenges due to narrated voices over muted scenes or dubbing in different languages. To address these challenges, we define and analyze the problem of dubbing detection in broadcast data, which has not been explored before. We propose a method to represent the temporal relationship between the auditory and visual streams. This method consists of canonical correlation analysis to learn a joint multimodal space, and long short term memory (LSTM) networks to model cross-modality temporal dependencies. Our contributions also include the introduction of a newly acquired dataset of face-speech segments from TV data, which we have made publicly available. The proposed method achieves promising performance on this real world dataset as compared to several baselines. Nam Do-Hoang Le, Jean-Marc Odobez |
ACM Multimedia | 2 |
| 2016 | Gaze Estimation in the 3D Space Using RGB-D Sensors - Towards Head-Pose and User Invariance
Kenneth Alberto Funes Mora, Jean-Marc Odobez |
Int. J. Comput. Vis. | 2 |
| 2016 | Deep Dynamic Neural Networks for Multimodal Gesture Segmentation and RecognitionabstractThis paper describes a novel method called Deep Dynamic Neural Networks (DDNN) for multimodal gesture recognition. A semi-supervised hierarchical dynamic framework based on a Hidden Markov Model (HMM) is proposed for simultaneous gesture segmentation and recognition where skeleton joint information, depth and RGB images, are the multimodal input observations. Unlike most traditional approaches that rely on the construction of complex handcrafted features, our approach learns high-level spatio-temporal representations using deep neural networks suited to the input modality: a Gaussian-Bernouilli Deep Belief Network (DBN) to handle skeletal dynamics, and a 3D Convolutional Neural Network (3DCNN) to manage and fuse batches of depth and RGB images. This is achieved through the modeling and learning of the emission probabilities of the HMM required to infer the gesture sequence. This purely data driven approach achieves a Jaccard index score of 0.81 in the ChaLearn LAP gesture spotting challenge. The performance is on par with a variety of state-of-the-art hand-tuned feature-based approaches and other learning-based methods, therefore opening the door to the use of deep learning techniques in order to further explore multimodal time series data. Di Wu 0009, Lionel Pigou, Pieter-Jan Kindermans, Nam Do-Hoang Le, Ling Shao 0001, Joni Dambre, Jean-Marc Odobez |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2015 | Deciphering the Silent Participant: On the Use of Audio-Visual Cues for the Classification of Listener Categories in Group DiscussionsabstractEstimating a silent participant's degree of engagement and his role within a group discussion can be challenging, as there are no speech related cues available at the given time. Having this information available, however, can provide important insights into the dynamics of the group as a whole. In this paper, we study the classification of listeners into several categories (attentive listener, side participant and bystander). We devised a thin-sliced perception test where subjects were asked to assess listener roles and engagement levels in 15-second video-clips taken from a corpus of group interviews. Results show that humans are usually able to assess silent participant roles. Using the annotation to identify from a set of multimodal low-level features, such as past speaking activity, backchannels (both visual and verbal), as well as gaze patterns, we could identify the features which are able to distinguish between different listener categories. Moreover, the results show that many of the audio-visual effects observed on listeners in dyadic interactions, also hold for multi-party interactions. A preliminary classifier achieves an accuracy of 64 %. Catharine Oertel, Kenneth Alberto Funes Mora, Joakim Gustafson, Jean-Marc Odobez |
ICMI | 4 |
| 2015 | Combining dynamic head pose-gaze mapping with the robot conversational state for attention recognition in human-robot interactions
Samira Sheikhi, Jean-Marc Odobez |
Pattern Recognit. Lett. | 2 |
| 2014 | Geometric Generative Gaze Estimation (G3E) for Remote RGB-D CamerasabstractWe propose a head pose invariant gaze estimation model for distant RGB-D cameras. It relies on a geometric understanding of the 3D gaze action and generation of eye images. By introducing a semantic segmentation of the eye region within a generative process, the model (i) avoids the critical feature tracking of geometrical approaches requiring high resolution images, (ii) decouples the person dependent geometry from the ambient conditions, allowing adaptation to different conditions without retraining. Priors in the generative framework are adequate for training from few samples. In addition, the model is capable of gaze extrapolation allowing for less restrictive training schemes. Comparisons with state of the art methods validate these properties which make our method highly valuable for addressing many diverse tasks in sociology, HRI and HCI. Kenneth Alberto Funes Mora, Jean-Marc Odobez |
CVPR | 2 |
| 2014 | EYEDIAP: a database for the development and evaluation of gaze estimation algorithms from RGB and RGB-D camerasabstractThe lack of a common benchmark for the evaluation of the gaze estimation task from RGB and RGB-D data is a serious limitation for distinguishing the advantages and disadvantages of the many proposed algorithms found in the literature. This paper intends to overcome this limitation by introducing a novel database along with a common framework for the training and evaluation of gaze estimation approaches. In particular, we have designed this database to enable the evaluation of the robustness of algorithms with respect to the main challenges associated to this task: i) Head pose variations; ii) Person variation; iii) Changes in ambient and sensing conditions and iv) Types of target: screen or 3D object. Kenneth Alberto Funes Mora, Florent Monay, Jean-Marc Odobez |
ETRA | 3 |
| 2014 | A conditional random field approach for audio-visual people diarizationabstractWe investigate the problem of audio-visual (AV) person diarization in broadcast data. That is, automatically associate the faces and voices of people and determine when they appear or speak in the video. The contributions are twofolds. First, we formulate the problem within a novel CRF framework that simultaneously performs the AV association of voices and face clusters to build AV person models, and the joint segmentation of the audio and visual streams using a set of AV cues and their association strength. Secondly, we use for this AV association strength a score that does not only rely on lips activity, but also on contextual visual information (face size, position, number of detected faces,...) that leads to more reliable association measures. Experiments on 6 hours of broadcast data show that our framework is able to improve the AV-person diarization especially for speaker segments erroneously labeled in the mono-modal case. Paul Gay, Elie Khoury 0001, Sylvain Meignier, Jean-Marc Odobez, Paul Deléglise |
ICASSP | 4 |
| 2014 | A conditional random field approach for face identification in broadcast news using overlaid textabstractWe investigate the problem of face identification in broadcast programs where people names are obtained from text overlays automatically processed with Optical Character Recognition (OCR) and further linked to the faces throughout the video. To solve the face-name association and propagation, we propose a novel approach that combines the positive effects of two Conditional Random Field (CRF) models: a CRF for person diarization (joint temporal segmentation and association of voices and faces) that benefit from the combination of multiple cues including as main contributions the use of identification sources (OCR appearances) and recurrent local face visual background (LFB) playing the role of a namedness feature; a second CRF for the joint identification of the person clusters that improves identification performance thanks to the use of further diarization statistics. Experiments conducted on a recent and substantial public dataset of 7 different shows demonstrate the interest and complementarity of the different modeling steps and information sources, leading to state of the art results. Paul Gay, Elie Khoury 0001, Sylvain Meignier, Jean-Marc Odobez, Paul Deléglise |
ICIP | 4 |
| 2014 | Improving head and body pose estimation through semi-supervised manifold alignmentabstractIn this paper, we explore the use of a semi-supervised manifold alignment method for domain adaptation in the context of human body and head pose estimation in videos. We build upon an existing state-of-the-art system that leverages on external labelled datasets for the body and head features, and on the unlabelled test data with weak velocity labels to do a coupled estimation of the body and head pose. While this previous approach showed promising results, the learning of the underlying manifold structure of the features in the train and target data and the need to align them were not explored despite the fact that the pose features between two datasets may vary according to the scene, e.g. due to different camera point of view or perspective. In this paper, we propose to use a semi-supervised manifold alignment method to bring the train and target samples closer within the resulting embedded space. To this end, we consider an adaptation set from the target data and rely on (weak) labels, given for example by the velocity direction whenever they are reliable. These labels, along with the training labels are used to bias the manifold distance within each manifold and to establish correspondences for alignment. Alexandre Heili, Jagannadan Varadarajan, Bernard Ghanem, Narendra Ahuja, Jean-Marc Odobez |
ICIP | 5 |
| 2014 | Automated bobbing and phase analysis to measure walking entrainment to musicabstractIn this paper, we investigate the influence of music on human walking behaviors in a public setting monitored by surveillance cameras. To this end, we propose a novel algorithm to characterize the frequency and phase of the walk. It relies on a human-by-detection tracking framework, along with a robust fitting of the human head bobbing motion. Preliminary experiments conducted on more than 100 tracks show that an accuracy greater than 85% for foot strike estimation can be achieved, suggesting that large scale analysis is at reach for finer music/walking behavior relationship studies. Adolfo López, Carina Westling, Rémi Emonet, M. Easteal, L. Lavia, Harry J. Witchel, Jean-Marc Odobez |
ICIP | 7 |
| 2014 | Automatic Maya hieroglyph retrieval using shape and context informationabstractWe propose an automatic Maya hieroglyph retrieval method integrating shape and glyph context information. Two recent local shape descriptors, Gradient Field Histogram of Orientation Gradient (GF-HOG) and Histogram of Orientation Shape Context (HOOSC), are evaluated. To encode the context information, we propose to convert each Maya glyph block into a first-order Markov chain and apply the co-occurrence of neighbouring glyphs. The retrieval results obtained based on visual matching are therefore re-ranked. Experimental results show that our method can significantly improve the glyph retrieval accuracy even with a basic co-occurrence model. Furthermore, two unique glyph datasets are contributed which can be used as novel shape benchmarks in future research. Rui Hu 0010, Carlos Pallan, Guido Krempel, Jean-Marc Odobez, Daniel Gatica-Perez |
ACM Multimedia | 4 |
| 2014 | Temporal Analysis of Motif Mixtures Using Dirichlet ProcessesabstractIn this paper, we present a new model for unsupervised discovery of recurrent temporal patterns (or motifs) in time series (or documents). The model is designed to handle the difficult case of multivariate time series obtained from a mixture of activities, that is, our observations are caused by the superposition of multiple phenomena occurring concurrently and with no synchronization. The model uses nonparametric Bayesian methods to describe both the motifs and their occurrences in documents. We derive an inference scheme to automatically and simultaneously recover the recurrent motifs (both their characteristics and number) and their occurrence instants in each document. The model is widely applicable and is illustrated on datasets coming from multiple modalities, mainly videos from static cameras and audio localization data. The rich semantic interpretation that the model offers can be leveraged in tasks such as event counting or for scene analysis. The approach is also used as a mean of doing soft camera calibration in a camera network. A thorough study of the model parameters is provided and a cross-platform implementation of the inference algorithm will be made publicly available. Rémi Emonet, Jagannadan Varadarajan, Jean-Marc Odobez |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Leveraging colour segmentation for upper-body detection
Stefan Duffner, Jean-Marc Odobez |
Pattern Recognit. | 2 |
| 2014 | Exploiting Long-Term Connectivity and Visual Motion in CRF-Based Multi-Person TrackingabstractWe present a conditional random field approach to tracking-by-detection in which we model pairwise factors linking pairs of detections and their hidden labels, as well as higher order potentials defined in terms of label costs. To the contrary of previous papers, our method considers long-term connectivity between pairs of detections and models similarities as well as dissimilarities between them, based on position, color, and as novelty, visual motion cues. We introduce a set of feature-specific confidence scores, which aim at weighting feature contributions according to their reliability. Pairwise potential parameters are then learned in an unsupervised way from detections or from tracklets. Label costs are defined so as to penalize the complexity of the labeling, based on prior knowledge about the scene like the location of entry/exit zones. Experiments on PETS’09, TUD, CAVIAR, Parking Lot, and Town Center public data sets show the validity of our approach, and similar or better performance than recent state-of-the-art algorithms. Alexandre Heili, Adolfo López, Jean-Marc Odobez |
IEEE Trans. Image Process. | 3 |
| 2013 | Localized anomaly detection via hierarchical integrated activity discoveryabstractWith the increasing number and variety of camera installations, unsupervised methods that learn typical activities have become popular for anomaly detection. In this article, we consider recent methods based on temporal probabilistic models and improve them in multiple ways. Our contributions are the following: (i) we integrate the low level processing and the temporal activity modeling, showing how this feedback improves the overall quality of the captured information, (ii) we show how the same approach can be taken to do hierarchical multi-camera processing, (iii) we use spatial analysis of the anomalies both to perform local anomaly detection and to frame automatically the detected anomalies. We illustrate the approach on both traffic data and videos coming from a metro station. Thiyagarajan Chockalingam, Rémi Emonet, Jean-Marc Odobez |
AVSS | 3 |
| 2013 | Given that, should i respond?: contextual addressee estimation in multi-party human-robot interactions
Dinesh Babu Jayagopi, Jean-Marc Odobez |
HRI | 2 |
| 2013 | The vernissage corpus: a conversational human-robot-interaction dataset
Dinesh Babu Jayagopi, Samira Sheikhi, David Klotz, Johannes Wienke, Jean-Marc Odobez, Sebastian Wrede 0001, Vasil Khalidov, Laurent Nyugen, Britta Wrede, Daniel Gatica-Perez |
HRI | 5 |
| 2013 | Person independent 3D gaze estimation from remote RGB-D camerasabstractWe address the problem of person independent 3D gaze estimation using a remote, low resolution, RGB-D camera. The approach relies on a sparse technique to reconstruct normalized eye test images from a gaze appearance model (a set of eye image/gaze pairs) and infer their gaze accordingly. In this context, the paper makes three contributions: (i) unlike most previous approaches, we exploit the coupling (and constraints) between both eyes to infer their gaze jointly; (ii) we show that a generic gaze appearance model built from the aggregation of person-specific models can be used to handle unseen users and compensate for appearance variations across people, since a test user eyes' appearance will be reconstructed from similar users within the generic model. (iii) we propose an automatic model selection method that leads to comparable performance with a reduced computational load. Kenneth Alberto Funes Mora, Jean-Marc Odobez |
ICIP | 2 |
| 2013 | Time-sensitive topic models for action recognition in videosabstractIn this paper, we postulate that temporal information is important for action recognition in videos. Keeping temporal information, videos are represented as word×time documents. We propose to use time-sensitive probabilistic topic models and we extend them for the context of supervised learning. Our time-sensitive approach is compared to both PLSA and Bag-of-Words. Our approach is shown to both capture semantics from data and yield classification performance comparable to other methods, outperforming them when the amount of training data is low. Romain Tavenard, Rémi Emonet, Jean-Marc Odobez |
ICIP | 3 |
| 2013 | A semi-automated system for accurate gaze coding in natural dyadic interactionsabstractIn this paper we propose a system capable of accurately coding gazing events in natural dyadic interactions. Contrary to previous works, our approach exploits the actual continuous gaze direction of a participant by leveraging on remote RGB-D sensors and a head pose-independent gaze estimation method. Our contributions are: i) we propose a system setup built from low-cost sensors and a technique to easily calibrate these sensors in a room with minimal assumptions; ii) we propose a method which, provided short manual annotations, can automatically detect gazing events in the rest of the sequence; iii) we demonstrate on substantially long, natural dyadic data that high accuracy can be obtained, showing the potential of our system. Our approach is non-invasive and does not require collaboration from the interactors. These characteristics are highly valuable in psychology and sociology research. Kenneth Alberto Funes Mora, Laurent Son Nguyen, Daniel Gatica-Perez, Jean-Marc Odobez |
ICMI | 4 |
| 2013 | Leveraging the robot dialog state for visual focus of attention recognitionabstractThe Visual Focus of Attention (what or whom a person is looking at) or VFOA is a fundamental cue in non-verbal communication and plays an important role when designing effective human-machine interaction systems. However, recognizing the VFOA of an interacting person is difficult for a robot, since due to low resolution imaging, eye gaze estimation is not possible. Rather, head pose cue is used as a substitute for gaze, but leads to ambiguities in its interpretation as VFOA indicator. In this paper, we investigate the use of the robot conversational state, which the robot is aware of, as contextual information to improve VFOA recognition from head pose. We propose a dynamic Bayesian model that accounts for the robot state (speaking status, person he addresses, reference to objects) along with a dynamic head-to-gaze mapping function. Experiments on a publicly available human-robot interaction dataset, where a humanoid robot plays the role of an art guide and quiz master, shows that using such conversational context is effective in improving VFOA. Samira Sheikhi, Vasil Khalidov, David Klotz, Britta Wrede, Jean-Marc Odobez |
ICMI | 5 |
| 2013 | Fusing matching and biometric similarity measures for face diarization in videoabstractThis paper addresses face diarization in videos, that is, deciding which face appears and when in the video. To achieve this face-track clustering task, we propose a hierarchical approach combining the strength of two complementary measures: (i) a pairwise matching similarity relying on local interest points allowing the accurate clustering of faces tracks captured in similar conditions, a situation typically found in temporally close shots of broadcast videos or in talk-shows; (ii) a biometric cross-likelihood ratio similarity measure relying on Gaussian Mixture Models (GMMs) modeling the distribution of densely sampled local features (Discrete Cosine Transform (DCT) coefficients), that better handle appearance variability. Experiments carried out on a public video dataset and on the data from the French REPERE challenge demonstrate the effectiveness of our approach in comparison with state-of-the-art methods. Elie Khoury 0001, Paul Gay, Jean-Marc Odobez |
ICMR | 3 |
| 2013 | A Sequential Topic Model for Mining Recurrent Activities from Long Term Video Logs
Jagannadan Varadarajan, Rémi Emonet, Jean-Marc Odobez |
Int. J. Comput. Vis. | 3 |
| 2013 | Track Creation and Deletion Framework for Long-Term Online Multiface TrackingabstractTo improve visual tracking, a large number of papers study more powerful features, or better cue fusion mechanisms, such as adaptation or contextual models. A complementary approach consists of improving the track management, that is, deciding when to add a target or stop its tracking, for example, in case of failure. This is an essential component for effective multiobject tracking applications, and is often not trivial. Deciding whether or not to stop a track is a compromise between avoiding erroneous early stopping while tracking is fine, and erroneous continuation of tracking when there is an actual failure. This decision process, very rarely addressed in the literature, is difficult due to object detector deficiencies or observation models that are insufficient to describe the full variability of tracked objects and deliver reliable likelihood (tracking) information. This paper addresses the track management issue and presents a real-time online multiface tracking algorithm that effectively deals with the above difficulties. The tracking itself is formulated in a multiobject state-space Bayesian filtering framework solved with Markov Chain Monte Carlo. Within this framework, an explicit probabilistic filtering step decides when to add or remove a target from the tracker, where decisions rely on multiple cues such as face detections, likelihood measures, long-term observations, and track state characteristics. The method has been applied to three challenging data sets of more than 9 h in total, and demonstrate a significant performance increase compared to more traditional approaches (Markov Chain Monte Carlo, reversible-jump Markov Chain Monte Carlo) only relying on head detection and likelihood for track management. Stefan Duffner, Jean-Marc Odobez |
IEEE Trans. Image Process. | 2 |
| 2013 | Observation of Vehicle Axles Through Pass-by Noise: A Strategy of Microphone Array DesignabstractThis paper focuses on road traffic monitoring using sounds and proposes, more specifically, a microphone array design methodology for observing vehicle trajectory from acoustic-based correlation functions. In a former work, authors have shown that combining generalized cross correlation (GCC) functions and a particle filter onto the audio signals simultaneously acquired by two sensors placed near the road allows the joint estimation of the speed and the wheelbase length of road vehicles as they pass by. This is mainly due to the broadband nature of the tire/road noise, which makes their spatial dissociation possible by means of an appropriate GCC processor. At the time, nothing has been said about the best distance to chose between the sensors. A methodology is proposed here to find this optimum, which is expected to improve the observation quality and, thus, the tracking performance. Theoretical developments of this paper are partially assessed with preliminary experiments. Patrick Marmaroli, Mikael Carmona, Jean-Marc Odobez, Xavier Falourd, Hervé Lissek |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2012 | We are not contortionists: Coupled adaptive learning for head and body orientation estimation in surveillance videoabstractIn this paper, we deal with the estimation of body and head poses (i.e orientations) in surveillance videos, and we make three main contributions. First, we address this issue as a joint model adaptation problem in a semi-supervised framework. Second, we propose to leverage the adaptation on multiple information sources (external labeled datasets, weak labels provided by the motion direction, data structure manifold), and in particular, on the coupling at the output level of the head and body classifiers, accounting for the restriction in the configurations that the head and body pose can jointly take. Third, we propose a kernel-formulation of this principle that can be efficiently solved using a global optimization scheme. The method is applied to body and head features computed from automatically extracted body and head location tracks. Thorough experiments on several datasets demonstrate the validity of our approach, the benefit of the coupled adaptation, and that the method performs similarly or better than a state-of-the-art algorithm. Cheng Chen 0023, Jean-Marc Odobez |
CVPR | 2 |
| 2012 | Bridging the past, present and future: Modeling scene activities from event relationships and global rulesabstractThis paper addresses the discovery of activities and learns the underlying processes that govern their occurrences over time in complex surveillance scenes. To this end, we propose a novel topic model that accounts for the two main factors that affect these occurrences: (1) the existence of global scene states that regulate which of the activities can spontaneously occur; (2) local rules that link past activity occurrences to current ones with temporal lags. These complementary factors are mixed in the probabilistic generative process, thanks to the use of a binary random variable that selects for each activity occurrence which one of the above two factors is applicable. All model parameters are efficiently inferred using a collapsed Gibbs sampling inference scheme. Experiments on various datasets from the literature show that the model is able to capture temporal processes at multiple scales: the scene-level first order Markovian process, and causal relationships amongst activities that can be used to predict which activity can happen after another one, and after what delay, thus providing a rich interpretation of the scene's dynamical content. Jagannadan Varadarajan, Rémi Emonet, Jean-Marc Odobez |
CVPR | 3 |
| 2012 | Using self-context for multimodal detection of head nods in face-to-face interactionsabstractHead nods occur in virtually every face-to-face discussion. As part of the backchannel domain, they are not only used to express a 'yes', but also to display interest or enhance communicative attention. Detecting head nods in natural interactions is a challenging task as head nods can be subtle, both in amplitude and duration. In this study, we make use of findings in psychology establishing that the dynamics of head gestures are conditioned on the person's speaking status. We develop a multimodal method using audio-based self-context to detect head nods in natural settings. We demonstrate that our multimodal approach using the speaking status of the person under analysis significantly improved the detection rate over a visual-only approach. Laurent Son Nguyen, Jean-Marc Odobez, Daniel Gatica-Perez |
ICMI | 2 |
| 2012 | Investigating the midline effect for visual focus of attention recognitionabstractThis paper addresses the recognition of people's visual focus of attention (VFOA), the discrete version of gaze indicating who is looking at whom or what. In absence of high definition images, we rely on people's head pose to recognize the VFOA. To the contrary of most previous works that assumed a fixed mapping between head pose directions and gaze target directions, we investigate novel gaze models documented in psychovision that produce a dynamic (temporal) mapping between them. This mapping accounts for two important factors affecting the head and gaze relationship: the shoulder orientation defining the gaze midline of a person varies over time; and gaze shifts from frontal to the side involve different head rotations than the reverse. Evaluated on a public dataset and on data recorded with the humanoid robot Nao, the method exhibit better adaptivity often producing better performance than state-of-the-art approach. Samira Sheikhi, Jean-Marc Odobez |
ICMI | 2 |
| 2011 | Combined estimation of location and body pose in surveillance videoabstractIn surveillance videos, cues such as head or body pose provide important information for analyzing people's behavior and interactions. In this paper we propose an approach that jointly estimates body location and body pose in monocular surveillance video. Our approach is based on tracks derived by multi-object tracking. First, body pose classification is conducted using sparse representation technique on each frame of the tracks, generating (noisy) observation on body poses. Then, both location and body pose in 3D space are estimated jointly in a particle filtering framework by utilizing a soft coupling of body pose with the movement. The experiments show that the proposed system successfully tracks body position and pose simultaneously in many scenarios. The output of the system can be used to perform further analysis on behaviors and interactions. Cheng Chen 0023, Alexandre Heili, Jean-Marc Odobez |
AVSS | 3 |
| 2011 | Multi-camera open space human activity discovery for anomaly detectionabstractWe address the discovery of typical activities in video stream contents and its exploitation for estimating the abnormality levels of these streams. Such estimates can be used to select the most interesting cameras to show to a human operator. Our contributions come from the following facets: i) the method is fully unsupervised and learns the activities from long term data; ii) the method is scalable and can efficiently handle the information provided by multiple un-calibrated cameras, jointly learning activities shared by them if it happens to be the case (e.g. when they have overlapping fields of view); iii) unlike previous methods which were mainly applied to structured urban traffic scenes, we show that ours performs well on videos from a metro environment where human activities are only loosely constrained. Rémi Emonet, Jagannadan Varadarajan, Jean-Marc Odobez |
AVSS | 3 |
| 2011 | Joint Adaptive Colour Modelling and Skin, Hair and Clothes Segmentation using Coherent Probabilistic Index MapsabstractWe address the joint segmentation of regions around faces into different classesskin, hair, clothing and background -and the learning of their respective colour models.To this end, we adopt a Bayesian framework with two main elements.First, the modelling and learning of prior beliefs over class colour distributions of different kinds, viz.discrete distributions for background and clothing, and continuous distributions for skin and hair.This component of the model allows the formulation of the colour adaptation task as the computation of the posterior beliefs over class colour distributions given the prior (general) class colour model, a colour likelihood function and observed pixel colours.Second, a spatial prior based on probabilistic index maps enhanced with Markov random field regularisation.The inference of the segmentation and of the adapted colour models is solved using a variational scheme.Segmentation results on a small database of annotated images demonstrate the impact of the different modelling choices. Carl Scheffler, Jean-Marc Odobez |
BMVC | 2 |
| 2011 | Extracting and locating temporal motifs in video scenes using a hierarchical non parametric Bayesian modelabstractIn this paper, we present an unsupervised method for mining activities in videos. From unlabeled video sequences of a scene, our method can automatically recover what are the recurrent temporal activity patterns (or motifs) and when they occur. Using non parametric Bayesian methods, we are able to automatically find both the underlying number of motifs and the number of motif occurrences in each document. The model's robustness is first validated on synthetic data. It is then applied on a large set of video data from state-of-the-art papers. We show that it can effectively recover temporal activities with high semantics for humans and strong temporal information. The model is also used for prediction where it is shown to be as efficient as other approaches. Although illustrated on video sequences, this model can be directly applied to various kinds of time series where multiple activities occur simultaneously. Rémi Emonet, Jagannadan Varadarajan, Jean-Marc Odobez |
CVPR | 3 |
| 2011 | Exploiting long-term observations for track creation and deletion in online multi-face trackingabstractIn many visual multi-object tracking applications, the question when to add or remove a target is not trivial due to, for example, erroneous outputs of object detectors or observation models that cannot describe the full variability of the objects to track. In this paper, we present a real-time, online multi-face tracking algorithm that effectively deals with missing or uncertain detections in a principled way. The tracking is formulated in a multi-object state-space Bayesian filtering framework solved with Markov Chain Monte Carlo. Within this framework, an explicit probabilistic filtering step relying on head detections, likelihood models, and long term observations as well as object track characteristics has been designed to take the decision on when to add or remove a target from the tracker. The proposed method applied on three challenging datasets of more than 9 hours shows a significant performance increase compared to a traditional approach relying on head detection and likelihood models only. Stefan Duffner, Jean-Marc Odobez |
FG | 2 |
| 2011 | Searching the past: an improved shape descriptor to retrieve maya hieroglyphsabstractArchaeologists often spend significant time looking at traditional printed catalogs to identify and classify historical images. Our collaborative efforts between archaeologists and multimedia researchers seek to develop a tool to retrieve two specific types of ancient Maya visual information: hieroglyphs and iconographic elements. Towards that goal we present two contributions in this paper. The first one is the introduction and analysis of a new dataset of 3400+ Maya hieroglyphs, whose compilation involved manual search, annotation and segmentation by experts. This dataset presents several challenges for visual description and automatic retrieval as it is rich in complex visual details. The second and main contribution is the in-depth analysis of the Histogram Of Orientation Shape Context (HOOSC), and more precisely, the development of 4 improvements that were designed to handle the visual complexity of Maya hieroglyphs: open contours, mixture of thick and thin lines, hatches, large instance variability, and a variety of internal details. Experiments demonstrate that the adequate combination of our improvements to retrieve Maya hieroglyphs, provides results with roughly 20% more precision compared to the original HOOSC descriptor. Complementary results with the MPEG-7 shape dataset validate (or not) the proposed improvements, showing that the design of appropriate descriptors depends on the nature of the shapes one deals with. Edgar Roman-Rangel, Carlos Pallan, Jean-Marc Odobez, Daniel Gatica-Perez |
ACM Multimedia | 3 |
| 2011 | Engagement-based Multi-party Dialog with a Humanoid Robot
David Klotz, Johannes Wienke, Julia Peltason, Britta Wrede, Sebastian Wrede 0001, Vasil Khalidov, Jean-Marc Odobez |
SIGDIAL Conference | 7 |
| 2011 | 3D human pose recovery from image by efficient visual feature selection
Cheng Chen 0023, Yi Yang 0001, Feiping Nie 0001, Jean-Marc Odobez |
Comput. Vis. Image Underst. | 4 |
| 2011 | Fast human detection from joint appearance and foreground feature subset covariances
Jian Yao 0007, Jean-Marc Odobez |
Comput. Vis. Image Underst. | 2 |
| 2011 | Analyzing Ancient Maya Glyph Collections with Contextual Shape Descriptors
Edgar Roman-Rangel, Carlos Pallan, Jean-Marc Odobez, Daniel Gatica-Perez |
Int. J. Comput. Vis. | 3 |
| 2011 | Multiperson Visual Focus of Attention from Head Pose and Meeting Contextual CuesabstractThis paper introduces a novel contextual model for the recognition of people's visual focus of attention (VFOA) in meetings from audio-visual perceptual cues. More specifically, instead of independently recognizing the VFOA of each meeting participant from his own head pose, we propose to jointly recognize the participants' visual attention in order to introduce context-dependent interaction models that relate to group activity and the social dynamics of communication. Meeting contextual information is represented by the location of people, conversational events identifying floor holding patterns, and a presentation activity variable. By modeling the interactions between the different contexts and their combined and sometimes contradictory impact on the gazing behavior, our model allows us to handle VFOA recognition in difficult task-based meetings involving artifacts, presentations, and moving people. We validated our model through rigorous evaluation on a publicly available and challenging data set of 12 real meetings (5 hours of data). The results demonstrated that the integration of the presentation and conversation dynamical context using our model can lead to significant performance improvements. Sileye O. Ba, Jean-Marc Odobez |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Probabilistic Latent Sequential Motifs: Discovering Temporal Activity Patterns in Video ScenesabstractThis paper introduces a novel probabilistic activity modeling approach that mines recurrent sequential patterns from documents given as word-time occurrences. In this model, documents are represented as a mixture of sequential activity motifs (or topics) and their starting occurrences. The novelties are threefold. First, unlike previous ap-proaches where topics only modeled the co-occurrence of words at a given time instant, our topics model the co-occurrence and temporal order in which the words occur within a temporal window. Second, our model accounts for the important case where activities occur concurrently in the document. And third, our method explicitly models with latent variables the starting time of the activities within the documents, enabling to implicitly align the occurrences of the same pattern during the joint inference of the temporal topics and their starting times. The model and its robustness to the presence of noise have been validated on synthetic data. Its effectiveness is also illustrated in video activity analysis from low-level motion features, where the discovered topics capture frequent patterns that implicitly represent typical trajectories of scene objects. 1 Jagannadan Varadarajan, Rémi Emonet, Jean-Marc Odobez |
BMVC | 3 |
| 2009 | Dynamic Partitioned Sampling For Tracking With Discriminative FeaturesabstractWe present a multi-cue fusion method for tracking with particle filters which relies on a novel hierarchical sampling strategy. Similarly to previous works, it tackles the problem of tracking in a relatively high-dimensional state space by dividing such a space into partitions, each one corresponding to a single cue, and sampling from them in a hierarchical manner. However, unlike other approaches, the order of partitions is not fixed a priori but changes dynamically depending on the reliability of each cue, i.e. more reliable cues are sampled first. We call this approach Dynamic Partitioned Sampling (DPS). The reliability of each cue is measured in terms of its ability to discriminate the object with respect to the background, where the background is not described by a fixed model or by random patches but is represented by a set of informative "background particles" which are tracked in order to be as similar as possible to the object. The effectiveness of this general framework is demonstrated on the specific problem of head tracking with three different cues: colour, edge and contours. Experimental results prove the robustness of our algorithm in several challenging video sequences. Stefan Duffner, Jean-Marc Odobez, Elisa Ricci 0001 |
BMVC | 2 |
| 2009 | Learning large margin likelihoods for realtime head pose tracking
Elisa Ricci 0001, Jean-Marc Odobez |
ICIP | 2 |
| 2009 | Visual activity context for focus of attention estimation in dynamic meetingsabstractWe address the problem of recognizing, in dynamic meetings in which people do not remain seated all the time, the visual focus of attention (VFOA) of seated people from their head pose and contextual activity cues. We propose a model that comprises the VFOA of a meeting participant as the hidden state, and his head pose as the observation. To account for the presence of moving visual targets due to the dynamic nature of the meeting, the locations of the visual targets are used as an input variables to the head pose observation model. Contextual information is introduced in the VFOA dynamics through a slide activity variable and speaking or visual activity variables that relate people's focus to the meeting activity context. The main novelty of this paper is the introduction of visual activity context for FOA recognition to account for the correlation between a person's focus and the other people's gestures, hand and body motions. We evaluate our model on a large dataset of 5 hours. Our results show that, for VFOA estimation in meetings, visual activity contextual information can be as effective as speaking context. Sileye O. Ba, Hayley Hung, Jean-Marc Odobez |
ICME | 3 |
| 2009 | Structure and appearance features for robust 3D facial actions trackingabstractThis paper presents a robust and accurate method for joint head pose and facial actions tracking, even under challenging conditions such as varying lighting, large head movements, and fast motion. This is made possible by the combination of two types of facial features. We use locations sampled from the facial texture whose appearance is initialized on the first frame and adapted over time, and also illumination-invariant patches located on characteristic points of the face such as the corners of the eyes or of the mouth. The first type of features contains rich information about the global appearance of the face and thus leads to an accurate tracking, while the second type guaranties robustness and stability by avoiding drift. We demonstrate our system on the Boston University Face Tracking benchmark, and show it outperforms state-of-the-art methods. Stéphanie Lefèvre, Jean-Marc Odobez |
ICME | 2 |
| 2009 | Investigating the use of visual focus of attention for audio-visual speaker diarisationabstractAudio-visual speaker diarisation is the task of estimating ``who spoke when'' using audio and visual cues. Giulia Garau, Sileye O. Ba, Hervé Bourlard, Jean-Marc Odobez |
ACM Multimedia | 4 |
| 2009 | Recognizing Visual Focus of Attention From Head Pose in Natural MeetingsabstractWe address the problem of recognizing the visual focus of attention (VFOA) of meeting participants based on their head pose. To this end, the head pose observations are modeled using a Gaussian mixture model (GMM) or a hidden Markov model (HMM) whose hidden states correspond to the VFOA. The novelties of this paper are threefold. First, contrary to previous studies on the topic, in our setup, the potential VFOA of a person is not restricted to other participants only. It includes environmental targets as well (a table and a projection screen), which increases the complexity of the task, with more VFOA targets spread in the pan as well as tilt gaze space. Second, we propose a geometric model to set the GMM or HMM parameters by exploiting results from cognitive science on saccadic eye motion, which allows the prediction of the head pose given a gaze target. Third, an unsupervised parameter adaptation step not using any labeled data is proposed, which accounts for the specific gazing behavior of each participant. Using a publicly available corpus of eight meetings featuring four persons, we analyze the above methods by evaluating, through objective performance measures, the recognition of the VFOA from head pose information obtained either using a magnetic sensor device or a vision-based tracking system. The results clearly show that in such complex but realistic situations, the VFOA recognition performance is highly dependent on how well the visual targets are separated for a given meeting participant. In addition, the results show that the use of a geometric model with unsupervised adaptation achieves better results than the use of training data to set the HMM parameters. Sileye O. Ba, Jean-Marc Odobez |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | Multi-party focus of attention recognition in meetings from head pose and multimodal contextual cuesabstractThis paper presents investigations on visual focus of attention (VFOA) recognition in meetings from audio-visual perceptual cues. Rather than independently recognizing the VFOA of each participant from his own head pose, we propose to recognize participants' VFOA jointly in order to introduce context dependent interaction models that relates to group activity and the social dynamics of communication. To this end, we designed an input-output hidden Markov model (IOHMM), whose hidden states are the joint VFOA of all participants, and whose main observations are the head poses. Interaction models are introduced in the form of contextual cues that affect the temporal evolution of the joint VFOA sequence, allowing us to model group dynamics that accounts for people's tendency to share the same focus, or to have their VFOA driven by contextual cues such as slide activity or the participant speaking activity. The model is rigorously evaluated on a publicly available dataset of 4 real meetings of 23min on average, showing an overall 10% relative performance increase w.r.t. the independent recognition case. Sileye O. Ba, Jean-Marc Odobez |
ICASSP | 2 |
| 2008 | Visual focus of attention estimation from head pose posterior probability distributionsabstractWe address the problem of recognizing the visual focus of attention (VFOA) of meeting participants from their head pose and contextual cues. The main contribution of the paper is the use of a head pose posterior distribution as a representation of the head pose information contained in the image data. This posterior encodes the probabilities of the different head poses given the image data, and constitute therefore a richer representation of the data than the mean or the mode of this distribution, as done in all previous work. These observations are exploited in a joint interaction model of all meeting participants pose observations, VFOAs, speaking status and of environmental contextual cues. Numerical experiments on a public database of 4 meetings of 22 min on average show that this change of representation allows for a 5.4% gain with respect to the standard approach using head pose as observation. Sileye O. Ba, Jean-Marc Odobez |
ICME | 2 |
| 2008 | Investigating automatic dominance estimation in groups from visual attention and speaking activityabstractWe study the automation of the visual dominance ratio (VDR); a classic measure of displayed dominance in social psychology literature, which combines both gaze and speaking activity cues. The VDR is modified to estimate dominance in multi-party group discussions where natural verbal exchanges are possible and other visual targets such as a table and slide screen are present. Our findings suggest that fully automated versions of these measures can estimate effectively the most dominant person in a meeting and can match the dominance estimation performance when manual labels of visual attention are used. Hayley Hung, Dinesh Babu Jayagopi, Sileye O. Ba, Jean-Marc Odobez, Daniel Gatica-Perez |
ICMI | 4 |
| 2008 | Predicting two facets of social verticality in meetings from five-minute time slices and nonverbal cuesabstractThis paper addresses the automatic estimation of two aspects of social verticality (status and dominance) in small-group meetings using nonverbal cues. The correlation of nonverbal behavior with these social constructs have been extensively documented in social psychology, but their value for computational models is, in many cases, still unknown. We present a systematic study of automatically extracted cues - including vocalic, visual activity, and visual attention cues - and investigate their relative effectiveness to predict both the most-dominant person and the high-status project manager from relative short observations. We use five hours of task-oriented meeting data with natural behavior for our experiments. Our work suggests that, although dominance and role-based status are related concepts, they are not equivalent and are thus not equally explained by the same nonverbal cues. Furthermore, the best cues can correctly predict the person with highest dominance or role-based status with an accuracy of 70% approximately. Dinesh Babu Jayagopi, Sileye O. Ba, Jean-Marc Odobez, Daniel Gatica-Perez |
ICMI | 3 |
| 2008 | Detecting queues at vending machines: A statistical layered approachabstractThis paper presents a method for monitoring activities at a ticket vending machine in a video-surveillance context. Rather than relying on the output of a tracking module, which is prone to errors, the events are directly recognized from image measurements. This especially does not require tracking. A statistical layered approach is proposed, where in the first layer, several sub-events are defined and detected using a discriminative approach. The second layer uses the result of the first and models the temporal relationships of the high-level event using a hidden Markov model (HMM). Results are assessed on 3h30 hours of real video footage coming from Turin metro station. Xavier Naturel, Jean-Marc Odobez |
ICPR | 2 |
| 2008 | Tracking the Visual Focus of Attention for a Varying Number of Wandering PeopleabstractWe define and address the problem of finding the visual focus of attention for a varying number of wandering people (VFOA-W), determining where the people's movement is unconstrained. VFOA-W estimation is a new and important problem with mplications for behavior understanding and cognitive science, as well as real-world applications. One such application, which we present in this article, monitors the attention passers-by pay to an outdoor advertisement. Our approach to the VFOA-W problem proposes a multi-person tracking solution based on a dynamic Bayesian network that simultaneously infers the (variable) number of people in a scene, their body locations, their head locations, and their head pose. For efficient inference in the resulting large variable-dimensional state-space we propose a Reversible Jump Markov Chain Monte Carlo (RJMCMC) sampling scheme, as well as a novel global observation model which determines the number of people in the scene and localizes them. We propose a Gaussian Mixture Model (GMM) and Hidden Markov Model (HMM)-based VFOA-W model which use head pose and location information to determine people's focus state. Our models are evaluated for tracking performance and ability to recognize people looking at an outdoor advertisement, with results indicating good performance on sequences where a moderate number of people pass in front of an advertisement. Kevin Smith 0001, Sileye O. Ba, Jean-Marc Odobez, Daniel Gatica-Perez |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Multi-Layer Background Subtraction Based on Color and TextureabstractIn this paper, we propose a robust multi-layer background subtraction technique which takes advantages of local texture features represented by local binary patterns (LBP) and photometric invariant color measurements in RGB color space. LBP can work robustly with respective to light variation on rich texture regions but not so efficiently on uniform regions. In the latter case, color information should overcome LBP's limitation. Due to the illumination invariance of both the LBP feature and the selected color feature, the method is able to handle local illumination changes such as cast shadows from moving objects. Due to the use of a simple layer-based strategy, the approach can model moving background pixels with quasi-periodic flickering as well as background scenes which may vary over time due to the addition and removal of long-time stationary objects. Finally, the use of a cross-bilateral filter allows to implicitly smooth detection results over regions of similar intensity and preserve object boundaries. Numerical and qualitative experimental results on both simulated and real data demonstrate the robustness of the proposed method. Jian Yao 0007, Jean-Marc Odobez |
CVPR | 2 |
| 2007 | A Cognitive and Unsupervised Map Adaptation Approach to the Recognition of the Focus of Attention from Head PoseabstractIn this paper, the recognition of the visual focus of attention (VFOA) of meeting participants (as defined by their eye gaze direction) from their head pose is addressed. To this end, the head pose observations are modeled using an hidden Markov model (HMM) whose hidden states corresponds to the VFOA. The novelties are threefold. First, contrary to previous studies on the topic, in our set-up, the potential VFOA of a person is not restricted to other participants only, but includes environmental targets (a table and a projection screen), which increases the complexity of the task, with more VFOA targets spread in the pan and tilt (as well) gaze space. Second, the HMM parameters are set by exploiting results from the cognitive science on saccadic eye motion, which allows to predict what the head pose should be given an actual gaze target. Third, an unsupervised parameter adaptation step is proposed which accounts for the specific gazing behaviour of each participant. Using a publicly available corpus of 8 meetings featuring 4 persons, we analyze the above methods by evaluating, through objective performance measures, the recognition of the VFOA from head pose information obtained either using a magnetic sensor device or a vision based tracking system. Jean-Marc Odobez, Sileye O. Ba |
ICME | 1 |
| 2007 | Using audio and video features to classify the most dominant person in a group meetingabstractThe automated extraction of semantically meaningful information from multi-modal data is becoming increasingly necessary due to the escalation of captured data for archival. A novel area of multi-modal data labelling, which has received relatively little attention, is the automatic estimation of the most dominant person in a group meeting. In this paper, we provide a framework for detecting dominance in group meetings using different audio and video cues. We show that by using a simple model for dominance estimation we can obtain promising results. Hayley Hung, Dinesh Babu Jayagopi, Chuohao Yeo, Gerald Friedland, Sileye O. Ba, Jean-Marc Odobez, Kannan Ramchandran, Nikki Mirghafori, Daniel Gatica-Perez |
ACM Multimedia | 6 |
| 2007 | A Thousand Words in a SceneabstractThis paper presents a novel approach for visual scene modeling and classification, investigating the combined use of text modeling methods and local invariant features. Our work attempts to elucidate (1) whether a text-like bag-of-visterms representation (histogram of quantized local visual features) is suitable for scene (rather than object) classification, (2) whether some analogies between discrete scene representations and text documents exist, and (3) whether unsupervised, latent space models can be used both as feature extractors for the classification task and to discover patterns of visual co-occurrence. Using several data sets, we validate our approach, presenting and discussing experiments on each of these issues. We first show, with extensive experiments on binary and multi-class scene classification tasks using a 9,500-image data set, that the bag-of-visterms representation consistently outperforms classical scene classification approaches. In other data sets we show that our approach competes with or outperforms other recent, more complex, methods. We also show that Probabilistic Latent Semantic Analysis (PLSA) generates a compact scene representation, discriminative for accurate classification, and more robust than the bag-of-visterms representation when less labeled training data is available. Finally, through aspect-based image ranking experiments, we show the ability of PLSA to automatically extract visually meaningful scene patterns, making such representation useful for browsing image collections. Pedro Quelhas, Florent Monay, Jean-Marc Odobez, Daniel Gatica-Perez, Tinne Tuytelaars |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Audiovisual Probabilistic Tracking of Multiple Speakers in MeetingsabstractTracking speakers in multiparty conversations constitutes a fundamental task for automatic meeting analysis. In this paper, we present a novel probabilistic approach to jointly track the location and speaking activity of multiple speakers in a multisensor meeting room, equipped with a small microphone array and multiple uncalibrated cameras. Our framework is based on a mixed-state dynamic graphical model defined on a multiperson state-space, which includes the explicit definition of a proximity-based interaction model. The model integrates audiovisual (AV) data through a novel observation model. Audio observations are derived from a source localization algorithm. Visual observations are based on models of the shape and spatial structure of human heads. Approximate inference in our model, needed given its complexity, is performed with a Markov Chain Monte Carlo particle filter (MCMC-PF), which results in high sampling efficiency. We present results-based on an objective evaluation procedure-that show that our framework 1) is capable of locating and tracking the position and speaking activity of multiple meeting participants engaged in real conversations with good accuracy, 2) can deal with cases of visual clutter and occlusion, and 3) significantly outperforms a traditional sampling-based approach Daniel Gatica-Perez, Guillaume Lathoud, Jean-Marc Odobez, Iain McCowan |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Short-Term Spatio-Temporal Clustering Applied to Multiple Moving SpeakersabstractDistant microphones permit to process spontaneous multiparty speech with very little constraints on speakers, as opposed to close-talking microphones. Minimizing the constraints on speakers permits a large diversity of applications, including meeting summarization and browsing, surveillance, hearing aids, and more natural human-machine interaction. Such applications of distant microphones require to determine where and when the speakers are talking. This is inherently a multisource problem, because of background noise sources, as well as the natural tendency of multiple speakers to talk over each other. Moreover, spontaneous speech utterances are highly discontinuous, which makes it difficult to track the multiple speakers with classical filtering approaches, such as Kalman filtering of particle filters. As an alternative, this paper proposes a probabilistic framework to determine the trajectories of multiple moving speakers in the short-term only, i.e., only while they speak. Instantaneous location estimates that are close in space and time are grouped into ldquoshort-term clustersrdquo in a principled manner. Each short-term cluster determines the precise start and end times of an utterance and a short-term spatial trajectory. Contrastive experiments clearly show the benefit of using short-term clustering, on real indoor recordings with seated speakers in meetings, as well as multiple moving speakers. Guillaume Lathoud, Jean-Marc Odobez |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Tracking the multi person wandering visual focus of attentionabstractEstimating the wandering visual focus of attention (WVFOA) for multiple people is an important problem with many applications in human behavior understanding. One such application, addressed in this paper, monitors the attention of passers-by to outdoor advertisements. This paper investigates the problem of tracking the wandering visual focus-of-attention (VFOA) of multiple people, an important problem with many applications in human behavior understanding. We address the specific problem of monitoring attention to outdoor advertisements. To solve the WVFOA problem, we propose a multi-person tracking approach based on a hybrid Dynamic Bayesian Network that simultaneously infers the number of people in the scene, their body and head locations, and their head pose, in a joint state-space formulation that is amenable for person interaction modeling. The model exploits both global measurements and individual observations for the VFOA. For inference in the resulting high-dimensional state-space, we propose a trans-dimensional Markov Chain Monte Carlo (MCMC) sampling scheme, which not only handles a varying number of people, but also efficiently searches the state-space by allowing person-part state updates. Our model was rigorously evaluated for tracking and its ability to recognize when people look at an outdoor advertisement using a realistic data set. Kevin Smith 0001, Sileye O. Ba, Daniel Gatica-Perez, Jean-Marc Odobez |
ICMI | 4 |
| 2006 | Embedding Motion in Model-Based Stochastic TrackingabstractParticle filtering is now established as one of the most popular methods for visual tracking. Within this framework, there are two important considerations. The first one refers to the generic assumption that the observations are temporally independent given the sequence of object states. The second consideration, often made in the literature, uses the transition prior as the proposal distribution. Thus, the current observations are not taken into account, requiring the noise process of this prior to be large enough to handle abrupt trajectory changes. As a result, many particles are either wasted in low likelihood regions of the state space, resulting in low sampling efficiency, or more importantly, propagated to distractor regions of the image, resulting in tracking failures. In this paper, we propose to handle both considerations using motion. We first argue that, in general, observations are conditionally correlated, and propose a new model to account for this correlation, allowing for the natural introduction of implicit and/or explicit motion measurements in the likelihood term. Second, explicit motion measurements are used to drive the sampling process towards the most likely regions of the state space. Overall, the proposed model handles abrupt motion changes and filters out visual distractors, when tracking objects with generic models based on shape or color distribution. Results were obtained on head tracking experiments using several sequences with moving camera involving large dynamics. When compared against the Condensation Algorithm, they have demonstrated the superior tracking performance of our approach. Jean-Marc Odobez, Daniel Gatica-Perez, Sileye O. Ba |
IEEE Trans. Image Process. | 1 |
| 2006 | Application of Information Retrieval Technologies to Presentation SlidesabstractPresentations are becoming an increasingly more common means of communication in working environments, and slides are often the necessary supporting material on which the presentations rely. In this paper, we describe a slide indexing and retrieval system in which the slides are captured as images (through a framegrabber) at the moment they are displayed during a presentation and then transcribed with an optical character recognition (OCR) system. In this context, we show that such an approach presents several advantages over the use of commercial software (API based) to obtain the slide transcriptions. We report a set of retrieval experiments conducted on a database of 26 real presentations (570 slides) collected at a workshop. The experiments show that the overall retrieval performance is close to that obtained using either a manual transcription of the slides or the API software. Moreover, the experiments show that the OCR-based approach outperforms significantly the API in extracting the text embedded in images and figures Alessandro Vinciarelli, Jean-Marc Odobez |
IEEE Trans. Multim. | 2 |
| 2005 | Using Particles to Track Varying Numbers of Interacting PeopleabstractIn this paper, we present a Bayesian framework for the fully automatic tracking of a variable number of interacting targets using a fixed camera. This framework uses a joint multi-object state-space formulation and a trans-dimensional Markov Chain Monte Carlo (MCMC) particle filter to recursively estimates the multi-object configuration and efficiently search the state-space. We also define a global observation model comprised of color and binary measurements capable of discriminating between different numbers of objects in the scene. We present results which show that our method is capable of tracking varying numbers of people through several challenging real-world tracking situations such as full/partial occlusion and entering/leaving the scene. Kevin Smith 0001, Daniel Gatica-Perez, Jean-Marc Odobez |
CVPR (1) | 3 |
| 2005 | Modeling Scenes with Local Descriptors and Latent AspectsabstractWe present a new approach to model visual scenes in image collections, based on local invariant features and probabilistic latent space models. Our formulation provides answers to three open questions:(l) whether the invariant local features are suitable for scene (rather than object) classification; (2) whether unsupennsed latent space models can be used for feature extraction in the classification task; and (3) whether the latent space formulation can discover visual co-occurrence patterns, motivating novel approaches for image organization and segmentation. Using a 9500-image dataset, our approach is validated on each of these issues. First, we show with extensive experiments on binary and multi-class scene classification tasks, that a bag-of-visterm representation, derived from local invariant descriptors, consistently outperforms state-of-the-art approaches. Second, we show that probabilistic latent semantic analysis (PLSA) generates a compact scene representation, discriminative for accurate classification, and significantly more robust when less training data are available. Third, we have exploited the ability of PLSA to automatically extract visually meaningful aspects, to propose new algorithms for aspect-based image ranking and context-sensitive image segmentation. Pedro Quelhas, Florent Monay, Jean-Marc Odobez, Daniel Gatica-Perez, Tinne Tuytelaars, Luc Van Gool |
ICCV | 3 |
| 2005 | OCR Based Slide RetrievalabstractThis paper addresses the problem of acquiring, indexing and retrieving slides in the context of automatic oral presentation processing. Since the most suitable acquisition technique, in such a context, is the use of a framegrabber (a device capturing as images the slides displayed on a screen), the slides must be transcribed with an optical character recognition system. Retrieval experiments performed on a corpus of 570 slides (26 presentations) gathered at a workshop show that performance obtained with the OCR transcriptions are close to those obtained by extracting the text from the electronic version (pdf or ppt) of the slides (through apposite APIs). N. Daddaoua, Jean-Marc Odobez, Alessandro Vinciarelli |
ICDAR | 2 |
| 2005 | Evaluation of Multiple Cue Head Pose Estimation Algorithms in Natural EnvironementsabstractHead pose estimation is a research area which has many applications, e.g. in human computer interfaces design or in the analysis of people's focus-of-attention. The paper addresses the issue of head pose estimation, and makes two contributions. First it introduces a database of more than 2 hours of video with head pose annotation involving people engaged in office activities or meeting discussion. The database is publicly available. The second is an algorithm which couples tracking and head pose estimation in a mixed-state particle filter. The approach combines the robustness of color-based tracking by exploiting skin head/face models with the localization accuracy of texture-based head models, as demonstrated by the reported experiments Sileye O. Ba, Jean-Marc Odobez |
ICME | 2 |
| 2005 | Sports Event Recognition Using Layered HMMSabstractThe recognition of events in video data is a subject of much current interest. In this paper, we address several issues related to this topic. The first one is overfitting when very large feature spaces are used and relatively small amounts of training data are available. The second is the use of a framework that can recognise events at different time scales, as standard hidden Markov model (HMM) do not model well long-term term temporal dependencies in the data. In this paper we propose a method combining layered HMMs and an unsupervised low level clustering of the features to address these issues. Experiments conducted on the recognition task of different events in 7 rugby games demonstrates the potential of our approach with respect to standard HMM techniques coupled with a feature size reduction technique. While the current focus of this work is on events in sports videos, we believe the techniques shown here are general enough to be applied to other sources of data Mark Barnard, Jean-Marc Odobez |
ICME | 2 |
| 2005 | Multimodal multispeaker probabilistic tracking in meetingsabstractTracking speakers in multiparty conversations constitutes a fundamental task for automatic meeting analysis. In this paper, we present a probabilistic approach to jointly track the location and speaking activity of multiple speakers in a multisensor meeting room, equipped with a small microphone array and multiple uncalibrated cameras. Our framework is based on a mixed-state dynamic graphical model defined on a multiperson state-space, which includes the explicit definition of a proximity-based interaction model. The model integrates audio-visual (AV) data through a novel observation model. Audio observations are derived from a source localization algorithm. Visual observations are based on models of the shape and spatial structure of human heads. Approximate inference in our model, needed given its complexity, is performed with a Markov Chain Monte Carlo particle filter (MCMC-PF), which results in high sampling efficiency. We present results -based on an objective evaluation procedure-that show that our framework (1) is capable of locating and tracking the position and speaking activity of multiple meeting participants engaged in real conversations with good accuracy; (2) can deal with cases of visual clutter and partial occlusion; and (3) significantly outperforms a traditional sampling-based approach. Daniel Gatica-Perez, Guillaume Lathoud, Jean-Marc Odobez, Iain McCowan |
ICMI | 3 |
| 2005 | Monte Carlo video text segmentationabstractThis paper presents a probabilistic algorithm for segmenting and recognizing text embedded in video sequences based on adaptive thresholding using a Bayes filtering method. The algorithm approximates the posterior distribution of segmentation thresholds of video text by a set of weighted samples. The set of samples is initialized by applying a classical segmentation algorithm on the first video frame and further refined by random sampling under a temporal Bayesian framework. This framework allows us to evaluate a text image segmentor on the basis of recognition result instead of visual segmentation result, which is directly relevant to our character recognition task. Results on a database of 6944 images demonstrate the validity of the algorithm. Datong Chen, Jean-Marc Odobez, Jean-Philippe Thiran |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2005 | Video text recognition using sequential Monte Carlo and error voting methods
Datong Chen, Jean-Marc Odobez |
Pattern Recognit. Lett. | 2 |
| 2004 | Text detection, recognition in images and video frames
Datong Chen, Jean-Marc Odobez, Hervé Bourlard |
Pattern Recognit. | 2 |
| 2004 | A localization/verification scheme for finding text in images and video frames based on contrast independent features and machine learning methods
Datong Chen, Jean-Marc Odobez, Jean-Philippe Thiran |
Signal Process. Image Commun. | 2 |
| 2003 | An implicit motion likelihood for tracking with particle filtersabstractParticle filters are now established as the most popular method for visual tracking. Within this framework, it is generally assumed that the data are temporally independent given the sequence of object states. In this paper, we argue that in general the data are correlated, and that modeling such dependency should improve tracking robustness. To take data correlation into account, we propose a new model which can be interpreted as introducing a likelihood on implicit motion measurements. The proposed model allows to filter out visual distractors when tracking objects with generic models based on shape or color distribution representations, as shown by the reported experiments. Jean-Marc Odobez, Sileye O. Ba, Daniel Gatica-Perez |
BMVC | 1 |
| 2003 | Sequential Monte Carlo video text segmentationabstractThis paper presents a probabilistic algorithm for segmenting and recognizing text embedded in video sequences. The algorithm approximates the posterior distribution of segmentation thresholds of video text by a set of weighted samples. After initialization the set of samples is recursively refined by random sampling under a temporal Bayesian framework. The proposed methodology allows us to estimate the optimal text segmentation parameters directly in function of the string recognition results instead of segmentation quality. Results on a database of 6944 images demonstrate the validity of the algorithm. Datong Chen, Jean-Marc Odobez |
ICIP (3) | 2 |
| 2003 | Audio-visual speaker tracking with importance particle filtersabstractWe present a probabilistic method for audio-visual (AV) speaker tracking, using an uncalibrated wide-angle camera and a micro- phone array. The algorithm fuses 2-D object shape and audio information via importance particle filters (I-PFs), allowing for the asymmetrical integration of AV information in a way that efficiently exploits the complementary features of each modality. Audio localization information is used to generate an importance sampling (IS) function, which guides the random search process of a particle filter towards regions of the configuration space likely to contain the true configuration (a speaker). The measurement process integrates contour-based and audio observations, which results in reliable head tracking in realistic scenarios. We show that imperfect single modalities can be combined into an algorithm that automatically initializes and tracks a speaker, switches between multiple speakers, tolerates visual clutter, and recovers from total AV object occlusion, in the context of a multimodal meeting room. Daniel Gatica-Perez, Guillaume Lathoud, Iain McCowan, Jean-Marc Odobez, Darren Moore |
ICIP (3) | 4 |
| 2003 | A Hierarchical Keyframe User Interface for Browsing Video over the Internet
Maël Guillemot, Pierre Wellner, Daniel Gatica-Perez, Jean-Marc Odobez |
INTERACT | 4 |
| 2002 | Robust video text segmentation and recognition with multiple hypothesesabstractA method for segmenting and recognizing text embedded in video and images is proposed. Multiple segmentation of the same text region is performed, thus producing multiple hypotheses of binary text images. The segmentation algorithm is stated as a statistical labeling and is based on a Markov random field (MRF) model of the label map. Background regions in each hypothesis are then removed by performing a connected component analysis and by enforcing a more stringent constraint (called GCC - grayscale consistency constraint) on the text characters' grayscale values using a robust 1D-median operator. Each text image hypothesis is then processed by optical character recognition (OCR) software. The final result is then selected from the set of output strings. Results show that both the use of multiple hypotheses and the GCC significantly improve the results. Jean-Marc Odobez, Datong Chen |
ICIP (2) | 1 |
| 2000 | Analysis of Doppler Ultrasound Time Frequency Images Using Deformable ModelsabstractDoppler ultrasound is a widely used technique to study the blood flow velocity and evaluate the severity of an arterial stenosis. The envelope of maximal frequencies in the power spectrum of the signal represents important information for the characterisation of the blood flow. Classical signal processing techniques usually rely on local observations (i.e. computed at a given time instant). Therefore, they are quite unreliable in the presence of low-level signals and/or in noisy environments. In this article, we propose a new automatic method to estimate the maximal frequency profile. It is derived from image processing techniques and based on deformable models. The application of this more global method to real noisy signals gives very promising results, according to a clinical expert and in comparison with the results that we obtain with other classical methods. Jean-Marc Odobez, Emmanuel Roy, Pierre Abraham |
ICIP | 1 |
| 1998 | Direct incremental model-based image motion segmentation for video analysis
Jean-Marc Odobez, Patrick Bouthemy |
Signal Process. | 1 |
| 1997 | Adaptive motion-compensated wavelet filtering for image sequence codingabstractThis paper deals with new advances made in the field of discrete spatio-temporal filters applied to digital image sequences. Within time-varying images, the temporal correlation of the information is folded within the spatio-temporal domain by motions originating from both camera and object displacements. The spatio-temporal video information can be therefore reformulated in terms of motion trajectories and intensity variations along these trajectories. To yield that signal description, the spatio-temporal scenes will be segmented according to motion. Motion-compensated temporal filters have been used as convolutional filters applied along the assumed motion trajectories. Spectral interpretations show the efficiency of motion-compensated filtering for video signals. As a matter of fact, the whole signal analysis performed in this paper also includes a spatial filtering to achieve a complete spatio-temporal (2-D+T) decomposition. Motion-compensated filtering leads to multiresolution applications. It leads to optimum and adaptive signal-to-noise decomposition procedures based on the temporal correlative content. Such properties allow enhancing tasks like temporal interpolation, image sequence smoothing, and restoration. Simulation results are presented in this paper to illustrate the field of image sequence coding. Jean-Pierre Leduc, Jean-Marc Odobez, Claude Labit |
IEEE Trans. Image Process. | 2 |
| 1995 | Direct Model-Based Image Motion Segmentation for Dynamic Scene Analysis
Jean-Marc Odobez, Patrick Bouthemy |
ACCV | 1 |
| 1995 | Motion-compensated adaptive wavelet filtering for image sequence processingabstractThe paper presents new approaches in the field of motion-compensated spatio-temporal filters applied to digital image sequences. In time-varying imagery, the temporal correlation information of pixel intensities is folded by motions which may originate from both camera and object displacements. Motion-compensated filters are defined as temporal filters applied along assumed motion trajectories. As a matter of fact, the paper deals with three-dimensional spatio-temporal filters and aims at generalizing the motion-compensated temporal filtering process as the product of two distinct operators. The first operator depends only on the estimated motion parameters derived from both motion-based image segmentations and parametric affine modelings of regions in motion. The second operator analyzes only the correlations of image-by-image intensities measured along the assumed motion trajectories. Multiresolution filters or wavelets may be consequently applied along the motion trajectories to produce optimum and adaptive resulting procedures for purposes like spatio-temporal prediction, interpolation and smoothing. In the paper, applications are provided to cover the field of image sequence coding and interpolation. Jean-Pierre Leduc, Jean-Marc Odobez, Claude Labit |
ICASSP | 2 |
| 1995 | Determination of singular points in 2D deformable flow fieldsabstractDigital image analysis appears to be more and more relevant to the study of physical phenomena involving fluid motion, and of their evolution over time. In that context, 2D deformable motion analysis is one of the important issues to be investigated. The interpretation of such deformable 2D flow fields can generally be stated as the characterization of linear models provided that first order approximations are considered in an adequate neighborhood of so-called singular points, where the velocity becomes null. This paper describes an efficient method, based on a statistical approach, which explicitly addresses these problems, and allows us to locate, characterize and track such singular points in an image sequence. It does not require the prior computation of the velocity field. The method has been validated by experiments carried out with synthetic and real examples corresponding to meteorological image sequences. In fact, the described approach can be of interest in different applications dealing with the characterization of vector fields. Mariette Maurizot, Patrick Bouthemy, Bernard Delyon, Anatoli B. Juditsky, Jean-Marc Odobez |
ICIP (3) | 5 |
| 1995 | MRF-based motion segmentation exploiting a 2D motion model robust estimationabstractThis paper deals with motion-segmentation, that is, with the partitioning of the image into regions of homogeneous motion. Here, homogeneous means that in each region a 2D polynomial model (e.g. an affine one) is able to describe at each location the underlying "true" motion with a predefined precision /spl eta/. However, no estimation of this true motion field is required. The motion models are computed using a multiresolution robust estimator. Therefore, as opposed to almost all other motion-segmentation scheme, the motion model of a given region only needs to be estimated once at a given time instant. Moreover, the determination of the boundaries between the different regions, which is stated as a statistical regularization based on a multiscale Markov random field (MRF) modeling, only requires one pass. Finally, thanks to the definition of an explicit detection step of areas where the error between the underlying motion and the one given by the estimated models is not within the precision /spl eta/, we are able to get a good segmentation from the very beginning of the sequence, and to manage the appearance of new objects in the scene, as well as the momentary increase in the complexity of motion in already existing regions. Results obtained on many real image sequences have validated our approach. Jean-Marc Odobez, Patrick Bouthemy |
ICIP (3) | 1 |
| 1995 | Robust Multiresolution Estimation of Parametric Motion Models
Jean-Marc Odobez, Patrick Bouthemy |
J. Vis. Commun. Image Represent. | 1 |
| 1994 | A ROI Approach for Hybrid Image Sequence CodingabstractIn this paper we present an approach of selective compression of an image sequence based on a priori selection of region(s) of interest (ROI). This method relies on a given motion-based segmentation analysis. The problem of selective compression based on the concept of inhomogeneous spatial reconstruction quality is considered in the context of hybrid DPCM subband coding. Both spatial and frequency localization of the subband representation are explicitly used. Hierarchical compression is applied through adaptive quantization in the spatio-frequency domain. A weighted distortion metric is used to introduce both a priori and velocity-based visual masking.> Eric Nguyen, Claude Labit, Jean-Marc Odobez |
ICIP (3) | 3 |
| 1994 | Detection of Multiple Moving Objects using Multiscale MRF with Camera Motion CompensationabstractWe address the problem of detecting moving objects from a moving camera. The apparent flow field induced by the camera motion is modeled by a 2D parametric motion model and compensated for using the values of the parameters estimated by a multiresolution robust method. Motion detection is achieved through a statistical regularization approach based on multiscale Markov random field (MRF) models. Particular attention has been paid to the definition of the energy function involved and to the considered observations. This method has been validated by experiments carried out on different real image sequences.> Jean-Marc Odobez, Patrick Bouthemy |
ICIP (2) | 1 |