EDBT 2026 Demo / reviewers in the wild / expert
Emmanuel Dellandréa
dblp:79/5140
· DBLP profile ↗
34ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0001-7346-228XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 since 2021Human-computer interaction and ubiquitous computing · 4Systems, architecture and hardware · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Image recognition and object detection · 28% Robot manipulation · 16% Transfer learning and domain adaptation · 14% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
exploration |
0.7 | 1 | 2023 | Cell-Free Latent Go-Explore · ICML 2023 |
Machine learning › Representation and self-supervised learning › representation learning
latent representation learning |
0.7 | 1 | 2023 | Cell-Free Latent Go-Explore · ICML 2023 |
Machine learning › Transfer learning and domain adaptation
knowledge transfer |
0.6 | 2 | 2018 | Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2018 Large Scale Semi-Supervised Object Detection Using Visual and Semantic Knowledge Transfer · CVPR 2016 |
Computer vision › Image recognition and object detection › object detection
semi-supervised object detection |
0.6 | 2 | 2018 | Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2018 Large Scale Semi-Supervised Object Detection Using Visual and Semantic Knowledge Transfer · CVPR 2016 |
Robotics › Robot manipulation › grasping
grasp detection |
0.5 | 1 | 2021 | Scoring Graspability based on Grasp Regression for Better Grasp Prediction · ICRA 2021 |
Robotics › Robot manipulation
grasping |
0.5 | 1 | 2021 | Scoring Graspability based on Grasp Regression for Better Grasp Prediction · ICRA 2021 |
Computer vision › Image recognition and object detection › object localization
instance localization |
0.4 | 1 | 2020 | Deep Multicameral Decoding for Localizing Unoccluded Object Instances from a Single RGB Image · Int. J. Comput. Vis. 2020 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.4 | 1 | 2020 | Deep Multicameral Decoding for Localizing Unoccluded Object Instances from a Single RGB Image · Int. J. Comput. Vis. 2020 |
Machine learning › Transfer learning and domain adaptation
cross-domain transfer |
0.3 | 1 | 2018 | Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2018 |
Computer vision › Vision and language › cross-modal alignment
visual-semantic similarity |
0.3 | 1 | 2018 | Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2018 |
Computer vision › Image recognition and object detection › object detection › part-based object detection
deformable part model |
0.3 | 1 | 2017 | Weakly Supervised Learning of Deformable Part-Based Models for Object Detection via Region Proposals · IEEE Trans. Multim. 2017 |
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection |
0.3 | 1 | 2017 | Weakly Supervised Learning of Deformable Part-Based Models for Object Detection via Region Proposals · IEEE Trans. Multim. 2017 |
Computer vision › Vision and language
image captioning |
0.2 | 1 | 2015 | Combining Geometric, Textual and Visual Features for Predicting Prepositions in Image Descriptions · EMNLP 2015 |
Computer vision › 3D vision › 3d scene understanding
spatial relation understanding |
0.2 | 1 | 2015 | Combining Geometric, Textual and Visual Features for Predicting Prepositions in Image Descriptions · EMNLP 2015 |
Computer vision › Image recognition and object detection › object detection
object proposal generation |
0.1 | 1 | 2017 | Weakly Supervised Learning of Deformable Part-Based Models for Object Detection via Region Proposals · IEEE Trans. Multim. 2017 |
Computer vision › Vision and language
visual entity recognition |
0.1 | 1 | 2015 | Combining Geometric, Textual and Visual Features for Predicting Prepositions in Image Descriptions · EMNLP 2015 |
Methods — techniques the papers use, named apart from their topics
latent representation learning · 0.7go-explore · 0.7CNN · 0.6loss function design · 0.5deep neural network · 0.5deep multicameral decoding · 0.4knowledge transfer · 0.3region proposal · 0.3deformable part-based model · 0.3semantic relatedness · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XmoPipe: A Pipeline for Large-Scale In-the-Wild Human Motion Dataset ConstructionabstractABSTRACT Large‐scale human motion datasets are essential for training robust motion models for analysis, synthesis, and understanding. While marker‐based motion capture provides precise data, it is costly and limited in scale and diversity. Recent advances in monocular motion capture and video‐language understanding open the way to extract plausible motion from unconstrained online videos. We present a scalable pipeline for constructing in‐the‐wild human motion datasets. From a few keywords, the system retrieves videos, extracts 3D body and facial motion, and generates high‐level textual descriptions. The pipeline is flexible, enabling targeted collection of various motions, multi‐person interactions, or expressive behaviors. We demonstrate its quality by training motion reconstruction and motion generation models, showing performance approaching that of models trained on traditional motion capture datasets, along with strong cross‐dataset generalization. Code and motion data are available at https://github.com/NatSalaz/xmopipe . Nathan Salazar, Emmanuel Dellandréa, Mathieu Lefort, Alexandre Meyer |
Comput. Animat. Virtual Worlds | 2 |
| 2025 | Online Continual Learning of Diffusion Models: Multi-Mode Adaptive Generative DistillationabstractContinual learning typically relies on storing real data, which is impractical in privacy-sensitive settings. Generative replay with diffusion models offers a high-fidelity alternative. However, in online continual learning (OCL), these models struggle with catastrophic forgetting and incur high computational costs from frequent updates and sampling. Existing distillation methods reduce generation steps but rely on a fixed teacher model, limiting their effectiveness as data distributions evolve. To address these, we introduce Multi-Mode Adaptive Generative Distillation (MAGD), which incorporates two innovative techniques: Noisy Intermediate Generative Distillation (NIGD) and SNR-Guided Generative Distillation (SGGD). NIGD leverages intermediate noisy images, created during the reverse process rather than by adding noise post-generation, to enhance knowledge transfer. SGGD uses a signal-to-noise ratio (SNR) based threshold to optimize the sampling of time steps, reducing unnecessary generation. Guided by an Exponential Moving Average (EMA) teacher, MAGD effectively mitigates catastrophic forgetting as it adapts to new data streams. Experiments on Fashion-MNIST, CIFAR-10, and CIFAR-100 show that MAGD reduces generation overhead by up to 25% relative to standard generative distillation and 92% compared to DDGR-1000, while maintaining generating quality. Furthermore, in class-conditioned diffusion models, MAGD outperforms memory-based methods in terms of classification accuracy. Matthieu Grard, Emmanuel Dellandréa, Liming Chen 0002 |
ICIP | 3 |
| 2024 | Imbalanced Data Robust Online Continual Learning Based on Evolving Class Aware Memory Selection and Built-In Contrastive Representation LearningabstractWe introduce Memory Selection with Contrastive Learning (MSCL), an advanced Continual Learning (CL) approach, addressing challenges in dynamic and imbalanced environments. MSCL combines Feature-Distance Based Sample Selection (FDBS) for memory management, focusing on inter-class similarities and intra-class diversity, with a contrastive learning loss (IWL) for adaptive data representation. Our evaluations on datasets like MNIST, Cifar-100, miniImageNet, PACS, and DomainNet show that MSCL not only competes with but often surpasses existing memory-based CL methods, particularly in imbalanced scenarios, enhancing both balanced and imbalanced learning performance. Emmanuel Dellandréa, Matthieu Grard, Liming Chen 0002 |
ICIP | 2 |
| 2023 | Cell-Free Latent Go-ExploreabstractIn this paper, we introduce Latent Go-Explore (LGE), a simple and general approach based on the Go-Explore paradigm for exploration in reinforcement learning (RL). Go-Explore was initially introduced with a strong domain knowledge constraint for partitioning the state space into cells. However, in most real-world scenarios, drawing domain knowledge from raw observations is complex and tedious. If the cell partitioning is not informative enough, Go-Explore can completely fail to explore the environment. We argue that the Go-Explore approach can be generalized to any environment without domain knowledge and without cells by exploiting a learned latent representation. Thus, we show that LGE can be flexibly combined with any strategy for learning a latent representation. Our results indicate that LGE, although simpler than Go-Explore, is more robust and outperforms state-of-the-art algorithms in terms of pure exploration on multiple hard-exploration environments including Montezuma's Revenge. The LGE implementation is available as open-source at https://github.com/qgallouedec/lge. Quentin Gallouédec, Emmanuel Dellandréa |
ICML | 2 |
| 2022 | Look Beyond Bias with Entropic Adversarial Data AugmentationabstractDeep neural networks do not discriminate between spurious and causal patterns, and will only learn the most predictive ones while ignoring the others. This shortcut learning behaviour is detrimental to a network’s ability to generalize to an unknown test-time distribution in which the spurious correlations do not hold anymore. Debiasing methods were developed to make networks robust to such spurious biases but require to know in advance if a dataset is biased and make heavy use of minority counter-examples that do not display the majority bias of their class. In this paper, we argue that such samples should not be necessarily needed because the “hidden” causal information is often also contained in biased images. To study this idea, we propose 3 publicly released synthetic classification benchmarks, exhibiting predictive classification shortcuts, each of a different and challenging nature, without any minority samples acting as counter-examples. First, we investigate the effectiveness of several state-of-the-art strategies on our benchmarks and show that they do not yield satisfying results on them. Then, we propose an architecture able to succeed on our benchmarks, despite their unusual properties, using an entropic adversarial data augmentation training scheme. An encoder-decoder architecture is tasked to produce images that are not recognized by a classifier, by maximizing the conditional entropy of its outputs, and keep as much as possible of the initial content. A precise control of the information destroyed, via a disentangling process, enables us to remove the shortcut and leave everything else intact. Furthermore, results competitive with the state-of-the-art on the BAR dataset ensure the applicability of our method in real-life situations. Thomas Duboudin, Emmanuel Dellandréa, Corentin Abgrall, Gilles Hénaff, Liming Chen 0002 |
ICPR | 2 |
| 2021 | Scoring Graspability based on Grasp Regression for Better Grasp PredictionabstractGrasping objects is one of the most important abilities that a robot needs to master in order to interact with its environment. Current state-of-the-art methods rely on deep neural networks trained to jointly predict a graspability score together with a regression of an offset with respect to grasp reference parameters. However, these two predictions are performed independently, which can lead to a decrease in the actual graspability score when applying the predicted offset. Therefore, in this paper, we extend a state-of-the-art neural network with a scorer that evaluates the graspability of a given position, and introduce a novel loss function which correlates regression of grasp parameters with graspability score. We show that this novel architecture improves performance from 82.13% for a state-of-the-art grasp detection network to 85.74% on Jacquard dataset. When the learned model is transferred onto a real robot, the proposed method correlating graspability and grasp regression achieves a 92.4% rate compared to 88.1% for the baseline trained without the correlation. Amaury Depierre, Emmanuel Dellandréa, Liming Chen 0002 |
ICRA | 2 |
| 2020 | Deep Multicameral Decoding for Localizing Unoccluded Object Instances from a Single RGB Image
Matthieu Grard, Emmanuel Dellandréa, Liming Chen 0002 |
Int. J. Comput. Vis. | 2 |
| 2018 | Jacquard: A Large Scale Dataset for Robotic Grasp DetectionabstractGrasping skill is a major ability that a wide number of real-life applications require for robotisation. State-of-the-art robotic grasping methods perform prediction of object grasp locations based on deep neural networks. However, such networks require huge amount of labeled data for training making this approach often impracticable in robotics. In this paper, we propose a method to generate a large scale synthetic dataset with ground truth, which we refer to as the Jacquard grasping dataset. Jacquard is built on a subset of ShapeNet, a large CAD models dataset, and contains both RGB-D images and annotations of successful grasping positions based on grasp attempts performed in a simulated environment. We carried out experiments using an off-the-shelf CNN, with three different evaluation metrics, including real grasping robot trials. The results show that Jacquard enables much better generalization skills than a human labeled dataset thanks to its diversity of objects and grasping positions. For the purpose of reproducible research in robotics, we are releasing along with the Jacquard dataset a web interface for researchers to evaluate the successfulness of their grasping position detections using our dataset. Amaury Depierre, Emmanuel Dellandréa, Liming Chen 0002 |
IROS | 2 |
| 2018 | Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object DetectionabstractDeep CNN-based object detection systems have achieved remarkable success on several large-scale object detection benchmarks. However, training such detectors requires a large number of labeled bounding boxes, which are more difficult to obtain than image-level annotations. Previous work addresses this issue by transforming image-level classifiers into object detectors. This is done by modeling the differences between the two on categories with both image-level and bounding box annotations, and transferring this information to convert classifiers to detectors for categories without bounding box annotations. We improve this previous work by incorporating knowledge about object similarities from visual and semantic domains during the transfer process. The intuition behind our proposed method is that visually and semantically similar categories should exhibit more common transferable properties than dissimilar categories, e.g. a better detector would result by transforming the differences between a dog classifier and a dog detector onto the cat class, than would by transforming from the violin class. Experimental results on the challenging ILSVRC2013 detection dataset demonstrate that each of our proposed object similarity based knowledge transfer methods outperforms the baseline methods. We found strong evidence that visual similarity and semantic relatedness are complementary for the task, and when combined notably improve detection, achieving state-of-the-art detection performance in a semi-supervised setting. Yuxing Tang, Josiah Wang, Boyang Gao, Emmanuel Dellandréa, Robert J. Gaizauskas, Liming Chen 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Affective Video Content Analysis: A Multidisciplinary InsightabstractIn our present society, the cinema has become one of the major forms of entertainment providing unlimited contexts of emotion elicitation for the emotional needs of human beings. Since emotions are universal and shape all aspects of our interpersonal and intellectual experience, they have proved to be a highly multidisciplinary research field, ranging from psychology, sociology, neuroscience, etc., to computer science. However, affective multimedia content analysis work from the computer science community benefits but little from the progress achieved in other research fields. In this paper, a multidisciplinary state-of-the-art for affective movie content analysis is given, in order to promote and encourage exchanges between researchers from a very wide range of fields. In contrast to other state-of-the-art papers on affective video content analysis, this work confronts the ideas and models of psychology, sociology, neuroscience, and computer science. The concepts of aesthetic emotions and emotion induction, as well as the different representations of emotions are introduced, based on psychological and sociological theories. Previous global and continuous affective video content analysis work, including video emotion recognition and violence detection, are also presented in order to point out the limitations of affective video content analysis work. Yoann Baveye, Christel Chamaret, Emmanuel Dellandréa, Liming Chen 0002 |
IEEE Trans. Affect. Comput. | 3 |
| 2018 | Discriminative Transfer Learning Using Similarities and DissimilaritiesabstractTransfer learning (TL) aims at solving the problem of learning an effective classification model for a target category, which has few training samples, by leveraging knowledge from source categories with far more training data. We propose a new discriminative TL (DTL) method, combining a series of hypotheses made by both the model learned with target training samples and the additional models learned with source category samples. Specifically, we use the sparse reconstruction residual as a basic discriminant and enhance its discriminative power by comparing two residuals from a positive and a negative dictionary. On this basis, we make use of similarities and dissimilarities by choosing both positively correlated and negatively correlated source categories to form additional dictionaries. A new Wilcoxon-Mann-Whitney statistic-based cost function is proposed to choose the additional dictionaries with unbalanced training data. Also, two parallel boosting processes are applied to both the positive and negative data distributions to further improve classifier performance. On two different image classification databases, the proposed DTL consistently outperforms other state-of-the-art TL methods while at the same time maintaining very efficient runtime. Ying Lu 0007, Liming Chen 0002, Alexandre Saidi, Emmanuel Dellandréa, Yunhong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Weakly Supervised Learning of Deformable Part-Based Models for Object Detection via Region ProposalsabstractThe success of deformable part-based models (DPMs) for visual object detection relies on a large number of labeled bounding boxes. With only image-level annotations, our goal is to propose a model enhancing the weakly supervised DPMs by emphasizing the importance of location and size of the initial class-specific root filter. To adaptively select a discriminative set of candidate bounding boxes as this root filter estimate, first, we explore the generic objectness measurement to combine the most salient regions and “good” region proposals. Second, we propose learning of the latent class label of each candidate window as a binary classification problem, by training category-specific classifiers used to coarsely classify a candidate window into either a target object or a nontarget class. Finally, we design a flexible enlarging-and-shrinking postprocessing procedure to modify the DPMs outputs, which can effectively match the approximative object aspect ratios and further improve final accuracy. Extensive experimental results on the challenging PASCAL Visual Object Class 2007 and the Microsoft Common Objects in Context 2014 dataset demonstrate that our proposed framework is effective for initialization of the DPM's root filter. It also shows competitive final localization performance with state-of-the-art weakly supervised object detection methods, particularly for the object categories that are relatively salient in the images and deformable in structures. Yuxing Tang, Emmanuel Dellandréa, Liming Chen 0002 |
IEEE Trans. Multim. | 3 |
| 2016 | Large Scale Semi-Supervised Object Detection Using Visual and Semantic Knowledge TransferabstractDeep CNN-based object detection systems have achieved remarkable success on several large-scale object detection benchmarks. However, training such detectors requires a large number of labeled bounding boxes, which are more difficult to obtain than image-level annotations. Previous work addresses this issue by transforming image-level classifiers into object detectors. This is done by modeling the differences between the two on categories with both imagelevel and bounding box annotations, and transferring this information to convert classifiers to detectors for categories without bounding box annotations. We improve this previous work by incorporating knowledge about object similarities from visual and semantic domains during the transfer process. The intuition behind our proposed method is that visually and semantically similar categories should exhibit more common transferable properties than dissimilar categories, e.g. a better detector would result by transforming the differences between a dog classifier and a dog detector onto the cat class, than would by transforming from the violin class. Experimental results on the challenging ILSVRC2013 detection dataset demonstrate that each of our proposed object similarity based knowledge transfer methods outperforms the baseline methods. We found strong evidence that visual similarity and semantic relatedness are complementary for the task, and when combined notably improve detection, achieving state-of-the-art detection performance in a semi-supervised setting. Yuxing Tang, Josiah Wang, Boyang Gao, Emmanuel Dellandréa, Robert J. Gaizauskas, Liming Chen 0002 |
CVPR | 4 |
| 2016 | Automatic 2.5-D Facial Landmarking and Emotion Annotation for Social Interaction AssistanceabstractPeople with low vision, Alzheimer's disease, and autism spectrum disorder experience difficulties in perceiving or interpreting facial expression of emotion in their social lives. Though automatic facial expression recognition (FER) methods on 2-D videos have been extensively investigated, their performance was constrained by challenges in head pose and lighting conditions. The shape information in 3-D facial data can reduce or even overcome these challenges. However, high expenses of 3-D cameras prevent their widespread use. Fortunately, 2.5-D facial data from emerging portable RGB-D cameras provide a good balance for this dilemma. In this paper, we propose an automatic emotion annotation solution on 2.5-D facial data collected from RGB-D cameras. The solution consists of a facial landmarking method and a FER method. Specifically, we propose building a deformable partial face model and fit the model to a 2.5-D face for localizing facial landmarks automatically. In FER, a novel action unit (AU) space-based FER method has been proposed. Facial features are extracted using landmarks and further represented as coordinates in the AU space, which are classified into facial expressions. Evaluated on three publicly accessible facial databases, namely EURECOM, FRGC, and Bosphorus databases, the proposed facial landmarking and expression recognition methods have achieved satisfactory results. Possible real-world applications using our algorithms have also been discussed. Xi Zhao 0001, Jianhua Zou, Huibin Li 0001, Emmanuel Dellandréa, Ioannis A. Kakadiaris, Liming Chen 0002 |
IEEE Trans. Cybern. | 4 |
| 2015 | Deep learning vs. kernel methods: Performance for emotion prediction in videosabstractRecently, mainly due to the advances of deep learning, the performances in scene and object recognition have been progressing intensively. On the other hand, more subjective recognition tasks, such as emotion prediction, stagnate at moderate levels. In such context, is it possible to make affective computational models benefit from the breakthroughs in deep learning? This paper proposes to introduce the strength of deep learning in the context of emotion prediction in videos. The two main contributions are as follow: (i) a new dataset, composed of 30 movies under Creative Commons licenses, continuously annotated along the induced valence and arousal axes (publicly available) is introduced, for which (ii) the performance of the Convolutional Neural Networks (CNN) through supervised fine-tuning, the Support Vector Machines for Regression (SVR) and the combination of both (Transfer Learning) are computed and discussed. To the best of our knowledge, it is the first approach in the literature using CNNs to predict dimensional affective scores from videos. The experimental results show that the limited size of the dataset prevents the learning or finetuning of CNN-based frameworks but that transfer learning is a promising solution to improve the performance of affective movie content analysis frameworks as long as very large datasets annotated along affective dimensions are not available. Yoann Baveye, Emmanuel Dellandréa, Christel Chamaret, Liming Chen 0002 |
ACII | 2 |
| 2015 | Combining Geometric, Textual and Visual Features for Predicting Prepositions in Image DescriptionsabstractWe investigate the role that geometric, textual and visual features play in the task of predicting a preposition that links two visual entities depicted in an image.The task is an important part of the subsequent process of generating image descriptions.We explore the prediction of prepositions for a pair of entities, both in the case when the labels of such entities are known and unknown.In all situations we found clear evidence that all three features contribute to the prediction task. Arnau Ramisa, Josiah Wang, Ying Lu 0007, Emmanuel Dellandréa, Francesc Moreno-Noguer, Robert J. Gaizauskas |
EMNLP | 4 |
| 2015 | LIRIS-ACCEDE: A Video Database for Affective Content AnalysisabstractResearch in affective computing requires ground truth data for training and benchmarking computational models for machine-based emotion understanding. In this paper, we propose a large video database, namely LIRIS-ACCEDE, for affective content analysis and related applications, including video indexing, summarization or browsing. In contrast to existing datasets with very few video resources and limited accessibility due to copyright constraints, LIRIS-ACCEDE consists of 9,800 good quality video excerpts with a large content diversity. All excerpts are shared under creative commons licenses and can thus be freely distributed without copyright issues. Affective annotations were achieved using crowdsourcing through a pair-wise video comparison protocol, thereby ensuring that annotations are fully consistent, as testified by a high inter-annotator agreement, despite the large diversity of raters' cultural backgrounds. In addition, to enable fair comparison and landmark progresses of future affective computational models, we further provide four experimental protocols and a baseline for prediction of emotions using a large set of both visual and audio features. The dataset (the video clips, annotations, features and protocols) is publicly available at: http://liris-accede.ec-lyon.fr/. Yoann Baveye, Emmanuel Dellandréa, Christel Chamaret, Liming Chen 0002 |
IEEE Trans. Affect. Comput. | 2 |
| 2014 | Fusing generic objectness and deformable part-based models for weakly supervised object detectionabstractIn the context of lack of object-level annotation, we propose a model that enhances the weakly supervised deformable part model (DPM) by emphasizing the importance of size and aspect ratio of the initial class-specific root filter. For each image, to extract a reliable bounding box as this root filter estimate, we explore the generic objectness measurement to obtain a reference window based on the most salient region, and select a small set of candidate windows by adaptive thresholding and greedy Non-Maximum Suppression (NMS). The initial root filter estimate is decided by optimizing the score of overlap between the reference box and candidate boxes, as well as their corresponding objectness score. Then the derived window is treated as a positive training window for DPM training. Finally, we design a flexible enlarging-and-shrinking post-processing procedure to modify the output of DPM, which can effectively fit to the aspect ratio of the object and further improve the final accuracy. Experimental results on the challenging PASCAL VOC 2007 database demonstrate that our proposed framework is effective and competitive with the state-of-the-arts. Yuxing Tang, Emmanuel Dellandréa, Simon Masnou, Liming Chen 0002 |
ICIP | 3 |
| 2014 | Evaluation of video activity localizations integrating quality and quantity measurements
Christian Wolf 0001, Eric Lombardi, Julien Mille, Oya Çeliktutan, Mingyuan Jiu, Emre Dogan, Gonen Eren, Moez Baccouche, Emmanuel Dellandréa, Charles-Edmond Bichot, Christophe Garcia, Bülent Sankur |
Comput. Vis. Image Underst. | 9 |
| 2013 | A Large Video Data Base for Computational Models of Induced EmotionabstractTo contribute to the need for emotional databases and affective tagging, the LIRIS-ACCEDE is proposed in this paper. LIRIS-ACCEDE is an Annotated Creative Commons Emotional DatabasE composed of 9800 video clips extracted from 160 movies shared under Creative Commons licenses. It allows to make this database publicly available without copyright issues. The 9800 video clips (each 8-12 seconds long) are sorted along the induced valence axis, from the video perceived the most negatively to the video perceived the most positively. The annotation was carried out by 1518 annotators from 89 different countries using crowd sourcing. A baseline late fusion scheme using ground truth from annotations is computed to predict emotion categories in video clips. Yoann Baveye, Jean-Noel Bettinelli, Emmanuel Dellandréa, Liming Chen 0002, Christel Chamaret |
ACII | 3 |
| 2013 | Multimodal recognition of visual concepts using histograms of textual concepts and selective weighted late fusion scheme
Ningning Liu, Emmanuel Dellandréa, Liming Chen 0002, Chao Zhu 0003, Yu Zhang 0052, Charles-Edmond Bichot, Stéphane Bres, Bruno Tellez |
Comput. Vis. Image Underst. | 2 |
| 2013 | A unified probabilistic framework for automatic 3D facial expression analysis based on a Bayesian belief inference and statistical feature models
Xi Zhao 0001, Emmanuel Dellandréa, Jianhua Zou, Liming Chen 0002 |
Image Vis. Comput. | 2 |
| 2011 | Associating Textual Features with Visual Ones to Improve Affective Image Classification
Ningning Liu, Emmanuel Dellandréa, Bruno Tellez, Liming Chen 0002 |
ACII (1) | 2 |
| 2011 | Reconstructive and Discriminative Sparse Representation for Visual Object CategorizationabstractInternational audience Huanzhang Fu, Emmanuel Dellandréa, Liming Chen 0002 |
BMVC | 2 |
| 2011 | Accurate Landmarking of Three-Dimensional Facial Data in the Presence of Facial Expressions and Occlusions Using a Three-Dimensional Statistical Facial Feature ModelabstractThree-dimensional face landmarking aims at automatically localizing facial landmarks and has a wide range of applications (e.g., face recognition, face tracking, and facial expression analysis). Existing methods assume neutral facial expressions and unoccluded faces. In this paper, we propose a general learning-based framework for reliable landmark localization on 3-D facial data under challenging conditions (i.e., facial expressions and occlusions). Our approach relies on a statistical model, called 3-D statistical facial feature model, which learns both the global variations in configurational relationships between landmarks and the local variations of texture and geometry around each landmark. Based on this model, we further propose an occlusion classifier and a fitting algorithm. Results from experiments on three publicly available 3-D face databases (FRGC, BU-3-DFE, and Bosphorus) demonstrate the effectiveness of our approach, in terms of landmarking accuracy and robustness, in the presence of expressions and occlusions. Xi Zhao 0001, Emmanuel Dellandréa, Liming Chen 0002, Ioannis A. Kakadiaris |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2010 | Bayesian GOETHE TrackingabstractOcclusions pose serious challenges when tracking multiple targets. By severly changing the measurement, they imply strong inter-target dependencies. Exact computation of these dependencies is not feasible. The GOETHE approximations preserve much of the information while staying computationally affordable. Sebastian J. Wirkert, Emmanuel Dellandréa, Liming Chen 0002 |
ICPR | 2 |
| 2010 | Automatic 3D Facial Expression Recognition Based on a Bayesian Belief Net and a Statistical Facial Feature ModelabstractAutomatic facial expression recognition on 3D face data is still a challenging problem. In this paper we propose a novel approach to perform expression recognition automatically and flexibly by combining a Bayesian Belief Net (BBN) and Statistical facial feature models (SFAM). A novel BBN is designed for the specific problem with our proposed parameter computing method. By learning global variations in face landmark configuration (morphology) and local ones in terms of texture and shape around landmarks, morphable Statistic Facial feature Model (SFAM) allows not only to perform an automatic landmarking but also to compute the belief to feed the BBN. Tested on the public 3D face expression database BU-3DFE, our automatic approach allows to recognize expressions successfully, reaching an average recognition rate over 82%. Xi Zhao 0001, Di Huang 0001, Emmanuel Dellandréa, Liming Chen 0002 |
ICPR | 3 |
| 2010 | Multi-stage classification of emotional speech motivated by a dimensional emotion model
Zhongzhe Xiao, Emmanuel Dellandréa, Weibei Dou, Liming Chen 0002 |
Multim. Tools Appl. | 2 |
| 2009 | Image Categorization Using ESFS: A New Embedded Feature Selection Method Based on SFS
Huanzhang Fu, Zhongzhe Xiao, Emmanuel Dellandréa, Weibei Dou, Liming Chen 0002 |
ACIVS | 3 |
| 2009 | A 3D Statistical Facial Feature Model and Its Application on Locating Facial Landmarks
Xi Zhao 0001, Emmanuel Dellandréa, Liming Chen 0002 |
ACIVS | 2 |
| 2009 | A People Counting System Based on Face Detection and Tracking in a VideoabstractVision-based people counting systems have wide potential applications including video surveillance and public resources management. Most works in the literature rely on moving object detection and tracking, assuming that all moving objects are people. In this paper, we present our people counting approach based on face detection, tracking and trajectory classification. While we have used a standard face detector, we achieve face tracking combining a new scale invariant Kalman filter with kernel based tracking algorithm. From each potential face trajectory an angle histogram of neighboring points is then extracted. Finally, an Earth Mover's Distance-based K-NN classification discriminates true face trajectories from the false ones. Experimented on a video dataset of more than 160 potential people trajectories, our approach displays an accuracy rate up to 93%. Xi Zhao 0001, Emmanuel Dellandréa, Liming Chen 0002 |
AVSS | 2 |
| 2009 | Visual Object Categorization via Sparse RepresentationabstractIn this paper, we consider the problem of classifying a real world image to the corresponding object class based on its visual content via sparse representation, which is originally used as a powerful tool for acquiring, representing and compressing high-dimensional signals. Assuming the intuitive hypothesis that an image could be represented by a linear combination of the training images from the same class, we propose a novel approach for visual object categorization in which a sparse representation of the image is first of all obtained by solving a L1 (or L0)-minimization problem and then fed into a traditional classifier such as Support Vector Machine (SVM) to finally perform the specified task. Experimental results obtained on the SIMPLIcity database have shown that this new approach can improve the classification performance compared to standard SVM using directly features extracted from the image. Huanzhang Fu, Chao Zhu 0003, Emmanuel Dellandréa, Charles-Edmond Bichot, Liming Chen 0002 |
ICIG | 3 |
| 2009 | Precise 2.5D facial landmarking via an analysis by synthesis approachabstract3D face landmarking aims at automatic localization of 3D facial features and has a wide range of applications, including face recognition, face tracking, facial expression analysis. Methods so far developed for pure 2D texture images were shown sensitive to lighting condition changes. In this paper, we present a statistical model-based technique for accurate 3D face landmarking, thus using an ¿analysis by synthesis¿ approach. Our model learns from a training set both variations of global face shapes as well as the local ones in terms of scale-free texture and range patches around each landmark. Given a shape instance, local regions of a new face can be approximated by synthesizing texture and range instances using respectively the texture and range models. By optimizing an objective function describing the similarity of the new face and instances, we can optimize the best shape in order to locate the landmarks. Experimented on more than 1860 face models from FRGC datasets, our method achieves an average of locating errors less than 7 mm for 15 feature points. Compared with a curvature analysis-based method also developed within our team, this learning-based method enables localization of more facial landmarks with a general better accuracy at the cost of a learning step. Xi Zhao 0001, Przemyslaw Szeptycki, Emmanuel Dellandréa, Liming Chen 0002 |
WACV | 3 |
| 2005 | Features extraction and selection for emotional speech classificationabstractThe classification of emotional speech is a topic in speech recognition with more and more interest, and it has giant prospect in applications in a wide variety of fields. It is an important preparation for automatic classification and recognition of emotions to select a proper feature set as a description to the emotional speech, and to find a proper definition to the emotions in speech. The speech samples used in this paper come from Berlin database which contains 7 kinds of emotions, with 207 speech samples of male voice and 287 speech samples of female voice. A feature set of 50 potentially features is extracted and analyzed, and the best features are selected. A definition of emotions as 3-states emotions is also proposed in this paper. Zhongzhe Xiao, Emmanuel Dellandréa, Weibei Dou, Liming Chen 0002 |
AVSS | 2 |