VLDB 2026 Research / reviewers in the wild / expert
Alice Caplier
dblp:44/6708
· DBLP profile ↗
58ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-5937-4627ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 21 · 3 since 2021Human-computer interaction and ubiquitous computing · 4Security and privacy · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Look around when in doubt: Adaptive contextual reasoning for conformal and explainable road object detection with CAESAR++abstractTrustworthy object detection is a cornerstone for autonomous vision systems operating in safety-critical applications. This paper introduces CAESAR++ (Context-Aware Explanations via Semantic Attribution and Refinement), an adaptive contextual reasoning framework designed to enhance trustworthiness in road object detection. CAESAR++ integrates a two-step conformal prediction approach to obtain statistically grounded uncertainty estimates that determine when contextual reasoning is required, dynamically calibrates context window sizes according to that uncertainty, and employs dual-color saliency maps to disentangle bottom-up sensory evidence from top-down contextual cues within a unified explanation. Unlike methods that treat uncertainty estimation and explainability as separate components, CAESAR++ couples them in a single detector-agnostic pipeline without requiring retraining of the core model. Comprehensive experiments on TJU-DHD-Traffic, BDD100K, and Pascal VOC with ten state-of-the-art detectors demonstrate consistent improvements in detection accuracy, robustness, uncertainty calibration, and explanation quality, while preserving competitive computational efficiency. These contributions advance the development of transparent, robust, and responsible vision systems for autonomous navigation safety. Anh-Thu Mai, Marina Nicolas, Patricia Ladret, Alice Caplier |
Neurocomputing | 4 |
| 2025 | Robust Road Object Detection with CAESAR: Context-Aware Explanations via Semantic Attribution and RefinementabstractContemporary road object detectors face entangled challenges of opaque decision-making, occlusions, and adverse weather, yet no existing method addresses all simultaneously. We introduce CAESAR (Context-Aware Explanations via Semantic Attribution and Refinement), a model-agnostic framework that improves baseline detectors’ performance while generating faithful, human-aligned, instance-level explanations. CAESAR combines top-down contextual cues with bottom-up features, adaptively balancing global semantics and local details according to prediction uncertainty. When integrated into state-of-the-art detectors, CAESAR consistently enhances detection accuracy and robustness across unfavorable conditions on TJU-DHD-Traffic and BDD100K datasets. Its modular context component supports standalone recalculation for domain adaptation without altering the core detector. Faithfulness evaluation confirms its alignment with the model’s internal reasoning, and an ablation study demonstrates its compatibility with complementary components for optimal performance. By leveraging human attention strategies, CAESAR advances trustworthy vision systems with robust road object detection for autonomous driving and beyond. Anh-Thu Mai, Marina Nicolas, Patricia Ladret, Alice Caplier |
VCIP | 4 |
| 2024 | Parallelized Nonlinear Scaled Transform for HEVCabstractTransform coding is widely used for compressing video data by compacting energy in the frequency domain. Conventional codecs like HEVC rely on the Discrete Cosine Transform, which has demonstrated a favorable balance between performance and complexity. The emergence of deep learning in video coding has prompted exploration into nonlinear transforms (NLT) to capitalize on nonlinearity for non-stationary signal sources. NLT, trained to minimize rate-distortion, surpass conventional linear transforms. However, these approaches require modifications to existing standards, making them incompatible with current codecs. Additionally, scalar quantization (SQ) can be substituted with an optimized scheme like rate-distortion optimized quantization (RDOQ). While RDOQ yields satisfactory results, its iterative nature poses challenges for hardware integration. Although DL-based RDOQ have been proposed to tackle this issue, they are not directly trained to minimize a rate-distortion function and the transform operation is not optimized. For that purpose, we propose a parallelized nonlinear scaled transform trained for replacing and improving forward DCT and SQ of HEVC while ensuring standard compliance. Our proposed method is tested on 4×4 and 8×8 blocks and reaches 1.13 % BD-rate reduction on luma, up to 1.85 %. Pierre-Alain Afro, Loïc Strus, Hugo Chauvet, Laurent Bonnaud, Alice Caplier, Frédéric Robin |
VCIP | 5 |
| 2023 | Multi-QP Rate Distortion Optimized Quantization Using Deep LearningabstractRDOQ (Rate Distortion Optimized Quantization) is an efficient encoding tool that can be used with several codecs such as H.264/AVC, H.265/HEVC or AV1. Although this algorithm can significantly reduce the bit rate, its complexity is a limitation for video coding hardware solutions. Studies have succeeded in simplifying RDOQ allowing its integration into HM, the reference software implementation of H.265/HEVC. However, its iterative and sequential behavior does not allow an efficient hardware implementation. With the advent of machine learning, neural networks have been introduced to mimic the RDOQ algorithm in a parallel way. Previous proposed Deep-Learning based RDOQ frameworks still need to be trained for each Quantization Parameter (QP) which is a huge limitation for hardware implementation. To address this issue, we propose a Multi-QP Deep Learning based RDOQ, trained with data extracted from 4 QPs and achieving targeted performances with QP ranging from 22 to 37. Our multi-QP model almost reaches HM-RDOQ performance in terms of BD-Rate savings. Pierre-Alain Afro, Loïc Strus, Laurent Bonnaud, Alice Caplier, Frédéric Robin |
VCIP | 4 |
| 2021 | Using Synthetic Corruptions to Measure Robustness to Natural Distribution Shifts
Alfred Laugros, Alice Caplier, Matthieu Ospici |
BMVC | 2 |
| 2021 | Using the Overlapping Score to Improve Corruption BenchmarksabstractNeural Networks are sensitive to various corruptions that usually occur in real-world applications such as blurs, noises, low-lighting conditions, etc. To estimate the robustness of neural networks to these common corruptions, we generally use a group of modeled corruptions gathered into a benchmark. Unfortunately, no objective criterion exists to determine whether a benchmark is representative of a large diversity of independent corruptions. In this paper, we propose a metric called corruption overlapping score, which can be used to reveal flaws in corruption benchmarks. Two corruptions overlap when the robustnesses of neural networks to these corruptions are correlated. We argue that taking into account overlappings between corruptions can help to improve existing benchmarks or build better ones. Alfred Laugros, Alice Caplier, Matthieu Ospici |
ICIP | 2 |
| 2021 | Deep Learning for Spatio-Temporal Modeling of Dynamic Spontaneous EmotionsabstractFacial expressions involve dynamic morphological changes in a face, conveying information about the expresser's feelings. Each emotion has a specific spatial deformation over the face and temporal profile with distinct time segments. We aim at modeling the human dynamic emotional behavior by taking into consideration the visual content of the face and its evolution. But emotions can both speed-up or slow-down, therefore it is important to incorporate information from the local neighborhood frames (short-term dependencies) and the global setting (long-term dependencies) to summarize the segment context despite of its time variations. A 3D-Convolutional Neural Networks (3D-CNN) is used to learn early local spatiotemporal features. The 3D-CNN is designed to capture subtle spatiotemporal changes that may occur on the face. Then, a Convolutional-Long-Short-Term-Memory (ConvLSTM) network is designed to learn semantic information by taking into account longer spatiotemporal dependencies. The ConvLSTM network helps considering the global visual saliency of the expression. That is locating and learning features in space and time that stand out from their local neighbors in order to signify distinctive facial expression features along the entire sequence. Non-variant representations based on aggregating global spatiotemporal features at increasingly fine resolutions are then done using a weighted Spatial Pyramid Pooling layer. Dawood Al Chanti, Alice Caplier |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | Study of naturalness in tone-mapped images
Quyet-Tien Le, Patricia Ladret, Huu-Tuan Nguyen, Alice Caplier |
Comput. Vis. Image Underst. | 4 |
| 2019 | Large Field/Close-Up Image Classification: From Simple to Very Complex Features
Quyet-Tien Le, Patricia Ladret, Huu-Tuan Nguyen, Alice Caplier |
CAIP (2) | 4 |
| 2018 | Motion-based countermeasure against photo and video spoofing attacks in face recognition
Taiamiti Edmunds, Alice Caplier |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Unsupervised joint face alignment with gradient correlation coefficient
Weiyuan Ni, Ngoc-Son Vu, Alice Caplier |
Pattern Anal. Appl. | 3 |
| 2015 | How to predict the global instantaneous feeling induced by a facial picture?
Arnaud Lienhard, Patricia Ladret, Alice Caplier |
Signal Process. Image Commun. | 3 |
| 2015 | Local Patterns of Gradients for Face RecognitionabstractWe present a novel feature extraction method named local patterns of gradients (LPOGs) for robust face recognition. LPOG uses block-wised elliptical local binary patterns (BELBP), a refined variant of ELBP, and local phase quantization (LPQ) operators directly on gradient images for capturing local texture patterns to build up a feature vector of a face image. From one input image, two directional gradient images are computed. A symmetric pair of BELBP and a LPQ operator are then separately applied upon each gradient image to generate local patterns images. Histogram sequences of local patterns images' nonoverlapped subregions are finally concatenated to form the LPOG vector for the given image. Based on LPOG descriptor, we propose a novel face recognition system which exploits whitened principal component analysis (WPCA) for dimension reduction and weighted angle-based distance for classification. Experimental results on three large public databases (FERET, AR, and SCface) prove that LPOG WPCA system is robust against a wide range of challenges, such as illumination, expression, occlusion, pose, time-lapse variations, and low resolution. In addition, comparison with other systems shows that LPOG WPCA significantly outperforms the state-of-the-art methods. Computationally, timing benchmarks also demonstrate that our LPOG method is faster than many advanced feature extraction algorithms and can be applied in real-world applications. Huu-Tuan Nguyen, Alice Caplier |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Patch based local phase quantization of monogenic components for face recognitionabstractIn this paper, we propose a novel feature extraction method for Face recognition called patch based Local Phase Quantization of Monogenic components (PLPQMC). From the input image, the directional Monogenic bandpass components are generated. Then, each pixel of a bandpass image is replaced by the mean value of its rectangular neighborhood. Next, LPQ histogram sequences are computed upon those images. Finally, these histogram sequences are concatenated for constituting a global representation of the face image. Using the proposed method for feature extraction, we construct a new face recognition system with Whitened Principal Component Analysis (WPCA) for dimensionality reduction, k-nearest neighbor classifier and weighted angle distance for classification. Performance evaluations on two public face databases FERET and SCface show that our method is efficient against some challenging issues, e.g. expressions, illumination, time-lapse, low resolution, and it is competing with state-of-the-art methods. Huu-Tuan Nguyen, Alice Caplier |
ICIP | 2 |
| 2014 | Retina enhanced SURF descriptors for spatio-temporal concept detection
Sabin Tiberius Strat, Alexandre Benoît, Patrick Lambert, Alice Caplier |
Multim. Tools Appl. | 4 |
| 2013 | Towards a full emotional systemabstractThis study proposes a system that is able to classify a facial expression in one of the six categories, namely, Joy, Disgust, Anger, Sadness, Fear and Surprise and also to assign to each expression its intensity in the range: High, Medium and Low. This is carried out in two independent and parallel processes. Permanent and transient facial features are detected from still images, and pertinent information about the presence of transient features on specific facial regions and about facial distances computed from permanent facial features is extracted. Both classification and quantification processes are based on transient and permanent features. The belief theory is used with the two processes because of its ability in fusing data coming from different sensors. The system outputs a recognised and quantified expression. The quantification process allows recognising a new subset of expressions deduced from the basic ones. Indeed, by associating to each expression three intensities low, medium and high, we deduce three facial expressions. Finally, a set of 18 facial expressions is categorised instead of the six ones. Experimental results are given to show the system classification accuracy. Khadoudja Ghanem, Alice Caplier |
Behav. Inf. Technol. | 2 |
| 2013 | Lip contour segmentation and tracking compliant with lip-reading application constraints
Sébastien Stillittano, Vincent Girondel, Alice Caplier |
Mach. Vis. Appl. | 3 |
| 2012 | Adaptive appearance face tracking with alignment feedbacksabstractAdaptive appearance approaches are popular for tracking non-rigid objects, such as faces. However, these approaches usually lack direct mechanisms for correcting spatial misalignments (e.g., translation, scaling and rotation errors) existing in the tracking outputs. The unwanted errors are then accumulated in the target's appearance model. This inevitably has negative effects on tracking performance. Besides, many of these approaches rely on video-specific parameter setting. In this paper, we first adopt a self-adaptive dynamical model to predict the candidates of target. Hence, our tracker is able to work with identical parameters for various situations. Moreover, we introduce a multi-view joint face alignment stage to decrease the impact of mis-alignment. Aligned faces are further used as feedbacks to update the appearance model. We test the proposed algorithm on outdoor surveillance videos and real-world YouTube videos. Experimental results prove the effectiveness of our method in tracking faces under uncontrolled conditions. Weiyuan Ni, Alice Caplier |
ICIP | 2 |
| 2012 | Multiple patterns of gradient magnitudes for face recognitionabstractDescribing efficiently faces is a task of mounting importance. Most of existing algorithms do not address all the three criteria, i.e., robustness, distinctiveness and low computational cost. Inspired by recent features so-called POEM (Patterns of Oriented Edge Magnitudes) which is argued balancing well the three concerns, we first provide an improvement to it and then propose novel features so-called self-POEM considering the relations between edge distributions along different directions of one region itself. Self-POEM provides the complementary strength which is not present in POEM. Combining them together, a more robust algorithm is obtained and by experiments we prove that our approach is more efficient than contemporary ones. Ngoc-Son Vu, Huu-Tuan Nguyen, Alice Caplier |
ICIP | 3 |
| 2012 | Face recognition using Multi-modal Binary Patterns
Thanh Phuong Nguyen 0001, Ngoc-Son Vu, Alice Caplier |
ICPR | 3 |
| 2012 | Lucas-Kanade based entropy congealing for joint face alignment
Weiyuan Ni, Ngoc-Son Vu, Alice Caplier |
Image Vis. Comput. | 3 |
| 2012 | Using retina modelling to characterize blinking: comparison between EOG and video analysis
Antoine Picot, Sylvie Charbonnier, Alice Caplier, Ngoc-Son Vu |
Mach. Vis. Appl. | 3 |
| 2012 | Face recognition using the POEM descriptor
Ngoc-Son Vu, Hannah M. Dee, Alice Caplier |
Pattern Recognit. | 3 |
| 2012 | Enhanced Patterns of Oriented Edge Magnitudes for Face Recognition and Image MatchingabstractA good feature descriptor is desired to be discriminative, robust, and computationally inexpensive in both terms of time and storage requirement. In the domain of face recognition, these properties allow the system to quickly deliver high recognition results to the end user. Motivated by the recent feature descriptor called Patterns of Oriented Edge Magnitudes (POEM), which balances the three concerns, this paper aims at enhancing its performance with respect to all these criteria. To this end, we first optimize the parameters of POEM and then apply the whitened principal-component-analysis dimensionality reduction technique to get a more compact, robust, and discriminative descriptor. For face recognition, the efficiency of our algorithm is proved by strong results obtained on both constrained (Face Recognition Technology, FERET) and unconstrained (Labeled Faces in the Wild, LFW) data sets in addition with the low complexity. Impressively, our algorithm is about 30 times faster than those based on Gabor filters. Furthermore, by proposing an additional technique that makes our descriptor robust to rotation, we validate its efficiency for the task of image matching. Ngoc-Son Vu, Alice Caplier |
IEEE Trans. Image Process. | 2 |
| 2012 | On-Line Detection of Drowsiness Using Brain and Visual InformationabstractA drowsiness detection system using both brain and visual activity is presented in this paper. The brain activity is monitored using a single electroencephalographic (EEG) channel. An EEG-based drowsiness detector using diagnostic techniques and fuzzy logic is proposed. Visual activity is monitored through blinking detection and characterization. Blinking features are extracted from an electrooculographic (EOG) channel. Features are merged using fuzzy logic to create an EOG-based drowsiness detector. The features used by the EOG-based detector are voluntary restricted to the features that can be automatically extracted from a video analysis of the same accuracy. Both detection systems are then merged using cascading decision rules according to a medical scale of drowsiness evaluation. Merging brain and visual information makes it possible to detect three levels of drowsiness: “awake,” “drowsy,” and “very drowsy.” One major advantage of the system is that it does not have to be tuned for each driver. The system was tested on driving data from 20 different drivers and reached 80.6% correct classifications on three drowsiness levels. The results show that EEG and EOG detectors are redundant: EEG-based detections are used to confirm EOG-based detection and thus enable the false alarm rate to be reduced to 5% while the true positive rate is not decreased, compared with a single EOG-based detector. Antoine Picot, Sylvie Charbonnier, Alice Caplier |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2011 | An Online Three-Stage Method for Facial Point Localization
Weiyuan Ni, Ngoc-Son Vu, Alice Caplier |
CAIP (2) | 3 |
| 2011 | Mining patterns of orientations and magnitudes for face recognitionabstractGood face recognition system is one which quickly de- livers high accurate results to the end user. For this purpose, face representation must be robust, discriminative and also of low computational cost in both terms of time and space. Inspired by recently proposed feature set so-called POEM (Patterns of Oriented Edge Magnitudes) which considers the relationships between edge distributions of different image patches and is argued balancing well the three concerns, this work proposes to further exploit patterns of both orientations and magnitudes for building more efficient algorithm. We first present novel features called Patterns of Dominant Orientations (PDO) which consider the relationships between "dominant" orientations of local image regions at different scales. We also propose to apply the whitened PCA technique upon both the POEM and PDO based representations to get more compact and discriminative face descriptors. We then show that the two methods have complementary strength and that by combining the two descriptors, one obtains stronger results than either of them considered separately. By experiments carried out on several common benchmarks, including both frontal and non- frontal FERET as well as the AR datasets, we prove that our approach is more efficient than contemporary ones. Ngoc-Son Vu, Alice Caplier |
IJCB | 2 |
| 2011 | Newton optimization based Congealing for facial image alignmentabstractCongealing is an unsupervised image alignment method for a set of images, and the transformation parameters are obtained by minimizing a sum-of-entropies function. In this paper, we provide a solution to improve the estimation of transformation parameters using a Newton optimization method, under the premise of maintaining the compatibility with feature descriptors. Besides, instead of SIFT descriptor in canonical Congealing, we combine Congealing with POEM (Patterns of Oriented Edge Magnitudes) which catches both edge in- formation and the relation between pixels at a neighboring region. The experiment results show that our alignment method has better ability for the removal of unwanted displacements, and also improves the performance of face recognition. Weiyuan Ni, Alice Caplier |
ICIP | 2 |
| 2010 | Face Recognition with Patterns of Oriented Edge Magnitudes
Ngoc-Son Vu, Alice Caplier |
ECCV (1) | 2 |
| 2010 | Crowd behaviour analysis using histograms of motion directionabstractA practical system for the automated analysis of crowded scenes will have to deal with multiple occlusions and tracking failures, in a context in which the cameras may move at any time to point in any direction, at any level of zoom. This paper presents a prototype component of such a system. Much work in crowd modelling assumes that the camera will be static for extended periods of time and that a model of the scene can therefore be learned; we do not make this assumption and instead build a simple representation of motion patterns that is applicable across different views and which learns motion scale rapidly. Our representation is based upon histograms of motion direction alongside an indication of motion speed. These can be used for detecting frames in which behaviour differs from the training set, and also for localisation of where in the image these anomalous events occur. We evaluate this work against five event-detection scenarios from the public PETS2009 crowd behaviour dataset. Hannah M. Dee, Alice Caplier |
ICIP | 2 |
| 2010 | Patch-Based Similarity HMMs for Face Recognition with a Single Reference ImageabstractIn this paper we present a new architecture for face recognition with a single reference image, which completely separates the training process from the recognition process. In the training stage, by using a database containing various individuals, the spatial relations between face components are represented by two Hidden Markov Models (HMMs), one modeling within-subject similarities, and the other modeling inter-subject differences. This allows us during the recognition stage to take a pair of face images, neither of which has been seen before, and to determine whether or not they come from the same individual. Whilst other face-recognition HMMs use Maximum Likelihood criterion, we test our approach using both Maximum Likelihood and Maximum a Posteriori (MAP) criterion, and find that MAP provides better results. Importantly, the training database can be entirely separated from the gallery and test images: this means that adding new individuals to the system can be done without re-training. We present results based upon models trained on the FERET training dataset, and demonstrate that these give satisfactory recognition rates on both the FERET database itself and more impressively the unseen AR database. When compared to other HMM based face recognition techniques, our algorithm is of much lower complexity due to the small size of our observation sequence. Ngoc-Son Vu, Alice Caplier |
ICPR | 2 |
| 2010 | Fusing bio-inspired vision data for simplified high level scene interpretation: Application to face motion analysis
Alexandre Benoît, Alice Caplier |
Comput. Vis. Image Underst. | 2 |
| 2010 | Using Human Visual System modeling for bio-inspired low level image processing
Alexandre Benoît, Alice Caplier, Barthélémy Durette, Jeanny Hérault |
Comput. Vis. Image Underst. | 2 |
| 2009 | A Generalization of the Pignistic Transform for Partial Bet
Thomas Burger, Alice Caplier |
ECSQARU | 2 |
| 2009 | Inner and outer lip contour tracking using cubic curve parametric modelsabstractThe first step in lip-reading applications is mouth contour extraction to provide a link between the lip shape and the oral message. In our approach, lip contours are detected in the first image with the two algorithms developed and for static images. On subsequent images of the sequence, several key points (mouth corners, inner and outer middle contour points) are tracked with the Lucas-Kanade method to define an initial parametric lip model of the mouth. According to a combined luminance and chrominance gradient, the model is optimized and precisely locked on to the lip contours. The algorithm performances are evaluated with respect to a lip-reading application. Sébastien Stillittano, Vincent Girondel, Alice Caplier |
ICIP | 3 |
| 2009 | Illumination-robust face recognition using retina modelingabstractIllumination variations that might occur on face images degrade the performance of face recognition systems. In this paper, we propose a novel method of illumination normalization based on retina modeling by combining two adaptive nonlinear functions and a Difference of Gaussians filter. The proposed algorithm is evaluated on the Yale B database and the Feret illumination database using two face recognition methods: PCA based and Local Binary Pattern based (LBP). Experimental results show that the proposed method achieves very high recognition rates even for the most challenging illumination conditions. Our algorithm has also a low computational complexity. Ngoc-Son Vu, Alice Caplier |
ICIP | 2 |
| 2009 | Comparison between EOG and high frame rate camera for drowsiness detectionabstractDrowsiness is responsible for a large number car crashes. Blinks analysis from electrooculogram (EOG) signal brings reliable information on drowsiness but EOG recording condition can be really disturbing for the driver. On the other hand, video approaches seem a lot more practical but the standard acquisition rate does not give the same accuracy than EOG for blinks analysis. So, a high frame rate camera seems a good compromise. The purpose of this paper is to study to what extent a high speed camera could replace the EOG for the extraction of blinks features in order to design a system to detect drowsiness. An original method to detect and characterize blinks from the video is presented. This method uses two energy signals extracted from the video analysis: one related to the contours of the eyes and the other one to the moving contours. A comparison between the different features extracted from the EOG and from the video is then performed. This study shows that duration, frequency, PERCLOS 80 and dynamic features extracted from the EOG and from the video signals are highly correlated. The frame rate influence on the accuracy of the different features extracted is also studied. Antoine Picot, Alice Caplier, Sylvie Charbonnier |
WACV | 2 |
| 2009 | The Natural Statistics of Audiovisual SpeechabstractHumans, like other animals, are exposed to a continuous stream of signals, which are dynamic, multimodal, extended, and time varying in nature. This complex input space must be transduced and sampled by our sensory systems and transmitted to the brain where it can guide the selection of appropriate actions. To simplify this process, it's been suggested that the brain exploits statistical regularities in the stimulus space. Tests of this idea have largely been confined to unimodal signals and natural scenes. One important class of multisensory signals for which a quantitative input space characterization is unavailable is human speech. We do not understand what signals our brain has to actively piece together from an audiovisual speech stream to arrive at a percept versus what is already embedded in the signal structure of the stream itself. In essence, we do not have a clear understanding of the natural statistics of audiovisual speech. In the present study, we identified the following major statistical features of audiovisual speech. First, we observed robust correlations and close temporal correspondence between the area of the mouth opening and the acoustic envelope. Second, we found the strongest correlation between the area of the mouth opening and vocal tract resonances. Third, we observed that both area of the mouth opening and the voice envelope are temporally modulated in the 2-7 Hz frequency range. Finally, we show that the timing of mouth movements relative to the onset of the voice is consistently between 100 and 300 ms. We interpret these data in the context of recent neural theories of speech which suggest that speech communication is a reciprocally coupled, multisensory event, whereby the outputs of the signaler are matched to the neural processes of the receiver. Chandramouli Chandrasekaran, Andrea Trubanova, Sébastien Stillittano, Alice Caplier, Asif A. Ghazanfar |
PLoS Comput. Biol. | 4 |
| 2009 | A belief-based sequential fusion approach for fusing manual signs and non-manual signals
Oya Aran, Thomas Burger, Alice Caplier, Lale Akarun |
Pattern Recognit. | 3 |
| 2009 | Multimodal focus attention and stress detection and feedback in an augmented driver simulator
Alexandre Benoît, Laurent Bonnaud, Alice Caplier, Phillipe Ngo, Jean-Yves Lionel Lawson, Daniela Gorski Trevisan, Vjekoslav Levacic, Céline Mancas, Guillaume Chanel |
Pers. Ubiquitous Comput. | 3 |
| 2008 | Open or Closed Mouth State Detection: Static Supervised Classification Based on Log-Polar Signature
Christian Bouvier, Alexandre Benoît, Alice Caplier, Pierre-Yves Coulon |
ACIVS | 3 |
| 2007 | Facial expression classification: An approach based on the fusion of facial deformations using the transferable belief model
Zakia Hammal, Laurent Couvreur, Alice Caplier, Michèle Rombaut |
Int. J. Approx. Reason. | 3 |
| 2006 | Extracting Static Hand Gestures in Dynamic ContextabstractCued speech is a specific visual coding that complements oral language lip-reading, by adding static hand gestures (a static gesture can be presented on a single photograph as it contains no motion). By nature, cued speech is simple enough to be believed as automatically recognizable. Unfortunately, despite its static definition, fluent cued speech has an important dynamic dimension due to co-articulation. Hence, the reduction from a continuous cued speech coding stream to the corresponding discrete chain of static gestures is really an issue for automatic cued speech processing. We present here how the biological motion analysis method are presented has been combined with a fusion strategy based on the belief theory in order to perform such a reduction. Thomas Burger, Alexandre Benoît, Alice Caplier |
ICIP | 3 |
| 2006 | Modeling Hesitation and Conflict: A Belief-Based Approach for Multi-class ProblemsabstractSupport vector machine (SVM) is a powerful tool for binary classification. Numerous methods are known to fuse several binary SVMs into multi-class (MC) classifiers. These methods are efficient, but an accurate study of the misclassified items leads to notice two sources of mistakes: (1) the response of each classifier does not use the entire information from the SVM, and (2) the decision method does not use the entire information from the classifier responses. In this paper, we present a method which partially prevents these two losses of information by applying belief theories (BTs) to SVM fusion, while keeping the efficient aspect of the classical methods Thomas Burger, Oya Aran, Alice Caplier |
ICMLA | 3 |
| 2006 | Parametric models for facial features segmentation
Zakia Hammal, Nicolas Eveno, Alice Caplier, Pierre-Yves Coulon |
Signal Process. | 3 |
| 2005 | Hypovigilence analysis: open or closed eye or mouth? Blinking or yawning frequency?abstractThis paper proposes a frequency method to estimate the state open or closed of eye and mouth and to detect associated motion events such as blinking and yawning. The context of that work is the detection of hypovigilence state of a user such as a driver, a pilot. In A. Benoit and Caplier (2005) we proposed a method for motion detection and estimation which is based on the processing achieved by the human visual system. The motion analysis algorithm the filtering step occurring at the retina level and the analysis done at the visual cortex level. This method is used to estimate the motion of eye and mouth: blinking is related to fast vertical motion of the eyelid and yawning is related to large vertical mouth opening. The detection of the open or closed state of the feature is based on the analysis of the total energy of the image at the output of the retina filter: this energy is higher for open features. The absolute level of energy associated to a specific state being different from a person to another and for different illumination conditions, the energy level associated to each state open or closed is adaptive and is updated each time a motion event (blinking or yawning) is detected. No constraint about motion is required. The system is working in real time and under all type of lighting conditions since the retina filtering is able to cope with illumination variations. This allows to estimate blinking and yawning frequencies which are clues of hypovigilance. Alexandre Benoît, Alice Caplier |
AVSS | 2 |
| 2005 | A belief theory-based static posture recognition systems for real-time video surveillance applicationsabstractThis paper presents a system that can automatically recognize four different static human body postures for video surveillance applications. The considered postures are standing, sitting, squatting, and lying. The data come from the persons 2D segmentation and from their face localization. It consists in distance measurements relative to a reference posture (standing, arms stretched horizontally). The recognition is based on data fusion using the belief theory, because this theory allows the modelling of imprecision and uncertainty. The efficiency and the limits of the recognition system are highlighted thanks to the processing of several thousands of frames. A considered application is the monitoring of elder people in hospitals or at home. This system allows real-time processing. Vincent Girondel, Alice Caplier, Laurent Bonnaud |
AVSS | 2 |
| 2005 | Head nods analysis: interpretation of non verbal communication gesturesabstractThis paper proposes a real time frequency method to detect 2D rigid rotations of pan or tilt of a moving head. We aim at interpreting head nods involved in the non verbal communication process in the same way as human being: direction of the rotation is estimated but not its precise amplitude. The idea of the method is to analyze the image spectrum in the log polar domain where global 2D head rotations are transformed into simple energy translations. In order to make the log polar spectrum easy to interpret, a prefiltering stage inspired from the biological model of the human retina is applied: mobile contours are enhanced and static contours are attenuated, high frequency noise is eliminated and variations of illumination are cancelled. Estimated rotations are integrated in a data fusion process able to detect and to interpret in real time head nods of approbation or negation. Alexandre Benoît, Alice Caplier |
ICIP (3) | 2 |
| 2005 | Static human body postures recognition in video sequences using the belief theoryabstractThis paper presents a system that can automatically recognize four different static human body postures in video sequences. The considered postures are standing, sitting, squatting, and lying. The recognition is based on data fusion using the belief theory. The data come from the persons 2D segmentation and from their face localization. It consists in distance measurements relative to a reference posture ("Da Vinci posture": standing, arms stretched horizontally). The segmentation is based on an adaptive background removal algorithm. The face localization process uses skin detection based on color information with an adaptive thresholding. The efficiency and the limits of the recognition system are highlighted thanks to the analysis of a great number of results. This system allows real-time processing. Vincent Girondel, Laurent Bonnaud, Alice Caplier, Michèle Rombaut |
ICIP (2) | 3 |
| 2004 | Accurate and quasi-automatic lip trackingabstractLip segmentation is an essential stage in many multimedia systems such as videoconferencing, lip reading, or low-bit-rate coding communication systems. In this paper, we propose an accurate and robust quasi-automatic lip segmentation algorithm. First, the upper mouth boundary and several characteristic points are detected in the first frame by using a new kind of active contour: the "jumping snake." Unlike classic snakes, it can be initialized far from the final edge and the adjustment of its parameters is easy and intuitive. Then, to achieve the segmentation, we propose a parametric model composed of several cubic curves. Its high flexibility enables accurate lip contour extraction even in the challenging case of a very asymmetric mouth. Compared to existing models, it brings a significant improvement in accuracy and realism. The segmentation in the following frames is achieved by using an interframe tracking of the keypoints and the model parameters. However, we show that, with a usual tracking algorithm, the keypoints' positions become unreliable after a few frames. We therefore propose an adjustment process that enables an accurate tracking even after hundreds of frames. Finally, we show that the mean keypoints' tracking errors of our algorithm are comparable to manual points' selection errors. Nicolas Eveno, Alice Caplier, Pierre-Yves Coulon |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | Jumping snakes and parametric model for lip segmentationabstractLip segmentation is an essential stage in many multimedia systems such as videoconferencing, lip reading, or low bit rate coding communication systems. In this paper, we propose an accurate and robust lip segmentation algorithm. First, the upper mouth boundary and several characteristic points are detected by using a new kind of active contour : the "jumping snake". Unlike classic snakes, it can be initialized far from the final edge and the adjustment of its parameters is easy and intuitive. In a second step, a parametric model composed of several cubic curves is fitted on the lips. This model is flexible enough to reproduce the specificities of very different lip shapes. Compared to existing models, it brings a significant accuracy and realism improvement. Nicolas Eveno, Alice Caplier, Pierre-Yves Coulon |
ICIP (2) | 2 |
| 2002 | A parametric model for realistic lip segmentationabstractLip segmentation is an essential stage in many multimedia systems such as videoconferencing, lip reading, or low bit rate coding communication systems. In this paper, we propose an accurate and robust lip segmentation algorithm. First, the mouth region and several characteristic points are detected by using "hybrid edges" (which combine colour and intensity information) and a priori knowledge about the lip morphology. Corners position, which is crucial, is provided by a coarse-to-fine process. Then, a parametric model is fitted on the lips. We consider that the lip boundary is composed of several independent cubic polynomial curves. It gives a low complexity global model that is flexible enough to reproduce the specificity of very different lip shapes. Compared to existing models, it brings a significant accuracy and realism improvement. Moreover, it ensures a robust convergence towards the edges because the different parts of the model are independent. Nicolas Eveno, Alice Caplier, Pierre-Yves Coulon |
ICARCV | 2 |
| 2002 | Key points based segmentation of lipsabstractLip segmentation is an essential stage in many multimedia systems such as videoconferencing, lip reading, or low bit rate coding communication systems. We propose an accurate and robust lip segmentation algorithm. First, characteristic points are found by using "hybrid edges" (which combine color and intensity information) and a priori knowledge about the lip morphology. Corners position, which is crucial, is provided by a coarse-to-fine process. Then, a parametric model is fitted on the lips. We consider that the lip boundary is composed of several independent cubic polynomial curves. It gives a low complexity model that is flexible enough to reproduce the specificity of very different lip shapes. Compared to existing models, it brings a significant accuracy improvement. Moreover, it ensures a robust convergence towards the edges. Nicolas Eveno, Alice Caplier, Pierre-Yves Coulon |
ICME (2) | 2 |
| 2001 | Robust fast extraction of video objects combining frame differences and adaptive reference imageabstractThis paper introduces a video object segmentation algorithm developed in the context of the European project Art.live where constraints on the quality of segmentation and the processing rate (at least 10 images/second) are required. In order to obtain a fine segmentation (no blocking effect, boundaries precision, temporal stability without flickering), the segmentation process is based on Markov random field (MRF) modelling which involves consecutive frame difference and a reference image in a unified way. Temporal changes of the luminance are predominant when the reference image is not yet available whereas the reference image prevails for low textured moving objects or for objects which stop moving for a while. The increased processing rate comes from the substitution of some Markovian iterations with morphological operations without loss of quality. Simulation results show the efficiency of the proposed method in term of accuracy and complexity (/spl sime/6 images/second for 352/spl times/288 pixels YUV images on a low-end processor). Alice Caplier, Laurent Bonnaud, Jean-Marc Chassery |
ICIP (2) | 1 |
| 2001 | New color transformation for lips segmentationabstractA robust pre-processing algorithm for lip segmentation is presented. We define a new transformation based on RGB color space: the chromatic curve map. It aims at increasing discrimination between lips and skin. It allows robust lips detection under non uniform lighting conditions and without any particular make-up. The motivation of the present work is to extract visual information for automatic speech recognition, videoconferencing and speaker's face synthesis under natural lighting conditions. Nicolas Eveno, Alice Caplier, Pierre-Yves Coulon |
MMSP | 2 |
| 1999 | Spatiotemporal MRF approach to video segmentation: Application to motion detection and lip segmentation
Franck Luthon, Alice Caplier, Marc Liévin |
Signal Process. | 2 |
| 1995 | A New Spatiotemporal Approach for Image Analysis. Application to Motion Detection
Alice Caplier, Franck Luthon |
CAIP | 1 |
| 1994 | An MRF Based Motion Detection Algorithm Implemented on Analog Resistive Network
Franck Luthon, George V. Popescu, Alice Caplier |
ECCV (1) | 3 |