Mohamed Daoudi

dblp:73/2417 · DBLP profile ↗
← Back
117ranked-venue papers
7as first author
32since 2021 · last 2026
0000-0003-4219-7860ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 78 · 4 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 75 · 5 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Region-Aware Latent Axis Discovery for Predictive Botulinum Toxin Facial Simulation
Estephe Arnaud, Mohamed Daoudi, Pierre Guerreschi
FG2
2026 ReactionMamba: Generating Short & Long Human Reaction Sequences
Hajra Anwar Beg, Baptiste Chopin, Hao Tang 0005, Mohamed Daoudi
FG4
2026 AmbiGest: A Dataset of Social Gestures with Inter-Class Similarity and Intra-Class Variability
Hajra Anwar Beg, Mohamed Daoudi, Angela Bartolo
FG2
2026 A Non-Invasive 3D Gait Analysis Framework for Quantifying Psychomotor Retardation in Major Depressive Disorder
Fouad Boutaleb, Emery Pierson, Mohamed Daoudi, Clémence Nineuil, Ali Amad, Fabien D'Hondt
FG3
2026 Beyond Fixed Topologies: Unregistered Training and Comprehensive Evaluation Metrics for 3D Talking Heads
Federico Nocentini, Thomas Besnier, Claudio Ferrari, Sylvain Arguillère, Mohamed Daoudi, Stefano Berretti
Int. J. Comput. Vis.5
2025 Measuring Anxiety Levels with Head Motion Patterns in Severe Depression Population
abstract
Depression and anxiety are prevalent mental health disorders that frequently cooccur, with anxiety significantly influencing both the manifestation and treatment of depression. An accurate assessment of anxiety levels in individuals with depression is crucial to develop effective and personalized treatment plans. This study proposes a new noninvasive method for quantifying anxiety severity by analyzing head movements -specifically speed, acceleration, and angular displacement during video-recorded interviews with patients suffering from severe depression. Using data from a new CALYPSO Depression Dataset, we extracted head motion characteristics and applied regression analysis to predict clinically evaluated anxiety levels. Our results demonstrate a high level of precision, achieving a mean absolute error (MAE) of 0.35 in predicting the severity of psychological anxiety based on head movement patterns. This indicates that our approach can enhance the understanding of anxiety’s role in depression and assist psychiatrists in refining treatment strategies for individuals.
Fouad Boualeb, Emery Pierson, Nicolas Doudeau, Clémence Nineuil, Ali Amad, Mohamed Daoudi
FG6
2025 Sheep Facial Pain Assessment Under Weighted Graph Neural Networks
abstract
Accurately recognizing and assessing pain in sheep is key to discern animal health and mitigating harmful situations. However, such accuracy is limited by the ability to manage automatic monitoring of pain in those animals. Facial expression scoring is a widely used and useful method to evaluate pain in both humans and other living beings. Researchers also analyzed the facial expressions of sheep to assess their health state and concluded that facial landmark detection and pain level prediction are essential. For this purpose, we propose a novel weighted graph neural network (WGNN) model to link sheep’s detected facial landmarks and define pain levels. Furthermore, we propose a new sheep facial landmarks dataset that adheres to the parameters of the Sheep Facial Expression Scale (SPFES). Currently, there is no comprehensive performance benchmark that specifically evaluates the use of graph neural networks (GNNs) on sheep facial landmark data to detect and measure pain levels. The YOLOv8n detector architecture achieves a mean average precision ($\mathbf{m A P}$) of $\mathbf{5 9. 3 0 \%}$ with the sheep facial landmarks dataset, among seven other detection models. The WGNN framework has an accuracy of $92.71 \%$ for tracking multiple facial parts expressions with the YOLOv8n lightweight on-board device deployment-capable model.
Alam Noor, Luís Almeida 0001, Mohamed Daoudi, Kai Li 0002, Eduardo Tovar
FG3
2025 Wearable-Derived Behavioral and Physiological Biomarkers for Classifying Unipolar and Bipolar Depression Severity
abstract
Depression is a complex mental disorder characterized by a range of observable and measurable indicators that go beyond traditional subjective assessments. Recent research has increasingly focused on objective, passive, and continuous monitoring using wearable devices to gain more precise insights into the physiological and behavioral aspects of depression. However, most existing studies primarily distinguish between healthy and depressed individuals, adopting a binary classification that fails to capture the heterogeneity of depressive disorders. In this study, we leverage wearable devices to predict depression subtypes—specifically unipolar and bipolar depression—aiming to identify distinctive biomarkers that could enhance diagnostic precision and support personalized treatment strategies. To this end, we introduce the CALYPSO dataset, designed for non-invasive detection of depression subtypes and symptomatology through physiological and behavioral signals, including blood volume pulse, electrodermal activity, body temperature, and three-axis acceleration. Additionally, we establish a benchmark on the dataset using well-known features and standard machine learning methods. Preliminary results indicate that features related to physical activity, extracted from accelerometer data, are the most effective in distinguishing between unipolar and bipolar depression, achieving an accuracy of 96.77%. Temperature-based features also showed high discriminative power, reaching an accuracy of 93.55%. These findings highlight the potential of physiological and behavioral monitoring for improving the classification of depressive subtypes, paving the way for more tailored clinical interventions.
Yassine Ouzar, Clémence Nineuil, Fouad Boualeb, Emery Pierson, Ali Amad, Mohamed Daoudi
FG6
2025 REACT 2025: the Third Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, a broad spectrum of human facial reactions might be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023 and REACT 2024 challenges, we are proposing the REACT 2025 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can be used to generate multiple appropriate, diverse, realistic and synchronised human-style facial reactions expressed by human listeners in response to an input stimulus (i.e., audio-visual behaviours expressed by their corresponding speakers). As a key of the challenge, we provide challenge participants with the first natural and large-scale multi-modal Multiple Appropriate Facial Reaction Generation (MAFRG) dataset (called MARS) recording 136 human-human dyadic interactions containing a total of 2856 interaction sessions covering five different topics. In addition, this paper also presents the challenge guidelines and the performance of our baselines on the two proposed sub-challenges: Offline MAFRG and Online MAFRG, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2025
Siyang Song, Micol Spitale, Xiangyu Kong 0001, Hengde Zhu, Cristina Palmero, Germán Barquero, Sergio Escalera, Michel F. Valstar, Mohamed Daoudi, Tobias Baur 0001, Fabien Ringeval, Andrew Howes 0001, Elisabeth André, Hatice Gunes
ACM Multimedia10
2025 Scanmove: Motion prediction and transfer for unregistered body meshes
Thomas Besnier, Sylvain Arguillère, Mohamed Daoudi
Comput. Graph.3
2025 Basis Restricted Elastic Shape Analysis on the Space of Unregistered Surfaces
Emmanuel Hartman, Emery Pierson, Martin Bauer 0004, Mohamed Daoudi, Nicolas Charon
Int. J. Comput. Vis.4
2024 ScanTalk: 3D Talking Heads from Unregistered Scans
Federico Nocentini, Thomas Besnier, Claudio Ferrari, Sylvain Arguillère, Stefano Berretti, Mohamed Daoudi
ECCV (29)6
2024 MGRFormer: A Multimodal Transformer Approach for Surgical Gesture Recognition
abstract
Automatic surgical gesture recognition has the potential to revolutionize the field of surgery by enhancing patient care, surgical training, and our understanding of surgical skills. By integrating kinematic data, which precisely captures hand movements, with video data for contextual understanding, multimodal machine learning can greatly enhance the accuracy of surgical gesture recognition systems by capturing complementary knowledge. Recent research has highlighted the capabilities of Transformer-based models for temporal action segmentation. A key component of these models is the iterative refinement module, which enhances predictions using contextual data. In this study, we propose MGRFormer, a novel multimodal framework that leverages the interaction between kinematics and visual data at the refinement stage for the task of surgical gesture recognition. We evaluated our MGRFormer on the VTS dataset, and the results demonstrated that our approach outperformed unimodal and multimodal state-of-the-art methods by a large margin.
Kevin Feghoul, Deise Santana Maia, Mehdi El Amrani, Mohamed Daoudi, Ali Amad
FG4
2024 GM-GAN: Geometric Generative Models Based on Morphological Equivariant PDEs and GANs
El Hadji S. Diop, Thierno Fall, Alioune Mbengue, Mohamed Daoudi
ICPR (25)4
2024 Bipartite Graph Diffusion Model for Human Interaction Generation
abstract
The generation of natural human motion interactions is a hot topic in computer vision and computer animation. It is a challenging task due to the diversity of possible human motion interactions. Diffusion models, which have already shown remarkable generative capabilities in other domains, are a good candidate for this task. In this paper, we introduce a novel bipartite graph diffusion method (BiGraphDiff) to generate human motion interactions between two persons. Specifically, bipartite node sets are constructed to model the inherent geometric constraints between skeleton nodes during interactions. The interaction graph diffusion model is transformer-based, combining some state-of-theart motion methods. We show that the proposed achieves new state-of-the-art results on leading benchmarks for the human interaction generation task. Code, pre-trained models and additional results are available at https://github.com/CRISTAL-3DSAM/BiGraphDiff.
Baptiste Chopin, Hao Tang 0005, Mohamed Daoudi
WACV3
2024 Generating Multiple 4D Expression Transitions by Learning Face Landmark Trajectories
abstract
In this work, we address the problem of 4D facial expressions generation. This is usually addressed by animating a neutral 3D face to reach an expression peak, and then get back to the neutral state. In the real world though, people show more complex expressions, and switch from one expression to another. We thus propose a new model that generates transitions between different expressions, and synthesizes long and composed 4D expressions. This involves three sub-problems: (i) modeling the temporal dynamics of expressions, (ii) learning transitions between them, and (iii) deforming a generic mesh. We propose to encode the temporal evolution of expressions using the motion of a set of 3D landmarks, that we learn to generate by training a manifold-valued GAN (Motion3DGAN). To allow the generation of composed expressions, this model accepts two labels encoding the starting and the ending expressions. The final sequence of meshes is generated by a Sparse2Dense mesh Decoder (S2D-Dec) that maps the landmark displacements to a dense, per-vertex displacement of a known mesh topology. By explicitly working with motion trajectories, the model is totally independent from the identity. Extensive experiments on five public datasets show that our proposed approach brings significant improvements with respect to previous solutions, while retaining good generalization to unseen data.
Naima Otberdout, Claudio Ferrari, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo
IEEE Trans. Affect. Comput.3
2024 Transformer-Based Self-Supervised Multimodal Representation Learning for Wearable Emotion Recognition
abstract
Recently, wearable emotion recognition based on peripheral physiological signals has drawn massive attention due to its less invasive nature and its applicability in real-life scenarios. However, how to effectively fuse multimodal data remains a challenging problem. Moreover, traditional fully-supervised based approaches suffer from overfitting given limited labeled data. To address the above issues, we propose a novel self-supervised learning (SSL) framework for wearable emotion recognition, where efficient multimodal fusion is realized with temporal convolution-based modality-specific encoders and a transformer-based shared encoder, capturing both intra-modal and inter-modal correlations. Extensive unlabeled data is automatically assigned labels by five signal transforms, and the proposed SSL model is pre-trained with signal transformation recognition as a pretext task, allowing the extraction of generalized multimodal representations for emotion-related downstream tasks. For evaluation, the proposed SSL model was first pre-trained on a large-scale self-collected physiological dataset and the resulting encoder was subsequently frozen or fine-tuned on three public supervised emotion recognition datasets. Ultimately, our SSL-based method achieved state-of-the-art results in various emotion classification tasks. Meanwhile, the proposed model was proved to be more accurate and robust compared to fully-supervised methods on low data regimes.
Yujin Wu, Mohamed Daoudi, Ali Amad
IEEE Trans. Affect. Comput.2
2023 Spatial-Temporal Graph Transformer for Surgical Skill Assessment in Simulation Sessions
Kevin Feghoul, Deise Santana Maia, Mehdi El Amrani, Mohamed Daoudi, Ali Amad
CIARP4
2023 The Florence 4D Facial Expression Dataset
abstract
Human facial expressions change dynamically, so their recognition / analysis should be conducted by accounting for the temporal evolution of face deformations either in 2D or 3D. While abundant 2D video data do exist, this is not the case in 3D, where few 3D dynamic (4D) datasets were released for public use. The negative consequence of this scarcity of data is amplified by current deep learning based-methods for facial expression analysis that require large quantities of variegate samples to be effectively trained. With the aim of smoothing such limitations, in this paper we propose a large dataset, named Florence 4D, composed of dynamic sequences of 3D face models, where a combination of synthetic and real identities exhibit an unprecedented variety of 4D facial expressions, with variations that include the classical neutral-apex transition, but generalize to expression-to-expression. All these characteristics are not exposed by any of the existing 4D datasets and they cannot even be obtained by combining more than one dataset. We strongly believe that making such a data corpora publicly available to the community will allow designing and experimenting new applications that were not possible to investigate till now. To show at some extent the difficulty of our data in terms of different identities and varying expressions, we also report a baseline experimentation on the proposed dataset that can be used as baseline.
Filippo Principi, Stefano Berretti, Claudio Ferrari, Naima Otberdout, Mohamed Daoudi, Alberto Del Bimbo
FG5
2023 BaRe-ESA: A Riemannian Framework for Unregistered Human Body Shapes
abstract
We present Basis Restricted Elastic Shape Analysis (BaRe-ESA), a novel Riemannian framework for human body scan representation, interpolation and extrapolation. BaRe-ESA operates directly on unregistered meshes, i.e., without the need to establish prior point to point correspondences or to assume a consistent mesh structure. Our method relies on a latent space representation, which is equipped with a Riemannian (non-Euclidean) metric associated to an invariant higher-order metric on the space of surfaces. Experimental results on the FAUST and DFAUST datasets show that BaRe-ESA brings significant improvements with respect to previous solutions in terms of shape registration, interpolation and extrapolation. The efficiency and strength of our model is further demonstrated in applications such as motion transfer and random generation of body shape and pose.
Emmanuel Hartman, Emery Pierson, Martin Bauer 0004, Nicolas Charon, Mohamed Daoudi
ICCV5
2023 Toward Mesh-Invariant 3D Generative Deep Learning with Geometric Measures
Thomas Besnier, Sylvain Arguillère, Emery Pierson, Mohamed Daoudi
Comput. Graph.4
2023 Guest Editorial : Learning with Manifolds in Computer Vision
Mohamed Daoudi, Mehrtash Harandi, Vittorio Murino
Image Vis. Comput.1
2023 Interaction Transformer for Human Reaction Generation
abstract
We address the challenging task of human reaction generation, which aims to generate a corresponding reaction based on an input action. Most of the existing works do not focus on generating and predicting the reaction and cannot generate the motion when only the action is given as input. To address this limitation, we propose a novel interaction Transformer (InterFormer) consisting of a Transformer network with both temporal and spatial attention. Specifically, temporal attention captures the temporal dependencies of the motion of both characters and of their interaction, while spatial attention learns the dependencies between the different body parts of each character and those which are part of the interaction. Moreover, we propose using graphs to increase the performance of spatial attention via an interaction distance module that helps focus on nearby joints from both characters. Extensive experiments on the SBU interaction, K3HI, and DuetDance datasets demonstrate the effectiveness of InterFormer. Our method is general and can be used to generate more complex and long-term interactions. We also provide videos of generated reactions and the code with pre-trained models athttps://github.com/CRISTAL-3DSAM/InterFormer
Baptiste Chopin, Hao Tang 0005, Naima Otberdout, Mohamed Daoudi, Nicu Sebe
IEEE Trans. Multim.4
2022 Sparse to Dense Dynamic 3D Facial Expression Generation
abstract
In this paper, we propose a solution to the task of generating dynamic 3D facial expressions from a neutral 3D face and an expression label. This involves solving two sub-problems: (i) modeling the temporal dynamics of expressions, and (ii) deforming the neutral mesh to obtain the expressive counterpart. We represent the temporal evolution of expressions using the motion of a sparse set of 3D landmarks that we learn to generate by training a manifold-valued GAN (Motion3DGAN). To better encode the expression-induced deformation and disentangle it from the identity information, the generated motion is represented as per-frame displacement from a neutral configuration. To generate the expressive meshes, we train a Sparse2Dense mesh Decoder (S2D-Dec) that maps the landmark displacements to a dense, per-vertex displacement. This allows us to learn how the motion of a sparse set of landmarks influences the deformation of the overall face surface, independently from the identity. Experimental results on the CoMA and D3DFACS datasets show that our solution brings significant improvements with respect to previous solutions in terms of both dynamic expression generation and mesh reconstruction, while retaining good generalization to unseen data. Code and models are available at https://github.com/CRISTAL-3DSAM/Sparse2Dense.
Naima Otberdout, Claudio Ferrari, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo
CVPR3
2022 3D Shape Sequence of Human Comparison and Classification Using Current and Varifolds
Emery Pierson, Mohamed Daoudi, Sylvain Arguillère
ECCV (3)2
2022 Fusion of Physiological and Behavioural Signals on SPD Manifolds with Application to Stress and Pain Detection
abstract
Existing multimodal stress/pain recognition approaches generally extract features from different modalities independently and thus ignore cross-modality correlations. This paper proposes a novel geometric framework for multimodal stress/pain detection utilizing Symmetric Positive Definite (SPD) matrices as a representation that incorporates the correlation relationship of physiological and behavioural signals from covariance and cross-covariance. Considering the non-linearity of the Riemannian manifold of SPD matrices, well-known machine learning techniques are not suited to classify these matrices. Therefore, a tangent space mapping method is adopted to map the derived SPD matrix sequences to the vector sequences in the tangent space where the LSTM-based network can be applied for classification. The proposed framework has been evaluated on two public multimodal datasets, achieving both the state-of-the-art results for stress and pain detection tasks.
Yujin Wu, Mohamed Daoudi, Ali Amad, Laurent Sparrow, Fabien D'Hondt
SMC2
2022 A Riemannian Framework for Analysis of Human Body Surface
abstract
We propose a novel framework for comparing 3D human shapes under the change of shape and pose. This problem is challenging since 3D human shapes vary significantly across subjects and body postures. We solve this problem by using a Riemannian approach. Our core contribution is the mapping of the human body surface to the space of metrics and normals. We equip this space with a family of Riemannian metrics, called Ebin (or DeWitt) metrics. We treat a human body surface as a point in a "shape space" equipped with a family of Riemannian metrics. The family of metrics is invariant under rigid motions and reparametrizations; hence it induces a metric on the "shape space" of surfaces. Using the alignment of human bodies with a given template, we show that this family of metrics allows us to distinguish the changes in shape and pose. The proposed framework has several advantages. First, we define a family of metrics with desired invariance properties for the comparison of human shape. Second, we present an efficient framework to compute geodesic paths between human shape given the chosen metric. Third, this framework provides some basic tools for statistical shape analysis of human body surfaces. Finally, we demonstrate the utility of the proposed frame-work in pose and shape retrieval of human body.
Emery Pierson, Mohamed Daoudi, Alice Barbara Tumpach
WACV2
2022 Foreword to the Special Section on 3D Object Retrieval 2022 Symposium (3DOR2022)
Stefano Berretti, Theoharis Theoharis, Mohamed Daoudi, Claudio Ferrari, Remco C. Veltkamp
Comput. Graph.3
2022 Projection-based classification of surfaces for 3D human mesh sequence retrieval
Emery Pierson, Juan Carlos Álvarez Paiva, Mohamed Daoudi
Comput. Graph.3
2022 Dynamic Facial Expression Generation on Hilbert Hypersphere With Conditional Wasserstein Generative Adversarial Nets
abstract
In this work, we propose a novel approach for generating videos of the six basic facial expressions given a neutral face image. We propose to exploit the face geometry by modeling the facial landmarks motion as curves encoded as points on a hypersphere. By proposing a conditional version of manifold-valued Wasserstein generative adversarial network (GAN) for motion generation on the hypersphere, we learn the distribution of facial expression dynamics of different classes, from which we synthesize new facial expression motions. The resulting motions can be transformed to sequences of landmarks and then to images sequences by editing the texture information using another conditional Generative Adversarial Network. To the best of our knowledge, this is the first work that explores manifold-valued representations with GAN to address the problem of dynamic facial expression generation. We evaluate our proposed approach both quantitatively and qualitatively on two public datasets; Oulu-CASIA and MUG Facial Expression. Our experimental results demonstrate the effectiveness of our approach in generating realistic videos with continuous motion, realistic appearance and identity preservation. We also show the efficiency of our framework for dynamic facial expressions generation, dynamic facial expression transfer and data augmentation for training improved emotion recognition models.
Naima Otberdout, Mohamed Daoudi, Anis Kacem 0001, Lahoucine Ballihi, Stefano Berretti
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Automatic Estimation of Self-Reported Pain by Trajectory Analysis in the Manifold of Fixed Rank Positive Semi-Definite Matrices
abstract
We propose an automatic method to estimate self-reported pain based on facial landmarks extracted from videos. For each video sequence, we decompose the face into four different regions and the pain intensity is measured by modeling the dynamics of facial movement using the landmarks of these regions. A formulation based on Gram matrices is used for representing the trajectory of landmarks on the Riemannian manifold of symmetric positive semi-definite matrices of fixed rank. A curve fitting algorithm is used to smooth the trajectories and temporal alignment is performed to compute the similarity between the trajectories on the manifold. A Support Vector Regression classifier is then trained to encode extracted trajectories into pain intensity levels consistent with self-reported pain intensity measurement. Finally, a late fusion of the estimation for each region is performed to obtain the final predicted pain level. The proposed approach is evaluated on two publicly available datasets, the UNBCMcMaster Shoulder Pain Archive and the Biovid Heat Pain dataset. We compared our method to the state-of-the-art on both datasets using different testing protocols, showing the competitiveness of the proposed approach.
Benjamin Szczapa, Mohamed Daoudi, Stefano Berretti, Pietro Pala, Alberto Del Bimbo, Zakia Hammal
IEEE Trans. Affect. Comput.2
2021 Human Motion Prediction Using Manifold-Aware Wasserstein GAN
abstract
Human motion prediction aims to forecast future human poses given a prior pose sequence. The discontinuity of the predicted motion and the performance deterioration in long-term horizons are still the main challenges encountered in current literature. In this work, we tackle these issues by using a compact manifold-valued representation of human motion. Specifically, we model the temporal evolution of the 3D human poses as trajectory, what allows us to map human motions to single points on a sphere manifold. To learn these non-Euclidean representations, we build a manifold-aware Wasserstein generative adversarial model that captures the temporal and spatial dependencies of human motion through different losses. Extensive experiments show that our approach outperforms the state-of-the-art on CMU MoCap and Human 3.6M datasets. Our qualitative results show the smoothness of the predicted motions.
Baptiste Chopin, Naima Otberdout, Mohamed Daoudi, Angela Bartolo
FG3
2020 Modelling the Statistics of Cyclic Activities by Trajectory Analysis on the Manifold of Positive-Semi-Definite Matrices
abstract
In this paper, a model is presented to extract statistical summaries to characterize the repetition of a cyclic body action, for instance a gym exercise, for the purpose of checking the compliance of the observed action to a template one and highlighting the parts of the action that are not correctly executed (if any). The proposed system relies on a Riemannian metric to compute the distance between two poses in such a way that the geometry of the manifold where the pose descriptors lie is preserved; a model to detect the begin and end of each cycle; a model to temporally align the poses of different cycles so as to accurately estimate the cross-sectional mean and variance of poses across different cycles. The proposed model is demonstrated using gym videos taken from the Internet.
Ettore Maria Celozzi, Luca Ciabini, Luca Cultrera, Pietro Pala, Stefano Berretti, Mohamed Daoudi, Alberto Del Bimbo
FG6
2020 Face and Gesture Analysis for Health Informatics
abstract
The goal of Face and Gesture Analysis for Health Informatics's workshop is to share and discuss the achievements as well as the challenges in using computer vision and machine learning for automatic human behavior analysis and modeling for clinical research and healthcare applications. The workshop aims to promote current research and support growth of multidisciplinary collaborations to advance this groundbreaking research. The meeting gathers scientists working in related areas of computer vision and machine learning, multi-modal signal processing and fusion, human centered computing, behavioral sensing, assistive technologies, and medical tutoring systems for healthcare applications and medicine.
Zakia Hammal, Di Huang 0001, Kevin Bailly, Liming Chen 0002, Mohamed Daoudi
ICMI5
2020 Hybrid Approach for 3D Head Reconstruction: Using Neural Networks and Visual Geometry
abstract
Recovering the 3D geometric structure of a face from a single input image is a challenging active research area in computer vision. In this paper, we present a novel method for reconstructing 3D heads from a single or multiple image(s) using a hybrid approach based on deep learning and geometric techniques. We propose an encoder-decoder network based on the U-net architecture and trained on synthetic data only. It predicts both pixel-wise normal vectors and landmarks maps from a single input photo. Landmarks are used for the pose computation and the initialization of the optimization problem, which, in turn, reconstructs the 3D head geometry by using a parametric morphable model and normal vector fields. State-of-the-art results are achieved through qualitative and quantitative evaluation tests on both single and multi-view settings. Despite the fact that the model was trained only on synthetic data, it successfully recovers 3D geometry and precise poses for realworld images.
Oussema Bouafif, Bogdan Khomutenko, Mohamed Daoudi
ICPR3
2020 Automatic Estimation of Self-Reported Pain by Interpretable Representations of Motion Dynamics
abstract
We propose an automatic method for pain intensity measurement from video. For each video, pain intensity was measured using the dynamics of facial movement using 66 facial points. Gram matrices formulation was used for facial points trajectory representations on the Riemannian manifold of symmetric positive semi-definite matrices of fixed rank. Curve fitting and temporal alignment were then used to smooth the extracted trajectories. A Support Vector Regression model was then trained to encode the extracted trajectories into ten pain intensity levels consistent with the Visual Analogue Scale for pain intensity measurement. The proposed approach was evaluated using the UNBC McMaster Shoulder Pain Archive and was compared to the state-of-the-art on the same data. Using both 5-fold cross-validation and leave-one-subject-out cross-validation, our results are competitive with respect to state-of-the-art methods.
Benjamin Szczapa, Mohamed Daoudi, Stefano Berretti, Pietro Pala, Alberto Del Bimbo, Zakia Hammal
ICPR2
2020 A Novel Geometric Framework on Gram Matrix Trajectories for Human Behavior Understanding
abstract
In this paper, we propose a novel space-time geometric representation of human landmark configurations and derive tools for comparison and classification. We model the temporal evolution of landmarks as parametrized trajectories on the Riemannian manifold of positive semidefinite matrices of fixed-rank. Our representation has the benefit to bring naturally a second desirable quantity when comparing shapes-the spatial covariance-in addition to the conventional affine-shape representation. We derived then geometric and computational tools for rate-invariant analysis and adaptive re-sampling of trajectories, grounding on the Riemannian geometry of the underlying manifold. Specifically, our approach involves three steps: (1) landmarks are first mapped into the Riemannian manifold of positive semidefinite matrices of fixed-rank to build time-parameterized trajectories; (2) a temporal warping is performed on the trajectories, providing a geometry-aware (dis-)similarity measure between them; (3) finally, a pairwise proximity function SVM is used to classify them, incorporating the (dis-)similarity measure into the kernel function. We show that such representation and metric achieve competitive results in applications as action recognition and emotion recognition from 3D skeletal data, and facial expression recognition from videos. Experiments have been conducted on several publicly available up-to-date benchmarks.
Anis Kacem 0001, Mohamed Daoudi, Boulbaba Ben Amor, Stefano Berretti, Juan Carlos Álvarez Paiva
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Automatic Analysis of Facial Expressions Based on Deep Covariance Trajectories
abstract
In this article, we propose a new approach for facial expression recognition (FER) using deep covariance descriptors. The solution is based on the idea of encoding local and global deep convolutional neural network (DCNN) features extracted from still images, in compact local and global covariance descriptors. The space geometry of the covariance matrices is that of symmetric positive definite (SPD) matrices. By conducting the classification of static facial expressions using a support vector machine (SVM) with a valid Gaussian kernel on the SPD manifold, we show that deep covariance descriptors are more effective than the standard classification with fully connected layers and softmax. Besides, we propose a completely new and original solution to model the temporal dynamic of facial expressions as deep trajectories on the SPD manifold. As an extension of the classification pipeline of covariance descriptors, we apply SVM with valid positive definite kernels derived from global alignment for deep covariance trajectories classification. By performing extensive experiments on the Oulu-CASIA, CK+, static facial expression in the wild (SFEW), and acted facial expressions in the wild (AFEW) data sets, we show that both the proposed static and dynamic approaches achieve the state-of-the-art performance for FER outperforming many recent approaches.
Naima Otberdout, Anis Kacem 0001, Mohamed Daoudi, Lahoucine Ballihi, Stefano Berretti
IEEE Trans. Neural Networks Learn. Syst.3
2019 When Computers Decode your Social Intention
abstract
In this demo session, we will propose our framework that is based on our paper [1] . In real time, we proposed to analyze the trajectories of the human arm to predict social intention (personal or social intention). The trajectories of different 3D markers acquired by Mocap system are defined in shape spaces of open curves, thus analyze in a Riemannian manifold. The results obtained in the experiments on a new dataset show an average recognition of about 68% for the proposed method, which is comparable with the average score produced by human evaluation. The experimental results show also that the classification rate could be used to improve social communication between human and virtual agents. To the best of our knowledge, this is the first demo in real time, which uses computer vision techniques to analyze the effect of social intention on motor action for improving the social communication between human and avatar. The main goal is to categorize the user intention among two classes denote {personal, social}. This experimentation contains 3 parts: a) data acquisitions and a learning step; b) classification; c) Kinematic analysis of the evolution of subjects to interact with the avatar. To successfully drive our study, all the using scripts are writing under Matlab and C/C++. Then the using equipments are: 1) Qualisys motion capture camera (qualisys system). The qualisys system is delivered with a desk computer with 8 GB, a processor Intel core i7-4770k (8 CPUs) at 3.5 GHz. The frequency of those cameras can varies from 100 to 500 Hz. A black glove equipped with infrared reflective markers, all those equipments are also provided by qualisys system. 2) A Matlab software (version R2014a) installed on a desk computer (qualisys system); the Qualisys system provide a specific driver that allow to couple all the Matlab scripts with their system. Thus, it is possible to command all the cameras directly from Matlab for real time analysis, see Fig. 1 .
Paul Audain Desrosiers, Mohamed Daoudi, Yann Coello
FG2
2019 2D Landmark-Based Facial Asymmetry Assessment in the Clinical Case of Facial Paralysis
abstract
In this paper, we propose a novel technique for quantifying the facial asymmetry from 2D videos to evaluate facial paralysis treatments based on Botulinum Toxin (BT) injections. Our approach uses 2D facial landmarks and barycentric coordinates to objectively quantify the facial asymmetry across 2D videos. To assess our approach, a new dataset of 2D videos, containing eighteen patients before and after the treatments have been collected. For each patient, we have collected nine facial expressions. Experimental results on the newly collected dataset show that the proposed approach provides promising results in concordance with clinical annotations.
Benjamin Szczapa, Mohamed Daoudi, Anis Kacem 0001, Pierre Guerreschi, Ludwig Gebert, Juan Carlos Álvarez Paiva
FG2
2019 Lip reading with Hahn Convolutional Neural Networks
Abderrahim Mesbah, Aissam Berrahou, Hicham Hammouchi, Hassan Berbia, Hassan Qjidaa, Mohamed Daoudi
Image Vis. Comput.6
2019 Magnifying Subtle Facial Motions for Effective 4D Expression Recognition
abstract
In this paper, an effective approach is proposed for automatic 4D Facial Expression Recognition (FER). It combines two growing but disparate ideas in the domain of computer vision, i.e., computing spatial facial deformations using a Riemannian method and magnifying them by a temporal filtering technique. Key frames highly related to facial expressions are first extracted from a long 4D video through a spectral clustering process, forming the Onset-Apex-Offset flow. It is then analyzed to capture the spatial deformations based on Dense Scalar Fields (DSF), where registration and comparison of neighboring 3D faces are jointly led. The generated temporal evolution of these deformations is further fed into a magnification method to amplify facial activities over time. The proposed approach allows revealing subtle deformations and thus improves the emotion classification performance. Experiments are conducted on the BU-4DFE and BP-4D databases, and competitive results are achieved compared to the state-of-the-art.
Qingkai Zhen, Di Huang 0001, Hassen Drira, Boulbaba Ben Amor, Yunhong Wang 0001, Mohamed Daoudi
IEEE Trans. Affect. Comput.6
2018 Deep Covariance Descriptors for Facial Expression Recognition
Naima Otberdout, Anis Kacem 0001, Mohamed Daoudi, Lahoucine Ballihi, Stefano Berretti
BMVC3
2018 A New Computational Approach to Identify Human Social Intention in Action
abstract
In this paper, we propose to analyze the trajectories of the human arm to predict social intention (personal or social intention). The trajectories of different 3D markers acquired by Mocap system, are defined in shape spaces of open curves. The results obtained in the experiments on a new dataset show an average recognition of about 68% for the proposed method, which is comparable with the average score produced by human evaluation. The experimental results show also that the classification rate could be used to improve social communication between human and virtual agents. To the best of our knowledge, this is the first paper which uses computer vision techniques to analyze the effect of social intention on motor action for improving the social communication between human and avatar.
Mohamed Daoudi, Yann Coello, Paul Audain Desrosiers, Laurent Ott
FG1
2018 Barycentric Representation and Metric Learning for Facial Expression Recognition
abstract
In this paper, we tackle the problem of dynamic facial expression recognition. An affine-invariant facial shape representation based on barycentric coordinates is proposed and related to the Grassmannian representation. Unlike the latter, the barycentric representation allows us to work directly on Euclidean space and apply a metric learning algorithm to find a suitable metric that is discriminative enough to compare facial shapes under different expressions. Finally, we exploit the learned metric in a machinery combining a Dynamic Time Warping (DTW) phase and a pairwise proximity function SVM classifier for a rate-invariant classification of the facial sequences. Experiments on the AFEW dataset show the effectiveness of our approach while exploiting only geometric features.
Anis Kacem 0001, Mohamed Daoudi, Juan Carlos Álvarez Paiva
FG2
2018 Detecting Depression Severity by Interpretable Representations of Motion Dynamics
abstract
Recent breakthroughs in deep learning using automated measurement of face and head motion have made possible the first objective measurement of depression severity. While powerful, deep learning approaches lack interpretability. We developed an interpretable method of automatically measuring depression severity that uses barycentric coordinates of facial landmarks and a Lie-algebra based rotation matrix of 3D head motion. Using these representations, kinematic features are extracted, preprocessed, and encoded using Gaussian Mixture Models (GMM) and Fisher vector encoding. A multi-class SVM is used to classify the encoded facial and head movement dynamics into three levels of depression severity. The proposed approach was evaluated in adults with history of chronic depression. The method approached the classification accuracy of state-of-the-art deep learning while enabling clinically and theoretically relevant findings. The velocity and acceleration of facial movement strongly mapped onto depression severity symptoms consistent with clinical data and theory.
Anis Kacem 0001, Zakia Hammal, Mohamed Daoudi, Jeffrey F. Cohn
FG3
2018 Spontaneous Expression Detection from 3D Dynamic Sequences by Analyzing Trajectories on Grassmann Manifolds
abstract
In this paper, we propose a framework for online spontaneous emotion detection, such as happiness or physical pain, from depth videos. Our approach consists on mapping the video streams onto a Grassmann manifold (i.e., space of k-dimensional linear subspaces) to form time-parameterized trajectories. To this end, depth videos are decomposed into short-time subsequences, each approximated by a k-dimensional linear subspace, which is in turn a point on the Grassmann manifold. Then, the temporal evolution of subspaces gives rise to a precise mathematical representation of trajectories on the underlying manifold. In the final step, extracted spatio-temporal features based on computing the velocity vectors along the trajectories, termed Geometric Motion History (GMH), are fed to an early event detector based on Structured Output SVM, which enables online emotion detection from partially-observed data. Experimental results obtained on the publicly available Cam3D Kinect and BP4D-spontaneous databases validate the proposed solution. The first database has served to exemplify the proposed framework using depth sequences of the upper part of the body collected using depth-consumer cameras, while the second database allowed the application of the same framework to physical pain detection from high-resolution and long 3D-face sequences.
Taleb Alashkar, Boulbaba Ben Amor, Mohamed Daoudi, Stefano Berretti
IEEE Trans. Affect. Comput.3
2018 Introduction to the Special Issue on Representation, Analysis, and Recognition of 3D Humans
abstract
No abstract available.
Stefano Berretti, Mohamed Daoudi, Pavan Turaga, Anup Basu
ACM Trans. Multim. Comput. Commun. Appl.2
2018 Representation, Analysis, and Recognition of 3D Humans: A Survey
abstract
Computer Vision and Multimedia solutions are now offering an increasing number of applications ready for use by end users in everyday life. Many of these applications are centered for detection, representation, and analysis of face and body. Methods based on 2D images and videos are the most widespread, but there is a recent trend that successfully extends the study to 3D human data as acquired by a new generation of 3D acquisition devices. Based on these premises, in this survey, we provide an overview on the newly designed techniques that exploit 3D human data and also prospect the most promising current and future research directions. In particular, we first propose a taxonomy of the representation methods, distinguishing between spatial and temporal modeling of the data. Then, we focus on the analysis and recognition of 3D humans from 3D static and dynamic data, considering many applications for body and face.
Stefano Berretti, Mohamed Daoudi, Pavan Turaga, Anup Basu
ACM Trans. Multim. Comput. Commun. Appl.2
2017 A Novel Space-Time Representation on the Positive Semidefinite Cone for Facial Expression Recognition
abstract
In this paper, we study the problem of facial expression recognition using a novel space-time geometric representation. We describe the temporal evolution of facial landmarks as parametrized trajectories on the Riemannian manifold of positive semidefinite matrices of fixed-rank. Our representation has the advantage to bring naturally a second desirable quantity when comparing shapes - the spatial covariance - in addition to the conventional affine-shape representation. We derive then geometric and computational tools for rate-invariant analysis and adaptive re-sampling of trajectories, grounding on the Riemannian geometry of the manifold. Specifically, our approach involves three steps: 1) facial landmarks are first mapped into the Riemannian manifold of positive semidefinite matrices of rank 2, to build time-parameterized trajectories; 2) a temporal alignment is performed on the trajectories, providing a geometry-aware (dis-)similarity measure between them; 3) finally, pairwise proximity function SVM (ppfSVM) is used to classify them, incorporating the latter (dis-)similarity measure into the kernel function. We show the effectiveness of the proposed approach on four publicly available benchmarks (CK+, MMI, Oulu-CASIA, and AFEW). The results of the proposed approach are comparable to or better than the state-of-the-art methods when involving only facial landmarks.
Anis Kacem 0001, Mohamed Daoudi, Boulbaba Ben Amor, Juan Carlos Álvarez Paiva
ICCV2
2017 Analyzing of facial paralysis by shape analysis of 3D face sequences
Paul Audain Desrosiers, Yasmine Bennis, Mohamed Daoudi, Boulbaba Ben Amor, Pierre Guerreschi
Image Vis. Comput.3
2017 Joint gender, ethnicity and age estimation from 3D faces: An experimental illustration of their correlations
Baiqiang Xia, Boulbaba Ben Amor, Mohamed Daoudi
Image Vis. Comput.3
2017 Motion segment decomposition of RGB-D sequences for human behavior understanding
Maxime Devanne, Stefano Berretti, Pietro Pala, Hazem Wannous, Mohamed Daoudi, Alberto Del Bimbo
Pattern Recognit.5
2016 Novel generative model for facial expressions based on statistical shape analysis of landmarks trajectories
abstract
We propose a novel geometric framework for analyzing spontaneous facial expressions, with the specific goal of comparing, matching, and averaging the shapes of landmarks trajectories. Here we represent facial expressions by the motion of the landmarks across the time. The trajectories are represented by curves. We use elastic shape analysis of these curves to develop a Riemannian framework for analyzing shapes of these trajectories. In terms of empirical evaluation, our results on two databases: UvA-NEMO and Cohn-Kanade CK+ are very promising. From a theoretical perspective, this framework allows formal statistical inferences, such as generation of facial expressions.
Paul Audain Desrosiers, Mohamed Daoudi, Maxime Devanne
ICPR2
2016 Learning shape variations of motion trajectories for gait analysis
abstract
The analysis of human gait is more and more investigated due to its large panel of potential applications in various domains, like rehabilitation, deficiency diagnosis, surveillance and movement optimization. In addition, the release of depth sensors offers new opportunities to achieve gait analysis in a non-intrusive context. In this paper, we propose a gait analysis method from depth sequences by analyzing separately each step so as to be robust to gait duration and incomplete cycles. We analyze the shape of the motion trajectory as signature of the gait and consider shape variations within a Riemannian manifold to learn step models. During classification, the derivation of each performed step is evaluated in an online manner to qualitatively analyze the gait. Experiments are carried out in the context of abnormal gait detection and person re-identification trough gait recognition. Results demonstrated the potential of the method in both scenarios.
Maxime Devanne, Hazem Wannous, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICPR3
2016 Magnifying subtle facial motions for 4D Expression Recognition
abstract
In this paper, we propose an effective approach for automatic 4D Facial Expression Recognition (FER). The flow of 3D facial scans is first modeled to capture spatial deformations based on the recently-developed Riemannian approach, namely Dense Scalar Fields (DSF), where registration and comparison of neighboring 3D face frames are jointly led. The deformations are then fed into a temporal filtering based magnification step to amplify the slight facial actions over time. The proposed method allows revealing subtle (hidden) deformations which enhances the performance in classification. We evaluate our approach on the BU-4DFE dataset, and the state-of-art accuracy up to 94.18% is achieved, which is superior to the top one so far reported, clearly demonstrating its effectiveness.
Qingkai Zhen, Di Huang 0001, Yunhong Wang 0001, Hassen Drira, Boulbaba Ben Amor, Mohamed Daoudi
ICPR6
2016 Gauge Invariant Framework for Shape Analysis of Surfaces
abstract
This paper describes a novel framework for computing geodesic paths in shape spaces of spherical surfaces under an elastic Riemannian metric. The novelty lies in defining this Riemannian metric directly on the quotient (shape) space, rather than inheriting it from pre-shape space, and using it to formulate a path energy that measures only the normal components of velocities along the path. In other words, this paper defines and solves for geodesics directly on the shape space and avoids complications resulting from the quotient operation. This comprehensive framework is invariant to arbitrary parameterizations of surfaces along paths, a phenomenon termed as gauge invariance. Additionally, this paper makes a link between different elastic metrics used in the computer science literature on one hand, and the mathematical literature on the other hand, and provides a geometrical interpretation of the terms involved. Examples using real and simulated 3D objects are provided to help illustrate the main ideas.
Alice Barbara Tumpach, Hassen Drira, Mohamed Daoudi, Anuj Srivastava
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 A Grassmann framework for 4D facial shape analysis
Taleb Alashkar, Boulbaba Ben Amor, Mohamed Daoudi, Stefano Berretti
Pattern Recognit.3
2015 Accurate 3D action recognition using learning on the Grassmann manifold
Rim Slama, Hazem Wannous, Mohamed Daoudi, Anuj Srivastava
Pattern Recognit.3
2015 Combining face averageness and symmetry for 3D-based gender classification
Baiqiang Xia, Boulbaba Ben Amor, Hassen Drira, Mohamed Daoudi, Lahoucine Ballihi
Pattern Recognit.4
2015 3-D Human Action Recognition by Shape Analysis of Motion Trajectories on Riemannian Manifold
abstract
Recognizing human actions in 3-D video sequences is an important open problem that is currently at the heart of many research domains including surveillance, natural interfaces and rehabilitation. However, the design and development of models for action recognition that are both accurate and efficient is a challenging task due to the variability of the human pose, clothing and appearance. In this paper, we propose a new framework to extract a compact representation of a human action captured through a depth sensor, and enable accurate action recognition. The proposed solution develops on fitting a human skeleton model to acquired data so as to represent the 3-D coordinates of the joints and their change over time as a trajectory in a suitable action space. Thanks to such a 3-D joint-based framework, the proposed solution is capable to capture both the shape and the dynamics of the human body, simultaneously. The action recognition problem is then formulated as the problem of computing the similarity between the shape of trajectories in a Riemannian manifold. Classification using k-nearest neighbors is finally performed on this manifold taking advantage of Riemannian geometry in the open curve shape space. Experiments are carried out on four representative benchmarks to demonstrate the potential of the proposed solution in terms of accuracy/latency for a low-latency action recognition. Comparative results with state-of-the-art methods are reported.
Maxime Devanne, Hazem Wannous, Stefano Berretti, Pietro Pala, Mohamed Daoudi, Alberto Del Bimbo
IEEE Trans. Cybern.5
2014 Grassmannian Representation of Motion Depth for 3D Human Gesture and Action Recognition
abstract
Recently developed commodity depth sensors open up new possibilities of dealing with rich descriptors, which capture geometrical features of the observed scene. Here, we propose an original approach to represent geometrical features extracted from depth motion space, which capture both geometric appearance and dynamic of human body simultaneously. In this approach, sequence features are modeled temporally as subspaces lying on the Grassmann manifold. Classification task is carried out via computation of probability density functions on tangent space of each class tacking benefit from the geometric structure of the Grassmann manifold. The experimental evaluation is performed on three existing datasets containing various challenges, including MSR-action 3D, UT-kinect and MSR-Gesture3D. Results reveal that our approach outperforms the state-of-the-art methods, with accuracy of 98.21% on MSR-Gesture3D and 95.25% on UT-kinect, and achieves a competitive performance of 86.21% on MSR-action 3D.
Rim Slama, Hazem Wannous, Mohamed Daoudi
ICPR3
2014 3D human motion analysis framework for shape similarity and retrieval
Rim Slama, Hazem Wannous, Mohamed Daoudi
Image Vis. Comput.3
2014 4-D Facial Expression Recognition by Learning Geometric Deformations
abstract
In this paper, we present an automatic approach for facial expression recognition from 3-D video sequences. In the proposed solution, the 3-D faces are represented by collections of radial curves and a Riemannian shape analysis is applied to effectively quantify the deformations induced by the facial expressions in a given subsequence of 3-D frames. This is obtained from the dense scalar field, which denotes the shooting directions of the geodesic paths constructed between pairs of corresponding radial curves of two faces. As the resulting dense scalar fields show a high dimensionality, Linear Discriminant Analysis (LDA) transformation is applied to the dense feature space. Two methods are then used for classification: 1) 3-D motion extraction with temporal Hidden Markov model (HMM) and 2) mean deformation capturing with random forest. While a dynamic HMM on the features is trained in the first approach, the second one computes mean deformations under a window and applies multiclass random forest. Both of the proposed classification schemes on the scalar fields showed comparable results and outperformed earlier studies on facial expression recognition from 3-D video sequences.
Boulbaba Ben Amor, Hassen Drira, Stefano Berretti, Mohamed Daoudi, Anuj Srivastava
IEEE Trans. Cybern.4
2013 3D Face Recognition under Expressions, Occlusions, and Pose Variations
abstract
We propose a novel geometric framework for analyzing 3D faces, with the specific goals of comparing, matching, and averaging their shapes. Here we represent facial surfaces by radial curves emanating from the nose tips and use elastic shape analysis of these curves to develop a Riemannian framework for analyzing shapes of full facial surfaces. This representation, along with the elastic Riemannian metric, seems natural for measuring facial deformations and is robust to challenges such as large facial expressions (especially those with open mouths), large pose variations, missing parts, and partial occlusions due to glasses, hair, and so on. This framework is shown to be promising from both--empirical and theoretical--perspectives. In terms of the empirical evaluation, our results match or improve upon the state-of-the-art methods on three prominent databases: FRGCv2, GavabDB, and Bosphorus, each posing a different type of challenge. From a theoretical perspective, this framework allows for formal statistical inferences, such as the estimation of missing facial parts using PCA on tangent spaces and computing average shapes.
Hassen Drira, Boulbaba Ben Amor, Anuj Srivastava, Mohamed Daoudi, Rim Slama
IEEE Trans. Pattern Anal. Mach. Intell.4
2013 A comparison of methods for non-rigid 3D shape retrieval
Zhouhui Lian, Afzal Godil, Benjamin Bustos, Mohamed Daoudi, Jeroen Hermans, Shun Kawamura, Yukinori Kurita, Guillaume Lavoué, Hien Van Nguyen, Ryutarou Ohbuchi, Yuki Ohkita, Yuya Ohishi, Fatih Porikli, Martin Reuter 0001, Ivan Sipiran, Dirk Smeets, Paul Suetens, Hedi Tabia, Dirk Vandermeulen
Pattern Recognit.4
2013 A parts-based approach for automatic 3D shape categorization using belief functions
abstract
Grouping 3D objects into (semantically) meaningful categories is a challenging and important problem in 3D mining and shape processing. Here, we present a novel approach to categorize 3D objects. The method described in this article, is a belief-function-based approach and consists of two stages: the training stage, where 3D objects in the same category are processed and a set of representative parts is constructed, and the labeling stage, where unknown objects are categorized. The experimental results obtained on the Tosca-Sumner and the Shrec07 datasets show that the system efficiently performs in categorizing 3D models.
Hedi Tabia, Mohamed Daoudi, Jean-Philippe Vandeborre, Olivier Colot
ACM Trans. Intell. Syst. Technol.2
2012 3D dynamic expression recognition based on a novel Deformation Vector Field and Random Forest
Hassen Drira, Boulbaba Ben Amor, Mohamed Daoudi, Anuj Srivastava, Stefano Berretti
ICPR3
2012 Indexed heat curves for 3D-model retrieval
Rachid El Khoury, Jean-Philippe Vandeborre, Mohamed Daoudi
ICPR3
2012 Place Recognition via 3D Modeling for Personal Activity Lifelog Using Wearable Camera
Hazem Wannous, Vladislavs Dovgalecs, Rémi Mégret, Mohamed Daoudi
MMM4
2012 Boosting 3-D-Geometric Features for Efficient Face Recognition and Gender Classification
abstract
We utilize ideas from two growing but disparate ideas in computer vision-shape analysis using tools from differential geometry and feature selection using machine learning-to select and highlight salient geometrical facial features that contribute most in 3-D face recognition and gender classification. First, a large set of geometries curve features are extracted using level sets (circular curves) and streamlines (radial curves) of the Euclidean distance functions of the facial surface; together they approximate facial surfaces with arbitrarily high accuracy. Then, we use the well-known Adaboost algorithm for feature selection from this large set and derive a composite classifier that achieves high performance with a minimal set of features. This greatly reduced set, consisting of some level curves on the nose and some radial curves in the forehead and cheeks regions, provides a very compact signature of a 3-D face and a fast classification algorithm for face recognition and gender selection. It is also efficient in terms of data storage and transmission costs. Experimental results, carried out using the FRGCv2 dataset, yield a rank-1 face recognition rate of 98% and a gender classification rate of 86% rate.
Lahoucine Ballihi, Boulbaba Ben Amor, Mohamed Daoudi, Anuj Srivastava, Driss Aboutajdine
IEEE Trans. Inf. Forensics Secur.3
2011 Non-rigid 3D shape classification using bag-of-feature techniques
abstract
In this paper, we present a new method for 3D-shape categorization using Bag-of-Feature techniques (BoF). This method is based on vector quantization of invariant descriptors of 3D-object patches. We analyze the performance of two wellknown classifiers: the Naïve Bayes and the SVM. The results show the effectiveness of our approach and prove that the method is robust to non-rigid and deformable shapes, in which the class of transformations may be very wide due to the capability of such shapes to bend and assume different forms.
Hedi Tabia, Olivier Colot, Mohamed Daoudi, Jean-Philippe Vandeborre
ICME3
2011 Joint ACM workshop on human gesture and behavior understanding: (J-HGBU'11)
abstract
The ability to understand social signals of a person we are communicating with is the core of social intelligence. Social Intelligence is a facet of human intelligence that has been argued to be indispensable and perhaps the most important for success in life. At the same time, human-centric multimedia applications for humans and about humans are becoming increasingly important. 3D modeled human-objects, like bodies, heads and faces are exploited for animation, security, and human computer interaction, while three dimensional motion of arms, legs and local body features is used for more complete human gesture, activity and behavior analysis. The Joint Human Gesture and Behavior Understanding (J-HGBU) workshop event consists of two parts focusing on these complementary challenges: the Workshop on Multimedia Access to 3D Human Objects (MA3HO'11) and the Workshop on Social Signal Processing (SSPW'11).
Maja Pantic, Alex Pentland, Alessandro Vinciarelli, Rita Cucchiara, Mohamed Daoudi, Alberto Del Bimbo
ACM Multimedia5
2011 Learning Boundary Edges for 3D-Mesh Segmentation
abstract
Abstract This paper presents a 3D‐mesh segmentation algorithm based on a learning approach. A large database of manually segmented 3D‐meshes is used to learn a boundary edge function. The function is learned using a classifier which automatically selects from a pool of geometric features the most relevant ones to detect candidate boundary edges. We propose a processing pipeline that produces smooth closed boundaries using this edge function. This pipeline successively selects a set of candidate boundary contours, closes them and optimizes them using a snake movement. Our algorithm was evaluated quantitatively using two different segmentation benchmarks and was shown to outperform most recent algorithms from the state‐of‐the‐art.
Halim Benhabiles, Guillaume Lavoué, Jean-Philippe Vandeborre, Mohamed Daoudi
Comput. Graph. Forum4
2011 Eurographics 2010 Workshop on 3D Object Retrieval (EG 3DOR'10) in cooperation with ACM SIGGRAPH
abstract
published
Mohamed Daoudi, Tobias Schreck
Comput. Graph. Forum1
2011 A New 3D-Matching Method of Nonrigid and Partially Similar Models Using Curve Analysis
abstract
The 3D-shape matching problem plays a crucial role in many applications, such as indexing or modeling, by example. Here, we present a novel approach to matching 3D objects in the presence of nonrigid transformation and partially similar models. In this paper, we use the representation of surfaces by 3D curves extracted around feature points. Indeed, surfaces are represented with a collection of closed curves, and tools from shape analysis of curves are applied to analyze and to compare curves. The belief functions are used to define a global distance between 3D objects. The experimental results obtained on the TOSCA and the SHREC07 data sets show that the system performs efficiently in retrieving similar 3D models.
Hedi Tabia, Mohamed Daoudi, Jean-Philippe Vandeborre, Olivier Colot
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Shape analysis of local facial patches for 3D facial expression recognition
Ahmed Maalej, Boulbaba Ben Amor, Mohamed Daoudi, Anuj Srivastava, Stefano Berretti
Pattern Recognit.3
2011 3D facial expression recognition using SIFT descriptors of automatically detected keypoints
Stefano Berretti, Boulbaba Ben Amor, Mohamed Daoudi, Alberto Del Bimbo
Vis. Comput.3
2010 Pose and Expression-Invariant 3D Face Recognition using Elastic Radial Curves
abstract
In this paper we explore the use of shapes of elastic radial curves to model 3D facial deformations, caused by changes in facial expressions. We represent facial surfaces by indexed collections of radial curves on them, emanating from the nose tips, and compare the facial shapes by comparing the shapes of their corresponding curves. Using a past approach on elastic shape analysis of curves, we obtain an algorithm for comparing facial surfaces. We also introduce a quality control module which allows our approach to be robust to pose variation and missing data. Comparative evaluation using a common experimental setup on GAVAB dataset, considered as the most expression-rich and noise-prone 3D face dataset, shows that our approach outperforms other state-of-the-art approaches.
Hassen Drira, Boulbaba Ben Amor, Mohamed Daoudi, Anuj Srivastava
BMVC3
2010 A Set of Selected SIFT Features for 3D Facial Expression Recognition
abstract
In this paper, the problem of person-independent facial expression recognition is addressed on 3D shapes. To this end, an original approach is proposed that computes SIFT descriptors on a set of facial landmarks of depth images, and then selects the subset of most relevant features. Using SVM classification of the selected features, an average recognition rate of 77.5% on the BU-3DFE database has been obtained. Comparative evaluation on a common experimental setup, shows that our solution is able to obtain state of the art results.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala, Boulbaba Ben Amor, Mohamed Daoudi
ICPR5
2010 Local 3D Shape Analysis for Facial Expression Recognition
abstract
We investigate the problem of facial expression recognition using 3D face data. Our approach is based on local shape analysis of several relevant regions of a given face scan. These regions or patches from facial surfaces are extracted and represented by sets of closed curves. A Riemannian framework is used to derive the shape analysis of the extracted patches. The applied framework permits to calculate a similarity (or dissimilarity) distances between patches, and to compute the optimal deformation between them. Once calculated, these measures are employed as inputs to a commonly used classification techniques such as AdaBoost and Support Vector Machines (SVM). A quantitative evaluation of our novel approach is conducted on a subset of the publicly available BU-3DFE database.
Ahmed Maalej, Boulbaba Ben Amor, Mohamed Daoudi, Anuj Srivastava, Stefano Berretti
ICPR3
2010 3D-Shape Retrieval Using Curves and HMM
abstract
In this paper, we propose a new approach for 3D-shape matching. This approach encloses an off-line step and an on-line step. In the off-line one, an alphabet, of which any shape can be composed, is constructed. First, 3D-objects are subdivided into a set of 3D-parts. The subdivision consists to extract from each object a set of feature points with associated curves. Then the whole set of 3D-parts is clustered into different classes from a semantic point of view. After that, each class is modeled by a Hidden Markov Model (HMM). The HMM, which represents a character in the alphabet, is trained using the set of curves corresponding to the class parts. Hence, any 3D-object can be represented by a set of characters. The on-line step consists to compare the set of characters representing the 3D-object query and that of each object in the given dataset. The experimental results obtained on the TOSCA dataset show that the system efficiently performs in retrieving similar 3D-models.
Hedi Tabia, Olivier Colot, Mohamed Daoudi, Jean-Philippe Vandeborre
ICPR3
2010 ACM workshop on 3d object retrieval: 3DOR'10 chair's welcome
abstract
3D media has emerged rapidly as a new type of content within the multimedia domain. The recent acceleration of 3D content production, witnessed across all fields up to user-generated content, is causing a huge amount of traffic and data stored and transmitted using Internet technologies. Recent advances in 3D acquisition and 3D graphics rendering technologies boosted the creation of 3D model archives for several application domains. These include archaeology and cultural heritage, computer-assisted design (CAD), medicine and bioinformatics, 3D face recognition and security, entertainment and serious gaming, spatial data and 3D city management.
Mohamed Daoudi, Michela Spagnuolo, Remco C. Veltkamp
ACM Multimedia1
2010 A subjective experiment for 3D-mesh segmentation evaluation
abstract
In this paper we present a subjective quality assessment experiment for 3D-mesh segmentation. For this end, we carefully designed a protocol with respect to several factors namely the rendering conditions, the possible interactions, the rating range, and the number of human subjects. To carry out the subjective experiment, more than 40 human observers have rated a set of 250 segmentation results issued from various algorithms. The obtained Mean Opinion Scores, which represent the human subjects' point of view toward the quality of each segmentation, have then been used to evaluate both the quality of automatic segmentation algorithms and the quality of similarity metrics used in recent mesh segmentation benchmarking systems.
Halim Benhabiles, Guillaume Lavoué, Jean-Philippe Vandeborre, Mohamed Daoudi
MMSP4
2010 A 3-D Search engine based on Fourier series
Elmustapha Ait Lmaati, Ahmed El Oirrak, Driss Aboutajdine, Mohamed Daoudi, Mohammed Najib Kaddioui
Comput. Vis. Image Underst.4
2010 A comparative study of existing metrics for 3D-mesh segmentation evaluation
Halim Benhabiles, Jean-Philippe Vandeborre, Guillaume Lavoué, Mohamed Daoudi
Vis. Comput.4
2009 A Riemannian analysis of 3D nose shapes for partial human biometrics
abstract
In this paper we explore the use of shapes of noses for performing partial human biometrics. The basic idea is to represent nasal surfaces using indexed collections of iso-curves, and to analyze shapes of noses by comparing their corresponding curves. We extend past work in Riemannian analysis of shapes of closed curves in R3to obtain a similar Riemannian analysis for nasal surfaces. In particular, we obtain algorithms for computing geodesics, computing statistical means, and stochastic clustering. We demonstrate these ideas in two application contexts : authentication and identification. We evaluate performances on a large database involving 2000 scans from FRGC v2 database, and present a hierarchical organization of nose databases to allow for efficient searches.
Hassen Drira, Boulbaba Ben Amor, Anuj Srivastava, Mohamed Daoudi
ICCV4
2009 Skin and non-skin probability approximation based on discriminative tree distribution
abstract
We investigate the probability tree models to approximate skin and non-skin distributions. These models have presented good results in solving the skin detection problem. However, there are two main disadvantages of the existing skin/non-skin tree distributions based models: (1) the structure of some tree distributions is predefined; and (2) the inter and the intra classes of skin/non-skin are not taken into account at the same time by the existing skin and/or non-skin tree models. To overcome these drawbacks, we propose a new classifier based on an image patch joint distribution approximation modelled by a discriminative skin/non-skin tree. On the Compaq database, we examine the performances of the proposed approach compared with the baseline model and two others based on dependency tree's distributions. Experimental results show that the new approach is a significant improvement over the others.
Sanaa El Fkihi, Mohamed Daoudi, Driss Aboutajdine
ICIP2
2009 Fast and efficient 3D face recognition using wavelet networks
abstract
3D shape of face has recently emerged as a major research in face biometrics. However, while it is reputed to be relatively invariant to lighting conditions and pose, one still needs to cope with facial expression variations for a reliable face recognition solution and running time of the matching algorithms for fast identification software. We present in this paper our solutions to overcome these limitations. We propose a new method of 3D facial recognition based on wavelet networks. Firstly, depth image is preprocessed in order to crop the useful area of the face image. Secondly, a compact and representative biometric signature is produced by means of wavelet networks. Finally, the matching of two faces is made by computing Euclidean distance between their two corresponding signatures. To show the efficiency and accuracy of our approach, a subset taken from FRGC v2 dataset is used to made evaluations.
Salwa Said, Boulbaba Ben Amor, Mourad Zaied, Chokri Ben Amar, Mohamed Daoudi
ICIP5
2009 A framework for the objective evaluation of segmentation algorithms using a ground-truth of human segmented 3D-models
abstract
In this paper, we present an evaluation method of 3D-mesh segmentation algorithms based on a ground-truth corpus. This corpus is composed of a set of 3D-models grouped in different classes (animals, furnitures, etc.) associated with several manual segmentations produced by human observers. We define a measure that quantifies the consistency between two segmentations of a 3D-model, whatever their granularity. Finally, we propose an objective quality score for the automatic evaluation of 3D-mesh segmentation algorithms based on these measures and on the ground-truth corpus. Thus the quality of segmentations obtained by automatic algorithms is evaluated in a quantitative way thanks to the quality score, and on an objective basis thanks to the groundtruth corpus. Our approach is illustrated through the evaluation of two recent 3D-mesh segmentation methods.
Halim Benhabiles, Jean-Philippe Vandeborre, Guillaume Lavoué, Mohamed Daoudi
Shape Modeling International4
2009 Partial 3D Shape Retrieval by Reeb Pattern Unfolding
abstract
Abstract This paper presents a novel approach for fast and efficient partial shape retrieval on a collection of 3D shapes. Each shape is represented by a Reeb graph associated with geometrical signatures. Partial similarity between two shapes is evaluated by computing a variant of their maximum common sub‐graph. By investigating Reeb graph theory, we take advantage of its intrinsic properties at two levels. First, we show that the segmentation of a shape by a Reeb graph provides charts with disk or annulus topology only. This topology control enables the computation of concise and efficient sub‐part geometrical signatures based on parameterisation techniques. Secondly, we introduce the notion of Reeb pattern on a Reeb graph along with its structural signature. We show this information discards Reeb graph structural distortion and still depicts the topology of the related sub‐parts. The number of combinations to evaluate in the matching process is then dramatically reduced by only considering the combinations of topology equivalent Reeb patterns. The proposed framework is invariant against rigid transformations and robust against non‐rigid transformations and surface noise. It queries the collection in interactive time (from 4 to 30 seconds for the largest queries). It outperforms the competing methods of the SHREC 2007 contest in term of NDCG vector and provides, respectively, a gain of 14.1% and 40.9% on the approaches by Biasotti et al.[ BMSF06 ]and Cornea et al.[ CDS*05 ]. As an application, we present an intelligent modelling‐by‐example system which enables a novice user to rapidly create new 3D shapes by composing shapes of a collection having similar sub‐parts.
Julien Tierny, Jean-Philippe Vandeborre, Mohamed Daoudi
Comput. Graph. Forum3
2009 An Intrinsic Framework for Analysis of Facial Surfaces
Chafik Samir, Anuj Srivastava, Mohamed Daoudi, Eric Klassen
Int. J. Comput. Vis.3
2008 Three-dimensional face recognition using elastic deformations of facial surfaces
abstract
We propose a pattern theoretic approach for studying variability in shapes of facial surfaces. Our idea is to impose a specific, yet natural, coordinate system, called a curvilinear coordinate system, on facial surfaces. In this system, one coordinate xi1measures the distance of a point from the tip of the nose and its level curves are called the facial curves. The other coordinate xi2measures distances along these curves; level curves of this coordinate are orthogonal to the facial curves. To compare two facial surfaces we use elastic deformations that use stretching, shrinking, and bending to optimally register points across two surfaces. We will demonstrate this idea on Florida State University (FSU) 3D face database.
Mohamed Daoudi, Lahoucine Ballihi, Chafik Samir, Anuj Srivastava
ICME1
2008 Fast and precise kinematic skeleton extraction of 3D dynamic meshes
abstract
Shape skeleton extraction is a fundamental pre-processing task in shape-based pattern recognition. This paper presents a new algorithm for fast and precise extraction of kinematic skeletons of 3D dynamic surface meshes. Unlike previous approaches, surface motions are characterized by the mesh edge-length deviation induced by its transformation through time. Then a static skeleton extraction algorithm based on Reeb graphs exploits this latter information to extract the kinematic skeleton. This hybrid static and dynamic shape analysis enables the precise detection of objects¿ articulations as well as shape topological transitions corresponding to possibly-articulated immobile objects¿ features. Experiments show that the proposed algorithm is faster than previous techniques and still achieves better accuracy.
Julien Tierny, Jean-Philippe Vandeborre, Mohamed Daoudi
ICPR3
2008 SHape REtrieval contest 2008: 3D face scans
abstract
Three-Dimensional face recognition is a challenging task with a large number of proposed solutions [1, 2]. With variations in pose and expression the identification of a face scan based on 3D geome-try is difficult. To improve on this task and to evaluate existing face
Frank Bart ter Haar, Mohamed Daoudi, Remco C. Veltkamp
Shape Modeling International2
2008 The mixture of K-Optimal-Spanning-Trees based probability approximation: Application to skin detection
Sanaa El Fkihi, Mohamed Daoudi, Driss Aboutajdine
Image Vis. Comput.2
2008 Enhancing 3D mesh topological skeletons with discrete contour constrictions
Julien Tierny, Jean-Philippe Vandeborre, Mohamed Daoudi
Vis. Comput.3
2007 Topology driven 3D mesh hierarchical segmentation
abstract
In this paper, we propose to address the semantic- oriented 3D mesh hierarchical segmentation problem, using enhanced topological skeletons. This high level information drives both the feature boundary computation as well as the feature hierarchy definition. Proposed hierarchical scheme is based on the key idea that the topology of a feature is a more important decomposition criterion than its geometry. First, the enhanced topological skeleton of the input triangulated surface is constructed. Then it is used to delimit the core of the object and to identify junction areas. This second step results in a fine segmentation of the object. Finally, a fine to coarse strategy enables a semantic- oriented hierarchical composition of features, subdividing human limbs into arms and hands for example. Method performance is evaluated according to seven criteria enumerated in latest segmentation surveys [3]. Thanks to the high level description it uses as an input, presented approach results, with low computation times, in robust and meaningful compatible hierarchical decompositions.
Julien Tierny, Jean-Philippe Vandeborre, Mohamed Daoudi
Shape Modeling International3
2007 A probabilistic approach for 3D shape retrieval by characteristic views
Saïd Mahmoudi, Mohamed Daoudi
Pattern Recognit. Lett.2
2007 A Bayesian 3-D Search Engine Using Adaptive Views Clustering
abstract
In this paper, we propose a method for three-dimensional (3D)-model indexing based on two-dimensional (2D) views, which we call adaptive views clustering (AVC). The goal of this method is to provide an "optimal" selection of 2D views from a 3D model, and a probabilistic Bayesian method for 3D-model retrieval from these views. The characteristic view selection algorithm is based on an adaptive clustering algorithm and uses statistical model distribution scores to select the optimal number of views. Starting from the fact that all views do not have equal importance, we also introduce a novel Bayesian approach to improve the retrieval. Finally, we present our results and compare our method to some state-of-the-art 3D retrieval descriptors on the Princeton 3D Shape Benchmark database and a 3D-CAD-models database supplied by the car manufacturer Renault
Tarik Filali Ansary, Mohamed Daoudi, Jean-Philippe Vandeborre
IEEE Trans. Multim.2
2006 Probability Approximation Using Best-Tree Distribution for Skin Detection
Sanaa El Fkihi, Mohamed Daoudi, Driss Aboutajdine
ACIVS2
2006 3D Face Recognition Using Shapes of Facial Curves
abstract
Recognition of human beings using shapes of their full facial surfaces is a difficult problem. Our approach is to approximate a facial surface using a collection of (closed) facial curves, and to compare surfaces by comparing their corresponding curves. The differences between shapes of curves are quantified using lengths of geodesic paths between them on a pre-defined curve shape space. The metric for comparing facial surfaces is a composition of the metric involving individual facial curves. These ideas are demonstrated in the context of face recognition using the nearest-neighbor classifier
Chafik Samir, Anuj Srivastava, Mohamed Daoudi
ICASSP (5)3
2006 Three-Dimensional Face Recognition Using Shapes of Facial Curves
abstract
We study shapes of facial surfaces for the purpose of face recognition. The main idea is to 1) represent surfaces by unions of level curves, called facial curves, of the depth function and 2) compare shapes of surfaces implicitly using shapes of facial curves. The latter is performed using a differential geometric approach that computes geodesic lengths between closed curves on a shape manifold. These ideas are demonstrated using a nearest-neighbor classifier on two 3D face databases: Florida State University and Notre Dame, highlighting a good recognition performance.
Chafik Samir, Anuj Srivastava, Mohamed Daoudi
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Automatic 3D Face Recognition Using Topological Techniques
abstract
In this paper, we use the three-dimensional topological shape information for human face identification. We propose a new method to represent 3D faces as a topological graph. Fine registration of surfaces is done by first automatically finding topological connected components, and then constructing its topological graph representing the important topological changes on the face. The similarity calculation between 3D faces is processed using coarse-to-fine strategy while preserving the consistency of the graph structures, which result in establishing a correspondence between the parts of faces. The experiments made with a 144 3D faces dataset show the efficiency of our approach
Chafik Samir, Jean-Philippe Vandeborre, Mohamed Daoudi
ICME3
2005 Skin detection using pairwise models
Bruno Jedynak, Huicheng Zheng, Mohamed Daoudi
Image Vis. Comput.3
2004 From Maximum Entropy to Belief Propagation: An application to Skin Detection
abstract
We build a maximum entropy model for skin detection. This model imposes constraints on various marginal distributions. Parameter estimation as well as optimization cannot be tackled without approximations. We propose to use a tree approximation of the pixel lattice. Parameter estimation is then reduced to the estimations of color histograms for neighbor pixels. Moreover, the belief propagation algorithm permits to obtain fast solution for skin probability at pixel locations. We assess the performance on the Compaq database. 1
Huicheng Zheng, Mohamed Daoudi, Bruno Jedynak
BMVC2
2004 Blocking objectionable images: adult images and harmful symbols
abstract
This paper describes a practical objectionable image filtering system, aimed at children's safer Web access. It includes two image filters: adult image filter and harmful symbol filter. In the adult image filter, we adopt a statistical model for skin detection and a neural network for adult image classification. The performance of the skin detection of our model outperforms that of the baseline model. Its elapsed time is about 0.18 second per image, which compares very well against previous systems. In the harmful symbol filter, we present an edge based Zernike moments method, which can capture the shape feature of a symbol object effectively. Its elapsed time is about 0.13 second per image. Experimental results on a large image database show that both of our filters can give promising performances
Huicheng Zheng, Mohamed Daoudi
ICME3
2003 Affine invariant descriptors for color images using Fourier series
Ahmed El Oirrak, Mohamed Daoudi, Driss Aboutajdine
Pattern Recognit. Lett.2
2002 Estimation of general 2D affine motion using Fourier descriptors
Ahmed El Oirrak, Mohamed Daoudi, Driss Aboutajdine
Pattern Recognit.2
2002 Affine invariant descriptors using Fourier series
Ahmed El Oirrak, Mohamed Daoudi, Driss Aboutajdine
Pattern Recognit. Lett.2
2001 Image indexing & retrieval using intermediate features
abstract
Visual information retrieval systems use low-level such as color, texture and shape for queries. Users usually have a more abstract notion of what will satisfy them. Using low-level to correspond to high-level abstractions is one aspect of the gap.In this paper, we introduce intermediate features. These are low-level semantic features and level image features. That is, in one hand, they can be arranged to produce high level concept and in another hand, they can be learned from a small annotated database. These can then be used in an retrieval system.We report experiments where intermediate are textures. These are learned from a small annotated database. The resulting indexing procedure is then demonstrated to be superior to a standard color histrogram indexing.
Mohamad Obeid, Bruno Jedynak, Mohamed Daoudi
ACM Multimedia3
2000 New Multiscale Planar Shape Invariant Representation under a General Affine Transformations
abstract
We introduce a set of invariant and local descriptors, which are independent under a general affine group transformation (GA(2): rotation, uniform scaling, translation, and stretching) of planar curves: a new invariant in affine scale-space (IASS). For their extraction, we propose a method based on multiple convolutions between the affine-parametrized curve and Gaussian kernel. The IASS representation of planar curves generalizes, in the affine case, a curvature scale space invariant description under similarity transformation (rotation, uniform scaling and translation).
Mohamed Daoudi, Stanislaw Matusiak
ICPR1
1999 Shape distances for contour tracking and motion estimation
Mohamed Daoudi, Faouzi Ghorbel, A. Mokadem, Olivier Avaro, Henri Sanson
Pattern Recognit.1
1998 Multiple blocs classification for fast encoding in fractal-based images compression
abstract
Increasing the search speed for matching range and domain blocs is the main challenge facing fractal-based images compression. One way to remedy at this problem is to classify image blocs into categories and only search among domain blocs which are in the same category as the target range bloc. Since image blocs with a simple edge are a very important portions of the perceptual information content in image, we propose a method to both identify and classify this kind of blocs according to their edge presentation. We refer to this method as forced classification. This method is combined with other suitable methods of blocs classification available in the literature to allow a fast and encoding of grey-scale images. The result obtained is good, the encoding time for a 512/spl times/512 image is reduced by a factor of 37.52% than using the Fisher classification only, while the loss of image quality is low.
Khalil Maalmi, Rachid Benslimane, Mohamed Daoudi
SMC3
1998 Planar closed contour representation by invariant under a general affine transformation
abstract
This paper presents a new multiscale and zero-crossing, curvature-based shape representation technique for planar curves with general affine transformation. The method consists of the concept of describing a curve at varying levels of detail using features that are invariant with respect to transformations that do not change the shape of the curve. The process of describing a curve at increasing levels of abstraction is referred to as the affine evolution of that curve. This affine evolution does not change the physical interpretation of planar curves and characterize the behaviors of inflexion points of the curves during its evolution.
Stanislaw Matusiak, Mohamed Daoudi, Faouzi Ghorbel
SMC2
1996 Global planar rigid motion estimation applied to object-oriented coding
abstract
The aim of this paper is to present two kinds of shape distances in a dynamic images context. The first distance is obtained by the complete and stable set of invariants under rigid motion for closed curves. The second distance is obtained by using a Hausdorff distance for parameter estimation. An original application serves as a useful test for evaluating these proposed distances in coding applications.
Faouzi Ghorbel, Mohamed Daoudi, A. Mokadem, Olivier Avaro, Henri Sanson
ICPR2
1996 A shape distance by complete and stable invariant descriptors for contour tracking
abstract
We consider the problem of comparing geometric objects in order to determine the extent to which one object resembles another. Invariant feature families are presented. A complete and stable set of invariant features has been applied to define all invariant distance in the shapes space. This distance allows us to detect and follow moving objects in a dynamic scene. In order to evaluate the performance of such a metric, experimental results are given.
A. Mokadem, Mohamed Daoudi, Faouzi Ghorbel
ICPR2