VLDB 2026 Research / reviewers in the wild / expert
Hazem Wannous
dblp:97/5331
· DBLP profile ↗
38ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0001-8475-4309ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 25 · 1 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GCD: Geometry-Constrained Contact-Aware Diffusion for Text-Driven 3D Hand-Object Motion Synthesis
Josue Adossehoun, Hazem Wannous |
ICPR (11) | 2 |
| 2026 | ICPR 2026 Competition on Privacy-Preserving Person Re-identification from Top-View RGB-Depth Camera (TVRID)
Raphaël Delécluse, Hazem Wannous, Laurent Guimas |
ICPR (16) | 2 |
| 2026 | Three-Step Hierarchical Transformer for Multi-pedestrian Trajectory Prediction
Raphaël Delécluse, Hazem Wannous, Laurent Grisoni, Laurent Guimas |
ICPR (3) | 2 |
| 2025 | AG-MAE: Anatomically Guided Spatio-Temporal Masked Auto-Encoder for Online Hand Gesture RecognitionabstractHand gesture recognition plays a crucial role in the domain of computer vision, as it enhances human-computer interaction by enabling intuitive, touch-free control and communication. While offline methods have made significant advances in isolated gesture recognition, real-world applications demand online and continuous processing. Skeleton-based methods, though effective, face challenges due to the intricate nature of hand joints and the diverse 3D motions they induce. This paper introduces AG-MAE, a novel approach that integrates anatomical constraints to guide the self-supervised training of a spatio-temporal masked autoencoder, enhancing the learning of 3D keypoint representations. By incorporating anatomical knowledge, AG-MAE learns more discriminative features for hand poses and movements, subsequently improving online gesture recognition. Evaluation on standard datasets demonstrates the superiority of our approach and its potential for real-world applications. Code is available at: https://github.com/o-ikne/AG-MAE.git. Omar Ikne, Benjamin Allaert, Hazem Wannous |
3DV | 3 |
| 2025 | SMSCI: Simultaneous Modeling of Social and Contextual Interactions for Multi Pedestrian Trajectory Prediction
Mayssa Zaier, Hazem Wannous, Hassen Drira |
CAIP (1) | 2 |
| 2025 | Privacy-Preserving Person Re-Identification from Temporal Sequences with Transformer and Hungarian OptimizationabstractPerson re-identification (Re-ID) is a crucial task in surveillance and human behavior analysis, often used in public spaces such as transport hubs. Traditional RGB-based Re-ID methods raise privacy concerns and are highly sensitive to lighting variations and occlusion. In this paper, we propose a novel Re-ID approach that leverages depth images, which inherently obscures facial and other identifiable features, making it a privacy-preserving solution. Our method addresses the association problem between multiple views of individuals by applying the Hungarian algorithm, optimizing the matching process through minimization of the global cost across the distance matrix. We further enhance the approach by introducing temporal sequences of frames as input to a Transformer encoder architecture, which exploits both RGB and depth modalities. This architecture captures dynamic movement patterns, improving feature extraction and re-identification accuracy. Additionally, we employ batch hard triplet loss to enhance discriminative feature learning by focusing on the hardest samples. We evaluate both depth-only and RGB-D models on several top-view datasets, including TVPR2, GODPR, and BIWI RGBD-ID. Our results demonstrate that depth-only reidentification can achieve competitive performance compared to state-of-the-art methods, as measured by standard metrics such as Cumulative Matching Characteristics (CMC) and Mean Average Precision (mAP), while prioritizing privacy preservation. Code is available at: https://github.com/RaphaelDel/PrivacyPreserving-ReID.git Raphaël Delécluse, Hazem Wannous, Laurent Guimas |
FG | 2 |
| 2025 | Geometry-Aware Deep Learning for 3D Skeleton-Based Motion PredictionabstractInternational audience Mayssa Zaier, Hazem Wannous, Hassen Drira |
WACV | 2 |
| 2025 | GNF: Gaussian Neural Fields for Multidimensional Signal Representation and ReconstructionabstractAbstract Neural fields have emerged as a powerful framework for representing continuous multidimensional signals such as images and videos, 3D and 4D objects and scenes, and radiance fields. While efficient, achieving high‐quality representation requires the use of wide and deep neural networks. These, however, are slow to train and evaluate. Although several acceleration techniques have been proposed, they either trade memory for faster training and/or inference, rely on thousands of fitted primitives with considerable optimization time, or compromise the smooth, continuous nature of neural fields. In this paper, we introduce Gaussian Neural Fields (GNF), a novel compact neural decoder that maps learned feature grids into continuous non‐linear signals, such as RGB images, Signed Distance Functions (SDFs), and radiance fields, using a single compact layer of Gaussian kernels defined in a high‐dimensional feature space. Our key observation is that neurons in traditional MLPs perform simple computations, usually a dot product followed by an activation function, necessitating wide and deep MLPs or high‐resolution feature grids to model complex functions. In this paper, we show that replacing MLP‐based decoders with Gaussian kernels whose centers are learned features yields highly accurate representations of 2D (RGB), 3D (geometry), and 5D (radiance fields) signals with just a single layer of such kernels. This representation is highly parallelizable, operates on low‐resolution grids, and trains in under 15 seconds for 3D geometry and under 11 minutes for view synthesis. GNF matches the accuracy of deep MLP‐based decoders with far fewer parameters and significantly higher inference throughput. The source code is publicly available at https://grbfnet.github.io/ . Abelaziz Bouzidi, Hamid Laga, Hazem Wannous, Ferdous Sohel |
Comput. Graph. Forum | 3 |
| 2025 | eMotion-GAN: A motion-based GAN for photorealistic and facial expression preserving frontal view synthesisabstractFacial expression recognition (FER) systems frequently suffer significant performance degradation when confronted with head pose variations, a pervasive challenge in real-world applications ranging from healthcare monitoring to human–computer interaction. While existing frontal view synthesis (FVS) methods attempt to address this issue, they predominantly operate in the appearance domain, often introducing artifacts that distort the subtle motion patterns crucial for accurate expression analysis. We present eMotion-GAN, a two-stage generative motion-domain framework that fundamentally rethinks frontalization by decomposing facial dynamics into two distinct components: (1) expression-related motion stemming from muscle activity, and (2) pose-related motion acting as noise. We conducted extensive evaluations using several widely recognized dynamic FER datasets, which encompass sequences exhibiting various degrees of head pose variations in both intensity and orientation. Our results demonstrate the effectiveness of our approach in significantly reducing the FER performance gap between frontal and non-frontal faces. Specifically, we achieved a FER improvement of up to +5% for small pose variations and up to +20% improvement for larger pose variations. Code and pre-trained models are available at: https://github.com/o-ikne/eMotion-GAN.git . • Treats head pose as structured noise in optical flow for robust frontalization. • Needs no landmarks; not affected by inaccurate facial landmark detection. • Splits facial motion into pose and expression; gains 20% FER accuracy for poses. • Enables expression transfer for animation and facial data augmentation. • Reduces artifacts and outperforms appearance-based frontalization methods. Omar Ikne, Benjamin Allaert, Ioan Marius Bilasco, Hazem Wannous |
Comput. Vis. Image Underst. | 4 |
| 2025 | Pedestrian trajectory prediction: a literature review and current trends
Mayssa Zaier, Hazem Wannous, Hassen Drira, Jacques Boonaert |
Neural Comput. Appl. | 2 |
| 2024 | Skeleton-Based Self-Supervised Feature Extraction for Improved Dynamic Hand Gesture RecognitionabstractHuman-computer interaction (HCI) has become integral to modern life, especially in digital environments. However, challenges persist in utilizing hand gestures due to factors such as the dynamic nature of gestures and the intricacies of intra and inter-finger movements. In this paper, we propose an innovative approach to improve skeleton-based hand gesture recognition by integrating self-supervised learning, a promising technique for acquiring distinctive representations directly from unlabeled data. The proposed method takes advantage of prior knowledge of hand topology, combining topology-aware self-supervised learning with a customized skeleton-based architecture to derive meaningful representations from skeleton data under different hand poses. We introduce customized masking strategies for skeletal hand data and design a model architecture that incorporates spatial connectivity information, improving the model's understanding of the interrelationships between hand joints. The extensive experiments demonstrate the effectiveness of the approach, with state-of-the-art performance on benchmark datasets. An exploration of the generalization of learned representations across datasets and a study of the impact of fine-tuning with limited labeled data are conducted, highlighting the adaptability and robustness of our approach. Omar Ikne, Benjamin Allaert, Hazem Wannous |
FG | 3 |
| 2024 | Spatio-Temporal Sparse Graph Convolution Network for Hand Gesture RecognitionabstractUnlike whole-body action recognition, hand gestures involve spatially closely distributed joints, promoting stronger collaboration. This needs to be taken into account in order to capture complex spatial and temporal features. In response to these challenges, this paper presents a Spatio-Temporal Sparse Graph Convolution Network (ST-SGCN) for dynamic recognition of hand gestures. Based on decoupled spatio-temporal processing, the ST-SGCN incorporates Graph Convolutional Networks, attention mechanism and asymmetric convolutions to capture the nuanced movements of hand joints. The key novelty is the introduction of sparse spatio-temporal directed interactions, overcoming the limitations associated with dense, undirected methods. The sparse aspect models essential interactions between hand joints selectively, improving computational efficiency and interpretability. Directed in-teractions capture asymmetrical dependencies between hand joints, improving discernment of joint influences. Experimental evaluations on three benchmark datasets, including Briareo, SHREC'17 and IPN Hand, demonstrate ST-SGCN's state-of-the-art performance for dynamic hand gesture recognition. Omar Ikne, Rim Slama, Hichem Saoudi, Hazem Wannous |
FG | 4 |
| 2024 | Motion-Lie Transformer: Geometric Attention For 3D Human Pose Motion PredictionabstractSkeletal motion prediction aims to forecast future movement based on 3D skeleton sequences, crucial for applications such as autonomous driving and virtual reality. However, anticipating the motion of 3D articulated objects is challenging due to their inherent non linearity and stochastic nature. Existing approaches often represent the skeleton as a set of 3D joints, which unfortunately ignores joint relationships and anatomical constraints. Moreover, conventional recurrent neural networks struggle with capturing long-term dependencies in motion contexts. To address these limitations, we propose encoding anatomical constraints through Lie algebra representation, integrating self-attention in transformer networks. Our Motion-Lie Transformer architecture, leveraging Transformers with self-attention, preserves human motion kinematics. Empirical evaluations on datasets like Human3.6M, GTA-IM, and PROX promise competitive performance and accurate 3D human pose estimation. Mayssa Zaier, Hazem Wannous, Hassen Drira, Jacques Boonaert |
ICIP | 2 |
| 2024 | SHREC 2024: Recognition of dynamic hand motions molding clayabstractGesture recognition is a tool to enable novel interactions with different techniques and applications, like Mixed Reality and Virtual Reality environments. With all the recent advancements in gesture recognition from skeletal data, it is still unclear how well state-of-the-art techniques perform in a scenario using precise motions with two hands. This paper presents the results of the SHREC 2024 contest organized to evaluate methods for their recognition of highly similar hand motions using the skeletal spatial coordinate data of both hands. The task is the recognition of 7 motion classes given their spatial coordinates in a frame-by-frame motion. The skeletal data has been captured using a Vicon system and pre-processed into a coordinate system using Blender and Vicon Shogun Post. We created a small, novel dataset with a high variety of durations in frames. This paper shows the results of the contest, showing the techniques created by the 5 research groups on this challenging task and comparing them to our baseline method. Ben Veldhuijzen, Remco C. Veltkamp, Omar Ikne, Benjamin Allaert, Hazem Wannous, Marco Emporio, Andrea Giachetti 0001, Joseph J. LaViola Jr., He Ruiwen, Halim Benhabiles, Adnane Cabani, Anthony Fleury, Karim Hammoudi, Konstantinos Gavalas, Christoforos Vlachos, Athanasios Papanikolaou, Ioannis Romanelis, Vlassis Fotis, Gerasimos Arvanitis, Konstantinos Moustakas, Martin Hanik, Esfandiar Nava-Yazdani, Christoph von Tycowicz |
Comput. Graph. | 5 |
| 2023 | Cross-Modal Attention for Accurate Pedestrian Trajectory Prediction
Mayssa Zaier, Hazem Wannous, Hassen Drira, Jacques Boonaert |
BMVC | 2 |
| 2023 | STr-GCN: Dual Spatial Graph Convolutional Network and Transformer Graph Encoder for 3D Hand Gesture RecognitionabstractSkeleton-based hand gesture recognition is a challenging task that sparked a lot of attention in recent years, especially with the rise of Graph Neural Networks. In this paper, we propose a new deep learning architecture for hand gesture recognition using 3D hand skeleton data and we call STr-GCN. It decouples the spatial and temporal learning of the gesture by leveraging Graph Convolutional Networks (GCN) and Transformers. The key idea is to combine two powerful networks: a Spatial Graph Convolutional Network unit that understands intra-frame interactions to extract powerful features from different hand joints and a Transformer Graph Encoder which is based on a Temporal Self-Attention module to incorporate inter-frame correlations. We evaluate the performance of our method on three benchmarks: the SHREC'17 Track dataset, Briareo dataset and the First Person Hand Action dataset. The experiments show the efficiency of our approach, which achieves or outperforms the state of the art. The code to reproduce our results is available in this link. Rim Slama, Wael Rabah, Hazem Wannous |
FG | 3 |
| 2023 | Capsule Transformer Network for Dynamic Hand Gesture Recognition Using Multimodal DataabstractIn recent years, deep learning techniques have achieved remarkable success in video analysis and more especially in action and gesture recognition. Even though convolutional neural networks (CNNs) remain the most widely used models, they have difficulty in capturing the global contextual information involving spatial and temporal domains or intermodality due to the local feature learning mechanism. This paper introduces a Capsule Transformer Network, which composed of a frame capsule module for extracting hand features and a gesture transformer module for modeling the temporal features and recognizing the dynamic gesture. Spatial attention is ensured through the capsule module to enhance the spatial information of the hand image, while the transformer module guarantees temporal attention through gesture sequence. We propose to use multimodal data, including RGB, depth and IR data, which improves the accuracy of our approach as it better captures the 3D structure of the hand and can distinguish between similar hand gestures. Testing on two datasets, Briareo and SHREC17, the proposed approach outperforms or equals previous methods. Alexandre Lebas, Rim Slama, Hazem Wannous |
ICIP | 3 |
| 2022 | Continuous Hand Gesture Recognition using Deep Coarse and Fine Hand Features
Hazem Wannous, Jean-Philippe Vandeborre |
BMVC | 1 |
| 2022 | Learning Co-occurrence Features Across Spatial and Temporal Domains for Hand Gesture RecognitionabstractHand gesture is the most natural modality for human-machine interaction and its recognition can be considered one of the most complicated and interesting challenges for computer vision community. In recent years, there has been a noticeable advancement in the field of machine learning and computer vision. However, providing a hand gesture recognition system robust enough to work in real-time applications remains challenging. Dynamic hand gestures can be seen as variations in shape or movement during hand motion and often both together. To tackle these challenges, we propose a dynamic hand gesture recognition approach based on hand skeletal sequences. In particular, we introduce a simple but effective deep network architecture to deal with Spatio-temporal co-occurrence features computed on 3D coordinates of hand joints along the gesture sequence. Experimental results show that our approach outperforms state-of-the-art methods on two public datasets, First Person Hand Action and SHREC’2017, with an efficient time computational model compared to most existing approaches. Mohammad Rehan, Hazem Wannous, Jafar Alkheir, Kinda Aboukassem |
CBMI | 2 |
| 2022 | Spatio-Temporal Analysis of Transformer based Architecture for Attention Estimation from EEGabstractFor many years now, understanding the brain mechanism has been a great research subject in many different fields. Brain signal processing and especially electroencephalogram (EEG) has recently known a growing interest both in academia and industry. One of the main examples is the increasing number of Brain-Computer Interfaces (BCI) aiming to link brains and computers. In this paper, we present a novel framework allowing us to retrieve the attention state, i.e degree of attention given to a specific task, from EEG signals. While previous methods often consider the spatial relationship in EEG through electrodes and process them in recurrent or convolutional based architecture, we propose here to also exploit the time and frequency information with a transformer-based network that has already shown its supremacy in many machine-learning (ML) related studies, e.g. machine translation. In addition to this novel architecture, an extensive study on the feature extraction methods, frequential bands and temporal windows length has also been carried out. The proposed network has been trained and validated on two public datasets and achieves higher results compared to state-of-the-art models. As well as proposing better results, the framework could be used in real applications, e.g. Attention Deficit Hyperactivity Disorder (ADHD) symptoms or vigilance during a driving assessment. Victor Delvigne, Hazem Wannous, Jean-Philippe Vandeborre, Laurence Ris, Thierry Dutoit |
ICPR | 2 |
| 2022 | RailSet: A Unique Dataset for Railway Anomaly DetectionabstractUnderstanding the driving environment is one of the key factors in achieving an autonomous vehicle. In particular, the detection of anomalies in the traffic lane is a high priority scenario, as it directly involves vehicle's safety. Recent state of the art image processing techniques for anomaly detection are all based on deep learning of neural networks. These algorithms require a considerable amount of annotated data for training and test purposes. While many datasets exist in the field of autonomous road vehicles, such datasets are extremely rare in the railway domain. In this work, we present a new innovative dataset relevant for railway anomaly detection called RailSet. It consists of 6600 high-quality manually annotated images containing normal situations and 1100 images of railway defects such as hole anomaly and rails discontinuity. Due to the lack of anomaly samples in public images and difficulties to create anomalies in the railway environment, we generate artificially images of abnormal scenes, using a deep learning algorithm named StyleMapGAN. This dataset is created as a contribution to the development of autonomous trains able to perceive tracks damage in front of the train. The dataset is available at this link. Arij Zouaoui, Ankur Mahtani, Mohamed Amine Hadded, Sebastien Ambellouis, Jacques Boonaert, Hazem Wannous |
IPAS | 6 |
| 2022 | PhyDAA: Physiological Dataset Assessing AttentionabstractAttention Deficit Hyperactivity Disorder (ADHD) is the most prevalent neurodevelopmental disorder among children. It affects patients’ lives in many ways: inattention, difficulty with stimuli inhibition or motor function regulation. Different treatments exist today, but these can present side effects or are not effective for all subgroups. Neurofeedback (NF) is an innovative treatment consisting of brain activity display. NF training could consist of a virtual reality (VR) video-game in which the participant’s attention affects the game. Attention being assessed through physiological signals, one of the main steps is to design an estimator for the attention state. We present a novel framework able to record physiological signals in specific attention states and able to estimate the corresponding attention state. We propose a database composed of electroencephalography signals (EEG), and an eye-tracker labelled with a score representing the attention span for 32 healthy participants. Different features are extracted from the signals and machine learning (ML) algorithms are proposed. Our approach exhibits high accuracy for attention estimation, which corroborates a correlation between attention state and physiological signals (i.e. EEG, eye-tracking signals). The dataset has been made publicly available to promote research in the domain and we encourage other scientists to use their own approach for attention estimation. Victor Delvigne, Hazem Wannous, Thierry Dutoit, Laurence Ris, Jean-Philippe Vandeborre |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Eye-Gaze Estimation using a Deep Capsule-based Regression NetworkabstractEye-gaze information is used in a variety of user platforms, such as driver monitoring systems and head-mounted interfaces. In order to estimate human eye-gaze, many solutions have been proposed, using different devices and techniques. However, achieving such estimation using only cheap devices like RGB cameras would enable gaze interactions on mobile devices and therefore generalise this kind of interaction. It could also enable behavior studies based on gaze and made on every day devices. We propose in this paper a new method for eye-gaze estimation using a new deep learning architecture based on the Capsule Neural Network. Capsule Networks have shown great results so far on classification tasks, but only a few works use them for regression tasks.By taking advantage of the Capsule Network architecture and its ability to reconstruct images, we are able to recreate simplified eye images and then estimate human gaze from them. Experiments are performed on two representative datasets for the task of eye-gaze estimation. Encouraging results are obtained for both the estimation and the reconstruction. Vivien Bernard, Hazem Wannous, Jean-Philippe Vandeborre |
CBMI | 2 |
| 2021 | Survey on Style in 3D Human Body Motion: Taxonomy, Data, Recognition and Its ApplicationsabstractThe meaning of the wordstyledepends on its context. While actions have already been quite studied for a while,stylein human body motion is a growing topic of interest. In the context of animation,styleis crucial as it brings realism and expressiveness to the motion of a character. Even though it is undoubtedly a key element in motions, its definition and the use of the wordstylein itself, among research works, lack consensus. Achieving realistic motions is tedious. It requires either a large motion capture dataset or the considerable work of artist animators. The lack of consistentstyledata is thus a challenge. Stylistic motion generation is quite studied in order to overcome this issue. This paper focuses on the study ofstylein human body motion from 3D human body skeletal data. It establishes a taxonomy of definitions ofstyle, describes the data that have been used up until now, introduces key notions about motion capture data as well as machine learning, and presents approaches aboutstylerecognition, person identification through theirstyleand motionstylegeneration. Sarah Ribet, Hazem Wannous, Jean-Philippe Vandeborre |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | 2D Deep Video Capsule Network with Temporal Shift for Action RecognitionabstractAction recognition in continuous video streams is a growing field since the past few years. Deep learning techniques and in particular Convolutional Neural Networks (CNNs) achieved good results in this topic. However, intrinsic CNNs limitations begin to cap the results since 2D CNN cannot capture temporal information and 3D CNN are to much resource demanding for real-time applications. Capsule Network, evolution of CNN, already proves its interesting benefits on small and low informational datasets like MNIST but yet its true potential has not emerged. In this paper we tackle the action recognition problem by proposing a new architecture combining Temporal Shift module over deep Capsule Network. Temporal Shift module permits us to insert temporal information over 2D Capsule Network with a zero computational cost to conserve the lightness of 2D capsules and their ability to connect spatial features. Our proposed approach outperforms or brings near state-of-the-art results on color and depth information on public datasets like First Person Hand Action and DHG 14/28 with a number of parameters 10 to 40 times less than existing approaches. Théo Voillemin, Hazem Wannous, Jean-Philippe Vandeborre |
ICPR | 2 |
| 2019 | Heterogeneous hand gesture recognition using 3D dynamic skeletal data
Quentin De Smedt, Hazem Wannous, Jean-Philippe Vandeborre |
Comput. Vis. Image Underst. | 2 |
| 2018 | Action Recognition from 3D Skeleton Sequences using Deep Networks on Lie Group FeaturesabstractThis paper addresses the problem of human action recognition from sequences of 3D skeleton data. For this purpose, we combine a deep learning network with geometric features extracted from data lie on a non-Euclidean space, which have been recently shown to be very effective to capture the geometric structure of the human pose. In particular, our approach claims to incorporate the intrinsic nature of the data characterized by Lie Group into deep neural networks and to learn more adequate geometric features for 3D action recognition problem. First, geometric features are extracted from 3D joints of skeleton sequences using the Lie group representation. Then, the network model is built from stacked units of 1-dimensional CNN across the temporal domain. Finally, CNN-features are then used to train an LSTM layer to model dependencies in the temporal domain, and to perform the action recognition. The experimental evaluation is performed on three public datasets containing various challenges: UT-Kinect, Florence 3D-Action and MSR-Action 3D. Results reveal that our approach achieves most of the state-of-the-art performance. Manel Rhif, Hazem Wannous, Imed Riadh Farah |
ICPR | 2 |
| 2017 | Motion segment decomposition of RGB-D sequences for human behavior understanding
Maxime Devanne, Stefano Berretti, Pietro Pala, Hazem Wannous, Mohamed Daoudi, Alberto Del Bimbo |
Pattern Recognit. | 4 |
| 2016 | Learning shape variations of motion trajectories for gait analysisabstractThe analysis of human gait is more and more investigated due to its large panel of potential applications in various domains, like rehabilitation, deficiency diagnosis, surveillance and movement optimization. In addition, the release of depth sensors offers new opportunities to achieve gait analysis in a non-intrusive context. In this paper, we propose a gait analysis method from depth sequences by analyzing separately each step so as to be robust to gait duration and incomplete cycles. We analyze the shape of the motion trajectory as signature of the gait and consider shape variations within a Riemannian manifold to learn step models. During classification, the derivation of each performed step is evaluated in an online manner to qualitatively analyze the gait. Experiments are carried out in the context of abnormal gait detection and person re-identification trough gait recognition. Results demonstrated the potential of the method in both scenarios. Maxime Devanne, Hazem Wannous, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo, Pietro Pala |
ICPR | 2 |
| 2015 | Accurate 3D action recognition using learning on the Grassmann manifold
Rim Slama, Hazem Wannous, Mohamed Daoudi, Anuj Srivastava |
Pattern Recognit. | 2 |
| 2015 | 3-D Human Action Recognition by Shape Analysis of Motion Trajectories on Riemannian ManifoldabstractRecognizing human actions in 3-D video sequences is an important open problem that is currently at the heart of many research domains including surveillance, natural interfaces and rehabilitation. However, the design and development of models for action recognition that are both accurate and efficient is a challenging task due to the variability of the human pose, clothing and appearance. In this paper, we propose a new framework to extract a compact representation of a human action captured through a depth sensor, and enable accurate action recognition. The proposed solution develops on fitting a human skeleton model to acquired data so as to represent the 3-D coordinates of the joints and their change over time as a trajectory in a suitable action space. Thanks to such a 3-D joint-based framework, the proposed solution is capable to capture both the shape and the dynamics of the human body, simultaneously. The action recognition problem is then formulated as the problem of computing the similarity between the shape of trajectories in a Riemannian manifold. Classification using k-nearest neighbors is finally performed on this manifold taking advantage of Riemannian geometry in the open curve shape space. Experiments are carried out on four representative benchmarks to demonstrate the potential of the proposed solution in terms of accuracy/latency for a low-latency action recognition. Comparative results with state-of-the-art methods are reported. Maxime Devanne, Hazem Wannous, Stefano Berretti, Pietro Pala, Mohamed Daoudi, Alberto Del Bimbo |
IEEE Trans. Cybern. | 2 |
| 2014 | Grassmannian Representation of Motion Depth for 3D Human Gesture and Action RecognitionabstractRecently developed commodity depth sensors open up new possibilities of dealing with rich descriptors, which capture geometrical features of the observed scene. Here, we propose an original approach to represent geometrical features extracted from depth motion space, which capture both geometric appearance and dynamic of human body simultaneously. In this approach, sequence features are modeled temporally as subspaces lying on the Grassmann manifold. Classification task is carried out via computation of probability density functions on tangent space of each class tacking benefit from the geometric structure of the Grassmann manifold. The experimental evaluation is performed on three existing datasets containing various challenges, including MSR-action 3D, UT-kinect and MSR-Gesture3D. Results reveal that our approach outperforms the state-of-the-art methods, with accuracy of 98.21% on MSR-Gesture3D and 95.25% on UT-kinect, and achieves a competitive performance of 86.21% on MSR-action 3D. Rim Slama, Hazem Wannous, Mohamed Daoudi |
ICPR | 2 |
| 2014 | 3D human motion analysis framework for shape similarity and retrieval
Rim Slama, Hazem Wannous, Mohamed Daoudi |
Image Vis. Comput. | 2 |
| 2012 | Place Recognition via 3D Modeling for Personal Activity Lifelog Using Wearable Camera
Hazem Wannous, Vladislavs Dovgalecs, Rémi Mégret, Mohamed Daoudi |
MMM | 1 |
| 2011 | Enhanced Assessment of the Wound-Healing Process by Accurate Multiview Tissue ClassificationabstractWith the widespread use of digital cameras, freehand wound imaging has become common practice in clinical settings. There is however still a demand for a practical tool for accurate wound healing assessment, combining dimensional measurements and tissue classification in a single user-friendly system. We achieved the first part of this objective by computing a 3-D model for wound measurements using uncalibrated vision techniques. We focus here on tissue classification from color and texture region descriptors computed after unsupervised segmentation. Due to perspective distortions, uncontrolled lighting conditions and view points, wound assessments vary significantly between patient examinations. The main contribution of this paper is to overcome this drawback with a multiview strategy for tissue classification, relying on a 3-D model onto which tissue labels are mapped and classification results merged. The experimental classification tests demonstrate that enhanced repeatability and robustness are obtained and that metric assessment is achieved through real area and volume measurements and wound outline extraction. This innovative tool is intended for use not only in therapeutic follow-up in hospitals but also for telemedicine purposes and clinical research, where repeatability and accuracy of wound assessment are critical. Hazem Wannous, Yves Lucas, Sylvie Treuillet |
IEEE Trans. Medical Imaging | 1 |
| 2010 | The IMMED project: wearable video monitoring of people with age dementiaabstractIn this paper, we describe a new application for multimedia indexing, using a system that monitors the instrumental activities of daily living to assess the cognitive decline caused by dementia. The system is composed of a wearable camera device designed to capture audio and video data of the instrumental activities of a patient, which is leveraged with multimedia indexing techniques in order to allow medical specialists to analyze several hour long observation shots efficiently. Rémi Mégret, Vladislavs Dovgalecs, Hazem Wannous, Svebor Karaman, Jenny Benois-Pineau, Elie Khoury 0001, Julien Pinquier, Philippe Joly, Régine André-Obrecht, Yann Gaëstel, Jean-François Dartigues |
ACM Multimedia | 3 |
| 2008 | Fusion of Multi-view Tissue Classification Based on Wound 3D Model
Hazem Wannous, Yves Lucas, Sylvie Treuillet, Benjamin Albouy-Kissi |
ACIVS | 1 |
| 2008 | A complete 3D wound assessment tool for accurate tissue classification and measurementabstractThis paper presents the complete 3D and color wound assessment tool, designed using a simple freely handled digital camera inside the ESCALE project. Combining a 3D model of the captured wound images using uncalibrated vision techniques with unsupervised tissue segmentation, it gives access to enhanced tissue classification and measurement. As a result, the tissue classification is directly mapped on the mesh surface of the wound to measure real tissue growth and changes. Clinical tests demonstrate that the monitoring of the healing process is very accurate compared to single view analysis. Hazem Wannous, Yves Lucas, Sylvie Treuillet, Benjamin Albouy-Kissi |
ICIP | 1 |