Rim Slama

dblp:129/6684 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-2723-9874ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Graph-based framework for temporal human action recognition and segmentation in industrial context
Toufik Benmessabih, Rim Slama, Vincent Havard, David Baudry
Eng. Appl. Artif. Intell.2
2024 Spatio-Temporal Sparse Graph Convolution Network for Hand Gesture Recognition
abstract
Unlike whole-body action recognition, hand gestures involve spatially closely distributed joints, promoting stronger collaboration. This needs to be taken into account in order to capture complex spatial and temporal features. In response to these challenges, this paper presents a Spatio-Temporal Sparse Graph Convolution Network (ST-SGCN) for dynamic recognition of hand gestures. Based on decoupled spatio-temporal processing, the ST-SGCN incorporates Graph Convolutional Networks, attention mechanism and asymmetric convolutions to capture the nuanced movements of hand joints. The key novelty is the introduction of sparse spatio-temporal directed interactions, overcoming the limitations associated with dense, undirected methods. The sparse aspect models essential interactions between hand joints selectively, improving computational efficiency and interpretability. Directed in-teractions capture asymmetrical dependencies between hand joints, improving discernment of joint influences. Experimental evaluations on three benchmark datasets, including Briareo, SHREC'17 and IPN Hand, demonstrate ST-SGCN's state-of-the-art performance for dynamic hand gesture recognition.
Omar Ikne, Rim Slama, Hichem Saoudi, Hazem Wannous
FG2
2024 Online human motion analysis in industrial context: A review
Toufik Benmessabih, Rim Slama, Vincent Havard, David Baudry
Eng. Appl. Artif. Intell.2
2023 STr-GCN: Dual Spatial Graph Convolutional Network and Transformer Graph Encoder for 3D Hand Gesture Recognition
abstract
Skeleton-based hand gesture recognition is a challenging task that sparked a lot of attention in recent years, especially with the rise of Graph Neural Networks. In this paper, we propose a new deep learning architecture for hand gesture recognition using 3D hand skeleton data and we call STr-GCN. It decouples the spatial and temporal learning of the gesture by leveraging Graph Convolutional Networks (GCN) and Transformers. The key idea is to combine two powerful networks: a Spatial Graph Convolutional Network unit that understands intra-frame interactions to extract powerful features from different hand joints and a Transformer Graph Encoder which is based on a Temporal Self-Attention module to incorporate inter-frame correlations. We evaluate the performance of our method on three benchmarks: the SHREC'17 Track dataset, Briareo dataset and the First Person Hand Action dataset. The experiments show the efficiency of our approach, which achieves or outperforms the state of the art. The code to reproduce our results is available in this link.
Rim Slama, Wael Rabah, Hazem Wannous
FG1
2023 Capsule Transformer Network for Dynamic Hand Gesture Recognition Using Multimodal Data
abstract
In recent years, deep learning techniques have achieved remarkable success in video analysis and more especially in action and gesture recognition. Even though convolutional neural networks (CNNs) remain the most widely used models, they have difficulty in capturing the global contextual information involving spatial and temporal domains or intermodality due to the local feature learning mechanism. This paper introduces a Capsule Transformer Network, which composed of a frame capsule module for extracting hand features and a gesture transformer module for modeling the temporal features and recognizing the dynamic gesture. Spatial attention is ensured through the capsule module to enhance the spatial information of the hand image, while the transformer module guarantees temporal attention through gesture sequence. We propose to use multimodal data, including RGB, depth and IR data, which improves the accuracy of our approach as it better captures the 3D structure of the hand and can distinguish between similar hand gestures. Testing on two datasets, Briareo and SHREC17, the proposed approach outperforms or equals previous methods.
Alexandre Lebas, Rim Slama, Hazem Wannous
ICIP2
2023 CG-MER: a card game-based multimodal dataset for emotion recognition
abstract
The field of affective computing has seen significant advancements in exploring the relationship between emotions and emerging technologies. This paper presents a novel and valuable contribution to this field with the introduction of a comprehensive French multimodal dataset designed specifically for emotion recognition. The dataset encompasses three primary modalities: facial expressions, speech, and gestures, providing a holistic perspective on emotions. Moreover, the dataset has the potential to incorporate additional modalities, such as Natural Language Processing (NLP) to expand the scope of emotion recognition research. The dataset was curated through engaging participants in card game sessions, where they were prompted to express a range of emotions while responding to diverse questions. The study included 10 sessions with 20 participants (9 females and 11 males). The dataset serves as a valuable resource for furthering research in emotion recognition and provides an avenue for exploring the intricate connections between human emotions and digital technologies.
Nisrine Farhat, Amine Bohi, Leila Ben Letaifa, Rim Slama
ICMV4
2023 MR-STGN: Multi-Residual Spatio Temporal Graph Network Using Attention Fusion for Patient Action Assessment
abstract
Accurate assessment of patient actions plays a crucial role in healthcare as it contributes significantly to disease progression monitoring and treatment effectiveness. However, traditional approaches to assess patient actions often rely on manual observation and scoring, which are subjective and time-consuming. In this paper, we propose an automated approach for patient action assessment using a Multi-Residual Spatio Temporal Graph Network (MR-STGN) that incorporates both angular and positional 3D skeletons. The MR-STGN is specifically designed to capture the spatio-temporal dynamics of patient actions. It achieves this by integrating information from multiple residual layers, with each layer extracting features at distinct levels of abstraction. Furthermore, we integrate an attention fusion mechanism into the network, which facilitates the adaptive weighting of various features. This empowers the model to concentrate on the most pertinent aspects of the patient's movements, offering precise instructions regarding specific body parts or movements that require attention. Ablation studies are conducted to analyze the impact of individual components within the proposed model. We evaluate our model on the UI-PRMD dataset demonstrating its performance in accurately predicting real-time patient action scores, surpassing state-of-the-art methods.
Youssef Mourchid, Rim Slama
MMSP2
2015 Accurate 3D action recognition using learning on the Grassmann manifold
Rim Slama, Hazem Wannous, Mohamed Daoudi, Anuj Srivastava
Pattern Recognit.1
2014 Grassmannian Representation of Motion Depth for 3D Human Gesture and Action Recognition
abstract
Recently developed commodity depth sensors open up new possibilities of dealing with rich descriptors, which capture geometrical features of the observed scene. Here, we propose an original approach to represent geometrical features extracted from depth motion space, which capture both geometric appearance and dynamic of human body simultaneously. In this approach, sequence features are modeled temporally as subspaces lying on the Grassmann manifold. Classification task is carried out via computation of probability density functions on tangent space of each class tacking benefit from the geometric structure of the Grassmann manifold. The experimental evaluation is performed on three existing datasets containing various challenges, including MSR-action 3D, UT-kinect and MSR-Gesture3D. Results reveal that our approach outperforms the state-of-the-art methods, with accuracy of 98.21% on MSR-Gesture3D and 95.25% on UT-kinect, and achieves a competitive performance of 86.21% on MSR-action 3D.
Rim Slama, Hazem Wannous, Mohamed Daoudi
ICPR1
2014 3D human motion analysis framework for shape similarity and retrieval
Rim Slama, Hazem Wannous, Mohamed Daoudi
Image Vis. Comput.1
2013 3D Face Recognition under Expressions, Occlusions, and Pose Variations
abstract
We propose a novel geometric framework for analyzing 3D faces, with the specific goals of comparing, matching, and averaging their shapes. Here we represent facial surfaces by radial curves emanating from the nose tips and use elastic shape analysis of these curves to develop a Riemannian framework for analyzing shapes of full facial surfaces. This representation, along with the elastic Riemannian metric, seems natural for measuring facial deformations and is robust to challenges such as large facial expressions (especially those with open mouths), large pose variations, missing parts, and partial occlusions due to glasses, hair, and so on. This framework is shown to be promising from both--empirical and theoretical--perspectives. In terms of the empirical evaluation, our results match or improve upon the state-of-the-art methods on three prominent databases: FRGCv2, GavabDB, and Bosphorus, each posing a different type of challenge. From a theoretical perspective, this framework allows for formal statistical inferences, such as the estimation of missing facial parts using PCA on tangent spaces and computing average shapes.
Hassen Drira, Boulbaba Ben Amor, Anuj Srivastava, Mohamed Daoudi, Rim Slama
IEEE Trans. Pattern Anal. Mach. Intell.5