Benjamin Allaert

dblp:165/8166 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-4291-9803ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SIP-DDoS framework based on federated learning for collaborative anomaly detection
abstract
Abstract With their inherent capacity for massive connectivity, ultra-low latency and high reliability, B5G networks provide an ideal infrastructure to support the diverse and dynamic requirements of IoT (Internet of Things) communication. Originally designed to initiate, modify and terminate multimedia sessions over IP networks, the Session Initiation Protocol (SIP), a standardized protocol developed by the 3GPP, has emerged as a promising protocol for enabling communication and coordination in IoT environments. Nonetheless, SIP encounters a multitude of Distributed Denial of Service (DDoS) threats, with INVITE flooding attacks emerging as a notable challenge. Traditional IoT-IDS (Intrusion Detection System) relies on machine learning models trained on local data of their deployment context. However, such a model may not detect some attack patterns observed in other deployment contexts. We propose a design approach based on federated learning, and in which different IDS collaborate to achieve early detection of any INVITE flooding attack faced by any of them. The results show the effectiveness of the framework in detecting and mitigating INVITE flooding attacks across a various of flow intensities under realistic SIP operational conditions. Performance evaluations under different relevant scenarios demonstrate the robustness of the proposed framework. Experiments show that federated learning (FL) enhances the analysis of SIP flooding attacks, improving accuracy from 47 to 99% through knowledge sharing. The FL model with a GRU architecture and a FedAvgM aggregation function delivers the best performance, even in varied scenarios.
Oussama Sbai, Benjamin Allaert, Patrick Sondi, Ahmed Meddahi
Cybersecur.2
2026 Survey on Vision-Based Risk Assessment for Automated Very Light Rail Transportation
abstract
In recent years, various Automated Driving Systems (ADS) leveraging machine learning (ML) and deep learning (DL) have been developed to enhance passenger comfort and safety. These systems can monitor and record accident-prone events, such as the behavior of road users and abnormal occurrences, using data from sensors placed at different locations on the vehicle. However, similar advancements are still in the early stages within the railway domain, where automotive applications must be adapted for trains, requiring appropriate sensor systems and vehicle-specific adjustments. This article reviews the primary approaches to vision-based live risk assessment (VLRA) for light rail or very light rail systems vehicles that bridge the gap between cars and trains, representing an intermediate stage in transportation technology. The survey categorizes the main approaches based on the type of risk factor, learning-based methods used, evaluation metrics, specific sensors or processes involved, and relevant datasets. Finally, we provide conclusions and propose directions for future research.
Justin Bescop, Benjamin Allaert, Jean-Philippe Vandeborre
IEEE Trans. Intell. Transp. Syst.2
2025 AG-MAE: Anatomically Guided Spatio-Temporal Masked Auto-Encoder for Online Hand Gesture Recognition
abstract
Hand gesture recognition plays a crucial role in the domain of computer vision, as it enhances human-computer interaction by enabling intuitive, touch-free control and communication. While offline methods have made significant advances in isolated gesture recognition, real-world applications demand online and continuous processing. Skeleton-based methods, though effective, face challenges due to the intricate nature of hand joints and the diverse 3D motions they induce. This paper introduces AG-MAE, a novel approach that integrates anatomical constraints to guide the self-supervised training of a spatio-temporal masked autoencoder, enhancing the learning of 3D keypoint representations. By incorporating anatomical knowledge, AG-MAE learns more discriminative features for hand poses and movements, subsequently improving online gesture recognition. Evaluation on standard datasets demonstrates the superiority of our approach and its potential for real-world applications. Code is available at: https://github.com/o-ikne/AG-MAE.git.
Omar Ikne, Benjamin Allaert, Hazem Wannous
3DV2
2025 TREB: Temporal Refinement of Egocentric Body Pose
abstract
Accurate egocentric body-pose estimation is essential for delivering immersive experiences in Virtual, Augmented, and Mixed Reality (VR/AR/MR) applications. A major challenge in this setting arises from the limited field of view provided by head-mounted cameras, which often leads to self-occlusion and poor visibility of certain body joints-particularly in the lower body. Current state-of-the-art methods typically perform pose estimation frame-by-frame in a view-aligned manner. While this is computationally efficient and performs well for visible joints, it struggles to estimate occluded joints due to the lack of contextual information. In this work, we explore the role of temporal information in improving joint estimation. We propose a lightweight model that refines pose predictions by leveraging a large temporal window. Our results show that incorporating such module improves the accuracy of joint estimation without incurring a large computation overhead.
Bruno Henriques, Benjamin Allaert, Nicolas Sutton-Charani, Pierre R. L. Slangen, Jean-Philippe Vandeborre
CBMI2
2025 eMotion-GAN: A motion-based GAN for photorealistic and facial expression preserving frontal view synthesis
abstract
Facial expression recognition (FER) systems frequently suffer significant performance degradation when confronted with head pose variations, a pervasive challenge in real-world applications ranging from healthcare monitoring to human–computer interaction. While existing frontal view synthesis (FVS) methods attempt to address this issue, they predominantly operate in the appearance domain, often introducing artifacts that distort the subtle motion patterns crucial for accurate expression analysis. We present eMotion-GAN, a two-stage generative motion-domain framework that fundamentally rethinks frontalization by decomposing facial dynamics into two distinct components: (1) expression-related motion stemming from muscle activity, and (2) pose-related motion acting as noise. We conducted extensive evaluations using several widely recognized dynamic FER datasets, which encompass sequences exhibiting various degrees of head pose variations in both intensity and orientation. Our results demonstrate the effectiveness of our approach in significantly reducing the FER performance gap between frontal and non-frontal faces. Specifically, we achieved a FER improvement of up to +5% for small pose variations and up to +20% improvement for larger pose variations. Code and pre-trained models are available at: https://github.com/o-ikne/eMotion-GAN.git . • Treats head pose as structured noise in optical flow for robust frontalization. • Needs no landmarks; not affected by inaccurate facial landmark detection. • Splits facial motion into pose and expression; gains 20% FER accuracy for poses. • Enables expression transfer for animation and facial data augmentation. • Reduces artifacts and outperforms appearance-based frontalization methods.
Omar Ikne, Benjamin Allaert, Ioan Marius Bilasco, Hazem Wannous
Comput. Vis. Image Underst.2
2024 Motion Consistency Constraint Map for Facial Expression Spotting
abstract
Facial expression spotting is an effective metric for categorizing human behavior changes. It refers to the precise localization of the temporal intervals in a sequence where a visual event occurs in a face. In this paper, we propose an innovative framework, which relies on the consistency in terms of orientation and intensity of the local facial motions. First, we build local facial motion consistency maps to differentiate expression-related facial motion from facial noise. Then, these maps are fed into a recurrent neural network to precisely delineate the temporal progression of facial expression activation. Extensive evaluations were undertaken on SNAP-2DFE dataset demonstrating the effectiveness of the proposed framework in temporally segmenting expression activation in presence of low or high head pose variations
Ouala Ben Jemaa, Amel Aissaoui, Benjamin Allaert, Ioan Marius Bilasco
CBMI3
2024 Skeleton-Based Self-Supervised Feature Extraction for Improved Dynamic Hand Gesture Recognition
abstract
Human-computer interaction (HCI) has become integral to modern life, especially in digital environments. However, challenges persist in utilizing hand gestures due to factors such as the dynamic nature of gestures and the intricacies of intra and inter-finger movements. In this paper, we propose an innovative approach to improve skeleton-based hand gesture recognition by integrating self-supervised learning, a promising technique for acquiring distinctive representations directly from unlabeled data. The proposed method takes advantage of prior knowledge of hand topology, combining topology-aware self-supervised learning with a customized skeleton-based architecture to derive meaningful representations from skeleton data under different hand poses. We introduce customized masking strategies for skeletal hand data and design a model architecture that incorporates spatial connectivity information, improving the model's understanding of the interrelationships between hand joints. The extensive experiments demonstrate the effectiveness of the approach, with state-of-the-art performance on benchmark datasets. An exploration of the generalization of learned representations across datasets and a study of the impact of fine-tuning with limited labeled data are conducted, highlighting the adaptability and robustness of our approach.
Omar Ikne, Benjamin Allaert, Hazem Wannous
FG2
2024 SHREC 2024: Recognition of dynamic hand motions molding clay
abstract
Gesture recognition is a tool to enable novel interactions with different techniques and applications, like Mixed Reality and Virtual Reality environments. With all the recent advancements in gesture recognition from skeletal data, it is still unclear how well state-of-the-art techniques perform in a scenario using precise motions with two hands. This paper presents the results of the SHREC 2024 contest organized to evaluate methods for their recognition of highly similar hand motions using the skeletal spatial coordinate data of both hands. The task is the recognition of 7 motion classes given their spatial coordinates in a frame-by-frame motion. The skeletal data has been captured using a Vicon system and pre-processed into a coordinate system using Blender and Vicon Shogun Post. We created a small, novel dataset with a high variety of durations in frames. This paper shows the results of the contest, showing the techniques created by the 5 research groups on this challenging task and comparing them to our baseline method.
Ben Veldhuijzen, Remco C. Veltkamp, Omar Ikne, Benjamin Allaert, Hazem Wannous, Marco Emporio, Andrea Giachetti 0001, Joseph J. LaViola Jr., He Ruiwen, Halim Benhabiles, Adnane Cabani, Anthony Fleury, Karim Hammoudi, Konstantinos Gavalas, Christoforos Vlachos, Athanasios Papanikolaou, Ioannis Romanelis, Vlassis Fotis, Gerasimos Arvanitis, Konstantinos Moustakas, Martin Hanik, Esfandiar Nava-Yazdani, Christoph von Tycowicz
Comput. Graph.4
2023 Spiking-Fer: Spiking Neural Network for Facial Expression Recognition With Event Cameras
abstract
Facial Expression Recognition (FER) is an active research domain that has shown great progress recently, notably thanks to the use of large deep learning models. However, such approaches are particularly energy intensive, which makes their deployment difficult for edge devices. To address this issue, Spiking Neural Networks (SNNs) coupled with event cameras are a promising alternative, capable of processing sparse and asynchronous events with lower energy consumption. In this paper, we establish the first use of event cameras for FER, named "Event-based FER", and propose the first related benchmarks by converting popular video FER datasets to event streams. To deal with this new task, we propose "Spiking-FER", a deep convolutional SNN model, and compare it against a similar Artificial Neural Network (ANN). Experiments show that the proposed approach achieves comparable performance to the ANN architecture, while consuming less energy by orders of magnitude (up to 65.39x). In addition, an experimental study of various event-based data augmentation techniques is performed to provide insights into the efficient transformations specific to event-based FER.
Sami Barchid, Benjamin Allaert, Amel Aissaoui, José Mennesson, Chaabane Djeraba
CBMI2
2023 Impact of Facial Landmark Localization on Facial Expression Recognition
abstract
Although facial landmark localization (FLL) approaches are becoming increasingly accurate in identifying facial components, one question remains unanswered: what is the impact of these approaches on subsequent, related tasks? In this paper, we focus on facial expression recognition (FER), where facial landmarks are used for face registration, which is a common usage. Since the common datasets for facial landmark localization do not allow for a proper measurement of performance according to the different difficulties (e.g., pose, expression, illumination, occlusion, motion blur), we also quantify the performance of recent approaches in the presence of head pose variations and facial expressions. Finally, we conduct a study of the impact of these approaches on FER. We show that the landmark accuracy achieved so far by optimizing the euclidean distance does not necessarily guarantee a gain in performance for FER. To deal with this issue, we propose a new evaluation metric for FLL that is more relevant to FER.
Romain Belmonte, Benjamin Allaert, Pierre Tirilly, Ioan Marius Bilasco, Chaabane Djeraba, Nicu Sebe
IEEE Trans. Affect. Comput.2
2022 A comparative study on optical flow for facial expression analysis
Benjamin Allaert, Isaac Ronald Ward, Ioan Marius Bilasco, Chaabane Djeraba, Mohammed Bennamoun
Neurocomputing1
2022 Micro and Macro Facial Expression Recognition Using Advanced Local Motion Patterns
abstract
In this paper, we develop a new method that recognizes facial expressions, on the basis of an innovative Local Motion Patterns (LMP) feature. The LMP feature analyzes locally the motion distribution in order to separate consistent mouvement patterns from noise. Indeed, facial motion extracted from the face is generally noisy and without specific processing, it can hardly cope with expression recognition requirements especially for micro-expressions. Direction and magnitude statistical profiles are jointly analyzed in order to filter out noise. This work presents three main contributions. The first one is the analysis of the face skin temporal elasticity and face deformations during expression. The second one is a unified approach for both macro and micro expression recognition leading the way to supporting a wide range of expression intensities. The third one is the step forward towards in-the-wild expression recognition, dealing with challenges such as various intensity and various expression activation patterns, illumination variations and small head pose variations. Our method outperforms state-of-the-art methods for micro expression recognition and positions itself among top-ranked state-of-the-art methods for macro expression recognition.
Benjamin Allaert, Ioan Marius Bilasco, Chaabane Djeraba
IEEE Trans. Affect. Comput.1
2022 Dynamic Facial Expression Recognition Under Partial Occlusion With Optical Flow Reconstruction
abstract
Video facial expression recognition is useful for many applications and received much interest lately. Although some methods give good results in controlled environments (no occlusion), recognition in the presence of partial facial occlusion remains a challenging task. To handle facial occlusions, methods based on the reconstruction of the occluded part of the face have been proposed. These methods are mainly based on the texture or the geometry of the face. However, the similarity of the face movement between different persons doing the same expression seems to be a real asset for the reconstruction. In this paper we exploit this asset and propose a new method based on an auto-encoder with skip connections to reconstruct the occluded part of the face in the optical flow domain. To the best of our knowledge, this is the first work that directly reconstructs the movement for facial expression recognition. We validated our approach in the controlled CK+ datasets on which different occlusions were generated. Our experiments show that the proposed method reduces the gap in the recognition accuracy between occluded and unoccluded situations. We also compare our approach with existing state-of-the-art approaches. In order to lay the basis of a reproducible and fair comparison in the future, we also propose a new experimental protocol that includes occlusion generation and reconstruction evaluation.
Delphine Poux, Benjamin Allaert, Nacim Ihaddadene, Ioan Marius Bilasco, Chaabane Djeraba, Mohammed Bennamoun
IEEE Trans. Image Process.2
2021 BAREM: A multimodal dataset of individuals interacting with an e-service platform
abstract
The use of e-service platforms has become essential for many applications (administrative documents, online shopping, reservations). Although these platforms have improved significantly the user experience, unexpected and stressful situations can occur. Navigation problems (latency, missing information, poor ergonomics) are not always reported to the designers. To address this problem, we propose a multimodal dataset (video, audio, and physiological data) to help implicitly quantify the impact of navigation problems on users when using an e-service platform. A scenario has been designed to generate various navigation problems which can lead to changes in user behaviour. A baseline is proposed to spot changes in user behaviour, opening the way towards automatically qualifying user experiences while using e-service platforms.
Romain Belmonte, Amel Aissaoui, Sofiane Mihoubi, Benjamin Allaert, José Mennesson, Ioan Marius Bilasco, Laurent Goncalves
CBMI4
2021 Facial expressions analysis under occlusions based on specificities of facial motion propagation
Delphine Poux, Benjamin Allaert, José Mennesson, Nacim Ihaddadene, Ioan Marius Bilasco, Chaabane Djeraba
Multim. Tools Appl.2
2018 Mastering Occlusions by Using Intelligent Facial Frameworks Based on the Propagation of Movement
abstract
In uncontrolled settings occlusions occur and interfere with facial expressions recognition task. It is interesting to limit the number of regions required for face expression recognition task in order to moderate the occlusion interference. We propose a weighting scheme that ranks the facial regions needed to recognize expressions. Weights are calculated based on the contribution of each region to boost recognition in presence of various occlusions. Intelligent facial frameworks, based on region ranks are computed in presence of static occlusions (such as glasses, hair, hand on the face). Evaluations conducted using motion information as the underlying descriptor show that our approach maintains, per expression, very good recognition rates under various static occlusions occurring in uncontrolled settings.
Delphine Poux, Benjamin Allaert, José Mennesson, Nacim Ihaddadene, Ioan Marius Bilasco, Chaabane Djeraba
CBMI2
2018 Impact of the face registration techniques on facial expressions recognition
Benjamin Allaert, José Mennesson, Ioan Marius Bilasco, Chaabane Djeraba
Signal Process. Image Commun.1