EDBT 2026 Demo / reviewers in the wild / expert
Enjie Ghorbel
dblp:173/8819
· DBLP profile ↗
26ranked-venue papers
4as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain adaptation for multi-label image classification: A discriminator-free approachabstractThis paper introduces a discriminator-free adversarial-based approach termed DDA-MLIC for Unsupervised Domain Adaptation (UDA) in the context of Multi-Label Image Classification (MLIC). While recent efforts have explored adversarial-based UDA methods for MLIC, they typically include an additional discriminator subnet. Nevertheless, decoupling the classification and the discrimination tasks may harm their task-specific discriminative power. Herein, we address this challenge by presenting a novel adversarial critic directly derived from the task-specific classifier. Specifically, we employ a two-component Gaussian Mixture Model (GMM) to model both source and target predictions, distinguishing between two distinct clusters. Instead of using the traditional Expectation Maximization (EM) algorithm, our approach utilizes a Deep Neural Network (DNN) to estimate the parameters of each GMM component. Subsequently, the source and target GMM parameters are leveraged to formulate an adversarial loss using the Fréchet distance. The proposed framework is therefore not only fully differentiable but is also cost-effective as it avoids the expensive iterative process usually induced by the standard EM method. The proposed method is evaluated on several multi-label image datasets covering three different types of domain shift. The obtained results demonstrate that DDA-MLIC outperforms existing state-of-the-art methods in terms of precision while requiring a lower number of parameters. The code is made publicly available at github.com/cvi2snt/DDA-MLIC . Inder Pal Singh, Enjie Ghorbel, Anis Kacem 0001, Djamila Aouada |
Expert Syst. Appl. | 2 |
| 2025 | Dimensionality Reduction on the SPD Manifold: A Comparative Study of Linear and Non-Linear Methods
Amal Araoud, Enjie Ghorbel, Faouzi Ghorbel |
ICAART (3) | 2 |
| 2025 | Audio-Visual Deepfake Detection With Local Temporal InconsistenciesabstractThis paper proposes an audio-visual deepfake detection approach that aims to capture fine-grained temporal inconsistencies between audio and visual modalities. To achieve this, both architectural and data synthesis strategies are introduced. From an architectural perspective, a temporal distance map, coupled with an attention mechanism, is designed to capture these inconsistencies while minimizing the impact of irrelevant temporal subsequences. Moreover, we explore novel pseudo-fake generation techniques to synthesize local inconsistencies. Our approach is evaluated against state-of-the-art methods using the DFDC and FakeAVCeleb datasets, demonstrating its effectiveness in detecting audio-visual deepfakes. Marcella Astrid, Enjie Ghorbel, Djamila Aouada |
ICASSP | 2 |
| 2025 | Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detectionabstractpeer reviewed Marcella Astrid, Anis Kacem 0001, Enjie Ghorbel, Djamila Aouada |
ICCV | 4 |
| 2025 | Uncertainty-Aware Knowledge Distillation for Compact and Efficient 6DoF Pose EstimationabstractCompact and efficient 6DoF object pose estimation is crucial in applications such as robotics, augmented reality, and space autonomous navigation systems, where lightweight models are critical for real-time accurate performance. This paper introduces a novel uncertainty-aware end-to-end Knowledge Distillation (KD) framework focused on keypoint-based 6DoF pose estimation. Keypoints predicted by a large teacher model exhibit varying levels of uncertainty that can be exploited within the distillation process to enhance the accuracy of the student model while ensuring its compactness. To this end, we propose a distillation strategy that aligns the student and teacher predictions by adjusting the knowledge transfer based on the uncertainty associated with each teacher keypoint prediction. Additionally, the proposed KD leverages this uncertainty-aware alignment of keypoints to transfer the knowledge at key locations of their respective feature maps. Experiments on the widely-used LINEMOD benchmark demonstrate the effectiveness of our method, achieving superior 6DoF object pose estimation with lightweight models compared to state-of-the-art approaches. Further validation on the SPEED+ dataset for spacecraft pose estimation highlights the robustness of our approach under diverse 6DoF pose estimation scenarios. Nassim Ali Ousalah, Anis Kacem 0001, Enjie Ghorbel, Emmanuel Koumandakis, Djamila Aouada |
IROS | 3 |
| 2025 | FPG-NAS: FLOPs-Aware Gated Differentiable Neural Architecture Search for Efficient 6DoF Pose EstimationabstractWe introduce FPG-NAS, a FLOPs-aware Gated Differentiable Neural Architecture Search framework for efficient 6DoF object pose estimation. Estimating 3D rotation and translation from a single image has been widely investigated yet remains computationally demanding, limiting applicability in resource-constrained scenarios. FPG-NAS addresses this by proposing a specialized differentiable NAS approach for 6DoF pose estimation, featuring a task-specific search space and a differentiable gating mechanism that enables discrete multi-candidate operator selection, thus improving architectural diversity. Additionally, a FLOPs regularization term ensures a balanced trade-off between accuracy and efficiency. The framework explores a vast search space of approximately 1092possible architectures. Experiments on the LINEMOD and SPEED+ datasets demonstrate that FPG-NAS-derived models outperform previous methods under strict FLOPs constraints. To the best of our knowledge, FPG-NAS is the first differentiable NAS framework specifically designed for 6DoF object pose estimation. Nassim Ali Ousalah, Peyman Rostami, Anis Kacem 0001, Enjie Ghorbel, Emmanuel Koumandakis, Djamila Aouada |
MMSP | 4 |
| 2024 | Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
Marcella Astrid, Enjie Ghorbel, Djamila Aouada |
BMVC | 2 |
| 2024 | LAA-Net: Localized Artifact Attention Network for Quality-Agnostic and Generalizable Deepfake DetectionabstractThis paper introduces a novel approach for high-quality deepfake detection called Localized Artifact Attention Net-work (LAA-Net). Existing methods for high-quality deep-fake detection are mainly based on a supervised binary classifier coupled with an implicit attention mechanism. As a result, they do not generalize well to unseen ma-nipulations. To handle this issue, two main contributions are made. First, an explicit attention mechanism within a multi-task learning framework is proposed. By combining heatmap-based and self-consistency attention strate-gies, LAA-Net is forced to focus on a few small artifact-prone vulnerable regions. Second, an Enhanced Feature Pyramid Network (E-FPN) is proposed as a simple and ef-fective mechanism for spreading discriminative low-level features into the final feature output, with the advantage of limiting redundancy. Experiments performed on sev-eral benchmarks show the superiority of our approach in terms of Area Under the Curve (AUC) and Average Preci-sion (AP). The code is available at https://github.com/10Ring/LAA-Net. Nesryne Mejri, Inder Pal Singh, Polina Kuleshova, Marcella Astrid, Anis Kacem 0001, Enjie Ghorbel, Djamila Aouada |
CVPR | 7 |
| 2024 | Statistics-Aware Audio-Visual Deepfake DetectorabstractIn this paper, we propose an enhanced audio-visual deep detection method. Recent methods in audio-visual deepfake detection mostly assess the synchronization between audio and visual features. Although they have shown promising results, they are based on the maximization/minimization of isolated feature distances without considering feature statistics. Moreover, they rely on cumbersome deep learning architectures and are heavily dependent on empirically fixed hyperparameters. Herein, to overcome these limitations, we propose: (1) a statistical feature loss to enhance the discrimination capability of the model, instead of relying solely on feature distances; (2) using the waveform for describing the audio as a replacement of frequency-based representations; (3) a post-processing normalization of the fakeness score; (4) the use of shallower network for reducing the computational complexity. Experiments on the DFDC and FakeAVCeleb datasets demonstrate the relevance of the proposed method. Marcella Astrid, Enjie Ghorbel, Djamila Aouada |
ICIP | 2 |
| 2024 | Facial Region-Based Ensembling for Unsupervised Temporal Deepfake LocalizationabstractThis paper addresses the challenge of temporal deepfake localization. Instead of classifying entire videos as real or fake, the goal is isolating forged frames in untrimmed videos that might be partially manipulated. Recently, few deepfake localization methods have emerged. They are mostly supervised, therefore relying on costly annotations and suffering from a lack of generalization to unseen manipulations. As an alternative, we propose reformulating deepfake localization as an unsupervised time-series anomaly detection problem. Hence, to investigate the relevance of the proposed formulation, recent state-of-the-art techniques in anomaly detection for timeseries are evaluated in the context of deepfake localization. To avoid using large architectures, geometric representations, e.g., facial landmarks, are used as input. Moreover, a facialregion based ensembling strategy is introduced for a better modelling of localized deepfake artifacts. Experiments performed on the ForgeryNet dataset demonstrate the effectiveness of the proposed ensembling method and highlight the suitability of the suggested formulation. Nesryne Mejri, Pavel Chernakov, Polina Kuleshova, Enjie Ghorbel, Djamila Aouada |
ICME | 4 |
| 2024 | A Hitchhiker's Guide to Fine-Grained Face Forgery Detection Using Common Sense ReasoningabstractExplainability in artificial intelligence is crucial for restoring trust, particularly in areas like face forgery detection, where viewers often struggle to distinguish between real and fabricated content. Vision and Large Language Models (VLLM) bridge computer vision and natural language, offering numerous applications driven by strong common-sense reasoning. Despite their success in various tasks, the potential of vision and language remains underexplored in face forgery detection, where they hold promise for enhancing explainability by leveraging the intrinsic reasoning capabilities of language to analyse fine-grained manipulation areas. For that reason, few works have recently started to frame the problem of deepfake detection as a Visual Question Answering (VQA) task, nevertheless omitting the realistic and informative open-ended multi-label setting. With the rapid advances in the field of VLLM, an exponential rise of investigations in that direction is expected. As such, there is a need for a clear experimental methodology that converts face forgery detection to a Visual Question Answering (VQA) task to systematically and fairly evaluate different VLLM architectures. Previous evaluation studies in deepfake detection have mostly focused on the simpler binary task, overlooking evaluation protocols for multi-label fine-grained detection and text-generative models. We propose a multi-staged approach that diverges from the traditional binary evaluation protocol and conducts a comprehensive evaluation study to compare the capabilities of several VLLMs in this context. In the first stage, we assess the models' performance on the binary task and their sensitivity to given instructions using several prompts. In the second stage, we delve deeper into fine-grained detection by identifying areas of manipulation in a multiple-choice VQA setting. In the third stage, we convert the fine-grained detection to an open-ended question and compare several matching strategies for the multi-label classification task. Finally, we qualitatively evaluate the fine-grained responses of the VLLMs included in the benchmark. We apply our benchmark to several popular models, providing a detailed comparison of binary, multiple-choice, and open-ended VQA evaluation across seven datasets. \url{https://nickyfot.github.io/hitchhickersguide.github.io/} Niki Maria Foteinopoulou, Enjie Ghorbel, Djamila Aouada |
NeurIPS | 2 |
| 2024 | Discriminator-free Unsupervised Domain Adaptation for Multi-label Image ClassificationabstractIn this paper, a discriminator-free adversarial-based Unsupervised Domain Adaptation (UDA) for Multi-Label Image Classification (MLIC) referred to as DDA-MLIC is proposed. Recently, some attempts have been made for introducing adversarial-based UDA methods in the context of MLIC. However, these methods, which rely on an additional discriminator subnet present one major shortcoming. The learning of domain-invariant features may harm their task-specific discriminative power, since the classification and discrimination tasks are decoupled. Herein, we propose to overcome this issue by introducing a novel adversarial critic that is directly deduced from the task-specific classifier. Specifically, a two-component Gaussian Mixture Model (GMM) is fitted on the source and target predictions in order to distinguish between two clusters. This allows extracting a Gaussian distribution for each component. The resulting Gaussian distributions are then used for formulating an adversarial loss based on a Fréchet distance. The proposed method is evaluated on several multi-label image datasets covering three different types of domain shift. The obtained results demonstrate that DDA-MLIC outperforms existing state-of-the-art methods in terms of precision while requiring a lower number of parameters. The code is publicly available at github.com/cvi2snt/DDA-MLIC.1 Inder Pal Singh, Enjie Ghorbel, Anis Kacem 0001, Arunkumar Rathinam, Djamila Aouada |
WACV | 2 |
| 2024 | Multi-label image classification using adaptive graph convolutional networks: From a single domain to multiple domainsabstractThis paper proposes an adaptive graph-based approach for multi-label image classification. Graph-based methods have been largely exploited in the field of multi-label classification, given their ability to model label correlations. Specifically, their effectiveness has been proven not only when considering a single domain but also when taking into account multiple domains. However, the topology of the used graph is not optimal as it is pre-defined heuristically. In addition, consecutive Graph Convolutional Network (GCN) aggregations tend to destroy the feature similarity. To overcome these issues, an architecture for learning the graph connectivity in an end-to-end fashion is introduced. This is done by integrating an attention-based mechanism and a similarity-preserving strategy. The proposed framework is then extended to multiple domains using an adversarial training scheme. Numerous experiments are reported on well-known single-domain and multi-domain benchmarks. The results demonstrate that our approach achieves competitive results in terms of mean Average Precision (mAP) and model size as compared to the state-of-the-art. The code will be made publicly available. Inder Pal Singh, Enjie Ghorbel, Oyebade K. Oyedotun, Djamila Aouada |
Comput. Vis. Image Underst. | 2 |
| 2024 | Unsupervised anomaly detection in time-series: An extensive evaluation and analysis of state-of-the-art methodsabstractpeer reviewed Nesryne Mejri, Laura Lopez-Fuentes, Kankana Roy, Pavel Chernakov, Enjie Ghorbel, Djamila Aouada |
Expert Syst. Appl. | 5 |
| 2024 | DermSynth3D: Synthesis of in-the-wild annotated dermatology images
Ashish Sinha, Jeremy Kawahara, Arezou Pakzad, Kumar Abhishek 0001, Matthieu Ruthven, Enjie Ghorbel, Anis Kacem 0001, Djamila Aouada, Ghassan Hamarneh |
Medical Image Anal. | 6 |
| 2023 | UNTAG: Learning Generic Features for Unsupervised Type-Agnostic Deepfake DetectionabstractThis paper introduces a novel framework for unsupervised type-agnostic deepfake detection called UNTAG. Existing methods are generally trained in a supervised manner at the classification level, focusing on detecting at most two types of forgeries; thus, limiting their generalization capability across different deepfake types. To handle that, we reformulate the deepfake detection problem as a one-class classification supported by a self-supervision mechanism. Our intuition is that by estimating the distribution of real data in a discriminative feature space, deepfakes can be detected as outliers regardless of their type. UNTAG involves two sequential steps. First, deep representations are learned based on a self-supervised pretext task focusing on manipulated regions. Second, a oneclass classifier fitted on authentic image embeddings is used to detect deepfakes. The results reported on several datasets show the effectiveness of UNTAG and the relevance of the proposed new paradigm. The code is publicly available. Nesryne Mejri, Enjie Ghorbel, Djamila Aouada |
ICASSP | 2 |
| 2023 | Multi-Label Deepfake ClassificationabstractIn this paper, we investigate the suitability of current multi-label classification approaches for deepfake detection. With the recent advances in generative modeling, new deepfake detection methods have been proposed. Nevertheless, they mostly formulate this topic as a binary classification problem, resulting in poor explainability capabilities. Indeed, a forged image might be induced by multi-step manipulations with different properties. For a better interpretability of the results, recognizing the nature of these stacked manipulations is highly relevant. For that reason, we propose to model deepfake detection as a multi-label classification task, where each label corresponds to a specific kind of manipulation. In this context, state-of-the-art multi-label image classification methods are considered. Extensive experiments are performed to assess the practical use case of deepfake detection. Inder Pal Singh, Nesryne Mejri, Enjie Ghorbel, Djamila Aouada |
MMSP | 4 |
| 2022 | Multi Label Image Classification using Adaptive Graph Convolutional Networks (ML-AGCN)abstractIn this paper, a novel graph-based approach for multi-label image classification called Multi-Label Adaptive Graph Convolutional Network (ML-AGCN) is introduced. Graph-based methods have shown great potential in the field of multi-label classification. However, these approaches heuristically fix the graph topology for modeling label dependencies, which might be not optimal. To handle that, we propose to learn the topology in an end-to-end manner. Specifically, we incorporate an attention-based mechanism for estimating the pairwise importance between graph nodes and a similarity-based mechanism for conserving the feature similarity between different nodes. This offers a more flexible way for adaptively modeling the graph. Experimental results are reported on two well-known datasets, namely, MS-COCO and VG-500. Results show that ML-AGCN outperforms state-of-the-art methods while reducing the number of model parameters. Inder Pal Singh, Enjie Ghorbel, Oyebade K. Oyedotun, Djamila Aouada |
ICIP | 2 |
| 2020 | DeepVI: A Novel Framework for Learning Deep View-Invariant Human Action Representations using a Single RGB CameraabstractIn this paper, we address the problem of cross-view action recognition from a monocular RGB camera. This topic has been considered extremely challenging due to the lack of 3D information in 2D images. Exploiting the advances in 3D pose estimation from a single RGB camera, we propose a new framework termed DeepVI, for cross-view action recognition without the need for pose alignment. Virtual viewpoints are used to augment the variability of training data along with the use of an end-to-end Deep Neural Network (DNN). The proposed network is composed of two modules. The first one, called SmoothNet, implicitly smooths skeleton joint trajectories using revisited temporal convolution in order to reduce the noise in the estimated 3D skeletons. The second module consists of a state-of-the-art approach designed for action recognition based on Spatial Temporal Graph Convolutional Networks (ST-GCN [40]). Experiments have been conducted in cross-view settings on two datasets, namely, NTU RGB-D and Northwestern-UCLA. The obtained results show the effectiveness of the proposed framework. Konstantinos Papadopoulos 0002, Enjie Ghorbel, Oyebade K. Oyedotun, Djamila Aouada, Björn Ottersten 0001 |
FG | 2 |
| 2020 | Vertex Feature Encoding and Hierarchical Temporal Modeling in a Spatio-Temporal Graph Convolutional Network for Action RecognitionabstractSpatio-temporal Graph Convolutional Networks (ST-GCNs) have shown great performance in the context of skeleton-based action recognition. Nevertheless, ST-GCNs use raw skeleton data as vertex features. Such features have low dimensionality and might not be optimal for action discrimination. Moreover, a single layer of temporal convolution is used to model short-term temporal dependencies but can be insufficient for capturing both long-term. In this paper, we extend the Spatio-Temporal Graph Convolutional Network for skeleton-based action recognition by introducing two novel modules, namely, the Graph Vertex Feature Encoder (GVFE) and the Dilated Hierarchical Temporal Convolutional Network (DH-TCN). On the one hand, the GVFE module learns appropriate vertex features for action recognition by encoding raw skeleton data into a new feature space. On the other hand, the DH-TCN module is capable of capturing both short-term and long-term temporal dependencies using a hierarchical dilated convolutional network. Experiments have been conducted on the challenging NTU RGB-D 60, NTU RGB-D 120 and Kinetics datasets. The obtained results show that our method competes with state-of-the-art approaches while using a smaller number of layers and parameters; thus reducing the required training time and memory. Konstantinos Papadopoulos 0002, Enjie Ghorbel, Djamila Aouada, Björn Ottersten 0001 |
ICPR | 2 |
| 2020 | Fast Adaptive Reparametrization (FAR) With Application to Human Action RecognitionabstractIn this letter, a fast approach for curve reparametrization, called Fast Adaptive Reparamterization (FAR), is introduced. Instead of computing an optimal matching between two curves such as Dynamic Time Warping (DTW) and elastic distance-based approaches, our method is applied to each curve independently, leading to linear computational complexity. It is based on a simple replacement of the curve parameter by a variable invariant under specific variations of reparametrization. The choice of this variable is heuristically made according to the application of interest. In addition to being fast, the proposed reparametrization can be applied not only to curves observed in Euclidean spaces but also to feature curves living in Riemannian spaces. To validate our approach, we apply it to the scenario of human action recognition using curves living in the Riemannian product Special Euclidean space$\mathbb {SE}(3)^n$. The obtained results on three benchmarks for human action recognition (MSRAction3D, Florence3D, and UTKinect) show that our approach competes with state-of-the-art methods in terms of accuracy and computational cost. Enjie Ghorbel, Girum G. Demisse, Djamila Aouada, Björn Ottersten 0001 |
IEEE Signal Process. Lett. | 1 |
| 2019 | Two-Stage RGB-Based Action Detection Using Augmented 3D Poses
Konstantinos Papadopoulos 0002, Enjie Ghorbel, Renato Baptista, Djamila Aouada, Björn Ottersten 0001 |
CAIP (1) | 2 |
| 2019 | View-invariant Action Recognition from RGB Data via 3D Pose EstimationabstractIn this paper, we propose a novel view-invariant action recognition method using a single monocular RGB camera. View-invariance remains a very challenging topic in 2D action recognition due to the lack of 3D information in RGB images. Most successful approaches make use of the concept of knowledge transfer by projecting 3D synthetic data to multiple viewpoints. Instead of relying on knowledge transfer, we propose to augment the RGB data by a third dimension by means of 3D skeleton estimation from 2D images using a CNN-based pose estimator. In order to ensure view-invariance, a pre-processing for alignment is applied followed by data expansion as a way for denoising. Finally, a Long-Short Term Memory (LSTM) architecture is used to model the temporal dependency between skeletons. The proposed network is trained to directly recognize actions from aligned 3D skeletons. The experiments performed on the challenging Northwestern-UCLA dataset show the superiority of our approach as compared to state-of-the-art ones. Renato Baptista, Enjie Ghorbel, Konstantinos Papadopoulos 0002, Girum G. Demisse, Djamila Aouada, Björn Ottersten 0001 |
ICASSP | 2 |
| 2018 | An extension of kernel learning methods using a modified Log-Euclidean distance for fast and accurate skeleton-based Human Action Recognition
Enjie Ghorbel, Jacques Boonaert, Rémi Boutteau, Stéphane Lecoeuche, Xavier Savatier |
Comput. Vis. Image Underst. | 1 |
| 2018 | Kinematic Spline Curves: A temporal invariant descriptor for fast action recognition
Enjie Ghorbel, Rémi Boutteau, Jacques Boonaert, Xavier Savatier, Stéphane Lecoeuche |
Image Vis. Comput. | 1 |
| 2016 | A fast and accurate motion descriptor for human action recognition applicationsabstractWith the availability of the recent human skeleton extraction algorithm introduced by Shotton et al. [1], an interest for skeleton-based action recognition methods has been renewed. Despite the importance of the low-latency aspect in applications, it can be noted that the majority of recent approaches has not been evaluated in terms of computational cost. In this paper, a novel fast and accurate human action descriptor named Kinematic Spline Curves (KSC) is introduced. This descriptor is built by interpolating the kinematics of joints (position, velocity and acceleration). To overcome the anthropometric and the execution rate variability, we respectively propose the use of a skeleton normalization and a temporal normalization. For this purpose, a new temporal normalization method based on the Normalized Accumulated kinetic Energy (NAE) of the human skeleton is suggested. Finally, the classification step is performed using a linear Support Vector Machine (SVM). Experimental results on challenging benchmarks show the efficiency of our approach in terms of recognition accuracy and computational latency. Enjie Ghorbel, Rémi Boutteau, Jacques Boonaert, Xavier Savatier, Stéphane Lecoeuche |
ICPR | 1 |