EDBT 2026 Demo / reviewers in the wild / expert
Djamila Aouada
dblp:20/7872
· DBLP profile ↗
100ranked-venue papers
7as first author
47since 2021 · last 2026
0000-0002-7576-2064ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 71 · 7 first-author · 31 since 2021Artificial intelligence and machine learning · 46 · 1 first-author · 26 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASTER: Latent Pseudo-Anomaly Generation for Unsupervised Time-Series Anomaly Detection
Romain Hermary, Samet Hicsonmez, Dan Pineau, Abd El Rahman Shabayek, Djamila Aouada |
ICPR (3) | 5 |
| 2026 | Training Free Zero-Shot Image Anomaly Localisation via Diffusion Inversion
Samet Hicsonmez, Abd El Rahman Shabayek, Djamila Aouada |
ICPR (1) | 3 |
| 2026 | VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with DiffusionabstractDetecting visual anomalies in diverse, multi-class real-world images is a significant challenge. We introduce VLMDiff, a novel unsupervised multi-class visual anomaly detection framework. It integrates a Latent Diffusion Model (LDM) with a Vision-Language Model (VLM) for enhanced anomaly localization and detection. Specifically, a pretrained VLM with a simple prompt extracts detailed image descriptions, serving as additional conditioning for LDM training. Current diffusion-based methods rely on synthetic noise generation, limiting their generalization and requiring per-class model training, which hinders scalability. VLMDiff, however, leverages VLMs to obtain normal captions without manual annotations or additional training. These descriptions condition the diffusion model, learning a robust normal image feature representation for multiclass anomaly detection. Our method achieves competitive performance, improving the pixel-level Per-Region-Overlap (PRO) metric by up to 25 points on the Real-IAD dataset and 8 points on the COCO-AD dataset, outperforming state-of-the-art diffusion-based approaches. Code is available at https://github.com/giddyyupp/VLMDiff. Samet Hicsonmez, Abd El Rahman Shabayek, Djamila Aouada |
WACV | 3 |
| 2026 | Domain adaptation for multi-label image classification: A discriminator-free approachabstractThis paper introduces a discriminator-free adversarial-based approach termed DDA-MLIC for Unsupervised Domain Adaptation (UDA) in the context of Multi-Label Image Classification (MLIC). While recent efforts have explored adversarial-based UDA methods for MLIC, they typically include an additional discriminator subnet. Nevertheless, decoupling the classification and the discrimination tasks may harm their task-specific discriminative power. Herein, we address this challenge by presenting a novel adversarial critic directly derived from the task-specific classifier. Specifically, we employ a two-component Gaussian Mixture Model (GMM) to model both source and target predictions, distinguishing between two distinct clusters. Instead of using the traditional Expectation Maximization (EM) algorithm, our approach utilizes a Deep Neural Network (DNN) to estimate the parameters of each GMM component. Subsequently, the source and target GMM parameters are leveraged to formulate an adversarial loss using the Fréchet distance. The proposed framework is therefore not only fully differentiable but is also cost-effective as it avoids the expensive iterative process usually induced by the standard EM method. The proposed method is evaluated on several multi-label image datasets covering three different types of domain shift. The obtained results demonstrate that DDA-MLIC outperforms existing state-of-the-art methods in terms of precision while requiring a lower number of parameters. The code is made publicly available at github.com/cvi2snt/DDA-MLIC . Inder Pal Singh, Enjie Ghorbel, Anis Kacem 0001, Djamila Aouada |
Expert Syst. Appl. | 4 |
| 2026 | Introduction to the Special Issue on Multimodal Video Understanding and Analysis with Foundation Models
Fan Liu 0008, Hanjia Lyu, Yinwei Wei, Hehe Fan, Djamila Aouada, Jiebo Luo 0001, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Audio-Visual Deepfake Detection With Local Temporal InconsistenciesabstractThis paper proposes an audio-visual deepfake detection approach that aims to capture fine-grained temporal inconsistencies between audio and visual modalities. To achieve this, both architectural and data synthesis strategies are introduced. From an architectural perspective, a temporal distance map, coupled with an attention mechanism, is designed to capture these inconsistencies while minimizing the impact of irrelevant temporal subsequences. Moreover, we explore novel pseudo-fake generation techniques to synthesize local inconsistencies. Our approach is evaluated against state-of-the-art methods using the DFDC and FakeAVCeleb datasets, demonstrating its effectiveness in detecting audio-visual deepfakes. Marcella Astrid, Enjie Ghorbel, Djamila Aouada |
ICASSP | 3 |
| 2025 | CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task SolversabstractWe propose CAD-Assistant, a general-purpose CAD agent for AI-assisted design. Our approach is based on a powerful Vision and Large Language Model (VLLM) as a planner and a tool-augmentation paradigm using CAD-specific tools. CAD-Assistant addresses multimodal user queries by generating actions that are iteratively executed on a Python interpreter equipped with the FreeCAD software, accessed via its Python API. Our framework is able to assess the impact of generated CAD commands on geometry and adapts subsequent actions based on the evolving state of the CAD design. We consider a wide range of CAD-specific tools including a sketch image parameterizer, rendering modules, a 2D cross-section generator, and other specialized routines. CAD-Assistant is evaluated on multiple CAD benchmarks, where it outperforms VLLM baselines and supervised task-specific methods. Beyond existing benchmarks, we qualitatively demonstrate the potential of tool-augmented VLLMs as general-purpose CAD solvers across diverse workflows. Dimitrios Mallis, Ahmet Serdar Karadeniz, Sebastian Cavada, Danila Rukhovich, Niki Maria Foteinopoulou, Kseniya Cherenkova, Anis Kacem 0001, Djamila Aouada |
ICCV | 8 |
| 2025 | Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detectionabstractpeer reviewed Marcella Astrid, Anis Kacem 0001, Enjie Ghorbel, Djamila Aouada |
ICCV | 5 |
| 2025 | CAD-Recode: Reverse Engineering CAD Code From Point CloudsabstractComputer-Aided Design (CAD) models are typically constructed by sequentially drawing parametric sketches and applying CAD operations to obtain a 3D model. The problem of 3D CAD reverse engineering consists of reconstructing the sketch and CAD operation sequences from 3D representations such as point clouds. In this paper, we address this challenge through novel contributions across three levels: CAD sequence representation, network design, and training dataset. In particular, we represent CAD sketch-extrude sequences as Python code. The proposed CAD-Recode translates a point cloud into Python code that, when executed, reconstructs the CAD model. Taking advantage of the exposure of pre-trained Large Language Models (LLMs) to Python code, we leverage a relatively small LLM as a decoder for CAD-Recode and combine it with a lightweight point cloud projector. CAD-Recode is trained on a procedurally generated dataset of one million CAD sequences. CAD-Recode significantly outperforms existing methods across the DeepCAD, Fusion360 and real-world CC3D datasets. Furthermore, we show that our CAD Python code output is interpretable by off-the-shelf LLMs, enabling CAD editing and CAD-specific question answering from point clouds. Danila Rukhovich, Elona Dupont, Dimitrios Mallis, Kseniya Cherenkova, Anis Kacem 0001, Djamila Aouada |
ICCV | 6 |
| 2025 | Multimae Meets Earth Observation: Pre-Training Multi-Modal Multi-Task Masked Autoencoders for Earth Observation TasksabstractMulti-modal data in Earth Observation (EO) presents a huge opportunity for improving transfer learning capabilities when pre-training deep learning models. Unlike prior work that often overlooks multi-modal EO data, recent methods have started to include it, resulting in more effective pre-training strategies. However, existing approaches commonly face challenges in effectively transferring learning to downstream tasks where the structure of available data differs from that used during pre-training. This paper addresses this limitation by exploring a more flexible multi-modal, multi-task pre-training strategy for EO data. Specifically, we adopt a Multi-modal Multi-task Masked Autoencoder (MultiMAE) that we pre-train by reconstructing diverse input modalities, including spectral, elevation, and segmentation data. The pre-trained model demonstrates robust transfer learning capabilities, outperforming state-of-the-art methods on various EO datasets for classification and segmentation tasks. Our approach exhibits significant flexibility, handling diverse input configurations without requiring modality-specific pre-trained models. Code will be available at: https://github.com/josesosajs/multimae-meets-eo. Jose Sosa, Danila Rukhovich, Anis Kacem 0001, Djamila Aouada |
ICIP | 4 |
| 2025 | Uncertainty-Aware Knowledge Distillation for Compact and Efficient 6DoF Pose EstimationabstractCompact and efficient 6DoF object pose estimation is crucial in applications such as robotics, augmented reality, and space autonomous navigation systems, where lightweight models are critical for real-time accurate performance. This paper introduces a novel uncertainty-aware end-to-end Knowledge Distillation (KD) framework focused on keypoint-based 6DoF pose estimation. Keypoints predicted by a large teacher model exhibit varying levels of uncertainty that can be exploited within the distillation process to enhance the accuracy of the student model while ensuring its compactness. To this end, we propose a distillation strategy that aligns the student and teacher predictions by adjusting the knowledge transfer based on the uncertainty associated with each teacher keypoint prediction. Additionally, the proposed KD leverages this uncertainty-aware alignment of keypoints to transfer the knowledge at key locations of their respective feature maps. Experiments on the widely-used LINEMOD benchmark demonstrate the effectiveness of our method, achieving superior 6DoF object pose estimation with lightweight models compared to state-of-the-art approaches. Further validation on the SPEED+ dataset for spacecraft pose estimation highlights the robustness of our approach under diverse 6DoF pose estimation scenarios. Nassim Ali Ousalah, Anis Kacem 0001, Enjie Ghorbel, Emmanuel Koumandakis, Djamila Aouada |
IROS | 5 |
| 2025 | FPG-NAS: FLOPs-Aware Gated Differentiable Neural Architecture Search for Efficient 6DoF Pose EstimationabstractWe introduce FPG-NAS, a FLOPs-aware Gated Differentiable Neural Architecture Search framework for efficient 6DoF object pose estimation. Estimating 3D rotation and translation from a single image has been widely investigated yet remains computationally demanding, limiting applicability in resource-constrained scenarios. FPG-NAS addresses this by proposing a specialized differentiable NAS approach for 6DoF pose estimation, featuring a task-specific search space and a differentiable gating mechanism that enables discrete multi-candidate operator selection, thus improving architectural diversity. Additionally, a FLOPs regularization term ensures a balanced trade-off between accuracy and efficiency. The framework explores a vast search space of approximately 1092possible architectures. Experiments on the LINEMOD and SPEED+ datasets demonstrate that FPG-NAS-derived models outperform previous methods under strict FLOPs constraints. To the best of our knowledge, FPG-NAS is the first differentiable NAS framework specifically designed for 6DoF object pose estimation. Nassim Ali Ousalah, Peyman Rostami, Anis Kacem 0001, Enjie Ghorbel, Emmanuel Koumandakis, Djamila Aouada |
MMSP | 6 |
| 2025 | MiCADangelo: Fine-Grained Reconstruction of Constrained CAD Models from 3D ScansabstractComputer-Aided Design (CAD) plays a foundational role in modern manufacturing and product development, often requiring designers to modify or build upon existing models. Converting 3D scans into parametric CAD representations—a process known as CAD reverse engineering—remains a significant challenge due to the high precision and structural complexity of CAD models. Existing deep learning-based approaches typically fall into two categories: bottom-up, geometry-driven methods, which often fail to produce fully parametric outputs, and top-down strategies, which tend to overlook fine-grained geometric details. Moreover, current methods neglect an essential aspect of CAD modeling: sketch-level constraints. In this work, we introduce a novel approach to CAD reverse engineering inspired by how human designers manually perform the task. Our method leverages multi-plane cross-sections to extract 2D patterns and capture fine parametric details more effectively. It enables the reconstruction of detailed and editable CAD models, outperforming state-of-the-art methods and, for the first time, incorporating sketch constraints directly into the reconstruction process. Ahmet Serdar Karadeniz, Dimitrios Mallis, Danila Rukhovich, Kseniya Cherenkova, Anis Kacem 0001, Djamila Aouada |
NeurIPS | 6 |
| 2025 | Removing Geometric Bias in One-Class Anomaly Detection with Adaptive Feature PerturbationabstractInternational audience Romain Hermary, Vincent Gaudillière, Abd El Rahman Shabayek, Djamila Aouada |
WACV | 4 |
| 2025 | PICASSO: A Feed-Forward Framework for Parametric Inference of CAD Sketches via Rendering Self-SupervisionabstractThis work introduces PICASSO, a framework for the parameterization of 2D CAD sketches from hand-drawn and precise sketch images. PICASSO converts a given CAD sketch image into parametric primitives that can be seamlessly integrated into CAD software. Our framework leverages rendering self-supervision to enable the pre-training of a CAD sketch parameterization network using sketch renderings only, thereby eliminating the need for corresponding CAD parameterization. Thus, we significantly reduce reliance on parameter-level annotations, which are often unavailable, particularly for hand-drawn sketches. The two primary components of PICASSO are (1) a Sketch Parameterization Network (SPN) that predicts a series of parametric primitives from CAD sketch images, and (2) a Sketch Rendering Network (SRN) that renders parametric CAD sketches in a differentiable manner and facilitates the computation of a rendering (image-level) loss for self-supervision. We demonstrate that the proposed PICASSO can achieve reasonable performance even when finetuned with only a small number of parametric CAD sketches. Extensive evaluation on the widely used SketchGraphs [37] and CAD as Language [14] datasets validates the effectiveness of the proposed approach on zero- and few-shot learning scenarios. Ahmet Serdar Karadeniz, Dimitrios Mallis, Nesryne Mejri, Kseniya Cherenkova, Anis Kacem 0001, Djamila Aouada |
WACV | 6 |
| 2025 | Information Theoretic Pruning of Coupled Channels in Deep Neural NetworksabstractVariational channel pruning approaches have obtained impressive results thanks to their stochastic nature, well established foundation in information theory, and the practically appealing structured sparsity pattern they offer. Despite their success in pruning Plain Networks (PlainNets), their application has faced certain limitations in networks with structurally coupled channels such as ResNets. In such scenarios, not only is it required to prune structurally coupled channels together, but it is also necessary to ensure that the whole coupled group is irrelevant before pruning is applied. This is an under-investigated problem as most existing methods are designed without taking these couplings into account. In this paper, we propose a novel approach based on Information Theoretic Pruning of structurally Coupled Channels (ITPCC) in neural networks. IT-PCC allows for learning the probabilistic distribution of coupled channel set importance and prunes the ones with the least relevant information to the task at hand. Experimental results for image classification on CIFAR10, CI-FAR100, and ImageNet datasets show that the proposed method outperforms the state-of-the-art, more significantly at high compression rates. Peyman Rostami, Nilotpal Sinha, Nidhal Eddine Chenni, Anis Kacem 0001, Abd El Rahman Shabayek, Carl Shneider, Djamila Aouada |
WACV | 7 |
| 2024 | Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
Marcella Astrid, Enjie Ghorbel, Djamila Aouada |
BMVC | 3 |
| 2024 | DAVINCI: A Single-Stage Architecture for Constrained CAD Sketch Inference
Ahmet Serdar Karadeniz, Dimitrios Mallis, Nesryne Mejri, Kseniya Cherenkova, Anis Kacem 0001, Djamila Aouada |
BMVC | 6 |
| 2024 | CAD-SIGNet: CAD Language Inference from Point Clouds Using Layer-Wise Sketch Instance Guided AttentionabstractReverse engineering in the realm of Computer-Aided Design (CAD) has been a longstanding aspiration, though not yet entirely realized. Its primary aim is to uncover the CAD process behind a physical object given its 3D scan. We propose CAD-SIGNet, an end-to-end trainable and aetoregressive architecture to recover the design history of a CAD model represented as a sequence of sketch-and-extrusion from an input point cloud. Our model learns CAD visual-language representations by layer-wise crossattention between point cloud and CAD language embedding. In particular, a new Sketch instance Guided Attention (SGA) module is proposed in order to reconstruct the finegrained details of the sketches. Thanks to its auto-regressive nature, CAD-SIGNet not only reconstructs a unique full design history of the corresponding CAD model given an input point cloud but also provides multiple plausible design choices. This allows for an interactive reverse engineering scenario by providing designers with multiple next step choices along with the design process. Extensive experiments on publicly available CAD datasets showcase the effectiveness of our approach against existing baseline models in two settings, namely, full design history recovery and conditional auto-completion from point clouds. Mohammad Sadil Khan, Elona Dupont, Sk Aziz Ali, Kseniya Cherenkova, Anis Kacem 0001, Djamila Aouada |
CVPR | 6 |
| 2024 | LAA-Net: Localized Artifact Attention Network for Quality-Agnostic and Generalizable Deepfake DetectionabstractThis paper introduces a novel approach for high-quality deepfake detection called Localized Artifact Attention Net-work (LAA-Net). Existing methods for high-quality deep-fake detection are mainly based on a supervised binary classifier coupled with an implicit attention mechanism. As a result, they do not generalize well to unseen ma-nipulations. To handle this issue, two main contributions are made. First, an explicit attention mechanism within a multi-task learning framework is proposed. By combining heatmap-based and self-consistency attention strate-gies, LAA-Net is forced to focus on a few small artifact-prone vulnerable regions. Second, an Enhanced Feature Pyramid Network (E-FPN) is proposed as a simple and ef-fective mechanism for spreading discriminative low-level features into the final feature output, with the advantage of limiting redundancy. Experiments performed on sev-eral benchmarks show the superiority of our approach in terms of Area Under the Curve (AUC) and Average Preci-sion (AP). The code is available at https://github.com/10Ring/LAA-Net. Nesryne Mejri, Inder Pal Singh, Polina Kuleshova, Marcella Astrid, Anis Kacem 0001, Enjie Ghorbel, Djamila Aouada |
CVPR | 8 |
| 2024 | TransCAD: A Hierarchical Transformer for CAD Sequence Inference from Point Clouds
Elona Dupont, Kseniya Cherenkova, Dimitrios Mallis, Gleb Gusev, Anis Kacem 0001, Djamila Aouada |
ECCV (61) | 6 |
| 2024 | Statistics-Aware Audio-Visual Deepfake DetectorabstractIn this paper, we propose an enhanced audio-visual deep detection method. Recent methods in audio-visual deepfake detection mostly assess the synchronization between audio and visual features. Although they have shown promising results, they are based on the maximization/minimization of isolated feature distances without considering feature statistics. Moreover, they rely on cumbersome deep learning architectures and are heavily dependent on empirically fixed hyperparameters. Herein, to overcome these limitations, we propose: (1) a statistical feature loss to enhance the discrimination capability of the model, instead of relying solely on feature distances; (2) using the waveform for describing the audio as a replacement of frequency-based representations; (3) a post-processing normalization of the fakeness score; (4) the use of shallower network for reducing the computational complexity. Experiments on the DFDC and FakeAVCeleb datasets demonstrate the relevance of the proposed method. Marcella Astrid, Enjie Ghorbel, Djamila Aouada |
ICIP | 3 |
| 2024 | Facial Region-Based Ensembling for Unsupervised Temporal Deepfake LocalizationabstractThis paper addresses the challenge of temporal deepfake localization. Instead of classifying entire videos as real or fake, the goal is isolating forged frames in untrimmed videos that might be partially manipulated. Recently, few deepfake localization methods have emerged. They are mostly supervised, therefore relying on costly annotations and suffering from a lack of generalization to unseen manipulations. As an alternative, we propose reformulating deepfake localization as an unsupervised time-series anomaly detection problem. Hence, to investigate the relevance of the proposed formulation, recent state-of-the-art techniques in anomaly detection for timeseries are evaluated in the context of deepfake localization. To avoid using large architectures, geometric representations, e.g., facial landmarks, are used as input. Moreover, a facialregion based ensembling strategy is introduced for a better modelling of localized deepfake artifacts. Experiments performed on the ForgeryNet dataset demonstrate the effectiveness of the proposed ensembling method and highlight the suitability of the suggested formulation. Nesryne Mejri, Pavel Chernakov, Polina Kuleshova, Enjie Ghorbel, Djamila Aouada |
ICME | 5 |
| 2024 | SPADES: A Realistic Spacecraft Pose Estimation Dataset using Event SensingabstractIn recent years, there has been a growing demand for improved autonomy for in-orbit operations such as rendezvous, docking, and proximity manoeuvres, leading to increased interest in employing Deep Learning-based Spacecraft Pose Estimation techniques. However, due to limited access to real target datasets, algorithms are often trained using synthetic data and applied in the real domain, resulting in a performance drop due to the domain gap. State-of-the-art approaches employ Domain Adaptation techniques to mitigate this issue. In the search for viable solutions, event sensing has been explored in the past and shown to reduce the domain gap between simulations and real-world scenarios. Event sensors have made significant advancements in hardware and software in recent years. Moreover, the characteristics of the event sensor offer several advantages in space applications compared to RGB sensors. To facilitate further training and evaluation of DL-based models, we introduce a new dataset, SPADES, comprising real event data acquired in a controlled laboratory environment and simulated event data using the same camera intrinsics. Furthermore, we introduce an image-based event representation that performs better than existing representations. In addition, we propose an effective data filtering method to improve the quality of training data, thus enhancing model performance. A multifaceted baseline evaluation was conducted using different event representations, event filtering strategies, and algorithmic frameworks, and the results are summarized. The dataset will be made available at http://cvi2.uni.lu/spades. Arunkumar Rathinam, Haytam Qadadri, Djamila Aouada |
ICRA | 3 |
| 2024 | SpelsNet: Surface Primitive Elements Segmentation by B-Rep Graph Structure SupervisionabstractWithin the realm of Computer-Aided Design (CAD), Boundary-Representation (B-Rep) is the standard option for modeling shapes. We present SpelsNet, a neural architecture for the segmentation of 3D point clouds into surface primitive elements under topological supervision of its B-Rep graph structure. We also propose a point-to-BRep adjacency representation that allows for adapting conventional Linear Algebraic Representation of B-Rep graph structure to the point cloud domain. Thanks to this representation, SpelsNet learns from both spatial and topological domains to enable accurate and topologically consistent surface primitive element segmentation. In particular, SpelsNet is composed of two main components; (1) a supervised 3D spatial segmentation head that outputs B-Rep element types and memberships; (2) a graph-based head that leverages the proposed topological supervision. To enable the learning of SpelsNet with the proposed point-to-BRep adjacency supervision, we extend two existing CAD datasets with the required annotations, and conduct a thorough experimental validation on them. The obtained results showcase the efficacy of SpelsNet and its topological supervision compared to a set of baselines and state-of-the-art approaches. Kseniya Cherenkova, Elona Dupont, Anis Kacem 0001, Gleb Gusev, Djamila Aouada |
NeurIPS | 5 |
| 2024 | A Hitchhiker's Guide to Fine-Grained Face Forgery Detection Using Common Sense ReasoningabstractExplainability in artificial intelligence is crucial for restoring trust, particularly in areas like face forgery detection, where viewers often struggle to distinguish between real and fabricated content. Vision and Large Language Models (VLLM) bridge computer vision and natural language, offering numerous applications driven by strong common-sense reasoning. Despite their success in various tasks, the potential of vision and language remains underexplored in face forgery detection, where they hold promise for enhancing explainability by leveraging the intrinsic reasoning capabilities of language to analyse fine-grained manipulation areas. For that reason, few works have recently started to frame the problem of deepfake detection as a Visual Question Answering (VQA) task, nevertheless omitting the realistic and informative open-ended multi-label setting. With the rapid advances in the field of VLLM, an exponential rise of investigations in that direction is expected. As such, there is a need for a clear experimental methodology that converts face forgery detection to a Visual Question Answering (VQA) task to systematically and fairly evaluate different VLLM architectures. Previous evaluation studies in deepfake detection have mostly focused on the simpler binary task, overlooking evaluation protocols for multi-label fine-grained detection and text-generative models. We propose a multi-staged approach that diverges from the traditional binary evaluation protocol and conducts a comprehensive evaluation study to compare the capabilities of several VLLMs in this context. In the first stage, we assess the models' performance on the binary task and their sensitivity to given instructions using several prompts. In the second stage, we delve deeper into fine-grained detection by identifying areas of manipulation in a multiple-choice VQA setting. In the third stage, we convert the fine-grained detection to an open-ended question and compare several matching strategies for the multi-label classification task. Finally, we qualitatively evaluate the fine-grained responses of the VLLMs included in the benchmark. We apply our benchmark to several popular models, providing a detailed comparison of binary, multiple-choice, and open-ended VQA evaluation across seven datasets. \url{https://nickyfot.github.io/hitchhickersguide.github.io/} Niki Maria Foteinopoulou, Enjie Ghorbel, Djamila Aouada |
NeurIPS | 3 |
| 2024 | Self-Supervised Learning for Place Representation Generalization across Appearance ChangesabstractVisual place recognition is a key to unlocking spatial navigation for animals, humans and robots. While state-of-the-art approaches are trained in a supervised manner and therefore hardly capture the information needed for generalizing to unusual conditions, we argue that self-supervised learning may help abstracting the place representation so that it can be foreseen, irrespective of the conditions. More precisely, in this paper, we investigate learning features that are robust to appearance modifications while sensitive to geometric transformations in a self-supervised manner. This dual-purpose training is made possible by combining the two self-supervision main paradigms, i.e. contrastive and predictive learning. Our results on standard benchmarks reveal that jointly learning such appearance-robust and geometry-sensitive image descriptors leads to competitive visual place recognition results across adverse seasonal and illumination conditions, without requiring any humanannotated labels.1. Mohamed Adel Musallam, Vincent Gaudillière, Djamila Aouada |
WACV | 3 |
| 2024 | Discriminator-free Unsupervised Domain Adaptation for Multi-label Image ClassificationabstractIn this paper, a discriminator-free adversarial-based Unsupervised Domain Adaptation (UDA) for Multi-Label Image Classification (MLIC) referred to as DDA-MLIC is proposed. Recently, some attempts have been made for introducing adversarial-based UDA methods in the context of MLIC. However, these methods, which rely on an additional discriminator subnet present one major shortcoming. The learning of domain-invariant features may harm their task-specific discriminative power, since the classification and discrimination tasks are decoupled. Herein, we propose to overcome this issue by introducing a novel adversarial critic that is directly deduced from the task-specific classifier. Specifically, a two-component Gaussian Mixture Model (GMM) is fitted on the source and target predictions in order to distinguish between two clusters. This allows extracting a Gaussian distribution for each component. The resulting Gaussian distributions are then used for formulating an adversarial loss based on a Fréchet distance. The proposed method is evaluated on several multi-label image datasets covering three different types of domain shift. The obtained results demonstrate that DDA-MLIC outperforms existing state-of-the-art methods in terms of precision while requiring a lower number of parameters. The code is publicly available at github.com/cvi2snt/DDA-MLIC.1 Inder Pal Singh, Enjie Ghorbel, Anis Kacem 0001, Arunkumar Rathinam, Djamila Aouada |
WACV | 5 |
| 2024 | Hardware Aware Evolutionary Neural Architecture Search using Representation Similarity MetricabstractHardware-aware Neural Architecture Search (HW-NAS) is a technique used to automatically design the architecture of a neural network for a specific task and target hardware. However, evaluating the performance of candidate architectures is a key challenge in HW-NAS, as it requires significant computational resources. To address this challenge, we propose an efficient hardware-aware evolution-based NAS approach called HW-EvRSNAS. Our approach re-frames the neural architecture search problem as finding an architecture with performance similar to that of a reference model for a target hardware, while adhering to a cost constraint for that hardware. This is achieved through a representation similarity metric known as Representation Mutual Information (RMI) employed as a proxy performance evaluator. It measures the mutual information between the hidden layer representations of a reference model and those of sampled architectures using a single training batch. We also use a penalty term that penalizes the search process in proportion to how far an architecture’s hardware cost is from the desired hardware cost threshold. This resulted in a significantly reduced search time compared to the literature that reached up to 8000× speedups resulting in lower CO2emissions. The proposed approach is evaluated on two different search spaces while using lower computational resources. Furthermore, our approach is thoroughly examined on six different edge devices under various hardware cost constraints. Nilotpal Sinha, Abd El Rahman Shabayek, Anis Kacem 0001, Peyman Rostami, Carl Shneider, Djamila Aouada |
WACV | 6 |
| 2024 | Multi-label image classification using adaptive graph convolutional networks: From a single domain to multiple domainsabstractThis paper proposes an adaptive graph-based approach for multi-label image classification. Graph-based methods have been largely exploited in the field of multi-label classification, given their ability to model label correlations. Specifically, their effectiveness has been proven not only when considering a single domain but also when taking into account multiple domains. However, the topology of the used graph is not optimal as it is pre-defined heuristically. In addition, consecutive Graph Convolutional Network (GCN) aggregations tend to destroy the feature similarity. To overcome these issues, an architecture for learning the graph connectivity in an end-to-end fashion is introduced. This is done by integrating an attention-based mechanism and a similarity-preserving strategy. The proposed framework is then extended to multiple domains using an adversarial training scheme. Numerous experiments are reported on well-known single-domain and multi-domain benchmarks. The results demonstrate that our approach achieves competitive results in terms of mean Average Precision (mAP) and model size as compared to the state-of-the-art. The code will be made publicly available. Inder Pal Singh, Enjie Ghorbel, Oyebade K. Oyedotun, Djamila Aouada |
Comput. Vis. Image Underst. | 4 |
| 2024 | Unsupervised anomaly detection in time-series: An extensive evaluation and analysis of state-of-the-art methodsabstractpeer reviewed Nesryne Mejri, Laura Lopez-Fuentes, Kankana Roy, Pavel Chernakov, Enjie Ghorbel, Djamila Aouada |
Expert Syst. Appl. | 6 |
| 2024 | DermSynth3D: Synthesis of in-the-wild annotated dermatology images
Ashish Sinha, Jeremy Kawahara, Arezou Pakzad, Kumar Abhishek 0001, Matthieu Ruthven, Enjie Ghorbel, Anis Kacem 0001, Djamila Aouada, Ghassan Hamarneh |
Medical Image Anal. | 8 |
| 2024 | Exploiting autoencoder's weakness to generate pseudo anomalies
Marcella Astrid, Muhammad Zaigham Zaheer, Djamila Aouada, Seung-Ik Lee |
Neural Comput. Appl. | 3 |
| 2023 | UNTAG: Learning Generic Features for Unsupervised Type-Agnostic Deepfake DetectionabstractThis paper introduces a novel framework for unsupervised type-agnostic deepfake detection called UNTAG. Existing methods are generally trained in a supervised manner at the classification level, focusing on detecting at most two types of forgeries; thus, limiting their generalization capability across different deepfake types. To handle that, we reformulate the deepfake detection problem as a one-class classification supported by a self-supervision mechanism. Our intuition is that by estimating the distribution of real data in a discriminative feature space, deepfakes can be detected as outliers regardless of their type. UNTAG involves two sequential steps. First, deep representations are learned based on a self-supervised pretext task focusing on manipulated regions. Second, a oneclass classifier fitted on authentic image embeddings is used to detect deepfakes. The results reported on several datasets show the effectiveness of UNTAG and the relevance of the proposed new paradigm. The code is publicly available. Nesryne Mejri, Enjie Ghorbel, Djamila Aouada |
ICASSP | 3 |
| 2023 | 3D-Aware Object Localization using Gaussian Implicit Occupancy FunctionabstractTo automatically localize a target object in an image is crucial for many computer vision applications. To represent the 2D object, ellipse labels have recently been identified as a promising alternative to axis-aligned bounding boxes. This paper further considers 3D-aware ellipse labels, i.e., ellipses which are projections of a 3D ellipsoidal approximation of the object, for 2D target localization. Indeed, projected ellipses carry more geometric information about the object geometry and pose (3D awareness) than traditional 3D-agnostic bounding box labels. Moreover, such a generic 3D ellipsoidal model allows for approximating known to coarsely known targets. We then propose to have a new look at ellipse regression and replace the discontinuous geometric ellipse parameters with the parameters of an implicit Gaussian distribution encoding object occupancy in the image. The models are trained to regress the values of this bivariate Gaussian distribution over the image pixels using a statistical loss function. We introduce a novel non-trainable differentiable layer, E-DSNT, to extract the distribution parameters. Also, we describe how to readily generate consistent 3D-aware Gaussian occupancy parameters using only coarse dimensions of the target and relative pose labels. We extend three existing spacecraft pose estimation datasets with 3D-aware Gaussian occupancy labels to validate our hypothesis. Labels and source code are publicly accessible here: https://cvi2.uni.lu/3d-aware-obj-loc/. Vincent Gaudillière, Leo Pauly, Arunkumar Rathinam, Albert García Sanchez, Mohamed Adel Musallam, Djamila Aouada |
IROS | 6 |
| 2023 | Multi-Label Deepfake ClassificationabstractIn this paper, we investigate the suitability of current multi-label classification approaches for deepfake detection. With the recent advances in generative modeling, new deepfake detection methods have been proposed. Nevertheless, they mostly formulate this topic as a binary classification problem, resulting in poor explainability capabilities. Indeed, a forged image might be induced by multi-step manipulations with different properties. For a better interpretability of the results, recognizing the nature of these stacked manipulations is highly relevant. For that reason, we propose to model deepfake detection as a multi-label classification task, where each label corresponds to a specific kind of manipulation. In this context, state-of-the-art multi-label image classification methods are considered. Extensive experiments are performed to assess the practical use case of deepfake detection. Inder Pal Singh, Nesryne Mejri, Enjie Ghorbel, Djamila Aouada |
MMSP | 5 |
| 2023 | A new perspective for understanding generalization gap of deep neural networks trained with large batch sizes
Oyebade K. Oyedotun, Konstantinos Papadopoulos 0002, Djamila Aouada |
Appl. Intell. | 3 |
| 2023 | CalcGraph: taming the high costs of deep learning using modelsabstractAbstract Models based on differential programming, like deep neural networks, are well established in research and able to outperform manually coded counterparts in many applications. Today, there is a rising interest to introduce this flexible modeling to solve real-world problems. A major challenge when moving from research to application is the strict constraints on computational resources(memory and time). It is difficult to determine and contain the resource requirements of differential models, especially during the early training and hyperparameter exploration stages. In this article, we address this challenge by introducingCalcGraph, a model abstraction of differentiable programming layers.CalcGraph allows to model the computational resources that should be used and thenCalcGraph’s model interpreter can automatically schedule the execution respecting the specifications made. We propose a novel way to efficiently switch models from storage to preallocated memory zones and vice versa to maximize the number of model executions given the available resources. We demonstrate the efficiency of our approach by showing that it consumes less resources than state-of-the-art frameworks like TensorFlow and PyTorch for single-model and multi-model execution. Joe Lorentz, Thomas Hartmann 0001, Assaad Moawad, François Fouquet, Djamila Aouada, Yves Le Traon |
Softw. Syst. Model. | 5 |
| 2023 | Why Is Everyone Training Very Deep Neural Network With Skip Connections?abstractRecent deep neural networks (DNNs) with several layers of feature representations rely on some form of skip connections to simultaneously circumnavigate optimization problems and improve generalization performance. However, the operations of these models are still not clearly understood, especially in comparison to DNNs without skip connections referred to as plain networks (PlainNets) that are absolutely untrainable beyond some depth. As such, the exposition of this article is the theoretical analysis of the role of skip connections in training very DNNs using concepts from linear algebra and random matrix theory. In comparison with PlainNets, the results of our investigation directly unravel the following: 1) why DNNs with skip connections are easier to optimize and 2) why DNNs with skip connections exhibit improved generalization. Our investigation results concretely show that the hidden representations of PlainNets progressively suffer from information loss via singularity problems with depth increase, thus making their optimization difficult. In contrast, as model depth increases, the hidden representations of DNNs with skip connections circumnavigate singularity problems to retain full information that reflects in improved optimization and generalization. For theoretical analysis, this article studies in relation to PlainNets two popular skip connection-based DNNs that are residual networks (ResNets) and residual network with aggregated features (ResNeXt). Oyebade K. Oyedotun, Kassem Al Ismaeil, Djamila Aouada |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | CADOps-Net: Jointly Learning CAD Operation Types and Steps from Boundary-Representationsabstract3D reverse engineering is a long sought-after, yet not completely achieved goal in the Computer-Aided Design (CAD) industry. The objective is to recover the construction history of a CAD model. Starting from a Boundary Representation (B-Rep) of a CAD model, this paper proposes a new deep neural network, CADOps-Net, that jointly learns the CAD operation types and the decomposition into different CAD operation steps. This joint learning allows to divide a B-Rep into parts that were created by various types of CAD operations at the same construction step; therefore providing relevant information for further recovery of the design history. Furthermore, we propose the novel CC3D-Ops dataset that includes over 37k CAD models annotated with CAD operation type labels and step labels. Compared to existing datasets, the complexity and variety of CC3D-Ops models are closer to those used for industrial purposes. Our experiments, conducted on the proposed CC3D-Ops and the publicly available Fusion360 datasets, demonstrate the competitive performance of CADOps-Net with respect to state-of-the-art, and confirm the importance of the joint learning of CAD operation types and steps. Elona Dupont, Kseniya Cherenkova, Anis Kacem 0001, Sk Aziz Ali, Ilya Arzhannikov, Gleb Gusev, Djamila Aouada |
3DV | 7 |
| 2022 | Leveraging Equivariant Features for Absolute Pose RegressionabstractWhile end-to-end approaches have achieved state-of-the-art performance in many perception tasks, they are not yet able to compete with 3D geometry-based methods in pose estimation. Moreover, absolute pose regression has been shown to be more related to image retrieval. As a result, we hypothesize that the statistical features learned by classical Convolutional Neural Networks do not carry enough geometric information to reliably solve this inherently geometric task. In this paper, we demonstrate how a translation and rotation equivariant Convolutional Neural Network directly induces representations of camera motions into the feature space. We then show that this geometric property allows for implicitly augmenting the training data under a whole group of image plane-preserving transformations. Therefore, we argue that directly learning equivariant features is preferable than learning data-intensive intermediate representations. Comprehensive experimental validation demonstrates that our lightweight model outperforms existing ones on standard datasets.1 Mohamed Adel Musallam, Vincent Gaudillière, Miguel Ortiz del Castillo, Kassem Al Ismaeil, Djamila Aouada |
CVPR | 5 |
| 2022 | A Closer Look at Autoencoders for Unsupervised Anomaly DetectionabstractUnsupervised anomaly detection is a challenging problem, where the aim is to detect irregular data instances. Interestingly, generative models can learn data distribution, and thus have been proposed for anomaly detection. In this direction, the variational autoencoder (VAE) is popular, as it enforces an explicit probabilistic interpretation of the latent space. We note that there are other generative autoencoders (AEs) such as the denoising AE (DAE) and contractive AE (CAE), which also model data generation process without enforcing an explicit probabilistic latent space interpretation. While it is intuitively straightforward to see the benefit of a latent space with explicit probabilistic interpretation for generative tasks, it is unclear how this can be crucial for anomaly detection problems. Consequently, our exposition in this paper is to investigate the extent to which different latent space attributes of AEs impact their performances for anomaly detection tasks. We take the conventional and deterministic AE that we refer to as plain AE (PAE) as the baseline for performance comparison. Our results obtained using five different datasets reveal that an explicit probabilistic latent space is not necessary for good performance. The best results on most of the datasets are obtained using CAE, which enjoys stable latent representations. Oyebade K. Oyedotun, Djamila Aouada |
ICASSP | 2 |
| 2022 | Multi Label Image Classification using Adaptive Graph Convolutional Networks (ML-AGCN)abstractIn this paper, a novel graph-based approach for multi-label image classification called Multi-Label Adaptive Graph Convolutional Network (ML-AGCN) is introduced. Graph-based methods have shown great potential in the field of multi-label classification. However, these approaches heuristically fix the graph topology for modeling label dependencies, which might be not optimal. To handle that, we propose to learn the topology in an end-to-end manner. Specifically, we incorporate an attention-based mechanism for estimating the pairwise importance between graph nodes and a similarity-based mechanism for conserving the feature similarity between different nodes. This offers a more flexible way for adaptively modeling the graph. Experimental results are reported on two well-known datasets, namely, MS-COCO and VG-500. Results show that ML-AGCN outperforms state-of-the-art methods while reducing the number of model parameters. Inder Pal Singh, Enjie Ghorbel, Oyebade K. Oyedotun, Djamila Aouada |
ICIP | 4 |
| 2021 | Explaining Defect Detection with Saliency Maps
Joe Lorentz, Thomas Hartmann 0001, Assaad Moawad, François Fouquet, Djamila Aouada |
IEA/AIE (2) | 5 |
| 2021 | Leveraging High-Frequency Components for Deepfake DetectionabstractIn the past years, RGB-based deepfake detection has shown notable progress thanks to the development of effective deep neural networks. However, the performance of deepfake detectors remains primarily dependent on the quality of the forged content and the level of artifacts introduced by the forgery method. To detect these artifacts, it is often necessary to separate and analyze the frequency components of an image. In this context, we propose to utilize the high-frequency components of color images by introducing an end-to-end trainable module that (a) extracts features from high-frequency components and (b) fuses them with the features of the RGB input. The module not only exploits the high-frequency anomalies present in manipulated images but also can be used with most RGB-based deepfake detectors. Experimental results show that the proposed approach boosts the performance of state-of-the-art networks, such as XceptionNet and EfficientNet, on a challenging deepfake dataset. Nesryne Mejri, Konstantinos Papadopoulos 0002, Djamila Aouada |
MMSP | 3 |
| 2021 | Deep network compression with teacher latent subspace learning and LASSO
Oyebade K. Oyedotun, Abd El Rahman Shabayek, Djamila Aouada, Björn Ottersten 0001 |
Appl. Intell. | 3 |
| 2021 | Training very deep neural networks: Rethinking the role of skip connections
Oyebade K. Oyedotun, Kassem Al Ismaeil, Djamila Aouada |
Neurocomputing | 3 |
| 2020 | DeepVI: A Novel Framework for Learning Deep View-Invariant Human Action Representations using a Single RGB CameraabstractIn this paper, we address the problem of cross-view action recognition from a monocular RGB camera. This topic has been considered extremely challenging due to the lack of 3D information in 2D images. Exploiting the advances in 3D pose estimation from a single RGB camera, we propose a new framework termed DeepVI, for cross-view action recognition without the need for pose alignment. Virtual viewpoints are used to augment the variability of training data along with the use of an end-to-end Deep Neural Network (DNN). The proposed network is composed of two modules. The first one, called SmoothNet, implicitly smooths skeleton joint trajectories using revisited temporal convolution in order to reduce the noise in the estimated 3D skeletons. The second module consists of a state-of-the-art approach designed for action recognition based on Spatial Temporal Graph Convolutional Networks (ST-GCN [40]). Experiments have been conducted in cross-view settings on two datasets, namely, NTU RGB-D and Northwestern-UCLA. The obtained results show the effectiveness of the proposed framework. Konstantinos Papadopoulos 0002, Enjie Ghorbel, Oyebade K. Oyedotun, Djamila Aouada, Björn Ottersten 0001 |
FG | 4 |
| 2020 | 3d Deformation Signature for Dynamic Face RecognitionabstractThis work proposes a novel 3D Deformation Signature (3DS) to represent a 3D deformation signal for 3D Dynamic Face Recognition. 3DS is computed given a non-linear 6D-space representation which guarantees physically plausible 3D deformations. A unique deformation indicator is computed per triangle in a triangulated mesh as a ratio derived from scale and in-plane deformation in the canonical space. These indicators, concatenated, construct the 3DS for each temporal instance. There is a pressing need of non-intrusive bio-metric measurements in domains like surveillance and security. By construction, 3DS is a non-intrusive facial measurement that is resistant to common security attacks like presentation, template and adversarial attacks. Two dynamic datasets (BU4DFE and COMA) were examined, in a standard classification framework, to evaluate 3DS. A first rank recognition accuracy of 99.9%, that outperforms existing literature, was achieved. Assuming an open-world setting, 99.97% accuracy was attained in detecting unseen distractors. Abd El Rahman Shabayek, Djamila Aouada, Kseniya Cherenkova, Gleb Gusev, Björn Ottersten 0001 |
ICASSP | 2 |
| 2020 | Pvdeconv: Point-Voxel Deconvolution for Autoencoding CAD Construction in 3DabstractWe propose a Point-Voxel DeConvolution (PVDeConv) module for 3D data autoencoder. To demonstrate its efficiency we learn to synthesize high-resolution point clouds of 10k points that densely describe the underlying geometry of Computer Aided Design (CAD) models. Scanning artifacts, such as protrusions, missing parts, smoothed edges and holes, inevitably appear in real 3D scans of fabricated CAD objects. Learning the original CAD model construction from a 3D scan requires a ground truth to be available together with the corresponding 3D scan of an object. To solve the gap, we introduce a new dedicated dataset, the CC3D, containing 50k+ pairs of CAD models and their corresponding 3D meshes. This dataset is used to learn a convolutional autoencoder for point clouds sampled from the pairs of 3D scans - CAD models. The challenges of this new dataset are demonstrated in comparison with other generative point cloud sampling models trained on ShapeNet. The CC3D autoencoder is efficient with respect to memory consumption and training time as compared to stateof-the-art models for 3D data generation. Kseniya Cherenkova, Djamila Aouada, Gleb Gusev |
ICIP | 2 |
| 2020 | Going Deeper With Neural Networks Without Skip ConnectionsabstractWe propose the training of very deep neural networks (DNNs) without shortcut connections known as PlainNets. Training such networks is a notoriously hard problem due to: (1) the relatively popular challenge of vanishing and exploding activations, and (2) the less studied ‘near singularity’ problem. We argue that if the aforementioned problems are tackled together, the training of deeper PlainNets becomes easier. Subsequently, we propose the training of very deep PlainNets by leveraging Leaky Rectified Linear Units (LReLUs), parameter constraint and strategic parameter initialization. Our approach is simple and allows to successfully train very deep PlainNets having up to 100 layers without employing shortcut connections. To validate this approach, we validate on five challenging datasets; namely, MNIST, CIFAR-10, CIFAR100, SVHN and ImageNet datasets. We report the best results known on the ImageNet dataset using a PlainNet with top-1 and top-5 error rates of 24.1% and 7.3%, respectively. Oyebade K. Oyedotun, Abd El Rahman Shabayek, Djamila Aouada, Björn Ottersten 0001 |
ICIP | 3 |
| 2020 | 3D Sparse Deformation Signature for Dynamic Face RecognitionabstractThis paper proposes a novel compact and memory efficient Sparse 3D Deformation Signature (S3DS) to represent a sparse 3D deformation signal for 3D Dynamic Face Recognition. S3DS is based on a non-linear 6D-space representation that secures physically plausible 3D deformations. A unique deformation indicator is computed per triangle in a triangulated mesh, thanks to a recent 3D Deformation Signature (3DS) that is based on Lie Bodies. The proposed S3DS sparsely concatenates unique triangular indicators to construct the facial signature for each temporal instance. The novel descriptor shall benefit domains like surveillance and security in providing non-intrusive bio-metric measurements. By construction, S3DS is resistant to common security attacks like presentation, template and adversarial attacks. Two dynamic datasets (BU4DFE and COMA) were examined in various sparse concatenation settings. Using high reduction rates of $\approx 500$, a first rank recognition accuracy similar to the state of the art was achieved. At low reduction rates of $\approx 40$, S3DS outperformed most existing literature on BU4DFE achieving 99.92%. On COMA, it achieved 99.93% which outperforms existing literature. In an open-world experimental setup, using thousands of distractors, the accuracy reached up to 100% in detecting unseen distractors with high reduction rates in the 3D facial descriptor size. Abd El Rahman Shabayek, Djamila Aouada, Kseniya Cherenkova, Gleb Gusev |
ICIP | 2 |
| 2020 | Why Do Deep Neural Networks with Skip Connections and Concatenated Hidden Representations Work?
Oyebade K. Oyedotun, Djamila Aouada |
ICONIP (3) | 2 |
| 2020 | Vertex Feature Encoding and Hierarchical Temporal Modeling in a Spatio-Temporal Graph Convolutional Network for Action RecognitionabstractSpatio-temporal Graph Convolutional Networks (ST-GCNs) have shown great performance in the context of skeleton-based action recognition. Nevertheless, ST-GCNs use raw skeleton data as vertex features. Such features have low dimensionality and might not be optimal for action discrimination. Moreover, a single layer of temporal convolution is used to model short-term temporal dependencies but can be insufficient for capturing both long-term. In this paper, we extend the Spatio-Temporal Graph Convolutional Network for skeleton-based action recognition by introducing two novel modules, namely, the Graph Vertex Feature Encoder (GVFE) and the Dilated Hierarchical Temporal Convolutional Network (DH-TCN). On the one hand, the GVFE module learns appropriate vertex features for action recognition by encoding raw skeleton data into a new feature space. On the other hand, the DH-TCN module is capable of capturing both short-term and long-term temporal dependencies using a hierarchical dilated convolutional network. Experiments have been conducted on the challenging NTU RGB-D 60, NTU RGB-D 120 and Kinetics datasets. The obtained results show that our method competes with state-of-the-art approaches while using a smaller number of layers and parameters; thus reducing the required training time and memory. Konstantinos Papadopoulos 0002, Enjie Ghorbel, Djamila Aouada, Björn Ottersten 0001 |
ICPR | 3 |
| 2020 | Revisiting the Training of Very Deep Neural Networks without Skip ConnectionsabstractDeep neural networks (DNNs) with many layers of feature representations yield state-of-the-art results on several difficult learning tasks. However, optimizing very deep DNNs without shortcut connections known as PlainNets, is a notoriously hard problem. Considering the growing interest in this area, this paper investigates holistically two scenarios that plague the training of very deep PlainNets: (1) the relatively popular challenge of `vanishing and exploding units' activations', and (2) the less investigated `singularity' problem, which is studied in details in this paper. In contrast to earlier works that study only the saturation and explosion of units' activations in isolation, this paper harmonizes the inconspicuous coexistence of the aforementioned problems for very deep PlainNets. Particularly, we argue that the aforementioned problems would have to be tackled simultaneously for the successful training of very deep PlainNets. Finally, different techniques that can be employed for tackling the optimization problem are discussed, and a specific combination of simple techniques that allows the successful training of PlainNets having up to 100 layers is demonstrated. Oyebade K. Oyedotun, Abd El Rahman Shabayek, Djamila Aouada, Björn Ottersten 0001 |
ICPR | 3 |
| 2020 | Structured Compression of Deep Neural Networks with Debiased Elastic Group LASSOabstractState-of-the-art Deep Neural Networks (DNNs) are typically too cumbersome to be practically useful in portable electronic devices. As such, several works pursue model compression that seeks to drastically reduce computational memory footprints, FLOPS and memory for storage. Many of these works achieve unstructured compression, where the compressed models are not directly useful since dedicated hardware and specialized algorithms are required for storage of sparse weights and fast sparse matrix-vector multi-plication respectively. In this paper, we propose structured compression of large DNNs using debiased elastic group LASSO (DEGL), which is motivated by different interesting characteristics of the individual components. That is, where group LASSO penalty enforces structured sparsity, l2-norm penalty promotes features grouping, and debiasing disentangles sparsity and shrinkage effects of group LASSO. We perform extensive experiments by applying DEGL to different DNN architectures including LeNet, VGG, AlexNet and ResNet on MNIST, CIFAR-10, CIFAR-100 and ImageNet datasets. Furthermore, we validate the effectiveness of our proposal on domain adaptation using Oxford-102 flower species and Food-5K datasets. Results show that DEGL can compress DNNs by several folds with small or no loss of performance. Particularly, DEGL outperforms conventional group LASSO and several other state-of-the-art methods that perform structured compression. Oyebade K. Oyedotun, Djamila Aouada, Björn Ottersten 0001 |
WACV | 2 |
| 2020 | Fast Adaptive Reparametrization (FAR) With Application to Human Action RecognitionabstractIn this letter, a fast approach for curve reparametrization, called Fast Adaptive Reparamterization (FAR), is introduced. Instead of computing an optimal matching between two curves such as Dynamic Time Warping (DTW) and elastic distance-based approaches, our method is applied to each curve independently, leading to linear computational complexity. It is based on a simple replacement of the curve parameter by a variable invariant under specific variations of reparametrization. The choice of this variable is heuristically made according to the application of interest. In addition to being fast, the proposed reparametrization can be applied not only to curves observed in Euclidean spaces but also to feature curves living in Riemannian spaces. To validate our approach, we apply it to the scenario of human action recognition using curves living in the Riemannian product Special Euclidean space$\mathbb {SE}(3)^n$. The obtained results on three benchmarks for human action recognition (MSRAction3D, Florence3D, and UTKinect) show that our approach competes with state-of-the-art methods in terms of accuracy and computational cost. Enjie Ghorbel, Girum G. Demisse, Djamila Aouada, Björn Ottersten 0001 |
IEEE Signal Process. Lett. | 3 |
| 2019 | Two-Stage RGB-Based Action Detection Using Augmented 3D Poses
Konstantinos Papadopoulos 0002, Enjie Ghorbel, Renato Baptista, Djamila Aouada, Björn Ottersten 0001 |
CAIP (1) | 4 |
| 2019 | View-invariant Action Recognition from RGB Data via 3D Pose EstimationabstractIn this paper, we propose a novel view-invariant action recognition method using a single monocular RGB camera. View-invariance remains a very challenging topic in 2D action recognition due to the lack of 3D information in RGB images. Most successful approaches make use of the concept of knowledge transfer by projecting 3D synthetic data to multiple viewpoints. Instead of relying on knowledge transfer, we propose to augment the RGB data by a third dimension by means of 3D skeleton estimation from 2D images using a CNN-based pose estimator. In order to ensure view-invariance, a pre-processing for alignment is applied followed by data expansion as a way for denoising. Finally, a Long-Short Term Memory (LSTM) architecture is used to model the temporal dependency between skeletons. The proposed network is trained to directly recognize actions from aligned 3D skeletons. The experiments performed on the challenging Northwestern-UCLA dataset show the superiority of our approach as compared to state-of-the-art ones. Renato Baptista, Enjie Ghorbel, Konstantinos Papadopoulos 0002, Girum G. Demisse, Djamila Aouada, Björn Ottersten 0001 |
ICASSP | 5 |
| 2019 | Learning to Fuse Latent Representations for Multimodal DataabstractMultimodal learning leverages data from different modalities to improve the performance of a trained model. Typically, latent representations extracted from multimodal data are provided via direct feature fusion for end-to-end training of a deep neural network towards a specific task. However, the informativeness of the different data modalities can easily vary across a collected dataset. As such, naively or directly fusing the latent representations obtained for one modality and the other, as is commonly done in state-of-the-art works, may burden the model in finding concise representations that are indeed useful for learning. In this paper, we propose to instead learn the fusion of latent representations for multimodal data by using a modality gating mechanism that allows the dynamic weighting of extracted latent representations based on their informativness. Extensive experiments using the BU-3DFE dataset for facial expression recognition and the Washington object classification multimodal RGB-D dataset show that learning the fusion of the latent representations for different data modalities leads to improved model generalization than the conventional naive fusion method. Oyebade K. Oyedotun, Djamila Aouada, Björn Ottersten 0001 |
ICASSP | 2 |
| 2019 | Bodyfitr: Robust Automatic 3D Human Body FittingabstractThis paper proposes BODYFITR, a fully automatic method to fit a human body model to static 3D scans with complex poses. Automatic and reliable 3D human body fitting is necessary for many applications related to healthcare, digital ergonomics, avatar creation and security, especially in industrial contexts for large-scale product design. Existing works either make prior assumptions on the pose, require manual annotation of the data or have difficulty handling complex poses. This work addresses these limitations by providing a novel automatic fitting pipeline with carefully integrated building blocks designed for a systematic and robust approach. It is validated on the 3DBodyTex dataset, with hundreds of high-quality 3D body scans, and shown to outperform prior works in static body pose and shape estimation, qualitatively and quantitatively. The method is also applied to the creation of realistic 3D avatars from the high-quality texture scans of 3DBodyTex, further demonstrating its capabilities. Alexandre Saint 0001, Abd El Rahman Shabayek, Kseniya Cherenkova, Gleb Gusev, Djamila Aouada, Björn Ottersten 0001 |
ICIP | 5 |
| 2018 | 3DBodyTex: Textured 3D Body DatasetabstractIn this paper, a dataset, named 3DBodyTex, of static 3D body scans with high-quality texture information is presented along with a fully automatic method for body model fitting to a 3D scan. 3D shape modelling is a fundamental area of computer vision that has a wide range of applications in the industry. It is becoming even more important as 3D sensing technologies are entering consumer devices such as smartphones. As the main output of these sensors is the 3D shape, many methods rely on this information alone. The 3D shape information is, however, very high dimensional and leads to models that must handle many degrees of freedom from limited information. Coupling texture and 3D shape alleviates this burden, as the texture of 3D objects is complementary to their shape. Unfortunately, high-quality texture content is lacking from commonly available datasets, and in particular in datasets of 3D body scans. The proposed 3DBodyTex dataset aims to fill this gap with hundreds of high-quality 3D body scans with high-resolution texture. Moreover, a novel fully automatic pipeline to fit a body model to a 3D scan is proposed. It includes a robust 3D landmark estimator that takes advantage of the high-resolution texture of 3DBodyTex. The pipeline is applied to the scans, and the results are reported and discussed, showcasing the diversity of the features in the dataset. Alexandre Saint 0001, Eman Ahmed, Abd El Rahman Shabayek, Kseniya Cherenkova, Gleb Gusev, Djamila Aouada, Björn Ottersten 0001 |
3DV | 6 |
| 2018 | Improving the Capacity of Very Deep Networks with Maxout UnitsabstractDeep neural networks inherently have large representational power for approximating complex target functions. However, models based on rectified linear units can suffer reduction in representation capacity due to dead units. Moreover, approximating very deep networks trained with dropout at test time can be more inexact due to the several layers of nonlinearities. To address the aforementioned problems, we propose to learn the activation functions of hidden units for very deep networks via maxout. However, maxout units increase the model parameters, and therefore model may suffer from overfitting; we alleviate this problem by employing elastic net regularization. In this paper, we propose very deep networks with maxout units and elastic net regularization and show that the features learned are quite linearly separable. We perform extensive experiments and reach state-of-the-art results on the USPS and MNIST datasets. Particularly, we reach an error rate of 2.19% on the USPS dataset, surpassing the human performance error rate of 2.5% and all previously reported results, including those that employed training data augmentation. On the MNIST dataset, we reach an error rate of 0.36% which is competitive with the state-of-the-art results. Oyebade K. Oyedotun, Abd El Rahman Shabayek, Djamila Aouada, Björn Ottersten 0001 |
ICASSP | 3 |
| 2018 | A Revisit of Action Detection Using Improved TrajectoriesabstractIn this paper, we revisit trajectory-based action detection in a potent and non-uniform way. Improved trajectories have been proven to be an effective model for motion description in action recognition. In temporal action localization, however, this approach is not efficiently exploited. Trajectory features extracted from uniform video segments result in significant performance degradation due to two reasons: (a) during uniform segmentation, a significant amount of noise is often added to the main action and (b) partial actions can have negative impact in classifier's performance. Since uniform video segmentation seems to be insufficient for this task, we propose a two-step supervised non-uniform segmentation, performed in an online manner. Action proposals are generated using either 2D or 3D data, therefore action classification can be directly performed on them using the standard improved trajectories approach. We experimentally compare our method with other approaches and we show improved performance on a challenging online action detection dataset. Konstantinos Papadopoulos 0002, Michel Antunes, Djamila Aouada, Björn Ottersten 0001 |
ICASSP | 3 |
| 2018 | Deformation Based Curved Shape RepresentationabstractIn this paper, we introduce a deformation based representation space for curved shapes in . Given an ordered set of points sampled from a curved shape, the proposed method represents the set as an element of a finite dimensional matrix Lie group. Variation due to scale and location are filtered in a preprocessing stage, while shapes that vary only in rotation are identified by an equivalence relationship. The use of a finite dimensional matrix Lie group leads to a similarity metric with an explicit geodesic solution. Subsequently, we discuss some of the properties of the metric and its relationship with a deformation by least action. Furthermore, invariance to reparametrization or estimation of point correspondence between shapes is formulated as an estimation of sampling function. Thereafter, two possible approaches are presented to solve the point correspondence estimation problem. Finally, we propose an adaptation of k-means clustering for shape analysis in the proposed representation space. Experimental results show that the proposed representation is robust to uninformative cues, e.g., local shape perturbation and displacement. In comparison to state of the art methods, it achieves a high precision on the Swedish and the Flavia leaf datasets and a comparable result on MPEG-7, Kimia99 and Kimia216 datasets. Girum G. Demisse, Djamila Aouada, Björn Ottersten 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Full 3D Reconstruction of Non-Rigidly Deforming ObjectsabstractIn this article, we discuss enhanced full 360° 3D reconstruction of dynamic scenes containing non-rigidly deforming objects using data acquired from commodity depth or 3D cameras. Several approaches for enhanced and full 3D reconstruction of non-rigid objects have been proposed in the literature. These approaches suffer from several limitations due to requirement of a template, inability to tackle large local deformations and topology changes, inability to tackle highly noisy and low-resolution data, and inability to produce online results. We target online and template-free enhancement of the quality of noisy and low-resolution full 3D reconstructions of dynamic non-rigid objects. For this purpose, we propose a view-independent recursive and dynamic multi-frame 3D super-resolution scheme for noise removal and resolution enhancement of 3D measurements. The proposed scheme tracks the position and motion of each 3D point at every timestep by making use of the current acquisition and the result of the previous iteration. The effects of system blur due to per-point tracking are subsequently tackled by introducing a novel and efficient multi-level 3D bilateral total variation regularization. These characteristics enable the proposed scheme to handle large deformations and topology changes accurately. A thorough evaluation of the proposed scheme on both real and simulated data is carried out. The results show that the proposed scheme improves upon the performance of the state-of-the-art methods and is able to accurately enhance the quality of low-resolution and highly noisy 3D reconstructions while being robust to large local deformations. Hassan Afzal, Djamila Aouada, Bruno Mirbach, Björn Ottersten 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Deformation-Based 3D Facial Expression RepresentationabstractWe propose a deformation-based representation for analyzing expressions from three-dimensional (3D) faces. A point cloud of a 3D face is decomposed into an ordered deformable set of curves that start from a fixed point. Subsequently, a mapping function is defined to identify the set of curves with an element of a high-dimensional matrix Lie group, specifically the direct product of SE(3). Representing 3D faces as an element of a high-dimensional Lie group has two main advantages. First, using the group structure, facial expressions can be decoupled from a neutral face. Second, an underlying non-linear facial expression manifold can be captured with the Lie group and mapped to a linear space, Lie algebra of the group. This opens up the possibility of classifying facial expressions with linear models without compromising the underlying manifold. Alternatively, linear combinations of linearised facial expressions can be mapped back from the Lie algebra to the Lie group. The approach is tested on the Binghamton University 3D Facial Expression (BU-3DFE) and the Bosphorus datasets. The results show that the proposed approach performed comparably, on the BU-3DFE dataset, without using features or extensive landmark points. Girum G. Demisse, Djamila Aouada, Björn Ottersten 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | Unsupervised Vanishing Point Detection and Camera Calibration from a Single Manhattan Image with Radial DistortionabstractThe article concerns the automatic calibration of a camera with radial distortion from a single image. It is known that, under the mild assumption of square pixels and zero skew, lines in the scene project into circles in the image, and three lines suffice to calibrate the camera up to an ambiguity between focal length and radial distortion. The calibration results highly depend on accurate circle estimation, which is hard to accomplish because lines tend to project into short circular arcs. To overcome this problem, we show that, given a short circular arc edge, it is possible to robustly determine a line that goes through the center of the corresponding circle. These lines, henceforth called Lines of Circle Centres (LCCs), are used in a new method that detects sets of parallel lines and estimates the calibration parameters, including the center and amount of distortion, focal length, and camera orientation with respect to the Manhattan frame. Extensive experiments in both semi-synthetic and real images show that our algorithm outperforms state-of-the-art approaches in unsupervised calibration from a single image, while providing more information. Michel Antunes, João Pedro Barreto 0001, Djamila Aouada, Björn Ottersten 0001 |
CVPR | 3 |
| 2017 | Enhanced trajectory-based action recognition using human poseabstractAction recognition using dense trajectories is a popular concept. However, many spatio-temporal characteristics of the trajectories are lost in the final video representation when using a single Bag-of-Words model. Also, there is a significant amount of extracted trajectory features that are actually irrelevant to the activity being analyzed, which can considerably degrade the recognition performance. In this paper, we propose a human-tailored trajectory extraction scheme, in which trajectories are clustered using information from the human pose. Two configurations are considered; first, when exact skeleton joint positions are provided, and second, when only an estimate thereof is available. In both cases, the proposed method is further strengthened by using the concept of local Bag-of-Words, where a specific codebook is generated for each skeleton joint group. This has the advantage of adding spatial human pose awareness in the video representation, effectively increasing its discriminative power. We experimentally compare the proposed method with the standard dense trajectories approach on two challenging datasets. Konstantinos Papadopoulos 0002, Michel Antunes, Djamila Aouada, Björn Ottersten 0001 |
ICIP | 3 |
| 2017 | Deformation transfer of 3D human shapes and poses on manifoldsabstractIn this paper, we introduce a novel method to transfer the deformation of a human body to another directly on a manifold. There exists a rich literature on transferring deformations based on Euclidean representations. However, a 3D human shape and pose live on a manifold and have a Riemannian structure. The proposed method uses the Lie Bodies manifold representation of 3D triangulated bodies. Its benefits are preserved, namely, minimum required degrees of freedom for any triangle deformation and no heuristics to constrain excessive ones. We give a closed form solution for deformation transfer directly on the Lie Bodies. The deformations have strictly positive determinants ensuring that non-physical deformations are removed. We show examples on three datasets, and highlight differences with the Euclidean deformation transfer. Abd El Rahman Shabayek, Djamila Aouada, Alexandre Saint 0001, Björn Ottersten 0001 |
ICIP | 2 |
| 2017 | Training Very Deep Networks via Residual Learning with Stochastic Input Shortcut Connections
Oyebade K. Oyedotun, Abd El Rahman Shabayek, Djamila Aouada, Björn Ottersten 0001 |
ICONIP (2) | 3 |
| 2017 | Real-Time Enhancement of Dynamic Depth Videos with Non-Rigid DeformationsabstractWe propose a novel approach for enhancing depth videos containing non-rigidly deforming objects. Depth sensors are capable of capturing depth maps in real-time but suffer from high noise levels and low spatial resolutions. While solutions for reconstructing 3D details in static scenes, or scenes with rigid global motions have been recently proposed, handling unconstrained non-rigid deformations in relative complex scenes remains a challenge. Our solution consists in a recursive dynamic multi-frame super-resolution algorithm where the relative local 3D motions between consecutive frames are directly accounted for. We rely on the assumption that these 3D motions can be decoupled into lateral motions and radial displacements. This allows to perform a simple local per-pixel tracking where both depth measurements and deformations are dynamically optimized. The geometric smoothness is subsequently added using a multi-level$L_1$minimization with a bilateral total variation regularization. The performance of this method is thoroughly evaluated on both real and synthetic data. As compared to alternative approaches, the results show a clear improvement in reconstruction accuracy and in robustness to noise, to relative large non-rigid deformations, and to topological changes. Moreover, the proposed approach, implemented on a CPU, is shown to be computationally efficient and working in real-time. Kassem Al Ismaeil, Djamila Aouada, Thomas Solignac, Bruno Mirbach, Björn Ottersten 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Similarity Metric for Curved Shapes in Euclidean SpaceabstractIn this paper, we introduce a similarity metric for curved shapes that can be described, distinctively, by ordered points. The proposed method represents a given curve as a point in the deformation space, the direct product of rigid transformation matrices, such that the successive action of the matrices on a fixed starting point reconstructs the full curve. In general, both open and closed curves are represented in the deformation space modulo shape orientation and orientation preserving diffeomorphisms. The use of direct product Lie groups to represent curved shapes led to an explicit formula for geodesic curves and the formulation of a similarity metric between shapes by the L2-norm on the Lie algebra. Additionally, invariance to reparametrization or estimation of point correspondence between shapes is performed as an intermediate step for computing geodesics. Furthermore, since there is no computation of differential quantities on the curves, our representation is more robust to local perturbations and needs no pre-smoothing. We compare our method with the elastic shape metric defined through the square root velocity (SRV) mapping, and other shape matching approaches. Girum G. Demisse, Djamila Aouada, Björn Ottersten 0001 |
CVPR | 2 |
| 2016 | A revisit to human action recognition from depth sequences: Guided SVM-sampling for joint selectionabstractThis paper revisits the problem of human action recognition from skeleton joint locations, and analyses the tradeoff of sampling the joint space with respect to the recognition performance and computational complexity. The provided insights led to the design of a new algorithm for automatically selecting the most appropriate set of joints for each action. During the training stage, the approach applies a guided joint sampling strategy for learning different SVM classifiers, selecting the classifier that maximizes confidence and ambiguity metrics. Experimental results on three action datasets show that pre-selecting the most varying skeleton joints for each action dramatically reduces the computational complexity while keeping competitive recognition rates. Michel Antunes, Djamila Aouada, Björn Ottersten 0001 |
WACV | 2 |
| 2016 | Enhancement of dynamic depth scenes by upsampling for precise super-resolution (UP-SR)
Kassem Al Ismaeil, Djamila Aouada, Bruno Mirbach, Björn Ottersten 0001 |
Comput. Vis. Image Underst. | 2 |
| 2016 | Feature engineering strategies for credit card fraud detection
Alejandro Correa 0003, Djamila Aouada, Aleksandar Stojanovic 0002, Björn Ottersten 0001 |
Expert Syst. Appl. | 2 |
| 2015 | View-Independent Enhanced 3D Reconstruction of Non-rigidly Deforming Objects
Hassan Afzal, Djamila Aouada, François Destelle, Bruno Mirbach, Björn Ottersten 0001 |
CAIP (2) | 2 |
| 2015 | Template-based statistical shape modelling on deformation spaceabstractA statistical model for shapes in R2or R3is proposed. Shape modelling is a difficult problem mainly due to the non-linear nature of its space. Our approach considers curves as shape contours, and models their deformations with respect to a de-formable template shape. Contours are uniformly sampled into a discrete sequence of points. Hence, the deformation of a shape is formulated as an action of transformation matrices on each of these points. A parametrized stochastic model based on Markov process is proposed to model shape variability in the deformation space. The model's parameters are estimated from a labeled training dataset. Moreover, a similarity metric based on the Mahalanobis distance is proposed. Subsequently, the model has been successfully tested for shape recognition, synthesis, and retrieval. Girum G. Demisse, Djamila Aouada, Björn Ottersten 0001 |
ICIP | 2 |
| 2015 | Detecting Credit Card Fraud Using Periodic FeaturesabstractWhen constructing a credit card fraud detection model, it is very important to extract the right features from transactional data. This is usually done by aggregating the transactions in order to observe the spending behavioral patterns of the customers. In this paper we propose to create a new set of features based on analyzing the periodic behavior of the time of a transaction using the von Mises distribution. Using a real credit card fraud dataset provided by a large European card processing company, we compare state-of-the-art credit card fraud detection models, and evaluate how the different sets of features have an impact on the results. By including the proposed periodic features into the methods, the results show an average increase in savings of 13%. The aforementioned card processing company is currently incorporating the methodology proposed in this paper into their fraud detection system. Alejandro Correa 0003, Djamila Aouada, Aleksandar Stojanovic 0002, Björn Ottersten 0001 |
ICMLA | 2 |
| 2015 | Example-dependent cost-sensitive decision trees
Alejandro Correa 0003, Djamila Aouada, Björn Ottersten 0001 |
Expert Syst. Appl. | 2 |
| 2015 | Unified multi-lateral filter for real-time depth map enhancement
Frederic Garcia, Djamila Aouada, Bruno Mirbach, Thomas Solignac, Björn Ottersten 0001 |
Image Vis. Comput. | 2 |
| 2014 | Surface UP-SR for an improved face recognition using low resolution depth camerasabstractWe address the limitation of low resolution depth cameras in the context of face recognition. Considering a face as a surface in 3-D, we reformulate the recently proposed Upsampling for Precise Super-Resolution algorithm as a new approach on three dimensional points. This reformulation allows an efficient implementation, and leads to a largely enhanced 3-D face reconstruction. Moreover, combined with a dedicated face detection and representation pipeline, the proposed method provides an improved face recognition system using low resolution depth cameras. We show experimentally that this system increases the face recognition rate as compared to directly using the low resolution raw data.1 Djamila Aouada, Kassem Al Ismaeil, Kedija Kedir Idris, Björn Ottersten 0001 |
AVSS | 1 |
| 2014 | SPN2: Single-sided privacy preserving nearest neighbor and its application to face recognitionabstractWe address the privacy concerns that raise when running a nearest neighbor (NN) search on confidential data in a surveillance system composed of a client and a server. The proposed privacy preserving NN search uses Boneh-Goh-Nissim encryption to hide both the query data captured by the client and the database records stored in the server. As opposed to state-of-the-art approaches which rely on a large number of interactions, this encryption enables the client to fully outsource the NN computation to the server; hence, ensuring a single-sided private computation, and resulting in a one-round protocol between the server and the client. We analyze the practical feasibility of this algorithm on a face recognition problem. We formally prove and experimentally show that the resulting system maintains the recognition rate while fully preserving the privacy of both the database and the acquired faces1. Djamila Aouada, Dalia Khader |
AVSS | 1 |
| 2014 | Example-Dependent Cost-Sensitive Logistic Regression for Credit ScoringabstractEveral real-world classification problems are example-dependent cost-sensitive in nature, where the costs due to misclassification vary between examples. Credit scoring is a typical example of cost-sensitive classification. However, it is usually treated using methods that do not take into account the real financial costs associated with the lending business. In this paper, we propose a new example-dependent cost matrix for credit scoring. Furthermore, we propose an algorithm that introduces the example-dependent costs into a logistic regression. Using two publicly available datasets, we compare our proposed method against state-of-the-art example-dependent cost-sensitive algorithms. The results highlight the importance of using real financial costs. Moreover, by using the proposed cost-sensitive logistic regression, significant improvements are made in the sense of higher savings. Alejandro Correa 0003, Djamila Aouada, Björn Ottersten 0001 |
ICMLA | 2 |
| 2014 | RGB-D Multi-view System Calibration for Full 3D Scene ReconstructionabstractOne of the most crucial requirements for building a multi-view system is the estimation of relative poses of all cameras. An approach tailored for a RGB-D cameras based multi-view system is missing. We propose BAICP+ which combines Bundle Adjustment (BA) and Iterative Closest Point (ICP) algorithms to take into account both 2D visual and 3D shape information in one minimization formulation to estimate relative pose parameters of each camera. BAICP+ is generic enough to take different types of visual features into account and can be easily adapted to varying quality of 2D and 3D data. We perform experiments on real and simulated data. Results show that with the right weighting factor BAICP+ has an optimal performance when compared to BA and ICP used independently or sequentially. Hassan Afzal, Djamila Aouada, David Font, Bruno Mirbach, Björn Ottersten 0001 |
ICPR | 2 |
| 2014 | Improving Credit Card Fraud Detection with Calibrated ProbabilitiesabstractPrevious analysis has shown that applying Bayes minimum risk to detect credit card fraud leads to better results measured by monetary savings, as compared with traditional methodologies. Nevertheless, this approach requires good probability estimates that not only separate well between positive and negative examples, but also assess the real probability of the event. Unfortunately, not all classification algorithms satisfy this restriction. In this paper, two different methods for calibrating probabilities are evaluated and analyzed in the context of credit card fraud detection, with the objective of finding the model that minimizes the real losses due to fraud. Even though under-sampling is often used in the context of classification with unbalanced datasets, it is shown that when probabilistic models are used to make decisions based on minimizing risk, using the full dataset provides significantly better results. In order to test the algorithms, a real dataset provided by a large European card processing company is used. It is shown that by calibrating the probabilities and then using Bayes minimum Risk the losses due to fraud are reduced. Furthermore, because of the good overall results, the aforementioned card processing company is currently incorporating the methodology proposed in this paper into their fraud detection system. Finally, the methodology has been tested on a different application, namely, direct marketing. Alejandro Correa 0003, Aleksandar Stojanovic 0002, Djamila Aouada, Björn Ottersten 0001 |
SDM | 3 |
| 2013 | Depth Super-Resolution by Enhanced Shift and Add
Kassem Al Ismaeil, Djamila Aouada, Bruno Mirbach, Björn Ottersten 0001 |
CAIP (2) | 2 |
| 2013 | Dynamic super resolution of depth sequences with non-rigid motionsabstractWe enhance the resolution of depth videos acquired with low resolution time-of-flight cameras. To that end, we propose a new dedicated dynamic super-resolution that is capable to accurately super-resolve a depth sequence containing one or multiple moving objects without strong constraints on their shape or motion, thus clearly outperforming any existing super-resolution techniques that perform poorly on depth data and are either restricted to global motions or not precise because of an implicit estimation of motion. The proposed approach is based on a new data model that leads to a robust registration of all depth frames after a dense upsampling. The textureless nature of depth images allows to robustly handle sequences with multiple moving objects as confirmed by our experiments. Kassem Al Ismaeil, Djamila Aouada, Bruno Mirbach, Björn Ottersten 0001 |
ICIP | 2 |
| 2013 | Cost Sensitive Credit Card Fraud Detection Using Bayes Minimum RiskabstractCredit card fraud is a growing problem that affects card holders around the world. Fraud detection has been an interesting topic in machine learning. Nevertheless, current state of the art credit card fraud detection algorithms miss to include the real costs of credit card fraud as a measure to evaluate algorithms. In this paper a new comparison measure that realistically represents the monetary gains and losses due to fraud detection is proposed. Moreover, using the proposed cost measure a cost sensitive method based on Bayes minimum risk is presented. This method is compared with state of the art algorithms and shows improvements up to 23% measured by cost. The results of this paper are based on real life transactional data provided by a large European card processing company. Alejandro Correa 0003, Aleksandar Stojanovic 0002, Djamila Aouada, Björn Ottersten 0001 |
ICMLA (1) | 3 |
| 2013 | Real-time depth enhancement by fusion for RGB-D camerasabstractThis study presents a real‐time refinement procedure for depth data acquired by RGB‐D cameras. Data from RGB‐D cameras suffer from undesired artefacts such as edge inaccuracies or holes owing to occlusions or low object remission. In this work, the authors use recent depth enhancement filters intended for time‐of‐flight cameras, and extend them to structured light‐based depth cameras, such as the Kinect camera. Thus, given a depth map and its corresponding two‐dimensional image, we correct the depth measurements by separately treating its undesired regions. To that end, the authors propose specific confidence maps to tackle areas in the scene that require a special treatment. Furthermore, in the case of filtering artefacts, the authors introduce the use of RGB images as guidance images as an alternative to real‐time state‐of‐the‐art fusion filters that use greyscale guidance images. The experimental results show that the proposed fusion filter provides dense depth maps with corrected erroneous or invalid depth measurements and adjusted depth edges. In addition, the authors propose a mathematical formulation that enables to use the filter in real‐time applications. Frederic Garcia, Djamila Aouada, Thomas Solignac, Bruno Mirbach, Björn Ottersten 0001 |
IET Comput. Vis. | 2 |
| 2012 | Spatio-temporal ToF data enhancement by fusionabstractWe propose an extension of our previous work on spatial domain Time-of-Flight (ToF) data enhancement to the temporal domain. Our goal is to generate enhanced depth maps at the same frame rate of the 2-D camera that, coupled with a ToF camera, constitutes a hybrid ToF multi-camera rig. To that end, we first estimate the motion between consecutive 2-D frames, and then use it to predict their corresponding depth maps. The enhanced depth maps result from the fusion between the recorded 2-D frames and the predicted depth maps by using our previous contribution on ToF data enhancement. The experimental results show that the proposed approach overcomes the ToF camera drawbacks; namely, low resolution in space and time and high level of noise within depth measurements, providing enhanced depth maps at video frame rate. Frederic Garcia, Djamila Aouada, Bruno Mirbach, Björn Ottersten 0001 |
ICIP | 2 |
| 2012 | Bilateral filter evaluation based on exponential kernels
Kassem Al Ismaeil, Djamila Aouada, Björn Ottersten 0001 |
ICPR | 2 |
| 2011 | A new multi-lateral filter for real-time depth enhancementabstractWe present an adaptive multi-lateral filter for real-time low-resolution depth map enhancement. Despite the great advantages of Time-of-Flight cameras in 3-D sensing, there are two main drawbacks that restricts their use in a wide range of applications; namely, their fairly low spatial resolution, compared to other 3-D sensing systems, and the high noise level within the depth measurements. We therefore propose a new data fusion method based upon a bilateral filter. The proposed filter is an extension the pixel weighted average strategy for depth sensor data fusion. It includes a new factor that allows to adaptively consider 2-D data or 3-D data as guidance information. Consequently, unwanted artefacts such as texture copying get almost entirely eliminated, outperforming alternative depth enhancement filters. In addition, our algorithm can be effectively and efficiently implemented for real-time applications. Frederic Garcia, Djamila Aouada, Bruno Mirbach, Thomas Solignac, Björn Ottersten 0001 |
AVSS | 2 |
| 2011 | Spiral colour model: Reduction from 3-D to 2-DabstractWe define a new reduced model to represent coloured images. We propose to use two components for a full definition of a colour instead of three. To that end we take advantage of the geometrical structure of the HCL conical colour space and approximate its circular base by a spiral. We thus write chroma as a function of hue. The resulting spiral is therefore defined by one parameter only. This parameter is then combined with luminance in order to represent all the colour information. Our experiments show that our proposed model ensures an accurate representation of coloured digital images. Further more, it preserves the perceptual properties of the original HCL representation. Frederic Garcia, Djamila Aouada, Bruno Mirbach, Björn Ottersten 0001 |
ICASSP | 2 |
| 2010 | Mahalanobis-based Adaptive Nonlinear Dimension ReductionabstractWe define a new adaptive embedding approach for data dimension reduction applications. Our technique entails a local learning of the manifold of the initial data, with the objective of defining local distance metrics that take into account the different correlations between the data points. We choose to illustrate the properties of our work on the isomap algorithm. We show through multiple simulations that the new adaptive version of isomap is more robust to noise than the original non-adaptive one. Djamila Aouada, Yuliy M. Baryshnikov, Hamid Krim |
ICPR | 1 |
| 2010 | Squigraphs for Fine and Compact Modeling of 3-D ShapesabstractWe propose to superpose global topological and local geometric 3-D shape descriptors in order to define one compact and discriminative representation for a 3-D object. While a number of available 3-D shape modeling techniques yield satisfactory object classification rates, there is still a need for a refined and efficient identification/recognition of objects among the same class. In this paper, we use Morse theory in a two-phase approach. To ensure the invariance of the final representation to isometric transforms, we choose the Morse function to be a simple and intrinsic global geodesic function defined on the surface of a 3-D object. The first phase is a coarse representation through a reduced topological Reeb graph. We use it for a meaningful decomposition of shapes into primitives. During the second phase, we add detailed geometric information by tracking the evolution of Morse function's level curves along each primitive. We then embed the manifold of these curves into [Formula: see text], and obtain a single curve. By combining phase one and two, we build new graphs rich in topological and geometric information that we refer to as squigraphs. Our experiments show that squigraphs are more general than existing techniques. They achieve similar classification rates to those achieved by classical shape descriptors. Their performance, however, becomes clearly superior when finer classification and identification operations are targeted. Indeed, while other techniques see their performances dropping, squigraphs maintain a performance rate of the order of 97%. Djamila Aouada, Hamid Krim |
IEEE Trans. Image Process. | 1 |
| 2009 | Novel similarity invariant for space curves using turning angles and its application to object recognitionabstractWe present a new similarity invariant signature for space curves. This signature is based on the information contained in the turning angles of both the tangent and the binormal vectors at each point on the curve. For an accurate comparison of these signatures, we define a Riemannian metric on the space of the invariant. We show through relevant examples that, unlike classical invariants, the one we define in this paper enjoys multiple important properties at the same time, namely, a high discrimination level, independence of any reference point, uniqueness property, as well as a good preservation of the correspondence between curves. Moreover, we illustrate how to match 3D objects by extracting and comparing the invariant signatures of their curved skeletons. Djamila Aouada, Hamid Krim |
ICASSP | 1 |
| 2009 | Meaningful 3D shape partitioning using Morse functionsabstractTo simplify the matching and recognition of 3D objects, we propose to decompose a complex 3D shape into simpler primitive parts. Our partitioning of objects relies on their topological Reeb graphs. Taking advantage of the properties of Morse theory, we detect the critical points of the global geodesic function. These points define the levels at which the segmentation happens. To preserve the geometry of objects, we choose to use level curves instead of intervals. To proceed with object matching, we propose a kernel-based technique to register Reeb graphs. This optimal positioning of two Reeb graphs prepares for a pairwise comparison of the geometry of their primitives. Djamila Aouada, Hamid Krim |
ICIP | 1 |
| 2007 | Statistical Analysis of the Global Geodesic Function for 3D Object ClassificationabstractThis paper presents a novel classification strategy for 3D objects. Our technique is based on using a global geodesic function to intrinsically describe the surface of an object. The choice of the global geodesic function ensures the invariance of the classification procedure to scaling and all isometric transformations. Using the Jensen-Shannon divergence, feature parameters are extracted from the probability distribution functions of the global geodesic function for each one of the classes. These parameters are used in the decision of a class membership of an object. This approach demonstrates low computational cost, efficiency, and robustness to resolution over many different data sets. Djamila Aouada, Hamid Krim |
ICASSP (1) | 1 |
| 2007 | 3D Mixed Invariant and its Application on Object ClassificationabstractA new integro-differential invariant for curves in 3D transformed by affine group action is presented in this paper. The derivatives involved are of the first order, and therefore this invariant is significantly less sensitive to noise than classical affine differential invariants, the simplest of which involves derivatives of order 5. A classification procedure based on characteristic curves of an object surface is considered using our proposed mixed invariants. Substantiating examples are provided to verify efficiency and discriminant power of the characteristic spatial curve based 3D object classification. Djamila Aouada, Hamid Krim, Irina A. Kogan |
ICASSP (1) | 2 |