Abd El Rahman Shabayek

dblp:69/10637 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-8730-3765ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021
YearPublicationVenuePosition
2026 ASTER: Latent Pseudo-Anomaly Generation for Unsupervised Time-Series Anomaly Detection
Romain Hermary, Samet Hicsonmez, Dan Pineau, Abd El Rahman Shabayek, Djamila Aouada
ICPR (3)4
2026 Training Free Zero-Shot Image Anomaly Localisation via Diffusion Inversion
Samet Hicsonmez, Abd El Rahman Shabayek, Djamila Aouada
ICPR (1)2
2026 VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion
abstract
Detecting visual anomalies in diverse, multi-class real-world images is a significant challenge. We introduce VLMDiff, a novel unsupervised multi-class visual anomaly detection framework. It integrates a Latent Diffusion Model (LDM) with a Vision-Language Model (VLM) for enhanced anomaly localization and detection. Specifically, a pretrained VLM with a simple prompt extracts detailed image descriptions, serving as additional conditioning for LDM training. Current diffusion-based methods rely on synthetic noise generation, limiting their generalization and requiring per-class model training, which hinders scalability. VLMDiff, however, leverages VLMs to obtain normal captions without manual annotations or additional training. These descriptions condition the diffusion model, learning a robust normal image feature representation for multiclass anomaly detection. Our method achieves competitive performance, improving the pixel-level Per-Region-Overlap (PRO) metric by up to 25 points on the Real-IAD dataset and 8 points on the COCO-AD dataset, outperforming state-of-the-art diffusion-based approaches. Code is available at https://github.com/giddyyupp/VLMDiff.
Samet Hicsonmez, Abd El Rahman Shabayek, Djamila Aouada
WACV2
2025 Removing Geometric Bias in One-Class Anomaly Detection with Adaptive Feature Perturbation
abstract
International audience
Romain Hermary, Vincent Gaudillière, Abd El Rahman Shabayek, Djamila Aouada
WACV3
2025 Information Theoretic Pruning of Coupled Channels in Deep Neural Networks
abstract
Variational channel pruning approaches have obtained impressive results thanks to their stochastic nature, well established foundation in information theory, and the practically appealing structured sparsity pattern they offer. Despite their success in pruning Plain Networks (PlainNets), their application has faced certain limitations in networks with structurally coupled channels such as ResNets. In such scenarios, not only is it required to prune structurally coupled channels together, but it is also necessary to ensure that the whole coupled group is irrelevant before pruning is applied. This is an under-investigated problem as most existing methods are designed without taking these couplings into account. In this paper, we propose a novel approach based on Information Theoretic Pruning of structurally Coupled Channels (ITPCC) in neural networks. IT-PCC allows for learning the probabilistic distribution of coupled channel set importance and prunes the ones with the least relevant information to the task at hand. Experimental results for image classification on CIFAR10, CI-FAR100, and ImageNet datasets show that the proposed method outperforms the state-of-the-art, more significantly at high compression rates.
Peyman Rostami, Nilotpal Sinha, Nidhal Eddine Chenni, Anis Kacem 0001, Abd El Rahman Shabayek, Carl Shneider, Djamila Aouada
WACV5
2024 Hardware Aware Evolutionary Neural Architecture Search using Representation Similarity Metric
abstract
Hardware-aware Neural Architecture Search (HW-NAS) is a technique used to automatically design the architecture of a neural network for a specific task and target hardware. However, evaluating the performance of candidate architectures is a key challenge in HW-NAS, as it requires significant computational resources. To address this challenge, we propose an efficient hardware-aware evolution-based NAS approach called HW-EvRSNAS. Our approach re-frames the neural architecture search problem as finding an architecture with performance similar to that of a reference model for a target hardware, while adhering to a cost constraint for that hardware. This is achieved through a representation similarity metric known as Representation Mutual Information (RMI) employed as a proxy performance evaluator. It measures the mutual information between the hidden layer representations of a reference model and those of sampled architectures using a single training batch. We also use a penalty term that penalizes the search process in proportion to how far an architecture’s hardware cost is from the desired hardware cost threshold. This resulted in a significantly reduced search time compared to the literature that reached up to 8000× speedups resulting in lower CO2emissions. The proposed approach is evaluated on two different search spaces while using lower computational resources. Furthermore, our approach is thoroughly examined on six different edge devices under various hardware cost constraints.
Nilotpal Sinha, Abd El Rahman Shabayek, Anis Kacem 0001, Peyman Rostami, Carl Shneider, Djamila Aouada
WACV2
2021 Deep network compression with teacher latent subspace learning and LASSO
Oyebade K. Oyedotun, Abd El Rahman Shabayek, Djamila Aouada, Björn Ottersten 0001
Appl. Intell.2
2020 3d Deformation Signature for Dynamic Face Recognition
abstract
This work proposes a novel 3D Deformation Signature (3DS) to represent a 3D deformation signal for 3D Dynamic Face Recognition. 3DS is computed given a non-linear 6D-space representation which guarantees physically plausible 3D deformations. A unique deformation indicator is computed per triangle in a triangulated mesh as a ratio derived from scale and in-plane deformation in the canonical space. These indicators, concatenated, construct the 3DS for each temporal instance. There is a pressing need of non-intrusive bio-metric measurements in domains like surveillance and security. By construction, 3DS is a non-intrusive facial measurement that is resistant to common security attacks like presentation, template and adversarial attacks. Two dynamic datasets (BU4DFE and COMA) were examined, in a standard classification framework, to evaluate 3DS. A first rank recognition accuracy of 99.9%, that outperforms existing literature, was achieved. Assuming an open-world setting, 99.97% accuracy was attained in detecting unseen distractors.
Abd El Rahman Shabayek, Djamila Aouada, Kseniya Cherenkova, Gleb Gusev, Björn Ottersten 0001
ICASSP1
2020 Going Deeper With Neural Networks Without Skip Connections
abstract
We propose the training of very deep neural networks (DNNs) without shortcut connections known as PlainNets. Training such networks is a notoriously hard problem due to: (1) the relatively popular challenge of vanishing and exploding activations, and (2) the less studied ‘near singularity’ problem. We argue that if the aforementioned problems are tackled together, the training of deeper PlainNets becomes easier. Subsequently, we propose the training of very deep PlainNets by leveraging Leaky Rectified Linear Units (LReLUs), parameter constraint and strategic parameter initialization. Our approach is simple and allows to successfully train very deep PlainNets having up to 100 layers without employing shortcut connections. To validate this approach, we validate on five challenging datasets; namely, MNIST, CIFAR-10, CIFAR100, SVHN and ImageNet datasets. We report the best results known on the ImageNet dataset using a PlainNet with top-1 and top-5 error rates of 24.1% and 7.3%, respectively.
Oyebade K. Oyedotun, Abd El Rahman Shabayek, Djamila Aouada, Björn Ottersten 0001
ICIP2
2020 3D Sparse Deformation Signature for Dynamic Face Recognition
abstract
This paper proposes a novel compact and memory efficient Sparse 3D Deformation Signature (S3DS) to represent a sparse 3D deformation signal for 3D Dynamic Face Recognition. S3DS is based on a non-linear 6D-space representation that secures physically plausible 3D deformations. A unique deformation indicator is computed per triangle in a triangulated mesh, thanks to a recent 3D Deformation Signature (3DS) that is based on Lie Bodies. The proposed S3DS sparsely concatenates unique triangular indicators to construct the facial signature for each temporal instance. The novel descriptor shall benefit domains like surveillance and security in providing non-intrusive bio-metric measurements. By construction, S3DS is resistant to common security attacks like presentation, template and adversarial attacks. Two dynamic datasets (BU4DFE and COMA) were examined in various sparse concatenation settings. Using high reduction rates of $\approx 500$, a first rank recognition accuracy similar to the state of the art was achieved. At low reduction rates of $\approx 40$, S3DS outperformed most existing literature on BU4DFE achieving 99.92%. On COMA, it achieved 99.93% which outperforms existing literature. In an open-world experimental setup, using thousands of distractors, the accuracy reached up to 100% in detecting unseen distractors with high reduction rates in the 3D facial descriptor size.
Abd El Rahman Shabayek, Djamila Aouada, Kseniya Cherenkova, Gleb Gusev
ICIP1
2020 Revisiting the Training of Very Deep Neural Networks without Skip Connections
abstract
Deep neural networks (DNNs) with many layers of feature representations yield state-of-the-art results on several difficult learning tasks. However, optimizing very deep DNNs without shortcut connections known as PlainNets, is a notoriously hard problem. Considering the growing interest in this area, this paper investigates holistically two scenarios that plague the training of very deep PlainNets: (1) the relatively popular challenge of `vanishing and exploding units' activations', and (2) the less investigated `singularity' problem, which is studied in details in this paper. In contrast to earlier works that study only the saturation and explosion of units' activations in isolation, this paper harmonizes the inconspicuous coexistence of the aforementioned problems for very deep PlainNets. Particularly, we argue that the aforementioned problems would have to be tackled simultaneously for the successful training of very deep PlainNets. Finally, different techniques that can be employed for tackling the optimization problem are discussed, and a specific combination of simple techniques that allows the successful training of PlainNets having up to 100 layers is demonstrated.
Oyebade K. Oyedotun, Abd El Rahman Shabayek, Djamila Aouada, Björn Ottersten 0001
ICPR2
2019 Bodyfitr: Robust Automatic 3D Human Body Fitting
abstract
This paper proposes BODYFITR, a fully automatic method to fit a human body model to static 3D scans with complex poses. Automatic and reliable 3D human body fitting is necessary for many applications related to healthcare, digital ergonomics, avatar creation and security, especially in industrial contexts for large-scale product design. Existing works either make prior assumptions on the pose, require manual annotation of the data or have difficulty handling complex poses. This work addresses these limitations by providing a novel automatic fitting pipeline with carefully integrated building blocks designed for a systematic and robust approach. It is validated on the 3DBodyTex dataset, with hundreds of high-quality 3D body scans, and shown to outperform prior works in static body pose and shape estimation, qualitatively and quantitatively. The method is also applied to the creation of realistic 3D avatars from the high-quality texture scans of 3DBodyTex, further demonstrating its capabilities.
Alexandre Saint 0001, Abd El Rahman Shabayek, Kseniya Cherenkova, Gleb Gusev, Djamila Aouada, Björn Ottersten 0001
ICIP2
2018 3DBodyTex: Textured 3D Body Dataset
abstract
In this paper, a dataset, named 3DBodyTex, of static 3D body scans with high-quality texture information is presented along with a fully automatic method for body model fitting to a 3D scan. 3D shape modelling is a fundamental area of computer vision that has a wide range of applications in the industry. It is becoming even more important as 3D sensing technologies are entering consumer devices such as smartphones. As the main output of these sensors is the 3D shape, many methods rely on this information alone. The 3D shape information is, however, very high dimensional and leads to models that must handle many degrees of freedom from limited information. Coupling texture and 3D shape alleviates this burden, as the texture of 3D objects is complementary to their shape. Unfortunately, high-quality texture content is lacking from commonly available datasets, and in particular in datasets of 3D body scans. The proposed 3DBodyTex dataset aims to fill this gap with hundreds of high-quality 3D body scans with high-resolution texture. Moreover, a novel fully automatic pipeline to fit a body model to a 3D scan is proposed. It includes a robust 3D landmark estimator that takes advantage of the high-resolution texture of 3DBodyTex. The pipeline is applied to the scans, and the results are reported and discussed, showcasing the diversity of the features in the dataset.
Alexandre Saint 0001, Eman Ahmed, Abd El Rahman Shabayek, Kseniya Cherenkova, Gleb Gusev, Djamila Aouada, Björn Ottersten 0001
3DV3
2018 Improving the Capacity of Very Deep Networks with Maxout Units
abstract
Deep neural networks inherently have large representational power for approximating complex target functions. However, models based on rectified linear units can suffer reduction in representation capacity due to dead units. Moreover, approximating very deep networks trained with dropout at test time can be more inexact due to the several layers of nonlinearities. To address the aforementioned problems, we propose to learn the activation functions of hidden units for very deep networks via maxout. However, maxout units increase the model parameters, and therefore model may suffer from overfitting; we alleviate this problem by employing elastic net regularization. In this paper, we propose very deep networks with maxout units and elastic net regularization and show that the features learned are quite linearly separable. We perform extensive experiments and reach state-of-the-art results on the USPS and MNIST datasets. Particularly, we reach an error rate of 2.19% on the USPS dataset, surpassing the human performance error rate of 2.5% and all previously reported results, including those that employed training data augmentation. On the MNIST dataset, we reach an error rate of 0.36% which is competitive with the state-of-the-art results.
Oyebade K. Oyedotun, Abd El Rahman Shabayek, Djamila Aouada, Björn Ottersten 0001
ICASSP2
2017 Deformation transfer of 3D human shapes and poses on manifolds
abstract
In this paper, we introduce a novel method to transfer the deformation of a human body to another directly on a manifold. There exists a rich literature on transferring deformations based on Euclidean representations. However, a 3D human shape and pose live on a manifold and have a Riemannian structure. The proposed method uses the Lie Bodies manifold representation of 3D triangulated bodies. Its benefits are preserved, namely, minimum required degrees of freedom for any triangle deformation and no heuristics to constrain excessive ones. We give a closed form solution for deformation transfer directly on the Lie Bodies. The deformations have strictly positive determinants ensuring that non-physical deformations are removed. We show examples on three datasets, and highlight differences with the Euclidean deformation transfer.
Abd El Rahman Shabayek, Djamila Aouada, Alexandre Saint 0001, Björn Ottersten 0001
ICIP1
2017 Training Very Deep Networks via Residual Learning with Stochastic Input Shortcut Connections
Oyebade K. Oyedotun, Abd El Rahman Shabayek, Djamila Aouada, Björn Ottersten 0001
ICONIP (2)2