VLDB 2026 Research / reviewers in the wild / expert
Abdenour Hadid
dblp:29/2765
· DBLP profile ↗
87ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0001-9092-735XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 60 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 4 first-author · 14 since 2021Security and privacy · 7Human-computer interaction and ubiquitous computing · 5Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainability-Guided Deepfake Detection for High-Fidelity Facial Edits
Bibek Das, Soumi Chattopadhyay, Chandranath Adak, Astitva Pandey, Ashutosh Parihar, Zahid Akhtar, Soumya Dutta, Abdenour Hadid |
ICPR (4) | 8 |
| 2026 | Diffusion-Latent Invisible Watermarking for Proactive Deepfake Provenance Verification
Bibek Das, Anurag Deo, Chandranath Adak, Soumi Chattopadhyay, Zahid Akhtar, Soumya Dutta, Abdenour Hadid |
ICPR (4) | 7 |
| 2026 | SPARK-IL: Spectral Retrieval-Augmented RAG for Knowledge-Driven Deepfake Detection via Incremental Learning
Hessen Bougueffa Eutamene, Abdellah Zakaria Sellam, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
ICPR (13) | 4 |
| 2026 | Does data augmentation help or hinder the generalization of deepfake video detection?
Bachir Kaddar, Sid Ahmed Fezza, Elhocine Boutellaa, Wassim Hamidouche, Abdenour Hadid |
Multim. Tools Appl. | 5 |
| 2026 | PE-CLIP: A Parameter-Efficient Fine-Tuning of Vision Language Models for Dynamic Facial Expression RecognitionabstractThe emergence of Vision-Language Models (VLMs) like Contrastive Language-Image Pretraining (CLIP) provides appealing solutions to various vision problems including Dynamic Facial Expression Recognition (DFER). However, most of the proposed approaches face major challenges, particularly related to inefficient full fine-tuning of the encoders and the complexity of the models. Moreover, some of the proposed methods seem to struggle with suboptimal performance due to (i) poor alignment between textual and visual representations, and (ii) ineffective temporal modeling. To address these challenges, we propose PE-CLIP, a parameter-efficient fine-tuning (PEFT) framework that elegantly adapts CLIP for dynamic facial expression recognition, requiring significantly reduced number of trainable parameters while maintaining high accuracy. At its core, to enhance efficiency and performance, PE-CLIP introduces two specialized adapters namely a Temporal Dynamic Adapter (TDA) and a Shared Adapter (ShA). The TDA is a GRU-based module with a dynamic scaling mechanism, capturing sequential dependencies while adaptively modulating the contribution of each temporal feature to emphasize the most informative ones while mitigating irrelevant variations. The ShA is a lightweight adapter refine representations within both textual and visual encoders, ensuring consistent feature processing while maintaining parameter efficiency. Additionally, we leverage Multi-modal Prompt Learning (MaPLe), which introduces learnable prompts to both visual and action unit-based textual description inputs, further improving the semantic alignment between modalities and enabling the efficient adaptation of CLIP for dynamic tasks. We evaluate our proposed PE-CLIP on two benchmark datasets, namely DFEW, FERV39K, and AFEW, achieving competitive performance compared to state-of-the-art methods while requiring fewer trainable parameters. By striking an optimal balance between parameter efficiency and performance, PE-CLIP sets a new benchmark in resource-efficient DFER. The source code of the proposed PE-CLIP will be publicly available at https://github.com/Ibtissam-SAADI/PE-CLIP . Ibtissam Saadi, Abdenour Hadid, Douglas W. Cunningham, Abdelmalik Taleb-Ahmed, Yassin Elhillali |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | DeeCLIP: A Robust and Generalizable Transformer-Based Framework for Detecting AI-Generated Images
Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
ACIVS | 5 |
| 2025 | SAViL-Det: Semantic-Aware Vision-Language Model for Multi-script Text Detection
Mohammed-En-nadhir Zighem, Abdenour Hadid |
ACIVS | 2 |
| 2025 | Bi-LORA: A Vision-Language Approach for Synthetic Image DetectionabstractABSTRACT Advancements in deep image synthesis techniques, such as generative adversarial networks (GANs) and diffusion models (DMs), have ushered in an era of generating highly realistic images. While this technological progress has captured significant interest, it has also raised concerns about the high challenge in distinguishing real images from their synthetic counterparts. This paper takes inspiration from the potent convergence capabilities between vision and language, coupled with the zero‐shot nature of vision‐language models (VLMs). We introduce an innovative method called Bi‐LORA that leverages VLMs, combined with low‐rank adaptation (LORA) tuning techniques, to enhance the precision of synthetic image detection for unseen model‐generated images. The pivotal conceptual shift in our methodology revolves around reframing binary classification as an image captioning task, leveraging the distinctive capabilities of cutting‐edge VLM, notably bootstrapping language image pre‐training (BLIP)2. Rigorous and comprehensive experiments are conducted to validate the effectiveness of our proposed approach, particularly in detecting unseen diffusion‐generated images from unknown diffusion‐based generative models during training, showcasing robustness to noise, and demonstrating generalisation capabilities to GANs. The experiments show that Bi‐LORA outperforms state of the art models in cross‐generator tasks because it leverages multi‐modal learning, open‐world visual knowledge, and benefits from robust, high‐level semantic understanding. By combining visual and textual knowledge, it can handle variations in the data distribution (such as those caused by different generators) and maintain strong performance across different domains. Its ability to transfer knowledge, robustly extract features and perform zero‐shot learning also contributes to its generalisation capabilities, making it more adaptable to new generators. The experimental results showcase an impressive average accuracy of 93.41% in synthetic image detection on unseen generation models. The code and models associated with this research can be publicly accessed at https://github.com/Mamadou‐Keita/VLM‐DETECT . Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, David Camacho, Abdenour Hadid |
Expert Syst. J. Knowl. Eng. | 6 |
| 2025 | Generative AI in the context of assistive technologies: Trends, limitations and future directionsabstractWith the tremendous successes of Large Language Models (LLMs) like ChatGPT for text generation and Dall-E for high-quality image generation, generative Artificial Intelligence (AI) models have shown a hype in our society. Generative AI seamlessly delved into different aspects of society ranging from economy, education, legislation, computer science, finance, and even healthcare. This article provides a comprehensive survey on the increased and promising use of generative AI in assistive technologies benefiting different parties, ranging from the assistive system developers, medical practitioners, care workforce, to the people who need the care and the comfort. Ethical concerns, biases, lack of transparency, insufficient explainability, and limited trustworthiness are major challenges when using generative AI in assistive technologies, particularly in systems that impact people directly. Key future research directions to address these issues include creating standardized rules, establishing commonly accepted evaluation metrics and benchmarks for explainability and reasoning processes, and making further advancements in understanding and reducing bias and its potential harms. Beyond showing the current trends of applying generative AI in the scope of assistive technologies in four identified key domains, which include care sectors, medical sectors, helping people in need, and co-working, the survey also discusses the current limitations and provides promising future research directions to foster better integration of generative AI in assistive technologies. • Presenting the current trends in using generative AI in building assistive systems. • Highlighting the risks and benefits of using generative AI in assistive systems. • Discussing open issues and appealing future research directions. Biying Fu, Abdenour Hadid, Naser Damer |
Image Vis. Comput. | 2 |
| 2025 | Skew-probabilistic neural networks for learning from imbalanced data
Shraddha M. Naik, Tanujit Chakraborty, Madhurima Panja, Abdenour Hadid, Bibhas Chakraborty |
Pattern Recognit. | 4 |
| 2024 | Knowledge-Based Convolutional Neural Network for the Simulation and Prediction of Two-Phase Darcy FlowsabstractPhysics-informed neural networks (PINNs) have gained significant prominence as a powerful tool in the field of scientific computing and simulations. Their ability to seamlessly integrate physical principles into deep learning architectures has revolutionized the approaches to solving complex problems in physics and engineering. However, a persistent challenge faced by mainstream PINNs lies in their handling of discontinuous input data, leading to inaccuracies in predictions. This study addresses these challenges by incorporating the discretized forms of the governing equations into the PINN framework. We propose to combine the power of neural networks with the dynamics imposed by the discretized differential equations. By discretizing the governing equations, the PINN learns to account for the discontinuities and accurately capture the underlying relationships between inputs and outputs, improving the accuracy compared to traditional interpolation techniques. Moreover, by leveraging the power of neural networks, the computational cost associated with numerical simulations is substantially reduced. We evaluate our model on a large-scale dataset for the prediction of pressure and saturation fields demonstrating high accuracies compared to non-physically aware models. Zakaria Elabid, Daniel Busby, Abdenour Hadid |
ICASSP | 3 |
| 2024 | TempoKGAT: A Novel Graph Attention Network Approach for Temporal Graph Analysis
Lena Sasal, Daniel Busby, Abdenour Hadid |
ICONIP (2) | 3 |
| 2024 | Face to Cartoon Incremental Super-Resolution Using Knowledge Distillation
Trinetra Devkatte, Shiv Ram Dubey, Satish Kumar Singh, Abdenour Hadid |
ICPR (7) | 4 |
| 2024 | FIDAVL: Fake Image Detection and Attribution Using Vision-Language Model
Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
ICPR (21) | 5 |
| 2024 | SEECAD: Semantic End-to-End Communication for Autonomous DrivingabstractSemantic communication is a key paradigm in future 6G systems, designed to revolutionize the physical layer of traditional communication, in order to enhance efficiency of data transmission. Fueled by advancements in deep learning architectures, image processing has made significant strides in segmentation and scenes analysis for autonomous vehicles (AVs). Motivated by these advancements, we present an innovative outlook on communication systems, by leveraging semantic data. Thus, we introduce a novel semantic end-to-end communication system, named SEECAD, specifically designed for image transmission in autonomous driving environments. SEECAD is based on a theoretical model, aligning with the semantic level concepts and leveraging a shared knowledge base to efficiently transmit meaningful image data. The semantic encoder and decoder of SEECAD are built upon deep learning architecture, empowered by Low-Density Parity-Check (LDPC) codes. This integration serves to minimize semantic error transmission and enhance the segmentation accuracy at the receiver. Our proposed semantic communication approach was extensively evaluated in various wireless image transmission scenarios over an AWGN channel, using different QAM modulations (4QAM and 16QAM). Our experimental results demonstrated that the proposed SEECAD achieves accurate and effective image transmission in noisy environments. Soheyb Ribouh, Abdenour Hadid |
IV | 2 |
| 2024 | TG-PhyNN: An Enhanced Physically-Aware Graph Neural Network Framework for Forecasting Spatio-Temporal Data
Zakaria Elabid, Lena Sasal, Daniel Busby, Abdenour Hadid |
PRICAI (2) | 4 |
| 2024 | When geoscience meets generative AI and large language models: Foundations, trends, and future challengesabstractAbstract Generative Artificial Intelligence (GAI) represents an emerging field that promises the creation of synthetic data and outputs in different modalities. GAI has recently shown impressive results across a large spectrum of applications ranging from biology, medicine, education, legislation, computer science, and finance. As one strives for enhanced safety, efficiency, and sustainability, generative AI indeed emerges as a key differentiator and promises a paradigm shift in the field. This article explores the potential applications of generative AI and large language models in geoscience. The recent developments in the field of machine learning and deep learning have enabled the generative model's utility for tackling diverse prediction problems, simulation, and multi‐criteria decision‐making challenges related to geoscience and Earth system dynamics. This survey discusses several GAI models that have been used in geoscience comprising generative adversarial networks (GANs), physics‐informed neural networks (PINNs), and generative pre‐trained transformer (GPT)‐based structures. These tools have helped the geoscience community in several applications, including (but not limited to) data generation/augmentation, super‐resolution, panchromatic sharpening, haze removal, restoration, and land surface changing. Some challenges still remain, such as ensuring physical interpretation, nefarious use cases, and trustworthiness. Beyond that, GAI models show promises to the geoscience community, especially with the support to climate change, urban science, atmospheric science, marine science, and planetary science through their extraordinary ability to data‐driven modelling and uncertainty quantification. Abdenour Hadid, Tanujit Chakraborty, Daniel Busby |
Expert Syst. J. Knowl. Eng. | 1 |
| 2024 | Driver's facial expression recognition: A comprehensive survey
Ibtissam Saadi, Douglas W. Cunningham, Abdelmalik Taleb-Ahmed, Abdenour Hadid, Yassin Elhillali |
Expert Syst. Appl. | 4 |
| 2024 | Deepfake Detection Using Spatiotemporal TransformerabstractRecent advances in generative models and the availability of large-scale benchmarks have made deepfake video generation and manipulation easier. Nowadays, the number of new hyper-realistic deepfake videos used for negative purposes is dramatically increasing, thus creating the need for effective deepfake detection methods. Although many existing deepfake detection approaches, particularly CNN-based methods, show promising results, they suffer from several drawbacks. In general, poor generalization results have been obtained under unseen/new deepfake generation methods. The crucial reason for the above defect is that CNN-based methods focus on the local spatial artifacts, which are unique for every manipulation method. Therefore, it is hard to learn the general forgery traces of different manipulation methods without considering the dependencies that extend beyond the local receptive field. To address this problem, this article proposes a framework that combines Convolutional Neural Network (CNN) with Vision Transformer (ViT) to improve detection accuracy and enhance generalizability. Our method, namedHCiT, exploits the advantages of CNNs to extract meaningful local features, as well as the ViT’s self-attention mechanism to learn discriminative global contextual dependencies in a frame-level image explicitly. In this hybrid architecture, the high-level feature maps extracted from the CNN are fed into the ViT model that determines whether a specific video is fake or real. Experiments were performed on Faceforensics++, DeepFake Detection Challenge preview, Celeb datasets, and the results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, the HCiT method shows a great capacity for generalization on datasets covering various techniques of deepfake generation. The source code is available at: https://github.com/KADDAR-Bachir/HCiT Bachir Kaddar, Sid Ahmed Fezza, Zahid Akhtar, Wassim Hamidouche, Abdenour Hadid, Joan Serra-Sagristà |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Probabilistic AutoRegressive Neural Networks for Accurate Long-Range Forecasting
Madhurima Panja, Tanujit Chakraborty, Uttam Kumar 0001, Abdenour Hadid |
ICONIP (13) | 4 |
| 2023 | Kinship recognition from faces using deep learning with imbalanced data
Alice Othmani, Duqing Han, Runpeng Ye, Abdenour Hadid |
Multim. Tools Appl. | 5 |
| 2022 | Knowledge-based Deep Learning for Modeling Chaotic SystemsabstractDeep Learning has received increased attention due to its unbeatable success in many fields, such as computer vision, natural language processing, recommendation systems, and most recently in simulating multiphysics problems and predicting nonlinear dynamical systems. However, modeling and forecasting the dynamics of chaotic systems remains an open research problem since training deep learning models requires big data, which is not always available in many cases. Such deep learners can be trained from additional information obtained from simulated results and by enforcing the physical laws of the chaotic systems. This paper considers extreme events and their dynamics and proposes elegant models based on deep neural networks, called knowledge-based deep learning (KDL). Our proposed KDL can learn the complex patterns governing chaotic systems by jointly training on real and simulated data directly from the dynamics and their differential equations. This knowledge is transferred to model and forecast real-world chaotic events exhibiting extreme behavior. We validate the efficiency of our model by assessing it on three real-world benchmark datasets: El Nino sea surface temperature, San Juan Dengue viral infection, and Bjørnøya daily precipitation, all governed by extreme events’ dynamics. Using prior knowledge of extreme events and physics-based loss functions to lead the neural network learning, we ensure physically consistent, generalizable, and accurate forecasting, even in a small data regime. Zakaria Elabid, Tanujit Chakraborty, Abdenour Hadid |
ICMLA | 3 |
| 2022 | W-Transformers: A Wavelet-based Transformer Framework for Univariate Time Series ForecastingabstractDeep learning utilizing transformers has recently achieved a lot of success in many vital areas such as natural language processing, computer vision, anomaly detection, and recommendation systems, among many others. Among several merits of transformers, the ability to capture long-range temporal dependencies and interactions is desirable for time series forecasting, leading to its progress in various time series applications. In this paper, we build a transformer model for non-stationary time series. The problem is challenging yet crucially important. We present a novel framework for univariate time series representation learning based on the wavelet-based transformer encoder architecture and call it W-Transformer. The proposed W-Transformers utilize a maximal overlap discrete wavelet transformation (MODWT) to the time series data and build local transformers on the decomposed datasets to vividly capture the nonstationarity and long-range nonlinear dependencies in the time series. Evaluating our framework on several publicly available benchmark time series datasets from various domains and with diverse characteristics, we demonstrate that it performs, on average, significantly better than the baseline forecasters for long-term forecasting, even for datasets that consist of only a few hundred training samples. Lena Sasal, Tanujit Chakraborty, Abdenour Hadid |
ICMLA | 3 |
| 2022 | Evaluation of Pre-Trained CNN Models for Geographic Fake Image DetectionabstractThanks to the remarkable advances in generative adversarial networks (GANs), it is becoming increasingly easy to generate/manipulate images. The existing works have mainly focused on deepfake in face images and videos. However, we are currently witnessing the emergence of fake satellite images, which can be misleading or even threatening to national security. Consequently, there is an urgent need to develop detection methods capable of distinguishing between real and fake satellite images. To advance the field, in this paper, we explore the suitability of several convolutional neural network (CNN) architectures for fake satellite image detection. Specifically, we benchmark four CNN models by conducting extensive experiments to evaluate their performance and robustness against various image distortions. This work allows the establishment of new baselines and may be useful for the development of CNN-based methods for fake satellite image detection. Sid Ahmed Fezza, Mohammed Yasser Ouis, Bachir Kaddar, Wassim Hamidouche, Abdenour Hadid |
MMSP | 5 |
| 2022 | A Deep Multiscale Spatiotemporal Network for Assessing Depression From Facial DynamicsabstractRecently, deep learning models have been successfully employed in many video-based affective computing applications (e.g., detecting pain, stress, and Alzheimer’s disease). One key application is automatic depression recognition – recognition of facial expressions associated with depressive behaviour. State-of-the-art deep learning algorithms to recognize depression typically explore spatial and temporal information individually, by using 2D convolutional neural networks (CNNs) to analyze appearance information, and then by either mapping facial feature variations or averaging the depression level over video frames. This approach has limitations in terms of its ability to represent dynamic information that can help to accurately discriminate between depression levels. In contrast, models based on 3D CNNs allow to directly encode the spatio-temporal relationships, although these models rely on temporal information with fixed range and single receptive field. This approach limits the ability to capture variations of facial expression with diverse ranges, and the exploitation of diverse facial areas. In this article, a novel 3D CNN architecture – the Multiscale Spatiotemporal Network (MSN) – is introduced to effectively represent facial information related to depressive behaviours from videos. The basic structure of the model is composed of parallel convolutional layers with different temporal depths and sizes of receptive field, which allows the MSN to explore a wide range of spatio-temporal variations in facial expressions. Experimental results on two benchmark datasets show that our MSN architecture is effective, outperforming state-of-the-art methods in automatic depression recognition. Wheidima C. Melo, Eric Granger, Abdenour Hadid |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | HCiT: Deepfake Video Detection Using a Hybrid Model of CNN features and Vision TransformerabstractThe number of new falsified video contents is dramatically increasing, making the need to develop effective deepfake detection methods more urgent than ever. Even though many existing deepfake detection approaches show promising results, the majority of them still suffer from a number of critical limitations. In general, poor generalization results have been obtained under unseen or new deepfake generation methods. Consequently, in this paper, we propose a deepfake detection method called HCiT, which combines Convolutional Neural Network (CNN) with Vision Transformer (ViT). The HCiT hybrid architecture exploits the advantages of CNN to extract local information with the ViT's self-attention mechanism to improve the detection accuracy. In this hybrid architecture, the feature maps extracted from the CNN are feed into ViT model that determines whether a specific video is fake or real. Experiments were performed on Faceforensics++ and DeepFake Detection Challenge preview datasets, and the results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, the HCiT method shows a great capacity for generalization on datasets covering various techniques of deepfake generation. The source code is available at: https://github.com/KADDAR-Bachir/HCiT Bachir Kaddar, Sid Ahmed Fezza, Wassim Hamidouche, Zahid Akhtar, Abdenour Hadid |
VCIP | 5 |
| 2020 | Multi-view Deep Features for Robust Facial Kinship VerificationabstractAutomatic kinship verification from facial images is an emerging research topic in machine learning community. In this paper, we proposed an effective facial features extraction model based on multi-view deep features. Thus, we used four pre-trained deep learning models using eight features layers (FC6 and FC7 layers of each VGG-F, VGG-M, VGG-S and VGG-Face models) to train the proposed Multilinear Side-Information based Discriminant Analysis integrating Within Class Covariance Normalization (MSIDA + WCCN) method. Furthermore, we show that how can metric learning methods based on WCCN method integration improves the Simple Scoring Cosine similarity (SSC) method. We refer that we used the SSC method in RFIW'20 competition using the eight deep features concatenation. Thus, the integration of WCCN in the metric learning methods decreases the intra-class variations effect introduced by the deep features weights. We evaluate our proposed method on two kinship benchmarks namely KinFaceW-I and KinFaceW-II databases using four Parent-Child relations (Father-Son, Father-Daughter, Mother-Son and Mother-Daughter). Thus, the proposed MSIDA + WCCN method improves the SSC method with 12.80% and 14.65% on KinFaceW-I and KinFaceW-II databases, respectively. The results obtained are positively compared with some modern methods, including those that rely on deep learning. Oualid Laiadi, Abdelmalik Ouamane, Abdelhamid Benakcha, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
FG | 5 |
| 2020 | Kinship Verification From Gait?abstractKinship verification aims to determine whether two persons are kin related or not. This is an emerging topic in computer vision due to its practical potential applications such as family album management. Most of previous works are based on checking kinship from face patterns and more recently from voices. We provide in this paper the first investigation in the literature on kinship verification from gait. The main purpose is to study whether family members do share some gait patterns. As this is a new topic, we started by collecting a new dataset for kinship verification from human gait containing several pairs of video sequences of celebrities and their relatives. The database will be released to the research community for research purposes. Along with the database, we provide results using baseline methods using silhouette and video based analysis. Moreover, we also propose a two-stream 3DCNN to tackle the problem. The preliminary experimental results point out the potential usefulness of gait information for kinship verification. Salah Eddine Bekhouche, Abdelhakim Chergui, Abdenour Hadid, Yassine Ruichek |
ICIP | 3 |
| 2020 | Age estimation from faces using deep learning: A comparative analysis
Alice Othmani, Abdul Rahman Taleb, Hazem Abdelkawy, Abdenour Hadid |
Comput. Vis. Image Underst. | 4 |
| 2020 | Tensor cross-view quadratic discriminant analysis for kinship verification in the wild
Oualid Laiadi, Abdelmalik Ouamane, Abdelhamid Benakcha, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
Neurocomputing | 5 |
| 2020 | Vision-based human activity recognition: a surveyabstractAbstract Human activity recognition (HAR) systems attempt to automatically identify and analyze human activities using acquired information from various types of sensors. Although several extensive review papers have already been published in the general HAR topics, the growing technologies in the field as well as the multi-disciplinary nature of HAR prompt the need for constant updates in the field. In this respect, this paper attempts to review and summarize the progress of HAR systems from the computer vision perspective. Indeed, most computer vision applications such as human computer interaction, virtual reality, security, video surveillance and home monitoring are highly correlated to HAR tasks. This establishes new trend and milestone in the development cycle of HAR systems. Therefore, the current survey aims to provide the reader with an up to date analysis of vision-based HAR related literature and recent progress in the field. At the same time, it will highlight the main challenges and future directions. Djamila Romaissa Beddiar, Brahim Nini, Mohammad Sabokrou, Abdenour Hadid |
Multim. Tools Appl. | 4 |
| 2019 | Face Anti-Spoofing via Sample Learning Based Recurrent Neural Network (RNN)
Tuomas Holmberg, Wheidima C. Melo, Abdenour Hadid |
BMVC | 4 |
| 2019 | Kinship Verification based Deep and Tensor Features through Extreme Learning MachineabstractChecking the kinship of facial images is a difficult research topic in computer vision that has attracted attention in recent years. The methods suggested so far are not strong enough to predict kinship relationships only by facial appearance. To mitigate this problem, we propose a new approach called Deep-Tensor+ELM to kinship verification based on deep (VGG-Face descriptor) and tensor (BSIF-Tensor & LPQ-Tensor using MSIDA method) features through Extreme Learning Machine (ELM). While ELM aims to deal with small size training features dimension, deep and tensor features are proven to provide significant enhancement over shallow features or vector-based counterparts. We evaluate our proposed method on the largest kinship benchmark namely FIW database using four Grandparent-Grandchild relations (GF-GD, GF-GS, GM-GD and GM-GS). The results obtained are positively compared with some modern methods, including those that rely on deep learning. Oualid Laiadi, Abdelmalik Ouamane, Abdelhamid Benakcha, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
FG | 5 |
| 2019 | Combining Global and Local Convolutional 3D Networks for Detecting Depression from Facial ExpressionsabstractDeep learning architectures have been successfully applied in video-based health monitoring, to recognize distinctive variations in the facial appearance of subjects. To detect patterns of variation linked to depressive behavior, deep neural networks (NNs) typically exploit spatial and temporal information separately by, e.g., cascading a 2D convolutional NN (CNN) with a recurrent NN (RNN), although the intrinsic spatio-temporal relationships can deteriorate. With the recent advent of 3D CNNs like the convolutional 3D (C3D) network, these spatio-temporal relationships can be modeled to improve performance. However, the accuracy of C3D networks remain an issue when applied to depression detection. In this paper, the fusion of diverse C3D predictions are proposed to improve accuracy, where spatio-temporal features are extracted from global (full-face) and local (eyes) regions of subject. This allows to increasingly focus on a local facial region that is highly relevant for analyzing depression. Additionally, the proposed network integrates 3D Global Average Pooling in order to efficiently summarize spatio-temporal features without using fully-connected layers, and thereby reduce the number of model parameters and potential over-fitting. Experimental results on the Audio Visual Emotion Challenge (AVEC 2013 and AVEC 2014) depression datasets indicates that combining the responses of global and local C3D networks achieves a higher level of accuracy than state-of-the-art systems. Wheidima C. Melo, Eric Granger, Abdenour Hadid |
FG | 3 |
| 2019 | Learning to Detect Genuine versus Posed Pain from Facial Expressions using Residual Generative Adversarial NetworksabstractWe present a novel approach based on Residual Generative Adversarial Network (R-GAN) to discriminate genuine pain expression from posed pain expression by magnifying the subtle changes in the face. In addition to the adversarial task, the discriminator network in R-GAN estimates the intensity level of the pain. Moreover, we propose a novel Weighted Spatiotemporal Pooling (WSP) to capture and encode the appearance and dynamic of a given video sequence into an image map. In this way, we are able to transform any video into an image map embedding subtle variations in the facial appearance and dynamics. This allows using any pre-trained model on still images for video analysis. Our extensive experiments show that our proposed framework achieves promising results compared to state-of-the-art approaches on three benchmark databases, i.e., UNBC-McMaster Shoulder Pain, BioVid Head Pain, and STOIC. Mohammad Tavakolian, Carlos Guillermo Bermudez Cruces, Abdenour Hadid |
FG | 3 |
| 2019 | AWSD: Adaptive Weighted Spatiotemporal Distillation for Video RepresentationabstractWe propose an Adaptive Weighted Spatiotemporal Distillation (AWSD) technique for video representation by encoding the appearance and dynamics of the videos into a single RGB image map. This is obtained by adaptively dividing the videos into small segments and comparing two consecutive segments. This allows using pre-trained models on still images for video classification while successfully capturing the spatiotemporal variations in the videos. The adaptive segment selection enables effective encoding of the essential discriminative information of untrimmed videos. Based on Gaussian Scale Mixture, we compute the weights by extracting the mutual information between two consecutive segments. Unlike pooling-based methods, our AWSD gives more importance to the frames that characterize actions or events thanks to its adaptive segment length selection. We conducted extensive experimental analysis to evaluate the effectiveness of our proposed method and compared our results against those of recent state-of-the-art methods on four benchmark datatsets, including UCF101, HMDB51, ActivityNet v1.3, and Maryland. The obtained results on these benchmark datatsets showed that our method significantly outperforms earlier works and sets the new state-of-the-art performance in video classification. Code is available at the project webpage: https://mohammadt68.github.io/AWSD/. Mohammad Tavakolian, Hamed Rezazadegan Tavakoli, Abdenour Hadid |
ICCV | 3 |
| 2019 | Depression Detection Based on Deep Distribution LearningabstractMajor depressive disorder is among the most common and harmful mental health problems. Several deep learning architectures have been proposed for video-based detection of depression based on the facial expressions of subjects. To predict the depression level, these architectures are often modeled for regression with Euclidean loss. Consequently, they do not leverage the data distribution, nor explore the ordinal relationship between facial images and depression levels, and have limited robustness to noisy and uncertain labeling. This paper introduces a deep learning architecture for accurately predicting depression levels through distribution learning. It relies on a new expectation loss function that allows to estimate the underlying data distribution over depression levels, where expected values of the distribution are optimized to approach the ground-truth levels. The proposed approach can produce accurate predictions of depression levels even under label uncertainty. Extensive experiments on the AVEC2013 and AVEC2014 datasets indicate that the proposed architecture represents an effective approach that can outperform state-of-the-art techniques. Wheidima C. Melo, Eric Granger, Abdenour Hadid |
ICIP | 3 |
| 2019 | Learning multi-view deep and shallow features through new discriminative subspace for bi-subject and tri-subject kinship verification
Oualid Laiadi, Abdelmalik Ouamane, Abdelhamid Benakcha, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
Appl. Intell. | 5 |
| 2019 | Bag of words KAZE (BoWK) with two-step classification for high-resolution remote sensing imagesabstractThe bag‐of‐words (BoW) model has been widely used for scene classification in recent state‐of‐the‐art methods. However, inter‐class similarity among scene categories and very high spatial resolution imagery makes its performance limited in the remote‐sensing domain. Therefore, this research presents a new KAZE‐based image descriptor that makes use of the BoW approach to substantially increase classification performance. Specifically, a novel multi‐neighbourhood KAZE is proposed for small image patches. Secondly, the spatial pyramid matching and BoW representation can be adopted to use the extracted features and make an innovative BoW KAZE (BoWK) descriptor. Third, two bags of multi‐neighbourhood KAZE features are selected in which each bag is regarded as separated feature descriptors. Next, canonical correlation analysis is introduced as a feature fusion strategy to further refine the BOWK features, which allows a more effective and robust fusion approach than the traditional feature fusion strategies. Experiments on three challenging remote‐sensing data sets show that the proposed BoWK descriptor not only surpasses the conventional KAZE descriptor but also yields significantly higher classification performance than the state‐of‐the‐art methods used now. Moreover, the proposed BoWK approach produces rich informative features to describe the scene images with low‐computational cost and a much lower dimension. Abdenour Hadid, Shahbaz Pervez |
IET Comput. Vis. | 3 |
| 2019 | A Spatiotemporal Convolutional Neural Network for Automatic Pain Intensity Estimation from Facial DynamicsabstractDevising computational models for detecting abnormalities reflective of diseases from facial structures is a novel and emerging field of research in automatic face analysis. In this paper, we focus on automatic pain intensity estimation from faces. This has a paramount potential diagnosis values in healthcare applications. In this context, we present a novel 3D deep model for dynamic spatiotemporal representation of faces in videos. Using several convolutional layers with diverse temporal depths, our proposed model captures a wide range of spatiotemporal variations in the faces. Moreover, we introduce a cross-architecture knowledge transfer technique for training 3D convolutional neural networks using a pre-trained 2D architecture. This strategy is a practical approach for training 3D models, especially when the size of the database is relatively small. Our extensive experiments and analysis on two benchmarking and publicly available databases, namely the UNBC-McMaster shoulder pain and the BioVid, clearly show that our proposed method consistently outperforms many state-of-the-art methods in automatic pain intensity estimation. Mohammad Tavakolian, Abdenour Hadid |
Int. J. Comput. Vis. | 2 |
| 2019 | Kinship verification from face images in discriminative subspaces of color components
Oualid Laiadi, Abdelmalik Ouamane, Elhocine Boutellaa, Abdelhamid Benakcha, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
Multim. Tools Appl. | 6 |
| 2019 | Replayed Video Attack Detection Based on Motion Blur AnalysisabstractFace presentation attacks are the main threats to face recognition systems, and many presentation attack detection (PAD) methods have been proposed in recent years. Although these methods have achieved significant performance in some specific intrusion modes, difficulties still exist in addressing replayed video attacks. That is because the replayed fake faces contain a variety of aliveness signals, such as eye blinking and facial expression changes. Replayed video attacks occur when attackers try to invade biometric systems by presenting face videos in front of the cameras, and these videos are often launched by a liquid-crystal display (LCD) screen. Due to the smearing effects and movements of LCD, videos captured from the real and replayed fake faces present different motion blurs, which are reflected mainly in blur intensity variation and blur width. Based on these descriptions, a motion blur analysis-based method is proposed to deal with the replayed video attack problem. We first present a 1D convolutional neural network (CNN) for motion blur intensity variation description in the time domain, which consists of a serial of 1D convolutional and pooling filters. Then, a local similar pattern (LSP) feature is introduced to extract blur width. Finally, features extracted from 1D CNN and LSP are fused to detect the replayed video attacks. Extensive experiments on two standard face PAD databases, i.e., relay-attack and OULU-NPU, indicate that our proposed method based on the motion blur analysis significantly outperforms the state-of-the-art methods and shows excellent generalization capability. Lei Li 0008, Zhaoqiang Xia, Abdenour Hadid, Xiaoyue Jiang, Haixi Zhang, Xiaoyi Feng |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Deep Discriminative Model for Video Classification
Mohammad Tavakolian, Abdenour Hadid |
ECCV (4) | 2 |
| 2018 | Contextual Weighting of Patches for Local Matching in Still-to-Video Face RecognitionabstractStill-to-video face recognition (FR) systems for watchlist screening seek to recognize individuals of interest given faces captured over a network of video surveillance cameras. Screening faces against a watchlist is a challenging application because only a limited number of reference stills is available per individual during enrollment, and the appearance of face captures in videos changes from camera to camera, due to variations in illumination, pose, blur, scale, expression and occlusion. In order to improve the robustness of FR systems, several local matching techniques have been proposed that rely on static or dynamic weighting of patches. However, these approaches are not suitable for watchlist screening applications where the capturing conditions vary significantly over different camera fields of view (FoV). In this paper, a new dynamic weighting technique is proposed for weighting facial patches based on video data collected a priori from the specific operational domain (camera FoV) and on image quality assessment. Results obtained on videos from the Chokepoint dataset indicate that the proposed approach can significantly outperform the reference local matching methods because patch weights tend to grow for discriminant facial regions. Ibtihel Amara, Eric Granger, Abdenour Hadid |
FG | 3 |
| 2018 | Deep Binary Representation of Facial Expressions: A Novel Framework for Automatic Pain Intensity RecognitionabstractAutomatic pain assessment is crucial in clinical diagnosis. Experiencing pain causes deformations in the facial structure resulting in different spontaneous facial expressions. In this paper, we aim to represent the facial expressions as a compact binary code for classification of different pain intensity levels. We divide a given face video into non-overlapping equal-length segments. Using a Convolutional Neural Network (CNN), we extract features from randomly sampled frames from all segments. The obtained features are aggregated by exploiting statistics to incorporate low-level visual patterns and high-level structural information. Finally, this processed information is encoded using a deep network to obtain a single binary code such that videos with the same pain intensity level have smaller Hamming distance than those of different levels. Extensive experiments on the publicly available UNBC-McMaster database demonstrates that our proposed method achieves superior performance compared to the state-of-the-art. Mohammad Tavakolian, Abdenour Hadid |
ICIP | 2 |
| 2018 | Deep Spatiotemporal Representation of the Face for Automatic Pain Intensity EstimationabstractAutomatic pain intensity assessment has a high value in disease diagnosis applications. Inspired by the fact that many diseases and brain disorders can interrupt normal facial expression formation, we aim to develop a computational model for automatic pain intensity assessment from spontaneous and micro facial variations. For this purpose, we propose a 3D deep architecture for dynamic facial video representation. The proposed model is built by stacking several convolutional modules where each module encompasses a 3D convolution kernel with a fixed temporal depth, several parallel 3D convolutional kernels with different temporal depths, and an average pooling layer. Deploying variable temporal depths in the proposed architecture allows the model to effectively capture a wide range of spatiotemporal variations on the faces. Extensive experiments on the UNBC-McMaster Shoulder Pain Expression Archive database show that our proposed model yields in a promising performance compared to the state-of-the-art in automatic pain intensity estimation. Mohammad Tavakolian, Abdenour Hadid |
ICPR | 2 |
| 2018 | Feature Fusion with Deep Supervision for Remote-Sensing Image Scene ClassificationabstractThe convolutional neural networks (CNNs) have shown an intrinsic ability to automatically extract high level representations for image classification, but there is a major hurdle to their deployment in the remote-sensing domain because of a relative lack of training data. Moreover, traditional fusion methods use either low-level features or score-based fusion to fuse the features. In order to address the aforementioned issues, we employed a deep supervision (DS) strategy to enhance the generalization performance in the intermediate layers of the AlexNet model for remote-sensing image scene classification. The proposed DS strategy not only prevents from overfitting, but also extracts the features more transparently. Secondly, the canonical correlation analysis (CCA) is adopted as a feature fusion strategy to further refine the features with more discriminative power. The fused AlexNet features achieved by the proposed framework have much higher discrimination than the pure features. Extensive experiments on two challenging datasets: 1) UC MERCED data set and 2) WHU-RS dataset demonstrate that the two proposed approaches both enhance the performance of the original AlexNet architecture, and also outperform several state-of-the-art methods currently in use. Abdenour Hadid |
ICTAI | 3 |
| 2018 | On the generalization of color texture-based face anti-spoofing
Zinelabidine Boulkenafet, Jukka Komulainen, Abdenour Hadid |
Image Vis. Comput. | 3 |
| 2018 | Face spoofing detection with local binary pattern network
Lei Li 0008, Xiaoyi Feng, Zhaoqiang Xia, Xiaoyue Jiang, Abdenour Hadid |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Kinship verification from facial images and videos: human versus machine
Miguel Bordallo López, Abdenour Hadid, Elhocine Boutellaa, Jorge Gonçalves 0001, Vassilis Kostakos, Simo Hosio |
Mach. Vis. Appl. | 2 |
| 2018 | A Survey on Computer Vision for Assistive Medical Diagnosis From FacesabstractAutomatic medical diagnosis is an emerging center of interest in computer vision as it provides unobtrusive objective information on a patient's condition. The face, as a mirror of health status, can reveal symptomatic indications of specific diseases. Thus, the detection of facial abnormalities or atypical features is at upmost importance when it comes to medical diagnostics. This survey aims to give an overview of the recent developments in medical diagnostics from facial images based on computer vision methods. Various approaches have been considered to assess facial symptoms and to eventually provide further help to the practitioners. However, the developed tools are still seldom used in clinical practice, since their reliability is still a concern due to the lack of clinical validation of the methodologies and their inadequate applicability. Nonetheless, efforts are being made to provide robust solutions suitable for healthcare environments, by dealing with practical issues such as real-time assessment or patients positioning. This survey provides an updated collection of the most relevant and innovative solutions in facial images analysis. The findings show that with the help of computer vision methods, over 30 medical conditions can be preliminarily diagnosed from the automatic detection of some of their symptoms. Furthermore, future perspectives, such as the need for interdisciplinary collaboration and collecting publicly available databases, are highlighted. Jérôme Thevenot, Miguel Bordallo López, Abdenour Hadid |
IEEE J. Biomed. Health Informatics | 3 |
| 2017 | Foreword
Bir Bhanu, Abdenour Hadid, Mark Nixon, Vitomir Struc |
FG | 2 |
| 2017 | OULU-NPU: A Mobile Face Presentation Attack Database with Real-World VariationsabstractThe vulnerabilities of face-based biometric systems to presentation attacks have been finally recognized but yet we lack generalized software-based face presentation attack detection (PAD) methods performing robustly in practical mobile authentication scenarios. This is mainly due to the fact that the existing public face PAD datasets are beginning to cover a variety of attack scenarios and acquisition conditions but their standard evaluation protocols do not encourage researchers to assess the generalization capabilities of their methods across these variations. In this present work, we introduce a new public face PAD database, OULU-NPU, aiming at evaluating the generalization of PAD methods in more realistic mobile authentication scenarios across three covariates: unknown environmental conditions (namely illumination and background scene), acquisition devices and presentation attack instruments (PAI). This publicly available database consists of 5940 videos corresponding to 55 subjects recorded in three different environments using high-resolution frontal cameras of six different smartphones. The high-quality print and video-replay attacks were created using two different printers and two different display devices. Each of the four unambiguously defined evaluation protocols introduces at least one previously unseen condition to the test set, which enables a fair comparison on the generalization capabilities between new and existing approaches. The baseline results using color texture analysis based face PAD method demonstrate the challenging nature of the database. Zinelabidine Boulkenafet, Jukka Komulainen, Lei Li 0008, Xiaoyi Feng, Abdenour Hadid |
FG | 5 |
| 2017 | A competition on generalized software-based face presentation attack detection in mobile scenariosabstractIn recent years, software-based face presentation attack detection (PAD) methods have seen a great progress. However, most existing schemes are not able to generalize well in more realistic conditions. The objective of this competition is to evaluate and compare the generalization performances of mobile face PAD techniques under some real-world variations, including unseen input sensors, presentation attack instruments (PAI) and illumination conditions, on a larger scale OULU-NPU dataset using its standard evaluation protocols and metrics. Thirteen teams from academic and industrial institutions across the world participated in this competition. This time typical liveness detection based on physiological signs of life was totally discarded. Instead, every submitted system relies practically on some sort of feature representation extracted from the face and/or background regions using hand-crafted, learned or hybrid descriptors. Interesting results and findings are presented and discussed in this paper. Zinelabidine Boulkenafet, Jukka Komulainen, Zahid Akhtar, Azeddine Benlamoudi, Djamel Samai, Salah Eddine Bekhouche, Abdelkrim Ouafi, Fadi Dornaika, Abdelmalik Taleb-Ahmed, Fei Peng 0001, L. B. Zhang, Min Long 0003, Shruti Bhilare, Vivek Kanhangad, Artur Costa-Pazo, Esteban Vázquez-Fernández, Daniel Pérez-Cabo, J. J. Moreira-Perez, Daniel González-Jiménez, Amir Mohammadi, Sushil Bhattacharjee, Sébastien Marcel, Svetlana Volkova, N. Abe, X. Feng, Z. Xia, Rui Shao 0001, Pong C. Yuen, Waldir R. de Almeida, Fernanda A. Andaló, Rafael Padilha, Gabriel Bertocco, William Dias, Jacques Wainer, Ricardo da Silva Torres, Anderson Rocha 0001, Marcus A. Angeloni, Guilherme Folego, Alan Godoy, Abdenour Hadid |
IJCB | 45 |
| 2017 | Face anti-spoofing via deep local binary patternsabstractConvolutional neural networks (CNNs) have achieved excellent performance in the field of pattern recognition when huge amount of training data is available. However, training a CNN model is less obvious when only a limited amount of data is given such as in the case of face anti-spoofing problem. It is indeed not easy to collect very large sets of fake faces. Especially for the fully-connected layers, tens of thousands of parameters need to be learned. To tackle this problem of lack of training data in face anti-spoofing, we propose to explore the incorporation of hand-crafted features in the CNN framework. In our proposed approach, the color local binary patterns (LBP) features are extracted from the convolutional feature maps, which are fine tuned based on the VGG-face model. These features are then fed into support vector machine (SVM) classifier. Extensive experiments are conducted on two benchmark and publicly available databases showing very interesting performance compared to state-of-the-art methods. Lei Li 0008, Xiaoyi Feng, Xiaoyue Jiang, Zhaoqiang Xia, Abdenour Hadid |
ICIP | 5 |
| 2017 | Pyramid multi-level features for facial demographic estimation
Salah Eddine Bekhouche, Abdelkrim Ouafi, Fadi Dornaika, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
Expert Syst. Appl. | 5 |
| 2017 | Deep convolutional hashing using pairwise multi-label supervision for large-scale visual search
Zhaoqiang Xia, Xiaoyi Feng, Jie Lin 0001, Abdenour Hadid |
Signal Process. Image Commun. | 4 |
| 2017 | Face Antispoofing Using Speeded-Up Robust Features and Fisher Vector EncodingabstractThe vulnerabilities of face biometric authentication systems to spoofing attacks have received a significant attention during the recent years. Some of the proposed countermeasures have achieved impressive results when evaluated on intratests, i.e., the system is trained and tested on the same database. Unfortunately, most of these techniques fail to generalize well to unseen attacks, e.g., when the system is trained on one database and then evaluated on another database. This is a major concern in biometric antispoofing research that is mostly overlooked. In this letter, we propose a novel solution based on describing the facial appearance by applying Fisher vector encoding on speeded-up robust features extracted from different color spaces. The evaluation of our countermeasure on three challenging benchmark face-spoofing databases, namely the CASIA face antispoofing database, the replay-attack database, and MSU mobile face spoof database, showed excellent and stable performance across all the three datasets. Most importantly, in interdatabase tests, our proposed approach outperforms the state of the art and yields very promising generalization capabilities, even when only limited training data are used. Zinelabidine Boulkenafet, Jukka Komulainen, Abdenour Hadid |
IEEE Signal Process. Lett. | 3 |
| 2017 | Efficient Tensor-Based 2D+3D Face VerificationabstractWe propose a novel approach for face verification by encoding 2D and 3D face images as a high order tensor. To perform tensor dimensionality reduction for both the unsupervised and supervised cases, we propose multilinear whitened principal component analysis (MWPCA) and tensor exponential discriminant analysis (TEDA), respectively. MWPCA is utilized to solve the small sample size problem in the high-dimensional space and to improve the discrimination power achieved by classical MPCA. In the supervised case, we extend multilinear discriminant analysis to TEDA in order to emphasize the discriminant data included in the null space of the within-class scatter matrix of each tensor's mode. Additionally, TEDA enlarges the margin between samples belonging to different classes via distance diffusion mappings. Our proposed approach can be seen as a novel data fusion method based on tensor representation. Indeed, the histograms of different local descriptors extracted from both 2D and 3D face modalities are combined through different tensor modes. The extensive experimental evaluation carried out on FRGC v2.0, Bosphorus, and CASIA 2D and 3D face databases indicates that the proposed approach performs significantly better than the state-of-the-art approaches. Abdelmalik Ouamane, Ammar Chouchane, Elhocine Boutellaa, Mebarka Belahcene, Salah Bourennane, Abdenour Hadid |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2016 | Kernel sparse modeling for prototype selection
Fadi Dornaika, Ihab Kamal Aldine, Abdenour Hadid |
Knowl. Based Syst. | 3 |
| 2016 | Audiovisual synchrony assessment for replay attack detection in talking face biometrics
Elhocine Boutellaa, Zinelabidine Boulkenafet, Jukka Komulainen, Abdenour Hadid |
Multim. Tools Appl. | 4 |
| 2016 | Comments on the "Kinship Face in the Wild" Data SetsabstractThe Kinship Face in the Wild data sets, recently published in TPAMI, are currently used as a benchmark for the evaluation of kinship verification algorithms. We recommend that these data sets are no longer used in kinship verification research unless there is a compelling reason that takes into account the nature of the images. We note that most of the image kinship pairs are cropped from the same photographs. Exploiting this cropping information, competitive but biased performance can be obtained using a simple scoring approach, taking only into account the nature of the image pairs rather than any features about kin information. To illustrate our motives, we provide classification results utilizing a simple scoring method based on the image similarity of both images of a kinship pair. Using simply the distance of the chrominance averages of the images in the Lab color space without any training or using any specific kin features, we achieve performance comparable to state-of-the-art methods. We provide the source code to prove the validity of our claims and ensure the repeatability of our experiments. Miguel Bordallo López, Elhocine Boutellaa, Abdenour Hadid |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Face Spoofing Detection Using Colour Texture AnalysisabstractResearch on non-intrusive software-based face spoofing detection schemes has been mainly focused on the analysis of the luminance information of the face images, hence discarding the chroma component, which can be very useful for discriminating fake faces from genuine ones. This paper introduces a novel and appealing approach for detecting face spoofing using a colour texture analysis. We exploit the joint colour-texture information from the luminance and the chrominance channels by extracting complementary low-level feature descriptions from different colour spaces. More specifically, the feature histograms are computed over each image band separately. Extensive experiments on the three most challenging benchmark data sets, namely, the CASIA face anti-spoofing database, the replay-attack database, and the MSU mobile face spoof database, showed excellent results compared with the state of the art. More importantly, unlike most of the methods proposed in the literature, our proposed approach is able to achieve stable performance across all the three benchmark data sets. The promising results of our cross-database evaluation suggest that the facial colour texture representation is more stable in unknown conditions compared with its gray-scale counterparts. Zinelabidine Boulkenafet, Jukka Komulainen, Abdenour Hadid |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2015 | Face anti-spoofing based on color texture analysisabstractResearch on face spoofing detection has mainly been focused on analyzing the luminance of the face images, hence discarding the chrominance information which can be useful for discriminating fake faces from genuine ones. In this work, we propose a new face anti-spoofing method based on color texture analysis. We analyze the joint color-texture information from the luminance and the chrominance channels using a color local binary pattern descriptor. More specifically, the feature histograms are extracted from each image band separately. Extensive experiments on two benchmark datasets, namely CASIA face anti-spoofing and Replay-Attack databases, showed excellent results compared to the state-of-the-art. Most importantly, our inter-database evaluation depicts that the proposed approach showed very promising generalization capabilities. Zinelabidine Boulkenafet, Jukka Komulainen, Abdenour Hadid |
ICIP | 3 |
| 2015 | On the use of Kinect depth data for identity, gender and ethnicity classification from facial images
Elhocine Boutellaa, Abdenour Hadid, Messaoud Bengherabi, Samy Ait-Aoudia |
Pattern Recognit. Lett. | 2 |
| 2015 | Special issue on "Soft Biometrics"
Paulo Lobato Correia, Abdenour Hadid, Thomas B. Moeslund |
Pattern Recognit. Lett. | 2 |
| 2015 | Gender and texture classification: A comparative analysis using 13 variants of local binary patterns
Abdenour Hadid, Juha Ylioinas, Messaoud Bengherabi, Mohammad Ghahramani, Abdelmalik Taleb-Ahmed |
Pattern Recognit. Lett. | 1 |
| 2015 | On soft biometrics
Mark S. Nixon, Paulo Lobato Correia, Kamal Nasrollahi, Thomas B. Moeslund, Abdenour Hadid, Massimo Tistarelli |
Pattern Recognit. Lett. | 5 |
| 2014 | Generalized textured contact lens detection by extracting BSIF description from Cartesian iris imagesabstractTextured contact lenses cause severe problems for iris biometric systems because they can be used to alter the appearance of iris texture in order to deliberately increase the false positive and, especially, false negative match rates. Many texture analysis based techniques have been proposed for detecting the presence of cosmetic contact lenses. However, it has been shown recently that the generalization capability of the existing approaches is not sufficient because they have been developed for detecting specific lens texture patterns and evaluated only on those same lens types seen during development phase. This scenario does not apply in unpredictable practical applications because unseen lens patterns will be definitely experienced in operation. In this paper, we address this issue by studying the effect of different iris image preprocessing techniques and introducing a novel approach formore generalized cosmetic contact lens detection using binarized statistical image features (BSIF).Our extensive experimental analysis on benchmark datasets shows that the BSIF description extracted from preprocessed Cartesian iris texture images yields to promising generalization capabilities across unseen texture patterns and different iris sensors with mean equal error rate of 0.14%and 0.88%, respectively. The findings support the intuition that the textural differences between genuine iris texture and fake ones are best described by preserving the regular structure of different printing signatures without transforming the iris images into polar coordinate system. Jukka Komulainen, Abdenour Hadid, Matti Pietikäinen |
IJCB | 2 |
| 2014 | Multi scale multi descriptor local binary features and exponential discriminant analysis for robust face authenticationabstractIn this paper we present an efficient face verification system based on the fusion of multi-scale multi-descriptor local binary features. First, the face is divided into regions and each region is divided into several patches. For each patch and at every specific scale, the statistics of the baseline Local Binary Pattern (LBP), the Local Phase Quantization (LPQ) and the recently proposed Binarized Statistical Image Feature (BSIF) are summarized by histograms. The histograms of different patches belonging to the same region are concatenated to form a highly dimensional feature vector representing a specific descriptor at a specific scale. Second, we propose an efficient dimensionality reduction technique based on Exponential Linear Discriminant Analysis EDA coupled with Within-Class Covariance Normalization (WCCN) to downgrade the effect of the directions of high intravariability and to enhance the discrimination power of the EDA. The projected histograms for each region are scored using the cosine similarity metric. Lastly, the different region scores corresponding to different descriptors at different scales are fused using support vector machine classifier (SVM). Experimental verification results demonstrate that the proposed authentication pipeline outperforms all the existing systems on the XM2VTS controlled database and interestingly compete with the top performing systems on the challenging LFW database. Abdelmalik Ouamane, Messaoud Bengherabi, Abderrezak Guessoum, Abdenour Hadid, Mohamed Cheriet |
ICIP | 4 |
| 2014 | An In-depth Examination of Local Binary Descriptors in Unconstrained Face RecognitionabstractAutomatic face recognition in unconstrained conditions is a difficult task which has recently attained increasing attention. In this domain, face verification methods have significantly improved since the release of the Labeled Faces in the Wild database, but the related problem of face identification, is still lacking considerations, which is partly because of the shortage of representative databases. Only recently, two new datasets called Remote Face and Point-and-Shoot Challenge were published providing appropriate benchmarks for the research community to investigate the problem of face recognition in challenging imaging conditions, in both, verification and identification modes. In this paper we provide an in-depth examination of three local binary description methods in unconstrained face recognition evaluating them on these two recently published datasets. In detail, we investigate three well established methods separately and fusing them at rank- and score-levels. We are using a well-defined evaluation protocol allowing a fair comparison of our results for future examinations. Juha Ylioinas, Abdenour Hadid, Juho Kannala, Matti Pietikäinen |
ICPR | 2 |
| 2014 | Learning local image descriptors using binary decision treesabstractIn this paper we propose a unified framework for learning such local image descriptors that describe pixel neighborhoods using binary codes. The descriptors are constructed using binary decision trees which are learnt from a set of training image patches. Our framework generalizes several previously proposed binary descriptors, such as BRIEF, LBP and their variants, and provides a principled way to learn new constructions which have not been previously studied. Further, the proposed framework can utilize both labeled or unlabeled training data, and hence fits to both supervised and unsupervised learning scenarios. We evaluate our framework using varying levels of supervision in the learning phase. The experiments show that our descriptor constructions perform comparably to benchmark descriptors in two different applications, namely texture categorization and age group classification from facial images. Juha Ylioinas, Juho Kannala, Abdenour Hadid, Matti Pietikäinen |
WACV | 3 |
| 2013 | Demographic classification from face videos using manifold learning
Abdenour Hadid, Matti Pietikäinen |
Neurocomputing | 1 |
| 2012 | Efficient Image Appearance Description Using Dense Sampling Based Local Binary Patterns
Juha Ylioinas, Abdenour Hadid, Yimo Guo, Matti Pietikäinen |
ACCV (3) | 2 |
| 2012 | Can gait biometrics be Spoofed?
Abdenour Hadid, Mohammad Ghahramani, Vili Kellokumpu, Matti Pietikäinen, John D. Bustard, Mark S. Nixon |
ICPR | 1 |
| 2012 | Age Classification in Unconstrained Conditions Using LBP Variants
Juha Ylioinas, Abdenour Hadid, Matti Pietikäinen |
ICPR | 2 |
| 2011 | Improving the recognition of faces occluded by facial accessoriesabstractFacial occlusions, due for example to sunglasses, hats, scarf, beards etc., can significantly affect the performance of any face recognition system. Unfortunately, the presence of facial occlusions is quite common in real-world applications especially when the individuals are not cooperative with the system such as in video surveillance scenarios. While there has been an enormous amount of research on face recognition under pose/illumination changes and image degradations, problems caused by occlusions are mostly overlooked. The focus of this paper is thus on facial occlusions, and particularly on how to improve the recognition of faces occluded by sunglasses and scarf. We propose an efficient approach which consists of first detecting the presence of scarf/sunglasses and then processing the non-occluded facial regions only. The occlusion detection problem is approached using Gabor wavelets, PCA and support vector machines (SVM), while the recognition of the non-occluded facial part is performed using block-based local binary patterns. Experiments on AR face database showed that the proposed method yields significant performance improvements compared to existing works for recognizing partially occluded and also non-occluded faces. Furthermore, the performance of the proposed approach is also assessed under illumination and extreme facial expression changes, demonstrating interesting results. Rui Min 0002, Abdenour Hadid, Jean-Luc Dugelay |
FG | 2 |
| 2011 | Competition on counter measures to 2-D facial spoofing attacksabstractSpoofing identities using photographs is one of the most common techniques to attack 2-D face recognition systems. There seems to exist no comparative studies of different techniques using the same protocols and data. The motivation behind this competition is to compare the performance of different state-of-the-art algorithms on the same database using a unique evaluation method. Six different teams from universities around the world have participated in the contest. Use of one or multiple techniques from motion, texture analysis and liveness detection appears to be the common trend in this competition. Most of the algorithms are able to clearly separate spoof attempts from real accesses. The results suggest the investigation of more complex attacks. Murali Mohan Chakka, André Anjos, Sébastien Marcel, Roberto Tronci, Daniele Muntoni, Gianluca Fadda, Maurizio Pili, Nicola Sirena, Gabriele Murgia, Marco Ristori, Fabio Roli, Dong Yi, Zhen Lei 0001, Stan Z. Li, William Robson Schwartz, Anderson Rocha 0001, Hélio Pedrini, Javier Lorenzo-Navarro, Modesto Castrillón-Santana, Jukka Komulainen, Abdenour Hadid, Matti Pietikäinen |
IJCB | 23 |
| 2011 | Face spoofing detection from single images using micro-texture analysisabstractCurrent face biometric systems are vulnerable to spoo ing attacks. A spoofing attack occurs when a person tries to masquerade as someone else by falsifying data and thereby gaining illegitimate access. Inspired by image quality assessment, characterization of printing artifacts, and differences in light reflection, we propose to approach the problem of spoofing detection from texture analysis point of view. Indeed, face prints usually contain printing quality defects that can be well detected using texture features. Hence, we present a novel approach based on analyzing facial image textures for detecting whether there is a live person in front of the camera or a face print. The proposed approach analyzes the texture of the facial images using multi-scale local binary patterns (LBP). Compared to many previous works, our proposed approach is robust, computationally fast and does not require user-cooperation. In addition, the texture features that are used for spoofing detection can also be used for face recognition. This provides a unique feature space for coupling spoofing detection and face recognition. Extensive experimental analysis on a publicly avail able database showed excellent results compared to existing works. Jukka Komulainen, Abdenour Hadid, Matti Pietikäinen |
IJCB | 2 |
| 2011 | Facial Deblur Inference Using Subspace Analysis for Recognition of Blurred FacesabstractThis paper proposes a novel method for recognizing faces degraded by blur using deblurring of facial images. The main issue is how to infer a Point Spread Function (PSF) representing the process of blur on faces. Inferring a PSF from a single facial image is an ill-posed problem. Our method uses learned prior information derived from a training set of blurred faces to make the problem more tractable. We construct a feature space such that blurred faces degraded by the same PSF are similar to one another. We learn statistical models that represent prior knowledge of predefined PSF sets in this feature space. A query image of unknown blur is compared with each model and the closest one is selected for PSF inference. The query image is deblurred using the PSF corresponding to that model and is thus ready for recognition. Experiments on a large face database (FERET) artificially degraded by focus or motion blur show that our method substantially improves the recognition performance compared to existing methods. We also demonstrate improved performance on real blurred images on the FRGC 1.0 face database. Furthermore, we show and explain how combining the proposed facial deblur inference with the local phase quantization (LPQ) method can further enhance the performance. Masashi Nishiyama, Abdenour Hadid, Hidenori Takeshima, Jamie Shotton, Tatsuo Kozakaya, Osamu Yamaguchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Recognition of Blurred Faces via Facial Deblurring Combined with Blur-Tolerant DescriptorsabstractBlur is often present in real-world images and significantly affects the performance of face recognition systems. To improve the recognition of blurred faces, we propose a new approach which inherits the advantages of two recent methods. The idea consists of first reducing the amount of blur in the images via deblurring and then extracting blur-tolerant descriptors for recognition. We assess our analysis on real blurred face images (FRGC 1.0 database) and also on face images artificially degraded by focus blur (FERET database), demonstrating significant performance enhancement compared to the state-of-the-art. Abdenour Hadid, Masashi Nishiyama, Yoichi Sato 0001 |
ICPR | 1 |
| 2009 | Combining appearance and motion for face and gender recognition from videos
Abdenour Hadid, Matti Pietikäinen |
Pattern Recognit. | 1 |
| 2008 | Combining motion and appearance for gender classification from video sequencesabstractWe investigate whether combining appearance (face structure) and motion (the way a person is talking and moving his/her facial features) boosts gender classification from face sequences. We propose and compare different schemes based on appearance only, motion only, and combination of appearance and motion. Experiments on various face video datasets of persons uttering phrases or expressing emotions show that combination of motion and appearance is useful for gender analysis of familiar faces, yielding in classification accuracy of 100%. However, for unfamiliar faces, motion seems to not provide additional discriminative information as the best performance (96.3%) is obtained using an appearance based approach with Local Binary Pattern (LBP) features and Support Vector Machines (SVMs). Abdenour Hadid, Matti Pietikäinen |
ICPR | 1 |
| 2006 | Face Description with Local Binary Patterns: Application to Face RecognitionabstractThis paper presents a novel and efficient facial image representation based on local binary pattern (LBP) texture features. The face image is divided into several regions from which the LBP feature distributions are extracted and concatenated into an enhanced feature vector to be used as a face descriptor. The performance of the proposed method is assessed in the face recognition problem under different challenges. Other applications and several extensions are also discussed. Timo Ahonen, Abdenour Hadid, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | A Novel Real Time System for Facial Expression Recognition
Xiaoyi Feng, Matti Pietikäinen, Abdenour Hadid, Hongmei Xie |
ACII | 3 |
| 2004 | A Discriminative Feature Space for Detecting and Recognizing Faces
Abdenour Hadid, Matti Pietikäinen, Timo Ahonen |
CVPR (2) | 1 |
| 2004 | Face Recognition with Local Binary Patterns
Timo Ahonen, Abdenour Hadid, Matti Pietikäinen |
ECCV (1) | 2 |