VLDB 2026 Research / reviewers in the wild / expert
Zahid Akhtar
dblp:52/10105
· DBLP profile ↗
41ranked-venue papers
3as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 15 · 10 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Security and privacy · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 3 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainability-Guided Deepfake Detection for High-Fidelity Facial Edits
Bibek Das, Soumi Chattopadhyay, Chandranath Adak, Astitva Pandey, Ashutosh Parihar, Zahid Akhtar, Soumya Dutta, Abdenour Hadid |
ICPR (4) | 6 |
| 2026 | Diffusion-Latent Invisible Watermarking for Proactive Deepfake Provenance Verification
Bibek Das, Anurag Deo, Chandranath Adak, Soumi Chattopadhyay, Zahid Akhtar, Soumya Dutta, Abdenour Hadid |
ICPR (4) | 5 |
| 2025 | CE-KD: Class-Wise Expert-Based Knowledge Distillation for Facial Expression RecognitionabstractKnowledge distillation (KD) is a model compression technique that transfers knowledge from a complex and well-trained teacher model to a compact student model, thereby enabling the student to mimic the performance and behavior of the teacher. However, traditional KD methods struggle with long-tailed facial expression recognition (FER), as FER datasets often exhibit severe class imbalance. For instance, certain expressions (e.g., happiness) are overrepresented, while others (e.g., fear) have significantly fewer samples. This innate class-imbalanced property of FER leads to suboptimal knowledge transfer for the underrepresented expressions (i.e. biased learning and poor generalization for underrepresented classes). To address this issue, this paper introduces the CE-KD, a Class-wise Expert-based Knowledge Distillation framework that enables the student model to effectively learn fine-grained expression details and high-level emotional concepts from specialized emotion experts. This improves the generalizability of the FER models. Extensive experiments on benchmark datasets — FERPlus and RAF-DB — demonstrate that our CE-KD framework provides a practical solution to implement efficient FER systems in real-world applications while maintaining robust performance across different emotion expressions. Khin Cho Win, Zahid Akhtar, C. Krishna Mohan |
AVSS | 2 |
| 2025 | IndicSideFace: A Dataset for Advancing Deepfake Detection on Side-Face Perspectives of Indian SubjectsabstractThe rapid advancement of generative models and their misuse have made deepfake detection a crucial area of research. However, existing datasets and detection techniques predominantly focus on frontal-face perspectives, leaving sideface views largely underexplored. To bridge this gap, we present IndicSideFace, a novel dataset specifically curated for advancing deepfake detection on side-face perspectives of Indian subjects. This dataset encompasses a diverse range of side-face angles, varying lighting conditions, and demographic attributes, providing a comprehensive benchmark for evaluating detection algorithms. Our experiments using state-of-the-art models highlight the unique challenges posed by side-face deepfakes, such as partial facial feature visibility and uncommon head poses. The findings reveal significant limitations in existing detection approaches when applied to side-face perspectives, underscoring the need for specialized solutions. With IndicSideFace, we aim to strengthen the resilience of deepfake detectors and stimulate further research in this critical yet underexplored domain. Anurag Deo, Aditya Bangar, Chandranath Adak, Rahul Verma, Deepak Nagar, Zahid Akhtar, Soumya Dutta, Soumi Chattopadhyay, Sukalpa Chanda |
FG | 6 |
| 2025 | Adversarial Attacks on Deepfake Detectors: A Challenge in the Era of AI-Generated Media (AADD-2025)abstractThe rapid proliferation of AI-generated media, particularly hyper-realistic deepfakes, has underscored the critical need for robust detection systems to mitigate risks such as misinformation and identity theft. However, state-of-the-art deepfake detectors remain vulnerable to adversarial attacks-subtle perturbations designed to evade classification. To address this gap, we organized the Adversarial Attacks on Deepfake Detectors (AADD-2025) challenge, a competitive evaluation aimed at advancing methodologies to expose and strengthen weaknesses in deepfake detection models. The challenge tasked participants with generating adversarial examples capable of evading four diverse classifiers (including ResNet, DenseNet, and two blind models) while preserving structural similarity to original deepfakes. A dataset comprising 16 subsets of high- and low-quality deepfake images generated by GAN-based and diffusion models (e.g., StableDiffusion, StyleGAN3) was provided. Participants were evaluated using a weighted combination of Structural Similarity Index (SSIM) and attack success rates across all classifiers. Thirteen teams proposed innovative solutions leveraging techniques such as latent-space manipulation, ensemble gradient optimization, surrogate modeling, and frequency-domain perturbation. Top-performing approaches, including MR-CAS (1st place), Safe AI (2nd place), and RoMa (3rd place), achieved high SSIM scores (0.74-0.93) while successfully misleading classifiers. Notably, MR-CAS's latent diffusion model inversion strategy and Safe AI's consensus-orthogonal gradient weighting framework demonstrated superior transferability across architectures, including Vision Transformers. The challenge revealed critical insights: latent-space attacks outperformed pixel-level methods, ensemble-based strategies enhanced cross-model robustness, and adversarial perturbations optimized for both CNNs and transformers proved most effective. However, gaps persist in generalizing attacks across heterogeneous models and maintaining perceptual fidelity, highlighting the urgency of developing adaptive defenses and hybrid detection mechanisms. By fostering collaboration and innovation, AADD-2025 provides a benchmark for evaluating adversarial robustness in deepfake detection and underscores the need for resilient systems in the era of AI-generated media. Sebastiano Battiato, Mirko Casu, Francesco Guarnera, Luca Guarnera, Giovanni Puglisi, Orazio Pontorno, Claudio Vittorio Ragaglia, Zahid Akhtar |
ACM Multimedia | 8 |
| 2025 | (DFF '25) 1st Deepfake Forensics Workshop: Detection, Attribution, Recognition, and Adversarial Challenges in the Era of AI-Generated MediaabstractThe proliferation of generative models, particularly Generative Adversarial Networks (GANs) and Diffusion Models, has reshaped multimedia content creation. Alongside creative and commercial opportunities, they have introduced unprecedented risks through the production of highly realistic synthetic content, or deepfakes. These artifacts challenge visual and auditory trust, with major implications for media, security, politics, and law. This workshop provides a forum to examine deepfake technology from forensic, technical, legal, and social perspectives. It will bring together experts to advance robust and explainable detection methods, define benchmarking practices, and address ethical and regulatory frameworks. Topics include detection and attribution, adversarial countermeasures, multimodal analysis, model traceability, legal admissibility of synthetic content, as well as real-world deployment challenges and dataset creation. Further information about the workshop is available at https://iplab.dmi.unict.it/mfs/acm-dff-ws-2025/ Sebastiano Battiato, Mirko Casu, Francesco Guarnera, Luca Guarnera, Giovanni Puglisi, Orazio Pontorno, Claudio Vittorio Ragaglia, Zahid Akhtar |
ACM Multimedia | 8 |
| 2025 | WaveletFusion: enhancing plant leaf disease classification with multi-scale feature extraction and explainable AI
Lakshmi Srinivas Panchananam, Praveen Kumar Chandaliya, Zahid Akhtar, Kishor P. Upla, Ramachandra Raghavendra |
Expert Syst. Appl. | 3 |
| 2024 | Towards Inclusive Face Recognition Through Synthetic Ethnicity AlterationabstractNumerous studies have shown that existing Face Recognition Systems (FRS), including commercial ones, often exhibit biases toward certain ethnicities due to under-represented data. In this work, we explore ethnicity alteration and skin tone modification using synthetic face image generation methods to increase the diversity of datasets. We conduct a detailed analysis by first constructing a balanced face image dataset representing three ethnicities: Asian, Black, and Indian. We then make use of existing Generative Adversarial Network-based (GAN) image-to-image translation and manifold learning models to alter the ethnicity from one to another. A systematic analysis is further conducted to assess the suitability of such datasets for FRS by studying the realistic skin-tone representation using Individual Typology Angle (ITA). Further, we also analyze the quality characteristics using existing Face image quality assessment (FIQA) approaches. We then provide a holistic FRS performance analysis using four different systems. Our findings pave the way for future research works in (i) developing both specific ethnicity and general (any to any) ethnicity alteration models, (ii) expanding such approaches to create databases with diverse skin tones, (iii) creating datasets representing various ethnicities which further can help in mitigating bias while addressing privacy concerns. Praveen Kumar Chandaliya, Kiran B. Raja, Ramachandra Raghavendra, Zahid Akhtar, Christoph Busch 0001 |
FG | 4 |
| 2024 | Robust Facial Emotion Recognition System via De-Pooling Feature Enhancement and Weighted Exponential Moving Average
Khin Cho Win, Zahid Akhtar, C. Krishna Mohan |
TENCON | 2 |
| 2024 | Deepfake Detection Using Spatiotemporal TransformerabstractRecent advances in generative models and the availability of large-scale benchmarks have made deepfake video generation and manipulation easier. Nowadays, the number of new hyper-realistic deepfake videos used for negative purposes is dramatically increasing, thus creating the need for effective deepfake detection methods. Although many existing deepfake detection approaches, particularly CNN-based methods, show promising results, they suffer from several drawbacks. In general, poor generalization results have been obtained under unseen/new deepfake generation methods. The crucial reason for the above defect is that CNN-based methods focus on the local spatial artifacts, which are unique for every manipulation method. Therefore, it is hard to learn the general forgery traces of different manipulation methods without considering the dependencies that extend beyond the local receptive field. To address this problem, this article proposes a framework that combines Convolutional Neural Network (CNN) with Vision Transformer (ViT) to improve detection accuracy and enhance generalizability. Our method, namedHCiT, exploits the advantages of CNNs to extract meaningful local features, as well as the ViT’s self-attention mechanism to learn discriminative global contextual dependencies in a frame-level image explicitly. In this hybrid architecture, the high-level feature maps extracted from the CNN are fed into the ViT model that determines whether a specific video is fake or real. Experiments were performed on Faceforensics++, DeepFake Detection Challenge preview, Celeb datasets, and the results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, the HCiT method shows a great capacity for generalization on datasets covering various techniques of deepfake generation. The source code is available at: https://github.com/KADDAR-Bachir/HCiT Bachir Kaddar, Sid Ahmed Fezza, Zahid Akhtar, Wassim Hamidouche, Abdenour Hadid, Joan Serra-Sagristà |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Writer Identification from Nordic Historical Manuscripts using Transformer NetworksabstractHandwriting has been used as a form of authentication for the last 1000 years. Forensic analysis of handwriting using computers has been in practice since the 1970s. With the evolution of deep-learning techniques over the last decade, such automated forensic analysis of handwritten text has become dependent on deep-learning techniques. In this paper, we investigate the prowess of transformer networks in the context of identifying the writer of a handwritten sample. We here propose a deep feature embedding-based transformer network, WiT, for writer identification. Experiments were conducted on a historical Nordic manuscript dataset comprising 9253 handwritten samples scribbled by 50 writers for the very first time, and encouraging results were obtained. Rigorous experiments were also conducted to check the noise / damage resiliency of WiT, and the outcomes were quite promising. Chandranath Adak, Batturi Jaswanth, Zahid Akhtar, Andre Kåsen, Sukalpa Chanda |
IJCB | 3 |
| 2023 | Smart Navigation and Energy Management Framework for Autonomous Electric Vehicles in Complex EnvironmentsabstractAutonomous electric vehicles (AEVs) are revolutionizing the world of smart city transportation due to their low-resource consumption, improved traffic efficiency, zero carbon emissions, and improved road safety. To ensure the safe passage of vehicles through a complex environment, it is essential to plan for safe and smart navigation and energy management for AEVs. This demands an effective model for locating the optimal electric charging stations (ECSs) for scheduling and recharging the AEVs when they run on low battery. Many research works, however, do not focus on navigation and scheduling policies for AEV charging that would occur in extreme events in complex environments. This article puts forth a collaborative optimal navigation and charge planning (CONCP) framework based on multiagent deep reinforcement learning (MADRL). To ensure the safe passage of vehicles through the complex environment, it is essential to plan for safe and smart navigation and energy management for AEVs. The CONCP framework aims to achieve the best route from the origin to the final destination for each AEV, scheduling the optimal ECS while avoiding obstacles, reducing traffic congestion, and maximizing energy efficiency, accordingly. The experimental results indicate that CONCP achieves 27% higher success rates, 31% fewer collision rates, and 37% higher reward per episode than the other state-of-the-art algorithms. Gunasekaran Raja, Gayathri Saravanan, Sahaya Beni Prathiba, Zahid Akhtar, Sunder Ali Khowaja, Kapal Dev |
IEEE Internet Things J. | 4 |
| 2023 | Panoramic image generation using deep neural networks
Izat Khamiyev, Dias Issa, Zahid Akhtar, M. Fatih Demirci |
Soft Comput. | 3 |
| 2023 | Guest Editorial: Cybersecurity Intelligence in the Healthcare System
Abhinav Kumar 0005, Zahid Akhtar, Muhammad Khurram Khan |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Gabor filter bank with deep autoencoder based face recognition system
Rabah Hammouche, Abdelouahab Attia, Samir Akhrouf, Zahid Akhtar |
Expert Syst. Appl. | 4 |
| 2022 | Incorporating deep learning into capacitive images for smartphone user authentication
Md. Shafaeat Hossain, Mohammad Tariqul Islam 0002, Zahid Akhtar |
J. Inf. Secur. Appl. | 3 |
| 2022 | Deep learning-driven palmprint and finger knuckle pattern-based multimodal Person recognition system
Abdelouahab Attia, Sofiane Maza, Zahid Akhtar, Youssef Chahir |
Multim. Tools Appl. | 3 |
| 2022 | Contactless person recognition using 2D and 3D finger knuckle patterns
Mourad Chaa, Zahid Akhtar, Abdelhai Lati |
Multim. Tools Appl. | 2 |
| 2022 | Face based person recognition mechanism using monogenic Binarized Statistical Image Features
Nour Elhouda Chalabi, Abdelouahab Attia, Abderraouf Bouziane, Zahid Akhtar |
Multim. Tools Appl. | 4 |
| 2022 | Local features enhancement using deep auto-encoder scheme for the recognition of the proposed handwritten Arabic-Maghrebi characters database
Soumia Djaghbellou, Abdelouahab Attia, Abderraouf Bouziane, Zahid Akhtar |
Multim. Tools Appl. | 4 |
| 2022 | Dental biometric systems: a comparative study of conventional descriptors and deep learning-based features
Ayse Betül Oktay, Zahid Akhtar, Anil Gürses |
Multim. Tools Appl. | 2 |
| 2021 | CCAP: Cooperative Context Aware Pruning for Neural Network Model CompressionabstractIn this paper, we propose a new cross-domain model compression technique to yield a compact target model. We utilize a Cooperative Context-Aware Pruning (CCAP) module to produce sparse attention maps. They are then used to transmit the source models’ parameters to the target model precisely. We also leverage a weight-regular loss to minimize the difference between the source models’ and the target models’ parameters. Our quantitatively empirical evaluation shows that our CCAP module plus the weight-regular loss achieves lower model complexity without having serious performance decreasing. Li-Yun Wang, Zahid Akhtar |
ISM | 2 |
| 2021 | Parametric Audio-visual Quality Measurement Based on Interpretable Fuzzy LogicabstractIn this paper, a fuzzy inference based system for the quality of experience assessment is presented. In particular, two models have been proposed to assess the perceived quality of experience. The first one is a global model which estimates the overall audiovisual quality without going through the individual evaluations of visual and auditory modalities while the other model has been created by merging fuzzy inference systems based on separate auditory and visual objective quality scores. Two different sets of parameters have been tested for the second model leading to two different measures. The contribution of this research lies in the application of the fuzzy inference logic to estimate the audiovisual quality of multimedia data. The experimental results on a publicly available quality dataset show competitive predictive performances of the proposed measures when compared to two existing audiovisual metrics based on random forests regression and multilayer perceptron. Fatima Boudjerida, Atidel Lahoulou, Tiago H. Falk, Zahid Akhtar |
QoMEX | 4 |
| 2021 | HCiT: Deepfake Video Detection Using a Hybrid Model of CNN features and Vision TransformerabstractThe number of new falsified video contents is dramatically increasing, making the need to develop effective deepfake detection methods more urgent than ever. Even though many existing deepfake detection approaches show promising results, the majority of them still suffer from a number of critical limitations. In general, poor generalization results have been obtained under unseen or new deepfake generation methods. Consequently, in this paper, we propose a deepfake detection method called HCiT, which combines Convolutional Neural Network (CNN) with Vision Transformer (ViT). The HCiT hybrid architecture exploits the advantages of CNN to extract local information with the ViT's self-attention mechanism to improve the detection accuracy. In this hybrid architecture, the feature maps extracted from the CNN are feed into ViT model that determines whether a specific video is fake or real. Experiments were performed on Faceforensics++ and DeepFake Detection Challenge preview datasets, and the results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, the HCiT method shows a great capacity for generalization on datasets covering various techniques of deepfake generation. The source code is available at: https://github.com/KADDAR-Bachir/HCiT Bachir Kaddar, Sid Ahmed Fezza, Wassim Hamidouche, Zahid Akhtar, Abdenour Hadid |
VCIP | 4 |
| 2021 | DeepSmoke: Deep learning model for smoke detection and segmentation in outdoor environments
Salman Khan 0004, Khan Muhammad 0001, Tanveer Hussain 0001, Javier Del Ser, Fabio Cuzzolin, Siddhartha Bhattacharyya 0001, Zahid Akhtar, Victor Hugo C. de Albuquerque |
Expert Syst. Appl. | 7 |
| 2021 | 3D Palmprint recognition using Tan and Triggs normalization technique and GIST descriptors
Mourad Chaa, Zahid Akhtar |
Multim. Tools Appl. | 2 |
| 2021 | Particle swarm optimization based block feature selection in face recognition system
Nour Elhouda Chalabi, Abdelouahab Attia, Abderraouf Bouziane, Zahid Akhtar |
Multim. Tools Appl. | 4 |
| 2021 | Feature Pooling of Modulation Spectrum Features for Improved Speech Emotion Recognition in the WildabstractInterest in affective computing is burgeoning, in great part due to its role in emerging affective human-computer interfaces (HCI). To date, the majority of existing research on automated emotion analysis has relied on data collected in controlled environments. With the rise of HCI applications on mobile devices, however, so-called “in-the-wild” settings have posed a serious threat for emotion recognition systems, particularly those based on voice. In this case, environmental factors such as ambient noise and reverberation severely hamper system performance. In this paper, we quantify the detrimental effects that the environment has on emotion recognition and explore the benefits achievable with speech enhancement. Moreover, we propose a modulation spectral feature pooling scheme that is shown to outperform a state-of-the-art benchmark system for environment-robust prediction of spontaneous arousal and valence emotional primitives. Experiments on an environment-corrupted version of the RECOLA dataset of spontaneous interactions show the proposed feature pooling scheme, combined with speech enhancement, outperforming the benchmark across different noise-only, reverberation-only and noise-plus-reverberation conditions. Additional tests with the SEWA database show the benefits of the proposed method for in-the-wild applications. Anderson R. Avila, Zahid Akhtar, João Felipe Santos, Douglas D. O'Shaughnessy, Tiago H. Falk |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | Accelerating deep reinforcement learning model for game strategy
Yifan Li 0007, Yuchun Fang, Zahid Akhtar |
Neurocomputing | 3 |
| 2020 | Human Behavior Understanding in Big Multimedia Data Using CNN based Facial Expression Recognition
Sana Zahir, Amin Ullah, Zahid Akhtar, Khan Muhammad 0001 |
Mob. Networks Appl. | 4 |
| 2019 | Low Dose Abdominal CT Image Reconstruction: An Unsupervised Learning Based ApproachabstractIn medical practice, the X-ray Computed tomography-based scans expose a high radiation dose and lead to the risk of prostate or abdomen cancers. On the other hand, the low-dose CT scan can reduce radiation exposure to the patient. But the reduced radiation dose degrades image quality for human perception, and adversely affects the radiologist's diagnosis and prognosis. In this paper, we introduce a GAN based auto-encoder network to de-noise the CT images. Our network first maps CT images to low dimensional manifolds and then restore the images from its corresponding manifold representations. Our reconstruction algorithm separately calculates perceptual similarity, learns the latent feature maps, and achieves more accurate and visually pleasing reconstructions. We also showed the effectiveness of our model on a number of patient abdomen CT images, and compare our results with existing deep learning and iterative reconstruction methods. Experimental results demonstrate that our model outperforms other state-of-the-art methods in terms of PSNR, SSIM, and statistical properties of the image regions. https://github.com/ShibaPrasad/CT-Image-Reconstruction. Shiba Kuanar, Vassilis Athitsos, Dwarikanath Mahapatra, Kamisetty Ramamohan Rao, Zahid Akhtar, Dipankar Dasgupta |
ICIP | 5 |
| 2019 | Generalizable Adversarial Examples Detection Based on Bi-model Decision MismatchabstractModern applications of artificial neural networks have yielded remarkable performance gains in a wide range of tasks. However, recent studies have discovered that such modelling strategy is vulnerable to Adversarial Examples, i.e. examples with subtle perturbations often too small and imperceptible to humans, but that can easily fool neural networks. Defense techniques against adversarial examples have been proposed, but ensuring robust performance against varying or novel types of attacks remains an open problem. In this work, we focus on the detection setting, in which case attackers become identifiable while models remain vulnerable. Particularly, we employ the decision layer of independently trained models as features for posterior detection. The proposed framework does not require any prior knowledge of adversarial examples generation techniques, and can be directly employed along with unmodified off-the-shelf models. Experiments on the standard MNIST and CIFAR10 datasets deliver empirical evidence that such detection approach generalizes well across not only different adversarial examples generation methods but also quality degradation attacks. Non-linear binary classifiers trained on top of our proposed features can achieve a high detection rate (>90%) in a set of white-box attacks and maintain such performance when tested against unseen attacks. João Monteiro 0002, Isabela Albuquerque, Zahid Akhtar, Tiago H. Falk |
SMC | 3 |
| 2019 | 3D palmprint recognition using unsupervised convolutional deep learning network and SVM classifierabstractSince past decade, efforts are afoot to design better hand‐based automatic person recognition systems. Among the various hand‐based biometric traits, palmprint as a biometric characteristic is now gaining increased attention from both the academic and industrial communities owing to its highly distinctive texture patterns, features richness, and stability. Here, the authors propose a new 3D palmprint recognition framework based on an unsupervised convolutional deep learning network named PCANet. Specifically, the proposed framework first reconstructs illumination‐invariant 3D palmprint images using Single Scale Retinex (SSR) algorithm. Then, PCANet topology is employed to extract discriminative features from SSR images. Finally, a multi‐class support vector machine (SVM) classification scheme is utilised to determine the identity of the person. Extensive experimental analysis on publicly available 3D palmprint PolyU dataset, which is composed of 8000 range images from 200 individuals, shows that proposed method outperforms existing approaches and is also able to attain 99.98% rank‐1 accuracy. Mourad Chaa, Zahid Akhtar, Abdelouahab Attia |
IET Image Process. | 2 |
| 2018 | Improved Audio-Visual Laughter Detection Via Multi-Scale Multi-Resolution Image Texture Features and Classifier FusionabstractEfforts are afoot to design better context-aware human-computer interaction techniques that have knowledge of both their surrounding and the affective state of the user. One of the most important nonverbal behavioural cues for affective human-machine interaction is laughter. Automatic detection of laughter is an interesting, yet challenging problem, which in recent years has gained increased attention from both the academic and industrial communities. The majority of existing laughter detection systems rely on either audio or video modalities. Humans, however, typically rely on audio-visual cues during conversation and/or interaction, thus it is expected that improved results can be achieved if both modalities are used. In this work, we propose a multimodal framework that analyzes audio and video channels separately, then fuses their decisions. Conventional speech spectral and prosodic features are used, whereas new multi -scale multiresolution binarized statistical image features are proposed due to their improved expressive power. Experiments with the publicly available MAHNOB Laughter database show that decision level fusion based on support vector machine classifiers leads to improved performance over single modality approaches, as well as over previously-proposed methods, all whilst requiring just a fraction of the computational power. Zahid Akhtar, Stefany Bedoya, Tiago H. Falk |
ICASSP | 1 |
| 2017 | A competition on generalized software-based face presentation attack detection in mobile scenariosabstractIn recent years, software-based face presentation attack detection (PAD) methods have seen a great progress. However, most existing schemes are not able to generalize well in more realistic conditions. The objective of this competition is to evaluate and compare the generalization performances of mobile face PAD techniques under some real-world variations, including unseen input sensors, presentation attack instruments (PAI) and illumination conditions, on a larger scale OULU-NPU dataset using its standard evaluation protocols and metrics. Thirteen teams from academic and industrial institutions across the world participated in this competition. This time typical liveness detection based on physiological signs of life was totally discarded. Instead, every submitted system relies practically on some sort of feature representation extracted from the face and/or background regions using hand-crafted, learned or hybrid descriptors. Interesting results and findings are presented and discussed in this paper. Zinelabidine Boulkenafet, Jukka Komulainen, Zahid Akhtar, Azeddine Benlamoudi, Djamel Samai, Salah Eddine Bekhouche, Abdelkrim Ouafi, Fadi Dornaika, Abdelmalik Taleb-Ahmed, Fei Peng 0001, L. B. Zhang, Min Long 0003, Shruti Bhilare, Vivek Kanhangad, Artur Costa-Pazo, Esteban Vázquez-Fernández, Daniel Pérez-Cabo, J. J. Moreira-Perez, Daniel González-Jiménez, Amir Mohammadi, Sushil Bhattacharjee, Sébastien Marcel, Svetlana Volkova, N. Abe, X. Feng, Z. Xia, Rui Shao 0001, Pong C. Yuen, Waldir R. de Almeida, Fernanda A. Andaló, Rafael Padilha, Gabriel Bertocco, William Dias, Jacques Wainer, Ricardo da Silva Torres, Anderson Rocha 0001, Marcus A. Angeloni, Guilherme Folego, Alan Godoy, Abdenour Hadid |
IJCB | 3 |
| 2017 | BAUM-1: A Spontaneous Audio-Visual Face Database of Affective and Mental StatesabstractIn affective computing applications, access to labeled spontaneous affective data is essential for testing the designed algorithms under naturalistic and challenging conditions. Most databases available today are acted or do not contain audio data. We present a spontaneous audio-visual affective face database of affective and mental states. The video clips in the database are obtained by recording the subjects from the frontal view using a stereo camera and from the half-profile view using a mono camera. The subjects are first shown a sequence of images and short video clips, which are not only meticulously fashioned but also timed to evoke a set of emotions and mental states. Then, they express their ideas and feelings about the images and video clips they have watched in an unscripted and unguided way in Turkish. The target emotions, include the six basic ones (happiness, anger, sadness, disgust, fear, surprise) as well as boredom and contempt. We also target several mental states, which are unsure (including confused, undecided), thinking, concentrating, and bothered. Baseline experimental results on the BAUM-1 database show that recognition of affective and mental states under naturalistic conditions is quite challenging. The database is expected to enable further research on audio-visual affect and mental state recognition under close-to-real scenarios. Sara Zhalehpour, Onur Onder, Zahid Akhtar, Çigdem Eroglu Erdem |
IEEE Trans. Affect. Comput. | 3 |
| 2017 | Investigating Apache Hama: a bulk synchronous parallel computing framework
Kamran Siddique, Zahid Akhtar, Yangwoo Kim, Young-Sik Jeong, Edward J. Yoon |
J. Supercomput. | 2 |
| 2016 | Mobile ocular biometrics in visible spectrum using local image descriptors: A preliminary studyabstractOcular biometrics refers to personal identification using iris, conjunctival vasculature, periocular or eye movements. Contrary to most of other biometric traits, ocular biometrics does not require high user cooperation and close capture distance. Biometrics is now adopted ubiquitously as an alternative to passwords on mobile devices. Especially, ocular biometrics in the visible spectrum has attracted a lot of attention owing to the fact that it can be acquired using the regular RGB cameras already available in all mobile devices. The use of local image descriptors (i.e., analysis of microtextural features) for ocular biometrics is gaining more and more popularity because of their compactness, computationally inexpensiveness, excellent performance and flexibility. In this work, we explore the possibility of performing large scale mobile ocular biometric recognition in the visible spectrum using local image descriptors. We design a weighted fusion scheme to combine the information originating from four different local descriptors. The experimental analysis of the devised scheme, on newly collected and publicly available large scale database using three different mobile devices, shows promising results. Zahid Akhtar, Christian Micheloni, Gian Luca Foresti |
ICIP | 1 |
| 2014 | MoBio_LivDet: Mobile biometric liveness detectionabstractBiometric authentication is now being used ubiquitously as an alternative to passwords on mobile devices. However, current biometric systems are vulnerable to simple spoofing attacks. Several liveness detection methods have been proposed to determine whether there is a live person or an artificial replica in front of the biometric sensor. Yet, the problem is unsolved due to hardship in finding discriminative and computationally inexpensive features for spoofing attacks. Moreover, previous liveness detection approaches are not explicitly aimed for mobile biometric, thus principally unsuited for portable devices. Therefore, we build a software-based multi-biometric prototype that detects face, iris and fingerprint spoofing attacks on mobile devices. We present MoBio_LivDet (Mobile Biometric Liveness Detection), a novel approach that analyzes local features and global structures of the biometric images using a set of low-level feature descriptors and decision level fusion. The system allows user to balance the security level (robustness against spoofing) and convenience that they want. The proposed method is highly fast, simple, efficient, robust and does not require user-cooperation, thus making it extremely apt for mobile devices. Experimental analysis on publicly available face, iris and fingerprint data sets with real spoofing attacks show promising results. Zahid Akhtar, Christian Micheloni, Claudio Piciarelli, Gian Luca Foresti |
AVSS | 1 |
| 2014 | Multimodal emotion recognition with automatic peak frame selectionabstractIn this paper we present an effective framework for multimodal emotion recognition based on a novel approach for automatic peak frame selection from audio-visual video sequences. Given a video with an emotional expression, peak frames are the ones at which the emotion is at its apex. The objective of peak frame selection is to make the training process for the automatic emotion recognition system easier by summarizing the expressed emotion over a video sequence. The main steps of the proposed framework consists of extraction of video and audio features based on peak frame selection, unimodal classification and decision level fusion of audio and visual results. We evaluated the performance of our approach on eNTERFACE'05 audio-visual database containing six basic emotional classes. Experimental results demonstrate the effectiveness and superiority of the proposed system over other methods in the literature. Sara Zhalehpour, Zahid Akhtar, Çigdem Eroglu Erdem |
INISTA | 2 |
| 2011 | Robustness of multi-modal biometric verification systems under realistic spoofing attacksabstractRecent works have shown that multi-modal biometric systems are not robust against spoofing attacks [12, 15,13]. However, this conclusion has been obtained under the hypothesis of a "worst case" attack, where the attacker is able to replicate perfectly the genuine biometric traits. Aim of this paper is to analyse the robustness of some multi-modal verification systems, combining fingerprint and face bio-metrics, under realistic spoofing attacks, in order to investigate the validity of the results obtained under the worst-case attack assumption. Battista Biggio, Zahid Akhtar, Giorgio Fumera, Gian Luca Marcialis, Fabio Roli |
IJCB | 2 |