VLDB 2026 Research / reviewers in the wild / expert
Abhijit Das 0001
dblp:73/1871-1
· DBLP profile ↗
36ranked-venue papers
15as first author
22since 2021 · last 2026
0000-0002-6793-0582ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 9 first-author · 18 since 2021Artificial intelligence and machine learning · 23 · 12 first-author · 11 since 2021Security and privacy · 12 · 4 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 10 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Narrating For You: Prompt-guided Audio-visual Narrating Face Generation Employing Multi-entangled Latent SpaceabstractWe present a novel approach for generating realistic speaking and talking faces by synthesizing a person’s voice and facial movements from a static image, a voice profile, and a target text. The model encodes the prompt/driving text, the driving image, and the voice profile of an individual and then combines them to pass them to the multi-entangled latent space to foster key-value pairs and queries for the audio and video modality generation pipeline. The multi-entangled latent space is responsible for establishing the spatiotemporal person-specific features between the modalities. Further, entangled features are passed to the respective decoder of each modality for output audio and video generation. Our experiments and analysis through standard metrics demonstrate the effectiveness of our model. All model checkpoints, code, and the proposed dataset can be found at https://github.com/narratingForYou/NarratingForYou. Aashish Chandra K, Aashutosh A V, Abhijit Das 0001 |
WACV | 3 |
| 2025 | ViM-Disparity: Bridging the Gap of Speed, Accuracy and Memory for Disparity Map GenerationabstractIn this work we propose a Visual Mamba (ViM) based architecture, to dissolve the existing trade-off for real-time and accurate model with low computation overhead for disparity map generation (DMG). Moreover, we proposed a performance measure that can jointly evaluate the inference speed, computation overhead and the accurateness of a DMG model. The code implementation and corresponding models are available at: https://github.com/MBora/ViM-Disparity. Maheswar Bora, Tushar Anand, Saurabh Atreya, Aritra Mukherjee, Abhijit Das 0001 |
ICASSP | 5 |
| 2025 | Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake DetectionabstractWhile Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored. In this work, we aim to explore the potential of MLLMs for audio deepfake detection. Combining audio inputs with a range of text prompts as queries to find out the viability of MLLMs to learn robust representations across modalities for audio deepfake detection. Therefore, we attempt to explore text-aware and context-rich, question-answer based prompts with binary decisions. We hypothesise that such a feature-guided reasoning will help in facilitating deeper multimodal understanding and enable robust feature learning for audio deepfake detection. We evaluate the performance of two MLLMs, Qwen2-Audio-7B-Instruct and SALMONN, in two evaluation modes: (a) zero-shot and (b) fine-tuned. Our experiments demonstrate that combining audio with a multi-prompt approach could be a viable way forward for audio deepfake detection. Our experiments show that the models perform poorly without task-specific training and struggle to generalise to out-of-domain data. However, they achieve good performance on in-domain data with minimal supervision, indicating promising potential for audio deepfake detection. Akanksha Chuchra, Shukesh Reddy, Sudeepta Mishra, Abhijit Das 0001, Abhinav Dhall |
IJCB | 4 |
| 2025 | Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025abstractThis paper presents a summary of the 2025 Sclera Segmentation Benchmarking Competition (SSBC), which focused on the development of privacy-preserving sclera-segmentation models trained using synthetically generated ocular images. The goal of the competition was to evaluate how well models trained on synthetic data perform in comparison to those trained on real-world datasets. The competition featured two tracks: (i) one relying solely on synthetic data for model development, and (ii) one combining/mixing synthetic with (a limited amount of) real-world data. A total of nine research groups submitted diverse segmentation models, employing a variety of architectural designs, including transformer-based solutions, lightweight models, and segmentation networks guided by generative frameworks. Experiments were conducted across three evaluation datasets containing both synthetic and real-world images, collected under diverse conditions. Results show that models trained entirely on synthetic data can achieve competitive performance, particularly when dedicated training strategies are employed, as evidenced by the top performing models that achieved F1scores of over 0.8 in the synthetic data track. Moreover, performance gains in the mixed track were often driven more by methodological choices rather than by the inclusion of real data, highlighting the promise of synthetic data for privacy-aware biometric development. The code and data for the competition is available at: https://github.com/dariant/SSBC_2025. Matej Vitek, Darian Tomasevic, Abhijit Das 0001, Sabari Nathan, Gökhan Özbulak, G. A. T. Özbulak, Jean-Paul Calbimonte, André Anjos, Hariohm Hemant Bhatt, Dhruv Dhirendra Premani, Jay Chaudhari, Caiyong Wang, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Divya Velayudan, Maregu Assefa, Naoufel Werghi, Zachary A. Daniels, Leeon John, Ritesh Vyas, Jalil Nourmohammadi Khiarak, Taher Akbari Saeed, Mahsa Nasehi, Ali Kianfar, Mobina Pashazadeh Panahi, Geetanjali Sharma, Pushp Raj Panth, Ramachandra Raghavendra, Aditya Nigam, Umapada Pal 0001, Peter Peer, Vitomir Struc |
IJCB | 3 |
| 2025 | KDC-MAE: Knowledge Distilled Contrastive Mask Auto-EncoderabstractIn this work, we attempted to extend the thought and showcase a way forward for the Self-supervised Learning (SSL) learning paradigm by combining contrastive learning, self-distillation (knowledge distillation) and masked data modelling, the three major SSL frameworks, to learn a joint and coordinated representation. The proposed technique of SSL learns by the collaborative power of different learning objectives of SSL. Hence to jointly learn the different SSL objectives we proposed a new SSL architecture KDC-MAE, a complementary masking strategy to learn the modular correspondence, and a weighted way to combine them coordinately. Experimental results conclude that the contrastive masking correspondence along with the KD learning objective has lent a hand to performing better learning for multiple modalities over multiple tasks. Maheswar Bora, Saurabh Atreya, Aritra Mukherjee, Abhijit Das 0001 |
WACV | 4 |
| 2024 | INDIFACE: Illuminating India's Deepfake Landscape with a Comprehensive Synthetic DatasetabstractDue to the recent progress in Deepfake generation, several datasets and manipulation techniques have been proposed in the recent literature with various effective face-swap and face-reenactment methods. Deepfake is an emerging threat to society and government as it can jeopardize law enforcement and cause personal loss. Investigations in the literature established that demographic variation had impacted the performance of Deepfake detection. To date, Deepfake detection has not been studied in the Indian context; hence, in this work, we proposed a Deepfake dataset INDIFACE entirely with Indian subjects. We have collected 101 original videos and used two different manipulation techniques for Deepfake generation. We provide detailed benchmarking with state-of-the-art methods on Deepfake datasets, showcasing that the existing model is insufficient to detect Deepfake detection for the Indian scenario. Hence, more attention is required to this area of research. The proposed dataset INDIFACE is publicly available at. Kartik Kuckreja, Ximi Hoque, Nishit Poddar, Shukesh Reddy, Abhinav Dhall, Abhijit Das 0001 |
FG | 6 |
| 2024 | Limited Data, Unlimited Potential: A Study on ViTs Augmented by Masked AutoencodersabstractVision Transformers (ViTs) have become ubiquitous in computer vision. Despite their success, ViTs lack inductive biases, which can make it difficult to train them with limited data. To address this challenge, prior studies suggest training ViTs with self-supervised learning (SSL) and fine-tuning sequentially. However, we observe that jointly optimizing ViTs for the primary task and a Self-Supervised Auxiliary Task (SSAT) is surprisingly beneficial when the amount of training data is limited. We explore the appropriate SSL tasks that can be optimized alongside the primary task, the training schemes for these tasks, and the data scale at which they can be most effective. Our findings reveal that SSAT is a powerful technique that enables ViTs to leverage the unique characteristics of both the self-supervised and primary tasks, achieving better performance than typical ViTs pre-training with SSL and fine-tuning sequentially. Our experiments, conducted on 10 datasets, demonstrate that SSAT significantly improves ViT performance while reducing carbon footprint. We also confirm the effectiveness of SSAT in the video domain for deepfake detection, showcasing its generalizability. Our code is available at https://github.com/dominickrei/Limited-data-vits. Srijan Das, Tanmay Jain, Dominick Reilly, Pranav Balaji, Soumyajit Karmakar, Shyam Marjit, Xiang Li 0109, Abhijit Das 0001, Michael S. Ryoo |
WACV | 8 |
| 2024 | A novel multi-task learning technique for offline handwritten short answer spotting and recognition
Abhijit Das 0001, Hemmaphan Suwanwiwat, Umapada Pal 0001 |
Multim. Tools Appl. | 1 |
| 2023 | Enhancing 3D-Air Signature by Pen Tip Tail Trajectory Awareness: Dataset and Featuring by Novel Spatio-temporal CNNabstractThis work proposes a novel process of using pen tip and tail 3D trajectory for air signature. To acquire the trajectories we developed a new pen tool and a stereo camera was used. We proposed SliT-CNN, a novel 2D spatial-temporal convolutional neural network (CNN) for better featuring of the air signature. In addition, we also collected an air signature dataset from 45 signers. Skilled forgery signatures per user are also collected. A detailed benchmarking of the proposed dataset using existing techniques and proposed CNN on existing and proposed dataset exhibit the effectiveness of our methodology. Saurabh Atreya, Maheswar Bora, Aritra Mukherjee, Abhijit Das 0001 |
IJCB | 4 |
| 2023 | Sclera Segmentation and Joint Recognition Benchmarking Competition: SSRBC 2023abstractThis paper presents the summary of the Sclera Segmentation and Joint Recognition Benchmarking Competition (SSRBC 2023) held in conjunction with IEEE International Joint Conference on Biometrics (IJCB 2023). Different from the previous editions of the competition, SSRBC 2023 not only explored the performance of the latest and most advanced sclera segmentation models, but also studied the impact of segmentation quality on recognition performance. Five groups took part in SSRBC 2023 and submitted a total of six segmentation models and one recognition technique for scoring. The submitted solutions included a wide variety of conceptually diverse deep-learning models and were rigorously tested on three publicly available datasets, i.e., MASD, SBVPI and MOBIUS. Most of the segmentation models achieved encouraging segmentation and recognition performance. Most importantly, we observed that better segmentation results always translate into better verification performance. Abhijit Das 0001, Saurabh Atreya, Aritra Mukherjee, Matej Vitek, Caiyong Wang, Guangzhe Zhao, Fadi Boutros, Patrick Siebke, Jan Niklas Kolf, Naser Damer, Sun Ye, Lu Hexin, Fan Aobo, You Sheng, Sabari Nathan, R. Suganya 0001, Rampriya Rajendran Shanthi, Geetanjali Sharma, P. Priyanka, Aditya Nigam, Peter Peer, Umapada Pal 0001, Vitomir Struc |
IJCB | 1 |
| 2023 | Recent Advancement in 3D Biometrics using Monocular CameraabstractRecent literature has witnessed significant interest towards 3D biometrics employing monocular vision for robust authentication methods. Motivated by this, in this work we seek to provide insight on recent development in the area of 3D biometrics employing monocular vision. We present the similarity and dissimilarity of 3D monocular biometrics and classical biometrics, listing the strengths and challenges. Further, we provide an overview of recent techniques in 3D biometrics with monocular vision, as well as application systems adopted by the industry. Finally, we discuss open research problems in this area of research. Aritra Mukherjee, Abhijit Das 0001 |
IJCB | 2 |
| 2023 | Depth-guided Robust Face Morphing Attack DetectionabstractRecently, morphing attack detection (MAD) solutions have achieved remarkable success with the aid of deep learning techniques. Despite the good performance achieved by binary label or binary pixel-wise supervised MAD models, the robustness of such models drops when facing variations in morphing attacks. In this work, we propose a novel process that leverages facial depth information to build a robust and generalized MAD. The depth map, representing the 3D shape of the face in a 2D image, is more informative compared to binary and binary pixel-wise map labels. To validate the idea we synthetically generated 3D depth map ground truth. Furthermore, we introduce a novel MAD architecture designed to capture subtle information from the 3D depth data. In addition, we analyze the training loss formulation to further enhance the MAD performance. Driven by the need for developing MAD solutions while preserving the privacy of individuals for legal and ethical reasons, we conduct our experiments on privacy-friendly synthetic training data and authentic evaluation data. The experimental results on existing public datasets in SYN-MAD 22 competition demonstrate the effectiveness of our proposed solution in terms of both robustness and generalization. Harsh Rachalwar, Meiling Fang, Naser Damer, Abhijit Das 0001 |
IJCB | 4 |
| 2023 | Face attribute analysis from structured light: an end-to-end approach
Vikas Thamizharasan, Abhijit Das 0001, Daniele Battaglino, François Brémond, Antitza Dantcheva |
Multim. Tools Appl. | 2 |
| 2023 | Exploring Bias in Sclera Segmentation Models: A Group Evaluation ApproachabstractBias and fairness of biometric algorithms have been key topics of research in recent years, mainly due to the societal, legal and ethical implications of potentially unfair decisions made by automated decision-making models. A considerable amount of work has been done on this topic across different biometric modalities, aiming at better understanding the main sources of algorithmic bias or devising mitigation measures. In this work, we contribute to these efforts and present the first study investigating bias and fairness of sclera segmentation models. Although sclera segmentation techniques represent a key component of sclera-based biometric systems with a considerable impact on the overall recognition performance, the presence of different types of biases in sclera segmentation methods is still underexplored. To address this limitation, we describe the results of a group evaluation effort (involving seven research groups), organized to explore the performance of recent sclera segmentation models within a common experimental framework and study performance differences (and bias), originating from various demographic as well as environmental factors. Using five diverse datasets, we analyze seven independently developed sclera segmentation models in different experimental configurations. The results of our experiments suggest that there are significant differences in the overall segmentation performance across the seven models and that among the considered factors, ethnicity appears to be the biggest cause of bias. Additionally, we observe that training with representative and balanced data does not necessarily lead to less biased results. Finally, we find that in general there appears to be a negative correlation between the amount of bias observed (due to eye color, ethnicity and acquisition device) and the overall segmentation performance, suggesting that advances in the field of semantic segmentation may also help with mitigating bias. Matej Vitek, Abhijit Das 0001, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Jalil Nourmohammadi-Khiarak, Mohsen Akbari Shahpar, Meysam Asgari-Chenaghlu, Farhang Jaryani, Juan E. Tapia, Andres Valenzuela, Caiyong Wang, Yunlong Wang 0003, Zhaofeng He 0001, Zhenan Sun, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Kiran B. Raja, Gourav Gupta, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, S. V. Aruna Kumar, B. S. Harish, Umapada Pal 0001, Peter Peer, Vitomir Struc |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Lip as biometric and beyond: a survey
Debbrota Paul Chowdhury, Ritu Kumari, Sambit Bakshi, Manmath Narayan Sahoo, Abhijit Das 0001 |
Multim. Tools Appl. | 5 |
| 2022 | A Spatio-Temporal Approach for Apathy ClassificationabstractApathy is characterized by symptoms such as reduced emotional response, lack of motivation, and limited social interaction. Current methods for apathy diagnosis require the patient’s presence in a clinic and time consuming clinical interviews, which are costly and inconvenient for both, patients and clinical staff, hindering among other large-scale diagnostics. In this work, we propose a novel spatio-temporal framework for apathy classification, which is streamlined to analyze facial dynamics and emotion in videos. Specifically, we divide the videos into smaller clips, and proceed to extract associated facial dynamics and emotion-based features. Statistical representations/descriptors based on each feature and clip serve as input of the proposed Gated Recurrent Unit (GRU)-architecture. Temporal representations of individual features at the lower level of the proposed architecture are combined at deeper layers of the proposed GRU architecture, in order to obtain the final feature-set for apathy classification. Based on extensive experiments, we show that fusion of characteristics such as emotion and facial dynamics in proposed deep-bi-directional GRU obtains an accuracy of 95.34% in apathy classification. Abhijit Das 0001, Xuesong Niu, Antitza Dantcheva, S. L. Happy, Hu Han 0001, Radia Zeghari, Philippe Robert, Shiguang Shan, François Brémond, Xilin Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Demystifying Attention Mechanisms for Deepfake DetectionabstractManipulated images and videos, i.e., deepfakes have become increasingly realistic due to the tremendous progress of deep learning methods. However, such manipulation has triggered social concerns, necessitating the introduction of robust and reliable methods for deepfake detection. In this work, we explore a set of attention mechanisms and adapt them for the task of deepfake detection. Generally, attention mechanisms in videos modulate the representation learned by a convolutional neural network (CNN) by focusing on the salient regions across space-time. In our scenario, we aim at learning discriminative features to take into account the temporal evolution of faces to spot manipulations. To this end, we address the two research questions ‘How to use attention mechanisms?’, and ‘What type of attention is effective for the task of deepfake detection?’ Towards answering these questions, we provide a detailed study and experiments on videos tampered by four manipulation techniques, as included in the FaceForensics++ dataset. We investigate three scenarios, where the networks are trained to detect (a) all manipulated videos, (b) each manipulation technique individually, as well as (c) the veracity of videos pertaining to manipulation techniques not included in the train set. Abhijit Das 0001, Srijan Das, Antitza Dantcheva |
FG | 1 |
| 2021 | BVPNet: Video-to-BVP Signal Prediction for Remote Heart Rate EstimationabstractIn this paper, we propose a new method for remote photoplethysmography (rPPG) based heart rate (HR) estimation. In particular, our proposed method BVPNet is streamlined to predict the blood volume pulse (BVP) signals from face videos. Towards this, we firstly define ROIs based on facial landmarks and then extract the raw temporal signal from each ROI. Then the extracted signals are pre-processed via first-order difference and Butterworth filter and combined to form a Spatial-Temporal map (STMap). We then propose to revise U-Net, in order to predict BVP signals from the STMap. BVPNet takes into account both temporal and frequency domain losses in order to learn better than conventional models. Our experimental results suggest that our BVPNet outperforms the state-of-the-art methods on two publicly available datasets (MMSE-HR and VIPL-HR). Abhijit Das 0001, Hao Lu 0009, Hu Han 0001, Antitza Dantcheva, Shiguang Shan, Xilin Chen 0001 |
FG | 1 |
| 2021 | ICDAR 2021 Competition on Script Identification in the Wild
Abhijit Das 0001, Miguel A. Ferrer, Aythami Morales, Moisés Díaz Cabrera, Umapada Pal 0001, Donato Impedovo, Wentao Yang 0003, Kensho Ota, Tadahito Yao, Le Quang Hung, Nguyen Quoc Cuong, Seungjae Kim, Abdeljalil Gattal |
ICDAR (4) | 1 |
| 2021 | Editorial to special issue on novel insights on ocular biometrics
Maria De Marsico, Hugo Proença 0001, Sambit Bakshi, Abhijit Das 0001 |
Image Vis. Comput. | 4 |
| 2021 | Benchmarked multi-script Thai scene text dataset and its multi-class detection solution
Hemmaphan Suwanwiwat, Abhijit Das 0001, Umapada Pal 0001 |
Multim. Tools Appl. | 2 |
| 2021 | The P-DESTRE: A Fully Annotated Dataset for Pedestrian Detection, Tracking, and Short/Long-Term Re-Identification From Aerial DevicesabstractOver the years, unmanned aerial vehicles (UAVs) have been regarded as a potential solution to surveil public spaces, providing a cheap way for data collection, while covering large and difficult-to-reach areas. This kind of solutions can be particularly useful to detect, track and identify subjects of interest in crowds, for security/safety purposes. In this context, various datasets are publicly available, yet most of them are only suitable for evaluating detection, tracking and short-term re-identification techniques. This paper announces the free availability of the P-DESTRE dataset, the first of its kind to provide video/UAV-based data for pedestrian long-term re-identification research, with ID annotations consistent across data collected in different days. As a secondary contribution, we provide the results attained by the state-of-the-art pedestrian detection, tracking, short/long term re-identification techniques in well-known surveillance datasets, used as baselines for the corresponding effectiveness observed in the P-DESTRE data. This comparison highlights the discriminating characteristics of P-DESTRE with respect to similar sets. Finally, we identify the most problematic data degradation factors and co-variates for UAV-based automated data analysis, which should be considered in subsequent technologic/conceptual advances in this field. The dataset and the full specification of the empirical evaluation carried out are freely available at http://p-destre.di.ubi.pt/. S. V. Aruna Kumar, Ehsan Yaghoubi, Abhijit Das 0001, B. S. Harish, Hugo Proença 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Apathy Classification by Exploiting Task RelatednessabstractApathy is characterized by symptoms such as reduced emotional response, lack of motivation, and limited social interaction. Current methods for apathy diagnosis require the patient's presence in a clinic and time consuming clinical interviews, which are costly and inconvenient for both patients and clinical staff, hindering among others large-scale diagnostics. In this work we propose a multi-task learning (MTL) framework for apathy classification based on facial analysis, entailing both emotion and facial movements. In addition, it leverages information from other auxiliary tasks (i.e., clinical scores), which might be closely or distantly related to the main task of apathy classification. Our proposed MTL approach (termed MTL+) improves apathy classification by jointly learning model weights and the relatedness of the auxiliary tasks to the main task in an iterative manner. Our results on 90 video sequences acquired from 45 subjects obtained an apathy classification accuracy of up to 80%, using the concatenated emotion and motion features. Our results further demonstrate the improved performance of MTL+ over MTL. S. L. Happy, Antitza Dantcheva, Abhijit Das 0001, François Brémond, Radia Zeghari, Philippe Robert |
FG | 3 |
| 2020 | SSBC 2020: Sclera Segmentation Benchmarking Competition in the Mobile EnvironmentabstractThe paper presents a summary of the 2020 Sclera Segmentation Benchmarking Competition (SSBC), the 7th in the series of group benchmarking efforts centred around the problem of sclera segmentation. Different from previous editions, the goal of SSBC 2020 was to evaluate the performance of sclera-segmentation models on images captured with mobile devices. The competition was used as a platform to assess the sensitivity of existing models to i) differences in mobile devices used for image capture and ii) changes in the ambient acquisition conditions. 26 research groups registered for SSBC 2020, out of which 13 took part in the final round and submitted a total of 16 segmentation models for scoring. These included a wide variety of deep-learning solutions as well as one approach based on standard image processing techniques. Experiments were conducted with three recent datasets. Most of the segmentation models achieved relatively consistent performance across images captured with different mobile devices (with slight differences across devices), but struggled most with low-quality images captured in challenging ambient conditions, i.e., in an indoor environment and with poor lighting. Matej Vitek, Abhijit Das 0001, Yann Pourcenoux, Alexandre Missler, C. Paumier, Sumanta Das, Ishita De Ghosh, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Junxing Hu, Yong He 0009, Caiyong Wang, Yunlong Wang 0003, Zhenan Sun, Dailé Osorio Roig, Christian Rathgeb, Christoph Busch 0001, Juan E. Tapia, Andres Valenzuela, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, Sabari Nathan, R. Suganya 0001, Vineet Mehta, Abhinav Dhall, Kiran B. Raja, Gourav Gupta, Jalil Nourmohammadi-Khiarak, Mohsen Akbari-Shahper, Farhang Jaryani, Meysam Asgari-Chenaghlu, Ritesh Vyas, Sristi Dakshit, Peter Peer, Umapada Pal 0001, Vitomir Struc |
IJCB | 2 |
| 2020 | ICFHR 2020 Competition on Short answer ASsessment and Thai student SIGnature and Name COMponents Recognition and Verification (SASIGCOM 2020)abstractThis paper describes the results of the competition on Short answer ASsessment and Thai student SIGnature and Name COMponents Recognition and Verification (SASIGCOM 2020) in conjunction with the 17th International Conference on Frontiers in Handwriting Recognition (ICFHR 2020). The competition was aimed to automate the evaluation process short answer-based examination and record the development and gain attention to such system. The proposed competition contains three elements which are short answer assessment (recognition and marking the answers to short-answer questions derived from examination papers), student name components (first and last names) and signature verification and recognition. Signatures and name components data were collected from 100 volunteers. For the Thai signature dataset, there are 30 genuine signatures, 12 skilled and 12 simple forgeries for each writer. With Thai name components dataset, there are 30 genuine and 12 skilfully forged name components for each writer. There are 104 exam papers in the short answer assessment dataset, 52 of which were written with cursive handwriting; the rest of 52 papers were written with printed handwriting. The exam papers contain ten questions, and the answers to the questions were designed to be a few words per question. Three teams from distinguished labs submitted their systems. For short answer assessment, word spotting task was also performed. This paper analysed the results produced by their algorithms using a performance measure and defines a way forward for this subject of research. Both the datasets, along with some of the accompanying ground truth/baseline mask will be made freely available for research purposes via the TC10/TC11. Abhijit Das 0001, Hemmaphan Suwanwiwat, Umapada Pal 0001, Michael Blumenstein |
ICFHR | 1 |
| 2019 | Characterizing the State of Apathy with Facial Expression and Motion AnalysisabstractReduced emotional response, lack of motivation, and limited social interaction comprise the major symptoms of apathy. Current methods for apathy diagnosis require the patient's presence in a clinic, and time consuming clinical interviews and questionnaires involving medical personnel, which are costly and logistically inconvenient for patients and clinical staff, hindering among other large scale diagnostics. In this paper we introduce a novel machine learning framework to classify apathetic and non-apathetic patients based on analysis of facial dynamics, entailing both emotion and facial movement. Our approach caters to the challenging setting of current apathy assessment interviews, which include short video clips with wide face pose variations, very low-intensity expressions, and insignificant inter-class variations. We test our algorithm on a dataset consisting of 90 video sequences acquired from 45 subjects and obtained an accuracy of 84% in apathy classification. Based on extensive experiments, we show that the fusion of emotion and facial local motion produces the best feature set for apathy classification. In addition, we train regression models to predict the clinical scores related to the mental state examination (MMSE) and the neuropsychiatric apathy inventory (NPI) using the motion and emotion features. Our results suggest that the performance can be further improved by appending the predicted clinical scores to the video-based feature representation. S. L. Happy, Antitza Dantcheva, Abhijit Das 0001, Radia Zeghari, Philippe Robert, François Brémond |
FG | 3 |
| 2019 | Robust Remote Heart Rate Estimation from Face Utilizing Spatial-temporal AttentionabstractIn this work, we propose an end-to-end approach for robust remote heart rate (HR) measurement gleaned from facial videos. Specifically the approach is based on remote photoplethysmography (rPPG), which constitutes a pulse triggered perceivable chromatic variation, sensed in RGB-face videos. Consequently, rPPGs can be affected in less-constrained settings. To unpin the shortcoming, the proposed algorithm utilizes a spatio-temporal attention mechanism, which places focus on the salient features included in rPPG-signals. In addition, we propose an effective rPPG augmentation approach, generating multiple rPPG signals with varying HRs from a single face video. Experimental results on the public datasets VIPL-HR and MMSE-HR show that the proposed method outperforms state-of-the-art algorithms in remote HR estimation. Xuesong Niu, Xingyuan Zhao, Hu Han 0001, Abhijit Das 0001, Antitza Dantcheva, Shiguang Shan, Xilin Chen 0001 |
FG | 4 |
| 2018 | ICFHR 2018 Competition on Thai Student Signatures and Name Components Recognition and Verification (TSNCRV2018)abstractThis paper summarises the results of the competition on the 1st Thai Student Signature and Name Components Recognition and Verification (TSNCRV 2018). It was organised in the context of the 16th International Conference on Frontiers in Handwriting Recognition (ICFHR 2018). The aim of this competition was to record the development and gain attention on Thai student signatures and name component recognition and verification. Two different types of datasets were used for the competition: the first dataset contains Thai student signatures and the second dataset contains Thai student name components. Signatures and name components from 100 volunteers each were included in the competition datasets. For Thai signature dataset, there are 30 genuine signatures, 12 skilled and 12 simple forgeries for each writer. For Thai name components, there are 30 genuine and 12 skilfully forged name components for each writer. For both the datasets the individuals were asked to write their name/signature in the given space on a white piece of paper for number of time (with a pause between each 10 samples). The skilled forgers were asked practice emitting the original signature for certain number of times till they fill skilled to forge. Five teams from distinguish labs submitted their systems. This paper analysed the results produced by these algorithms/systems using a performance measure and defined a way forward for this subject of research. Both the datasets along with some of the accompanying ground truth/baseline mask will be made freely available for research purposes via the TC10/TC11. Hemmaphan Suwanwiwat, Abhijit Das 0001, Umapada Pal 0001, Michael Blumenstein |
ICFHR | 2 |
| 2018 | More Realistic and Efficient Face-Based Mobile Authentication using CNNsabstractIn this work, we propose a more realistic and efficient face-based mobile authentication technique using CNNs. This paper discusses and explores an inevitable problem of using face images for mobile authentication, taken from varying distances with a front/selfie camera of the mobile phone. Incidentally, once an individual comes towards a certain distance from the camera, the face images get large and appear over-sized. Simultaneously sharp features of some portions of the face, such as forehead, cheek, and chin are changed completely. As a result, the face features change and the impact increases exponentially once the individual crosses a certain distance and gradually approaches towards the front camera. This work proposes a solution (achieving better accuracy and facial features, whereby face images were cropped and aligned around its close bounding box) to mitigate the aforementioned identified gap. The work investigated different frontier face detection and recognition techniques to justify the proposed solution. Among all the employed methods evaluated, CNNs worked best. For a quantitative comparison of the proposed method, manually cropped face images/annotations of the face images along with their close boundary were prepared. In turn, we have developed a database considering the above-mentioned scenario for 40 individuals, which will be publicly available for academic research purposes. The experimental results achieved indicate a successful implementation of the proposed method and the performance of the proposed technique is also found to be superior in comparison to the existing state-of-the-art. Abhijit Das 0001, Abira Sengupta, Umapada Pal 0001, Michael Blumenstein |
IJCNN | 1 |
| 2017 | A decision-level fusion strategy for multimodal ocular biometric in visible spectrum based on posterior probabilityabstractIn this work, we propose a posterior probability-based decision-level fusion strategy for multimodal ocular biometric in the visible spectrum employing iris, sclera and peri-ocular trait. To best of our knowledge this is the first attempt to design a multimodal ocular biometrics using all three ocular traits. Employing all these traits in combination can help to increase the reliability and universality of the system. For instance in some scenarios, the sclera and iris can be highly occluded or for completely closed eyes scenario, the peri-ocular trait can be relied on for the decision. The proposed system is constituted of three independent traits and their combinations. The classification output of the trait which produces highest posterior probability is to consider as the final decision. An appreciable reliability and universal applicability of ocular trait are achieved in experiments conducted employing the proposed scheme. Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein |
IJCB | 1 |
| 2017 | SSERBC 2017: Sclera segmentation and eye recognition benchmarking competitionabstractThis paper summarises the results of the Sclera Segmentation and Eye Recognition Benchmarking Competition (SSERBC 2017). It was organised in the context of the International Joint Conference on Biometrics (IJCB 2017). The aim of this competition was to record the recent developments in sclera segmentation and eye recognition in the visible spectrum (using iris, sclera and peri-ocular, and their fusion), and also to gain the attention of researchers on this subject. In this regard, we have used the Multi-Angle Sclera Dataset (MASD version 1). It is comprised of2624 images taken from both the eyes of 82 identities. Therefore, it consists of images of 164 (82×2) eyes. A manual segmentation mask of these images was created to baseline both tasks. Precision and recall based statistical measures were employed to evaluate the effectiveness of the segmentation and the ranks of the segmentation task. Recognition accuracy measure has been employed to measure the recognition task. Manually segmented sclera, iris and peri-ocular regions were used in the recognition task. Sixteen teams registered for the competition, and among them, six teams submitted their algorithms or systems for the segmentation task and two of them submitted their recognition algorithm or systems. The results produced by these algorithms or systems reflect current developments in the literature of sclera segmentation and eye recognition, employing cutting edge techniques. The MASD version 1 dataset with some of the ground truth will be freely available for research purposes. The success of the competition also demonstrates the recent interests of researchers from academia as well as industry on this subject. Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein, Dejan Stepec, Peter Rot, Ziga Emersic, Peter Peer, Vitomir Struc, S. V. Aruna Kumar, B. S. Harish |
IJCB | 1 |
| 2017 | Linking face images captured from the optical phenomenon in the wild for forensic scienceabstractThis paper discusses the possibility of use of some challenging face images scenario captured from optical phenomenon in the wild for forensic purpose towards individual identification. Occluded and under cover face images in surveillance scenario can be collected from its reflection on a surrounding glass or on a smooth wall that is under the coverage of the surveillance camera and such scenario of face images can be linked for forensic purposes. Another similar scenario that can also be used for forensic is the face images of an individual standing behind a transparent glass wall. To investigate the capability of these images for personal identification this study is conducted. This work investigated different types of features employed in the literature to establish individual identification by such degraded face images. Among them, local region based featured worked best. To achieve higher accuracy and better facial features face image were cropped manually along its close bounding box and noise removal was performed (reflection, etc.). In order to experiment we have developed a database considering the above mentioned scenario, which will be publicly available for academic research. Initial investigation substantiates the possibility of using such face images for forensic purpose. Abhijit Das 0001, Abira Sengupta, Miguel A. Ferrer, Umapada Pal 0001, Michael Blumenstein |
IJCB | 1 |
| 2016 | Fast and efficent multimodal eye biometrics using projective dictionary pair learningabstractThis work proposes a projective pairwise dictionary learning-based approach for fast and efficient multimodal eye biometrics. The work uses a faster Projective pairwise Discriminative Dictionary Learning (DL) in contrast to the traditional DL which uses synthesis DL. Projective Pairwise Discriminative Dictionary (PPDD) uses a synthesis dictionary and an analysis dictionary jointly to achieve the goal of pattern representation and discrimination. As the PPDD process of DL is in contrast to the use of l0or l1-norm sparsity constraints on the representation coefficients adopted in most traditional DL, it works faster than other DL. Moreover, the blending of synthesis dictionary and an analysis dictionary also enhance the feature representation of the complex eye patterns. We employed the combination of sclera and iris traits to establish multimodal biometrics. The experimental study and analysis conducted fulfill the hypothesis we considered. In this work we employed a part of the UBIRIS version 1 dataset to conduct the experiments. Abhijit Das 0001, Prabir Mondal, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein |
CEC | 1 |
| 2016 | A framework for liveness detection for direct attacks in the visible spectrum for multimodal ocular biometrics
Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein |
Pattern Recognit. Lett. | 1 |
| 2014 | Fuzzy logic based selera recognitionabstractIn this paper a sclera recognition and validation system is proposed. Here sclera segmentation was performed by Fuzzy logic-based clustering. Since the selera vessels are not prominent, image enhancement was required. A Fuzzy logic-based Brightness Preserving Dynamic Fuzzy Histogram Equalization and discrete Meyer wavelet was used to enhance the vessel patterns. For feature extraction, the Dense Local Binary Pattern (D-LBP) was used. D-LBP patch descriptors of each training image are used to form a bag of features, which is used to produce the training model. Support Vector Machines (SVMs) are used for classification. The UBIRIS version 1 dataset is used here for experimentation. An encouraging Equal Error Rate (EER) of 4.31% was achieved in our experiments. Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein |
FUZZ-IEEE | 1 |
| 2013 | Sclera recognition using dense-SIFTabstractIn this paper we propose a biometric sclera recognition and validation system. Here the sclera segmentation is performed bya time-adaptive active contour-based region growing technique. The sclera vessels are not prominent so image enhancement is required and hence a bank of 2D decomposition. A Haar wavelet multi-resolution filter is used to enhance the vessels pattern for better accuracy. For feature extraction, Dense Scale Invariant Feature Transform (D-SIFT) is used. D-SIFT patch descriptors of each training image are used to form bag of features by using k-means clustering and a spatial pyramid model, which is used to produce the training model. Support Vector Machines (SVMs) are used for classification. The UBIRIS version 1 dataset is used here for experimentation. Anencouraging Equal Error Rate (EER) of 0.66% is attained in the experiments presented. Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein |
ISDA | 1 |