EDBT 2026 Demo / reviewers in the wild / expert
Puneet Gupta 0002
dblp:06/1383-2
· DBLP profile ↗
32ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Cross-Attention Transformer for Non-contact Multimodal Pain Classification using Remote Physiological Signals and Visual Features
Anup Kumar Gupta 0001, Puneet Gupta 0002, Abhinav Dhall |
FG | 2 |
| 2026 | Exploring PPG-Guided Knowledge Distillation for Contactless Respiration Estimation
Trishna Saikia, Anup Kumar Gupta 0001, Puneet Gupta 0002, Pasi Liljeberg |
FG | 3 |
| 2026 | SHINE: Synergizing transformers with contrastive learning for thriving rPPG-based SpO2 estimation
Vaidehi Agarwal, Trishna Saikia, Anup Kumar Gupta 0001, Puneet Gupta 0002 |
Expert Syst. Appl. | 4 |
| 2026 | Elevating adversarial robustness by contrastive multitasking defence in medical image segmentation
Sneha Shukla, Puneet Gupta 0002 |
Neural Networks | 2 |
| 2025 | VISIONARY: Novel Spatial-Spectral Attention Mechanism for Hyperspectral Image DenoisingabstractImage denoising mitigates noise from the captured images and thereby, enhances the efficacy of high-demand vision applications, such as classification and segmentation. Hyperspectral Images (HSIs), with their multiple spectral bands, provide valuable information and make them highly applicable in real-world applications. Current Deep Learning methods mainly employ Transformers to denoise HSIs spatially and spectrally through self-attention (SA). However, SA focuses on individual samples and overlooks potential correlations within the images, indicating room for improvement. Moreover, existing Transformer-based denoising methods often fail to appropriately balance the importance of spatial and spectral features. This paper presents a novel method, VISIONARY, to address these issues by obtaining better HSI feature representation. To this end, it introduces the Spatial-Spectral-Cubic Transformer (SS-Cformer) block to address the shortcomings of Transformers in HSI denoising, particularly their inability to capture correlations within images of the same type, by introducing Global Feature Attention (GFA). Additionally, the SSC-former independently determines the optimal weightage for spatial and spectral features using attention mechanisms, leading to more effective denoising. Our method, VISIONARY is based on the integration of Transformer, U-Net and CNN architecture. Experimental results demonstrate that VISIONARY outperforms well-known methods on publicly available datasets, and our SSCformer block can be easily integrated with existing Transformer-based HSI denoising methods to improve their efficacy. Aditya Dixit, Nischit Hosamani, Puneet Gupta 0002, Ankur Garg |
WACV | 3 |
| 2025 | EVADE: A novel method to detect adversarial and OOD samples in Medical Image Segmentation
Sneha Shukla, Puneet Gupta 0002 |
Expert Syst. Appl. | 2 |
| 2025 | AVENUE: A Novel Deepfake Detection Method Based on Temporal Convolutional Network and rPPG InformationabstractIn Deep Learning (DL), an adversary creates Deepfakes by manipulating facial features to fool someone. The Deepfakes pose a security threat to anyone’s privacy and a primary concern for our society. It can be detected by utilizing the texture and physiological properties of the face, like eye and lip movements; however, such methods are incompetent when Deepfakes are created using recent Generative Adversarial Networks (GAN). Alternatively, Remote Photoplethysmography (rPPG) information can be used for Deepfake detection because GANs neglect human physiological information for Deepfake generation. Such detection can be inaccurate when rPPG signals are affected by the noises induced by facial deformation and illumination variations. Furthermore, the exiting Deepfake detections are usually performed using sequential models, and such models fail to process the long sequence of temporal information. These issues are mitigated by our proposed method AVENUE , that is, \(A\) no \(V\) el d \(E\) epfake detectio \(N\) method based on temporal convol \(U\) tion n \(E\) twork and rPPG information. For mitigating the noise issues in the rPPG signals, the proposed method detects and employs relatively stable clips of the input video for Deepfake detection. The stable clips are those clips that are least affected by facial deformations. Also, we use a modified Temporal convolutional network to model the long sequence of Deepfake information rather than the sequential architectures. We performed the experimental result on publicly available datasets of Deepfake videos. It demonstrates that our proposed method performs better than the existing rPPG-based Deepfake detection methods. Lokendra Birla, Trishna Saikia, Puneet Gupta 0002 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | HR-TRACK: An rPPG Method for Heartrate Monitoring Using Temporal Convolution Networks
Lokendra Birla, Sneha Shukla, Trishna Saikia, Puneet Gupta 0002 |
ICPR (13) | 4 |
| 2024 | Exploring the feasibility of adversarial attacks on medical image segmentation
Sneha Shukla, Anup Kumar Gupta 0001, Puneet Gupta 0002 |
Multim. Tools Appl. | 3 |
| 2023 | ALPINE: Improving Remote Heart Rate Estimation using Contrastive LearningabstractHeart rate (HR) is a crucial physiological indicator of human health and can be used to detect cardiovascular disorders. The traditional HR estimation methods, such as electrocardiograms (ECG) and photoplethysmographs, require skin contact. Due to the increased risk of viral in- fection from skin contact, these approaches are avoided in the ongoing COVID-19 pandemic. Alternatively, one can use the non-contact HR estimation technique, remote photo- plethysmography (rPPG), wherein HR is estimated from the facial videos of a person. Unfortunately, the existing rPPG methods perform poorly in the presence of facial deformations. Recently, there has been a proliferation of deep learning networks for rPPG. However, these networks require large-scale labelled data for better generalization. To alleviate these shortcomings, we propose a method ALPINE, that is, A noveL rPPG technique for Improving the remote heart rate estimatioN using contrastive lEarning. ALPINE utilizes the contrastive learning framework during training to address the issue of limited labelled data and introduces diversity in the data samples for better network generalization. Additionally, we introduce a novel hybrid loss comprising contrastive loss, signal-to-noise ratio (SNR) loss and data fidelity loss. Our novel contrastive loss maximizes the similarity between the rPPG information from different facial regions, thereby minimizing the effect of local noise. The SNR loss improves the quality of temporal signals, and the data fidelity loss ensures that the correct rPPG signal is extracted. Our extensive experiments on publicly available datasets demonstrate that the proposed method, ALPINE outperforms the previous well-known rPPG methods. Lokendra Birla, Sneha Shukla, Anup Kumar Gupta 0001, Puneet Gupta 0002 |
WACV | 4 |
| 2023 | RADIANT: Better rPPG estimation using signal embeddings and TransformerabstractRemote photoplethysmography can provide non-contact heart rate (HR) estimation by analyzing the skin color variations obtained from face videos. These variations are subtle, imperceptible to human eyes, and easily affected by noise. Existing deep learning-based rPPG estimators are incompetent due to three reasons. Firstly, they suppress the noise by utilizing information from the whole face even though different facial regions contain different noise characteristics. Secondly, local noise characteristics inherently affect the convolutional neural network (CNN) architectures. Lastly, the CNN sequential architectures fail to preserve long temporal dependencies. To address these issues, we propose RADIANT, that is, rPPG estimation using Signal Embeddings and Transformer. Our architecture utilizes a multi-head attention mechanism that facilitates feature subspace learning to extract the multiple correlations among the color variations corresponding to the periodic pulse. Also, its global information processing ability helps to suppress local noise characteristics. Furthermore, we propose novel signal embedding to enhance the rPPG feature representation and suppress noise. We have also improved the generalization of our architecture by adding a new training set. To this end, the effectiveness of synthetic temporal signals and data augmentations were explored. Experiments on extensively utilized rPPG datasets demonstrate that our architecture outperforms previous well-known architectures. Code: https://github.com/Deep-Intelligence-Lab/RADIANT.git Anup Kumar Gupta 0001, Lokendra Birla, Puneet Gupta 0002 |
WACV | 4 |
| 2023 | PERSIST: Improving micro-expression spotting using better feature encodings and multi-scale Gaussian TCN
Puneet Gupta 0002 |
Appl. Intell. | 1 |
| 2023 | Trustworthy Medical Image Segmentation with improved performance for in-distribution samples
Sneha Shukla, Lokendra Birla, Anup Kumar Gupta 0001, Puneet Gupta 0002 |
Neural Networks | 4 |
| 2023 | MERASTC: Micro-Expression Recognition Using Effective Feature Encodings and 2D Convolutional Neural NetworkabstractFacial micro-expression (ME) can disclose genuine and concealed human feelings. It makes MEs extensively useful in real-world applications pertaining to affective computing and psychology. Unfortunately, they are induced by subtle facial movements for a short duration of time, which makes the ME recognition, a highly challenging problem even for human beings. In automatic ME recognition, the well-known features encode either incomplete or redundant information, and there is a lack of sufficient training data. The proposed method, Micro-Expression Recognition by Analysing Spatial and Temporal Characteristics,$MERASTC$mitigates these issues for improving the ME recognition. It compactly encodes the subtle deformations using action units (AUs), landmarks, gaze, and appearance features of all the video frames while preserving most of the relevant ME information. Furthermore, it improves the efficacy by introducing a novel neutral face normalization for ME and initiating the utilization of gaze features in deep learning-based ME recognition. The features are provided to the 2D convolutional neural network that jointly analyses the spatial and temporal behavior for correct ME classification. Experimental results1on publicly available datasets indicate that the proposed method exhibits better performance than the well-known methods. Puneet Gupta 0002 |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | SUNRISE: Improving 3D Mask Face Anti-Spoofing for Short Videos Using Pre-Emptive Split and MergeabstractIn the current digital world, face analytic systems become an integral part of daily life. However, todays cutting-edge technology and readily available social media information make these systems vulnerable through face spoofing attacks. These attacks can be mitigated using remote Photoplethysmography (rPPG), which remotely detects cardiovascular signals. Unfortunately, the illumination variation and face deformations can easily corrupt the pulse signals and thereby degrade the performance of rPPG-based anti-spoofing methods, even when longer-duration face videos are employed. This paper proposes a face anti-spoofing methodSUNRISE, that is,Short videosUsiNg pRe-emptIveSplit and mErge. It is an rPPG-based face anti-spoofing for short duration videos wherein facial deformations are removed by introducing the split and merge mechanism. It splits the video into several clips, provides low importance to the clips containing facial deformations, and eventually merges the results using quality-based fusion. The efficacy of existing rPPG-based methods is limited because they employed high dimensional features for training using the limited training data. We mitigate this limitation by utilizing the statistical features of clips. The experimental results on publicly available datasets reveal that the proposed method exhibits significantly better performance than the well-known existing methods for a short time duration Lokendra Birla, Puneet Gupta 0002, Shravan Kumar |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | UNFOLD: 3-D U-Net, 3-D CNN, and 3-D Transformer-Based Hyperspectral Image DenoisingabstractHyperspectral Images (HSIs) encompass data across numerous spectral bands, making them valuable in various practical fields such as remote sensing, agriculture, and marine monitoring. Unfortunately, inevitable noise introduction during sensing restricts their applicability, necessitating denoising for optimal utilization. The existing Deep Learning based denoising methods suffer from various limitations. For instance, Convolutional Neural Networks (CNNs) struggle with long-range dependencies, while Vision Transformers struggle to capture local details. This paper introduces a novel method,UNFOLD, that addresses these inherent limitations by harmoniously integrating the strengths of 3D U-Net, 3D CNN, and 3D Transformer architectures. Unlike several existing methods that predominantly capture dependencies either along the spatial or the spectral dimension,UNFOLDaddresses HSI denoising as a 3D task, synergizing spatial and spectral information through the utilization of 3D Transformer and 3D CNN. It employs the self-attention mechanism of Transformers to capture the global dependencies and model long-range relationships across spatial and spectral dimensions. To overcome the limitations of 3D Transformer in capturing fine-grained local and spatial features,UNFOLDcomplements it by incorporating 3D CNN. Moreover,UNFOLDutilize a modified form of 3D U-Net architecture for HSI denoising, wherein it employs a 3D Transformer based encoder instead of the conventional 3D CNN-based encoder. It further capitalizes on the property of U-Net to integrate features across various scales, thereby enhancing efficacy by preserving intricate structural details. Results from extensive experiments demonstrate thatUNFOLDoutperforms the state-of-the-art HSI denoising methods. Aditya Dixit, Anup Kumar Gupta 0001, Puneet Gupta 0002, Saurabh Srivastava 0003, Ankur Garg |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | FATALRead - Fooling visual speech recognition models
Anup Kumar Gupta 0001, Puneet Gupta 0002, Esa Rahtu |
Appl. Intell. | 2 |
| 2022 | PATRON: Exploring respiratory signal derived from non-contact face videos for face anti-spoofing
Lokendra Birla, Puneet Gupta 0002 |
Expert Syst. Appl. | 2 |
| 2021 | Contextual Emotion Learning ChallengeabstractEmotion recognition via vision has been deeply associated with facial expressions, and the inference of emotions has, more often than not, been based on the same. However, context, both environmental and social, plays an imperative role in emotion recognition but has not been incorporated widely so far. The meaning of emotion might entirely switch when shifted from one setting to another if only facial expressions are taken into account. Moreover, there exists no study in the Indian context about the same. To cater to this issue, we generate and introduce the Indian Contextual Emotion Recognition (ICER) dataset based on the multi-ethnic Indian context. This paper summarises the Contextual Emotion Learning Challenge (CELC 2021) organized in conjunction with the 16th IEEE Conference on Automatic Face and Gesture Recognition (FG) 2021. We outline the tasks posed in the challenge, the novel dataset, along with its challenges and the evaluation method. Lastly, we conclude by discussing the possible future directions. Jainendra Shukla, Puneet Gupta 0002, Aniket Bera, Arka Sarkar, Prakhar Goel, Shubhangi Butta, Anup Kumar Gupta 0001, Snehil Sanyal, Debanga Raj Neog, Manas Kamal Bhuyan, Kalyani Marathe, Linda G. Shapiro, Alex Colbrn, Varchita Lalwani |
FG | 2 |
| 2021 | DARE: Deceiving Audio-Visual speech Recognition model
Saumya Mishra, Anup Kumar Gupta 0001, Puneet Gupta 0002 |
Knowl. Based Syst. | 3 |
| 2018 | Real Time Hand Segmentation on Frugal Headmounted Device for Gestural InterfaceabstractWith the resurgence of Head Mounted Displays (HMDs), in-air gestures form a natural and intuitive interaction mode of communication. HMDs such as Microsoft Hololens, Daqri smart-glasses etc., have on-board processors with additional sensors, making the device expensive. Our goal, therefore, is to enable mass-market reach by extending the interaction space around mobile devices for Augmented/Virtual reality: with just frugal head-mounts such as Google Cardboard and Wearality with a smartphone. One necessary step for a feasible human-computer interaction is gesture recognition, preceded by a reliable hand segmentation in egocentric view. We propose a technique for real-time hand segmentation that utilize only RGB camera in an off-the-shelf mobile device. The novelty of our work lies in coming up with a filtering technique that term Multi Orientation Matched filter, used for hand segmentation that works on-device, even in situations of skin-like background. We have extensively tested our hand segmentation method on public datasets. We provide comparisons of hand segmentation with the existing methods under the same limitations. We demonstrate that our method outperforms in term of computational time and comparable results in term of accuracy. Further, we demonstrate that using our solution a zoom gesture classification can work in real-time on android smartphones. Jitender Maurya, Ramya Hebbalaguppe, Puneet Gupta 0002 |
ICIP | 3 |
| 2018 | Robust Adaptive Heart-Rate Monitoring Using Face VideosabstractHeart rate (HR) monitoring is indispensable for several real-world scenarios, especially when acquired in a non-contact manner. It can be accomplished using face videos acquired from ubiquitous cameras in an inexpensive, non-invasive and unobtrusive manner. But the HR monitoring can be erroneous when the video contains facial expressions, out-of-plane movements, change in camera parameters (like focus) and variations in environmental factors (like illumination). The proposed system mitigates these problems for improving the HR monitoring. For this, it defines an adaptive temporal signal selection mechanism which identifies and removes the facial areas affected by facial expressions. Moreover, it introduces a novel post-processing mechanism which perform HR monitoring by utilizing face reconstruction and quality. The post-processing is used when the face video contains facial movements. Experimental results reveal that incorporation of adaptive temporal signal selection and post-processing mechanisms can significantly improve the HR monitoring. It depicts that the Pearson correlation between actual and estimated HR is 0.95 while the average absolute error is 1.63 beats per minute, which indicates that the proposed system provides good HR monitoring. Puneet Gupta 0002, Brojeshwar Bhowmick, Arpan Pal 0001 |
WACV | 1 |
| 2017 | Accurate heart-rate estimation from face videos using quality-based fusionabstractEstimating heart rate (HR) accurately using face videos acquired from a low cost camera in contactless manner is of paramount importance for many real-world applications. Such existing systems perform spuriously due to change in camera parameters, respiration, facial expressions and environmental factors. This paper mitigates the issues for accurate HR estimation. The face video consisting of frontal, profile or multiple faces is divided into multiple overlapping fragments to determine HR estimates. The HR estimates are fused using quality-based fusion which aims to minimize illumination and face deformations. Experimental results demonstrate that the proposed system exhibit better performance than the state of the art systems and establishes the efficacy of the quality-based fusion in HR estimation. Puneet Gupta 0002, Brojeshwar Bhowmick, Arpan Pal 0001 |
ICIP | 1 |
| 2017 | An efficient slap image matching system based on dynamic classifier selection and aggregation
Puneet Gupta 0002, Phalguni Gupta |
Inf. Sci. | 1 |
| 2016 | An accurate slap fingerprint based verification system
Puneet Gupta 0002, Phalguni Gupta |
Neurocomputing | 1 |
| 2016 | An accurate infrared hand geometry and vein pattern based authentication system
Puneet Gupta 0002, Saurabh Srivastava 0003, Phalguni Gupta |
Knowl. Based Syst. | 1 |
| 2015 | Fingerprint Orientation Modeling Using Symmetric FiltersabstractAccurate fingerprint orientation is a prerequisite in fingerprint based recognition system. This paper proposes an algorithm for modeling the fingerprint orientation field by using a model based algorithm based on the weighted Legendre basis. Weights required in the modeling are obtained by using symmetric filters, such that: i) high weights should be assigned to the areas near singular points, ii) areas having uniform ridge-valley flow should be given high weights, and iii) areas containing bad quality due to dry/ wet fingerprints, scars, bruises or sensor condition should be given low weights. These conditions ensure accurate reconstruction of fingerprint orientation field for bad quality areas while preserving the true orientation field near singular points. The proposed algorithm has been evaluated on a publicly available database, FVC2004 DB1A. Experimental results reveal that it has better orientation field estimation compared to the various state of the art algorithms. Puneet Gupta 0002, Phalguni Gupta |
WACV | 1 |
| 2015 | Multi-modal fusion of palm-dorsa vein pattern for accurate personal authentication
Puneet Gupta 0002, Phalguni Gupta |
Knowl. Based Syst. | 1 |
| 2014 | A Dynamic Slap Fingerprint Based Verification System
Puneet Gupta 0002, Phalguni Gupta |
ICIC (1) | 1 |
| 2014 | An efficient slap fingerprint segmentation and hand classification algorithm
Puneet Gupta 0002, Phalguni Gupta |
Neurocomputing | 1 |
| 2014 | Abductive Analysis of Administrative Policies in Rule-Based Access ControlabstractIn large organizations, access control policies are managed by multiple users (administrators). An administrative policy specifies how each user in an enterprise may change the policy. Fully understanding the consequences of an administrative policy in an enterprise system can be difficult, because of the scale and complexity of the access control policy and the administrative policy, and because sequences of changes by different users may interact in unexpected ways. Administrative policy analysis helps by answering questions such as user-permission reachability, which asks whether specified users can together change the policy in a way that achieves a specified goal, namely, granting a specified permission to a specified user. This paper presents a rule-based access control policy language, a rule-based administrative policy model that controls addition and removal of facts and rules, and an abductive analysis algorithm for user-permission reachability. Abductive analysis means that the algorithm can analyze policy rules even if the facts initially in the policy (e.g., information about users) are unavailable. The algorithm does this by computing minimal sets of facts that, if present in the initial policy, imply reachability of the goal. Puneet Gupta 0002, Scott D. Stoller, Zhongyuan Xu |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2012 | Four Slap Fingerprint Segmentation
Nishant Singh, Aditya Nigam, Puneet Gupta 0002, Phalguni Gupta |
ICIC (2) | 3 |