VLDB 2026 Research / reviewers in the wild / expert
Anup Kumar Gupta 0001
dblp:305/5359
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-1090-6036ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Cross-Attention Transformer for Non-contact Multimodal Pain Classification using Remote Physiological Signals and Visual Features
Anup Kumar Gupta 0001, Puneet Gupta 0002, Abhinav Dhall |
FG | 1 |
| 2026 | Exploring PPG-Guided Knowledge Distillation for Contactless Respiration Estimation
Trishna Saikia, Anup Kumar Gupta 0001, Puneet Gupta 0002, Pasi Liljeberg |
FG | 2 |
| 2026 | SHINE: Synergizing transformers with contrastive learning for thriving rPPG-based SpO2 estimation
Vaidehi Agarwal, Trishna Saikia, Anup Kumar Gupta 0001, Puneet Gupta 0002 |
Expert Syst. Appl. | 3 |
| 2024 | Exploring the feasibility of adversarial attacks on medical image segmentation
Sneha Shukla, Anup Kumar Gupta 0001, Puneet Gupta 0002 |
Multim. Tools Appl. | 2 |
| 2023 | ALPINE: Improving Remote Heart Rate Estimation using Contrastive LearningabstractHeart rate (HR) is a crucial physiological indicator of human health and can be used to detect cardiovascular disorders. The traditional HR estimation methods, such as electrocardiograms (ECG) and photoplethysmographs, require skin contact. Due to the increased risk of viral in- fection from skin contact, these approaches are avoided in the ongoing COVID-19 pandemic. Alternatively, one can use the non-contact HR estimation technique, remote photo- plethysmography (rPPG), wherein HR is estimated from the facial videos of a person. Unfortunately, the existing rPPG methods perform poorly in the presence of facial deformations. Recently, there has been a proliferation of deep learning networks for rPPG. However, these networks require large-scale labelled data for better generalization. To alleviate these shortcomings, we propose a method ALPINE, that is, A noveL rPPG technique for Improving the remote heart rate estimatioN using contrastive lEarning. ALPINE utilizes the contrastive learning framework during training to address the issue of limited labelled data and introduces diversity in the data samples for better network generalization. Additionally, we introduce a novel hybrid loss comprising contrastive loss, signal-to-noise ratio (SNR) loss and data fidelity loss. Our novel contrastive loss maximizes the similarity between the rPPG information from different facial regions, thereby minimizing the effect of local noise. The SNR loss improves the quality of temporal signals, and the data fidelity loss ensures that the correct rPPG signal is extracted. Our extensive experiments on publicly available datasets demonstrate that the proposed method, ALPINE outperforms the previous well-known rPPG methods. Lokendra Birla, Sneha Shukla, Anup Kumar Gupta 0001, Puneet Gupta 0002 |
WACV | 3 |
| 2023 | RADIANT: Better rPPG estimation using signal embeddings and TransformerabstractRemote photoplethysmography can provide non-contact heart rate (HR) estimation by analyzing the skin color variations obtained from face videos. These variations are subtle, imperceptible to human eyes, and easily affected by noise. Existing deep learning-based rPPG estimators are incompetent due to three reasons. Firstly, they suppress the noise by utilizing information from the whole face even though different facial regions contain different noise characteristics. Secondly, local noise characteristics inherently affect the convolutional neural network (CNN) architectures. Lastly, the CNN sequential architectures fail to preserve long temporal dependencies. To address these issues, we propose RADIANT, that is, rPPG estimation using Signal Embeddings and Transformer. Our architecture utilizes a multi-head attention mechanism that facilitates feature subspace learning to extract the multiple correlations among the color variations corresponding to the periodic pulse. Also, its global information processing ability helps to suppress local noise characteristics. Furthermore, we propose novel signal embedding to enhance the rPPG feature representation and suppress noise. We have also improved the generalization of our architecture by adding a new training set. To this end, the effectiveness of synthetic temporal signals and data augmentations were explored. Experiments on extensively utilized rPPG datasets demonstrate that our architecture outperforms previous well-known architectures. Code: https://github.com/Deep-Intelligence-Lab/RADIANT.git Anup Kumar Gupta 0001, Lokendra Birla, Puneet Gupta 0002 |
WACV | 1 |
| 2023 | Trustworthy Medical Image Segmentation with improved performance for in-distribution samples
Sneha Shukla, Lokendra Birla, Anup Kumar Gupta 0001, Puneet Gupta 0002 |
Neural Networks | 3 |
| 2023 | UNFOLD: 3-D U-Net, 3-D CNN, and 3-D Transformer-Based Hyperspectral Image DenoisingabstractHyperspectral Images (HSIs) encompass data across numerous spectral bands, making them valuable in various practical fields such as remote sensing, agriculture, and marine monitoring. Unfortunately, inevitable noise introduction during sensing restricts their applicability, necessitating denoising for optimal utilization. The existing Deep Learning based denoising methods suffer from various limitations. For instance, Convolutional Neural Networks (CNNs) struggle with long-range dependencies, while Vision Transformers struggle to capture local details. This paper introduces a novel method,UNFOLD, that addresses these inherent limitations by harmoniously integrating the strengths of 3D U-Net, 3D CNN, and 3D Transformer architectures. Unlike several existing methods that predominantly capture dependencies either along the spatial or the spectral dimension,UNFOLDaddresses HSI denoising as a 3D task, synergizing spatial and spectral information through the utilization of 3D Transformer and 3D CNN. It employs the self-attention mechanism of Transformers to capture the global dependencies and model long-range relationships across spatial and spectral dimensions. To overcome the limitations of 3D Transformer in capturing fine-grained local and spatial features,UNFOLDcomplements it by incorporating 3D CNN. Moreover,UNFOLDutilize a modified form of 3D U-Net architecture for HSI denoising, wherein it employs a 3D Transformer based encoder instead of the conventional 3D CNN-based encoder. It further capitalizes on the property of U-Net to integrate features across various scales, thereby enhancing efficacy by preserving intricate structural details. Results from extensive experiments demonstrate thatUNFOLDoutperforms the state-of-the-art HSI denoising methods. Aditya Dixit, Anup Kumar Gupta 0001, Puneet Gupta 0002, Saurabh Srivastava 0003, Ankur Garg |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | FATALRead - Fooling visual speech recognition models
Anup Kumar Gupta 0001, Puneet Gupta 0002, Esa Rahtu |
Appl. Intell. | 1 |
| 2021 | Contextual Emotion Learning ChallengeabstractEmotion recognition via vision has been deeply associated with facial expressions, and the inference of emotions has, more often than not, been based on the same. However, context, both environmental and social, plays an imperative role in emotion recognition but has not been incorporated widely so far. The meaning of emotion might entirely switch when shifted from one setting to another if only facial expressions are taken into account. Moreover, there exists no study in the Indian context about the same. To cater to this issue, we generate and introduce the Indian Contextual Emotion Recognition (ICER) dataset based on the multi-ethnic Indian context. This paper summarises the Contextual Emotion Learning Challenge (CELC 2021) organized in conjunction with the 16th IEEE Conference on Automatic Face and Gesture Recognition (FG) 2021. We outline the tasks posed in the challenge, the novel dataset, along with its challenges and the evaluation method. Lastly, we conclude by discussing the possible future directions. Jainendra Shukla, Puneet Gupta 0002, Aniket Bera, Arka Sarkar, Prakhar Goel, Shubhangi Butta, Anup Kumar Gupta 0001, Snehil Sanyal, Debanga Raj Neog, Manas Kamal Bhuyan, Kalyani Marathe, Linda G. Shapiro, Alex Colbrn, Varchita Lalwani |
FG | 7 |
| 2021 | DARE: Deceiving Audio-Visual speech Recognition model
Saumya Mishra, Anup Kumar Gupta 0001, Puneet Gupta 0002 |
Knowl. Based Syst. | 2 |