Daksh Thapar

dblp:206/6241 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-2871-0376ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 6 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Gait Recognition via Pristine Feature Learning
Anuj Rathore, Daksh Thapar, Mahesh Chandran
ACIVS2
2025 Elemental Composite Prototypical Network: Few-Shot Object Detection on Outdoor 3D Point Cloud Scenes
abstract
This paper introduces the Elemental Composite Prototypical Network (ECPN), a novel approach to few-shot learning (FSL) in outdoor 3D point cloud object detection. Such point clouds are inherently non-uniformly packed and show marked intra-class variations due to aberrations in lidar scanning methods. Due to the limited availability of examples in the FSL setting, the intra-class variations serve as a much more formidable challenge to traditional detection algorithms. ECPN employs a novel prototypical learning method that solves the issues mentioned above. We generate and leverage multiple elemental prototypes for each class to capture essential geometric features from limited examples. These elemental prototypes are then combined in a weighted manner to arrive at composite prototypes that score relevant and irrelevant features in the elemental prototypes with respect to the query point cloud scene. Moreover, we introduce a novel feature-similarity-discrimination loss to refine the model's ability to distinguish between relevant objects and their background, significantly improving object detection accuracy in FSL scenarios. Our extensive testing on the nuScenes dataset demonstrates that ECPN significantly outperforms existing baselines, offering a robust solution to the complexities of outdoor few-shot 3D object detection (O-FS3D) and setting a new standard for future research.
Arkadipta De, Vartika Sengar, Daksh Thapar, Mahesh Chandran, Manohar Kaul
WACV3
2025 Learning Semantic Part-Based Graph Structure for 3D Point Cloud Domain Generalization
abstract
In 3D data analysis, point clouds provide detailed geometric insights for applications like computer vision and geospatial analysis. However, their irregularity and diversity make classification challenging, especially in domain generalization, where models must generalize to new data distributions. Our research introduces a novel 3D Domain Generalization (3DDG) method using Unsupervised Part Decomposition (UPD) and Graph Structure Induction (GSI). The UPD module employs spectral clustering and a modified Shannon entropy method to segment point clouds into meaningful parts. The GSI module constructs a graph of these parts' spatial relationships, processed by a Graph Neural Network (GNN) to understand complex geometries. Our approach enhances part-based analysis, improving classification accuracy on the PointDA-10 and GraspNetPC-10 datasets by 1.25% and 2.6%, respectively. These results highlight our advancements in 3D domain generalization, enabling more robust classification models for diverse point cloud data.
G. Ujwal Sai, Arkadipta De, Vartika Sengar, Anuj Rathore, Daksh Thapar, Manohar Kaul
WACV5
2024 GaitW: Enhancing Gait Recognition in the Wild Using Dynamic Information
Daksh Thapar, Jayesh Chaudhari, Sunny Manchanda, Aditya Nigam, Chetan Arora 0001
ACCV (1)1
2023 Parts Based Attention for Highly Occluded Pedestrian Detection with Transformers
abstract
Despite the significant progress made in pedestrian detection in last decade, detecting pedestrians under heavy occlusion still remains a challenging problem. In state of the art (SOTA), convolutional neural network (CNN) based models, the reason is attributed to non-maximal-suppression (NMS), which often erroneously deletes true positives when one pedestrian is occluding other. SOTA transformer based models do not have such NMS step, yet fail to detect highly occluded pedestrians. In this paper, we study the reasons for such failures. We observe that such models first predict key-points, and then compute the attention at the specific key-points. Our analysis reveals that the key-points do not have any preference towards semantically important body parts. Under heavy occlusion, such key-points end up attending to non-discriminative regions or background, leading to false negatives. We take inspiration from the conventional wisdom of detecting objects using their parts, and bias the attention of proposed transformer architecture towards semantically important, and highly discriminative human body parts. The intervention leads to SOTA results on benchmark Citypersons and Caltech datasets, achieving 30.75%, and 32.96% miss-rate (lower is better) respectively, against 32.6%, and 38.2% by the current SOTA. Code is available at https://ajayshastry08.github.io/pa_dino
K. N. Ajay Shastry, Jayesh Chaudhari, Daksh Thapar, Aditya Nigam, Chetan Arora 0001
ICIP3
2023 From Forks to Forceps: A New Framework for Instance Segmentation of Surgical Instruments
abstract
Minimally invasive surgeries and related applications demand surgical tool classification and segmentation at the instance level. Surgical tools are similar in appearance and are long, thin, and handled at an angle. The fine-tuning of state-of-the-art (SOTA) instance segmentation models trained on natural images for instrument segmentation has difficulty discriminating instrument classes. Our research demonstrates that while the bounding box and segmentation mask are often accurate, the classification head misclassifies the class label of the surgical instrument. We present a new neural network framework that adds a classification module as a new stage to existing instance segmentation models. This module specializes in improving the classification of instrument masks generated by the existing model. The module comprises multi-scale mask attention, which attends to the instrument region and masks the distracting background features. We propose training our classifier module using metric learning with arc loss to handle low inter-class variance of surgical instruments. We conduct exhaustive experiments on the benchmark datasets EndoVis2017 and EndoVis2018. We demonstrate that our method outperforms all (more than 18) SOTA methods compared with, and improves the SOTA performance by at least 12 points (20%) on the EndoVis2017 benchmark challenge and generalizes effectively across the datasets. Project page with source code is available at nets-iitd.github.io/s3net.
Britty Baby, Daksh Thapar, Mustafa Chasmai, Tamajit Banerjee, Kunal Dargan, Ashish Suri, Subhashis Banerjee, Chetan Arora 0001
WACV2
2022 Merry Go Round: Rotate a Frame and Fool a DNN
abstract
A large proportion of videos captured today are first person videos shot from wearable cameras. Similar to other computer vision tasks, Deep Neural Networks (DNNs) are the workhorse for most state-of-the-art (SOTA) egocentric vision techniques. On the other hand DNNs are known to be susceptible to Adversarial Attacks (AAs) which add imperceptible noise to the input. Both black-box, as well as white-box attacks on image as well as video analysis tasks have been shown. We observe that most AA techniques basically add intensity perturbation to an image. Even for videos, the same process is essentially repeated for each frame independently. We note that definition of imperceptibility used for images may not be applicable for videos, where a small intensity change happening randomly in two consecutive frames may still be perceptible. In this paper we make a key novel suggestion to use perturbation in optical flow to carry out AAs on a video analysis system. Such perturbation is especially useful for egocentric videos, because there is lot of shake in the egocentric videos anyways, and adding a little more, keeps it highly imperceptible. In general our idea can be seen as adding structured, parametric noise as the adversarial perturbation. Our implementation of the idea by adding 3D rotations to the frames, reveal that using our technique, one can mount a black-box AA on an egocentric activity detection system in one-third of the queries compared to the SOTA AA technique.
Daksh Thapar, Aditya Nigam, Chetan Arora 0001
CVPR1
2021 Anonymizing Egocentric Videos
abstract
In egocentric videos, the face of a wearer capturing the video is never captured. This gives a false sense of security that the wearer’s privacy is preserved while sharing such videos. However, egocentric cameras are typically harnessed to wearer’s head, and hence, also capture wearer’s gait. Recent works have shown that wearer gait signatures can be extracted from egocentric videos, which can be used to determine if two egocentric videos have the same wearer. In a more damaging scenario, one can even recognize a wearer using hand gestures from egocentric videos, or identify a wearer in third person videos such as from a surveillance camera. We believe, this could be a death knell in sharing of egocentric videos, and fatal for egocentric vision research. In this work, we suggest a novel technique to anonymize egocentric videos, which create carefully crafted, but small, and imperceptible optical flow perturbations in an egocentric video’s frames. Importantly, these perturbations do not affect object detection or action/activity recognition from egocentric videos but are strong enough to dis-balance the gait recovery process. In our experiments on benchmark EPIC-Kitchens dataset, the proposed perturbation degrades the wearer recognition performance of [42], from 66.3% to 13.4%, while preserving the activity recognition performance of [10] from 89.6% to 87.4%. To test our anonymization with more wearer recognition techniques, we also developed a stronger, and more generalizable wearer recognition method based on camera egomotion cues. The approach achieves state-ofthe-art (SOTA) performance of 59.67% on EPIC-Kitchens, compared to 55.06% by [42]. However, the accuracy of our recognition technique also drops to 12% using the proposed anonymizing perturbations.
Daksh Thapar, Aditya Nigam, Chetan Arora 0001
ICCV1
2020 Hierarchical X-Ray Report Generation via Pathology Tags and Multi Head Attention
Preethi Srinivasan, Daksh Thapar, Arnav Bhavsar, Aditya Nigam
ACCV (5)2
2020 Is Sharing of Egocentric Video Giving Away Your Biometric Signature?
Daksh Thapar, Chetan Arora 0001, Aditya Nigam
ECCV (17)1
2020 Recognizing Camera Wearer from Hand Gestures in Egocentric Videos: https: //egocentricbiometric.github.io/
abstract
Wearable egocentric cameras are typically harnessed to a wearer's head, giving them the unique advantage of capturing their points of view. Hoshen and Peleg have shown that egocentric cameras indirectly capture the wearer's gait, which can be used to identify a wearer based on their egocentric videos. The authors have shown a wearer recognition accuracy of up to 77% over 32 subjects. However, an important limitation of their work is that such gait features can be extracted only from walking sequences of a wearer. In this work, we take the privacy threat a notch higher and show that even the wearer's hand gestures, as seen through an egocentric video, leak wearer's identity. We have designed a model to extract and match hand gesture signatures from egocentric videos. We demonstrate the threat on the EPIC kitchen dataset containing 55 hours of the egocentric videos acquired from 32 subjects doing various activities. We show that: (1) Our model can recognize a wearer with an accuracy of up to 73% based on the same activity, i.e., the model has seen 'cut' activity by a wearer in the train set, and recognizes the wearer based on another 'cut' activity by him/her while testing. (2) The hand gesture signatures transfer across activities, i.e., even if our model does not see 'cut' activity of a wearer at the train time, but sees other activities such as 'wash', 'mix' etc., the model can still recognize a wearer with an accuracy of up to 60%, by matching hand gesture signatures of 'cut' at test time with train time signatures of 'wash' or 'mix'. (3) The hand gesture features even transfer across subjects, i.e., even if the model has not seen any activity by some subject, one can still verify a wearer (open-set), and predict that the same wearer has performed both activities with an Equal Error Rate of 15.21%. The code, trained models are available at https://egocentricbiometric.github.io/
Daksh Thapar, Aditya Nigam, Chetan Arora 0001
ACM Multimedia1
2019 Learning Domain Specific Features using Convolutional Autoencoder: A Vein Authentication Case Study using Siamese Triplet Loss Network
abstract
Recently, deep hierarchically learned models (such as CNN) have achieved superior performance in various computer vision tasks but limited attention has been paid to biometrics till now. This is major because of the number of samples available in biometrics are limited and are not enough to train CNN efficiently. However, deep learning often requires a lot of training data because of the huge number of parameters to be tuned by the learning algorithm. How about designing an end-to-end deep learning network to match the biometric features when the number of training samples is limited? To address this problem, we propose a new way to design an end-to-end deep neural network that works in two major steps: first an auto-encoder has been trained for learning domain specific features followed by a Siamese network trained via. triplet loss function for matching. A publicly available vein image data set has been utilized as a case study to justify our proposal. We observed that transformations learned from such a network provide domain specific and most discriminative vascular features. Subsequently, the corresponding traits are matched using multimodal pipelined end-to-end network in which the convolutional layers are pre-trained in an unsupervised fashion as an autoencoder. Thorough experimental studies suggest that the proposed framework consistently outperforms several state-of-the-art vein recognition approaches.
Manish Agnihotri, Aditya Rathod, Daksh Thapar, Gaurav Jaswal, Kamlesh Tiwari, Aditya Nigam
ICPRAM3
2019 FKIMNet: A Finger Dorsal Image Matching Network Comparing Component (Major, Minor and Nail) Matching with Holistic (Finger Dorsal) Matching
abstract
Current finger knuckle image recognition systems, often require users to place fingers' major or minor joints flatly towards the capturing sensor. To extend these systems for user non-intrusive application scenarios, such as consumer electronics, forensic, defence etc, we suggest matching the full dorsal fingers, rather than the major/ minor region of interest (ROI) alone. In particular, this paper makes a comprehensive study on the comparisons between full finger and fusion of finger ROI's for finger knuckle image recognition. These experiments suggest that using full-finger, provides a more elegant solution. Addressing the finger matching problem, we propose a CNN (convolutional neural network) which creates a 128-D feature embedding of an image. It is trained via. triplet loss function, which enforces the L2 distance between the embeddings of the same subject to be approaching zero, whereas the distance between any 2 embeddings of different subjects to be at least a margin. For precise training of the network, we use dynamic adaptive margin, data augmentation, and hard negative mining. In distinguished experiments, the individual performance of finger, as well as weighted sum score level fusion of major knuckle, minor knuckle, and nail modalities have been computed, justifying our assumption to consider full finger as biometrics instead of its counterparts. The proposed method is evaluated using two publicly available finger knuckle image datasets i.e., PolyU FKP dataset and PolyU Contactless FKI Datasets.
Daksh Thapar, Gaurav Jaswal, Aditya Nigam
IJCNN1
2019 Gait metric learning siamese network exploiting dual of spatio-temporal 3D-CNN intra and LSTM based inter gait-cycle-segment features
Daksh Thapar, Gaurav Jaswal, Aditya Nigam, Chetan Arora 0001
Pattern Recognit. Lett.1
2018 All-Conv Net for Bird Activity Detection: Significance of Learned Pooling
Arjun Pankajakshan, Anshul Thakur, Daksh Thapar, Padmanabhan Rajan, Aditya Nigam
INTERSPEECH3