Chetan Arora 0001

dblp:19/1006-1 · DBLP profile ↗
← Back
98ranked-venue papers
8as first author
56since 2021 · last 2026
0000-0003-0155-0250ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 83 · 7 first-author · 47 since 2021Artificial intelligence and machine learning · 42 · 6 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
abstract
Open-vocabulary object detectors (OVODs) unify vision and language to detect arbitrary object categories based on text prompts, enabling strong zero-shot generalization to novel concepts. As these models gain traction in high-stakes applications such as robotics, autonomous driving, and surveillance, understanding their security risks becomes crucial. In this work, we conduct the first study of backdoor attacks on OVODs and reveal a new attack surface introduced by prompt tuning. We propose TrAP (Trigger-Aware Prompt tuning), a multi-modal backdoor injection strategy that jointly optimizes prompt parameters in both image and text modalities along with visual triggers. TrAP enables the attacker to implant malicious behavior using lightweight, learnable prompt tokens without retraining the base model weights, thus preserving generalization while embedding a hidden backdoor. We adopt a curriculum-based training strategy that progressively shrinks the trigger size, enabling effective backdoor activation using small trigger patches at inference. Experiments across multiple datasets show that TrAP achieves high attack success rates for both object misclassification and object disappearance attacks, while also improving clean image performance on downstream datasets compared to the zero-shot setting.
Ankita Raj, Chetan Arora 0001
AAAI2
2026 Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score
abstract
Vision-language models (VLMs) like CLIP excel in zeroshot learning by aligning image and text representations through contrastive pretraining. Existing approaches to unsupervised adaptation (UA) for fine-grained classification with VLMs either rely on fixed alignment scores that may not capture evolving, subtle class distinctions or on computationally expensive pseudo-labeling strategies that limit scalability. In contrast, we show that modeling fine-grained cross-modal interactions during adaptation produces more accurate, class-discriminative pseudo-labels and substantially improves performance over state-of-the-art (SOTA) methods. We introduce Fine-grained Alignment and Interaction Refinement (FAIR), an innovative approach that dynamically aligns localized image features with descriptive language embeddings through a set of Class Description Anchors (CDA). This enables the definition of a Learned Alignment Score (LAS), which incorporates CDA as an adaptive classifier, facilitating cross-modal interactions to improve self-training in unsupervised adaptation. Furthermore, we propose a self-training weighting mechanism designed to refine pseudo-labels in the presence of inter-class ambiguities. Our approach, FAIR, delivers a substantial performance boost in fine-grained unsupervised adaptation, achieving a notable overall gain of 2.78% across 13 fine-grained datasets compared to SOTA methods.1
Eman Ali, Sathira Silva, Chetan Arora 0001, Muhammad Haris Khan
WACV3
2026 Domain Generalizing DINO for Visual Regression via Latent Distractor Subspace Consistency
abstract
Vision Foundation Models, such as DINO [20], have demonstrated remarkable generalization in classification; however, their application to out-of-domain visual regression tasks remains a significant and underexplored challenge. Unlike classification, domain generalization in regression poses distinct challenges: regression produces continuous outputs and is particularly sensitive to high-variance, label-irrelevant factors (e.g., illumination, blur, or contrast). These factors can entangle with task-relevant features and induce spurious correlations. While recent regression methods [11], [15], [24], [38], [39] have shown promise, they often rely on CNN backbones and require the pre-specification of known distractors. This demands significant domain expertise and fails to address spurious correlations that emerge during training. To address these challenges, we propose LDSC, a Latent Distractor Subspace Consistency framework that disentangles intermediate feature representation into task-relevant and latent distractor subspaces, and regularizes the latter under photometric perturbations to suppress spurious correlations while preserving discriminative features during training. Our proposed method, LDSC, is the first to effectively adapt the powerful DINO backbone for domain generalized visual regression. LDSC achieves state-of-the-art results on seven benchmark regression datasets, demonstrating its strong performance in domain generalization for visual regression with percentage improvements of (41.75%, 20.12%, 52.05%, 8.27%, 22.21%, 3.55%) over state-of-the-art DG regression methods, respectively. Project page is available: ldsc-iitd.github.io.
Nikhil Reddy, Chetan Arora 0001, Mahsa Baktash
WACV2
2025 PEFTDiff: Diffusion-Guided Transferability Estimation for Parameter-Efficient Fine-Tuning
Prafful Kumar Khoba, Zijian Wang 0009, Chetan Arora 0001, Mahsa Baktash
ICCV3
2025 Focus on Texture: Rethinking Pre-training in Masked Autoencoders for Medical Image Classification
Chetan Madan, Aarjav Satia, Soumen Basu, Pankaj Gupta 0005, Usha Dutta, Chetan Arora 0001
MICCAI (4)6
2025 Prompting without Panic: Attribute-Aware, Zero-Shot, Test-Time Calibration
Ramya Hebbalaguppe, Tamoghno Kandar, Abhinav Nagpal, Chetan Arora 0001
ECML/PKDD (6)4
2025 Feature Space Perturbation: A Panacea to Enhanced Transferability Estimation
abstract
Leveraging a transferability estimation metric facilitates the non-trivial challenge of selecting the optimal model for the downstream task from a pool of pre-trained models. Most existing metrics primarily focus on identifying the statistical relationship between feature embeddings and the corresponding labels within the target dataset, but over-look crucial aspect of model robustness. This oversight may limit their effectiveness in accurately ranking pre-trained models. To address this limitation, we introduce a feature perturbation method that enhances the transferability estimation process by systematically altering the feature space. Our method includes a Spread operation that increases intra-class variability, adding complexity within classes, and an Attract operation that minimizes the distances between different classes, thereby blurring the class boundaries. Through extensive experimentation, we demonstrate the efficacy of our feature perturbation method in providing a more precise and robust estimation of model transferability. Notably, the existing LogMe method exhibited a significant improvement, showing a 28.84% increase in performance after applying our feature perturbation method. The implementation is available at https://github.com/prafful-kumar/enhancing_TE.git
Prafful Kumar Khoba, Zijian Wang 0009, Chetan Arora 0001, Mahsa Baktash
WACV3
2025 LQ-Adapter: ViT-Adapter with Learnable Queries for Gallbladder Cancer Detection from Ultrasound Images
abstract
We focus on the problem of Gallbladder Cancer (GBC) detection from Ultrasound (US) images. The problem presents unique challenges to modern Deep Neural Network (DNN) techniques due to low image quality arising from noise, textures, and viewpoint variations. Tackling such challenges would necessitate precise localization performance by the DNN to identify the discerning features for the downstream malignancy prediction. While several techniques have been proposed in the recent years for the problem, all of these methods employ complex custom architectures. Inspired by the success of foundational models for natural image tasks, along with the use of adapters to fine-tune such models for the custom tasks, we investigate the merit of one such design, ViT-Adapter, for the GBC detection problem. We observe that ViT-Adapter relies pre-dominantly on a primitive CNN-based spatial prior module to inject the localization information via cross-attention, which is inefficient for our problem due to the small pathology sizes, and variability in their appearances due to non-regular structure of the malignancy. In response, we propose, LQ-Adapter, a modified Adapter design for ViT, which improves localization information by leveraging learnable content queries over the basic spatial prior module. Our method surpasses existing approaches, enhancing the mean IoU (mIoU) scores by 5.4%, 5.8%, and 2.7% over ViT-Adapters, DINO, and FocalNet-DINO, respectively on the US image-based GBC detection dataset, and establishing a new state-of-the-art (SOTA). Additionally, we validate the applicability and effectiveness of LQ-Adapter on the Kvasir-Seg dataset for polyp detection from colonoscopy images. Superior performance of our design on this problem as well showcases its capability to handle diverse medical imaging tasks across different datasets. Source code and trained models are publicly released.
Chetan Madan, Mayuna Gupta, Soumen Basu, Pankaj Gupta 0005, Chetan Arora 0001
WACV5
2024 Calibration Transfer via Knowledge Distillation
Ramya Hebbalaguppe, Mayank Baranwal, Kartik Anand, Chetan Arora 0001
ACCV (8)4
2024 GaitW: Enhancing Gait Recognition in the Wild Using Dynamic Information
Daksh Thapar, Jayesh Chaudhari, Sunny Manchanda, Aditya Nigam, Chetan Arora 0001
ACCV (1)5
2024 Examining the Threat Landscape: Foundation Models and Model Stealing
Ankita Raj, Deepankar Varma, Chetan Arora 0001
BMVC3
2024 FocusMAE: Gallbladder Cancer Detection from Ultrasound Videos with Focused Masked Autoencoders
abstract
In recent years, automated Gallbladder Cancer (GBC) detection has gained the attention of researchers. Current state-of-the-art (SOTA) methodologies relying on ultra-sound sonography (US) images exhibit limited generalization, emphasizing the need for transformative approaches. We observe that individual US frames may lack sufficient information to capture disease manifestation. This study advocates for a paradigm shift towards video-based GBC detection, leveraging the inherent advantages of spatiotemporal representations. Employing the Masked Autoencoder (MAE) for representation learning, we address shortcomings in conventional image-based methods. We propose a novel design called FocusMAE to systematically bias the selection of masking tokens from high-information regions, fostering a more refined representation of malignancy. Additionally, we contribute the most extensive US video dataset for GBC detection. We also note that, this is the first study on US video-based GBC detection. We validate the proposed methods on the curated dataset, and report a new SOTA accuracy of 96.4% for the GBC detection problem, against an accuracy of 84% by current Image-based SOTA – GBCNet and RadFormer, and 94.7% by Video-based SOTA – AdaMAE. We further demonstrate the generality of the proposed FocusMAE on a public CT-based Covid detection dataset, reporting an improvement in accuracy by 3.3% over current baselines. Project page with source code, trained models, and data is available at: https://gbc-iitd.github.io/focusmae.
Soumen Basu, Mayuna Gupta, Chetan Madan, Pankaj Gupta 0005, Chetan Arora 0001
CVPR5
2024 ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth Estimation
abstract
In the absence of parallax cues, a learning based single image depth estimation (SIDE) model relies heavily on shading and contextual cues in the image. While this simplicity is attractive, it is necessary to train such models on large and varied datasets, which are difficult to capture. It has been shown that using embeddings from pretrained foundational models, such as CLIP, improves zero shot transfer in several applications. Taking inspiration from this, in our paper we explore the use of global image priors generated from a pretrained ViT model to provide more detailed contextual information. We argue that the embedding vector from a ViT model, pretrained on a large dataset, captures greater relevant information for SIDE than the usual route of generating pseudo image captions, followed by CLIP based text embeddings. Based on this idea, we propose a new SIDE model using a diffusion backbone which is conditioned on ViT embeddings. Our proposed design establishes a new state-of-the-art (SOTA) for SIDE on NYU Depth v2 dataset, achieving Abs Rel error of 0.059(14% improvement) compared to 0.069 by the current SOTA (VPD). And on KITTI dataset, achieving Sq Rel error of 0.139 (2% improvement) compared to 0.142 by the current SOTA (GED). For zero shot transfer with a model trained on NYU Depth v2, we report mean relative improvement of (20%, 23%,81%, 25%) over NeWCRF on (Sun-RGBD, iBimsl, DIODE, HyperSim) datasets, compared to (16%, 18%, 45%, 9%) by ZoEDepth. The code is available in our project page.
Suraj Patni, Aradhye Agarwal, Chetan Arora 0001
CVPR3
2024 Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach
abstract
With the increased importance of autonomous navigation systems has come an increasing need to protect the safety of Vulnerable Road Users (VRUs) such as pedestrians. Predicting pedestrian intent is one such challenging task, where prior work predicts the binary cross/no-cross intention with a fusion of visual and motion features. However, there has been no effort so far to hedge such predictions with human-understandable reasons. We address this issue by introducing a novel problem setting of exploring the intuitive reasoning behind a pedestrian’s intent. In particular, we show that predicting the ‘WHY’ can be very useful in understanding the ‘WHAT’. To this end, we propose a novel, reason-enriched PIE++ dataset consisting of multi-label textual explanations/reasons for pedestrian intent. We also introduce a novel multi-task learning framework called MINDREAD, which leverages a cross-modal representation learning framework for predicting pedestrian intent as well as the reason behind the intent. Our comprehensive experiments show significant improvement of 5.6% and 7% in accuracy and F1-score for the task of intent prediction on the PIE++ dataset using MINDREAD. We also achieved a 4.4% improvement in accuracy on a commonly used JAAD dataset. Extensive evaluation using quantitative/qualitative metrics and user studies shows the effectiveness of our approach.
Vaishnavi Khindkar, Vineeth N. Balasubramanian, Chetan Arora 0001, Anbumani Subramanian, C. V. Jawahar
IROS3
2024 D-MASTER: Mask Annealed Transformer for Unsupervised Domain Adaptation in Breast Cancer Detection from Mammograms
Tajamul Ashraf, Krithika Rangarajan, Mohit Gambhir, Richa Gauba, Chetan Arora 0001
MICCAI (11)5
2024 VideoCutMix: Temporal Segmentation of Surgical Videos in Scarce Data Scenarios
Rohan Raju Dhanakshirur, Mrinal Tyagi, Britty Baby, Ashish Suri, Prem Kumar Kalra, Chetan Arora 0001
MICCAI (6)6
2024 MMBCD: Multimodal Breast Cancer Detection from Mammograms with Clinical History
Kshitiz Jain, Aditya Bansal, Krithika Rangarajan, Chetan Arora 0001
MICCAI (1)4
2024 Follow the Radiologist: Clinically Relevant Multi-view Cues for Breast Cancer Detection from Mammograms
Kshitiz Jain, Krithika Rangarajan, Chetan Arora 0001
MICCAI (1)3
2024 Assessing Risk of Stealing Proprietary Models for Medical Imaging Tasks
Ankita Raj, Harsh Swaika, Deepankar Varma, Chetan Arora 0001
MICCAI (11)4
2024 United We Stand, Divided We Fall: UnityGraph for Unsupervised Procedure Learning from Videos
abstract
Given multiple videos of the same task, procedure learning addresses identifying the key-steps and determining their order to perform the task. For this purpose, existing approaches use the signal generated from a pair of videos. This makes key-steps discovery challenging as the algorithms lack inter-videos perspective. Instead, we propose an unsupervised Graph-based Procedure Learning (GPL) framework. GPL consists of the novel UnityGraph that represents all the videos of a task as a graph to obtain both intra-video and inter-videos context. Further, to obtain similar embeddings for the same key-steps, the embeddings of UnityGraph are updated in an unsupervised manner using the Node2Vec algorithm. Finally, to identify the key-steps, we cluster the embeddings using KMeans. We test GPL on benchmark ProceL, CrossTask, and EgoProceL datasets and achieve an average improvement of 2% on third-person datasets and 3.6% on EgoProceL over the state-of-the-art.
Siddhant Bansal, Chetan Arora 0001, C. V. Jawahar
WACV2
2024 FinderNet: A Data Augmentation Free Canonicalization aided Loop Detection and Closure technique for Point clouds in 6-DOF separation
abstract
We focus on the problem of LiDAR point cloud based loop detection (or Finding) and closure (LDC) for mobile robots. State-of-the-art (SOTA) methods directly generate learned embeddings from a given point cloud, require large data augmentation, and are not robust to wide viewpoint variations in 6 Degrees-of-Freedom (DOF). Moreover, the absence of strong priors in an unstructured point cloud leads to highly inaccurate LDC. In this original approach, we propose independent roll and pitch canonicalization of point clouds using a common dominant ground plane. We discretize the canonicalized point clouds along the axis perpendicular to the ground plane leads to images similar to digital elevation maps (DEMs), which expose strong spatial priors in the scene. Our experiments show that LDC based on learnt embeddings from such DEMs is not only data efficient but also significantly more robust, and generalizable than the current SOTA. We report an (average precision for loop detection, mean absolute translation/rotation error) improvement of (8.4, 16.7/5.43)% on the KITTI08 sequence, and (11.0, 34.0/25.4)% on GPR10 sequence, over the current SOTA. To further test the robustness of our technique on point clouds in 6-DOF motion we create and opensource a custom dataset called Lidar-UrbanFly Dataset (LUF) which consists of point clouds obtained from a LiDAR mounted on a quadrotor. More details on our website https://gsc2001.github.io/FinderNet/
Sudarshan S. Harithas, Gurkirat Singh, Aneesh Chavan, Sarthak Sharma, Suraj Patni, Chetan Arora 0001, K. Madhava Krishna
WACV6
2024 Army of Thieves: Enhancing Black-Box Model Extraction via Ensemble based sample selection
abstract
Machine Learning (ML) models become vulnerable to Model Stealing Attacks (MSA) when they are deployed as a service. In such attacks, the deployed model is queried repeatedly to build a labelled dataset. This dataset allows the attacker to train a thief model that mimics the original model. To maximize query efficiency, the attacker has to select the most informative subset of data points from the pool of available data. Existing attack strategies utilize approaches like Active Learning and Semi-Supervised learning to minimize costs. However, in the black-box setting, these approaches may select sub-optimal samples as they train only one thief model. Depending on the thief model’s capacity and the data it was pretrained on, the model might even select noisy samples that harm the learning process. In this work, we explore the usage of an ensemble of deep learning models as our thief model. We call our attack Army of Thieves(AOT) as we train multiple models with varying complexities to leverage the crowd’s wisdom. Based on the ensemble’s collective decision, uncertain samples are selected for querying, while the most confident samples are directly included in the training data. Our approach is the first one to utilize an ensemble of thief models to perform model extraction. We outperform the base approaches of existing state-of-the-art methods by at least 3% and achieve a 21% higher adversarial sample transferability than previous work for models trained on the CIFAR-10 dataset. Code is available at: https://github.com/akshitjindal1/AOT_WACV.
Akshit Jindal, Vikram Goyal, Saket Anand, Chetan Arora 0001
WACV4
2024 SEMA: Semantic Attention for Capturing Long-Range Dependencies in Egocentric Lifelogs
abstract
Transformer architecture is a defacto standard for modeling global dependency in long sequences. However, quadratic space and time complexity for self-attention prohibits transformers from scaling to extremely long sequences (> 10k). Low-rank decomposition as a non-negative matrix factorization (NMF) of self-attention demonstrates remarkable performance in linear space and time complexity with strong theoretical guarantees. However, our analysis reveals that NMF-based works struggle to capture the rich spatio-temporal visual cues scattered across the long sequences resulting from egocentric lifelogs. To capture such cues, we propose a novel attention mechanism named SEMantic Atention (SEMA), which factorizes the self-attention matrix into a semantically meaningful subspace. We demonstrate SEMA in a representation learning setting, aiming to recover activity patterns in extremely long (weeks-long) egocentric lifelogs using a novel self-supervised training pipeline. Compared to the current state-of-the-art, we report significant improvement in terms of (NMI, AMI, and F-Score) for EgoRoutine, UTE, and Epic Kitchens datasets. Furthermore, to underscore the efficacy of SEMA, we extend its application to conventional video tasks such as online action detection, video recognition, and action localization. Code is available at https://github.com/Pravin74/Semantic_attention/
Pravin Nagar, K. N. Ajay Shastry, Jayesh Chaudhari, Chetan Arora 0001
WACV4
2024 Domain-Aware Knowledge Distillation for Continual Model Generalization
abstract
Generalization on unseen domains is critical for Deep Neural Networks (DNNs) to perform well in real-world applications such as autonomous navigation. However, catastrophic forgetting limits the ability of domain generalization and unsupervised domain adaption approaches to adapt to constantly changing target domains. To overcome these challenges, We propose DoSe framework, a Domain-aware Self-Distillation method based on batch normalization prototypes to facilitate continual model generalization across varying target domains. Specifically, we enforce the consistency of batch normalization statistics between two batches of images sampled from the same target domain distribution between the student and teacher models. To alleviate catastrophic forgetting, we introduce a novel exemplar-based replay buffer to identify difficult samples for the model to retain the knowledge. Specifically, we demonstrate that identifying difficult samples and updating the model periodically using them can help in preserving knowledge learned from previously seen domains. We conduct extensive experiments on two real-world datasets ACDC, C-Driving, and one synthetic dataset SHIFT to verify the efficiency of the proposed DoSe framework. On ACDC, our method outperforms existing SOTA in Domain Generalization, Unsupervised Domain Adaptation, and Daytime settings by 26%, 14%, and 70% respectively.
Nikhil Reddy, Mahsa Baktash, Chetan Arora 0001
WACV3
2024 Favoring One Among Equals - Not a Good Idea: Many-to-one Matching for Robust Transformer based Pedestrian Detection
abstract
We investigate the reasons for lower performance of transformer based pedestrian detection models compared to convolutional neural network (CNN) based ones. CNN models generate dense pedestrian proposals, refine each proposal individually, and follow it up with non-maximal-suppression (NMS) to generate sparse predictions. In contrast, transformer models select one proposal per groundtruth (GT) pedestrian box and backpropagate positive gradient from them. All other proposals, many of them highly similar to the selected ones, are passed negative gradient. Though this leads to sparse predictions, obviating the need of NMS, the arbitrary selection of one among many similar proposals, hinders effective training, and lower accuracy of pedestrian detection. To mitigate the problem, instead of commonly used Kuhn-Munkres matching algorithm, we propose Min-cost-flow based formulation, and incorporate constraints such as, each ground truth box is matched to atleast one proposal, and many equally good proposals can be matched to a single ground truth box. We propose first transformer based pedestrian detection model incorporating our matching algorithm. Extensive experiments reveal that our approach achieves a miss rate (lower is better) of 3.7 / 17.4 / 21.8 / 8.3 / 2.0 on Eurocity / TJU-traffic / TJUcampus / Cityperson / Caltech datasets compared to 4.7 /18.7 / 24.8 / 8.5 / 3.1 by the current SOTA. Code is available at https://ajayshastry08.github.io/flow_matcher
K. N. Ajay Shastry, K. Ravi Sri Teja, Aditya Nigam, Chetan Arora 0001
WACV4
2023 UTRNet: High-Resolution Urdu Text Recognition in Printed Documents
Arjun Ghosh, Chetan Arora 0001
ICDAR (5)3
2023 From Feline Classification to Skills Evaluation: A Multitask Learning Framework for Evaluating Micro Suturing Neurosurgical Skills
abstract
Automated skill evaluation of a trainee is key to the utility of the surgical training system. The focus of this paper is to develop an automated tool for the assessment of trainees for micro-suturing task. The real-life training datasets for the micro-suturing task are often small, with long-tailed distribution, making it difficult to develop machine-learning-based tools for automated assessment. Further, micro-suturing is often performed at various magnifications and suture sizes, which makes the automated assessment more challenging compared to macro-suturing. Hence, currently, assessment is done manually by an expert using the final outcome image. In this paper, we propose a multi-task learning-based convolutional-neural-network regression model to score the effectualness of the micro-suturing task from the final outcome image. We propose a novel equivalent of the logit-adjustment (used in classification) applicable to regression formulation which effectively handles the problems associated with the long-tail distribution of the data. Additionally, we contribute the largest open-access dataset for suturing images and the first dataset pertaining to the micro-suturing task. We also demonstrate that the performance of the proposed algorithm surpasses the performance of human experts and also other state-of-the-art (SOTA) algorithms. The dataset and the code are available at: https: //aineurosurgery.github.io/microsuturing
Rohan Raju Dhanakshirur, Varidh Katiyar, Ashish Suri, Prem Kumar Kalra, Chetan Arora 0001
ICIP6
2023 Parts Based Attention for Highly Occluded Pedestrian Detection with Transformers
abstract
Despite the significant progress made in pedestrian detection in last decade, detecting pedestrians under heavy occlusion still remains a challenging problem. In state of the art (SOTA), convolutional neural network (CNN) based models, the reason is attributed to non-maximal-suppression (NMS), which often erroneously deletes true positives when one pedestrian is occluding other. SOTA transformer based models do not have such NMS step, yet fail to detect highly occluded pedestrians. In this paper, we study the reasons for such failures. We observe that such models first predict key-points, and then compute the attention at the specific key-points. Our analysis reveals that the key-points do not have any preference towards semantically important body parts. Under heavy occlusion, such key-points end up attending to non-discriminative regions or background, leading to false negatives. We take inspiration from the conventional wisdom of detecting objects using their parts, and bias the attention of proposed transformer architecture towards semantically important, and highly discriminative human body parts. The intervention leads to SOTA results on benchmark Citypersons and Caltech datasets, achieving 30.75%, and 32.96% miss-rate (lower is better) respectively, against 32.6%, and 38.2% by the current SOTA. Code is available at https://ajayshastry08.github.io/pa_dino
K. N. Ajay Shastry, Jayesh Chaudhari, Daksh Thapar, Aditya Nigam, Chetan Arora 0001
ICIP5
2023 Gall Bladder Cancer Detection from US Images with only Image Level Labels
Soumen Basu, Ashish Papanai, Pankaj Gupta 0005, Chetan Arora 0001
MICCAI (1)5
2023 Learnable Query Initialization for Surgical Instrument Instance Segmentation
Rohan Raju Dhanakshirur, K. N. Ajay Shastry, Kaustubh Borgavi, Ashish Suri, Prem Kumar Kalra, Chetan Arora 0001
MICCAI (9)6
2023 How Reliable are the Metrics Used for Assessing Reliability in Medical Imaging?
Soumen Basu, Chetan Arora 0001
MICCAI (3)3
2023 Attention Attention Everywhere: Monocular Depth Prediction with Skip Attention
abstract
Monocular Depth Estimation (MDE) aims to predict pixel-wise depth given a single RGB image. For both, the convolutional as well as the recent attention-based models, encoder-decoder-based architectures have been found to be useful due to the simultaneous requirement of global context and pixel-level resolution. Typically, a skip connection module is used to fuse the encoder and decoder features, which comprises of feature map concatenation followed by a convolution operation. Inspired by the demonstrated benefits of attention in a multitude of computer vision problems, we propose an attention-based fusion of encoder and decoder features. We pose MDE as a pixel query refinement problem, where coarsest-level encoder features are used to initialize pixel-level queries, which are then refined to higher resolutions by the proposed Skip Attention Module (SAM). We formulate the prediction problem as ordinal regression over the bin centers that discretize the continuous depth range and introduce a Bin Center Predictor (BCP) module that predicts bins at the coarsest level using pixel queries. Apart from the benefit of image adaptive depth binning, the proposed design helps learn improved depth embedding in initial pixel queries via direct supervision from the ground truth. Extensive experiments on the two canonical datasets, NYUV2 and KITTI, show that our architecture outperforms the state-of-the-art by 5.3% and 3.9%, respectively, along with an improved generalization performance by 9.4% on the SUNRGBD dataset. Code is available at https://github.com/ashutosh1807/PixelFormer.git.
Ashutosh Agarwal, Chetan Arora 0001
WACV2
2023 Reducing Annotation Effort by Identifying and Labeling Contextually Diverse Classes for Semantic Segmentation Under Domain Shift
abstract
In Active Domain Adaptation (ADA), one uses Active Learning (AL) to select a subset of images from the target domain, which are then annotated and used for supervised domain adaptation (DA). Given the large performance gap between supervised and unsupervised DA techniques, ADA allows for an excellent trade-off between annotation cost and performance. Prior art makes use of measures of uncertainty or disagreement of models to identify ‘regions' to be annotated by the human oracle. However, these regions frequently comprise of pixels at object boundaries which are hard and tedious to annotate. Hence, even if the fraction of image pixels annotated reduces, the overall annotation time and the resulting cost still remain high. In this work, we propose an ADA strategy, which given a frame, identifies a set of classes that are hardest for the model to predict accurately, thereby recommending semantically meaningful regions to be annotated in a selected frame. We show that these set of ‘hard' classes are context-dependent and typically vary across frames, and when annotated help the model generalize better. We propose two ADA techniques: the Anchor-based and Augmentation-based approaches to select complementary and diverse regions in the context of the current training set. Our approach achieves 66.6 mIoU on GTA5 →Cityscapes dataset with an annotation budget of 4.7% in comparison to 64.9 mIoU by MADA [22] using 5% of annotations. Our technique can also be used as a decorator for any existing frame-based AL technique, e.g., we report 1.5% performance improvement for CDAL [1] on Cityscapes using our approach.
Sharat Agarwal, Saket Anand, Chetan Arora 0001
WACV3
2023 From Forks to Forceps: A New Framework for Instance Segmentation of Surgical Instruments
abstract
Minimally invasive surgeries and related applications demand surgical tool classification and segmentation at the instance level. Surgical tools are similar in appearance and are long, thin, and handled at an angle. The fine-tuning of state-of-the-art (SOTA) instance segmentation models trained on natural images for instrument segmentation has difficulty discriminating instrument classes. Our research demonstrates that while the bounding box and segmentation mask are often accurate, the classification head misclassifies the class label of the surgical instrument. We present a new neural network framework that adds a classification module as a new stage to existing instance segmentation models. This module specializes in improving the classification of instrument masks generated by the existing model. The module comprises multi-scale mask attention, which attends to the instrument region and masks the distracting background features. We propose training our classifier module using metric learning with arc loss to handle low inter-class variance of surgical instruments. We conduct exhaustive experiments on the benchmark datasets EndoVis2017 and EndoVis2018. We demonstrate that our method outperforms all (more than 18) SOTA methods compared with, and improves the SOTA performance by at least 12 points (20%) on the EndoVis2017 benchmark challenge and generalizes effectively across the datasets. Project page with source code is available at nets-iitd.github.io/s3net.
Britty Baby, Daksh Thapar, Mustafa Chasmai, Tamajit Banerjee, Kunal Dargan, Ashish Suri, Subhashis Banerjee, Chetan Arora 0001
WACV8
2023 RadFormer: Transformers with global-local attention for interpretable and accurate Gallbladder Cancer detection
Soumen Basu, Pratyaksha Rana, Pankaj Gupta 0005, Chetan Arora 0001
Medical Image Anal.5
2023 Generating Personalized Summaries of Day Long Egocentric Videos
abstract
The popularity of egocentric cameras and their always-on nature has lead to the abundance of day long first-person videos. The highly redundant nature of these videos and extreme camera-shakes make them difficult to watch from beginning to end. These videos require efficient summarization tools for consumption. However, traditional summarization techniques developed for static surveillance videos or highly curated sports videos and movies are either not suitable or simply do not scale for such hours long videos in the wild. On the other hand, specialized summarization techniques developed for egocentric videos limit their focus to important objects and people. This paper presents a novel unsupervised reinforcement learning framework to summarize egocentric videos both in terms of length and the content. The proposed framework facilitates incorporating various prior preferences such as faces, places, or scene diversity and interactive user choice in terms of including or excluding the particular type of content. This approach can also be adapted to generate summaries of various lengths, making it possible to view even 1-minute summaries of one's entire day. When using the facial saliency-based reward, we show that our approach generates summaries focusing on social interactions, similar to the current state-of-the-art (SOTA). The quantitative comparisons on the benchmark Disney dataset show that our method achieves significant improvement in Relaxed F-Score (RFS) (29.60 compared to 19.21 from SOTA), BLEU score (0.68 compared to 0.67 from SOTA), Average Human Ranking (AHR), and unique events covered. Finally, we show that our technique can be applied to summarize traditional, short, hand-held videos as well, where we improve the SOTA F-score on benchmark SumMe and TVSum datasets from 41.4 to 46.40 and 57.6 to 58.3 respectively. We also provide a Pytorch implementation and a web demo at https://pravin74.github.io/Int-sum/index.html.
Pravin Nagar, Anuj Rathore, C. V. Jawahar, Chetan Arora 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 ScribbleNet: Efficient interactive annotation of urban city scenes for semantic segmentation
Bhavani Sambaturu, Ashutosh Gupta 0009, C. V. Jawahar, Chetan Arora 0001
Pattern Recognit.4
2022 Surpassing the Human Accuracy: Detecting Gallbladder Cancer from USG Images with Curriculum Learning
abstract
We explore the potential of CNN-based models for gall-bladder cancer (GBC) detection from ultrasound (USG) images as no prior study is known. USG is the most common diagnostic modality for GB diseases due to its low cost and accessibility. However, USG images are challenging to analyze due to low image quality, noise, and varying viewpoints due to the handheld nature of the sensor. Our exhaustive study of state-of-the-art (SOTA) image classification techniques for the problem reveals that they often fail to learn the salient GB region due to the presence of shadows in the USG images. SOTA object detection techniques also achieve low accuracy because of spurious textures due to noise or adjacent organs. We propose GBCNet to tackle the challenges in our problem. GBCNet first extracts the regions of interest (ROIs) by detecting the GB (and not the cancer), and then uses a new multi-scale, second-order pooling architecture specializing in classifying GBC. To effectively handle spurious textures, we propose a curriculum inspired by human visual acuity, which reduces the texture biases in GBCNet. Experimental results demonstrate that GBC-Net significantly outperforms SOTA CNN models, as well as the expert radiologists. Our technical innovations are generic to other USG image analysis tasks as well. Hence, as a validation, we also show the efficacy of GBCNet in detecting breast cancer from USG images. Project page with source code, trained models, and data is available at https://GBC-iitd.github.io/GBCnet.
Soumen Basu, Pratyaksha Rana, Pankaj Gupta 0005, Chetan Arora 0001
CVPR5
2022 A Stitch in Time Saves Nine: A Train-Time Regularizing Loss for Improved Neural Network Calibration
abstract
Deep Neural Networks (dnns) are known to make over-confident mistakes, which makes their use problematic in safety-critical applications. State-of-the-art (sota) calibration techniques improve on the confidence of predicted labels alone, and leave the confidence of non-max classes (e.g. top-2, top-5) uncalibrated. Such calibration is not suitable for label refinement using post-processing. Further, most sota techniques learn a few hyper-parameters post-hoc, leaving out the scope for image, or pixel specific calibration. This makes them unsuitable for calibration under domain shift, or for dense prediction tasks like semantic segmentation. In this paper, we argue for intervening at the train time itself, so as to directly produce calibrated dnn models. We propose a novel auxiliary loss function: Multi-class Difference in Confidence and Accuracy (mdca), to achieve the same. mdca can be used in conjunction with other application/task specific loss functions. We show that training with mdca leads to better calibrated models in terms of Expected Calibration Error (ece), and Static Calibration Error (sce) on image classification, and segmentation tasks. We report ece (sce) score of 0.72 (1.60) on the cifar 100 dataset, in comparison to 1.90 (1.71) by the sota. Under domain shift, a ResNet-18 model trained on pacs dataset using mdca gives a average ece (sce) score of 19.7 (9.7) across all domains, compared to 24.2 (11.8) by the sota. For segmentation task, we report a 2× reduction in calibration error on pascal-voc dataset in comparison to Focal Loss [32]. Finally, mdca training improves calibration even on imbalanced data, and for natural language classification tasks.
Ramya Hebbalaguppe, Jatin Prakash, Neelabh Madan, Chetan Arora 0001
CVPR4
2022 Merry Go Round: Rotate a Frame and Fool a DNN
abstract
A large proportion of videos captured today are first person videos shot from wearable cameras. Similar to other computer vision tasks, Deep Neural Networks (DNNs) are the workhorse for most state-of-the-art (SOTA) egocentric vision techniques. On the other hand DNNs are known to be susceptible to Adversarial Attacks (AAs) which add imperceptible noise to the input. Both black-box, as well as white-box attacks on image as well as video analysis tasks have been shown. We observe that most AA techniques basically add intensity perturbation to an image. Even for videos, the same process is essentially repeated for each frame independently. We note that definition of imperceptibility used for images may not be applicable for videos, where a small intensity change happening randomly in two consecutive frames may still be perceptible. In this paper we make a key novel suggestion to use perturbation in optical flow to carry out AAs on a video analysis system. Such perturbation is especially useful for egocentric videos, because there is lot of shake in the egocentric videos anyways, and adding a little more, keeps it highly imperceptible. In general our idea can be seen as adding structured, parametric noise as the adversarial perturbation. Our implementation of the idea by adding 3D rotations to the frames, reveal that using our technique, one can mount a black-box AA on an egocentric activity detection system in one-third of the queries compared to the SOTA AA technique.
Daksh Thapar, Aditya Nigam, Chetan Arora 0001
CVPR3
2022 Use of Metric Learning for the Recognition of Handwritten Digits, and its Application to Increase the Outreach of Voice-based Communication Platforms
abstract
Initiation, monitoring, and evaluation of development programmes can involve field-based data collection about project activities. This data collection through digital devices may not always be feasible though, for reasons such as unaffordability of smartphones and tablets by field-based cadre, or shortfalls in their training and capacity building. Paper-based data collection has been argued to be more appropriate in several contexts, with automated digitization of the paper forms through OCR (Optical Character Recognition) and OMR (Optical Mark Recognition) techniques. We contribute with providing a large dataset of handwritten digits, and deep learning based models and methods built using this data, that are effective in real-world environments. We demonstrate the deployment of these tools in the context of a maternal and child health and nutrition awareness project, which uses IVR (Interactive Voice Response) systems to provide awareness information to rural women SHG (Self Help Group) members in north India. Paper forms were used to collect phone numbers of the SHG members at scale, which were digitized using the OCR tools developed by us, and used to push almost 4 million phone calls. The data, model, and code have been released in the open-source domain.
Devesh Pant, Dibyendu Talukder, Deepak Kumar 0013, Rachit Pandey, Aaditeshwar Seth, Chetan Arora 0001
COMPASS6
2022 My View is the Best View: Procedure Learning from Egocentric Videos
Siddhant Bansal, Chetan Arora 0001, C. V. Jawahar
ECCV (13)2
2022 Master of All: Simultaneous Generalization of Urban-Scene Segmentation to All Adverse Weather Conditions
Nikhil Reddy, Abhinav Singhal, Mahsa Baktash, Chetan Arora 0001
ECCV (39)5
2022 Depthformer: Multiscale Vision Transformer for Monocular Depth Estimation with Global Local Information Fusion
abstract
Attention-based models such as transformers have shown outstanding performance on dense prediction tasks, such as semantic segmentation, owing to their capability of capturing long-range dependency in an image. However, the benefit of transformers for monocular depth prediction has seldom been explored so far. This paper benchmarks var-ious transformer-based models for the depth estimation task on an indoor NYUV2 dataset and an outdoor KITTI dataset. We propose a novel attention-based architecture, Depthformer for monocular depth estimation that uses multi-head self-attention to produce the multiscale feature maps, which are effectively combined by our proposed de-coder network. We also propose a Transbins module that divides the depth range into bins whose center value is estimated adaptively per image. The final depth estimated is a linear combination of bin centers for each pixel. Trans-bins module takes advantage of the global receptive field using the transformer module in the encoding stage. Experimental results on NYUV2 and KITTI depth estimation benchmark demonstrate that our proposed method improves the state-of-the-art by 3.3%, and 3.3% respectively in terms of Root Mean Squared Error (RMSE). Code is available at https://github.com/ashutosh1807/Depthformer.git.
Ashutosh Agarwal, Chetan Arora 0001
ICIP2
2022 Representation Learning Using Rank Loss for Robust Neurosurgical Skills Evaluation
abstract
Surgical simulators provide hands-on training and learning of the necessary psychomotor skills. Automated skill evaluation of the trainee doctors based on the video of a task being performed by them is an important key step for the optimal utilization of such simulators. However, current skill evaluation techniques require accurate tracking information of the instruments which restricts their applicability to robot assisted surgeries only. In this paper, we propose a novel neural network architecture that can perform skill evaluation using video data alone (and no tracking information). Given the small dataset available for training such a system, the network trained using ℓ2regression loss easily overfits the training data. We propose a novel rank loss to help learn robust representation, leading to 5% improvement for skill score prediction on the benchmark JIGSAWS dataset. To demonstrate the applicability of our method on non-robotic surgeries, we contribute a new neuro-endoscopic technical skills (NETS) training dataset comprising of 100 short videos of 12 subjects. Our method achieved 27% improvement over the state of the art on the NETS dataset. Project page with source code, and data is available at nets-iitd.github.io/nets-v1.
Britty Baby, Mustafa Chasmai, Tamajit Banerjee, Ashish Suri, Subhashis Banerjee, Chetan Arora 0001
ICIP6
2022 New Objects on the Road? No Problem, We'll Learn Them Too
abstract
Object detection plays an essential role in providing localization, path planning, and decision making capabilities in autonomous navigation systems. However, existing object detection models are trained and tested on a fixed number of known classes. This setting makes the object detection model difficult to generalize well in real-world road scenarios while encountering an unknown object. We address this problem by introducing our framework that handles the issue of unknown object detection and updates the model when unknown object labels are available. Next, our solution includes three major components that address the inherent problems present in the road scene datasets. The novel components are a) Feature-Mix that improves the unknown object detection by widening the gap between known and unknown classes in latent feature space, b) Focal regression loss handling the problem of improving small object detection and intra-class scale variation, and c) Curriculum learning further enhances the detection of small objects. We use Indian Driving Dataset (IDD) and Berkeley Deep Drive (BDD) dataset for evaluation. Our solution provides state-of-the-art performance on open-world evaluation metrics. We hope this work will create new directions for open-world object detection for road scenes, making it more reliable and robust autonomous systems.
Shyam Nandan Rai, K. J. Joseph, Rohit Saluja, Vineeth N. Balasubramanian, Chetan Arora 0001, Anbumani Subramanian, C. V. Jawahar
IROS6
2022 Unsupervised Contrastive Learning of Image Representations from Ultrasound Videos with Hard Negative Mining
Soumen Basu, Somanshu Singla, Pratyaksha Rana, Pankaj Gupta 0005, Chetan Arora 0001
MICCAI (4)6
2022 A Novel Data Augmentation Technique for Out-of-Distribution Sample Detection Using Compounded Corruptions
Ramya Hebbalaguppe, Soumya Suvra Ghosal, Jatin Prakash, Harshad Khadilkar, Chetan Arora 0001
ECML/PKDD (3)5
2022 Does Data Repair Lead to Fair Models? Curating Contextually Fair Data To Reduce Model Bias
abstract
Contextual information is a valuable cue for Deep Neural Networks (DNNs) to learn better representations and improve accuracy. However, co-occurrence bias in the training dataset may hamper a DNNmodel’s generalizabil- ity to unseen scenarios in the real world. For example, in COCO [26], many object categories have a much higher cooccurrence with men compared to women, which can bias a DNN’s prediction in favor of men. Recent works have focused on task-specific training strategies to handle bias in such scenarios, but fixing the available data is often ignored. In this paper, we propose a novel and more generic solution to address the contextual bias in the datasets by selecting a subset of the samples, which is fair in terms of the co-occurrence with various classes for a protected attribute. We introduce a data repair algorithm using the coefficient of variation( cv), which can curate fair and contextually balanced data for a protected class(es). This helps in training a fair model irrespective of the task, architecture or training methodology. Our proposed solution is simple, effective and can even be used in an active learning setting where the data labels are not present or being generated incrementally. We demonstrate the effectiveness of our algorithm for the task of object detection and multi-label image classification across different datasets. Through a series of experiments, we validate that curating contextually fair data helps make model predictions fair by balancing the true positive rate for the protected class across groups without compromising on the model’s overall performance. Code: https://github.com/sumanyumuku98/contextual-bias
Sharat Agarwal, Sumanyu Muku, Saket Anand, Chetan Arora 0001
WACV4
2022 Multi-Domain Incremental Learning for Semantic Segmentation
abstract
Recent efforts in multi-domain learning for semantic segmentation attempt to learn multiple geographical datasets in a universal, joint model. A simple fine-tuning experiment performed sequentially on three popular road scene segmentation datasets demonstrates that existing segmentation frameworks fail at incrementally learning on a series of visually disparate geographical domains. When learning a new domain, the model catastrophically forgets previously learned knowledge. In this work, we pose the problem of multi-domain incremental learning for semantic segmentation. Given a model trained on a particular geographical domain, the goal is to (i) incrementally learn a new geographical domain, (ii) while retaining performance on the old domain, (iii) given that the previous domain’s dataset is not accessible. We propose a dynamic architecture that assigns universally shared, domain-invariant parameters to capture homogeneous semantic features present in all domains, while dedicated domain-specific parameters learn the statistics of each domain. Our novel optimization strategy helps achieve a good balance between retention of old knowledge (stability) and acquiring new knowledge (plasticity). We demonstrate the effectiveness of our proposed solution on domain incremental settings pertaining to real-world driving scenes from roads of Germany (Cityscapes), the United States (BDD100k), and India (IDD).1
Prachi Garg, Rohit Saluja, Vineeth N. Balasubramanian, Chetan Arora 0001, Anbumani Subramanian, C. V. Jawahar
WACV4
2022 To miss-attend is to misalign! Residual Self-Attentive Feature Alignment for Adapting Object Detectors
abstract
Advancements in adaptive object detection can lead to tremendous improvements in applications like autonomous navigation, as they alleviate the distributional shifts along the detection pipeline. Prior works adopt adversarial learning to align image features at global and local levels, yet the instance-specific misalignment persists. Also, adaptive object detection remains challenging due to visual diversity in background scenes and intricate combinations of objects. Motivated by structural importance, we aim to attend prominent instance-specific regions, overcoming the feature misalignment issue. We propose a novel resIduaL seLf-attentive featUre alignMEnt (ILLUME) method for adaptive object detection. ILLUME comprises Self-Attention Feature Map (SAFM) module that enhances structural attention to object-related regions and thereby generates domain invariant features. Our approach significantly reduces the domain distance with the improved feature alignment of the instances. Qualitative results demonstrate the ability of ILLUME to attend important object instances required for alignment. Experimental results on several benchmark datasets show that our method outperforms the existing state-of-the-art approaches.
Vaishnavi Khindkar, Chetan Arora 0001, Vineeth N. Balasubramanian, Anbumani Subramanian, Rohit Saluja, C. V. Jawahar
WACV2
2022 FLUID: Few-Shot Self-Supervised Image Deraining
abstract
Self-supervised methods have shown promising results in denoising and dehazing tasks, where the collection of the paired dataset is challenging and expensive. However, we find that these methods fail to remove the rain streaks when applied for image deraining tasks. The method’s poor performance is due to the explicit assumptions: (i) the distribution of noise or haze is uniform and (ii) the value of a noisy or hazy pixel is independent of its neighbors. The rainy pixels are non-uniformly distributed, and it is not necessarily dependant on its neighboring pixels. Hence, we conclude that the self-supervised method needs to have some prior knowledge about rain distribution to perform the deraining task. To provide this knowledge, we hypothesize a network trained with minimal supervision to estimate the likelihood of rainy pixels. This leads us to our proposed method called FLUID: Few Shot Sel f-Supervised Image Deraining.We perform extensive experiments and comparisons with existing image deraining and few-shot image-to-image translation methods on Rain 100L and DDN-SIRR datasets containing real and synthetic rainy images. In addition, we use the Rainy Cityscapes dataset to show that our method trained in a few-shot setting can improve semantic segmentation and object detection in rainy conditions. Our approach obtains a mIoU gain of 51.20 over the current best-performing deraining method. [Project Page]
Shyam Nandan Rai, Rohit Saluja, Chetan Arora 0001, Vineeth N. Balasubramanian, Anbumani Subramanian, C. V. Jawahar
WACV3
2021 Anonymizing Egocentric Videos
abstract
In egocentric videos, the face of a wearer capturing the video is never captured. This gives a false sense of security that the wearer’s privacy is preserved while sharing such videos. However, egocentric cameras are typically harnessed to wearer’s head, and hence, also capture wearer’s gait. Recent works have shown that wearer gait signatures can be extracted from egocentric videos, which can be used to determine if two egocentric videos have the same wearer. In a more damaging scenario, one can even recognize a wearer using hand gestures from egocentric videos, or identify a wearer in third person videos such as from a surveillance camera. We believe, this could be a death knell in sharing of egocentric videos, and fatal for egocentric vision research. In this work, we suggest a novel technique to anonymize egocentric videos, which create carefully crafted, but small, and imperceptible optical flow perturbations in an egocentric video’s frames. Importantly, these perturbations do not affect object detection or action/activity recognition from egocentric videos but are strong enough to dis-balance the gait recovery process. In our experiments on benchmark EPIC-Kitchens dataset, the proposed perturbation degrades the wearer recognition performance of [42], from 66.3% to 13.4%, while preserving the activity recognition performance of [10] from 89.6% to 87.4%. To test our anonymization with more wearer recognition techniques, we also developed a stronger, and more generalizable wearer recognition method based on camera egomotion cues. The approach achieves state-ofthe-art (SOTA) performance of 59.67% on EPIC-Kitchens, compared to 55.06% by [42]. However, the accuracy of our recognition technique also drops to 12% using the proposed anonymizing perturbations.
Daksh Thapar, Aditya Nigam, Chetan Arora 0001
ICCV3
2021 Identifying Physically Realizable Triggers for Backdoored Face Recognition Networks
abstract
Backdoor attacks embed a hidden functionality into deep neural networks, causing the network to display anomalous behavior when activated by a predetermined pattern in the input (Trigger), while behaving well otherwise on public test data. Recent works have shown that backdoored face recognition (FR) systems can respond to natural-looking triggers like a particular pair of sunglasses. Such attacks pose a serious threat to the applicability of FR systems in high-security applications. We propose a novel technique to (1) detect whether an FR network is compromised with a natural, physically realizable trigger, and (2) identify such triggers given a compromised network. We demonstrate the effectiveness of our methods with a compromised FR network, where we are able to identify the trigger (e.g. green-sunglasses or redbowtie) with a top-5 accuracy of 74%, whereas a naïve brute force baseline achieves 56% accuracy.
Ankita Raj, Ambar Pal, Chetan Arora 0001
ICIP3
2021 Efficient and Generic Interactive Segmentation Framework to Correct Mispredictions During Clinical Evaluation of Medical Images
Bhavani Sambaturu, Ashutosh Gupta 0009, C. V. Jawahar, Chetan Arora 0001
MICCAI (2)4
2021 VmAP: A Fair Metric for Video Object Detection
abstract
Video object detection is the task of detecting objects in a sequence of frames, typically, with a significant overlap in content among consecutive frames. Mean Average Precision (mAP) was originally proposed for evaluating object detection techniques in independent frames, but has been used for evaluating video based object detectors as well. This is undesirable since the average precision over all frames masks the biases that a certain object detector might have against certain types of objects depending on the number of frames for which the object is present in a video sequence. In this paper we show several disadvantages of mAP as a metric for evaluating video based object detection. Specifically, we show that: (a) some object detectors could be severely biased against some specific kind of objects, such as small, blurred, or low contrast objects, and such differences may not reflect in mAP based evaluation, (b) operating a video based object detector at the best frame based precision/recall value (high F1 score) may lead to many false positives without a significant increase in the number of objects detected. (c) mAP does not take into account that tracking can be potentially used to recover missed detections in the temporal neighborhood while this can be account for while evaluating detectors. As an alternate, we suggest a novel evaluation metric (VmAP) which takes the focus away from evaluating detections on every frame. Unlike mAP, VmAP rewards a high recall of different object views throughout the video. We form sets of bounding boxes having similar views of an object in a temporal neighborhood and use a set-level recall for evaluation. We show that VmAP is able to address all the challenges with the mAP listed above. Our experiments demonstrate hidden biases in object detectors, shows upto 99% reduction in false positives while maintaining similar object recall and shows a 9% improvement in correlation with post-tracking performance.
Anupam Sobti, Vaibhav Mavi, M. Balakrishnan, Chetan Arora 0001
ACM Multimedia4
2020 Contextual Diversity for Active Learning
Sharat Agarwal, Himanshu Arora, Saket Anand, Chetan Arora 0001
ECCV (16)4
2020 An Inference Algorithm for Multi-label MRF-MAP Problems with Clique Size 100
Ishant Shanu, Siddhant Bharti, Chetan Arora 0001, S. N. Maheshwari
ECCV (20)3
2020 Is Sharing of Egocentric Video Giving Away Your Biometric Signature?
Daksh Thapar, Chetan Arora 0001, Aditya Nigam
ECCV (17)2
2020 Robust and Adaptive Traffic Signal Control for Unstructured Driving Scenarios in the Developing World
abstract
s- Adaptive traffic signal control (ATSC) systems are key to efficient traffic management in smart cities. While RADAR or inductive loop based systems are regularly used for the task in the developed world, financial and infrastructure constraints precludes use of sophisticated and expensive sensors in the developing countries. RGB cameras are cheap and easy to deploy, making their usage attractive for ATSC. However, unstructured driving scenarios, and the requirement to work 24×7 pose a significant challenge to the camera based systems. In this paper we propose to use traffic tailback length (TTL) for training a Deep Reinforcement Learning (RL) Agent in an ATSC system. We show that TTL can be calculated by simply performing a `background subtraction' on the input video captured from an RGB camera. The background subtraction operation is well studied in computer vision and can be robustly performed in low illumination, and rainy, snowy, or hazy weather. The operation does not require detecting vehicles explicitly, thus making the proposed system applicable in unstructured driving conditions, involving unseen vehicles. We show, in a simulated environment using SUMO, that our technique outperform the state of the art, which uses explicit vehicle detection information as well as speed of the vehicle, in terms of vehicle throughput, queue length, and average delay.
Ishan Maheshwari, Ujwal Padam Tewari, Chetan Arora 0001
IV3
2020 Concept Drift Detection for Multivariate Data Streams and Temporal Segmentation of Daylong Egocentric Videos
abstract
The long and unconstrained nature of egocentric videos makes it imperative to use temporal segmentation as an important pre-processing step for many higher-level inference tasks. Activities of the wearer in an egocentric video typically span over hours and are often separated by slow, gradual changes. Furthermore, the change of camera viewpoint due to the wearer's head motion causes frequent and extreme, but, spurious scene changes. The continuous nature of boundaries makes it difficult to apply traditional Markov Random Field (MRF) pipelines relying on temporal discontinuity, whereas deep Long Short Term Memory (LSTM) networks gather context only upto a few hundred frames, rendering them ineffective for egocentric videos. In this paper, we present a novel unsupervised temporal segmentation technique especially suited for day-long egocentric videos. We formulate the problem as detecting concept drift in a time-varying, non i.i.d. sequence of frames. Statistically bounded thresholds are calculated to detect concept drift between two temporally adjacent multivariate data segments with different underlying distributions while establishing guarantees on false positives. Since the derived threshold indicates confidence in the prediction, it can also be used to control the granularity of the output segmentation. Using our technique, we report significantly improved state of the art f-measure for daylong egocentric video datasets, as well as photostream datasets derived from them: HUJI~(73.01%, 59.44%), UTEgo~(58.41%, 60.61%) and Disney~(67.63%, 68.83%).
Pravin Nagar, Mansi Khemka, Chetan Arora 0001
ACM Multimedia3
2020 Recognizing Camera Wearer from Hand Gestures in Egocentric Videos: https: //egocentricbiometric.github.io/
abstract
Wearable egocentric cameras are typically harnessed to a wearer's head, giving them the unique advantage of capturing their points of view. Hoshen and Peleg have shown that egocentric cameras indirectly capture the wearer's gait, which can be used to identify a wearer based on their egocentric videos. The authors have shown a wearer recognition accuracy of up to 77% over 32 subjects. However, an important limitation of their work is that such gait features can be extracted only from walking sequences of a wearer. In this work, we take the privacy threat a notch higher and show that even the wearer's hand gestures, as seen through an egocentric video, leak wearer's identity. We have designed a model to extract and match hand gesture signatures from egocentric videos. We demonstrate the threat on the EPIC kitchen dataset containing 55 hours of the egocentric videos acquired from 32 subjects doing various activities. We show that: (1) Our model can recognize a wearer with an accuracy of up to 73% based on the same activity, i.e., the model has seen 'cut' activity by a wearer in the train set, and recognizes the wearer based on another 'cut' activity by him/her while testing. (2) The hand gesture signatures transfer across activities, i.e., even if our model does not see 'cut' activity of a wearer at the train time, but sees other activities such as 'wash', 'mix' etc., the model can still recognize a wearer with an accuracy of up to 60%, by matching hand gesture signatures of 'cut' at test time with train time signatures of 'wash' or 'mix'. (3) The hand gesture features even transfer across subjects, i.e., even if the model has not seen any activity by some subject, one can still verify a wearer (open-set), and predict that the same wearer has performed both activities with an Equal Error Rate of 15.21%. The code, trained models are available at https://egocentricbiometric.github.io/
Daksh Thapar, Aditya Nigam, Chetan Arora 0001
ACM Multimedia3
2019 Multi-sensor Energy Efficient Obstacle Detection
abstract
With the improvement in technology, both the cost and the power requirement of cameras, as well as other sensors have come down significantly. It has allowed these sensors to be integrated into portable as well as wearable systems. Such systems are usually operated in a hands-free and always-on manner where they need to function continuously in a variety of scenarios. In such situations, relying on a single sensor or a fixed sensor combination can be detrimental to both performance as well as energy requirements. Consider the case of an obstacle detection task. Here using an RGB camera helps in recognizing the obstacle type but takes much more energy than an ultrasonic sensor. Infrared cameras can perform better than RGB camera at night but consume twice the energy. Therefore, an efficient system must use a combination of sensors, with an adaptive control that ensures the use of the sensors appropriate to the context. In this adaptation, one needs to consider both performance and energy and their trade-off. In this paper, we explore the strengths of different sensors as well their trade-off for developing a deep neural network based wearable device. We choose a specific case study in the context of a mobility assistance device for the visually impaired. The device detects obstacles in the path of a visually impaired person and is required to operate both at day and night with minimal energy to increase the usage time on a single charge. The device employs multiple sensors: ultrasonic sensor, RGB Camera, and NIR Camera along with a deep neural network accelerator for speeding up computation. We show that by adaptively choosing the appropriate sensor for the context, we can achieve up to 90% reduction in energy while maintaining comparable performance to a single sensor system.
Anupam Sobti, M. Balakrishnan, Chetan Arora 0001
DSD3
2019 Generating 1 Minute Summaries of Day Long Egocentric Videos
abstract
The popularity of egocentric cameras and their always-on nature has lead to the abundance of day-long first-person videos. Because of the extreme shake and highly redundant nature, these videos are difficult to watch from beginning to end and often require summarization tools for their efficient consumption. However, traditional summarization techniques developed for static surveillance videos, or highly curated sports videos and movies are, either, not suitable or simply do not scale for such hours long videos in the wild. On the other hand, specialized summarization techniques developed for egocentric videos limit their focus to important objects and people. In this paper, we present a novel unsupervised reinforcement learning technique to generate video summaries from day long egocentric videos. Our approach can be adapted to generate summaries of various lengths making it possible to view even 1-minute summaries of one's entire day. The technique can also be adapted to various rewards, such as distinctiveness and indicativeness of the summary. When using the facial saliency-based reward, we show that our approach generates summaries focusing on social interactions, similar to the current state-of-the-art (SOTA). Quantitative comparison on the benchmark Disney dataset shows that our method achieves significant improvement in Relaxed F-Score (RFS) (32.56 vs. 19.21) and BLEU score (12.12 vs. 10.64). Finally, we show that our technique can be applied for summarizing traditional, short, hand-held videos as well, where we improve the SOTA F-score on benchmark SumMe and TVSum datasets from 41.4 to 45.6 and 57.6 to 59.1 respectively.
Anuj Rathore, Pravin Nagar, Chetan Arora 0001, C. V. Jawahar
ACM Multimedia3
2019 EGO-SLAM: A Robust Monocular SLAM for Egocentric Videos
abstract
Regardless of the tremendous progress, a truly general purpose pipeline for Simultaneous Localization and Mapping (SLAM) remains a challenge. We investigate the reported failure of state of the art (SOTA) SLAM techniques on egocentric videos. We find that the dominant 3D rotations, low parallax between successive frames, and primarily forward motion in egocentric videos are the most common causes of failures. The incremental nature of SOTA SLAM, in the presence of unreliable pose and 3D estimates in egocentric videos, with no opportunities for global loop closures, generates drifts and leads to the eventual failures of such techniques. Taking inspiration from batch mode Structure from Motion (SFM) techniques, we propose to solve SLAM as an SFM problem over the sliding temporal windows. This makes the problem well constrained. Further, we propose to initialize the camera poses using 2D rotation averaging, followed by translation averaging before structure estimation using bundle adjustment. This helps in stabilizing the camera poses when 3D estimates are not reliable. We show that the proposed SLAM technique, incorporating the two key ideas works successfully for long, shaky egocentric videos where other SOTA techniques have been reported to fail. Qualitative and quantitative comparisons on publicly available egocentric video datasets validate our results.
Suvam Patra, Kartikeya Gupta, Faran Ahmad, Chetan Arora 0001, Subhashis Banerjee
WACV4
2019 Gait metric learning siamese network exploiting dual of spatio-temporal 3D-CNN intra and LSTM based inter gait-cycle-segment features
Daksh Thapar, Gaurav Jaswal, Aditya Nigam, Chetan Arora 0001
Pattern Recognit. Lett.4
2018 Inference in Higher Order MRF-MAP Problems With Small and Large Cliques
abstract
Higher Order MRF-MAP formulation has been a popular technique for solving many problems in computer vision. Inference in a general MRF-MAP problem is NP Hard, but can be performed in polynomial time for the special case when potential functions are submodular. Two popular combinatorial approaches for solving such formulations are flow based and polyhedral approaches. Flow based approaches work well with small cliques and in that mode can handle problems with millions of variables. Polyhedral approaches can handle large cliques but in small numbers. We show in this paper that the variables in these seemingly disparate techniques can be mapped to each other. This allows us to combine the two styles in a joint framework exploiting the strength of both of them. Using the proposed joint framework, we are able to perform tractable inference in MRF-MAP problems with millions of variables and a mix of small and large cliques, a formulation which can not be solved by either of the two styles individually. We show applicability of this hybrid framework on object segmentation problem as an example of a situation where quality of results is significantly better than systems which are based only on the use of small or large cliques.
Ishant Shanu, Chetan Arora 0001, S. N. Maheshwari
CVPR2
2018 Exploiting Texture Cues for Clothing Parsing in Fashion Images
abstract
We focus on the problem of parsing fashion images for detecting various types of clothing and style. The current state-of-the-art techniques for the problem are mostly based on variations of the SegNet model. The techniques formulate the problem as segmentation and typically rely on geometrical shapes and position to segment the image. However, specifically for fashion images, each clothing item is made of specific type of materials with characteristic visual texture patterns. Exploiting the texture for recognizing the clothing type is an important cue which has been ignored so far by the state-of-the-art. In this paper, we propose a two-stream deep neural network architecture for fashion image parsing. While the first stream uses the regular fully convolutional network segmentation architecture to give accurate spatial segments, the second stream provides texture features learned from hand-crafted Gabor feature maps used as input, and helps in determining the clothing type resulting in improved recognition of the various segments. Our experiments show that, the proposed two-stream architecture successfully reduces the confusion between the clothing types, having similar visual shapes in the images but different material. Our approach achieves state-of-the-art results on the standard benchmark datasets, such as Fashionista and CFPD.
Tarasha Khurana, Kushagra Mahajan, Chetan Arora 0001, Atul Rai
ICIP3
2018 U-Segnet: Fully Convolutional Neural Network Based Automated Brain Tissue Segmentation Tool
abstract
Automated brain tissue segmentation into white matter (WM), gray matter (GM), and cerebro-spinal fluid (CSF) from magnetic resonance images (MRI) is helpful in the diagnosis of neuro-disorders such as epilepsy, Alzheimer's, multiple sclerosis, etc. However, thin GM structures at the periphery of cortex and smooth transitions on tissue boundaries such as between GM and WM, or WM and CSF pose difficulty in building a reliable segmentation tool. This paper proposes a Fully Convolutional Neural Network (FCN) tool, that is a hybrid of two widely used deep learning segmentation architectures SegNet and U-Net, for improved brain tissue segmentation. We propose a skip connection inspired from U-Net, in the SegNet architetcure, to incorporate fine multiscale information for better tissue boundary identification. We show that the proposed U-SegNet architecture, improves segmentation performance, as measured by average dice ratio, to 89.74% on the widely used IBSR dataset consisting of T-1 weighted MRI volumes of 18 subjects.
Pulkit Kumar, Pravin Nagar, Chetan Arora 0001, Anubha Gupta
ICIP3
2018 Pose Aware Fine-Grained Visual Classification Using Pose Experts
abstract
We focus on the problem of fine-grained visual classification (FGVC). We posit that unreasonable effectiveness of the state-of-the-art in this area is because of similar object categories present in the ImageNet dataset, which allows such models to be pretrained on a much larger set of samples and learn generic features for those object categories. We observe an important and often ignored additional structure present in an FGVC problem: the objects are captured from a small set of viewing angles only. We notice that subtle differences between object categories are difficult to pick from an arbitrary angle but easier to identify from a similar pose. We show in this paper that training specialized pose experts, focusing on classification from a single, fixed pose, and combining them in an ensemble style framework successfully exploits the structure in the problem. We demonstrate the effectiveness of the proposed approach on the benchmark Stanford Cars, FGVC-Aircrafts, and DeepFashion datasets. To highlight the contribution when the target category features may not be available in a pretrained network, we test on footwear class. We contribute a new 1000 object, 12 category footwear dataset, each object captured from 4 different poses and show significant improvement on this dataset.
Kushagra Mahajan, Tarasha Khurana, Ayush Chopra, Isha Gupta, Chetan Arora 0001, Atul Rai
ICIP5
2018 Making Deep Neural Network Fooling Practical
abstract
With the success of deep neural networks (DNNs), the robustness of such models under adversarial or fooling attacks has become extremely important. It has been shown that a simple perturbation of the image, invisible to a human observer, is sufficient to fool a deep network. Building on top of such work, methods have been proposed to generate adversarial samples which are robust to natural perturbations (camera noise, rotation, shift, scaling etc.). In this paper, we review multiple such fooling algorithms and show that the generated adversarial samples exhibit distributions largely different from the true distribution of the training samples, and thus are easily detectable by a simple meta classifier. We argue that for truly practical DNN fooling, not only should the adversarial samples be robust against various distortions, but must also follow the training set distribution and be undetectable from such meta classifiers. Finally we propose a new adversarial sample generation technique that outperforms commonly known methods when evaluated simultaneously on robustness and detectability.
Ambar Pal, Chetan Arora 0001
ICIP2
2018 Diversity in Fashion Recommendation Using Semantic Parsing
abstract
Developing recommendation system for fashion images is challenging due to the inherent ambiguity associated with what criterion a user is looking at. Suggesting multiple images where each output image is similar to the query image on the basis of a different feature or part is one way to mitigate the problem. Existing works for fashion recommendation have used Siamese or Triplet network to learn features between a similar pair and a similar-dissimilar triplet respectively. However, these methods do not provide basic information such as, how two clothing images are similar, or which parts present in the two images make them similar. In this paper, we propose to recommend images by explicitly learning and exploiting part based similarity. We propose a novel approach of learning discriminative features from weakly-supervised data by using visual attention over the parts and a texture encoding network. We show that the learned features surpass the state-of-the-art in retrieval task on DeepFashion dataset. We then use the proposed model to recommend fashion images having an explicit variation with respect to similarity of any of the parts.
Sagar Verma, Sukhad Anand, Chetan Arora 0001, Atul Rai
ICIP3
2018 Making Third Person Techniques Recognize First-Person Actions in Egocentric Videos
abstract
We focus on first-person action recognition from egocentric videos. Unlike third person domain, researchers have divided first-person actions into two categories: involving hand-object interactions and the ones without, and developed separate techniques for the two action categories. Further, it has been argued that traditional cues used for third person action recognition do not suffice, and egocentric specific features, such as head motion and handled objects have been used for such actions. Unlike the state-of-the-art approaches, we show that a regular two stream Convolutional Neural Network (CNN) with Long Short-Term Memory (LSTM) architecture, having separate streams for objects and motion, can generalize to all categories of first-person actions. The proposed approach unifies the feature learned by all action categories, making the proposed architecture much more practical. In an important observation, we note that the size of the objects visible in the egocentric videos is much smaller. We show that the performance of the proposed model improves after cropping and resizing frames to make the size of objects comparable to the size of ImageNet's objects. Our experiments on the standard datasets: GTEA, EGTEA Gaze+, HUJI, ADL, UTE, and Kitchen, proves that our model significantly outperforms various state-of-the-art techniques.
Sagar Verma, Pravin Nagar, Divam Gupta, Chetan Arora 0001
ICIP4
2018 Stabilizing First Person 360 Degree Videos
abstract
The use of 360 degree cameras, enabling one to record and share full-spherical 360° X 180° view without any cropping in the viewing angle, is on the rise. Shake in such videos is problematic, especially when used in conjunction with VR headsets causing cybersickness to the viewer. The current state-of-the-art video stabilization algorithm [17] designed specifically for 360 degree videos considers the special geometrical constraints in such videos. However, the specific steps in the algorithm can abruptly change the viewing direction in a video leading to unnatural experience for the viewer. In this paper, we propose to fix this anomaly by the use of L1 smoothness constraints on the camera path, as suggested by Grundmann et al. [7]. The modified algorithm is generic and our experiments indicate that the proposed algorithm not only gives a more natural and smoother stabilization for 360 degree videos but can be used for stabilizing normal field of view videos as well.
Chetan Arora 0001, Vivek Kwatra
WACV1
2018 Learning Higher Order Potentials for MRFs
abstract
Higher order MRF-MAP formulation has been shown to improve solutions in many popular computer vision problems. Most of these approaches have considered hand tuned clique potentials only. Over the last few years, while there has been steady improvement in inference techniques making it possible to perform tractable inference for clique sizes even up to few hundreds, the learning techniques for such clique potentials have been limited to clique size of merely 3 or 4. In this paper, we investigate learning of higher order clique potentials up to clique size of 16. We use structural support vector machine (SSVM), a large-margin learning framework, to learn higher order potential functions from data. It formulates the training problem as a quadratic programming problem (QP) that requires solving MAP inference problems in the inner iteration. We introduce multiple innovations in the formulation by introducing soft submodularity constraints which keep QP constraints manageable and at the same time makes MAP inference tractable. Unlike contemporary approaches to solving the original problem using the cutting plane technique, we propose to solve the problem using subgradient descent. This allows us to scale for problems with clique size even up to 16. We give indicative experiments to show the improvement gained in real applications using learned potentials instead of hand tuned ones.
Dinesh Khandelwal, Parag Singla, Chetan Arora 0001
WACV3
2018 A Joint 3D-2D Based Method for Free Space Detection on Roads
abstract
In this paper, we address the problem of road segmentation and free space detection in the context of autonomous driving. Traditional methods either use 3-dimensional (3D) cues such as point clouds obtained from LIDAR, RADAR or stereo cameras or 2-dimensional (2D) cues such as lane markings, road boundaries and object detection. Typical 3D point clouds do not have enough resolution to detect fine differences in heights such as between road and pavement. Image based 2D cues fail when encountering uneven road textures such as due to shadows, potholes, lane markings or road restoration. We propose a novel free road space detection technique combining both 2D and 3D cues. In particular, we use CNN based road segmentation from 2D images and plane/box fitting on sparse depth data obtained from SLAM as priors to formulate an energy minimization using conditional random field (CRF), for road pixels classification. While the CNN learns the road texture and is unaffected by depth boundaries, the 3D information helps in overcoming texture based classification failures. Finally, we use the obtained road segmentation with the 3D depth data from monocular SLAM to detect the free space for the navigation purposes. Our experiments on KITTI odometry dataset [12], Camvid dataset [7] as well as videos captured by us validate the superiority of the proposed approach over the state of the art.
Suvam Patra, Pranjal Maheshwari, Shashank Yadav, Subhashis Banerjee, Chetan Arora 0001
WACV5
2018 Object Detection in Real-Time Systems: Going Beyond Precision
abstract
Applications like autonomous driving, industrial robotics, surveillance, and wearable assistive technology rely on object detectors as an integral part of the system. Thus, an increase in performance of object detectors directly affects the quality of such systems. In the recent years, convolutional neural networks (CNNs) and its variants emerged as the state of art in object detection, where performance is usually measured either in terms of mean average precision (mAP) or number of frames processed per second (fps). Many applications which use object detectors are resource constrained in practice. Even though it is clear from the published results, that a frame-level analysis of the system in terms of mAP or fps proves the superiority of one algorithm over the other, we observe that such metrics do not necessarily apply to real time applications with resource constraints. A slower algorithm even though highly accurate may need to drop frames to maintain the necessary frame rate and lose on the accuracy. We propose a closer look at the metrics used for performance in real-time applications, and suggest some new evaluation criterion. Our comparison of state of the art detectors on these metrics has also thrown some surprises in terms of conventional wisdom, which we present in this paper. Our framework is available at https://www.github.com/anupamsobti/object-detectionreal-time-systems.
Anupam Sobti, Chetan Arora 0001, M. Balakrishnan
WACV2
2018 EgoSampling: Wide View Hyperlapse From Egocentric Videos
abstract
The possibility of sharing one's point of view makes the use of wearable cameras compelling. These videos are often long, boring, and coupled with extreme shaking, as the camera is worn on a moving person. Fast-forwarding (i.e., frame sampling) is a natural choice for quick video browsing. However, this accentuates the shake caused by natural head motion in an egocentric video, making the fast-forwarded video useless. We propose EgoSampling, an adaptive frame sampling that gives stable, fast-forwarded, hyperlapse videos. Adaptive frame sampling is formulated as an energy minimization problem, whose optimal solution can be found in polynomial time. We further turn the camera shake from a drawback into a feature, enabling the increase in field of view of the output video. This is obtained when each output frame is mosaiced from several input frames. The proposed technique also enables the generation of a single hyperlapse video from multiple egocentric videos, allowing even faster video consumption.
Tavi Halperin, Yair Poleg, Chetan Arora 0001, Shmuel Peleg
IEEE Trans. Circuits Syst. Video Technol.3
2017 Deep CNN with color lines model for unmarked road segmentation
abstract
Road detection from a monocular camera is an important perception module in any advanced driver assistance or autonomous driving system. Traditional techniques [1, 2, 3, 4, 5, 6] work reasonably well for this problem, when the roads are well maintained and the boundaries are clearly marked. However, in many developing countries or even for the rural areas in the developed countries, the assumption does not hold which leads to failure of such techniques. In this paper we propose a novel technique based on the combination of deep convolutional neural networks (CNNs), along with color lines model [7] based prior in a conditional random field (CRF) framework. While the CNN learns the road texture, the color lines model allows to adapt to varying illumination conditions. We show that our technique outperforms the state of the art segmentation techniques on the unmarked road segmentation problem. Though, not a focus of this paper, we show that even on the standard benchmark datasets like KITTI [8] and CamVid [9], where the road boundaries are well marked, the proposed technique performs competitively to the contemporary techniques.
Shashank Yadav, Suvam Patra, Chetan Arora 0001, Subhashis Banerjee
ICIP3
2017 Unsupervised Learning of Deep Feature Representation for Clustering Egocentric Actions
abstract
Popularity of wearable cameras in life logging, law enforcement, assistive vision and other similar applications is leading to explosion in generation of egocentric video content. First person action recognition is an important aspect of automatic analysis of such videos. Annotating such videos is hard, not only because of obvious scalability constraints, but also because of privacy issues often associated with egocentric videos. This motivates the use of unsupervised methods for egocentric video analysis. In this work, we propose a robust and generic unsupervised approach for first person action clustering. Unlike the contemporary approaches, our technique is neither limited to any particular class of actions nor requires priors such as pre-training, fine-tuning, etc. We learn time sequenced visual and flow features from an array of weak feature extractors based on convolutional and LSTM autoencoder networks. We demonstrate that clustering of such features leads to the discovery of semantically meaningful actions present in the video. We validate our approach on four disparate public egocentric actions datasets amounting to approximately 50 hours of videos. We show that our approach surpasses the supervised state of the art accuracies without using the action labels.
Bharat Lal Bhatnagar, Suriya Singh, Chetan Arora 0001, C. V. Jawahar
IJCAI3
2017 Computing Egomotion with Local Loop Closures for Egocentric Videos
abstract
Finding the camera pose is an important step in many egocentric video applications. It has been widely reported that, state of the art SIAM algorithms fail on egocentric videos [1, 2, 3, 4]. In this paper, we propose a robust method for camera pose estimation, designed specifically for egocentric videos. In an egocentric video, the camera views the same scene point multiple times as the wearer's head sweeps back and forth. We use this specific motion profile to perform short loop closures aligned with wearer's footsteps. For egocentric videos, depth estimation is usually noisy. In an important departure, we use 2D computations for rotation averaging which do not rely upon depth estimates. The two modification results in much more stable algorithm as is evident from our experiments on various egocentric video datasets for different egocentric applications. The proposed algorithm resolves a long standing problem in egocentric vision and unlocks new usage scenarios for future applications.
Suvam Patra, Himanshu Aggarwal, Himani Arora, Subhashis Banerjee, Chetan Arora 0001
WACV5
2017 Trajectory aligned features for first person action recognition
Suriya Singh, Chetan Arora 0001, C. V. Jawahar
Pattern Recognit.2
2016 Min Norm Point Algorithm for Higher Order MRF-MAP Inference
abstract
Many tasks in computer vision and machine learning can be modelled as the inference problems in an MRF-MAP formulation and can be reduced to minimizing a submodular function. Using higher order clique potentials to model complex dependencies between pixels improves the performance but the current state of the art inference algorithms fail to scale for larger clique sizes. We adapt a well known Min Norm Point algorithm from mathematical optimization literature to exploit the sum of submodular structure found in the MRF-MAP formulation. Unlike some contemporary methods, we do not make any assumptions (other than submodularity) on the type of the clique potentials. Current state of the art inference algorithms for general submodular function takes many hours for problems with clique size 16, and fail to scale beyond. On the other hand, our algorithm is highly efficient and can perform optimal inference in few seconds even on clique size an order of magnitude larger. The proposed algorithm can even scale to clique sizes of many hundreds, unlocking the usage of really large size cliques for MRF-MAP inference problems in computer vision. We demonstrate the efficacy of our approach by experimenting on synthetic as well as real datasets.
Ishant Shanu, Chetan Arora 0001, Parag Singla
CVPR2
2016 First Person Action Recognition Using Deep Learned Descriptors
abstract
We focus on the problem of wearer's action recognition in first person a.k.a. egocentric videos. This problem is more challenging than third person activity recognition due to unavailability of wearer's pose and sharp movements in the videos caused by the natural head motion of the wearer. Carefully crafted features based on hands and objects cues for the problem have been shown to be successful for limited targeted datasets. We propose convolutional neural networks (CNNs) for end to end learning and classification of wearer's actions. The proposed network makes use of egocentric cues by capturing hand pose, head motion and saliency map. It is compact. It can also be trained from relatively small number of labeled egocentric videos that are available. We show that the proposed network can generalize and give state of the art performance on various disparate egocentric action datasets.
Suriya Singh, Chetan Arora 0001, C. V. Jawahar
CVPR2
2016 Compact CNN for indexing egocentric videos
abstract
While egocentric video is becoming increasingly popular, browsing it is very difficult. In this paper we present a compact 3D Convolutional Neural Network (CNN) architecture for long-term activity recognition in egocentric videos. Recognizing long-term activities enables us to temporally segment (index) long and unstructured egocentric videos. Existing methods for this task are based on hand tuned features derived from visible objects, location of hands, as well as optical flow. Given a sparse optical flow volume as input, our CNN classifies the camera wearer's activity. We obtain classification accuracy of 89%, which outperforms the current state-of-the-art by 19%. Additional evaluation is performed on an extended egocentric video dataset, classifying twice the amount of categories than current state-of-the-art. Furthermore, our CNN is able to recognize whether a video is egocentric or not with 99.2% accuracy, up by 24% from current state-of-the-art. To better understand what the network actually learns, we propose a novel visualization of CNN kernels as flow fields.
Yair Poleg, Ariel Ephrat, Shmuel Peleg, Chetan Arora 0001
WACV4
2016 Lazy Generic Cuts
Dinesh Khandelwal, Kush Bhatia, Chetan Arora 0001, Parag Singla
Comput. Vis. Image Underst.3
2015 EgoSampling: Fast-forward and stereo for egocentric videos
abstract
While egocentric cameras like GoPro are gaining popularity, the videos they capture are long, boring, and difficult to watch from start to end. Fast forwarding (i.e. frame sampling) is a natural choice for faster video browsing. However, this accentuates the shake caused by natural head motion, making the fast forwarded video useless. We propose EgoSampling, an adaptive frame sampling that gives more stable fast forwarded videos. Adaptive frame sampling is formulated as energy minimization, whose optimal solution can be found in polynomial time. In addition, egocentric video taken while walking suffers from the left-right movement of the head as the body weight shifts from one leg to another. We turn this drawback into a feature: Stereo video can be created by sampling the frames from the left most and right most head positions of each step, forming approximate stereo-pairs.
Yair Poleg, Tavi Halperin, Chetan Arora 0001, Shmuel Peleg
CVPR3
2015 Generalized Flows for Optimal Inference in Higher Order MRF-MAP
abstract
Use of higher order clique potentials in MRF-MAP problems has been limited primarily because of the inefficiencies of the existing algorithmic schemes. We propose a new combinatorial algorithm for computing optimal solutions to 2 label MRF-MAP problems with higher order clique potentials. The algorithm runs in time O(2(k)n(3)) in the worst case (k is size of clique and n is the number of pixels). A special gadget is introduced to model flows in a higher order clique and a technique for building a flow graph is specified. Based on the primal dual structure of the optimization problem, the notions of the capacity of an edge and a cut are generalized to define a flow problem. We show that in this flow graph, when the clique potentials are submodular, the max flow is equal to the min cut, which also is the optimal solution to the problem. We show experimentally that our algorithm provides significantly better solutions in practice and is hundreds of times faster than solution schemes like Dual Decomposition [1], TRWS [2] and Reduction [3], [4], [5]. The framework represents a significant advance in handling higher order problems making optimal inference practical for medium sized cliques.
Chetan Arora 0001, Subhashis Banerjee, Prem Kumar Kalra, S. N. Maheshwari
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Head Motion Signatures from Egocentric Videos
Yair Poleg, Chetan Arora 0001, Shmuel Peleg
ACCV (3)2
2014 Fast Approximate Inference in Higher Order MRF-MAP Labeling Problems
abstract
Use of higher order clique potentials for modeling inference problems has exploded in last few years. The algorithmic schemes proposed so far do not scale well with increasing clique size, thus limiting their use to cliques of size at most 4 in practice. Generic Cuts (GC) of Arora et al. [9] shows that when potentials are submodular, inference problems can be solved optimally in polynomial time for fixed size cliques. In this paper we report an algorithm called Approximate Cuts (AC) which uses a generalization of the gadget of GC and provides an approximate solution to inference in 2-label MRF-MAP problems with cliques of size k ≥ 2. The algorithm gives optimal solution for submodular potentials. When potentials are non-submodular, we show that important properties such as weak persistency hold for solution inferred by AC. AC is a polynomial time primal dual approximation algorithm for fixed clique size. We show experimentally that AC not only provides significantly better solutions in practice, it is an order of magnitude faster than message passing schemes like Dual Decomposition [19] and GTRWS [17] or Reduction based techniques like [10, 13, 14].
Chetan Arora 0001, Subhashis Banerjee, Prem Kumar Kalra, S. N. Maheshwari
CVPR1
2014 Multi-label Generic Cuts: Optimal Inference in Multi-label Multi-clique MRF-MAP Problems
abstract
We propose an algorithm called Multi Label Generic Cuts (MLGC) for computing optimal solutions to MRF-MAP problems with submodular multi label multi-clique potentials. A transformation is introduced to convert a m-label k-clique problem to an equivalent 2-label (mk)-clique problem. We show that if the original multi-label problem is submodular then the transformed 2-label multi-clique problem is also submodular. We exploit sparseness in the feasible configurations of the transformed 2-label problem to suggest an improvement to Generic Cuts [3] to solve the 2-label problems efficiently. The algorithm runs in time O(mk n3) in the worst case (n is the number of pixels) generalizing O(2k n3) running time of Generic Cuts. We show experimentally that MLGC is an order of magnitude faster than the current state of the art [17, 20]. While the result of MLGC is optimal for submodular clique potential it is significantly better than the compared methods even for problems with non-submodular clique potential.
Chetan Arora 0001, S. N. Maheshwari
CVPR1
2014 Temporal Segmentation of Egocentric Videos
abstract
The use of wearable cameras makes it possible to record life logging egocentric videos. Browsing such long unstructured videos is time consuming and tedious. Segmentation into meaningful chapters is an important first step towards adding structure to egocentric videos, enabling efficient browsing, indexing and summarization of the long videos. Two sources of information for video segmentation are (i) the motion of the camera wearer, and (ii) the objects and activities recorded in the video. In this paper we address the motion cues for video segmentation. Motion based segmentation is especially difficult in egocentric videos when the camera is constantly moving due to natural head movement of the wearer. We propose a robust temporal segmentation of egocentric videos into a hierarchy of motion classes using a new Cumulative Displacement Curves. Unlike instantaneous motion vectors, segmentation using integrated motion vectors performs well even in dynamic and crowded scenes. No assumptions are made on the underlying scene structure and the method works in indoor as well as outdoor situations. We demonstrate the effectiveness of our approach using publicly available videos as well as choreographed videos. We also suggest an approach to detect the fixation of wearer's gaze in the walking portion of the egocentric videos.
Yair Poleg, Chetan Arora 0001, Shmuel Peleg
CVPR2
2014 Optical flow for non Lambertian surfaces by cancelling illuminant chromaticity
abstract
Optical flow, the pixel level correspondences between a pair of images is an important problem in computer vision. Standard optical flow computation algorithms assume constant brightness and fail on specular surfaces. Earlier work to alleviate problems with specularity evaluate the illuminant chromaticity using a few correspondences in the images and then jointly optimize flow and appearance under the dichromatic model. We argue that the correspondences obtained by these methods are mostly pairs of pixels that are Lambertian thus giving a noisy estimate of the illuminant chromaticity. We suggest a new approach to evaluate the illuminant chromaticity which does not require exact correspondences and gives a better estimate of illuminant chromaticity. We use the evaluated chromaticity to project the input images on to a specular invariant color space and show that standard optical flow algorithms on this color space significantly improves the flow results. The suggested approach is simple, efficient and more importantly can utilize existing algorithms to compute optical flow on non Lambertian surfaces.
Chetan Arora 0001, Michael Werman
ICIP1
2013 Efficient representation of distributions for background subtraction
abstract
Multi dimensional probability distributions are used in many surveillance tasks such as modeling color distribution of background pixels for Background Subtraction. Accurate representation of such distributions, e.g. in a histogram, requires much memory that may not be available when a histogram is computed for each pixel. Parametric representations such as Gaussian Mixture Models (GMM) are very efficient in memory but may not be accurate enough when the distribution is not from the assumed model. We propose a memory efficient representation for distributions. Histograms cells usually have equal width, and count the hits in each cell (Equi-width histograms). In most cases a 1D distribution can be represented more efficiently when cell sizes change so that each cell will have same number of hits (Equi-depth histograms). We propose to describe compactly multi-dimensional distributions (e.g. color) using an equi-depth histograms. Online computation of such histograms is described, and examples are given for background subtraction.
Yedid Hoshen, Chetan Arora 0001, Yair Poleg, Shmuel Peleg
AVSS2
2013 Higher Order Matching for Consistent Multiple Target Tracking
abstract
This paper addresses the data assignment problem in multi frame multi object tracking in video sequences. Traditional methods employing maximum weight bipartite matching offer limited temporal modeling. It has recently been shown [6, 8, 24] that incorporating higher order temporal constraints improves the assignment solution. Finding maximum weight matching with higher order constraints is however NP-hard and the solutions proposed until now have either been greedy [8] or rely on greedy rounding of the solution obtained from spectral techniques [15]. We propose a novel algorithm to find the approximate solution to data assignment problem with higher order temporal constraints using the method of dual decomposition and the MPLP message passing algorithm [21]. We compare the proposed algorithm with an implementation of [8] and [15] and show that proposed technique provides better solution with a bound on approximation factor for each inferred solution.
Chetan Arora 0001, Amir Globerson
ICCV1
2012 Generic Cuts: An Efficient Algorithm for Optimal Inference in Higher Order MRF-MAP
Chetan Arora 0001, Subhashis Banerjee, Prem Kumar Kalra, S. N. Maheshwari
ECCV (5)1
2010 An Efficient Graph Cut Algorithm for Computer Vision Problems
Chetan Arora 0001, Subhashis Banerjee, Prem Kumar Kalra, S. N. Maheshwari
ECCV (3)1
2000 Rectified Mosaicing: Mosaics without the Curl
abstract
When images captured by a tilted camera are mosaiced into a panorama, the resulting mosaic is curled. This happens, for example, with a panning camera that is not perfectly horizontal, and with a translating camera facing a tilted planar surface. The tilt of the camera causes differences in image velocity between the top and bottom parts of the image, causing the curled mosaic. In rectified mosaicing these distortions are overcome by warping the strips into rectangles, while keeping some image feature invariant. This warping equalizes the image motion at the different image parts, and the resulting mosaic is straight. Mosaicing is done without camera calibration or knowledge of the scene, and the process adapts automatically to smooth changes in the scene and the imaging conditions.
Shmuel Peleg, Assaf Zomet, Chetan Arora 0001
CVPR3