Anirban Chakraborty 0001

dblp:73/2286-1 · DBLP profile ↗
← Back
52ranked-venue papers
7as first author
27since 2021 · last 2026
0000-0002-6946-9152ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 4 first-author · 20 since 2021Artificial intelligence and machine learning · 25 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
abstract
While Large Vision Language Models (LVLMs) are increasingly deployed in real-world applications, their ability to interpret abstract visual inputs remains limited. Specifically, they struggle to comprehend hand-drawn sketches, a modality that offers an intuitive means of expressing concepts that are difficult to describe textually. We identify the primary bottleneck as the absence of a large-scale dataset that jointly models sketches, photorealistic images, and corresponding natural language instructions. To address this, we present two key contributions: (1) a new, large-scale dataset of image-sketch-instruction triplets designed to facilitate both pretraining and instruction tuning, and (2) O3SLM, an LVLM trained on this dataset. Comprehensive evaluations on multiple sketch-based tasks: (a) object localization, (b) counting, (c) image retrieval i.e., (SBIR and fine-grained SBIR), and (d) visual question answering (VQA); while incorporating the three existing sketch datasets, namely QuickDraw!, Sketchy, and Tu-Berlin, along with our generated SketchVCL dataset, show that O3SLM achieves state-of-the-art performance, substantially outperforming existing LVLMs in sketch comprehension and reasoning.
Rishi Gupta, Mukilan Karuppasamy, Shyam Marjit, Aditay Tripathi, Anirban Chakraborty 0001
AAAI5
2026 FedSCAl: Leveraging Server and Client Alignment for Unsupervised Federated Source-Free Domain Adaptation
abstract
We address the Federated source-Free Domain Adaptation (FFreeDA) problem, with clients holding unlabeled data with significant inter-client domain gaps. The FFreeDA setup constrains the FL frameworks to employ only a pre-trained server model as the setup restricts access to the source dataset during the training rounds. Often, this source domain dataset has a distinct distribution to the clients’ domains. To address the challenges posed by the FFreeDA setup, adaptation of the Source-Free Domain Adaptation (SFDA) methods to FL struggles with client-drift in real-world scenarios due to extreme data heterogeneity caused by the aforementioned domain gaps, resulting in unreliable pseudo-labels. In this paper, we introduce FedSCAl, an FL framework leveraging our proposed Server-Client Alignment (SCAl) mechanism to regularize client updates by aligning the clients’ and server model’s predictions. We observe an improvement in the clients’ pseudo-labeling accuracy post alignment, as the SCAl mechanism helps to mitigate the client-drift. Further, we present extensive experiments on benchmark vision datasets showcasing how FedSCAl consistently outperforms state-of-the-art FL methods in the FFreeDA setup for classification tasks.
M. Yashwanth, Sampath Koti, Arunabh Singh, Shyam Marjit, Anirban Chakraborty 0001
WACV5
2025 Optimizing Federated Learning for Scalable Power-demand Forecasting in Microgrids
abstract
Real-time monitoring of power consumption in cities and micro-grids through the Internet of Things (IoT) can help forecast future demand and optimize grid operations. But moving all consumer-level usage data to the cloud for predictions and analysis at fine time scales can expose activity patterns. Federated Learning (FL) is a privacy-sensitive collaborative DNN training approach that retains data on edge devices, trains the models on private data locally, and aggregates the local models in the cloud. But key challenges exist: (i) clients can have non-independently identically distributed (non-IID) data, and (ii) the learning should be computationally cheap while scaling to 1000s of (unseen) clients. In this paper, we develop and evaluate several optimizations to FL training across edge and cloud for time-series demand forecasting in micro-grids and city-scale utilities using DNNs to achieve a high prediction accuracy while minimizing the training cost. We showcase the benefit of using exponentially weighted loss while training and show that it further improves the prediction of the final model. Finally, we evaluate these strategies by validating over 1000s of clients for three states in the US from the OpenEIA corpus, and performing FL both in a pseudo-distributed setting and a Pi edge cluster. The results highlight the benefits of the proposed methods over baselines like ARIMA and DNNs trained for individual consumers, which are not scalable.
Roopkatha Banerjee, Sampath Koti, Gyanendra Singh, Anirban Chakraborty 0001, Gurunath Gurrala, Bhushan Jagyasi, Yogesh L. Simmhan
eScience4
2025 DAD++: Improved data-free test time adversarial defense
Gaurav Kumar Nayak, Inder Khatri, Shubham Randive, Ruchit Rawal, Anirban Chakraborty 0001
Neurocomputing5
2025 CARL: Cost-Optimized Online Container Placement on VMs Using Adversarial Reinforcement Learning
abstract
Containerization has become popular for the deployment of applications on public clouds. Large enterprises may host 100 s of applications on 1000 s containers that are placed onto Virtual Machines (VMs). Such placement decisions happen continuously as applications are updated by DevOps pipelines that deploy the containers. Managing the placement of container resource requests onto the available capacities of VMs needs to be cost-efficient. This is well-studied, and usually modelled as a multi-dimensional Vector Bin-packing Problem (VBP). Many heuristics, and recently machine learning approaches, have been developed to solve this NP-hard problem for real-time decisions. We propose CARL, a novel approach to solve VBP through Adversarial Reinforcement Learning (RL) for cost minimization. It mimics the placement behavior of an offline semi-optimal VBP solver (teacher), while automatically learning a reward function for reducing the VM costs which out-performs the teacher. It requires limited historical container workload traces to train, and is resilient to changes in the workload distribution during inferencing. We extensively evaluate CARL on workloads derived from realistic traces from Google and Alibaba for the placement of 5 k–10 k container requests onto 2 k–8 k VMs, and compare it with classic heuristics and state-of-the-art RL methods. (1) CARL isfast, e.g., making placement decisions at$\approx 1900$requests/sec onto 8,900 candidate VMs. (2) It isefficient, achieving$\approx 16\%$lower VM costs than classic and contemporary RL methods. (3) It isrobustto changes in the workload, offering competitive results even when the resource needs or inter-arrival time of the container requests skew from the training workload.
Prathamesh Saraf Vinayak, Saswat Subhajyoti Mallick, Lakshmi Jagarlamudi, Anirban Chakraborty 0001, Yogesh L. Simmhan
IEEE Trans. Cloud Comput.4
2025 Robust Few-Shot Learning Without Using Any Adversarial Samples
abstract
The high cost of acquiring and annotating samples has made the "few-shot" learning problem of prime importance. Existing works mainly focus on improving performance on clean data and overlook robustness concerns on the data perturbed with adversarial noise. Recently, a few efforts have been made to combine the few-shot problem with the robustness objective using sophisticated meta-learning techniques. These methods rely on the generation of adversarial samples in every episode of training, which further adds to the computational burden. To avoid such time-consuming and complicated procedures, we propose a simple but effective alternative that does not require any adversarial samples. Inspired by the cognitive decision-making process in humans, we enforce high-level feature matching between the base class data and their corresponding low-frequency samples in the pretraining stage via self distillation. The model is then fine-tuned on the samples of novel classes where we additionally improve the discriminability of low-frequency query set features via cosine similarity. On a one-shot setting of the CIFAR-FS dataset, our method yields a massive improvement of 60.55% and 62.05% in adversarial accuracy on the projected gradient descent (PGD) and state-of-the-art auto attack, respectively, with a minor drop in clean accuracy compared to the baseline. Moreover, our method only takes of the standard training time while being faster than thestate-of-the-art adversarial meta-learning methods. The code is available at https://github.com/vcl-iisc/robust-few-shot-learning.
Gaurav Kumar Nayak, Ruchit Rawal, Inder Khatri, Anirban Chakraborty 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Dynamic Data Selection for Efficient SSL via Coarse-to-Fine Refinement
Aditay Tripathi, Pradeep Shenoy, Anirban Chakraborty 0001
ECCV (58)3
2024 Enhancing 3D Referential Grounding by Learning Coarse Spatial Relationships
Soham Joshi, Aditay Tripathi, Viswanath Gopalakrishnan, Anirban Chakraborty 0001
ICPR (30)4
2024 Query-guided Attention in Vision Transformers for Localizing Objects Using a Single Sketch
abstract
In this study, we explore sketch-based object localization on natural images. Given a crude hand-drawn object sketch, the task is to locate all instances of that object in the target image. This problem proves difficult due to the abstract nature of hand-drawn sketches, variations in the style and quality of sketches, and the large domain gap between the sketches and the natural images. Existing solutions address this using attention-based frameworks to merge query information into image features. Yet, these methods often integrate query features after independently learning image features, causing inadequate alignment and as a result incorrect localization. In contrast, we propose a novel sketch-guided vision transformer encoder that uses cross-attention after each block of the transformer-based image encoder to learn query-conditioned image features, leading to stronger alignment with the query sketch. Further, at the decoder’s output, object and sketch features are refined better to align the representation of objects with the sketch query, thereby improving localization. The proposed model also generalizes to the object categories not seen during training, as the target image features learned by the proposed model are query-aware. Our framework can utilize multiple sketch queries via a trainable novel sketch fusion strategy. The model is evaluated on the images from the public benchmark, MS-COCO, using the sketch queries from QuickDraw! and Sketchy datasets. Compared with existing localization methods, the proposed approach gives a 6.6% and 8.0% improvement in mAP for seen objects using sketch queries from QuickDraw! and Sketchy datasets, respectively, and a 12.2% improvement in AP@50 for large objects that are ‘unseen’ during training. The code is available at https://vcl-iisc.github.io/locformer/.
Aditay Tripathi, Anand Mishra 0001, Anirban Chakraborty 0001
WACV3
2024 Minimizing Layerwise Activation Norm Improves Generalization in Federated Learning
abstract
Federated Learning (FL) is an emerging machine learning framework that enables multiple clients (coordinated by a server) to collaboratively train a global model by aggregating the locally trained models without sharing any client’s training data. It has been observed in recent works that learning in a federated manner may lead the aggregated global model to converge to a ‘sharp minimum’ thereby adversely affecting the generalizability of this FL-trained model. Therefore, in this work, we aim to improve the generalization performance of models trained in a federated setup by introducing a ‘flatness’ constrained FL optimization problem. This flatness constraint is imposed on the top eigenvalue of the Hessian computed from the training loss. As each client trains a model on its local data, we further re-formulate this complex problem utilizing the client loss functions and propose a new computationally efficient regularization technique1, dubbed ‘MAN,’ which Minimizes Activation’s Norm of each layer on client-side models. We also theoretically show that minimizing the activation norm reduces the top eigenvalue of the layer-wise Hessian of the client’s loss, which in turn decreases the overall Hessian’s top eigenvalue, ensuring convergence to a flat minimum. We apply our proposed flatness-constrained optimization to the existing FL techniques and obtain significant improvements, thereby establishing new state-of-the-art.
M. Yashwanth, Gaurav Kumar Nayak, Harsh Rangwani, Arya Singh, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001
WACV6
2024 Multimodal query-guided object localization
Aditay Tripathi, Rajath R. Dani, Anand Mishra 0001, Anirban Chakraborty 0001
Multim. Tools Appl.4
2023 Edges to Shapes to Concepts: Adversarial Augmentation for Robust Vision
abstract
Recent work has shown that deep vision models tend to be overly dependent on low-level or “texture” features, leading to poor generalization. Various data augmentation strategies have been proposed to overcome this so-called texture bias in DNNs. We propose a simple, lightweight adversarial augmentation technique that explicitly incentivizes the network to learn holistic shapes for accurate prediction in an object classification setting. Our augmentations superpose edgemaps from one image onto another image with shuffled patches, using a randomly determined mixing proportion, with the image label of the edgemap image. To classify these augmented images, the model needs to not only detect and focus on edges but distinguish between relevant and spurious edges. We show that our augmentations significantly improve classification accuracy and robustness measures on a range of datasets and neural architectures. As an example, for ViT-S, We obtain absolute gains on classification accuracy gains up to 6%. We also obtain gains of up to 28% and 8.5% on natural adversarial and out-of-distribution datasets like ImageNet-A (for ViT-B) and ImageNet-R (for ViT-S), respectively. Analysis using a range of probe datasets shows substantially increased shape sensitivity in our trained models, explaining the observed improvement in robustness and classification accuracy.
Aditay Tripathi, Rishubh Singh, Anirban Chakraborty 0001, Pradeep Shenoy
CVPR3
2023 CoNMix for Source-free Single and Multi-target Domain Adaptation
abstract
This work introduces the novel task of Source-free Multi-target Domain Adaptation and proposes adaptation framework comprising of Consistency with Nuclear-Norm Maximization and MixUp knowledge distillation (CoNMix) as a solution to this problem. The main motive of this work is to solve for Single and Multi target Domain Adaptation (SMTDA) for the source-free paradigm, which enforces a constraint where the labeled source data is not available during target adaptation due to various privacy-related restrictions on data sharing. The source-free approach leverages target pseudo labels, which can be noisy, to improve the target adaptation. We introduce consistency between label preserving augmentations and utilize pseudo label refinement methods to reduce noisy pseudo labels. Further, we propose novel MixUp Knowledge Distillation (MKD) for better generalization on multiple target domains using various source-free STDA models. We also show that the Vision Transformer (VT) backbone gives better feature representation with improved domain transferability and class discriminability. Our proposed framework achieves the state-of-the-art (SOTA) results in various paradigms of source-free STDA and MTDA settings on popular domain adaptation datasets like Office-Home, Office-Caltech, and DomainNet. Project Page: https://sites.google.com/view/conmix-vcl
Rohit Lal, Himanshu Patil, Anirban Chakraborty 0001
WACV4
2023 DE-CROP: Data-efficient Certified Robustness for Pretrained Classifiers
abstract
Certified defense using randomized smoothing is a popular technique to provide robustness guarantees for deep neural networks against l2adversarial attacks. Existing works use this technique to provably secure a pretrained non-robust model by training a custom denoiser network on entire training data. However, access to the training set may be restricted to a handful of data samples due to constraints such as high transmission cost and the proprietary nature of the data. Thus, we formulate a novel problem of "how to certify the robustness of pretrained models using only a few training samples". We observe that training the custom denoiser directly using the existing techniques on limited samples yields poor certification. To overcome this, our proposed approach (DE-CROP)1generates class-boundary and interpolated samples corresponding to each training sample, ensuring high diversity in the feature space of the pretrained classifier. We train the denoiser by maximizing the similarity between the denoised output of the generated sample and the original training sample in the classifier’s logit space. We also perform distribution level matching using domain discriminator and maximum mean discrepancy that yields further benefit. In white box setup, we obtain significant improvements over the baseline on multiple benchmark datasets and also report similar performance under the challenging black box setup.
Gaurav Kumar Nayak, Ruchit Rawal, Anirban Chakraborty 0001
WACV3
2023 Grounding Scene Graphs on Natural Images via Visio-Lingual Message Passing
abstract
This paper presents a framework for jointly grounding objects that follow certain semantic relationship constraints given in a scene graph. A typical natural scene contains several objects, often exhibiting visual relationships of varied complexities between them. These inter-object relationships provide strong contextual cues towards improving grounding performance compared to a traditional object query-only-based localization task. A scene graph is an efficient and structured way to represent all the objects and their semantic relationships in the image. In an attempt towards bridging these two modalities representing scenes and utilizing contextual information for improving object localization, we rigorously study the problem of grounding scene graphs on natural images. To this end, we propose a novel graph neural network-based approach referred to as Visio-Lingual Message Passing Graph Neural Network (VL-MPAG Net). In VL-MPAG Net, we first construct a directed graph with object proposals as nodes and an edge between a pair of nodes representing a plausible relation between them. Then a three-step inter-graph and intra-graph message passing is performed to learn the context- dependent representation of the proposals and query objects. These object representations are used to score the proposals to generate object localization. The proposed method significantly outperforms the baselines on four public datasets.
Aditay Tripathi, Anand Mishra 0001, Anirban Chakraborty 0001
WACV3
2023 Improving Domain Adaptation Through Class Aware Frequency Transformation
Himanshu Patil, Rohit Lal, Anirban Chakraborty 0001
Int. J. Comput. Vis.4
2022 Uncertainty-Aware Adaptation for Self-Supervised 3D Human Pose Estimation
abstract
The advances in monocular 3D human pose estimation are dominated by supervised techniques that require large-scale 2D/3D pose annotations. Such methods often behave erratically in the absence of any provision to discard unfamiliar out-of-distribution data. To this end, we cast the 3D human pose learning as an unsupervised domain adaptation problem. We introduce MRP-Net11Project page: https://sites.google.com/view/mrp-net that constitutes a common deep network backbone with two output heads subscribing to two diverse configurations; a) model-free Joint localization and b) model-based parametric regression. Such a design allows us to derive suitable measures to quantify prediction uncertainty at both pose and Joint level granularity. While supervising only on labeled synthetic samples, the adaptation process aims to minimize the uncertainty for the unlabeled target images while maximizing the same for an extreme out-of-distribution dataset (backgrounds). Alongside synthetic-to-real 3D pose adaptation, the Joint-uncertainties allow expanding the adaptation to work on in-the-wild images even in the presence of occlusion and truncation scenarios. We present a comprehensive evaluation of the proposed approach and demonstrate state-of-the-art performance on benchmark datasets.
Jogendra Kundu, Siddharth Seth, Pradyumna YM, Varun Jampani, Anirban Chakraborty 0001, Venkatesh Babu Radhakrishnan
CVPR5
2022 DAD: Data-free Adversarial Defense at Test Time
abstract
Deep models are highly susceptible to adversarial attacks. Such attacks are carefully crafted imperceptible noises that can fool the network and can cause severe consequences when deployed. To encounter them, the model requires training data for adversarial training or explicit regularization-based techniques. However, privacy has become an important concern, restricting access to only trained models but not the training data (e.g. biometric data). Also, data curation is expensive and companies may have proprietary rights over it. To handle such situations, we propose a completely novel problem of ‘test-time adversarial defense in absence of training data and even their statistics’. We solve it in two stages: a) detection and b) correction of adversarial samples. Our adversarial sample detection framework is initially trained on arbitrary data and is subsequently adapted to the unlabelled test data through unsupervised domain adaptation. We further correct the predictions on detected adversarial samples by transforming them in Fourier domain and obtaining their low frequency component at our proposed suitable radius for model prediction. We demonstrate the efficacy of our proposed technique via extensive experiments against several adversarial attacks and for different model architectures and datasets. For a non-robust Resnet-18 model pretrained on CIFAR-10, our detection method correctly identifies 91.42% adversaries. Also, we significantly improve the adversarial accuracy from 0% to 37.37% with a minimal drop of 0.02% in clean accuracy on state-of-the-art ‘Auto Attack’ without having to retrain the model.
Gaurav Kumar Nayak, Ruchit Rawal, Anirban Chakraborty 0001
WACV3
2022 Mining Data Impressions From Deep Models as Substitute for the Unavailable Training Data
abstract
Pretrained deep models hold their learnt knowledge in the form of model parameters. These parameters act as "memory" for the trained models and help them generalize well on unseen data. However, in absence of training data, the utility of a trained model is merely limited to either inference or better initialization towards a target task. In this paper, we go further and extract synthetic data by leveraging the learnt model parameters. We dub them Data Impressions, which act as proxy to the training data and can be used to realize a variety of tasks. These are useful in scenarios where only the pretrained models are available and the training data is not shared (e.g., due to privacy or sensitivity concerns). We show the applicability of data impressions in solving several computer vision tasks such as unsupervised domain adaptation, continual learning as well as knowledge distillation. We also study the adversarial robustness of lightweight models trained via knowledge distillation using these data impressions. Further, we demonstrate the efficacy of data impressions in generating data-free Universal Adversarial Perturbations (UAPs) with better fooling rates. Extensive experiments performed on benchmark datasets demonstrate competitive performance achieved using data impressions in absence of original training data.
Gaurav Kumar Nayak, Konda Reddy Mopuri, Saksham Jain, Anirban Chakraborty 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 MMD-ReID: A Simple but Effective Solution for Visible-Thermal Person ReID
Chaitra Jambigi, Ruchit Rawal, Anirban Chakraborty 0001
BMVC3
2021 Beyond Classification: Knowledge Distillation using Multi-Object Impressions
Gaurav Kumar Nayak, Monish Keswani, Sharan Seshadri, Anirban Chakraborty 0001
BMVC4
2021 Incremental Learning for Animal Pose Estimation using RBF k-DPP
Gaurav Kumar Nayak, Het Shah, Anirban Chakraborty 0001
BMVC3
2021 Pose-Transformation and Radial Distance Clustering for Unsupervised Person Re-identification
Siddharth Seth, Akash Sonth, Anirban Chakraborty 0001
BMVC3
2021 Real-Time Object Detection and Localization in Compressive Sensed Video
abstract
Typically a 1-2MP CCTV camera generates around 7-12GB of data per day. Frame-by-frame processing of such an enormous amount of data requires hefty computational resources. In recent years, compressive sensing approaches have shown impressive compression results by reducing the sampling bandwidth. Different sampling mechanisms were developed to incorporate compressive sensing in image and video acquisition. Though all-CMOS [1], [2] sensor cameras that perform compressive sensing can help save a lot of bandwidth on sampling and minimize the memory required to store videos, the traditional signal processing, and deep learning models can realize operations only on the reconstructed data. To realize the original uncompressed domain, most reconstruction techniques are computationally expensive and time-consuming. To bridge this gap, we propose a novel task of detection and localization of objects directly on the compressed frames. Thereby mitigating the need to reconstruct the frames and reducing the search rate up to $20 \times$ (compression rate). We achieved an accuracy of 46.27% mAP with the proposed model on a GeForce GTX 1080 Ti. We were also able to show real-time inference on an NVIDIA TX2 embedded board with 45.11% mAP, thereby achieving the best balance between the accuracy, inference time, and memory constraints.
Yeshwanth Bethi, Sathyaprakash Narayanan, Venkat Rangan, Anirban Chakraborty 0001, Chetan Singh Thakur
ICIP4
2021 Non-local Latent Relation Distillation for Self-Adaptive 3D Human Pose Estimation
abstract
Available 3D human pose estimation approaches leverage different forms of strong (2D/3D pose) or weak (multi-view or depth) paired supervision. Barring synthetic or in-studio domains, acquiring such supervision for each new target environment is highly inconvenient. To this end, we cast 3D pose learning as a self-supervised adaptation problem that aims to transfer the task knowledge from a labeled source domain to a completely unpaired target. We propose to infer image-to-pose via two explicit mappings viz. image-to-latent and latent-to-pose where the latter is a pre-learned decoder obtained from a prior-enforcing generative adversarial auto-encoder. Next, we introduce relation distillation as a means to align the unpaired cross-modal samples i.e., the unpaired target videos and unpaired 3D pose sequences. To this end, we propose a new set of non-local relations in order to characterize long-range latent pose interactions, unlike general contrastive relations where positive couplings are limited to a local neighborhood structure. Further, we provide an objective way to quantify non-localness in order to select the most effective relation set. We evaluate different self-adaptation settings and demonstrate state-of-the-art 3D human pose estimation performance on standard benchmarks.
Jogendra Kundu, Siddharth Seth, Anirudh Jamkhandi, Pradyumna YM, Varun Jampani, Anirban Chakraborty 0001, Venkatesh Babu Radhakrishnan
NeurIPS6
2021 Effectiveness of Arbitrary Transfer Sets for Data-free Knowledge Distillation
abstract
Knowledge Distillation is an effective method to transfer the learning across deep neural networks. Typically, the dataset originally used for training the Teacher model is chosen as the "Transfer Set" to conduct the knowledge transfer to the Student. However, this original training data may not always be freely available due to privacy or sensitivity concerns. In such scenarios, existing approaches either iteratively compose a synthetic set representative of the original training dataset, one sample at a time or learn a generative model to compose such a transfer set. However, both these approaches involve complex optimization (GAN training or several backpropagation steps to synthesize one sample) and are often computationally expensive. In this paper, as a simple alternative, we investigate the effectiveness of "arbitrary transfer sets" such as random noise, publicly available synthetic, and natural datasets, all of which are completely unrelated to the original training dataset in terms of their visual or semantic contents. Through extensive experiments on multiple benchmark datasets such as MNIST, FMNIST, CIFAR-10 and CIFAR-100, we discover and validate surprising effectiveness of using arbitrary data to conduct knowledge distillation when this dataset is "target-class balanced". We believe that this important observation can potentially lead to designing base-lines for the data-free knowledge distillation task.
Gaurav Kumar Nayak, Konda Reddy Mopuri, Anirban Chakraborty 0001
WACV3
2021 Robust gait based human identification on incomplete and multi-view sequences
Utkarsh Shreemali, Anirban Chakraborty 0001
Multim. Tools Appl.2
2020 DeGAN: Data-Enriching GAN for Retrieving Representative Samples from a Trained Classifier
abstract
In this era of digital information explosion, an abundance of data from numerous modalities is being generated as well as archived everyday. However, most problems associated with training Deep Neural Networks still revolve around lack of data that is rich enough for a given task. Data is required not only for training an initial model, but also for future learning tasks such as Model Compression and Incremental Learning. A diverse dataset may be used for training an initial model, but it may not be feasible to store it throughout the product life cycle due to data privacy issues or memory constraints. We propose to bridge the gap between the abundance of available data and lack of relevant data, for the future learning tasks of a given trained network. We use the available data, that may be an imbalanced subset of the original training dataset, or a related domain dataset, to retrieve representative samples from a trained classifier, using a novel Data-enriching GAN (DeGAN) framework. We demonstrate that data from a related domain can be leveraged to achieve state-of-the-art performance for the tasks of Data-free Knowledge Distillation and Incremental Learning on benchmark datasets. We further demonstrate that our proposed framework can enrich any data, even from unrelated domains, to make it more useful for the future learning tasks of a given network.
Sravanti Addepalli, Gaurav Kumar Nayak, Anirban Chakraborty 0001, Venkatesh Babu Radhakrishnan
AAAI3
2020 Kinematic-Structure-Preserved Representation for Unsupervised 3D Human Pose Estimation
abstract
Estimation of 3D human pose from monocular image has gained considerable attention, as a key step to several human-centric applications. However, generalizability of human pose estimation models developed using supervision on large-scale in-studio datasets remains questionable, as these models often perform unsatisfactorily on unseen in-the-wild environments. Though weakly-supervised models have been proposed to address this shortcoming, performance of such models relies on availability of paired supervision on some related task, such as 2D pose or multi-view image pairs. In contrast, we propose a novel kinematic-structure-preserved unsupervised 3D pose estimation framework, which is not restrained by any paired or unpaired weak supervisions. Our pose estimation framework relies on a minimal set of prior knowledge that defines the underlying kinematic 3D structure, such as skeletal joint connectivity information with bone-length ratios in a fixed canonical scale. The proposed model employs three consecutive differentiable transformations namely forward-kinematics, camera-projection and spatial-map transformation. This design not only acts as a suitable bottleneck stimulating effective pose disentanglement, but also yields interpretable latent pose representations avoiding training of an explicit latent embedding to pose mapper. Furthermore, devoid of unstable adversarial setup, we re-utilize the decoder to formalize an energy-based loss, which enables us to learn from in-the-wild videos, beyond laboratory settings. Comprehensive experiments demonstrate our state-of-the-art unsupervised and weakly-supervised pose estimation performance on both Human3.6M and MPI-INF-3DHP datasets. Qualitative results on unseen environments further establish our superior generalization ability.
Jogendra Kundu, Siddharth Seth, Rahul M. V., Mugalodi Rakesh, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001
AAAI6
2020 WAMDA: Weighted Alignment of Sources for Multi-source Domain Adaptation
Surbhi Aggarwal, Jogendra Kundu, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001
BMVC4
2020 Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image Synthesis
abstract
Camera captured human pose is an outcome of several sources of variation. Performance of supervised 3D pose estimation approaches comes at the cost of dispensing with variations, such as shape and appearance, that may be useful for solving other related tasks. As a result, the learned model not only inculcates task-bias but also dataset-bias because of its strong reliance on the annotated samples, which also holds true for weakly-supervised models. Acknowledging this, we propose a self-supervised learning framework to disentangle such variations from unlabeled video frames. We leverage the prior knowledge on human skeleton and poses in the form of a single part-based 2D puppet model, human pose articulation constraints, and a set of unpaired 3D poses. Our differentiable formalization, bridging the representation gap between the 3D pose and spatial part maps, not only facilitates discovery of interpretable pose disentanglement, but also allows us to operate on videos with diverse camera movements. Qualitative results on unseen in-the-wild datasets establish our superior generalization across multiple tasks beyond the primary tasks of 3D pose estimation and part segmentation. Furthermore, we demonstrate state-of-the-art weakly-supervised 3D pose estimation performance on both Human3.6M and MPI-INF-3DHP datasets.
Jogendra Kundu, Siddharth Seth, Varun Jampani, Mugalodi Rakesh, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001
CVPR6
2020 Sketch-Guided Object Localization in Natural Images
Aditay Tripathi, Rajath R. Dani, Anand Mishra 0001, Anirban Chakraborty 0001
ECCV (6)4
2020 Text-based Person Search via Attribute-aided Matching
abstract
Text-based person search aims to retrieve the pedestrian images that best match a given text query. Existing methods utilize class-id information to get discriminative and identity-preserving features. However, it is not well-explored whether it is beneficial to explicitly ensure that the semantics of the data are retained. In the proposed work, we aim to create semantics-preserving embeddings through an additional task of attribute prediction. Since attribute annotation is typically unavailable in text-based person search, we first mine them from the text corpus. These attributes are then used as a means to bridge the modality gap between the image-text inputs, as well as to improve the representation learning. In summary, we propose an approach for text-based person search by learning an attribute-driven space along with a class-information driven space, and utilize both for obtaining the retrieval results. Our experiments on benchmark dataset, CUHK-PEDES, show that learning the attribute-space not only helps in improving performance, giving us state-of-the-art Rank-1 accuracy of 56.68%, but also yields humanly-interpretable features.
Surbhi Aggarwal, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001
WACV3
2020 QUICKSAL: A small and sparse visual saliency model for efficient inference in resource constrained hardware
abstract
Visual saliency is an important problem in the field of cognitive science and computer vision with applications such as surveillance, adaptive compressing, detecting unknown objects and scene understanding. In this paper, we propose a small and sparse neural network model for performing salient object segmentation that is suitable for use in mobile and embedded applications. Our model is built using depthwise separable convolutions and bottleneck inverted residuals which have been proven to perform very memory efficient inference and can be easily implemented using standard functions available in all deep learning frameworks. The multiscale features extracted along the layers with deep residuals allow our network to learn high quality saliency maps. We present the quantitative results of our QUICKSAL model with multiple levels of model sparsity ranging from 0% to ~96%, with the non-zero parameter count varying from ~3.3M to ~0.14M respectively - on publicly available benchmark datasets - showing that our highly constrained approach is comparable to other state-of-the-art approaches (parameter count ~35M). We also present qualitative results on camouflage images and show that our model can successfully distinguish between the salient and non-salient parts even when both seem blended together.
Vignesh Ramanathan, Pritesh Dwivedi, Bharath Katabathuni, Anirban Chakraborty 0001, Chetan Singh Thakur
WACV4
2020 PerSeg : segmenting salient objects from bag of single image perturbations
Avishek Majumder 0001, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001
Multim. Tools Appl.3
2020 Operator-in-the-Loop Deep Sequential Multi-Camera Feature Fusion for Person Re-Identification
abstract
Given a target image as query, person re-identification systems retrieve a ranked list of candidate matches on a per-camera basis. In deployed systems, a human operator scans these lists and labels sighted targets by touch or mouse-based selection. However, classical re-id approaches generate per-camera lists independently. Therefore, target identifications by operator in a subset of cameras cannot be utilized to improve ranking of the target in remaining set of network cameras. To address this shortcoming, we propose a novel sequential multi-camera re-id approach. The proposed approach can accommodate human operator inputs and provides early gains via a monotonic improvement in target ranking. At the heart of our approach is a fusion function which operates on deep feature representations of query and candidate matches. We formulate an optimization procedure custom-designed to incrementally improve query representation. Since existing evaluation methods cannot be directly adopted to our setting, we also propose two novel evaluation protocols. The results on two large-scale re-id datasets (Market-1501, DukeMTMC-reID) demonstrate that our multi-camera method significantly outperforms baselines and other popular feature fusion schemes. Additionally, we conduct a comparative subject-based study of human operator performance. The superior operator performance enabled by our approach makes a compelling case for its integration into deployable video-surveillance systems.
Navaneet K. L., Ravi Kiran Sarvadevabhatla, Shashank Shekhar 0006, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001
IEEE Trans. Inf. Forensics Secur.5
2019 All for One: Frame-wise Rank Loss for Improving Video-based Person Re-identification
abstract
Person re-identification involves retrieving correct matches for a target image (query) from a set of gallery images, while video based re-identification extends this to the case of query and gallery videos. Typical video-based re-id methods ignore the temporal evolution of the intermediate representations of the video sequences. We propose a novel loss function, termed rank loss, to explicitly ensure that the learnt representations achieve enhanced performance and robustness as the sequence progresses and that better intermediate representations result in an improved final representation. Experiments indicate that the addition of rank loss indeed helps in improving the re-id performance while achieving performance comparable to state-of-the-art approaches.
Navaneet K. L., Vasudha Todi, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001
ICASSP4
2019 From Strings to Things: Knowledge-Enabled VQA Model That Can Read and Reason
abstract
Text present in images are not merely strings, they provide useful cues about the image. Despite their utility in better image understanding, scene texts are not used in traditional visual question answering (VQA) models. In this work, we present a VQA model which can read scene texts and perform reasoning on a knowledge graph to arrive at an accurate answer. Our proposed model has three mutually interacting modules: i. proposal module to get word and visual content proposals from the image, ii. fusion module to fuse these proposals, question and knowledge base to mine relevant facts, and represent these facts as multi-relational graph, iii. reasoning module to perform a novel gated graph neural network based reasoning on this graph. The performance of our knowledge-enabled VQA model is evaluated on our newly introduced dataset, viz. text-KVQA. To the best of our knowledge, this is the first dataset which identifies the need for bridging text recognition with knowledge graph based reasoning. Through extensive experiments, we show that our proposed method outperforms traditional VQA as well as question-answering over knowledge base-based methods on text-KVQA.
Ajeet Kumar Singh, Anand Mishra 0001, Shashank Shekhar 0006, Anirban Chakraborty 0001
ICCV4
2019 OCR-VQA: Visual Question Answering by Reading Text in Images
abstract
The problem of answering questions about an image is popularly known as visual question answering (or VQA in short). It is a well-established problem in computer vision. However, none of the VQA methods currently utilize the text often present in the image. These "texts in images" provide additional useful cues and facilitate better understanding of the visual content. In this paper, we introduce a novel task of visual question answering by reading text in images, i.e., by optical character recognition or OCR. We refer to this problem as OCR-VQA. To facilitate a systematic way of studying this new problem, we introduce a large-scale dataset, namely OCRVQA-200K. This dataset comprises of 207,572 images of book covers and contains more than 1 million question-answer pairs about these images. We judiciously combine well-established techniques from OCR and VQA domains to present a novel baseline for OCR-VQA-200K. The experimental results and rigorous analysis demonstrate various challenges present in this dataset leaving ample scope for the future research. We are optimistic that this new task along with compiled dataset will open-up many exciting research avenues both for the document image analysis and the VQA communities.
Anand Mishra 0001, Shashank Shekhar 0006, Ajeet Kumar Singh, Anirban Chakraborty 0001
ICDAR4
2019 Efficient Person Re-Identification in Videos Using Sequence Lazy Greedy Determinantal Point Process (SLGDPP)
abstract
Given a sequence of observations for each person in each camera, identifying or re-identifying the same person across different cameras is one of the objectives of video surveillance systems. In the case of video based person re-id, the challenge is to handle the high correlation between temporally adjacent frames. The presence of non-informative frames results in high redundancy which needs to be removed for an efficient re-id. We propose a novel method to handle this challenge using Determinantal Point Process (DPP) to select the most diverse and informative subset of frames from a given sequence. Since subset selection problem is NP-Hard, we propose to use an approximate solution called Lazy Greedy DPP (LGDPP) and further extend it to utilize the temporal information of sequences with our proposed Sequential LGDPP (SLGDPP) for video-based person re-id. The major advantages of the proposed DPP variants are their simplicity and plug and play nature, which make it possible to use them atop any pretrained re-id model followed by a feature fusion module. The effectiveness of proposed frameworks is demonstrated on two popular video re-id benchmark datasets through improvements over state-of-the-art methods and naive baseline sampling methods.
Gaurav Kumar Nayak, Utkarsh Shreemali, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001
ICIP4
2019 Zero-Shot Knowledge Distillation in Deep Networks
abstract
Knowledge distillation deals with the problem of training a smaller model (Student) from a high capacity source model (Teacher) so as to retain most of its performance. Existing approaches use either the training data or meta-data extracted from it in order to train the Student. However, accessing the dataset on which the Teacher has been trained may not always be feasible if the dataset is very large or it poses privacy or safety concerns (e.g., bio-metric or medical data). Hence, in this paper, we propose a novel data-free method to train the Student from the Teacher. Without even using any meta-data, we synthesize the Data Impressions from the complex Teacher model and utilize these as surrogates for the original training data samples to transfer its learning to Student via knowledge distillation. We, therefore, dub our method “Zero-Shot Knowledge Distillation" and demonstrate that our framework results in competitive generalization performance as achieved by distillation using the actual training data samples on multiple benchmark datasets.
Gaurav Kumar Nayak, Konda Reddy Mopuri, Vaisakh Shaj, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001
ICML5
2019 N-HAR: A Neuromorphic Event-Based Human Activity Recognition System using Memory Surfaces
abstract
In recent years, a new generation of low-power, neuromorphic, event-based vision sensors has been gaining popularity for their very low latency and data sparsity. Though the conventional frame-based cameras have advanced in a lot of ways, they suffer from data redundancy and temporal latency. The bio-inspired artificial retinas eliminate the data redundancy by capturing only the change in illumination at each pixel and asynchronously communicating in binary spikes. In this work, we propose a system to achieve the task of human activity recognition based on the event-based camera data. We show that such tasks, which generally need high frame rate sensors for accurate predictions, can be achieved by adapting existing computer vision techniques to the spiking domain. We used event memory surfaces to make the sparse event data compatible with deep convolutional neural networks (CNNs). We leverage upon the recent advances in deep convolutional networks based video analysis and adapt such frameworks onto the neuromorphic domain. We also provide the community with a new dataset consisting of five categories of human activities captured in real world without any simulations. We achieved an accuracy of 94.3% using event memory surfaces on our activity recognition dataset.
Bibrat Ranjan Pradhan, Yeshwanth Bethi, Sathyaprakash Narayanan, Anirban Chakraborty 0001, Chetan Singh Thakur
ISCAS4
2017 Optimal Landmark Selection for Registration of 4D Confocal Image Stacks in Arabidopsis
abstract
Technologically advanced imaging techniques have allowed us to generate and study the internal part of a tissue over time by capturing serial optical images that contain spatio-temporal slices of hundreds of tightly packed cells. Image registration of such live-imaging datasets of developing multicelluar tissues is one of the essential components of all image analysis pipelines. In this paper, we present a fully automated 4D(X-Y-Z-T) registration method of live imaging stacks that takes care of both temporal and spatial misalignments. We present a novel landmark selection methodology where the shape features of individual cells are not of high quality and highly distinguishable. The proposed registration method finds the best image slice correspondence from consecutive image stacks to account for vertical growth in the tissue and the discrepancy in the choice of the starting focal point. Then, it uses local graph-based approach to automatically find corresponding landmark pairs, and finally the registration parameters are used to register the entire image stack. The proposed registration algorithm combined with an existing tracking method is tested on multiple image stacks of tightly packed cells of Arabidopsis shoot apical meristem and the results show that it significantly improves the accuracy of cell lineages and division statistics.
Katya Mkrtchyan, Anirban Chakraborty 0001, Amit K. Roy-Chowdhury
IEEE ACM Trans. Comput. Biol. Bioinform.2
2017 Person Reidentification Using Multiple Egocentric Views
abstract
Development of a robust and scalable multicamera surveillance system is the need of the hour to ensure public safety and security. Being able to reidentify and track one or more targets over multiple nonoverlapping camera field of views in a crowded environment remains an important and challenging problem because of occlusions, large change in the viewpoints, and illumination across cameras. However, the rise of wearable imaging devices has led to new avenues in solving the reidentification (re-id) problem. Unlike static cameras, where the views are often restricted or low resolution and occlusions are common scenarios, egocentric/first person views (FPVs) mostly get zoomed in, unoccluded face images. In this paper, we present a person re-id framework designed for a network of multiple wearable devices. The proposed framework builds on commonly used facial feature extraction and similarity computation methods between camera pairs and utilizes a data association method to yield globally optimal and consistent re-id results with much improved accuracy. Moreover, to ensure its utility in practical applications where a large amount of observations are available every instant, an online scheme is proposed as a direct extension of the batch method. This can dynamically associate new observations to already observed and labeled targets in an iterative fashion. We tested both the offline and online methods on realistic FPV video databases, collected using multiple wearable cameras in a complex office environment and observed large improvements in performance when compared with the state of the arts.
Anirban Chakraborty 0001, Bappaditya Mandal, Junsong Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.1
2016 A poisson process model for activity forecasting
abstract
Activity forecasting has recently become an active research area for its importance in critical applications like automated navigation and human-computer interaction. However, for a video observed upto a certain time, all of the existing forecasting works focus on predicting the activity label, i.e., predicting what the next unobserved activity is. To the best of our knowledge, no work has answered the crucial question yet: when the next unobserved activity will occur. In this paper, we propose an approach for predicting the starting time of the next unobserved activity without assuming that we know its label. We model activities occurring at a variable rate using a Log-Gaussian Cox Process (LGCP) and learn the rate function from the training data. Then the starting time is predicted using importance sampling algorithm. In our experiments on the challenging MPII-Cooking dataset, we find that both the label of the last observed activity and the label of the activity being predicted affect the time prediction accuracy.
Tahmida Mahmud, Mahmudul Hasan 0003, Anirban Chakraborty 0001, Amit K. Roy-Chowdhury
ICIP3
2016 Person re-identification using multiple first-person-views on wearable devices
abstract
The rise of wearable devices has led to many new ways of re-identifying an individual. Unlike static cameras, where the views are often restricted or zoomed out and occlusions are common scenarios, first-person-views (FPVs) or ego-centric views see people closely and mostly get un-occluded face images. In this paper, we propose a face re-identification framework designed for a network of multiple wearable devices. This framework utilizes a global data association method termed as Network Consistent Reidentification (NCR) that not only helps in maintaining consistency in association results across the network, but also improves the pair-wise face re-identification accuracy. To test the proposed pipeline, we collected a database of FPV videos of 72 persons using multiple wearable devices (such as Google Glasses) in a multi-storied office environment. Experimental results indicate that NCR is able to consistently achieve large performance gains when compared to the state-of-the-art methodologies.
Anirban Chakraborty 0001, Bappaditya Mandal, Hamed Kiani Galoogahi
WACV1
2016 Network Consistent Data Association
abstract
Existing data association techniques mostly focus on matching pairs of data-point sets and then repeating this process along space-time to achieve long term correspondences. However, in many problems such as person re-identification, a set of data-points may be observed at multiple spatio-temporal locations and/or by multiple agents in a network and simply combining the local pairwise association results between sets of data-points often leads to inconsistencies over the global space-time horizons. In this paper, we propose a Novel Network Consistent Data Association (NCDA) framework formulated as an optimization problem that not only maintains consistency in association results across the network, but also improves the pairwise data association accuracies. The proposed NCDA can be solved as a binary integer program leading to a globally optimal solution and is capable of handling the challenging data-association scenario where the number of data-points varies across different sets of instances in the network. We also present an online implementation of NCDA method that can dynamically associate new observations to already observed data-points in an iterative fashion, while maintaining network consistency. We have tested both the batch and the online NCDA in two application areas-person re-identification and spatio-temporal cell tracking and observed consistent and highly accurate data association results in all the cases.
Anirban Chakraborty 0001, Abir Das, Amit K. Roy-Chowdhury
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Context aware spatio-temporal cell tracking in densely packed multilayer tissues
Anirban Chakraborty 0001, Amit K. Roy-Chowdhury
Medical Image Anal.1
2014 Context-Aware Activity Forecasting
Anirban Chakraborty 0001, Amit K. Roy-Chowdhury
ACCV (5)1
2014 Consistent Re-identification in a Camera Network
Abir Das, Anirban Chakraborty 0001, Amit K. Roy-Chowdhury
ECCV (2)2
2014 A conditional random field model for tracking in densely packed cell structures
abstract
Automated tracking of plant and animal cells in time lapse live-imaging datasets of developing multicellular tissues is required for quantitative, high throughput analysis of cell division, migration and cell growth. In this paper, we present a novel cell tracking method that exploits the tight spatial topology of neighboring cells in a multicellular field as contextual information and combines it with physical features of individual cells for generating reliable cell lineages. The 2D image slices of multicellular tissues are modeled as CRFs and spatio-temporal cell to cell correspondences are obtained by performing inference on this CRF using loopy belief propagation. We present results on a (3D+t) confocal image stack of Arabidopsis shoot meristem and show that the method can handle many visual analysis challenges associated with such cell tracking problems, viz. poor feature quality of individual cells, low SNR in parts of images, variable number of cells across slices and cell division detection.
Anirban Chakraborty 0001, Amit K. Roy-Chowdhury
ICIP1
2011 Cell Resolution 3D Reconstruction of Developing Multilayer Tissues from Sparsely Sampled Volumetric Microscopy Images
abstract
Understanding of the growth dynamics in developmental biology is often pursued through the analysis of cell sizes and shapes obtained from CLSM based imaging at cell resolution of multi-layer tissues. This necessitates the development of robust 3D reconstruction methods using such images. However, all of the current methods of 3D reconstruction using CLSM imaging require large number of cell slices. But in the case of live cell imaging, i.e., imaging a growing tissue, such high depth resolution is not feasible in order to avoid photodynamic damage to the growing cells from prolonged exposure to laser radiation. In this work, we have addressed the problem of 3D reconstruction at cell resolution of a developing multi-layer tissue in the plant meristem when the amount of data is as limited as two to four slices per cell. This introduces significant image analysis challenges in terms of sparsity of the data, low signal-to-noise ratio, and a wide range of shapes and sizes. Motivated by the physical structure of the cells, we propose to reconstruct a cell cluster as a packing of truncated ellipsoids representing the individual cells. We test the proposed computational method on time-lapse CLSM images of Shoot Apical Meristem (SAM) cells of model plant Arabidopsis Thaliana. We show that the 3D reconstruction can lead to 3D shape models of complete cell clusters, which is an essential first step towards obtaining growth statistics for individual cells.
Anirban Chakraborty 0001, Ram Kishor Yadav, G. Venugopala Reddy, Amit K. Roy-Chowdhury
BIBM1