VLDB 2026 Research / reviewers in the wild / expert
Venkatesh Babu Radhakrishnan
dblp:20/6289 · also R. Venkatesh Babu, Radhakrishnan Venkatesh Babu, Venkatesh Babu R.
· DBLP profile ↗
171ranked-venue papers
24as first author
55since 2021 · last 2026
0000-0002-1926-1804ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 138 · 16 first-author · 40 since 2021Artificial intelligence and machine learning · 106 · 10 first-author · 48 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Harnessing Diffusion-Generated Synthetic Images for Fair Image ClassificationabstractImage classification systems often inherit biases from uneven group representation in training data. For example, in face datasets for hair color classification, blond hair may be disproportionately associated with females, reinforcing stereotypes. A recent approach leverages the Stable Diffusion model to generate balanced training data, but these models often struggle to preserve the original data distribution. In this work, we explore multiple diffusion-finetuning techniques, e.g., LoRA and DreamBooth, to generate images that more accurately represent each training group by learning directly from their samples. Additionally, in order to prevent a single DreamBooth model from being overwhelmed by excessive intra-group variations, we explore a technique of clustering images within each group and train a DreamBooth model per cluster. These models are then used to generate group-balanced data for pretraining, followed by fine-tuning on real data. Experiments on multiple benchmarks demonstrate that the studied finetuning approaches outperform vanilla Stable Diffusion on average and achieve results comparable to SOTA debiasing techniques like Group-DRO, while surpassing them as the dataset bias severity increases. Abhipsa Basu, Aviral Gupta, Abhijnya Bhat, Venkatesh Babu Radhakrishnan |
AAAI | 4 |
| 2026 | UniC-Lift: Unified 3D Instance Segmentation via Contrastive Learningabstract3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF) have advanced novel-view synthesis. Recent methods extend multi-view 2D segmentation to 3D, enabling instance/semantic segmentation for better scene understanding. A key challenge is the inconsistency of 2D instance labels across views, leading to poor 3D predictions. Existing methods use a two-stage approach in which some rely on contrastive learning with hyperparameter-sensitive clustering, while others preprocess labels for consistency. We propose a unified framework that merges these steps, reducing training time and improving performance by introducing a learnable feature embedding for segmentation in Gaussian primitives. This embedding is then efficiently decoded into instance labels through a novel "Embedding-to-Label" process, effectively integrating the optimization. While this unified framework offers substantial benefits, we observed artifacts at the object boundaries. To address the object boundary issues, we propose hard-mining samples along these boundaries. However, directly applying hard mining to the feature embeddings proved unstable. Therefore, we apply a linear layer to the rasterized feature embeddings before calculating the triplet loss, which stabilizes training and significantly improves performance. Our method outperforms baselines qualitatively and quantitatively on the ScanNet, Replica3D, and Messy-Rooms datasets. Ankit Dhiman, R. Srinath, Jaswanth Reddy, Lokesh R. Boregowda, Venkatesh Babu Radhakrishnan |
AAAI | 5 |
| 2025 | Reflecting Reality: Enabling Diffusion Models to Produce Faithful Mirror ReflectionsabstractWe tackle the problem of generating highly realistic and plausible mirror reflections using diffusion-based generative models. We formulate this problem as an image inpainting task, allowing for more user control over the placement of mirrors during the generation process. To enable this, we create SynMirror, a large-scale dataset of diverse synthetic scenes with objects placed in front of mirrors. SynMirror contains around 198K samples rendered from 66K unique 3D objects, along with their associated depth maps, normal maps and instance-wise segmentation masks, to capture relevant geometric properties of the scene. Using this dataset, we propose a novel depth-conditioned inpainting method called MirrorFusion, which generates high-quality geometrically consistent and photo-realistic mirror reflections given an input image and a mask depicting the mirror region. MirrorFusion outperforms state-of-the-art methods on SynMirror, as demonstrated by extensive quantitative and qualitative analysis. To the best of our knowledge, we are the first to successfully tackle the challenging problem of generating controlled and faithful mirror reflections of an object in a scene using diffusion based models. Syn-Mirror and MirrorFusion open up new avenues for image editing and augmented reality applications for practitioners and researchers alike. The project page is available at: https://val.cds.iisc.ac.in/reflecting-reality.github.io/. Ankit Dhiman, Manan Shah, Rishubh Parihar, Yash Bhalgat, Lokesh R. Boregowda, Venkatesh Babu Radhakrishnan |
3DV | 6 |
| 2025 | MirrorVerse: Pushing Diffusion Models to Realistically Reflect the WorldabstractDiffusion models have become central to various image editing tasks, yet they often fail to fully adhere to physical laws, particularly with effects like shadows, reflections, and occlusions. In this work, we address the challenge of generating photorealistic mirror reflections using diffusion-based generative models. Despite extensive training data, existing diffusion models frequently overlook the nuanced details crucial to authentic mirror reflections. Recent approaches have attempted to resolve this by creating synthetic datasets and framing reflection generation as an in-painting task; however, they struggle to generalize across different object orientations and positions relative to the mirror. Our method overcomes these limitations by introducing key augmentations into the synthetic data pipeline: (1) random object positioning, (2) randomized rotations, and (3) grounding of objects, significantly enhancing generalization across poses and placements. To further address spatial relationships and occlusions in scenes with multiple objects, we implement a strategy to pair objects during dataset generation, resulting in a dataset robust enough to handle these complex scenarios. Achieving generalization to real-world scenes remains a challenge, so we introduce a three-stage training curriculum to develop the MirrorFusion 2.0 model to improve real-world performance. We provide extensive qualitative and quantitative evaluations to support our approach. The project page is available at: https://mirror-verse.github.io/. Ankit Dhiman, Manan Shah, Venkatesh Babu Radhakrishnan |
CVPR | 3 |
| 2025 | Compass Control: Multi Object Orientation Control for Text-to-Image GenerationabstractExisting approaches for controlling text-to-image diffusion models, while powerful, do not allow for explicit 3D object-centric control, such as precise control of object orientation. In this work, we address the problem of multi-object orientation control in text-to-image diffusion models. This enables the generation of diverse multi-object scenes with precise orientation control for each object. The key idea is to condition the diffusion model with a set of orientation-aware compass tokens, one for each object, along with text tokens. A light-weight encoder network predicts these compass tokens taking object orientation as the input. The model is trained on a synthetic dataset of procedurally generated scenes, each containing one or two 3D assets on a plain background. However, direct training this framework results in poor orientation control as well as leads to entanglement among objects. To mitigate this, we intervene in the generation process and constrain the cross-attention maps of each compass token to its corresponding object regions. The trained model is able to achieve precise orientation control for a) complex objects not seen during training and b) multi-object scenes with more than two objects, indicating strong generalization capabilities. Further, when combined with personalization methods, our method precisely controls the orientation of the new object in diverse contexts. Our method achieves state-of-the-art orientation control and text alignment quantified with extensive evaluations and a user study. Project website Rishubh Parihar, Vaibhav Agrawal, Sachidanand VS, Venkatesh Babu Radhakrishnan |
CVPR | 4 |
| 2025 | MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular DetectionabstractCurrent monocular 3D detectors are held back by the limited diversity and scale of real-world datasets. While data augmentation certainly helps, it’s particularly difficult to generate realistic scene-aware augmented data for outdoor settings. Most current approaches to synthetic data generation focus on realistic object appearance through improved rendering techniques. However, we show that where and how objects are positioned is just as crucial for training effective 3D monocular detectors. The key obstacle lies in automatically determining realistic object placement parameters - including position, dimensions, and directional alignment when introducing synthetic objects into actual scenes. To address this, we introduce MonoPlace3D, a novel system that considers the 3D scene content to create realistic augmentations. Specifically, given a background scene, Mono-Place3D learns a distribution over plausible 3D bounding boxes. Subsequently, we render realistic objects and place them according to the locations sampled from the learned distribution. Our comprehensive evaluation on two standard datasets KITTI and NuScenes, demonstrates that MonoPlace3D significantly improves the accuracy of multiple existing monocular 3D detectors while being highly data efficient. Project Page Rishubh Parihar, Srinjay Sarkar, Sarthak Vora, Jogendra Kundu, Venkatesh Babu Radhakrishnan |
CVPR | 5 |
| 2025 | Composing Parts for Expressive Object GenerationabstractImage composition and generation are processes where the artists need control over various parts of the generated images. However, the current state-of-the-art generation models, like Stable Diffusion, cannot handle fine-grained part-level attributes in the text prompts. Specifically, when additional attribute details are added to the base text prompt, these text-to-image models either generate an image vastly different from the image generated from the base prompt or ignore the attribute details. To mitigate these issues, we introduce PartComposer, a training-free method that enables image generation based on fine-grained part-level attributes specified for objects in the base text prompt. This allows more control for artists and enables novel object compositions by combining distinctive object parts. PartComposer first localizes object parts by denoising the object region from a specific diffusion process. This enables each part token to be localized to the right region. After obtaining part masks, we run a localized diffusion process in each part region based on fine-grained part attributes and combine them to produce the final image. All stages of PartComposer are based on repurposing a pre-trained diffusion model, which enables it to generalize across domains. We demonstrate the effectiveness of part-level control provided by PartComposer through qualitative visual examples and quantitative comparisons with contemporary baselines. Harsh Rangwani, Aishwarya Agarwal, Kuldeep Kulkarni, Venkatesh Babu Radhakrishnan, Srikrishna Karanam |
CVPR | 4 |
| 2025 | Zero-Shot Depth Aware Image Editing With Diffusion Models
Rishubh Parihar, Sachidanand VS, Venkatesh Babu Radhakrishnan |
ICCV | 3 |
| 2025 | Discovering a Zero (Zero-Vector Class of Machine Learning)abstractIn Machine learning, separating data into classes is a very fundamental problem. A mathematical framework around the classes is presented in this work to deepen the understanding of classes. The classes are defined as vectors in a Vector Space, where addition corresponds to the union of classes, and scalar multiplication resembles set complement of classes. The Zero-Vector in the vector space corresponds to a class referred to as the Metta-Class. This discovery enables numerous applications. One such application, termed 'clear learning' in this work, focuses on learning the true nature (manifold) of the data instead of merely learning a boundary sufficient for classification. Another application, called 'unary class learning', involves learning a single class in isolation rather than learning by comparing two or more classes. Additionally, 'set operations on classes' is another application highlighted in this work. Furthermore, Continual Learning of classes is facilitated by smaller networks. The Metta-Class enables neural networks to learn only the data manifold; therefore, it can also be used for generation of new data. Results for the key applications are shown using the MNIST dataset. To further strengthen the claims, some results are also produced using the CIFAR-10 and ImageNet-1k embeddings. The code supporting these applications is publicly available at: github.com/hm-4/Metta-Class. Harikrishna Metta, Venkatesh Babu Radhakrishnan |
ICML | 2 |
| 2025 | ChromaDistill: Colorizing Monochrome Radiance Fields with Knowledge DistillationabstractColorization is a well-explored problem in the domains of image and video processing. However, extending colorization to 3D scenes presents significant challenges. Re-cent Neural Radiance Field (NeRF) and Gaussian-Splatting (3DGS) methods enable high-quality novel-view synthe-sis for multi-view images. However, the question arises: How can we colorize these 3D representations? This work presents a method for synthesizing colorized novel views from input grayscale multi-view images. Using image or video colorization methods to colorize novel views from these 3D representations naively will yield output with se-vere inconsistencies. We introduce a novel method to use powerful image colorization models for colorizing 3D representations. We propose a distillation-based method that transfers color from these networks trained on natural images to the target 3D representation. Notably, this strat-egy does not add any additional weights or computational overhead to the original representation during inference. Extensive experiments demonstrate that our method produces high-quality colorized views for indoor and outdoor scenes, showcasing significant cross-view consistency advantages over baseline approaches. Our method is agnos-tic to the underlying 3D representation and easily gener-alizable to NeRF and 3DGS methods. Further, we vali-date the efficacy of our approach in several diverse applications: 1.) Infra-Red (IR) multi-view images and 2.) Legacy grayscale multi-view image sequences. Project Webpage: https://val.cds.iisc.ac.in/chroma-distill.github.io/ Ankit Dhiman, R. Srinath, Srinjay Sarkar, Lokesh R. Boregowda, Venkatesh Babu Radhakrishnan |
WACV | 5 |
| 2025 | Attribute Diffusion: Diffusion Driven Diverse Attribute EditingabstractImage attribute editing is a widely researched area fueled by the recent advancements in deep generative models. Existing methods treat semantic attributes as binary and do not allow the user to generate multiple variations of the attribute edits. This limits the applications of editing methods in the real world, e.g., exploring multiple eyeglass variations on an e-commerce platform. In this work, we present a technique to generate a collection of diverse attribute edits and a principled way to explore them. Generation and controlled exploration of attribute variations is challenging as it requires fine control over the attribute styles while preserving other attributes and the identity of the subject. Capitalizing on the attribute disentanglement property of the latent spaces of pre-trained GANs, we represent the attribute edits in this space. Next, we train a diffusion model to model these latent directions of edits. We propose a coarse-to-fine sampling strategy to explore these variations in a controlled manner. Extensive experiments on various datasets establish the effectiveness and generalization of the proposed approach for the generation and controlled exploration of diverse attribute edits. Code is available at - project page. Rishubh Parihar, Prasanna Balaji, Raghav Magazine, Sarthak Vora, Varun Jampani, Venkatesh Babu Radhakrishnan |
WACV | 6 |
| 2024 | Leveraging Vision-Language Models for Improving Domain Generalization in Image ClassificationabstractVision-Language Models (VLMs) such as CLIP are trained on large amounts of image-text pairs, resulting in remarkable generalization across several data distributions. However, in several cases, their expensive training and data collection/curation costs do not justify the end application. This motivates a vendor-client paradigm, where a vendor trains a large-scale VLM and grants only input-output access to clients on a pay-per-query basis in a black-box setting. The client aims to minimize inference cost by distilling the VLM to a student model using the limited available task-specific data, and further deploying this student model in the downstream application. While naive distillation largely improves the In-Domain (ID) accuracy of the student, it fails to transfer the superior out-of-distribution (OOD) generalization of the VLM teacher using the limited available labeled images. To mitigate this, we propose Vision-Language to Vision - Align, Distill, Predict (VL2V-ADiP), which first aligns the vision and language modalities of the teacher model with the vision modality of a pre-trained student model, and further distills the aligned VLM representations to the student. This maximally retains the pre-trained features of the student, while also incorporating the rich representations of the VLM image encoder and the superior generalization of the text embeddings. The proposed approach achieves state-of-the-art results on the standard Domain Generalization benchmarks in a black-box teacher setting as well as a white-box setting where the weights of the VLM are accessible. Project page: http://val.cds.iisc.ac.in/VL2V-ADiP/ Sravanti Addepalli, Ashish Ramayee Asokan, Lakshay Sharma, Venkatesh Babu Radhakrishnan |
CVPR | 4 |
| 2024 | Balancing Act: Distribution-Guided Debiasing in Diffusion ModelsabstractDiffusion Models (DMs) have emerged as powerful generative models with unprecedented image generation capability. These models are widely used for data augmentation and creative applications. However, DMs reflect the biases present in the training datasets. This is especially concerning in the context of faces, where the DM prefers one demographic subgroup vs others (eg. female vs male). In this work, we present a method for debiasing DMs without relying on additional reference data or model retraining. Specifically, we propose Distribution Guidance, which enforces the generated images to follow the prescribed attribute distribution. To realize this, we build on the key insight that the latent features of denoising UNet hold rich demographic semantics, and the same can be leveraged to guide debiased generation. We train Attribute Distribution Predictor (ADP) - a small mlp that maps the latent features to the distribution of attributes. ADP is trained with pseudo labels generated from existing attribute classifiers. The proposed Distribution Guidance with ADP enables us to do fair generation. Our method reduces bias across single/multiple attributes and outperforms the baseline by a significant margin for unconditional and text-conditional diffusion models. Further, we present a downstream task of training a fair attribute classifier by augmenting the training set with our generated data. Code is available at - project page. Rishubh Parihar, Abhijnya Bhat, Abhipsa Basu, Saswat Mallick, Jogendra Kundu, Venkatesh Babu Radhakrishnan |
CVPR | 6 |
| 2024 | DeiT-LT: Distillation Strikes Back for Vision Transformer Training on Long-Tailed DatasetsabstractVision Transformer (ViT) has emerged as a prominent architecture for various computer vision tasks. In ViT, we divide the input image into patch tokens and process them through a stack of self-attention blocks. However, unlike Convolutional Neural Network (CNN), ViT's simple architecture has no informative inductive bias (e.g., locality, etc.). Due to this, ViT requires a large amount of data for pre-training. Various data-efficient approaches (DeiT) have been proposed to train ViT on balanced datasets effectively. However, limited literature discusses the use of ViT for datasets with long-tailed imbalances. In this work, we introduce DeiT-LT to tackle the problem of training ViTs from scratch on long-tailed datasets. In DeiT-LT, we introduce an efficient and effective way of distillation from CNN via distillation DIST token by using out-of-distribution images and re-weighting the distillation loss to enhance focus on tail classes. This leads to the learning of local CNN-like features in early ViT blocks, improving generalization for tail classes. Further, to mitigate overfitting, we propose distilling from a flat CNN teacher, which leads to learning low-rank generalizable features for DIST tokens across all ViT blocks. With the proposed DeiT-LT scheme, the distillation DIST token becomes an expert on the tail classes, and the classifier CLS token becomes an expert on the head classes. The experts help to effectively learn features corresponding to both the majority and minority classes using a distinct set of tokens within the same ViT architecture. We show the effectiveness of DeiT-LT for training ViT from scratch on datasets ranging from small-scale CIFAR-10 LT to large-scale iNaturalist-2018. Project Page: https://rangwani-harsh.github.io/DeiT-LT. Harsh Rangwani, Pradipto Mondal, Ashish Ramayee Asokan, Venkatesh Babu Radhakrishnan |
CVPR | 5 |
| 2024 | Text2Place: Affordance-Aware Text Guided Human Placement
Rishubh Parihar, Sachidanand VS, Venkatesh Babu Radhakrishnan |
ECCV (3) | 4 |
| 2024 | PreciseControl: Enhancing Text-to-Image Diffusion Models with Fine-Grained Attribute Control
Rishubh Parihar, Sachidanand VS, Sabariswaran Mani, Tejan Karmali, Venkatesh Babu Radhakrishnan |
ECCV (82) | 5 |
| 2024 | Selective Mixup Fine-Tuning for Optimizing Non-Decomposable ObjectivesabstractThe rise in internet usage has led to the generation of massive amounts of data, resulting in the adoption of various supervised and semi-supervised machine learning algorithms, which can effectively utilize the colossal amount of data to train models. However, before deploying these models in the real world, these must be strictly evaluated on performance measures like worst-case recall and satisfy constraints such as fairness. We find that current state-of-the-art empirical techniques offer sub-optimal performance on these practical, non-decomposable performance objectives. On the other hand, the theoretical techniques necessitate training a new model from scratch for each performance objective. To bridge the gap, we propose SelMix, a selective mixup-based inexpensive fine-tuning technique for pre-trained models, to optimize for the desired objective. The core idea of our framework is to determine a sampling distribution to perform a mixup of features between samples from particular classes such that it optimizes the given objective. We comprehensively evaluate our technique against the existing empirical and theoretically principled methods on standard benchmark datasets for imbalanced classification. We find that proposed SelMix fine-tuning significantly improves the performance for various practical non-decomposable objectives across benchmarks. Shrinivas Ramasubramanian, Harsh Rangwani, Sho Takemori, Kunal Samanta, Yuhei Umeda, Venkatesh Babu Radhakrishnan |
ICLR | 6 |
| 2024 | Mitigating Biases in Blackbox Feature Extractors for Image Classification TasksabstractIn image classification, it is common to utilize a pretrained model to extract meaningful features of the input images, and then to train a classifier on top of it to make predictions for any downstream task. Trained on enormous amounts of data, these models have been shown to contain harmful biases which can hurt their performance when adapted for a downstream classification task. Further, very often they may be blackbox, either due to scale, or because of unavailability of model weights or architecture. Thus, during a downstream task, we cannot debias such models by updating the weights of the feature encoder, as only the classifier can be finetuned. In this regard, we investigate the suitability of some existing debiasing techniques and thereby motivate the need for more focused research towards this problem setting. Furthermore, we propose a simple method consisting of a clustering-based adaptive margin loss with a blackbox feature encoder, with no knowledge of the bias attribute. Our experiments demonstrate the effectiveness of our method across multiple benchmarks. Abhipsa Basu, Saswat Subhajyoti Mallick, Venkatesh Babu Radhakrishnan |
NeurIPS | 3 |
| 2024 | Aligning Non-Causal Factors for Transformer-Based Source-Free Domain AdaptationabstractConventional domain adaptation algorithms aim to achieve better generalization by aligning only the task-discriminative causal factors between a source and target domain. However, we find that retaining the spurious correlation between causal and non-causal factors plays a vital role in bridging the domain gap and improving target adaptation. Therefore, we propose to build a framework that disentangles and supports causal factor alignment by aligning the non-causal factors first. We also investigate and find that the strong shape bias of vision transformers, coupled with its multi-head attentions, make it a suitable architecture for realizing our proposed disentanglement. Hence, we propose to build a Causality-enforcing Source-Free Transformer framework (C-SFTrans1) to achieve disentanglement via a novel two-stage alignment approach: a) non-causal factor alignment: non-causal factors are aligned using a style classification task which leads to an overall global alignment, b) task-discriminative causal factor alignment: causal factors are aligned via target adaptation. We are the first to investigate the role of vision transformers (ViTs) in a privacy-preserving source-free setting. Our approach achieves state-of-the-art results in several DA benchmarks. Sunandini Sanyal, Ashish Ramayee Asokan, Suvaansh Bhambri, Pradyumna YM, Akshay R. Kulkarni, Jogendra Kundu, Venkatesh Babu Radhakrishnan |
WACV | 7 |
| 2024 | Minimizing Layerwise Activation Norm Improves Generalization in Federated LearningabstractFederated Learning (FL) is an emerging machine learning framework that enables multiple clients (coordinated by a server) to collaboratively train a global model by aggregating the locally trained models without sharing any client’s training data. It has been observed in recent works that learning in a federated manner may lead the aggregated global model to converge to a ‘sharp minimum’ thereby adversely affecting the generalizability of this FL-trained model. Therefore, in this work, we aim to improve the generalization performance of models trained in a federated setup by introducing a ‘flatness’ constrained FL optimization problem. This flatness constraint is imposed on the top eigenvalue of the Hessian computed from the training loss. As each client trains a model on its local data, we further re-formulate this complex problem utilizing the client loss functions and propose a new computationally efficient regularization technique1, dubbed ‘MAN,’ which Minimizes Activation’s Norm of each layer on client-side models. We also theoretically show that minimizing the activation norm reduces the top eigenvalue of the layer-wise Hessian of the client’s loss, which in turn decreases the overall Hessian’s top eigenvalue, ensuring convergence to a flat minimum. We apply our proposed flatness-constrained optimization to the existing FL techniques and obtain significant improvements, thereby establishing new state-of-the-art. M. Yashwanth, Gaurav Kumar Nayak, Harsh Rangwani, Arya Singh, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001 |
WACV | 5 |
| 2023 | RMLVQA: A Margin Loss Approach For Visual Question Answering with Language BiasesabstractVisual Question Answering models have been shown to suffer from language biases, where the model learns a correlation between the question and the answer, ignoring the image. While early works attempted to use question-only models or data augmentations to reduce this bias, we propose an adaptive margin loss approach having two components. The first component considers the frequency of answers within a question type in the training data, which addresses the concern of the class-imbalance causing the language biases. However, it does not take into account the answering difficulty of the samples, which impacts their learning. We address this through the second component, where instance-specific margins are learnt, allowing the model to distinguish between samples of varying complexity. We introduce a bias-injecting component to our model, and compute the instance-specific margins from the confidence of this component. We combine these with the estimated margins to consider both answer-frequency and task-complexity in the training loss. We show that, while the margin loss is effective for out-of-distribution (ood) data, the bias-injecting component is essential for generalising to in-distribution (id) data. Our proposed approach, Robust Margin Loss for Visual Question Answering (RMLVQA)11Code available at https://github.com/val-iisc/RMLVQA improves upon the existing state-of-the-art results when compared to augmentation-free methods on benchmark VQA datasets suffering from language biases, while maintaining competitive performance on id data, making our method the most robust one among all comparable methods. Abhipsa Basu, Sravanti Addepalli, Venkatesh Babu Radhakrishnan |
CVPR | 3 |
| 2023 | DART: Diversify-Aggregate-Repeat Training Improves Generalization of Neural NetworksabstractGeneralization of Neural Networks is crucial for deploying them safely in the real world. Common training strategies to improve generalization involve the use of data augmentations, ensembling and model averaging. In this work, we first establish a surprisingly simple but strong benchmark for generalization which utilizes diverse augmentations within a training minibatch, and show that this can learn a more balanced distribution of features. Further, we propose Diversify-Aggregate-Repeat Training (DART) strategy that first trains diverse models using different augmentations (or domains) to explore the loss basin, and further Aggregates their weights to combine their expertise and obtain improved generalization. We find that Repeating the step of Aggregation throughout training improves the over-all optimization trajectory and also ensures that the individual models have sufficiently low loss barrier to obtain improved generalization on combining them. We theoretically justify the proposed approach and show that it indeed generalizes better. In addition to improvements in In-Domain generalization, we demonstrate SOTA performance on the Domain Generalization benchmarks in the popular DomainBed framework as well. Our method is generic and can easily be integrated with several base training algorithms to achieve performance gains. Our code is available here: https://github.com/val-iisc/DART. Samyak Jain, Sravanti Addepalli, Pawan Kumar Sahu, Priyam Dey, Venkatesh Babu Radhakrishnan |
CVPR | 5 |
| 2023 | NoisyTwins: Class-Consistent and Diverse Image Generation Through StyleGANsabstractStyleGANs are at the forefront of controllable image generation as they produce a latent space that is semantically disentangled, making it suitable for image editing and manipulation. However, the performance of StyleGANs severely degrades when trained via class-conditioning on large-scale long-tailed datasets. We find that one reason for degradation is the collapse of latents for each class in the$\mathcal{W}$latent space. With NoisyTwins, we first introduce an effective and inexpensive augmentation strategy for class embeddings, which then decorrelates the latents based on self-supervision in the$\mathcal{W}$space. This decorrelation mitigates collapse, ensuring that our method preserves intra-class diversity with class-consistency in image generation. We show the effectiveness of our approach on large-scale real-world long-tailed datasets of ImageNet-LT and iNaturalist 2019, where our method outperforms other methods by ∼ 19% on FID, establishing a new state-of-the-art. Harsh Rangwani, Lavish Bansal, Kartik Sharma, Tejan Karmali, Varun Jampani, Venkatesh Babu Radhakrishnan |
CVPR | 6 |
| 2023 | Inspecting the Geographical Representativeness of Images from Text-to-Image ModelsabstractRecent progress in generative models has resulted in models that produce both realistic as well as relevant images for most textual inputs. These models are being used to generate millions of images everyday, and hold the potential to drastically impact areas such as generative art, digital marketing and data augmentation. Given their outsized impact, it is important to ensure that the generated content reflects the artifacts and surroundings across the globe, rather than over-representing certain parts of the world. In this paper, we measure the geographical representativeness of common nouns (e.g., a house) generated through DALL•E 2 and Stable Diffusion models using a crowdsourced study comprising 540 participants across 27 countries. For deliberately underspecified inputs without country names, the generated images most reflect the surroundings of the United States followed by India, and the top generations rarely reflect surroundings from all other countries (average score less than 3 out of 5). Specifying the country names in the input increases the representativeness by 1.44 points on average on a 5 − point Likert scale for DALL•E 2 and 0.75 for Stable Diffusion, however, the overall scores for many countries still remain low, highlighting the need for future models to be more geographically inclusive. Lastly, we examine the feasibility of quantifying the geographical representativeness of generated images without conducting user studies.1 Abhipsa Basu, Venkatesh Babu Radhakrishnan, Danish Pruthi |
ICCV | 2 |
| 2023 | Strata-NeRF : Neural Radiance Fields for Stratified ScenesabstractNeural Radiance Field (NeRF) approaches learn the underlying 3D representation of a scene and generate photo-realistic novel views with high fidelity. However, most proposed settings concentrate on modelling a single object or a single level of a scene. However, in the real world, we may capture a scene at multiple levels, resulting in a layered capture. For example, tourists usually capture a monument’s exterior structure before capturing the inner structure. Modelling such scenes in 3D with seamless switching between levels can drastically improve immersive experiences. However, most existing techniques struggle in modelling such scenes. We propose Strata-NeRF, a single neural radiance field that implicitly captures a scene with multiple levels. Strata-NeRF achieves this by conditioning the NeRFs on Vector Quantized (VQ) latent representations which allow sudden changes in scene structure. We evaluate the effectiveness of our approach in multi-layered synthetic dataset comprising diverse scenes and then further validate its generalization on the real-world RealEstate10K dataset. We find that Strata-NeRF effectively captures stratified scenes, minimizes artifacts, and synthesizes high-fidelity views compared to existing approaches. https://ankitatiisc.github.io/Strata-NeRF/ Ankit Dhiman, R. Srinath, Harsh Rangwani, Rishubh Parihar, Lokesh R. Boregowda, Srinath Sridhar 0002, Venkatesh Babu Radhakrishnan |
ICCV | 7 |
| 2023 | Domain-Specificity Inducing Transformers for Source-Free Domain AdaptationabstractConventional Domain Adaptation (DA) methods aim to learn domain-invariant feature representations to improve the target adaptation performance. However, we motivate that domain-specificity is equally important since in-domain trained models hold crucial domain-specific properties that are beneficial for adaptation. Hence, we propose to build a framework that supports disentanglement and learning of domain-specific factors and task-specific factors in a unified model. Motivated by the success of vision transformers in several multi-modal vision problems, we find that queries could be leveraged to extract the domain-specific factors. Hence, we propose a novel Domain-Specificity inducing Transformer (DSiT) framework1for disentangling and learning both domain-specific and task-specific factors. To achieve disentanglement, we propose to construct novel Domain-Representative Inputs (DRI) with domain-specific information to train a domain classifier with a novel domain token. We are the first to utilize vision transformers for domain adaptation in a privacy-oriented source-free setting, and our approach achieves state-of-the-art performance on single-source, multi-source, and multi-target benchmarks. Sunandini Sanyal, Ashish Ramayee Asokan, Suvaansh Bhambri, Akshay R. Kulkarni, Jogendra Kundu, Venkatesh Babu Radhakrishnan |
ICCV | 6 |
| 2023 | Feature Reconstruction From Outputs Can Mitigate Simplicity Bias in Neural Networks
Sravanti Addepalli, Anshul Nasery, Venkatesh Babu Radhakrishnan, Praneeth Netrapalli, Prateek Jain 0002 |
ICLR | 3 |
| 2022 | Amplitude Spectrum Transformation for Open Compound Domain Adaptive Semantic SegmentationabstractOpen compound domain adaptation (OCDA) has emerged as a practical adaptation setting which considers a single labeled source domain against a compound of multi-modal unlabeled target data in order to generalize better on novel unseen domains. We hypothesize that an improved disentanglement of domain-related and task-related factors of dense intermediate layer features can greatly aid OCDA. Prior-arts attempt this indirectly by employing adversarial domain discriminators on the spatial CNN output. However, we find that latent features derived from the Fourier-based amplitude spectrum of deep CNN features hold a more tractable mapping with domain discrimination. Motivated by this, we propose a novel feature space Amplitude Spectrum Transformation (AST). During adaptation, we employ the AST auto-encoder for two purposes. First, carefully mined source-target instance pairs undergo a simulation of cross-domain feature stylization (AST-Sim) at a particular layer by altering the AST-latent. Second, AST operating at a later layer is tasked to normalize (AST-Norm) the domain content by fixing its latent to a mean prototype. Our simplified adaptation technique is not only clustering-free but also free from complex adversarial alignment. We achieve leading performance against the prior arts on the OCDA scene segmentation benchmarks. Jogendra Kundu, Akshay R. Kulkarni, Suvaansh Bhambri, Varun Jampani, Venkatesh Babu Radhakrishnan |
AAAI | 5 |
| 2022 | Beyond Learning Features: Training a Fully-Functional Classifier with ZERO Instance-Level Labels
Deepak Babu Sam, Abhinav Agarwalla, Venkatesh Babu Radhakrishnan |
AAAI | 3 |
| 2022 | Uncertainty-Aware Adaptation for Self-Supervised 3D Human Pose EstimationabstractThe advances in monocular 3D human pose estimation are dominated by supervised techniques that require large-scale 2D/3D pose annotations. Such methods often behave erratically in the absence of any provision to discard unfamiliar out-of-distribution data. To this end, we cast the 3D human pose learning as an unsupervised domain adaptation problem. We introduce MRP-Net11Project page: https://sites.google.com/view/mrp-net that constitutes a common deep network backbone with two output heads subscribing to two diverse configurations; a) model-free Joint localization and b) model-based parametric regression. Such a design allows us to derive suitable measures to quantify prediction uncertainty at both pose and Joint level granularity. While supervising only on labeled synthetic samples, the adaptation process aims to minimize the uncertainty for the unlabeled target images while maximizing the same for an extreme out-of-distribution dataset (backgrounds). Alongside synthetic-to-real 3D pose adaptation, the Joint-uncertainties allow expanding the adaptation to work on in-the-wild images even in the presence of occlusion and truncation scenarios. We present a comprehensive evaluation of the proposed approach and demonstrate state-of-the-art performance on benchmark datasets. Jogendra Kundu, Siddharth Seth, Pradyumna YM, Varun Jampani, Anirban Chakraborty 0001, Venkatesh Babu Radhakrishnan |
CVPR | 6 |
| 2022 | Towards Data-Free Model Stealing in a Hard Label SettingabstractMachine learning models deployed as a service (MLaaS) are susceptible to model stealing attacks, where an adversary attempts to steal the model within a restricted access framework. While existing attacks demonstrate near-perfect clone-model performance using softmax predictions of the classification network, most of the APIs allow access to only the top-1 labels. In this work, we show that it is indeed possible to steal Machine Learning models by accessing only top-1 predictions (Hard Label setting) as well, without access to model gradients (Black-Box setting) or even the training dataset (Data-Free setting) within a low query budget. We propose a novel GAN-based framework11Project Page: https://sites.google.com/view/dfms-hl that trains the student and generator in tandem to steal the model effectively while overcoming the challenge of the hard label setting by utilizing gradients of the clone network as a proxy to the victim's gradients. We propose to overcome the large query costs associated with a typical Data-Free setting by utilizing publicly available (potentially unrelated) datasets as a weak image prior. We additionally show that even in the absence of such data, it is possible to achieve state-of-the-art results within a low query budget using synthetically crafted samples. We are the first to demonstrate the scalability of Model Stealing in a restricted access setting on a 100 class dataset as well. Sunandini Sanyal, Sravanti Addepalli, Venkatesh Babu Radhakrishnan |
CVPR | 3 |
| 2022 | Towards Efficient and Effective Self-supervised Learning of Visual Representations
Sravanti Addepalli, Kaushal Bhogale, Priyam Dey, Venkatesh Babu Radhakrishnan |
ECCV (31) | 4 |
| 2022 | Scaling Adversarial Training to Large Perturbation Bounds
Sravanti Addepalli, Samyak Jain, Gaurang Sriramanan, Venkatesh Babu Radhakrishnan |
ECCV (5) | 4 |
| 2022 | Hierarchical Semantic Regularization of Latent Spaces in StyleGANs
Tejan Karmali, Rishubh Parihar, Susmit Agrawal, Harsh Rangwani, Varun Jampani, Maneesh Kumar Singh 0001, Venkatesh Babu Radhakrishnan |
ECCV (15) | 7 |
| 2022 | Concurrent Subsidiary Supervision for Unsupervised Source-Free Domain Adaptation
Jogendra Kundu, Suvaansh Bhambri, Akshay R. Kulkarni, Hiran Sarkar, Varun Jampani, Venkatesh Babu Radhakrishnan |
ECCV (30) | 6 |
| 2022 | Improving GANs for Long-Tailed Data Through Group Spectral Regularization
Harsh Rangwani, Naman Jaswani, Tejan Karmali, Varun Jampani, Venkatesh Babu Radhakrishnan |
ECCV (15) | 5 |
| 2022 | Completely Self-supervised Crowd Counting via Distribution Matching
Deepak Babu Sam, Abhinav Agarwalla, Jimmy Joseph, Vishwanath A. Sindagi, Venkatesh Babu Radhakrishnan, Vishal M. Patel |
ECCV (31) | 5 |
| 2022 | Balancing Discriminability and Transferability for Source-Free Domain AdaptationabstractConventional domain adaptation (DA) techniques aim to improve domain transferability by learning domain-invariant representations; while concurrently preserving the task-discriminability knowledge gathered from the labeled source data. However, the requirement of simultaneous access to labeled source and unlabeled target renders them unsuitable for the challenging source-free DA setting. The trivial solution of realizing an effective original to generic domain mapping improves transferability but degrades task discriminability. Upon analyzing the hurdles from both theoretical and empirical standpoints, we derive novel insights to show that a mixup between original and corresponding translated generic samples enhances the discriminability-transferability trade-off while duly respecting the privacy-oriented source-free setting. A simple but effective realization of the proposed insights on top of the existing source-free DA approaches yields state-of-the-art performance with faster convergence. Beyond single-source, we also outperform multi-source prior-arts across both classification and semantic segmentation benchmarks. Jogendra Kundu, Akshay R. Kulkarni, Suvaansh Bhambri, Deepesh Mehta, Shreyas Anand Kulkarni, Varun Jampani, Venkatesh Babu Radhakrishnan |
ICML | 7 |
| 2022 | A Closer Look at Smoothness in Domain Adversarial TrainingabstractDomain adversarial training has been ubiquitous for achieving invariant representations and is used widely for various domain adaptation tasks. In recent times, methods converging to smooth optima have shown improved generalization for supervised learning tasks like classification. In this work, we analyze the effect of smoothness enhancing formulations on domain adversarial training, the objective of which is a combination of task loss (eg. classification, regression etc.) and adversarial terms. We find that converging to a smooth minima with respect to (w.r.t.) task loss stabilizes the adversarial training leading to better performance on target domain. In contrast to task loss, our analysis shows that converging to smooth minima w.r.t. adversarial loss leads to sub-optimal generalization on the target domain. Based on the analysis, we introduce the Smooth Domain Adversarial Training (SDAT) procedure, which effectively enhances the performance of existing domain adversarial methods for both classification and object detection tasks. Our analysis also provides insight into the extensive usage of SGD over Adam in the community for domain adversarial training. Harsh Rangwani, Sumukh K. Aithal, Arihant Jain, Venkatesh Babu Radhakrishnan |
ICML | 5 |
| 2022 | Everything is There in Latent Space: Attribute Editing and Attribute Style Manipulation by StyleGAN Latent Space ExplorationabstractUnconstrained Image generation with high realism is now possible using recent Generative Adversarial Networks (GANs). However, it is quite challenging to generate images with a given set of attributes. Recent methods use style-based GAN models to perform image editing by leveraging the semantic hierarchy present in the layers of the generator. We present Few-shot Latent-based Attribute Manipulation and Editing (FLAME), a simple yet effective framework to perform highly controlled image editing by latent space manipulation. Specifically, we estimate linear directions in the latent space (of a pre-trained StyleGAN) that controls semantic attributes in the generated image. In contrast to previous methods that either rely on large-scale attribute labeled datasets or attribute classifiers, FLAME uses minimal supervision of a few curated image pairs to estimate disentangled edit directions. FLAME can perform both individual and sequential edits with high precision on a diverse set of images while preserving identity. Further, we propose a novel task of Attribute Style Manipulation to generate diverse styles for attributes such as eyeglass and hair. We first encode a set of synthetic images of the same identity but having different attribute styles in the latent space to estimate an attribute style manifold. Sampling a new latent from this manifold will result in a new attribute style in the generated image. We propose a novel sampling method to sample latent from the manifold, enabling us to generate a diverse set of attribute styles beyond the styles present in the training set. FLAME can generate diverse attribute styles in a disentangled manner. We illustrate the superior performance of FLAME against previous image editing methods by extensive qualitative and quantitative comparisons. FLAME generalizes well on out-of-distribution images from art domain as well as on other datasets such as cars and churches. Rishubh Parihar, Ankit Dhiman, Tejan Karmali, Venkatesh Babu Radhakrishnan |
ACM Multimedia | 4 |
| 2022 | Efficient and Effective Augmentation Strategy for Adversarial TrainingabstractAdversarial training of Deep Neural Networks is known to be significantly more data-hungry when compared to standard training. Furthermore, complex data augmentations such as AutoAugment, which have led to substantial gains in standard training of image classifiers, have not been successful with Adversarial Training. We first explain this contrasting behavior by viewing augmentation during training as a problem of domain generalization, and further propose Diverse Augmentation-based Joint Adversarial Training (DAJAT) to use data augmentations effectively in adversarial training. We aim to handle the conflicting goals of enhancing the diversity of the training dataset and training with data that is close to the test distribution by using a combination of simple and complex augmentations with separate batch normalization layers during training. We further utilize the popular Jensen-Shannon divergence loss to encourage the \emph{joint} learning of the \emph{diverse augmentations}, thereby allowing simple augmentations to guide the learning of complex ones. Lastly, to improve the computational efficiency of the proposed method, we propose and utilize a two-step defense, Ascending Constraint Adversarial Training (ACAT), that uses an increasing epsilon schedule and weight-space smoothing to prevent gradient masking. The proposed method DAJAT achieves substantially better robustness-accuracy trade-off when compared to existing methods on the RobustBench Leaderboard on ResNet-18 and WideResNet-34-10. The code for implementing DAJAT is available here: https://github.com/val-iisc/DAJAT Sravanti Addepalli, Samyak Jain, Venkatesh Babu Radhakrishnan |
NeurIPS | 3 |
| 2022 | Subsidiary Prototype Alignment for Universal Domain AdaptationabstractUniversal Domain Adaptation (UniDA) deals with the problem of knowledge transfer between two datasets with domain-shift as well as category-shift. The goal is to categorize unlabeled target samples, either into one of the "known" categories or into a single "unknown" category. A major problem in UniDA is negative transfer, i.e. misalignment of "known" and "unknown" classes. To this end, we first uncover an intriguing tradeoff between negative-transfer-risk and domain-invariance exhibited at different layers of a deep network. It turns out we can strike a balance between these two metrics at a mid-level layer. Towards designing an effective framework based on this insight, we draw motivation from Bag-of-visual-Words (BoW). Word-prototypes in a BoW-like representation of a mid-level layer would represent lower-level visual primitives that are likely to be unaffected by the category-shift in the high-level features. We develop modifications that encourage learning of word-prototypes followed by word-histogram based classification. Following this, subsidiary prototype-space alignment (SPA) can be seen as a closed-set alignment problem, thereby avoiding negative transfer. We realize this with a novel word-histogram-related pretext task to enable closed-set SPA, operating in conjunction with goal task UniDA. We demonstrate the efficacy of our approach on top of existing UniDA techniques, yielding state-of-the-art performance across three standard UniDA and Open-Set DA object recognition benchmarks. Jogendra Kundu, Suvaansh Bhambri, Akshay R. Kulkarni, Hiran Sarkar, Varun Jampani, Venkatesh Babu Radhakrishnan |
NeurIPS | 6 |
| 2022 | Escaping Saddle Points for Effective Generalization on Class-Imbalanced DataabstractReal-world datasets exhibit imbalances of varying types and degrees. Several techniques based on re-weighting and margin adjustment of loss are often used to enhance the performance of neural networks, particularly on minority classes. In this work, we analyze the class-imbalanced learning problem by examining the loss landscape of neural networks trained with re-weighting and margin based techniques. Specifically, we examine the spectral density of Hessian of class-wise loss, through which we observe that the network weights converges to a saddle point in the loss landscapes of minority classes. Following this observation, we also find that optimization methods designed to escape from saddle points can be effectively used to improve generalization on minority classes. We further theoretically and empirically demonstrate that Sharpness-Aware Minimization (SAM), a recent technique that encourages convergence to a flat minima, can be effectively used to escape saddle points for minority classes. Using SAM results in a 6.2\% increase in accuracy on the minority classes over the state-of-the-art Vector Scaling Loss, leading to an overall average increase of 4\% across imbalanced datasets. The code is available at https://github.com/val-iisc/Saddle-LongTail. Harsh Rangwani, Sumukh K. Aithal, Venkatesh Babu Radhakrishnan |
NeurIPS | 4 |
| 2022 | Cost-Sensitive Self-Training for Optimizing Non-Decomposable MetricsabstractSelf-training based semi-supervised learning algorithms have enabled the learning of highly accurate deep neural networks, using only a fraction of labeled data. However, the majority of work on self-training has focused on the objective of improving accuracy whereas practical machine learning systems can have complex goals (e.g. maximizing the minimum of recall across classes, etc.) that are non-decomposable in nature. In this work, we introduce the Cost-Sensitive Self-Training (CSST) framework which generalizes the self-training-based methods for optimizing non-decomposable metrics. We prove that our framework can better optimize the desired non-decomposable metric utilizing unlabeled data, under similar data distribution assumptions made for the analysis of self-training. Using the proposed CSST framework, we obtain practical self-training methods (for both vision and NLP tasks) for optimizing different non-decomposable metrics using deep neural networks. Our results demonstrate that CSST achieves an improvement over the state-of-the-art in majority of the cases across datasets and objectives. Harsh Rangwani, Shrinivas Ramasubramanian, Sho Takemori, Kato Takashi, Yuhei Umeda, Venkatesh Babu Radhakrishnan |
NeurIPS | 6 |
| 2022 | LEAD: Self-Supervised Landmark Estimation by Aligning Distributions of Feature SimilarityabstractIn this work, we introduce LEAD, an approach to dis-cover landmarks from an unannotated collection of category-specific images. Existing works in self-supervised landmark detection are based on learning dense (pixel-level) feature representations from an image, which are further used to learn landmarks in a semi-supervised manner. While there have been advances in self-supervised learning of image features for instance-level tasks like classification, these methods do not ensure dense equivariant representations. The property of equivariance is of interest for dense prediction tasks like landmark estimation. In this work, we introduce an approach to enhance the learning of dense equivariant representations in a self-supervised fashion. We follow a two-stage training approach: first, we train a network using the BYOL [13] objective which operates at an instance level. The correspondences obtained through this network are further used to train a dense and compact representation of the image using a lightweight network. We show that having such a prior in the feature extractor helps in landmark detection, even under drastically limited number of annotations while also improving generalization across scale variations. Tejan Karmali, Abhinav Atrishi, Sai Sree Harsha, Susmit Agrawal, Varun Jampani, Venkatesh Babu Radhakrishnan |
WACV | 6 |
| 2021 | Few-Shot Domain Adaptation for Low Light RAW Image Enhancement
K. Ram Prabhakar, Vishal Vinod, Nihar R. Sahoo, Venkatesh Babu Radhakrishnan |
BMVC | 4 |
| 2021 | Labeled From Unlabeled: Exploiting Unlabeled Data for Few-Shot Deep HDR DeghostingabstractHigh Dynamic Range (HDR) deghosting is an indispensable tool in capturing wide dynamic range scenes without ghosting artifacts. Recently, convolutional neural networks (CNNs) have shown tremendous success in HDR deghosting. However, CNN-based HDR deghosting methods require collecting large datasets with ground truth, which is a tedious and time-consuming process. This paper proposes a pioneering work by introducing zero and few-shot learning strategies for data-efficient HDR deghosting. Our approach consists of two stages of training. In stage one, we train the model with few labeled (5 or less) dynamic samples and a pool of unlabeled samples with a self-supervised loss. We use the trained model to predict HDRs for the unlabeled samples. To derive data for the next stage of training, we propose a novel method for generating corresponding dynamic inputs from the predicted HDRs of unlabeled data. The generated artificial dynamic inputs and predicted HDRs are used as paired labeled data. In stage two, we finetune the model with the original few labeled data and artificially generated labeled data. Our few-shot approach outperforms many fully-supervised methods in two publicly available datasets, using as little as five labeled dynamic samples. K. Ram Prabhakar, Gowtham Senthil, Susmit Agrawal, Venkatesh Babu Radhakrishnan, Rama Krishna Sai S. Gorthi |
CVPR | 4 |
| 2021 | Generalize then Adapt: Source-Free Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation (DA) has gained substantial interest in semantic segmentation. However, almost all prior arts assume concurrent access to both labeled source and unlabeled target, making them unsuitable for scenarios demanding source-free adaptation. In this work1, we enable source-free DA by partitioning the task into two: a) source-only domain generalization and b) source-free target adaptation. Towards the former, we provide theoretical insights to develop a multi-head framework trained with a virtually extended multi-source dataset, aiming to balance generalization and specificity. Towards the latter, we utilize the multi-head framework to extract reliable target pseudo-labels for self-training. Additionally, we introduce a novel conditional prior-enforcing auto-encoder that discourages spatial irregularities, thereby enhancing the pseudo-label quality. Experiments on the standard GTA5→Cityscapes and SYNTHIA→Cityscapes benchmarks show our superiority even against the non-source-free prior-arts. Further, we show our compatibility with online adaptation enabling deployment in a sequentially changing environment. Jogendra Kundu, Akshay R. Kulkarni, Varun Jampani, Venkatesh Babu Radhakrishnan |
ICCV | 5 |
| 2021 | S3VAADA: Submodular Subset Selection for Virtual Adversarial Active Domain AdaptationabstractUnsupervised domain adaptation (DA) methods have focused on achieving maximal performance through aligning features from source and target domains without using labeled data in the target domain. Whereas, in the real-world scenario’s it might be feasible to get labels for a small proportion of target data. In these scenarios, it is important to select maximally-informative samples to label and find an effective way to combine them with the existing knowledge from source data. Towards achieving this, we propose S3VAADA which i) introduces a novel submodular criterion to select a maximally informative subset to label and ii) enhances a cluster-based DA procedure through novel improvements to effectively utilize all the available data for improving generalization on target. Our approach consistently outperforms the competing state-of-the-art approaches on datasets with varying degrees of domain shifts. The project page with additional details is available here: https://sites.google.com/iisc.ac.in/s3vaada-iccv2021/. Harsh Rangwani, Arihant Jain, Sumukh K. Aithal, Venkatesh Babu Radhakrishnan |
ICCV | 4 |
| 2021 | Deep Implicit Surface Point Prediction NetworksabstractDeep neural representations of 3D shapes as implicit functions have been shown to produce high fidelity models surpassing the resolution-memory trade-off faced by the explicit representations using meshes and point clouds. However, most such approaches focus on representing closed shapes. Unsigned distance function (UDF) based approaches have been proposed recently as a promising alternative to represent both open and closed shapes. However, since the gradients of UDFs vanish on the surface, it is challenging to estimate local (differential) geometric properties like the normals and tangent planes which are needed for many downstream applications in vision and graphics. There are additional challenges in computing these properties efficiently with a low-memory footprint. This paper presents a novel approach that models such surfaces using a new class of implicit representations called the closest surface-point (CSP) representation. We show that CSP allows us to represent complex surfaces of any topology (open or closed) with high fidelity. It also allows for accurate and efficient computation of local geometric properties. We further demonstrate that it leads to efficient implementation of downstream algorithms like sphere-tracing for rendering the 3D surface as well as to create explicit mesh-based representations. Extensive experimental evaluation on the ShapeNet dataset validate the above contributions with results surpassing the state-of-the-art. Code and data are available at https://sites.google.com/view/cspnet. Rahul Venkatesh, Tejan Karmali, Sarthak Sharma, Aurobrata Ghosh, Venkatesh Babu Radhakrishnan, László A. Jeni, Maneesh Kumar Singh 0001 |
ICCV | 5 |
| 2021 | Non-local Latent Relation Distillation for Self-Adaptive 3D Human Pose EstimationabstractAvailable 3D human pose estimation approaches leverage different forms of strong (2D/3D pose) or weak (multi-view or depth) paired supervision. Barring synthetic or in-studio domains, acquiring such supervision for each new target environment is highly inconvenient. To this end, we cast 3D pose learning as a self-supervised adaptation problem that aims to transfer the task knowledge from a labeled source domain to a completely unpaired target. We propose to infer image-to-pose via two explicit mappings viz. image-to-latent and latent-to-pose where the latter is a pre-learned decoder obtained from a prior-enforcing generative adversarial auto-encoder. Next, we introduce relation distillation as a means to align the unpaired cross-modal samples i.e., the unpaired target videos and unpaired 3D pose sequences. To this end, we propose a new set of non-local relations in order to characterize long-range latent pose interactions, unlike general contrastive relations where positive couplings are limited to a local neighborhood structure. Further, we provide an objective way to quantify non-localness in order to select the most effective relation set. We evaluate different self-adaptation settings and demonstrate state-of-the-art 3D human pose estimation performance on standard benchmarks. Jogendra Kundu, Siddharth Seth, Anirudh Jamkhandi, Pradyumna YM, Varun Jampani, Anirban Chakraborty 0001, Venkatesh Babu Radhakrishnan |
NeurIPS | 7 |
| 2021 | Aligning Silhouette Topology for Self-Adaptive 3D Human Pose RecoveryabstractArticulation-centric 2D/3D pose supervision forms the core training objective in most existing 3D human pose estimation techniques. Except for synthetic source environments, acquiring such rich supervision for each real target domain at deployment is highly inconvenient. However, we realize that standard foreground silhouette estimation techniques (on static camera feeds) remain unaffected by domain-shifts. Motivated by this, we propose a novel target adaptation framework that relies only on silhouette supervision to adapt a source-trained model-based regressor. However, in the absence of any auxiliary cue (multi-view, depth, or 2D pose), an isolated silhouette loss fails to provide a reliable pose-specific gradient and requires to be employed in tandem with a topology-centric loss. To this end, we develop a series of convolution-friendly spatial transformations in order to disentangle a topological-skeleton representation from the raw silhouette. Such a design paves the way to devise a Chamfer-inspired spatial topological-alignment loss via distance field computation, while effectively avoiding any gradient hindering spatial-to-pointset mapping. Experimental results demonstrate our superiority against prior-arts in self-adapting a source trained model to diverse unlabeled target domains, such as a) in-the-wild datasets, b) low-resolution image domains, and c) adversarially perturbed image domains (via UAP). Mugalodi Rakesh, Jogendra Kundu, Varun Jampani, Venkatesh Babu Radhakrishnan |
NeurIPS | 4 |
| 2021 | Towards Efficient and Effective Adversarial TrainingabstractThe vulnerability of Deep Neural Networks to adversarial attacks has spurred immense interest towards improving their robustness. However, present state-of-the-art adversarial defenses involve the use of 10-step adversaries during training, which renders them computationally infeasible for application to large-scale datasets. While the recent single-step defenses show promising direction, their robustness is not on par with multi-step training methods. In this work, we bridge this performance gap by introducing a novel Nuclear-Norm regularizer on network predictions to enforce function smoothing in the vicinity of data samples. While prior works consider each data sample independently, the proposed regularizer uses the joint statistics of adversarial samples across a training minibatch to enhance optimization during both attack generation and training, obtaining state-of-the-art results amongst efficient defenses. We achieve further gains by incorporating exponential averaging of network weights over training iterations. We finally introduce a Hybrid training approach that combines the effectiveness of a two-step variant of the proposed defense with the efficiency of a single-step defense. We demonstrate superior results when compared to multi-step defenses such as TRADES and PGD-AT as well, at a significantly lower computational cost. Gaurang Sriramanan, Sravanti Addepalli, Arya Baburaj, Venkatesh Babu Radhakrishnan |
NeurIPS | 4 |
| 2021 | Class balancing GAN with a classifier in the loopabstractGenerative Adversarial Networks (GANs) have swiftly evolved to imitate increasingly complex image distributions. However, majority of the developments focus on performance of GANs on balanced datasets. We find that the existing GANs and their training regimes which work well on balanced datasets fail to be effective in case of imbalanced (i.e. long-tailed) datasets. In this work we introduce a novel theoretically motivated Class Balancing regularizer for training GANs. Our regularizer makes use of the knowledge from a pre-trained classifier to ensure balanced learning of all the classes in the dataset. This is achieved via modelling the effective class frequency based on the exponential forgetting observed in neural networks and encouraging the GAN to focus on underrepresented classes. We demonstrate the utility of our regularizer in learning representations for long-tailed distributions via achieving better performance than existing approaches over multiple datasets. Specifically, when applied to an unconditional GAN, it improves the FID from $13.03$ to $9.01$ on the long-tailed iNaturalist-$2019$ dataset. Harsh Rangwani, Konda Reddy Mopuri, Venkatesh Babu Radhakrishnan |
UAI | 3 |
| 2021 | Locate, Size, and Count: Accurately Resolving People in Dense Crowds via DetectionabstractWe introduce a detection framework for dense crowd counting and eliminate the need for the prevalent density regression paradigm. Typical counting models predict crowd density for an image as opposed to detecting every person. These regression methods, in general, fail to localize persons accurate enough for most applications other than counting. Hence, we adopt an architecture that locates every person in the crowd, sizes the spotted heads with bounding box and then counts them. Compared to normal object or face detectors, there exist certain unique challenges in designing such a detection system. Some of them are direct consequences of the huge diversity in dense crowds along with the need to predict boxes contiguously. We solve these issues and develop our LSC-CNN model, which can reliably detect heads of people across sparse to dense crowds. LSC-CNN employs a multi-column architecture with top-down feature modulation to better resolve persons and produce refined predictions at multiple resolutions. Interestingly, the proposed training regime requires only point head annotation, but can estimate approximate size information of heads. We show that LSC-CNN not only has superior localization than existing density regressors, but outperforms in counting as well. The code for our approach is available at https://github.com/val-iisc/lsc-cnn. Deepak Babu Sam, Skand Vishwanath Peri, Mukuntha N. S. 0001, Amogh Kamath, Venkatesh Babu Radhakrishnan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | DeGAN: Data-Enriching GAN for Retrieving Representative Samples from a Trained ClassifierabstractIn this era of digital information explosion, an abundance of data from numerous modalities is being generated as well as archived everyday. However, most problems associated with training Deep Neural Networks still revolve around lack of data that is rich enough for a given task. Data is required not only for training an initial model, but also for future learning tasks such as Model Compression and Incremental Learning. A diverse dataset may be used for training an initial model, but it may not be feasible to store it throughout the product life cycle due to data privacy issues or memory constraints. We propose to bridge the gap between the abundance of available data and lack of relevant data, for the future learning tasks of a given trained network. We use the available data, that may be an imbalanced subset of the original training dataset, or a related domain dataset, to retrieve representative samples from a trained classifier, using a novel Data-enriching GAN (DeGAN) framework. We demonstrate that data from a related domain can be leveraged to achieve state-of-the-art performance for the tasks of Data-free Knowledge Distillation and Incremental Learning on benchmark datasets. We further demonstrate that our proposed framework can enrich any data, even from unrelated domains, to make it more useful for the future learning tasks of a given network. Sravanti Addepalli, Gaurav Kumar Nayak, Anirban Chakraborty 0001, Venkatesh Babu Radhakrishnan |
AAAI | 4 |
| 2020 | Kinematic-Structure-Preserved Representation for Unsupervised 3D Human Pose EstimationabstractEstimation of 3D human pose from monocular image has gained considerable attention, as a key step to several human-centric applications. However, generalizability of human pose estimation models developed using supervision on large-scale in-studio datasets remains questionable, as these models often perform unsatisfactorily on unseen in-the-wild environments. Though weakly-supervised models have been proposed to address this shortcoming, performance of such models relies on availability of paired supervision on some related task, such as 2D pose or multi-view image pairs. In contrast, we propose a novel kinematic-structure-preserved unsupervised 3D pose estimation framework, which is not restrained by any paired or unpaired weak supervisions. Our pose estimation framework relies on a minimal set of prior knowledge that defines the underlying kinematic 3D structure, such as skeletal joint connectivity information with bone-length ratios in a fixed canonical scale. The proposed model employs three consecutive differentiable transformations namely forward-kinematics, camera-projection and spatial-map transformation. This design not only acts as a suitable bottleneck stimulating effective pose disentanglement, but also yields interpretable latent pose representations avoiding training of an explicit latent embedding to pose mapper. Furthermore, devoid of unstable adversarial setup, we re-utilize the decoder to formalize an energy-based loss, which enables us to learn from in-the-wild videos, beyond laboratory settings. Comprehensive experiments demonstrate our state-of-the-art unsupervised and weakly-supervised pose estimation performance on both Human3.6M and MPI-INF-3DHP datasets. Qualitative results on unseen environments further establish our superior generalization ability. Jogendra Kundu, Siddharth Seth, Rahul M. V., Mugalodi Rakesh, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001 |
AAAI | 5 |
| 2020 | WAMDA: Weighted Alignment of Sources for Multi-source Domain Adaptation
Surbhi Aggarwal, Jogendra Kundu, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001 |
BMVC | 3 |
| 2020 | Towards Achieving Adversarial Robustness by Enforcing Feature Consistency Across Bit PlanesabstractAs humans, we inherently perceive images based on their predominant features, and ignore noise embedded within lower bit planes. On the contrary, Deep Neural Networks are known to confidently misclassify images corrupted with meticulously crafted perturbations that are nearly imperceptible to the human eye. In this work, we attempt to address this problem by training networks to form coarse impressions based on the information in higher bit planes, and use the lower bit planes only to refine their prediction. We demonstrate that, by imposing consistency on the representations learned across differently quantized images, the adversarial robustness of networks improves significantly when compared to a normally trained model. Present state-of-the-art defenses against adversarial attacks require the networks to be explicitly trained using adversarial samples that are computationally expensive to generate. While such methods that use adversarial training continue to achieve the best results, this work paves the way towards achieving robustness without having to explicitly train on adversarial samples. The proposed approach is therefore faster, and also closer to the natural learning process in humans. Sravanti Addepalli, Vivek B. S., Arya Baburaj, Gaurang Sriramanan, Venkatesh Babu Radhakrishnan |
CVPR | 5 |
| 2020 | Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image SynthesisabstractCamera captured human pose is an outcome of several sources of variation. Performance of supervised 3D pose estimation approaches comes at the cost of dispensing with variations, such as shape and appearance, that may be useful for solving other related tasks. As a result, the learned model not only inculcates task-bias but also dataset-bias because of its strong reliance on the annotated samples, which also holds true for weakly-supervised models. Acknowledging this, we propose a self-supervised learning framework to disentangle such variations from unlabeled video frames. We leverage the prior knowledge on human skeleton and poses in the form of a single part-based 2D puppet model, human pose articulation constraints, and a set of unpaired 3D poses. Our differentiable formalization, bridging the representation gap between the 3D pose and spatial part maps, not only facilitates discovery of interpretable pose disentanglement, but also allows us to operate on videos with diverse camera movements. Qualitative results on unseen in-the-wild datasets establish our superior generalization across multiple tasks beyond the primary tasks of 3D pose estimation and part segmentation. Furthermore, we demonstrate state-of-the-art weakly-supervised 3D pose estimation performance on both Human3.6M and MPI-INF-3DHP datasets. Jogendra Kundu, Siddharth Seth, Varun Jampani, Mugalodi Rakesh, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001 |
CVPR | 5 |
| 2020 | Towards Inheritable Models for Open-Set Domain AdaptationabstractThere has been a tremendous progress in Domain Adaptation (DA) for visual recognition tasks. Particularly, open-set DA has gained considerable attention wherein the target domain contains additional unseen categories. Existing open-set DA approaches demand access to a labeled source dataset along with unlabeled target instances. However, this reliance on co-existing source and target data is highly impractical in scenarios where data-sharing is restricted due to its proprietary nature or privacy concerns. Addressing this, we introduce a practical DA paradigm where a source-trained model is used to facilitate adaptation in the absence of the source dataset in future. To this end, we formalize knowledge inheritability as a novel concept and propose a simple yet effective solution to realize inheritable models suitable for the above practical paradigm. Further, we present an objective way to quantify inheritability to enable the selection of the most suitable source model for a given target domain, even in the absence of the source data. We provide theoretical insights followed by a thorough empirical evaluation demonstrating state-of-the-art open-set domain adaptation performance. Jogendra Kundu, Naveen Venkat, Ambareesh Revanur, Rahul M. V., Venkatesh Babu Radhakrishnan |
CVPR | 5 |
| 2020 | Universal Source-Free Domain AdaptationabstractThere is a strong incentive to develop versatile learning techniques that can transfer the knowledge of class-separability from a labeled source domain to an unlabeled target domain in the presence of a domain-shift. Existing domain adaptation (DA) approaches are not equipped for practical DA scenarios as a result of their reliance on the knowledge of source-target label-set relationship (e.g. Closed-set, Open-set or Partial DA). Furthermore, almost all prior unsupervised DA works require coexistence of source and target samples even during deployment, making them unsuitable for real-time adaptation. Devoid of such impractical assumptions, we propose a novel two-stage learning process. 1) In the Procurement stage, we aim to equip the model for future source-free deployment, assuming no prior knowledge of the upcoming category-gap and domain-shift. To achieve this, we enhance the model's ability to reject out-of-source distribution samples by leveraging the available source data, in a novel generative classifier framework. 2) In the Deployment stage, the goal is to design a unified adaptation algorithm capable of operating across a wide range of category-gaps, with no access to the previously seen source samples. To this end, in contrast to the usage of complex adversarial training regimes, we define a simple yet effective source-free adaptation objective by utilizing a novel instance-level weighting mechanism, named as Source Similarity Metric (SSM). A thorough evaluation shows the practical usability of the proposed learning framework with superior DA performance even over state-of-the-art source-dependent approaches. Jogendra Kundu, Naveen Venkat, Rahul M. V., Venkatesh Babu Radhakrishnan |
CVPR | 4 |
| 2020 | From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksabstractReconstructing 3D models from 2D images is one of the fundamental problems in computer vision. In this work, we propose a deep learning technique for 3D object reconstruction from a single image. Contrary to recent works that either use 3D supervision or multi-view supervision, we use only single view images with no pose information during training as well. This makes our approach more practical requiring only an image collection of an object category and the corresponding silhouettes. We learn both 3D point cloud reconstruction and pose estimation networks in a self-supervised manner, making use of differentiable point cloud renderer to train with 2D supervision. A key novelty of the proposed technique is to impose 3D geometric reasoning into predicted 3D point clouds by rotating them with randomly sampled poses and then enforcing cycle consistency on both 3D reconstructions and poses. In addition, using single-view supervision allows us to do test-time optimization on a given test image. Experiments on the synthetic ShapeNet and real-world Pix3D datasets demonstrate that our approach, despite using less supervision, can achieve competitive performance compared to pose-supervised and multi-view supervised approaches. Navaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung, Varun Jampani, Venkatesh Babu Radhakrishnan |
CVPR | 6 |
| 2020 | Single-Step Adversarial Training With Dropout SchedulingabstractDeep learning models have shown impressive performance across a spectrum of computer vision applications including medical diagnosis and autonomous driving. One of the major concerns that these models face is their susceptibility to adversarial attacks. Realizing the importance of this issue, more researchers are working towards developing robust models that are less affected by adversarial attacks. Adversarial training method shows promising results in this direction. In adversarial training regime, models are trained with mini-batches augmented with adversarial samples. Fast and simple methods (e.g., single-step gradient ascent) are used for generating adversarial samples, in order to reduce computational complexity. It is shown that models trained using single-step adversarial training method (adversarial samples are generated using non-iterative method) are pseudo robust. Further, this pseudo robustness of models is attributed to the gradient masking effect. However, existing works fail to explain when and why gradient masking effect occurs during single-step adversarial training. In this work, (i) we show that models trained using single-step adversarial training method learn to prevent the generation of single-step adversaries, and this is due to over-fitting of the model during the initial stages of training, and (ii) to mitigate this effect, we propose a single-step adversarial training method with dropout scheduling. Unlike models trained using existing single-step adversarial training methods, models trained using the proposed single-step adversarial training method are robust against both single-step and multi-step adversarial attacks, and the performance is on par with models trained using computationally expensive multi-step adversarial training methods, in white-box and black-box settings. Vivek B. S., Venkatesh Babu Radhakrishnan |
CVPR | 2 |
| 2020 | Appearance Consensus Driven Self-supervised Human Mesh Recovery
Jogendra Kundu, Mugalodi Rakesh, Varun Jampani, Rahul M. V., Venkatesh Babu Radhakrishnan |
ECCV (1) | 5 |
| 2020 | Unsupervised Cross-Modal Alignment for Multi-person 3D Pose Estimation
Jogendra Kundu, Ambareesh Revanur, Govind Vitthal Waghmare, Rahul M. V., Venkatesh Babu Radhakrishnan |
ECCV (13) | 5 |
| 2020 | Class-Incremental Domain Adaptation
Jogendra Kundu, Rahul M. V., Naveen Venkat, Ambareesh Revanur, Venkatesh Babu Radhakrishnan |
ECCV (13) | 5 |
| 2020 | Towards Practical and Efficient High-Resolution HDR Deghosting with CNN
K. Ram Prabhakar, Susmit Agrawal, Durgesh Kumar Singh, Balraj Ashwath, Venkatesh Babu Radhakrishnan |
ECCV (21) | 5 |
| 2020 | Learning to Count in the Crowd from Limited Labeled Data
Vishwanath A. Sindagi, Rajeev Yasarla, Deepak Babu Sam, Venkatesh Babu Radhakrishnan, Vishal M. Patel |
ECCV (11) | 4 |
| 2020 | Saliency-Driven Class Impressions For Feature Visualization Of Deep Neural NetworksabstractIn this paper, we propose a data-free method of extracting Impressions of each class from the classifier's memory. The Deep Learning regime empowers classifiers to extract distinct patterns (or features) of a given class from training data, which is the basis on which they generalize to unseen data. Before deploying these models on critical applications, it is very useful to visualize the features considered to be important for classification. Existing visualization methods develop high confidence images consisting of both background and foreground features. This makes it hard to judge what the important features of a given class are. In this work, we propose a saliency-driven approach to visualize discriminative features that are considered most important for a given task. Another drawback of existing methods is that, confidence of the generated visualizations is increased by creating multiple instances of the given class. We restrict the algorithm to develop a single object per image, which helps further in extracting features of high confidence, and also results in better visualizations. We further demonstrate the generation of negative images as naturally fused images of two or more classes. Our code is available at: https://giChub.com/val-iisc/Saliency-driven-Class-Impressions. Sravanti Addepalli, Dipesh Tamboli, Venkatesh Babu Radhakrishnan, Biplab Banerjee |
ICIP | 3 |
| 2020 | CDNet++: Improved Change Detection with Deep Neural Network Feature CorrelationabstractIn this paper, we present a deep convolutional neural network (CNN) architecture for segmenting semantic changes between two images. The main objective is to segment changes at the semantic level than detecting background changes, which are irrelevant to the application. The difficulties include seasonal changes, lighting differences, artifacts due to alignment and occlusion. The existing approaches fail to address all the problems together; thus, none of them achieve state-of-the-art performance in three publicly available change detection datasets: VL-CMU-CD [1], TSUNAMI [2] and GSV [2]. Our proposed approach is a simple yet effective method that can handle even adverse challenges. In our approach, we leverage the correlation between high-level abstract CNN features to segment the changes. Compared with several traditional and other deep learning-based change detection methods, our proposed method achieves state-of-the-art performance in all three datasets. K. Ram Prabhakar, Akshaya Ramaswamy, Suvaansh Bhambri, Jayavardhana Gubbi, Venkatesh Babu Radhakrishnan, P. Balamuralidhar |
IJCNN | 5 |
| 2020 | Guided Adversarial Attack for Evaluating and Enhancing Adversarial DefensesabstractAdvances in the development of adversarial attacks have been fundamental to the progress of adversarial defense research. Efficient and effective attacks are crucial for reliable evaluation of defenses, and also for developing robust models. Adversarial attacks are often generated by maximizing standard losses such as the cross-entropy loss or maximum-margin loss within a constraint set using Projected Gradient Descent (PGD). In this work, we introduce a relaxation term to the standard loss, that finds more suitable gradient-directions, increases attack efficacy and leads to more efficient adversarial training. We propose Guided Adversarial Margin Attack (GAMA), which utilizes function mapping of the clean image to guide the generation of adversaries, thereby resulting in stronger attacks. We evaluate our attack against multiple defenses and show improved performance when compared to existing attacks. Further, we propose Guided Adversarial Training (GAT), which achieves state-of-the-art performance amongst single-step defenses by utilizing the proposed relaxation term for both attack generation and training. Gaurang Sriramanan, Sravanti Addepalli, Arya Baburaj, Venkatesh Babu Radhakrishnan |
NeurIPS | 4 |
| 2020 | Your Classifier can Secretly Suffice Multi-Source Domain AdaptationabstractMulti-Source Domain Adaptation (MSDA) deals with the transfer of task knowledge from multiple labeled source domains to an unlabeled target domain, under a domain-shift. Existing methods aim to minimize this domain-shift using auxiliary distribution alignment objectives. In this work, we present a different perspective to MSDA wherein deep models are observed to implicitly align the domains under label supervision. Thus, we aim to utilize implicit alignment without additional training objectives to perform adaptation. To this end, we use pseudo-labeled target samples and enforce a classifier agreement on the pseudo-labels, a process called Self-supervised Implicit Alignment (SImpAl). We find that SImpAl readily works even under category-shift among the source domains. Further, we propose classifier agreement as a cue to determine the training convergence, resulting in a simple training algorithm. We provide a thorough evaluation of our approach on five benchmarks, along with detailed insights into each component of our approach. Naveen Venkat, Jogendra Kundu, Durgesh Kumar Singh, Ambareesh Revanur, Venkatesh Babu Radhakrishnan |
NeurIPS | 5 |
| 2020 | Text-based Person Search via Attribute-aided MatchingabstractText-based person search aims to retrieve the pedestrian images that best match a given text query. Existing methods utilize class-id information to get discriminative and identity-preserving features. However, it is not well-explored whether it is beneficial to explicitly ensure that the semantics of the data are retained. In the proposed work, we aim to create semantics-preserving embeddings through an additional task of attribute prediction. Since attribute annotation is typically unavailable in text-based person search, we first mine them from the text corpus. These attributes are then used as a means to bridge the modality gap between the image-text inputs, as well as to improve the representation learning. In summary, we propose an approach for text-based person search by learning an attribute-driven space along with a class-information driven space, and utilize both for obtaining the retrieval results. Our experiments on benchmark dataset, CUHK-PEDES, show that learning the attribute-space not only helps in improving performance, giving us state-of-the-art Rank-1 accuracy of 56.68%, but also yields humanly-interpretable features. Surbhi Aggarwal, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001 |
WACV | 2 |
| 2020 | Cross-Conditioned Recurrent Networks for Long-Term Synthesis of Inter-Person Human Motion InteractionsabstractModeling dynamics of human motion is one of the most challenging sequence modeling problem, with diverse applications in animation industry, human-robot interaction, motion-based surveillance, etc. Available attempts to use auto-regressive techniques for long-term single-person motion generation usually fails, resulting in stagnated motion or divergence to unrealistic pose patterns. In this paper, we propose a novel cross-conditioned recurrent framework targeting long-term synthesis of inter-person interactions beyond several minutes. We carefully integrate positive implications of both auto-regressive and encoder-decoder recurrent architecture, by interchangeably utilizing two separate fixed-length cross person motion prediction models for long-term generation in a novel hierarchical fashion. As opposed to prior approaches, we guarantee structural plausibility of 3D pose by training the recurrent model to regress latent representation of a separately trained generative pose embedding network. Different variants of the proposed frameworks are evaluated through extensive experiments on SBU-interaction, CMU-MoCAP and an inhouse collection of duet-dance dataset. Qualitative and quantitative evaluation on several tasks, such as Short-term motion prediction, Long-term motion synthesis and Interaction-based motion retrieval against prior state-of-the-art approaches clearly highlight superiority of the proposed framework. Jogendra Kundu, Himanshu Buckchash, Priyanka Mandikal, Rahul M. V., Anirudh Jamkhandi, Venkatesh Babu Radhakrishnan |
WACV | 6 |
| 2020 | VRT-Net: Real-Time Scene Parsing via Variable Resolution TransformabstractUrban scene parsing is a basic requirement for various autonomous navigation systems especially in self-driving. Most of the available approaches employ generic image parsing architectures designed for segmentation of object focused scene captured in indoor setups. However, images captured in car-mounted cameras exhibit an extreme effect of perspective geometry, causing a significant scale disparity between near and farther objects. Recognizing this, we formalize a unique Variable Resolution Transform (VRT) technique motivated from the foveal magnification in human eye. Following this, we design a Fovea Estimation Network (FEN) which is trained to estimate a single most convenient fixation location along with the associated magnification factor, best suited for a given input image. The proposed framework is designed to enable its usage as a wrapper over the available real-time scene parsing models, thereby demonstrating a superior trade-off between speed and quality as compared to the prior state-of-the-arts. Jogendra Kundu, Gaurav Singh Rajput, Venkatesh Babu Radhakrishnan |
WACV | 3 |
| 2020 | Unsupervised Cross-Dataset Adaptation via Probabilistic Amodal 3D Human Pose CompletionabstractDespite remarkable success of supervised deep learning models for 3D human pose estimation, performance of such models is mostly limited to constrained laboratory settings. Such models not only exhibit an alarming level of dataset bias, but also fail to operate on unconstrained videos in the presence of external variations such as camera motion, partial body visibility, occlusion, etc. Acknowledging these shortcomings, firstly, we aim to formalize a motion representation learning framework by effectively utilizing both constrained and artificially generated unconstrained video samples for datasets with 3D pose annotation. Without ignoring the inherent uncertainty in pose estimation for the truncated video frames, we devise a novel probabilistic amodal pose completion framework to enable generation of multiple plausible pose-filling outcomes. Secondly, to address dataset bias, the probabilistic amodal framework is reutilized to design novel self-supervised objectives. This not only enables adaptation of the model to target unannotated datasets (wild YouTube videos) but also encourages learning of generic motion representations beyond the available supervised data even in unconstrained scenarios. Such a training regime helps us achieve state-of-the art performance on unsupervised cross-dataset pose estimation, with a significant improvement in partially-visible unconstrained scenarios. Jogendra Kundu, Rahul M. V., Jay Patravali, Venkatesh Babu Radhakrishnan |
WACV | 4 |
| 2020 | Going Beyond the Regression Paradigm with Accurate Dot Prediction for Dense CrowdsabstractWe present an alternative to the paradigm of density regression widely being employed for tackling crowd counting. In the prevalent regression approach, a model is trained for mapping images to its crowd density rather than counting by detecting every person. This framework is motivated from the difficulty to discriminate humans in highly dense crowds where unfavorable perspective, occlusion and clutter are prevalent. Though regression methods estimate overall crowd counts pretty well, localization of individual persons suffers and varies considerably across the entire density spectrum. Moreover, individual detection of people aids more explainable practical systems than predicting blind crowd count or density map. Hence, we move away from density regression and reformulate the task as localized dot prediction in dense crowds. Our dot detection model, DD-CNN, is trained for pixel-wise binary classification to detect people instead of regressing local crowd density. In order to handle severe scale variation and detect people of all scales with accurate dots, we use a novel multi-scale architecture which does not require any ground truth scale information. This training regime, which incorporates top-down feedback, helps our model to localize people in sparse as well as dense crowds. Our model delivers superior counting performance on major crowd datasets. We also evaluate on some additional metrics and evidence superior localization of the dot detection formulation. Deepak Babu Sam, Skand Vishwanath Peri, Mukuntha N. S. 0001, Venkatesh Babu Radhakrishnan |
WACV | 4 |
| 2020 | PerSeg : segmenting salient objects from bag of single image perturbations
Avishek Majumder 0001, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001 |
Multim. Tools Appl. | 2 |
| 2020 | Pictionary-Style Word Guessing on Hand-Drawn Object Sketches: Dataset, Analysis and Deep Network ModelsabstractThe ability of intelligent agents to play games in human-like fashion is popularly considered a benchmark of progress in Artificial Intelligence. In our work, we introduce the first computational model aimed at Pictionary, the popular word-guessing social game. We first introduce Sketch-QA, a guessing task. Styled after Pictionary, Sketch-QA uses incrementally accumulated sketch stroke sequences as visual data. Sketch-QA involves asking a fixed question ("What object is being drawn?") and gathering open-ended guess-words from human guessers. We analyze the resulting dataset and present many interesting findings therein. To mimic Pictionary-style guessing, we propose a deep neural model which generates guess-words in response to temporally evolving human-drawn object sketches. Our model even makes human-like mistakes while guessing, thus amplifying the human mimicry factor. We evaluate our model on the large-scale guess-word dataset generated via Sketch-QA task and compare with various baselines. We also conduct a Visual Turing Test to obtain human impressions of the guess-words generated by humans and our model. Experimental results demonstrate the promise of our approach for Pictionary and similarly themed games. Ravi Kiran Sarvadevabhatla, Shiv Surya, Trisha Mittal, Venkatesh Babu Radhakrishnan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Operator-in-the-Loop Deep Sequential Multi-Camera Feature Fusion for Person Re-IdentificationabstractGiven a target image as query, person re-identification systems retrieve a ranked list of candidate matches on a per-camera basis. In deployed systems, a human operator scans these lists and labels sighted targets by touch or mouse-based selection. However, classical re-id approaches generate per-camera lists independently. Therefore, target identifications by operator in a subset of cameras cannot be utilized to improve ranking of the target in remaining set of network cameras. To address this shortcoming, we propose a novel sequential multi-camera re-id approach. The proposed approach can accommodate human operator inputs and provides early gains via a monotonic improvement in target ranking. At the heart of our approach is a fusion function which operates on deep feature representations of query and candidate matches. We formulate an optimization procedure custom-designed to incrementally improve query representation. Since existing evaluation methods cannot be directly adopted to our setting, we also propose two novel evaluation protocols. The results on two large-scale re-id datasets (Market-1501, DukeMTMC-reID) demonstrate that our multi-camera method significantly outperforms baselines and other popular feature fusion schemes. Additionally, we conduct a comparative subject-based study of human operator performance. The superior operator performance enabled by our approach makes a compelling case for its integration into deployable video-surveillance systems. Navaneet K. L., Ravi Kiran Sarvadevabhatla, Shashank Shekhar 0006, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | BiHMP-GAN: Bidirectional 3D Human Motion Prediction GANabstractHuman motion prediction model has applications in various fields of computer vision. Without taking into account the inherent stochasticity in the prediction of future pose dynamics, such methods often converges to a deterministic undesired mean of multiple probable outcomes. Devoid of this, we propose a novel probabilistic generative approach called Bidirectional Human motion prediction GAN, or BiHMP-GAN. To be able to generate multiple probable human-pose sequences, conditioned on a given starting sequence, we introduce a random extrinsic factor r, drawn from a predefined prior distribution. Furthermore, to enforce a direct content loss on the predicted motion sequence and also to avoid mode-collapse, a novel bidirectional framework is incorporated by modifying the usual discriminator architecture. The discriminator is trained also to regress this extrinsic factor r, which is used alongside with the intrinsic factor (encoded starting pose sequence) to generate a particular pose sequence. To further regularize the training, we introduce a novel recursive prediction strategy. In spite of being in a probabilistic framework, the enhanced discriminator architecture allows predictions of an intermediate part of pose sequence to be used as a conditioning for prediction of the latter part of the sequence. The bidirectional setup also provides a new direction to evaluate the prediction quality against a given test sequence. For a fair assessment of BiHMP-GAN, we report performance of the generated motion sequence using (i) a critic model trained to discriminate between real and fake motion sequence, and (ii) an action classifier trained on real human motion dynamics. Outcomes of both qualitative and quantitative evaluations, on the probabilistic generations of the model, demonstrate the superiority of BiHMP-GAN over previously available methods. Jogendra Kundu, Maharshi Gor, Venkatesh Babu Radhakrishnan |
AAAI | 3 |
| 2019 | CAPNet: Continuous Approximation Projection for 3D Point Cloud Reconstruction Using 2D SupervisionabstractKnowledge of 3D properties of objects is a necessity in order to build effective computer vision systems. However, lack of large scale 3D datasets can be a major constraint for datadriven approaches in learning such properties. We consider the task of single image 3D point cloud reconstruction, and aim to utilize multiple foreground masks as our supervisory data to alleviate the need for large scale 3D datasets. A novel differentiable projection module, called ‘CAPNet’, is introduced to obtain such 2D masks from a predicted 3D point cloud. The key idea is to model the projections as a continuous approximation of the points in the point cloud. To overcome the challenges of sparse projection maps, we propose a loss formulation termed ‘affinity loss’ to generate outlierfree reconstructions. We significantly outperform the existing projection based approaches on a large-scale synthetic dataset. We show the utility and generalizability of such a 2D supervised approach through experiments on a real-world dataset, where lack of 3D data can be a serious concern. To further enhance the reconstructions, we also propose a test stage optimization procedure to obtain reconstructions that display high correspondence with the observed input image. Navaneet K. L., Priyanka Mandikal, Mayank Agarwal, Venkatesh Babu Radhakrishnan |
AAAI | 4 |
| 2019 | Almost Unsupervised Learning for Dense Crowd CountingabstractWe present an unsupervised learning method for dense crowd count estimation. Marred by large variability in appearance of people and extreme overlap in crowds, enumerating people proves to be a difficult task even for humans. This implies creating large-scale annotated crowd data is expensive and directly takes a toll on the performance of existing CNN based counting models on account of small datasets. Motivated by these challenges, we develop Grid Winner-Take-All (GWTA) autoencoder to learn several layers of useful filters from unlabeled crowd images. Our GWTA approach divides a convolution layer spatially into a grid of cells. Within each cell, only the maximally activated neuron is allowed to update the filter. Almost 99.9% of the parameters of the proposed model are trained without any labeled data while the rest 0.1% are tuned with supervision. The model achieves superior results compared to other unsupervised methods and stays reasonably close to the accuracy of supervised baseline. Furthermore, we present comparisons and analyses regarding the quality of learned features across various models. Deepak Babu Sam, Neeraj N. Sajjan, Himanshu Maurya, Venkatesh Babu Radhakrishnan |
AAAI | 4 |
| 2019 | All for One: Frame-wise Rank Loss for Improving Video-based Person Re-identificationabstractPerson re-identification involves retrieving correct matches for a target image (query) from a set of gallery images, while video based re-identification extends this to the case of query and gallery videos. Typical video-based re-id methods ignore the temporal evolution of the intermediate representations of the video sequences. We propose a novel loss function, termed rank loss, to explicitly ensure that the learnt representations achieve enhanced performance and robustness as the sequence progresses and that better intermediate representations result in an improved final representation. Experiments indicate that the addition of rank loss indeed helps in improving the re-id performance while achieving performance comparable to state-of-the-art approaches. Navaneet K. L., Vasudha Todi, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001 |
ICASSP | 3 |
| 2019 | A Fast, Scalable, and Reliable Deghosting Method for Extreme Exposure FusionabstractHDR fusion of extreme exposure images with complex camera and object motion is a challenging task. Existing patch-based optimization techniques generate noisy and/or blurry results with undesirable artifacts for difficult scenarios. Additionally, they are computationally intensive and have high execution times. Recently proposed CNN-based methods offer fast alternatives, but still fail to generate artifact-free results for extreme exposure images. Furthermore, they do not scale to an arbitrary number of input images. To address these issues, we propose a simple, yet effective CNN-based multi-exposure image fusion method that produces artifact-free HDR images. Our method is fast, and scales to an arbitrary number of input images. Additionally, we prepare a large dataset of 582 varying exposure images with corresponding deghosted HDR images to train our model. We test the efficacy of our algorithm on publicly available datasets, and achieve significant improvements over existing state-of-the-art methods. Through experimental results, we demonstrate that our method produces artifact-free results, and offers a speed-up of around 54× over existing state-of-the-art HDR fusion methods. K. Ram Prabhakar, Rajat Arora 0002, Adhitya Swaminathan, Kunal Pratap Singh, Venkatesh Babu Radhakrishnan |
ICCP | 5 |
| 2019 | FDA: Feature Disruptive Attack
Aditya Ganeshan, Vivek B. S., Venkatesh Babu Radhakrishnan |
ICCV | 3 |
| 2019 | GAN-Tree: An Incrementally Learned Hierarchical Generative Framework for Multi-Modal Data DistributionsabstractDespite the remarkable success of generative adversarial networks, their performance seems less impressive for diverse training sets, requiring learning of discontinuous mapping functions. Though multi-mode prior or multi-generator models have been proposed to alleviate this problem, such approaches may fail depending on the empirically chosen initial mode components. In contrast to such bottom-up approaches, we present GAN-Tree, which follows a hierarchical divisive strategy to address such discontinuous multi-modal data. Devoid of any assumption on the number of modes, GAN-Tree utilizes a novel mode-splitting algorithm to effectively split the parent mode to semantically cohesive children modes, facilitating unsupervised clustering. Further, it also enables incremental addition of new data modes to an already trained GAN-Tree, by updating only a single branch of the tree structure. As compared to prior approaches, the proposed framework offers a higher degree of flexibility in choosing a large variety of mutually exclusive and exhaustive tree nodes called GAN-Set. Extensive experiments on synthetic and natural image datasets including ImageNet demonstrate the superiority of GAN-Tree against the prior state-of-the-art. Jogendra Kundu, Maharshi Gor, Dakshit Agrawal, Venkatesh Babu Radhakrishnan |
ICCV | 4 |
| 2019 | UM-Adapt: Unsupervised Multi-Task Adaptation Using Adversarial Cross-Task DistillationabstractAiming towards human-level generalization, there is a need to explore adaptable representation learning methods with greater transferability. Most existing approaches independently address task-transferability and cross-domain adaptation, resulting in limited generalization. In this paper, we propose UM-Adapt - a unified framework to effectively perform unsupervised domain adaptation for spatially-structured prediction tasks, simultaneously maintaining a balanced performance across individual tasks in a multi-task setting. To realize this, we propose two novel regularization strategies; a) Contour-based content regularization (CCR) and b) exploitation of inter-task coherency using a cross-task distillation module. Furthermore, avoiding a conventional ad-hoc domain discriminator, we re-utilize the cross-task distillation loss as output of an energy function to adversarially minimize the input domain discrepancy. Through extensive experiments, we demonstrate superior generalizability of the learned representations simultaneously for multiple tasks under domain-shifts from synthetic to natural environments. UM-Adapt yields state-of-the-art transfer learning results on ImageNet classification and comparable performance on PASCAL VOC 2007 detection task, even with a smaller backbone-net. Moreover, the resulting semi-supervised framework outperforms the current fully-supervised multi-task learning state-of-the-art on both NYUD and Cityscapes dataset. Jogendra Kundu, Nishank Lakkakula, Venkatesh Babu Radhakrishnan |
ICCV | 3 |
| 2019 | Efficient Person Re-Identification in Videos Using Sequence Lazy Greedy Determinantal Point Process (SLGDPP)abstractGiven a sequence of observations for each person in each camera, identifying or re-identifying the same person across different cameras is one of the objectives of video surveillance systems. In the case of video based person re-id, the challenge is to handle the high correlation between temporally adjacent frames. The presence of non-informative frames results in high redundancy which needs to be removed for an efficient re-id. We propose a novel method to handle this challenge using Determinantal Point Process (DPP) to select the most diverse and informative subset of frames from a given sequence. Since subset selection problem is NP-Hard, we propose to use an approximate solution called Lazy Greedy DPP (LGDPP) and further extend it to utilize the temporal information of sequences with our proposed Sequential LGDPP (SLGDPP) for video-based person re-id. The major advantages of the proposed DPP variants are their simplicity and plug and play nature, which make it possible to use them atop any pretrained re-id model followed by a feature fusion module. The effectiveness of proposed frameworks is demonstrated on two popular video re-id benchmark datasets through improvements over state-of-the-art methods and naive baseline sampling methods. Gaurav Kumar Nayak, Utkarsh Shreemali, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001 |
ICIP | 3 |
| 2019 | Zero-Shot Knowledge Distillation in Deep NetworksabstractKnowledge distillation deals with the problem of training a smaller model (Student) from a high capacity source model (Teacher) so as to retain most of its performance. Existing approaches use either the training data or meta-data extracted from it in order to train the Student. However, accessing the dataset on which the Teacher has been trained may not always be feasible if the dataset is very large or it poses privacy or safety concerns (e.g., bio-metric or medical data). Hence, in this paper, we propose a novel data-free method to train the Student from the Teacher. Without even using any meta-data, we synthesize the Data Impressions from the complex Teacher model and utilize these as surrogates for the original training data samples to transfer its learning to Student via knowledge distillation. We, therefore, dub our method “Zero-Shot Knowledge Distillation" and demonstrate that our framework results in competitive generalization performance as achieved by distillation using the actual training data samples on multiple benchmark datasets. Gaurav Kumar Nayak, Konda Reddy Mopuri, Vaisakh Shaj, Venkatesh Babu Radhakrishnan, Anirban Chakraborty 0001 |
ICML | 4 |
| 2019 | Unsupervised Feature Learning of Human Actions As Trajectories in Pose Embedding ManifoldabstractAn unsupervised human action modeling framework can provide useful pose-sequence representation, which can be utilized in a variety of pose analysis applications. In this work we propose a novel temporal pose-sequence modeling framework, which can embed the dynamics of 3D human-skeleton joints to a latent space in an efficient manner. In contrast to an end-to-end framework explored by previous works, we disentangle the task of individual pose representation learning from the task of learning actions as a sequence of pose embeddings. In order to realize a continuous pose embedding manifold along with better reconstructions, we propose an unsupervised, manifold learning procedure named Encoder GAN, (or EnGAN). Further we use the pose embeddings generated by EnGAN to model human actions using an RNN auto-encoder architecture, PoseRNN. We introduce first-order gradient loss to explicitly enforce temporal regularity in the predicted motion sequence. A hierarchical feature fusion technique is also investigated for simultaneous modeling of local skeleton joints along with global pose variations. We demonstrate state-of-the-art transfer-ability of the learned representation against other supervisedly and unsupervisedly learned motion embeddings for the task of fine-grained action recognition on SBU interaction dataset. Further, we show the qualitative strengths of the proposed framework by visualizing skeleton pose reconstructions and interpolations in pose-embedding space, and low dimensional principal component projections of the reconstructed pose trajectories. Jogendra Kundu, Maharshi Gor, Phani Krishna Uppala, Venkatesh Babu Radhakrishnan |
WACV | 4 |
| 2019 | Dense 3D Point Cloud Reconstruction Using a Deep Pyramid NetworkabstractReconstructing a high-resolution 3D model of an object is a challenging task in computer vision. Designing scalable and light-weight architectures is crucial while addressing this problem. Existing point-cloud based reconstruction approaches directly predict the entire point cloud in a single stage. Although this technique can handle low-resolution point clouds, it is not a viable solution for generating dense, high-resolution outputs. In this work, we introduce DensePCR, a deep pyramidal network for point cloud reconstruction that hierarchically predicts point clouds of increasing resolution. Towards this end, we propose an architecture that first predicts a low-resolution point cloud, and then hierarchically increases the resolution by aggregating local and global point features to deform a grid. Our method generates point clouds that are accurate, uniform and dense. Through extensive quantitative and qualitative evaluation on synthetic and real datasets, we demonstrate that DensePCR outperforms the existing state-of-the-art point cloud reconstruction works, while also providing a light-weight and scalable architecture for predicting high-resolution outputs. Priyanka Mandikal, Venkatesh Babu Radhakrishnan |
WACV | 2 |
| 2019 | Generalizable Data-Free Objective for Crafting Universal Adversarial PerturbationsabstractMachine learning models are susceptible to adversarial perturbations: small changes to input that can cause large changes in output. It is also demonstrated that there exist input-agnostic perturbations, called universal adversarial perturbations, which can change the inference of target model on most of the data samples. However, existing methods to craft universal perturbations are (i) task specific, (ii) require samples from the training data distribution, and (iii) perform complex optimizations. Additionally, because of the data dependence, fooling ability of the crafted perturbations is proportional to the available training data. In this paper, we present a novel, generalizable and data-free approach for crafting universal adversarial perturbations. Independent of the underlying task, our objective achieves fooling via corrupting the extracted features at multiple layers. Therefore, the proposed objective is generalizable to craft image-agnostic perturbations across multiple vision tasks such as object recognition, semantic segmentation, and depth estimation. In the practical setting of black-box attack scenario (when the attacker does not have access to the target model and it's training data), we show that our objective outperforms the data dependent objectives to fool the learned models. Further, via exploiting simple priors related to the data distribution, our objective remarkably boosts the fooling ability of the crafted perturbations. Significant fooling rates achieved by our objective emphasize that the current deep learning models are now at an increased risk, since our objective generalizes across multiple tasks without the requirement of training data for crafting the perturbations. To encourage reproducible research, we have released the codes for our proposed algorithm.1. Konda Reddy Mopuri, Aditya Ganeshan, Venkatesh Babu Radhakrishnan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | CNN Fixations: An Unraveling Approach to Visualize the Discriminative Image RegionsabstractDeep convolutional neural networks (CNNs) have revolutionized the computer vision research and have seen unprecedented adoption for multiple tasks, such as classification, detection, and caption generation. However, they offer little transparency into their inner workings and are often treated as black boxes that deliver excellent performance. In this paper, we aim at alleviating this opaqueness of CNNs by providing visual explanations for the network's predictions. Our approach can analyze a variety of CNN-based models trained for computer vision applications, such as object recognition and caption generation. Unlike the existing methods, we achieve this via unraveling the forward pass operation. The proposed method exploits feature dependencies across the layer hierarchy and uncovers the discriminative image locations that guide the network's predictions. We name these locations CNN fixations, loosely analogous to human eye fixations. Our approach is a generic method that requires no architectural changes, additional training, or gradient computation, and computes the important image locations (CNN fixations). We demonstrate through a variety of applications that our approach is able to localize the discriminative image locations across different network architectures, diverse vision tasks, and data modalities. Konda Reddy Mopuri, Utsav Garg, Venkatesh Babu Radhakrishnan |
IEEE Trans. Image Process. | 3 |
| 2018 | Top-Down Feedback for Crowd Counting Convolutional Neural NetworkabstractCounting people in dense crowds is a demanding task even for humans. This is primarily due to the large variability in appearance of people. Often people are only seen as a bunch of blobs. Occlusions, pose variations and background clutter further compound the difficulty. In this scenario, identifying a person requires larger spatial context and semantics of the scene. But the current state-of-the-art CNN regressors for crowd counting are feedforward and use only limited spatial context to detect people. They look for local crowd patterns to regress the crowd density map, resulting in false predictions. Hence, we propose top-down feedback to correct the initial prediction of the CNN. Our architecture consists of a bottom-up CNN along with a separate top-down CNN to generate feedback. The bottom-up network, which regresses the crowd density map, has two columns of CNN with different receptive fields. Features from various layers of the bottom-up CNN are fed to the top-down network. The feedback, thus generated, is applied on the lower layers of the bottom-up network in the form of multiplicative gating. This masking weighs activations of the bottom-up network at spatial as well as feature levels to correct the density prediction. We evaluate the performance of our model on all major crowd datasets and show the effectiveness of top-down feedback. Deepak Babu Sam, Venkatesh Babu Radhakrishnan |
AAAI | 2 |
| 2018 | Game of Sketches: Deep Recurrent Models of Pictionary-Style Word GuessingabstractThe ability of machine-based agents to play games in human-like fashion is considered a benchmark of progress in AI. In this paper, we introduce the first computational model aimed at Pictionary, the popular word-guessing social game. We first introduce Sketch-QA, an elementary version of Visual Question Answering task. Styled after Pictionary, Sketch-QA uses incrementally accumulated sketch stroke sequences as visual data. Notably, Sketch-QA involves asking a fixed question ("What object is being drawn?") and gathering open-ended guess-words from human guessers. To mimic Pictionary-style guessing, we propose a deep neural model which generates guess-words in response to temporally evolving human-drawn sketches. Our model even makes human-like mistakes while guessing, thus amplifying the human mimicry factor. We evaluate our model on the large-scale guess-word dataset generated via Sketch-QA task and compare with various baselines. We also conduct a Visual Turing Test to obtain human impressions of the guess-words generated by humans and our model. Experimental results demonstrate the promise of our approach for Pictionary and similarly themed games. Ravi Kiran Sarvadevabhatla, Shiv Surya, Trisha Mittal, Venkatesh Babu Radhakrishnan |
AAAI | 4 |
| 2018 | 3D-LMNet: Latent Embedding Matching for Accurate and Diverse 3D Point Cloud Reconstruction from a Single Image
Priyanka Mandikal, Navaneet K. L., Mayank Agarwal, Venkatesh Babu Radhakrishnan |
BMVC | 4 |
| 2018 | AdaDepth: Unsupervised Content Congruent Adaptation for Depth EstimationabstractSupervised deep learning methods have shown promising results for the task of monocular depth estimation; but acquiring ground truth is costly, and prone to noise as well as inaccuracies. While synthetic datasets have been used to circumvent above problems, the resultant models do not generalize well to natural scenes due to the inherent domain shift. Recent adversarial approaches for domain adaption have performed well in mitigating the differences between the source and target domains. But these methods are mostly limited to a classification setup and do not scale well for fully-convolutional architectures. In this work, we propose AdaDepth - an unsupervised domain adaptation strategy for the pixel-wise regression task of monocular depth estimation. The proposed approach is devoid of above limitations through a) adversarial learning and b) explicit imposition of content consistency on the adapted target representation. Our unsupervised approach performs competitively with other established approaches on depth estimation tasks and achieves state-of-the-art results in a semi-supervised setting. Jogendra Kundu, Phani Krishna Uppala, Anuj Pahuja, Venkatesh Babu Radhakrishnan |
CVPR | 4 |
| 2018 | NAG: Network for Adversary GenerationabstractAdversarial perturbations can pose a serious threat for deploying machine learning systems. Recent works have shown existence of image-agnostic perturbations that can fool classifiers over most natural images. Existing methods present optimization approaches that solve for a fooling objective with an imperceptibility constraint to craft the perturbations. However, for a given classifier, they generate one perturbation at a time, which is a single instance from the manifold of adversarial perturbations. Also, in order to build robust models, it is essential to explore the manifold of adversarial perturbations. In this paper, we propose for the first time, a generative approach to model the distribution of adversarial perturbations. The architecture of the proposed model is inspired from that of GANs and is trained using fooling and diversity objectives. Our trained generator network attempts to capture the distribution of adversarial perturbations for a given classifier and readily generates a wide variety of such perturbations. Our experimental evaluation demonstrates that perturbations crafted by our model (i) achieve state-of-the-art fooling rates, (ii) exhibit wide variety and (iii) deliver excellent cross model generalizability. Our work can be deemed as an important step in the process of inferring about the complex manifolds of adversarial perturbations. Konda Reddy Mopuri, Utkarsh Ojha, Utsav Garg, Venkatesh Babu Radhakrishnan |
CVPR | 4 |
| 2018 | Divide and Grow: Capturing Huge Diversity in Crowd Images With Incrementally Growing CNNabstractAutomated counting of people in crowd images is a challenging task. The major difficulty stems from the large diversity in the way people appear in crowds. In fact, features available for crowd discrimination largely depend on the crowd density to the extent that people are only seen as blobs in a highly dense scene. We tackle this problem with a growing CNN which can progressively increase its capacity to account for the wide variability seen in crowd scenes. Our model starts from a base CNN density regressor, which is trained in equivalence on all types of crowd images. In order to adapt with the huge diversity, we create two child regressors which are exact copies of the base CNN. A differential training procedure divides the dataset into two clusters and fine-tunes the child networks on their respective specialties. Consequently, without any hand-crafted criteria for forming specialties, the child regressors become experts on certain types of crowds. The child networks are again split recursively, creating two experts at every division. This hierarchical training leads to a CNN tree, where the child regressors are more fine experts than any of their parents. The leaf nodes are taken as the final experts and a classifier network is then trained to predict the correct specialty for a given test image patch. The proposed model achieves higher count accuracy on major crowd datasets. Further, we analyse the characteristics of specialties mined automatically by our method. Deepak Babu Sam, Neeraj N. Sajjan, Venkatesh Babu Radhakrishnan, Mukundhan Srinivasan |
CVPR | 3 |
| 2018 | Ask, Acquire, and Attack: Data-Free UAP Generation Using Class Impressions
Konda Reddy Mopuri, Phani Krishna Uppala, Venkatesh Babu Radhakrishnan |
ECCV (9) | 3 |
| 2018 | Gray-Box Adversarial Training
Vivek B. S., Konda Reddy Mopuri, Venkatesh Babu Radhakrishnan |
ECCV (15) | 3 |
| 2018 | iSPA-Net: Iterative Semantic Pose Alignment NetworkabstractUnderstanding and extracting 3D information of objects from monocular 2D images is a fundamental problem in computer vision. In the task of 3D object pose estimation, recent data driven deep neural network based approaches suffer from scarcity of real images with 3D keypoint and pose annotations. Drawing inspiration from human cognition, where the annotators use a 3D CAD model as structural reference to acquire ground-truth viewpoints for real images; we propose an iterative Semantic Pose Alignment Network, called iSPA-Net. Our approach focuses on exploiting semantic 3D structural regularity to solve the task of fine-grained pose estimation by predicting viewpoint difference between a given pair of images. Such image comparison based approach also alleviates the problem of data scarcity and hence enhances scalability of the proposed approach for novel object categories with minimal annotation. The fine-grained object pose estimator is also aided by correspondence of learned spatial descriptor of the input image pair. The proposed pose alignment framework enjoys the faculty to refine its initial pose estimation in consecutive iterations by utilizing an online rendering setup along with effectiveness of a non-uniform bin classification of pose-difference. This enables iSPA-Net to achieve state-of-the-art performance on various real image viewpoint estimation datasets. Further, we demonstrate effectiveness of the approach for multiple applications. First, we show results for active object viewpoint localization to capture images from similar pose considering only a single image as pose reference. Second, we demonstrate the ability of the learned semantic correspondence to perform unsupervised part-segmentation transfer using only a single part-annotated 3D template model per object class. To encourage reproducible research, we have released the codes for our proposed algorithm. Jogendra Kundu, Aditya Ganeshan, Rahul M. V., Aditya Prakash 0008, Venkatesh Babu Radhakrishnan |
ACM Multimedia | 5 |
| 2017 | Fast Feature Fool: A data independent approach to universal adversarial perturbations
Konda Reddy Mopuri, Utsav Garg, Venkatesh Babu Radhakrishnan |
BMVC | 3 |
| 2017 | DeLiGAN: Generative Adversarial Networks for Diverse and Limited DataabstractA class of recent approaches for generating images, called Generative Adversarial Networks (GAN), have been used to generate impressively realistic images of objects, bedrooms, handwritten digits and a variety of other image modalities. However, typical GAN-based approaches require large amounts of training data to capture the diversity across the image modality. In this paper, we propose DeLiGAN - a novel GAN-based architecture for diverse and limited training data scenarios. In our approach, we reparameterize the latent generative space as a mixture model and learn the mixture models parameters along with those of GAN. This seemingly simple modification to the GAN framework is surprisingly effective and results in models which enable diversity in generated samples although trained with limited data. In our work, we show that DeLiGAN can generate images of handwritten digits, objects and hand-drawn sketches, all using limited amounts of data. To quantitatively characterize intra-class diversity of generated samples, we also introduce a modified version of inception-score, a measure which has been found to correlate well with human assessment of generated samples. Swaminathan Gurumurthy, Ravi Kiran Sarvadevabhatla, Venkatesh Babu Radhakrishnan |
CVPR | 3 |
| 2017 | Switching Convolutional Neural Network for Crowd Counting
Deepak Babu Sam, Shiv Surya, Venkatesh Babu Radhakrishnan |
CVPR | 3 |
| 2017 | DeepFuse: A Deep Unsupervised Approach for Exposure Fusion with Extreme Exposure Image PairsabstractWe present a novel deep learning architecture for fusing static multi-exposure images. Current multi-exposure fusion (MEF) approaches use hand-crafted features to fuse input sequence. However, the weak hand-crafted representations are not robust to varying input conditions. Moreover, they perform poorly for extreme exposure image pairs. Thus, it is highly desirable to have a method that is robust to varying input conditions and capable of handling extreme exposure without artifacts. Deep representations have known to be robust to input conditions and have shown phenomenal performance in a supervised setting. However, the stumbling block in using deep learning for MEF was the lack of sufficient training data and an oracle to provide the ground-truth for supervision. To address the above issues, we have gathered a large dataset of multi-exposure image stacks for training and to circumvent the need for ground truth images, we propose an unsupervised deep learning framework for MEF utilizing a no-reference quality metric as loss function. The proposed approach uses a novel CNN architecture trained to learn the fusion operation without reference ground truth image. The model fuses a set of common low level features extracted from each image to generate artifact-free perceptually pleasing results. We perform extensive quantitative and qualitative evaluation and show that the proposed technique outperforms existing state-of-the-art approaches for a variety of natural images. K. Ram Prabhakar, V. Sai Srikar, Venkatesh Babu Radhakrishnan |
ICCV | 3 |
| 2017 | SketchParse: Towards Rich Descriptions for Poorly Drawn Sketches using Multi-Task Hierarchical Deep NetworksabstractThe ability to semantically interpret hand-drawn line sketches, although very challenging, can pave way for novel applications in multimedia. We propose SKETCHPARSE, the first deep-network architecture for fully automatic parsing of freehand object sketches. SKETCHPARSE is configured as a two-level fully convolutional network. The first level contains shared layers common to all object categories. The second level contains a number of expert sub-networks. Each expert specializes in parsing sketches from object categories which contain structurally similar parts. Effectively, the two-level configuration enables our architecture to scale up efficiently as additional categories are added. We introduce a router layer which (i) relays sketch features from shared layers to the correct expert (ii) eliminates the need to manually specify object category during inference. To bypass laborious part-level annotation, we sketchify photos from semantic object-part image datasets and use them for training. Our architecture also incorporates object pose prediction as a novel auxiliary task which boosts overall performance while providing supplementary information regarding the sketch. We demonstrate SKETCHPARSE's abilities (i) on two challenging large-scale sketch datasets (ii) in parsing unseen, semantically related object categories (iii) in improving fine-grained sketch-based image retrieval. As a novel application, we also outline how SKETCHPARSE's output can be used to generate caption-style descriptions for hand-drawn sketches. Ravi Kiran Sarvadevabhatla, Isht Dwivedi, Abhijat Biswas, Sahil Manocha, Venkatesh Babu Radhakrishnan |
ACM Multimedia | 5 |
| 2017 | Anomaly detection via short local trajectories
Sovan Biswas, Venkatesh Babu Radhakrishnan |
Neurocomputing | 2 |
| 2017 | Evaluating Multiexposure Fusion Using Image InformationabstractMultiexposure fusion (MEF) refers to image fusion methods that capture high dynamic range natural scenes from a set of low dynamic range camera images. In this letter, we study the problem of designing quality assessment (QA) algorithms to estimate the perceptual quality of images generated by different MEF algorithms. We develop our quality index by evaluating individual quality maps between a given fused test image and individual over-/underexposed images at multiple scales and orientations and then combining these maps across the over-/underexposed images. Our approach works on the premise that the true undistorted reference is contained across the over-/underexposed source images. We identify this true reference based on the notion of perceived image information using natural scene statistical models. It is shown that our approach outperforms the state of the art QA algorithms in terms of correlation with human perception of quality on a publicly available MEF database. Hisham Rahman, Rajiv Soundararajan, Venkatesh Babu Radhakrishnan |
IEEE Signal Process. Lett. | 3 |
| 2017 | DeepFix: A Fully Convolutional Neural Network for Predicting Human Eye FixationsabstractUnderstanding and predicting the human visual attention mechanism is an active area of research in the fields of neuroscience and computer vision. In this paper, we propose DeepFix, a fully convolutional neural network, which models the bottom-up mechanism of visual attention via saliency prediction. Unlike classical works, which characterize the saliency map using various hand-crafted features, our model automatically learns features in a hierarchical fashion and predicts the saliency map in an end-to-end manner. DeepFix is designed to capture semantics at multiple scales while taking global context into account, by using network layers with very large receptive fields. Generally, fully convolutional nets are spatially invariant-this prevents them from modeling location-dependent patterns (e.g., centre-bias). Our network handles this by incorporating a novel location-biased convolutional layer. We evaluate our model on multiple challenging saliency data sets and show that it achieves the state-of-the-art results. Srinivas S. Kruthiventi, Kumar Ayush, Venkatesh Babu Radhakrishnan |
IEEE Trans. Image Process. | 3 |
| 2017 | Object Category Understanding via Eye Fixations on Freehand SketchesabstractThe study of eye gaze fixations on photographic images is an active research area. In contrast, the image sub-category of freehand sketches has not received as much attention for such studies. In this paper, we analyze the results of a free-viewing gaze fixation study conducted on 3904 freehand sketches distributed across 160 object categories. Our analysis shows that fixation sequences exhibit marked consistency within a sketch, across sketches of a category and even across suitably grouped sets of categories. This multi-level consistency is remarkable given the variability in depiction and extreme image content sparsity that characterizes hand-drawn object sketches. In this paper, we show that the multi-level consistency in the fixation data can be exploited to 1) predict a test sketch's category given only its fixation sequence and 2) build a computational model which predicts part-labels underlying fixations on objects. We hope that our findings motivate the community to deem sketch-like representations worthy of gaze-based studies vis-a-vis photographic images. Ravi Kiran Sarvadevabhatla, Sudharshan Suresh, Venkatesh Babu Radhakrishnan |
IEEE Trans. Image Process. | 3 |
| 2016 | 'Part'ly First Among Equals: Semantic Part-Based Benchmarking for State-of-the-Art Object Recognition Systems
Ravi Kiran Sarvadevabhatla, Shanthakumar Venkatraman, Venkatesh Babu Radhakrishnan |
ACCV (5) | 3 |
| 2016 | Learning Neural Network Architectures using Backpropagation
Suraj Srinivas, Venkatesh Babu Radhakrishnan |
BMVC | 2 |
| 2016 | Saliency Unified: A Deep Architecture for simultaneous Eye Fixation Prediction and Salient Object SegmentationabstractHuman eye fixations often correlate with locations of salient objects in the scene. However, only a handful of approaches have attempted to simultaneously address the related aspects of eye fixations and object saliency. In this work, we propose a deep convolutional neural network (CNN) capable of predicting eye fixations and segmenting salient objects in a unified framework. We design the initial network layers, shared between both the tasks, such that they capture the object level semantics and the global contextual aspects of saliency, while the deeper layers of the network address task specific aspects. In addition, our network captures saliency at multiple scales via inception-style convolution blocks. Our network shows a significant improvement over the current state-of-the-art for both eye fixation prediction and salient object segmentation across a number of challenging datasets. Srinivas S. Kruthiventi, Vennela Gudisa, Jaley H. Dholakiya, Venkatesh Babu Radhakrishnan |
CVPR | 4 |
| 2016 | Ghosting-free multi-exposure image fusion in gradient domainabstractThis paper presents an algorithm to produce ghosting-free High Dynamic Range (HDR) image by fusing set of multiple exposed images in gradient domain. Recently proposed Gradient domain based exposure fusion method provides high quality result but the scope of which is limited to static camera without foreground object motion. The presence of moving objects/hand shake produces a set of misaligned images. The result of gradient domain approach on misaligned images suffers from ghosting artifacts. In order to produce better HDR image without image registration, we propose to create an aligned image set from input image set by photometric calibration. The gradient of aligned image set is then used to reconstruct the fused final image. The proposed algorithm tested on several publicly available dynamic image sets shows that resultant HDR image is ghosting-free and well exposed. Additionally, the proposed method is fast and thus can be used in consumer appliances such as mobile phones, portable devices with digital cameras. K. Ram Prabhakar, Venkatesh Babu Radhakrishnan |
ICASSP | 2 |
| 2016 | Object saliency using a background priorabstractWhat makes an object salient? Almost all the works so far determine object saliency based on the amount of the contrast of a patch or super pixel with it's surrounding. Due to this approach, objects consisting of multiple colors, which is usually the case with a majority of natural objects, are allocated varying saliency values. Hence, post-processing for it's application is another problem. Taking note of this and keeping in mind the ease of extension to different applications, we provide a new perspective to this problem. We propose a simple yet powerful method for modelling the background for salient object detection. As a corollary of "Rule of Thirds" we model the background as the most occurring super pixels lying along the image border. Saliency is determined based on the distance of other super pixels from the background super pixels. Comparison of the proposed approach with the state of the art shows how our approach can provide more consistent saliency values throughout an object. Chintak Sheth, Venkatesh Babu Radhakrishnan |
ICASSP | 2 |
| 2016 | CrowdNet: A Deep Convolutional Network for Dense Crowd CountingabstractOur work proposes a novel deep learning framework for estimating crowd density from static images of highly dense crowds. We use a combination of deep and shallow, fully convolutional networks to predict the density map for a given crowd image. Such a combination is used for effectively capturing both the high-level semantic information (face/body detectors) and the low-level features (blob detectors), that are necessary for crowd counting under large scale variations. As most crowd datasets have limited training samples (<100 images) and deep learning based approaches require large amounts of training data, we perform multi-scale data augmentation. Augmenting the training samples in such a manner helps in guiding the CNN to learn scale invariant representations. Our method is tested on the challenging UCF_CC_50 dataset, and shown to outperform the state of the art methods. Lokesh Boominathan, Srinivas S. Kruthiventi, Venkatesh Babu Radhakrishnan |
ACM Multimedia | 3 |
| 2016 | Analyzing Structural Characteristics of Object Category Representations From Their Semantic-part DistributionsabstractStudies from neuroscience show that part-mapping computations are employed by human visual system in the process of object recognition. In this paper, we present an approach for analyzing semantic-part characteristics of object category representations. For our experiments, we use category-epitome, a recently proposed sketch-based spatial representation for objects. To enable part-importance analysis, we first obtain semantic-part annotations of hand-drawn sketches originally used to construct the epitomes. We then examine the extent to which the semantic-parts are present in the epitomes of a category and visualize the relative importance of parts as a word cloud. Finally, we show how such word cloud visualizations provide an intuitive understanding of category-level structural trends that exist in the category-epitome object representations. Our method is general in applicability and can also be used to analyze part-based visual object representations for other depiction methods such as photographic images. Ravi Kiran Sarvadevabhatla, Venkatesh Babu Radhakrishnan |
ACM Multimedia | 2 |
| 2016 | SwiDeN: Convolutional Neural Networks For Depiction Invariant Object RecognitionabstractCurrent state of the art object recognition architectures achieve impressive performance but are typically specialized for a single depictive style (e.g. photos only, sketches only). In this paper, we present SwiDeN: our Convolutional Neural Network (CNN) architecture which recognizes objects regardless of how they are visually depicted (line drawing, realistic shaded drawing, photograph etc.). In SwiDeN, we utilize a novel `deep' depictive style-based switching mechanism which appropriately addresses the depiction-specific and depiction-invariant aspects of the problem. We compare SwiDeN with alternative architectures and prior work on a 50-category Photo-Art dataset containing objects depicted in multiple styles. Experimental results show that SwiDeN outperforms other approaches for the depiction-invariant object recognition problem. Ravi Kiran Sarvadevabhatla, Shiv Surya, Srinivas S. Kruthiventi, Venkatesh Babu Radhakrishnan |
ACM Multimedia | 4 |
| 2016 | Enabling My Robot To Play Pictionary: Recurrent Neural Networks For Sketch RecognitionabstractFreehand sketching is an inherently sequential process. Yet, most approaches for hand-drawn sketch recognition either ignore this sequential aspect or exploit it in an ad-hoc manner. In our work, we propose a recurrent neural network architecture for sketch object recognition which exploits the long-term sequential and structural regularities in stroke data in a scalable manner. Specifically, we introduce a Gated Recurrent Unit based framework which leverages deep sketch features and weighted per-timestep loss to achieve state-of-the-art results on a large database of freehand object sketches across a large number of object categories. The inherently online nature of our framework is especially suited for on-the-fly recognition of objects as they are being drawn. Thus, our framework can enable interesting applications such as camera-equipped robots playing the popular party game Pictionary with human players and generating sparsified yet recognizable sketches of objects. Ravi Kiran Sarvadevabhatla, Jogendra Kundu, Venkatesh Babu Radhakrishnan |
ACM Multimedia | 3 |
| 2016 | A survey on compressed domain video analysis techniques
Venkatesh Babu Radhakrishnan, Manu Tom, Paras Wadekar |
Multim. Tools Appl. | 1 |
| 2015 | Data-free Parameter Pruning for Deep Neural NetworksabstractDeep Neural nets (NNs) with millions of parameters are at the heart of many state-of-the-art computer vision systems today. However, recent works have shown that much smaller models can achieve similar levels of performance. In this work, we address the problem of pruning parameters in a trained NN model. Instead of removing individual weights one at a time as done in previous works, we remove one neuron at a time. We show how similar neurons are redundant, and propose a systematic way to remove them. Our experiments in pruning the densely connected layers show that we can remove upto 85\% of the total parameters in an MNIST-trained network, and about 35\% for AlexNet without significantly affecting performance. Our method can be applied on top of most networks with a fully connected layer to give a smaller network. Suraj Srinivas, Venkatesh Babu Radhakrishnan |
BMVC | 2 |
| 2015 | Attribute-Graph: A Graph Based Approach to Image RankingabstractWe propose a novel image representation, termed Attribute-Graph, to rank images by their semantic similarity to a given query image. An Attribute-Graph is an undirected fully connected graph, incorporating both local and global image characteristics. The graph nodes characterise objects as well as the overall scene context using mid-level semantic attributes, while the edges capture the object topology. We demonstrate the effectiveness of Attribute-Graphs by applying them to the problem of image ranking. We benchmark the performance of our algorithm on the 'rPascal' and 'rImageNet' datasets, which we have created in order to evaluate the ranking performance on complex queries containing multiple objects. Our experimental evaluation shows that modelling images as Attribute-Graphs results in improved ranking performance over existing techniques. Nikita Prabhu, Venkatesh Babu Radhakrishnan |
ICCV | 2 |
| 2015 | Crowd flow segmentation in compressed domain using CRFabstractCrowd flow segmentation is an important step in many video surveillance tasks. In this work, we propose an algorithm for segmenting flows in H.264 compressed videos in a completely unsupervised manner. Our algorithm works on motion vectors which can be obtained by partially decoding the compressed video without extracting any additional features. Our approach is based on modelling the motion vector field as a Conditional Random Field (CRF) and obtaining oriented motion segments by finding the optimal labelling which minimises the global energy of CRF. These oriented motion segments are recursively merged based on gradient across their boundaries to obtain the final flow segments. This work in compressed domain can be easily extended to pixel domain by substituting motion vectors with motion based features like optical flow. The proposed algorithm is experimentally evaluated on a standard crowd flow dataset and its superior performance in both accuracy and computational time are demonstrated through quantitative results. Srinivas S. Kruthiventi, Venkatesh Babu Radhakrishnan |
ICIP | 2 |
| 2015 | Salient object detection via objectness measureabstractSalient object detection has become an important task in many image processing applications. The existing approaches exploit background prior and contrast prior to attain state of the art results. In this paper, instead of using background cues, we estimate the foreground regions in an image using objectness proposals and utilize it to obtain smooth and accurate saliency maps. We propose a novel saliency measure called ‘foreground connectivity’ which determines how tightly a pixel or a region is connected to the estimated foreground. We use the values assigned by this measure as foreground weights and integrate these in an optimization framework to obtain the final saliency maps. We extensively evaluate the proposed approach on two benchmark databases and demonstrate that the results obtained are better than the existing state of the art approaches. R. Sai Srivatsa, Venkatesh Babu Radhakrishnan |
ICIP | 2 |
| 2015 | Eye of the Dragon: Exploring Discriminatively Minimalist Sketch-based Abstractions for Object CategoriesabstractAs a form of visual representation, freehand line sketches are typically studied as an end product of the sketching process. However, from a recognition point of view, one can also study various orderings and properties of the primitive strokes that compose the sketch. Studying sketches in this manner has enabled us to create novel sparse yet discriminative sketch-based representations for object categories which we term category-epitomes. Concurrently, the epitome construction provides a natural measure for quantifying the sparseness underlying the original sketch, which we term epitome-score. We analyze category-epitomes and epitome-scores for hand-drawn sketches from a sketch dataset of 160 object categories commonly encountered in daily life. Our analysis provides a novel viewpoint for examining the complexity of representation for visual object categories. Ravi Kiran Sarvadevabhatla, Venkatesh Babu Radhakrishnan |
ACM Multimedia | 2 |
| 2015 | Robust tracking with interest points: A sparse representation approach
Venkatesh Babu Radhakrishnan, Priti Parate, Acharya K. Aniruddha |
Image Vis. Comput. | 1 |
| 2015 | Anomaly detection in compressed H.264/AVC video
Sovan Biswas, Venkatesh Babu Radhakrishnan |
Multim. Tools Appl. | 2 |
| 2015 | Synthetic image super resolution using FeatureMatch
S. Avinash Ramakanth, Venkatesh Babu Radhakrishnan |
Multim. Tools Appl. | 2 |
| 2015 | Compressed domain human action recognition in H.264/AVC video streams
Manu Tom, Venkatesh Babu Radhakrishnan, Gnana Praveen Rajasekhar |
Multim. Tools Appl. | 2 |
| 2014 | SeamSeg: Video Object Segmentation Using Patch SeamsabstractIn this paper, we propose a technique for video object segmentation using patch seams across frames. Typically, seams, which are connected paths of low energy, are utilised for retargeting, where the primary aim is to reduce the image size while preserving the salient image contents. Here, we adapt the formulation of seams for temporal label propagation. The energy function associated with the proposed video seams provides temporal linking of patches across frames, to accurately segment the object. The proposed energy function takes into account the similarity of patches along the seam, temporal consistency of motion and spatial coherency of seams. Label propagation is achieved with high fidelity in the critical boundary regions, utilising the proposed patch seams. To achieve this without additional overheads, we curtail the error propagation by formulating boundary regions as rough-sets. The proposed approach out-perform state-of-the-art supervised and unsupervised algorithms, on benchmark datasets. S. Avinash Ramakanth, Venkatesh Babu Radhakrishnan |
CVPR | 2 |
| 2014 | Visual object tracking via random ferns based classificationabstractDesigning a robust algorithm for visual object tracking has been a challenging task since many years. There are trackers in the literature that are reasonably accurate for many tracking scenarios but most of them are computationally expensive. This narrows down their applicability as many tracking applications demand real time response. In this paper, we present a tracker based on random ferns. Tracking is posed as a classification problem and classification is done using ferns. We used ferns as they rely on binary features and are extremely fast at both training and classification as compared to other classification algorithms. Our experiments show that the proposed tracker performs well on some of the most challenging tracking datasets and executes much faster than one of the state-of-the-art trackers, without much difference in tracking accuracy. Acharya K. Aniruddha, Venkatesh Babu Radhakrishnan |
ICASSP | 2 |
| 2014 | Optical flow estimation using Approximate Nearest Neighbor Field fusionabstractThis paper proposes an optical flow algorithm by adapting Approximate Nearest Neighbor Fields (ANNF) to obtain a pixel level optical flow between image sequence. Patch similarity based coherency is performed to refine the ANNF maps. Further improvement in mapping between the two images are obtained by fusing bidirectional ANNF maps between pair of images. Thus a highly accurate pixel level flow is obtained between the pair of images. Using pyramidal cost optimization, the pixel level optical flow is further optimized to a sub-pixel level. The proposed approach is evaluated on the mid-dlebury dataset and the performance obtained is comparable with the state of the art approaches. Furthermore, the proposed approach can be used to compute large displacement optical flow as evaluated using MPI Sintel dataset. O. U. Nirmal Jith, S. Avinash Ramakanth, Venkatesh Babu Radhakrishnan |
ICASSP | 3 |
| 2014 | Sparse representation based anomaly detection with enhanced local dictionariesabstractIn this paper, we propose a novel approach for anomaly detection by modeling the usual behaviour with enhanced dictionary. The corresponding sparse reconstruction error indicates the anomaly. We compute the dictionaries, for each local region, from feature descriptors obtained from usual behavior. The novelty of the proposed work is in enhancing the local dictionaries based on the similarity of usual behavior with its spatial neighbors. Dictionary enhancement is achieved by appending `transformed dictionary' to the `local dictionary'. This `transformed dictionary' is learned based on the transformations of behavior patterns across two neighboring regions. We conduct experiments on widely used UCSD Ped1 and Ped2 datasets to compare with the existing algorithms and demonstrate the improvement in anomaly detection with enhanced dictionaries compared to typically learned local dictionary. Sovan Biswas, Venkatesh Babu Radhakrishnan |
ICIP | 2 |
| 2014 | Super-pixel based crowd flow segmentation in H.264 compressed videosabstractIn this paper, we have proposed a simple yet robust novel approach for segmentation of high density crowd flows based on super-pixels in H.264 compressed videos. The collective representation of the motion vectors of the compressed video sequence is transformed to color map and super-pixel segmentation is performed at various scales for clustering the coherent motion vectors. The number of dynamically meaningful flow segments is determined by measuring the confidence score of the accumulated multi-scale super-pixel boundaries. The final crowd flow segmentation is obtained from the edges that are consistent across all the super-pixel resolutions. Hence, our major contribution involves obtaining the flow segmentation by clustering the motion vectors and determination of number of flow segments using only motion super-pixels without any prior assumption of the number of flow segments. The proposed approach was bench-marked on standard crowd flow dataset. Experiments demonstrated better accuracy and speedup for the proposed approach compared to the state-of-the-art methods. Sovan Biswas, Gnana Praveen Rajasekhar, Venkatesh Babu Radhakrishnan |
ICIP | 3 |
| 2014 | FeatureMatch: A General ANNF Estimation Technique and its ApplicationsabstractIn this paper, we propose FeatureMatch, a generalised approximate nearest-neighbour field (ANNF) computation framework, between a source and target image. The proposed algorithm can estimate ANNF maps between any image pairs, not necessarily related. This generalisation is achieved through appropriate spatial-range transforms. To compute ANNF maps, global colour adaptation is applied as a range transform on the source image. Image patches from the pair of images are approximated using low-dimensional features, which are used along with KD-tree to estimate the ANNF map. This ANNF map is further improved based on image coherency and spatial transforms. The proposed generalisation, enables us to handle a wider range of vision applications, which have not been tackled using the ANNF framework. We illustrate two such applications namely: 1) optic disk detection and 2) super resolution. The first application deals with medical imaging, where we locate optic disks in retinal images using a healthy optic disk image as common target image. The second application deals with super resolution of synthetic images using a common source image as dictionary. We make use of ANNF mappings in both these applications and show experimentally that our proposed approaches are faster and accurate, compared with the state-ofthe-art techniques. S. Avinash Ramakanth, Venkatesh Babu Radhakrishnan |
IEEE Trans. Image Process. | 2 |
| 2013 | H.264 compressed video classification using Histogram of Oriented Motion Vectors (HOMV)abstractIn this paper, we have proposed a simple and effective approach to classify H.264 compressed videos, by capturing orientation information from the motion vectors. Our major contribution involves computing Histogram of Oriented Motion Vectors (HOMV) for overlapping hierarchical Space-Time cubes. The Space-Time cubes selected are partially overlapped. HOMV is found to be very effective to define the motion characteristics of these cubes. We then use Bag of Features (BOF) approach to define the video as histogram of HOMV keywords, obtained using k-means clustering. The video feature, thus computed, is found to be very effective in classifying videos. We demonstrate our results with experiments on two large publicly available video database. Sovan Biswas, Venkatesh Babu Radhakrishnan |
ICASSP | 2 |
| 2013 | A near optimal projection for Sparse representation based classificationabstractSparse representation based classification (SRC) is one of the most successful methods that has been developed in recent times for face recognition. Optimal projection for Sparse representation based classification (OPSRC) provides a dimensionality reduction map that is supposed to give optimum performance for SRC framework. However, the computational complexity involved in this method is too high. Here, we propose a new projection technique using the data scatter matrix which is computationally superior to the optimal projection method with comparable classification accuracy with respect OPSRC. The performance of the proposed approach is benchmarked with various publicly available face database. Sreekanth Raja, Venkatesh Babu Radhakrishnan |
ICASSP | 2 |
| 2013 | Context-aware real-time tracking in sparse representation frameworkabstractReal-time object tracking is a difficult task in unconstrained environment due to the variations in factors such as pose, size, illumination, partial occlusion and motion blur. In this paper, we propose a novel approach based on local sparse representation for robust object tracking to address the above issues. In the proposed approach, a search window, including the object and the surrounding information is used to create dictionary of overlapping patches. This context-based method is used to discriminate the confusing patches of the foreground with those in the background. Candidate patches are sparsely represented in the dictionary and foreground/background classification is done by computing the confidence map based on the distribution of sparse coefficients. Pyramidal structure of the object window that depicts object at different scales is used to create a dictionary that can handle scale changes. Object is localized by seeking the mode of the confidence map. A suitable dictionary update strategy is used to alleviate the drift problem during tracking. Numerous experiments on challenging videos demonstrate that the proposed tracker outperforms several state-of-the-art algorithms. The proposed approach tracks at a high processing speed and is suitable for real-time applications. M. J. Ashwini, Venkatesh Babu Radhakrishnan, K. R. Ramakrishnan |
ICIP | 2 |
| 2013 | Real-time robust tracking via sparse representation: A mode-seeking approachabstractIn this paper, we propose a robust realtime tracking as a mode seeking process over likelihood map via sparse representation. In order to estimate the likelihood map, the target and candidate models were represented as overlapping patches. The likelihood map of the target candidate is obtained by sparsely representing the candidate patches in the space spanned by target patches and trivial bases. The object is localized by iteratively seeking the mode of this likelihood map. Since the mode-seeking process localizes the object in few iterations, it achieves realtime speed. Since the local patches are less sensitive to global appearance and illumination changes, the proposed approach shows robustness to the aforementioned challenges. We quantify the performance of the proposed tracker on many video sequences with various challenges involving occlusion, illumination change and pose variations. The proposed approach shows excellent performance in terms of robustness and speed compared to other trackers. Venkatesh Babu Radhakrishnan |
ICIP | 1 |
| 2013 | Interest points based object tracking via sparse representationabstractIn this paper, we propose an interest point based object tracker in sparse representation (SR) framework. In the past couple of years, there have been many proposals for object tracking in sparse framework exhibiting robust performance in various challenging scenarios. One of the major issues with these SR trackers is its slow execution speed mainly attributed to the particle filter framework. In this paper, we propose a robust interest point based tracker in l1minimization framework that runs at real-time with better performance compared to the state of the art trackers. In the proposed tracker, the target dictionary is obtained from the patches around target interest points. Next, the interest points from the candidate window of the current frame are obtained. The correspondence between target and candidate points are obtained via solving the proposed l1minimization problem. A robust matching criterion is proposed to prune the noisy matches. The object is localized by measuring the displacement of these interest points. The reliable candidate patches are used for updating the target dictionary. The performance of the proposed tracker is bench marked with several complex video sequences and found to be fast and robust compared to reported state of the art trackers. Venkatesh Babu Radhakrishnan, Priti Parate |
ICIP | 1 |
| 2013 | Rapid human action recognition in H.264/AVC compressed domain for video surveillanceabstractThis paper discusses a novel high-speed approach for human action recognition in H.264/AVC compressed domain. The proposed algorithm utilizes cues from quantization parameters and motion vectors extracted from the compressed video sequence for feature extraction and further classification using Support Vector Machines (SVM). The ultimate goal of our work is to portray a much faster algorithm than pixel domain counterparts, with comparable accuracy, utilizing only the sparse information from compressed video. Partial decoding rules out the complexity of full decoding, and minimizes computational load and memory usage, which can effect in reduced hardware utilization and fast recognition results. The proposed approach can handle illumination changes, scale, and appearance variations, and is robust in outdoor as well as indoor testing scenarios. We have tested our method on two benchmark action datasets and achieved more than 85% accuracy. The proposed algorithm classifies actions with speed (>2000 fps) approximately 100 times more than existing state-of-the-art pixel-domain algorithms. Manu Tom, Venkatesh Babu Radhakrishnan |
VCIP | 2 |
| 2013 | Subject independent human action recognition using spatio-depth information and meta-cognitive RBF network
Venkatesh Babu Radhakrishnan, Ramaswamy Savitha, Suresh Sundaram 0002, Bhuvnesh Agarwal |
Eng. Appl. Artif. Intell. | 1 |
| 2012 | Human detection using sparse representationabstractThe problem of human detection is challenging, more so, when faced with adverse conditions such as occlusion and background clutter. This paper addresses the problem of human detection by representing an extracted feature of an image using a sparse linear combination of chosen dictionary atoms. The detection along with the scale finding, is done by using the coefficients obtained from sparse representation. This is of particular interest as we address the problem of scale using a scale-embedded dictionary where the conventional methods detect the object by running the detection window at all scales. G. Krishna Vinay, Sk. Mohammadul Haque, Venkatesh Babu Radhakrishnan, K. R. Ramakrishnan |
ICASSP | 3 |
| 2012 | Fragment-based real-time object tracking: A sparse representation approachabstractReal-time object tracking is a critical task in many computer vision applications. Achieving rapid and robust tracking while handling changes in object pose and size, varying illumination and partial occlusion, is a challenging task given the limited amount of computational resources. In this paper we propose a real-time object tracker in l1framework addressing these issues. In the proposed approach, dictionaries containing templates of overlapping object fragments are created. The candidate fragments are sparsely represented in the dictionary fragment space by solving the l1regularized least squares problem. The non zero coefficients indicate the relative motion between the target and candidate fragments along with a fidelity measure. The final object motion is obtained by fusing the reliable motion information. The dictionary is updated based on the object likelihood map. The proposed tracking algorithm is tested on various challenging videos and found to outperform earlier approach. M. S. Naresh Kumar, Priti Parate, Venkatesh Babu Radhakrishnan |
ICIP | 3 |
| 2012 | Optic disk localization using L1 minimizationabstractAutomatic eye screening for conditions like diabetic retinopathy critically hinges on detection and localization of Optic disk (OD). In this paper, we present a novel scale-embedded dictionary-based method that poses the problem of OD localization as that of classification, carried out in sparse representation framework. A dictionary is created with manually marked fixed-sized sub-images that contain OD at the center, for multiple scales. For a given test image, all subimages are sparsely represented as a linear combination of OD dictionary elements. A confidence measure indicating the likelihood of the presence of OD is obtained from these coefficients. Red channel and gray intensity images are processed independently, and their respective confidence measures are fused to form a confidence map. A blob detector is run on the confidence map, whose peak response is considered to be at the location of the OD. The proposed method is evaluated on publicly available databases such as DIARETDB0, DIARETDB1 and DRIVE. The OD was correctly localized in 253 out of 259 images, with an average computation time of 3.8 seconds/image and accuracy of 97.6%. Comparisons with two existing techniques are also discussed. Neelam Sinha, Venkatesh Babu Radhakrishnan |
ICIP | 2 |
| 2012 | Meta-Cognitive Neuro-Fuzzy Inference System for human emotion recognitionabstractIn this paper, we propose a Meta-Cognitive Neuro-Fuzzy Inference System (McFIS) for recognition of emotions from facial features. Local binary patterns have been proven to effectively describe the statistical characteristics of face image as it contains information related to edges, spots, etc. The aim of McFIS is to approximate the functional relationship between the facial features and various emotions. McFIS classifier and its sequential learning algorithm is developed based on the principles of self-regulation observed in human meta-cognition. McFIS decides on what-to-learn, when-to-learn and how-to-learn based on the knowledge stored in the classifier and the information contained in the new training samples. The sequential learning algorithm of McFIS is controlled and monitored by the meta-cognitive components which uses class-specific, knowledge based criteria along with self-regulatory thresholds to decide on one of the following strategies: a) sample deletion b) sample learning and c) sample reserve. Performance of proposed McFIS based facial emotion recognition is evaluated on LBP features extracted from JAFFE database. The simulation results are compared with support vector machine classifier and other results available in literature. The results indicate the superior performance of McFIS in comparison to other algorithms. K. Subramanian 0001, Suresh Sundaram 0002, Venkatesh Babu Radhakrishnan |
IJCNN | 3 |
| 2012 | Human action recognition using a fast learning fully complex-valued classifier
Venkatesh Babu Radhakrishnan, Suresh Sundaram 0002, Ramaswamy Savitha |
Neurocomputing | 1 |
| 2011 | Fully complex-valued ELM classifiers for human action recognitionabstractIn this paper, we present a fast learning neural network classifier for human action recognition. The proposed classifier is a fully complex-valued neural network with a single hidden layer. The neurons in the hidden layer employ the fully complex-valued hyperbolic secant as an activation function. The parameters of the hidden layer are chosen randomly and the output weights are estimated analytically as a minimum norm least square solution to a set of linear equations. The fast leaning fully complex-valued neural classifier is used for recognizing human actions accurately. Optical flow-based features extracted from the video sequences are utilized to recognize 10 different human actions. The feature vectors are computationally simple first order statistics of the optical flow vectors, obtained from coarse to fine rectangular patches centered around the object. The results indicate the superior performance of the complex-valued neural classifier for action recognition. The superior performance of the complex neural network for action recognition stems from the fact that motion, by nature, consists of two components, one along each of the axes. Venkatesh Babu Radhakrishnan, Suresh Sundaram 0002 |
IJCNN | 1 |
| 2010 | Online adaptive radial basis function networks for robust object tracking
Venkatesh Babu Radhakrishnan, Suresh Sundaram 0002, Anamitra Makur |
Comput. Vis. Image Underst. | 1 |
| 2008 | Robust object tracking with background-weighted local kernels
Jaideep Jeyakar, Venkatesh Babu Radhakrishnan, K. R. Ramakrishnan |
Comput. Vis. Image Underst. | 2 |
| 2008 | Evaluation and monitoring of video quality for UMA enabled video streaming systems
Venkatesh Babu Radhakrishnan, Andrew Perkis, Odd Inge Hillestad |
Multim. Tools Appl. | 1 |
| 2007 | Kernel-Based Spatial-Color Modeling for Fast Moving Object TrackingabstractVisual tracking has been a challenging problem in computer vision over the decades. The applications of Visual Tracking are far-reaching, ranging from surveillance and monitoring to smart rooms. Mean-shift (MS) tracker, which gained more attention recently, is known for tracking objects in a cluttered environment and its low computational complexity. The major problem encountered in histogram-based MS is its inability to track rapidly moving objects. In order to track fast moving objects, we propose a new robust mean-shift tracker that uses both spatial similarity measure and color histogram-based similarity measure. The inability of MS tracker to handle large displacements is circumvented by the spatial similarity-based tracking module, which lacks robustness to object's appearance change. The performance of the proposed tracker is better than the individual trackers for tracking fast-moving objects with better accuracy. Venkatesh Babu Radhakrishnan, Anamitra Makur |
ICASSP (1) | 1 |
| 2007 | Robust Object Tracking with Radial Basis Function NetworksabstractVisual tracking has been a challenging problem in computer vision over the decades. The applications of visual tracking are far-reaching, ranging from surveillance and monitoring to smart rooms. In this paper we present a novel object tracker based on fast learning radial basis function (RBF) networks. Here, the object and background pixel-based color features are used to develop object/non-object RBF classifiers. The posterior probability information of these classifiers are used for developing an efficient object model for tracking in the subsequent frames. The performance of the proposed tracker is tested with many video sequences of real-life complexity and compared against the color-based mean-shift tracker. The proposed tracker is illustrated to be suitable for real-time robust object tracking due to its low computational complexity. Venkatesh Babu Radhakrishnan, Suresh Sundaram 0002, Anamitra Makur |
ICASSP (1) | 1 |
| 2007 | Robust Object Tracking using Local Kernels and Background InformationabstractThe mean shift algorithm has been proved to be efficient for tracking 2D blobs through a video sequence. Even so, this algorithm has certain inherent disadvantages. In this paper, we propose a robust tracking algorithm which overcomes the drawbacks of global color histogram based tracking. We incorporate tracking based only on reliable colors by separating the object from its background. A fast yet robust model update is employed to overcome illumination changes. This algorithm is computationally simple enough to be executed real time and was tested on several complex video sequences. The proposed technique could be easily extended to other tracking algorithms too. Jaideep Jeyakar, Venkatesh Babu Radhakrishnan, K. R. Ramakrishnan |
ICIP (5) | 2 |
| 2007 | Robust tracking with motion estimation and local Kernel-based color modeling
Venkatesh Babu Radhakrishnan, Patrick Pérez, Patrick Bouthemy |
Image Vis. Comput. | 1 |
| 2007 | Compressed domain video retrieval using object and global motion descriptors
Venkatesh Babu Radhakrishnan, K. R. Ramakrishnan |
Multim. Tools Appl. | 1 |
| 2007 | No-reference JPEG-image quality assessment using GAP-RBF
Venkatesh Babu Radhakrishnan, Suresh Sundaram 0002, Andrew Perkis |
Signal Process. | 1 |
| 2006 | Kernel-Based Robust Tracking for Objects Undergoing Occlusion
Venkatesh Babu Radhakrishnan, Patrick Pérez, Patrick Bouthemy |
ACCV (2) | 1 |
| 2006 | Object-based Surveillance Video Compression using Foreground Motion CompensationabstractVideo surveillance is currently one of the most active area of research in both academia and industry. Though much work has been done in the area of smart surveillance, relatively little work has been reported to compress the surveillance videos. In this paper, we propose an object based video compression system using foreground motion compensation for applications such as archival and transmission of surveillance video. The proposed system segments independently moving objects from the video and codes them with respect to the previously reconstructed frame. The error resulting from object-based motion compensation is coded using SA-DCT procedure. The proposed system codes the surveillance video using far lesser bits compared to conventional video compression techniques Venkatesh Babu Radhakrishnan, Anamitra Makur |
ICARCV | 1 |
| 2006 | Image Quality Measurement Using Sparse Extreme Learning Machine ClassifierabstractIn this paper, we present a machine learning approach to measure the visual quality of JPEG-coded images. The features for predicting the perceived image quality are extracted by considering key human visual sensitivity factors such as edge amplitude, edge length, background activity and background luminance. Image quality estimation involves computation of functional relationship between HVS features and subjective test scores. The subjective test scores for modified images are obtained with-out referring to their original images (called 'no reference'). Here, the problem of quality estimation is transformed to a sparse data classification problem using a sparse extreme learning machine (S-ELM). The S-ELM classifier estimate the posterior probability of a given image. Here, the mean opinion score ('visual quality') of an image is derived using the predicted class number and their estimated posterior probability. The experimental results prove that the estimated visual quality emulate the mean opinion score very well. The experimental results are compared with the existing JPEG no-reference image quality index and full-reference structural similarity image quality index. The result clearly shows the machine learning approach outperform the existing algorithms in the literature Suresh Sundaram 0002, Venkatesh Babu Radhakrishnan, Narasimhan Sundararajan |
ICARCV | 2 |
| 2005 | An HVS-based no-reference perceptual quality assessment of JPEG coded images using neural networksabstractIn this paper, we present a novel no-reference (NR) metric to assess the quality of JPEG-coded images. The features for predicting the perceived image quality are extracted by considering the key human visual sensitivity factors such as, edge amplitude, edge length, background activity and background luminance. The extracted features with the subjective test results are used to train a multi-layer perceptron (MLP) neural network. Experimental results show that the prediction of the trained neural network is very close to the mean opinion score (MOS). The subjective test results of the proposed metric are compared with the Wang-Bovik's NR blockiness metric. Further, this metric can be extended to assess the quality of the MPLG/H.26x compressed videos. Venkatesh Babu Radhakrishnan, Andrew Perkis |
ICIP (1) | 1 |
| 2005 | Robust tracking with motion estimation and kernel-based color modellingabstractVisual tracking is still a challenging problem in computer vision. The applications of visual tracking are far-reaching, ranging from surveillance and monitoring to smart rooms. In this work, we propose a new method to track arbitrary objects using both sum-of-squared differences (SSD) and color-based mean-shift (MS) trackers in the Kalman filter framework. The SSD and the MS trackers complement each other by overcoming their respective disadvantages. The rapid model change in SSD tracker is overcome by the MS tracker module, while the inability of MS tracker to handle large displacements and occlusions is circumvented by the SSD module. In addition, rapid scale changes of the object generated by camera ego-motion or zooming are measured by a global affine motion estimation. Finally, the global appearance model on which MS relies is updated, based on the Bhattacharyya distance between this target model and current candidate model. This permits to tackle global appearance changes of the object. The performance of the proposed tracker is better than the individual SSD and MS trackers. Venkatesh Babu Radhakrishnan, Patrick Pérez, Patrick Bouthemy |
ICIP (1) | 1 |
| 2004 | Recognition of human actions using motion history information extracted from the compressed video
Venkatesh Babu Radhakrishnan, K. R. Ramakrishnan |
Image Vis. Comput. | 1 |
| 2004 | Video object segmentation: a compressed domain approachabstractThis paper addresses the problem of extracting video objects from MPEG compressed video. The only cues used for object segmentation are the motion vectors which are sparse in MPEG. A method for automatically estimating the number of objects and extracting independently moving video objects using motion vectors is presented here. First, the motion vectors are accumulated over a few frames to enhance the motion information, which are further spatially interpolated to get dense motion vectors. The final segmentation, using the dense motion vectors, is obtained by applying the expectation maximization (EM) algorithm. A block-based affine clustering method is proposed for determining the number of appropriate motion models to be used for the EM step and the segmented objects are temporally tracked to obtain the video objects. Finally, a strategy for edge refinement is proposed to extract the precise object boundaries. Illustrative examples are provided to demonstrate the efficacy of the approach. A prominent application of the proposed method is that of object-based coding, which is part of the MPEG-4 standard. Venkatesh Babu Radhakrishnan, K. R. Ramakrishnan, S. H. Srinivasan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Compressed domain human motion recognition using motion history informationabstractIn this paper we present a system for classifying various human actions in compressed domain video framework. We introduce the notion of quantifying the motion involved, through what we call "Motion Flow History" (MFH). The encoded motion information readily available in the compressed MPEG stream is used to construct the coarse Motion History Image (MHI) and the corresponding MFH. The features extracted from the static MHI and MFH compactly characterize the temporal and motion information of the action. Since the features are extracted from the partially decoded sparse motion data, the computational load is minimized to a great extent. The extracted features are used to train the KNN, Neural network and the Bayes classifiers for recognizing a set of seven human actions. Experimental results show that the proposed method efficiently recognizes the set of actions considered. Venkatesh Babu Radhakrishnan, K. R. Ramakrishnan |
ICASSP (3) | 1 |
| 2003 | Compressed domain human motion recognition using motion history informationabstractIn this paper we present a system for classifying various human actions in compressed domain video framework. We introduce the notion of quantifying the motion involved, through what we call "motion flow history" (MFH). The encoded motion information readily available in the compressed MPEG stream is used to construct the coarse motion history image (MHI) and the corresponding MFH. The features extracted from the static MHI and MFH compactly characterize the temporal and motion information of the action. Since the features are extracted from the partially decoded sparse motion data, the computational load is minimized to a great extent. The extracted features are used to train the KNN, neural network, SVM and the Bayes classifiers for recognizing a set of seven human actions. Experimental results show that the proposed method efficiently recognizes the set of actions considered. Venkatesh Babu Radhakrishnan, K. R. Ramakrishnan |
ICIP (3) | 1 |
| 2002 | Compressed domain motion segmentation for video object extractionabstractThis paper addresses the problem of extracting video objects from MPEG compressed video. The only cues used for object segmentation are the motion vectors which are sparse in MPEG. A method for automatically estimating the number of objects and extracting independently moving video objects using motion vectors is presented here. First, the motion vectors are accumulated over few frames to enhance the motion information, which are further spatially interpolated to get a dense motion vectors. The final segmentation from the dense motion vectors is obtained by applying Expectation Maximization (EM) algorithm. A block based affine clustering method is proposed for determining the number of appropriate motion models to be used for the EM step. Finally, the segmented objects are temporally tracked to obtain the video objects. This work has been carried out in the context of the emerging MPEG-4 standard which aims at interactivity at object level. Venkatesh Babu Radhakrishnan, K. R. Ramakrishnan |
ICASSP | 1 |
| 2002 | Compressed domain action classification using HMM
Venkatesh Babu Radhakrishnan, B. Anantharaman, K. R. Ramakrishnan, S. H. Srinivasan |
Pattern Recognit. Lett. | 1 |