EDBT 2026 Demo / reviewers in the wild / expert
Pravendra Singh
dblp:160/8743
· DBLP profile ↗
52ranked-venue papers
18as first author
36since 2021 · last 2026
0000-0003-1001-2219ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 13 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning with less: A survey of deep learning in medical imaging under varying supervision levels
Suruchi Kumari, Pravendra Singh |
Artif. Intell. Medicine | 2 |
| 2026 | Multi-scale hypergraphs for trajectory imputation and prediction in real-world scenarios
Pranav Singh Chib, Pravendra Singh |
Expert Syst. Appl. | 2 |
| 2026 | Leveraging language prior for infrared small target detection
Pranav Singh Chib, Pravendra Singh |
Neurocomputing | 2 |
| 2026 | Annotation noise-aware semi-supervised and domain adaptive learning for 3D medical imaging
Suruchi Kumari, Pravendra Singh |
Neurocomputing | 2 |
| 2026 | Dual-view diffusion for pedestrian trajectory imputation
Pranav Singh Chib, Pravendra Singh |
Neural Networks | 2 |
| 2026 | Integrating orthogonal supervision for sparse semi-supervised 3D medical image segmentation
Suruchi Kumari, Pravendra Singh |
Neural Networks | 2 |
| 2025 | A Unified Degradation-Robust Approach to SSL and UDA for 3D Medical ImagesabstractMedical image segmentation often faces the dual challenges of limited annotations and domain shifts, further complicated by degraded images in practical scenarios. Traditional methods tend to underperform when these issues occur simultaneously, as they are typically designed for specific tasks. To address this, we propose a unified framework that effectively handles limited annotations and domain shifts while also managing both clean and degraded images during inference. Overcoming these challenges requires focusing on three critical aspects: First, the model must be robust to various noise conditions. Second, it should excel at capturing domain-invariant features. Third, it should effectively utilize unlabeled data. We propose three major components in our approach to tackle these challenges. First, the Wavelet-based Cross-Component Exchange (WCCE) swaps high-frequency wavelet components between labeled and unlabeled images to enhance robustness. Second, we employ a diffusion VNet architecture with a reweighting mechanism to capture domain-invariant features. Finally, we utilize Cross-Decoder Pseudo (CDP) training to effectively leverage unlabeled data. Evaluations on three publicly available medical datasets and across four types of degraded image scenarios demonstrate that our method outperforms state-of-the-art (SOTA) techniques, consistently delivering superior performance across varying image qualities. Our approach not only addresses annotation scarcity and domain shift but also effectively manages noisy and blurred conditions, setting a new benchmark in medical image segmentation. Suruchi Kumari, Pravendra Singh |
AAAI | 2 |
| 2025 | Annotation Ambiguity Aware Semi-Supervised Medical Image SegmentationabstractDespite the remarkable progress of deep learning-based methods in medical image segmentation, their use in clinical practice remains limited for two main reasons. First, obtaining a large medical dataset with precise annotations to train segmentation models is challenging. Secondly, most current segmentation techniques generate a single deterministic segmentation mask for each image. However, in real-world scenarios, there is often significant uncertainty regarding what defines the "correct" segmentation, and various expert annotators might provide different segmentations for the same image. To tackle both of these problems, we propose Annotation Ambiguity Aware Semi-Supervised Medical Image Segmentation (AmbiSSL). AmbiSSL combines a small amount of multi-annotator labeled data and a large set of unlabeled data to generate diverse and plausible segmentation maps. Our method consists of three key components: (1) The Diverse Pseudo-Label Generation (DPG) module utilizes multiple decoders, created by performing randomized pruning on the original backbone decoder. These pruned decoders enable the generation of a diverse pseudo-label set; (2) a Semi-Supervised Latent Distribution Learning (SSLDL) module constructs a common latent space by utilizing both ground truth annotations and pseudo-label set; and (3) a Cross-Decoder Supervision (CDS) module, which enables pruned decoders to guide each other’s learning. We evaluated the proposed method on two publicly available datasets. Extensive experiments demonstrate that AmbiSSL can generate diverse segmentation maps using only a small amount of labeled data and abundant unlabeled data, offering a more practical solution for medical image segmentation by reducing reliance on large labeled datasets. Suruchi Kumari, Pravendra Singh |
CVPR | 2 |
| 2025 | Addressing Label Scarcity and Domain Shift in Medical Image Segmentation
Suruchi Kumari, Pravendra Singh |
MICCAI (7) | 2 |
| 2025 | Improving Medical Image Segmentation with Implicit Representation and Noisy Label Robustness
Suruchi Kumari, Harshdeep Singh, Pravendra Singh |
MICCAI (16) | 3 |
| 2025 | Small and dim target detection in infrared imagery: A review, current techniques and future directions
Pravendra Singh |
Neurocomputing | 2 |
| 2024 | MWIRSTD: A MWIR Small Target Detection DatasetabstractThis paper presents a novel mid-wave infrared (MWIR) small target detection dataset (MWIRSTD) comprising 14 video sequences containing approximately 1053 images with annotated targets of three distinct classes of small objects. Captured using cooled MWIR imagers, the dataset offers a unique opportunity for researchers to develop and evaluate state-of-the-art methods for small object detection in realistic MWIR scenes. Unlike existing datasets, which primarily consist of uncooled thermal images or synthetic data with targets superimposed onto the background or vice versa, MWIRSTD provides authentic MWIR data with diverse targets and environments. Extensive experiments on various traditional methods and deep learning-based techniques for small target detection are performed on the proposed dataset, providing valuable insights into their efficacy. The dataset and code are available at https://github.com/avinres/MWIRSTD. Avinash Upadhyay, Pravendra Singh |
ICIP | 5 |
| 2024 | MS-TIP: Imputation Aware Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction aims to predict future trajectories based on observed trajectories. Current state-of-the-art methods often assume that the observed sequences of agents are complete, which is a strong assumption that overlooks inherent uncertainties. Understanding pedestrian behavior when dealing with missing values in the observed sequence is crucial for enhancing the performance of predictive models. In this work, we propose the MultiScale hypergraph for Trajectory Imputation and Prediction (MS-TIP), a novel approach that simultaneously addresses the imputation of missing observations and the prediction of future trajectories. Specifically, we leverage transformers with diagonal masked self-attention to impute incomplete observations. Further, our approach promotes complex interaction modeling through multi-scale hypergraphs, optimizing our trajectory prediction module to capture different types of interactions. With the inclusion of scenic attention, we learn contextual scene information, instead of sole reliance on coordinates. Additionally, our approach utilizes an intermediate control point and refinement module to infer future trajectories accurately. Extensive experiments validate the efficacy of MS-TIP in precisely predicting pedestrian future trajectories. Code is publicly available at https://github.com/Pranav-chib/MS-TIP. Pranav Singh Chib, Achintya Nath, Paritosh Kabra, Ishu Gupta, Pravendra Singh |
ICML | 5 |
| 2024 | Enhancing Trajectory Prediction through Self-Supervised Waypoint Distortion PredictionabstractTrajectory prediction is an important task that involves modeling the indeterminate nature of agents to forecast future trajectories given the observed trajectory sequences. The task of predicting trajectories poses significant challenges, as agents not only move individually through time but also interact spatially. The learning of complex spatio-temporal representations stands as a fundamental challenge in trajectory prediction. To this end, we propose a novel approach called SSWDP (Self-Supervised Waypoint Distortion Prediction). We propose a simple yet highly effective self-supervised task of predicting distortion present in the observed trajectories to improve the representation learning of the model. Our approach can complement existing trajectory prediction methods. The experimental results highlight a significant improvement with relative percentage differences of 22.7%/38.9%, 33.8%/36.4%, and 16.60%/23.20% in ADE/FDE for the NBA, TrajNet++, and ETH-UCY datasets, respectively, compared to the baseline methods. Our approach also demonstrates a significant improvement over baseline methods with relative percentage differences of 76.8%/82.5% and 61.0%/36.1% in ADE/FDE for TrajNet++ and NBA datasets in distorted environments, respectively. Pranav Singh Chib, Pravendra Singh |
ICML | 2 |
| 2024 | STNet: Small Target Detection Network for IR Imagery
Pranav Singh Chib, Pravendra Singh |
ICPR (30) | 3 |
| 2024 | Pedestrian Trajectory Prediction with Missing Data: Datasets, Imputation, and BenchmarkingabstractPedestrian trajectory prediction is crucial for several applications such as robotics and self-driving vehicles. Significant progress has been made in the past decade thanks to the availability of pedestrian trajectory datasets, which enable trajectory prediction methods to learn from pedestrians' past movements and predict future trajectories. However, these datasets and methods typically assume that the observed trajectory sequence is complete, ignoring real-world issues such as sensor failure, occlusion, and limited fields of view that can result in missing values in observed trajectories. To address this challenge, we present TrajImpute, a pedestrian trajectory prediction dataset that simulates missing coordinates in the observed trajectory, enhancing real-world applicability. TrajImpute maintains a uniform distribution of missing data within the observed trajectories. In this work, we comprehensively examine several imputation methods to reconstruct the missing coordinates and benchmark them for imputing pedestrian trajectories. Furthermore, we provide a thorough analysis of recent trajectory prediction methods and evaluate the performance of these models on the imputed trajectories. Our experimental evaluation of the imputation and trajectory prediction methods offers several valuable insights. Our dataset provides a foundational resource for future research on imputation-aware pedestrian trajectory prediction, potentially accelerating the deployment of these methods in real-world applications. Publicly accessible links to the datasets and code files are available at https://github.com/Pranav-chib/TrajImpute. Pranav Singh Chib, Pravendra Singh |
NeurIPS | 2 |
| 2024 | Improving trajectory prediction in dynamic multi-agent environment by dropping waypoints
Pranav Singh Chib, Pravendra Singh |
Knowl. Based Syst. | 2 |
| 2024 | Rectification-Based Knowledge Retention for Task Incremental LearningabstractIn the task incremental learning problem, deep learning models suffer from catastrophic forgetting of previously seen classes/tasks as they are trained on new classes/tasks. This problem becomes even harder when some of the test classes do not belong to the training class set, i.e., the task incremental generalized zero-shot learning problem. We propose a novel approach to address the task incremental learning problem for both the non zero-shot and zero-shot settings. Our proposed approach, called Rectification-based Knowledge Retention (RKR), applies weight rectifications and affine transformations for adapting the model to any task. During testing, our approach can use the task label information (task-aware) to quickly adapt the network to that task. We also extend our approach to make it task-agnostic so that it can work even when the task label information is not available during testing. Specifically, given a continuum of test data, our approach predicts the task and quickly adapts the network to the predicted task. We experimentally show that our proposed approach achieves state-of-the-art results on several benchmark datasets for both non zero-shot and zero-shot task incremental learning. Pratik Mazumder, Pravendra Singh, Piyush Rai, Vinay P. Namboodiri |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | A Host Kernel-Based Approach for Tracing and Analyzing vCPUs in Virtual MachinesabstractVirtual machines (VMs) are a crucial technology for cloud computing, permitting multiple operating systems to run on a single physical host. Here, we propose a method for analyzing the performance of virtual machines in a cloud computing environment. The current approach of tracing both the host and virtual machines can be inefficient and may not be feasible for various reasons, such as security and privacy issues. We propose a host-only kernel tracing technique that uses only trace data generated on the host system. The trace data is then analyzed using the TraceCompass tool and two algorithms: Virtual CPU State Finder (VSF) and State Time Analysis (STA). The VSF algorithm generates a graphical view of the state of each thread running inside virtual machines for the identification of latency and performance degradation. The STA algorithm calculates the amount of time each virtual CPU spends in different states, providing important insights into the resource utilization and performance of virtual resources. In addition, our usage of the Trace Compass tool and EASE scripting allows more manageable analysis and visualization of the tracing data, making it accessible to a broader range of users. This method provides cloud providers with valuable information to maintain QoS and SLA parameters and improve the overall performance of running VMs. Ravjot Singh, Prakhar Gupta, Naman Jain, Neetesh Kumar, Pravendra Singh |
GLOBECOM | 5 |
| 2023 | Leveraging joint incremental learning objective with data ensemble for class incremental learning
Pratik Mazumder, Mohammed Asad Karim, Indu Joshi, Pravendra Singh |
Neural Networks | 4 |
| 2023 | Mitigate forgetting in few-shot class-incremental learning using different image views
Pratik Mazumder, Pravendra Singh |
Neural Networks | 2 |
| 2023 | Clustering Single-Cell RNA Sequence Data Using Information Maximized and Noise-Invariant RepresentationsabstractSingle-cell RNA sequencing (scRNA-seq) is a revolutionary methodology that helps to analyze transcriptome or genome information from a single cell. However, high dimensionality and sparsity in data due to dropout events pose computational challenges for existing state-of-the-art scRNA-seq clustering methods. Learning efficient representations becomes even more challenging due to the presence of noise in scRNA-seq data. To overcome the effect of noise and learn effective representations, this paper proposes sc-INDC (Single-Cell Information Maximized Noise-Invariant Deep Clustering), a deep neural network that facilitates learning of informative and noise-invariant representations of scRNA-seq data. Furthermore, the time complexity of the proposed sc-INDC is significantly lower compared to state-of-the-art scRNA-seq clustering methods. Extensive experimentation on fourteen publicly available scRNA-seq datasets illustrates the efficacy of the proposed model. Additionally, visualizations of t-SNE plots and several ablation studies are also conducted to provide insights into the improved representation ability of sc-INDC. Code of the proposed sc-INDC will be available at: https://github.com/arnabkmondal/sc-INDC. Arnab Kumar Mondal, Indu Joshi, Pravendra Singh, Prathosh A. P. |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Attaining Class-Level Forgetting in Pretrained Model Using Few Samples
Pravendra Singh, Pratik Mazumder, Mohammed Asad Karim |
ECCV (13) | 1 |
| 2022 | Fair Visual Recognition in Limited Data Regime using Self-Supervision and Self-DistillationabstractDeep learning models generally learn the biases present in the training data. Researchers have proposed several approaches to mitigate such biases and make the model fair. Bias mitigation techniques assume that a sufficiently large number of training examples are present. However, we observe that if the training data is limited, then the effectiveness of bias mitigation methods is severely degraded. In this paper, we propose a novel approach to address this problem. Specifically, we adapt self-supervision and self-distillation to reduce the impact of biases on the model in this setting. Self-supervision and self-distillation are not used for bias mitigation. However, through this work, we demonstrate for the first time that these techniques are very effective in bias mitigation. We empirically show that our approach can significantly reduce the biases learned by the model. Further, we experimentally demonstrate that our approach is complementary to other bias mitigation strategies. Our approach significantly improves their performance and further reduces the model biases in the limited data regime. Specifically, on the L-CIFAR-10S skewed dataset, our approach significantly reduces the bias score of the baseline model by 78.22% and outperforms it in terms of accuracy by a significant absolute margin of 8.89%. It also significantly reduces the bias score for the state-of-the-art domain independent bias mitigation method by 59.26% and improves its performance by a significant absolute margin of 7.08%. Pratik Mazumder, Pravendra Singh, Vinay P. Namboodiri |
WACV | 2 |
| 2022 | Few-shot image classification with composite rotation based self-supervised auxiliary task
Pratik Mazumder, Pravendra Singh, Vinay P. Namboodiri |
Neurocomputing | 2 |
| 2022 | Protected attribute guided representation learning for bias mitigation in limited data
Pratik Mazumder, Pravendra Singh |
Knowl. Based Syst. | 2 |
| 2022 | Dual class representation learning for few-shot image classification
Pravendra Singh, Pratik Mazumder |
Knowl. Based Syst. | 1 |
| 2022 | On restoration of degraded fingerprints
Indu Joshi, Ayush Utkarsh, Pravendra Singh, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra |
Multim. Tools Appl. | 3 |
| 2022 | Context extraction module for deep convolutional neural networks
Pravendra Singh, Pratik Mazumder, Vinay P. Namboodiri |
Pattern Recognit. | 1 |
| 2021 | Few-Shot Lifelong LearningabstractMany real-world classification problems often have classes with very few labeled training samples. Moreover, all possible classes may not be initially available for training, and may be given incrementally. Deep learning models need to deal with this two-fold problem in order to perform well in real-life situations. In this paper, we propose a novel Few-Shot Lifelong Learning (FSLL) method that enables deep learning models to perform lifelong/continual learning on few-shot data. Our method selects very few parameters from the model for training every new set of classes instead of training the full model. This helps in preventing overfitting. We choose the few parameters from the model in such a way that only the currently unimportant parameters get selected. By keeping the important parameters in the model intact, our approach minimizes catastrophic forgetting. Furthermore, we minimize the cosine similarity between the new and the old class prototypes in order to maximize their separation, thereby improving the classification performance. We also show that integrating our method with self-supervision improves the model performance significantly. We experimentally show that our method significantly outperforms existing methods on the miniImageNet, CIFAR-100, and CUB-200 datasets. Specifically, we outperform the state-of-the-art method by an absolute margin of 19.27% for the CUB dataset. Pratik Mazumder, Pravendra Singh, Piyush Rai |
AAAI | 2 |
| 2021 | Rectification-Based Knowledge Retention for Continual LearningabstractDeep learning models suffer from catastrophic forgetting when trained in an incremental learning setting. In this work, we propose a novel approach to address the task incremental learning problem, which involves training a model on new tasks that arrive in an incremental manner. The task incremental learning problem becomes even more challenging when the test set contains classes that are not part of the train set, i.e., a task incremental generalized zero-shot learning problem. Our approach can be used in both the zero-shot and non zero-shot task incremental learning settings. Our proposed method uses weight rectifications and affine transformations in order to adapt the model to different tasks that arrive sequentially. Specifically, we adapt the network weights to work for new tasks by "rectifying" the weights learned from the previous task. We learn these weight rectifications using very few parameters. We additionally learn affine transformations on the outputs generated by the network in order to better adapt them for the new task. We perform experiments on several datasets in both zero-shot and non zero-shot task incremental learning settings and empirically show that our approach achieves state-of-the-art results. Specifically, our approach outperforms the state-of-the-art non zero-shot task incremental learning method by over 5% on the CIFAR-100 dataset. Our approach also significantly outperforms the state-of-the-art task incremental generalized zero-shot learning method by absolute margins of 6.91% and 6.33% for the AWA1 and CUB datasets, respectively. We validate our approach using various ablation studies. Pravendra Singh, Pratik Mazumder, Piyush Rai, Vinay P. Namboodiri |
CVPR | 1 |
| 2021 | Knowledge Consolidation based Class Incremental Online Learning with Limited DataabstractWe propose a novel approach for class incremental online learning in a limited data setting. This problem setting is challenging because of the following constraints: (1) Classes are given incrementally, which necessitates a class incremental learning approach; (2) Data for each class is given in an online fashion, i.e., each training example is seen only once during training; (3) Each class has very few training examples; and (4) We do not use or assume access to any replay/memory to store data from previous classes. Therefore, in this setting, we have to handle twofold problems of catastrophic forgetting and overfitting. In our approach, we learn robust representations that are generalizable across tasks without suffering from the problems of catastrophic forgetting and overfitting to accommodate future classes with limited samples. Our proposed method leverages the meta-learning framework with knowledge consolidation. The meta-learning framework helps the model for rapid learning when samples appear in an online fashion. Simultaneously, knowledge consolidation helps to learn a robust representation against forgetting under online updates to facilitate future learning. Our approach significantly outperforms other methods on several benchmarks. Mohammed Asad Karim, Vinay Kumar Verma, Pravendra Singh, Vinay P. Namboodiri, Piyush Rai |
IJCAI | 3 |
| 2021 | Improving Few-Shot Learning using Composite Rotation based Auxiliary TaskabstractIn this paper, we propose an approach to improve few-shot classification performance using a composite rotation based auxiliary task. Few-shot classification methods aim to produce neural networks that perform well for classes with a large number of training samples and classes with less number of training samples. They employ techniques to enable the network to produce highly discriminative features that are also very generic. Generally, the better the quality and generic-nature of the features produced by the network, the better is the performance of the network on few-shot learning. Our approach aims to train networks to produce such features by using a self-supervised auxiliary task. Our proposed composite rotation based auxiliary task performs rotation at two levels, i.e., rotation of patches inside the image (inner rotation) and rotation of the whole image (outer rotation) and assigns one out of 16 rotation classes to the modified image. We then simultaneously train for the composite rotation prediction task along with the original classification task, which forces the network to learn high-quality generic features that help improve the few-shot classification performance. We experimentally show that our approach performs better than existing few-shot learning methods on multiple benchmark datasets. Pratik Mazumder, Pravendra Singh, Vinay P. Namboodiri |
WACV | 2 |
| 2021 | RNNP: A Robust Few-Shot Learning ApproachabstractLearning from a few examples is an important practical aspect of training classifiers. Various works have examined this aspect quite well. However, all existing approaches assume that the few examples provided are always correctly labeled. This is a strong assumption, especially if one considers the current techniques for labeling using crowd-based labeling services. We address this issue by proposing a novel robust few-shot learning approach. Our method relies on generating robust prototypes from a set of few examples. Specifically, our method refines the class prototypes by producing hybrid features from the support examples of each class. The refined prototypes help to classify the query images better. Our method can replace the evaluation phase of any few-shot learning method that uses a nearest neighbor prototype-based evaluation procedure to make them robust. We evaluate our method on standard mini-ImageNet and tiered-ImageNet datasets. We perform experiments with various label corruption rates in the support examples of the few-shot classes. We obtain significant improvement over widely used few-shot learning methods that suffer significant performance degeneration in the presence of label noise. We finally provide extensive ablation experiments to validate our method. Pratik Mazumder, Pravendra Singh, Vinay P. Namboodiri |
WACV | 2 |
| 2021 | AVGZSLNet: Audio-Visual Generalized Zero-Shot Learning by Reconstructing Label Features from Multi-Modal EmbeddingsabstractIn this paper, we propose a novel approach for generalized zero-shot learning in a multi-modal setting, where we have novel classes ofaudio/video during testing that are not seen during training. We use the semantic relatedness oftext embeddings as a means for zero-shot learning by aligning audio and video embeddings with the corresponding class label text feature space. Our approach uses a cross-modal decoder and a composite triplet loss. The cross-modal decoder enforces a constraint that the class label text features can be reconstructed from the audio and video embeddings of data points. This helps the audio and video embeddings to move closer to the class label text embedding. The composite triplet loss makes use of the audio, video, and text embeddings. It helps bring the embeddings from the same class closer and push away the embeddings from different classes in a multi-modal setting. This helps the network to perform better on the multi-modal zero-shot learning task. Importantly, our multi-modal zero-shot learning approach works even if a modality is missing at test time. We test our approach on the generalized zero-shot classification and retrieval tasks and show that our approach outperforms other models in the presence of a single modality as well as in the presence of multiple modalities. We validate our approach by comparing it with previous approaches and using various ablations. Pratik Mazumder, Pravendra Singh, Kranti Kumar Parida, Vinay P. Namboodiri |
WACV | 2 |
| 2021 | Calibrating feature maps for deep CNNs
Pravendra Singh, Pratik Mazumder, Mohammed Asad Karim, Vinay P. Namboodiri |
Neurocomputing | 1 |
| 2020 | CPWC: Contextual Point Wise Convolution for Object RecognitionabstractConvolutional layers are a major driving force behind the successes of deep learning. Pointwise convolution (PWC) is a 1 × 1 convolutional filter that is primarily used for parameter reduction. However, the PWC ignores the spatial information around the points it is processing. This design is by choice, in order to reduce the overall parameters and computations. However, we hypothesize that this shortcoming of PWC has a significant impact on the network performance. We propose an alternative design for pointwise convolution, which uses spatial information from the input efficiently. Our design significantly improves the performance of the networks without substantially increasing the number of parameters and computations. We experimentally show that our design results in significant improvement in the performance of the network for classification as well as detection. Pratik Mazumder, Pravendra Singh, Vinay P. Namboodiri |
ICASSP | 2 |
| 2020 | Passive Batch Injection Training Technique: Boosting Network Performance by Injecting Mini-Batches from a different Data DistributionabstractThis work presents a novel training technique for deep neural networks that makes use of additional data from a distribution that is different from that of the original input data. This technique aims to reduce overfitting and improve the generalization performance of the network. Our proposed technique, namely Passive Batch Injection Training Technique (PBITT), even reduces the level of overfitting in networks that already use the standard techniques for reducing overfitting such as L2regularization and batch normalization, resulting in significant accuracy improvements. Passive Batch Injection Training Technique (PBITT) introduces a few passive mini-batches into the training process that contain data from a distribution that is different from the input data distribution. This technique does not increase the number of parameters in the final model and also does not increase the inference (test) time but still improves the performance of deep CNNs. To the best of our knowledge, this is the first work that makes use of different data distribution to aid the training of convolutional neural networks (CNNs). We thoroughly evaluate the proposed approach on standard architectures: VGG, ResNet, and WideResNet, and on several popular datasets: CIFAR-10, CIFAR-100, SVHN, and ImageNet. We observe consistent accuracy improvement by using the proposed technique. We also show experimentally that the model trained by our technique generalizes well to other tasks such as object detection on the MS-COCO dataset using Faster R-CNN. We present extensive ablations to validate the proposed approach. Our approach improves the accuracy of VGG-16 by a significant margin of 2.1% over the CIFAR-100 dataset. Pravendra Singh, Pratik Mazumder, Vinay P. Namboodiri |
IJCNN | 1 |
| 2020 | SkipConv: Skip Convolution for Computationally Efficient Deep CNNsabstractConvolution operation in deep convolutional neural networks is the most computationally expensive as compared to other operations. Most of the model computation (FLOPS) in the deep architecture belong to convolution operation. In this paper, we are proposing a novel skip convolution operation that employs significantly fewer computation as compared to the traditional one without sacrificing model accuracy. Skip convolution operation produces structured sparsity in the output feature maps without requiring sparsity in the model parameters for computation reduction. The existing convolution operation performs the redundant computation for object feature representation while the proposed convolution skips redundant computation. Our empirical evaluation for various deep models (VGG, ResNet, MobileNet, and Faster R-CNN) over various benchmarked datasets (CIFAR-10, CIFAR-100, ImageNet, and MS-COCO) show that skip convolution reduces the computation significantly while preserving feature representational capacity. The proposed approach is model-agnostic and can be applied over any architecture. The proposed approach does not require a pretrained model and does train from scratch. Hence we achieve significant computation reduction at training and test time. We are also able to reduce computation in an already compact model such as MobileNet using skip convolution. We also show empirically that the proposed convolution works well for other tasks such as object detection. Therefore, SkipConv can be a widely usable and efficient way of reducing computation in deep CNN models. Pravendra Singh, Vinay P. Namboodiri |
IJCNN | 1 |
| 2020 | Calibrating CNNs for Lifelong LearningabstractWe present an approach for lifelong/continual learning of convolutional neural networks (CNN) that does not suffer from the problem of catastrophic forgetting when moving from one task to the other. We show that the activation maps generated by the CNN trained on the old task can be calibrated using very few calibration parameters, to become relevant to the new task. Based on this, we calibrate the activation maps produced by each network layer using spatial and channel-wise calibration modules and train only these calibration parameters for each new task in order to perform lifelong learning. Our calibration modules introduce significantly less computation and parameters as compared to the approaches that dynamically expand the network. Our approach is immune to catastrophic forgetting since we store the task-adaptive calibration parameters, which contain all the task-specific knowledge and is exclusive to each task. Further, our approach does not require storing data samples from the old tasks, which is done by many replay based methods. We perform extensive experiments on multiple benchmark datasets (SVHN, CIFAR, ImageNet, and MS-Celeb), all of which show substantial improvements over state-of-the-art methods (e.g., a 29% absolute increase in accuracy on CIFAR-100 with 10 classes at a time). On large-scale datasets, our approach yields 23.8% and 9.7% absolute increase in accuracy on ImageNet-100 and MS-Celeb-10K datasets, respectively, by employing very few (0.51% and 0.35% of model parameters) task-adaptive calibration parameters. Pravendra Singh, Vinay Kumar Verma, Pratik Mazumder, Lawrence Carin, Piyush Rai |
NeurIPS | 1 |
| 2020 | Accuracy Booster: Performance Boosting using Feature Map Re-calibrationabstractConvolution Neural Networks (CNN) have been extremely successful in solving intensive computer vision tasks. The convolutional filters used in CNNs have played a major role in this success, by extracting useful features from the inputs. Recently researchers have tried to boost the performance of CNNs by re-calibrating the feature maps produced by these filters, e.g., Squeeze-and-Excitation Networks (SENets). These approaches have achieved better performance by Exciting up the important channels or feature maps while diminishing the rest. However, in the process, architectural complexity has increased. We propose an architectural block that introduces much lower complexity than the existing methods of CNN performance boosting while performing significantly better than them. We carry out experiments on the CIFAR, ImageNet and MS-COCO datasets, and show that the proposed block can challenge the state-of-the-art results. Our method boosts the ResNet-50 architecture to perform comparably to the ResNet-152 architecture, which is a three times deeper network, on classification. We also show experimentally that our method is not limited to classification but also generalizes well to other tasks such as object detection. Pravendra Singh, Pratik Mazumder, Vinay P. Namboodiri |
WACV | 1 |
| 2020 | Cooperative Initialization based Deep Neural Network TrainingabstractResearchers have proposed various activation functions. These activation functions help the deep network to learn non-linear behavior with a significant effect on training dynamics and task performance. The performance of these activations also depends on the initial state of the weight parameters, i.e., different initial state leads to a difference in the performance of a network. In this paper, we have proposed a cooperative initialization for training the deep network using ReLU activation function to improve the network performance. Our approach uses multiple activation functions in the initial few epochs for the update of all sets of weight parameters while training the network. These activation functions cooperate to overcome their drawbacks in the update of weight parameters, which in effect learn better "feature representation" and boost the network performance later. Cooperative initialization based training also helps in reducing the overfitting problem and does not increase the number of parameters, inference (test) time in the final model while improving the performance. Experiments show that our approach outperforms various baselines and, at the same time, performs well over various tasks such as classification and detection. The Top-1 classification accuracy of the model trained using our approach improves by 2.8% for VGG-16 and 2.1% for ResNet-56 on CIFAR-100 dataset. Pravendra Singh, Munender Varshney, Vinay P. Namboodiri |
WACV | 1 |
| 2020 | Leveraging Filter Correlations for Deep Model CompressionabstractWe present a filter correlation based model compression approach for deep convolutional neural networks. Our approach iteratively identifies pairs of filters with the largest pairwise correlations and drops one of the filters from each such pair. However, instead of discarding one of the filters from each such pair naïvely, the model is re-optimized to make the filters in these pairs maximally correlated, so that discarding one of the filters from the pair results in minimal information loss. Moreover, after discarding the filters in each round, we further finetune the model to recover from the potential small loss incurred by the compression. We evaluate our proposed approach using a comprehensive set of experiments and ablation studies. Our compression method yields state-of-the-art FLOPs compression rates on various benchmarks, such as LeNet-5, VGG-16, and ResNet-50,56, while still achieving excellent predictive performance for tasks such as object detection on benchmark datasets. Pravendra Singh, Vinay Kumar Verma, Piyush Rai, Vinay P. Namboodiri |
WACV | 1 |
| 2020 | A "Network Pruning Network" Approach to Deep Model CompressionabstractWe present a filter pruning approach for deep model compression, using a multitask network. Our approach is based on learning a a pruner network to prune a pre-trained target network. The pruner is essentially a multitask deep neural network with binary outputs that help identify the filters from each layer of the original network that do not have any significant contribution to the model and can therefore be pruned. The pruner network has the same architecture as the original network except that it has a multitask/multi-output last layer containing binary-valued outputs (one per filter), which indicate which filters have to be pruned. The pruner's goal is to minimize the number of filters from the original network by assigning zero weights to the corresponding output feature-maps. In contrast to most of the existing methods, instead of relying on iterative pruning, our approach can prune the network (original network) in one go and, moreover, does not require specifying the degree of pruning for each layer (and can learn it instead). The compressed model produced by our approach is generic and does not need any special hardware/software support. Moreover, augmenting with other methods such as knowledge distillation, quantization, and connection pruning can increase the degree of compression for the proposed approach. We show the efficacy of our proposed approach for classification and object detection tasks. Vinay Kumar Verma, Pravendra Singh, Vinay P. Namboodiri, Piyush Rai |
WACV | 2 |
| 2020 | HetConv: Beyond Homogeneous Convolution Kernels for Deep CNNs
Pravendra Singh, Vinay Kumar Verma, Piyush Rai, Vinay P. Namboodiri |
Int. J. Comput. Vis. | 1 |
| 2020 | GIFSL - grafting based improved few-shot learning
Pratik Mazumder, Pravendra Singh, Vinay P. Namboodiri |
Image Vis. Comput. | 2 |
| 2020 | FALF ConvNets: Fatuous auxiliary loss based filter-pruning for efficient deep CNNs
Pravendra Singh, Vinay Sameer Raja Kadi, Vinay P. Namboodiri |
Image Vis. Comput. | 1 |
| 2020 | EDS pooling layer
Pravendra Singh, Prem Raj, Vinay P. Namboodiri |
Image Vis. Comput. | 1 |
| 2019 | HetConv: Heterogeneous Kernel-Based Convolutions for Deep CNNsabstractWe present a novel deep learning architecture in which the convolution operation leverages heterogeneous kernels. The proposed HetConv (Heterogeneous Kernel-Based Convolution) reduces the computation (FLOPs) and the number of parameters as compared to standard convolution operation while still maintaining representational efficiency. To show the effectiveness of our proposed convolution, we present extensive experimental results on the standard convolutional neural network (CNN) architectures such as VGG and ResNet. We find that after replacing the standard convolutional filters in these architectures with our proposed HetConv filters, we achieve 3X to 8X FLOPs based improvement in speed while still maintaining (and sometimes improving) the accuracy. We also compare our proposed convolutions with group/depth wise convolutions and show that it achieves more FLOPs reduction with significantly higher accuracy. Pravendra Singh, Vinay Kumar Verma, Piyush Rai, Vinay P. Namboodiri |
CVPR | 1 |
| 2019 | Play and Prune: Adaptive Filter Pruning for Deep Model CompressionabstractWhile convolutional neural networks (CNN) have achieved impressive performance on various classification/recognition tasks, they typically consist of a massive number of parameters. This results in significant memory requirement as well as computational overheads. Consequently, there is a growing need for filter-level pruning approaches for compressing CNN based models that not only reduce the total number of parameters but reduce the overall computation as well. We present a new min-max framework for filter-level pruning of CNNs. Our framework, called Play and Prune (PP), jointly prunes and fine-tunes CNN model parameters, with an adaptive pruning rate, while maintaining the model's predictive performance. Our framework consists of two modules: (1) An adaptive filter pruning (AFP) module, which minimizes the number of filters in the model; and (2) A pruning rate controller (PRC) module, which maximizes the accuracy during pruning. Moreover, unlike most previous approaches, our approach allows directly specifying the desired error tolerance instead of pruning level. Our compressed models can be deployed at run-time, without requiring any special libraries or hardware. Our approach reduces the number of parameters of VGG-16 by an impressive factor of 17.5X, and number of FLOPS by 6.43X, with no loss of accuracy, significantly outperforming other state-of-the-art filter pruning methods. Pravendra Singh, Vinay Kumar Verma, Piyush Rai, Vinay P. Namboodiri |
IJCAI | 1 |
| 2019 | Stability Based Filter Pruning for Accelerating Deep CNNsabstractConvolutional neural networks (CNN) have achieved impressive performance on the wide variety of tasks (classification, detection, etc.) across multiple domains at the cost of high computational and memory requirements. Thus, leveraging CNNs for real-time applications necessitates model compression approaches that not only reduce the total number of parameters but reduce the overall computation as well. In this work, we present a stability-based approach for filter-level pruning of CNNs. We evaluate our proposed approach on different architectures (LeNet, VGG-16, ResNet, and Faster RCNN) and datasets and demonstrate its generalizability through extensive experiments. Moreover, our compressed models can be used at run-time without requiring any special libraries or hardware. Our model compression method reduces the number of FLOPS by an impressive factor of 6.03X and GPU memory footprint by more than 17X, significantly outperforming other state-of-the-art filter pruning methods. Pravendra Singh, Vinay Sameer Raja Kadi, Nikhil Verma, Vinay P. Namboodiri |
WACV | 1 |
| 2019 | Multi-Layer Pruning Framework for Compressing Single Shot MultiBox DetectorabstractWe propose a framework for compressing state-of-the-art Single Shot MultiBox Detector (SSD). The framework addresses compression in the following stages: Sparsity Induction, Filter Selection, and Filter Pruning. In the Sparsity Induction stage, the object detector model is sparsified via an improved global threshold. In Filter Selection & Pruning stage, we select and remove filters using sparsity statistics of filter weights in two consecutive convolutional layers. This results in the model with the size smaller than most existing compact architectures. We evaluate the performance of our framework with multiple datasets and compare over multiple methods. Experimental results show that our method achieves state-of-the-art compression of 6.7X and 4.9X on PASCAL VOC dataset on models SSD300 and SSD512 respectively. We further show that the method produces maximum compression of 26X with SSD512 on German Traffic Sign Detection Benchmark (GTSDB). Additionally, we also empirically show our method's adaptability for classification based architecture VGG16 on datasets CIFAR and German Traffic Sign Recognition Benchmark (GTSRB) achieving a compression rate of 125X and 200X with the reduction in flops by 90.50% and 96.6% respectively with no loss of accuracy. In addition to this, our method does not require any special libraries or hardware support for the resulting compressed models. Pravendra Singh, Manikandan R, Neeraj Matiyali, Vinay P. Namboodiri |
WACV | 1 |