VLDB 2026 Research / reviewers in the wild / expert
Rahul Bhotika
dblp:28/2609
· DBLP profile ↗
20ranked-venue papers
2as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DPL: Diverse Preference Learning Without A Reference ModelabstractAbhijnan Nath, Andrey Volozin, Saumajit Saha, Albert Aristotle Nanda, Galina Grunin, Rahul Bhotika, Nikhil Krishnaswamy. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Abhijnan Nath, Andrey Volozin, Saumajit Saha, Albert Nanda, Galina Grunin, Rahul Bhotika, Nikhil Krishnaswamy |
NAACL (Long Papers) | 6 |
| 2023 | Masked Vision and Language Modeling for Multi-modal Representation Learning
Gukyeong Kwon, Zhaowei Cai, Avinash Ravichandran, Erhan Bas, Rahul Bhotika, Stefano Soatto |
ICLR | 5 |
| 2023 | Relaxing Contrastiveness in Multimodal Representation LearningabstractMultimodal representation learning for images with paired raw texts can improve the usability and generality of the learned semantic concepts while significantly reducing annotation costs. In this paper, we explore the design space of loss functions in visual-linguistic pretraining frameworks and propose a novel Relaxed Contrastive (ReCo) objective, which act as a drop-in replacement of the widely used InfoNCE loss. The key insight of ReCo is to allow a relaxed negative space by not penalizing unpaired multimodal samples (i.e., negative pairs) that are already orthogonal or negatively correlated. Unlike the widely-used InfoNCE, which keeps repelling negative pairs as long as they are not anti-correlated, ReCo by design embraces more diversity and flexibility of the learned embeddings. We conduct exten-sive experiments using ReCo with state-of-the-art models by pretraining on the MIMIC-CXR dataset that consists of chest radiographs and free-text radiology reports, and eval-uating on the CheXpert dataset for multimodal retrieval and disease classification. Our ReCo achieves an absolute improvement of 2.9% over the InfoNCE baseline on the CheXpert Retrieval dataset in average retrieval precision and re-ports better or comparable performance in the linear evaluation and finetuning for classification. We further show that ReCo outperforms InfoNCE on the Flickr30K dataset by 1.7% in retrieval Recall@1, demonstrating the generalizability of our approach to natural images. Zudi Lin, Erhan Bas, Kunwar Yashraj Singh, Gurumurthy Swaminathan, Rahul Bhotika |
WACV | 5 |
| 2022 | Task Adaptive Parameter Sharing for Multi-Task LearningabstractAdapting pre-trained models with broad capabilities has become standard practice for learning a wide range of downstream tasks. The typical approach of fine-tuning different models for each task is performant, but incurs a substantial memory cost. To efficiently learn multiple down-stream tasks we introduce Task Adaptive Parameter Sharing (TAPS), a simple method for tuning a base model to a new task by adaptively modifying a small, task-specific subset of layers. This enables multi-task learning while minimizing the resources used and avoids catastrophic forgetting and competition between tasks. TAPS solves a joint optimization problem which determines both the layers that are shared with the base model and the value of the task-specific weights. Further, a sparsity penalty on the number of active layers promotes weight sharing with the base model. Compared to other methods, TAPS retains a high accuracy on the target tasks while still introducing only a small number of task-specific parameters. Moreover, TAPS is agnostic to the particular architecture used and requires only minor changes to the training scheme. We evaluate our method on a suite of fine-tuning tasks and architectures (ResNet, DenseNet, ViT) and show that it achieves state-of-the-art performance while being simple to implement. Matthew Wallingford, Alessandro Achille, Avinash Ravichandran, Charless C. Fowlkes, Rahul Bhotika, Stefano Soatto |
CVPR | 6 |
| 2022 | Class-Incremental Learning with Strong Pre-trained ModelsabstractClass-incremental learning (CIL) has been widely stud-ied under the setting of starting from a small number of classes (base classes). Instead, we explore an understud-ied real-world setting of CIL that starts with a strong model pre-trained on a large number of base classes. We hypoth-esize that a strong base model can provide a good repre-sentation for novel classes and incremental learning can be done with small adaptations. We propose a 2-stage training scheme, i) feature augmentation - cloning part of the backbone and fine-tuning it on the novel data, and ii) fusion - combining the base and novel classifiers into a unified classifier. Experiments show that the proposed method sig-nificantly outperforms state-of-the-art CIL methods on the large-scale ImageNet dataset (e.g. + 10% overall accuracy than the best). We also propose and analyze understudied practical CIL scenarios, such as base-novel overlap with distribution shift. Our proposed method is robust and gen-eralizes to all analyzed CIL settings. Tz-Ying Wu, Gurumurthy Swaminathan, Zhizhong Li 0001, Avinash Ravichandran, Nuno Vasconcelos, Rahul Bhotika, Stefano Soatto |
CVPR | 6 |
| 2022 | X-DETR: A Versatile Architecture for Instance-wise Vision-Language Tasks
Zhaowei Cai, Gukyeong Kwon, Avinash Ravichandran, Erhan Bas, Zhuowen Tu, Rahul Bhotika, Stefano Soatto |
ECCV (36) | 6 |
| 2022 | Semi-supervised Vision Transformers at ScaleabstractWe study semi-supervised learning (SSL) for vision transformers (ViT), an under-explored topic despite the wide adoption of the ViT architectures to different tasks. To tackle this problem, we use a SSL pipeline, consisting of first un/self-supervised pre-training, followed by supervised fine-tuning, and finally semi-supervised fine-tuning. At the semi-supervised fine-tuning stage, we adopt an exponential moving average (EMA)-Teacher framework instead of the popular FixMatch, since the former is more stable and delivers higher accuracy for semi-supervised vision transformers. In addition, we propose a probabilistic pseudo mixup mechanism to interpolate unlabeled samples and their pseudo labels for improved regularization, which is important for training ViTs with weak inductive bias. Our proposed method, dubbed Semi-ViT, achieves comparable or better performance than the CNN counterparts in the semi-supervised classification setting. Semi-ViT also enjoys the scalability benefits of ViTs that can be readily scaled up to large-size models with increasing accuracy. For example, Semi-ViT-Huge achieves an impressive 80\% top-1 accuracy on ImageNet using only 1\% labels, which is comparable with Inception-v4 using 100\% ImageNet labels. The code is available at https://github.com/amazon-science/semi-vit. Zhaowei Cai, Avinash Ravichandran, Paolo Favaro, Manchen Wang, Davide Modolo, Rahul Bhotika, Zhuowen Tu, Stefano Soatto |
NeurIPS | 6 |
| 2021 | End-to-end Piece-wise Unwarping of Document ImagesabstractDocument unwarping attempts to undo physical deformations of the paper and recover a ’flatbed’ scanned document-image for downstream tasks such as OCR. Current state-of-the-art relies on global unwarping of the document which is not robust to local deformation changes. Moreover, a global unwarping often produces spurious warping artifacts in less warped regions to compensate for severe warps present in other parts of the document. In this paper, we propose the first end-to-end trainable piece-wise unwarping1method that predicts local deformation fields and stitches them together with global information to obtain an improved unwarping. The proposed piece-wise formulation results in 4% improvement in terms of multi-scale structural similarity (MS-SSIM) and shows better performance in terms of OCR metrics, character error rate (CER) and word error rate (WER) compared to the state-of-the-art. Sagnik Das, Kunwar Yashraj Singh, Jon Wu, Erhan Bas, Vijay Mahadevan, Rahul Bhotika, Dimitris Samaras |
ICCV | 6 |
| 2021 | Estimating informativeness of samples with Smooth Unique Information
Hrayr Harutyunyan, Alessandro Achille, Giovanni Paolini, Orchid Majumder, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto |
ICLR | 6 |
| 2020 | Incremental Few-Shot Meta-learning via Indirect Discriminant Alignment
Orchid Majumder, Alessandro Achille, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto |
ECCV (7) | 5 |
| 2020 | Rethinking the Hyperparameters for Fine-tuning
Pratik Chaudhari, Hao Yang 0043, Michael Lam, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto |
ICLR | 6 |
| 2020 | Predicting Training Time Without TrainingabstractWe tackle the problem of predicting the number of optimization steps that a pre-trained deep network needs to converge to a given value of the loss function. To do so, we leverage the fact that the training dynamics of a deep network during fine-tuning are well approximated by those of a linearized model. This allows us to approximate the training loss and accuracy at any point during training by solving a low-dimensional Stochastic Differential Equation (SDE) in function space. Using this result, we are able to predict the time it takes for Stochastic Gradient Descent (SGD) to fine-tune a model to a given loss without having to perform any training. In our experiments, we are able to predict training time of a ResNet within a 20\% error margin on a variety of datasets and hyper-parameters, at a 30 to 45-fold reduction in cost compared to actual training. We also discuss how to further reduce the computational and memory cost of our method, and in particular we show that by exploiting the spectral properties of the gradients' matrix it is possible to predict training time on a large dataset while processing only a subset of the samples. Luca Zancato, Alessandro Achille, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto |
NeurIPS | 4 |
| 2019 | Few-Shot Learning With Embedded Class Models and Shot-Free Meta TrainingabstractWe propose a method for learning embeddings for few-shot learning that is suitable for use with any number of shots (shot-free). Rather than fixing the class prototypes to be the Euclidean average of sample embeddings, we allow them to live in a higher-dimensional space (embedded class models) and learn the prototypes along with the model parameters. The class representation function is defined implicitly, which allows us to deal with a variable number of shots per class with a simple constant-size architecture. The class embedding encompasses metric learning, that facilitates adding new classes without crowding the class representation space. Despite being general and not tuned to the benchmark, our approach achieves state-of-the-art performance on the standard few-shot benchmark datasets. Avinash Ravichandran, Rahul Bhotika, Stefano Soatto |
ICCV | 2 |
| 2009 | Nonparametric Intensity Priors for Level Set Segmentation of Low Contrast Structures
Sokratis Makrogiannis, Rahul Bhotika, James V. Miller, John Skinner, Melissa Vass |
MICCAI (1) | 2 |
| 2007 | A Probabilistic Model for Haustral Curvatures with Applications to Colon CAD
John Melonakos, Paulo R. S. Mendonça, Rahul Bhotika, Saad A. Sirohey |
MICCAI (2) | 3 |
| 2006 | Joint Recognition of Complex Events and Track MatchingabstractWe present a novel method for jointly performing recognition of complex events and linking fragmented tracks into coherent, long-duration tracks. Many event recognition methods require highly accurate tracking, and may fail when tracks corresponding to event actors are fragmented or partially missing. However, these conditions occur frequently from occlusions, traffic and tracking errors. Recently, methods have been proposed for linking track fragments from multiple objects under these difficult conditions. Here, we develop a method for solving these two problems jointly. A hypothesized event model, represented as a Dynamic Bayes Net, supplies data-driven constraints on the likelihood of proposed track fragment matches. These event-guided constraints are combined with appearance and kinematic constraints used in the previous track linking formulation. The result is the most likely track linking solution given the event model, and the highest event score given all of the track fragments. The event model with the highest score is determined to have occurred, if the score exceeds a threshold. Results demonstrated on a busy scene of airplane servicing activities, where many non-event movers and long fragmented tracks are present, show the promise of the approach to solving the joint problem. Michael T. Chan, Anthony Hoogs, Rahul Bhotika, A. G. Amitha Perera, John Schmiederer, Gianfranco Doretto |
CVPR (2) | 3 |
| 2006 | Part-Based Local Shape Models for Colon Polyp Detection
Rahul Bhotika, Paulo R. S. Mendonça, Saad A. Sirohey, Wesley D. Turner, Ying-Lin Lee, Julie M. McCoy, Rebecca E. B. Brown, James V. Miller |
MICCAI (2) | 1 |
| 2005 | Model-Based Analysis of Local Shape for Lesion Detection in CT Scans
Paulo R. S. Mendonça, Rahul Bhotika, Saad A. Sirohey, Wesley D. Turner, James V. Miller, Ricardo S. Avila |
MICCAI | 2 |
| 2002 | A Probabilistic Theory of Occupancy and Emptiness
Rahul Bhotika, David J. Fleet, Kiriakos N. Kutulakos |
ECCV (3) | 1 |
| 2001 | Flexible flow for 3D nonrigid tracking and shape recoveryabstractWe introduce linear methods for model-based tracking of nonrigid 3D objects and for acquiring such models from video. 3D motions and flexions are calculated directly from image intensities without information-lossy intermediate results. Measurement uncertainty is quantified and fully propagated through the inverse model to yield posterior mean (PM) and/or mode (MAP) pose estimates. A Bayesian framework manages uncertainty, accommodates priors, and gives confidence measures. We obtain highly accurate and robust closed-form estimators by minimizing information loss from non-reversible (inner-product and least-squares) operations, and, when unavoidable, performing such operations with the appropriate error norm. For model acquisition, we show how to refine a crude or generic model to fit the video subject. We demonstrate with tracking, model refinement, and super-resolution texture lifting from low-quality low-resolution video. Matthew Brand, Rahul Bhotika |
CVPR (1) | 2 |