VLDB 2026 Research / reviewers in the wild / expert
Suraj Kothawade
dblp:220/3896
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2024
0000-0001-9405-6862ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 19% Learning paradigms · 16% Image recognition and object detection · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 20 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms
class imbalance |
0.8 | 1 | 2024 | SCoRe: Submodular Combinatorial Representation Learning · ICML 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | SCoRe: Submodular Combinatorial Representation Learning · ICML 2024 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
0.8 | 1 | 2024 | Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model › personalized image generation
subject-driven text-to-image generation |
0.8 | 1 | 2024 | Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning · NeurIPS 2024 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.7 | 1 | 2023 | DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation · ACL (1) 2023 |
Machine learning › Efficient and distributed learning
data-efficient learning |
0.7 | 1 | 2023 | DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation · ACL (1) 2023 |
Computer vision › Image recognition and object detection › object detection › detector training
active learning for object detection |
0.6 | 1 | 2022 | Talisman: Targeted Active Learning for Object Detection with Rare Classes and Slices Using Submodular Mutual Information · ECCV (38) 2022 |
Machine learning › Transfer learning and domain adaptation
few-shot classification |
0.6 | 1 | 2022 | PLATINUM: Semi-Supervised Model Agnostic Meta-Learning using Submodular Mutual Information · ICML 2022 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.6 | 1 | 2022 | PLATINUM: Semi-Supervised Model Agnostic Meta-Learning using Submodular Mutual Information · ICML 2022 |
Computer vision › Image recognition and object detection
object detection |
0.6 | 1 | 2022 | Talisman: Targeted Active Learning for Object Detection with Rare Classes and Slices Using Submodular Mutual Information · ECCV (38) 2022 |
Machine learning › Learning paradigms › class imbalance
rare class detection |
0.6 | 1 | 2022 | Talisman: Targeted Active Learning for Object Detection with Rare Classes and Slices Using Submodular Mutual Information · ECCV (38) 2022 |
Data mining
pattern mining |
0.6 | 1 | 2022 | PRISM: A Rich Class of Parameterized Submodular Information Measures for Guided Data Subset Selection · AAAI 2022 |
Machine learning › Efficient and distributed learning
active learning |
0.5 | 1 | 2021 | SIMILAR: Submodular Information Measures Based Active Learning In Realistic Scenarios · NeurIPS 2021 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning · NeurIPS 2024 |
Mathematical optimization
submodular optimization |
0.2 | 1 | 2024 | SCoRe: Submodular Combinatorial Representation Learning · ICML 2024 |
Machine learning › Learning paradigms
semi-supervised learning |
0.2 | 1 | 2022 | PLATINUM: Semi-Supervised Model Agnostic Meta-Learning using Submodular Mutual Information · ICML 2022 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 1 | 2021 | SIMILAR: Submodular Information Measures Based Active Learning In Realistic Scenarios · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.1 | 1 | 2021 | SIMILAR: Submodular Information Measures Based Active Learning In Realistic Scenarios · NeurIPS 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.1 | 1 | 2021 | SIMILAR: Submodular Information Measures Based Active Learning In Realistic Scenarios · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
submodular information measures · 2.0contrastive loss · 1.5submodular mutual information · 1.1λ-harmonic reward · 0.8u-net fine-tuning · 0.8reward preference optimization · 0.8bradley-terry preference model · 0.8subset selection · 0.7fairness-aware selection · 0.7submodular optimization · 0.6active learning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SCoRe: Submodular Combinatorial Representation LearningabstractIn this paper we introduce the **SCoRe** (**S**ubmodular **Co**mbinatorial **Re**presentation Learning) framework, a novel approach in representation learning that addresses inter-class bias and intra-class variance. SCoRe provides a new combinatorial viewpoint to representation learning, by introducing a family of loss functions based on set-based submodular information measures. We develop two novel combinatorial formulations for loss functions, using the *Total Information* and *Total Correlation*, that naturally minimize intra-class variance and inter-class bias. Several commonly used metric/contrastive learning loss functions like supervised contrastive loss, orthogonal projection loss, and N-pairs loss, are all instances of SCoRe, thereby underlining the versatility and applicability of SCoRe in a broad spectrum of learning scenarios. Novel objectives in SCoRe naturally model class-imbalance with up to 7.6% improvement in classification on CIFAR-10-LT, CIFAR-100-LT, MedMNIST, 2.1% on ImageNet-LT, and 19.4% in object detection on IDD and LVIS (v1.0), demonstrating its effectiveness over existing approaches. Anay Majee, Suraj Kothawade, KrishnaTeja Killamsetty, Rishabh Iyer 0001 |
ICML | 2 |
| 2024 | Subject-driven Text-to-Image Generation via Preference-based Reinforcement LearningabstractText-to-image generative models have recently attracted considerable interest, enabling the synthesis of high-quality images from textual prompts. However, these models often lack the capability to generate specific subjects from given reference images or to synthesize novel renditions under varying conditions. Methods like DreamBooth and Subject-driven Text-to-Image (SuTI) have made significant progress in this area. Yet, both approaches primarily focus on enhancing similarity to reference images and require expensive setups, often overlooking the need for efficient training and avoiding overfitting to the reference images. In this work, we present the $\lambda$-Harmonic reward function, which provides a reliable reward signal and enables early stopping for faster training and effective regularization. By combining the Bradley-Terry preference model, the $\lambda$-Harmonic reward function also provides preference labels for subject-driven generation tasks. We propose Reward Preference Optimization (RPO), which offers a simpler setup (requiring only 3\% of the negative samples used by DreamBooth) and fewer gradient steps for fine-tuning. Unlike most existing methods, our approach does not require training a text encoder or optimizing text embeddings and achieves text-image alignment by fine-tuning only the U-Net component. Empirically, $\lambda$-Harmonic proves to be a reliable approach for model selection in subject-driven generation tasks. Based on preference labels and early stopping validation from the $\lambda$-Harmonic reward function, our algorithm achieves a state-of-the-art CLIP-I score of 0.833 and a CLIP-T score of 0.314 on DreamBench. Yanting Miao, William Loh, Suraj Kothawade, Pascal Poupart, Abdullah Rashwan, Yeqing Li |
NeurIPS | 3 |
| 2024 | Beyond Active Learning: Leveraging the Full Potential of Human Interaction via Auto-Labeling, Human Correction, and Human VerificationabstractActive Learning (AL) is a human-in-the-loop framework to interactively and adaptively label data instances, thereby enabling significant gains in model performance compared to random sampling. AL approaches function by selecting the hardest instances to label, often relying on notions of diversity and uncertainty. However, we believe that these current paradigms of AL do not leverage the full potential of human interaction granted by automated label suggestions. Indeed, we show that for many classification tasks and datasets, most people verifying if an automatically suggested label is correct take 3× to 4× less time than they do changing an incorrect suggestion to the correct label (or labeling from scratch without any suggestion). Utilizing this result, we propose Clarifier (aCtive LeARnIng From tIEred haRdness), an Interactive Learning framework that admits more effective use of human interaction by leveraging the reduced cost of verification. By targeting the hard (uncertain) instances with existing AL methods, the intermediate instances with a novel label suggestion scheme using submodular mutual information functions on a per-class basis, and the easy (confident) instances with highest-confidence auto-labeling, Clarifier can improve over the performance of existing AL approaches on multiple datasets – particularly on those that have a large number of classes – by almost 1.5× to 2× in terms of relative labeling cost. Nathan Beck, KrishnaTeja Killamsetty, Suraj Kothawade, Rishabh Iyer 0001 |
WACV | 3 |
| 2023 | DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent AdaptationabstractSuraj Kothawade, Anmol Mekala, D.Chandra Sekhara Hetha Havya, Mayank Kothyari, Rishabh Iyer, Ganesh Ramakrishnan, Preethi Jyothi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Suraj Kothawade, Anmol Reddy Mekala, D. Chandra Sekhara Hetha Havya, Mayank Kothyari, Rishabh Iyer 0001, Ganesh Ramakrishnan, Preethi Jyothi |
ACL (1) | 1 |
| 2022 | PRISM: A Rich Class of Parameterized Submodular Information Measures for Guided Data Subset SelectionabstractWith ever-increasing dataset sizes, subset selection techniques are becoming increasingly important for a plethora of tasks. It is often necessary to guide the subset selection to achieve certain desiderata, which includes focusing or targeting certain data points, while avoiding others. Examples of such problems include: i)targeted learning, where the goal is to find subsets with rare classes or rare attributes on which the model is under performing, and ii)guided summarization, where data (e.g., image collection, text, document or video) is summarized for quicker human consumption with specific additional user intent. Motivated by such applications, we present PRISM, a rich class of PaRameterIzed Submodular information Measures. Through novel functions and their parameterizations, PRISM offers a variety of modeling capabilities that enable a trade-off between desired qualities of a subset like diversity or representation and similarity/dissimilarity with a set of data points. We demonstrate how PRISM can be applied to the two real-world problems mentioned above, which require guided subset selection. In doing so, we show that PRISM interestingly generalizes some past work, therein reinforcing its broad utility. Through extensive experiments on diverse datasets, we demonstrate the superiority of PRISM over the state-of-the-art in targeted learning and in guided image-collection summarization. PRISM is available as a part of the SUBMODLIB (https://github.com/decile-team/submodlib) and TRUST (https://github.com/decile-team/trust) toolkits. Suraj Kothawade, Vishal Kaushal, Ganesh Ramakrishnan, Jeff A. Bilmes, Rishabh Iyer 0001 |
AAAI | 1 |
| 2022 | Talisman: Targeted Active Learning for Object Detection with Rare Classes and Slices Using Submodular Mutual Information
Suraj Kothawade, Saikat Ghosh, Rishabh Iyer 0001 |
ECCV (38) | 1 |
| 2022 | PLATINUM: Semi-Supervised Model Agnostic Meta-Learning using Submodular Mutual InformationabstractFew-shot classification (FSC) requires training models using a few (typically one to five) data points per class. Meta-learning has proven to be able to learn a parametrized model for FSC by training on various other classification tasks. In this work, we propose PLATINUM (semi-suPervised modeL Agnostic meTa learnIng usiNg sUbmodular Mutual information ), a novel semi-supervised model agnostic meta learning framework that uses the submodular mutual in- formation (SMI) functions to boost the perfor- mance of FSC. PLATINUM leverages unlabeled data in the inner and outer loop using SMI func- tions during meta-training and obtains richer meta- learned parameterizations. We study the per- formance of PLATINUM in two scenarios - 1) where the unlabeled data points belong to the same set of classes as the labeled set of a cer- tain episode, and 2) where there exist out-of- distribution classes that do not belong to the la- beled set. We evaluate our method on various settings on the miniImageNet, tieredImageNet and CIFAR-FS datasets. Our experiments show that PLATINUM outperforms MAML and semi- supervised approaches like pseduo-labeling for semi-supervised FSC, especially for small ratio of labeled to unlabeled samples. Changbin Li, Suraj Kothawade, Feng Chen 0001, Rishabh Iyer 0001 |
ICML | 2 |
| 2022 | Object-Level Targeted Selection via Deep Template MatchingabstractRetrieving images with objects that are semantically similar to objects of interest (OOI) in a query image has many practical use cases. A few examples include fixing failures like false negatives/positives of a learned model or mitigating class imbalance in a dataset. The targeted selection task requires finding the relevant data from a large-scale pool of unlabeled data. Manual mining at this scale is infeasible. Further, the OOI are often small and occupy less than 1% of image area, are occluded, and co-exist with many semantically different objects in cluttered scenes. Existing semantic image retrieval methods often focus on mining for larger sized geographical landmarks, and/or require extra labeled data, such as images/image-pairs with similar objects, for mining images with generic objects. We propose a fast and robust template matching algorithm in the DNN feature space, that retrieves semantically similar images at the object-level from a large unlabeled pool of data. We project the region(s) around the OOI in the query image to the DNN feature space for use as the template. This enables our method to focus on the semantics of the OOI without requiring extra labeled data. In the context of autonomous driving, we evaluate our system for targeted selection by using failure cases of object detectors as OOI. We demonstrate its efficacy on a large unlabeled dataset with 2.2M images and show high recall in mining for images with small-sized OOI. We compare our method against a well-known semantic image retrieval method, which also does not require extra labeled data. Lastly, we show that our method is flexible and retrieves images with one or more semantically different co-occurring OOI seamlessly. Suraj Kothawade, Donna Roy, Michele Fenzi, Elmar Haussmann, José M. Álvarez 0004, Christoph Angerer |
IV | 1 |
| 2021 | Robotic Lime Picking by Considering Leaves as Permeable ObstaclesabstractThe problem of robotic lime picking is challenging; lime plants have dense foliage which makes it difficult for a robotic arm to grasp a lime without coming in contact with leaves. Existing approaches either do not consider leaves, or treat them as obstacles and completely avoid them, often resulting in undesirable or infeasible plans. We focus on reaching a lime in the presence of dense foliage by considering the leaves of a plant as permeable obstacles with a collision cost. We then adapt the rapidly exploring random tree star (RRT*) algorithm for the problem of fruit harvesting by incorporating the cost of collision with leaves into the path cost. To reduce the time required for finding low-cost paths to goal, we bias the growth of the tree using an artificial potential field (APF). We compare our proposed method with prior work in a 2-D environment and a 6-DOF robot simulation. Our experiments and a real-world demonstration on a robotic lime picking task demonstrate the applicability of our approach. Heramb Nemlekar, Ziang Liu 0002, Suraj Kothawade, Sherdil Niyaz, Barath Raghavan, Stefanos Nikolaidis |
IROS | 3 |
| 2021 | SIMILAR: Submodular Information Measures Based Active Learning In Realistic ScenariosabstractActive learning has proven to be useful for minimizing labeling costs by selecting the most informative samples. However, existing active learning methods do not work well in realistic scenarios such as imbalance or rare classes,out-of-distribution data in the unlabeled set, and redundancy. In this work, we propose SIMILAR (Submodular Information Measures based actIve LeARning), a unified active learning framework using recently proposed submodular information measures (SIM) as acquisition functions. We argue that SIMILAR not only works in standard active learning but also easily extends to the realistic settings considered above and acts as a one-stop solution for active learning that is scalable to large real-world datasets. Empirically, we show that SIMILAR significantly outperforms existing active learning algorithms by as much as ~5%−18%in the case of rare classes and ~5%−10%in the case of out-of-distribution data on several image classification tasks like CIFAR-10, MNIST, and ImageNet. Suraj Kothawade, Nathan Beck, KrishnaTeja Killamsetty, Rishabh Iyer 0001 |
NeurIPS | 1 |
| 2019 | Learning Collaborative Action Plans from YouTube Videos
Po-Jen Lai, Sayan Paul, Suraj Kothawade, Stefanos Nikolaidis |
ISRR | 4 |
| 2019 | Demystifying Multi-Faceted Video Summarization: Tradeoff Between Diversity, Representation, Coverage and ImportanceabstractThis paper addresses automatic summarization of videos in a unified manner. In particular, we propose a framework for multi-faceted summarization for extractive, query base and entity summarization (summarization at the level of entities like objects, scenes, humans and faces in the video). We investigate several summarization models which capture notions of diversity, coverage, representation and importance, and argue the utility of these different models depending on the application. While most of the prior work on submodular summarization approaches has focused on combining several models and learning weighted mixtures, we focus on the explainability of different models and featurizations, and how they apply to different domains. We also provide implementation details on summarization systems and the different modalities involved. We hope that the study from this paper will give insights into practitioners to appropriately choose the right summarization models for the problems at hand. Vishal Kaushal, Rishabh Iyer 0001, Khoshrav Doctor, Anurag Sahoo, Pratik Dubal, Suraj Kothawade, Rohan Mahadev, Kunal Dargan, Ganesh Ramakrishnan |
WACV | 6 |
| 2019 | Learning From Less Data: A Unified Data Subset Selection and Active Learning Framework for Computer VisionabstractSupervised machine learning based state-of-the-art computer vision techniques are in general data hungry. Their data curation poses the challenges of expensive human labeling, inadequate computing resources and larger experiment turn around times. Training data subset selection and active learning techniques have been proposed as possible solutions to these challenges. A special class of subset selection functions naturally model notions of diversity, coverage and representation and can be used to eliminate redundancy thus lending themselves well for training data subset selection. They can also help improve the efficiency of active learning in further reducing human labeling efforts by selecting a subset of the examples obtained using the conventional uncertainty sampling based techniques. In this work, we empirically demonstrate the effectiveness of two diversity models, namely the Facility-Location and Dispersion models for training-data subset selection and reducing labeling effort. We demonstrate this across the board for a variety of computer vision tasks including Gender Recognition, Face Recognition, Scene Recognition, Object Detection and Object Recognition. Our results show that diversity based subset selection done in the right way can increase the accuracy by upto 5 - 10% over existing baselines, particularly in settings in which less training data is available. This allows the training of complex machine learning models like Convolutional Neural Networks with much less training data and labeling costs while incurring minimal performance loss. Vishal Kaushal, Rishabh Iyer 0001, Suraj Kothawade, Rohan Mahadev, Khoshrav Doctor, Ganesh Ramakrishnan |
WACV | 3 |
| 2019 | A Framework Towards Domain Specific Video SummarizationabstractIn the light of exponentially increasing video content, video summarization has attracted a lot of attention recently due to its ability to optimize time and storage. Characteristics of a good summary of a video depend on the particular domain under question. We propose a novel framework for domain specific video summarization. Given a video of a particular domain, our system can produce a summary based on what is important for that domain in addition to possessing other desired characteristics like representativeness, coverage, diversity etc. as suitable to that domain. Past related work has focused either on using supervised approaches for ranking the snippets to produce summary or on using unsupervised approaches of generating the summary as a subset of snippets with the above characteristics. We look at the joint problem of learning domain specific importance of segments as well as the desired summary characteristic for that domain. Our studies show that the more efficient way of incorporating domain specific relevances into a summary is by obtaining ratings of shots as opposed to binary inclusion/exclusion information. We also argue that ratings can be seen as unified representation of all possible ground truth summaries of a video, taking us one step closer in dealing with challenges associated with multiple ground truth summaries of a video. We also propose a novel evaluation measure which is more naturally suited in assessing the quality of video summary for the task at hand than F1 like measures. It leverages the ratings information and is richer in appropriately modeling desirable and undesirable characteristics of a summary. Lastly, we release a gold standard dataset for furthering research in domain specific video summarization, which to our knowledge is the first dataset with long videos across several domains with rating annotations. Vishal Kaushal, Sandeep Subramanian, Suraj Kothawade, Rishabh Iyer 0001, Ganesh Ramakrishnan |
WACV | 3 |