Samarth Mishra

dblp:194/2977 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-3425-2647ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Transfer learning and domain adaptation · 32% Representation and self-supervised learning · 19% Video understanding and tracking · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
1.122022
How Transferable are Video Representations Based on Synthetic Data? · NeurIPS 2022
Task2Sim: Towards Effective Pre-training and Transfer from Synthetic Data · CVPR 2022
Machine learning › Transfer learning and domain adaptation › zero-shot learning
multi-label zero-shot learning
0.912025
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models · CVPR 2025
Computer vision › Vision and language
vision-language model
0.912025
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning
compositional representation
0.812024
Interpretable Compositional Representations for Robust Few-Shot Generalization · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot generalization
0.812024
Interpretable Compositional Representations for Robust Few-Shot Generalization · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Interpretable Compositional Representations for Robust Few-Shot Generalization · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Transfer learning and domain adaptation
zero-shot learning
0.812024
Interpretable Compositional Representations for Robust Few-Shot Generalization · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › Video understanding and tracking › action recognition
human action recognition
0.712023
Learning Human Action Recognition Representations Without Real Humans · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder
0.712023
Learning Human Action Recognition Representations Without Real Humans · NeurIPS 2023
Computer vision › Video understanding and tracking › action recognition
privacy-preserving action recognition
0.712023
Learning Human Action Recognition Representations Without Real Humans · NeurIPS 2023
Computer vision › Video understanding and tracking
action recognition
0.612022
How Transferable are Video Representations Based on Synthetic Data? · NeurIPS 2022
Natural language and speech › Language models and text generation › large language model training
pretraining data selection
0.612022
Task2Sim: Towards Effective Pre-training and Transfer from Synthetic Data · CVPR 2022
Machine learning › Representation and self-supervised learning › pre-training
task-adaptive pretraining
0.612022
Task2Sim: Towards Effective Pre-training and Transfer from Synthetic Data · CVPR 2022
Information retrieval › search engines › semantic search › entity retrieval
attribute-based retrieval
0.512021
Effectively Leveraging Attributes for Visual Similarity · ICCV 2021
Machine learning › Learning paradigms
continual learning
0.412019
Incremental Object Learning From Contiguous Views · CVPR 2019
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.212024
Interpretable Compositional Representations for Robust Few-Shot Generalization · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Trustworthy machine learning
robustness
0.212024
Interpretable Compositional Representations for Robust Few-Shot Generalization · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Privacy and data protection
privacy-preserving machine learning
0.212023
Learning Human Action Recognition Representations Without Real Humans · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › semantic representation learning
attribute representation
0.112021
Effectively Leveraging Attributes for Visual Similarity · ICCV 2021
Computer vision › 3D vision
3d object dataset
0.112019
Incremental Object Learning From Contiguous Views · CVPR 2019

Methods — techniques the papers use, named apart from their topics

synthetic data · 1.9masked autoencoder · 1.3human-removed real data · 1.3score standardization · 0.9prompt engineering · 0.9adaptive fusion · 0.9prototype learning · 0.8part decomposition · 0.8crowdsourcing evaluation · 0.8graphics simulation · 0.6attribute embedding · 0.5
YearPublicationVenuePosition
2025 SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
abstract
Zero-shot multi-label recognition (MLR) with Vision-Language Models (VLMs) faces significant challenges without training data, model tuning, or architectural modifications. Existing approaches require prompt tuning or architectural adaptations, limiting zero-shot applicability. Our work proposes a novel solution treating VLMs as black boxes, leveraging scores without training data or ground truth. We make two contributions. First, we find that VLM scores suffer from image- and prompt-specific biases, and that simple standardization is surprisingly effective at removing these and boosting MLR performance. And second, we introduce compound prompts grounded in realistic object combinations. Our analysis reveals "AND"/"OR" signal ambiguities that cause maximum compound scores to be surprisingly suboptimal compared to second-highest scores. We introduce an adaptive fusion method to address this issue. Our method enhances other zero-shot approaches, consistently improving their results. Experiments show superior mean Average Precision (mAP) compared to methods requiring training data, achieved through refined object ranking for robust zero-shot MLR. Code can be found at https://github.com/kjmillerCURIS/SPARC.
Aditya Gangrade, Samarth Mishra, Kate Saenko, Venkatesh Saligrama
CVPR3
2024 Interpretable Compositional Representations for Robust Few-Shot Generalization
abstract
We propose Recognition as Part Composition (RPC), an image encoding approach inspired by human cognition. It is based on the cognitive theory that humans recognize complex objects by components, and that they build a small compact vocabulary of concepts to represent each instance with. RPC encodes images by first decomposing them into salient parts, and then encoding each part as a mixture of a small number of prototypes, each representing a certain concept. We find that this type of learning inspired by human cognition can overcome hurdles faced by deep convolutional networks in low-shot generalization tasks, like zero-shot learning, few-shot learning and unsupervised domain adaptation. Furthermore, we find a classifier using an RPC image encoder is fairly robust to adversarial attacks, that deep neural networks are known to be prone to. Given that our image encoding principle is based on human cognition, one would expect the encodings to be interpretable by humans, which we find to be the case via crowd-sourcing experiments. Finally, we propose an application of these interpretable encodings in the form of generating synthetic attribute annotations for evaluating zero-shot learning methods on new datasets.
Samarth Mishra, Pengkai Zhu, Venkatesh Saligrama
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Fine-grained Few-shot Recognition by Deep Object Parsing
Ruizhao Zhu, Pengkai Zhu, Samarth Mishra, Venkatesh Saligrama
BMVC3
2023 Learning Human Action Recognition Representations Without Real Humans
abstract
Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets contain images of people and hence are accompanied with issues related to privacy, ethics, and data protection, often preventing them from being publicly shared for reproducible research. Existing work has attempted to alleviate these problems by blurring faces, downsampling videos, or training on synthetic data. On the other hand, analysis on the {\em transferability} of privacy-preserving pre-trained models to downstream tasks has been limited. In this work, we study this problem by first asking the question: can we pre-train models for human action recognition with data that does not include real humans? To this end, we present, for the first time, a benchmark that leverages real-world videos with {\em humans removed} and synthetic data containing virtual humans to pre-train a model. We then evaluate the transferability of the representation learned on this data to a diverse set of downstream action recognition benchmarks. Furthermore, we propose a novel pre-training strategy, called Privacy-Preserving MAE-Align, to effectively combine synthetic data and human-removed real data. Our approach outperforms previous baselines by up to 5\% and closes the performance gap between human and no-human action recognition representations on downstream tasks, for both linear probing and fine-tuning. Our benchmark, code, and models are available at https://github.com/howardzh01/PPMA.
Howard Zhong, Samarth Mishra, Donghyun Kim 0006, SouYoung Jin, Rameswar Panda, Hilde Kuehne, Leonid Karlinsky, Venkatesh Saligrama, Aude Oliva, Rogério Feris
NeurIPS2
2022 Task2Sim: Towards Effective Pre-training and Transfer from Synthetic Data
abstract
Pre-training models on Imagenet or other massive datasets of real images has led to major advances in Computer vision, albeit accompanied with shortcomings related to curation cost, privacy, usage rights, and ethical issues. In this paper, for the first time, we study the transferability of pre-trained models based on synthetic data generated by graphics simulators to downstream tasks from very different domains. In using such synthetic data for pre-training, we find that downstream performance on different tasks are fa-vored by different configurations of simulation parameters (e.g. lighting, object pose, backgrounds, etc.), and that there is no one-size-fits-all solution. It is thus better to tailor syn-thetic pre-training data to a specific downstream task, for best performance. We introduce Task2Sim, a unified model mapping downstream task representations to optimal sim-ulation parameters to generate synthetic pre-training data for them. Task2Sim learns this mapping by training to find the set of best parameters on a set of “seen” tasks. Once trained, it can then be used to predict best simulation pa-rameters for novel “unseen” tasks in one shot, without re-quiring additional training. Given a budget in number of images per class, our extensive experiments with 20 di-verse downstream tasks show Task2Sim's task-adaptive pre-training data results in significantly better downstream per-formance than non-adaptively choosing simulation param-eters on both seen and unseen tasks. It is even competitive with pre-training on real images from Imagenet.
Samarth Mishra, Rameswar Panda, Cheng Perng Phoo, Chun-Fu Chen 0001, Leonid Karlinsky, Kate Saenko, Venkatesh Saligrama, Rogério Feris
CVPR1
2022 How Transferable are Video Representations Based on Synthetic Data?
abstract
Action recognition has improved dramatically with massive-scale video datasets. Yet, these datasets are accompanied with issues related to curation cost, privacy, ethics, bias, and copyright. Compared to that, only minor efforts have been devoted toward exploring the potential of synthetic video data. In this work, as a stepping stone towards addressing these shortcomings, we study the transferability of video representations learned solely from synthetically-generated video clips, instead of real data. We propose SynAPT, a novel benchmark for action recognition based on a combination of existing synthetic datasets, in which a model is pre-trained on synthetic videos rendered by various graphics simulators, and then transferred to a set of downstream action recognition datasets, containing different categories than the synthetic data. We provide an extensive baseline analysis on SynAPT revealing that the simulation-to-real gap is minor for datasets with low object and scene bias, where models pre-trained with synthetic data even outperform their real data counterparts. We posit that the gap between real and synthetic action representations can be attributed to contextual bias and static objects related to the action, instead of the temporal dynamics of the action itself. The SynAPT benchmark is available at https://github.com/mintjohnkim/SynAPT.
Yo-whan Kim, Samarth Mishra, SouYoung Jin, Rameswar Panda, Hilde Kuehne, Leonid Karlinsky, Venkatesh Saligrama, Kate Saenko, Aude Oliva, Rogério Feris
NeurIPS2
2021 Surprisingly Simple Semi-Supervised Domain Adaptation with Pretraining and Consistency
Samarth Mishra, Kate Saenko, Venkatesh Saligrama
BMVC1
2021 Effectively Leveraging Attributes for Visual Similarity
Samarth Mishra, Zhongping Zhang, Yuan Shen 0001, Ranjitha Kumar, Venkatesh Saligrama, Bryan A. Plummer
ICCV1
2019 Incremental Object Learning From Contiguous Views
abstract
In this work, we present CRIB (Continual Recognition Inspired by Babies), a synthetic incremental object learning environment that can produce data that models visual imagery produced by object exploration in early infancy. CRIB is coupled with a new 3D object dataset, Toys-200, that contains 200 unique toy-like object instances, and is also compatible with existing 3D datasets. Through extensive empirical evaluation of state-of-the-art incremental learning algorithms, we find the novel empirical result that repetition can significantly ameliorate the effects of catastrophic forgetting. Furthermore, we find that in certain cases repetition allows for performance approaching that of batch learning algorithms. Finally, we propose an unsupervised incremental learning task with intriguing baseline results.
Stefan Stojanov, Samarth Mishra, Ngoc Anh Thai, Nikhil Dhanda, Ahmad Humayun, Chen Yu 0001, Linda B. Smith, James M. Rehg
CVPR2
2017 Faster Algorithms for Weighted Recursive State Machines
Krishnendu Chatterjee, Bernhard Kragl, Samarth Mishra, Andreas Pavlogiannis
ESOP3