VLDB 2026 Research / reviewers in the wild / expert
Siddharth Joshi 0004
dblp:63/6495-4
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Representation and self-supervised learning · 59% Learning theory · 17% Trustworthy machine learning · 15% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
contrastive learning |
2.1 | 3 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023 Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least · ICML 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
1.6 | 2 | 2025 | Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025 Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Machine learning › Trustworthy machine learning
robustness |
1.0 | 2 | 2024 | Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024 Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Machine learning › Efficient and distributed learning
dataset distillation |
0.9 | 1 | 2025 | Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025 |
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning |
0.8 | 1 | 2024 | Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning
projection head |
0.8 | 1 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation analysis
representation learning theory |
0.8 | 1 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift |
0.8 | 1 | 2024 | Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning
feature suppression |
0.7 | 1 | 2023 | Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023 |
Machine learning › Learning theory
generalization bounds |
0.7 | 1 | 2023 | Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least · ICML 2023 |
Machine learning › Learning theory › implicit bias
implicit bias of gradient descent |
0.7 | 1 | 2023 | Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023 |
Machine learning › Learning theory › inductive bias
simplicity bias |
0.7 | 1 | 2023 | Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.3 | 1 | 2025 | Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025 |
Machine learning › Representation and self-supervised learning
representation matching |
0.3 | 1 | 2025 | Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
theoretical analysis · 1.4trajectory matching · 0.9self-supervised learning · 0.9knowledge distillation · 0.9layer-wise feature weighting analysis · 0.8intra-class contrasting · 0.8inter-class feature sharing · 0.8contrastive loss · 0.8contrastive learning · 0.7augmentation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep NetworksabstractDataset distillation (DD) generates small synthetic datasets that can efficiently train deep networks with a limited amount of memory and compute. Despite the success of DD methods for supervised learning, DD for self-supervised pre-training of deep models has remained unaddressed. Pre-training on unlabeled data is crucial for efficiently generalizing to downstream tasks with limited labeled data. In this work, we propose the first effective DD method for SSL pre-training. First, we show, theoretically and empirically, that naiive application of supervised DD methods to SSL fails, due to the high variance of the SSL gradient. Then, we address this issue by relying on insights from knowledge distillation (KD) literature. Specifically, we train a small student model to match the representations of a larger teacher model trained with SSL. Then, we generate a small synthetic dataset by matching the training trajectories of the student models. As the KD objective has considerably lower variance than SSL, our approach can generate synthetic datasets that can successfully pre-train high-quality encoders. Through extensive experiments, we show that our distilled sets lead to up to 13% higher accuracy than prior work, on a variety of downstream tasks, in the presence of limited labeled data. Code at https://github.com/BigML-CS-UCLA/MKDT. Siddharth Joshi 0004, Jiayi Ni, Baharan Mirzasoleiman |
ICLR | 1 |
| 2024 | Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over QuantityabstractContrastive Language-Image Pre-training (CLIP) on large-scale image-caption datasets learns representations that can achieve remarkable zero-shot generalization. However, such models require a massive amount of pre-training data. Improving the quality of the pre-training data has been shown to be much more effective in improving CLIP’s performance than increasing its volume. Nevertheless, finding small subsets of training data that provably generalize best has remained an open question. In this work, we propose the first theoretically rigorous data selection method for CLIP. We show that subsets that closely preserve the cross-covariance of the images and captions of the full data provably achieve a superior generalization performance.Our extensive experiments on ConceptualCaptions3M and ConceptualCaptions12M demonstrate that subsets found by \textsc{ClipCov} achieve over 2.7x and 1.4x the accuracy of the next best baseline on ImageNet and its shifted versions. Moreover, we show that our subsets obtain 1.5x the average accuracy across 11 downstream datasets, of the next best baseline. The code is available at: \url{https://github.com/BigML-CS-UCLA/clipcov-data-efficient-clip}. Siddharth Joshi 0004, Arnav Jain, Ali Payani, Baharan Mirzasoleiman |
AISTATS | 1 |
| 2024 | Investigating the Benefits of Projection Head for Representation LearningabstractAn effective technique for obtaining high-quality representations is adding a projection head on top of the encoder during training, then discarding it and using the pre-projection representations. Despite its proven practical effectiveness, the reason behind the success of this technique is poorly understood. The pre-projection representations are not directly optimized by the loss function, raising the question: what makes them better? In this work, we provide a rigorous theoretical answer to this question. We start by examining linear models trained with self-supervised contrastive loss. We reveal that the implicit bias of training algorithms leads to layer-wise progressive feature weighting, where features become increasingly unequal as we go deeper into the layers. Consequently, lower layers tend to have more normalized and less specialized representations. We theoretically characterize scenarios where such representations are more beneficial, highlighting the intricate interplay between data augmentation and input features. Additionally, we demonstrate that introducing non-linearity into the network allows lower layers to learn features that are completely absent in higher layers. Finally, we show how this mechanism improves the robustness in supervised contrastive learning and supervised learning. We empirically validate our results through various experiments on CIFAR-10/100, UrbanCars and shifted versions of ImageNet. We also introduce a potential alternative to projection head, which offers a more interpretable and controllable design. Yihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi 0004, Baharan Mirzasoleiman |
ICLR | 4 |
| 2024 | Understanding the Robustness of Multi-modal Contrastive Learning to Distribution ShiftabstractRecently, multimodal contrastive learning (MMCL) approaches, such as CLIP, have achieved a remarkable success in learning representations that are robust against distribution shift and generalize to new domains. Despite the empirical success, the mechanism behind learning such generalizable representations is not understood. In this work, we rigorously analyze this problem and
uncover two mechanisms behind MMCL's robustness: \emph{intra-class contrasting}, which allows the model to learn features with a high variance, and \emph{inter-class feature sharing}, where annotated details in one class help learning other classes better. Both mechanisms prevent spurious features that are over-represented in the training data to overshadow the generalizable core features. This yields superior zero-shot classification accuracy under distribution shift. Furthermore, we theoretically demonstrate the benefits of using rich captions on robustness and explore the effect of annotating different types of details in the captions. We validate our theoretical findings through experiments, including a well-designed synthetic experiment and an experiment involving training CLIP models on MSCOCO/Conceptual Captions and evaluating them on shifted ImageNets. Yihao Xue, Siddharth Joshi 0004, Baharan Mirzasoleiman |
ICLR | 2 |
| 2023 | Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the LeastabstractSelf-supervised learning (SSL) learns high-quality representations from large pools of unlabeled training data. As datasets grow larger, it becomes crucial to identify the examples that contribute the most to learning such representations. This enables efficient SSL by reducing the volume of data required. Nevertheless, quantifying the value of examples for SSL has remained an open question. In this work, we address this problem for the first time, by proving that examples that contribute the most to contrastive SSL are those that have the most similar augmentations to other examples, in expectation. We provide rigorous guarantees for the generalization performance of contrastive learning on such subsets. Through extensive experiments, we show that we can safely exclude 20% of examples from CIFAR100 and 40% from STL10 and TinyImageNet, without affecting downstream task performance. In general, subsets selected by our method outperform random subsets by over 3% across these datasets. Interestingly, we also discover the subsets that contribute the most to contrastive learning are those that contribute the least to supervised learning. Siddharth Joshi 0004, Baharan Mirzasoleiman |
ICML | 1 |
| 2023 | Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature SuppressionabstractContrastive learning (CL) has emerged as a powerful technique for representation learning, with or without label supervision. However, supervised CL is prone to collapsing representations of subclasses within a class by not capturing all their features, and unsupervised CL may suppress harder class-relevant features by focusing on learning easy class-irrelevant features; both significantly compromise representation quality. Yet, there is no theoretical understanding of class collapse or feature suppression at test time. We provide the first unified theoretically rigorous framework to determine which features are learnt by CL. Our analysis indicate that, perhaps surprisingly, bias of (stochastic) gradient descent towards finding simpler solutions is a key factor in collapsing subclass representations and suppressing harder class-relevant features. Moreover, we present increasing embedding dimensionality and improving the quality of data augmentations as two theoretically motivated solutions to feature suppression. We also provide the first theoretical explanation for why employing supervised and unsupervised CL together yields higher-quality representations, even when using commonly-used stochastic gradient methods. Yihao Xue, Siddharth Joshi 0004, Eric Gan, Baharan Mirzasoleiman |
ICML | 2 |