Siddharth Joshi 0004

dblp:63/6495-4 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Representation and self-supervised learning · 59% Learning theory · 17% Trustworthy machine learning · 15%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
2.132024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least · ICML 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
1.622025
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Machine learning › Trustworthy machine learning
robustness
1.022024
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Machine learning › Efficient and distributed learning
dataset distillation
0.912025
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning
0.812024
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024
Machine learning › Representation and self-supervised learning › contrastive learning
projection head
0.812024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Machine learning › Representation and self-supervised learning › representation analysis
representation learning theory
0.812024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.812024
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift · ICLR 2024
Machine learning › Representation and self-supervised learning › contrastive learning
feature suppression
0.712023
Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023
Machine learning › Learning theory
generalization bounds
0.712023
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least · ICML 2023
Machine learning › Learning theory › implicit bias
implicit bias of gradient descent
0.712023
Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023
Machine learning › Learning theory › inductive bias
simplicity bias
0.712023
Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression · ICML 2023
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.312025
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025
Machine learning › Representation and self-supervised learning
representation matching
0.312025
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025

Methods — techniques the papers use, named apart from their topics

theoretical analysis · 1.4trajectory matching · 0.9self-supervised learning · 0.9knowledge distillation · 0.9layer-wise feature weighting analysis · 0.8intra-class contrasting · 0.8inter-class feature sharing · 0.8contrastive loss · 0.8contrastive learning · 0.7augmentation · 0.7
YearPublicationVenuePosition
2025 Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks
abstract
Dataset distillation (DD) generates small synthetic datasets that can efficiently train deep networks with a limited amount of memory and compute. Despite the success of DD methods for supervised learning, DD for self-supervised pre-training of deep models has remained unaddressed. Pre-training on unlabeled data is crucial for efficiently generalizing to downstream tasks with limited labeled data. In this work, we propose the first effective DD method for SSL pre-training. First, we show, theoretically and empirically, that naiive application of supervised DD methods to SSL fails, due to the high variance of the SSL gradient. Then, we address this issue by relying on insights from knowledge distillation (KD) literature. Specifically, we train a small student model to match the representations of a larger teacher model trained with SSL. Then, we generate a small synthetic dataset by matching the training trajectories of the student models. As the KD objective has considerably lower variance than SSL, our approach can generate synthetic datasets that can successfully pre-train high-quality encoders. Through extensive experiments, we show that our distilled sets lead to up to 13% higher accuracy than prior work, on a variety of downstream tasks, in the presence of limited labeled data. Code at https://github.com/BigML-CS-UCLA/MKDT.
Siddharth Joshi 0004, Jiayi Ni, Baharan Mirzasoleiman
ICLR1
2024 Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
abstract
Contrastive Language-Image Pre-training (CLIP) on large-scale image-caption datasets learns representations that can achieve remarkable zero-shot generalization. However, such models require a massive amount of pre-training data. Improving the quality of the pre-training data has been shown to be much more effective in improving CLIP’s performance than increasing its volume. Nevertheless, finding small subsets of training data that provably generalize best has remained an open question. In this work, we propose the first theoretically rigorous data selection method for CLIP. We show that subsets that closely preserve the cross-covariance of the images and captions of the full data provably achieve a superior generalization performance.Our extensive experiments on ConceptualCaptions3M and ConceptualCaptions12M demonstrate that subsets found by \textsc{ClipCov} achieve over 2.7x and 1.4x the accuracy of the next best baseline on ImageNet and its shifted versions. Moreover, we show that our subsets obtain 1.5x the average accuracy across 11 downstream datasets, of the next best baseline. The code is available at: \url{https://github.com/BigML-CS-UCLA/clipcov-data-efficient-clip}.
Siddharth Joshi 0004, Arnav Jain, Ali Payani, Baharan Mirzasoleiman
AISTATS1
2024 Investigating the Benefits of Projection Head for Representation Learning
abstract
An effective technique for obtaining high-quality representations is adding a projection head on top of the encoder during training, then discarding it and using the pre-projection representations. Despite its proven practical effectiveness, the reason behind the success of this technique is poorly understood. The pre-projection representations are not directly optimized by the loss function, raising the question: what makes them better? In this work, we provide a rigorous theoretical answer to this question. We start by examining linear models trained with self-supervised contrastive loss. We reveal that the implicit bias of training algorithms leads to layer-wise progressive feature weighting, where features become increasingly unequal as we go deeper into the layers. Consequently, lower layers tend to have more normalized and less specialized representations. We theoretically characterize scenarios where such representations are more beneficial, highlighting the intricate interplay between data augmentation and input features. Additionally, we demonstrate that introducing non-linearity into the network allows lower layers to learn features that are completely absent in higher layers. Finally, we show how this mechanism improves the robustness in supervised contrastive learning and supervised learning. We empirically validate our results through various experiments on CIFAR-10/100, UrbanCars and shifted versions of ImageNet. We also introduce a potential alternative to projection head, which offers a more interpretable and controllable design.
Yihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi 0004, Baharan Mirzasoleiman
ICLR4
2024 Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift
abstract
Recently, multimodal contrastive learning (MMCL) approaches, such as CLIP, have achieved a remarkable success in learning representations that are robust against distribution shift and generalize to new domains. Despite the empirical success, the mechanism behind learning such generalizable representations is not understood. In this work, we rigorously analyze this problem and uncover two mechanisms behind MMCL's robustness: \emph{intra-class contrasting}, which allows the model to learn features with a high variance, and \emph{inter-class feature sharing}, where annotated details in one class help learning other classes better. Both mechanisms prevent spurious features that are over-represented in the training data to overshadow the generalizable core features. This yields superior zero-shot classification accuracy under distribution shift. Furthermore, we theoretically demonstrate the benefits of using rich captions on robustness and explore the effect of annotating different types of details in the captions. We validate our theoretical findings through experiments, including a well-designed synthetic experiment and an experiment involving training CLIP models on MSCOCO/Conceptual Captions and evaluating them on shifted ImageNets.
Yihao Xue, Siddharth Joshi 0004, Baharan Mirzasoleiman
ICLR2
2023 Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
abstract
Self-supervised learning (SSL) learns high-quality representations from large pools of unlabeled training data. As datasets grow larger, it becomes crucial to identify the examples that contribute the most to learning such representations. This enables efficient SSL by reducing the volume of data required. Nevertheless, quantifying the value of examples for SSL has remained an open question. In this work, we address this problem for the first time, by proving that examples that contribute the most to contrastive SSL are those that have the most similar augmentations to other examples, in expectation. We provide rigorous guarantees for the generalization performance of contrastive learning on such subsets. Through extensive experiments, we show that we can safely exclude 20% of examples from CIFAR100 and 40% from STL10 and TinyImageNet, without affecting downstream task performance. In general, subsets selected by our method outperform random subsets by over 3% across these datasets. Interestingly, we also discover the subsets that contribute the most to contrastive learning are those that contribute the least to supervised learning.
Siddharth Joshi 0004, Baharan Mirzasoleiman
ICML1
2023 Which Features are Learnt by Contrastive Learning? On the Role of Simplicity Bias in Class Collapse and Feature Suppression
abstract
Contrastive learning (CL) has emerged as a powerful technique for representation learning, with or without label supervision. However, supervised CL is prone to collapsing representations of subclasses within a class by not capturing all their features, and unsupervised CL may suppress harder class-relevant features by focusing on learning easy class-irrelevant features; both significantly compromise representation quality. Yet, there is no theoretical understanding of class collapse or feature suppression at test time. We provide the first unified theoretically rigorous framework to determine which features are learnt by CL. Our analysis indicate that, perhaps surprisingly, bias of (stochastic) gradient descent towards finding simpler solutions is a key factor in collapsing subclass representations and suppressing harder class-relevant features. Moreover, we present increasing embedding dimensionality and improving the quality of data augmentations as two theoretically motivated solutions to feature suppression. We also provide the first theoretical explanation for why employing supervised and unsupervised CL together yields higher-quality representations, even when using commonly-used stochastic gradient methods.
Yihao Xue, Siddharth Joshi 0004, Eric Gan, Baharan Mirzasoleiman
ICML2