Jeff Z. HaoChen

dblp:267/5319 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Optimization for machine learning · 32% Representation and self-supervised learning · 31% Learning theory · 11%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 18 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
1.942023
A theoretical study of inductive biases in contrastive learning · ICLR 2023
Connect, Not Collapse: Explaining Contrastive Learning for Unsupervised Domain Adaptation · ICML 2022
Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss · NeurIPS 2021
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
1.122022
Beyond Separability: Analyzing the Linear Transferability of Contrastive Representations to Related Subpopulations · NeurIPS 2022
Connect, Not Collapse: Explaining Contrastive Learning for Unsupervised Domain Adaptation · ICML 2022
Machine learning › Representation and self-supervised learning
inductive biases
0.712023
A theoretical study of inductive biases in contrastive learning · ICLR 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
Diagnosing and Rectifying Vision Models using Language · ICLR 2023
Machine learning › Representation and self-supervised learning › contrastive learning › contrastive feature learning
contrastive embedding
0.612022
Beyond Separability: Analyzing the Linear Transferability of Contrastive Representations to Related Subpopulations · NeurIPS 2022
Machine learning › Optimization for machine learning
learned optimizer
0.612022
Amortized Proximal Optimization · NeurIPS 2022
Machine learning › Optimization for machine learning › optimization
meta-optimization
0.612022
Amortized Proximal Optimization · NeurIPS 2022
Machine learning › Trustworthy machine learning
robustness
0.612022
Self-supervised Learning is More Robust to Dataset Imbalance · ICLR 2022
Machine learning › Optimization for machine learning
implicit regularization
0.512021
Shape Matters: Understanding the Implicit Bias of the Noise Covariance · COLT 2021
Machine learning › Learning theory
over-parameterization
0.512021
Shape Matters: Understanding the Implicit Bias of the Noise Covariance · COLT 2021
Machine learning › Learning theory
provable guarantees
0.512021
Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss · NeurIPS 2021
Machine learning › Representation and self-supervised learning › contrastive learning › contrastive loss
spectral contrastive loss
0.512021
Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss · NeurIPS 2021
Machine learning › Optimization for machine learning
stochastic gradient descent
0.512021
Shape Matters: Understanding the Implicit Bias of the Noise Covariance · COLT 2021
Machine learning › Optimization for machine learning
convergence analysis
0.412019
Random Shuffling Beats SGD after Finite Epochs · ICML 2019
Machine learning › Optimization for machine learning › convergence guarantees
non-asymptotic convergence
0.412019
Random Shuffling Beats SGD after Finite Epochs · ICML 2019
Machine learning › Optimization for machine learning
stochastic optimization
0.412019
Random Shuffling Beats SGD after Finite Epochs · ICML 2019
Machine learning › Learning theory
generalization bounds
0.112021
Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss · NeurIPS 2021
Mathematical optimization › continuous optimization
convex optimization
0.112019
Random Shuffling Beats SGD after Finite Epochs · ICML 2019

Methods — techniques the papers use, named apart from their topics

language-guided diagnosis · 0.7theoretical analysis · 0.6self-supervised pretraining · 0.6natural gradient descent · 0.6linear classifier · 0.6domain adversarial training · 0.6contrastive pre-training · 0.6contrastive learning · 0.6amortized proximal optimization · 0.6K-FAC · 0.6stochastic gradient descent · 0.4random shuffling · 0.4
YearPublicationVenuePosition
2023 A theoretical study of inductive biases in contrastive learning
Jeff Z. HaoChen, Tengyu Ma 0001
ICLR1
2023 Diagnosing and Rectifying Vision Models using Language
Jeff Z. HaoChen, Shih-Cheng Huang, Kuan-Chieh Wang, James Zou 0001, Serena Yeung-Levy
ICLR2
2022 Self-supervised Learning is More Robust to Dataset Imbalance
Jeff Z. HaoChen, Adrien Gaidon, Tengyu Ma 0001
ICLR2
2022 Connect, Not Collapse: Explaining Contrastive Learning for Unsupervised Domain Adaptation
abstract
We consider unsupervised domain adaptation (UDA), where labeled data from a source domain (e.g., photos) and unlabeled data from a target domain (e.g., sketches) are used to learn a classifier for the target domain. Conventional UDA methods (e.g., domain adversarial training) learn domain-invariant features to generalize from the source domain to the target domain. In this paper, we show that contrastive pre-training, which learns features on unlabeled source and target data and then fine-tunes on labeled source data, is competitive with strong UDA methods. However, we find that contrastive pre-training does not learn domain-invariant features, diverging from conventional UDA intuitions. We show theoretically that contrastive pre-training can learn features that vary subtantially across domains but still generalize to the target domain, by disentangling domain and class information. We empirically validate our theory on benchmark vision datasets.
Kendrick Shen, Robbie Jones, Ananya Kumar, Sang Michael Xie, Jeff Z. HaoChen, Tengyu Ma 0001, Percy Liang
ICML5
2022 Amortized Proximal Optimization
abstract
We propose a framework for online meta-optimization of parameters that govern optimization, called Amortized Proximal Optimization (APO). We first interpret various existing neural network optimizers as approximate stochastic proximal point methods which trade off the current-batch loss with proximity terms in both function space and weight space. The idea behind APO is to amortize the minimization of the proximal point objective by meta-learning the parameters of an update rule. We show how APO can be used to adapt a learning rate or a structured preconditioning matrix. Under appropriate assumptions, APO can recover existing optimizers such as natural gradient descent and KFAC. It enjoys low computational overhead and avoids expensive and numerically sensitive operations required by some second-order optimizers, such as matrix inverses. We empirically test APO for online adaptation of learning rates and structured preconditioning matrices for regression, image reconstruction, image classification, and natural language translation tasks. Empirically, the learning rate schedules found by APO generally outperform optimal fixed learning rates and are competitive with manually tuned decay schedules. Using APO to adapt a structured preconditioning matrix generally results in optimization performance competitive with second-order methods. Moreover, the absence of matrix inversion provides numerical stability, making it effective for low-precision training.
Juhan Bae, Paul Vicol, Jeff Z. HaoChen, Roger B. Grosse
NeurIPS3
2022 Beyond Separability: Analyzing the Linear Transferability of Contrastive Representations to Related Subpopulations
abstract
Contrastive learning is a highly effective method for learning representations from unlabeled data. Recent works show that contrastive representations can transfer across domains, leading to simple state-of-the-art algorithms for unsupervised domain adaptation. In particular, a linear classifier trained to separate the representations on the source domain can also predict classes on the target domain accurately, even though the representations of the two domains are far from each other. We refer to this phenomenon as linear transferability. This paper analyzes when and why contrastive representations exhibit linear transferability in a general unsupervised domain adaptation setting. We prove that linear transferability can occur when data from the same class in different domains (e.g., photo dogs and cartoon dogs) are more related with each other than data from different classes in different domains (e.g., photo dogs and cartoon cats) are. Our analyses are in a realistic regime where the source and target domains can have unbounded density ratios and be weakly related, and they have distant representations across domains.
Jeff Z. HaoChen, Colin Wei, Ananya Kumar, Tengyu Ma 0001
NeurIPS1
2021 Shape Matters: Understanding the Implicit Bias of the Noise Covariance
abstract
The noise in stochastic gradient descent (SGD) provides a crucial implicit regularization effect for training overparameterized models. Prior theoretical work largely focuses on spherical Gaussian noise, whereas empirical studies demonstrate the phenomenon that parameter-dependent noise — induced by mini-batches or label perturbation — is far more effective than Gaussian noise. This paper theoretically characterizes this phenomenon on a quadratically-parameterized model introduced by Vaskevicius et al. and Woodworth et al. We show that in an over-parameterized setting, SGD with label noise recovers the sparse ground-truth with an arbitrary initialization, whereas SGD with Gaussian noise or gradient descent overfits to dense solutions with large norms. Our analysis reveals that parameter-dependent noise introduces a bias towards local minima with smaller noise variance, whereas spherical Gaussian noise does not.
Jeff Z. HaoChen, Colin Wei, Jason D. Lee, Tengyu Ma 0001
COLT1
2021 Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss
abstract
Recent works in self-supervised learning have advanced the state-of-the-art by relying on the contrastive learning paradigm, which learns representations by pushing positive pairs, or similar examples from the same class, closer together while keeping negative pairs far apart. Despite the empirical successes, theoretical foundations are limited -- prior analyses assume conditional independence of the positive pairs given the same class label, but recent empirical applications use heavily correlated positive pairs (i.e., data augmentations of the same image). Our work analyzes contrastive learning without assuming conditional independence of positive pairs using a novel concept of the augmentation graph on data. Edges in this graph connect augmentations of the same data, and ground-truth classes naturally form connected sub-graphs. We propose a loss that performs spectral decomposition on the population augmentation graph and can be succinctly written as a contrastive learning objective on neural net representations. Minimizing this objective leads to features with provable accuracy guarantees under linear probe evaluation. By standard generalization bounds, these accuracy guarantees also hold when minimizing the training contrastive loss. In all, this work provides the first provable analysis for contrastive learning where the guarantees can apply to realistic empirical settings.
Jeff Z. HaoChen, Colin Wei, Adrien Gaidon, Tengyu Ma 0001
NeurIPS1
2019 Random Shuffling Beats SGD after Finite Epochs
abstract
A long-standing problem in stochastic optimization is proving that \rsgd, the without-replacement version of \sgd, converges faster than the usual with-replacement \sgd. Building upon \citep{gurbuzbalaban2015random}, we present the first (to our knowledge) non-asymptotic results for this problem by proving that after a reasonable number of epochs \rsgd converges faster than \sgd. Specifically, we prove that for strongly convex, second-order smooth functions, the iterates of \rsgd converge to the optimal solution as $\mathcal{O}(\nicefrac{1}{T^2} + \nicefrac{n^3}{T^3})$, where $n$ is the number of components in the objective, and $T$ is number of iterations. This result implies that after $\mathcal{O}(\sqrt{n})$ epochs, \rsgd is strictly better than \sgd (which converges as $\mathcal{O}(\nicefrac{1}{T})$). The key step toward showing this better dependence on $T$ is the introduction of $n$ into the bound; and as our analysis shows, in general a dependence on $n$ is unavoidable without further changes. To understand how \rsgd works in practice, we further explore two empirically useful settings: data sparsity and over-parameterization. For sparse data, \rsgd has the rate $\mathcal{O}\left(\frac{1}{T^2}\right)$, again strictly better than \sgd. Under a setting closely related to over-parameterization, \rsgd is shown to converge faster than \sgd after any arbitrary number of iterations. Finally, we extend the analysis of \rsgd to smooth non-convex and convex functions.
Jeff Z. HaoChen, Suvrit Sra
ICML1