Alexander C. Li

dblp:243/3349 · also Alexander Cong Li · DBLP profile ↗
← Back
9ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-9884-6383ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Generative modeling · 17% Transfer learning and domain adaptation · 16% Trustworthy machine learning · 16%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › generative model
generative classifier
0.912025
Generative Classifiers Avoid Shortcut Solutions · ICLR 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Generative Classifiers Avoid Shortcut Solutions · ICLR 2025
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.912025
Generative Classifiers Avoid Shortcut Solutions · ICLR 2025
Machine learning › Trustworthy machine learning › robustness
shortcut learning
0.912025
Generative Classifiers Avoid Shortcut Solutions · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.922023
Your Diffusion Model is Secretly a Zero-Shot Classifier · ICCV 2023
Test-time Adaptation of Discriminative Models via Diffusion Generative Feedback · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › deep transfer learning
attention transfer
0.812024
On the Surprising Effectiveness of Attention Transfer for Vision Transformers · NeurIPS 2024
Machine learning › Representation and self-supervised learning
pre-training
0.812024
On the Surprising Effectiveness of Attention Transfer for Vision Transformers · NeurIPS 2024
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.812024
On the Surprising Effectiveness of Attention Transfer for Vision Transformers · NeurIPS 2024
Machine learning › Efficient and distributed learning › active learning
active data collection
0.712023
Internet Explorer: Targeted Representation Learning on the Open Web · ICML 2023
Machine learning › Efficient and distributed learning
data-efficient learning
0.712023
Internet Explorer: Targeted Representation Learning on the Open Web · ICML 2023
Machine learning › Generative modeling › diffusion model
diffusion model adaptation
0.712023
Test-time Adaptation of Discriminative Models via Diffusion Generative Feedback · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation
test-time adaptation
0.712023
Test-time Adaptation of Discriminative Models via Diffusion Generative Feedback · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification
0.712023
Your Diffusion Model is Secretly a Zero-Shot Classifier · ICCV 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.612022
Understanding Collapse in Non-contrastive Siamese Representation Learning · ECCV (31) 2022
Machine learning › Representation and self-supervised learning
siamese representation learning
0.612022
Understanding Collapse in Non-contrastive Siamese Representation Learning · ECCV (31) 2022
Machine learning › Representation and self-supervised learning › feature transformation
fourier features
0.512021
Functional Regularization for Reinforcement Learning via Learned Fourier Features · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning
functional regularization
0.512021
Functional Regularization for Reinforcement Learning via Learned Fourier Features · NeurIPS 2021
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.512021
Functional Regularization for Reinforcement Learning via Learned Fourier Features · NeurIPS 2021
Machine learning › Reinforcement learning
sample efficiency
0.512021
Functional Regularization for Reinforcement Learning via Learned Fourier Features · NeurIPS 2021
Computer vision › Image recognition and object detection
image classification
0.522025
Generative Classifiers Avoid Shortcut Solutions · ICLR 2025
On the Surprising Effectiveness of Attention Transfer for Vision Transformers · NeurIPS 2024
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.412020
Sub-policy Adaptation for Hierarchical Reinforcement Learning · ICLR 2020
Machine learning › Reinforcement learning
hindsight relabeling
0.412020
Generalized Hindsight for Reinforcement Learning · NeurIPS 2020
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.412020
Generalized Hindsight for Reinforcement Learning · NeurIPS 2020
Machine learning › Reinforcement learning
multi-task reinforcement learning
0.412020
Generalized Hindsight for Reinforcement Learning · NeurIPS 2020
Machine learning › Generative modeling
generative feedback
0.212023
Test-time Adaptation of Discriminative Models via Diffusion Generative Feedback · NeurIPS 2023
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.212023
Your Diffusion Model is Secretly a Zero-Shot Classifier · ICCV 2023
Computer vision › Vision and language
vision-language model
0.212023
Your Diffusion Model is Secretly a Zero-Shot Classifier · ICCV 2023

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.5autoregressive model · 0.9knowledge distillation · 0.8fine-tuning · 0.8ensembling · 0.8attention distillation · 0.8text-query image search · 0.7self-supervised training · 0.7diffusion classifier · 0.7conditional density estimation · 0.7
YearPublicationVenuePosition
2025 Generative Classifiers Avoid Shortcut Solutions
abstract
Discriminative approaches to classification often learn shortcuts that hold in-distribution but fail even under minor distribution shift. This failure mode stems from an overreliance on features that are spuriously correlated with the label. We show that generative classifiers, which use class-conditional generative models, can avoid this issue by modeling all features, both core and spurious, instead of mainly spurious ones. These generative classifiers are simple to train, avoiding the need for specialized augmentations, strong regularization, extra hyperparameters, or knowledge of the specific spurious correlations to avoid. We find that diffusion-based and autoregressive generative classifiers achieve state-of-the-art performance on five standard image and text distribution shift benchmarks and reduce the impact of spurious correlations in realistic applications, such as medical or satellite datasets. Finally, we carefully analyze a Gaussian toy setting to understand the inductive biases of generative classifiers, as well as the data properties that determine when generative classifiers outperform discriminative ones.
Alexander C. Li, Ananya Kumar, Deepak Pathak
ICLR1
2024 On the Surprising Effectiveness of Attention Transfer for Vision Transformers
abstract
Conventional wisdom suggests that pre-training Vision Transformers (ViT) improves downstream performance by learning useful representations. Is this actually true? We investigate this question and find that the features and representations learned during pre-training are not essential. Surprisingly, using only the attention patterns from pre-training (i.e., guiding how information flows between tokens) is sufficient for models to learn high quality features from scratch and achieve comparable downstream performance. We show this by introducing a simple method called attention transfer, where only the attention patterns from a pre-trained teacher ViT are transferred to a student, either by copying or distilling the attention maps. Since attention transfer lets the student learn its own features, ensembling it with a fine-tuned teacher also further improves accuracy on ImageNet. We systematically study various aspects of our findings on the sufficiency of attention maps, including distribution shift settings where they underperform fine-tuning. We hope our exploration provides a better understanding of what pre-training accomplishes and leads to a useful alternative to the standard practice of fine-tuning.
Alexander C. Li, Yuandong Tian, Beidi Chen, Deepak Pathak, Xinlei Chen
NeurIPS1
2023 Your Diffusion Model is Secretly a Zero-Shot Classifier
abstract
The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhibit impressive compositional generalization abilities. Almost all use cases thus far have solely focused on sampling; however, diffusion models can also provide conditional density estimates, which are useful for tasks beyond image generation. In this paper, we show that the density estimates from large-scale text-to-image diffusion models like Stable Diffusion can be leveraged to perform zero-shot classification without any additional training. Our generative approach to classification, which we call Diffusion Classifier, attains strong results on a variety of benchmarks and outperforms alternative methods of extracting knowledge from diffusion models. Although a gap remains between generative and discriminative approaches on zero-shot recognition tasks, our diffusion-based approach has stronger multimodal compositional reasoning abilities than competing discriminative approaches. Finally, we use Diffusion Classifier to extract standard classifiers from class-conditional diffusion models trained on ImageNet. These models approach the performance of SOTA discriminative classifiers and exhibit strong "effective robustness" to distribution shift. Overall, our results are a step toward using generative over discriminative models for downstream tasks. Results and visualizations on our website: diffusion-classifier.github.io/
Alexander C. Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown, Deepak Pathak
ICCV1
2023 Internet Explorer: Targeted Representation Learning on the Open Web
abstract
Vision models typically rely on fine-tuning general-purpose models pre-trained on large, static datasets. These general-purpose models only capture the knowledge within their pre-training datasets, which are tiny, out-of-date snapshots of the Internet—where billions of images are uploaded each day. We suggest an alternate approach: rather than hoping our static datasets transfer to our desired tasks after large-scale pre-training, we propose dynamically utilizing the Internet to quickly train a small-scale model that does extremely well on a target dataset. Our approach, called Internet Explorer, explores the web in a self-supervised manner to progressively find relevant examples that improve performance on a desired target dataset. It cycles between searching for images on the Internet with text queries, self-supervised training on downloaded images, determining which images were useful, and prioritizing what to search for next. We evaluate Internet Explorer across several datasets and show that it outperforms or matches CLIP oracle performance using just a single GPU desktop to actively query the Internet for 30-40 hours.
Alexander C. Li, Ellis Brown, Alexei A. Efros, Deepak Pathak
ICML1
2023 Test-time Adaptation of Discriminative Models via Diffusion Generative Feedback
Mihir Prabhudesai, Tsung-Wei Ke, Alexander C. Li, Deepak Pathak, Katerina Fragkiadaki
NeurIPS3
2022 Understanding Collapse in Non-contrastive Siamese Representation Learning
Alexander C. Li, Alexei A. Efros, Deepak Pathak
ECCV (31)1
2021 Functional Regularization for Reinforcement Learning via Learned Fourier Features
abstract
We propose a simple architecture for deep reinforcement learning by embedding inputs into a learned Fourier basis and show that it improves the sample efficiency of both state-based and image-based RL. We perform infinite-width analysis of our architecture using the Neural Tangent Kernel and theoretically show that tuning the initial variance of the Fourier basis is equivalent to functional regularization of the learned deep network. That is, these learned Fourier features allow for adjusting the degree to which networks underfit or overfit different frequencies in the training data, and hence provide a controlled mechanism to improve the stability and performance of RL optimization. Empirically, this allows us to prioritize learning low-frequency functions and speed up learning by reducing networks' susceptibility to noise in the optimization process, such as during Bellman updates. Experiments on standard state-based and image-based RL benchmarks show clear benefits of our architecture over the baselines.
Alexander C. Li, Deepak Pathak
NeurIPS1
2020 Sub-policy Adaptation for Hierarchical Reinforcement Learning
Alexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter Abbeel
ICLR1
2020 Generalized Hindsight for Reinforcement Learning
abstract
One of the key reasons for the high sample complexity in reinforcement learning (RL) is the inability to transfer knowledge from one task to another. In standard multi-task RL settings, low-reward data collected while trying to solve one task provides little to no signal for solving that particular task and is hence effectively wasted. However, we argue that this data, which is uninformative for one task, is likely a rich source of information for other tasks. To leverage this insight and efficiently reuse data, we present Generalized Hindsight: an approximate inverse reinforcement learning technique for relabeling behaviors with the right tasks. Intuitively, given a behavior generated under one task, Generalized Hindsight returns a different task that the behavior is better suited for. Then, the behavior is relabeled with this new task before being used by an off-policy RL optimizer. Compared to standard relabeling techniques, Generalized Hindsight provides a substantially more efficient re-use of samples, which we empirically demonstrate on a suite of multi-task navigation and manipulation tasks.
Alexander C. Li, Lerrel Pinto, Pieter Abbeel
NeurIPS1