VLDB 2026 Research / reviewers in the wild / expert
Alexander C. Li
dblp:243/3349 · also Alexander Cong Li
· DBLP profile ↗
9ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-9884-6383ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Generative modeling · 17% Transfer learning and domain adaptation · 16% Trustworthy machine learning · 16% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › generative model
generative classifier |
0.9 | 1 | 2025 | Generative Classifiers Avoid Shortcut Solutions · ICLR 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Generative Classifiers Avoid Shortcut Solutions · ICLR 2025 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift |
0.9 | 1 | 2025 | Generative Classifiers Avoid Shortcut Solutions · ICLR 2025 |
Machine learning › Trustworthy machine learning › robustness
shortcut learning |
0.9 | 1 | 2025 | Generative Classifiers Avoid Shortcut Solutions · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 2 | 2023 | Your Diffusion Model is Secretly a Zero-Shot Classifier · ICCV 2023 Test-time Adaptation of Discriminative Models via Diffusion Generative Feedback · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation › deep transfer learning
attention transfer |
0.8 | 1 | 2024 | On the Surprising Effectiveness of Attention Transfer for Vision Transformers · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
pre-training |
0.8 | 1 | 2024 | On the Surprising Effectiveness of Attention Transfer for Vision Transformers · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.8 | 1 | 2024 | On the Surprising Effectiveness of Attention Transfer for Vision Transformers · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › active learning
active data collection |
0.7 | 1 | 2023 | Internet Explorer: Targeted Representation Learning on the Open Web · ICML 2023 |
Machine learning › Efficient and distributed learning
data-efficient learning |
0.7 | 1 | 2023 | Internet Explorer: Targeted Representation Learning on the Open Web · ICML 2023 |
Machine learning › Generative modeling › diffusion model
diffusion model adaptation |
0.7 | 1 | 2023 | Test-time Adaptation of Discriminative Models via Diffusion Generative Feedback · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation
test-time adaptation |
0.7 | 1 | 2023 | Test-time Adaptation of Discriminative Models via Diffusion Generative Feedback · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification |
0.7 | 1 | 2023 | Your Diffusion Model is Secretly a Zero-Shot Classifier · ICCV 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.6 | 1 | 2022 | Understanding Collapse in Non-contrastive Siamese Representation Learning · ECCV (31) 2022 |
Machine learning › Representation and self-supervised learning
siamese representation learning |
0.6 | 1 | 2022 | Understanding Collapse in Non-contrastive Siamese Representation Learning · ECCV (31) 2022 |
Machine learning › Representation and self-supervised learning › feature transformation
fourier features |
0.5 | 1 | 2021 | Functional Regularization for Reinforcement Learning via Learned Fourier Features · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning
functional regularization |
0.5 | 1 | 2021 | Functional Regularization for Reinforcement Learning via Learned Fourier Features · NeurIPS 2021 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
0.5 | 1 | 2021 | Functional Regularization for Reinforcement Learning via Learned Fourier Features · NeurIPS 2021 |
Machine learning › Reinforcement learning
sample efficiency |
0.5 | 1 | 2021 | Functional Regularization for Reinforcement Learning via Learned Fourier Features · NeurIPS 2021 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 2 | 2025 | Generative Classifiers Avoid Shortcut Solutions · ICLR 2025 On the Surprising Effectiveness of Attention Transfer for Vision Transformers · NeurIPS 2024 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.4 | 1 | 2020 | Sub-policy Adaptation for Hierarchical Reinforcement Learning · ICLR 2020 |
Machine learning › Reinforcement learning
hindsight relabeling |
0.4 | 1 | 2020 | Generalized Hindsight for Reinforcement Learning · NeurIPS 2020 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.4 | 1 | 2020 | Generalized Hindsight for Reinforcement Learning · NeurIPS 2020 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
0.4 | 1 | 2020 | Generalized Hindsight for Reinforcement Learning · NeurIPS 2020 |
Machine learning › Generative modeling
generative feedback |
0.2 | 1 | 2023 | Test-time Adaptation of Discriminative Models via Diffusion Generative Feedback · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.2 | 1 | 2023 | Your Diffusion Model is Secretly a Zero-Shot Classifier · ICCV 2023 |
Computer vision › Vision and language
vision-language model |
0.2 | 1 | 2023 | Your Diffusion Model is Secretly a Zero-Shot Classifier · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 1.5autoregressive model · 0.9knowledge distillation · 0.8fine-tuning · 0.8ensembling · 0.8attention distillation · 0.8text-query image search · 0.7self-supervised training · 0.7diffusion classifier · 0.7conditional density estimation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generative Classifiers Avoid Shortcut SolutionsabstractDiscriminative approaches to classification often learn shortcuts that hold in-distribution but fail even under minor distribution shift. This failure mode stems from an overreliance on features that are spuriously correlated with the label. We show that generative classifiers, which use class-conditional generative models, can avoid this issue by modeling all features, both core and spurious, instead of mainly spurious ones. These generative classifiers are simple to train, avoiding the need for specialized augmentations, strong regularization, extra hyperparameters, or knowledge of the specific spurious correlations to avoid. We find that diffusion-based and autoregressive generative classifiers achieve state-of-the-art performance on five standard image and text distribution shift benchmarks and reduce the impact of spurious correlations in realistic applications, such as medical or satellite datasets. Finally, we carefully analyze a Gaussian toy setting to understand the inductive biases of generative classifiers, as well as the data properties that determine when generative classifiers outperform discriminative ones. Alexander C. Li, Ananya Kumar, Deepak Pathak |
ICLR | 1 |
| 2024 | On the Surprising Effectiveness of Attention Transfer for Vision TransformersabstractConventional wisdom suggests that pre-training Vision Transformers (ViT) improves downstream performance by learning useful representations. Is this actually true? We investigate this question and find that the features and representations learned during pre-training are not essential. Surprisingly, using only the attention patterns from pre-training (i.e., guiding how information flows between tokens) is sufficient for models to learn high quality features from scratch and achieve comparable downstream performance. We show this by introducing a simple method called attention transfer, where only the attention patterns from a pre-trained teacher ViT are transferred to a student, either by copying or distilling the attention maps. Since attention transfer lets the student learn its own features, ensembling it with a fine-tuned teacher also further improves accuracy on ImageNet. We systematically study various aspects of our findings on the sufficiency of attention maps, including distribution shift settings where they underperform fine-tuning. We hope our exploration provides a better understanding of what pre-training accomplishes and leads to a useful alternative to the standard practice of fine-tuning. Alexander C. Li, Yuandong Tian, Beidi Chen, Deepak Pathak, Xinlei Chen |
NeurIPS | 1 |
| 2023 | Your Diffusion Model is Secretly a Zero-Shot ClassifierabstractThe recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhibit impressive compositional generalization abilities. Almost all use cases thus far have solely focused on sampling; however, diffusion models can also provide conditional density estimates, which are useful for tasks beyond image generation. In this paper, we show that the density estimates from large-scale text-to-image diffusion models like Stable Diffusion can be leveraged to perform zero-shot classification without any additional training. Our generative approach to classification, which we call Diffusion Classifier, attains strong results on a variety of benchmarks and outperforms alternative methods of extracting knowledge from diffusion models. Although a gap remains between generative and discriminative approaches on zero-shot recognition tasks, our diffusion-based approach has stronger multimodal compositional reasoning abilities than competing discriminative approaches. Finally, we use Diffusion Classifier to extract standard classifiers from class-conditional diffusion models trained on ImageNet. These models approach the performance of SOTA discriminative classifiers and exhibit strong "effective robustness" to distribution shift. Overall, our results are a step toward using generative over discriminative models for downstream tasks. Results and visualizations on our website: diffusion-classifier.github.io/ Alexander C. Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown, Deepak Pathak |
ICCV | 1 |
| 2023 | Internet Explorer: Targeted Representation Learning on the Open WebabstractVision models typically rely on fine-tuning general-purpose models pre-trained on large, static datasets. These general-purpose models only capture the knowledge within their pre-training datasets, which are tiny, out-of-date snapshots of the Internet—where billions of images are uploaded each day. We suggest an alternate approach: rather than hoping our static datasets transfer to our desired tasks after large-scale pre-training, we propose dynamically utilizing the Internet to quickly train a small-scale model that does extremely well on a target dataset. Our approach, called Internet Explorer, explores the web in a self-supervised manner to progressively find relevant examples that improve performance on a desired target dataset. It cycles between searching for images on the Internet with text queries, self-supervised training on downloaded images, determining which images were useful, and prioritizing what to search for next. We evaluate Internet Explorer across several datasets and show that it outperforms or matches CLIP oracle performance using just a single GPU desktop to actively query the Internet for 30-40 hours. Alexander C. Li, Ellis Brown, Alexei A. Efros, Deepak Pathak |
ICML | 1 |
| 2023 | Test-time Adaptation of Discriminative Models via Diffusion Generative Feedback
Mihir Prabhudesai, Tsung-Wei Ke, Alexander C. Li, Deepak Pathak, Katerina Fragkiadaki |
NeurIPS | 3 |
| 2022 | Understanding Collapse in Non-contrastive Siamese Representation Learning
Alexander C. Li, Alexei A. Efros, Deepak Pathak |
ECCV (31) | 1 |
| 2021 | Functional Regularization for Reinforcement Learning via Learned Fourier FeaturesabstractWe propose a simple architecture for deep reinforcement learning by embedding inputs into a learned Fourier basis and show that it improves the sample efficiency of both state-based and image-based RL. We perform infinite-width analysis of our architecture using the Neural Tangent Kernel and theoretically show that tuning the initial variance of the Fourier basis is equivalent to functional regularization of the learned deep network. That is, these learned Fourier features allow for adjusting the degree to which networks underfit or overfit different frequencies in the training data, and hence provide a controlled mechanism to improve the stability and performance of RL optimization. Empirically, this allows us to prioritize learning low-frequency functions and speed up learning by reducing networks' susceptibility to noise in the optimization process, such as during Bellman updates. Experiments on standard state-based and image-based RL benchmarks show clear benefits of our architecture over the baselines. Alexander C. Li, Deepak Pathak |
NeurIPS | 1 |
| 2020 | Sub-policy Adaptation for Hierarchical Reinforcement Learning
Alexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter Abbeel |
ICLR | 1 |
| 2020 | Generalized Hindsight for Reinforcement LearningabstractOne of the key reasons for the high sample complexity in reinforcement learning (RL) is the inability to transfer knowledge from one task to another. In standard multi-task RL settings, low-reward data collected while trying to solve one task provides little to no signal for solving that particular task and is hence effectively wasted. However, we argue that this data, which is uninformative for one task, is likely a rich source of information for other tasks. To leverage this insight and efficiently reuse data, we present Generalized Hindsight: an approximate inverse reinforcement learning technique for relabeling behaviors with the right tasks. Intuitively, given a behavior generated under one task, Generalized Hindsight returns a different task that the behavior is better suited for. Then, the behavior is relabeled with this new task before being used by an off-policy RL optimizer. Compared to standard relabeling techniques, Generalized Hindsight provides a substantially more efficient re-use of samples, which we empirically demonstrate on a suite of multi-task navigation and manipulation tasks. Alexander C. Li, Lerrel Pinto, Pieter Abbeel |
NeurIPS | 1 |