Kaichao You

dblp:238/1508 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0002-1955-3743ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Transfer learning and domain adaptation · 46% Efficient and distributed learning · 23% Segmentation and scene understanding · 9%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 87% Hardware accelerators and domain-specific architectures · 13%
Computer graphics and multimedia
2 papers
Image and video processing · 87% Computational photography and imaging · 13%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
fine-tuning
1.432022
Ranking and Tuning Pre-trained Models: A New Paradigm for Exploiting Model Hubs · J. Mach. Learn. Res. 2022
Co-Tuning for Transfer Learning · NeurIPS 2020
Stochastic Normalization · NeurIPS 2020
Image and video processing
video frame interpolation
1.122022
Video Interpolation by Event-Driven Anisotropic Adjustment of Optical Flow · ECCV (7) 2022
TimeReplayer: Unlocking the Potential of Event Cameras for Video Interpolation · CVPR 2022
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
partial domain adaptation
1.022023
From Big to Small: Adaptive Learning to Partial-Set Domains · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Learning to Transfer Examples for Partial Domain Adaptation · CVPR 2019
Machine learning › Efficient and distributed learning
deep learning compilers
0.912025
depyf: Open the Opaque Box of PyTorch Compiler for Machine Learning Researchers · J. Mach. Learn. Res. 2025
Machine learning › Efficient and distributed learning
KV cache management
0.912025
Jenga: Effective Memory Management for Serving LLM with Heterogeneity · SOSP 2025
Machine learning › Efficient and distributed learning › inference serving
large language model serving
0.912025
Jenga: Effective Memory Management for Serving LLM with Heterogeneity · SOSP 2025
Memory systems › memory management
heterogeneous memory management
0.912025
Jenga: Effective Memory Management for Serving LLM with Heterogeneity · SOSP 2025
Memory systems
memory management
0.912025
Jenga: Effective Memory Management for Serving LLM with Heterogeneity · SOSP 2025
Computer vision › Segmentation and scene understanding › semantic segmentation › non-RGB semantic segmentation
event-based semantic segmentation
0.712023
Event-Based Semantic Segmentation With Posterior Attention · IEEE Trans. Image Process. 2023
Computer vision › Segmentation and scene understanding
semantic segmentation
0.712023
Event-Based Semantic Segmentation With Posterior Attention · IEEE Trans. Image Process. 2023
Machine learning › Reinforcement learning
deep reinforcement learning
0.612022
Tianshou: A Highly Modularized Deep Reinforcement Learning Library · J. Mach. Learn. Res. 2022
Computer vision › 3D vision
event-based vision
0.612022
TimeReplayer: Unlocking the Potential of Event Cameras for Video Interpolation · CVPR 2022
Machine learning › Transfer learning and domain adaptation › pre-training and adaptation
pretrained model reuse
0.612022
Ranking and Tuning Pre-trained Models: A New Paradigm for Exploiting Model Hubs · J. Mach. Learn. Res. 2022
Machine learning › Transfer learning and domain adaptation › pre-trained models
pre-trained model selection
0.512021
LogME: Practical Assessment of Pre-trained Models for Transfer Learning · ICML 2021
Machine learning › Transfer learning and domain adaptation
transferability estimation
0.512021
LogME: Practical Assessment of Pre-trained Models for Transfer Learning · ICML 2021
Machine learning › Deep learning architectures and training › normalization
batch normalization
0.412020
Stochastic Normalization · NeurIPS 2020
Computer vision › Image recognition and object detection
image classification
0.412020
Co-Tuning for Transfer Learning · NeurIPS 2020
Machine learning › Deep learning architectures and training
normalization
0.412020
Stochastic Normalization · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation › fine-tuning
regularization for fine-tuning
0.412020
Stochastic Normalization · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.412019
Universal Domain Adaptation · CVPR 2019
Machine learning › Learning theory
model selection
0.412019
Towards Accurate Model Selection in Deep Unsupervised Domain Adaptation · ICML 2019
Machine learning › Transfer learning and domain adaptation › domain adaptation
open-set domain adaptation
0.412019
Universal Domain Adaptation · CVPR 2019
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
universal domain adaptation
0.412019
Universal Domain Adaptation · CVPR 2019
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.412019
Towards Accurate Model Selection in Deep Unsupervised Domain Adaptation · ICML 2019
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.312025
Jenga: Effective Memory Management for Serving LLM with Heterogeneity · SOSP 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.212023
Event-Based Semantic Segmentation With Posterior Attention · IEEE Trans. Image Process. 2023
Machine learning › Deep learning architectures and training
transformer
0.212023
Event-Based Semantic Segmentation With Posterior Attention · IEEE Trans. Image Process. 2023
Computational photography and imaging
event camera
0.212022
Video Interpolation by Event-Driven Anisotropic Adjustment of Optical Flow · ECCV (7) 2022
Machine learning › Transfer learning and domain adaptation
domain-invariant representation learning
0.112019
Learning to Transfer Examples for Partial Domain Adaptation · CVPR 2019
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.112019
Universal Domain Adaptation · CVPR 2019

Methods — techniques the papers use, named apart from their topics

memory management · 1.7bytecode decompilation · 1.7cycle consistency · 1.1convolution · 0.8batch normalization · 0.8transformer · 0.7self-training · 0.7posterior attention · 0.7bi-level selection · 0.7adversarial training · 0.7unsupervised learning · 0.6optical flow · 0.6neural network · 0.6anisotropic adjustment · 0.6
YearPublicationVenuePosition
2025 Jenga: Effective Memory Management for Serving LLM with Heterogeneity
Chen Zhang 0001, Kuntai Du, Woosuk Kwon, Xiangxi Mo, Kaichao You, Zhuohan Li 0001, Mingsheng Long, Jidong Zhai, Joseph Gonzalez 0001, Ion Stoica
SOSP8
2025 depyf: Open the Opaque Box of PyTorch Compiler for Machine Learning Researchers
abstract
PyTorch 2.x introduces a compiler designed to accelerate deep learning programs. However, for machine learning researchers, fully leveraging the PyTorch compiler can be challenging due to its operation at the Python bytecode level, making it appear as an opaque box. To address this, we introduce depyf, a tool designed to demystify the inner workings of the PyTorch compiler. depyf decompiles the bytecode generated by PyTorch back into equivalent source code and establishes connections between the code objects in the memory and their counterparts in source code format on the disk. This feature enables users to step through the source code line by line using debuggers, thus enhancing their understanding of the underlying processes. Notably, depyf is non-intrusive and user-friendly, primarily relying on two convenient context managers for its core functionality. The project is openly available at https://github.com/thuml/depyf and is recognized as a PyTorch ecosystem project at https://pytorch.org/blog/introducing-depyf.
Kaichao You, Runsheng Bai, Jianmin Wang 0001, Ion Stoica, Mingsheng Long
J. Mach. Learn. Res.1
2024 Efficient ConvBN Blocks for Transfer Learning and Beyond
abstract
Convolution-BatchNorm (ConvBN) blocks are integral components in various computer vision tasks and other domains. A ConvBN block can operate in three modes: Train, Eval, and Deploy. While the Train mode is indispensable for training models from scratch, the Eval mode is suitable for transfer learning and beyond, and the Deploy mode is designed for the deployment of models. This paper focuses on the trade-off between stability and efficiency in ConvBN blocks: Deploy mode is efficient but suffers from training instability; Eval mode is widely used in transfer learning but lacks efficiency. To solve the dilemma, we theoretically reveal the reason behind the diminished training stability observed in the Deploy mode. Subsequently, we propose a novel Tune mode to bridge the gap between Eval mode and Deploy mode. The proposed Tune mode is as stable as Eval mode for transfer learning, and its computational efficiency closely matches that of the Deploy mode. Through extensive experiments in object detection, classification, and adversarial example generation across $5$ datasets and $12$ model architectures, we demonstrate that the proposed Tune mode retains the performance while significantly reducing GPU memory footprint and training time, thereby contributing efficient ConvBN blocks for transfer learning and beyond. Our method has been integrated into both PyTorch (general machine learning framework) and MMCV/MMEngine (computer vision framework). Practitioners just need one line of code to enjoy our efficient ConvBN blocks thanks to PyTorch's builtin machine learning compilers.
Kaichao You, Guo Qin, Anchang Bao, Jiulong Shan, Mingsheng Long
ICLR1
2023 Test-Time Training-Free Domain Adaptation
abstract
Deploying deep learning models to new environments is very challenging. Domain adaptation (DA) is a promising paradigm to solve the problem by collecting and adapting to unlabeled data in new environments. Though research efforts have led to steady performance improvement over the past decade, DA algorithms are still hard to deploy, as training on unlabeled new data makes tuning difficult and not feasible for inference-only devices. To make DA practical, in this paper we study a new problem named Test-time Training-Free Domain Adaptation (TTDA), where trained models must adapt to a single input (mimicking the test-time scenario) without training. By exploiting spatial activation that was previously overlooked and simply averaged out, we propose a simple method based on Feature Statistics Transformation (FST) on-the-fly for each test example. The proposed algorithm is tested in the TTDA setting on two standard DA benchmarks. Surprisingly, it surpasses or performs on par with state-of-the-art DA methods, even though they require additional training. We envision that this training-free paradigm has the potential to bring DA to embedded devices and would be of interest to audience of community.
Yongxiang Feng, Weihua He, Kaichao You, Yaoyuan Wang, Yihang Lou, Jianxing Liao
ICASSP3
2023 From Big to Small: Adaptive Learning to Partial-Set Domains
abstract
Domain adaptation targets at knowledge acquisition and dissemination from a labeled source domain to an unlabeled target domain under distribution shift. Still, the common requirement of identical class space shared across domains hinders applications of domain adaptation to partial-set domains. Recent advances show that deep pre-trained models of large scale endow rich knowledge to tackle diverse downstream tasks of small scale. Thus, there is a strong incentive to adapt models from large-scale domains to small-scale domains. This paper introduces Partial Domain Adaptation (PDA), a learning paradigm that relaxes the identical class space assumption to that the source class space subsumes the target class space. First, we present a theoretical analysis of partial domain adaptation, which uncovers the importance of estimating the transferable probability of each class and each instance across domains. Then, we propose Selective Adversarial Network (SAN and SAN++) with a bi-level selection strategy and an adversarial adaptation mechanism. The bi-level selection strategy up-weighs each class and each instance simultaneously for source supervised training, target self-training, and source-target adversarial adaptation through the transferable probability estimated alternately by the model. Experiments on standard partial-set datasets and more challenging tasks with superclasses show that SAN++ outperforms several domain adaptation methods.
Zhangjie Cao, Kaichao You, Jianmin Wang 0001, Mingsheng Long
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Event-Based Semantic Segmentation With Posterior Attention
abstract
In the past years, attention-based Transformers have swept across the field of computer vision, starting a new stage of backbones in semantic segmentation. Nevertheless, semantic segmentation under poor light conditions remains an open problem. Moreover, most papers about semantic segmentation work on images produced by commodity frame-based cameras with a limited framerate, hindering their deployment to auto-driving systems that require instant perception and response at milliseconds. An event camera is a new sensor that generates event data at microseconds and can work in poor light conditions with a high dynamic range. It looks promising to leverage event cameras to enable perception where commodity cameras are incompetent, but algorithms for event data are far from mature. Pioneering researchers stack event data as frames so that event-based segmentation is converted to frame-based segmentation, but characteristics of event data are not explored. Noticing that event data naturally highlight moving objects, we propose a posterior attention module that adjusts the standard attention by the prior knowledge provided by event data. The posterior attention module can be readily plugged into many segmentation backbones. Plugging the posterior attention module into a recently proposed SegFormer network, we get EvSegFormer (the event-based version of SegFormer) with state-of-the-art performance in two datasets (MVSEC and DDD-17) collected for event-based segmentation. Code is available at https://github.com/zexiJia/EvSegFormer to facilitate research on event-based vision.
Zexi Jia, Kaichao You, Weihua He, Yang Tian 0002, Yongxiang Feng, Yaoyuan Wang, Xu Jia 0012, Yihang Lou, Guoqi Li 0002
IEEE Trans. Image Process.2
2022 TimeReplayer: Unlocking the Potential of Event Cameras for Video Interpolation
abstract
Recording fast motion in a high FPS (frame-per-second) requires expensive high-speed cameras. As an alternative, interpolating low-FPS videos from commodity cameras has attracted significant attention. If only low-FPS videos are available, motion assumptions (linear or quadratic) are necessary to infer intermediate frames, which fail to model complex motions. Event camera, a new camera with pixels producing events of brightness change at the temporal resolution of μs (10–6second), is a game-changing device to enable video interpolation at the presence of arbitrarily complex motion. Since event camera is a novel sensor, its potential has not been fulfilled due to the lack of processing algorithms. The pioneering work Time Lens introduced event cameras to video interpolation by designing optical devices to collect a large amount of paired training data of high-speed frames and events, which is too costly to scale. To fully unlock the potential of event cameras, this paper proposes a novel TimeReplayer algorithm to interpolate videos captured by commodity cameras with events. It is trained in an unsupervised cycleconsistent style, canceling the necessity of high-speed training data and bringing the additional ability of video extrapolation. Its state-of-the-art results and demo videos in supplementary reveal the promising future of event-based vision.
Weihua He, Kaichao You, Zhendong Qiao, Xu Jia 0012, Wenhui Wang 0001, Huchuan Lu, Yaoyuan Wang, Jianxing Liao
CVPR2
2022 Video Interpolation by Event-Driven Anisotropic Adjustment of Optical Flow
Kaichao You, Weihua He, Yaoyuan Wang, Jianxing Liao
ECCV (7)2
2022 Tianshou: A Highly Modularized Deep Reinforcement Learning Library
abstract
In this paper, we present Tianshou, a highly modularized Python library for deep reinforcement learning (DRL) that uses PyTorch as its backend. Tianshou intends to be research-friendly by providing a flexible and reliable infrastructure of DRL algorithms. It supports online and offline training with more than 20 classic algorithms through a unified interface. To facilitate related research and prove Tianshou's reliability, we have released Tianshou's benchmark of MuJoCo environments, covering eight classic algorithms with state-of-the-art performance. We open-sourced Tianshou at https://github.com/thu-ml/tianshou/.
Jiayi Weng, Huayu Chen, Kaichao You, Alexis Duburcq, Hang Su 0006, Jun Zhu 0001
J. Mach. Learn. Res.4
2022 Ranking and Tuning Pre-trained Models: A New Paradigm for Exploiting Model Hubs
abstract
Model hubs with many pre-trained models (PTMs) have become a cornerstone of deep learning. Although built at a high cost, they remain under-exploited---practitioners usually pick one PTM from the provided model hub by popularity and then fine-tune the PTM to solve the target task. This naïve but common practice poses two obstacles to full exploitation of pre-trained model hubs: first, the PTM selection by popularity has no optimality guarantee, and second, only one PTM is used while the remaining PTMs are ignored. An alternative might be to consider all possible combinations of PTMs and extensively fine-tune each combination, but this would not only be prohibitive computationally but may also lead to statistical over-fitting. In this paper, we propose a new paradigm for exploiting model hubs that is intermediate between these extremes. The paradigm is characterized by two aspects: (1) We use an evidence maximization procedure to estimate the maximum value of label evidence given features extracted by pre-trained models. This procedure can rank all the PTMs in a model hub for various types of PTMs and tasks before fine-tuning. (2) The best ranked PTM can either be fine-tuned and deployed if we have no preference for the model's architecture or the target PTM can be tuned by the top $K$ ranked PTMs via a Bayesian procedure that we propose. This procedure, which we refer to as B-Tuning, not only improves upon specialized methods designed for tuning homogeneous PTMs, but also applies to the challenging problem of tuning heterogeneous PTMs where it yields a new level of benchmark performance.
Kaichao You, Yong Liu 0007, Jianmin Wang 0001, Michael I. Jordan, Mingsheng Long
J. Mach. Learn. Res.1
2021 LogME: Practical Assessment of Pre-trained Models for Transfer Learning
abstract
This paper studies task adaptive pre-trained model selection, an underexplored problem of assessing pre-trained models for the target task and select best ones from the model zoo \emph{without fine-tuning}. A few pilot works addressed the problem in transferring supervised pre-trained models to classification tasks, but they cannot handle emerging unsupervised pre-trained models or regression tasks. In pursuit of a practical assessment method, we propose to estimate the maximum value of label evidence given features extracted by pre-trained models. Unlike the maximum likelihood, the maximum evidence is \emph{immune to over-fitting}, while its expensive computation can be dramatically reduced by our carefully designed algorithm. The Logarithm of Maximum Evidence (LogME) can be used to assess pre-trained models for transfer learning: a pre-trained model with a high LogME value is likely to have good transfer performance. LogME is \emph{fast, accurate, and general}, characterizing itself as the first practical method for assessing pre-trained models. Compared with brute-force fine-tuning, LogME brings at most $3000\times$ speedup in wall-clock time and requires only $1%$ memory footprint. It outperforms prior methods by a large margin in their setting and is applicable to new settings. It is general enough for diverse pre-trained models (supervised pre-trained and unsupervised pre-trained), downstream tasks (classification and regression), and modalities (vision and language). Code is available at this repository: \href{https://github.com/thuml/LogME}{https://github.com/thuml/LogME}.
Kaichao You, Yong Liu 0007, Jianmin Wang 0001, Mingsheng Long
ICML1
2020 Stochastic Normalization
abstract
Fine-tuning pre-trained deep networks on a small dataset is an important component in the deep learning pipeline. A critical problem in fine-tuning is how to avoid over-fitting when data are limited. Existing efforts work from two aspects: (1) impose regularization on parameters or features; (2) transfer prior knowledge to fine-tuning by reusing pre-trained parameters. In this paper, we take an alternative approach by refactoring the widely used Batch Normalization (BN) module to mitigate over-fitting. We propose a two-branch design with one branch normalized by mini-batch statistics and the other branch normalized by moving statistics. During training, two branches are stochastically selected to avoid over-depending on some sample statistics, resulting in a strong regularization effect, which we interpret as ``architecture regularization.'' The resulting method is dubbed stochastic normalization (\textbf{StochNorm}). With the two-branch architecture, it naturally incorporates pre-trained moving statistics in BN layers during fine-tuning, exploiting more prior knowledge of pre-trained networks. Extensive empirical experiments show that StochNorm is a powerful tool to avoid over-fitting in fine-tuning with small datasets. Besides, StochNorm is readily pluggable in modern CNN backbones. It is complementary to other fine-tuning methods and can work together to achieve stronger regularization effect.
Zhi Kou, Kaichao You, Mingsheng Long, Jianmin Wang 0001
NeurIPS2
2020 Co-Tuning for Transfer Learning
abstract
Fine-tuning pre-trained deep neural networks (DNNs) to a target dataset, also known as transfer learning, is widely used in computer vision and NLP. Because task-specific layers mainly contain categorical information and categories vary with datasets, practitioners only \textit{partially} transfer pre-trained models by discarding task-specific layers and fine-tuning bottom layers. However, it is a reckless loss to simply discard task-specific parameters who take up as many as $20\%$ of the total parameters in pre-trained models. To \textit{fully} transfer pre-trained models, we propose a two-step framework named \textbf{Co-Tuning}: (i) learn the relationship between source categories and target categories from the pre-trained model and calibrated predictions; (ii) target labels (one-hot labels), as well as source labels (probabilistic labels) translated by the category relationship, collaboratively supervise the fine-tuning process. A simple instantiation of the framework shows strong empirical results in four visual classification tasks and one NLP classification task, bringing up to $20\%$ relative improvement. While state-of-the-art fine-tuning techniques mainly focus on how to impose regularization when data are not abundant, Co-Tuning works not only in medium-scale datasets (100 samples per class) but also in large-scale datasets (1000 samples per class) where regularization-based methods bring no gains over the vanilla fine-tuning. Co-Tuning relies on a typically valid assumption that the pre-trained dataset is diverse enough, implying its broad application area.
Kaichao You, Zhi Kou, Mingsheng Long, Jianmin Wang 0001
NeurIPS1
2019 Learning to Transfer Examples for Partial Domain Adaptation
abstract
Domain adaptation is critical for learning in new and unseen environments. With domain adversarial training, deep networks can learn disentangled and transferable features that effectively diminish the dataset shift between the source and target domains for knowledge transfer. In the era of Big Data, large-scale labeled datasets are readily available, stimulating the interest in partial domain adaptation (PDA), which transfers a recognizer from a large labeled domain to a small unlabeled domain. It extends standard domain adaptation to the scenario where target labels are only a subset of source labels. Under the condition that target labels are unknown, the key challenges of PDA are how to transfer relevant examples in the shared classes to promote positive transfer and how to ignore irrelevant ones in the source domain to mitigate negative transfer. In this work, we propose a unified approach to PDA, Example Transfer Network (ETN), which jointly learns domain-invariant representations across domains and a progressive weighting scheme to quantify the transferability of source examples. A thorough evaluation on several benchmark datasets shows that ETN consistently achieves state-of-the-art results for various partial domain adaptation tasks.
Zhangjie Cao, Kaichao You, Mingsheng Long, Jianmin Wang 0001, Qiang Yang 0001
CVPR2
2019 Universal Domain Adaptation
abstract
Domain adaptation aims to transfer knowledge in the presence of the domain gap. Existing domain adaptation methods rely on rich prior knowledge about the relationship between the label sets of source and target domains, which greatly limits their application in the wild. This paper introduces Universal Domain Adaptation (UDA) that requires no prior knowledge on the label sets. For a given source label set and a target label set, they may contain a common label set and hold a private label set respectively, bringing up an additional category gap. UDA requires a model to either (1) classify the target sample correctly if it is associated with a label in the common label set, or (2) mark it as ``unknown'' otherwise. More importantly, a UDA model should work stably against a wide spectrum of commonness (the proportion of the common label set over the complete label set) so that it can handle real-world problems with unknown target label sets. To solve the universal domain adaptation problem, we propose Universal Adaptation Network (UAN). It quantifies sample-level transferability to discover the common label set and the label sets private to each domain, thereby promoting the adaptation in the automatically discovered common label set and recognizing the ``unknown'' samples successfully. A thorough evaluation shows that UAN outperforms the state of the art closed set, partial and open set domain adaptation methods in the novel UDA setting.
Kaichao You, Mingsheng Long, Zhangjie Cao, Jianmin Wang 0001, Michael I. Jordan
CVPR1
2019 Towards Accurate Model Selection in Deep Unsupervised Domain Adaptation
abstract
Deep unsupervised domain adaptation (Deep UDA) methods successfully leverage rich labeled data in a source domain to boost the performance on related but unlabeled data in a target domain. However, algorithm comparison is cumbersome in Deep UDA due to the absence of accurate and standardized model selection method, posing an obstacle to further advances in the field. Existing model selection methods for Deep UDA are either highly biased, restricted, unstable, or even controversial (requiring labeled target data). To this end, we propose Deep Embedded Validation (DEV), which embeds adapted feature representation into the validation procedure to obtain unbiased estimation of the target risk with bounded variance. The variance is further reduced by the technique of control variate. The efficacy of the method has been justified both theoretically and empirically.
Kaichao You, Ximei Wang, Mingsheng Long, Michael I. Jordan
ICML1