Kaiyang Zhou

dblp:203/3155 · DBLP profile ↗
← Back
38ranked-venue papers
14as first author
33since 2021 · last 2026
0000-0002-8153-3903ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 13 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models
abstract
Domain adaptation and generalization are crucial for real-world applications, such as autonomous driving and medical imaging where the model must operate reliably across environments with distinct data distributions. However, these tasks are challenging because the model needs to overcome various domain gaps caused by variations in, for example, lighting, weather, sensor configurations, and so on. Addressing domain gaps simultaneously in different modalities, known as multimodal domain adaptation and generalization, is even more challenging due to unique challenges in different modalities. Over the past few years, significant progress has been made in these areas, with applications ranging from action recognition to semantic segmentation, and more. Recently, the emergence of large-scale pre-trained multimodal foundation models, such as CLIP, has inspired numerous research studies, which leverage these models to enhance downstream adaptation and generalization. This survey summarizes recent advances in multimodal adaptation and generalization, particularly how these areas evolve from traditional approaches to foundation models. Specifically, this survey covers (1) multimodal domain adaptation, (2) multimodal test-time adaptation, (3) multimodal domain generalization, (4) domain adaptation and generalization with the help of multimodal foundation models, and (5) adaptation of multimodal foundation models. For each topic, we formally define the problem and give a thorough review of existing methods. Additionally, we analyze relevant datasets and applications, highlighting open challenges and potential future research directions.
Hao Dong 0011, Moru Liu, Kaiyang Zhou, Eleni N. Chatzi, Juho Kannala, Cyrill Stachniss, Olga Fink
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Bootstrapping Grounded Chain-of-Thought in Multimodal Llms for Data-Efficient Model Adaptation
Jiaer Xia, Bingkui Tong, Yuhang Zang, Kaiyang Zhou
ICCV5
2025 Fair Allocation of Divisible Goods under Non-Linear Valuations
Haris Aziz 0001, Zixu He, Xinhang Lu, Kaiyang Zhou
AAMAS4
2025 EmoSym: A Symbiotic Framework for Unified Emotional Understanding and Generation via Latent Reasoning
abstract
Current affective computing paradigms often treat emotional understanding and generation as separate tasks, yet they inherently possess symbiotic potential for mutual enhancement. In this paper, we aim to bridge the gap by developing a unified framework. The primary challenge lies in the extraction of precise and semantically rich representations of abstract emotions, which are crucial for both tasks. To address this, we harness the Chain-of-Thought reasoning at the latent space of multimodal large language models and propose EmoSym, a unified framework built upon this advanced foundation. Our framework is executed through three key steps: 1) Emotional reasoning knowledge compression. To enable efficient transfer of emotional reasoning priors, we design specialized reasoning tokens to compact emotion-aware contexts from external reasoning knowledge bases into latent representations. 2) Verifiable reinforcement reasoning optimization. To ensure more reliable and consistent emotional reasoning, we develop a verifiable reinforcement learning paradigm to further enhance the reasoning token by emotion-specific verifiable reward signals. Processed through the above two steps, the reasoning token simultaneously enhances emotional understanding while enriching semantic representations, benefiting subsequent emotional generation tasks. 3) Reasoning-augmented generation and online feedback. We then fuse it with emotional representations and feed them into a diffusion model to generate emotion-evoking images. Additionally, to create a generative-to-understanding enhancement feedback, we propose an Online Emotional Memory Bank (OEMB). It leverages newly generated images to progressively update the training dataset in the training process to reinforce understanding. Extensive experiments demonstrate the superior capabilities of our framework in both emotional understanding and generation tasks.
Yibo Lyu, Zitong Yu, Rui Shao 0001, Kaiyang Zhou, Liqiang Nie
ACM Multimedia5
2025 Contextual Object Detection with Multimodal Large Language Models
Yuhang Zang, Wei Li 0319, Kaiyang Zhou, Chen Change Loy
Int. J. Comput. Vis.4
2025 Neural Prompt Search
abstract
The size of vision models has grown exponentially over the last few years, especially after the emergence of Vision Transformer. This has motivated the development of parameter-efficient tuning methods, such as learning adapter layers or visual prompt tokens, which allow a tiny portion of model parameters to be trained whereas the vast majority obtained from pre-training are frozen. However, designing a proper tuning method is non-trivial: one might need to try out a lengthy list of design choices, not to mention that each downstream dataset often requires custom designs. In this paper, we view the existing parameter-efficient tuning methods as “prompt modules” and propose Neural prOmpt seArcH (NOAH), a novel approach that learns, for large vision models, the optimal design of prompt modules through a neural architecture search algorithm, specifically for each downstream dataset. By conducting extensive experiments on over 20 vision datasets, we demonstrate that NOAH (i) is superior to individual prompt modules, (ii) has good few-shot learning ability, and (iii) is domain-generalizable.
Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Polarity Prompting Vision Foundation Models for Pathology Image Analysis
abstract
The sharp rise in non-alcoholic fatty liver disease (NAFLD) cases has become a major health concern in recent years. Accurately identifying tissue alteration regions is crucial for NAFLD diagnosis but challenging with small-scale pathology datasets. Recently, prompt tuning has emerged as an effective strategy for adapting vision models to small-scale data analysis. However, current prompting techniques, designed primarily for general image classification, use generic cues that are inadequate when dealing with the intricacies of pathological tissue analysis. To solve this problem, we introduce Quantitative Attribute-based Polarity Visual Prompting (Q-PoVP), a new prompting method for pathology image analysis. Q-PoVP introduces two types of measurable attributes: K-function-based spatial attributes and histogram-based morphological attributes. Both help to measure tissue conditions quantitatively. We develop a quantitative attribute-based polarity visual prompt generator that converts quantitative visual attributes into positive and negative visual prompts, facilitating a more comprehensive and nuanced interpretation of pathological images. To enhance feature discrimination, we introduce a novel orthogonal-based polarity visual prompt tuning technique that disentangles and amplifies positive visual attributes while suppressing negative ones. We extensively tested our method on three different tasks. Our task-specific prompting demonstrates superior performance in both diagnostic accuracy and interpretability compared to existing methods. This dual advantage makes it particularly valuable for clinical settings, where healthcare providers require not only reliable results but also transparent reasoning to support informed patient care decisions. Code is available at https://github.com/7LFB/Q-PoVP.
Chong Yin, Si-Qi Liu 0003, Kaiyang Zhou, Vincent Wai-Sun Wong, Pong C. Yuen
IEEE Trans. Medical Imaging3
2024 Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
abstract
With the emergence of pre-trained vision-language models like CLIP, how to adapt them to various downstream classification tasks has garnered significant attention in re-cent research. The adaptation strategies can be typically categorized into three paradigms: zero-shot adaptation, few-shot adaptation, and the recently-proposed training-free few-shot adaptation. Most existing approaches are tai-lored for a specific setting and can only cater to one or two of these paradigms. In this paper, we introduce a versa-tile adaptation approach that can effectively work under all three settings. Specifically, we propose the dual memory networks that comprise dynamic and static memory components. The static memory caches training data knowledge, enabling training-free few-shot adaptation, while the dynamic memory preserves historical test features online during the testing process, allowing for the exploration of additional data insights beyond the training set. This novel capability enhances model performance in the few-shot setting and enables model usability in the absence of training data. The two memory networks employ the same flexible memory interactive strategy, which can operate in a training-free mode and can be further enhanced by in-corporating learnable projection layers. Our approach is tested across 11 datasets under the three task settings. Re-markably, in the zero-shot scenario, it outperforms existing methods by over 3% and even shows superior results against methods utilizing external training data. Addition-ally, our method exhibits robust performance against nat-ural distribution shifts. Codes are available at https://github.com/YBZh/DMN.
Yabin Zhang 0001, Wenjie Zhu 0003, Zhiyuan Ma 0002, Kaiyang Zhou, Lei Zhang 0006
CVPR5
2024 Prompting Vision Foundation Models for Pathology Image Analysis
abstract
The rapid increase in cases of non-alcoholic fatty liver disease (NAFLD) in recent years has raised significant public concern. Accurately identifying tissue alteration regions is crucial for the diagnosis of NAFLD, but this task presents challenges in pathology image analysis, particularly with small-scale datasets. Recently, the paradigm shift from full fine-tuning to prompting in adapting vision foundation models has offered a new perspective for small-scale data analysis. However, existing prompting methods based on task-agnostic prompts are mainly developed for generic image recognition, which fall short in providing instructive cues for complex pathology images. In this paper, we propose Quantitative Attribute-based Prompting (QAP), a novel prompting method specifically for liver pathology image analysis. QAP is based on two quantitative attributes, namely K-function-based spatial attributes and histogram-based morphological attributes, which are aimed for quantitative assessment of tissue states. Moreover, a conditional prompt generator is designed to turn these instance-specific attributes into visual prompts. Extensive experiments on three diverse tasks demonstrate that our task-specific prompting method achieves better diagnostic performance as well as better interpretability. Code is available at https://github.com/7LFBIQAP.
Chong Yin, Si-Qi Liu 0003, Kaiyang Zhou, Vincent Wai-Sun Wong, Pong C. Yuen
CVPR3
2024 Octopus: Embodied Vision-Language Programmer from Environmental Feedback
Yuhao Dong, Shuai Liu 0002, Bo Li 0080, Haoran Tan, Chencheng Jiang, Jiamu Kang, Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu 0002
ECCV (1)10
2024 Open-Vocabulary Calibration for Fine-tuned CLIP
abstract
Vision-language models (VLMs) have emerged as formidable tools, showing their strong capability in handling various open-vocabulary tasks in image recognition, text-driven visual content generation, and visual chatbots, to name a few. In recent years, considerable efforts and resources have been devoted to adaptation methods for improving downstream performance of VLMs, particularly on parameter-efficient fine-tuning methods like prompt learning. However, a crucial aspect that has been largely overlooked is the confidence calibration problem in fine-tuned VLMs, which could greatly reduce reliability when deploying such models in the real world. This paper bridges the gap by systematically investigating the confidence calibration problem in the context of prompt learning and reveals that existing calibration methods are insufficient to address the problem, especially in the open-vocabulary setting. To solve the problem, we present a simple and effective approach called Distance-Aware Calibration (DAC), which is based on scaling the temperature using as guidance the distance between predicted text labels and base classes. The experiments with 7 distinct prompt learning methods applied across 11 diverse downstream datasets demonstrate the effectiveness of DAC, which achieves high efficacy without sacrificing the inference speed.
Shuoyuan Wang, Jindong Wang 0001, Bob Zhang 0001, Kaiyang Zhou, Hongxin Wei
ICML5
2024 Generalized Out-of-Distribution Detection: A Survey
Kaiyang Zhou, Yixuan Li 0001, Ziwei Liu 0002
Int. J. Comput. Vis.2
2024 Guest Editorial: Special Issue on the Promises and Dangers of Large Vision Models
Kaiyang Zhou, Ziwei Liu 0002, Xiaohua Zhai, Chunyuan Li, Kate Saenko
Int. J. Comput. Vis.1
2024 MixStyle Neural Networks for Domain Generalization and Adaptation
Kaiyang Zhou, Yongxin Yang, Yu Qiao 0001, Tao Xiang 0002
Int. J. Comput. Vis.1
2023 Panoptic Video Scene Graph Generation
abstract
Towards building comprehensive real-world visual perception systems, we propose and study a new problem called panoptic scene graph generation (PVSG). PVSG is related to the existing video scene graph generation (VidSGG) problem, which focuses on temporal interactions between humans and objects localized with bounding boxes in videos. However, the limitation of bounding boxes in detecting non-rigid objects and backgrounds often causes VidSGG systems to miss key details that are crucial for comprehensive video understanding. In contrast, PVSG requires nodes in scene graphs to be grounded by more precise, pixel-level segmentation masks, which facilitate holistic scene understanding. To advance research in this new area, we contribute a high-quality PVSG dataset, which consists of 400 videos (289 third-person + 111 egocentric videos) with totally 150K frames labeled with panoptic segmentation masks as well as fine, temporal scene graphs. We also provide a variety of baseline methods and share useful design practices for future work.
Wenxuan Peng, Xiangtai Li, Zujin Guo, Liangyu Chen 0005, Bo Li 0080, Zheng Ma 0008, Kaiyang Zhou, Wayne Zhang 0001, Chen Change Loy, Ziwei Liu 0002
CVPR8
2023 4D Panoptic Scene Graph Generation
abstract
We are living in a three-dimensional space while moving forward through a fourth dimension: time. To allow artificial intelligence to develop a comprehensive understanding of such a 4D environment, we introduce **4D Panoptic Scene Graph (PSG-4D)**, a new representation that bridges the raw visual data perceived in a dynamic 4D world and high-level visual understanding. Specifically, PSG-4D abstracts rich 4D sensory data into nodes, which represent entities with precise location and status information, and edges, which capture the temporal relations. To facilitate research in this new area, we build a richly annotated PSG-4D dataset consisting of 3K RGB-D videos with a total of 1M frames, each of which is labeled with 4D panoptic segmentation masks as well as fine-grained, dynamic scene graphs. To solve PSG-4D, we propose PSG4DFormer, a Transformer-based model that can predict panoptic segmentation masks, track masks along the time axis, and generate the corresponding scene graphs via a relation component. Extensive experiments on the new dataset show that our method can serve as a strong baseline for future research on PSG-4D. In the end, we provide a real-world application example to demonstrate how we can achieve dynamic scene understanding by integrating a large language model into our PSG-4D system.
Jun Cen, Wenxuan Peng, Shuai Liu 0002, Fangzhou Hong, Xiangtai Li, Kaiyang Zhou, Qifeng Chen 0001, Ziwei Liu 0002
NeurIPS7
2023 What Makes Good Examples for Visual In-Context Learning?
abstract
Large vision models with billions of parameters and trained on broad data have great potential in numerous downstream applications. However, these models are typically difficult to adapt due to their large parameter size and sometimes lack of accesss to their weights---entities able to develop large vision models often provide APIs only. In this paper, we study how to better utilize large vision models through the lens of in-context learning, a concept that has been well-known in natural language processing but has only been studied very recently in computer vision. In-context learning refers to the ability to perform inference on tasks never seen during training by simply conditioning on in-context examples (i.e., input-output pairs) without updating any internal model parameters. To demystify in-context learning in computer vision, we conduct an extensive research and identify a critical problem: downstream performance is highly sensitivie to the choice of visual in-context examples. To address this problem, we propose a prompt retrieval framework specifically for large vision models, allowing the selection of in-context examples to be fully automated. Concretely, we provide two implementations: (i) an unsupervised prompt retrieval method based on nearest example search using an off-the-shelf model, and (ii) a supervised prompt retrieval method, which trains a neural network to choose examples that directly maximize in-context learning performance. Both methods do not require access to the internal weights of large vision models. Our results demonstrate that our methods can bring non-trivial improvements to visual in-context learning in comparison to the commonly-used random selection. Code and models will be released.
Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu 0002
NeurIPS2
2023 Full-Spectrum Out-of-Distribution Detection
Kaiyang Zhou, Ziwei Liu 0002
Int. J. Comput. Vis.2
2023 Semi-Supervised and Long-Tailed Object Detection with CascadeMatch
Yuhang Zang, Kaiyang Zhou, Chen Huang 0001, Chen Change Loy
Int. J. Comput. Vis.2
2023 Semi-Supervised Domain Generalization with Stochastic StyleMatch
Kaiyang Zhou, Chen Change Loy, Ziwei Liu 0002
Int. J. Comput. Vis.1
2023 Domain Generalization: A Survey
abstract
Generalization to out-of-distribution (OOD) data is a capability natural to humans yet challenging for machines to reproduce. This is because most learning algorithms strongly rely on the i.i.d. assumption on source/target data, which is often violated in practice due to domain shift. Domain generalization (DG) aims to achieve OOD generalization by using only source data for model learning. Over the last ten years, research in DG has made great progress, leading to a broad spectrum of methodologies, e.g., those based on domain alignment, meta-learning, data augmentation, or ensemble learning, to name a few; DG has also been studied in various application areas including computer vision, speech recognition, natural language processing, medical imaging, and reinforcement learning. In this paper, for the first time a comprehensive literature review in DG is provided to summarize the developments over the past decade. Specifically, we first cover the background by formally defining DG and relating it to other relevant fields like domain adaptation and transfer learning. Then, we conduct a thorough review into existing methods and theories. Finally, we conclude this survey with insights and discussions on future research directions.
Kaiyang Zhou, Ziwei Liu 0002, Yu Qiao 0001, Tao Xiang 0002, Chen Change Loy
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Learning to Augment via Implicit Differentiation for Domain Generalization
Tingwei Wang, Da Li 0001, Kaiyang Zhou, Tao Xiang 0002, Yi-Zhe Song
BMVC3
2022 Conditional Prompt Learning for Vision-Language Models
abstract
With the rise of powerful pre-trained vision-language models like CLIP, it becomes essential to investigate ways to adapt these models to downstream datasets. A recently proposed method named Context Optimization (CoOp) introduces the concept of prompt learning—a recent trend in NLP—to the vision domain for adapting pre-trained vision-language models. Specifically, CoOp turns context words in a prompt into a set of learnable vectors and, with only a few labeled images for learning, can achieve huge improvements over intensively-tuned manual prompts. In our study we identify a critical problem of CoOp: the learned context is not generalizable to wider unseen classes within the same dataset, suggesting that CoOp overfits base classes observed during training. To address the problem, we propose Conditional Context Optimization (CoCoOp), which extends CoOp by further learning a lightweight neural network to generate for each image an input-conditional token (vector). Compared to CoOp's static prompts, our dynamic prompts adapt to each instance and are thus less sensitive to class shift. Extensive experiments show that CoCoOp generalizes much better than CoOp to unseen classes, even showing promising transferability beyond a single dataset; and yields stronger domain generalization performance as well. Code is available at https://github.com/KaiyangZhou/CoOp.
Kaiyang Zhou, Chen Change Loy, Ziwei Liu 0002
CVPR1
2022 Panoptic Scene Graph Generation
Yi Zhe Ang, Zujin Guo, Kaiyang Zhou, Wayne Zhang 0001, Ziwei Liu 0002
ECCV (27)4
2022 Open-Vocabulary DETR with Conditional Matching
Yuhang Zang, Wei Li 0319, Kaiyang Zhou, Chen Huang 0001, Chen Change Loy
ECCV (9)3
2022 OpenOOD: Benchmarking Generalized Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection is vital to safety-critical machine learning applications and has thus been extensively studied, with a plethora of methods developed in the literature. However, the field currently lacks a unified, strictly formulated, and comprehensive benchmark, which often results in unfair comparisons and inconclusive results. From the problem setting perspective, OOD detection is closely related to neighboring fields including anomaly detection (AD), open set recognition (OSR), and model uncertainty, since methods developed for one domain are often applicable to each other. To help the community to improve the evaluation and advance, we build a unified, well-structured codebase called OpenOOD, which implements over 30 methods developed in relevant fields and provides a comprehensive benchmark under the recently proposed generalized OOD detection framework. With a comprehensive comparison of these methods, we are gratified that the field has progressed significantly over the past few years, where both preprocessing methods and the orthogonal post-hoc methods show strong potential.
Pengyun Wang, Dejian Zou, Zitang Zhou, Kunyuan Ding, Wenxuan Peng, Bo Li 0080, Yiyou Sun, Xuefeng Du, Kaiyang Zhou, Wayne Zhang 0001, Dan Hendrycks, Yixuan Li 0001, Ziwei Liu 0002
NeurIPS12
2022 Learning to Prompt for Vision-Language Models
Kaiyang Zhou, Chen Change Loy, Ziwei Liu 0002
Int. J. Comput. Vis.1
2022 Learning Generalisable Omni-Scale Representations for Person Re-Identification
abstract
An effective person re-identification (re-ID) model should learn feature representations that are both discriminative, for distinguishing similar-looking people, and generalisable, for deployment across datasets without any adaptation. In this paper, we develop novel CNN architectures to address both challenges. First, we present a re-ID CNN termed omni-scale network (OSNet) to learn features that not only capture different spatial scales but also encapsulate a synergistic combination of multiple scales, namely omni-scale features. The basic building block consists of multiple convolutional streams, each detecting features at a certain scale. For omni-scale feature learning, a unified aggregation gate is introduced to dynamically fuse multi-scale features with channel-wise weights. OSNet is lightweight as its building blocks comprise factorised convolutions. Second, to improve generalisable feature learning, we introduce instance normalisation (IN) layers into OSNet to cope with cross-dataset discrepancies. Further, to determine the optimal placements of these IN layers in the architecture, we formulate an efficient differentiable architecture search algorithm. Extensive experiments show that, in the conventional same-dataset setting, OSNet achieves state-of-the-art performance, despite being much smaller than existing re-ID models. In the more challenging yet practical cross-dataset setting, OSNet beats most recent unsupervised domain adaptation methods without using any target data.
Kaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao Xiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Dynamic Instance Domain Adaptation
abstract
Most existing studies on unsupervised domain adaptation (UDA) assume that each domain's training samples come with domain labels (e.g., painting, photo). Samples from each domain are assumed to follow the same distribution and the domain labels are exploited to learn domain-invariant features via feature alignment. However, such an assumption often does not hold true-there often exist numerous finer-grained domains (e.g., dozens of modern painting styles have been developed, each differing dramatically from those of the classic styles). Therefore, forcing feature distribution alignment across each artificially-defined and coarse-grained domain can be ineffective. In this paper, we address both single-source and multi-source UDA from a completely different perspective, which is to view each instance as a fine domain. Feature alignment across domains is thus redundant. Instead, we propose to perform dynamic instance domain adaptation (DIDA). Concretely, a dynamic neural network with adaptive convolutional kernels is developed to generate instance-adaptive residuals to adapt domain-agnostic deep features to each individual instance. This enables a shared classifier to be applied to both source and target domain data without relying on any domain annotation. Further, instead of imposing intricate feature alignment losses, we adopt a simple semi-supervised learning paradigm using only a cross-entropy loss for both labeled source and pseudo labeled target data. Our model, dubbed DIDA-Net, achieves state-of-the-art performance on several commonly used single-source and multi-source UDA datasets including Digits, Office-Home, DomainNet, Digit-Five, and PACS.
Zhongying Deng, Kaiyang Zhou, Da Li 0001, Junjun He, Yi-Zhe Song, Tao Xiang 0002
IEEE Trans. Image Process.2
2021 Domain Attention Consistency for Multi-Source Domain Adaptation
Zhongying Deng, Kaiyang Zhou, Yongxin Yang, Tao Xiang 0002
BMVC2
2021 Energy-Based Open-World Uncertainty Modeling for Confidence Calibration
abstract
Confidence calibration is of great importance to the reliability of decisions made by machine learning systems. However, discriminative classifiers based on deep neural networks are often criticized for producing overconfident predictions that fail to reflect the true correctness likelihood of classification accuracy. We argue that such an inability to model uncertainty is mainly caused by the closed-world nature in softmax: a model trained by the cross-entropy loss will be forced to classify input into one of K pre-defined categories with high probability. To address this problem, we for the first time propose a novel K+1-way softmax formulation, which incorporates the modeling of open-world uncertainty as the extra dimension. To unify the learning of the original K-way classification task and the extra dimension that models uncertainty, we 1) propose a novel energy-based objective function, and moreover, 2) theoretically prove that optimizing such an objective essentially forces the extra dimension to capture the marginal data distribution. Extensive experiments show that our approach, Energy-based Open-World Softmax (EOW-Softmax), is superior to existing state-of-the-art methods in improving confidence calibration.
Yezhen Wang, Bo Li 0080, Tong Che, Kaiyang Zhou, Ziwei Liu 0002, Dongsheng Li 0002
ICCV4
2021 Domain Generalization with MixStyle
Kaiyang Zhou, Yongxin Yang, Yu Qiao 0001, Tao Xiang 0002
ICLR1
2021 Domain Adaptive Ensemble Learning
abstract
The problem of generalizing deep neural networks from multiple source domains to a target one is studied under two settings: When unlabeled target data is available, it is a multi-source unsupervised domain adaptation (UDA) problem, otherwise a domain generalization (DG) problem. We propose a unified framework termed domain adaptive ensemble learning (DAEL) to address both problems. A DAEL model is composed of a CNN feature extractor shared across domains and multiple classifier heads each trained to specialize in a particular source domain. Each such classifier is an expert to its own domain but a non-expert to others. DAEL aims to learn these experts collaboratively so that when forming an ensemble, they can leverage complementary information from each other to be more effective for an unseen target domain. To this end, each source domain is used in turn as a pseudo-target-domain with its own expert providing supervisory signal to the ensemble of non-experts learned from the other sources. To deal with unlabeled target data under the UDA setting where real expert does not exist, DAEL uses pseudo labels to supervise the ensemble learning. Extensive experiments on three multi-source UDA datasets and two DG datasets show that DAEL improves the state of the art on both problems, often by significant margins.
Kaiyang Zhou, Yongxin Yang, Yu Qiao 0001, Tao Xiang 0002
IEEE Trans. Image Process.1
2020 Deep Domain-Adversarial Image Generation for Domain Generalisation
abstract
Machine learning models typically suffer from the domain shift problem when trained on a source dataset and evaluated on a target dataset of different distribution. To overcome this problem, domain generalisation (DG) methods aim to leverage data from multiple source domains so that a trained model can generalise to unseen domains. In this paper, we propose a novel DG approach based on Deep Domain-Adversarial Image Generation (DDAIG). Specifically, DDAIG consists of three components, namely a label classifier, a domain classifier and a domain transformation network (DoTNet). The goal for DoTNet is to map the source training data to unseen domains. This is achieved by having a learning objective formulated to ensure that the generated data can be correctly classified by the label classifier while fooling the domain classifier. By augmenting the source training data with the generated unseen domain data, we can make the label classifier more robust to unknown domain changes. Extensive experiments on four DG datasets demonstrate the effectiveness of our approach.
Kaiyang Zhou, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002
AAAI1
2020 Learning to Generate Novel Domains for Domain Generalization
Kaiyang Zhou, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002
ECCV (16)1
2019 Omni-Scale Feature Learning for Person Re-Identification
abstract
As an instance-level recognition problem, person re-identification (ReID) relies on discriminative features, which not only capture different spatial scales but also encapsulate an arbitrary combination of multiple scales. We callse features of both homogeneous and heterogeneous scales omni-scale features. In this paper, a novel deep ReID CNN is designed, termed Omni-Scale Network (OSNet), for omni-scale feature learning. This is achieved by designing a residual block composed of multiple convolutional feature streams, each detecting features at a certain scale. Importantly, a novel unified aggregation gate is introduced to dynamically fuse multi-scale features with input-dependent channel-wise weights. To efficiently learn spatial-channel correlations and avoid overfitting, the building block uses both pointwise and depthwise convolutions. By stacking such blocks layer-by-layer, our OSNet is extremely lightweight and can be trained from scratch on existing ReID benchmarks. Despite its small model size, our OSNet achieves state-of-the-art performance on six person-ReID datasets. Code and models are available at: https://github.com/KaiyangZhou/deep-person-reid.
Kaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao Xiang 0002
ICCV1
2018 Deep Reinforcement Learning for Unsupervised Video Summarization With Diversity-Representativeness Reward
abstract
Video summarization aims to facilitate large-scale video browsing by producing short, concise summaries that are diverse and representative of original videos. In this paper, we formulate video summarization as a sequential decision-making process and develop a deep summarization network (DSN) to summarize videos. DSN predicts for each video frame a probability, which indicates how likely a frame is selected, and then takes actions based on the probability distributions to select frames, forming video summaries. To train our DSN, we propose an end-to-end, reinforcement learning-based framework, where we design a novel reward function that jointly accounts for diversity and representativeness of generated summaries and does not rely on labels or user interactions at all. During training, the reward function judges how diverse and representative the generated summaries are, while DSN strives for earning higher rewards by learning to produce more diverse and more representative summaries. Since labels are not required, our method can be fully unsupervised. Extensive experiments on two benchmark datasets show that our unsupervised method not only outperforms other state-of-the-art unsupervised methods, but also is comparable to or even superior than most of published supervised approaches.
Kaiyang Zhou, Yu Qiao 0001, Tao Xiang 0002
AAAI1
2018 Video Summarisation by Classification with Deep Reinforcement Learning
Kaiyang Zhou, Tao Xiang 0002, Andrea Cavallaro
BMVC1