Jiahao Xie 0002

dblp:217/4325-2 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0001-9237-2802ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Representation and self-supervised learning · 56% Image recognition and object detection · 10% Segmentation and scene understanding · 9%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › self-supervised visual representation learning
self-supervised visual pre-training
1.222023
Correlational Image Modeling for Self-Supervised Visual Pre-Training · CVPR 2023
UniVIP: A Unified Framework for Self-Supervised Visual Pre-training · CVPR 2022
Machine learning › Generative modeling
diffusion model
0.912025
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation · Int. J. Comput. Vis. 2025
Computer vision › Segmentation and scene understanding
instance segmentation
0.912025
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation · Int. J. Comput. Vis. 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling
0.712023
Masked Frequency Modeling for Self-Supervised Visual Pre-Training · ICLR 2023
Machine learning › Representation and self-supervised learning › pre-training
visual pre-training
0.712023
Masked Frequency Modeling for Self-Supervised Visual Pre-Training · ICLR 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
UniVIP: A Unified Framework for Self-Supervised Visual Pre-training · CVPR 2022
Machine learning › Representation and self-supervised learning › contrastive learning
instance discrimination
0.612022
UniVIP: A Unified Framework for Self-Supervised Visual Pre-training · CVPR 2022
Machine learning › Trustworthy machine learning › out-of-distribution generalization
invariant learning
0.612022
Delving into Inter-Image Invariance for Unsupervised Visual Representations · Int. J. Comput. Vis. 2022
Machine learning › Learning paradigms › unsupervised learning
unsupervised visual representation learning
0.612022
Delving into Inter-Image Invariance for Unsupervised Visual Representations · Int. J. Comput. Vis. 2022
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.512021
Unsupervised Object-Level Representation Learning from Scene Images · NeurIPS 2021
Computer vision › Image recognition and object detection
object detection
0.512021
Unsupervised Object-Level Representation Learning from Scene Images · NeurIPS 2021
Computer vision › Image recognition and object detection
object discovery
0.512021
Unsupervised Object-Level Representation Learning from Scene Images · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
deep clustering
0.412020
Online Deep Clustering for Unsupervised Representation Learning · CVPR 2020
Machine learning › Time series and sequential data
online clustering
0.412020
Online Deep Clustering for Unsupervised Representation Learning · CVPR 2020
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.412020
Online Deep Clustering for Unsupervised Representation Learning · CVPR 2020
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning
0.212023
Correlational Image Modeling for Self-Supervised Visual Pre-Training · CVPR 2023
Machine learning › Representation and self-supervised learning
pre-training
0.112021
Unsupervised Object-Level Representation Learning from Scene Images · NeurIPS 2021
Machine learning › Representation and self-supervised learning › pre-training
unsupervised pre-training
0.112021
Unsupervised Object-Level Representation Learning from Scene Images · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.7data augmentation · 1.4diffusion model · 0.9self-supervised pretraining · 0.7masked frequency modeling · 0.7cross-attention · 0.7bootstrap learning · 0.7optimal transport · 0.6online deep clustering · 0.4memory module · 0.4
YearPublicationVenuePosition
2025 MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
Jiahao Xie 0002, Wei Li 0319, Xiangtai Li, Ziwei Liu 0002, Yew-Soon Ong, Chen Change Loy
Int. J. Comput. Vis.1
2023 Correlational Image Modeling for Self-Supervised Visual Pre-Training
abstract
We introduce Correlational Image Modeling (CIM), a novel and surprisingly effective approach to self-supervised visual pre-training. Our CIM performs a simple pretext task: we randomly crop image regions (exemplars) from an input image (context) and predict correlation maps between the exemplars and the context. Three key designs enable correlational image modeling as a nontrivial and meaningful self-supervisory task. First, to generate useful exemplar-context pairs, we consider cropping image regions with various scales, shapes, rotations, and transformations. Second, we employ a bootstrap learning framework that involves online and target encoders. During pre-training, the former takes exemplars as inputs while the latter converts the context. Third, we model the output correlation maps via a simple cross-attention block, within which the context serves as queries and the exemplars offer values and keys. We show that CIM performs on par or better than the current state of the art on self-supervised and transfer benchmarks. Code is available at https://github.com/weivision/Correlational-Image-Modeling.git.
Wei Li 0319, Jiahao Xie 0002, Chen Change Loy
CVPR2
2023 Masked Frequency Modeling for Self-Supervised Visual Pre-Training
Jiahao Xie 0002, Wei Li 0319, Xiaohang Zhan, Ziwei Liu 0002, Yew-Soon Ong, Chen Change Loy
ICLR1
2022 UniVIP: A Unified Framework for Self-Supervised Visual Pre-training
abstract
Self-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images like those in ImageNet and ignores the correlation among the scene and instances, as well as the semantic difference of instances in the scene. To address the above problems, we propose a Unified Self-supervised Visual Pre-training (UniVIP), a novel self-supervised framework to learn versatile visual representations on either single-centric-object or non-iconic dataset. The framework takes into account the representation learning at three levels: 1) the similarity of scene-scene, 2) the correlation of scene-instance, 3) the discrimination of instance-instance. During the learning, we adopt the optimal transport algorithm to automatically measure the discrimination of instances. Massive experiments show that Uni-VIP pre-trained on non-iconic COCO achieves state-of-the-art transfer performance on a variety of downstream tasks, such as image classification, semi-supervised learning, object detection and segmentation. Furthermore, our method can also exploit single-centric-object dataset such as ImageNet and outperforms BYOL by 2.5% with the same pre-training epochs in linear probing, and surpass current self-supervised object detection methods on COCO dataset, demonstrating its universality and potential.
Zhaowen Li, Yousong Zhu, Fan Yang 0089, Wei Li 0314, Chaoyang Zhao, Yingying Chen 0003, Zhiyang Chen 0002, Jiahao Xie 0002, Rui Zhao 0001, Ming Tang 0001, Jinqiao Wang
CVPR8
2022 Delving into Inter-Image Invariance for Unsupervised Visual Representations
Jiahao Xie 0002, Xiaohang Zhan, Ziwei Liu 0002, Yew-Soon Ong, Chen Change Loy
Int. J. Comput. Vis.1
2021 Unsupervised Object-Level Representation Learning from Scene Images
abstract
Contrastive self-supervised learning has largely narrowed the gap to supervised pre-training on ImageNet. However, its success highly relies on the object-centric priors of ImageNet, i.e., different augmented views of the same image correspond to the same object. Such a heavily curated constraint becomes immediately infeasible when pre-trained on more complex scene images with many objects. To overcome this limitation, we introduce Object-level Representation Learning (ORL), a new self-supervised learning framework towards scene images. Our key insight is to leverage image-level self-supervised pre-training as the prior to discover object-level semantic correspondence, thus realizing object-level representation learning from scene images. Extensive experiments on COCO show that ORL significantly improves the performance of self-supervised learning on scene images, even surpassing supervised ImageNet pre-training on several downstream tasks. Furthermore, ORL improves the downstream performance when more unlabeled scene images are available, demonstrating its great potential of harnessing unlabeled data in the wild. We hope our approach can motivate future research on more general-purpose unsupervised representation learning from scene data.
Jiahao Xie 0002, Xiaohang Zhan, Ziwei Liu 0002, Yew-Soon Ong, Chen Change Loy
NeurIPS1
2020 Online Deep Clustering for Unsupervised Representation Learning
abstract
Joint clustering and feature learning methods have shown remarkable performance in unsupervised representation learning. However, the training schedule alternating between feature clustering and network parameters update leads to unstable learning of visual representations. To overcome this challenge, we propose Online Deep Clustering (ODC) that performs clustering and network update simultaneously rather than alternatingly. Our key insight is that the cluster centroids should evolve steadily in keeping the classifier stably updated. Specifically, we design and maintain two dynamic memory modules, i.e., samples memory to store samples' labels and features, and centroids memory for centroids evolution. We break down the abrupt global clustering into steady memory update and batch-wise label re-assignment. The process is integrated into network update iterations. In this way, labels and the network evolve shoulder-to-shoulder rather than alternatingly. Extensive experiments demonstrate that ODC stabilizes the training process and boosts the performance effectively.
Xiaohang Zhan, Jiahao Xie 0002, Ziwei Liu 0002, Yew-Soon Ong, Chen Change Loy
CVPR2