VLDB 2026 Research / reviewers in the wild / expert
Jiahao Xie 0002
dblp:217/4325-2
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0001-9237-2802ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Representation and self-supervised learning · 56% Image recognition and object detection · 10% Segmentation and scene understanding · 9% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › self-supervised visual representation learning
self-supervised visual pre-training |
1.2 | 2 | 2023 | Correlational Image Modeling for Self-Supervised Visual Pre-Training · CVPR 2023 UniVIP: A Unified Framework for Self-Supervised Visual Pre-training · CVPR 2022 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation · Int. J. Comput. Vis. 2025 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.9 | 1 | 2025 | MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation · Int. J. Comput. Vis. 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling |
0.7 | 1 | 2023 | Masked Frequency Modeling for Self-Supervised Visual Pre-Training · ICLR 2023 |
Machine learning › Representation and self-supervised learning › pre-training
visual pre-training |
0.7 | 1 | 2023 | Masked Frequency Modeling for Self-Supervised Visual Pre-Training · ICLR 2023 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.6 | 1 | 2022 | UniVIP: A Unified Framework for Self-Supervised Visual Pre-training · CVPR 2022 |
Machine learning › Representation and self-supervised learning › contrastive learning
instance discrimination |
0.6 | 1 | 2022 | UniVIP: A Unified Framework for Self-Supervised Visual Pre-training · CVPR 2022 |
Machine learning › Trustworthy machine learning › out-of-distribution generalization
invariant learning |
0.6 | 1 | 2022 | Delving into Inter-Image Invariance for Unsupervised Visual Representations · Int. J. Comput. Vis. 2022 |
Machine learning › Learning paradigms › unsupervised learning
unsupervised visual representation learning |
0.6 | 1 | 2022 | Delving into Inter-Image Invariance for Unsupervised Visual Representations · Int. J. Comput. Vis. 2022 |
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning |
0.5 | 1 | 2021 | Unsupervised Object-Level Representation Learning from Scene Images · NeurIPS 2021 |
Computer vision › Image recognition and object detection
object detection |
0.5 | 1 | 2021 | Unsupervised Object-Level Representation Learning from Scene Images · NeurIPS 2021 |
Computer vision › Image recognition and object detection
object discovery |
0.5 | 1 | 2021 | Unsupervised Object-Level Representation Learning from Scene Images · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
deep clustering |
0.4 | 1 | 2020 | Online Deep Clustering for Unsupervised Representation Learning · CVPR 2020 |
Machine learning › Time series and sequential data
online clustering |
0.4 | 1 | 2020 | Online Deep Clustering for Unsupervised Representation Learning · CVPR 2020 |
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning |
0.4 | 1 | 2020 | Online Deep Clustering for Unsupervised Representation Learning · CVPR 2020 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.2 | 1 | 2023 | Correlational Image Modeling for Self-Supervised Visual Pre-Training · CVPR 2023 |
Machine learning › Representation and self-supervised learning
pre-training |
0.1 | 1 | 2021 | Unsupervised Object-Level Representation Learning from Scene Images · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning › pre-training
unsupervised pre-training |
0.1 | 1 | 2021 | Unsupervised Object-Level Representation Learning from Scene Images · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 1.7data augmentation · 1.4diffusion model · 0.9self-supervised pretraining · 0.7masked frequency modeling · 0.7cross-attention · 0.7bootstrap learning · 0.7optimal transport · 0.6online deep clustering · 0.4memory module · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
Jiahao Xie 0002, Wei Li 0319, Xiangtai Li, Ziwei Liu 0002, Yew-Soon Ong, Chen Change Loy |
Int. J. Comput. Vis. | 1 |
| 2023 | Correlational Image Modeling for Self-Supervised Visual Pre-TrainingabstractWe introduce Correlational Image Modeling (CIM), a novel and surprisingly effective approach to self-supervised visual pre-training. Our CIM performs a simple pretext task: we randomly crop image regions (exemplars) from an input image (context) and predict correlation maps between the exemplars and the context. Three key designs enable correlational image modeling as a nontrivial and meaningful self-supervisory task. First, to generate useful exemplar-context pairs, we consider cropping image regions with various scales, shapes, rotations, and transformations. Second, we employ a bootstrap learning framework that involves online and target encoders. During pre-training, the former takes exemplars as inputs while the latter converts the context. Third, we model the output correlation maps via a simple cross-attention block, within which the context serves as queries and the exemplars offer values and keys. We show that CIM performs on par or better than the current state of the art on self-supervised and transfer benchmarks. Code is available at https://github.com/weivision/Correlational-Image-Modeling.git. Wei Li 0319, Jiahao Xie 0002, Chen Change Loy |
CVPR | 2 |
| 2023 | Masked Frequency Modeling for Self-Supervised Visual Pre-Training
Jiahao Xie 0002, Wei Li 0319, Xiaohang Zhan, Ziwei Liu 0002, Yew-Soon Ong, Chen Change Loy |
ICLR | 1 |
| 2022 | UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingabstractSelf-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images like those in ImageNet and ignores the correlation among the scene and instances, as well as the semantic difference of instances in the scene. To address the above problems, we propose a Unified Self-supervised Visual Pre-training (UniVIP), a novel self-supervised framework to learn versatile visual representations on either single-centric-object or non-iconic dataset. The framework takes into account the representation learning at three levels: 1) the similarity of scene-scene, 2) the correlation of scene-instance, 3) the discrimination of instance-instance. During the learning, we adopt the optimal transport algorithm to automatically measure the discrimination of instances. Massive experiments show that Uni-VIP pre-trained on non-iconic COCO achieves state-of-the-art transfer performance on a variety of downstream tasks, such as image classification, semi-supervised learning, object detection and segmentation. Furthermore, our method can also exploit single-centric-object dataset such as ImageNet and outperforms BYOL by 2.5% with the same pre-training epochs in linear probing, and surpass current self-supervised object detection methods on COCO dataset, demonstrating its universality and potential. Zhaowen Li, Yousong Zhu, Fan Yang 0089, Wei Li 0314, Chaoyang Zhao, Yingying Chen 0003, Zhiyang Chen 0002, Jiahao Xie 0002, Rui Zhao 0001, Ming Tang 0001, Jinqiao Wang |
CVPR | 8 |
| 2022 | Delving into Inter-Image Invariance for Unsupervised Visual Representations
Jiahao Xie 0002, Xiaohang Zhan, Ziwei Liu 0002, Yew-Soon Ong, Chen Change Loy |
Int. J. Comput. Vis. | 1 |
| 2021 | Unsupervised Object-Level Representation Learning from Scene ImagesabstractContrastive self-supervised learning has largely narrowed the gap to supervised pre-training on ImageNet. However, its success highly relies on the object-centric priors of ImageNet, i.e., different augmented views of the same image correspond to the same object. Such a heavily curated constraint becomes immediately infeasible when pre-trained on more complex scene images with many objects. To overcome this limitation, we introduce Object-level Representation Learning (ORL), a new self-supervised learning framework towards scene images. Our key insight is to leverage image-level self-supervised pre-training as the prior to discover object-level semantic correspondence, thus realizing object-level representation learning from scene images. Extensive experiments on COCO show that ORL significantly improves the performance of self-supervised learning on scene images, even surpassing supervised ImageNet pre-training on several downstream tasks. Furthermore, ORL improves the downstream performance when more unlabeled scene images are available, demonstrating its great potential of harnessing unlabeled data in the wild. We hope our approach can motivate future research on more general-purpose unsupervised representation learning from scene data. Jiahao Xie 0002, Xiaohang Zhan, Ziwei Liu 0002, Yew-Soon Ong, Chen Change Loy |
NeurIPS | 1 |
| 2020 | Online Deep Clustering for Unsupervised Representation LearningabstractJoint clustering and feature learning methods have shown remarkable performance in unsupervised representation learning. However, the training schedule alternating between feature clustering and network parameters update leads to unstable learning of visual representations. To overcome this challenge, we propose Online Deep Clustering (ODC) that performs clustering and network update simultaneously rather than alternatingly. Our key insight is that the cluster centroids should evolve steadily in keeping the classifier stably updated. Specifically, we design and maintain two dynamic memory modules, i.e., samples memory to store samples' labels and features, and centroids memory for centroids evolution. We break down the abrupt global clustering into steady memory update and batch-wise label re-assignment. The process is integrated into network update iterations. In this way, labels and the network evolve shoulder-to-shoulder rather than alternatingly. Extensive experiments demonstrate that ODC stabilizes the training process and boosts the performance effectively. Xiaohang Zhan, Jiahao Xie 0002, Ziwei Liu 0002, Yew-Soon Ong, Chen Change Loy |
CVPR | 2 |