Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chengxiang Fan

dblp:353/0691 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0000-2555-4112ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Segmentation and scene understanding · 45% Video understanding and tracking · 18% Deep learning architectures and training · 13%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
instance segmentation
2.232024
Generative Active Learning for Long-tailed Instance Segmentation · ICML 2024
DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data · CVPR 2024
SegPrompt: Boosting Open-world Segmentation via Category-level Prompt Learning · ICCV 2023
Machine learning › Generative modeling › diffusion model › diffusion model training
diffusion model fine-tuning
0.912025
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? · ICLR 2025
Computer vision › Segmentation and scene understanding
image segmentation
0.912025
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? · ICLR 2025
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.912025
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? · ICLR 2025
Machine learning › Efficient and distributed learning
active learning
0.812024
Generative Active Learning for Long-tailed Instance Segmentation · ICML 2024
Machine learning › Deep learning architectures and training
data augmentation
0.812024
DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data · CVPR 2024
Machine learning › Deep learning architectures and training › data augmentation
generative data augmentation
0.812024
DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data · CVPR 2024
Computer vision › Segmentation and scene understanding › instance segmentation
long-tailed instance segmentation
0.812024
Generative Active Learning for Long-tailed Instance Segmentation · ICML 2024
Computer vision › Segmentation and scene understanding › object segmentation
class-agnostic segmentation
0.712023
SegPrompt: Boosting Open-world Segmentation via Category-level Prompt Learning · ICCV 2023
Computer vision › Video understanding and tracking
instance embedding
0.712023
CTVIS: Consistent Training for Online Video Instance Segmentation · ICCV 2023
Computer vision › Video understanding and tracking › video instance segmentation
online video instance segmentation
0.712023
CTVIS: Consistent Training for Online Video Instance Segmentation · ICCV 2023
Computer vision › Segmentation and scene understanding › instance segmentation
open-world instance segmentation
0.712023
SegPrompt: Boosting Open-world Segmentation via Category-level Prompt Learning · ICCV 2023
Computer vision › Video understanding and tracking
video instance segmentation
0.712023
CTVIS: Consistent Training for Online Video Instance Segmentation · ICCV 2023
Machine learning › Transfer learning and domain adaptation › cross-domain transfer
cross-dataset transfer
0.212023
SegPrompt: Boosting Open-world Segmentation via Category-level Prompt Learning · ICCV 2023

Methods — techniques the papers use, named apart from their topics

one-step fine-tuning · 0.9diffusion prior · 0.9prompt diversity · 0.8gradient cache · 0.8generative model · 0.8generative data selection · 0.8momentum-averaged embedding · 0.7memory bank · 0.7contrastive learning · 0.7adversarial training · 0.7
YearPublicationVenuePosition
2025 What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
abstract
Extensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I diffusion models for a few dense perception tasks. However, several crucial design decisions in this process still lack comprehensive justification, encompassing the necessity of the multi-step diffusion mechanism, training strategy, inference ensemble strategy, and fine-tuning data quality. In this work, we conduct a thorough investigation into critical factors that affect transfer efficiency and performance when using diffusion priors. Our key findings are: 1) High-quality fine-tuning data is paramount for both semantic and geometry perception tasks. 2) As a special case of the diffusion scheduler by setting its hyper-parameters, the multi-step generation can be simplified to a one-step fine-tuning paradigm without any loss of performance, while significantly speeding up inference. 3) Apart from fine-tuning the diffusion model with only latent space supervision, task-specific supervision can be beneficial to enhance fine-grained details. These observations culminate in the development of GenPercept, an effective deterministic one-step fine-tuning paradigm tailored for dense visual perception tasks exploiting diffusion priors. Different from the previous multi-step methods, our paradigm offers a much faster inference speed, and can be seamlessly integrated with customized perception decoders and loss functions for task-specific supervision, which can be critical for improving the fine-grained details of predictions. Comprehensive experiments on a diverse set of dense visual perceptual tasks, including monocular depth estimation, surface normal estimation, image segmentation, and matting, are performed to demonstrate the remarkable adaptability and effectiveness of our proposed method. Code: https://github.com/aim-uofa/GenPercept
Guangkai Xu, Yongtao Ge, Chengxiang Fan, Kangyang Xie, Zhiyue Zhao, Hao Chen 0041, Chunhua Shen
ICLR4
2024 DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data
abstract
Instance segmentation is data-hungry, and as model capacity increases, data scale becomes crucial for improving the accuracy. Most instance segmentation datasets today require costly manual annotation, limiting their data scale. Models trained on such data are prone to overfitting on the training set, especially for those rare categories. While recent works have delved into exploiting generative models to create synthetic datasets for data augmentation, these approaches do not efficiently harness the full potential of generative models. To address these issues, we introduce a more efficient strategy to construct generative datasets for data augmentation, termed DiverGen. Firstly, we provide an explanation of the role of generative data from the perspective of distribution discrepancy. We investigate the impact of different data on the distribution learned by the model. We argue that generative data can expand the data distribution that the model can learn, thus mitigating overfitting. Additionally, we find that the diversity of generative data is crucial for improving model performance and enhance it through various strategies, including category diversity, prompt diversity, and generative model diversity. With these strategies, we can scale the data to millions while maintaining the trend of model performance improvement. On the LVIS dataset, DiverGen significantly outperforms the strong model X-Paste, achieving +1.1 box AP and +1.1 mask AP across all categories, and +1.9 box AP and +2.5 mask AP for rare categories. Our codes are available at https://github.com/aim-uofa/DiverGen.
Chengxiang Fan, Muzhi Zhu, Hao Chen 0041, Yang Liu 0357, Weijia Wu 0001, Huaqi Zhang, Chunhua Shen
CVPR1
2024 Dual-Domain Multi-Model GAN Fingerprint Restoration for Compressed Fake Face Attribution
abstract
Recent advances in GAN fingerprint have shown increasing success in fake face attribution. However, the fake faces are usually compressed during network transmission, which causes the degradation of GAN fingerprint and the decrease of attribution accuracy. To this issue, a dual-domain multi-model GAN fingerprint restoration method for compressed fake face attribution is proposed in this paper. Firstly, considering that image-domain and fingerprint-domain are directly and indirectly affected by compression respectively, we propose a dual-domain parallel restoration architecture that enhances GAN fingerprint using direct image-domain and indirect fingerprint-domain restoration, thereby improving attribution performance by mining the cross-domain complementarity. Secondly, since real and fake GAN-speciffc restoration models can describe GAN fingerprint from different aspects, we first enhance GAN fingerprint by multiple restoration models, and then improve attribution performance by exploiting the cross-model complementarity through the multi-model restoration fusion strategy. Experiments demonstrate the superiority of our method under different compression qualities.
Chengxiang Fan, Aohong Shen, Zhen Han 0002, Cai Tong, Zhongyuan Wang 0001, Dekang Yi
ICME1
2024 Generative Active Learning for Long-tailed Instance Segmentation
abstract
Recently, large-scale language-image generative models have gained widespread attention and many works have utilized generated data from these models to further enhance the performance of perception tasks. However, not all generated data can positively impact downstream models, and these methods do not thoroughly explore how to better select and utilize generated data. On the other hand, there is still a lack of research oriented towards active learning on generated data. In this paper, we explore how to perform active learning specifically for generated data in the long-tailed instance segmentation task. Subsequently, we propose BSGAL, a new algorithm that estimates the contribution of the current batch-generated data based on gradient cache. BSGAL is meticulously designed to cater for unlimited generated data and complex downstream segmentation tasks. BSGAL outperforms the baseline approach and effectually improves the performance of long-tailed segmentation.
Muzhi Zhu, Chengxiang Fan, Hao Chen 0041, Yang Liu 0357, Weian Mao, Xiaogang Xu 0002, Chunhua Shen
ICML2
2023 CTVIS: Consistent Training for Online Video Instance Segmentation
abstract
The discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/negative embeddings. Recent online VIS methods leverage CIs sourced from one reference frame only, which we argue is insufficient for learning highly discriminative embeddings. Intuitively, a possible strategy to enhance CIs is replicating the inference phase during training. To this end, we propose a simple yet effective training strategy, called Consistent Training for Online VIS (CTVIS), which devotes to aligning the training and inference pipelines in terms of building CIs. Specifically, CTVIS constructs CIs by referring inference the momentum-averaged embedding and the memory bank storage mechanisms, and adding noise to the relevant embeddings. Such an extension allows a reliable comparison between embeddings of current instances and the stable representations of historical instances, thereby conferring an advantage in modeling VIS challenges such as occlusion, re-identification, and deformation. Empirically, CTVIS outstrips the SOTA VIS models by up to +5.0 points on three VIS benchmarks, including YTVIS19 (55.1% AP), YTVIS21 (50.1% AP) and OVIS (35.5% AP). Furthermore, we find that pseudo-videos transformed from images can train robust models surpassing fully-supervised ones.
Kaining Ying, Weian Mao, Zhenhua Wang 0003, Hao Chen 0041, Lin Wu 0001, Yifan Liu 0001, Chengxiang Fan, Yunzhi Zhuge, Chunhua Shen
ICCV8
2023 SegPrompt: Boosting Open-world Segmentation via Category-level Prompt Learning
abstract
Current closed-set instance segmentation models rely on pre-defined class labels for each mask during training and evaluation, largely limiting their ability to detect novel objects. Open-world instance segmentation (OWIS) models address this challenge by detecting unknown objects in a class-agnostic manner. However, previous OWIS approaches completely erase category information during training to keep the model’s ability to generalize to unknown objects. In this work, we propose a novel training mechanism termed SegPrompt that uses category information to improve the model’s class-agnostic segmentation ability for both known and unknown categories. In addition, the previous OWIS training setting exposes the unknown classes to the training set and brings information leakage, which is unreasonable in the real world. Therefore, we provide a new open-world benchmark closer to a real-world scenario by dividing the dataset classes into known-seen-unseen parts. For the first time, we focus on the model’s ability to discover objects that never appear in the training set images.Experiments show that SegPrompt can improve the overall and unseen detection performance by 5.6% and 6.1% in AR on our new benchmark without affecting the inference efficiency. We further demonstrate the effectiveness of our method on existing cross-dataset transfer and strongly supervised settings, leading to 5.5% and 12.3% relative improvement. Code and data are released at: https://github.com/aim-uofa/SegPrompt
Muzhi Zhu, Hengtao Li, Hao Chen 0041, Chengxiang Fan, Weian Mao, Chenchen Jing, Yifan Liu 0001, Chunhua Shen
ICCV4