Zhichao Sun 0004

dblp:152/6104-4 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
13since 2021 · last 2026
0009-0006-3038-6006ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Coarse-to-fine crack cue for robust crack detection
Zelong Liu, Yuliang Gu, Zhichao Sun 0004, Huachao Zhu, Xin Xiao 0010, Bo Du 0001, Laurent Najman, Yongchao Xu
Pattern Recognit.3
2026 Anomaly or Characteristic: Memory-Based Coarse-to-Fine Feature Fusion for Industrial Anomaly Detection
abstract
Unsupervised anomaly detection methods primarily focus on modeling the distribution of normal samples at image/feature level. Significant deviation from the modeled distribution is then considered as anomaly. Yet, each normal sample may have its own unique characteristic drifting from the idea distribution, making it difficult to distinguish between anomaly and characteristic. In this paper, we propose a memory-based Coarse-to-Fine Feature Fusion (C3F) module to tackle this challenge. Specifically, we construct a group of memory banks that model the feature distribution of normal samples at various levels of granularity. The memory-based C3F is applied to each skip-connection between the encoder and the decoder, and progressively removes anomaly while maintaining characteristic. This helps to reconstruct defect-free image with characteristic preserved, encouraging large (respsmall) deviation from the modeled distribution for anomaly (respcharacteristic). Besides, we also introduce a novel rough anomaly score map Guided Segmentation (GS) module to achieve precise anomaly localization. Extensive experiments on widely used VisA and MVTec-AD benchmarks demonstrate the wide-ranging applicability of the proposed method termed C3FGS on industrial components of various forms. The implementation code is publicly available athttps://github.com/LZL501/c3f_industrial_anomaly_detection.
Huachao Zhu, Zelong Liu, Zhichao Sun 0004, Xin Xiao 0010, Yongchao Xu
IEEE Trans. Multim.3
2025 Beyond Pixel Uncertainty: Bounding the OoD Objects in Road Scenes
Huachao Zhu, Zelong Liu, Zhichao Sun 0004, Yuda Zou, Gui-Song Xia, Yongchao Xu
ICCV3
2025 Test-Time Training with Local Contrast-Preserving Copy-Pasted Image for Domain Generalization in Retinal Vessel Segmentation
Yuliang Gu, Zhichao Sun 0004, Zelong Liu, Yongchao Xu
MICCAI (7)2
2025 Neighborhood-Consistent Binary Transformation for Domain-Invariant Chest X-Ray Diagnosis
Zelong Liu, Huachao Zhu, Zhichao Sun 0004, Yuda Zou, Yuliang Gu, Bo Du 0001, Yongchao Xu
MICCAI (5)3
2025 Dual structure-aware image filterings for semi-supervised medical image segmentation
Yuliang Gu, Zhichao Sun 0004, Xin Xiao 0010, Yepeng Liu 0002, Yongchao Xu, Laurent Najman
Medical Image Anal.2
2025 MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching
abstract
Many keypoint detection and description methods have been proposed for image matching or registration. While these methods demonstrate promising performance for single-modality image matching, they often struggle with multimodal data because the descriptors trained on single-modality data tend to lack robustness against the non-linear variations present in multimodal data. Extending such methods to multimodal image matching often requires well-aligned multimodal data to learn modality-invariant descriptors. However, acquiring such data is often costly and impractical in many real-world scenarios. To address this challenge, we propose a modality-invariant feature learning network (MIFNet) to compute modality-invariant features for keypoint descriptions in multimodal image matching using only single-modality training data. Specifically, we propose a novel latent feature aggregation module and a cumulative hybrid aggregation module to enhance the base keypoint descriptors trained on single-modality data by leveraging pre-trained features from Stable Diffusion models. We validate our method with recent keypoint detection and description methods in three multimodal retinal image datasets (CF-FA, CF-OCT, EMA-OCTA) and two remote sensing datasets (Optical-SAR and Optical-NIR). Extensive experiments demonstrate that the proposed MIFNet is able to learn modality-invariant feature for multimodal image matching without accessing the targeted modality and has good zero-shot generalization ability. The code will be released at https://github.com/lyp-deeplearning/MIFNet.
Yepeng Liu 0002, Zhichao Sun 0004, Baosheng Yu, Yitian Zhao, Bo Du 0001, Yongchao Xu, Jun Cheng 0003
IEEE Trans. Image Process.2
2024 Shape Transformation Driven by Active Contour for Class-Imbalanced Semi-Supervised Medical Image Segmentation
abstract
Annotating 3D medical images demands expert knowledge and is time-consuming. As a result, semi-supervised learning (SSL) approaches have gained significant interest in 3D medical image segmentation. The significant size differences among various organs in the human body lead to imbalanced class distribution, which is a major challenge in the real-world application of these SSL approaches. To address this issue, we develop a novel Shape Transformation driven by Active Contour (STAC), that enlarges smaller organs to alleviate imbalanced class distribution across different organs. Inspired by curve evolution theory in active contour methods, STAC employs a signed distance function (SDF) as the level set function, to implicitly represent the shape of organs, and deforms voxels in the direction of the steepest descent of SDF (i.e., the normal vector). To ensure that the voxels far from expansion organs remain unchanged, we design an SDF-based weight function to control the degree of deformation for each voxel. We then use STAC as a data-augmentation process during the training stage. Experimental results on two benchmark datasets demonstrate that the proposed method significantly outperforms some state-of-the-art methods. Source code is publicly available at https://github.com/GuGuLL123/STAC.
Yuliang Gu, Yepeng Liu 0002, Zhichao Sun 0004, Jinchi Zhu, Yongchao Xu, Laurent Najman
BIBM3
2024 Shifted Autoencoders for Point Annotation Restoration in Object Counting
Yuda Zou, Xin Xiao 0010, Peilin Zhou, Zhichao Sun 0004, Bo Du 0001, Yongchao Xu
ECCV (25)4
2024 P2A: Transforming Proposals to Anomaly Masks
Huachao Zhu, Zhichao Sun 0004, Zelong Liu, Yongchao Xu
ICPR (33)2
2024 Position-Guided Prompt Learning for Anomaly Detection in Chest X-Rays
Zhichao Sun 0004, Yuliang Gu, Yepeng Liu 0002, Yongchao Xu
MICCAI (1)1
2024 Spatial-Aware Attention Generative Adversarial Network for Semi-supervised Anomaly Detection in Medical Image
Zhichao Sun 0004, Zelong Liu, Rui Yu 0002, Bo Du 0001, Yongchao Xu
MICCAI (5)2
2024 Nighttime image semantic segmentation with retinex theory
Zhichao Sun 0004, Huachao Zhu, Xin Xiao 0010, Yuliang Gu, Yongchao Xu
Image Vis. Comput.1