Zhikang Zhao

dblp:308/4229 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-6515-1454ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 ITJP: Image and Text Joint Prompts for Few-Shot Whole Slide Image Classification
abstract
Multiple instance learning has achieved remarkable results in whole slide image (WSI) classification. However, constrained by patient privacy and the scarcity of cancer data, the insufficient quantity of data poses a challenge in training models with the extensive parameters required for large-pixel WSIs, giving rise to the requirement of few-shot WSI classification. Recently, pre-trained vision-language models (preVLM) have demonstrated great superiority in few-shot WSI classification tasks due to their good transfer-learning and few-shot capabilities. Therefore, we propose an Image and Text Joint Prompts for few-shot whole slide image classification method, ITJP, to construct a few-shot WSI classification model under the multiple instance learning framework. ITJP utilizes the pre-VLM to extract instance features under the unsupervised setting, constructs image prompts to guide the aggregation of instance features into bag features, and subsequently guides the classification of bag features by text prompts. Specifically, we propose an image prototype guided aggregation strategy, where image prototype is obtained by clustering patches extracted from representative WSIs. The image prototype guides the aggregation of instance features and is directly compared with the image patches to achieve more accurate similarity, thereby enhancing the focus on classification-relevant features. Furthermore, low-rank linear transformation is designed to obtain variable image prototype, enhancing adaptability in specific WSI classification task and enabling effective adaptation to the few-shot scenario. We conduct extensive experiments on two WSI datasets to demonstrate the significant performance of the ITJP for few-shot WSI classification.
Xinzhu Zhang, Zhikang Zhao, Jing Zhao 0015
ICME3
2025 RPMIL: Rethinking Uncertainty-Aware Probabilistic Multiple Instance Learning for Whole Slide Pathology Diagnosis
abstract
Whole slide images (WSIs) are gigapixel digital scans of traditional pathology slides, offering substantial support for cancer diagnosis. Current multiple instance learning (MIL) methods for WSIs typically extract instance features and aggregate these into a single bag feature for prediction. We observe that these MIL methods rely on point estimation, where each bag is mapped to a deterministic embedding. Such MIL methods based on point estimation fail to capture the full spectrum of data variability due to the reliance on fixed embedding, especially when the number of trainable bags is limited. In this paper, we rethink probabilistic modeling in MIL and propose RPMIL, an uncertainty-aware probabilistic MIL method for whole slide pathology diagnosis. RPMIL learns a probabilistic aggregator to consolidate instance features into dynamic bag feature distributions instead of a deterministic bag feature. Specifically, we employ a variational autoencoder approach to compress multiple instance features into a low-dimension space with probabilistic representation and obtain the bag feature distribution formulated by the mean and variance. Furthermore, we drive the prediction by jointly leveraging the instance feature distribution and bag feature distribution. We evaluate the WSI classification performance on two public datasets: Camelyon16 and TCGA-NSCLC. Extensive experiments demonstrate that our method surpasses point estimation methods in MIL, achieving state-of-the-art levels.
Zhikang Zhao, Kaitao Chen, Jing Zhao 0015
IJCAI1
2025 Multiple instance learning with hierarchical discrimination and smoothing attention for histopathological diagnosis
Jing Zhao 0015, Zhikang Zhao, Xueru Song, Shiliang Sun
Appl. Intell.2
2025 Local to Global: A Sparse Transformer-Based Small Object Detector for Remote Sensing Images
abstract
Object detection plays a crucial role in remote sensing due to the urgent demands of various applications, such as urban planning and environmental monitoring. Despite notable progress, current methods still struggle with detecting challenging small objects. At the object level, the limited pixel representation, blurred details, and background interference of small objects impose greater demands on feature extractors. At the network level, resource bias fails to provide adequate learning signals for these objects. In this paper, we propose a Sparse Transformer-based detector (STDet) to tackle these challenges. Specifically, we design a Local-to-Global Transformer network (LGFormer) to explore essential feature representations. The Local Transformer Block establishes correlations between tokens and their surrounding data, while the Global Transformer Block captures long-distance dependencies related to the objects. Meanwhile, we introduce a Scale-Balanced Label Assignment (SBLA) strategy that considers more samples to small objects. SBLA dynamically shifts the learning focus to easily overlooked objects and alleviates the issue of sample imbalance. Extensive experiments on three large-scale remote sensing datasets demonstrate the effectiveness of STDet and its superiority in small object detection.
Zheng Li 0027, Yongcheng Wang 0001, Hao Feng 0008, Chi Chen 0003, Dongdong Xu 0001, Yunxiao Gao, Zhikang Zhao
IEEE Trans. Geosci. Remote. Sens.8
2024 A method of degradation mechanism-based unsupervised remote sensing image super-resolution
Zhikang Zhao, Yongcheng Wang 0001, Ning Zhang 0025, Yuxi Zhang 0003, Zheng Li 0027, Chi Chen 0003
Image Vis. Comput.1
2024 Corrigendum to "A method of degradation mechanism-based unsupervised remote sensing image super-resolution" [Image and Vision Computing, Vol 148 (2024), 105108]
Zhikang Zhao
Image Vis. Comput.1
2024 Context Feature Integration and Balanced Sampling Strategy for Small Weak Object Detection in Remote Sensing Imagery
abstract
Deep learning has made significant achievements in remote sensing object detection tasks. However, small weak objects located in complex scenes are still not effectively addressed. The lack of feature information and negligible contributions during the optimization stage are the main reasons. To solve the indicated issues, a novel remote-sensing object detection method is proposed in this letter. Firstly, the context feature integration module (CFIM) is designed to extract implicit clues co-occurring with the object to compensate for the lack of features in small weak objects. The receptive field expansive deformable convolution (RFConv) constructed in CFIM can adaptively adjust the information extraction range based on the object’s characteristics, thereby capturing suitable context features. Secondly, to make small weak objects competitive during the optimization process, we propose tailored optimization functions: the balanced sampling strategy (BSS) and the modulated loss function (MLF). BSS makes up for the sample deficiency of small weak objects by balancing sampling. More specifically, BSS dynamically mines potential positive samples from the ignored set in order to increase the chances of matching. MLF progressively adjusts the loss proportion of the object and pushes the detector to be more sensitive to small weak objects during the training stage. Our proposed method achieves 94.79% and 71.23% mAP on the NWPU VHR-10 and DIOR datasets, which also proves the effectiveness.
Zheng Li 0027, Yongcheng Wang 0001, Yuxi Zhang 0003, Yunxiao Gao, Zhikang Zhao, Hao Feng 0008
IEEE Geosci. Remote. Sens. Lett.5
2024 Remote Sensing Hyperspectral Image Super-Resolution via Multidomain Spatial Information and Multiscale Spectral Information Fusion
abstract
Hyperspectral image super-resolution technology has made remarkable progress due to the development of deep learning. However, the technique still faces two challenges, i.e., the imbalance between spectral and spatial information extraction, and the parameter deviation and high computational effort associated with 3D convolution. In this article, we propose a super-resolution method for remote sensing hyperspectral images based on multi-domain spatial information and multi-scale spectral information fusion (MSSR). Specifically, inspired by the high degree of self-similarity of remote sensing hyperspectral images, a spatial-spectral attention module based on dilated convolution (DSSA) for capturing global spatial information is proposed. The extraction of local spatial information is then accomplished by residual blocks using small-size convolution kernels. Meanwhile, we propose the 3D Inception module to efficiently mine multi-scale spectral information. The module only retains the scale of the 3D convolution kernel in spectral dimension, which greatly reduces the high computational cost caused by 3D convolution. Comparative experimental results on four benchmark datasets demonstrate that compared with the current cutting-edge models, our method achieves state-of-the-art results and the model computation is greatly reduced.
Chi Chen 0003, Yongcheng Wang 0001, Yuxi Zhang 0003, Zhikang Zhao, Hao Feng 0008
IEEE Trans. Geosci. Remote. Sens.4
2024 Hyperspectral Image Classification Framework Based on Multichannel Graph Convolutional Networks and Class-Guided Attention Mechanism
abstract
Graph convolutional networks (GCNs) can extract features of samples in non-Euclidean space, which can be used for hyperspectral image (HSI) classification in collaboration with convolutional neural networks (CNNs). The features of GCNs and CNNs are incompatible to a certain extent, and traditional graph convolution methods use a single channel two-dimensional matrix to extract features. As a result, it is difficult to explore the relationships fully and flexibly between samples. To further exploit the potential of these two networks for collaborative extraction of HSI features, we propose a fusion framework based on multichannel GCNs and class-guided attention mechanism (MG2A). Specifically, a multichannel graph convolutional network (MGCN) module is designed for batchwise network training, where the adjacent matrix of each channel contains different information between samples. In addition, we develop a class-guided attention mechanism to adaptively fuse the features of multiple MGCN modules and learn the transformation process of the features. Finally, the features of CNNs and MGCN modules are fused at multiple layers through a fusion framework. Experimental results on four benchmark HSI datasets show that MG2A achieves better classification performance compared to other state-of-the-art methods.
Hao Feng 0008, Yongcheng Wang 0001, Chi Chen 0003, Dongdong Xu 0001, Zhikang Zhao
IEEE Trans. Geosci. Remote. Sens.5
2022 A Multi-Degradation Aided Method for Unsupervised Remote Sensing Image Super Resolution With Convolution Neural Networks
abstract
In remote sensing, it is desirable to improve image resolution by using the image super-resolution (SR) technique. However, there are two challenges: the first one is that high-resolution (HR) images are insufficient or unavailable; another one is that the single degradation model such as bicubic (BIC) cannot super-resolve favorable images in the real world. To address the above two problems, this article presents a multi-degradation, unsupervised SR method based on deep learning. This framework consists of a degrader$ {D}$to fit the image degradation model and a generator$ {G}$to generate SR image. By introducing$ {D}$, calculating the loss function between SR image and HR image as supervised SR methods did can be converted into calculating loss between low resolution (LR) image and image degraded by SR image, thereby realizing unsupervised learning. Experiments on several degradation models show that our method renders the state-of-the-art results compared with existing unsupervised SR methods, and achieves competitive results in contrast with supervised SR methods. Moreover, for real remote sensing images obtained by the Jilin-1 satellite, our method obtained more plausible results visually, which demonstrate the potential in real-world applications.
Ning Zhang 0025, Yongcheng Wang 0001, Xin Zhang 0062, Dongdong Xu 0001, Xiaodong Wang 0019, Guangli Ben, Zhikang Zhao, Zheng Li 0027
IEEE Trans. Geosci. Remote. Sens.7