VLDB 2026 Research / reviewers in the wild / expert
Zhizhuo Jiang
dblp:297/1842
· DBLP profile ↗
16ranked-venue papers
1as first author
16since 2021 · last 2025
0000-0002-5269-2753ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Novel Split Deep Unfolding Transformer for Pan-SharpeningabstractPan-sharpening is a commonly employed strategy to obtain high-resolution multispectral (HRMS) images. Existing deep unfolding networks for pan-sharpening suffer from ineffectively establishing the relationship between panchromatic (PAN) images and generated noisy HRMS (GN-HRMS) images in PAN-guided image denoising, lacking the support of physical models. In this paper, we first design a degradation-fusion-aware unfolding framework (DF-UF) to separate the processing of PAN-prior in PAN-guided image denoising into an individual module, PAN-prior processor, for better integrating physical models. Then, we derive a flexible intensity-hue-saturation (F-IHS) to act as the PAN-prior processor, which models the relationship between PAN images and GN-HRMS images in terms of intensity components through the intensity-hue-saturation (IHS) theory. Finally, plugging F-IHS into DF-UF, we propose a degradation-intensity-aware unfolding transformer (DIUT) to address the problem of incomplete utilization of PAN images in the denoising process. Extensive experiments on diverse scenes show that the performance of DIUT surpasses existing state-of-the-art methods. Zhizhuo Jiang, Xueqian Wang 0002, Yaowen Li, Huajie Wang, Yu Liu 0005 |
ICASSP | 2 |
| 2025 | Spatial-Spectral Consistency: A Semi-Supervised Approach for Multispectral Scene ClassificationabstractMultispectral remote sensing images, with their richer spectral information, can achieve better scene classification performance compared to RGB images. However, high annotation costs remain a significant challenge. To reduce these costs, we propose a spatial-spectral consistency (SSC) semi-supervised learning method that fully leverages abundant unlabeled data and effectively exploits spectral information from multispectral images. Our method employs two branches to extract spatial and spectral features, respectively. The predictions from the two branches for the weakly augmented input are first fused to generate pseudo-labels, which are then used to supervise the branches in predicting the strongly augmented input. Additionally, we introduce a spectral attention module into the network to enhance its ability to extract spectral information. We conduct extensive experiments on the EuroSAT and SEN12MS datasets, demonstrating that our method outperforms other semi-supervised approaches, achieving state-of-the-art (SOTA) performance. Jin Li 0069, Huajie Wang, Zhizhuo Jiang, Yu Liu 0005 |
ICIP | 3 |
| 2025 | GCBF: Grouped Cross-Band Fusion Network for Multispectral Scene ClassificationabstractRemote sensing scene classification is a crucial task for remote sensing image interpretation. Existing multispectral scene classification methods have overlooked the interrelationships between different spectral bands, which limits the mining of complementary information within the images. Addressing this issue, we propose a grouped cross-band fusion (GCBF) network for remote sensing multispectral scene classification to take full advantage of complementary information between various spectral bands. Firstly, we separate the various bands of the given multispectral image into different groups to better capture the characteristics of each spectral band. Then, we use the existing UniFormer as a feature extractor to learn the representations of red, green, and blue (RGB) bands. For the spectral bands other than RGB, we propose a new network called multi-stage grouped spectral feature extraction (MGSFE) network to learn discriminative representations. We also draw inspiration from the band combination in the field of remote sensing and introduce a cross-band attention fusion (CBAF) module designed to adaptively merge features from both the RGB bands and other spectral bands. Extensive experiments on three widely used remote sensing multispectral scene classification datasets of BigEarthNet, SEN12MS, and EuroSAT demonstrate the superiority of our proposed method compared with several state-of-the-art (SOTA) methods. Jin Li 0069, Yu Liu 0005, Wenda Zhao 0003, Zhizhuo Jiang, Xueqian Wang 0002, Bolun Zheng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | CADDN: A Content-Aware Downsampling-Based Detection Method for Small Objects in Remote Sensing ImagesabstractA key issue of existing deep-learning-based object detection methods in remote sensing images is that they often struggle to differentiate the background and small object regions due to multi-level downsampling operations therein. Downsampling operations help extract high-level semantic features but result in excessive loss of spatial features of small objects. In this paper, we propose a new small object detector using multispectral remote sensing images, named content-aware downsampling-based detection network (CADDN), where we newly design a content-aware downsampling-based module (CADM). Unlike conventional downsampling operations that apply uniform downsampling parameters across the entire feature map, CADM adaptively assigns higher weights to feature elements that are critical for distinguishing objects from the background, and this assignment is guided by the contextual awareness of object locations during the downsampling process. Experiments based on multispectral remote sensing images with small ships and vehicles demonstrate that CADM can accurately identify and preserve the locations of important object-related features, and CADDN correspondingly achieves superior small object detection performance than state-of-the-art methods. Linping Zhang, Yu Liu 0005, Xueqian Wang 0002, You He 0002, Gang Li 0008, Chang Liu 0053, Zhizhuo Jiang, Yang Liu 0119 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | A Benchmark and Frequency Compression Method for Infrared Few-Shot Object DetectionabstractInfrared few-shot object detection (IFSOD) aims to detect infrared objects with limited labeled examples. Current infrared datasets, however, suffer from limited diversity in object types and classes, hindering robust evaluation of model generalization on novel classes. To systematically assess dataset quality, we propose metrics for class diversity, instance variability, and object density. By integrating three widely used infrared datasets, we construct the first dataset specifically tailored for IFSOD, increasing instance density to 4.8 (a 1.1 improvement) and expanding the number of classes to 18 (a 5-class increase) compared to the source datasets. Furthermore, frequency analysis of spatial features reveals that sparse annotations introduce spectral bias in the frequency domain. Directly transforming spatial features to the frequency domain, however, mixes background noise with object features, causing spectral leakage and impairing the learning of discriminative features for novel classes. To address these issues, we propose the frequency compression few-shot detection (FC-fsd) method, which incorporates a frequency compression (FC) module. The FC module leverages Discrete Cosine Transform (DCT) within localized windows to reduce spectral leakage and enhance feature clarity. With minimal additional computational overhead, FC-fsd significantly outperforms state-of-the-art methods, achieving nAP50 scores of 28.57 (+13.37) and 35.63 (+2.59) in 1-shot and 2-shot settings, respectively. Our dataset is published athttps://github.com/RuihengZhang/IFSOD-dataset. Ruiheng Zhang 0001, Biwen Yang, Lixin Xu 0001, Yan Huang 0023, Qi Zhang 0070, Zhizhuo Jiang, Yu Liu 0005 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | A Cross-modal Fusion Method for Multispectral Small Ship DetectionabstractThe fusion module of RGB and infrared (IR) remote sensing images is the key of multispectral ship detection. Existing works have shown that the cross-attention-based feature fusion can achieve good performance by extracting the complementary information of RGB and IR modalities. However, the existing commonly used cross-attention mechanisms introduce lots of redundancy parameters and mainly focus on global feature interaction of multispectral images, ignoring local detail information that is also important for small ship detection. In this paper, we propose a novel multispectral ship detection approach named LoGFusion. In LoGFusion, we design the cross stage partial module with partial convolution (CSPMPC) to reduce feature redundancy and utilize the local cross-modal fusion module (LoCFM) and global cross-modal fusion module (GCFM) to capture both local and global cross-modal features. Furthermore, we introduce a Multispectral Small Ship Dataset (MSSD) containing over 5k ship targets for small target detection. Experiments on MSSD validate the effectiveness of our method in terms of small ship detection in multispectral images. Yang Liu 0119, Yu Liu 0005, Xueqian Wang 0002, Linping Zhang, Zhizhuo Jiang, Yaowen Li, Chenggang Yan 0001, Ying Fu 0001, Tao Zhang 0042 |
FUSION | 5 |
| 2024 | Language-Assisted Siamese Contrastive Framework for Fine-Grained Remote Sensing Ship Image RetrievalabstractAs the number of remote sensing (RS) images increases, it is crucial to retrieval ship targets according to specific demands. The existing ship image retrieval methods only extract features from the image modality, which may not fully utilize the rich text information available and ignore the high-level hierarchical relations between ship classes. In this paper, we propose a language-assisted siamese contrastive framework, namely LASCF, for fine-grained ship retrieval in RS images. In the new LASCF, the siamese vision models are employed to measure the similarity between images. Moreover, a label text encoder with a pretrained language model is designed to extract the high-level semantic information from labels, and thus the information of the hierarchical relations between ship classes are fused in LASCF. Finally, the multimodal similarity measurement module based on contrastive learning is proposed to optimize the siamese vision models. The experimental results show that the proposed LASCF outperforms several existing state-of-the-art methods. Zhizhuo Jiang, Yu Liu 0005, Yaowen Li, Xueqian Wang 0002, Chenggang Yan 0001 |
IGARSS | 2 |
| 2024 | CPDTD: Content-Perception Downsampling-Based Small Target Detector in Remote Sensing ImagesabstractExisting deep neural network (DNN)-based target detectors in remote sensing images (RSIs) often face challenges in distinguishing small targets from the background. This is mainly because the downsampling process in DNN-based target detectors results in excessive loss of small-target-related features. This paper proposes a new small target detector in RSIs named content-perception downsampling-based target detector (CPDTD), where a novel content-perception downsampling module (CPDM) is designed to replace standard downsampling methods (e.g. pooling and convolution with stride greater than 1). CPDM encodes the input feature map and predicts the location of important features that distinguish targets from backgrounds, assigning larger weights to critical features according to the perception of the position of targets in the content during the downsampling process. Experiments on measured multispectral RSIs regarding small ship and vehicle targets demonstrate the superiorities of our proposed CPDTD in comparison with existing methods. Linping Zhang, Yu Liu 0005, Xueqian Wang 0002, Lihui Xue, Gang Li 0008, Yang Liu 0119, Zhizhuo Jiang |
IGARSS | 7 |
| 2024 | Body Joint Boundary Prototype Match for Few-Shot Remote Sensing Semantic SegmentationabstractDeep networks require a large number of samples for optimization, so few-shot segmentation in remote sensing scenes is still an open problem. However, this challenge is exacerbated by the feature blurring and aliasing of bodies (low frequency) and boundaries (high frequency). The existing methods usually only focus on the body part of the class, that is, the low-frequency part, and ignore the critical role of boundary information, that is, high-frequency details, on feature representation. In this letter, we propose a novel body joint boundary prototype match (B2PM) approach that aims to enable prior learning of low- and high-frequency information by explicitly modeling the body and boundary features of objects. First, body-aware prototype learning (BodyPL) realizes the adaptive modeling of the body part of the object through a precise farthest point sampling (FPS) initialization algorithm and an adaptive part shift (APS) strategy, which alleviates the feature ambiguity of the body. Second, boundary-aware prototype learning (BoundPL) explicitly models boundary prototypes by building a patch division and assignment strategy to alleviate feature aliasing at boundaries. Finally, prototype match performs prior knowledge aggregation by computing the affinity between query features and support prototypes. Extensive experiments on commonly used benchmarks (iSAID and PASCAL VOC) demonstrate that B2PM improves the state of the art by significant margins. Yongqiang Mao, Zhizhuo Jiang, Yu Liu 0005, Yaowen Li, Chenggang Yan 0001, Bolun Zheng |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | A Novel Method for Maneuvering Extended Vehicle Tracking with Automotive RadarabstractIn high-resolution automotive radar tracking systems, vehicle targets are often regarded as extended targets, which means multiple measurements originated from scattering centers of vehicle targets can be detected at each scan and thus the traditional point target tracking schemes are unsuitable. Meanwhile, vehicle maneuvers, e.g., braking and swerving, cause serious degradation of the classical extended target tracking methods. In this paper, a novel method is proposed for maneuvering extended vehicle tracking with automotive radar. The data-region association (DRA) strategy is adopted to handle the vehicle extension effect, which is superior in describing the complex spatial distribution of vehicle target measurements. The interacting multiple model (IMM) method is combined with this DRA strategy to describe the evolution of target motion models. Accordingly, the proposed DRA-IMM method achieves satisfying tracking performance of extended vehicles and also guarantees the robustness in case of maneuvers. Furthermore, in view of the correlation between vehicle extension and its kinematic state, a ray-based strategy is devised to improve the prior distribution of the data-region association of the basic DRA-IMM, and accordingly an enhanced DRA-IMM (EDRA-IMM) method is proposed. Simulation result validates the effectiveness of the proposed DRA-IMM method for maneuvering extended vehicle tracking and the further improvement of the proposed EDRA-IMM method. Hongfei Xu, Yaowen Li, Yuxin Ke, Zhizhuo Jiang, Yu Liu 0005 |
FUSION | 4 |
| 2023 | Video Diffusion Models with Local-Global Context GuidanceabstractDiffusion models have emerged as a powerful paradigm in video synthesis tasks including prediction, generation, and interpolation. Due to the limitation of the computational budget, existing methods usually implement conditional diffusion models with an autoregressive inference pipeline, in which the future fragment is predicted based on the distribution of adjacent past frames. However, only the conditions from a few previous frames can't capture the global temporal coherence, leading to inconsistent or even outrageous results in long-term video prediction. In this paper, we propose a Local-Global Context guided Video Diffusion model (LGC-VD) to capture multi-perception conditions for producing high-quality videos in both conditional/unconditional settings. In LGC-VD, the UNet is implemented with stacked residual blocks with self-attention units, avoiding the undesirable computational cost in 3D Conv. We construct a local-global context guidance strategy to capture the multi-perceptual embedding of the past fragment to boost the consistency of future prediction. Furthermore, we propose a two-stage training strategy to alleviate the effect of noisy frames for more stable predictions. Our experiments demonstrate that the proposed method achieves favorable performance on video prediction, interpolation, and unconditional video generation. We release code at https://github.com/exisas/LGC-VD. Lu Zhang 0053, Yu Liu 0005, Zhizhuo Jiang, You He 0002 |
IJCAI | 4 |
| 2023 | Multi-agent Perception via Co-attentive Communication Mechanism
Ning Gong, Yuxin Ke, Zhizhuo Jiang, Yaowen Li |
PRCV (6) | 5 |
| 2023 | Caps-SSENet: An Improved Estimation Method for SAR Ship SizeabstractAccurate estimation of the sizes of ship targets plays a critical role in the task of ship classification in synthetic aperture radar (SAR) images. Existing deep neural networks (DNNs)-based methods for SAR ship size estimation (SSE) often adopt a fully connected structure that has limited capability in accurately modeling the relationships of features extracted from SAR images, leading to degraded performance of size estimation. It has been demonstrated that capsule networks provide new guidelines to capture relationships of image features by replacing traditional neurons with capsules, where the dynamic routing strategy is used to calculate correlations among capsules. In this letter, we propose an improved method for SAR SSE based on the capsule network named Caps-SSE network (SSENet). In our Caps-SSENet, a capsule-neural-mixing size mapping module is designed to transform the extracted image features into capsules and complete the estimation of ship sizes using informative feature correlations from dynamic routing. In addition, an average scaled mean square error (ASMSE) loss is proposed to improve the size estimation performance of small ships. Experimental results based on measured SAR data show that the proposed method reduces the estimation error of ship sizes in SAR images in comparison with the existing state-of-the-art method. Yu Liu 0005, Xueqian Wang 0002, Zhizhuo Jiang, Gang Li 0008, Bolun Zheng, Jiyong Zhang 0001, You He 0003 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Prospects for multi-agent collaboration and gaming: challenge, technology, and applicationabstractIn this study, we presented the prospects for multi-agent system research with a special focus on agent collaboration and gaming tasks. We briefly introduced some open issues and task challenges from three major perspectives: the multi-agent environment, collaboration, and gaming. Then we provided a related outlook for the technology directions that may create some research challenge insights. Finally, we discussed the outlook for the multi-agent collaboration and gaming application areas. Yu Liu 0005, Zhi Li 0057, Zhizhuo Jiang, You He 0002 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2022 | A Novel Loss Function for Optical and SAR Image Matching: Balanced Positive and Negative SamplesabstractImage matching is a primary technology for optical and synthetic aperture radar (SAR) image fusion but often shows limited performance due to the highly nonlinear differences between optical and SAR modalities. Recently, deep neural networks (DNNs) have been investigated to effectively extract nonlinear features for image matching tasks, where DNNs are trained based on the elaborated design of loss functions and a low loss value is often expected to obtain better image matching performance. In this letter, we first theoretically demonstrate that when the value of a state-of-the-art loss function decreases, the corresponding matching performance may not consistently improve due to the imbalanced effect of positive and negative samples. To tackle this issue, we proposed an improved loss function to train DNNs for image matching of SAR and optical images. We theoretically prove that the improved loss function ensures the improvement of the matching performance when the loss value decreases based on Taylor’s series expansion analysis. Experimental results on an open dataset with extensive optical and SAR image pairs show that 1) the proposed loss function is better than the original one in terms of image matching performance and 2) the combination of our loss function and existing multiscale convolutional gradient feature (MCGF)-based network provides better matching performance than other state-of-art approaches. Yueping He, Xueqian Wang 0002, Yu Liu 0005, Zhizhuo Jiang, Gang Li 0008, You He 0003 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Robust STAP Detection Based on Volume Cross-Correlation Function in Heterogeneous EnvironmentsabstractThe performance of moving target detection in heterogeneous environments with the traditional space-time adaptive processing (STAP) may degrade when the real clutter environments deviate from the prior assumption on the clutter distribution. In this letter, a new detector for STAP applications based on volume cross-correlation function (VCF), namely VCF-STAP, is proposed to achieve robust performance of moving target detection in heterogeneous environments. In the new VCF-STAP, the VCF is used to form a distance measure between the sample signal subspace and the target subspace without modeling the clutter distribution. Then, a new robust STAP detection statistic is constructed using this distance measure. Simulation and experimental results show that the proposed VCF-STAP achieves robust performance of moving target detection in heterogeneous environments, especially it achieves much superior detection performance compared with existing STAP methods when the real clutter environments do not satisfy their prior assumptions. Besides, it is also shown that VCF-STAP has the constant false alarm rate (CFAR) property. Zhizhuo Jiang, You He 0002, Gang Li 0008, Xiao-Ping Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |