VLDB 2026 Research / reviewers in the wild / expert
Zhiyu Jiang
dblp:166/0578
· DBLP profile ↗
23ranked-venue papers
7as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spectral consistency learning for cross-domain hyperspectral image classification
Zhiyu Jiang, Dandan Ma, Yuan Yuan 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Learning from the waist: Aspect-ratio-aware positive sampling for oriented ship detection
Dandan Ma, Zhiyu Jiang |
Expert Syst. Appl. | 3 |
| 2026 | Deep color constancy via a color shift aware conditional diffusion model
Haonan Su, Haiyan Jin, Yuanlin Zhang 0003, Bin Wang 0046, Zhiyu Jiang |
J. Vis. Commun. Image Represent. | 6 |
| 2026 | Noise perturbation augmentation based dual-branch alignment network for cross-domain hyperspectral image classification
Zhiyu Jiang, Dandan Ma |
Pattern Recognit. | 1 |
| 2025 | Structural consistency learning for unsupervised domain adaptive object detection
Zhiyu Jiang |
Neural Networks | 1 |
| 2025 | Cross-domain hyperspectral image classification
Zhiyu Jiang, Zhuozhao Liu, Dandan Ma |
Pattern Recognit. | 1 |
| 2025 | Variation Autoencoder of Spatial-Spectral Joint Mask for Hyperspectral Anomaly DetectionabstractIn recent years, autoencoders and their variants have emerged as effective tools for hyperspectral anomaly detection. Nevertheless, owing to the complex distribution of anomalous regions and the similarity in spatial-spectral features, these models often reconstruct anomalies and backgrounds simultaneously, hindering their ability to distinguish between them and reducing detection accuracy. To address this issue, we propose a novel hyperspectral anomaly detection method based on a spatial-spectral joint mask variational autoencoder (VAE). By combining the probabilistic modeling capabilities of VAEs with a masking-based attention mechanism, our method enables more precise extraction of essential background information in localized regions. Specifically, the spatial-spectral joint masking technique is proposed to guide the network to concentrate on background features across multiple dimensions, tackling issues of spatial structure approximation and spectral redundancy. To further enhance robustness in noisy and complex environments, we iteratively refine the reconstructed residual image through recursive filtering. Extensive comparative experiments and ablation studies on multiple public datasets demonstrate that our approach consistently outperforms existing methods in detection accuracy. Dandan Ma, Zhuozhao Liu, Zhiyu Jiang |
IEEE Signal Process. Lett. | 3 |
| 2025 | Implicit CLIP Prior Decoupling for Few-Shot Remote Sensing Image SegmentationabstractFew-Shot Segmentation (FSS) in remote sensing aims to achieve segmentation of novel categories in query images using limited annotated support images. Despite extensive research, the significant intra-class differences of remote sensing targets continue to hinder progress in this field. Pre-trained vision-language models (VLMs) possess strong generalization capabilities, and their cross-modal information can effectively mitigate intra-class variance issues. However, VLMs rarely focus on dense prediction tasks, and the complexity of remote sensing imagery limits the effectiveness of existing attempts on FSS tasks. To address this issue, this article proposes an Implicit CLIP Prior Decoupling Network (ICPD-Net), which mines effective cross-modal priors from VLMs and leverages ranking information to improve visual metric strategies. Specifically, the Implicit Prior Decoupling Module (IPDM) utilizes ambiguous foreground-background vision-language similarities to construct class-agnostic prompts, while employing a prior learner to mine implicit vision-language priors that alleviate intra-class differences. To fully leverage cross-modal information, the Reliable Feature Fusion Module (RFFM) utilizes vision-language priors to obtain high-confidence query features for fusion with support features, and further mitigating intra-class differences through self-support paradigm. Finally, the Dual Visual Priors Module (DVPM) introduces a novel rank information prior for visual feature measurement. This approach constructs an effective metric learning method by combining the ranking relationships of Euclidean distances between support-query features with the Normalized Discounted Cumulative Gain (NDCG) algorithm, while comprehensively exploring visual metric relationships through traditional cosine similarity prior. Extensive experiments on iSAID-5iand DLRSD-5idemonstrate that our method achieves significant improvements. Particularly under the 1-shot setting, our approach shows exceptional effectiveness, outperforming state-of-the-art methods by up to 11.48% on the iSAID-5i dataset. The code of ICPD-Net is available at https://github.com/yeh15/ICPD-Net. Zhiyu Jiang, Ye Yuan 0001, Dandan Ma, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | MiKA-HAD: Minority-Augmented KAN-Based Hybrid Self-Supervised Framework for Hyperspectral Anomaly DetectionabstractIn recent years, autoencoder-based models have demonstrated significant potential in hyperspectral anomaly detection. However, the inherent class imbalance between anomaly and background samples, coupled with substantial intra-class heterogeneity within background components, often leads traditional models to excessively capture majority-class background features, thereby weakening their discriminative capabilities for minority-class backgrounds and anomaly targets. Concurrently, conventional methods fail to sufficiently model nonlinear features and adapt to complex environments, undermining their detection performance and robustness in intricate environmental interference. To address these challenges, we propose MiKA-HAD, a Minority-augmented KAN-based Hybrid self-supervised framework for hyperspectral Anomaly Detection. It incorporates a density-based clustering-guided sample balancing strategy that dynamically synthesizes minority background samples with spectral-spatial diversity through self-supervised learning. This approach achieves feature space rebalancing while effectively suppressing noise interference on anomaly boundaries. To overcome high-dimensional data complexity and nonlinear feature extraction challenges, we construct a Hybrid Self-Supervised network with a KAN-based multi-path collaborative module that orchestrates synergy between local sensitivity preservation, nonlinear feature modeling, and global consistency maintenance. This tripartite architecture establishes dynamic equilibrium that enhances representation capability and environmental robustness. Extensive experiments show MiKA-HAD achieves superior detection accuracy and stability compared to existing approaches, particularly in complex environments with varying noise conditions. The framework establishes a new paradigm for robust hyperspectral anomaly detection by addressing both class imbalance bias and nonlinear feature modeling limitations. Dandan Ma, Zhuozhao Liu, Zhiyu Jiang, Yi Zheng 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Multimodal Difference Augmentation Learning for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD) plays a crucial role in applications such as environmental monitoring and urban planning. With the emergence of foundational vision-language models like CLIP, there is growing interest in integrating textual information into vision tasks. However, in the RSCD domain, limited efforts have been made to effectively leverage textual cues, and challenges persist in capturing differential features. To address these issues, this study proposes the Multimodal Difference Augmentation learning for remote sensing change detection model (MdaCD) that fully exploits textual information and enhances differential feature learning. MdaCD introduces a CLIP-Guided Masking process to direct textual descriptions toward image differences, and a Multimodal Fusion and Difference Augmentation process to integrate and refine differential features across modalities. The CLIP-Guided Masking process applies masking to bi-temporal image pairs before generating text prompts, enabling a more targeted analysis of changes. Meanwhile, the Multimodal Fusion and Difference Augmentation process computes a fused attention map to integrate visual and textual cues, effectively amplifying relevant differences. By applying Difference Augment functions, the differential features from both visual and textual embeddings are further refined and strengthened. The effectiveness of the proposed MdaCD model is validated through extensive experiments on two public RSCD datasets, where it achieves state-of-the-art performance with IoU scores of 84.88% on LEVIR-CD and 71.96% on SYSU-CD. The code and pretrained models of this work will be publicly available at https://github.com/haoyangofficial/MdaCD. Zhiyu Jiang, Dandan Ma, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Prototypical Metric Segment Anything Model for Data-Free Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation (FSS) is crucial for image interpretation, yet it is constrained by requirements for extensive base data and a narrow focus on foreground-background differentiation. This work introduces Data-free Few-shot Semantic Segmentation (DFSS), a task that requires limited labeled images and forgoes the need for extensive base data, allowing for comprehensive image segmentation. The proposed method utilizes the Segment Anything Model (SAM) for its generalization capabilities. The Prototypical Metric Segment Anything Model is introduced, featuring an initial segmentation phase followed by prototype matching, effectively addressing the learning challenges posed by limited data. To enhance discrimination in multi-class segmentation, the Supervised Prototypical Contrastive Loss (SPCL) is designed to refine prototype features, ensuring intra-class cohesion and inter-class separation. To further accommodate intra-class variability, the Adaptive Prototype Update (APU) strategy dynamically refines prototypes, adapting the model to class heterogeneity. The method's effectiveness is demonstrated through superior performance over existing techniques on the DFSS task, marking a significant advancement in UAV image segmentation. Zhiyu Jiang |
IEEE Signal Process. Lett. | 1 |
| 2022 | Prototype Queue Learning for Multi-Class Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation aims to undertake the segmentation task of novel classes with only a few annotated images. However, most existing methods tend to segment the foreground and background in the image, which limits practical application. In this paper, we present a Prototype Queue Network, which performs few-shot segmentation on multiclass in the images by aggregating binary classes into multiple classes. A prototype queue learning module is proposed to achieve multi-class segmentation by mining the relationship among features of different classes with queue and pseudo labels. In addition, a background latent class distribution refinement module is proposed to prevent the latent novel class in the background from being incorrectly predicted, which refines the boundary among different classes. Furthermore, we propose a two-steps segmentation module to optimize the process of extracting feature representation by adding progressive constraints, which can further improve the accuracy of segmentation. Experiments on the UDD and Vaihingen datasets demonstrate that our method achieves state-of-the-art performance. Zichao Wang 0005, Zhiyu Jiang, Yuan Yuan 0001 |
ICIP | 2 |
| 2022 | Adaptive open domain recognition by coarse-to-fine prototype-based network
Yuan Yuan 0001, Xinxing He, Zhiyu Jiang |
Pattern Recognit. | 3 |
| 2022 | Proxy-Based Deep Learning Framework for Spectral-Spatial Hyperspectral Image Classification: Efficient and RobustabstractDeep convolutional networks have been extensively deployed in hyperspectral image (HSI) classification. Reaching for high accuracy, the existing deep-learning-based methods commonly deepen or widen their networks for better performance, which brings higher computational complexity and the risk of overfitting. Although the introduction of the residual module and batch-normalization reduces the generalization degradation in complex networks, the mainstream methods still suffer from low robustness to the noise. To tackle these issues, a compact proxy-based deep learning framework is proposed to perform highly accurate HSI classification with superb efficiency and robustness. In this article: 1) novel deep proxies are integrated to replace the dense classifier layers in conventional networks, which represents specific classes in deep embedding space and enables fast and reliable convergence; 2) the proxy-based feature embedding is studied in distance metric and similarity metric, and compatible dual-metric loss functions are designed for further optimized embedding distribution, which leads to more robust generalization; and 3) state-of-the-art performance and robustness are demonstrated by the proposed framework on mainstream HSI data sets with the minimal network scale and time complexity. Yuan Yuan 0001, Chengze Wang, Zhiyu Jiang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Task-Related Self-Supervised Learning For Remote Sensing Image Change DetectionabstractChange detection for remote sensing images is widely applied for urban change detection, disaster assessment and other fields. However, most of the existing CNN-based change detection methods still suffer from the problem of inadequate pseudo-changes suppression and insufficient feature representation. In this work, an unsupervised change detection method based on Task-related Self-supervised Learning Change Detection network with smooth mechanism(TSLCD) is proposed to eliminate it. The main contributions include: (1) the task-related self-supervised learning module is introduced to extract spatial features more effectively. (2) a hard-sample-mining loss function is applied to pay more attention to the hard-to-classify samples. (3) a smooth mechanism is utilized to remove some of pseudo-changes and noise. Experiments on four remote sensing change detection datasets reveal that the proposed TSLCD method achieves the state-of-the-art for change detection task. Zhinan Cai, Zhiyu Jiang, Yuan Yuan 0001 |
ICASSP | 2 |
| 2021 | Deep Feature Selection-And-Fusion for RGB-D Semantic SegmentationabstractScene depth information can help visual information for more accurate semantic segmentation. However, how to effectively integrate multi-modality information into representative features is still an open problem. Most of the existing work uses DCNNs to implicitly fuse multi-modality information. But as the network deepens, some critical distinguishing features may be lost, which reduces the segmentation performance. This work proposes a unified and efficient feature selection-and-fusion network (FSFNet), which contains a symmetric cross-modality residual fusion module used for explicit fusion of multi-modality information. Besides, the network includes a detailed feature propagation module, which is used to maintain low-level detailed information during the forward process of the network. Compared with the state-of-the-art methods, experimental evaluations demonstrate that the proposed model achieves competitive performance on two public datasets. Yuejiao Su, Yuan Yuan 0001, Zhiyu Jiang |
ICME | 3 |
| 2021 | Self-Supervised Spectral Matching Network for Hyperspectral Target DetectionabstractHyperspectral target detection is a pixel-level recognition problem. Given a few target samples, it aims to identify the specific target pixels such as airplane, vehicle, ship, from the entire hyperspectral image. In general, the background pixels take the majority of the image and complexly distributed. As a result, the datasets are weak annotated and extremely imbalanced. To address these problems, a spectral mixing based self-supervised paradigm is designed for hyperspectral data to obtain an effective feature representation. The model adopts a spectral similarity based matching network framework. In order to learn more discriminative features, a pair-based loss is adopted to minimize the distance between target pixels while maximizing the distances between target and background. Furthermore, through a background separated step, the complex unlabeled spectra are downsampled into different sub-categories. The experimental results on three real hyperspectral datasets demonstrate that the proposed framework achieves better results compared with the existing detectors. Can Yao, Yuan Yuan 0001, Zhiyu Jiang |
IGARSS | 3 |
| 2020 | Open Set Domain Recognition via Attention-Based GCN and Semantic Matching OptimizationabstractOpen set domain recognition has got the attention in recent years. The task aims to specifically classify each sample in the practical unlabeled target domain, which consists of all known classes in the manually labeled source domain and target-specific unknown categories. The absence of annotated training data or auxiliary attribute information for unknown categories makes this task especially difficult. Moreover, exiting domain discrepancy in label space and data distribution further distracts the knowledge transferred from known classes to unknown classes. To address these issues, this work presents an end-to-end model based on attention-based GCN and semantic matching optimization, which first employs the attention mechanism to enable the central node to learn more discriminating representations from its neighbors in the knowledge graph. Moreover, a coarse-to-fine semantic matching optimization approach is proposed to progressively bridge the domain gap. Experimental results validate that the proposed model not only has superiority on recognizing the images of known and unknown classes, but also can adapt to various openness of the target domain. Xinxing He, Yuan Yuan 0001, Zhiyu Jiang |
ICPR | 3 |
| 2020 | Instance-Aware Remote Sensing Image Captioning with Cross-Hierarchy AttentionabstractThe spatial attention is a straightforward approach to enhance the performance for remote sensing image captioning. However, conventional spatial attention approaches consider only the attention distribution on one fixed coarse grid, resulting in the semantics of tiny objects can be easily ignored or disturbed during the visual feature extraction. Worse still, the fixed semantic level of conventional spatial attention limits the image understanding in different levels and perspectives, which is critical for tackling the huge diversity in remote sensing images. To address these issues, we propose a remote sensing image caption generator with instance-awareness and cross-hierarchy attention. 1) The instances awareness is achieved by introducing a multi-level feature architecture that contains the visual information of multi-level instance-possible regions and their surroundings. 2) Moreover, based on this multi-level feature extraction, a cross-hierarchy attention mechanism is proposed to prompt the decoder to dynamically focus on different semantic hierarchies and instances at each time step. The experimental results on public datasets demonstrate the superiority of proposed approach over existing methods. Chengze Wang, Zhiyu Jiang, Yuan Yuan 0001 |
IGARSS | 2 |
| 2020 | Weighted Hierarchical Sparse Representation for Hyperspectral Target DetectionabstractHyperspectral target detection has been widely studied in the field of remote sensing. However, background dictionary building issue and the correlation analysis of target and background dictionary issue have not been well studied. To tackle these issues, a Weighted Hierarchical Sparse Representation for hyperspectral target detection is proposed. The main contributions of this work are listed as follows. 1) Considering the insufficient representation of the traditional background dictionary building by dual concentric window structure, a hierarchical background dictionary is built considering the local and global spectral information simultaneously. 2) To reduce the impureness impact of background dictionary, target scores from target dictionary and background dictionary are weighted considered according to the dictionary quality. Three hyperspectral target detection data sets are utilized to verify the effectiveness of the proposed method. And the experimental results show a better performance when compared with the state-of-the-arts. Chenlu Wei, Zhiyu Jiang, Yuan Yuan 0001 |
IGARSS | 2 |
| 2018 | Contour-aware network for semantic segmentation via adaptive depth
Zhiyu Jiang, Yuan Yuan 0001, Qi Wang 0009 |
Neurocomputing | 1 |
| 2017 | HDPA: Hierarchical deep probability analysis for scene parsingabstractScene parsing is an important task in computer vision and many issues still need to be solved. One problem is about the non-unified framework for predicting things and stuff and the other one refers to the inadequate description of contextual information. In this paper, we address these issues by proposing a Hierarchical Deep Probability Analysis(HDPA) method which particularly exploits the power of probabilistic graphical model and deep convolutional neural network on pixel-level scene parsing. To be specific, an input image is initially segmented and represented through a CNN framework under Gaussian pyramid. Then the graphical models are built under each scale and the labels are ultimately predicted by structural analysis. Three contributions are claimed: unified framework for scene labeling, hierarchical probabilistic graphical modeling and adequate contextual information consideration. Experiments on three benchmarks show that the proposed method outperforms the state-of-the-arts in scene parsing. Yuan Yuan 0001, Zhiyu Jiang, Qi Wang 0009 |
ICME | 2 |
| 2015 | Video-based road detection via online structural learning
Yuan Yuan 0001, Zhiyu Jiang, Qi Wang 0009 |
Neurocomputing | 2 |