EDBT 2026 Demo / reviewers in the wild / expert
Xiuwei Zhang 0001
dblp:47/3693
· DBLP profile ↗
40ranked-venue papers
12as first author
25since 2021 · last 2026
0000-0001-7230-1476ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SOMA: Feature Gradient Enhanced Affine-Flow Matching for SAR-Optical RegistrationabstractAchieving pixel-level registration between SAR and optical images remains a challenging task due to their fundamentally different imaging mechanisms and visual characteristics. Although deep learning has achieved great success in many cross-modal tasks, its performance on SAR-Optical registration tasks is still unsatisfactory. Gradient-based information has traditionally played a crucial role in handcrafted descriptors by highlighting structural differences. However, such gradient cues have not been effectively leveraged in deep learning frameworks for SAR-Optical image matching. To address this gap, we propose SOMA, a dense registration framework that integrates structural gradient priors into deep features and refines alignment through a hybrid matching strategy. Specifically, we introduce the Feature Gradient Enhancer (FGE), which embeds multi-scale, multi-directional gradient filters into the feature space using attention and reconstruction mechanisms to boost feature distinctiveness. Furthermore, we propose the Global-Local Affine-Flow Matcher (GLAM), which combines affine transformation and flow-based refinement within a coarse-to-fine architecture to ensure both structural consistency and local accuracy. Experimental results demonstrate that SOMA significantly improves registration precision, increasing the CMR@1px by 12.29% on the SEN1-2 dataset and 18.50% on the GFGE_SO dataset. In addition, SOMA exhibits strong robustness and generalizes well across diverse scenes and resolutions. Tao Zhuo, Xiuwei Zhang 0001, Hanlin Yin, Wencong Wu, Yanning Zhang 0001 |
AAAI | 3 |
| 2026 | CDFNet: Cross-dimension fusion network with dual feature enhancement for multimodal object detection
Wencong Wu, Xiuwei Zhang 0001, Hanlin Yin, Haorui Zeng, Chenxu Wei |
Expert Syst. Appl. | 2 |
| 2026 | Lightweight modal-guided cross-attention fusion network for visible-infrared object detection
Wencong Wu, Hongxi Zhang, Xiuwei Zhang 0001, Hanlin Yin, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2025 | AdaSemiCD: An Adaptive Semi-Supervised Change Detection Method Based on Pseudo-Label EvaluationabstractChange detection (CD) is an essential field in remote sensing, with a primary focus on identifying areas of change in bitemporal image pairs captured at varying intervals of the same region. The data annotation process for CD tasks is both time-consuming and labor-intensive. To better utilize the scarce labeled data and abundant unlabeled data, we introduce an adaptive semi-supervised learning (SSL) method, AdaSemiCD, to improve pseudo-label usage and optimize the training process. Initially, due to the extreme class imbalance inherent in CD, the model is more inclined to focus on the background class, and it is easy to confuse the boundary of the target object. Considering these two points, we develop a measurable evaluation metric for pseudo-labels that enhances the representation of information entropy by class rebalancing and amplification of ambiguous areas, assigning greater weights to prospective change objects. Subsequently, to enhance the reliability of sample wise pseudo-labels, we introduce the AdaFusion module, to dynamically identify the most uncertain region and substitute it with more trustworthy content. Lastly, to ensure better training stability, we introduce the AdaEMA module, which updates the teacher model using only batches of trusted samples. Experimental results on ten public CD datasets validate the efficacy and generalizability of our proposed adaptive training framework. Lingyan Ran, Wen Dongcheng, Tao Zhuo, Shizhou Zhang, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | River Ice Fine-Grained Segmentation: A GF-2 Satellite Image Dataset and Deep Learning BenchmarkabstractSemantic segmentation of river ice image serves as a critical technological foundation for hydrological monitoring and ice flood early warning system. Current publicly available river ice datasets predominantly utilize UAV-captured image and ground-based photographic observations. To address the limitations of spatial coverage in existing datasets, we present NWPU_YRCC_GFICE - a satellite remote sensing dataset constructed from multi-spectral GF-2 satellite images. The dataset innovatively categorizes river ice into six fine-grained classes across freeze-thaw cycles and covers river ice data from Yellow River (Ningxia-Inner Mongolia section) spanning the past 10 years. We further establish a comprehensive deep learning benchmark, which evaluates 33 state-of-the-art segmentation models and two improved segmentation models based on YOLO and Segformer architecture, separately. Experiments are conducted on the NWPU_YRCC_GFICE dataset and three public river ice datasets (NWPU_YRCC_EX, NWPU_YRCC2, and Alberta river ice segmentation dataset). The proposed models exhibit excellent performance, surpassing the state-of-the-art methods. The presented NWPU_YRCC_GFICE dataset and benchmark enriches the river ice dataset and favors in promoting fine-grained river ice segmentation research from satellite view. Our dataset and code is available at https://github.com/ASGOLabMultisourceCooperationGroup/NWPU_YRCC_GFICE. Chenxu Wei, Haohao Zhou, Omirzhan Taukebayev, Wencong Wu, Amirkhan Temirbayev, Lingyan Ran, Hanlin Yin, Peng Wang 0015, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 13 |
| 2024 | Complementary Fusion Network Based on Frequency Hybrid Attention for PansharpeningabstractPansharpening is a feasible way to obtain the high-resolution (HR) multispectral (MS) images by using panchromatic (PAN) images to sharpen low-resolution MS images. Despite its great advances, most existing pansharpening methods neglect the importance of integrating local and non-local characteristics of images, resulting in the imbalance of spatial and spectral distribution. In this paper, we propose a complementary fusion network (CFNet) based on frequency hybrid attention mechanism for pansharpening. By introducing the frequency transformation and the deformable cross-attention, our model takes image-wide receptive field into consideration to explore global feature learning. Combined with the convolutional layers with local receptive field, CFNet can well capture local and non-local features. Experimental results demonstrate that the proposed method outperforms the comparison methods in terms of visual and quantitative qualities. Yinghui Xing, Litao Qu, Kai Zhang 0010, Yan Zhang 0127, Xiuwei Zhang 0001, Yanning Zhang 0001 |
ICASSP | 5 |
| 2024 | Edge-Guided Detector-Free Network for Robust and Accurate Visible-Thermal Image MatchingabstractRecent detector-free models strive to leverage both local and global context for image matching, showcasing enhanced robustness, particularly in scenarios with weak-textured scenes. Despite these advancements, automatically establishing feature correspondences between visible and thermal images still introduces additional challenges. Differences in radiation and geometry between these modalities often result in degraded performance for the majority of existing methods. To this end, we propose edge-guided detector-free model termed EDMatcher for visible-thermal image matching. Besides local and global context in the images, EDMatcher also leverages modality-robust structural information in image edges, which demonstrates promising robustness to images with distinct modalities. Moreover, an edge-masked ground-truth matrix generation strategy is introduced during the training, which helps EDMatcher to further focus on more salient regions while leaving out texture-less regions, leading to more efficient learning. Extensive experiments show that EDMatcher has strong generalization and achieves excellent matching performances. Zhaoshuai Qi, Xiuwei Zhang 0001, Tao Zhuo, Yanning Zhang 0001 |
ICME | 3 |
| 2024 | RGB-T Object Detection via Group Shuffled Multi-receptive Attention and Multi-modal Supervision
Jinzhong Wang, Xuetao Tian, Shun Dai, Tao Zhuo, Haorui Zeng, Hongjuan Liu, Xiuwei Zhang 0001, Yanning Zhang 0001 |
ICPR (17) | 8 |
| 2024 | Hierarchical Shared Architecture Search for Real-Time Semantic Segmentation of Remote Sensing ImagesabstractReal-time semantic segmentation of remote-sensing images demands a trade-off between speed and accuracy, which makes it challenging. Apart from manually designed networks, researchers seek to adopt neural architecture search (NAS) to discover a real-time semantic segmentation model with optimal performance automatically. Most existing NAS methods stack up no more than two types of searched cells, omitting the characteristics of resolution variation. This paper proposes the Hierarchical shared Architecture Search (HAS) method to automatically build a real-time semantic segmentation model for remote sensing images. Our model contains a lightweight backbone and a multi-scale feature fusion module. The lightweight backbone is carefully designed with low computational cost. The multi-scale feature fusion module is searched using the NAS method, where only the blocks from the same layer share identical cells. Extensive experiments reveal that our searched real-time semantic segmentation model of remote sensing images achieves the state-of-the-art trade-off between accuracy and speed. Specifically, on the LoveDA, Potsdam, and Vaihingen datasets, the searched network achieves 54.5% mIoU, 87.8% mIoU, and 84.1% mIoU, respectively, with an inference speed of 132.7 FPS. Besides, our searched network achieves 72.6% mIoU at 164.0 FPS on the CityScapes dataset and 72.3% mIoU at 186.4 FPS on the CamVid dataset. Wenna Wang, Lingyan Ran, Hanlin Yin, Mingjun Sun, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Empower Generalizability for Pansharpening Through Text-Modulated Diffusion ModelabstractPansharpening is crucial to remote sensing applications by fusing high-resolution (HR) panchromatic (PAN) images with low-resolution multispectral (LRMS) images to generate HR multispectral (HRMS) images. Recently, diffusion probabilistic models (DPMs) have provided high-quality results than regression-based methods when trained on specific pairwise data for their specific purpose. However, their performance degrades when applied to a new satellite dataset, which represents different imaging properties and spectral ranges, limiting the generalization ability of them. For better generalizability of pansharpening, in this article, we propose a text-modulated diffusion model (TMDiff) for unified pansharpening of different satellites. TMDiff takes a text-modulated 3-D UNet (TM3DU) as denoising network to gradually recover HRMS through iterative refinement over multiple time steps. By introducing satellite’s physical properties as text prompts, TM3DU is able to learn meta-knowledge across different satellites and thus can sharpen LRMS images with diverse spatial and spectral attributes. Extensive experiments on various satellite datasets demonstrate the state-of-the-art performance of our model in both qualitative and quantitative metrics. Furthermore, our model exhibits superior generalization ability to unseen datasets, highlighting its practical significance. Code is available athttps://github.com/codgodtao/TMDiff. Yinghui Xing, Litao Qu, Shizhou Zhang, Jiapeng Feng, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Improving Reliability of Heterogeneous Change Detection by Sample Synthesis and Knowledge TransferabstractDetecting changes in heterogeneous images without the supervision of changed label is a challenging yet critical task for quick responding natural disaster relief. Nevertheless, most of available unsupervised heterogeneous change detection methods strong rely on the quality of pseudo labels, and they suffer from performance degradation, even irreversible model collapse, when encounter the low-quality pseudo labels, leading to unreliable detection results. In order to improve the reliability of unsupervised heterogeneous change detection, in this paper, we propose a novel change detection paradigm based on sample synthesis and knowledge transfer. We address the issue of label reliability by artificially creating a changed region and assigning labels rather than constructing pseudo labels. These constructed labels guide the network in automatically learning the correspondence between heterogeneous images, confirming the reliability of changed regions. Moreover, an augmentation with synthetic samples on real samples makes it possible to generate more transferable samples while reducing the domain gap coarsely. A dual-branch joint training with feature contrastive learning is further developed to transfer the knowledge of changes from the synthetic sample domain to real sample domain. Experimental results on five public datasets demonstrate that our proposed method has superior performance when compared with available state-of-the-art methods. Our code is available at https://github.com/zhangqiiii/SS-KT. Yinghui Xing, Lingyan Ran, Xiuwei Zhang 0001, Hanlin Yin, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Remote Sensing Image Semantic Change Detection Boosted by Semi-Supervised Contrastive Learning of Semantic SegmentationabstractSemantic change detection (SCD) is a challenging task in remote sensing image (RSI) interpretation, which adopts multitemporal images to detect, locate, and analyze pixel-level land-cover “from-to” changes. In SCD, the severe class imbalance problem and the occurrence of confusing categories are very typical, making it challenging to accurately distinguish the easily confused categories with limited semantic context information. However, previous works did not address these issues in depth. This article proposes a novel SCD method named semi-supervised contrastive learning (SSCLNet), in which a simple and effective SCD network is designed as a strong baseline, and a semi-supervised contrastive learning module of semantic segmentation (SS) is presented to enhance the distinguishability of categories. Our baseline extracts semantic context through high-resolution network (HRNet), gets change information simply through an absolute difference, and then directly performs SCD based on the fusion of semantic context and change information. To utilize the semantic context information of the unlabeled non-changed regions, we employ a self-training (ST) method for semi-supervised SS. To learn distinguishable feature representations for easily confused categories, we present contrastive learning with an adaptive sampling strategy for SS. It selects challenging negative samples for each category from the other categories that exhibit similar features or attributes. The sampling space includes both the labeled changed samples and the non-changed samples predicted by ST. The comprehensive experiments on the SECOND and the Landsat-SCD dataset demonstrate that the proposed SSCLNet achieves the state-of-the-art (SOTA) performance, with a significant improvement of 2.07% and 4.15% in the score value, respectively. Xiuwei Zhang 0001, Yizhe Yang, Lingyan Ran, Kangwei Wang, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | MS-DETR: Multispectral Pedestrian Detection Transformer With Loosely Coupled Fusion and Modality-Balanced OptimizationabstractMultispectral pedestrian detection is an important task for many around-the-clock applications, since the visible and thermal modalities can provide complementary information especially under low light conditions. Due to the presence of two modalities, misalignment and modality imbalance are the most significant issues in multispectral pedestrian detection. In this paper, we propose MultiSpectral pedestrian DEtection TRansformer (MS-DETR) to fix above issues. MS-DETR consists of two modality-specific backbones and Transformer encoders, followed by a multi-modal Transformer decoder, and the visible and thermal features are fused in the multi-modal Transformer decoder. To well resist the misalignment between multi-modal images, we design a loosely coupled fusion strategy by sparsely sampling some keypoints from multi-modal features independently and fusing them with adaptively learned attention weights. Moreover, based on the insight that not only different modalities, but also different pedestrian instances tend to have different confidence scores to final detection, we further propose an instance-aware modality-balanced optimization strategy, which preserves visible and thermal decoder branches and aligns their predicted slots through an instance-wise dynamic loss. Our end-to-end MS-DETR shows superior performance on the challenging KAIST, CVC-14 and LLVIP benchmark datasets. The source code is available athttps://github.com/YinghuiXing/MS-DETR. Yinghui Xing, Song Wang 0002, Shizhou Zhang, Guoqiang Liang 0001, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | SSML-QNet: Scale-Separative Metric Learning Quadruplet Network for Multi-modal Image Patch MatchingabstractMulti-modal image matching is very challenging due to the significant diversities in visual appearance of different modal images. Typically, the existing well-performed methods mainly focus on learning invariant and discriminative features for measuring the relation between multi-modal image pairs. However, these methods often take the features as a whole and largely overlook the fact that different scale features for a same image pair may have different similarity, which may lead to sub-optimal results only. In this work, we propose a Scale-Separative Metric Learning Quadruplet network (SSML-QNet) for multi-modal image patch matching. Specifically, SSML-QNet can extract both relevant and irrelevant features of imaging modality with the proposed quadruplet network architecture. Then, the proposed Scale-Separative Metric Learning module separately encodes the similarity of different scale features with the pyramid structure. And for each scale, cross-modal consistent features are extracted and measured by coordinate and channel-wise attention sequentially. This makes our network robust to appearance divergence caused by different imaging mechanism. Experiments on the benchmark dataset (VIS-NIR, VIS-LWIR, Optical-SAR, and Brown) have verified that the proposed SSML-QNet is able to outperform other state-of-the-art methods. Furthermore, the cross-dataset transferring experiments on these four datasets also have shown that the proposed method has powerful ability of cross-dataset transferring. Xiuwei Zhang 0001, Hanlin Yin, Yinghui Xing, Yanning Zhang 0001 |
IJCAI | 1 |
| 2023 | Automatic Network Architecture Search for RGB-D Semantic SegmentationabstractRecent RGB-D semantic segmentation networks are usually manually designed. However, due to limited human efforts and time costs, their performance might be inferior for complex scenarios. To address this issue, we propose the first Neural Architecture Search (NAS) method that designs the network automatically. Specifically, the target network consists of an encoder and a decoder. The encoder is designed with two independent branches, where each branch specializes in extracting features from RGB and depth images, respectively. The decoder fuses the features and generates the final segmentation result. Besides, for automatic network design, we design a grid-like network-level search space combined with a hierarchical cell-level search space. By further developing an effective gradient-based search strategy, the network structure with hierarchical cell architectures is discovered. Extensive results on two datasets show that the proposed method outperforms the state-of-the-art approaches, which achieves a mIoU score of 55.1% on the NYU-Depth v2 dataset and 50.3% on the SUN-RGBD dataset. Wenna Wang, Tao Zhuo, Xiuwei Zhang 0001, Mingjun Sun, Hanlin Yin, Yinghui Xing, Yanning Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | FP-DARTS: Fast parallel differentiable neural architecture search for image classification
Wenna Wang, Xiuwei Zhang 0001, Hengfei Cui, Hanlin Yin, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2023 | AugFCOS: Augmented fully convolutional one-stage object detection network
Xiuwei Zhang 0001, Yinghui Xing, Wenna Wang, Hanlin Yin, Yanning Zhang 0001 |
Pattern Recognit. | 1 |
| 2023 | FastICENet: A real-time and accurate semantic segmentation model for aerial remote sensing river ice image
Xiuwei Zhang 0001, Lingyan Ran, Yinghui Xing, Wenna Wang, Zeze Lan, Hanlin Yin, Houjun He, Qixing Liu, Baosen Zhang, Yanning Zhang 0001 |
Signal Process. | 1 |
| 2023 | Pansharpening via Frequency-Aware Fusion Network With Explicit Similarity ConstraintsabstractThe process of fusing a high spatial resolution (HR) panchromatic (PAN) image and a low spatial resolution (LR) multispectral (MS) image to obtain an HRMS image is known as pansharpening. With the development of convolutional neural networks, the performance of pansharpening methods has been improved, however, the blurry effects and the spectral distortion still exist in their fusion results due to the insufficiency in details learning and the frequency mismatch between MS and PAN. Therefore, the improvement of spatial details at the premise of reducing spectral distortion is still a challenge. In this paper, we propose a frequency-aware fusion network (FAFNet) together with a novel high-frequency feature similarity loss to address above mentioned problems. FAFNet is mainly composed of two kinds of blocks, where the frequency aware blocks aim to extract features in the frequency domain with the help of discrete wavelet transform (DWT) layers, and the frequency fusion blocks reconstruct and transform the features from frequency domain to spatial domain with the assistance of inverse DWT (IDWT) layers. Finally, the fusion results are obtained through a convolutional block. In order to learn the correspondence, we also propose a high-frequency feature similarity loss to constrain the HF features derived from PAN and MS branches, so that HF features of PAN can reasonably be used to supplement that of MS. Experimental results on three datasets at both reduced- and full-resolution demonstrate the superiority of the proposed method compared with several state-of-the-art pansharpening models. The codes are available at https://github.com/YinghuiXing/FAFNet. Yinghui Xing, Yan Zhang 0127, Houjun He, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Progressive Modality-Alignment for Unsupervised Heterogeneous Change DetectionabstractChange detection based on heterogeneous images is of great importance in some applications, such as disaster monitoring and damage assessment. However, due to the huge modality discrepancy in heterogeneous images, it is difficult to accurately detect the changed regions. In this paper, we analyze the interference of modality-alignment and changed areas to each other, and propose a progressive modality-alignment based unsupervised change detection model for heterogeneous images. Specifically, the modality alignment is achieved in an iterative manner, which can improve the detection accuracy progressively. To reduce the influence of modality discrepancy and the changed regions to each other, a pseudo-label self-learning strategy is designed, where the pseudo-labels learned by the model itself are used to act as a guidance of change detection, and they are in turn refined by the proposed progressive model. Experimental results on different real heterogeneous images verify the effectiveness and robustness of proposed method. Yinghui Xing, Lingyan Ran, Xiuwei Zhang 0001, Hanlin Yin, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SSA-Net: Spatial Scale Attention Network for Image-Based Geo-LocalizationabstractImage-based geo-localization is estimating the location of a query image by matching it to a large amount of images in geo-tagged database. This matching task is very challenging due to the vast differences in visual appearance or modality of image pairs on different platforms, for example, one image from the RGB camera, the other from the light detection and ranging (LiDAR) sensor. The spatial layout of the scene can provide important clues and significantly reduce matching ambiguity. Therefore, we propose a novel deep network that embeds spatial configuration of the scenes into feature representation. Specifically, we design a spatial-scale attention (SSA) module to highlight the salience correspondence layout features at different scales. The encoded features not only represent the emergence of certain objects, but also reflect the relative locations of the objects. By this way, we learn more discriminative deep feature representations, leading to a higher recall. The experimental results on two standard cross-view benchmark datasets (CVUSA and CVACT) and a cross-modal dataset (GRAL) demonstrate that our method performs better than the state-of-the-art methods. Remarkably, the recall rate@top-1 improves from 27.6% to 40.5% on the GRAL dataset. Xiuwei Zhang 0001, Xiangchuang Meng, Hanlin Yin, Yuanzeng Yue, Yinghui Xing, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | DifUnet++: A Satellite Images Change Detection Network Based on Unet++ and Differential PyramidabstractChange detection (CD) is one of the most important topics in the field of remote sensing. In this letter, we propose an effective satellite images CD network named DifUnet++. As the presentation of explicit difference is more conducive to extract change features, we design a differential pyramid of two input images as the input of Unet++. Considering the scale diversity of changed regions in remote sensing images, a multiply side-outs fusion strategy is adopted to predict the detection results of different scales. Furthermore, a learning upsampling method is utilized to refine the details of CD. The proposed architecture is evaluated on two public satellite image CD data sets. The experimental results show that our method performs much better than state-of-the-art methods. Xiuwei Zhang 0001, Yuanzeng Yue, Wenxiang Gao, Shuai Yun, Qian Su, Hanlin Yin, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | ADHR-CDNet: Attentive Differential High-Resolution Change Detection Network for Remote Sensing ImagesabstractWith the development of deep learning, change detection technology has gained great progress. However, how to effectively extract multi-scale substantive changed features and accurately detect small changed objects as well as the accurate details is still a challenge. To solve the problem, we propose Attentived Differential High-Resolution Change Detection Network (ADHR-CDNet) for remote sensing images. In ADHR-CDNet, a novel high-resolution backbone with a Differential Pyramid Module (DPM) is proposed to extract multi-level and multi-scale substantive changed features. The backbone structure with four interconnected sub-network branches of different resolution is helpful to extract multi-level and multi-scale features. DPM is capable of distinguishing between substantive changes and pseudo changes induced by illumination, shadow, seasonal variation, and so on. Then, a novel Multi-Scale Spatial feature Attention Module (MSSAM) is presented to effectively fuse the spatial detail information of different scale features produced by our backbone to generate finer prediction. We conduct quantitative and qualitative experiments on three public change detection datasets: the Lebedev, the LEVIR-CD, and the WHU Building dataset. The proposed ADHR-CDNet reaches F1-score of 97.2% (improved 3.1%) on the Lebedev dataset, 91.4% (improved 1.6%) on the LEVIR-CD dataset, and 90.9% (improved 1.2%) on the WHU Building dataset. The experimental results demonstrate that our method performs much better than the state-of-the-art methods. The visualization comparison results show that our method can effectively detect small changed objects and significantly improve the details of detected changed objects. Our code is available at https://github.com/w-here/ASGO-113lab/tree/main/ADHR-CDNet. Xiuwei Zhang 0001, Mu Tian, Yinghui Xing, Yuanzeng Yue, Hanlin Yin, Runliang Xia, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | River ice monitoring and change detection with multi-spectral and SAR images: application over yellow river
Xiuwei Zhang 0001, Yuanzeng Yue, Fei Li 0011, Xiuzhong Yuan, Minhao Fan, Yanning Zhang 0001 |
Multim. Tools Appl. | 1 |
| 2021 | Attend to the Difference: Cross-Modality Person Re-Identification via Contrastive CorrelationabstractThe problem of cross-modality person re-identification has been receiving increasing attention recently, due to its practical significance. Motivated by the fact that human usually attend to the difference when they compare two similar objects, we propose a dual-path cross-modality feature learning framework which preserves intrinsic spatial structures and attends to the difference of input cross-modality image pairs. Our framework is composed by two main components: a Dual-path Spatial-structure-preserving Common Space Network (DSCSN) and a Contrastive Correlation Network (CCN). The former embeds cross-modality images into a common 3D tensor space without losing spatial structures, while the latter extracts contrastive features by dynamically comparing input image pairs. Note that the representations generated for the input RGB and Infrared images are mutually dependant to each other. We conduct extensive experiments on two public available RGB-IR ReID datasets, SYSU-MM01 and RegDB, and our proposed method outperforms state-of-the-art algorithms by a large margin with both full and simplified evaluation modes. Shizhou Zhang, Peng Wang 0015, Guoqiang Liang 0001, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | An Improved Camouflage Target Detection Using Hyperspectral Image Based on Block-Diagonal and Low-Rank Representation
Fei Li 0011, Xiuwei Zhang 0001, Lei Zhang 0054, Yanning Zhang 0001, Dongmei Jiang, Genping Zhao |
PRCV (4) | 2 |
| 2018 | Visible and infrared image registration based on region features and edginess
Yanjia Chen, Xiuwei Zhang 0001, Yanning Zhang 0001, Stephen J. Maybank, Zhipeng Fu |
Mach. Vis. Appl. | 2 |
| 2018 | Exploiting Structured Sparsity for Hyperspectral Anomaly DetectionabstractSparse representation-based background modeling facilitates much recent progress in hyperspectral anomaly detection (AD). The sparse representation of background often exhibits underlying structure, which is crucial to distinguish between background and anomaly. However, how to exploit such underlying structure is still challenging. To address this problem, we present a novel hyperspectral AD method, which can exploit the structured sparsity in modeling the background more accurately. With the plausible background area detected by a local RX detector, a robust background spectrum dictionary is learned in a principal component analysis way. A reweighted Laplace prior-based structured sparse representation model is then employed to reconstruct the spectrum of each pixel. With considering the structured sparsity in representation, the background pixels can be reconstructed more accurately than the anomaly ones, which thus can be detected based on the reconstruction error. To further improve the detection performance, an intracluster reconstruction model is developed to exploit the spatial similarity among the background pixels in the same cluster. The anomaly pixels can then be detected based on the cost of intracluster reconstruction error. By linearly combining these two detection results, improvement is obviously achieved on detection accuracy. Experimental results on both simulated and real-world data sets demonstrate that the proposed method outperforms several state-of-the-art hyperspectral AD methods. Fei Li 0011, Xiuwei Zhang 0001, Lei Zhang 0054, Dongmei Jiang, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Multi-modal Image Registration Based on Modified-SURF and Consensus Inliers Recovery
Yanjia Chen, Xiuwei Zhang 0001, Fei Li 0011, Yanning Zhang 0001 |
ICIG (2) | 2 |
| 2016 | Hyperspectral anomaly detection using background learning and structured sparse representationabstractA novel background dictionary learning and structured sparse representation based anomaly detection method is proposed for hyperspectral imagery. First, a robust PCA spectrum dictionary is learned from the plausible background area detected by the local RX detector. With the learned dictionary, the reweighted Laplace prior based structured sparse representation model is then employed to reconstruct the spectrum of each pixel in the image. Due to considering the structured sparsity in representation, the background spectra can be reconstructed more accurately than anomaly ones. Thus, reconstruction error is utilized to separate the anomaly pixels and background ones. Experimental results on both simulated and real-world datasets demonstrate that the proposed method outperforms several state-of-the-art hyperspectral anomaly detection methods. Fei Li 0011, Yanning Zhang 0001, Lei Zhang 0054, Xiuwei Zhang 0001, Dongmei Jiang |
IGARSS | 4 |
| 2014 | A multi-modal moving object detection method based on GrowCut segmentationabstractCommonly-used motion detection methods, such as background subtraction, optical flow and frame subtraction are all based on the differences between consecutive image frames. There are many difficulties, including similarities between objects and background, shadows, low illumination, thermal halo. Visible light images and thermal images are complementary. Many difficulties in motion detection do not occur simultaneously in visible and thermal images. The proposed multimodal detection method combines the advantages of multi-modal image and GrowCut segmentation, overcomes the difficulties mentioned above and works well in complicated outdoor surveillance environments. Experiments showed our method yields better results than commonly-used fusion methods. Xiuwei Zhang 0001, Yanning Zhang 0001, Stephen J. Maybank |
CIMSIVP | 1 |
| 2014 | Web Service Recommendation Based on Watchlist via Temporal and Tag Preference FusionabstractWith the increasing number of Web services available on the Internet, how to recommend Web services to interested users effectively and efficiently remains to be a big challenge. At present, collaborative filtering (CF) is the most widely used technique in the design of recommender systems to handle information overload. For Web services, however, it is difficult for user to collect personalized QoS (Quality of Service)data and other explicit feedbacks such as ratings. In most cases, only a part of the implicit feedbacks (e.g., watchlist) is available in service registry. In this paper, we leverage implicit feedback from user's watchlist to build a CF-based recommender system for Web service. Our main contribution is to transform implicit feedbacks into explicit ratings to improve the accuracy of service recommendation. More specifically, we first construct binary user-service rating matrix according to the implicit feedback from the watchlist. Then, temporal and tag preference are combined into the original rating matrix to generate a more accurate pseudo rating matrix, which can reflect users' different preference on services in their own watchlists. Finally, we use traditional user-based CF method to produce a personalized service recommendation list with corresponding pseudo ratings. Moreover, the empirical experiments based on ProgrammableWeb show that compared with traditional log-based CF method, the recommender system with temporal and tag preference is more accurate and precise. Xiuwei Zhang 0001, Keqing He 0002, Jian Wang 0018, Chong Wang 0004, Gang Tian, Jianxiao Liu |
ICWS | 1 |
| 2013 | An IR and visible image sequence automatic registration method based on optical flow
Yanning Zhang 0001, Xiuwei Zhang 0001, Stephen J. Maybank |
Mach. Vis. Appl. | 2 |
| 2012 | Business Rule Engine-based Framework for SaaS Application Development
Xiuwei Zhang 0001, Keqing He 0002, Jian Wang 0018, Chong Wang 0004 |
CLOSER | 1 |
| 2012 | A novel multi-object detection method in complex scene using synthetic aperture imaging
Zhao Pei, Yanning Zhang 0001, Tao Yang 0006, Xiuwei Zhang 0001, Yee-Hong Yang |
Pattern Recognit. | 4 |
| 2011 | A Practical Architecture of Cloudification of Legacy ApplicationsabstractCloud computing has been attracting much attention since its birth. How to cloudify software systems especially legacy applications in the cloud era is becoming increasingly important. Based on RGPS meta-model framework and International standards-ISO/IEC 19763, an architecture for cloudification of legacy applications is proposed, which consists of three parts: a Web portal, a SaaS service supermarket, and a SaaS application development platform. In this paper, we take an open source software as an example to illustrate the proposed approach. Based on the architecture and supporting techniques on software virtualization and multi-tenancy, we develop a prototype Cloud CRM to demonstrate the basic procedure for cloudification of legacy applications, as well as the feasibility of the proposed approach. Dunhui Yu, Jian Wang 0018, Bo Hu 0013, Jianxiao Liu, Xiuwei Zhang 0001, Keqing He 0002, Liang-Jie Zhang |
SERVICES | 5 |
| 2009 | A Novel Multi-planar Homography Constraint Algorithm for Robust Multi-people Location with Severe OcclusionabstractMulti-view approach has been proposed to solve occlusion and lack of visibility in crowded scenes. However, the problem is that too much redundancy information might bring about false alarm. Although researchers have done many efforts on how to use the multi-view information to track people accurately, it is particularly hard to wipe off the false alarm. Our approach is to use multiple views cooperatively to detect objects and use objects silhouette on planes of different height to remove false alarm. To achieve this we adopt a novel multi-planar homography constraint to resolve occlusions and false alarm. Experimental results show that our algorithm is able to accurately locate people in crowded scene maintaining correct correspondences across views. Moreover, the false alarm rate is obviously reduced. Xiaomin Tong, Tao Yang 0006, Runping Xi, Dapei Shao, Xiuwei Zhang 0001 |
ICIG | 5 |
| 2009 | Image Registration Based on Rectangle PatternabstractTo against the complexity in finding feature-pair in image registration caused by traditional features: corners, lines and image edge, a novel method for image registration based on rectangle pattern is proposed. Unlike traditional features, a rectangle pattern can be described as its four vertexes and center, which can afford five pair-wise points for any kind of image transformation, and it holds stable in different weather condition, time and imaging way, which formed by building angular, widely found in aero-image. Firstly, image edge is detected by canny, and then distance transform is applied on the result of canny, by follows, a thresholding and mask convoluting are used on the distance transform result to avoid the interfering complex lines and edge. The center of the rectangles is obtained by clustering on the prior result, and then the four vertexes is calculated by geometric restrict on rectangle pattern. Consequently, we calculate the centers pair of the rectangles by their slope difference which aims to get the correct pair-wise points set between two images. Finally, we use the four vertexes pair and center pair of the rectangle pattern as the input of RANSAC algorithm, to solve an affine transform. The proposed algorithm is proved to be effective and accurate on the translation, rotation and scale between electro-optic(EO) images pair and SAR-EO images pair. 1.3 pixels of registration accuracy result is obtained in the experiment. Xingong Zhang, Runping Xi, Xiuwei Zhang 0001, Tao Yang 0006 |
ICIG | 4 |
| 2009 | A Convenient Multi-camera Self-Calibration Method Based on Human Body Motion AnalysisabstractA novel and convenient multi-camera self-calibration method is proposed in this paper. Different from other calibration methods, our method is done by analyzing human body motion. The only constraint is that several people of different heights are needed to walk around the experimental environment one by one in the calibration period. By this way, two kinds of corresponding points are extracted from synchronous video sequences. One is the centroid of the moving human body. The other is points on the floor, which is extracted by matching floor planes in video sequences. The floor planes registration is based on shadow detection and co-motion feature. Based on these corresponding points, camera parameters and 3D points observed are estimated. The proposed method is tested in our own experimental environment. Experimental results show the accuracy of our calibration method. Our method can satisfy many applications of multi-view computer vision. Xiuwei Zhang 0001, Yanning Zhang 0001, Xingong Zhang, Tao Yang 0006, Xiaomin Tong, Haichao Zhang 0001 |
ICIG | 1 |
| 2008 | Pedestrian detection based on multi-modal cooperationabstractPedestrian detection plays an important role in automated surveillance system. However, it is challenging to detect pedestrian robustly and accurately in a cluttered environment. In this paper, we propose a new cooperative pedestrian detection method using both colour and thermal image sequences, which is compared with the method using only colour image sequence and that using multi-modal fusion. Experiment results show that our cooperative detection mechanism could get more accurate pedestrian areas, a lower false alarm rate and a higher detection precision. Therefore, it has broad application prospects in the field of industry and military. Yanning Zhang 0001, Xiaomin Tong, Xiuwei Zhang 0001, Jiangbin Zheng 0001, Siwei You |
MMSP | 3 |