Zhiguo Jiang 0001

dblp:08/1760-1 · also Zhi-Guo Jiang 0001 · DBLP profile ↗
← Back
95ranked-venue papers
1as first author
35since 2021 · last 2026
0000-0001-8786-2540ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 60 · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 2 since 2021Systems, architecture and hardware · 2
YearPublicationVenuePosition
2026 Adapting pathology foundation models for continual cross-center WSI retrieval
Zhiguo Jiang 0001, Kun Wu 0010, Jun Shi 0006, Yushan Zheng
Medical Image Anal.2
2026 Lifelong content-based histopathology image retrieval via bilevel coreset selection and distance consistency rehearsal
Zhiguo Jiang 0001, Kun Wu 0010, Jun Shi 0006, Yushan Zheng
Pattern Recognit.2
2025 Promptable Representation Distribution Learning and Data Augmentation for Gigapixel Histopathology WSI Analysis
abstract
Gigapixel image analysis, particularly for whole slide images (WSIs), often relies on multiple instance learning (MIL). Under the paradigm of MIL, patch image representations are extracted and then fixed during the training of the MIL classifiers for efficiency consideration. However, the invariance of representations makes it difficult to perform data augmentation for WSI-level model training, which significantly limits the performance of the downstream WSI analysis. The current data augmentation methods for gigapixel images either introduce additional computational costs or result in a loss of semantic information, which is hard to meet the requirements for efficiency and stability needed for WSI model training. In this paper, we propose a Promptable Representation Distribution Learning framework (PRDL) for both patch-level representation learning and WSI-level data augmentation. Meanwhile, we explore the use of prompts to guide data augmentation in feature space, which achieves promptable data augmentation for training robust WSI-level models. The experimental results have demonstrated that the proposed method stably outperforms state-of-the-art methods.
Kunming Tang, Zhiguo Jiang 0001, Jun Shi 0006, Wei Wang 0380, Yushan Zheng
AAAI2
2025 RescueADI: Adaptive Disaster Interpretation in Remote Sensing Images With Autonomous Agents
abstract
Current methods for disaster scene interpretation in remote sensing images (RSIs) mostly focus on isolated tasks such as segmentation, detection, or visual question-answering (VQA). However, these methods often fail to provide comprehensive and actionable insights, particularly in scenarios that demand the integration of multiple perception methods and specialized tools to address complex, multilayered challenges in geophysical disaster analysis. To fill this gap, this article introduces adaptive disaster interpretation (ADI), a novel task designed to solve requests by planning and executing multiple sequentially correlative interpretation tasks to provide a comprehensive analysis of disaster scenes. To facilitate research and application in this area, we present a new dataset named RescueADI, which contains high-resolution RSIs with annotations for three connected aspects: planning, perception, and recognition. The dataset includes 4044 RSIs, 16949 semantic masks, 14483 object bounding boxes, and 13424 interpretation requests across nine challenging request types. Moreover, we propose a new disaster interpretation method employing autonomous agents driven by large language models (LLMs) for task planning and execution, proving its efficacy in handling complex disaster interpretations. The proposed agent-based method solves various complex interpretation requests such as counting, area calculation, and path finding without human intervention, which traditional single-task approaches cannot handle effectively. Experimental results on RescueADI demonstrate the feasibility of the proposed task and show that our method achieves an accuracy 9% higher than existing VQA methods, highlighting its advantages over conventional disaster interpretation approaches.
Zhuoran Liu 0006, Danpei Zhao, Bo Yuan 0009, Zhiguo Jiang 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Partial-Label Contrastive Representation Learning for Fine-Grained Biomarkers Prediction From Histopathology Whole Slide Images
abstract
In the domain of histopathology analysis, existing representation learning methods for biomarkers prediction from whole slide images (WSIs) face challenges due to the complexity of tissue subtypes and label noise problems. This paper proposed a novel partial-label contrastive representation learning approach to enhance the discrimination of histopathology image representations for fine-grained biomarkers prediction. We designed a partial-label contrastive clustering (PLCC) module for partial-label disambiguation and a dynamic clustering algorithm to sample the most representative features of each category to the clustering queue during the contrastive learning process. We conducted comprehensive experiments on three gene mutation prediction datasets, including USTC-EGFR, BRCA-HER2, and TCGA-EGFR. The results show that our method outperforms 9 existing methods in terms of Accuracy, AUC, and F1 Score. Specifically, our method achieved an AUC of 0.950 in EGFR mutation subtyping of TCGA-EGFR and an AUC of 0.853 in HER2 0/1+/2+/3+ grading of BRCA-HER2, which demonstrates its superiority in fine-grained biomarkers prediction from histopathology whole slide images.
Yushan Zheng, Kun Wu 0010, Jun Li 0106, Kunming Tang, Jun Shi 0006, Zhiguo Jiang 0001, Wei Wang 0380
IEEE J. Biomed. Health Informatics7
2025 Slide-Based Graph Collaborative Training for Histopathology Whole Slide Image Analysis
abstract
The development of computational pathology lies in the consensus that pathological characteristics of tumors are significant guidance for cancer diagnostics. Most existing research focuses on the inner-contextual information within each WSI yet ignores the possible inter-correlations between slides. As the development of tumors is a continuous process involving a series of histological, morphological, and genetic changes that accumulate over time, the similarities and differences between WSIs across various stages, grades, locations and patients should potentially contribute to the representation of WSIs and deserve to be taken into account in WSI modeling. To verify the advancement of introducing the slide inter-correlations into the representation learning of WSIs, we proposed a generic WSI analysis pipeline SlideGCD that can be adapted to any existing Multiple Instance Learning (MIL) frameworks and improve their performance. With the new paradigm, the prior knowledge of cancer development can participate in the end-to-end workflow, which concurrently initializes and refines the slide representation, as a guide for message passing in the slide-based graph. Extensive comparisons and experiments are conducted to validate the effectiveness and robustness of the proposed pipeline across 4 different tasks, including cancer subtyping, cancer staging, survival prediction, and gene mutation prediction, with 8 representative SOTA WSI analysis frameworks as backbones. The code is available at https://github.com/HFUT-miaLab/SlideGCD.
Jun Shi 0006, Tong Shu, Zhiguo Jiang 0001, Wei Wang 0380, Yushan Zheng
IEEE Trans. Medical Imaging3
2025 Self-Supervised Representation Distribution Learning for Reliable Data Augmentation in Histopathology WSI Classification
abstract
Multiple instance learning (MIL) based whole slide image (WSI) classification is often carried out on the representations of patches extracted from WSI with a pre-trained patch encoder. The performance of classification relies on both patch-level representation learning and MIL classifier training. Most MIL methods utilize a frozen model pre-trained on ImageNet or a model trained with self-supervised learning on histopathology image dataset to extract patch image representations and then fix these representations in the training of the MIL classifiers for efficiency consideration. However, the invariance of representations cannot meet the diversity requirement for training a robust MIL classifier, which has significantly limited the performance of the WSI classification. In this paper, we propose a Self-Supervised Representation Distribution Learning framework (SSRDL) for patch-level representation learning with an online representation sampling strategy (ORS) for both patch feature extraction and WSI-level data augmentation. The proposed method was evaluated on three datasets under three MIL frameworks. The experimental results have demonstrated that the proposed method achieves the best performance in histopathology image representation learning and data augmentation and outperforms state-of-the-art methods under different WSI classification frameworks. The code is available at https://github.com/lazytkm/SSRDL.
Kunming Tang, Zhiguo Jiang 0001, Kun Wu 0010, Jun Shi 0006, Fengying Xie, Wei Wang 0380, Yushan Zheng
IEEE Trans. Medical Imaging2
2025 Pan-Cancer Histopathology WSI Pre-Training With Position-Aware Masked Autoencoder
abstract
Large-scale pre-training models have promoted the development of histopathology image analysis. However, existing self-supervised methods for histopathology images primarily focus on learning patch features, while there is a notable gap in the availability of pre-training models specifically designed for WSI-level feature learning. In this paper, we propose a novel self-supervised learning framework for pan-cancer WSI-level representation pre-training with the designed position-aware masked autoencoder (PAMA). Meanwhile, we propose the position-aware cross-attention (PACA) module with a kernel reorientation (KRO) strategy and an anchor dropout (AD) mechanism. The KRO strategy can capture the complete semantic structure and eliminate ambiguity in WSIs, and the AD contributes to enhancing the robustness and generalization of the model. We evaluated our method on 7 large-scale datasets from multiple organs for pan-cancer classification tasks. The results have demonstrated the effectiveness and generalization of PAMA in discriminative WSI representation learning and pan-cancer WSI pre-training. The proposed method was also compared with 8 WSI analysis methods. The experimental results have indicated that our proposed PAMA is superior to the state-of-the-art methods. The code and checkpoints are available at https://github.com/WkEEn/PAMA.
Kun Wu 0010, Zhiguo Jiang 0001, Kunming Tang, Jun Shi 0006, Fengying Xie, Wei Wang 0380, Yushan Zheng
IEEE Trans. Medical Imaging2
2024 Report-Guided Cross-Modal Representation Learning for Predicting EGFR Mutations by Whole Slide Image
abstract
Traditional PCR/NGS-based multigene panel testing is time-consuming and costly. Predicting EGFR mutations directly from H&E stained whole slide images (WSIs) can alleviate these limitations. Furthermore, histopathological reports contain valuable textual information that correlates with tissue areas in WSIs. However, recent research mainly analyses EGFR mutation status only from a single modality, ignoring rich information contained in reports. In this paper, we propose a report-guided cross-modal representation learning method for predicting EGFR mutations by WSIs. Specifically, we reconstruct report-level embeddings through exploring intrinsic relationships between diagnostic words in histopathological reports and tissue areas in WSIs. Finally, reconstructed histopathological report embedding and aggregated WSI embedding are fused for final prediction. More importantly, molecular testing report is also introduced as prior supervision information at the training stage to guarantee semantic consistency of fused feature and molecular report embedding. We evaluate our method on the TCGA-EGFR public benchmark dataset and an in-house clinical dataset (USTC-EGFR). Experimental results demonstrate that our method outperforms existing approaches in EGFR mutation prediction, highlighting the benefits of cross-modal learning in enhancing feature representational ability. The code is available at https://github.com/HFUT-miaLab/RCRL.
Qi Qiao, Jun Shi 0006, Zhiguo Jiang 0001, Wei Wang 0380, Yushan Zheng
BIBM3
2024 Satellite Video Super-Resolution via Unidirectional Recurrent Network and Various Degradation Modeling
abstract
Satellite video images contain temporal contextual information that is unavailable in single-frame images. Therefore, using a sequence of frames for super-resolution can significantly enhance the reconstruction effect. However, most existing satellite Video Super-Resolution (VSR) methods focus on improving the network’s presentation ability, overlooking the complex degradation processes present in real-world satellite videos which appear as a blind SR problem. In this paper, we propose an effective satellite VSR method based on a unidirectional recurrent network named URD-VSR. Simultaneously, a network independent of the SR structure is utilized to model the degradation process. Experiments on real satellite video datasets and integration with object detection demonstrate the effectiveness of the proposed method.
Xiaoyuan Wei, Haopeng Zhang 0001, Zhiguo Jiang 0001
IGARSS3
2024 SlideGCD: Slide-Based Graph Collaborative Training with Knowledge Distillation for Whole Slide Image Classification
Tong Shu, Jun Shi 0006, Dongdong Sun, Zhiguo Jiang 0001, Yushan Zheng
MICCAI (4)4
2024 Lifelong Histopathology Whole Slide Image Retrieval via Distance Consistency Rehearsal
Zhiguo Jiang 0001, Kun Wu 0010, Jun Shi 0006, Yushan Zheng
MICCAI (4)2
2024 Thick Cloud Removal in Multitemporal Remote Sensing Images Using a Coarse-to-Fine Framework
abstract
Abstract—Remote sensing (RS) images are widely used for Earth observation. However, cloud contamination greatly degrades the quality of RS images and limits their applications. In this letter, we propose a coarse-to-fine thick cloud removal method for a single pair of multitemporal RS images. First, we perform a global color transformation on a cloud-free reference image using linear regression coefficients between the pixels in the cloudy target image and the reference image in the same cloud-free regions, and obtain a coarse result. Then, a convolutional neural network (CNN) based on internal constraint is used to refine the coarse result, which does not require any construction of additional external training dataset in advance. We further design a multiscale feature extraction and fusion module and an auxiliary loss involving cloud regions to improve the performance of the CNN. Finally, Poisson image fusion is employed to generate a seamless cloud-free result. On a simulated test set containing 500 pairs of multitemporal RS images, the proposed method achieves satisfactory results with 25.1277 dB in peak signal-to-noise ratio (PSNR), 0.9077 in structural similarity (SSIM), and 0.9342 in correlation coefficient (CC). Qualitative and quantitative comparisons of our proposed against several state-of-the-art methods on the simulated and real cloudy images demonstrate the superiority of the proposed method.
Yue Zi, Xuedong Song, Fengying Xie, Zhiguo Jiang 0001
IEEE Geosci. Remote. Sens. Lett.4
2024 Histopathology language-image representation learning for fine-grained digital pathology cross-modal retrieval
Dingyi Hu, Zhiguo Jiang 0001, Jun Shi 0006, Fengying Xie, Kun Wu 0010, Kunming Tang, Jianguo Huai, Yushan Zheng
Medical Image Anal.2
2024 Weakly Supervised Remote Sensing Image Semantic Segmentation With Pseudo-Label Noise Suppression
abstract
Semantic segmentation of remote sensing images (RSIs) plays a crucial role in various applications, including urban planning and environmental monitoring. However, the high cost and complexity of obtaining detailed annotations for RSIs pose significant challenge. This issue necessitates the exploration of weakly supervised learning as an effective alternative, which utilizes more readily available, less granular forms of labeling. Yet, weakly supervised approaches face their own set of challenges, primarily due to scarcity of precise pixel-level labels which significantly hampers the model’s ability to learn accurate representations. In this article, we introduce a weakly supervised semantic segmentation (WSSS) approach for RSIs that leverages self-supervised learning (SSL) and pseudo-label noise mitigation to address these challenges. Our method leverages a self-supervised encoder for providing similarity information, which enhances feature representation in RSIs and enables the generation of more accurate pseudo-labels, thus reducing the noise in the pseudo-labels. Furthermore, we propose a refined loss function that incorporates gradient clipping and label smoothing to mitigate the impact of noisy labels, thereby improving the robustness and accuracy of the segmentation results. Extensive experiments on the ISPRS Potsdam, ISPRS Vaihingen, and iSAID datasets demonstrate that our approach achieves state-of-the-art (SOTA) performance, closely matching that of fully supervised methods. Our method not only reduces the dependency on expensive pixel-level annotations but also showcases the potential of SSL in enhancing WSSS tasks.
Zhiguo Jiang 0001, Haopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 OPODet: Toward Open World Potential Oriented Object Detection in Remote Sensing Images
abstract
Despite recent advances in object detection, closed-set detectors with fixed training classes often overlook or misclassify unannotated objects during testing. To address this, open world object detection (OWOD) algorithms identify and label these objects as unknown, better aligning with real-world scenarios and human learning. However, remote sensing images, with their arbitrary object orientations and large interclass feature disparities, pose significant challenges for these algorithms. To tackle this, we propose OPODet, an Open-world Potential Oriented object Detection framework for remote sensing images. Specifically, we incorporate the oriented unknown-aware region proposal network (OUA-RPN) into traditional oriented object detection models, enabling the network to predict potential oriented objects. To address the significant interclass feature differences among potential unknown classes, we propose a multiunknown-class clustering aligning prototype (MCAP) learning method to prevent feature collapse in the feature space. In addition, to address the lack of rotation information for potential objects, we introduce a rotation potential target consistency (RPTC) algorithm to impose explicit rotation constraints for generating more accurate potential unknown proposals. Extensive experiments on DIOR-R, DOTA-v1.0, and HRSC2016 datasets demonstrate the effectiveness of our approach in detecting potential oriented objects.
Zhiwen Tan, Zhiguo Jiang 0001, Zheming Yuan, Haopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Label Evolution Based on Local Contrast Measure for Single-Point Supervised Infrared Small-Target Detection
abstract
In recent years, the implementation of infrared small-target detection using convolutional neural networks (CNNs) has garnered widespread attention due to its high performance. Since this issue is often addressed through fully supervised image segmentation networks, the training process requires significant human effort and time to annotate pixel-level mask labels. Consequently, employing single-point supervision as a form of weak supervision for model training has aroused widespread interest in saving annotation costs. However, the class imbalance issue caused by single-point supervision in the early stages of training, along with inaccuracies in pseudo-label updates, and the difficulty in achieving convergence simultaneously, have made it challenging for such methods to attain satisfactory performance. In this article, we introduce a label evolution framework based on local contrast measure (LELCM) to address these issues. Before training, we expand the single-point labels into initial pseudo-labels based on the inherent information of the targets, which mitigates the problem of class imbalance. Furthermore, in the process of updating pseudo-labels, we employ a strategy that utilizes confidence contrast for updates, not only enabling more stable updates of pseudo-labels based on target characteristics but also facilitating adaptive cessation of updates. Our experimental results reveal that our approach not only attains target detection rates (Pd) on par with full supervision models but also achieves 80% of the full supervisory effect in terms of intersection over union (IoU).
Dongning Yang, Haopeng Zhang 0001, Zhiguo Jiang 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Unsupervised Multi-Spectral Image Super-Resolution Based on Conditional Variational Autoencoder
abstract
Unsupervised super-resolution aims to enhance the quality of images without high-resolution (HR) labels during the training stage, making it applicable to real-world scenarios. However, unsupervised super-resolution methods face the challenge of effectively learning the internal structure of images due to the absence of high-quality HR images as references. Moreover, multi-spectral remote sensing images often contain stochastic features caused by cloud and fog occlusions. These occlusions make it difficult to achieve accurate reconstruction of occluded areas through direct modeling of deep features in multi-spectral images. In this paper, we propose a method inspired by conditional variational autoencoders to address the issue of stochastic features in unsupervised multi-spectral super-resolution. Additionally, we introduce a channel attention feature fusion module to combine two types of features. We evaluated our unsupervised multi-spectral image super-resolution method using a real satellite remote sensing dataset. Experimental results demonstrate the qualitative and quantitative effectiveness of our approach.
Zhexin Han, Haopeng Zhang 0001, Zhiguo Jiang 0001
IGARSS4
2023 Rotated Ship Detection Based On Dense Points in High Resolution Remote Sensing Images
abstract
The rapid development of remote sensing technology has provided convenient conditions for obtaining abundant research data. The use of visible light remote sensing images for ship detection has profound significance in the fields of port management, maritime rescue, and military investigation. Our paper focuses on the shortcomings of current way of ship positioning expression and uses deep learning to study the rotated ship detection task. A kind of ship detection algorithm based on dense points related to the position of ships is proposed to solve the problem of angle boundary and vertex sorting ambiguity in current rotated object detection methods. Our method achieves 92.4 mAP on the HRSC2016 dataset, ranking among the top in similar studies.
Haopeng Zhang 0001, Zhiguo Jiang 0001
IGARSS4
2023 Position-Aware Masked Autoencoder for Histopathology WSI Representation Learning
Kun Wu 0010, Yushan Zheng, Jun Shi 0006, Fengying Xie, Zhiguo Jiang 0001
MICCAI (6)5
2023 WSODet: A Weakly Supervised Oriented Detector for Aerial Object Detection
abstract
In contrast to natural objects, aerial targets are usually non-axis aligned with arbitrary orientations. However, mainstream weakly supervised object detection (WSOD) methods can only predict horizontal bounding boxes (HBBs) from existing proposals generated by offline algorithms. To predict oriented bounding boxes (OBBs) for aerial targets while testing images end-to-end without proposals, WSODet is designed leveraging on layerwise relevance propagation (LRP) and point set representation (RepPoints). To be specific, based on the mainstream WSOD framework, LRP on multiple instance learning branch (MIL-LRP) is conducted to decrease the uncertainty and ambiguity of feature map. Then, a pseudo oriented label generation algorithm is designed to obtain OBB pseudolabels, which serve as supervision to train an oriented RepPoint Net under the guidance of improved oriented loss function (IOLF). During the test, input images are sent to oriented RepPoint branch (ORB) to obtain OBB predictions without proposals. Extensive experiments on the detection in optical remote sensing images (DIOR), Northwestern Polytechnical University (NWPU) VHR-10.v2, and HRSC2016 datasets demonstrate the effectiveness of our method to predict precise oriented aerial objects, achieving 22.2%, 46.5%, and 43.3% mAP, respectively. Moreover, training jointly with ORB boosts the results of the original WSOD framework compared with the existing WSOD methods even if there is no specific design for the original structure.
Zhiwen Tan, Zhiguo Jiang 0001, Haopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Kernel Attention Transformer for Histopathology Whole Slide Image Analysis and Assistant Cancer Diagnosis
abstract
Transformer has been widely used in histopathology whole slide image analysis. However, the design of token-wise self-attention and positional embedding strategy in the common Transformer limits its effectiveness and efficiency when applied to gigapixel histopathology images. In this paper, we propose a novel kernel attention Transformer (KAT) for histopathology WSI analysis and assistant cancer diagnosis. The information transmission in KAT is achieved by cross-attention between the patch features and a set of kernels related to the spatial relationship of the patches on the whole slide images. Compared to the common Transformer structure, KAT can extract the hierarchical context information of the local regions of the WSI and provide diversified diagnosis information. Meanwhile, the kernel-based cross-attention paradigm significantly reduces the computational amount. The proposed method was evaluated on three large-scale datasets and was compared with 8 state-of-the-art methods. The experimental results have demonstrated the proposed KAT is effective and efficient in the task of histopathology WSI analysis and is superior to the state-of-the-art methods.
Yushan Zheng, Jun Li 0106, Jun Shi 0006, Fengying Xie, Jianguo Huai, Zhiguo Jiang 0001
IEEE Trans. Medical Imaging7
2022 Histopathology Cross-Modal Retrieval based on Dual-Transformer Network
abstract
Computer-aided cancer diagnosis (CAD) methods based on the histopathological images have achieved great development. The content-based whole slide image (WSI) retrieval is one of the important application that can search for the informative data to assist clinical diagnosis. It is notable that the current retrieval system are mainly developed based on the image content and image labels. The diagnosis report for the WSIs given by the pathologists are also valuable data, but have not yet been adequately considered in modeling. In this paper, we propose a cross-modal retrieval framework based on histopathology WSIs and diagnosis report, which can simultaneously achieve four retrieval tasks for histopathology database across WSIs and diagnosis reports. The compact binary features from both WSIs and diagnosis reports are first extracted, and then built in a common vision-language semantic feature space by the constraint of the designed cross hashing loss function. The method was verified on a gastric histopathology dataset that contains 932 gastric cases with 4 lesion categories. Experimental results have demonstrated the effectiveness of the proposed method in the cross-modal retrieval tasks for digital pathology system.
Dingyi Hu, Fengying Xie, Zhiguo Jiang 0001, Yushan Zheng, Jun Shi 0006
BIBE3
2022 Multi-Frame Super-Resolution With Raw Images Via Modified Deformable Convolution
abstract
In this paper we propose a novel model towards multi-frame super-resolution, which leverages multiple RAW images and yields a super-resolved RGB image. To facilitate the pixel misalignment in burst photography, we apply a refined Pyramid Cascading and Deformable Convolution (PCD) feature alignment module. A new 3D deformable convolution fusion module is proposed subsequently to merge the information from all frames adaptively. In addition, we employ an encoder-decoder network to restore color and details in sRGB space after super-resolving images in linear space. Extensive experiments demonstrate the superiority of our architecture and the strength of multi-frame super-resolution with RAW images.
Gongzhe Li, Linwei Qiu, Haopeng Zhang 0001, Fengying Xie, Zhiguo Jiang 0001
ICASSP5
2022 Lesion-Aware Contrastive Representation Learning for Histopathology Whole Slide Images Analysis
Jun Li 0106, Yushan Zheng, Kun Wu 0010, Jun Shi 0006, Fengying Xie, Zhiguo Jiang 0001
MICCAI (2)6
2022 Kernel Attention Transformer (KAT) for Histopathology Whole Slide Image Classification
Yushan Zheng, Jun Li 0106, Jun Shi 0006, Fengying Xie, Zhiguo Jiang 0001
MICCAI (2)5
2022 Thin Cloud Removal for Remote Sensing Images Using a Physical-Model-Based CycleGAN With Unpaired Data
abstract
Thin cloud removal from remote sensing (RS) images is challenging. Recently, deep-learning-based methods have achieved excellent results using supervised training on paired image data. However, in practice, real paired image data are unavailable. Therefore, in this letter, we propose a novel thin cloud removal method, a physical-model-based CycleGAN (PM-CycleGAN), which can be trained using only unpaired data. The PM-CycleGAN training process comprises forward and backward loops. The forward loop first decomposes a cloudy image into a cloud-free image, thin cloud thickness map, and thickness coefficient using three generators. Then, it combines these three components using a physical model to reconstruct the original cloudy image to obtain the cycle consistency constraint. The backward loop first uses the physical model to synthesize a cloud-free image, thin cloud thickness map, and thickness coefficient into a cloudy image, which are then decomposed into the original three components using the three generators. Visual and quantitative comparisons against several state-of-the-art (SOTA) methods on a cloudy image dataset demonstrated the superiority of PM-CycleGAN.
Yue Zi, Fengying Xie, Xuedong Song, Zhiguo Jiang 0001, Haopeng Zhang 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Encoding histopathology whole slide images with location-aware graphs for diagnostically relevant regions retrieval
Yushan Zheng, Zhiguo Jiang 0001, Jun Shi 0006, Fengying Xie, Haopeng Zhang 0001, Dingyi Hu, Shujiao Sun, Zhongmin Jiang, Chenghai Xue
Medical Image Anal.2
2022 Hyperspectral Image Classification Using Feature Fusion Hypergraph Convolution Neural Network
abstract
Convolution neural networks (CNNs) and graph representation learning are two common methods for hyperspectral image (HSI) classification. Recently, graph convolutional neural networks, a combination of CNN and graph representation learning, have shown great potential in the HSI classification problem. However, the existing graph convolution network (GCN)-based methods have many problems, such as overdependence on the adjacency matrix, usage of a single modal feature, and lower accuracy than the mature CNN method. In this article, we propose a feature fusion hypergraph neural network (F2HNN) for HSI classification. F2HNN first generates hyperedges from features of different modalities to construct a hypergraph representing multimodal features in HSI. Then, the HSI and the extracted hypergraph are input into the hypergraph convolutional neural network for learning. In addition, we propose three feature fusion strategies. The first strategy is the most basic spatial and spectral feature fusion. The second strategy fuses the spectral features extracted by a pretrained multilayer perceptron (MLP) with the spatial features to reduce the redundant information of the original spectral features. The third strategy uses the fusion of CNN features, spectral features, and spatial features to explore the capabilities of F2HNN. Sufficient experiments on four datasets have proved the effectiveness of F2HNN.
Zhongtian Ma, Zhiguo Jiang 0001, Haopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 Frequency-Based Convolutional Neural Network for Efficient Segmentation of Histopathology Whole Slide Images
Yushan Zheng, Dingyi Hu, Jun Li 0106, Chenghai Xue, Zhiguo Jiang 0001
ICIG (2)6
2021 Self-Attention Fusion Module for Single Remote Sensing Image Super-Resolution
abstract
Single image super-resolution (SISR) is an important procedure to improve many remote sensing applications. Global features play an important role in pixel generation of SISR. In this paper, we proposed a self-attention fusion module named as SAF module which combines spatial attention and channel attention in parallel to handle this problem. Our self-attention fusion module can be flexibly added to many popular deep-learning-based SISR models to further improve their representation ability and learn global features. Experiments on UC Merced dataset indicate that SAF module can improve the performance of classic SISR models and achieve state-of-the-art super-resolution results.
Han Mei, Haopeng Zhang 0001, Zhiguo Jiang 0001
IGARSS3
2021 Selective focus saliency model driven by object class-awareness
abstract
Abstract Current many salient object detection (SOD) models only focus on highlighting visual conspicuous region but fail to make saliency detection for specific targets. In this paper, a selective focus saliency model driven by object class‐awareness (SF‐OCA) to run saliency detection is proposed. The framework consists of a visual saliency detection flow, a segmentation‐classification flow, and a class‐awareness selection module. It combines bottom‐up visual perception with a top‐down task‐driven manner, which is capable of detecting specific category salient targets and eliminating the interference from other saliency areas, providing a new idea for saliency detection. Experimental results show that the method achieves comparable performance with state‐of‐the‐art models on four public saliency datasets. In addition, a new dataset was also built to test the proposed framework for the selective focus saliency detection. Compared with other SOD methods, the method not only highlights visual saliency regions but can choose more important or more noteworthy targets in a class‐awareness manner. The method also shows better robustness under a variety of conditions including multi‐targets, small targets and complex background.
Danpei Zhao, Bo Yuan 0009, Zhenwei Shi 0001, Zhiguo Jiang 0001
IET Image Process.4
2021 Nonpairwise-Trained Cycle Convolutional Neural Network for Single Remote Sensing Image Super-Resolution
abstract
Single image super-resolution (SISR) is to recover the high spatial resolution image from a single low spatial resolution one, which is a useful procedure for many remote sensing applications. Most previous convolutional neural network (CNN)-based methods adopt supervised learning. However, paired high-resolution and low-resolution remote sensing images are actually hard to acquire for supervised learning SR methods. To handle this problem, we propose a novel cycle convolutional neural network (Cycle-CNN). Our network consists of two generative CNNs for down-sampling and SR separately and can be trained with unpaired data. We perform comprehensive experiments on panchromatic and multispectral images of the GaoFen-2 satellite and the UC Merced land use data set. Experimental results indicate that our method achieves state-of-the-art CNN-based SR results and is robust against noise and blur in remote sensing images. Comprehensively considering super-resolved image quality and time costs, our proposed method outperforms the compared learning-based SISR approaches.
Haopeng Zhang 0001, Pengrui Wang, Zhiguo Jiang 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Stain Standardization Capsule for Application-Driven Histopathological Image Normalization
abstract
Color consistency is crucial to developing robust deep learning methods for histopathological image analysis. With the increasing application of digital histopathological slides, the deep learning methods are probably developed based on the data from multiple medical centers. This requirement makes it a challenging task to normalize the color variance of histopathological images from different medical centers. In this paper, we propose a novel color standardization module named stain standardization capsule based on the capsule network and the corresponding dynamic routing algorithm. The proposed module can learn and generate uniform stain separation outputs for histopathological images in various color appearance without the reference to manually selected template images. The proposed module is light and can be jointly trained with the application-driven CNN model. The proposed method was validated on three histopathology datasets and a cytology dataset, and was compared with state-of-the-art methods. The experimental results have demonstrated that the SSC module is effective in improving the performance of histopathological image analysis and has achieved the best performance in the compared methods.
Yushan Zheng, Zhiguo Jiang 0001, Haopeng Zhang 0001, Fengying Xie, Dingyi Hu, Shujiao Sun, Jun Shi 0006, Chenghai Xue
IEEE J. Biomed. Health Informatics2
2021 Diagnostic Regions Attention Network (DRA-Net) for Histopathology WSI Recommendation and Retrieval
abstract
The development of whole slide imaging techniques and online digital pathology platforms have accelerated the popularization of telepathology for remote tumor diagnoses. During a diagnosis, the behavior information of the pathologist can be recorded by the platform and then archived with the digital case. The browsing path of the pathologist on the WSI is one of the valuable information in the digital database because the image content within the path is expected to be highly correlated with the diagnosis report of the pathologist. In this article, we proposed a novel approach for computer-assisted cancer diagnosis named session-based histopathology image recommendation (SHIR) based on the browsing paths on WSIs. To achieve the SHIR, we developed a novel diagnostic regions attention network (DRA-Net) to learn the pathology knowledge from the image content associated with the browsing paths. The DRA-Net does not rely on the pixel-level or region-level annotations of pathologists. All the data for training can be automatically collected by the digital pathology platform without interrupting the pathologists' diagnoses. The proposed approaches were evaluated on a gastric dataset containing 983 cases within 5 categories of gastric lesions. The quantitative and qualitative assessments on the dataset have demonstrated the proposed SHIR framework with the novel DRA-Net is effective in recommending diagnostically relevant cases for auxiliary diagnosis. The MRR and MAP for the recommendation are respectively 0.816 and 0.836 on the gastric dataset. The source code of the DRA-Net is available at https://github.com/zhengyushan/dpathnet.
Yushan Zheng, Zhiguo Jiang 0001, Fengying Xie, Jun Shi 0006, Haopeng Zhang 0001, Jianguo Huai, Xiaomiao Yang
IEEE Trans. Medical Imaging2
2020 Tracing Diagnosis Paths on Histopathology WSIs for Diagnostically Relevant Case Recommendation
Yushan Zheng, Zhiguo Jiang 0001, Haopeng Zhang 0001, Fengying Xie, Jun Shi 0006
MICCAI (5)2
2020 Out-of-region keypoint localization for 6D pose estimation
Zhiguo Jiang 0001, Haopeng Zhang 0001
Image Vis. Comput.2
2020 Finding Arbitrary-Oriented Ships From Remote Sensing Images Using Corner Detection
abstract
Ship detection in remote sensing images is a challenging task. In this letter, a novel anchor-free framework is proposed for detecting arbitrary-oriented ships in remote sensing images. First, an end-to-end fully convolutional network is designed to detect the three key points, including the bow, stern, and center of the ship, as well as its angle. Second, the key points of the bow and stern are combined to generate possible rotated bounding boxes. Third, the predicted center and angle information of the ship are used to confirm the bounding box. In the designed network, feature fusion and feature enhancement modules are introduced to improve the performance in complex scenes. The proposed method avoids complicated anchor design compared with anchor-based methods. The experimental results show that with good robustness to haze occlusion, scale variation, and adjacent ship disturbances, our method outperforms other state-of-the-art methods.
Fengying Xie, Yuanyao Lu, Zhiguo Jiang 0001
IEEE Geosci. Remote. Sens. Lett.4
2020 A Local Flatness Based Variational Approach to Retinex
abstract
A topic of continued interest in Retinex over the years has been finding ways to implement it with computational models of improved accuracy and efficiency. We have devised a new approach to digitally implementing the Retinex using a local deviation based variational model. The new model leads to improvements in the computed image quality with respect to illumination correction and image enhancement. Several contributions are made: 1) a new prior constraint, which we call local flatness, is proposed, and a new measure of Local Deviation (LD) is developed to quantify the degree of local illumination flatness; 2) a variational problem is defined and the solution is found by a logical sequence of steps; 3) discrete implementation of the variational solution is shown to effectively estimate and remove uneven illumination, yielding an accurate recovered image. Unlike other physical prior based variational Retinex models, which use the L2 norm of the illumination gradient to enforce smoothness of illumination, our LD prior selectively imposes local flatness on illumination by calculating the deviation between the estimated illumination surface to a reference plane. In the experiments, pseudo ground truth images are created by superimposing uneven illumination on real scenes, providing an effective way to objectively assess algorithm performance. The experimental results show that our method can reconstruct more accurate recovered images than other state-of-the-art methods, while maintaining good contrast.
Fengying Xie, Rui Zhang 0069, Zhiguo Jiang 0001, Alan C. Bovik
IEEE Trans. Image Process.4
2019 Unsupervised Remote Sensing Image Super-Resolution Using Cycle CNN
abstract
Single image super-resolution (SISR) is a useful procedure for many remote sensing applications. However, paired high-resolution and low-resolution remote sensing images are actually hard to acquire for supervised learning SR methods. In this paper, we propose an unsupervised network named Cycle-CNN to handle this problem. Our network consists of two generative CNNs for down-sampling and super-resolution separately, and can be trained with unpaired data. Experiments on panchromatic and multi-spectral images of GaoFen-2 satellite indicate that our method achieves state-of-the-art SR results and is robust against noise and blur in the remote sensing images.
Pengrui Wang, Haopeng Zhang 0001, Zhiguo Jiang 0001
IGARSS4
2019 Real-time 6D pose estimation from a single RGB image
Zhiguo Jiang 0001, Haopeng Zhang 0001
Image Vis. Comput.2
2019 Unsupervised Oil Tank Detection by Shape-Guide Saliency Model
abstract
In this letter, a novel oil tank detection framework based on a shape-guide saliency (SGS) model is proposed. Beyond the low-level visual stimuli, SGS focuses more on simulating the selective visual searching, which is dominated by the goal in human minds. Using a top–down strategy, SGS breaks the limitation of the low-level visual features and introduces the high-level task concept to measure saliency. For the oil tank detection, SGS model skillfully extracts the contour shape cue (CSC) as the target-oriented information and uses CSC to guide the selective saliency value calculation. Specifically, a sparse reconstruction with the target-specific dictionary is implemented to generate the saliency map. This saliency map only assigns high values to oil tank regions instead of highlighting all high-contrast regions. Consequently, SGS model is capable of accurately locating oil tanks and eliminating the interferences of high-contrast backgrounds. Experimental results on a remote sensing data set demonstrate that the proposed SGS model outperforms five class-independent saliency models. Comparisons with the state-of-the-art oil tank detection approaches demonstrate the effectiveness of the proposed method.
Minhao Jing, Danpei Zhao, Yue Gao 0008, Zhiguo Jiang 0001, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.5
2018 Inshore Ship Detection Based on Mask R-CNN
abstract
Inshore ship detection is a popular research domain for optical remote sensing image understanding with many applications in harbor management. However, recent approaches on inshore ship detection depend heavily on hand-crafted features, which need a complicated procedure. In this paper, we propose a new method to achieve inshore ship detection based on Mask R-CNN. We introduce Soft-Non-Maximum Suppression (Soft-NMS) into our framework to improve the robustness to nearby inshore ships. Both battleships and merchantships can be detected in our framework. Furthermore, our framework can also obtain the binary masks of inshore ships. Experimental results on a dataset collected from Google Earth have quantitatively and qualitatively demonstrated the effectiveness of our approach.
Shanlan Nie, Zhiguo Jiang 0001, Haopeng Zhang 0001, Bowen Cai 0001
IGARSS2
2018 Star Image Simulation and Subpixel Centroiding for an Earth Observing Sensor
abstract
In this paper, a novel solution is introduced for accurate subpixel star centroiding of the focused geostationary Earth observing sensor on a three-axis stabilized satellite. A small 2-dimensional array is utilized to better capture the star spot than the linear array of detectors. The popular center of mass method is used to compute star centroid in a single frame. Then the subpixel accuracy of star centroiding can be improved by fitting the linear trajectory of the observed star according to exact imaging time produced from time awarding system of the satellite. Experimental results on simulated star images in various conditions validate the effectiveness and robustness of our star centroiding method.
Haopeng Zhang 0001, Bowen Cai 0001, Zhiguo Jiang 0001
IGARSS4
2018 Online Exemplar-Based Fully Convolutional Network for Aircraft Detection in Remote Sensing Images
abstract
Convolutional neural network obtains remarkable achievements on target detection, due to its prominent capability on feature extraction. However, it still needs further study for aircraft detection task, since intraclass variation still restricts the accuracy of aircraft detection in remote sensing images. In this letter, we adopt regularity of aircraft circle response to design our end-to-end fully convolutional network (FCN), and embed online exemplar mining into our network to handle intraclass variation. The mined exemplars are employed to capture different intraclass characteristics, which effectively reduces the burden of network training. Specifically, we first select basic exemplars based on labeled information and initialize the relationships between exemplars and aircraft examples. Then, these relationships will be updated by the similarity of these examples in high-level features space. Finally, aircraft examples will be used to train different exemplar detectors according to updated relationships. Motivated by the geometric shape of aircraft, a circle response map is developed to construct our FCN to achieve more efficient aircraft detection. The comparative experiments indicate that superior performance of our network in accurate and efficient aircraft detection.
Bowen Cai 0001, Zhiguo Jiang 0001, Haopeng Zhang 0001, Shanlan Nie
IEEE Geosci. Remote. Sens. Lett.2
2018 Size-Scalable Content-Based Histopathological Image Retrieval From Database That Consists of WSIs
abstract
Content-based image retrieval (CBIR) has been widely researched for histopathological images. It is challenging to retrieve contently similar regions from histopathological whole slide images (WSIs) for regions of interest (ROIs) in different size. In this paper, we propose a novel CBIR framework for database that consists of WSIs and size-scalable query ROIs. Each WSI in the database is encoded into a matrix of binary codes. When retrieving, a group of region proposals that have similar size with the query ROI are firstly located in the database through an efficient table-lookup approach. Then, these regions are ranked by a designed multi-binary-code-based similarity measurement. Finally, the top relevant regions and their locations in the WSIs as well as the corresponding diagnostic information are returned to assist pathologists. The effectiveness of the proposed framework is evaluated on a fine-annotated WSI database of epithelial breast tumors. The experimental results have proved that the proposed framework is effective for retrieval from database that consists of WSIs. Specifically, for query ROIs of 4096 4096 pixels, the retrieval precision of the top 20 return has reached 96% and the retrieval time is less than 1.5 s.
Yushan Zheng, Zhiguo Jiang 0001, Haopeng Zhang 0001, Fengying Xie, Yibing Ma, Huaqiang Shi, Yu Zhao 0029
IEEE J. Biomed. Health Informatics2
2018 Histopathological Whole Slide Image Analysis Using Context-Based CBIR
abstract
Histopathological image classification (HIC) and content-based histopathological image retrieval (CBHIR) are two promising applications for the histopathological whole slide image (WSI) analysis. HIC can efficiently predict the type of lesion involved in a histopathological image. In general, HIC can aid pathologists in locating high-risk cancer regions from a WSI by providing a cancerous probability map for the WSI. In contrast, CBHIR was developed to allow searches for regions with similar content for a region of interest (ROI) from a database consisting of historical cases. Sets of cases with similar content are accessible to pathologists, which can provide more valuable references for diagnosis. A drawback of the recent CBHIR framework is that a query ROI needs to be manually selected from a WSI. An automatic CBHIR approach for a WSI-wise analysis needs to be developed. In this paper, we propose a novel aided-diagnosis framework of breast cancer using whole slide images, which shares the advantages of both HIC and CBHIR. In our framework, CBHIR is automatically processed throughout the WSI, based on which a probability map regarding the malignancy of breast tumors is calculated. Through the probability map, the malignant regions in WSIs can be easily recognized. Furthermore, the retrieval results corresponding to each sub-region of the WSIs are recorded during the automatic analysis and are available to pathologists during their diagnosis. Our method was validated on fully annotated WSI data sets of breast tumors. The experimental results certify the effectiveness of the proposed method.
Yushan Zheng, Zhiguo Jiang 0001, Haopeng Zhang 0001, Fengying Xie, Yibing Ma, Huaqiang Shi, Yu Zhao 0029
IEEE Trans. Medical Imaging2
2017 No Reference Assessment of Image Visibility for Dehazing
Manjun Qin, Fengying Xie, Zhiguo Jiang 0001
ICIG (1)3
2017 Kernel estimation for motion blur removal using deep convolutional neural network
abstract
Blind deblurring can restore the sharp image from the blur version when the blur kernel is unknown, which is a challenging task. Kernel estimation is crucial for blind deblurring. In this paper, a novel blur kernel estimation method based on regression model is proposed for motion blur. The motion blur features are firstly mined through convolutional neural network (CNN), and then mapped to motion length and orientation by support vector regression (SVR). Experiments show that the proposed model, namely CNNSVR, can give more accurate kernel estimation and generate better deblurring result compared with other state-of-the-art algorithms.
Yanan Lu, Fengying Xie, Zhiguo Jiang 0001
ICIP3
2017 Training deep convolution neural network with hard example mining for airport detection
abstract
The geometrical characteristic and low-level manually designed features are usually used to detect airports in optical remote sensing images. But it is insufficient to describe airport in low resolution and illumination environment. This paper presents a hard example mining algorithm to train the end-to-end deep convolutional neural network for airport detection in complex situation. Compared with conventional airport detection methods which design specific low-level manually designed features for high-resolution remote sensing images, an end-to-end network can mine the general characteristic among the training samples and learn high-level features in multi-scale and multi-view remote sensing images. Meanwhile, an automatic hard example mining principle is introduced to make training more efficiently and accurately. The proposed method is validated on a multi-scale and multi-view dataset collected from Google Earth. The experimental results demonstrate that the proposed method is robust and efficient, and superior to the state-of-the-art airport detection models.
Bowen Cai 0001, Zhiguo Jiang 0001, Haopeng Zhang 0001
IGARSS2
2017 Region proposal for ship detection based on structured forests edge method
abstract
Remote sensing images are with the characteristics of large width and sparse distribution of specific targets, so that the extraction of region proposal is necessary before detection. In this paper, we propose a new ship detection method on sea-background remote sensing images, which are generally influenced by clouds, waves and other inhomogeneities. Instead of exhaustive search, the core of our method is that the region proposals are obtained from edge detection based on structured forests, which makes our method accurate and efficient. This edge detection method only demands a small training set and then produces contours with the background suppressed. After some morphological processing on the contours, we obtained ship proposals by connected domain detection. Adopting support vector machine(SVM) as classifier, we finally acquire ship detection results. The remote sensing images in our datasets are downloaded from Google Earth map. In our experiments, the proposed method is feasible and effective, and it shows better performance than other methods especially in various illumination and interference conditions.
Zhiguo Jiang 0001, Haopeng Zhang 0001, Bowen Cai 0001
IGARSS2
2017 Chimney and condensing tower detection based on faster R-CNN in high resolution remote sensing images
abstract
The persistent haze weather in North China has aroused extensive attention to environmental protection. Among all pollution resources, the anthropogenic emission by fossil fuel power plants plays an important role. To assist the environmental protection administration monitoring fossil fuel power plants, we propose an effective approach in this paper to learn an integrated model for chimney and condensing tower detection based on Faster R-CNN in high resolution remote sensing images. Our method can detect chimneys and condensing towers under different imaging condition efficiently and accurately. Experimental results on a self-collected dataset demonstrate the effectiveness of the proposed method.
Zhiguo Jiang 0001, Haopeng Zhang 0001, Bowen Cai 0001, Gang Meng, Deshan Zuo
IGARSS2
2017 Orientation judgment for abstract paintings
Weiming Dong, Xiaopeng Zhang 0001, Zhiguo Jiang 0001
Multim. Tools Appl.4
2017 Feature extraction from histopathological images based on nucleus-guided convolutional neural network for breast lesion classification
Yushan Zheng, Zhiguo Jiang 0001, Fengying Xie, Haopeng Zhang 0001, Yibing Ma, Huaqiang Shi, Yu Zhao 0029
Pattern Recognit.2
2017 Breast Histopathological Image Retrieval Based on Latent Dirichlet Allocation
abstract
In the field of pathology, whole slide image (WSI) has become the major carrier of visual and diagnostic information. Content-based image retrieval among WSIs can aid the diagnosis of an unknown pathological image by finding its similar regions in WSIs with diagnostic information. However, the huge size and complex content of WSI pose several challenges for retrieval. In this paper, we propose an unsupervised, accurate, and fast retrieval method for a breast histopathological image. Specifically, the method presents a local statistical feature of nuclei for morphology and distribution of nuclei, and employs the Gabor feature to describe the texture information. The latent Dirichlet allocation model is utilized for high-level semantic mining. Locality-sensitive hashing is used to speed up the search. Experiments on a WSI database with more than 8000 images from 15 types of breast histopathology demonstrate that our method achieves about 0.9 retrieval precision as well as promising efficiency. Based on the proposed framework, we are developing a search engine for an online digital slide browsing and retrieval platform, which can be applied in computer-aided diagnosis, pathology education, and WSI archiving and management.
Yibing Ma, Zhiguo Jiang 0001, Haopeng Zhang 0001, Fengying Xie, Yushan Zheng, Huaqiang Shi, Yu Zhao 0029
IEEE J. Biomed. Health Informatics2
2017 Melanoma Classification on Dermoscopy Images Using a Neural Network Ensemble Model
abstract
We develop a novel method for classifying melanocytic tumors as benign or malignant by the analysis of digital dermoscopy images. The algorithm follows three steps: first, lesions are extracted using a self-generating neural network (SGNN); second, features descriptive of tumor color, texture and border are extracted; and third, lesion objects are classified using a classifier based on a neural network ensemble model. In clinical situations, lesions occur that are too large to be entirely contained within the dermoscopy image. To deal with this difficult presentation, new border features are proposed, which are able to effectively characterize border irregularities on both complete lesions and incomplete lesions. In our model, a network ensemble classifier is designed that combines back propagation (BP) neural networks with fuzzy neural networks to achieve improved performance. Experiments are carried out on two diverse dermoscopy databases that include images of both the xanthous and caucasian races. The results show that classification accuracy is greatly enhanced by the use of the new border features and the proposed classifier model.
Fengying Xie, Haidi Fan, Zhiguo Jiang 0001, Rusong Meng, Alan C. Bovik
IEEE Trans. Medical Imaging4
2016 Sparsity-constrained probabilistic latent semantic analysis for land cover classification
abstract
Land cover classification can be regarded as topic assignment that the pixels can be classified into different kinds of regions (e.g. road, tree, grass) according to the semantics of topics in topic model. In this paper, we present a novel probabilistic latent semantic analysis (pLSA) model based on sparsity constraint for classifying different kinds of land cover. In contrast with conventional topic model which usually assumes each local feature descriptor is only related to one visual word of the dictionary, our method uses sparse coding to characterize the potential relationship between the descriptor and multiple words. Therefore each descriptor can be represented by a small set of words. More importantly, we further apply sparse coding to mine the correlation of documents (i.e. image) in pLSA model. Consequently, our model can generate the more discriminative latent topics and benefit land cover classification. Experimental results on high-resolution remote sensing images demonstrate the excellent superiority of our method.
Jun Shi 0006, Xilan Tian, Zhiguo Jiang 0001, Danpei Zhao
IGARSS3
2016 Semi-supervised Conditional Random Field for hyperspectral remote sensing image classification
abstract
Conditional Random Field(CRF) has been successfully applied to the hyperspectral image classification. However, it suffers from the availability of large amount of labeled pixels, which is labor- and time-consuming to obtain in practice. In this paper, a semi-supervised CRF(ssCRF) is proposed for hyperspectral image classification with limited labeled pixels. Laplacian Support Vector Machine(LapSVM), after extended into the composite kernel type, is defined as the association potential. And the Potts model is utilized as the interaction potential. The ssCRF is evaluated on the two benchmarks and the results show the effectiveness of ssCRF.
Zhiguo Jiang 0001, Haopeng Zhang 0001, Bowen Cai 0001, Quanmao Wei
IGARSS2
2016 High-resolution optical satellite image simulation of ship target in large sea scenes
abstract
Ship target detection in optical remote sensing images has attracted more and more attention in the field of remote sensing. The ship target detection technology of optical remote sensing images is vulnerable to many factors, while the real data are difficult to contain various elements. In order to obtain the various situations in the large sea scenes, we develop a simulation system for high-resolution optical remote sensing image of ship targets. The simulated images with different sea states, cloud conditions, target types and imaging conditions can support the evaluation and comparison of ship detection algorithms as well as other tasks in remote sensing image analysis.
Zhiguo Jiang 0001, Haopeng Zhang 0001
IGARSS2
2016 Hierarchical reinforcement learning for saliency detection of low-resolution airports
abstract
The traditional airport detection methods usually utilize geometric characteristics, which are limited by large amount of data and low-resolution of the remote sensing images. In this paper, we present a novel hierarchical reinforcement learning (HRL) saliency model for quickly airports detecting in large cover area. In contrast with conventional saliency models which usually are effective for high-resolution nature images, our method learns hierarchically high-level features via multi-scale superpixels segmentation and Least Absolute Shrinkage and Selection Operator (LASSO). More importantly, we introduce back-propagation theory for hierarchical learning to adaptively control and generate saliency map. Therefore our unsupervised saliency model is more simple and effective for low-resolution airport detection. Compared with 18 state-of-the-art saliency models, experimental results demonstrate the excellent performance of our method on the remote sensing image datasets. It is more robust and accurate for long-range airports detection.
Danpei Zhao, Zhiguo Jiang 0001
IGARSS4
2016 A DAISY descriptor based multi-view stereo method for large-scale scenes
Bindang Xue, Donghai Han, Xiangzhi Bai, Fugen Zhou, Zhiguo Jiang 0001
J. Vis. Commun. Image Represent.6
2016 No-Reference Assessment on Haze for Remote-Sensing Images
abstract
Assessment on haze can filter out images with dense haze to improve the reliability of remote-sensing image interpretation. In this letter, a novel no-reference haze assessment method based on haze distribution is proposed for remote-sensing images. First, range channel of an image is defined and the haze distribution map (HDM) is extracted from the hazy image. Then, the haze assessment metric HDM-based haze assessment (HDMHA) is designed according to the HDM. Finally, the degree of haze in remote-sensing images is predicted using the proposed metric. In order to objectively verify the effectiveness of the proposed metric HDMHA, a method of simulating hazy remote-sensing images based on the haze imaging model is proposed in this letter, and the simulated hazy images are greatly similar to real ones in vision. A series of experiments are done on both real images and simulated images, and the results show that the proposed metric achieves good consistency when compared with subjective experiments and outperforms typical blind image quality assessment methods.
Xiaoxi Pan, Fengying Xie, Zhiguo Jiang 0001, Zhenwei Shi 0001, Xiaoyan Luo
IEEE Geosci. Remote. Sens. Lett.3
2015 Pattern Classification for Dermoscopic Images Based on Structure Textons and Bag-of-Features Model
Fengying Xie, Zhiguo Jiang 0001, Rusong Meng
ICIG (3)3
2015 Factorization of view-object manifolds for joint object recognition and pose estimation
Haopeng Zhang 0001, Tarek El-Gaaly, Ahmed M. Elgammal, Zhiguo Jiang 0001
Comput. Vis. Image Underst.4
2015 Adaptive segmentation based on multi-classification model for dermoscopy images
Fengying Xie, Yefen Wu, Zhiguo Jiang 0001, Rusong Meng
Frontiers Comput. Sci.4
2015 No Reference Quality Assessment for Multiply-Distorted Images Based on an Improved Bag-of-Words Model
abstract
Multiple distortion assessment is a big challenge in image quality assessment (IQA). In this letter, a no reference IQA model for multiply-distorted images is proposed. The features, which are sensitive to each distortion type even in the presence of other distortions, are first selected from three kinds of NSS features. An improved Bag-of-Words (BoW) model is then applied to encode the selected features. Lastly, a simple yet effective linear combination is used to map the image features to the quality score. The combination weights are obtained through lasso regression. A series of experiments show that the feature selection strategy and the improved BoW model are effective in improving the accuracy of quality prediction for multiple distortion IQA. Compared with other algorithms, the proposed method delivers the best result for multiple distortion IQA.
Yanan Lu, Fengying Xie, Tongliang Liu, Zhiguo Jiang 0001, Dacheng Tao
IEEE Signal Process. Lett.4
2015 No Reference Uneven Illumination Assessment for Dermoscopy Images
abstract
For the dermoscopy image, uneven illumination will influence segmentation accuracy and lead to wrong aided diagnosis result. In this paper, a no reference uneven illumination assessment metric is proposed for dermoscopy images. Firstly, the distorted image is decomposed to illumination and reflectance components through variational framework for Retinex (VFR). Then, the illumination component is extracted by basis function fitting. Lastly, average gradient of the illumination component (AGIC) is calculated as the uneven illumination metric. A series of experiments show that, the proposed illumination extraction method is insensitive to the image content, and the proposed metric delivers an accurate illumination assessment result.
Yanan Lu, Fengying Xie, Yefen Wu, Zhiguo Jiang 0001, Rusong Meng
IEEE Signal Process. Lett.4
2015 Haze Removal for a Single Remote Sensing Image Based on Deformed Haze Imaging Model
abstract
The contrast of remote sensing images captured in haze condition is poor, which influences their interpretation. In this letter, a novel dehazing algorithm based on the deformed haze imaging model is proposed. First, the model is deformed by introducing a translation term. Second, the atmospheric light and transmission are estimated according to the new model combined with dark channel prior. Lastly, the haze is successfully removed from remote sensing images using the proposed estimation algorithm. The estimated transmission is insensitive to the texture of ground objects, and the dehazing effect for nonuniform haze is more satisfactory than the compared method. Moreover, our approach can be used for general haze removal through adjusting the translation term. Experimental results reveal that the proposed method can recover the real scene clearly from haze remote sensing images along with the advantage of good color consistency.
Xiaoxi Pan, Fengying Xie, Zhiguo Jiang 0001, Jihao Yin
IEEE Signal Process. Lett.3
2014 Retrieval of pathology image for breast cancer using PLSA model based on texture and pathological features
abstract
Content-based image retrieval (CBIR) for digital pathology slides is of clinical use for breast cancer aided diagnosis. One of the largest challenges in CBIR is feature extraction. In this paper, we propose a novel pathology image retrieval method for breast cancer, which aims to characterize the pathology image content through texture and pathological features and further discover the latent high-level semantics. Specifically, the proposed method utilizes block Gabor features to describe the texture structure, and simultaneously designs nucleus-based pathological features to describe morphological characteristics of nuclei. Based on these two kinds of local feature descriptors, two codebooks are built to learn the probabilistic latent semantic analysis (pLSA) models. Consequently, each image is represented by the topics of pLSA models which can reveal the semantic concepts. Experimental results on the digital pathology image database for breast cancer demonstrate the feasibility and effectiveness of our method.
Yushan Zheng, Zhiguo Jiang 0001, Jun Shi 0006, Yibing Ma
ICIP2
2014 Shadow detection in remote sensing images based on weighted edge gradient ratio
abstract
This paper presents a novel shadow detection method in remote sensing images based on edge feature description of candidate regions. Edge gradient ratio is defined and used to represent the inherent properties of shadow regions. To improve the detection result, weighted edge gradient ratio (WEGR) is addressed, where the weight of a region is determined by the number of pixels belonging to shadow in the region, and edge gradient ratio is proposed to describe the edge feature surrounding the region. Experiments and comparisons indicate that our method achieves better accuracy both on high and low quality remote sensing images.
Bin Pan, Zhiguo Jiang 0001, Xiaoyan Luo
IGARSS3
2014 Parts-probability-based vehicle detection
Zhiguo Jiang 0001
Sci. China Inf. Sci.2
2014 Adaptive Graph Embedding Discriminant Projections
Jun Shi 0006, Zhiguo Jiang 0001
Neural Process. Lett.2
2014 Subspace Matching Pursuit for Sparse Unmixing of Hyperspectral Data
abstract
Sparse unmixing assumes that each mixed pixel in the hyperspectral image can be expressed as a linear combination of only a few spectra (endmembers) in a spectral library, known a priori. It then aims at estimating the fractional abundances of these endmembers in the scene. Unfortunately, because of the usually high correlation of the spectral library, the sparse unmixing problem still remains a great challenge. Moreover, most related work focuses on the l1convex relaxation methods, and little attention has been paid to the use of simultaneous sparse representation via greedy algorithms (GAs) (SGA) for sparse unmixing. SGA has advantages such as that it can get an approximate solution for the l0problem directly without smoothing the penalty term in a low computational complexity as well as exploit the spatial information of the hyperspectral data. Thus, it is necessary to explore the potential of using such algorithms for sparse unmixing. Inspired by the existing SGA methods, this paper presents a novel GA termed subspace matching pursuit (SMP) for sparse unmixing of hyperspectral data. SMP makes use of the low-degree mixed pixels in the hyperspectral image to iteratively find a subspace to reconstruct the hyperspectral data. It is proved that, under certain conditions, SMP can recover the optimal endmembers from the spectral library. Moreover, SMP can serve as a dictionary pruning algorithm. Thus, it can boost other sparse unmixing algorithms, making them more accurate and time efficient. Experimental results on both synthetic and real data demonstrate the efficacy of the proposed algorithm.
Zhenwei Shi 0001, Wei Tang 0016, Zhana Duren, Zhiguo Jiang 0001
IEEE Trans. Geosci. Remote. Sens.4
2014 Ship Detection in High-Resolution Optical Imagery Based on Anomaly Detector and Local Shape Feature
abstract
Ship detection in high-resolution optical imagery is a challenging task due to the variable appearances of ships and background. This paper aims at further investigating this problem and presents an approach to detect ships in a “coarse-to-fine” manner. First, to increase the separability between ships and background, we concentrate on the pixels in the vicinities of ships. We rearrange the spatially adjacent pixels into a vector, transforming the panchromatic image into a “fake” hyperspectral form. Through this procedure, each produced vector is endowed with some contextual information, which amplifies the separability between ships and background. Afterward, for the “fake” hyperspectral image, a hyperspectral algorithm is applied to extract ship candidates preliminarily and quickly by regarding ships as anomalies. Finally, to validate real ships out of ship candidates, an extra feature is provided with histograms of oriented gradients (HOGs) to generate a hypothesis using AdaBoost algorithm. This extra feature focuses on the gray values rather than the gradients of an image and includes some information generated by very near but not closely adjacent pixels, which can reinforce HOG to some degree. Experimental results on real database indicate that the hyperspectral algorithm is robust, even for the ships with low contrast. In addition, in terms of the shape of ships, the extended HOG feature turns out to be better than HOG itself as well as some other features such as local binary pattern.
Zhenwei Shi 0001, Xinran Yu, Zhiguo Jiang 0001
IEEE Trans. Geosci. Remote. Sens.3
2013 Joint Object and Pose Recognition Using Homeomorphic Manifold Analysis
abstract
Object recognition is a key precursory challenge in the fields of object manipulation and robotic/AI visual reasoning in general. Recognizing object categories, particular instances of objects and viewpoints/poses of objects are three critical subproblems robots must solve in order to accurately grasp/manipulate objects and reason about their environ- ments. Multi-view images of the same object lie on intrinsic low-dimensional manifolds in descriptor spaces (e.g. visual/depth descriptor spaces). These object manifolds share the same topology despite being geometrically different. Each object manifold can be represented as a deformed version of a unified manifold. The object manifolds can thus be parametrized by its homeomorphic mapping/reconstruction from the unified manifold. In this work, we construct a manifold descriptor from this mapping between homeomorphic manifolds and use it to jointly solve the three challenging recognition sub-problems. We extensively experiment on a challenging multi-modal (i.e. RGBD) dataset and other object pose datasets and achieve state-of-the-art results.
Haopeng Zhang 0001, Tarek El-Gaaly, Ahmed M. Elgammal, Zhiguo Jiang 0001
AAAI4
2013 Local and Non-local Graph Regularized Sparse Coding for Face Recognition
abstract
The recent emerging sparse coding (SC) algorithms do not take local manifold structure of samples into consideration, while graph regularized sparse coding (GraphSC) algorithm only constrains the locality consistency of samples. Furthermore, the graph construction approach based on k-nearest-neighbor usually pre-defines the number of neighbors for all the samples, which may fails to fit the intrinsic structure of each sample. To address these issues, we propose an local and nonlocal graph regularized sparse coding (LN-GraphSC) algorithm. LN-GraphSC incorporates both local and nonlocal information of samples at the same time. On the other hand, to alleviate the problem of neighbor parameter selection, we use average distance of each sample to wisely determine its own local and nonlocal samples. To verify the effectiveness of our proposed method, we evaluate our method on the task of face recognition. The experimental results on ORL and Yale face databases show our method has competitive performance when compared to SC and GraphSC.
Danpei Zhao, Jun Shi 0006, Zhiguo Jiang 0001
ICIG4
2013 Pathological Image Retrieval for Breast Cancer with pLSA Model
abstract
Pathological image retrieval contributes to computer-aided diagnosis for breast cancer due to the fact that the retrieval results generally contain detailed diagnostic information (e.g. abnormal regions and diagnostic opinion from other doctors) which can offer some reference and assistance to the doctor during diagnosis process. In this paper, we present a novel pathological image retrieval approach based on probabilistic latent semantic analysis (pLSA) model. The method respectively utilizes SIFT features after visual saliency detection, and block Gabor features for the construction of two semantic codebooks, which not only can characterize the salient local invariant features and texture information under different scales and orientations in the pathological images, but also consider the high-level semantic features. Furthermore, we apply pLSA model to discover the latent topics in each codebook. Finally each pathological image is represented by the combination of topics from these two codebooks. The proposed method is evaluated on the pathological image database for breast cancer, which includes 5 categories (mucinous cystadenocarcinoma, invasive lobular carcinoma, basal-like carcinoma, invasive breast cancer and low-grade adenosquamous carcinoma) and 110 cases for each category. Experimental results demonstrate the feasibility and effectiveness of our method.
Jun Shi 0006, Yibing Ma, Zhiguo Jiang 0001, Yu Zhao 0029
ICIG3
2013 Automatic Skin Lesion Segmentation Based on Supervised Learning
abstract
The accuracy of automatic skin lesion detection is important in the computer-aided diagnosis (CAD) of skin cancers. In this paper, a novel method of automatic skin lesion segmentation to get the accurate border is proposed. The initial lesion is extracted by the Otsu's threshold firstly. Secondly, the outer peripheral region around the initial lesion is obtained with the affinity propagation clustering method (AP). The outer periphery is divided into small homogeneous sub-regions using simple linear iterative clustering (SLIC). Finally, the homogeneous sub-regions are classified into the background skin and lesion by supervised learning and the accuracy border is obtained. A series of experiments done on the proposed method and the other four state-of-the-art automatic methods show that the proposed method delivers better accuracy and robust segmentation results.
Yefen Wu, Fengying Xie, Zhiguo Jiang 0001, Rusong Meng
ICIG3
2013 Optical Image Simulation System for Space Surveillance
abstract
Acquiring optical images is a basic task of space based surveillance system. However, these images are hard to obtain in real space condition or are classified for security reason. To solve this problem, a simulation system, based on STK and OpenGL, is established. Using data generated by STK, 3D models of satellites and a star catalogue, we can render the celestial background and space object in OpenGL. Then some post-processing is done to the acquired images to make the final results look real. The simulation results, with high reality and fidelity, can provide data support for space information processing technology, such as object detection, tracking, recognition or for the space surveillance system design.
Zhiguo Jiang 0001, Haopeng Zhang 0001, Jianwei Luo
ICIG2
2013 Sparse coding-based topic model for remote sensing image segmentation
abstract
Land cover segmentation can be viewed as topic assignment that the pixels are grouped into homogeneous regions according to different semantic topics in topic model. In this paper, we propose a novel topic model based on sparse coding for segmenting different kinds of land covers. Different from conventional topic models which generally assume each local feature descriptor is related to only one visual word of the codebook, our method utilizes sparse coding to characterize the potential correlation between the descriptor and multiple words. Therefore each descriptor can be represented by a small set of words. Furthermore, in this paper probabilistic Latent Semantic Analysis (pLSA) is applied to learn the latent relation among word, topic and document due to its simplicity and low computational cost. Experimental results on remote sensing image segmentation demonstrate the excellent superiority of our method over k-means clustering and conventional pLSA model.
Jun Shi 0006, Zhiguo Jiang 0001, Yibing Ma
IGARSS2
2013 A quasi-Newton-based spatial multiple materials detector for hyperspectral imagery
Zhenwei Shi 0001, Zhiguo Jiang 0001
Neural Comput. Appl.3
2013 Efficient sparse unmixing analysis for hyperspectral imagery based on random projection
Zhenwei Shi 0001, Xinya Zhai, Zhiguo Jiang 0001
Neural Comput. Appl.4
2012 An efficient hardware architecture of the optimised SIFT descriptor generation
abstract
Scale Invariant Feature Transform (SIFT) algorithm has the potential of detecting a large number of features from images, which makes the feature descriptor generation become a bottleneck of the processing speed and hence degrade the overall performance of the algorithm. To tackle this problem, we propose an efficient hardware architecture based on the polar sampled descriptor in this paper. It takes only 7.57us to generate a feature descriptor of 72 dimensions with a system frequency of 100MHz, which is equivalent to approximately 132100 feature descriptors per second. It can generate feature descriptors for VGA (640×480 pixels) resolution video at 60 frames per second (fps), provided that there are no more than 2200 features per frame. As far as we know, our hardware architecture has the highest processing speed for descriptor generation, compared with other existing architectures.
Wenjuan Deng, Yiqun Zhu, Zhiguo Jiang 0001
FPL4
2012 SIFT-based Elastic sparse coding for image retrieval
abstract
Bag-of-features (BoF) model based on SIFT generally assumes each descriptor is related to only one visual word of the codebook. Therefore, the potential correlation between the descriptor and other visual words is ignored. On the other hand, sparse coding through l1-norm regularization fails to generate optimal sparse representations since l1-norm regularization randomly selected one variable from a group of highly correlated variables. In this study we propose a novel bag-of-features model for image retrieval called SIFT-based Elastic sparse coding. The method utilizes a large number of SIFT descriptors to construct the codebook. The Elastic Net regression framework, which combines both l1-norm and l2-norm penalties, is then used to obtain the sparse-coefficient vector corresponding to the SIFT descriptor. Finally each image can be represented by a unified sparse-coefficient vector. Experimental results on Coil20 dataset demonstrate the consistent superiority of the proposed method over the state-of-the-art algorithms including original SIFT matching, conventional BoF strategy and BoF model based on l1-norm sparse coding.
Jun Shi 0006, Zhiguo Jiang 0001
ICIP2
2012 Roof-top detection based on structural elements combination
abstract
A novel roof-top extraction method for satellite images based on probabilistic topic model is presented. We model roof-top as the connected structural elements. The proposed method contains two major steps: 1) Detect structural elements, different from earlier structure detector, the proposed method automatically learn the types of elements from unlabeled samples; 2) Connect these elements to form roof-top boundary, where the relationships between elements are estimated by hierarchical topic model. This approach belongs to generative method where only a small number of roof-top samples are required. The experimental results demonstrate the effectiveness of the proposed approach.
Zhiguo Jiang 0001, Jihao Yin
IGARSS2
2012 Easy modeling of realistic trees from freehand sketches
Zhiguo Jiang 0001, Hongjun Li 0002, Xiaopeng Zhang 0001
Frontiers Comput. Sci.2
2011 Image segmentation with hierarchical topic assignment
abstract
Image segmentation can be viewed as the problem of topic assignment, in which pixels are grouped into regions with respect to the topic model of textures and real-world semantic meanings. In this paper, we apply a novel hierarchical topic model, which builds a multi-level image representation covers texture and semantic region, to image segmentation. Specifically, the proposed method firstly assigns texture-topic to each pixel according to the low-level model defines on visual word and neighboring constraint. Furthermore, inspired by the connection between image patch and linguistic sentence, we model semantic segments of natural scene as the combinations of texture topics. Finally, the segmentation result is achieved by finding homogeneous regions in topic field. Experimental results on natural scene images demonstrate the effectiveness of our method.
Zhiguo Jiang 0001
ICIP2
2011 Texture segmentation for remote sensing image based on texture-topic model
abstract
Textures of land covers provide significant evidences for segmentation and classification. Inspired by resent researches on topic model, we work on a novel texture segmentation method for very high resolution (VHR) remote sensing images based on Latent Dirichlet Allocation (LDA). In order to model spatial relationship between words in LDA, a constraint random variable which is used to control the selection of neighboring features of each specific texture is introduced to the model. The proposed method is evaluated on segmenting remote sensing images by finding the homogeneous regions in texture-topic map. The experimental results show our method has great potential for remote sensing image segmentation.
Zhiguo Jiang 0001, Xingmin Han
IGARSS2
2011 A Hierarchical Connection Graph Algorithm for Gable-Roof Detection in Aerial Image
abstract
In this letter, we present a hierarchical connection graph (HCG) algorithm based on a self-avoiding polygon (SAP) model for detecting and extracting gable roofs from aerial imagery. The SAP model is a deformable shape model that is capable of representing gable roofs of various shapes and appearances. The model is composed of a sequence of roof-corner templates that are connected into a SAP, which serves as a flexible shape prior. An energy function that combines features from three channels (corner, boundary, and interior area) is defined over the sequence to quantify the variability in appearances of gable roofs. To infer the most probable state of the corner sequence for an input image, we use an efficient algorithm-called HCG algorithm. The algorithm converts the solution space of a SAP model into a directed graph (which we call “HCG”) and searches for the best path using dynamic programming (DP). It is efficient for two reasons: 1) By constructing an HCG, the algorithm can quickly prune out a large amount of invalid solutions using only geometric constraints, which are inexpensive to compute, and 2) by employing DP, the algorithm decomposes the searching problem into smaller overlapping subproblems and reuses energy scores, which are expensive to compute. Experimental results on a set of challenging gable roofs show that our algorithm has good performance and is computationally effective.
Qiongchen Wang, Zhiguo Jiang 0001, Junli Yang, Danpei Zhao, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.2
2011 Blind Source Separation Using Quadratic form Innovation
Zhenwei Shi 0001, Hongjuan Zhang, Xueyan Tan, Zhiguo Jiang 0001
Neural Process. Lett.4
2009 Gable Roof Description by Self-Avoiding Polygon
Qiongchen Wang, Zhiguo Jiang 0001, Junli Yang, Danpei Zhao, Zhenwei Shi 0001
ACCV (3)2
2009 An architecture of optimised SIFT feature detection for an FPGA implementation of an image matcher
abstract
This paper has proposed an architecture of optimised SIFT (scale invariant feature transform) feature detection for an FPGA implementation of an image matcher. In order for SIFT based image matcher to be implemented on an FPGA efficiently, in terms of speed and hardware resource usage, the original SIFT algorithm has been significantly optimised in the following aspects: 1) upsampling has been replaced with downsampling to save the interpolation operation. 2) Only four scales with two octaves are needed for our image matcher with moderate degradation of matching performance. 3) The total dimension of the feature descriptor has been reduced to 72 from 128 of the original SIFT, which leads to significantly simplify the image matching operation. With the optimisation above, the proposed FPGA implementation is able to detect the features of a typical image of 640 × 480 pixels within 31 milliseconds. Therefore, compared with the existing SIFT FPGA implementation, which requires 33 milliseconds for an image of 320 × 240 pixels, a significant improvement has been achieved for our proposed architecture.
Lifan Yao, Yiqun Zhu, Zhiguo Jiang 0001, Danpei Zhao, Wenquan Feng
FPT4
2009 Maneuvering Target Tracking in Cluttered Background Based on Color Invariance and Support Vector Machine
abstract
Maneuvering targets tracking in cluttered environment is a challenging problem in computer vision because of the difficulty of distinguishing the target from the background. In this paper, we treat tracking as a binary classification problem and employ support vector machine to suppress the background. In order to enhance the robustness against illumination changes, we propose to combine color invariance with traditional RGB values to train the SVM. First, we use expectation maximization algorithm to extract the target from the environment; then, RGB and color invariance values are used to train SVM. In the incoming frames, pixels in regions of interest are classified by SVM and the confidence map is produced, which will afterward be used by traditional tracking approach to track the target, in this paper, we employ particle filter. Experimental results on challenging sequences validate the effectiveness of the proposed method in cluttered background target tracking.
Gang Meng, Zhiguo Jiang 0001, Danpei Zhao, Yue Gao 0008
ICIG2
2009 A Grammatical Framework for Building Rooftop Extraction
abstract
Roof detection has been studied for several decades, one of the big challenge is its structure and appearance diversity. In this paper, we present a grammatical framework to account for these diversities and a multiple way compositional algorithm to extract rooftops from aerial images. We represent rooftops by a context sensitive graph grammar consisting of 5 production rules and 3 types of commonly shared quadrilateral primitives. Each production rule includes a number of equations that constrain the attributes of a parent node and those of its children. In addition, a set of horizontal links are defined between peer nodes at the same level that account for spatial/appearance constraints. The graph grammar produces a large number of valid configurations and can be used to represent the wide structural variability of rooftops. Our rooftop extraction algorithm starts with a lower level bottom-up step that generates hypothesis of quadrilateral primitives by grouping edgelets hierarchically into bigger structures (straight lines, parallel lines and junctions). The grouping process repeats multiple times following alternative paths to reduce missing detections caused by partial occlusion and/or background clutter. Then the higher level relations served as context for lower level elements evaluates each hypothesis according to the graph grammar model and prunes incompatible ones to arrive at an optimal solution.
Qiongchen Wang, Zhiguo Jiang 0001
IGARSS (3)2
2004 A wavelet based algorithm for multi-focus micro-image fusion
abstract
In this paper, extending the depth of focus from microscopy images is investigated by using multi-focus image fusion techniques. The principle of image fusion based on depth of focus is analyzed, and comparisons of different spatial domain fusion methods are given. Furthermore, a multi-focus micro-image fusion algorithm based on area wavelet transform, which offers improved performance over the existing algorithms, is studied and presented. The performance of the spatial domain and wavelet-based algorithms is compared. Advantages and disadvantages of this algorithm are analyzed by experiments, and the main features of this algorithm are summarized and applicable areas are suggested.
Zhiguo Jiang 0001, Dong-bing Han, Xiao-kuan Zhou
ICIG1