Jihao Yin

dblp:67/11238 · DBLP profile ↗
← Back
68ranked-venue papers
17as first author
28since 2021 · last 2026
0000-0002-6773-0193ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 43 · 11 first-author · 18 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 5 since 2021Computer networks · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 STF-Net: A unified spatio-temporal-frequency learning framework for robust small-target detection from event streams
Xiaofei Yin, Yu Zhang 0026, Yuxi Guo, Jihao Yin
Neurocomputing5
2025 DefMamba: Deformable Visual State Space Model
abstract
Recently, state space models (SSM), particularly Mamba, have attracted significant attention from scholars due to their ability to effectively balance computational efficiency and performance. However, most existing visual Mamba methods flatten images into 1D sequences using predefined scan orders, which results the model being less capable of utilizing the spatial structural information of the image during the feature extraction process. To address this issue, we proposed a novel visual foundation model called Def-Mamba. This model includes a multi-scale backbone structure and deformable mamba (DM) blocks, which dynamically adjust the scanning path to prioritize important information, thus enhancing the capture and processing of relevant input features. By combining a deformable scanning (DS) strategy, this model significantly improves its ability to learn image structures and detects changes in object details. Numerous experiments have shown that Def-Mamba achieves state-of-the-art performance in various visual tasks, including image classification, object detection, instance segmentation, and semantic segmentation. The code is open source on DefMamba .
Leiye Liu, Miao Zhang 0004, Jihao Yin, Tingwei Liu, Wei Ji 0011, Yongri Piao, Huchuan Lu
CVPR3
2025 Distant-to-Close Novel View Synthesis for Asteroid Surface Imaging
abstract
Predictively synthesizing high-quality, close-range asteroid surface views from distant optical remote sensing imagery is critical for mission planning and landing-site selection in asteroid exploration missions. However, distant observations inherently lack sufficient resolution and surface detail, limiting existing novel view synthesis methods. To address this, we introduce, to the best of our knowledge, the first framework for distant-to-close novel view synthesis, tailored for asteroid surface imaging. Our method features two key innovations. First, a 3D Gaussian Splatting (3D-GS) super-resolution module applies 2D super-resolution to generate high-resolution virtual close-range views from distant images, enriching the 3D scene model with finer details. Second, an entropy-driven residual refinement strategy adaptively emphasizes structurally complex regions by assigning higher loss weights based on residual image entropy. This strategy triggers targeted subdivisions of 3D Gaussians in areas of high structural complexity. Experiments conducted on datasets from Hayabusa (Itokawa), Dawn (Vesta), Rosetta (67P/Churyumov-Gerasimenko), Hayabusa2 (Ryugu), and OSIRIS-REx (Bennu) missions demonstrate substantial improvements over baseline methods in quantitative metrics such as PSNR, SSIM, and LPIPS.
Linyan Cui, Gangzheng Ai, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.5
2025 Fast radiance field reconstruction from sparse inputs
Song Lai 0001, Linyan Cui, Jihao Yin
Pattern Recognit.3
2025 3-D Gaussian Splatting for Lunar Surface With Limited View Input
abstract
Accurate novel view synthesis of the lunar surface is critical for enhancing remote rover operation efficiency and supporting the creation of virtual environments for lunar exploration. Given the limited images captured by lunar rover, we propose Moon-GS, a novel framework designed specifically for challenging lunar scenarios based on 3D Gaussian Splatting (3DGS). The core methodology involves initializing dense Gaussian primitives using pixel-wise feature matching to address sparse and incomplete point cloud issues, augmenting the training with virtual view constraints to propagate limited ground truth information across multiple perspectives, and employing geometric regularization strategies to ensure consistent 3D representation. These innovations enable Moon-GS to model complex terrain with high fidelity, improve geometric consistency, and enhance rendering quality across unseen viewpoints. Experimental results on real and synthetic lunar datasets demonstrate its superiority over state-of-the-art methods in terms of rendering detail and structural integrity. This research offers a robust solution for lunar 3D visualization, advancing the capabilities of remote rover operations and future exploration missions.
Linyan Cui, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.4
2025 Orthogonal Subspace Representation for Generative Adversarial Networks
abstract
Disentanglement learning aims to separate explanatory factors of variation so that different attributes of the data can be well characterized and isolated, which promotes efficient inference for downstream tasks. Mainstream disentanglement approaches based on generative adversarial networks (GANs) learn interpretable data representation. However, most typical GAN-based works lack the discussion of the latent subspace, causing insufficient consideration of the variation of independent factors. Although some recent research analyzes the latent space on pretrained GANs for image editing, they do not emphasize learning representation directly from the subspace perspective. Appropriate subspace properties could facilitate corresponding feature representation learning to satisfy the independent variation requirements of the obtained explanatory factors, which is crucial for better disentanglement. In this work, we propose a unified framework for ensuring disentanglement, which fully investigates latent subspace learning (SL) in GAN. The novel GAN-based architecture explores orthogonal subspace representation (OSR) on vanilla GAN, named OSRGAN. To guide a subspace with strong correlation, less redundancy, and robust distinguishability, our OSR includes three stages, self-latent-aware, orthogonal subspace-aware, and structure representation-aware, respectively. First, the self-latent-aware stage promotes the latent subspace strongly correlated with the data space to discover interpretable factors, but with poor independence of variation. Second, the following orthogonal subspace-aware stage adaptively learns some 1-D linear subspace spanned by a set of orthogonal bases in the latent space. There is less redundancy between them, expressing the corresponding independence. Third, the structure representation-aware stage aligns the projection on the orthogonal subspace and the latent variables. Accordingly, feature representation in each linear subspace can be distinguishable, enhancing the independent expression of interpretable factors. In addition, we design an alternating optimization step, achieving a tradeoff training of OSRGAN on different properties. Despite it strictly constrains orthogonality, the loss weight coefficient of distinguishability induced by orthogonality could be adjusted and balanced with correlation constraint. To elucidate, this tradeoff training prevents our OSRGAN from overemphasizing any property and damaging the expressiveness of the feature representation. It takes into account both interpretable factors and their independent variation characteristics. Meanwhile, alternating optimization could keep the cost and efficiency of forward inference unchanged and will not burden the computational complexity. In theory, we clarify the significance of OSR, which brings better independence of factors, along with interpretability as correlation could converge to a high range faster. Moreover, through the convergence behavior analysis, including the objective functions under different constraints and the evaluation curve with iterations, our model demonstrates enhanced stability and definitely converges toward a higher peak for disentanglement. To depict the performance in downstream tasks, we compared the state-of-the-art GAN-based and even VAE-based approaches on different datasets. Our OSRGAN achieves higher disentanglement scores on FactorVAE, SAP, MIG, and VP metrics. All the experimental results illustrate that our novel GAN-based framework has considerable advantages on disentanglement.
Hongxiang Jiang, Xiaoyan Luo, Jihao Yin, Huazhu Fu, Fuxiang Wang
IEEE Trans. Neural Networks Learn. Syst.3
2025 Learning Discriminative Representation for Co-Salient Object Detection
abstract
Co-salient object detection (CoSOD) is the task of identifying and emphasizing the common salient objects in a collection of images. The current co-salient object detection frameworks often extract features and model interimage relations separately. Although these methods achieve promising performance in many scenes, separating the feature extraction and relation modeling falls short of obtaining discriminative features for co-salient objects, resulting in subperformance, especially in some complex and cluttered real-world scenes. In this article, we introduce a novel CoSOD framework to unify feature extraction and interimage relation modeling. We design an early token interaction module (ETIM) that bridges information flow between branches to simultaneously realize feature extraction and interimage information interaction. To further enhance our network's capability to distinguish co-salient objects from other irrelevant foreground objects, we introduce a pixel-to-group contrastive (PGC) learning method. This approach aids in eliminating the need for additional interaction modules while preserving features' discriminative power for co-salient objects. Our proposed CoSOD framework only includes a backbone embedded with ETIM, a decoder without interaction modules and a project head only used during the training phase. Extensive experiments on three challenging benchmarks, that is, CoCA, CoSOD3k, and Cosal2015, demonstrate that our proposed method can outperform current leading-edge models and achieve the new state-of-the-art. The source code is available at https://github.com/zhiwang98/LDRNet.
Yongri Piao, Tingwei Liu, Jihao Yin, Miao Zhang 0004, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.4
2024 Neural Radiance Fields for Unbounded Lunar Surface Scene
abstract
Accurate understanding of lunar surface topography is vital for effective decision-making and remote control of lunar rovers during exploration missions. Conventional sensing methods often struggle to capture the intricate details of the lunar landscape. In response, we propose an innovative approach that leverages NeRF to synthesize new viewpoints within the expansive lunar environment. By blending 3D hash grids and 2D plane grids representations, our approach provides a comprehensive scene representation. We employ the technique of spiral sampling and feature rendering to enhance rendering quality while simultaneously reducing training time. Additionally, we leverage sparse point cloud to aid the model in better learning the geometric structure of the lunar environment. Through experimentation, we have demonstrated that our method is capable of synthesizing realistic images of lunar environments.
Linyan Cui, Jihao Yin
ICRA3
2024 Prototype-Guided Structural Learning from Visual Foundation Model for Few-Shot Aerial Image Semantic Segmentation
abstract
Few-shot aerial image semantic segmentation aims to segment query images with few annotated support samples. It is challenging due to intra-class variations and complex object details in remote aeiral images. However, these two issues are inadequately addressed in existing few-shot segmentation methods. In this paper, we propose a novel Prototype-Guided structural learning (PGSL) framework based on recently proposed segment anything model (SAM). Specifically, to accommodate intra-class variation in aerial image, a novel Prototype-Guided transformer is designed to interact the multiple prototypes from support images with query images, yielding initial segmentation map. Moreover, to improve the performance on object contours, we propose a refine branch based on the SAM, which adopts initial segmentation maps as prompt. This integrates the structural knowledge inherent in SAM into our model. Experiment on iSAID-5i dataset demonstrates the proposed PGSL framework outperforms other state-of-the-art methods.
Qixiong Wang, Hongxiang Jiang, Jiaqi Feng 0001, Guangyun Zhang, Jihao Yin
IGARSS5
2024 Focal Perception Transformer for Light Field Salient Object Detection
Miao Zhang 0004, Yongri Piao, Jihao Yin, Huchuan Lu
PRCV (8)4
2024 S2JO: Spatial-Spectral Joint Optimization for Hyperspectral Image Classification With Noisy Labels
abstract
Hyperspectral image (HSI) annotation often suffers from noisy labels, which brings challenges in classification models training. Existing methods typically address this issue through two separate stages: noisy labels cleaning and robust model design. However, model performance is constrained by insufficient noisy label filtering. To explore an end-to-end joint optimization framework for HSI classification with noisy labels, we propose a spatial–spectral joint optimization (S2JO) network. The S2JO consists of a spatial nonuniform sampling (SNS) module and a spectral prototypes learning (SPL) module. The SNS module filters noisy labels dynamically by picking out the top-N confident points, purifying the input samples for the network. Meanwhile, the SPL module uses multicenter spectral prototypes to extract discriminate features accurately even if the input contains some noisy labels. Extensive experiments on the C2Seg-Beijing HSI subdataset demonstrate the superiority of the proposed S2JO over other state-of-the-art methods.
Jiaqi Feng 0001, Qixiong Wang, Hongxiang Jiang, Guangyun Zhang, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.5
2024 Scene-Object Holistic Relation Network for Fine-Grained Airplane Detection
abstract
The airplane detection and fine-grained recognition in the remote sensing images are challenging due to high interclass indistinction. The subtle distinctions between classes make it difficult to accurately classify objects based purely on bounding box features without considering the broader context. However, recent studies on remote sensing object detection focuses on refining the representation of bounding boxes while ignoring holistic context knowledge in remote sensing scenarios. This letter addresses this gap by introducing the scene-object holistic relation (SOHR) network for fine-grained airplane detection. Specifically, the SOHR network distinctively exploits global scene-object context information through a novel lightweight scene context attention (SCA) module, which aggregates scene context feature and object position information. Furthermore, the object relation transformer (ORT) is designed to model interactions among all objects within the scene explicitly, thereby increasing the model performance for ambiguous hard samples. The experimental results obtained from the FAIR1M dataset demonstrate that the proposed SOHR-Net achieves a state-of-the-art detection accuracy of 56.110% mean average precision (mAP). Compared with the baseline, SOHR-Net exhibits an increase of 2.517%.
Weiyu Ning, Qixiong Wang, Jiaqi Feng 0001, Hongxiang Jiang, Guangyun Zhang, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.6
2024 CAT: Center Attention Transformer With Stratified Spatial-Spectral Token for Hyperspectral Image Classification
abstract
Most hyperspectral image (HSI) classification methods rely on square patch sampling to incorporate spatial information, thereby facilitating the label prediction of the center pixel. However, square patch sampling introduces numerous heterogeneous pixels, which could distort the label prediction of center pixel. Moreover, it generates fixed training patch sample for each center pixel, hampering the performance of transformer-based models requiring a large number of training data. To address the above problems, we proposed Center Attention Transformer (CAT) with stratified spatial-spectral token generated by superpixel sampling for HSI classification. Firstly, to mitigate the inference of heterogeneous pixels, we propose Sampling From Superpixel Region mechanism to generate purer image cubes than traditional square neighborhood. Secondly, to expand the training data for transformer, we propose Multiple Stratified Random Sampling mechanism, which generates ample training samples without introducing additional labels. Finally, to more effectively extract information from the sampled patch tokens, we propose Spatial Spectral Token Generation mechanism and Center Attention Transformer structure with Gaussian Positional Embedding. This framework can extract long-range correlations of spectral information and pay more attention on the center pixel in spatial dimension. Experimental results on three HSI datasets demonstrate the performance of our proposed method CAT outperforms several state-of-the-art methods. The code of this work is available at https://github.com/fengjiaqi927/CAT-Center_Attention_Transformer.
Jiaqi Feng 0001, Qixiong Wang, Guangyun Zhang, Xiuping Jia, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.5
2024 Balanced Orthogonal Subspace Separation Detector for Few-Shot Object Detection in Aerial Imagery
abstract
Few-shot object detection (FSOD) in remote sensing images (RSIs) aims to achieve object location and classification with only a few training samples. Currently, mainstream transfer-learning methods employ a two-stage approach: pretraining on data-abundant base classes and fine-tuning on few-shot novel classes. However, existing approaches suffer notable degradation in both base and novel classes during fine-tuning, because of gradient conflict and class imbalance. To address this, we construct the balanced orthogonal subspace separation (BOSS) detector, a novel two-stage framework for FSOD. Specifically, to avoid contradictory gradients, BOSS distinctly isolates the training of base and novel classes at both structural and feature levels. For structural separation, a low-rank subspace adapter (LoSA) is introduced to ensure network optimization for novel classes without hampering base classes’ pretraining performance, effectively addressing over-fitting in few-shot scenarios. For feature disentanglement, an orthogonal subspace extractor (OSE) is presented, enhancing class separability by learning class-specific, orthogonal basis-spanned subspace. Finally, a balanced classifier (BC) is proposed to equalize the imbalanced loss, with its dual-component design mitigating bias toward predicting background or base classes. Comparative evaluations on diverse remote sensing datasets demonstrate BOSS’s superiority, outperforming state-of-the-art methods in mean average precision (mAP). These results underscore BOSS’s effectiveness in FSOD, particularly in challenging remote sensing contexts.
Hongxiang Jiang, Qixiong Wang, Jiaqi Feng 0001, Guangyun Zhang, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.5
2024 Disentangled Foreground-Semantic Adapter Network for Generalized Aerial Image Few-Shot Semantic Segmentation
abstract
Semantic segmentation of remote sensing imagery requires extensive annotated samples for training, facing challenges in adapting to novel classes with few annotations. Few-shot semantic segmentation (FS-Seg) employs a support-to-query paradigm, which encounters many practical constraints. Recently, generalized few-shot semantic segmentation (GFS-Seg) has been proposed to align with general semantic segmentation paradigms, enabling the segmentation of all classes (both base and novel classes) in the image. However, existing GFS-Seg methods struggle with a large intra-class variance of background, degradation on base classes, and overfitting on novel classes during fine-tuning in aerial imagery. To address the above issues, we propose the disentangled foreground-semantic adapter network (DFSA-Net) for generalized aerial image FS-Seg. Specifically, to reduce the interference from background features, DFSA-Net employs a foreground-semantic decoder (FSD) to decompose semantic segmentation into foreground aggregation and multiclass refinement. To mitigate the base classes degradation and novel class overfitting during fine-tuning, we propose disentangled low-rank adapter (DLA) for fine-tuning phase, designed to preserve the base parameters while ensuring efficient adaptation to novel classes. Finally, we introduce an inference ensemble strategy that merges base and novel decoder prediction to achieve final output. Experimental results on NWPU and iSAID datasets demonstrate the superiority of our DFSA-Net over other compared methods.
Qixiong Wang, Jihao Yin, Hongxiang Jiang, Jiaqi Feng 0001, Guangyun Zhang
IEEE Trans. Geosci. Remote. Sens.2
2024 Integrating Neural Radiance Fields With Deep Reinforcement Learning for Autonomous Mapping of Small Bodies
abstract
Surface 3-D mapping of small bodies is crucial for planning in situ deep-space explorations, yet current approaches rely heavily on manual processing and significantly increase time costs. To achieve autonomous and efficient 3-D mapping of small bodies, we propose a visual-based autonomous 3-D mapping framework that integrates neural radiance fields (NeRFs) with deep reinforcement learning (DRL). We introduce a NeRF variant, which is built within the optimization loop of DRL, for instant 3-D mapping and quantitative uncertainty estimation of the global mapping. We employ the uncertainty estimation to optimize the DRL policy, thereby predicting the next best state that minimizes uncertainty in surface mapping. Moreover, we introduce a new six degrees of freedom (6 DoF) simulation environment that simulates the visual-based 3-D mapping process of small bodies, and demonstrate that our framework significantly enhance 3-D mapping quality. Our framework provides a new technological direction for small body exploration missions, potentially enhancing efficiency while also advancing operational reliability.
Linyan Cui, Chuankai Liu, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.4
2024 Lunar Rover Cross-View Localization Through Integration of Rover and Orbital Images
abstract
Efficient visual localization of lunar rovers is essential for long-range autonomous exploration missions and the construction of the International Lunar Research Station. Given the communication delay and bandwidth limitations between the Earth and the Moon, we propose a novel cross-view localization framework that autonomously registers a single rover image to an orbital image. We developed a bird’s-eye-view (BEV) feature synthesis method that integrates geometric projection, cross-scale feature transfer (CSFT), and contour guidance mechanism (CGM). The basic principle involves first projecting rover features into the BEV perspective through geometric projection, then reconstructing BEV features in reference to orbital features using CSFT. This process is guided by CGM, enhancing the expression of rich terrain contours in BEV and orbital features, thereby improving the viewpoint and scale consistency of cross-view features. By conducting dense spatial correlation searches between BEV and orbital features, we can accurately estimate the position of the lunar rover. Additionally, we introduce a lunar surface simulation environment and construct the lunar cross-view localization (LCVL) simulation dataset based on this environment to demonstrate the framework’s effectiveness. Our research offers a new solution for rover localization, potentially improving the efficiency of future exploration missions.
Linyan Cui, Chuankai Liu, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.5
2023 Representation Disparity-aware Distillation for 3D Object Detection
abstract
In this paper, we focus on developing knowledge distillation (KD) for compact 3D detectors. We observe that off-the-shelf KD methods manifest their efficacy only when the teacher model and student counterpart share similar intermediate feature representations. This might explain why they are less effective in building extreme-compact 3D detectors where significant representation disparity arises due primarily to the intrinsic sparsity and irregularity in 3D point clouds. This paper presents a novel representation disparity-aware distillation (RDD) method to address the representation disparity issue and reduce performance gap between compact students and over-parameterized teachers. This is accomplished by building our RDD from an innovative perspective of information bottleneck (IB), which can effectively minimize the disparity of proposal region pairs from student and teacher in features and logits. Extensive experiments are performed to demonstrate the superiority of our RDD over existing KD methods. For example, our RDD increases mAP of CP-Voxel-S to 57.1% on nuScenes dataset, which even surpasses teacher performance while taking up only 42% FLOPs.
Yanjing Li, Sheng Xu 0007, Mingbao Lin, Jihao Yin, Baochang Zhang 0001, Xianbin Cao 0001
ICCV4
2023 Learning from small data for hyperspectral image classification
Xiaoyan Luo, Jihao Yin
Signal Process.4
2023 Multiscale Prototype Contrast Network for High-Resolution Aerial Imagery Semantic Segmentation
abstract
Semantic segmentation of high-resolution aerial images is a challenging task on account of complex scene-variation and large scale-difference. However, these two issues are inadequately addressed in general semantic segmentation methods. In this paper, we propose a Multi-scale Prototype Contrast Network (MPCNet) to improve the adaptive capability for different scenes and scales. Specifically, a novel multi-scale prototype transformer decoder (MPTD) is designed to extract dynamic scene-specific prototypes as pixel classifier by fusing information of feature maps and learnable class tokens. To exploit cross-scene context information and accommodate the large scale-difference in aerial image, we build a multi-scale prototype memory queue to store these multi-scale prototypes during training. Upon the multi-scale prototype memory queue, a novel multi-scale prototype contrastive loss is proposed to increase object feature discriminability across multiple scale, which brings better consistency of intermediate feature and boosts the convergence of network. Extensive experimental results on three publicly available datasets demonstrate the effectiveness and efficiency of our MPCNet over other state-of-the-art methods. The code is available at https://github.com/qixiong-wang/mmsegmentation-mpcnet.
Qixiong Wang, Xiaoyan Luo, Jiaqi Feng 0001, Guangyun Zhang, Xiuping Jia, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.6
2023 Mesh-Based DGCNN: Semantic Segmentation of Textured 3-D Urban Scenes
abstract
Textured 3D mesh is one of the final user products in photogrammetry and remote sensing. However, research on the semantic segmentation of complex urban scenes represented by textured 3D meshes is in its infancy. We present a mesh-based dynamic graph CNN (DGCNN) for the semantic segmentation of textured 3D meshes. To represent each mesh facet, composite input feature vectors are constructed by concatenating the face-inherent features, i.e., XYZ coordinates of the center of gravity (CoG), texture values, and normal vectors. A texture fusion module is embedded into the proposed mesh-based DGCNN to generate high-level semantic features of the high-resolution texture information, which is useful for semantic segmentation. We achieve competitive accuracies when the proposed method is applied to the SUM mesh datasets. The overall accuracy (OA), Kappa coefficient (Kap), mean precision (mP), mean recall (mR), mean F1 score (mF1), and mean intersection over union (mIoU) are 93.3%, 88.7%, 79.6%, 83.0%, 80.7%, and 69.6%, respectively. In particular, the OA, mean class accuracy (mAcc), mIoU, and mF1 increase by 0.3%, 12.4%, 3.4%, and 6.9%, respectively, compared to the state-of-the-art method.
Guangyun Zhang, Jihao Yin, Xiuping Jia, Ajmal Mian
IEEE Trans. Geosci. Remote. Sens.3
2022 Spectral Transformer with Dynamic Spatial Sampling and Gaussian Positional Embedding for Hyperspectral Image Classification
abstract
Owing to the global information extraction ability, transformers have been tentatively applied to hyperspectral image(HSI) classification. However, the existing transformer-based methods have not made full use of the flexible characteristics of spatial sampling nor considered the importance of the central pixel to the classification of HSI cubes. In order to enhance adaptability of transformers for HSI classification, we have proposed a novel spectral transformer with dynamic spatial sampling and gaussian positional embedding. To improve the effectiveness of spatial neighborhood information, Spatial Sample Selection(3S) mechanism generates image cube from super pixel region, making image cube more pure for classification. To extract long-range information in spectral dimension, Spectral Feature Extraction(SFE) network splits spectral bands into several slices and calculates the attention between them. To stress the importance of the central pixel to the classification of image cube, Gaussian Positional Embedding(GPE) reduces the weight of surrounding pixels during feature embedding stage. Experimental results demonstrate the performance of our proposed method. The code of this work is available at https://github.com/fengjiaqi927/HSI_transformer.
Jiaqi Feng 0001, Xiaoyan Luo, Qixiong Wang, Jihao Yin
IGARSS5
2022 Class-Balanced Contrastive Learning for Fine-Grained Airplane Detection
abstract
Airplane detection and fine-grained recognition in remote sensing images are challenging due to class imbalance and high inter-class indistinction. To alleviate these issues, we propose class-balanced contrastive learning (CBCL) approach for airplane detection to exploit the correlation between samples in different images, which is rarely explored in previous research. Specifically, we first dynamically build class-balanced memory queues during training, which mitigates class imbalance by memorizing training samples. Upon class-balanced memory queues, hard triplet contrastive learning is introduced to increase the inter-class discriminability, which enforces the maximum distance of the positive sample pair to be smaller than the minimum distance of the negative sample pair. We integrate the proposed CBCL strategy into oriented object detection frameworks for fine-grained airplane detection. The experimental results on the FAIR1M dataset reveal that several state-of-the-art algorithms with CBCL achieve significantly improvements.
Qixiong Wang, Xiaoyan Luo, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.4
2022 H2AN: Hierarchical Homogeneity-Attention Network for Hyperspectral Image Classification
abstract
Recently, a self-attention network (SAN) is developed as an effective strategy to extract features from attention areas for image classification. However, for hyperspectral image (HSI), the lack of position supervision of object regions and inefficient similarity computation lead to unsatisfactory classification performance on mixed pixels. To alleviate the above two problems for HSI image classification, we propose a novel hierarchical homogeneity-attention network (H2AN) in this article. First, we design a homogeneity-attention block (HAB) to depict the feature correlation with the homogeneous mask. Using the supervision of homogeneity mask, we can calculate the attention guided by the predefined number of homogeneity embeddings, which can reduce the heavy computation instead of the global search in self-attention block (SAB) of SAN. Second, we propose a hierarchical convolutional neural network (HCNN) inserting the HAB into different levels of network cells for highly efficient feature extraction of target regions, named H2AN. Because of the transferring of homogeneity property from shallow layer to deep layer, our H2AN outperforms the state-of-the-art methods in qualitative and quantitative experiments on three typical datasets.
Xiaoyan Luo, Qixiong Wang, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.5
2021 Hyperspectral Classification Using Cooperative Spatial-Spectral Attention Network With Tensor Low-Rank Reconstruction
abstract
Spatial and spectral attention networks have been both well introduced to Hyperspectral image (HSI) classification. However, in previous works, they are seldom considered jointly. To obtain a 3D spatial-spectral attention map, which is beneficial for extracting discriminative spatial-spectral features, we propose a novel cooperative spatial-spectral attention network with tensor low-rank reconstruction. Firstly, a tensor low-rank reconstruction (TLRR) block is designed to learn a spatial-spectral attention map tensor, which adaptively emphasizes the attention features of the salient spatial positions and informative spectral bands simultaneously. Secondly, these attention features are merged into simple convolutional features which are more discriminative for classification. Finally, the experimental results demonstrate that our proposed method outperforms some state-of-the-art methods on two typical HSI datasets.
Xiaoyan Luo, Qixiong Wang, Weifa Shen, Jihao Yin
ICIP6
2021 Unsupervised Domain Adaptation for Semantic Segmentation via Self-Supervision
abstract
Recently, deep learning (DL) methods have been widely used for semantic segmentation of remote sensing and achieved significant progress. However, DL-based methods are time-consuming and labor intensive for the networks requiring abundant data with accurate labeling. To solve this issue, unsupervised domain adaption (UDA) has recently been used to transfer the information from labeled source domain to unlabeled target domain. In this paper, we propose a novel UDA approach based on the self-supervised theory for remote sensing image. Specifically, we firstly utilize the inter-domain adaptation to reduce the gap between the source and target domain. Secondly, based on our proposed spatial-frequency (SF) index, we detach the target domain into an easy and hard split. Ultimately, we adopt the intra-domain adaptation by self-supervised adaptation to improve the performance of hard split. Experimental results on ISPRS Vaihingen and Potsdam datasets demonstrate the effectiveness and rationality of our methods against the other state-of-the-art approaches.
Weifa Shen, Qixiong Wang, Hongxiang Jiang, Jihao Yin
IGARSS5
2021 Multibranch Spatial-Channel Attention for Semantic Labeling of Very High-Resolution Remote Sensing Images
abstract
Very high-resolution (VHR) remote sensing images can provide fine but sometimes trivial ground object details; thus, the semantic labeling of VHR images is a challenging task. To improve the VHR labeling performance, spatial multiscale information and channel attention have been employed recently. However, the exploitation of global object features is still limited, which leads to the loss of capturing within-class variation from location to location. In this letter, we present a multibranch spatial-channel attention (MSCA) model to efficiently extract global dependency and combine it with multiscale and channel attention methods. In the spatial multiscale attention block, a multibranch feature fusion model is established to exploit the global relationship captured by self-attention and the multiscale correlation learned from dilated convolutions. To alleviate the computational cost of pixel-by-pixel self-attention operation, a spatial pyramid compressing method is also designed. In the channel attention block, average and max global pooling strategies are applied, respectively, in two channel attention branches to generalize global information from different perspectives. Those two blocks are then adaptively united by learnable weighting parameters. Experiments on two VHR image data sets demonstrate that the proposed network can yield better performance in comparison with state-of-the-art labeling methods tested.
Bingnan Han, Jihao Yin, Xiaoyan Luo, Xiuping Jia
IEEE Geosci. Remote. Sens. Lett.2
2021 Joint Spatial-Spectral Attention Network for Hyperspectral Image Classification
abstract
Hyperspectral images (HSIs) contain rich context information in the spatial domain and spectral domain. To fully explore that information, a data-driven joint spatial–spectral attention network (JSSAN) is proposed in this letter. Specifically, we first design a spatial–spectral attention ($\text{S}^{2}\text{A}$) block to simultaneously capture long-range interdependency of spatial and spectral data via the similarity evaluation. Then we adopt a weighted sum operation of features at all spatial positions and channels to selectively aggregate discriminative spatial–spectral features. Second, the$\text{S}^{2}\text{A}$block is inserted into simple convolutional neural network (CNN) structure to extract more representative features for classification, by adaptively emphasizing features of informative land covers and spectral bands which contribute more to class identification. The experimental results reveal that our proposed method outperforms several state-of-the-art algorithms.
Jihao Yin, Xiuping Jia, Bingnan Han
IEEE Geosci. Remote. Sens. Lett.2
2020 Few-Shot Learning With Attention-Weighted Graph Convolutional Networks For Hyperspectral Image Classification
abstract
In this paper, to alleviate the demand for enormous labeled data in the classification task, an Attention-weighted Graph Convolutional Networks (AwGCN) model for hyperspectral image (HSI) few-shot classification is proposed, which aims to explore the internal relationships of data for semi-supervised label propagation. To be specific, the attention-weighted graph is exploited to fully quantify the relationships of all samples, which is potential to solve the HSI few-shot learning problems. Subsequently, Graph Convolutional Networks (GCN) are applied to spread the labels, which ascertain the categories of samples based on the trained attention-weighted graph. The robust prediction of our proposed approach is validated on the real HSI and the experimental results show a competitive good performance, which demonstrates the superior ability of AwGCN in HSI few-shot classification.
Xinyi Tong 0001, Jihao Yin, Bingnan Han, Hui Qv
ICIP2
2020 Neural Network Pruning for Hyperspectral Image Band Selection
abstract
Neural network pruning attempts to reduce parameters without hurting original performance by inducing connection matrix sparsity of network. Inspired by this idea, we proposed an effective pruning-based band selection strategy, which is a potent feature extraction tool in hyperspectral image (HSI) classification. At first, we take the whole HSI bands as input to train original network parameters. For each band, all parameters in network are integrated to measure the band importance. With the novel band signification factor constraining, then the convolutional neural network (CNN) is pruned and remains some representative weights to retrain the compact sub-network, which can finally deal with the hyperspectral band selection problem. Experimental results on the real HSI dataset demonstrate that network pruning-based method can outperform the original CNN in classification accuracy. Also, it can achieve the superiority over filter-based and other CNN-based band selection algorithms in classification accuracy. Our code is available at https://github.com/qixiong-wang/Network-pruning-for-HSI-band-selection.
Qixiong Wang, Xiaoyan Luo, Jihao Yin
IGARSS4
2020 LG: A clustering framework supported by point proximity relations
Hui Qv, Jihao Yin, Xiaoyan Luo
Pattern Recognit.2
2019 Hyperspectral Image Classification Based on Generative Adversarial Network with Dropblock
abstract
Deep learning (DL) algorithms are widely applied in hyperspectral images (HSIs) classification. However, the insufficient utilization in spatial semantic information and inadequate number of HSIs samples both restrict the classification performance of DL-based HSIs algorithms. In this paper, we propose a novel method based on generative adversarial network (GAN) with DropBlock structure (DBGAN). Specifically, DropBlock enforces each unit in convolution neural network (CNN) to learn features by dropping contiguous regions of feature maps, therefore more spatial semantic information is capable to contribute in HSIs classification. Furthermore, GAN model can generate realistic samples by an adversarial game to mitigate HSIs data shortage. Extensive experimental comparisons demonstrate the effectiveness of the proposed method.
Jihao Yin, Bingnan Han
ICIP1
2019 Generative Adversarial Network with Folded Spectrum for Hyperspectral Image Classification
abstract
Hyperspectral image (HSIs) with abundant spectral information but limited labeled dataset endows the rationality and necessity of semi-supervised spectral-based classification methods. Where, the utilizing approach of spectral information is significant to classification accuracy. In this paper, we propose a novel semi-supervised method based on generative adversarial network (GAN) with folded spectrum (FS-GAN). Specifically, the original spectral vector is folded to 2D square spectrum as input of GAN, which can generate spectral texture and provide larger receptive field over both adjacent and non-adjacent spectral bands for deep feature extraction. The generated fake folded spectrum, the labeled and unlabeled real folded spectrum are then fed to the discriminator for semi-supervised learning. A feature matching strategy is applied to prevent model collapse. Extensive experimental comparisons demonstrate the effectiveness of the proposed method.
Jihao Yin, Bingnan Han, Hongmei Zhu
IGARSS2
2019 Global Self-Labeled Distribution Analysis for Hyperspectral Band Selection
abstract
A global self-labeled distribution analysis (GSLDA) for hyperspectral image (HSI) band selection is proposed in this paper, which focuses on an unsupervised method to ascertain the band discrimination. In order to generate the band labels for further analysis, the concept of the local minimum spanning forest (LMSF) is introduced into the construction of the global self-labeled band partitions based on graph theory. Meanwhile, the novel scoring strategy of triple-density indexes is applied to analyze the labeled-band distribution for determining the selected band subset with clear discrimination. The feasibility of the proposed method is evaluated on real hyperspectral data and the experiment results show a competitive good performance, which demonstrates that the selected bands hold apparent global discrimination and robust noise immunity.
Xinyi Tong 0001, Jihao Yin, Limin Wu, Hui Qv
IGARSS2
2019 Hyperspectral Image Classification Using CapsNet With Well-Initialized Shallow Layers
abstract
In this letter, an alternative data-driven HSI classification model based on CapsNet is proposed rather than recently predominant convolutional neural network (CNN)-based models. To adjust the CapsNet to HSI classification, we tune a new CapsNet architecture with three convolutional layers. The added shallow layer provides higher level features to the primary capsules, which indirectly speeds up the following routing procedure. To guarantee a good convergence of the whole CapsNet, the three shallow layers are initialized by transferring convolutional parameters from a pretrained CNN model. The improved CapsNet-based models with and without vote strategy both achieve significantly superior performance in HSI classification to the state-of-the-art CNN-based methods on real hyperspectral data sets.
Jihao Yin, Hongmei Zhu, Xiaoyan Luo
IEEE Geosci. Remote. Sens. Lett.1
2018 Spectral Diversity Enhancement for Pansharpening
abstract
Pansharpening is to generate a synthetic image with high spatial resolution and high spectral resolution via fusing panchromatic (PAN) and multispectral (MS) images. In most traditional pansharpening methods, the original MS image is firstly interpolated to the same size of PAN image by analytical interpolations. However, these interpolated methods could cause false spectral information due to ignoring mixed spectral characteristics in the MS image. To enhance the spectral diversity of upsampled MS image, we use the spatial structure information in PAN image to support the pansharpening in this paper. By introducing the superpixel structure for PAN image, the processible mixed pixels can be screened out in corresponding MS locations, and the other MS locations are considered to be occupied by pure pixels. For the pure and mixed pixels, their upsampling MS results can be obtained via a directly expending manner and a sparse representation manner respectively. Two different detail injection strategies are used for assessing the performance of analytical interpolations and our approach for pansharpening. Experimental results demonstrate that our method achieves the appreciable improvements with respect to analytical interpolations.
Liangyu Zhou, Xiaoyan Luo, Jihao Yin
ICIP3
2018 Gaussian Attractive Force-Based Alternative Parametric Active Contour Model for 3D Lunar Crater Detection
abstract
In this paper, we present an alternative parametric active contour (APAC) model for 3D lunar crater detection with shadow and overexposure problem. Compared with the traditional parametric active contour model, the main difference is that we construct a Gaussian attractive force field between two initial curves for each crater proposal, which enables the two initial curves to mutually convey message to each other and avoids local optimum. In addition, we introduce the elevation information estimated from CCD stereo images into the external energy term of the APAC model, which helps to remove false craters by providing the geometric properties of craters. The proposed method is evaluated on eight pairs of stereo images with different numbers, scales and illuminations captured by Chang'E-I satellite, which demonstrates its effectiveness and efficiency.
Shangbin Huang, Jihao Yin, Hongmei Zhu
IGARSS2
2018 Band Dual Density Discrimination Analysis for Hyperspectral Image Classification
abstract
A novel band discrimination analysis framework for hyperspectral image (HSI) supervised classification is proposed based on dual density (DD). Different from the popular supervised band selection (BS) approaches which measure the discrimination among classes under multivariate normal distribution hypothesis, our work infers the class discrimination degree (overlapping extent) for valid extraction of band subset without any assumed distribution. In the proposed framework, it is crucial to find indexes to measure the discrimination degree of each band, and therefore we develop the DD indexes, including the homogeneity density and the heterogeneity density. Viewing each band of the HSI as a data set, i.e., the data points in each data set are 1-D, and we first obtain the DD value pairs for all data points in each data set. Then, for each data set, we determine its discrimination degree using DD-based zone ratio or score quantify strategy. Finally, the bands, which are determined as the nonoverlapped or have high scores, are chosen as the band subset for the subsequent classification. Superiorities of the proposed BS are demonstrated on the three real-world HSIs over several well-known BS algorithms in terms of classification accuracy and speed.
Hui Qv, Jihao Yin, Xiaoyan Luo, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.2
2018 Corrections to "Segment-Oriented Depiction and Analysis for Hyperspectral Image Data"
abstract
In[1], information regarding the corresponding author is missing. The information is updated here. The updated footnote below shows that Xiaoyan Luo is the corresponding author for this paper.
Jihao Yin, Hui Qv, Xiaoyan Luo, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.1
2017 A hierarchical superpixel aggregation model for hyperspectral image
abstract
Superpixel has been widely applied in hyperspectral image processing as a pre-processing step for over-segmentation. However, most superpixel algorithms are difficult to control the segmentation balance between fragmentation and accuracy. In this paper, we propose a superpixel aggregation model to cluster the over-segmentations. Based on the own importance and interrelationship of superpixels, a two-step merging procedure is designed in the hierarchical wise from local to global comparisons. Aiding by a density peak metric, which is to exploit the spectral correlation in hyperspectral image, the similar neighbor superpixels are merged firstly, and then the similar regions in discontinuous spatial location are gathered. Experimental results show that the proposed model can achieve high accuracy in low region number compared with original superpixel algorithm, and the performance for unsupervised classification application is also remarkable.
Bingnan Han, Jihao Yin, Xiaoyan Luo, Hui Qv
IGARSS2
2017 3D lunar craters detection based on stereo matching
abstract
In this paper, we focus on the 3D crater detection problem on lunar surface, which helps high-precision spacecraft landing and rover navigation in moon exploration projects. A random structured forests method is firstly applied to detect the 2D edges of craters, and then dense correspondence between CCD stereo images estimates the elevations of craters. Finally, we propose a 3D crater detection model, which is solved by an iterative optimization algorithm with the initial 2D edge. Our method is evaluated on three pairs of stereo images captured by Chang'E-I satellite. Experimental results show that the 3D craters in each pair are quickly and accurately located, which demonstrate the effectiveness and efficiency.
Hongmei Zhu, Jihao Yin, Ding Yuan 0001
IGARSS2
2017 SVCV: segmentation volume combined with cost volume for stereo matching
abstract
Stereo matching between binocular stereo images is fundamental to many computer vision tasks, such as three‐dimensional (3D) reconstruction and robot navigation. Various structures of real 3D scenes lead stereo matching to be an old yet still challenging problem. In this study, the authors proposed a novel adaptive support weights technique which exploits the hierarchical information provided by multilevel segmentation to preserve the robustness to imaging conditions and spatial proximity in cost aggregation. Besides, a generalisable cost refinement strategy is designed to remove the matching ambiguity in large weakly textured regions. The proposed strategy utilises both the fluctuation of the filtered cost volume and the colour information to further improve the matching accuracy. Experimental results of 50 stereo images demonstrate the effectiveness and efficiency of the proposed method. Furthermore, a systematic evaluation is developed to assess the conventional steps in local stereo methods and then reliable suggestions are given to the beginners and researchers outside the stereo matching field.
Hongmei Zhu, Jihao Yin, Ding Yuan 0001
IET Comput. Vis.2
2017 Information-Assisted Density Peak Index for Hyperspectral Band Selection
abstract
Band selection has become an effective method to reduce hyperspectral dimensionality. In this letter, an information-assisted density peak index (IaDPI) is proposed to prioritize the bands. Based on a clustering method by finding density peaks, IaDPI introduces the intraband information entropy into the local density and intercluster distance to ensure cluster centers with a high quality. Also, the band distance is integrated with channel proximity to control the compactness of local density. Owing to the intraband entropy and the interband weighted dissimilarity, the selected band set with top-ranked IaDPI scores can hold high local density, clear global distinction, and good informative quality. Experimental results on real hyperspectral data indicate the advantages of the proposed IaDPI in good selection quality, robust noise immunity, and high classification accuracy.
Xiaoyan Luo, Rui Xue 0003, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.3
2017 Sparse representation over discriminative dictionary for stereo matching
Jihao Yin, Hongmei Zhu, Ding Yuan 0001, Tianfan Xue
Pattern Recognit.1
2017 Segment-Oriented Depiction and Analysis for Hyperspectral Image Data
abstract
A novel segment-oriented dictionary learning (SeODL) framework for hyperspectral image (HSI) classification is proposed. Differing from existing HSI classification methods which directly process the original whole spectral curves of pixels, our work focuses on local segment analysis to achieve fine depiction and effective exploitation. Viewing the separated segment as a basic processing unit, we first cluster them into two sets with the homogeneity in trend and fluctuation, and then two small dictionaries can be quickly learned. Second, to get meticulous and discriminability enhanced segment-oriented representations (SORs), the segments of the training and test pixels are coded on a novel binary-separated coding strategy. The coding stage for obtaining SORs is sped up by the employment of our proposed enhanced orthogonal matching pursuit technique. A characteristic splicing classifier with high performance can be trained using these SORs of the training pixels. Finally, a spiral searching strategy and a multiple majority-voting method are adopted for fully spatial information incorporation of the test pixels whose final SORs will be embedded into the trained characteristics splicing classifier to ascertain the labels. Experimental results on three real HSI data sets demonstrate the superiority of the proposed SeODL framework over several well-known classification algorithms in terms of classification accuracies.
Jihao Yin, Hui Qv, Xiaoyan Luo, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.1
2016 3-D point cloud normal estimation based on fitting algebraic spheres
abstract
In this paper, we proposed a novel method to estimate the normal information of the unorganized point cloud, which plays an essential part in 3D reconstruction. The original point cloud is firstly divided into cubes with different sizes by the octree method. Then, we fit algebraic sphere in each cube instead of planar surface to improve the accuracy of normal estimation. Finally, the raw normals are refined by a weighting function which increases along with the depth of octree. For evaluation, we compute the intersection angles between the estimated normals and the corresponding groundtruth. Besides, the estimated normals are also plugged into the Poisson surface reconstruction algorithm for intuitive comparison. Experimental results demonstrate the effectiveness of our normal estimating methods. Moreover, the strategy that normal estimation after division saves much more computing time, which promises the efficiency of our method.
Ding Yuan 0001, Hongmei Zhu, Jihao Yin
ICIP4
2016 Cloud detection of remote sensing images by deep learning
abstract
Cloud detection plays a major role for remote sensing image processing. Most of the existed cloud detection methods use the low-level feature of the cloud, which often cause error result especially for thin cloud and complex scene. In this paper, a novel cloud detection method based on deep learning framework is proposed. The designed deep Convolutional Neural Networks (CNNs) consists of four convolutional layers and two fully-connected layers, which can mine the deep features of cloud. The image is firstly clustered into superpixels as sub-region through simple linear iterative cluster (SLIC) method. Through the designed network model, the probability of each superpixel that belongs to cloud region is predicted, so that the cloud probability map of the image is generated. Lastly, the cloud region is obtained according to the gradient of the cloud map. Through the proposed method, both thin cloud and thick cloud can be detected well, and the result is insensitive to complex scene. Experimental results indicate that the proposed method is more robust and effective than compared methods.
Mengyun Shi, Fengying Xie, Yue Zi, Jihao Yin
IGARSS4
2016 Planet mineral distribution detection via clustering-aware nonnegative matrix factorization
abstract
Spectral unmixing is an important technique to exploit mineral distribution through remote sensing image. In this paper, we propose an unmixing algorithm combining clustering-aware method with the sparsity-constrained nonnegative matrix factorization (SNMF) algorithm. Pixels with similar spectra have high possibility to share similar typical endmembers, therefore we preprocess the image using K-means cluster algorithm and then optimizes the initial endmember spectra by selecting the typical ones of each cluster as the initial endmember value. Due to the local convergence feature of NMF, the optimal initial value can accelerate the convergence of the algorithm and obtain more accurate results. Meanwhile, we use the sparsity-constrained NMF in global unmixing to control the sparse property of abundance distribution. The experiments on synthetic data and Chang'e-1 hyperspectral data show that K-means nonnegative matrix factorization (KNMF) is superior to the other unmixing methods.
Jihao Yin, Xiaoyan Luo, Hui Qv, Bingnan Han
IGARSS1
2016 Dem-based shadow detection and removal for lunar craters
abstract
In this paper, we focus on the shadow problem of the lunar surface, which hinders the implementation of visual tasks and visual processing for moon exploration projects. A random walker model is firstly applied to detect the shadowed pixels in lunar craters. Then, the detected shadows are removed by rectifying the illumination coefficients and detail coefficients, which are obtained by using the multi-scale decomposition technique. The proposed algorithm is evaluated on three groups of CCD images and the corresponding DEM data. Satisfactory results, which achieve comparative illumination yet preserve enough details, demonstrate the effectiveness of the proposed algorithm.
Hongmei Zhu, Jihao Yin, Ding Yuan 0001, Guangyun Zhang
IGARSS2
2016 A Simple and Accurate TDOA-AOA Localization Method Using Two Stations
abstract
This letter focuses on locating passively a point source in the three-dimensional (3D) space, using the hybrid measurements of time difference of arrival (TDOA) and angle of arrival (AOA) observed at two stations. We propose a simple closed-form solution method by constructing new relationships between the hybrid measurements and the unknown source position. The mean-square error (MSE) matrix of the proposed solution is derived under the small error condition. Theoretical analysis discloses that the performance of the proposed solution can attain the Cramér-Rao bound (CRB) for Gaussian noise over the small error region where the bias compared to variance is small to be ignored. The proposed solution can be extended directly to more than two observing stations with CRB performance maintained theoretically. Simulations validate the performance of the proposed method.
Jihao Yin, Qun Wan, Shiwen Yang, K. C. Ho 0001
IEEE Signal Process. Lett.1
2015 Fluctuations of disparity space image for stereo matching in untextured regions
abstract
Allocating the disparities to the untextured regions in stereo image still remains an intractable and challenging problem. In this paper, we present a novel local stereo matching algorithm for large untextured regions. The core ideas behind our method are from two aspects: 1) the fluctuating characteristics of cost volume are first exploited to distinguish ambiguous and unambiguous image regions; 2) the matching costs of pixels in ambiguous regions are regularized with an adaptive cost aggregation. The WTA strategy is performed on the regularized cost volume followed by postprocessing to obtain accurate disparity map. Comparative experiments are conducted on different data sets and the results demonstrate the effectiveness and efficiency of our method.
Hongmei Zhu, Jihao Yin, Ding Yuan 0001, Wei Sui
ICIP2
2015 Design of augmented dictionary for sparse representation based on neural network
abstract
An efficient and flexible dictionary designing algorithm is proposed for sparse and redundant signal representation. The proposed Augmented Dictionary (AD) is based on a new dictionary model with an augmented form compared to the conventional model. With this model, we can bridge the gap between the classic dictionary learning approaches, which have general structure yet lack computational efficiency, and the artificial neural network theory, which has potential high parallel computational efficiency but poor universality of structure. In this paper, we discuss the advantages of augmented dictionary, and interpret how the augmented dictionary can be trained with labeled samples. The proposed neural network based augmented dictionary designing method enjoys some important features, such as high accuracy, strong robustness and desired computational efficiency. As a demonstration of these benefits, we present high-quality hyperspectral image classification results based on the new algorithm.
Hui Qv, Jihao Yin, Charles A. DiMarzio
IGARSS2
2015 Sub-block PCA-wavelet image sharpening approach for hyperspectral images
abstract
One of the most crucial issues to judge image quality is the spatial resolution. Hyperspectral image (HSI) sharpening is the process of combining spatial information to enhance spatial resolution of HSIs. Huge volumes of HSI data cause difficulties during the sharpening process. This paper proposed a practical and effective strategy to deal with HSI sharpening. We utilized sub-block method and combined PCA and wavelet fusion approaches to achieve the proposed scheme. Sub-block method helped reduce the calculation complexity and promote the efficiency. PCA and wavelet image sharpening contributed to enhance spatial resolution of HSI with less spectral distortion. The experiment demonstrated an efficient processing result and a good visualized effect. Qualitative and quantitative assessments were both used to evaluate the proposed approach.
Jianying Sun, Qunbo Lv, Jihao Yin
IGARSS4
2015 Hyperspectral image target detection based on exponential smoothing method
abstract
In this paper, we proposed a new hyperspectral image target detector based on time series analysis, named as Exponential Smoothing Target Detector (ES-TD). As a classical method of time series analysis, the exponential smoothing method can choose the motional weight in the different part of data, which accelerates the reconstruction and forecast of the unknown data. The proposed method has a three-step process. Firstly, we select the applicable smoothing parameter according to the shape of the data curve. Then, given the reference and test spectral curves, we use the exponential smoothing method to obtain two new smoothing curves. Finally, we calculate the similarity between the two smoothing curves using SAM to determine whether the test spectral curve is the target or not. The proposed method has the feature of high computational efficiency and robustness. Experimental results on two real hyperspectral data sets demonstrate the advantages of the new method.
Jihao Yin, Bingnan Han, Wanke Yu
IGARSS1
2015 Camera motion estimation through monocular normal flow vectors
Ding Yuan 0001, Jihao Yin, Jiankun Hu
Pattern Recognit. Lett.3
2015 Haze Removal for a Single Remote Sensing Image Based on Deformed Haze Imaging Model
abstract
The contrast of remote sensing images captured in haze condition is poor, which influences their interpretation. In this letter, a novel dehazing algorithm based on the deformed haze imaging model is proposed. First, the model is deformed by introducing a translation term. Second, the atmospheric light and transmission are estimated according to the new model combined with dark channel prior. Lastly, the haze is successfully removed from remote sensing images using the proposed estimation algorithm. The estimated transmission is insensitive to the texture of ground objects, and the dehazing effect for nonuniform haze is more satisfactory than the compared method. Moreover, our approach can be used for general haze removal through adjusting the translation term. Experimental results reveal that the proposed method can recover the real scene clearly from haze remote sensing images along with the advantage of good color consistency.
Xiaoxi Pan, Fengying Xie, Zhiguo Jiang 0001, Jihao Yin
IEEE Signal Process. Lett.4
2014 Intelligent search optimized edge potential function (EPF) approach to synthetic aperture radar (SAR) scene matching
abstract
Research on synthetic aperture radar (SAR) scene matching in the aircraft end-guidance has a significant value for both research and real-world application. The conventional scene matching methods, however, suffer many disadvantages such as heavy computation burden and low convergence rate so that these methods cannot meet the requirement of end-guidance system in terms of fast and real-time data processing. Furthermore, there are complex noises in the SAR image, which also compromise the effectiveness of using the conventional scene matching methods. To address the above issues, in this paper, the intelligent optimization method, Free Search with Adaptive Differential Evolution Exploitation and Quantum-Inspired Exploration, has been introduced to tackle the SAR scene matching problem. We first establish the effective similarity measurement function for target edge feature matching through introducing the edge potential function (EPF) model. Then, a new method, ADEQFS-EPF, has been proposed for SAR scene matching. In ADEQFS-EPF, the previous studied theoretical model, ADEQFS, is combined with EPF model. We also employed three recent proposed evolutionary algorithms to compare against the proposed method on optical and SAR datasets. The experiments based on Matlab simulation have verified the effectiveness of the application of ADEQFS and EPF model to the field of SAR scene matching.
Jihao Yin
IEEE Congress on Evolutionary Computation2
2014 Crater detection based on local non-negative matrix factorization
abstract
Due to the variations in the terrain, illumination and scale, it is difficult to detect craters from remote sensing image of planet surface. This paper proposes a novel automatic crater detection method by introducing the local non-negative matrix factorization (LNMF) for remote sensing images of Martian surface. LNMF is aimed at learning localized, part-based features from global samples, which has shown considerable prospect in feature extraction. Our detection algorithm contains three key procedures. Firstly, the crater candidates are detected by geometry approaches. Secondly, LNMF is applied in subspace learning for all crater samples and candidates. At last, we get the final detection results by discarding non-craters in candidates. The LNMF-based method has achieved satisfied results in the experiments conducted on the Mars Orbiter Camera (MOC) dataset.
Jihao Yin, Zetong Gu
IGARSS2
2014 Segmentation and classfication of hyperspectral images using Kendall Concordant Coefficient
abstract
As the abundant spectral information of hyperspectral image, traditional pixel-wise classification methods is time-consuming in hyperspectral images. And purely pixel-wise classification methods often ignore lots of space information. In this paper, we investigate the usage of Kendall Concordant Coefficient (KCC) for region-dependent segmentation of the original hyperspectral data cube. The KCC-based method could combine spectral and spatial information effectively, and it has strong robustness with low complexity because it is a nonparametric method. We conduct a series of experiments, and draw conclusions that KCC-based method could obtain better segmentation and classification results than purely pixel-wise methods.
Jihao Yin, Wanke Yu, Zetong Gu
IGARSS1
2013 A novel method of crater detection on digital elevation models
abstract
As an essential geomorphological structures on planetary surface, impact craters can provide significant information in determining the planetary chronology. This paper proposes a novel automatic crater detection algorithm by using digital elevation model (DEM) data. The method includes: 1) pre-processing of the original DEM data, which can eliminate the effect of other landform; 2) iterative crater detections, which can eliminate small objects and analyze roundness. We use the DEMs instead of the imagery data because the DEMs could be unaffected by the solar altitude and atmospheric conditions, etc. Our detection algorithm is evaluated using several test sets of Martian DEM data obtained by the Mars Obiter Laser Altimeter (MOLA) boarded on the Mars Global Surveyor. The experimental results show the high true detection rate and low false detection rate of our algorithm according to the Barlow catalogue.
Jihao Yin, Yueshan Liu
IGARSS1
2013 Wavelet Packet Analysis and Gray Model for Feature Extraction of Hyperspectral Data
abstract
Wavelet packet analysis (WPA) and gray model (GM) are investigated for nonlinear unsupervised feature extraction of hyperspectral remote sensing data in this letter. Treated as derivative series, a hyperspectral response curve of each pixel is decomposed into an approximation and various detailed compositions by WPA, and then, GM is continuously applied to find the relationship among those detailed compositions. Cluster-space representation is used for determining the optimal wavelet. New extracted features can reveal the intrinsic identities of hyperspectral data. Experimental results show the feasibility and reliability of our proposed method in terms of classification accuracy.
Jihao Yin, Xiuping Jia
IEEE Geosci. Remote. Sens. Lett.1
2012 Roof-top detection based on structural elements combination
abstract
A novel roof-top extraction method for satellite images based on probabilistic topic model is presented. We model roof-top as the connected structural elements. The proposed method contains two major steps: 1) Detect structural elements, different from earlier structure detector, the proposed method automatically learn the types of elements from unlabeled samples; 2) Connect these elements to form roof-top boundary, where the relationships between elements are estimated by hierarchical topic model. This approach belongs to generative method where only a small number of roof-top samples are required. The experimental results demonstrate the effectiveness of the proposed approach.
Zhiguo Jiang 0001, Jihao Yin
IGARSS3
2012 Cointegration theory for adaptive target detection in hyperspectral images
abstract
This paper investigates the usage of Johansen Cointegration Test for adaptive target detection with hyperspectral remote sensing data. Johansen Cointegration Test aims at mining long-term equilibrium relationship, which refers to the condition that if pairs of non-stationary series share similar tendencies, their linear combination could be stationary. Hyperspectral data are highly non-stationary series, but there should be similar patterns among the hyperspectral response curves of same materials. To be treated as derivative series, given hyperspectral response curves will be matched with the standard spectrum via Johansen Cointegration Test. The test statistics will be compared to a preset threshold to judge whether they are target or not. Quantitative experiments show that the proposed method performs better than a few other adaptive detection methods tested.
Jihao Yin, Xiuping Jia
IGARSS1
2012 Using incremental subspace and contour template for object tracking
Jihao Yin, Chongyang Fu, Jiankun Hu
J. Netw. Comput. Appl.1
2012 Free Search with Adaptive Differential Evolution Exploitation and Quantum-Inspired Exploration
Jihao Yin, Jiankun Hu
J. Netw. Comput. Appl.1
2012 Using Hurst and Lyapunov Exponent For Hyperspectral Image Feature Extraction
abstract
Hyperspectral image processing has attracted high attention in remote sensing fields. One of the main issues is to develop efficient methods for dimensionality reduction via feature extraction. This letter proposes a new nonlinear unsupervised feature extraction algorithm using Hurst and Lyapunov exponents to reveal local and general spectral profiles, respectively. A hyperspectral reflectance curve from each pixel is regarded as a time series, and it is represented by Hurst and Lyapunov exponents. These two new features are then used to overcome the Hughes problem for reliable classification. Experimental results show that the proposed method performs better than a few other feature extraction methods tested.
Jihao Yin, Xiuping Jia
IEEE Geosci. Remote. Sens. Lett.1
2012 A New Dimensionality Reduction Algorithm for Hyperspectral Image Using Evolutionary Strategy
abstract
Reducing the redundancy of spectral information is an important technique in classification of hyperspectral image. The existing methods are classified into two categories: feature extraction and band selection. Compared with the feature extraction, the band selection method preserves most of the characteristics of the original data without losing valuable details. However, the choice of the effective band remains challenging, especially when considering the computational burden, which makes many enumerative methods infeasible. Recently, immune clonal strategy (ICS) has been applied to solve complex computation problems. The major advantages of algorithms based on ICS are that they are highly paralleled, distributed, adaptive, and self-organizing. Therefore, in this paper, we convert the band selection problem into an optimization issue and propose a new algorithm, ICS-based effective band selection (ICS-EBS), to select effective band combinations. Then, the selected bands are used in classification of hyperspectral image. We evaluated the proposed algorithm by using two data sets collected from the Washington DC Mall and Northwest Tippecanoe County. ICS-EBS was compared against one latest proposed band selection algorithm, interclass separability index Algorithm (ICSIA). We also compared the results with those achieved by other stochastic algorithms such as genetic algorithm (GA) and ant colony optimization (ACO). The experimental results indicate that our proposed algorithm outperforms ICSIA, GA-EBS, and ACO-EBS for hyperspectral image classification.
Jihao Yin, Jiankun Hu
IEEE Trans. Ind. Informatics1
2006 Chinese Organization Name Recognition Using Chunk Analysis
Jihao Yin, Xiaozhong Fan, Jiangde Yu
PACLIC1