VLDB 2026 Research / reviewers in the wild / expert
Liang Xiao 0001
dblp:x/LiangXiao1
· DBLP profile ↗
214ranked-venue papers
8as first author
127since 2021 · last 2026
0009-0003-4111-3781ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 122 · 85 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 3 first-author · 24 since 2021Artificial intelligence and machine learning · 32 · 2 first-author · 18 since 2021Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIRA: Evaluating Multimodal AI on Complex Clinical Reasoning in Interventional RadiologyabstractWe present MIRA (Multimodal Interventional RAdiology evaluation), a comprehensive benchmark for evaluating large multimodal models in expert-level interventional radiology tasks requiring specialized domain knowledge and advanced visual reasoning capabilities. Unlike existing medical benchmarks that primarily provide binary labels without contextual depth, MIRA offers diverse question formats, including open-ended, closed-ended, single-choice, and multiple-choice categories, each accompanied by detailed expert-validated explanations. The benchmark incorporates approximately 184K high-quality medical images spanning multiple imaging modalities with 1.2M meticulously generated question-answer pairs across various anatomical regions. These pairs were created through a sophisticated cascade methodology involving expert interventional radiologists at both the data collection and validation stages. Our comprehensive evaluation, encompassing zero-shot testing and fine-tuning experiments of large multimodal models, revealing significant performance gaps between AI systems and human specialists. Fine-tuning experiments demonstrate substantial improvements, with models achieving up to 0.80 accuracy on single-choice questions. MIRA establishes a challenging benchmark that suggests promising directions for developing specialized clinical AI systems for interventional radiology. Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Yuxuan Sun 0002, Yixuan Si, Lin Yang 0002, Liang Xiao 0001 |
AAAI | 10 |
| 2026 | Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent DiffusionabstractDenoising Diffusion Probabilistic Models (DDPMs) have shown success in robust 3D object detection tasks. Existing methods often rely on the score matching from 3D boxes or pre-trained diffusion priors. However, they typically require multi-step iterations in inference, which limits efficiency. To address this, we propose a Robust single-stage fully Sparse 3D object Detection Network with a Detachable Latent Framework (DLF) of DDPMs, named RSDNet. Specifically, RSDNet learns the denoising process in latent feature spaces through lightweight denoising networks like multi-level denoising autoencoders (DAEs). This enables RSDNet to effectively understand scene distributions under multi-level perturbations, achieving robust and reliable detection. Meanwhile, we reformulate the noising and denoising mechanisms of DDPMs, enabling DLF to construct multi-type and multi-level noise samples and targets, enhancing RSDNet robustness to multiple perturbations. Furthermore, a semantic-geometric conditional guidance is introduced to perceive the object boundaries and shapes, alleviating the center feature missing problem in sparse representations, enabling RSDNet to perform in a fully sparse detection pipeline. Moreover, the detachable denoising network design of DLF enables RSDNet to perform single-step detection in inference, further enhancing detection efficiency. Extensive experiments on public benchmarks show that RSDNet can outperform existing methods, achieving state-of-the-art detection. Wentao Qu, Guofeng Mei, Jing Wang 0201, Yujiao Wu, Xiaoshui Huang, Liang Xiao 0001 |
AAAI | 6 |
| 2026 | Gradient-Guided RC Weighting for Timing-Driven Global RoutingabstractAs a critical step in electronic design automation (EDA), global routing provides a guide to subsequent steps and provides valuable feedback to previous steps, including congestion, timing, and power estimation. However, given the complexity of timing and power calculation, it is difficult to estimate the impact on timing and power during the routing process. To address this issue, we propose a gradient-guided framework that computes the ''capacity sensitivity'' and ''resistance sensitivity'' of each segment to estimate their influence on the timing objectives. Integrating these two values as weights to constrain the changes in capacitance and resistance of the wire segments, we develop a timing-driven global router with superior performance. Power is also considered by optimizing the cells' switching power. Tested on ISPD25 Contest benchmarks, we can achieve 14.3% and 18.5% improvements in worst negative slack and total negative slack, respectively, with comparable congestion. With power optimization, we can further improve switching power by 10.6%. Liang Xiao 0001, Qinkai Duan, Leilei Jin, Tsung-Yi Ho, Evangeline F. Y. Young, Martin D. F. Wong |
ISPD | 1 |
| 2026 | OpenSGG-VL: Open-Vocabulary 3DSGG with Orthogonal Residual Fusion and Iterative Relation RefinementabstractOpen-vocabulary 3D scene graph generation (3DSGG) aims to predict object categories and relation triplets from 3D scans while generalizing beyond a fixed label set. Prior open-vocabulary methods commonly rely on 2D features as an intermediate bridge to connect 3D representations with language, which can introduce misalignment and unstable multimodal fusion. In this work, we propose OpenSGG-VL, a framework that learns text-aligned 3D instance embeddings via large-scale 3D-Text contrastive learning, further enhanced by lightweight pose features to retain spatial cues. At inference, we extract 2D embeddings from RGB views and fuse them with 3D embeddings using Orthogonal Residual Fusion (ORF), which preserves the dominant semantic direction while injecting complementary geometric residuals. For open-set relation prediction, we employ an LLM and formulate relation inference as an iterative, scene-consistent refinement process with grouped relation decoding and multi-stage optimization. Experiments on 3DSSG demonstrate that our method achieves strong open-vocabulary performance and produces more coherent scene graphs than prior baselines. Liang Xiao 0001, Zhiyong Su |
ICMR | 3 |
| 2026 | CASA-SDF: Curriculum-aware spatial adaptation with curvature-guided density for neural implicit surface reconstruction
Zhiyong Su, Liang Xiao 0001 |
Neurocomputing | 4 |
| 2026 | Composite fractal scanning enhanced mamba for small object detection in remote sensing imagery
Wenjing Zhan, Yongke Li, Fang Liu 0034, Liang Xiao 0001 |
Neurocomputing | 4 |
| 2026 | D&D-Net: A diffusion and deep priors regularized network for hyperspectral reconstruction
Jingxiang Yang, Tian Lin 0001, Wenxiu Diao, Fang Liu 0034, Jia Liu 0020, Hongyi Liu 0001, Liang Xiao 0001 |
Signal Process. | 7 |
| 2026 | A Novel Panchromatic-Guided Tensor Low-Rank Model for Multispectral Image SharpeningabstractIn this letter, based on tensor modeling, we propose a novel panchromatic (Pan)-guided tensor low-rank (PGTLR) model for multispectral image (MSI) sharpening, which aims to fuse the low resolution (LR) MSI and Pan image to output the high resolution (HR) MSI. On one hand, we novelly exploit the tensor low-fibered-rank prior of HR MSI to model its global three-dimensional spatial-spectral correlations, which is constructed as the tensor nuclear norm (TNN) prior term. On the other hand, we further novelly exploit the Pan-guided tensor low-fibered-rank prior to model the spatial link between HR MSI and Pan, which is constructed as the novel Pan-guided TNN prior term. Furthermore, the proposed PGTLR model is optimized by an efficient alternative algorithm. Moreover, the experimental results on reduced-scale and full-scale datasets quantitatively and visually validate the superiority of PGTLR. Pengfei Liu 0002, Yihang Du, Nan Huang 0001, Zhizhong Zheng, Liang Xiao 0001 |
IEEE Signal Process. Lett. | 6 |
| 2026 | InstantGR: Scalable GPU Parallelization for 3-D Global RoutingabstractGlobal routing plays a crucial role in electronic design automation (EDA), serving not only as a means of optimizing routing but also as a tool for estimating routability in earlier stages such as logic synthesis and physical planning. However, these scenarios often require global routing on unpartitioned large designs, posing unique challenges in scalability, both in terms of runtime and design size. To tackle this issue, this paper introduces useful techniques for parallelizing large-scale global routing that can significantly increase parallelism and thus reduce runtime. We also propose a new flexible layer transition technique to increase the flexibility and routing quality of directed acyclic graph (DAG) routing. Building upon these techniques, we have developed an open-source GPU-based global router that achieves state-of-the-art results in the latest ISPD’24 Contest benchmarks, thereby showcasing the effectiveness of our methods. Liang Xiao 0001, Shiju Lin, Qinkai Duan, Tsung-Yi Ho, Evangeline F. Y. Young |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | Spectral-Guided Multiscale Feature-Aware Transformer for Hyperspectral Image ClassificationabstractTransformer-based methods have recently shown remarkable success in hyperspectral image classification (HSIC). However, their applications, in practice, still face two significant challenges. First, although the multihead mechanism in self-attention improves model robustness during training, it may overlook the continuity of spectral bands. Second, existing methods often struggle to effectively balance global and local information during multiscale feature extraction, limiting further improvements in classification performance. To address these issues, we propose a novel spectral-guided multiscale feature-aware Transformer (SMFAT) framework for HSIC. Specifically, a global low-rank spectral learning (GLSL) module is introduced to project hyperspectral image patches into a low-rank subspace, reducing spectral redundancy and capturing global spectral correlations. Furthermore, we introduce the multiscale feature-aware self-attention (MFASA) mechanism, which dynamically integrates fine- and coarse-grained features to enhance multiscale feature modeling. Finally, a spectral-guided fusion (SGF) module leverages the global spectral information extracted by the GLSL module to guide MFASA in more effectively capturing interspectral correlations and spectral continuity. This approach facilitates a more effective integration of spectral and spatial features in HSIs. Experiments on three well-known HSI datasets verify that the proposed SMFAT method significantly outperforms several state-of-the-art approaches in real-world HSIC tasks. The source code for this work is available at https://github.com/stellaZ77/SMFAT. Zhenqiu Shu, Kexin Zeng, Songze Tang, Zhengtao Yu 0001, Liang Xiao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2026 | MALT: ML Assisted Shallow-Light Tree ConstructionabstractTiming is a critical issue in electronic design automation (EDA). To reduce the delay of a net, an important strategy is to minimize the path lengths from the source to the sinks. However, minimizing the path lengths will inevitably sacrifice the total wirelength. To balance the two objectives, researchers use shallow-light tree (SLT) to model and optimize the problem. In this article, we introduce MALT, a novel approach that uses a neural network to guide the construction of Steiner shallow-light trees. The constructed trees are further refined by a dynamic programming-based branch merging algorithm, which improves the wirelength without sacrificing the path lengths of any sinks. Our experimental results demonstrate that the proposed framework achieves significant improvements over both state-of-the-art traditional SLT generation algorithms and existing machine learning enhanced methods. Liang Xiao 0001, Qijing Wang, Evangeline F. Y. Young, Martin D. F. Wong |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2025 | An End-to-End Robust Point Cloud Semantic Segmentation Network with Single-Step Conditional Diffusion ModelsabstractExisting conditional Denoising Diffusion Probabilistic Models (DDPMs) with a Noise-Conditional Framework (NCF) remain challenging for 3D scene understanding tasks, as the complex geometric details in scenes increase the difficulty of fitting the gradients of the data distribution (the scores) from semantic labels. This also results in longer training and inference time for DDPMs compared to non-DDPMs. From a different perspective, we delve deeply into the model paradigm dominated by the Conditional Network. In this paper, we propose an end-to-end robust semantic Segmentation Network based on a Conditional-Noise Framework (CNF) of DDPMs, named CDSegNet. Specifically, CDSegNet models the Noise Network (NN) as a learnable noise-feature generator. This enables the Conditional Network (CN) to understand 3D scene semantics under multi-level feature perturbations, enhancing the generalization in unseen scenes. Meanwhile, benefiting from the noise system of DDPMs, CDSegNet exhibits strong robustness for data noise and sparsity in experiments. Moreover, thanks to CNF, CDSegNet can generate the semantic labels in a single-step inference like non-DDPMs, due to avoiding directly fitting the scores from semantic labels in the dominant network of CDSegNet. On public indoor and outdoor benchmarks, CDSegNet significantly outperforms existing methods, achieving state-of-the-art performance. Wentao Qu, Jing Wang 0201, Yongshun Gong, Xiaoshui Huang, Liang Xiao 0001 |
CVPR | 5 |
| 2025 | Multi-scale Feature Interaction and Adaptive Experts for Panoptic Segmentation in Remote Sensing ImagesabstractPanoptic segmentation unifies the traditional tasks of instance and semantic segmentation. It plays a crucial role in the field of remote sensing; however, it encounters challenges in recognizing small objects and in the model’s ability to generalize across complex scenes. In this paper, we introduce the MFIAE framework to address two specific challenges: a multi-scale interactive attention fusion (MSIAF) module and an adaptive disturbance sparse mixture-of-experts (ADSMoE) module based on Transformer. The MSIAF module is designed to fully utilize the rich contextual information captured by low-resolution features while simultaneously utilizing the advantages of high-resolution features to enhance small object segmentation. In the ADSMoE module, adaptive noise is introduced to disturb the expert selection in order to enhance the randomness and exploration of the model when it comes to selecting experts, thereby improving its capacity for generalization and robustness. Additionally, it also reduces the computational overhead and the model complexity. The experimental results demonstrate that our approach achieves state-of-the-art performance on the BSB Aerial dataset. Zhenkun Sun, Jia Liu 0020, Jingxiang Yang, Liang Xiao 0001 |
ICASSP | 6 |
| 2025 | Mask-guided Multi-scale Spatial-Spectral Transformer for Snapshot Compressive ImagingabstractEffectively reconstructing 3D hyperspectral images (HSIs) from 2D measurements presents a significant challenge in Coded Aperture Snapshot Spectral Imaging (CASSI) systems. While recent transformers exhibit potential in HSI reconstruction, they often suffer from inadequate exploration of multi-scale spatial-spectral self-similarity, leading to mean effects and information loss. Additionally, these methods struggle with insufficient modeling of the degradation inherent in the compressive imaging process. To address these issues, we propose a novel Mask-guided Multi-scale Spatial-Spectral Transformer (MMSST). Specifically, we introduce a Degradation Aware Mask Attention (DAMA) module to incorporate degradation information of the compressive imaging process. Furthermore, MMSST leverages Local-Regional SpAtial attention (LRSA) and Global-Regional SpEctral attention (GRSE) to effectively exploit multi-scale self-similarity across spatial and spectral dimensions. Extensive experimental results demonstrate the effectiveness of our MMSST. Heyuan Yin, Jingxiang Yang, Jia Liu 0020, Liang Xiao 0001 |
ICASSP | 4 |
| 2025 | Part in Part Embedding Network for Zero-Shot LearningabstractZero-shot learning (ZSL) seeks to utilize semantic information from seen classes encountered during training to effectively recognize unseen classes during testing. When dealing with fine-grained images, capturing local features heavily influences the accuracy of semantic descriptions. Meanwhile, local features are represented at various scales across different layers of a neural network, making it hard to capture their local details fully. To address these challenges, we propose a novel part in part embedding network, termed PPEN. Specifically, PPEN consists of two key modules: the cross-layer aggregation (CLA) module and the part in part attention (PIPA) module. The CLA module is designed to fuse and preserve features from multiple layers of the network, thereby maintaining the richness of information across different scales. Further, the PIPA module focuses on identifying local features that are most pertinent to the class semantic vectors, enhancing the alignment between visual features and semantic descriptions. We evaluate our approach on three ZSL benchmarks, i.e., CUB, SUN, and AWA2, and demonstrate the superiority and competitiveness of our proposed approach. Code is available at https://github.com/zhou834177226/PPENet. Zhexian Zhou, Liang Xiao 0001, Guosen Xie |
ICASSP | 2 |
| 2025 | ChronoTE: Crosstalk-Aware Timing Estimation for Routing Optimization via Edge-Enhanced GNNsabstractAccurate timing estimation during the routing stage is critical for modern VLSI design closure, especially under increasing crosstalk effects in advanced technology nodes. During the routing process, the crosstalk effect is usually modeled by predicting coupling capacitance with congestion information. However, such estimations are often overly pessimistic, as crosstalk-induced delay is influenced not only by coupling capacitance but also by the relative arrival times of signals. In this work, we propose ChronoTE, a novel edge-enhanced graph neural network (GNN) framework that performs crosstalk-aware net delay estimation by jointly modeling physical topology and timing characteristics. By embedding timing-window-aware features into edge representations, ChronoTE enables accurate delay prediction without requiring full routing or parasitic extraction. Experimental results on industrial-scale open-source designs demonstrate that ChronoTE, by delivering sign-off quality delay estimation in the early global routing stage, significantly accelerates design closure and contributes to area reduction. Leilei Jin, Rongliang Fu, Zhen Zhuang, Liang Xiao 0001, Fangzhou Liu 0005, Bei Yu 0001, Tsung-Yi Ho |
ICCAD | 4 |
| 2025 | Multi-Scale Tubularity-Aware U-NetabstractU-Net architectures have made great progress in dealing with semantic segmentation tasks. However, existing frameworks have not yet possessed the ability of capturing sufficient local and contextual dependencies of tubular structures. The reasons are two-fold. First, traditional square convolutions are inherently limited to model irregular pixel changes due to their fixed geometric structures. Second, there exist semantic gaps among the multi-scale tubularity features and between stages of their encoding and the decoding. To mitigate these issues, we propose multi-scale tubularity-aware U-Net, by coupling a novel tubularity deformable convolution (TdConv) embedding and a dual attention Transformer (DaTrans) alternative to skip connection. On the one hand, TdConv embedding iteratively learns the deformation offsets of convolution itself in both directions along the tubular structure. On the other hand, DaTrans connection endows skip connections with attention mechanism from both multi-scale local pixel and cross-scale global semantic perspectives. Hinging on the local irregularity perception and the global semantic association, our method enables to analyze tubular structures appeared in complex contexts and at different scales. Extensive experiments show that our approach outperforms state-of-the-art techniques, including different U-Net variants, for various datasets on several tasks including road extraction and vessel segmentation. Jie Song 0014, Ziyun Cai, Liang Xiao 0001, Yawen Huang |
ICME | 5 |
| 2025 | Invited: AI-assisted RoutingabstractRouting is an important but complicated step in physical synthesis. Considering the potential of leveraging AI to seek higher efficiency and better quality in solving routing problems, we study in this work the methodology of AI-assisted routing in a systematic way. Decoupling the functionalities of different routing components will give a high flexibility in determining where and how AI can be used in an effective manner, while maintaining a high degree of interpretability. Two applications along this direction are presented, aiming at tackling the difficulties in routing with AI assistance. These provide examples of how to implement the methodology in practice, while revealing its effectiveness and potential. Qijing Wang, Liang Xiao 0001, Evangeline F. Y. Young |
ISPD | 2 |
| 2025 | Simultaneous Vision-Language Knowledge Transfer for Zero-Shot Human-Object Interaction Detection
Zhiyong Su, Liang Xiao 0001 |
PRCV (12) | 4 |
| 2025 | Adversarial purification of information maskingabstractAdversarial attacks meticulously generate minuscule, imperceptible perturbations that add to images to deceive neural networks . Adversarial purification methods seek to remove perturbations using generative models to achieve defense. However, residual perturbations lead to less-than-ideal results. Under the premise that perturbations are difficult to remove completely, we are the first to quantify the hazards of residual perturbations and explore how to achieve more robust defenses by reducing perturbations and resisting the impact of residual perturbations. Motivated by this, we propose a novel adversarial purification approach named Information Mask Purification (IMPure). Our method utilizes informative masks and a regional intersection reconstruction to generate images to reduce the perturbation residues. During training, we use the combination module to guide the generative model in recovering feature representations. Finally, we establish a combined constraint of pixel loss and perceptual loss to augment the model’s reconstruction adaptability. Extensive experiments on the complex dataset ImageNet with classifier models demonstrate that our approach achieves state-of-the-art results in defending against adversarial attack methods. Implementation code and pre-trained weights can be accessed at https://github.com/NoWindButRain/IMPure . Zhichao Lian, Shuangquan Zhang, Liang Xiao 0001 |
Neurocomputing | 4 |
| 2025 | Segmentation of 3D neuronal morphologies in microscopy images utilizing flexible open-curve snakes
Amir Vatani, Jie Song 0014, Liang Xiao 0001 |
Neurocomputing | 3 |
| 2025 | Tracking Mamba for Road Extraction From Satellite ImageryabstractAutomated road extraction from satellite imagery for dynamic map updating has become a crucial research focus in remote sensing, where existing state-of-the-art Transformer-based methods exhibit two critical limitations: (1) inadequate precision in capturing tubular road patterns and (2) suboptimal computational efficiency on standard GPUs. To address these challenges, we propose TrMamba, a new Tracking-based Mamba architecture that combines the original Mamba’s efficiency with enhanced tubular road pattern recognition through two key innovations: a tubular road pattern tracking mechanism for continuous road feature extraction and a tracking selective scanning module for directional context modeling via adaptive attention. By integrating a novel tubular tracking mechanism into the Mamba’s selective scanning process, TrMamba fundamentally improves the original paradigm and achieves superior road topology encoding, as demonstrated by extensive experiments showing state-of-the-art performance across multiple remote sensing benchmarks in both accuracy and computational efficiency. The source code is available at: https://github.com/Apheliosa/TrMamba. Jie Song 0014, Ziyun Cai, Liang Xiao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | A boundary evidence controlled level set inference method for nuclei instance segmentation in histopathology images
Amir Vatani, Jie Song 0014, Liang Xiao 0001 |
Multim. Tools Appl. | 3 |
| 2025 | Deep one-class probability learning for end-to-end image classification
Jia Liu 0020, Jingxiang Yang, Liang Xiao 0001 |
Neural Networks | 5 |
| 2025 | Diffusion-Based Adversarial Purification With Feature DistillationabstractAdversarial purification is a defense strategy that utilizes generative models to neutralize adversarial perturbations. Diffusion models stand out for their powerful generative ability, making them the latest choice for generative models in adversarial purification methods. It is crucial to keep the balance between robustness against attacks and the integrity of the image content. However, these diffusion models are typically trained exclusively on clean data. We experimentally demonstrate that in the adversarial purification task, the data distribution operated by the diffusion model is often contaminated by adversarial perturbations. Adjusting the diffusion length to effectively mask adversarial perturbations while preserving image label semantics proves challenging. In this paper, we propose an innovative distillation-based diffusion approach for adversarial purification, enabling the diffusion model to operate effectively on the contaminated data distribution. Our approach involves a novel training process that integrates adversarial samples into the diffusion model's training by considering adversarial perturbations as part of the diffusion noise predicted by the model. The key to resisting perturbations is to align the feature representation of the purified image more closely with that of the clean image. To accomplish this, we develop a feature distillation technique to empower the diffusion model to learn and extract clean feature representations from adversarial samples. Extensive experimentation shows that our method achieves state-of-the-art performance against various adaptive attack benchmarks. Zhichao Lian, Liang Xiao 0001 |
IEEE Trans. Big Data | 3 |
| 2025 | FaceGCN: Structured Priors Inspired Graph Convolutional Networks for Face Restoration With Unknown DegradationsabstractFacial image restoration has gained a tremendous progress since the increasing boom of the deep learning methods. Owing to its nature of strong ill-posedness, different categories of a-priori constraints have been harnessed or embedded in the existing deep architectures. While, as it turns to blind face restoration with more complicated degradations, the challenge becomes greater. In this paper, a further insightful step is taken by exploring the potentials of the graph convolutional networks (GCN) in conjunction with the structured priors for the blind problem. Specifically, a lightweight yet physically more intuitive model termed FaceGCN is proposed. On the one hand, a dynamic generator of facial adjacency matrices is constructed assisted by two self-supervised losses, allowing a sparse, accurate, and adaptive construction of case-specific face graphs with facial feature components as nodes. On the other hand, to model well the joint local-nonlocal correlations among various facial feature components, a kind of novel strip-attention GCN modules is correspondingly developed by splitting facial feature maps into intra- and inter-strips in both horizontal and vertical orientations, respectively. Extensive experimental results show that FaceGCN has achieved comparable or even superior performance to state-of-the-art methods, yet at a considerably less computational cost. Weidan Yan, Wenze Shao, Dengyin Zhang, Liang Xiao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Hierarchical Contrastive Learning for Multigranularity Ship Classification With Learnable Class QueriesabstractShip targets in remote sensing images can be categorized at various granularities due to variations in image quality, ranging from general ship categories to fine-grained classes like Nimitz-class carriers. Traditional studies mainly focus on fine-grained ship classification, often neglecting samples observed at coarser-grained levels. Samples distributed across multiple granularity levels exhibit semantic relationships among their annotated classes, enabling hierarchical knowledge transfer during model training. This paper incorporates two semantic relationships into deep learning-based representation learning and class prediction: parent-child relationships across levels and mutual exclusivity among sibling categories. For hierarchical representation learning, the proposed hierarchical contrastive learning algorithm extracts category-specific representations from input images and aligns them with their semantic relationships, ensuring that parent and child categories share similarities while sibling categories remain distinct. For hierarchical class predictions, a novel consistency loss ensures coherence in probability distributions between parent and child categories. Specially, cross-entropy loss is employed to impose mutual exclusivity among sibling categories. In this paper, a multi-modal dataset is also designedly developed for hierarchical classification, which integrates optical and synthetic aperture radar (SAR) images across multiple hierarchical levels. Experiments on two popular datasets and a multi-modal dataset demonstrate that the proposed method outperforms state-of-the-art approaches in hierarchical multi-granularity ship classification. Jingzhou Chen, Fengchao Xiong, Yuntao Qian, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Dual-Conditionally Guided Diffusion Models for Fusion of Unregistered Multisource Remote Sensing ImagesabstractRemote sensing image fusion is a critical technique for enhancing the quality of remote sensing images. Typically, it is presumed that images have been accurately registered to facilitate effective fusion. Traditional methods involve a sequential process of registration followed by fusion, or a parallel approach to registration and fusion, which do not fully leverage the interplay between these two stages. To overcome this limitation, we propose a novel joint learning method termed the Dual-Conditionally Guided Registration and Fusion Diffusion Network (RFDifNet) for unregistered remote sensing images. This method integrates image registration and fusion into a unified framework. The RFDifNet comprises two conditional diffusion-based subnetworks: the Registration Diffusion Module (RegDM) and the Fusion Diffusion Module (FusDM). In this architecture, the RegDM corrects the misalignment of the unregistered image and provides it as a conditional input to the FusDM to generate the fused image. Conversely, the fused image is also fed back into the RegDM as a conditional input, enabling a closed-loop iteration of the registration and fusion processes. To further refine the image registration by reducing noise interference and preserving edge details within the RegDM, we introduce a joint learning loss function based on fractional-order derivatives, demonstrating superior performance in geometric and detail preservation compared to traditional methods. Experimental results validate the outstanding performance of the proposed RFDifNet in both image registration and fusion tasks. The source code is available at: https://github.com/DDXNJUST/RFDifNet. Wenxiu Diao, Ling Hu 0003, Kai Zhang 0010, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Multimodal Feature Interactive Learning for Few-Shot Hyperspectral Image ClassificationabstractRecently auxiliary cross-scene information has been widely utilized to improve the hypersperctral image classification performance by knowledge transfer. However, recognition of different objects with the same semantic category is difficult when the object types in similar scenes are different or only limited similarity knowledge is provided. In this paper, a multi-modal feature interactive learning (MMFI) method is proposed based on both hyperspectral image modality and textual modality to distinguish similar objects, which enhances the transfer capability by utilizing the semantic prior from the textual modality. First, the adversarial domain mapping (ADM) module is designed to realize cross-domain knowledge transfer across different scenes in an adversarial learning manner. In particular, the noise is simulated as data distribution in different domains through domain mapping and aggregated with source and target domain data, which is then reconstructed and optimized to learn discriminative and conducive information for transfer. Then, the adaptive interactive learning (AIL) module acts on the latent features of the encoder to mine latent associations among the aggregated features and facilitate the expression of consistent features. In addition, few-shot learning with textual embedding enables more powerful semantic priors for few-shot prototypes, making up for insufficient recognition capability in the presence of hyperspectral image modality only. Experimental results on three datasets demonstrate the superiority of our method. Fang Liu 0034, Wenfei Gao, Jia Liu 0020, Xu Tang 0004, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Edge-Object Co-Driven Learning for Remote Sensing Change DetectionabstractRemote sensing change detection (CD) aims to accurately reveal surface changes by comparing two temporally separated images of the same area. However, in complex environments, insufficient edge detail recognition and limited feature extraction often affect the accuracy of CD. For this purpose, we propose a novel method named the edge-object co-driven learning network (EOCLNet), which employs a combination of the Pyramid Vision Transformer (PVT) and the Fast Segment Anything Model (FastSAM) as parallel feature extractors to capture rich multilevel features. Specifically, it includes three key components which are the edge extraction module (EEM), the object revelation module (ORM), and the edge-object learning (EOL). EEM explicitly captures edge details by combining low-level spatial features with high-level semantic features, providing essential edge knowledge. ORM reveals changed objects by aggregating the highest two levels of semantic features, providing initial change guidance. EOL is designed to implicitly mine edge clues by establishing relationships between edges and changed objects across multiple levels, receiving outputs from both EEM and ORM. Furthermore, during the training process, the uncertainty from the previous level’s change map is utilized to guide the learning at the next level, thereby achieving a transition from uncertainty to certainty. The effectiveness of EOCLNet is validated on three public datasets, where it outperforms several state-of-the-art CD methods. Yangguang Liu, Fang Liu 0034, Jia Liu 0020, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Gradient Subspace-Regularized Hyperspectral Image and Stripe-Coupled Nonconvex Tensor Low-Rank Priors for Destriping and DenoisingabstractIn this article, we propose a novel, unified, and effective hyperspectral image (HSI) destriping and denoising method with gradient subspace-regularized HSI and stripe-coupled nonconvex tensor low-rank priors (GSHSNTLRs). First, by exploiting the mode-3 low-rank properties of the gradients of HSI (i.e., the spatial horizontal gradient, spatial vertical gradient, and spectral gradient of HSI) along the spectral dimension, we apply the mode-3 low-rank decomposition of the gradients of HSI to obtain their representation coefficient tensors (RCTs), and further study the tensor low-tubal-rank properties of the RCTs in the gradient subspace. Thus, we propose the unified gradient subspace-regularized log tensor nuclear norm (LogTNN)-based nonconvex tensor low-rank prior term of the RCTs. Moreover, by fully considering the structural speciality of stripe noise, which has strong tensor low-tubal-rank property, we particularly study the HSI-guided tensor low-rank modeling for the stripe noise by exploring the tensor low-tubal-rank property of HSI plus stripe and propose the unified HSI and stripe-coupled LogTNN-based nonconvex tensor low-rank prior term of HSI and stripe simultaneously. Subsequently, the proposed GSHSNTLR model is solved by using the alternating direction method of multipliers (ADMMs). Finally, lots of experimental results and analysis fully demonstrate the destriping and denoising performance and superiority of GSHSNTLR. Pengfei Liu 0002, Haijian Long, Zhizhong Zheng, Nan Huang 0001, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Multiscale Self-Supervised Constraints and Change-Masks-Guided Network for Weakly Supervised Change DetectionabstractRemote sensing change detection (CD) is a highly significant subtask within the field of Earth observation. Recently, weakly supervised CD (WSCD) methods based on image-level annotations have attracted interest, it is challenging to generate a clear margin between changed and unchanged regions with a lack of detailed annotation. In this article, based on class activation maps (CAMs), we propose a novel WSCD network based on self-supervised learning and change mask guidance (SSCMNet). First, we design a multiscale self-supervised constraint (MSC) module to narrow the gap between weak supervision and full supervision and compensate for the inherent shortcomings of CAMs. Second, a change mask guidance (CMG) module is proposed to further guide the network to keep the integrity of changed objects according to the consistency within unchanged regions and inconsistency within changed regions. Finally, to address the challenge of transferring commonly used post-processing methods in semantic segmentation to CD, an adaptive post-processing (APP) module is designed to adaptively select one of the input images for post-processing. We conduct experiments on three publicly available remote sensing CD datasets. Quantitative metrics and visualized results demonstrate the outstanding performance of the proposed method. Jia Liu 0020, Hejun Luo, Fang Liu 0001, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Hyperspectral Image Denoising and Destriping via Gradient Tensor Subspace Low-Rank Learning and Along-Across Stripe Directional ConstraintsabstractThis paper proposes a new hyperspectral image (HSI) denoising and destriping method via gradient tensor subspace low-rank learning and along-across stripe directional constraints (GTSL2A2SDC) under the unified framework of tensor representation modeling. On one hand, for the modeling of stripe noise, based on the inherently directional and structural attributes of the stripe noise, we mainly investigate the mode-1 gradient tensor of stripe along the stripe direction which holds the preferable “zero plane” constraint as well as the mode-2 gradient tensor of stripe across the stripe direction which holds the preferable tensor low-fibered-rank attribute along the mode-3 spectral dimension. Therefore, we novelly propose the unified along-across stripe directional gradient tensor constraints for the stripe noise, which can simultaneously characterize the directional and structural attributes of the stripe noise. On the other hand, for the modeling of HSI, based on the preferably spectral low-rankness attributes of the multi-mode gradient tensors of HSI, namely, mode-1 gradient tensor, mode-2 gradient tensor and mode-3 gradient tensor, we further utilize the spectral low-rank factorization of the gradient tensors of HSI to get the corresponding representation tensors, and particularly investigate the nonlocal self-similarities-based low-rankness of the representation tensors under the gradient tensor-based subspace low-rank learning framework. Moreover, we optimize the proposed GTSL2A2SDC model via an efficiently alternative and iterative algorithm. Lastly, extensive experiments comprehensively validate the denoising and destriping capacity and superiority of GTSL2A2SDC. Pengfei Liu 0002, Haijian Long, Zhizhong Zheng, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Box2Change: A Novel Weakly Supervised Way for Change Detection via Consistency Instance SegmentationabstractChange detection in remote sensing images aims at revealing interesting changes about the earth surface and has been one of the most important issues in earth observation. In recent years, lots of fully-supervised change detection methods have achieved good performance with the help of deep learning architectures, which rely on large amounts of pixel-level labels. However, obtaining high-quality pixel-level labels is laborious and expensive. To alleviate this problem, we propose a novel weakly-supervised change detection way via consistency instance segmentation called Box2Change, which requires only box-level labels and achieves competitive results to fully-supervised change detection method. Compared with pixel-level label, it is much more efficient to get box-level label, which locates the potential changed area by a rectangle box. There are two key components in the proposed method, the Changed Instance Segmentation (CIS) and the Self-Supervised Consistency Learning (SSCL) in affine space. The former generates multi-scale changed instances, which learns positional information from box-level labels and segments the instance boundaries within a given bounded region. The latter introduces affine transform and employs consistency constraints in a self-supervised manner to increases the robustness to pseudo-change situations caused by light or noise. In experiments, three popular public change detection datasets are tested and both visual and numerical assessment are discussed, where the proposed method exhibits competitive performance to fully-supervised methods and achieves the state-of-the-art results compared with the other weakly-supervised change detection methods. Fang Liu 0034, Kanghua Yin, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Background-Driven and Foreground-Refined Network for Weakly Supervised Change DetectionabstractChange detection (CD) in remote sensing aims to reveal meaningful surface changes and has been flourishing in recent years. Compared with fully-supervised methods based on pixel-level labels, image-level labels are easy to acquire, which reduces manual labor to a large extent. However, image-level labels lack spatial-and-shape information while containing the least semantic information, which poses a great challenge to the weakly-supervised CD task. Motivated by the prior that bi-temporal images have background semantic consistency, we propose Background-Driven and Foreground-Refined (BDFR-Net) to ameliorate the above problem. Specifically, there are two key components in the proposed method: the Background-Driven Reconstruction (BDR) with image-level supervision and the Foreground-Refined Learning (FRL) with affinity learning. The former generates changed regions of foreground and background separation, which activates the foreground from image-level supervision and constrains the foreground by maintaining spatial and semantic consistency in background regions. The latter introduces Complementary Fusion and Label Adaption (CFLA) strategies to further refine the foreground, which can mine complementary information from foreground sequences and suppress false activations. In addition, affinity learning is proposed to stabilize and supervise the above process. Complementary relationships between foreground and background are fully utilized. Tested on two popular CD datasets, the results demonstrate that our proposed BDFR-Net produces completely changed regions with clear boundaries and outperforms state-of-the-art weakly-supervised methods. Fang Liu 0034, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | A Biobjective Model-Driven Autocoder for Blind Hyperspectral UnmixingabstractHyperspectral unmixing decomposes hyperspectral images (HSIs) into pure spectral signatures (endmembers) and their proportions (abundances). Existing deep learning methods for this task predominantly focus on either linear or nonlinear relationships, often struggling to achieve an optimal equilibrium between the robustness of linear estimations and the precision of nonlinear models. Furthermore, these methods often neglect the integration of spatial information and the interpretability of the results. This article proposes a biobjective model-driven autoencoder network for blind hyperspectral unmixing that simultaneously addresses both linear and nonlinear relationships. By combining linear and nonlinear kernel models within a model-driven deep learning framework, we aim to enhance the interpretability of the results. To effectively capture spatial information, we introduce an Adaptive Composite Kernel that integrates traditional and spatial-spectral kernels, using a dynamic convolution-like mechanism to optimize their respective weights. Our method establishes linear and nonlinear reconstruction losses, enabling the simultaneous estimation of endmembers from both types of relationships, thus capturing the complex and high-dimensional data structures inherent in HSIs. Experiments on real and synthetic datasets demonstrate the superiority of the proposed method. Hongru Zong, Zebin Wu 0001, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | RSIC-GMamba: A State-Space Model With Genetic Operations for Remote Sensing Image CaptioningabstractRecent advancements indicate that the novel Mamba framework, with its linear reasoning capabilities and performance on par with the Transformer framework, has become a popular alternative to Transformers. However, there is still a lack of exploration in remote sensing image captioning (RSIC). One important challenge is that the 1-D selective state-space model (SSM) is not suitable for 2-D remote sensing images (RSIs). To mitigate this challenge, this article proposes an SSM with genetic operations for RSIC (RSIC-GMamba) that integrates the search capabilities of heuristic genetic algorithm and the global modeling capabilities of Transformer into the Mamba framework. To comprehensively capture the multiscale global context of RSIs, we design a genetic SSM by incorporating a dilated convolution, genetic operations (crossover and mutation), and self-attention mechanism into the selective SSM. Specifically, dilated convolutions are employed to extract multiscale visual features through varying dilation rates. Crossover operation expands the spatial arrangement of image regions to thoroughly capture contextual information and mutation operation introduces randomness to enhance the model’s robustness. The self-attention mechanism is integrated to model relationships among SSM hidden states, thereby enhancing visual context. In addition, to fully utilize high- and low-level visual-semantic information, we propose a vision and scene text aggregation (ViSTA) module based on gating mechanisms. Experimental results on four RSIC datasets demonstrate the effectiveness of the proposed RSIC-GMamba. The code will be publicly available athttps://github.com/One-paper-luck/RSIC-GMamba. Lingwu Meng, Jing Wang 0201, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Learning Cross-Task Features With Mamba for Remote Sensing Image Multitask PredictionabstractMultitask learning (MTL) for remote sensing (RS) image is a rapidly evolving field that requires simultaneous predictions across several related tasks. However, many existing MTL methods often overlook the exploring of cross-task features, while the strong interdependencies among tasks are critical for MTL. In this article, we propose RSMTMamba, an innovative MTL framework that integrates Mamba for multitask prediction in RS images. Our network simultaneously performs semantic segmentation, height estimation, and boundary detection within a unified architecture. The proposed architecture prioritizes the decoder, with a shared encoder for feature extraction. Specifically, a Mamba-based cross-task feature learning (MCFL) module is introduced to capture the interrelations among different tasks. Unlike transformer-based architecture, which requires significant computational resources, the MCFL module can model both local and global cross-task relationships for RS image with linear complexity. Additionally, Mamba-integrated refine decoders are utilized to aggregate features from the encoder, preliminary decoders, and the MCFL module, which enhances multitask prediction performance. The experimental results on three RS datasets demonstrate that our proposed CFLMamba achieves the state-of-the-art prediction performance, outperforming several deep neural networks in RS image analysis. The code is available athttps://github.com/sycs-2024/RSMultitaskMamba. Liang Xiao 0001, Jianyu Chen 0003, Qian Du 0001, Qiaolin Ye |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Topological Information Aggregation Network for Few-Shot Cross-Domain Hyperspectral Image ClassificationabstractIn recent advancements, hyperspectral image (HSI) classification through few-shot learning (FSL) has significantly progressed. Domain adaptation, integrated with FSL, effectively utilizes transferable knowledge from a source domain (SD) with abundant labeled data to excel in classification tasks within a target domain (TD) with scarce labels. However, most existing methods usually use traditional convolutional neural networks (CNNs) to extract local spatial information to characterize and mine feature and distribution information while ignoring the underlying topological relationships among feature classes. Therefore, we propose a topology graph perception cross-domain FSL (TGP-CFSL) framework that leverages graph information aggregation. Specifically, to construct the extended topological relationships of the target, we have designed a topological graph-based multiscale fusion (TGMF) feature extraction module, which is adept at fully mining the topological spatial neighborhood information of the target. Meanwhile, a dual-graph information perception (DGIP) module is designed, which is able to characterize and aggregate intradomain topological relationships in terms of both feature representations and interdomain distribution similarities and to extract higher order domain distribution information for realizing domain alignment. Experimental results on three public HSI datasets demonstrate that the proposed method outperforms existing methods. Kai Shi 0001, Wenzhen Wang, Qichao Liu, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | DUSA-UNet: Dual Sparse Attentive U-Net for Multiscale Road Network ExtractionabstractThe challenges of road network segmentation demand an algorithm capable of adapting to the sparse and irregular shapes, as well as the diverse context, which often leads traditional encoding-decoding methods and simple Transformer embeddings to failure. We introduce a computationally efficient and powerful framework for elegant road-aware segmentation. Our method, called DUSA-UNet, effectively encodes fine-grained local road connectivity and holistic global topological semantics while decoding multiscale road network information. DUSA-UNet offers a novel alternative to the U-Net architecture by integrating connectivity attention, which can exploit intra-road interactions across multi-level sampling features with reduced computational complexity. This local interaction serves as valuable prior information for learning global interactions between road networks and the background through another integrality attention mechanism. The two forms of sparse attention are arranged alternatively and complementarily, and trained jointly, resulting in performance improvements without significant increases in computational complexity. Extensive experiments on various datasets with different resolutions, including Massachusetts, DeepGlobe, SpaceNet, and Large-Scale remote sensing images, demonstrate that DUSA-UNet outperforms state-of-the-art techniques. Our approach represents a significant advancement in the field of road network extraction, providing a computationally feasible solution that achieves high-quality segmentation results. Jie Song 0014, Ziyun Cai, Liang Xiao 0001, Yawen Huang, Yefeng Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Text-Driven Adaptive Semantic Alignment Network for Cross-Scene Hyperspectral Image ClassificationabstractLand cover in different scenes generally exhibits scene-invariant category semantic, typically represented and described consistently in a textual modality. Traditional cross-scene classification methods often treat categories as discrete class labels, neglecting their semantic information, or use category names merely as auxiliary textual modalities to enhance the discriminative representations of land cover. However, the cross-scene consistency of category semantic for land cover remains underexplored and underutilized. To address this issue, the text-driven adaptive semantic alignment network (TASA-Net) is proposed in this article for cross-scene hyperspectral image classification (HSIC). TASA-Net employs hand-crafted template prompts for stable category descriptions and vision-guided fine semantic prompts (VG-FSPs) for dynamic scene adaptation. Through a dual-gated adaptive mechanism, TASA-Net optimally weights coarse- and fine-grained semantics in a shared space, ensuring stable yet discriminative semantic representation. Additionally, cross-modal semantic alignment projects visual features into the shared semantic space, while a soft alignment strategy dynamically adjusts category correlations to enhance intraclass consistency and mitigate domain shifts. Ultimately, by leveraging text-driven semantic consistency representation, TASA-Net achieves zero-shot cross-scene transfer for unsupervised classification. Experiments demonstrate superior performance across multiple hyperspectral datasets, validating the critical role of textual modality in enhancing model robustness and cross-scene generalization ability. Wenzhen Wang, Fang Liu 0034, Hongyuan Zhu 0002, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Boosting Adversarial Transferability via Relative Feature Importance-Aware AttacksabstractModern deep neural networks are known highly vulnerable to adversarial examples. As a pioneering work, the fast gradient sign method (FGSM) is proved more transferable in black-box attacks than its multi-small-step extension, i.e., iterative-FGSM, particularly being restricted by a limited number of iterations. This paper revisits their early, representative successor MI-FGSM as a baseline, i.e., iterative-FGSM with momentum, and introduces an innovative boosting idea different from either FGSM-inspired algorithms or other mainstream methods. For one thing, during gradient backpropogation of MI-FGSM, the proposed approach merely requires amending the chain rule with respect to adversarial images using the counterpart original images. For another, a credible analysis has revealed that such a naively boosted MI-FGSM essentially performs a special kind of intermediate-layer attacks. In specific, the notable finding in the paper is a new principle of adversarial transferability guided by the relative feature importance, emphasizing the significance of semantically non-critical information for the first time in the literature, although originally thought to be weak in large. Experimental results on various leading victim models, both undefended and defended, demonstrate that the new approach incorporating robust gradients has indeed attained stronger adversarial transferability than state-of-the-art works. The code is available at:https://github.com/ljwooo/RFIA-main. Wenze Shao, Yubao Sun, Li-Qian Wang, Qi Ge, Liang Xiao 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | A Conditional Denoising Diffusion Probabilistic Model for Point Cloud UpsamplingabstractPoint cloud upsampling (PCU) enriches the representation of raw point clouds, significantly improving the performance in downstream tasks such as classification and reconstruction. Most of the existing point cloud upsampling methods focus on sparse point cloud feature extraction and upsampling module design. In a different way, we dive deeper into directly modelling the gradient of data distribution from dense point clouds. In this paper, we proposed a conditional denoising diffusion probabilistic model (DDPM) for point cloud upsampling, called PUDM. Specifically, PUDM treats the sparse point cloud as a condition, and iteratively learns the transformation relationship between the dense point cloud and the noise. Simultaneously, PUDM aligns with a dual mapping paradigm to further improve the discernment of point features. In this context, PUDM enables learning complex geometry details in the ground truth through the dominant features, while avoiding an additional upsampling module design. Furthermore, to generate high-quality arbitrary-scale point clouds during inference, PUDM exploits the prior knowledge of the scale between sparse point clouds and dense point clouds during training by parameterizing a rate factor. Moreover, PUDM exhibits strong noise robustness in experimental results. In the quantitative and qualitative evaluations on PU1K and PUGAN, PUDM significantly outperformed existing methods in terms of Chamfer Distance (CD) and Hausdorff Distance (HD), achieving state of the art (SOTA) performance. Wentao Qu, Yuantian Shao, Lingwu Meng, Xiaoshui Huang, Liang Xiao 0001 |
CVPR | 5 |
| 2024 | InstantGR: Scalable GPU Parallelization for Global RoutingabstractGlobal routing plays a crucial role in electronic design automation (EDA), serving not only as a means of optimizing routing but also as a tool for estimating routability in earlier stages such as logic synthesis and physical planning. However, these scenarios often require global routing on unpartitioned large designs, posing unique challenges in scalability, both in terms of runtime and design size. To tackle this issue, this paper introduces useful techniques for parallelizing large-scale global routing that can significantly increase parallelism and thus reduce runtime. Building upon these techniques, we have developed an open-source GPU-based global router that achieves the state-of-the-art results in the latest ISPD'24 Contest benchmarks, thereby showcasing the effectiveness of our methods. The source code of this work is available at https://github.com/cuhk-eda/InstantGR. Shiju Lin, Liang Xiao 0001, Evangeline F. Y. Young |
ICCAD | 2 |
| 2024 | Vision-Language Joint Learning for Box-Supervised Change Detection in Remote SensingabstractChange detection (CD) in remote sensing aims at revealing land cover changes according to the category of the ground objects. However, the category information is always missing in current popular vision-based CD methods. Considering that language analysis is really good at identifying different categories, a vision-language joint learning method is proposed in this paper, which consists of two vision-language joint representation (VLJR) modules and a changed instance segmentation (CIS) module. The former combines image features and language features with the help of text encoder and Transformer. The latter generates the final pixel-level CD result with only box-level labeled samples by level-set evolution and box matching supervision, which reduces manual-labor to a large extent. Tested on representative WHU datasets, the proposed method achieves comparable results to fully-supervised CD methods and is ahead of the other weakly-supervised methods. Kanghua Yin, Jia Liu 0020, Liang Xiao 0001 |
IGARSS | 4 |
| 2024 | Spectral-Spatial Attentions and Deep Supervision for Change Detection in Remote Sensing ImagesabstractThe task of remote sensing image change detection involves identifying differences between images captured in the same geographical area but at different times. When dealing with dual-time-series images, lighting and seasonal variations often make recognition challenging. To address the challenges, based on Unet++, we innovatively introduce the Spectral-Spatial Attention Module (SSAM) to better focus on fine-grained details. SSAM uses different frequency components to allocate differential weights to channels, allowing the network to pay more attention to the features relevant to the current task. Moreover, to better capture the change details, a multi-level deep supervision strategy is introduced to enhance the discriminative ability and robustness of early features. Our proposed method is named as SSUNet and has been validated on the CDD and LEVIRE-CD datasets, demonstrating significant advantages in detail recognition. Jia Liu 0020, Fang Liu 0001, Jingxiang Yang, Liang Xiao 0001 |
IGARSS | 6 |
| 2024 | TSBP: Improving Object Detection in Histology Images via Test-Time Self-guided Bounding-Box Propagation
Liang Xiao 0001, Yizhe Zhang 0001 |
MICCAI (4) | 2 |
| 2024 | MSTSENet: Multiscale Spectral-Spatial Transformer with Squeeze and Excitation network for hyperspectral image classification
Irfan Ahmad 0009, Ghulam Farooque, Qichao Liu, Fazal Hadi, Liang Xiao 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Stair Fusion Network With Context-Refined Attention for Remote Sensing Image Semantic SegmentationabstractSemantic segmentation of remote sensing images is essential in various fields, such as earth resource census, environmental pollution monitoring, and land use planning. The segmentation performance has been significantly improved recently with the development of deep learning. However, there are still some challenges in dealing with remote sensing images. One of the main issues is that features within the same category in remote sensing images could vary significantly, while features between different categories could be more similar, leading to confusion in segmentation. Moreover, the presence of large shadow areas narrows the feature differences between categories, making segmentation even more difficult. To address these challenges, one way is to leverage contextual and multi-scale information for accurate segmentation. As a consequence, in this paper, we propose a stair fusion network with context refined attention (SFCRNet). A context-based attention embedding module is proposed to enhance the representation of the processed features by utilizing the context to maximize information retention in the channel and spatial dimensions. It can retain the information on the original channel and the association between it and other channels. Furthermore, we present a stair fusion network where a stair shaped architecture and corresponding fusion module are designed to ensure that rich semantic information from high-level features is continuously transmitted to low-level layers. The experimental results on three datasets demonstrate the effectiveness of our proposed method. Jia Liu 0020, Wenyi Hua, Fang Liu 0034, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | MIMO-SST: Multi-Input Multi-Output Spatial-Spectral Transformer for Hyperspectral and Multispectral Image FusionabstractThe current advanced hyperspectral super-resolution methods utilize Convolutional Neural Networks (CNNs) that are either deeper or wider. These networks are designed to acquire end-to-end mapping capability, facilitating the transformation from Low-Resolution Hyperspectral Images (LR-HSI) and High-Resolution Multispectral Images (HR-MSI) to High-Resolution Hyperspectral Images (HR-HSI). The existing methods lack the capability to capture details and structures in the image effectively, while multi-input and multi-output methods can address this issue efficiently. Therefore, this paper proposes a novel network architecture named Multi-Input Multi-Output Spatial-Spectral Transformer (MIMO-SST). To apply the multi-input and multi-output methods in HSI fusion, specifically integrating the spatial-spectral information of LR-HSI and HR-MSI, we introduce multi-head feature map attention, multi-head feature channel attention, and a multi-scale convolutional gated feedforward network, constructing the proposed Mixture spatial-spectral Transformer. Moreover, to enhance the expressive power of image edges and recover the sharpened structure details, this study incorporates a novel wavelet-based high-frequency loss into the ultimate comprehensive loss, with the objective of refining the reconstruction of high-frequency details. Experimental studies on three simulated datasets and one real-world dataset demonstrate that the proposed method in this study outperforms contemporary state-of-the-art methods in terms of performance. It is noteworthy that our method exhibits a 0.85 dB improvement in terms of the PSNR metric on the CAVE dataset compared to state-of-the-art methods. Our code is publicly available at https://github.com/Freelancefangjian/MIMO-SST. Jingxiang Yang, Abdolraheem Khader, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Deep Unfolding Network Enhanced by Transformer Priors for Unregistered Hyperspectral and Multispectral Image FusionabstractIn satellite remote sensing, the complementary nature of hyperspectral (HSI) and multispectral (MSI) imagery necessitates their fusion to enhance both spatial and spectral resolution. However, the inherent misalignment between these datasets, due to differences in acquisition conditions, poses a significant challenge. This study presents a novel approach, the deep unfolding network enhanced by Transformer priors (DUNET), to address the simultaneous registration and fusion of HSI and MSI. Unlike conventional deep fusion methods, which are often treated as opaque “black boxes,” DUNET incorporates the deep unfolding method, leveraging mutual information and deep priors to facilitate a better degradation model-informed fusion process. The proposed network incorporates hybrid attention Transformers (HATs) and spatial-frequency modules to fully exploit the spatial-spectral information of HSI, resulting in a more accurate and detailed representation of the scene. We conducted extensive quantitative and visual experiments on three standard HSI datasets. The results demonstrate that our proposed DUNET method outperforms the existing mainstream algorithms in the field of remote sensing image fusion, showcasing its effectiveness. Specifically, our proposed method achieves the improvements of 3.4, 5.1, and 8.2 dB in terms of peak signal-to-noise ratio (PSNR) compared with the latest methods on the ICVL, Chikusei, and Houston datasets, respectively. Jingxiang Yang, Abdolraheem Khader, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Conjoint Cross-Attention Modeling and Joint Feature Calibrating for Remote Sensing Image Change Detection via a Triple-Double NetworkabstractRemote sensing (RS) image change detection (CD) based on deep learning (DL), has received increasing attention recently. However, the general independent learning of bi-temporal images ignores the relationship between them, falling short in learning of the change information. In this paper, a Triple-Double (TD) framework with ability of conjoint cross-attention modeling and joint feature calibrating is proposed for CD. Specifically, the TD framework composed of Triple-branch encoder and Double-branch decoder is constructed to extract diverse features and acquire changed maps with the guidance of original edge cues. To enhance the perception of the connection between the bi-temporal features, the multi-scale difference guidance (MDG) module and conjoint cross-attention (CCA) module are designed for the dual-branch encoder, wherein the CCA introduces a novel and efficient rule for modeling the affinity in spatial and channel dimension simultaneously. Furthermore, a joint feature calibration (JFC) module is introduced to enhance the expression of feature diversity in the joint features within the single-branch encoder. Experimental results on three public datasets demonstrate the superiority of the proposed method compared to the state-of-the-art (SOTA) methods. Fang Liu 0034, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Candidate-Aware and Change-Guided Learning for Remote Sensing Change DetectionabstractChange detection (CD) in remote sensing images aims at revealing earth surface changes between co-registered bitemporal images. A common way to reveal changed areas is to directly mix bitemporal features and generate CD results through supervised learning. However, a certain change usually corresponds to a real object in either of the two images, which exhibits coarse/fine shape in different scales. Therefore, a coarser-to-finer method called candidate-aware and change-guided network (CACG-Net) is proposed to effectively detect changes, where candidate objects are revealed and associated with interesting changes. Specifically, there are three key components. They are multistage change decoder (MCD), candidate-aware learning (CAL) and change guidance module (CGM). MCD reveals the most important changed objects in the coarse shape from the basic features extracted by the backbone (ResNet-18). To capture changes of interest, CAL is designed to select candidate objects in each temporal image, where a segmenter is utilized with variant change-losses. CGM intends to enrich the change details step-by-step through combining coarser change results and finer features, so that changed objects are gradually revealed in a coarser-to-finer way. Furthermore, deep supervision is employed throughout the layers of CACG-Net in the training procedure, which mitigates the learning difficulty in both deep and shallow layers. Test results on four popular datasets indicate that the proposed method outperforms several state-of-the-art CD algorithms in terms of accuracy and efficiency. Fang Liu 0034, Yangguang Liu, Jia Liu 0020, Xu Tang 0004, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Difference Guidance Learning With Feature Alignment for Change DetectionabstractChange detection (CD) in remote sensing aims at identifying changes of specific categories from multitemporal images acquired at different moments of a given scene. Due to seasonal alteration and light variation, there are always pseudo-changes hard to be recognized. To this end, we propose a difference guidance learning way to mitigate the effects of pseudo-change, which benefits capturing more discriminative information and identifying real changes. Specifically, it combines difference information with fused features in a guidance way and generates discriminative features in multiple scales. Besides that, feature alignment is conducted in the highest stage to learn feature correlations between bitemporal images, which benefits identifying semantic changes by information exchange. Therefore, the proposed method is named feature alignment and difference guidance network (FADG-Net). Furthermore, a set of convolutional layers with different receptive field sizes is also utilized to capture spatial information across different scales and enhance texture features accordingly. Tested on three public CD datasets, the effectiveness of the proposed FADG-Net is verified, where pseudo-change problem is mitigated and our method is superior to other comparison methods. Yangguang Liu, Fang Liu 0034, Jia Liu 0020, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Content-Guided and Class-Oriented Learning for VHR Image Semantic SegmentationabstractWith the flourishing of remote sensing (RS) platform techniques, very high-resolution (VHR) images have become more and more popular in recent years, which benefit the task of semantic segmentation but bring new challenges as well. Small objects, such as cars and trees, only occupy a few pixels in VHR images and are usually hard to segment. Moreover, the overlap problem about similar ground objects, such as low vegetation and trees, always results in underperformance. In this article, a content-guided and class-oriented network (CGCO-Net) for VHR image semantic segmentation is proposed to tackle this problem. Specifically, an adaptive content-guided fusion (ACGF) module with deformable convolution is introduced to capture long-distance dependencies and spatial aggregation effectively. With the guidance of the high-level features, the semantic content knowledge is gradually aggregated into low-level features and the details of the original features could be preserved. In addition, a multiscale channel alignment module is introduced into the encoder–decoder structure to further extract the long-range context information and reduce the calculation consumption. In order to improve the ability of pixel-level classification, a class-oriented representation learning (CORL) way is designed with transformer blocks by class embedding and deep supervision, which gradually enhance the discrimination and benefit the final segmentation. Furthermore, a weighted loss function and a threshold optimization strategy are employed to alleviate the sample imbalance problem. Tested on three public datasets and compared with several state-of-the-art methods, the proposed CGCO-net achieves good performance in both qualitative and quantitative analysis. Fang Liu 0034, Keming Liu, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | A Spectral Diffusion Prior for Unsupervised Hyperspectral Image Super-ResolutionabstractHyperspectral image (HSI) super-resolution with an auxiliary multispectral image (MSI) belongs to the class of inverse problems, where prior knowledge is essential for obtaining the target. Various hand-crafted or deep priors have been developed to enforce the desired solutions. Nevertheless, the spectral distribution knowledge is ignored and still not exploited as a prior. To this end, we design a spectral diffusion model (SDM) to capture the spectral distribution of HSIs and thereby exploit it as a prior for the problem of unsupervised HSI super-resolution. Specifically, we first investigate the spectrum generation problem and extend the diffusion model to fit the 1-D spectral data. Then, we transfer the spectral distribution knowledge of the trained SDM by means of keeping its transition information and thereby induce a regularization term in the framework of maximum a posteriori. At last, we integrate the iterative solving and diffusion generation processes together and employ the Adam to solve the final optimization problem by following the reverse spectral generative sequence. Experimental results conducted on both synthetic and real datasets demonstrate the effectiveness of the proposed approach. The code of the proposed approach is available onhttps://github.com/liuofficial/SDP. Zebin Wu 0001, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Adaptive Spatial Structure-Aware and Spectral Gradient Structure Tensor-Guided Model for PansharpeningabstractIn this article, we propose a novel adaptive spatial structure-aware and spectral gradient structure tensor-guided model (AS3GSTM) for pansharpening, which realizes the process of fusing the low-resolution multispectral (LRMS) image and the paired panchromatic (Pan) image to output the high-resolution multispectral (HRMS) image. Specifically, based on the basic spectral fidelity term between HRMS and LRMS obtained from the spatial degradation model for spectral fidelity, we also enforce the radiometric ratio-guided high-frequency detail fidelity term between HRMS, LRMS, and Pan for high-frequency detail fidelity. Moreover, considering that the HRMS image and the Pan image actually not only have strong spatial structure similarities, but also differ from each other, we further propose a novel Pan-guided adaptive spatial structure-aware prior term for the HRMS image to guide the fusion process. Besides, we particularly exploit the structure tensor of the spectral gradient of HRMS for simultaneously spectral-spatial prior modeling, and propose a novel spectral gradient-guided structure tensor total variation prior term for the HRMS image. Subsequently, we design an efficiently alternating algorithm to optimize the proposed AS3GSTM model. Finally, lots of fusion experiments comprehensively validate the superiority of AS3GSTM. Pengfei Liu 0002, Zhizhong Zheng, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | A Multiscale Grouping Transformer With CLIP Latents for Remote Sensing Image CaptioningabstractRecent progress has shown that integrating multiscale visual features with advanced Transformer architectures is a promising approach for remote sensing image captioning (RSIC). However, the lack of local modeling ability in self-attention may potentially lead to inaccurate contextual information. Moreover, the scarcity of trainable image-caption pairs poses challenges in effectively harnessing the semantic alignment between images and texts. To mitigate these issues, we propose a Multiscale Grouping Transformer with Contrastive Language-Image Pre-training (CLIP) latents (MG-Transformer) for RSIC. First of all, a CLIP image embedding and a set of region features are extracted within a Multi-level Feature Extraction module. To achieve a comprehensive image representation, a Semantic Correlation module is designed to integrate the image embedding and region features with an attention gate. Subsequently, the integrated image features are fed into a Transformer model. The Transformer encoder utilizes dilated convolutions with different dilation rates to obtain multiscale visual features. To enhance the local modeling ability of the self-attention mechanism in the encoder, we introduce a Global Grouping Attention mechanism. This mechanism incorporates a grouping operation into self-attention, allowing each attention head to focus on different contextual information. The Transformer decoder then adopts the Meshed Cross-Attention mechanism to establish relationships between various scales of visual features and text features. This facilitates the generation of captions for images by the decoder. Experimental results on three RSIC datasets demonstrate the superiority of the proposed MG-Transformer. The code will be publicly available at https://github.com/One-paper-luck/MG-Transformer. Lingwu Meng, Jing Wang 0201, Ran Meng, Yang Yang 0074, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | I2MEP-Net: Inter- and Intra-Modality Enhancing Prototypical Network for Few-Shot Hyperspectral Image ClassificationabstractPrototypical network, celebrated for its flexible network structure and metric computation capability, has become a prevailing strategy in addressing challenges associated with few-shot hyperspectral image classification. However, factors such as insufficient training samples, spectral mixing, and noise interference in complex scenarios severely impact the stability of its prototypes, ultimately leading to a degradation in classification performance. Therefore, this paper proposes a novel method called I2MEP-Net, which incorporates two auxiliary modalities to facilitate both inter- and intra-modality enhancing prototype learning with the base modality. Specifically, I2MEP-Net employs the auxiliary LiDAR modality with base HS modality from the same scene for inter-modality enhancing prototype learning, which offsets the sparsity of few-shot features through a cross-modal approach. In addition, it utilizes the target few-shot labeled data as the auxiliary HS modality for intra-modality enhancing prototype learning on the enhanced prototypes, in a way that adaptively generates diversity features, thereby further enriching the prototype embedding space and achieving more fine-grained and stable prototypes. Comprehensive experiments are conducted on the publicly available hyperspectral image datasets. These experiments indicate that the proposed I2MEP-Net outshines the existing state-of-the-art deep learning techniques and few-shot classification methodologies. Wenzhen Wang, Fang Liu 0034, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Hyperspectral Reconstruction From RGB Images via Physically Guided Graph Deep Prior LearningabstractRecovering the latent hyperspectral image (HSI) from RGB or multispectral image (MSI), which is dubbed spectral super-resolution (SSR), has demonstrated outstanding performance owing to the advancements in convolutional neural networks (CNNs). However, most of the current algorithms concentrate on the pursuit of networks with more expensive or complex structures, while ignoring the significant role of physical degradation models in SSR. In addition, the inherent defects of CNN make these networks focus more on the local correlation, while their ability to model the long-range correlations in the spectral and spatial domains still has room to improve. To overcome this shortcoming, we propose a physical degradation-guided deep prior learning network (PGDL-Net) for SSR via unfolding the optimization process of the blind SSR model, in which the priors of unknown spectral response function (SRF) and latent HSI are learned explicitly and represented by proximal operators. To jointly extract the local and non-local information, we design a hybrid graph Transformer as the proximal operator to solve the latent HSI. Furthermore, to ensure efficient learning of SRF and HSI, we also propose a novel loss function constraining the reconstruction error, degradation consistency, and observation fidelity for the learned SRF and HSI. Experimental results on multiple datasets illustrate the improved performance and stability of our method in SSR. Jingxiang Yang, Tian Lin 0001, Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Composite Neighbor-Aware Convolutional Metric Networks for Hyperspectral Image ClassificationabstractSupervised classification of hyperspectral image (HSI) is generally required to obtain better performance in spectral-spatial feature learning by fully using complex pixel- and superpixel-level interdependencies with small labeled samples. Limited by the local regular convolutions, convolutional neural networks (CNNs) can only exploit information from the short-range Euclidean neighbors of a target, hindering the effectiveness of feature representation. In contrast, graph convolutional networks (GCNs) can learn long-range dependencies between non-Euclidean neighbors but usually require the input of a full graph constructed from a whole HSI, making GCNs must be trained in a full-batch manner with tremendous computational consumption. In this work, we propose a composite neighbor-aware convolutional metric network (CNCMN), aiming to learn each target's representation from its composite neighbors (i.e., both Euclidean and non-Euclidean neighbors) in a batchwise manner. Specifically, for each target in an HSI, its Euclidean neighbors are the pixels in the local square region centered on itself, and its non-Euclidean neighbors are several related nodes selected from the constructed full graph. Correspondingly, a composite convolution (CoConv) is proposed by coupling an image convolution and a graph convolution, which can perform flexible convolutions on those composite neighbors and extract adaptively fused features from them. Besides, to further boost classification, we also propose a mini-batch metric classifier to dynamically optimize interclass and intraclass distances of samples batch by batch, which is then combined with the CoConv to form the mini-batch CNCMN. Extensive experiments on three real-world HSIs demonstrate the advantages of the proposed method over mini-batch deep learning algorithms and have obtained the state-of-the-art performance in these fields. The code is available at: https://github.com/qichaoliu/HSI-CNCMN. Qichao Liu, Liang Xiao 0001, Nan Huang 0001, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Unsupervised Deep Tensor Network for Hyperspectral-Multispectral Image FusionabstractFusing low-resolution (LR) hyperspectral images (HSIs) with high-resolution (HR) multispectral images (MSIs) is a significant technology to enhance the resolution of HSIs. Despite the encouraging results from deep learning (DL) in HSI-MSI fusion, there are still some issues. First, the HSI is a multidimensional signal, and the representability of current DL networks for multidimensional features has not been thoroughly investigated. Second, most DL HSI-MSI fusion networks need HR HSI ground truth for training, but it is often unavailable in reality. In this study, we integrate tensor theory with DL and propose an unsupervised deep tensor network (UDTN) for HSI-MSI fusion. We first propose a tensor filtering layer prototype and further build a coupled tensor filtering module. It jointly represents the LR HSI and HR MSI as several features revealing the principal components of spectral and spatial modes and a sharing code tensor describing the interaction among different modes. Specifically, the features on different modes are represented by the learnable filters of tensor filtering layers, the sharing code tensor is learned by a projection module, in which a co-attention is proposed to encode the LR HSI and HR MSI and then project them onto the sharing code tensor. The coupled tensor filtering module and projection module are jointly trained from the LR HSI and HR MSI in an unsupervised and end-to-end way. The latent HR HSI is inferred with the sharing code tensor, the features on spatial modes of HR MSIs, and the spectral mode of LR HSIs. Experiments on simulated and real remote-sensing datasets demonstrate the effectiveness of the proposed method. Jingxiang Yang, Liang Xiao 0001, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Learnable Snake R-CNN for Instance-Level Biomedical Image SegmentationabstractPrecisely knowing each instance’s position and extents is a critical first step in many biological applications. State-of-the-art techniques rely either on deep learning models designed to predict segmentation masks on each Region of Interest (RoI) or on classic active contour methods. The former struggles to precisely delineating boundaries and tends to output masks at low resolutions when the cells/nuclei are very irregular while the latter often needs good initialization and manual setting of parameters, thus limiting their usefulness. To bridge this gap, we introduce Snake R-CNN, a new level of the learnable active contour model that predict boundary on each RoI in a sequent way. To do so, for each RoI, we reformulate the contour deformation task in terms of a hidden state evolution problem and update the evolution process using energy minimization. We learn snake parameterizations per instance in an end-to-end manner, and demonstrate its effectiveness for contour inferences of various cell/nucleus types where consistently higher performances were obtained for comparison against state-of-the-arts. Jie Song 0014, Ziyun Cai, Yurong Song, Guoping Jiang, Zhichao Lian, Liang Xiao 0001 |
ICIP | 6 |
| 2023 | Spatial-Preserving and Edge-Orienting High-Resolution Network for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD), with a view to probing surface changes between bi-temporal images, makes a spurt of progress with the continuous innovation of deep learning. However, the extraction of multi-scale features and the detection of small domain of variation as well as the detail information in RSCD task still has large development space. Besides, current existing methods mostly focus on learning regional information but pay less regard to boundary identification, which leads to inaccurate detection results. Therefore, a spatial-preserving and edge-orienting high-resolution network is proposed to address the problems. In the overall architecture, a dual-branch encoder consists of a pyramid feature extracted branch and an enhanced HR network branch is designed to extract muti-scale bi-temporal features and small change objectives, while two edge-orienting modules (EOM) are embedded in order to utilize edge prior knowledge for further improving the accuracy of change detection. Moreover, spatial-preserving module (SPM) based on the self-attention calculation in spatial dimension is applied in the pyramid part to alleviate the poor location information of the high-level features. The experimental results demonstrate that the proposed network outperforms the cited state-of-the-art methods on LEVIR change detection datasets (LEVIR-CD). Fang Liu 0001, Jia Liu 0020, Liang Xiao 0001, Xu Tang 0004 |
IGARSS | 4 |
| 2023 | Domain-Specific and Domain-Common Feature Enhancement for Cross-Domain Few-Shot Hyperspectral Image ClassificationabstractThere is a small sample problem in hyperspectral image (HSI) classification task due to the difficulty of labeling samples. It is generally solved using a combination of few-shot learning and cross-domain method. In the paper, we propose a domain-specific and domain-common feature enhancement method for cross-domain few-shot HSI classification. It consists of a domain adaptation module and a feature enhancement module. The former is used to learn domain-specific features of both domains from the beginning of the network, and the latter is used to reduce domain differences by learning domain-common features through feature enhancement. The experimental results indicate that our proposed method performs better than the advanced classification methods. Wenfei Gao, Fang Liu 0001, Jia Liu 0020, Liang Xiao 0001, Xu Tang 0004 |
IGARSS | 4 |
| 2023 | Siamese Recurrent Residual Refinement Network for High-Resolution Image Change DetectionabstractIn this study, we propose a siamese recurrent residual refinement network (SR3Net) for change detection in high-resolution remote sensing images. SR3Net uses the siamese network to fully extract multi-scale features of bi-temporal images. The obtained multi-scale feature maps are input to the multi-level difference module (MDM) to generate difference feature maps. The residual refinement module (RRM) with residual refinement blocks (RRBs) learns the residual between the intermediate change map and the ground truth by alternately exploiting low-level and high-level integrated difference features. Moreover, RRBs can obtain complementary information of the intermediate predictions and add residuals to the intermediate prediction to refine the change map. Experiments on the WHU-CD dataset show that the proposed method outperforms state-of-the-art methods. Chengwei Huang, Ling Hu 0003, Wenzi Liao, Liang Xiao 0001 |
IGARSS | 4 |
| 2023 | A Multi-Scale Deep Feature Learning and Semantic Enhancement Approach for Remote Sensing Scene ClassificationabstractDeep learning has made great success in remote sensing scene classification since the powerful feature representation and complex nonlinear relationship learning. However, existing methods ignore the information redundancy and semantic ambiguity among them. To cope with this problem, we propose a multi-scale deep feature learning and semantic enhancement approach (MDFL-SE). First, we employ Pyramid Convolution PyConvResNet as the backbone to extract multilayer convolutional features. Then, a progressive deep feature aggregation module (PDFA) is designed to use high-level features to guide the low-level ones to choose the discriminative features. Finally, a global multiscale semantic extraction module (GMSE) and a grouped semantic extraction module (GSE) are combined to extract the channel and spatial information of multilayer fusion features. Experiments performed on AID and NWPURESISC45 RSSC datasets demonstrate that the proposed framework can obtain outstanding performance compared with state-of-the-art approaches. Hengyi Huang, Wenzhen Wang, Wenzi Liao, Liang Xiao 0001 |
IGARSS | 4 |
| 2023 | Unsupervised Domain Adaption for Remote Sensing Semantic Segmentation with Self-Attention MechanismabstractThe domain shift between the source and target domains limits the performance of traditional convolutional neural networks (CNNs) for feature extraction in remote sensing tasks. We propose an image translation network that uses generative adversarial networks (GANs) to transfer spectral distributions from training to test data, enhancing cross-domain semantic segmentation. Our approach fine-tunes the DeepLab-V3 framework on synthetic training data generated by the proposed network. Experimental results show improved performance in cross-domain semantic segmentation tasks for remote sensing images. Keming Liu, Fang Liu 0001, Jia Liu 0020, Liang Xiao 0001, Xu Tang 0004 |
IGARSS | 4 |
| 2023 | Edge-Guided Feature Dense Fusion Network for Remote Sensing Image Change DetectionabstractRemote sensing change detection (CD) is of great importance to Earth observation. Recently, Deep Learning (DL) has been increasingly used to extract useful features and make accurate decisions in a large number of remote sensing images, due to its ability to automatically learn semantic features. However, insufficient fusion of bitemporal images and the lack of prior knowledge of edge structures in current DL methods will result in inaccurate CD results, especially for building boundaries. To alleviate these problems, an edge-guided feature-densely-fused network (EGFDFN) is proposed in this paper. In contrast to conventional Siamese networks, EGFDFN extracts bitemporal features from an extra dual decoder instead of a dual encoder to obtain more accurate change features. In addition, an attention and dense fusion module (ADFM) and an edge guidance module (EGM) are used to enhance features and make full use of edge information. Experimental results demonstrate that the proposed method outperforms on LEVIR-CD dataset among other representative methods. Hejun Luo, Jia Liu 0020, Fang Liu 0001, Jingxiang Yang, Liang Xiao 0001 |
IGARSS | 6 |
| 2023 | Swin transformer with multiscale 3D atrous convolution for hyperspectral image classification
Ghulam Farooque, Qichao Liu, Allah Bux Sargano, Liang Xiao 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Intrinsic Decomposition Model-Guided Two-Stream Coupled Autoencoder for Unsupervised Hyperspectral Image Change DetectionabstractHyperspectral image change detection (HSI-CD) is one of the main research topics in remote sensing. Theoretically, the ground objects can be considered to have changed when their spectral features behave differently. However, in practical scenarios, this presupposition does not hold as the collected features are affected by many factors, such as illumination conditions, atmospheric effects, and topographic changes. Inspired by the intrinsic image decomposition, we propose a novel unsupervised deep learning framework based on two-stream coupled autoencoder (TSCA) to cope with bi-temporal co-registered HSI-CD. The network consists of two symmetric encoders and a decoder, which can jointly decompose bi-temporal images into abundance coefficients corresponding to the same set of spectral bases. As our network separates the component of spectral variation from multiple images, the extracted abundance features with inherent properties of materials can provide better performance for change detection. Moreover, to enforce alignment of the feature space, a reasonable consistency loss is devised to constrain the solution space, by cross-reconstruction in both branches. Experimental results demonstrate its superiority over the recently developed state of the arts. Jia Sun 0010, Jia Liu 0020, Liang Xiao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Spatial-Spectral Adaptive Learning With Pixelwise Filtering for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is significant in remote sensing applications. However, most methods focus on the spectral and spatial correlation information in the neighborhood while ignoring the feature difference among global different pixels. In this article, we propose a spatial–spectral adaptive learning with pixelwise filtering (SSALPF) method to fully consider the discriminative information of pixels in different spatial locations, which mainly consists of a parallel spatial–spectral adaptive learning (SSAL) module and a pixelwise filtering (PF) module. Specifically, the former aims to obtain joint spatial–spectral discriminative features of each pixel point in a parallel manner and is used as a guide for adaptive selection of filter kernel. The latter uses the adaptive filter kernel to implement pixel-level filtering on HSI, in order to learn the discriminative features contained in different pixel points for classification. The adaptive filter kernel is generated by a linear combination of a predefined dictionary containing multiple filter bases. Experiments demonstrate that the proposed method is superior to other methods on popular hyperspectral datasets. Wenfei Gao, Fang Liu 0034, Jia Liu 0020, Liang Xiao 0001, Xu Tang 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | PRBCD-Net: Predict-Refining-Involved Bidirectional Contrastive Difference Network for Unsupervised Change DetectionabstractHeterogeneous bi-temporal images have different visual appearances and inconsistent data distribution for the same scene, making it challenging to detect changes, which need to align the shared information and reduce various unwanted sensor-related noises for comparability. Mainstream methods usually adopt two types of techniques: feature transformation and image translation. The former relies on handcrafted priors while the latter lacks constraints on unwanted backgrounds, leading to limitations such as a lack of robustness to non-intrinsic changes (e.g., seasonal and atmospheric changes, and sensor-related noise) and unsatisfactory detection performance. To overcome these drawbacks, we propose a novel unsupervised predict-refining-involved bidirectional contrastive difference network (PRBCD-Net) composed of a coarse prediction module and iterative refining modules. Each refining module utilizes feature extractors with a cross-reconstruction constraint and bidirectional contrastive constraint to extract discriminative features, and then generate a refined change map by change map optimizers. Two advantages of the proposed PRBCD-Net are: 1) the cross-reconstruction constraint is used to promote the feature distribution consistency of the bi-temporal images by using the forward and backward transformations; 2) the bidirectional contrastive constraint is used to improve the discriminability of features by narrowing the gap between non-intrinsic changes while widening intrinsic changes under the guidance of a coarse change map. Thus, the refining module can generate a finer change map than the coarse one, and the performance can be further improved through multiple iterations. Experimental results demonstrate the effectiveness and robustness of the proposed method compared with state-of-the-art methods. Ling Hu 0003, Qichao Liu, Jia Liu 0020, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | S2DMSC: A Self-Supervised Deep Multilevel Subspace Clustering Approach for Large Hyperspectral ImagesabstractSubspace clustering (SC) has achieved remarkable success in hyperspectral images (HSIs) due to the powerful representation ability of handling high-dimensional complex data. However, most of the existing SC methods focus on linear subspace representation and ignore the more effective nonlinear representation. Besides, SC suffers from the bottlenecks, such as high computation load and memory capacity, due to the spectral decomposition of adjacency matrix for large HSIs. To overcome these limitations, we propose an end-to-end learnable network framework for large HSIs, called self-supervised deep multi-level subspace clustering (S2DMSC), which incorporates the convolutional neural network (CNN) module, multi-level subspace clustering (MSC) module, and high-quality pseudo-label-based self-supervised learning module into a unified learning framework. More concretely, the deep multi-level spatial-spectral representation from hierarchical superpixels is modeled as a sparsity-constrained self-expression module for SC to construct high-quality pseudo-labels to learn the network parameters and produce better clusters for hyperspectral pixels. Experimental results on four classical HSIs demonstrate the effectiveness of S2DMSC and exhibit superior clustering performance compared to the representative clustering methods. Nan Huang 0001, Liang Xiao 0001, Qichao Liu, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | MANet: An Efficient Multidimensional Attention-Aggregated Network for Remote Sensing Image Change DetectionabstractDeep learning has significantly advanced the change detection in remote sensing image with its excellent performance. For change detection tasks, there are two critical issues. First, with scale variance of different objects in remote sensing images, effectively aggregating multi-scale features helps to generate fine-grained change objects. Second, it is critical but challenging to fully exploit the variance information between bi-temporal images to avoid pseudo-variation and region blurring. To alleviate the above issues, this paper proposes an efficient multi-dimensional attention-aggregation network (MANet), which keeps better feature aggregation while maintaining excellent differential attention ability. This paper carries three main contributions. First, we propose a multiscale asymmetric convolutional attention (MACA) module. Due to the asymmetric convolution’s ability to focus on feature contours effectively, the MACA can not only aggregate multi-scale features effectively, but also refine the edge information of features. Second, we propose a dual-dimensional attention (DDA) module for adaptively fusing shallow and deep features, which is used to generate rich feature representations. Third, the difference guidance (DG) module is exploited for enhancing the attention of changed regions to mitigate the influence of uncorrelated changes on the change detection result. Experiments on four popular change detection datasets show that our network can accomplish higher detection accuracy than the state-of-the-art networks. Kaixuan Jiang, Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Adversarial Domain Alignment With Contrastive Learning for Hyperspectral Image ClassificationabstractRecently, deep learning-based hyperspectral image (HSI) classification techniques are flourishing and exhibit good performance, where cross domain information is usually utilized to reduce the dependency on large labeled samples. However, the gap between source domain and target domain makes it difficult to carry out knowledge transfer directly. In this paper, an adversarial domain alignment with contrastive learning method is designed for the HSI classification task to achieve feature consistency that benefits transferring knowledge. In details, spectral alignment and semantic alignment are conducted in local and global levels respectively in an adversarial learning way, and the adversarial loss acts on both source and target domains. In order to learn specific features for objects with different spatial scales, a multi-scale selection module is constructed in semantic alignment to select channel features adaptively. Moreover, contrastive learning is employed to increase both robustness and sensitiveness, where augmented data from the same/different samples are forced to be similar/dissimilar with each other. The training process is conducted in a few-shot learning way then the few-shot classification loss, the adversarial loss and the contrastive loss is optimized together. Tested on one source dataset and four target datasets, the experimental results show that the proposed method outperforms the other comparisons. Fang Liu 0034, Wenfei Gao, Jia Liu 0020, Xu Tang 0004, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Multiresolution Analysis-Inspired Spatial and Spectral Details Preserved Model for Variational PansharpeningabstractPansharpening, which is also known as the fusion of low resolution multispectral (LRMS) and panchromatic (PAN) images, refers to producing a high resolution multispectral (HRMS) image by preserving the spectral detail from the LRMS image while extracting the spatial detail from the PAN image. In this article, we revisit and novelly reinterpret the multi-resolution analysis (MRA)-based pansharpening framework as the fusion framework of “spectral detail + spatial detail" and can obtain two alternative formulations of spectral detail and spatial detail respectively, and hence propose a novel variational pansharpening method with MRA-inspired spatial and spectral details preserved model. Firstly, the spatial degradation relationship between HRMS and LRMS is imposed as the spectral fidelity term. Secondly, based on the new reinterpretation of “spectral detail + spatial detail" of MRA fusion framework, we propose to use the structure tensor to model the spatial detail image, and propose a new structure tensor total variation (STV)-guided spatial detail preserved prior term. Moreover, to model the spectral detail image, we propose to impose the spectral detail preserved constraint between the two alternative formulations of spectral detail as the MRA-inspired spectral detail preserved prior term. Then, we optimize the proposed model via the alternating direction method of multipliers (ADMM). Furthermore, variously experimental results on the reduced-scale and full-scale datasets validate the superiority of proposed method. Pengfei Liu 0002, Liang Xiao 0001, Zhizhong Zheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Learning Transferable Discriminative Knowledge From Attribute-Aligned Hyperspectral ImagesabstractHyperspectral image (HSI) classification faces the inherent challenge of small sample learning, primarily due to the difficulty in labeling vast land covers. Meta-learning, with its ability to learn transferable meta-knowledge from existing HSIs, is seen as a promising solution. However, different HSIs usually have varying distributions manifested as differing spectral wavelengths and reflectance shifts, which is often neglected in existing methods, stalling the acquisition of transferable features. To address this issue, we introduce an attribute-driven spectral alignment (ADSA) method, which parameterizes and embeds spectral attributes (i.e., spectral wavelengths and reflectance shifts) into a domain-adaptation model, aiming to decouple the domain-specific attributes of different HSIs in an unsupervised manner. After training, the parameterized attributes of source-domain (SD) HSIs can be substituted with those of the target domain (TD), allowing the decoder of ADSA to rebuild new HSIs sharing identical spectral attributes. By this means, numerous distribution-consistent labeled samples preserving inherent spectral–spatial structures can be obtained. Then, a 3-D residual prototypical network (RPN) based on 3-D convolutions and metric learning is designed to model complex structures of HSIs, which in combination with the few-shot learning (FSL) framework can extract valuable discriminative knowledge from these auxiliary samples. Finally, by applying this learned knowledge to the TD HSI, only a small number of labeled samples are required to obtain satisfactory performance. Extensive experiments on four real-world HSIs demonstrate the effectiveness of our method, and the performance outperforms several state-of-the-art methods. Qichao Liu, Liang Xiao 0001, Nan Huang 0001, Jinhui Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Prior Knowledge-Guided Transformer for Remote Sensing Image CaptioningabstractRemote sensing image captioning aims to generate meaningful and grammatically accurate sentences for remote sensing images. However, in comparison to natural image captioning, remote sensing image captioning encounters additional challenges due to the unique characteristics of remote sensing images. The first challenge arises from the abundance of objects present in these images. As the number of objects increases, it becomes increasingly difficult to determine the main focus of the description. Moreover, the objects in remote sensing images often share similar appearances, which further complicates the generation of accurate descriptions. To overcome these challenges, we propose a Prior Knowledge-guided Transformer for remote sensing image captioning. Firstly, scene-level and object-level features are extracted in a Multi-level Feature Extraction module. To further refine and enhance the extracted multi-level features, we introduce a Feature Enhancement module. This module utilizes a combination of graph neural networks and attention mechanisms to capture the correlation and difference between different objects or scene regions. Moreover, we propose a Prior Knowledge augmented Attention mechanism to select the objects that are more relevant to the scene regions by establishing the relationships between them. This attention mechanism is seamlessly integrated into the Transformer structure, providing valuable prior knowledge that promotes the caption generation process. Extensive experiments on three remote sensing image captioning datasets verify the superiority of the proposed method. Compared with the baseline methods, the proposed method achieves more impressive performance. The code will be publicly available at https://github.com/One-paper-luck/PKG-Transformer. Lingwu Meng, Jing Wang 0201, Yang Yang 0074, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Cross-Domain Few-Shot Hyperspectral Image Classification With Class-Wise AttentionabstractFew-shot learning (FSL) is an effective method to solve the problem of hyperspectral image (HSI) classification with few labeled samples. It learns transferable knowledge from sufficient labeled auxiliary data to classify unseen classes with limited labeled samples for training. However, the distribution difference between auxiliary data and unseen classes results in the learned transferable knowledge not being well applied to the new task. Therefore, a class-wise attentive cross-domain FSL (CA-CFSL) framework is proposed in this article, in which a feature extractor is learned to extract data features with discriminability and domain invariance. The class-wise attention metric module (CAMM) introduces class-wise attention on the FSL framework to learn more discriminative features, which improves the interclass decision boundaries. Furthermore, an asymmetric domain adversarial module (ADAM) is designed to enhance the ability of extracting domain-invariant representations, which combines asymmetric adversarial training with embedded domain-specific information. Experimental results on four public HSI datasets demonstrate that the proposed method outperforms the existing methods. Wenzhen Wang, Fang Liu 0001, Jia Liu 0020, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Cross-Modal Graph Knowledge Representation and Distillation Learning for Land Cover ClassificationabstractComplementary multimodal remote sensing (RS) data often leads to more robust and accurate classification performance. However, not all modal data can be available at the time of inference due to imaging conditions. To mitigate this issue, cross-modal knowledge distillation becomes an effective method, as it can leverage the complementary characteristics of multimodal data to guide cross-modal classification in cases with missing data. Therefore, this paper examines the shortcomings of traditional CNN cross-modal distillation methods in land cover classification: 1) insufficient knowledge representation; and 2) unstable knowledge transfer. Moreover, a novel cross-modal graph knowledge representation and distillation learning (CGKR-DL) framework is proposed to enhance land cover classification performance. The proposed CGKR-DL designs a single-stream joint feature learning network with convolutional neural network and graph convolutional network (CNN-GCN) to effectively construct the remote topology of data based on the strong correlation between land objects, thus enhancing the knowledge representation ability of the network. In addition, a multi-granularity graph distillation method is proposed to compensate for the inability of traditional CNN distillation in handling graph-structured information, where a feature distillation module based on graph discrimination (FD-GDM) is designed for stable graph feature distillation. We evaluate CGKR-DL on three publicly available multimodal RS datasets (HS-LiDAR, HS-SAR and HS-SAR-DSM) and achieve a significant improvement in comparison with several state-of-the-art methods. Wenzhen Wang, Fang Liu 0034, Wenzi Liao, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Multiple Deep Proximal Learning for Hyperspectral-Multispectral Image FusionabstractFusing low resolution (LR) hyperspectral image (HSI) with a high resolution (HR) multispectral image (MSI) could enhance the spatial resolution and quality of HSI. Current deep learning (DL) HSI-MSI fusion networks have achieved encouraging results, but their performance relies on large number of training images with known degradations consistent with the testing data. The trained DL model may fail on data with unseen degradations during inference. In this study, we propose a multiple deep proximal learning network (MDPro-Net) for HSI-MSI fusion, the unknown spatial-spectral degradations and latent HR HSI can be adaptively inferred. We first propose a joint variational fusion model with both the degradations and HR HSI as to-be-solved variables, which are regularized by multiple deep priors. Then we optimize the fusion model using quadratic splitting and alternative optimization strategy. The unknown blurring kernel, spectral degradation, and HR HSI are explicitly solved by three deep proximal operators. Through unrolling the solutions into a DL network, we build MDPro-Net, in which the deep proximal operators for degradations and HR HSI are learned in an end-to-end manner. Furthermore, in the deep proximal operator for latent HR HSI, a multi-scale transformer is designed to exploit the local and non-local dependencies. Experiments demonstrate that the proposed MDPro-Net is competitive with state-of-the-art fusion methods, in particular, it is robust in inferring the unseen degradations. Jingxiang Yang, Tian Lin 0001, Xiaoyang Chen 0005, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Learning Degradation-Aware Deep Prior for Hyperspectral Image ReconstructionabstractReconstructing the 3D hyperspectral image (HSI) from 2D snapshot measurements is a key task in spectral snapshot compressive imaging (SCI). Traditional model-based HSI reconstruction methods rely on hand-crafted priors. Recently, deep unfolding networks (DUNs) learn the priors using convolutional neural networks (CNNs) and have achieved satisfactory results. Most of DUNs assume the degradations of SCI are known. However, due to the phase aberration and distortion problems in real imaging process, there is a certain gap between the ideal and real degradation patterns, which may hinder the accurate HSI reconstruction. In this study, we propose a degradation-aware deep prior learning network (D2PL-Net), which tries to adaptively learn the practical degradation matrix during HSI reconstruction, thus bridges the gap between the ideal and real degradations. Specifically, we first propose a joint variational compressive reconstruction model, both of the latent HSI and unknown degradation can be explicitly solved. By unfolding the solutions into a deep network, D2PL-Net is built, which mainly consists of two parts, Degradation Matrix Learning (DML) mechanism and Degradation-guided Spectral-Spatial Transformer (DSST) in each stage. The former learns the degradation that approximates the real one; the latter represents the deep prior of latent HSI, it could exploit the spectral-wise and spatial-wise long-range dependencies of HSI under the guidance of learned degradation, and then reconstructs the HSI. To ensure an effective training of D2PL-Net, we propose a joint loss function constraining the HSI reconstruction errors, degradation-fidelity and degradation-consistency. Experiments on simulated and real-life datasets show that the proposed method is competitive with the state-of-the-art methods. Jingxiang Yang, Tian Lin 0001, Fang Liu 0034, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Deep Open-Curve Snake for Discriminative 3D Neuron TrackingabstractOpen-Curve Snake (OCS) has been successfully used in three-dimensional tracking of neurites. However, it is limited when dealing with noise-contaminated weak filament signals in real-world applications. In addition, its tracking results are highly sensitive to initial seeds and depend only on image gradient-derived forces. To address these issues and boost the canonical OCS tracker to a new level of learnable deep learning algorithms, we present Deep Open-Curve Snake (DOCS), a novel discriminative 3D neuron tracking framework that simultaneously learns a 3D distance-regression discriminator and a 3D deeply-learned tracker under the energy minimization, which can promote each other. In particular, the open curve tracking process in DOCS is formed as convolutional neural network prediction procedures of new deformation fields, stretching directions, and local radii and iteratively updated by minimizing a tractable energy function containing fitting forces and curve length. By sharing the same deep learning architectures in an end-to-end trainable framework, DOCS is able to fully grasp the information available in the volumetric neuronal data to address segmentation, tracing, and reconstruction of complete neuron structures in the wild. We demonstrated the superiority of DOCS by evaluating it on both the BigNeuron and Diadem datasets where consistently state-of-the-art performances were achieved for comparison against current neuron tracing and tracking approaches. Our method improves the average overlap score and distance score about 1.7% and 17% in the BigNeuron challenge data set, respectively, and the average overlap score about 4.1% in the Diadem dataset. Jie Song 0014, Zhichao Lian, Liang Xiao 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Prototype Learning for Automatic Check-OutabstractThe basic goal of Automatic Check-Out (ACO) task is to accurately predict the categories and quantities of products selected by customers in the check-out images. However, there is a significant domain gap between the single-product exemplars as training data and the check-out images as testing data. To mitigate the domain gap, we propose a novel method termed as Prototype Learning for Automatic Check-Out (PLACO). In PLACO, prototype learning is designed to reach the goal in two ways. Specifically, in the prototype-based classifier learning module, to fully exploit the invariance of category prototypes, the prototypes obtained from the single-product exemplars are employed to generate classifiers for classifying the proposals of check-out image. On the other side, in prototype alignment module, prototypes for both the single-product exemplar and check-out image domains are entered simultaneously to ensure intra-category compactness and inter-category sparsity. Moreover, to further improve the performance of PLACO, we develop a discriminative re-ranking module to both adjust the predicted scores of product proposals for bringing more discriminative ability in classifier learning and provide a reasonable sorting possibility by considering the fine-grained nature. Experiments are conducted on the large-scale RPC dataset for evaluations. Our PLACO obtains the optimal results in both traditional ACO task setting and incremental task setting. Hao Chen 0052, Xiu-Shen Wei, Liang Xiao 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | A Novel Video Stabilization Model With Motion Morphological Component PriorsabstractVideo stabilization is the process of improving the video quality by removing annoying fluctuant motion caused by camera jittering. A key issue of a successful solution is the temporal adaptability to motion and the overall robustness with respect to different motion types. However, most previous methods usually produce non-motion adaptive stabilized videos. In other words, under-smoothing in slow motion segments and over-smoothing in rapid motion segments will be produced for complex shaky videos. To overcome these drawbacks, we propose a novel video stabilization approach using a motion morphological component (MMC) decomposition. Specifically, the observed motion is decomposed into three MMCs: low-frequency smoothed (LFS) motion, high-frequency compensatory (HFC) motion, and shaky motion. LFS motion helps to largely stabilize videos, and HFC motion helps to recover missing motion to deal with over-smoothing. Subsequently, we present an MMC-based model to retrieve the desired smoothed motion, in which weighted nuclear norm and autoregression priors are used for LFS motion, while a sparsity prior is adopted for HFC motion. In addition, we design an adaptive weight setting scheme to detect rapid motions and to calculate the optimal weights. Finally, we develop a stabilization algorithm under the Alternating Direction Method of Multipliers (ADMM) framework. Experimental results demonstrate that our method can achieve high-quality results compared with that of other state-of-the-art stabilization methods in terms of robustness and efficiency, both quantitatively and qualitatively. Huicong Wu, Liang Xiao 0001, Le Sun 0002, Byeungwoo Jeon |
IEEE Trans. Multim. | 2 |
| 2022 | Learning Texture Enhancement Prior with Deep Unfolding Network for Snapshot Compressive Imaging
Mengying Jin, Zhihui Wei, Liang Xiao 0001 |
ACCV (3) | 3 |
| 2022 | Automatic Check-Out via Prototype-Based Classifier Learning from Single-Product Exemplars
Hao Chen 0052, Xiu-Shen Wei, Faen Zhang, Yang Shen 0006, Liang Xiao 0001 |
ECCV (25) | 6 |
| 2022 | Learning Transformations between Heterogeneous SAR and Optical Images for Change DetectionabstractChange detection based on heterogeneous images is challenging because of the distribution variance caused by imaging properties of different types of sensor. Most methods deal with this problem by transforming features into a common space. However, the lack of available labeled data limits the training of complex models and representation of heterogeneous distributions. In this paper, we propose to train a network via abundant unlabeled data by adopting cyclic adversarial pre-training in order to learn the relationship between heterogeneous distributions. After pre-training, for change detection, we introduce a constraint to maintain consistency of image content, to avoid the participation of changed pixels in training. Experiments on heterogeneous optical and SAR images prove the effectiveness of our proposed method. Zhenqing Chen, Jia Liu 0020, Liang Xiao 0001, Jiao Shi |
IGARSS | 5 |
| 2022 | An Orientation-Aware Anchor-Free Detector for Aerial Object DetectionabstractMost aerial object detectors mainly adopt anchor-based methods, which require manually designed rough anchors. However, these anchors usually contain background redundancy unrelated to the objects, decreasing detection performance and introducing additional calculations. In this paper, we propose an orientation-aware anchor-free network (OAF-Net) for object detection. OAF-Net first produces coarse oriented boxes by coarse location network based on pixel-level regression and then refines them with aligned features. In particular, we also apply a new metric to measure the orientation-aware center-ness, which is a weighting strategy of positive samples to reduce the different contributions of the low-quality detection boxes. Experiments on the DOTA1.0 and HRSC2016 datasets reveal that OAF-Net outperforms the state-of-the-art in terms of comprehensive detection performance for various objects with different scales and orientations in aerial images. Mudi Duan, Ran Meng, Liang Xiao 0001 |
IGARSS | 3 |
| 2022 | High Speed and Robust RGB-Thermal Tracking via Dual Attentive Stream Siamese NetworkabstractTo meet the need of robust tracking in some cases (e.g., extremely illumination and thermal crossover), plenty of RGB-T tracking methods have been proposed in recent years. However, many of them can hardly meet the real-time standard since they are difficult to balance the robust-ness and the speed. To address such problems, we propose a new Dual Attentive Siamese Network (DASN) for RGB-T tracking. Specifically, we use a dual-stream siamese deep learning network to model tracking as a similarity measure task. In addition, to promote the information propagation procedure between two modalities and suppress the inter-ference of background, we design the channel attention module (CAFE Module) and the channel-spatial attention module (CSAFE Module). What's more, the dual modality region proposal sub-network and the strategy of selecting proposal are constructed to boost the performance. The proposed DASN is trained end-to-end offline. Extensive ex-periments on three real RGB-T tracking datasets show that our tracker achieves very competive results with a high tracking speed over 140 frames per second. Code is released at https://github.com/easycodesniper-atk/SiamCSR.git. Chaoyang Guo, Liang Xiao 0001 |
IGARSS | 2 |
| 2022 | Spatial-Adaptive and Feature-Enhanced Siamese Network for Change DetectionabstractChange detection (CD) plays an increasingly important role in earth observation and reveals surface changes according to multi-temporal images. Although deep learning-based CD methods work well for their excellent modeling ability, objects in different size and shape are generally processed by the same filter kernels in feature extraction, which leads to spatial blurring and degrades the CD performance. In this paper, a spatial adaptive and feature enhanced (SAFE) siamese network is proposed to tackle this problem, where the SAFE consists of a spatial-adaptive (SA) part and a feature-enhanced (FE) part. Specifically, pixel belonging to different objects possesses its own spatial knowledge, which is captured by a soft fusion of multi-scale difference images (DIs) called SA part. Changed and unchanged areas are strengthened or weakened by the FE, which combines object features with each DI accordingly. Moreover, since there are more unchanged pixels than changed pixels, a weight-pair is introduced to balance changed and unchanged objects in the training process. The experimental results verify that compared with four representative CD algorithms, our proposed method performs best on the Change Detection Dataset (CDD). Yangguang Liu, Fang Liu 0034, Jia Liu 0020, Xu Tang 0004, Kaixuan Jiang, Liang Xiao 0001 |
IGARSS | 6 |
| 2022 | Manifold Augmentation Based Self-Supervised Contrastive Learning for Few-Shot Remote Sensing Scene ClassificationabstractDeep learning (DL)-based methods have achieved great success in the field of remote sensing scene images classification for the past few years. However, such methods usually require large amounts of labeled data. Compared with natural images, remote sensing images are relatively scarce and expensive to obtain, and DL methods lead to overfitting with very few samples. To address such problem, we propose a method which integrates manifold augmentation and self-supervised contrastive learning under meta-learning framework to cope with few-shot scene classification. The novelty of the paper is twofold: 1) we use manifold augmentation to expand labeled data and finetune the model to the specific task, and 2) self-supervised contrastive learning enforces the model to learn feature invariant to various scale and orientation of scene images. Extensive experimental results show that our proposed methods achieve remarkable few-shot classification performance on NWPU-RESISC45 datasets Yunrui Sheng, Liang Xiao 0001 |
IGARSS | 2 |
| 2022 | Siamese High-Resolution Network for Change DetectionabstractDeep learning for change detection can provide effective guidance in many applications, such as agricultural development, urban planning, disaster avoidance, etc. In this study, a Siamese deep learning network based on High-Resolution Network (HRNet) is proposed to generate accurate results. HRNet can integrate multi-dimensional features and output high-resolution results which have attracted attention due to its reliable feature extraction ability. In this paper, we extract the feature pairs of several different dimensions, including the two features behind the down-sampling in the stem stage which is an important part of HRNet. Moreover, feature ex-traction and intensive up-sampling tasks are completed by using a variety of feature fusion sub-networks, which are used to enhance the learning ability. Experiments show the superiority of the proposed Siamese HRNet on a widely used change detection dataset. Jia Liu 0020, Liang Xiao 0001, Jiao Shi |
IGARSS | 5 |
| 2022 | Learning a Coupled Multilinear Network for Unsupervised Hyperspectral-Multispectral Image FusionabstractFusing low resolution (LR) HSI with high resolution (HR) multispectral image (MSI) is an important technology to obtain HR hypersepctral image (HSI), which is hard to directly acquire due to the hardware limitation. Deep learning (DL) has been applied in HSI-MSI fusion, but the representability of DL networks for multidimensional (i.e., spectral-spatial) features still need improvement. And most DL HSI-MSI fusion networks are in supervised fashion, HR ground truth HSI is required for training, which is unavailable in reality. In this work, we investigate tensor theory, and propose a coupled multilinear network (CMuNet) for unsupervised HSI-MSI fusion, where deep image prior and degradation model can be jointly learned. CMuNet consists of coupled multilinear filtering subnets, it jointly represents the LR HSI and HR MSI as a random code and multidimensional features on spatial and spectral modes. The HR HSI is inferred with the random code, features on spatial modes of HR MSI and features on spectral mode of LR HSI. Experiments on several HSIs demonstrate the effectiveness of the proposed method. Jingxiang Yang, Liang Xiao 0001 |
IGARSS | 2 |
| 2022 | An Unsupervised Hyperspectral Image Fusion Method Based on Spectral Unmixing and Deep LearningabstractDue to the limitations of various hardware conditions, in practice, only high resolution multispectral and low resolution hyperspectral images are usually captured. In order to apply hyperspectral images in various fields, better quality hyperspectral images have become a problem to be solved. In this paper, we propose an image fusion method based on spectral unmixing, which effectively combines the advantages of multispectral images and low-resolution hyperspectral images to generate high-resolution hyperspectral images. To be specific, a deep learning model based on spectral decomposition is constructed, using multiplication iterative rules based on the traditional gradient descent algorithm to get initial high-resolution abundance and define degeneration networks to describe the spatial and spectral downsampling operations. Experiments show that this method can get fusion images better quality than other methods. Kexin Zheng, Abdolraheem Khader, Liang Xiao 0001 |
IGARSS | 3 |
| 2022 | Prototype-based classifier learning for long-tailed visual recognition
Xiu-Shen Wei, Shu-Lin Xu, Hao Chen 0052, Liang Xiao 0001, Yuxin Peng 0001 |
Sci. China Inf. Sci. | 4 |
| 2022 | Learning Deep Subspace Projection Prior for Dual-Camera Compressive Hyperspectral ImagingabstractCoded aperture snapshot spectral imaging (CASSI) captures the 3-D hyperspectral images (HSI) in the form of 2-D coded images. The dual-camera compressive hyperspectral imaging (DCCHI) can effectively improve the reconstruction quality by adding a parallel complementary panchromatic camera. Several regularization-based methods have been proposed for dual-camera reconstruction. However, the handcrafted priors of these methods are limited in representing the complex intrinsic structure of HSI. In this letter, we propose to learn deep subspace projection prior for dual-camera compressive reconstruction. We first design a deep subspace projection prior regularized dual-camera compressive reconstruction model and minimize it with alternative optimization. Then, we unfold the optimization process into a network. Specifically, the deep subspace projection prior learning leads to features with low-rank characteristics, which could efficiently exploit the spectral correlation of HSI. The dual-camera compressive reconstruction network is learned in an end-to-end manner. Extensive experiments substantiate the performance and efficiency of other start-of-the-art algorithms. Xiaoyang Chen 0005, Jingxiang Yang, Liang Xiao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | MC-JAFN: Multilevel Contexts-Based Joint Attentive Fusion Network for PansharpeningabstractPansharpening refers to a spatial–spectral contexts fusion procedure to produce high-quality multispectral (MS) images by retaining the fine spatial resolution of the panchromatic (PAN) images and the high spectral content of the MS images. This letter presents a novel end-to-end dual-branch deep learning-based fusion framework, exploiting the network to extract spatial and spectral contexts progressively in two separate branches level by level. For each level contexts extraction layer, a dual-branch weighted attentive fusion module is integrated to boost the important contexts aggregation and details injection while suppressing unimportant ones. Experimental results on two real datasets show that our method outperforms state-of-the-art methods in both objective metrics and image quality by visual appearance. Zhikang Xiang, Liang Xiao 0001, Wenzi Liao, Wilfried Philips |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Domain-Adaptive Few-Shot Learning for Hyperspectral Image ClassificationabstractRecently, hyperspectral image (HSI) classification by deep learning is flourishing. However, only a few labeled samples are available in practice since it is time-and-labor-consuming to label pixels in HSI (called target domain). This paper proposes a domain-adaptive few-shot learning (DAFSL) method to tackle this problem. Specifically, some other HSIs (called source domain) with large labeled samples are fully used as complementary information and a generative architecture is employed to adapt embedded features in source domain to that of target domain. We first perform domain adaptation with unsupervised learning. In details, the embedded features are generated by the encoder of an autoencoder, where both source and target samples could be well recovered and the reconstruction loss is used to measure the gap between source domain and target domain. At the same time, the embedded features are put into a metric space for classification in source domain and the encoder parameter is fine-tuned together with the classifier in target domain with few labels, so that both general and discriminative features are well captured. The experiment results show that DAFSL outperforms the other mainstream methods with limited labeled samples. Andi Zhang 0003, Fang Liu 0034, Jia Liu 0020, Xu Tang 0004, Wenfei Gao, Liang Xiao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2022 | Evolving Connections in Group of Neurons for Robust LearningabstractArtificial neural networks inspired from the learning mechanism of the brain have achieved great successes in machine learning, especially those with deep layers. The commonly used neural networks follow the hierarchical multilayer architecture with no connections between nodes in the same layer. In this article, we propose a new group architectures for neural-network learning. In the new architecture, the neurons are assigned irregularly in a group and a neuron may connect to any neurons in the group. The connections are assigned automatically by optimizing a novel connecting structure learning probabilistic model which is established based on the principle that more relevant input and output nodes deserve a denser connection between them. In order to efficiently evolve the connections, we propose to directly model the architecture without involving weights and biases which significantly reduce the computational complexity of the objective function. The model is optimized via an improved particle swarm optimization algorithm. After the architecture is optimized, the connecting weights and biases are then determined and we find the architecture is robust to corruptions. From experiments, the proposed architecture significantly outperforms existing popular architectures on noise-corrupted images when trained only by pure images. Jia Liu 0020, Maoguo Gong, Liang Xiao 0001, Fang Liu 0034 |
IEEE Trans. Cybern. | 3 |
| 2022 | A Total Variation Regularized Bipartite Network for Unsupervised Change DetectionabstractDetecting changes in complicated remote sensing images have been gaining much attention. One of the main challenges lies in how to detect intrinsic changes robustly while avoiding the false alarms caused by various challenging factors, such as spatial illumination variations, small viewpoint differences, noises and outliers between multi-temporal remote sensing images. To reduce the influence of these factors, in this paper, we propose an unsupervised joint learning model based on a total variation regularization and bipartite deep convolutional neural network, called total variation regularized bipartite network (TVRBN). In this model, parametric feature differences are initialized by a bipartite autoencoder. An objective function is defined as a parametric feature difference term integrated with a change pre-detection constraint term and a total variation regularization to the change probability map with intrinsic changes. Then the unified objective function is optimized jointly to learn the bipartite network parameters and detect an intrinsic change map. Due to the intrinsic difference constraints in feature space and total variation regularization, the proposed TVRBN method can compute higher-quality and smoother change maps, suppress non-intrinsic changes, and overcome small viewpoint differences in comparable images. Extensive experiments on both homogeneous and heterogeneous images demonstrate the robustness of the proposed method by comparing it with state-of-the-art methods. Ling Hu 0003, Jia Liu 0020, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Graph Convolutional Sparse Subspace Coclustering With Nonnegative Orthogonal Factorization for Large Hyperspectral ImagesabstractSparse subspace clustering (SSC) is a representative data clustering paradigm that has been broadly applied in the unsupervised classification of hyperspectral images (HSIs). Existing SSC methods usually produce a subspace affinity matrix between representations of hyperspectral pixels first, followed by spectral clustering for the affinity matrix. To this end, the separated framework fails to exploit the dualities contained in both features and pixels or higher order entities at the same time, and thus, it is difficult to compute coclusters simultaneously. In addition, SSC methods often require expensive computational consumption and memory capacity to approximate the spectral decomposition of the affinity matrix, thus hindering the applicability of SSC methods for large HSIs. To overcome these limitations, we propose a novel graph convolutional sparse subspace coclustering (GCSSC) model with nonnegative orthogonal factorization for large HSIs in which affinity matrix learning and spectral coclustering are integrated into a unified optimizing model to obtain the optimal clustering results. Specifically, to form a more compact self-representation, the superpixel-based adaptive dictionary construction strategy is proposed instead of the global dictionary to precisely represent the pixels. To explore the spatial–contextual and spectral neighboring characteristics between dictionary atoms, graph convolution is incorporated into the dictionary atoms to aggregate the local neighborhood information, and the affinity matrix in the proposed coclustering framework is constructed under a joint sparsity constrained representation model. To reduce high computational consumption and memory capacity, a nonnegative orthogonal factorization constraint is proposed to offer an alternative spectral clustering for hyperspectral pixels and dictionary atoms simultaneously. The clustering performance of the proposed method is evaluated for three classical HSIs, and the experimental results illustrate that the proposed method is memory and computationally efficient and outperforms the state-of-the-art HSI clustering methods. Nan Huang 0001, Liang Xiao 0001, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Bipartite Graph Partition-Based Coclustering Approach With Graph Nonnegative Matrix Factorization for Large Hyperspectral ImagesabstractClustering large hyperspectral images (HSIs) is a very challenging problem because large HSIs have high dimensionality, large spectral variability, and large computational and memory consumption. Recently, sparse subspace clustering (SSC) has achieved remarkable success in HSI clustering. However, most SSC-based methods suffer from the following bottlenecks for large HSIs: 1) high computational consumption and memory space during the construction of the similarity matrix and decomposition of the graph Laplacian matrix and 2) failure to capture the relationships among dictionary atoms, sparse coefficients, and hyperspectral pixels. To address these challenges, we propose a novel algorithm that extends SSC to cocluster large HSIs, called bipartite graph partition with graph nonnegative matrix factorization (BGP-GNMF). Specifically, to fully explore the characteristics of the spectral and spatial contexts in HSIs, we propose a novel superpixel and pixel coclustering framework with bipartite graph partitioning in the joint sparse representation domain, where superpixel-based dictionary atoms are defined as disjoint vertex sets of the bipartite graph and the joint sparsity representation is mapped into the adjacency matrix of the undirected bipartite graph. To overcome the challenges of high computational consumption and large memory space for large HSIs, the bipartite graph partition with orthonormal constrained nonnegative matrix factorization is proposed to simultaneously cluster the structured dictionary atoms and hyperspectral pixels with an indicator matrix. Finally, to exploit the intrinsic geometry of HSIs, we incorporate manifold regularization into the bipartite graph partition to improve final clustering accuracy. The effectiveness and efficiency of the proposed method are verified on three classical HSIs, and the experimental results illustrate the superiority of the proposed method compared with other state-of-the-art HSI clustering methods. Nan Huang 0001, Liang Xiao 0001, Yang Xu 0006, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Joint Variation Learning of Fusion and Difference Features for Change Detection in Remote Sensing ImagesabstractRemote sensing (RS) image change detection (CD) is an earth observation technique for detecting surface changes in the same area during a period. With the rapid development of deep learning, various deep neural networks especially Siamese ones have been widely used in the field of CD. However, they have the deficiency of insufficient contextual information aggregation, resulting in false and missed detections, and it is difficult to refine the detection of change edges. To alleviate these problems and obtain more accurate results, we propose an efficient self-weighted spatial-temporal attention network (SSANet). In contrast to the Siamese structure, our network is a novel joint learning framework composed of fusion sub-network, difference sub-network, and decoder. Fusion sub-network is used to extract multiscale object features where we propose a multi-core channel-aligning attention (MCA) module to capture the long-range semantic information for multi-scale context aggregation. Difference sub-network is used to extract the difference variation features, where we propose a feature differential reconfiguration (FDR) module to learn the temporal change information. FDR can effectively filter change information and reconstruct features to improve the perception of changed regions. To better balance the MCA and FDR modules, an asymmetric weighting (AW) module is proposed in the decoder to self-weight the multi-scale features and generate the change map. Experiments demonstrate the efficiency of proposed sub-networks and modules, and the state-of-the-art performance of SSANet. Kaixuan Jiang, Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | A Unified Pansharpening Method With Structure Tensor Driven Spatial Consistency and Deep Plug-and-Play PriorsabstractPansharpening is to generate a high resolution multispectral (HRMS) image by preserving the spectral information from a low resolution multispectral (LRMS) image and the spatial content from a panchromatic (PAN) image. This article proposes a unified pansharpening method with structure tensor driven spatial consistency and deep plug and play priors. First, the spectral fidelity constraint between HRMS and LRMS is imposed for preserving spectral information. Second, the structure tensor is applied to characterize the spatial geometric information of HRMS and PAN images, thus the structure tensor driven spatial consistency prior between HRMS and PAN is particularly exploited for preserving spatial content. Moreover, by generalizing the convolution neural network (CNN) fusion method into a unified variational framework, a novel CNN-based deep plug and play prior between the HRMS and CNN-based fused MS images is also proposed to generate more image characteristics for further preserving spectral information and spatial content. Besides, the proposed model is solved by the alternating direction method of multipliers (ADMM) algorithm. Finally, extensive experiments on both reduced and full resolution by comparing with various representative approaches exhibit the excellent performance of the proposed method. Pengfei Liu 0002, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Adaptive Graph Convolutional Network for PolSAR Image ClassificationabstractPolarimetric synthetic aperture radar (PolSAR) image classification is one of the hottest issues in remote sensing, where studies on pixel-level information and relationship are of great significance. In this article, graph convolutional network (GCN) is employed to accomplish this pixel-level task benefiting from its excellent capability in structure exploration and information propagation between different pixels. To reduce the communication burden between various PolSAR pixels and high computational cost for the whole PolSAR image, an adaptive GCN (AdapGCN) consisting of pixel-centered subgraphs is proposed in this article. In the AdapGCN, a data-adaptive kernel and a spatial-adaptive kernel are introduced to, respectively, model data structure and spatial structure for PolSAR image. Moreover, a multiscale learning structure is integrated to further explore complicated relations between pixels. Extensive comparative evaluations validate the superiority of our new AdapGCN model for PolSAR image classification over a wide range of state-of-the-art methods on three challenging benchmarks. Fang Liu 0034, Jingya Wang 0001, Xu Tang 0004, Jia Liu 0020, Xiangrong Zhang, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Model Inspired Autoencoder for Unsupervised Hyperspectral Image Super-ResolutionabstractThis article focuses on hyperspectral image (HSI) super-resolution that aims to fuse a low-spatial-resolution HSI and a high-spatial-resolution multispectral image to form a high-spatial-resolution HSI (HR-HSI). Existing deep learning-based approaches are mostly supervised that rely on a large number of labeled training samples, which is unrealistic. The commonly used model-based approaches are unsupervised and flexible but rely on handcrafted priors. Inspired by the specific properties of model, we make the first attempt to design a model-inspired deep network for HSI super-resolution in an unsupervised manner. This approach consists of an implicit autoencoder network built on the target HR-HSI that treats each pixel as an individual sample. The nonnegative matrix factorization (NMF) of the target HR-HSI is integrated into the autoencoder network, where the two NMF parts, spectral and spatial matrices, are treated as decoder parameters and hidden outputs, respectively. In the encoding stage, we present a pixelwise fusion model to estimate hidden outputs directly and then reformulate and unfold the model’s algorithm to form the encoder network. With the specific architecture, the proposed network is similar to a manifold prior-based model and can be trained patch by patch rather than the entire images. Moreover, we propose an additional unsupervised network to estimate the point spread function and spectral response function. Experimental results conducted on both synthetic and real datasets demonstrate the effectiveness of the proposed approach. Zebin Wu 0001, Liang Xiao 0001, Xiaojun Wu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Nonconvex Pansharpening Model With Spatial and Spectral Gradient Difference-Induced Nonconvex Sparsity PriorsabstractThis article proposed a nonconvex variational model for pansharpening with spatial and spectral gradient difference-induced nonconvex sparsity priors (PSSGDNSP), which can fuse the panchromatic (Pan) and low-resolution (LR) multispectral (MS) images to generate the high-resolution (HR) MS image. More particularly, the proposed PSSGDNSP model exploits the spatial gradient difference-induced nonconvex$l_{1/2}$sparsity prior between HR MS and Pan, and the spectral gradient difference-induced nonconvex$l_{1/2}$sparsity prior between HR and LR MS. Consequently, our proposed PSSGDNSP model well preserves both the spatial and spectral information. In fact, our proposed band-coupled model treats the MS image like a third-order tensor so that the intrinsic band correlation of the MS image can be fully kept. Moreover, we solve our proposed PSSGDNSP model by applying the alternating direction method of multipliers (ADMM) method. Finally, the experiments fully validate the superiority and performance of our proposed PSSGDNSP method. Pengfei Liu 0002, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Multilevel Superpixel Structured Graph U-Nets for Hyperspectral Image ClassificationabstractLimited by the shape-fixed kernels, convolutional neural networks (CNNs) are usually difficult to model difform land covers in hyperspectral images (HSIs), leading to inadequate land use. Recently, benefiting from the ability to conduct shape-adaptive convolutions and model complex patterns in graph-structured data, graph convolutional networks (GCNs) have been applied to HSI classification. However, due to the massive computation in GCNs, HSI is usually pretreated into a graph based on a specific superpixel segmentation, which limits the modeling of spatial topologies to the same scale. To break this limitation, we propose a multilevel superpixel structured graph U-Net (MSSGU) to learn multiscale features on multilevel graphs. Specifically, we construct several hierarchical segmentations from fine to coarse by progressively merging adjacent superpixels and then convert them into multilevel graphs. Meanwhile, based on the merging relations between hierarchical superpixels, we establish the pooling and unpooling functions to transfer features from one graph to another, thereby enabling different-level graphs to collaborate in a single network. Different from concatenating different-scale features straightforwardly in the feature fusion stage, MSSGU fuses them in a coarse-to-fine progressive manner, which can generate subtler fusion features adaptive to the pixelwise classification task. Moreover, we use a CNN instead of GCN to extract and fuse the pixel-level features, which greatly reduces the computation. Such a hybrid U-Net can exploit features of HSIs from a multiscale hierarchical perspective, and its performance has been proven competitive with other deep-learning-based methods by extensive experiments on three benchmark datasets. Qichao Liu, Liang Xiao 0001, Jingxiang Yang, Zhihui Wei |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Probabilistic Model Based on Bipartite Convolutional Neural Network for Unsupervised Change DetectionabstractThis article presents a probabilistic model based on a bipartite convolutional architecture for unsupervised change detection. We aim to develop a robust change detection method that can adapt to different types of data and scenarios for multitemporal coregistered remote sensing images of the same spatial resolution. On the premise of coregistration, unsupervised change detection usually suffers from the distinct appearances (different intensities or data structures) of the same object in multitemporal images, such as images obtained in different climatic conditions (season, illumination, and so on), and by different and even heterogeneous sensors. Since change detection in heterogeneous images can also adapt to other scenarios, many methods have been proposed recently focusing on such data, but most of them are limited by the need for labeled data or by specific assumptions. With the excellent and flexible feature learning capability of neural networks, we model the change detection into a Gibbs probabilistic model based on a bipartite neural network. The model is driven by an energy function defined as the squared feature distance, which is the core of change detection. Via optimizing the model, the difference degree of each pixel is automatically obtained for further identification. The probabilistic model learns to capture the distribution in an unsupervised way. Therefore, the proposed method can adapt to various scenarios without being trained by labeled data. Experiments on different types of data and scenarios demonstrate the superiority of the proposed method. Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | ADMM-HFNet: A Matrix Decomposition-Based Deep Approach for Hyperspectral Image FusionabstractHyperspectral image (HSI) fusion refers to the reconstruction of a high-resolution HSI by fusing a low-resolution HSI (LR-HSI) and a high-resolution multispectral image (HR-MSI) over the same scene. Recently, researchers have proposed many approaches to handle this issue. However, most of them assume that both the spatial and spectral degradation functions are known, which are often limited or unavailable in reality. This article presents a novel model-driven deep network based on matrix decomposition, which considers spectral correlations and reasonably embeds the well-known observation models. Specifically, the proposed method decomposes the desired HSI into spectral basis and coefficients. The spectral basis can be estimated from the LR-HSI via singular value decomposition. To learn the coefficients, a learning model is constructed by merging the observation models, matrix decomposition, and sparsity into a concise single formulation. For solving the proposed model, a deep framework is built by unrolling the alternating direction method of multipliers (ADMM), dubbed as ADMM-HFNet, where the involved parameters can be learned adaptively. It is worth noting that the spectral basis cannot fully represent the desired HSI. Therefore, another model is constructed here to supplement the approximation error, which can also be embedded in the deep network. After checking on three datasets, it is found that the proposed method stands out from advanced competing techniques in both quality measures and visual effects. Dunbin Shen, Zebin Wu 0001, Jinlong Yang 0002, Liang Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | BDANet: Multiscale Convolutional Neural Network With Cross-Directional Attention for Building Damage Assessment From Satellite ImagesabstractFast and effective responses are required when a natural disaster (e.g., earthquake and hurricane) strikes. Building damage assessment from satellite imagery is critical before relief effort is deployed. With a pair of predisaster and postdisaster satellite images, building damage assessment aims at predicting the extent of damage to buildings. With the powerful ability of feature representation, deep neural networks have been successfully applied to building damage assessment. Most existing works simply concatenate predisaster and postdisaster images as input of a deep neural network without considering their correlations. In this article, we propose a novel two-stage convolutional neural network for building damage assessment, called BDANet. In the first stage, a U-Net is used to extract the locations of buildings. Then, the network weights from the first stage are shared in the second stage for building damage assessment. In the second stage, a two-branch multiscale U-Net is employed as the backbone, where predisaster and postdisaster images are fed into the network separately. A cross-directional attention module is proposed to explore the correlations between predisaster and postdisaster images. Moreover, CutMix data augmentation is exploited to tackle the challenge of difficult classes. The proposed method achieves state-of-the-art performance on a large-scale dataset—xBD. The code is available athttps://github.com/ShaneShen/BDANet-Building-Damage-Assessment. Sijie Zhu, Taojiannan Yang, Chen Chen 0001, Delu Pan, Jianyu Chen 0003, Liang Xiao 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Detail-Injection-Model-Inspired Deep Fusion Network for PansharpeningabstractPansharpening is an image fusion procedure, which aims to produce a high spatial resolution multispectral image by combining a low spatial resolution multispectral image and a high spatial resolution panchromatic image. The most popular and successful paradigm for pansharpening is the framework known as detail injection, while it cannot fully exploit complex and non-linear complementary features of both images. In this paper, we propose a detail injection model inspired deep fusion network for pansharpening (DIM-FuNet). Firstly, by treating pansharpening as a complicated and non-linear details learning and injection problem, we establish a unified optimizing detail-injection model with triple detail fidelity terms: 1) a band-dependent spatial detail fidelity term, 2) a local detail fidelity term and 3) a complicated details synthesis term. Secondly, the model is optimized via the iterative gradient descent and unfolded into a deep convolutional neural network. Subsequently, the unrolling network has triple branches, in which, a point-wise convolutional sub-network, a depth-wise convolutional sub-network are corresponding to the former two detail constrained terms, and an adaptive weighted reconstruction module with a fusion sub-network to aggregate details of two branches and synthesis the final complicated details. Finally, the deep unrolling network is trained in end-to-end manners. Different from traditional deep fusion networks, the architecture design of DIM-FuNet is guided by the optimizing model and thus promotes better interpretability. Experimental results on reduced and full-resolution demonstrate the effectiveness of the proposed DIM-FuNet which achieves the best performance compared with the state-of-the-art pansharpening method. Zhikang Xiang, Liang Xiao 0001, Jingxiang Yang, Wenzi Liao, Wilfried Philips |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Variational Regularization Network With Attentive Deep Prior for Hyperspectral-Multispectral Image FusionabstractHyperspectral–multispectral image (HSI-MSI) fusion relies on a robust degradation model and data prior, where the former describes the degeneration of HSI in the spectral and spatial domains, and the latter reveals the latent statistics of the expected high-resolution (HR) HSI. In practice, the degradation model is often unknown, and the data prior is usually too complicated to be expressed analytically. In this study, we propose a variational network for HSI-MSI fusion (VaFuNet), in which the degradation model and data prior are implicitly represented by a deep learning network and jointly learned from the training data. A variational fusion model regularized by deep prior is first proposed, and then, it is optimized via a half-quadratic splitting and unfolded into a deep network. The deep prior is implicitly represented by a proximity operator. Due to the structural self-similarity, HSI possesses structural recurrences across different scales. To exploit such nonlocal prior and enhance the representability of network, we also propose a multiscale nonlocal attention and embed it into the deep prior proximity. The degradation model and deep prior proximity are jointly learned via end-to-end training. Experimental results on simulated and real-life HSI datasets demonstrate the effectiveness of the proposed VaFuNet HSI-MSI fusion method. Jingxiang Yang, Liang Xiao 0001, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Self-Supervised Multi-Category Counting Networks for Automatic Check-OutabstractThe practical task of Automatic Check-Out (ACO) is to accurately predict the presence and count of each product in an arbitrary product combination. Beyond the large-scale and the fine-grained nature of product categories as its main challenges, products are always continuously updated in realistic check-out scenarios, which is also required to be solved in an ACO system. Previous work in this research line almost depends on the supervisions of labor-intensive bounding boxes of products by performing a detection paradigm. While, in this paper, we propose a Self-Supervised Multi-Category Counting (S2MC2) network to leverage the point-level supervisions of products in check-out images to both lower the labeling cost and be able to return ACO predictions in a class incremental setting. Specifically, as a backbone, our S2MC2 is built upon a counting module in a class-agnostic counting fashion. Also, it consists of several crucial components including an attention module for capturing fine-grained patterns and a domain adaptation module for reducing the domain gap between single product images as training and check-out images as test. Furthermore, a self-supervised approach is utilized in S2MC2 to initialize the parameters of its backbone for better performance. By conducting comprehensive experiments on the large-scale automatic check-out dataset RPC, we demonstrate that our proposed S2MC2 achieves superior accuracy in both traditional and incremental settings of ACO tasks over the competing baselines. Hao Chen 0052, Yangzhun Zhou, Jun Li 0027, Xiu-Shen Wei, Liang Xiao 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Spatial-Spectral Total Variation Constrained Collaborative Tensor Regularization for Dual-Camera Compressive Hyperspectral ImagingabstractIn this paper, we propose a novel tensor-based approach to improve the reconstruction performance for dual-camera compressive hyperspectral imaging. We formulate a coupled tensor decomposition model to maintain the consistency of the spatial structure of the HSI and the panchromatic image. We introduce a regularizer over the core tensor to collaboratively promote global spatial-spectral correlations in HSI. Besides, we incorporate an anisotropic spatial-spectral total variation (SSTV) regularization to characterize the piecewise smooth structure of the HSI. Then the alternating direction method of multipliers (ADMM) algorithm is applied to the optimization problem. Experimental results on a public dataset demonstrate the superiority of the proposed approach. Zhenghui Liang, Yang Xu 0006, Liang Xiao 0001, Zhihui Wei |
IGARSS | 3 |
| 2021 | Multi-Supervised Recursive-CNN for Hyperspectral and Multispectral Image FusionabstractDeep learning has been widely used in remote sensing images fusion in recent years. However, many deep learning based methods use interpolation or a deconvolutional layer to upsample the low-resolution hyperspectral(LRHS) image, then the upsampled image is concated with the high-resolution multispectral(HRMS) image to fed the network, which lead to the spatial-spectral information loss. In this study, we propose a multi-supervised recursive convolutional neural network(MRCNN) for HS/MS images fusion. Specifically, we use the upsampling recursive sub-net(URSN) to upsample the LRHS image, which can effectively avoid the spatial-spectral information loss. In addition, to make our deep network much lighter, we introduce recursive learning to the network by using a residual block as the recursive unit for several recursions. Finally, a multi-supervised learning strategy is adopted for enhancing the gradient propagation and avoiding vanishing and exploding gradients. Simulated experiments on Cave and Moffett Field datasets show that the proposed network outperforms many state-of-the-art ones. Yuda Lu, Jingxiang Yang, Liang Xiao 0001 |
IGARSS | 3 |
| 2021 | Hyperspectral Image Denoising with Collaborative Total Variation and Low Rank RegularizationabstractVariational regularization methods are the mainstream methods typically adopted for hyperspectral images (HSIs) denoising, which borrow architectures originally developed for RG-B images, exhibiting limitations when cope with HSI data cubes. To overcome this limitation, this paper proposes a new collaborative total variation and low-rank regularization model (LRCTV) to remove mixed noise from HSI data. Specifically, the proposed method unfolds the HSI cube into 2D extended spectral-matrix, then obtains the horizontal and vertical gradient matrices, and applies 2D collaborative norm to the gradient matrix to model the directional selective smoothness, while the matrix nuclear norm is used to model the low rank structure. Experimental results on both simulated and real HSI datasets validated that the proposed method outperformed several state-of-the-art methods. Jinhuan Xu, Liang Xiao 0001 |
IGARSS | 3 |
| 2021 | Intra- and Inter-frame Iterative Temporal Convolutional Networks for Video StabilizationabstractVideo jitter is an uncomfortable product of irregular lens motion in time sequence. How to extract motion state information in a period of continuous video frames is a major issue for video stabilization. In this paper, we propose a novel sequence model, Intra- and Inter-frame Iterative Temporal Convolutional Networks (I3TC-Net), which alternatively transfer the spatial-temporal correlation of motion within and between frames. We hypothesize that the motion state information can be represented by transmission states. Specifically, we employ combination of Convolutional Long Short-Term Memory (ConvLSTM) and embedded encoder-decoder to generate the latent stable frame, which are used to update transmission states iteratively and learn a global homography transformation effectively for each unstable frame to generate the corresponding stabilized result along the time axis. Furthermore, we create a video dataset to solve the lack of stable data and improve the training effect. Experimental results show that our method outperforms state-of-the-art results on publicly available videos, such as 5.4 points improvements in stability score. The project page is available at https://github.com/root2022IIITC/IIITC. Haopeng Xie, Liang Xiao 0001, Huicong Wu |
MMAsia | 2 |
| 2021 | Deep associative learning for neural networks
Jia Liu 0020, Fang Liu 0001, Liang Xiao 0001 |
Neurocomputing | 4 |
| 2021 | Hypergraph-Regularized Low-Rank Subspace Clustering Using Superpixels for Unsupervised Spatial-Spectral Hyperspectral ClassificationabstractLow-rank subspace representations have been observed to be well-suited to hyperspectral imagery, which tends to have a global structure composed of a small number of ground-cover signatures, and additional graph-based regularization can further incorporate local information. However, in the context of unsupervised classification, existing approaches typically limit consideration to simple graphs built on spectral information alone. In contrast, a hypergraph-based low-rank subspace clustering is proposed to capture a more complex manifold structure. In addition, basing the hypergraph on a superpixel segmentation of the image exploits structure that is meaningful both spatially as well as spectrally. The experimental results reveal performance for the proposed superpixel-hypergraph approach superior to that of competing techniques representative of several prominent classes of unsupervised classification for hyperspectral imagery. Jinhuan Xu, James E. Fowler, Liang Xiao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Hybrid Local and Nonlocal 3-D Attentive CNN for Hyperspectral Image Super-ResolutionabstractA deep convolutional neural network (CNN) has shown its great potential in hyperspectral image (HSI) super-resolution (SR). Integrating CNN with attention mechanism is expected to boost the SR performance. However, how to learn attention along the spectral, spatial, and channel dimensions of HSI is still an open issue, and the current attention mechanism is not efficient in capturing long-range interdependency in HSI. In this letter, we first design a local 3-D attention module to learn the spectral-spatial-channel attention by exploiting local contextual information in HSI. Then, we propose a nonlocal 3-D attention module, in which the long-range interdependency in HSI can be exploited for attention learning. By jointly embedding the local and nonlocal attention in a residual 3-D CNN, a hybrid local and nonlocal 3-D attentive CNN can be built for HSI SR. The experimental results show that local and nonlocal attention formulation leads to competitive SR performance. Jingxiang Yang, Liang Xiao 0001, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | CNN-Enhanced Graph Convolutional Network With Pixel- and Superpixel-Level Feature Fusion for Hyperspectral Image ClassificationabstractRecently, the graph convolutional network (GCN) has drawn increasing attention in the hyperspectral image (HSI) classification. Compared with the convolutional neural network (CNN) with fixed square kernels, GCN can explicitly utilize the correlation between adjacent land covers and conduct flexible convolution on arbitrarily irregular image regions; hence, the HSI spatial contextual structure can be better modeled. However, to reduce the computational complexity and promote the semantic structure learning of land covers, GCN usually works on superpixel-based nodes rather than pixel-based nodes; thus, the pixel-level spectral–spatial features cannot be captured. To fully leverage the advantages of the CNN and GCN, we propose a heterogeneous deep network called CNN-enhanced GCN (CEGCN), in which CNN and GCN branches perform feature learning on small-scale regular regions and large-scale irregular regions, and generate complementary spectral–spatial features at pixel and superpixel levels, respectively. To alleviate the structural incompatibility of the data representation between the Euclidean data-oriented CNN and non-Euclidean data-oriented GCN, we propose the graph encoder and decoder to propagate features between image pixels and graph nodes, thus enabling the CNN and GCN to collaborate in a single network. In contrast to other GCN-based methods that encode HSI into a graph during preprocessing, we integrate the graph encoding process into the network and learn edge weights from training data, which can promote the node feature learning and make the graph more adaptive to HSI content. Extensive experiments on three data sets demonstrate that the proposed CEGCN is both qualitatively and quantitatively competitive compared with other state-of-the-art methods. Qichao Liu, Liang Xiao 0001, Jingxiang Yang, Zhihui Wei |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Efficient Deep Learning of Nonlocal Features for Hyperspectral Image ClassificationabstractDeep-learning-based methods, such as convolution neural network (CNN), have demonstrated their efficiency in hyperspectral image (HSI) classification. These methods can automatically learn spectral-spatial discriminative features within local patches. However, for each pixel in an HSI, it is not only related to its nearby pixels but also has connections to pixels far away from itself. Therefore, to incorporate the long-range contextual information, a deep fully convolutional network (FCN) with an efficient nonlocal module, named ENL-FCN, is proposed for HSI classification. In the proposed framework, a deep FCN considers an entire HSI as input and extracts spectral-spatial information in a local receptive field. The efficient nonlocal module is embedded in the network as a learning unit to capture the long-range contextual information. Different from the traditional nonlocal neural networks, the long-range contextual information is extracted in a specially designed criss-cross path for computation efficiency. Furthermore, using a recurrent operation, each pixel's response is aggregated from all pixels of HSI. The benefits of our proposed ENL-FCN are threefold: 1) the long-range contextual information is incorporated effectively; 2) the efficient module can be freely embedded in a deep neural network in a plug-and-play fashion; and 3) it has much fewer learning parameters and requires less computational resources. The experiments conducted on three popular HSI data sets demonstrate that the proposed method achieves state-of-the-art classification performance with lower computational cost in comparison with several leading deep neural networks for HSI. Sijie Zhu, Chen Chen 0001, Qian Du 0001, Liang Xiao 0001, Jianyu Chen 0003, Delu Pan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Sparse Coding Driven Deep Decision Tree Ensembles for Nucleus Segmentation in Digital Pathology ImagesabstractAutomating generalized nucleus segmentation has proven to be non-trivial and challenging in digital pathology. Most existing techniques in the field rely either on deep neural networks or on shallow learning-based cascading models. The former lacks theoretical understanding and tends to degrade performance when only limited amounts of training data are available while the latter often suffers from limitations for generalization. To address these issues, we propose sparse coding driven deep decision tree ensembles (ScD2TE), an easily trained yet powerful representation learning approach with performance highly competitive to deep neural networks in the generalized nucleus segmentation task. We explore the possibility of stacking several layers based on fast convolutional sparse coding–decision tree ensemble pairwise modules and generate a layer-wise encoder–decoder architecture with intra-decoder and inter-encoder dense connectivity patterns. Under this architecture, all the encoders share the same assumption across the different layers to represent images and interact with their decoders to give fast convergence. Compared with deep neural networks, our proposed ScD2TE does not require back-propagation computation and depends on less hyper-parameters. ScD2TE is able to achieve a fast end-to-end pixel-wise training in a layer-wise manner. We demonstrated the superiority of our segmentation method by evaluating it on the multi-disease state and multi-organ dataset where consistently higher performances were obtained for comparison against other state-of-the-art deep learning techniques and cascading methods with various connectivity patterns. Jie Song 0014, Liang Xiao 0001, Mohsen Molaei, Zhichao Lian |
IEEE Trans. Image Process. | 2 |
| 2021 | Simultaneous Video Stabilization and Rolling Shutter RemovalabstractDue to the delay in the row-wise exposure and the lack of stable support when a photographer holds a CMOS camera, video jitter and rolling shutter distortion are closely coupled degradations in the captured videos. However, previous methods have rarely considered both phenomena and usually treat them separately, with stabilization approaches that are unable to handle the rolling shutter effect and rolling shutter removal algorithms that are incapable of addressing motion shake. To tackle this problem, we propose a novel method that simultaneously stabilizes and rectifies a rolling shutter shaky video. The key issue is to estimate both inter-frame motion and intra-frame motion. Specifically, for each pair of adjacent frames, we first estimate a set of spatially variant inter-frame motions using a neighbor-motion-aware local motion model, where the classical mesh-based model is improved by introducing a new constraint to enhance the neighbor motion consistency. Then, different from other 2D rolling shutter removal methods that assume the pixels in the same row have a single intra-frame motion, we build a novel mesh-based intra-frame motion calculation model to cope with the depth variation in a mesh row and obtain more faithful estimation results. Finally, temporal and spatial motion constraints and an adaptive weight assignment strategy are considered together to generate the optimal warping transformations for different motion situations. Experimental results demonstrate the effectiveness and superiority of the proposed method when compared with other state-of-the-art methods. Huicong Wu, Liang Xiao 0001, Zhihui Wei |
IEEE Trans. Image Process. | 2 |
| 2020 | Video Deblurring Via 3d CNN and Fourier Accumulation LearningabstractCamera shake and target movement often leads to undesirable image blurring in videos. How to exploit spatial-temporal information of adjacent frames and reduce the processing time of deblurring are two major issues in video deblurring. In this paper, we propose a simple yet effective Fourier accumulation embedded 3D convolutional encoder-decoder network for video deblurring. Firstly, a 3D convolutional encoder-decoder module is constructed to extract multiscale spatial-temporal deep features and generate intermediate deblurred frames with complementary information which is beneficial for the deblurring of each frame. Then we embed a Fourier accumulation module following the 3D convolutional encoder-decoder, the Fourier accumulation module could fuse intermediate deblurred frames with learned weights in Fourier domain and then produce shaper deblurred frames. Experimental results show that our method has competitive performance compared with other state-of-the-art methods. Liang Xiao 0001, Jingxiang Yang |
ICASSP | 2 |
| 2020 | Locally Constrained Collaborative Representation Based Fisher's LDA for Clustering of Hyperspectral ImagesabstractClustering of hyperspectral images (HSIs) is a challenging task, due to high dimensional, unbalance data distribution and complex spectral-spatial structures. Existing methods mainly use either data spectral attribute only or spatial similarity only without gathering them appropriately. In this paper, under an extended linear discriminant analysis (LDA) framework, a new method, locally constrained collaborative representation based Fisher's LDA for clustering (LCR-FLDA), is proposed for the clustering of HSIs. We propose a New Fisher's LDA (FLDA) model by optimizing the integration of K-means-Laplacian and N-cut clustering to fully utilize the spectral-spatial discriminative representation information both in data and feature domain. Then, using locally constrained collaboration representation method to build a similarity matrix in FLDA to incorporate both local relationship structure and attribute information. Experimental results were conducted to illustrate the effectiveness of the proposed method. Nan Huang 0001, Liang Xiao 0001 |
IGARSS | 3 |
| 2020 | PERONA-MALIK DIFFUSION DRIVEN CNN FOR SUPERVISED CLASSIFICATION OF HYPERSPECTRAL IMAGESabstractWe present a novel auto-machine learning partial differential equations (PDE) driven deep learning framework for the classification of hyperspectral images (HSIs). The work is inspired by the famous PDE in image processing, namely, the Perona-Malik (PM) equation, which can form a scale-space and is capable of edge-preserving denoising using anisotropic diffusion (PM diffusion). In this framework, we firstly propose auto-machine learning-based trainable PM diffusion blocks (TPM-blocks) and then cascade them into a deep convolutional neural networks (CNN). Specifically, the 1 ×1 convolution layer and the trainable PM diffusion unit (TPMDU) are integrated as the TPM-block, and then multiple TPM-blocks are stacked to form a novel end-to-end deep learning architecture. We show that our deep learning method has the capacity of learning discriminative spectral and spatial features of HSIs. Experimental results on several popular datasets demonstrate that the proposed method achieves state-of-the-art performance compared with the several existing deep learning-based methods. Ning Wen, Qichao Liu, Liang Xiao 0001 |
IGARSS | 3 |
| 2020 | A Directional Message Propagation Convolutional Neural Network for Hyperspectral Images ClassificationabstractConvolutional neural networks (CNNs) have emerged as a powerful tool in remote sensing image analysis. However, the layer-by-layer convolutions (L2Convolutions) in CNNs cannot fully exploit the relativities of pixels in the 3D cube data, especially for hyperspectral images (HSIs). In this paper, a directional message propagation convolutional neural network framework (MPCNN), is proposed for the supervised classification of HSIs. In the proposed framework, we integrate a novel multi-directional message propagation mechanism, namely slice-by-slice convolutions (S2Convolutions), into the hidden feature layers which are generated by L2Convolutions to propagate feature information between feature maps in the same layer. Owing to the S2Convolutions, abundant and discriminative spectral-spatial feature learning can be enhanced compared with traditional CNNs without S2Convolutions. The performance of the proposed MPCNN is evaluated on benchmark dataset of HSIs, and quantitative and qualitative experiments show that the performance of the proposed method outperforms several state-of-the-art methods. Qichao Liu, Liang Xiao 0001, Zhihui Wei |
IGARSS | 3 |
| 2020 | Bipartite Residual Network for Change Detection in Heterogeneous Optical and Radar ImagesabstractThis paper presents a novel bipartite unsupervised residual networks (ResNet) for change detection based on two heterogeneous images acquired by optical and radars sensors on different dates. Most previous change detection methods in the unsupervised field use the shallow network and detect image changes at the pixel level. The proposed method detects image changes at the feature level via a bipartite deep residual networks. The network compares the difference information of the two images by fully considering the joint information of multi-temporal images. The generated image features can well represent the different information between the two images. ResNet is used to learn features from images. A new loss function is proposed for training the ResNet. Experimental results of homogeneous and heterogeneous images show that our method has promising performance compared to existing change detection methods for optical and radar images. Haocheng Zhang, Jia Liu 0020, Liang Xiao 0001 |
IGARSS | 3 |
| 2020 | Generalized Tensor Regression for Hyperspectral Image ClassificationabstractIn this article, we propose a novel tensorial approach, namely, generalized tensor regression, for hyperspectral image classification. First, a simple and effective classifier, i.e., the ridge regression for multivariate labels, is extended to its tensorial version by taking advantages of tensorial representation. Then, the discrimination information of different modes is exploited to further strengthen the capacity of the model. Moreover, the model can be simplified and solved easily. Different from traditional tensorial methods, the proposed model can be utilized to capture not only the intrinsic structure of data in a physical sense but also the generalized relationship of data in a logical sense. Our proposed approach is shown to be effective for different classification purposes on a series of instantiations. Specifically, our experiment results with hyperspectral images collected by the airborne visible/infrared imaging spectrometer, the reflective optics spectrographic imaging system and the ITRES CASI-1500 demonstrate the effectiveness of the proposed approach as compared to other tensor-based classifiers and multiple kernel learning methods. Zebin Wu 0001, Liang Xiao 0001, Jun Sun 0008, Hong Yan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Content-Guided Convolutional Neural Network for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) are of great interest and have demonstrated remarkable performance in hyperspectral images (HSIs) classification. However, due to the current configuration of the convolution layers with a fixed kernel shape, regular CNNs are inherently limited in modeling the diverse land-cover structures, particularly in the cross-classes edge regions, where irregular class boundaries would lead to high classification errors. To address this issue, we propose a content-guided CNN (CGCNN) for HSI classification. Compared with the shape-fixed kernel in the traditional CNN, the proposed content-guided convolution adaptively adjusts its kernel shape according to the spatial distribution of land covers. The content pattern is reflected by a latent guide map automatically learned from HSI. Such content-adaptive kernel with CGCNN could suppress the irregularity and unexpected features in class boundaries and, thus, improve the feature learning in cross-classes regions. Based on the content-guided convolution, a novel guided feature extraction unit (GFEU) is constructed for spectral-spatial feature learning of HSI. Finally, the CGCNN classification framework is established by stacking multiple GFEUs with dense connection, which is helpful for mitigating the gradient vanishing and increasing the robustness to overfitting. Extensive experiments on several HSIs demonstrate that the proposed approach possesses great details' preserving ability and its performance outperforms other state-of-the-art methods. Qichao Liu, Liang Xiao 0001, Jingxiang Yang, Jonathan Cheung-Wai Chan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | A Truncated Matrix Decomposition for Hyperspectral Image Super-ResolutionabstractHyperspectral image super-resolution addresses the problem of fusing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to produce a high-resolution hyperspectral image (HR-HSI). In this paper, we propose a novel fusion approach for hyperspectral image super-resolution by exploiting the specific properties of matrix decomposition, which consists of four main steps. First, an endmember extraction algorithm is used to extract an initial spectral matrix from LR-HSI. Then, with the initial spectral matrix, we estimate the spatial matrix, i.e., the spatial-contextual information, from the degraded observations of HR-HSI. Third, the spatial matrix is further utilized to estimate the spectral matrix from LR-HSI by solving a least squares (LS)-based problem. Finally, the target HR-HSI is constructed by combing the estimated spectral and spatial matrixes. In particular, two models are proposed to estimate the spatial matrix. One is a simple case that involves a LS-based problem, and the other is an elaborate case that consists of two fidelity terms and a spatial regularizer, where the spatial regularizer aiming to restrain the range of solutions is achieved by exploiting the superpixel-level low-rank characteristics of HR-HSI. Experiment results conducted on both synthetic and real data sets demonstrate the effectiveness of the proposed approach as compared to other hyperspectral image super-resolution methods. Zebin Wu 0001, Liang Xiao 0001, Jun Sun 0008, Hong Yan 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Learning Multiple Parameters for Kernel Collaborative Representation ClassificationabstractIn this article, the problem of automatically learning multiple parameters for kernel collaborative representation classification (KCRC) is considered. We investigate the KCRC and measure its generalization error via leave-one-out cross-validation (LOO-CV). By taking advantage of the specific properties of KCRC, a closed-form expression is derived for the outputs of LOO-CV. Then, a simple classification rule that provides probabilistic outputs is adopted, and thereby, an effective loss function that is an explicit function with respect to the parameters is proposed as the generalization error. The gradients of the loss function are calculated, and the parameters are learned by minimizing the loss function using a gradient-based optimization algorithm. Furthermore, the proposed approach makes it possible to solve the multiple kernel/feature learning problems of KCRC effectively. Experiment results on six data sets taken from different scenes demonstrate the effectiveness of the proposed approach. Zebin Wu 0001, Liang Xiao 0001, Hong Yan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Cross-Domain NER using Cross-Domain Language ModelingabstractDue to limitation of labeled resources, crossdomain named entity recognition (NER) has been a challenging task.Most existing work considers a supervised setting, making use of labeled data for both the source and target domains.A disadvantage of such methods is that they cannot train for domains without NER data.To address this issue, we consider using cross-domain LM as a bridge cross-domains for NER domain adaptation, performing crossdomain and cross-task knowledge transfer by designing a novel parameter generation network.Results show that our method can effectively extract domain differences from crossdomain LM contrast, allowing unsupervised domain adaptation while also giving state-ofthe-art results among supervised domain adaptation methods. Liang Xiao 0001, Yue Zhang 0004 |
ACL (1) | 2 |
| 2019 | A Hybrid Convolutional Neural Network with Anisotropic Diffusion for Hyperspectral Image Classification
Qichao Liu, Mohsen Molaei, Liang Xiao 0001 |
ICIG (3) | 4 |
| 2019 | Clustering Hyperspectral Images Via Sparse Dictionary Learning with Joint Sparsity and Shared WaveletsabstractSparse subspace clustering (SSC) algorithm has achieved an impressive performances in hyperspectral images clustering. However, the raw samples contained noises were used to construct the dictionary. Moreover, SSC represented each signal individually ignoring the relationship among hyperspectral pixels. To overcome these problems, we propose a sparse dictionary learning method for hyperspectral images clustering, in which joint sparsity and shared Wavelets are integrated to improve the expressive power of the learnt dictionary. First, we incorporate the shared Wavelets as a base dictionary into a unified joint sparsity constrained optimizing model to learn a structured sparse dictionary from both spectral and contextual characteristics of hyperspectral images. Then, the sparse representation coefficients based on the learnt sparse dictionary are adopted to construct a non-negative affinity matrix of graph. Finally, spectral clustering is employed to the affinity matrix to obtain the final clustering result. Experimental results clearly demonstrate that the proposed algorithm outperforms other state-of-the-art methods on the hyperspectral dataset. Nan Huang 0001, Liang Xiao 0001, Songze Tang, Qichao Liu |
IGARSS | 2 |
| 2019 | Data Augmentation and Refining with Steering Stencils for Supervised Classification of Hyperspectral ImageabstractLimited and expensive availability of labeled training samples resulted in the development of methods defining the hyperspectral classification task in the form of data augmentation based supervised learning. However, most of the methods just implicitly utilize the spectral-spatial information in the isotropic neighborhood, instead of explicitly indicating the anisotropic or steering neighborhood system. In this paper, we apply steering stencils for estimating the local directional homogenous regions and exploiting more valuable spectral-spatial contexts. By using a best steering stencil matching method, we propose a data augmentation and refining method to improve the performance of any spectral-spatial classifier with limited labeled samples. Experiments show that the proposed method is very effective for many spectral-spatial classifiers. Qichao Liu, Liang Xiao 0001, Pengfei Liu 0002, Nan Huang 0001 |
IGARSS | 2 |
| 2019 | Supervised Hyperspectral Image Classification Via Sparse Separable Convolutional Feature LearningabstractGenerally, the traditional supervised hyperspectral image (HSI) classification cannot fully exploit spatial and spectral features simultaneously. In this paper, we reformulate HSI feature learning in terms of sparse separable convolutional filter learning problem and propose a sparse separable convolutional classification model (SSCCM). In the proposed SSCCM, the sparse separable convolutional learning module (SSCLM) is used to extract robust spatial-spectral features and utilizes rank-one tensor decomposion learning mechanism to accelerate feature computation. While the SVM classification module (SVMCM) employs the 3D spatial-spectral feature array to represent the HSI for classification. Experimental results on the widely used HSI datasets demonstrate the superior performance of our proposed approach over the state-of-theart classification methods. Mengfei Song, Jie Song 0014, Liang Xiao 0001 |
IGARSS | 3 |
| 2019 | A Multi-Scale Densely Deep Learning Method for PansharpeningabstractPansharpening aims to produce a higher resolution multi-spectral (HRMS) image by fusing the spectral information in lower resolution multispectral (LRMS) image and the spatial information in corresponding high resolution panchromatic (PAN) image. In this work, we propose a multi-scale densely deep learning based pansharpening method. Following an end-to-end learning architecture, the proposed deep neural network contains three modules: 1) a parallel multi-scale convolutional layer is used to extract multiscale features of PAN image; 2) a global identity branch structure is adopted to preserve spectral structures; and 3) a dense learning block is integrated to improve the spectral-spatial expressive power. Compared with other state-of-the-art methods, experimental results obtained with our proposed method achieve high pansharpening quality in visualization and quantification. Zhikang Xiang, Liang Xiao 0001, Pengfei Liu 0002 |
IGARSS | 2 |
| 2019 | Hyperspectral Unmixing Via Simultaneous Dictionary Refining and Enhanced Sparse RegressionabstractThe dictionary-aided sparse regression (SR) approach has been developed in hyperspectral unmixing (HU) in remote sensing. By using an available spectral library as a dictionary, the SR approaches unmix the spectral image by selecting the endmember matrix from the library which best represent the image and its fractional abundance. In this paper, we proposed a simultaneous dictionary refining and enhanced sparse regression method for hyperspectral unmixing(DRESR). The proposed method not only enhances the sparsity of fractional abundance through double weighted sparse regularization, and also improves the spectral signature mismatches between an actual spectra and its corresponding endmember in the spectral library by using dictionary sparse refining. Experimental results on both synthetic and real hyperspectral data sets demonstrate better performance compared with several state-of-art algorithms. Yalei Gao, Zhizhong Zheng, Liang Xiao 0001 |
IGARSS | 4 |
| 2019 | Hyperspectral image clustering via sparse dictionary-based anchored regressionabstractClustering for hyperspectral images (HSIs) is a very challenging task because HSIs usually have large spectral variability, high dimensionality, and complex structures. The main issue of this study is to develop an improved sparse subspace clustering (SSC) method for HSIs. As an extension of spectral clustering, SSC algorithm has achieved great success; however, the direct self‐representation dictionary which is created by raw samples has poor representation power and also the widely used dictionary learning (DL) such as K‐Singular Value Decomposition (K‐SVD) faces with the problems of high computational complexity. In this study, the authors propose a novel HSI clustering method based on sparse DL and anchored regression. The proposed method follows three stages: (i) sparse DL; (ii) anchored subspace construction and regression; and (iii) representation‐based spectral clustering. Specifically, we adopt a fast sparse DL method under a double sparsity constrained optimising model to capture the intrinsic HSIs. To establish a compact subspace for collaborative representation, we present an anchored subspace construction method by using atoms clustering and grouping methods. Owing to the anchored subspace, we can fast compute the representation coefficients with a predefined projection matrix. Experimental results demonstrate that the proposed method achieves the best performance for the HSIs clustering. Nan Huang 0001, Liang Xiao 0001 |
IET Image Process. | 2 |
| 2019 | Multi-layer boosting sparse convolutional model for generalized nuclear segmentation from histopathology images
Jie Song 0014, Liang Xiao 0001, Mohsen Molaei, Zhichao Lian |
Knowl. Based Syst. | 2 |
| 2019 | Locally Low-Rank Regularized Video Stabilization With Motion Diversity ConstraintsabstractThis paper presents a novel motion aware regularization model with diversity constraints for motion smoothing in video stabilization. Differing from the global path optimization methods, the proposed model puts an emphasis on the relations of inter-frame motions and incorporates sliding windowed low-rank and smoothness constraints. The rationale behind it is cinematography rules which assume that camera motions can be divided into diverse patterns: zero velocity, constant velocity, and acceleration motion. Firstly, a locally motion aware fidelity term is adopted in light of the local motion stationarity. Secondly, to improve the robustness of the model for different motion patterns, a locally low-rank constrained regularization term is further introduced by considering the motion correlation in a local temporal window. Moreover, to cope with the over-smoothing problem in rapid motion situations with extreme acceleration, a motion steering kernel and varying window length are employed to enhance the flexibility of the proposed model. The experimental results demonstrate the superiority of the proposed optimization model and the efficiency to suppress over-smoothing when rapid motions occur. Meanwhile, we compare the results with some state-of-the-art methods quantitatively and qualitatively, and our method can achieve a comparable or even better stabilization effect. Huicong Wu, Liang Xiao 0001, Zhichao Lian, Hiuk Jae Shim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Hyperspectral Denoising Via Cross Total Variation-Regularized Unidirectional Nonlocal Low-Rank Tensor ApproximationabstractIn this paper, we propose a novel cross total variation regularized unidirectional nonlocal low rank tensor approximation method for hyperspectral image denoising. It fully explores the spectral-spatial correlation and non-local self-similarity simultaneously in tensor case and points out that the nonlocal self-similarity is the most important for precisely restoring the HSI. Following the research line in [1], we propose to embed the cross total variation (CrTV) regularization into the unidirectional low rank tensor framework to alleviate the common consistency issue of pixels in overlapped regions. CrTV shows great power to explore the spatial-spectral correlation and has great ability to keep the fine spatial details and preserve the spectra in the course of HSI denoising. The final model can be effectively solved by the alternating direction methods of multipliers (ADMM). Experimental results on HSI data sets validate that the complementary priors (i.e., spatial-spectral correlation and non local self-similarity) really contribute to the performance and also illustrate the superiority of the proposed method when compared with other state-of-the-art denoising methods. Le Sun 0002, Byeungwoo Jeon, Zebin Wu 0001, Liang Xiao 0001 |
ICIP | 4 |
| 2018 | Discriminative Pixel-Pairwise Constraint-Guided Extreme Learning Machine for Semi-Supervised Hyperspectral Image ClassificationabstractGenerally, the traditional semi-supervised extreme learning machine (S2-ELM) method cannot fully exploit the limited label information in hyperspectral image (HSI) classification. In this paper, we propose a discriminative S2-ELM method, called pixel-pairwise constrained S2-ELM (P2S2- ELM) method. Both the manifold regularization to leverage unlabeled data and the pixel-pairwise constraint between the labeled pixels are incorporated into a unified minimizing framework, thus the proposed P2S2-ELM method is able to learn a more effective and discriminative projection. Experimental results on several real hyperspectral data sets exhibit its efficiency and superiority to the counterparts, when only a small number of labeled samples are available. Jinhuan Xu, Pengfei Liu 0002, Le Sun 0002, Liang Xiao 0001 |
ICIP | 4 |
| 2018 | Simultaneous Dictionary Sparse Pruning and Collaborative Sparse Regression for Hyperspectral Image UnmixingabstractRecently, the dictionary-aided sparse regression (SR) method for hyperspectral unmixing has received much attention in the field of remote sensing. However, under the assumption that each pixel in the hyperspectral scene can be viewed as a combination of endmembers in the spectral library, most of SR methods ignore the spectral signature mismatches between an actual spectral signature and its corresponding endmember in spectral library. To overcome this problem, we proposed a joint optimizing unmixing model called DSPCSR which includes dictionary sparse pruning and collaborative sparse regression. By exploiting the sparse property of spectral mismatch error and the collaborative sparse property of the abundance matrix, the DSPCSR can provide good robustness and performance. Experiments on the synthetic and real datasets show that the proposed DSPCSR can achieve better performance compared with several state-of-art algorithms. Shengfu Li, Liang Xiao 0001, Zhihui Wei, Ling Qian |
IGARSS | 2 |
| 2018 | Pan-Sharpening with Hessian Nuclear Norm Induced Spatial ConsistencyabstractIn this paper, we propose a new variational pan-sharpening method with Hessian nuclear norm induced spatial consistency, which aims to fuse a low resolution (LR) multispectral (MS) image and a high resolution (HR) panchromatic (Pan) image into an HR MS image. In addition to using the fidelity based local spectral consistency term for preserving the spectral information, we particularly exploit the Hessian feature consistence between the HR MS image and Pan image, and propose a new Hessian nuclear norm induced spatial consistency term for preserving the spatial information. Then, the proposed model is solved by an efficient algorithm under the FISTA framework. Finally, the experimental results demonstrate that the proposed method outperforms various pan-sharpening methods in terms of higher spectral and spatial qualities. Pengfei Liu 0002, Liang Xiao 0001, Songze Tang |
IGARSS | 2 |
| 2018 | Multivariate Regression-Based Pan-Sharpening With Low Rank RegularizationabstractPan-sharpening is a fusion task of exploiting the spectral information in low resolution multispectral images (LRMS) with spatial information in a corresponding high resolution panchromatic image (PAN). Under the component substitution framework, a multivariate regression based fidelity term is presented to enforce high resolution spatial detail injection, while a low rank regularization term is proposed to capture intrinsic structures of latent high-resolution multispectral images (HRMS). To this end, a joint optimizing pan-sharpening model is proposed to establish a trade-off mechanism between the spatial detail injection and spectral-spatial preserving capacity. Finally, the Augmented Lagrangian Multiplier (ALM) method is used to develop a pan-sharpening algorithm. Experiments demonstrate that the proposed algorithm can achieve higher spatial and spectral resolution than several state-of-the-art methods. Heng Li 0012, Liang Xiao 0001 |
IGARSS | 3 |
| 2018 | Normal curvature-induced variational model for image restorationabstractIn this study, a novel normal curvature‐induced variational model which involves a higher‐order regulariser based on the normal curvature prior information of image surface is proposed for image restoration. Furthermore, the authors derive a preferably equivalent formulation for the proposed normal curvature‐induced higher‐order regulariser. Then, they design an efficient algorithm to solve the proposed model by using the famous alternating direction method of multipliers technique. Finally, they assess the performance of the proposed method on both natural images and biomedical cell images by comparing it with the famous fast total variation (TV) method, fractional‐order TV method and Hessian‐nuclear‐norm regularisation method. Specifically, the proposed method can achieve better and more balanced results in terms of peak‐signal‐to‐noise ratio, convergence rate and restoration quality. Pengfei Liu 0002, Liang Xiao 0001, Tao Li 0001 |
IET Image Process. | 2 |
| 2018 | Spatial rich model steganalysis feature normalization on random feature-subsets
Zhihui Wei, Liang Xiao 0001 |
Soft Comput. | 3 |
| 2018 | A Variational Pan-Sharpening Method Based on Spatial Fractional-Order Geometry and Spectral-Spatial Low-Rank PriorsabstractPan-sharpening refers to the fusion of a low-resolution (LR) multispectral (MS) image and a high-resolution (HR) panchromatic (PAN) image to obtain an HR MS image (i.e., pan-sharpened MS image). From the point of view of variational complementary data fusion, it becomes an optimization problem with geometry and spectral preserving constraints. In this paper, a novel unified optimizing pan-sharpening model is proposed by integrating a data-generative fidelity term and a compound prior term, which incorporates both spatial fractional-order geometry and spectral-spatial low-rank priors. Specifically, the proposed model consists of three important ingredients: 1) data-generative fidelity term, which models the degradation relationship between the LR and HR MS images to enforce the geometry and spectral preserving constraints; 2) fractional-order total variation-based spatial fractional-order geometry prior term, which especially exploits the spatial fractional-order gradient feature consistence between the PAN and pan-sharpened MS images to transfer the spatial structure information of the PAN image into the pan-sharpened MS image; and 3) weighted nuclear norm-based spectral-spatial low-rank prior term, which exploits the nonlocal patches-based low-rank structural sparsity simultaneously in the pan-sharpened MS image and the LR MS image for further preserving image spatial structures and spectral information. Thus, the main novelty behind the proposed model is an optimizing mechanism by fully taking advantage of the spatial details and texture expressive power of the spatial fractional-order geometry prior as well as the spectral-spatial correlation preserving capacity of the low-rank prior. Finally, the proposed model can be implemented in an alternating direction method of multipliers framework, and thus, an efficient algorithm is presented. To verify the validity, the new proposed method is systematically compared with some state-of-the-art techniques using the Pleiades, GeoEye-1, QuickBird, and WorldView2 satellite data sets in the subjective, objective, and efficiency aspects. The results show that the proposed method performs better than the compared methods in terms of higher spatial and spectral qualities. Pengfei Liu 0002, Liang Xiao 0001, Tao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Contour-Seed Pairs Learning-Based Framework for Simultaneously Detecting and Segmenting Various Overlapping Cells/Nuclei in Microscopy ImagesabstractIn this paper, we propose a novel contour-seed pairs learning-based framework for robust and automated cell/nucleus segmentation. Automated granular object segmentation in microscopy images has significant clinical importance for pathology grading of the cell carcinoma and gene expression. The focus of the past literature is dominated by either segmenting a certain type of cells/nuclei or simply splitting the clustered objects without contours inference of them. Our method addresses these issues by formulating the detection and segmentation tasks in terms of a unified regression problem, where a cascade sparse regression chain model is trained and then applied to return object locations and entire boundaries of clustered objects. In particular, we first learn a set of online convolutional features in each layer. Then, in the proposed cascade sparse regression chain, with the input from the learned features, we iteratively update the locations and clustered object boundaries until convergence. In this way, the boundary evidences of each individual object can be easily delineated and be further fed to a complete contour inference procedure optimized by the minimum description length principle. For any probe image, our method enables to analyze free-lying and overlapping cells with complex shapes. Experimental results show that the proposed method is very generic and performs well on contour inferences of various cell/nucleus types. Compared with the current segmentation techniques, our approach achieves state-of-the-art performances on four challenging datasets, i.e., the kidney renal cell carcinoma histopathology dataset, Drosophila Kc167 cellular dataset, differential interference contrast red blood cell dataset, and cervical cytology dataset. Jie Song 0014, Liang Xiao 0001, Zhichao Lian |
IEEE Trans. Image Process. | 2 |
| 2017 | Pansharpening via Locality-Constrained Sparse Representation
Songze Tang, Liang Xiao 0001 |
BMVC | 3 |
| 2017 | Color demosaicking via nonlocal tensor representationabstractA single sensor camera can capture scenes by means of color filter array. Each pixel samples only one of the three primary colors. Color demosaicking (CDM) is a process of reconstruction a full color image from this sensor data. In this paper, we propose a novel CDM scheme based on learned simultaneous sparse coding over nonlocal tensor representation. First, similar 2D patches are grouped to form a three-order tensor, that is, 3D array. Then, three sub-dictionaries, which characterize the coherent structures that appear in each dimension of the grouped tensor, are learned jointly by using Tucker decomposition. The consequent coefficient tensor is imposed by the grouped-block-sparsity constraint, which forces the similar patches to share the same atoms of the dictionaries in their sparse decomposition. Experimental results demonstrate the effectiveness both in the average CPSNR and visual quality. Wenze Shao, Hongyi Liu 0001, Zhihui Wei, Liang Xiao 0001 |
ICASSP | 6 |
| 2017 | Supervised classification of hyperspectral images using local-receptive-fields-based kernel extreme learning machineabstractIn this paper, we propose a local receptive fields (LRF) based kernel extreme learning machine (KELM) method for hyperspectral image (HSI) classification. As a single-hidden-layer feed-forward neural networks, kernel ELM has been used with success for classification of HSI. Considering local correlations in the spatial domain of HSI, the local random convolution nodes are introduced as LRF in the input layer. By using the LRF-based convolutional ELM feature learning on the principle component images, the proposed method can automatically learn rich feature hierarchies. All the learnt rich feature hierarchies are compressed to compact features, and the kernel ELM is applied to obtain the final classification results. Experimental results on the widely used real HSI data set indicate that the proposed approach outperforms several well-known classification methods. Jianyu Chen 0003, Liang Xiao 0001 |
ICIP | 3 |
| 2017 | PAN-Sharpening via residual deep learningabstractOne significant advantage of the deep convolutional neural networks (DCNN) is their representational ability for local complex structures. Inspired by this observation, a DCNN based residual learning model is proposed to learn a nonlinear mapping function between the high-resolution (HR) and low-resolution (LR) image patches. The DCNN is trained based on image patches, which are only sampled from the HR/LR panchromatic (PAN) image without other training images. We train the DCNN to obtain a nonlinear mapping function with HR/LR PAN patch pairs using mini-batch gradient descent based on back-propagation. By assuming HR/LR multispectral (MS) image shares the same mapping function between HR/LR PAN image patches from the viewpoint of transfer learning methodology, the HR MS image can be reconstructed from the observed LR MS image using the trained DCNN. Owing to the advantage of the residual learning mechanism, the proposed method can achieve a good geometrical details injection while preserves the spectral features. Experimental results show that the proposed method provides a better performance in both visual perception and numerical measures compared with the conventional methods. Nie Li, Nan Huang 0001, Liang Xiao 0001 |
IGARSS | 3 |
| 2017 | Supervised classification of hyperspectral images via heterogeneous deep neural networksabstractIn this paper, a new heterogeneous neural networks based deep learning method, named HNNDL, is presented for supervised classification of hyperspectral image (HSI) with a small number of labeled samples. Specifically, a deep neural Network (DNN) and a convolutional neural network (CNN) are combined to build a HNNDL architecture. The proposed architecture contains three modules: 1) dimension reduction and feature extraction, 2) training pixel-wise DNN and CNN, 3) bilateral filtering based decision level fusion on two soft probability maps which is produced by above classifiers. The rationale behind this heterogeneous deep learning architecture is their ability to learn more abstract and robust local spectral-spatial information by taking full advantages of complementary ability of each networks, and thus boost the performance of HSI classifier. Experimental results on the widely used HSI indicate that the proposed approach outperforms several well-known classification methods in terms of classification accuracy. Nan Huang 0001, Liang Xiao 0001 |
IGARSS | 4 |
| 2017 | Spectral-spatial subspace clustering for hyperspectral images VIA modulated low-rank representationabstractIn this paper, a novel spectral-spatial low-rank subspace clustering (SS-LRSC) algorithm is presented for clustering of hyperspectral images (HSI). Generally, employing the traditional LRSC framework directly cannot fully exploit the sample correlations in original spatial domain. Therefore, the proposed method utilizes a novel modulation strategy to modify the low rank representation matrix, which largely exploits the structure correlations. Specifically, a spectral and representation similarity weighted matrix is first applied to modulate the representation matrix; another local spatial bilateral filtering based modulation is further incorporated. Finally, the modulation method is integrated into the LRSC framework. Benefitting from the ability of the modulation method, SS-LRSC can capture both the structure correlations and inherent feature information of the data, which provides a competitive subspace clustering for HSI. Several experiments were conducted to illustrate the performance of the proposed algorithm. Jinhuan Xu, Nan Huang 0001, Liang Xiao 0001 |
IGARSS | 3 |
| 2017 | Video stabilisation with total warping variation modelabstractThis study proposes a robust approach to stabilise videos with a new variational minimising model. In video stabilisation, accumulation error often occurs in cascaded transformation chain‐based methods. To alleviate accumulation error, a new total warping variation (TWV) model is proposed, which describes the smoothness of stabilised camera motion and calculates all the warping transformations efficiently. After estimating original motion parameters based on a 2D similarity transformation model, the corresponding warping parameters are calculated under the TWV minimising framework, where the separable property of the motion parameters is utilised to obtain a closed‐form solution. The proposed method provides robust, smooth and precise motion trajectories after stabilisation. Furthermore, an iterative TWV method is introduced to reduce high‐frequency jitters as well as low‐frequency motions. Moreover, an online TWV method is presented for a long video sequence streaming by adopting a sliding windowed approach. Experimental results on various shaky video sequences show the effectiveness of the proposed method. Huicong Wu, Liang Xiao 0001, Hiuk Jae Shim, Songze Tang |
IET Image Process. | 2 |
| 2017 | Fast projections of spatial rich model feature for digital image steganalysis
Zhihui Wei, Liang Xiao 0001 |
Soft Comput. | 3 |
| 2017 | Structure-Based Low-Rank Model With Graph Nuclear Norm Regularization for Noise RemovalabstractNonlocal image representation methods, including group-based sparse coding and block-matching 3-D filtering, have shown their great performance in application to low-level tasks. The nonlocal prior is extracted from each group consisting of patches with similar intensities. Grouping patches based on intensity similarity, however, gives rise to disturbance and inaccuracy in estimation of the true images. To address this problem, we propose a structure-based low-rank model with graph nuclear norm regularization. We exploit the local manifold structure inside a patch and group the patches by the distance metric of manifold structure. With the manifold structure information, a graph nuclear norm regularization is established and incorporated into a low-rank approximation model. We then prove that the graph-based regularization is equivalent to a weighted nuclear norm and the proposed model can be solved by a weighted singular-value thresholding algorithm. Extensive experiments on additive white Gaussian noise removal and mixed noise removal demonstrate that the proposed method achieves a better performance than several state-of-the-art algorithms. Qi Ge, Xiaoyuan Jing, Fei Wu 0004, Zhihui Wei, Liang Xiao 0001, Wenze Shao, Dong Yue 0001, Haibo Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2017 | Boundary-to-Marker Evidence-Controlled Segmentation and MDL-Based Contour Inference for Overlapping NucleiabstractThis paper presents a novel method for automated morphology delineation and analysis of cell nuclei in histopathology images. Combining the initial segmentation information and concavity measurement, the proposed method first segments clusters of nuclei into individual pieces, avoiding segmentation errors introduced by the scale-constrained Laplacian-of-Gaussian filtering. After that a nuclear boundary-to-marker evidence computing is introduced to delineate individual objects after the refined segmentation process. The obtained evidence set is then modeled by the periodic B-splines with the minimum description length principle, which achieves a practical compromise between the complexity of the nuclear structure and its coverage of the fluorescence signal to avoid the underfitting and overfitting results. The algorithm is computationally efficient and has been tested on the synthetic database as well as 45 real histopathology images. By comparing the proposed method with several state-of-the-art methods, experimental results show the superior recognition performance of our method and indicate the potential applications of analyzing the intrinsic features of nuclei morphology. Jie Song 0014, Liang Xiao 0001, Zhichao Lian |
IEEE J. Biomed. Health Informatics | 2 |
| 2016 | ELM-based classification of ADHD patients using a novel local feature extraction methodabstractRecently, it has been an increasing interest in modeling abnormal temporal dynamics of functional interactions in psychiatric disorders. However, the accuracy of differentiating attention-deficit/hyperactivity disorder (ADHD) children form normal children has still much space for improvement. To further improve the accuracy, the key issue is to extract more effective features from original fMRI data. In this paper, we propose a novel local feature extraction method named Local Binary Encoding Method (LBEM) that can effectively characterize functional interaction patterns (FIPs). In particular, we show that the proposed method can well discriminate the functional interaction abnormalities, which is composed of a Bayesian connectivity change point model, a local feature extraction method and a kernel Extreme Learning Machine (ELM)-based classifier. The experiment on a real dataset of 23 ADHD children and 45 normal control (NC) children has shown that our method achieved better classification performance compared to the existing methods. Zhichao Lian, Min Li 0009, Zhonggeng Liu, Liang Xiao 0001, Zhihui Wei |
BIBM | 5 |
| 2016 | Hyperspectral image supervised classification via multi-view nuclear norm based 2D PCA feature extraction and kernel ELMabstractIn this paper, we propose a novel flexible framework for hyperspectral image (HSI) classification using multi-view spectral-spatial feature extracted by nuclear norm based 2D PCA. We first use the multihyphonthesis (MH) prediction method based on ridge regression to generate the 3D spatial-feature array from the HSI. Then, we apply the nuclear norm based 2D PCA to multi-view slices (the image with the spatial width and spectral dimension or with the spatial height and spectral dimension) of the former feature array, which can provide a structured spatial-spectral characterization for the reconstruction error slice and further extract the spatial-spectral feature. Finally, the 3D spatial-spectral feature array is used to represent the HSI for classification by extreme learning machine (ELM) based on Radial Basis Function (RBF) kernal. Finally, majority voting procedure is used to further improve the classification accuracy. The efficiency of the proposed method is demonstrated by experimental results with real hyperspectral dataset. Jue Jiang, Heng Li 0012, Liang Xiao 0001 |
IGARSS | 4 |
| 2016 | Fractional order variational pan-sharpeningabstractIn this paper, we propose a new fractional order variational method for pan-sharpening, which aims to obtain a high resolution multi-spectral (MS) image from a low resolution MS image and a high resolution panchromatic (PAN) image. On one hand, we use the data generative constraint for preserving the spectral information. More specifically, on the other hand, we exploit the fractional order gradient feature consistence between the high resolution MS image and PAN image for preserving the spatial information. Based on these assumptions, a new fractional order variational model is proposed and an efficient algorithm is designed to solve the proposed model. Experimental results show that the proposed method outperforms various well-known pan-sharpening methods in terms of higher spatial and spectral qualities. Pengfei Liu 0002, Liang Xiao 0001, Songze Tang, Le Sun 0002 |
IGARSS | 2 |
| 2016 | Hyperspectral image classification via region-based composite kernelsabstractThis paper presents a region-based composite kernel framework for spatial-spectral hyperspectral image classification, referred as RCK, by exploiting the local similarities of both the spectral and spatial features via superpixel segmentation. The proposed framework consists of three steps. In the first step, the original hyperspectral image together with its spatial feature image are segmented into several nonoverlapping regions by using an efficient superpixel segmentation algorithm. In the second step, a mean filtering is performed within each region of both the spectral feature image and the spatial feature image to generate the corresponding region-based spectral and spatial features, respectively. In the final step, both the obtained region-based features are combined and incorporated into a probabilistic kernel collaboration classifier by taking advantage of a composite kernel framework. Experimental results on two real hyperspectral images demonstrate the improvement of RCK over the traditional composite kernel framework, as well as its effectiveness as compared to some popular spatial-spectral techniques. Xiaoqian Shi, Zebin Wu 0001, Liang Xiao 0001, Zhiyong Xiao 0001, Yun-Hao Yuan 0001 |
IGARSS | 4 |
| 2016 | ELM-based spectral-spatial classification of hyperspectral images using bilateral filtering information on spectral band-subsetsabstractAs single-layer feed-forward neural networks, extreme learning machine (ELM) has recently been used with success for the classification of hyperspectral images (HSIs). However, the results of pure pixel-wise spectral classifiers often appear very noisy with limited training samples. To further improve the accuracy, we propose a novel spectral-spatial information integrating scheme for pixel-wise kernel ELM-based classifier. In particular, we show that a spatial bilateral filtering information on spectral band-subsets can significantly improve the accuracy of the pixel-wise kernel ELM based classifier. The benefits of the proposed method are twofold: 1) spectral structural similarity guided band-subsets partition and 2) incorporating the spectral-spatial information by bilateral filtering. Experiments on the widely used real HSI demonstrate that the proposed approach outperforms several well-known classification methods in terms of classification accuracy and low computational cost. Jinhuan Xu, Heng Li 0012, Liang Xiao 0001 |
IGARSS | 4 |
| 2016 | Non-local Spectral-spatial Centralized Sparse Representation for hyperspectral image classificationabstractThis paper presents a unified Non-local Spectral-spatial Centralized Sparse Representation (NL-CSR) model for the hyper-spectral image classification. The proposed model integrates local sparsity and non-local mean centralized induced sparsity. To achieve rich spectral-spatial information, the centralized sparsity enforces the sparse coding vector towards its non-local structural self-similar mean which is obtained via Image Patch Distance (IPD). Experiments validated that our NL-CSR model achieves convincing improvement over conventional sparsity based methods. Bushra Naz, Liang Xiao 0001, Shahzad Hyder Soomro |
IGARSS | 2 |
| 2016 | Surrogating circuit design solutions with robustness metrics
Jin Sun 0006, Liang Xiao 0001, Jiangshan Tian, Janet Roveda |
Integr. | 2 |
| 2016 | Pure spatial rich model features for digital image steganalysis
Zhihui Wei, Liang Xiao 0001 |
Multim. Tools Appl. | 3 |
| 2016 | Bi-component decomposition based hybrid regularization method for partly-textured CS-MR image reconstruction
Jun Zhang 0024, Zhihui Wei, Liang Xiao 0001 |
Signal Process. | 3 |
| 2016 | Spatial-Hessian-Feature-Guided Variational Model for Pan-SharpeningabstractIn this paper, we propose a new spatial-Hessian-feature-guided variational model for pan-sharpening, which aims at obtaining a pan-sharpened multispectral (MS) image with both high spatial and spectral resolutions from a low-resolution MS image and a high-resolution panchromatic (PAN) image. First, we assume that the low-resolution MS image corresponds to the blurred and downsampled version of the high-resolution pan-sharpened MS image. Since the pan-sharpened MS image and the PAN image are two images of the same scene, the pan-sharpened MS image shares similar geometric correspondence with the PAN image. To this end, the geometric correspondence between the PAN image and the pan-sharpened MS image is learnt as spatial position consistency by interest point detection. Second, a new vectorial Hessian Frobenius norm term based on the image spatial Hessian feature is presented to constrain the special correspondence between the PAN image and the pan-sharpened MS image, as well as the intracorrelations among different bands of the pan-sharpened MS image. Based on these assumptions, a novel variational model is proposed for pan-sharpening. Accordingly, an efficient algorithm for the proposed model is designed under the operator splitting framework. Finally, the results on both simulated data and real data demonstrate the effectiveness of the proposed method in producing pan-sharpened results with high spectral quality and high spatial quality. Pengfei Liu 0002, Liang Xiao 0001, Jun Zhang 0024, Bushra Naz |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | Coupled learning based on singular-values-unique and hog for face hallucinationabstractThis paper proposed a novel method for face hallucination based on a neighbor embedding technique. Traditional neighbor embedding approaches often offer counterintuitive results because consistency between high resolution images and low resolution images cannot be preserved without taking the intrinsic features of the image patches into account. In order to reinforce the consistency, on the one hand, we exploit the singular-values-unique (SVU) features inspired by singular values decomposition (SVD) successfully applied in image processing. On the other hand, we introduced the Histograms of Oriented Gradients (HOG) features to characterize the local geometric structure of the image patches to alleviate the effects of noise. At last, the learning space is extended to a coupled feature space that combines the SVU and HOG features. Simulation experiments show that this proposed approach could provide competitive results in simulation experiments in subjective and objective quality. Songze Tang, Liang Xiao 0001, Pengfei Liu 0002, Huicong Wu |
ICASSP | 2 |
| 2015 | Pan-Sharpening via Coupled Unitary Dictionary Learning
Shumiao Chen, Liang Xiao 0001, Zhihui Wei, Wei Huang 0013 |
ICIG (3) | 2 |
| 2015 | Blind Motion Deblurring Based on Fused l0-l1 Regularization
Liang Xiao 0001, Zhihui Wei, Linxue Sheng |
ICIG (2) | 3 |
| 2015 | Automatic Segmentation of White Matter Lesions Using SVM and RSF Model in Multi-channel MRI
Renping Yu, Liang Xiao 0001, Zhihui Wei, Xuan Fei |
ICIG (1) | 2 |
| 2015 | A new variational method for pan-sharpeningabstractIn this paper, we present a new variational method for pan-sharpening, which aims to obtain a high resolution multi-spectral (MS) image from a low resolution MS image and a high resolution panchromatic (PAN) image. Firstly, we assume that the desired high resolution MS image after down-sampling should be close to the low resolution MS image. More specifically, the intensity maps of PAN image and high resolution MS image bands are treated as three-dimensional (3D) differential surfaces. Then, we constrain that the surfaces of PAN image and high resolution MS image band should have the same bending directions at each point in 3D space. Based on these assumptions, a variational model is proposed and an efficient algorithm is designed to solve this variational model. Experimental results demonstrate that the proposed method outperforms various pan-sharpening methods in terms of both excellent spatial and spectral qualities. Pengfei Liu 0002, Liang Xiao 0001, Songze Tang |
IGARSS | 2 |
| 2015 | Joint dictionary learning with ridge regression for pansharpeningabstractA novel pansharpening method is proposed for creating a fused image of high spatial and spectral resolutions through merging a panchromatic (PAN) image with a multispectral (MS) image. To replace the patch pairs sampled from the images directly as the dictionary pairs, a joint learning model is proposed to learn a pair of compact dictionaries. Meanwhile, instead of restricting the coding coefficients of low resolution (LR) MS and high resolution (HR) MS image patches to be equal, ridge regression model is employed to describe their relation. Then, the fused MS image is calculated by combining the mapped sparse coefficients and the dictionary for the HR MS image. By comparing with some well-known methods in terms of several universal quality evaluation indexes, the simulated experimental results demonstrate the superiority of our method. Songze Tang, Liang Xiao 0001, Bushra Naz, Pengfei Liu 0002 |
IGARSS | 2 |
| 2015 | Image dehazing using two-dimensional canonical correlation analysisabstractImage dehazing is an important issue that interests both image processing and computer vision. In this study, image dehazing is modelled as an example‐based learning problem, and a novel dehazing algorithm using two‐dimensional (2D) canonical correlation analysis (CCA) is proposed. By assuming that the hazy‐free image patches are smooth and the pixel intensities in the same patch are approximate to constant, the authors deduce an underlying linear correlation between the observed hazy image patches and corresponding transmission patches. By maximising the correlation between the patch‐pairs of hazy image and corresponding transmission map, 2D CCA is able to learn a subspace to reconstruct the reliable transmission. Thus, given a test hazy image, the transmission map is aggregated by the nearest neighbour patches in the subspace and then globally refined by a local mean adaptive guided filter. The final hazy‐free image is obtained by using the dichromatic atmospheric model. Experimental results demonstrate the efficiency of the proposed method in single image dehazing. Liqian Wang, Liang Xiao 0001, Zhihui Wei |
IET Comput. Vis. | 2 |
| 2015 | Automatic method for white matter lesion segmentation based on T1-fluid-attenuated inversion recovery imagesabstractThe authors propose a fast and effective solution for automatic segmentation of white matter lesions by using T1 and fluid‐attenuated inversion recovery (FLAIR) image modalities with no need for manual segmentation and atlas registration. Initially, a brain tissue segmentation method is used to segment the T1 image into cerebrospinal fluid (CSF), grey matter and white matter. Based on the obtained tissue segmentation results, the region of interest (ROI) of the FLAIR image is created by subtracting the CSF from the FLAIR image. Subsequently, the authors calculate the z ‐score of the intensities in the ROI and define a threshold to perform a preliminary identification of abnormalities from normal tissues. The abnormalities obtained at this stage are used as the prior knowledge for the modified level‐set technique. The proposed level set method here is applied based on local Gaussian distribution to precisely detect the boundaries of the white matter lesions in the ROI. The level set method based on local Gaussian distribution fitting energy is robust to the intensity inhomogeneity of MR data and therefore capable of precisely extracting the boundaries of white matter lesions. Experimental analysis and quantitative comparisons with the peak‐seeking and state‐of‐the‐art white matter lesion segmentation (WMLS) techniques demonstrate that the algorithm is a stable and effective approach which significantly outperforms other trusted solutions for white matter lesion segmentation. Tianming Zhan, Liang Xiao 0001, Zhihui Wei |
IET Comput. Vis. | 4 |
| 2015 | Local brightness adaptive image colour enhancement with Wasserstein distanceabstractColour image enhancement is an important preprocessing phase of many image analysis tasks such as image segmentation, pattern recognition and so on. This study presents a new local brightness adaptive variational model using Wasserstein distance for colour image enhancement. Under the perceptually inspired variational framework, the proposed energy functional consists of an improved contrast energy term and a Wasserstein dispersion energy term. To better adjust image dynamic range, the authors propose a local brightness adaptive contrast energy term using the average brightness of image local patch as the local brightness indicator. To restore image true colours, a Wasserstein distance‐based dispersion energy term is used to measure the statistical similarity between the original image and the enhanced image. The proposed energy functional is minimised by using a gradient descent algorithm. Two objective measures are used to quantitatively measure the enhancement quality. Experimental results demonstrate the efficiency of the proposed model for removing colour cast and haze, enhancing contrast, recovering details and equalising low key images. Liqian Wang, Liang Xiao 0001, Hongyi Liu 0001, Zhihui Wei |
IET Image Process. | 2 |
| 2015 | A New Pan-Sharpening Method With Deep Neural NetworksabstractA deep neural network (DNN)-based new pansharpening method for the remote sensing image fusion problem is proposed in this letter. Research on representation learning suggests that the DNN can effectively model complex relationships between variables via the composition of several levels of nonlinearity. Inspired by this observation, a modified sparse denoising autoencoder (MSDA) algorithm is proposed to train the relationship between high-resolution (HR) and low-resolution (LR) image patches, which can be represented by the DNN. The HR/LR image patches only sample from the HR/LR panchromatic (PAN) images at hand, respectively, without requiring other training images. By connecting a series of MSDAs, we obtain a stacked MSDA (S-MSDA), which can effectively pretrain the DNN. Moreover, in order to better train the DNN, the entire DNN is again trained by a back-propagation algorithm after pretraining. Finally, assuming that the relationship between HR/LR multispectral (MS) image patches is the same as that between HR/LR PAN image patches, the HR MS image will be reconstructed from the observed LR MS image using the trained DNN. Comparative experimental results with several quality assessment indexes show that the proposed method outperforms other pan-sharpening methods in terms of visual perception and numerical measures. Wei Huang 0013, Liang Xiao 0001, Zhihui Wei, Hongyi Liu 0001, Songze Tang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Cartoon-texture composite regularization based non-blind deblurring method for partly-textured blurred images with Poisson noise
Zhengrong Zhang, Jun Zhang 0024, Zhihui Wei, Liang Xiao 0001 |
Signal Process. | 4 |
| 2015 | Supervised Spectral-Spatial Hyperspectral Image Classification With Weighted Markov Random FieldsabstractThis paper presents a new approach for hyperspectral image classification exploiting spectral-spatial information. Under the maximum a posteriori framework, we propose a supervised classification model which includes a spectral data fidelity term and a spatially adaptive Markov random field (MRF) prior in the hidden field. The data fidelity term adopted in this paper is learned from the sparse multinomial logistic regression (SMLR) classifier, while the spatially adaptive MRF prior is modeled by a spatially adaptive total variation (SpATV) regularization to enforce a spatially smooth classifier. To further improve the classification accuracy, the true labels of training samples are fixed as an additional constraint in the proposed model. Thus, our model takes full advantage of exploiting the spatial and contextual information present in the hyperspectral image. An efficient hyperspectral image classification algorithm, named SMLR-SpATV, is then developed to solve the final proposed model using the alternating direction method of multipliers. Experimental results on real hyperspectral data sets demonstrate that the proposed approach outperforms many state-of-the-art methods in terms of the overall accuracy, average accuracy, and kappa (k) statistic. Le Sun 0002, Zebin Wu 0001, Liang Xiao 0001, Zhihui Wei |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2014 | Spatial-spectral compressive sensing for hyperspectral images super-resolution over learned dictionaryabstractThis paper proposes a new hyperspectral images superresolution (HSI-SR) method based on compressive sensing (CS) theory, spatial sparsity and spectral similarity prior. First, according to sparsity and incoherence of CS theory, we propose a new dictionary learning method, ensuring that the learned dictionary not only has less dimensionality to speed up the sparse decomposition, but also satisfies sparsity well. Then, we introduce the spatial sparsity and spectral similarity regularizations into HSI-SR model, which can recover the spatial information effectively and preserve the spectral information well. The experimental results show the proposed method outperforms other well-known methods in terms of both objective measurements and visual evaluation. Wei Huang 0013, Zebin Wu 0001, Hongyi Liu 0001, Liang Xiao 0001, Zhihui Wei |
IGARSS | 4 |
| 2014 | Adaptive tensor matrix based kernel regression for hyperspectral image denoisingabstractKernel regression has been shown to be a powerful image denoising technique. In this paper, a three-dimensional (3-D) kernel regression hyperspectral image (HSI) denoising mechanism is proposed. The main contributions of this paper can be summarized as follows: Three orientation vectors and the corresponding coefficients are presented, which are adaptive for each pixel based on the innovation of 2-D structure tensor. An adaptive-driven 3-D tensor matrix is proposed for kernel regression, in which the spatial geometric structure and spectrum continuity are both considered. The proposed adaptive kernel regression is applied to HSI denoising. Both stimulated and real data experiments indicate that the proposed method can work well in detail preservation and noise removal. Hongyi Liu 0001, Zhengrong Zhang, Liang Xiao 0001, Zhihui Wei |
IGARSS | 3 |
| 2014 | Hyperspectral Image Classification Using Kernel Sparse Representation and Semilocal Spatial Graph RegularizationabstractThis letter presents a postprocessing algorithm for a kernel sparse representation (KSR)-based hyperspectral image classifier, which is based on the integration of spatial and spectral information. A pixelwise KSR is first used to find the sparse coefficient vectors of the hyperspectral image. Then, a sparsity concentration index (SCI) rule-guided semilocal spatial graph regularization (SSG), called SSG+SCI, is proposed to determine refined sparse coefficient vectors that promote spatial continuity within each class. Finally, these refined coefficient vectors are used to obtain the final classification map. Compared with previous approaches based on similar spatial-spectral postprocessing strategies, SSG+SCI clearly outperforms their results in terms of accuracy and the number of training samples, as it is demonstrated with two real hyperspectral images. Zebin Wu 0001, Le Sun 0002, Zhihui Wei, Liang Xiao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2014 | A fast adaptive reweighted residual-feedback iterative algorithm for fractional-order total variation regularized multiplicative noise removal of partly-textured images
Jun Zhang 0024, Zhihui Wei, Liang Xiao 0001 |
Signal Process. | 3 |
| 2014 | Variational Bayesian Method for RetinexabstractIn this paper, we propose a variational Bayesian method for Retinex to simulate and interpret how the human visual system perceives color. To construct a hierarchical Bayesian model, we use the Gibbs distributions as prior distributions for the reflectance and the illumination, and the gamma distributions for the model parameters. By assuming that the reflection function is piecewise continuous and illumination function is spatially smooth, we define the energy functions in the Gibbs distributions as a total variation function and a smooth function for the reflectance and the illumination, respectively. We then apply the variational Bayes approximation to obtain the approximation of the posterior distribution of unknowns so that the unknown images and hyperparameters are estimated simultaneously. Experimental results demonstrate the efficiency of the proposed method for providing competitive performance without additional information about the unknown parameters, and when prior information is added the proposed method outperforms the non-Bayesian-based Retinex methods we compared. Liqian Wang, Liang Xiao 0001, Hongyi Liu 0001, Zhihui Wei |
IEEE Trans. Image Process. | 2 |
| 2014 | Image-Based Quantitative Analysis of Gold Immunochromatographic Strip via Cellular Neural Network ApproachabstractGold immunochromatographic strip assay provides a rapid, simple, single-copy and on-site way to detect the presence or absence of the target analyte. This paper aims to develop a method for accurately segmenting the test line and control line of the gold immunochromatographic strip (GICS) image for quantitatively determining the trace concentrations in the specimen, which can lead to more functional information than the traditional qualitative or semi-quantitative strip assay. The canny operator as well as the mathematical morphology method is used to detect and extract the GICS reading-window. Then, the test line and control line of the GICS reading-window are segmented by the cellular neural network (CNN) algorithm, where the template parameters of the CNN are designed by the switching particle swarm optimization (SPSO) algorithm for improving the performance of the CNN. It is shown that the SPSO-based CNN offers a robust method for accurately segmenting the test and control lines, and therefore serves as a novel image methodology for the interpretation of GICS. Furthermore, quantitative comparison is carried out among four algorithms in terms of the peak signal-to-noise ratio. It is concluded that the proposed CNN algorithm gives higher accuracy and the CNN is capable of parallelism and analog very-large-scale integration implementation within a remarkably efficient time. Nianyin Zeng, Zidong Wang 0001, Bachar Zineddin, Min Du 0001, Liang Xiao 0001, Xiaohui Liu 0001, Terry Young |
IEEE Trans. Medical Imaging | 6 |
| 2013 | A hybrid active contour model with structure feature for image segmentationabstractWe propose a structured feature active contour model based on the level set method for image segmentation. We make the following two contributions. First, an adaptive data fitting term detects intensity variation based on the direction and global information. Second, integrated with a structured gradient vector flow (SGVF) method, we formulate a new regularization term with respect to the level set function via the duality formulation to penalize the length of active contour. We compare the proposed method to the classical active contour methods and demonstrate through the experiments on synthetic and medical images. Qi Ge, Liang Xiao 0001, Liqian Wang, Zhengrong Zhang, Zhihui Wei |
ICIP | 2 |
| 2013 | Compressive sensing ISAR imaging with stepped frequency continuous wave via Gini sparsityabstractIn this paper, we propose an improved version of CS-based model for inverse synthetic aperture radar (ISAR) imaging, which can sustain strong clutter noise and provide high quality images with extremely limited measurements. Different from traditional l1norm based CS ISAR imaging models, the essential of our model is to use the Gini index to measure the sparsity of signals. We also develop an iteratively re-weighted algorithm to find the solution of our model and reconstruct sparse signals from compressed samples. Experimental results of point targets and complex scene show that our approach significantly reduces the number of measurements needed for exact reconstruction and effectively suppresses the noise and outperforms l1norm based methods. Can Feng, Liang Xiao 0001, Zhihui Wei |
IGARSS | 2 |
| 2013 | A novel compound regularization and fast algorithm for compressive sensing deconvolution
Liang Xiao 0001, Zhihui Wei |
Neurocomputing | 1 |
| 2013 | New image restoration method associated with tetrolets shrinkage and weighted anisotropic total variation
Liqian Wang, Liang Xiao 0001, Jun Zhang 0024, Zhihui Wei |
Signal Process. | 2 |
| 2013 | Iterative Directional Total Variation Refinement for Compressive Sensing Image ReconstructionabstractWe propose a novel compressive sensing (CS) image reconstruction method based on iterative directional total variation (TV) refinement. As is generally known, classical TV-based CS reconstruction methods tend to produce over-smoothed image edges and texture details, since they favor piece-wise constant solutions. Hence, directional TV is introduced to describe the sparsity of the image gradient in order to overcome this drawback. However, it is difficult to estimate orientation field robustly and accurately from CS measurements. Inspired by vectorial ROF model, orientation field refinement model is presented and introduced into CS reconstruction. Extended experiments show that the proposed CS reconstruction method has a better improvement in the quality of the reconstructed image details over related TV-based CS reconstruction methods. Xuan Fei, Zhihui Wei, Liang Xiao 0001 |
IEEE Signal Process. Lett. | 3 |
| 2012 | A Relaxed Split Bregman Iteration for Total Variation Regularized Image Denoising
Jun Zhang 0024, Zhihui Wei, Liang Xiao 0001 |
ICIC (2) | 3 |
| 2012 | A novel sparsity constrained nonnegative matrix factorization for hyperspectral unmixingabstractSparsity is an intrinsic property of hyperspectral images, which means that the collected pixels can be represented by a part of materials. In this paper, a new sparsity based method for hyperspectral unmixing is proposed, referred to as the constrained sparse nonnegative matrix factorization (CSNMF). First, a novel sparse term which is explored to measure the sparsity of hyperspectral images is introduced to restrict the abundances. Second, minimum distance constraint which is convex is applied to restrict the endmembers. Then the alternating direction method of multipliers (ADMM) is used to solve the proposed CSNMF. The experimental results based on both synthetic mixtures and a real image scene demonstrate the effectiveness of the proposed approach. Zebin Wu 0001, Zhihui Wei, Liang Xiao 0001, Le Sun 0002 |
IGARSS | 4 |
| 2012 | An improved region-based model with local statistical features for image segmentation
Qi Ge, Liang Xiao 0001, Jun Zhang 0024, Zhihui Wei |
Pattern Recognit. | 2 |
| 2012 | A robust patch-statistical active contour model for image segmentation
Qi Ge, Liang Xiao 0001, Jun Zhang 0024, Zhihui Wei |
Pattern Recognit. Lett. | 2 |
| 2012 | Perceptual image quality assessment based on structural similarity and visual masking
Xuan Fei, Liang Xiao 0001, Yubao Sun, Zhihui Wei |
Signal Process. Image Commun. | 2 |
| 2011 | Perceptual Saliency Driven Total Variation for Image Denoising Using Tensor VotingabstractA nature image often contains various regions such as flat regions, ramps and edges with different singularities. A new perceptual saliency indicator is firstly proposed to distinguish edges and ramps. The proposed indicator is designed by a tensor voting approach with perceptual grouping performance. Using the perceptual saliency indicator, we propose a new variational model with an adaptive regularization term and a saliency weighted fidelity term. Experimental results demonstrate that our method has better performance in the staircase effect alleviation, the ramps and ridges preserving when compared with the state-of-the-art. Liang Xiao 0001, Fanbiao Zhang |
ICIG | 1 |
| 2011 | Compounded Regularization and Fast Algorithm for Compressive Sensing DeconvolutionabstractCompressive Sensing Deconvolution (CS Deconvolution) is a new challenge problem encountered in a wide variety of image processing fields. A compound variational regularization model which combined total variation and curve let-based sparsity prior is proposed to recovery blurred image from compressive measurements. We propose a novel fast algorithm using variable-splitting and Dual Douglas-Rachford operator splitting methods. Experiments demonstrate our proposed algorithm can obtain high-resolution data from highly incomplete measurements. Liang Xiao 0001, Zhihui Wei |
ICIG | 1 |
| 2011 | An improved region-based model with local statistical featureabstractIn this paper, a new region-based active contour model is proposed for image segmentation. Different from the general region-based active contour models, this model partitions the regions of interests in images depending on the local statistics of the intensity and the magnitude of gradient in the neighborhood of the contour. Inspired by the structure tensor method, an improved regularization term is defined through the duality formulation to penalize the length of region boundaries. Experiments on medical images demonstrate the proposed model outperforms the classical segmentation models in terms of efficiency and accuracy. Qi Ge, Zhihui Wei, Liang Xiao 0001, Jun Zhang 0024 |
ICIP | 3 |
| 2011 | Variational image restoration based on Poisson singular integral and curvelet-type decomposition space regularizationabstractImage restoration is a core topic of image processing. In this paper, we consider a variational restoration model consisting of Poisson singular integral (PSI) and curvelet-type decomposition space seminorm as regularizer. The PSI is used to impose a priori constraint on appropriate Lipschitz spaces, wherein a wide class of nonsmooth images can be accommodated. The seminorm of curvelet-type decomposition space is equivalent to the weighted curvelet coefficients which optimal represent smooth and edge parts of image with sparsity. We propose efficient algorithm to solve the optimization problem based on the Douglas-Rachford splitting (DRS) technique. Experimental results demonstrate that our proposed method can preserve important image features, such as edges and textures. Liang Xiao 0001, Zhihui Wei, Zhengrong Zhang |
ICIP | 2 |
| 2010 | Comments on "Staircase effect alleviation by coupling gradient fidelity term"
Liang Xiao 0001, Zhihui Wei |
Image Vis. Comput. | 1 |
| 2009 | A Nonlinear Inverse Scale Space Method for Multiplicative Noise Removal Based on Weberized Total VariationabstractMultiplicative noise removal has been drawn a greatly attention recently. Firstly, this paper proposes a new non-convex variational model for multiplicative noise removal under the Weberized TV regularization framework. Then we propose and study another surrogate strictly convex objective functional for Weberized TV regularization based multiplicative noise removal model. Finally, we adopt the recently proposed inverse scale space approach to estimate the underlying image under total variation (TV) regularization, in which a relaxation technique with two evolution equations is applied. Our experimental results show that the quality of images denoised is quite good and the detail information of restored images is well preserved by the proposed algorithm. Liang Xiao 0001, Zhihui Wei |
ICIG | 2 |
| 2009 | Compressed sensing image reconstruction based on morphological component analysisabstractCompressed sensing (CS) is a new area of signal processing for simultaneous signal sampling and compression. Most of existing methods for CS image reconstruction are suitable for piecewise smooth image, but do not behave well on texture-rich natural image. In this paper, a new optimization problem for CS image reconstruction is proposed, in which different regularization terms are introduced for different morphological components of image. Furthermore, an alternating iterative algorithm is presented to solve the relevant optimization problem. Experimental results show that the proposed method can be applied to reconstruct texture-rich images besides piecewise smooth ones, and outperforms the existing methods on preserving detail feature. Xingxiu Li, Zhihui Wei, Liang Xiao 0001, Yubao Sun, Jian Yang 0003 |
ICIP | 3 |
| 2005 | Collaborative design environment based on multi-agentabstractIt is necessary to provide collaborative design environment for preliminary design of complex products. For the specific requirements in preliminary design, this paper proposes the collaborative design environment for preliminary design of complex products(CDEPD). Agent-based architecture is adopted in CDEPD and it can support cooperative work for distributed designers from multiple disciplines. Based on GPGP from K. Decker, belief commitment-based coordination mechanism is proposed in this paper and it has been applied in CDEPD. Shenglei Chen, Huizhong Wu, Xianglan Han, Liang Xiao 0001 |
CSCWD (1) | 4 |
| 2004 | Inverse image warping without searchingabstractFor the depth information of desired view is unknown, a per-pixel searching step is often inevitable in methods of inverse image warping. A novel approach is proposed in this paper called "cross-segment algorithm (CSA)". Different from other existing methods, CSA tickles the corresponding problem by solving the equations of crossed segments instead of searching per-pixel. CSA pays more attention to the relationship between the different reference images including depth information than that between the reference and desired images. Because of eliminating the searching cost, CSA is proved to be an accelerated method of inverse image warping by experiments. Huizhong Wu, Fu Xiao 0001, Liang Xiao 0001 |
ICARCV | 4 |
| 2004 | Compute visibility without depth in multi-reference imagesabstractOne of the key problems in image-based rendering is to decide the visibility of objects in arbitrary viewplane. And reducing the need to compute depth for rendering has attracted a wide research interests. Most previous methods either work only on one reference image, which is unable to provide sufficient information for generating new views, or still rely partly on the depth information. In this paper, a novel approach is proposed to this problem based on multi-reference images, Where we deduce the warping function and a rendering order for image pixels onto the intermediary "aided-view", and in turn, to the target plane. By two-steps warping, it's proved a feasible method to compute visibility without depth in multi-reference images. Huizhong Wu, Fu Xiao 0001, Liang Xiao 0001 |
ICARCV | 4 |
| 2004 | Generalized Mumford-Shah model and fast algorithm for color image restoration and edge detectionabstractA generalized Mumford-Shah model for color image restoration and edge detection is established in this paper. First the color images are considered as the manifold based on the "surface based method", then a new physical quantity in the form of vector product, which describes the differences of the gradient between different channels, is introduced into the regularization term of the "object", then a generalized energy functional is proposed. Finally, we proposed a fast PDE's numerical iterative algorithm utilizing the steepest descent method and half-point scheme. Experimental results show the proposed model has good ability to overcoming the bad phenomena of color fluctuations along edges caused in the straightforward extension of the original Mumford-Shah model in vectorial cases. Liang Xiao 0001, Huizhong Wu, Zhihui Wei |
ICIG | 1 |