Rui Song 0003

dblp:01/2743-3 · DBLP profile ↗
← Back
65ranked-venue papers
3as first author
49since 2021 · last 2026
0000-0002-2790-1752ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 31 · 2 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 17 · 13 since 2021
YearPublicationVenuePosition
2026 Toward Memory-Efficient Hyperspectral Image Reconstruction via Consistency Learning
abstract
Spectral reconstruction (SR) aims to recover high-quality hyperspectral images (HSIs) from more readily available RGB or multispectral images (MSIs). While supervised SR has shown promising results, it is hindered by the difficulty of collecting abundant, well-registered RGB-HSI or MSI-HSI pairs. Semi-supervised SR (Semi-SR) offers a more practical solution by exploiting plentiful RGBs/MSIs together with limited HSIs. However, existing Semi-SR approaches still suffer from cross-domain discrepancies, cross-modality inconsistency, and unreliable pseudo-labels. To tackle these challenges, we propose a Manifold-aware Teacher-Student Semi-SR (MTSSR) framework, which seamlessly integrates labeled and unlabeled domains through a teacher-student paradigm and memory-efficient consistency learning. At its core, a Flexible Cross-attention Spectral Reconstruction (FCSR) network extracts scene-related spatial cues via customized self-attention and models scene-agnostic priors through dynamic quantization, thereby enhancing spectral fidelity. Furthermore, a manifold-aware dimensionality analysis derives a latent space that jointly captures spatial and spectral structures across modalities. This enables a manifold-aware alignment loss to enforce cross-modality consistency and a manifold-aware contrastive loss to progressively refine pseudo-label reliability. In addition, we develop a Threshold-adjusted Memory Bank Update (TMBU) strategy, which generates reliable negative samples by storing network-driven representations instead of memory-consuming HSIs, significantly reducing memory consumption. Extensive experiments on three visual and two remote sensing benchmarks demonstrate that MTSSR consistently outperforms state-of-the-art SR methods, achieving robust and memory-efficient spectral reconstruction.
Yihong Leng, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Image Process.3
2026 Two-Timer-KAN: Dual-Exclusive Fourier KANs With Gaussian Fusion for Few-Shot Multimodal Remote Sensing Imagery Classification
abstract
Multimodal remote sensing imagery classification (MRSIC) aims to synergistically leverage complementary information from heterogeneous data sources, enabling precise land-cover classification. Existing MRSIC approaches predominantly rely on abundant annotated samples, facing critical performance degradation under data-scarce scenarios that are particularly exacerbated by the inherent complexity of heterogeneous multimodal data. Furthermore, effectively extracting spatial-spectral information of multimodal data and fusing the cross-modal heterogeneous features persists as a significant challenge. To address these obstacles, we propose a pioneering few-shot MRSIC network, Two-timer-KAN, which integrates modality-specific feature extraction for spectral- and spatial-dominant data. Specifically, leveraging the nonlinear power of Kolmogorov-Arnold Networks (KANs), we develop the Dual-Exclusive Fourier KAN (DEF-KAN) encoder, which captures modality-specific global features in the frequency domain, bridging spectral and spatial gaps across various datasets. Following this, a Multivariate-Gaussian-based Cross-KAN (MG-Cross-KAN) is dedicated to enhancing the robustness of cross-modality fusion by capturing modality-shared features in a distribution-based manner. Finally, to further tackle classification ambiguity under limited annotated samples, we present a visual-textual bidirectional alignment strategy, which leverages textual descriptions as supplementary semantical knowledge to clarify class feature centers. Extensive experiments demonstrate that the proposed two-timer-KAN achieves superior performance, outperforming the state-of-the-art methods in both accuracy and robustness.
Jiaojiao Li 0001, Hailong Wu, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Image Process.4
2026 Leveraging Spatiotemporal Cues for Self-Supervised Stereo Depth Estimation in Endoscopic Videos
abstract
Self-supervised endoscopic depth estimation seeks to reconstruct dense depth information from clinical endoscopic video sequences without the necessity of ground-truth depth annotations, thereby offering significant potential for widespread application in surgical environments. Nevertheless, the majority of current approaches treat each video frame as an independent static image, neglecting the inherent temporal correlations and dynamic scene changes present in endoscopic videos. This limitation frequently results in depth estimation artifacts such as flickering, geometric inconsistencies, and structural distortions. In this study, we introduce a Siamese self-supervised framework that simultaneously leverages stereo and temporal information to enhance depth estimation in dynamic endoscopic videos. Central to our approach are two spatiotemporal feature aggregation modules. The Deformable Spatiotemporal Cross-View Fusion (DSCF) module explicitly constructs stereo and temporal cost volumes and performs integrated fusion at the cost-volume level. The Multi-Scale Selective Spatiotemporal Aggregation (MSSA) module captures long-term state memory through cross-resolution and cross-frame hidden-state propagation, while refining features by selectively integrating multi-scale inputs and multi-receptive-field spatiotemporal representations. To facilitate efficient deployment on endoscopic edge devices, we propose a three-stage training protocol that incrementally incorporates temporal supervision and ultimately distills the model into a monocular depth estimation network. Comprehensive evaluations conducted on four publicly available endoscopic datasets (SCARED, SERV-CT, EndoNeRF, and Hamlyn) demonstrate that our method attains state-of-the-art performance in both depth accuracy and image reconstruction quality, achieving an average reduction in root mean square error (RMSE) exceeding 9.1% relative to the strongest existing method.
Rongqi Wang, Rui Song 0003, Jingang Zhang, Yunfeng Nie
IEEE Trans. Medical Imaging2
2026 Physics-Guided Time-Interactive-Frequency Network for Cross-Domain Few-Shot Hyperspectral Image Classification
abstract
Recently, domain alignment and metric-based few-shot learning (FSL) have been introduced into hyperspectral image classification (HSIC) to solve the issues of uneven data distribution and scarcity of annotated data faced in practical applications. However, existing cross-domain few-shot methods ignore pivotal frequency priors of the complex field, which contribute to better category discrimination and knowledge transfer. To address this issue, we propose a novel physics-guided time-interactive-frequency network (PTFNet) for cross-domain few-shot HSIC, enabling the extraction of both frequency priors and spatial features (termed "time domain" following Fourier convention) simultaneously through a lightweight time-interactive-frequency module (TiF-Module) as a pioneering effort. Meanwhile, a spectral Fourier-based augmentation module (SFA-Module) is designed to decouple the frequency priors and enhance the diversity of distribution of physical attributes to imitate the domain shift. Then, the physics consistency loss is introduced to regularize the diverse embeddings to approximate the center of each category's physical attributes, guiding the network to excavate more transferable knowledge of source domain (SD). Furthermore, to fully exploit the discriminant time-frequency information and further improve the accuracy of boundary pixels, a set of multiorientation homogeneous prototypes is adopted to represent each class comprehensively, and an intuitive and flexible uncertainty-rectified bidirectional random walk strategy is applied to replace the Euclidean metric for more reliable classification. The experimental results on four public datasets demonstrate the prominent performance of the proposed PTFNet.
Jiaojiao Li 0001, Hailong Wu, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 SCFlow2: Plug-and-Play Object Pose Refiner with Shape-Constraint Scene Flow
abstract
We introduce SCFlow2, a plug-and-play refinement framework for 6D object pose estimation. Most recent 6D object pose methods rely on refinement to get accurate results. However, most existing refinement methods either suffer from noises in establishing correspondences, or rely on retraining for novel objects. SCFlow2 is based on the SCFlow model designed for refinement with shape constraint, but formulates the additional depth as a regularization in the iteration via 3D scene flow for RGBD frames. The key design of SCFlow2 is an introduction of geometry constraints into the training of recurrent matching network, by combining the rigid-motion embeddings in 3D scene flow and 3D shape prior of the target. We train SCFlow2 on a combination of dataset Objaverse, GSO and ShapeNet, and evaluate on BOP datasets with novel objects. After using our method as a post-processing, most state-of-the-art methods produce significantly better results, without any retraining or fine-tuning. The source code is available at https://scflow2.github.io.
Rui Song 0003, Jiaojiao Li 0001, Kerui Cheng, David Ferstl, Yinlin Hu
CVPR2
2025 Trafficloc: Localizing Traffic Surveillance Cameras in 3D Scenes
abstract
We tackle the problem of localizing traffic cameras within a 3D reference map and propose a novel image-to-point cloud registration (I2P) method, TrafficLoc, in a coarse-tofine matching fashion. To overcome the lack of large-scale real-world intersection datasets, we first introduce Carla Intersection, a new simulated dataset with 75 urban and rural intersections in Carla. We find that current I2P methods struggle with cross-modal matching under large viewpoint differences, especially at traffic intersections. TrafficLoc thus employs a novel Geometry-guided Attention Loss (GAL) to focus only on the corresponding geometric regions under different viewpoints during 2D-3D feature fusion. To address feature inconsistency in paired image patch-point groups, we further propose Inter-intra Contrastive Learning (ICL) to enhance separating 2D patch/3D group features within each intra-modality and introduce Dense Training Alignment (DTA) with soft-argmax for improving position regression. Extensive experiments show our TrafficLoc greatly improves the performance over the SOTA I2P methods (up to 86%) on Carla Intersection and generalizes well to real-world data. TrafficLoc also achieves new SOTA performance on KITTI and NuScenes datasets, demonstrating the superiority across both in-vehicle and traffic cameras. Our project page is publicly available at https://tum-luk.github.io/projects/trafficloc/.
Yan Xia 0003, Yunxiang Lu, Rui Song 0003, Oussema Dhaouadi, João F. Henriques, Daniel Cremers
ICCV3
2025 IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals
abstract
Semantic Scene Completion (SSC) has emerged as a pivotal approach for jointly learning scene geometry and semantics, enabling downstream applications such as navigation in mobile robotics. The recent generalization to Panoptic Scene Completion (PSC) advances the SSC domain by integrating instance-level information, thereby enhancing object-level sensitivity in scene understanding. While PSC was introduced using LiDAR modality, methods based on camera images remain largely unexplored. Moreover, recent Transformer-based approaches utilize a fixed set of learned queries to reconstruct objects within the scene volume. Although these queries are typically updated with image context during training, they remain static at test time, limiting their ability to dynamically adapt specifically to the observed scene. To overcome these limitations, we propose IPFormer, the first method that leverages context-adaptive instance proposals at train and test time to address vision-based 3D Panoptic Scene Completion. Specifically, IPFormer adaptively initializes these queries as panoptic instance proposals derived from image context and further refines them through attention-based encoding and decoding to reason about semantic instance-voxel relationships. Extensive experimental results show that our approach achieves state-of-the-art in-domain performance, exhibits superior zero-shot generalization on out-of-domain data, and achieves a runtime reduction exceeding 14$\times$. These results highlight our introduction of context-adaptive instance proposals as a pioneering effort in addressing vision-based 3D Panoptic Scene Completion. Code available at https://github.com/markus-42/ipformer.
Markus Gross 0003, Aya Fahmy, Danit Niwattananan, Dominik Muhle, Rui Song 0003, Daniel Cremers, Henri Meess
NeurIPS5
2025 Multi-Feature Interaction and Degradation Estimation Transformer for Spectral Compressive Imaging
abstract
Coded Aperture Snapshot Spectral Imaging (CASSI) systems provide an efficient approach to acquiring Hyperspectral Images (HSI), yet the reconstruction process still presents challenges. Traditional Deep Unfolding Networks (DUN) applied to CASSI often face constraints due to inadequate feature utilization and poor handling of multi-scale frequency-domain information, leading to the loss of image detail and global information. Furthermore, most DUN methodologies oversimplify degrading factors and fail to account for issues such as distortions found in actual imaging, thus affecting accuracy and robustness. This paper presents MIDET, a novel DUN tailored for CASSI systems, which integrates the fusion of band information, spatial information, and multi-scale information to meaningfully improve feature utilization and information interaction efficiency. Additionally, MIDET introduces a degradation-guided learning strategy and a frequency feature extraction module, enhancing the capability to handle real imaging distortions and preserve more details in HSI reconstruction. Experimental results demonstrate that MIDET significantly outperforms existing technologies on both simulated and real datasets, effectively enhancing the quality of HSI reconstruction.
Jiaojiao Li 0001, Ding Zhu, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Contrastive MLP Network Based on Adjacent Coordinates for Cross-Domain Zero-Shot Hyperspectral Image Classification
abstract
With the breakthrough of transfer learning and meta-learning, cross-domain few-shot hyperspectral image classification (CDFSL HSIC) technology has recently achieved satisfactory performance under limited annotations. Nevertheless, the most practical applications are zero-shot scenarios, which are intractable for CDFSL technology, such as the extraterrestrial detection scene, where unexplored objects are recognized by scientists to be more valuable for research. To conquer the zero-shot problem under domain shift, a two-stage contrastive MLP network (MAC-CDZS) is proposed, which constitutes a pioneering effort in the cross-domain zero-shot (CDZS) HSIC task. Firstly, given the remarkable performance of MLPs within a diminutive model size and their enhanced capacity for extracting spatial-spectral features of HSIs, the MLP framework has been strategically chosen as the foundational backbone of the first stage in the MAC-CDZS for facilitating efficient feature extraction. Secondly, to alleviate the potential category collapse, the second-stage fine-tuning framework is introduced, which extends the first-stage backbone by incorporating the elaborate adjacent coordinate module and contrastive learning paradigm for more harmonious classification performance. Specifically, the adjacent coordinate module is creatively designed to adequately mine the adjacent coordinates among samples for ameliorating category collapse from the perspective of grasping more reliable priors. Furthermore, a contrastive learning paradigm is innovatively constructed, comprising a Spatial Augmentation (SA) module tailored for hyperspectral patches and a construction strategy of sample pair under zero-shot conditions, which aims to boost the representation capability and alleviate the class collapse. The superior performance of the MAC-CDZS is demonstrated by experimental results on four benchmark datasets.
Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 NukesFormers: Unpaired Hyperspectral Image Generation With Nonuniform Domain Alignment
abstract
The persistent challenge of acquiring precisely coregistered RGB-hyperspectral image (HSI) pairs has significantly impeded the practical deployment of current data-driven Hyperspectral Image Generation (HIG) networks in engineering applications. Gleichzeitig, the ill-posed nature of the aligning constraints, compounded with the complexities of mining cross-domain features, also hinders the advancement of unpaired HIG (UnHIG) tasks. In this paper, we conquer these challenges by reformulating the UnHIG through Range-Null Space Decomposition (RND), modeling range-space feature interaction and null-space compensations. Specifically, the introduced contrastive learning effectively aligns the geometric and spectral distributions of unpaired data by building the interaction of range space, considering the consistent feature in degradation process. Our Dual-Dimensional Contrastive Prior Module (DCPM) captures mutual information within RGB and HSI domains of a single scene while modeling relationships between internal representations across different scenes, thereby constructing comprehensive cross-domain constraints. Furthermore, the Gabor kernel-based multi-head self-attention (G-MSA) adaptively separates high-frequency components, guiding subsequent modules to concentrate on relevant frequency intervals for target objects. Then, we propose a novel Non-uniform Kolmogorov-Arnold Networks (Nukes) to exhaustively excavate null-space components(degraded/high-frequency representations) through dual-domain frequency mapping. The proposed method was evaluated on three established datasets: NTIRE 2020 ’Clean’ track, NTIRE 2022, CAVE and Grss_dfc_2018. To assess real-world applicability, ratio experiments were conducted by changing the proportion of RGB and HSI. These experiments demonstrate that our approach achieves state-of-the-art performance in UnHIG.
Jiaojiao Li 0001, Shiyao Duan, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 UAT: Exploring Latent Uncertainty for Semi-Supervised Object Detection in Remote-Sensing Imagery
abstract
Object detectors for remote sensing imagery (RSI) are plagued by data-driven training approaches, underperforming when confronted with inadequate object information. Thus, semi-supervised object detection (SSOD) is introduced to alleviate high data dependence. However, it’s unfeasible to directly apply existing SSOD methods, considering the classification and regression uncertainties resulting from unique characteristics of RSI: 1) inter-class uncertainty: objects in remote sensing scenes are usually densely arranged, leading to severe feature fusion between different categories, 2) intra-class uncertainty: for same-category objects, the complicated contexts induced by various scenarios increase intra-class feature diversity, posing a challenge to clustering and 3) regression uncertainty: the inherent scale variation in RSI gives rise to inconsistent convergence rate, hindering the regression progress from effective converging. Therefore, we propose a Teacher-Student Model(TSM)-based SSOD method aiming at remote sensing scenes, termed Uncertainty-Aware Teacher (UAT), to quantify and promote the detector’s confidence, which is composed of Entropy-based Uncertain-pair Softening (EUS) policy, Multi-Gaussian Sample Fitting (MSF) module and Scale-adaptive Adjunct Loss (SAL). Specifically, EUS identifies the uncertain category pairs with class-wise information entropy, MSF remodels the intra-class distribution and sets the optimal threshold dynamically for pseudo-labels, and SAL adaptively re-scales bboxes based on the original size to accelerate the regression convergence. We have proven the efficiency of our method through extensive experiments on two public remote sensing datasets, DOTA and DIOR.
Jiaojiao Li 0001, Yuqing Ji, Kerui Cheng, Rui Song 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Background Suppression Network With Attention Collapse Inhibited Transformer for Optical Remote Sensing Object Detection
abstract
Object detection for remote sensing imagery (RSI) has been extensively exploited in practical applications. However, similar and multiscale objects in RSI, especially small objects, pose challenges to RSI object detection methods. Particularly, existing approaches ignore irrelevant background in RSI leading to hardship in discriminative feature extraction, resulting in instances of false positive (FP) and false negative (FN) of similar objects. In this article, we propose an irrelevant background suppression network (IBS-Net), which employs the structure of a convolutional neural network (CNN) in series with a Transformer to efficiently capture local and global information in images to respond the challenge of multiscale object detection. Primarily, a background detach module (BDM) is designed behind the backbone to suppress the irrelevant background and enhance the foreground to minimize the interference of irrelevant background for object detection. Furthermore, a composite-sampler (C-S) is devised to sample the vectors describing the foreground and the context, which expands the limited receptive field of the detector to better distinguish similar objects. Especially, considering that transformer-based object detection methods suffer from an attention collapse issue that leads to a degradation of the network representation. An attention collapse inhibited transformer (ACI-former) is presented by designing a partial residual connection, which induces the network to perceive more target information and reduces the loss of small target features thus improving the detection accuracy of small objects. Ultimately, we have conducted related experiments on two benchmarks, which demonstrate that our method has achieved prominent results compared with other mainstream detection methods.
Jiaojiao Li 0001, Haile Li, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 ASMF: A Self-Supervised Atmospheric Scatter Model-Based Fusion Network for Infrared Image Enhancement
abstract
Infrared camera sensors typically capture thermal radiation data, which can be influenced by random scattering and reflection during transmission. As a result, the infrared images produced often exhibit low contrast and clarity. This problem is especially evident in uncooled long-wave infrared (LWIR) images, which experience a decline in image quality, making it more difficult to enhance contrast. To tackle this issue, we propose a self-supervised network model that utilizes an atmospheric scattering model (ASM) specifically designed for the distinct characteristics of infrared imaging. Moreover, leveraging the analysis of ASM and the characteristics of infrared imaging, we introduce an image enhancement framework that integrates multi-scale dehazing with fusion techniques to facilitate self-supervised learning. Our method includes two self-supervised sub-networks, ASM-Net and F-Net, each equipped with tailored loss functions to address the challenges arising from the limited availability of high-quality datasets and the variability in imaging from different manufacturers. We improve self-supervised learning by integrating degradation models into the network architecture and incorporating prior knowledge into the loss functions. To evaluate our proposed method, we have developed a new LWIR dataset featuring real-world scenes. Extensive experiments conducted on established benchmark datasets, as well as our dataset, demonstrate that our approach significantly surpasses existing methods in both qualitative and quantitative assessments. The source code and dataset are available at https://github.com/ANDYLLY/Infrared-image-enhancement-method-ASMF-and-LWIR-Dataset-LGC.
Luyuan Liu, Rui Song 0003, Jiaojiao Li 0001, Wenman Wang, Qingrong Wen
IEEE Trans. Geosci. Remote. Sens.2
2025 HyperTaFOR: Task-Adaptive Few-Shot Open-Set Recognition With Spatial-Spectral Selective Transformer for Hyperspectral Imagery
abstract
Open-set recognition (OSR) aims to accurately classify known categories while effectively rejecting unknown negative samples. Existing methods for OSR in hyperspectral images (HSI) can be generally divided into two categories: reconstruction-based and distance-based methods. Reconstruction-based approaches focus on analyzing reconstruction errors during inference, whereas distance-based methods determine the rejection of unknown samples by measuring their distance to each prototype. However, these techniques often require a substantial amount of training data, which can be both time-consuming and expensive to gather, and they require manual threshold setting, which can be difficult for different tasks. Furthermore, effectively utilizing spectral-spatial information in HSI remains a significant challenge, particularly in open-set scenarios. To tackle these challenges, we introduce a few-shot OSR framework for HSI named HyperTaFOR, which incorporates a novel spatial-spectral selective transformer (S3Former). This framework employs a meta-learning strategy to implement a negative prototype generation module (NPGM) that generates task-adaptive rejection scores, allowing flexible categorization of samples into various known classes and anomalies for each task. Additionally, the S3Former is designed to extract spectral-spatial features, optimizing the use of central pixel information while reducing the impact of irrelevant spatial data. Comprehensive experiments conducted on three benchmark hyperspectral datasets show that our proposed method delivers competitive classification and detection performance in open-set environments when compared to state-of-the-art methods. The code is available online at https://github.com/B-Xi/TIP_2025_HyperTaFOR.
Bobo Xi, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001
IEEE Trans. Image Process.4
2025 HyperCASR: Spectral-Spatial Open-Set Recognition With Category-Aware Semantic Reconstruction for Hyperspectral Imagery
abstract
Open-set recognition (OSR) in hyperspectral imagery (HSI) focuses on accurately classifying known classes while effectively rejecting unknown negative samples. Most existing reconstruction-based approaches are susceptible to noise interference in the input images, and known classes can easily lead to inter-class confusion during the reconstruction process. Moreover, effectively utilizing the abundant spectral-spatial information in HSI within an open-set context presents significant challenges. To address these issues, we propose HyperCASR, an innovative framework for HSI OSR that integrates a grouped spectral-spatial retentive transformer (GSSRT) and a class-aware semantic reconstruction (CASR) module. This method begins by designing the GSSRT to extract features from HSI, enhancing the extraction capability of spatial-spectral information by introducing a grouped pixel embedding (GPE) module and a novel spatial retentive attention (SRA) mechanism. Subsequently, an independent autoencoder (AE) is assigned to each known class to reconstruct semantic features, which helps to mitigate noise interference and inter-class confusion. Additionally, by minimizing reconstruction errors to estimate class affiliation, the framework effectively identifies unknown classes. Experimental results across three benchmark datasets indicate that the HyperCASR framework significantly enhances classification performance for both known and unknown classes when compared to existing state-of-the-art methods. The code is available at https://github.com/B-Xi/TIP_2025_HyperCASR.
Bobo Xi, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001
IEEE Trans. Image Process.4
2025 Uncertainty-Guided Discriminative Priors Mining for Flexible Unsupervised Spectral Reconstruction
abstract
Existing supervised spectral reconstruction (SR) methods adopt paired RGB images and hyperspectral images (HSIs) to drive the overall paradigms. Nonetheless, in practice, "paired" requires higher device requirements such as specific well-calibrated dual cameras or more complex and exact registration processes among images with different time phases, widths, and spatial resolution. To tackle the above challenges, we propose a flexible uncertainty-aware unsupervised SR paradigm, which dynamically establishes the forceful and potent constraints with RGBs for driving unsupervised learning. As a specific plug-and-play tail in our paradigm, the uncertainty-aware saliency alignment module (USAM) calculates pixel- and spectralwise information entropy for uncertainty estimation, which attempts to represent the corresponding reflectivity or radiance to the light among different objects in various scenes, forcing the paradigm to adaptively explore the scene-agnostic prominent features. Furthermore, a progressively parallel network under our unsupervised paradigm is conducted to excavate discriminate structural and semantic priors of RGBs to assist in recovering dependable HSIs: 1) a learnable rank-guided structural representation (LRSR) flow is leveraged to characterize the latent structural priors via excavating nonzero elements in the full-rank matrix and further preserve evident boundaries in HSIs; and 2) a coarse-to-fine bandwise semantic perception (CBSP) flow is conducted to propagate perceptual bandwise affinity for aggregating and strengthening intrinsic interband dependencies, and further extract delicate semantic priors, which can recover plentiful contiguous spectral information in HSIs. Comprehensive quantitative and qualitative experimental results on three visual and two remote sensing benchmarks have shown the superiority and robustness of our method. We also conducted nine existing SR methods in our unsupervised paradigm to recover HSIs without any manual intervention, which proves the generality of our paradigm to some extent.
Yihong Leng, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Progressive Spatial Information-Guided Deep Aggregation Convolutional Network for Hyperspectral Spectral Super-Resolution
abstract
Fusion-based spectral super-resolution aims to yield a high-resolution hyperspectral image (HR-HSI) by integrating the available high-resolution multispectral image (HR-MSI) with the corresponding low-resolution hyperspectral image (LR-HSI). With the prosperity of deep convolutional neural networks, plentiful fusion methods have made breakthroughs in reconstruction performance promotions. Nevertheless, due to inadequate and improper utilization of cross-modality information, the most current state-of-the-art (SOTA) fusion-based methods cannot produce very satisfactory recovery quality and only yield desired results with a small upsampling scale, thus affecting the practical applications. In this article, we propose a novel progressive spatial information-guided deep aggregation convolutional neural network (SIGnet) for enhancing the performance of hyperspectral image (HSI) spectral super-resolution (SSR), which is decorated through several dense residual channel affinity learning (DRCA) blocks cooperating with a spatial-guided propagation (SGP) module as the backbone. Specifically, the DRCA block consists of an encoding part and a decoding part connected by a channel affinity propagation (CAP) module and several cross-layer skip connections. In detail, the CAP module is customized by exploiting the channel affinity matrix to model correlations among channels of the feature maps for aggregating the channel-wise interdependencies of the middle layers, thereby further boosting the reconstruction accuracy. Additionally, to efficiently utilize the two cross-modality information, we developed an innovative SGP module equipped with a simulation of the degradation part and a deformable adaptive fusion part, which is capable of refining the coarse HSI feature maps at pixel-level progressively. Extensive experimental results demonstrate the superiority of our proposed SIGnet over several SOTA fusion-based algorithms.
Jiaojiao Li 0001, Songcheng Du, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 PCViT: A Pyramid Convolutional Vision Transformer Detector for Object Detection in Remote-Sensing Imagery
abstract
Remote sensing object detection (RSOD) is a fundamental and valuable task in Earth monitoring. However, remote sensing images (RSIs) are typically acquired from a bird’s eye perspective, resulting in intrinsic properties such as the complex backgrounds, random and dense distribution of objects, and multiscale objects. These properties hinder the direct application of well-performed detection methods in the natural images (NIs) domain to the RSIs domain, thereby limiting the attainment of desired performance. To address this, we propose a pyramid convolutional vision transformer (PCViT) that gets rid of the limitations of existing transformer methods. Firstly, we employ a pyramid architecture to effectively capture the multiscale information present in RSIs. To enhance the feature extraction capabilities of the transformer, we introduce a parallel convolution module (PCM) that complements the local information that may be missed by the transformer. Furthermore, we propose a self-supervised pretraining strategy called multi-perspective pretraining (MPP) to pretrain the model and subsequently finetune it on the downstream detection task. During the finetuning stage, we introduce a Local/globalk-NN attention (LGKA) to improve the token relationship establishment. In the neck part, we propose a feature-reflowing pyramid network (FRPN) to facilitate contextual information interaction and further enhance our PCViT’s ability to process multiscale information. Experimental results on two representative datasets, namely NWPU VHR-10 and DIOR, demonstrate the effectiveness of our PCViT, as it achieves outstanding performance. These results highlight the suitability of PCViT for RSOD tasks.
Jiaojiao Li 0001, Penghao Tian, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Model-Driven Deep Pipeline With Uncertainty-Aware Bundle Adjustment for Satellite Photogrammetry
abstract
Bundle adjustment (BA), a vital technology in satellite photogrammetry, directly determines the quality of geographic information mapping. However, the existing BA methods suffer from bottlenecks in the cases of limited stereo views caused by input guidance inadequacy and biased modeling. To conquer these issues, a model-driven satellite photogrammetry deep pipeline (SPDP) is proposed in this article. Specifically, for the triplets of remote sensing images (RSIs), the fusion feature maps are extracted by our attention-driven multiscale feature extractor (AMFE), which emphasizes the image information and provides guidance for the subsequent multiview geometric processing. Following that, with the feature error volume as input, a dedicated feature-metric error perceptron module (FEPM) is built to infer the observation uncertainty and predict the pixel-wise compensations. Furthermore, a novel uncertainty-aware BA (UBA) is implemented to derive accurate and robust 3-D point clouds, which introduces the BA model transformation and the specialized iterative refinement to enhance the observation error elimination capability. The detailed experimental results demonstrate the feasibility and effectiveness of the proposed pipeline, which is significant for remote sensing surveys and mapping.
Kailang Cao, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Residual Mask in Cascaded Convolutional Transformer for Spectral Reconstruction
abstract
A significant challenge of spectral reconstruction (SR) task is the lower performance reconstructed in foreground regions compared to background regions, which can be attributed to the marked difference in diversity of objects and disparity of adjacent scene characteristics. Moreover, the reconstruction of edge regions is often fraught with substantial errors due to the transitional nature of these regions, an issue conventional single convolutional neural networks (CNNs) and transformers struggle to handle. To address these challenges, we introduce the residual mask in cascaded convolutional transformer (RC2T) to iteratively improve the reconstruction of hyperspectral images (HSIs). Specifically, we propose a residual-predict mask generator (RMG) to generate a residual mask that retains band properties to separate feature with different complexities. Meanwhile, to achieve band expansion of mask features within the autoencoder, we approximate it to a Markov process and exploit the multistage spectral-aware Markov transfer (MMT) for its lightweight implementation. Next, we introduce the parallel convolutional multihead self-attention module (PSM), in which CNN runs parallel to the transformer to handle simple and complex features separately. Additionally, the residual mask loss function uses the established relationship between complexity of feature and reconstruction accuracy to generate residual mask in a self-supervised manner for providing complex high-frequency prior. We have validated our approach using three published datasets (NTIRE 2020 “Clean” track, NTIRE 2022, and CAVE). Additionally, we also conducted experiments with the proposed method on remote sensing dataset grss_dfc_2018 and a satellite-borne remote sensing dataset, achieving optimal performance. The experimental results demonstrate that our RC2T method is state-of-the-art (SOTA) in the field of SR.
Jiaojiao Li 0001, Shiyao Duan, Yihong Leng, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 HyperMLP: Superpixel Prior and Feature Aggregated Perceptron Networks for Hyperspectral and LiDAR Hybrid Classification
abstract
Hyperspectral images have excellent spectral combining capabilities and LiDAR images have fine stereooscopic elevation information. Therefore, the multi-modal fusion classification of hyperspectral and LiDAR images is inevitably improves the interpretation ability of remote sensing images. In recent years, the MLP-Mixer, an image processing network based on MLP, has flourished in the field of image processing. In this work, we propose an innovative HyperMLP network based on the deep learning framework MLP-Mixer architecture to address the lack of spatial feature construction capability and the locality of multi-modal feature fusion in naive networks. Specifically,(1) The adoption of unsupervised superpixel embedding provides additional shallow morphological spatial feature information for the network, reduces the pressure of the feature extraction network, and enhances feature discrimination capabilities. (2) The feature scrambling strategy improves the diversity of features and strengthens generalization of the network by enhancing interactions between different spatial features. (3) By implementing the bilateral modulation strategy, feature fusion is applied at every stage of the deep network, reducing semantic drift between features. On three fiducial remote sensing datasets, classification tests are performed on the proposed HyperMLP network to verify its performance, and the results are definitely impressive.
Jiaojiao Li 0001, Rui Song 0003, Wei Li 0032, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 SWFormer: Stochastic Windows Convolutional Transformer for Hybrid Modality Hyperspectral Classification
abstract
Joint classification of hyperspectral images with hybrid modality can significantly enhance interpretation potentials, particularly when elevation information from the LiDAR sensor is integrated for outstanding performance. Recently, the transformer architecture was introduced to the HSI and LiDAR classification task, which has been verified as highly efficient. However, the existing naive transformer architectures suffer from two main drawbacks: 1) Inadequacy extraction for local spatial information and multi-scale information from HSI simultaneously. 2) The matrix calculation in the transformer consumes vast amounts of computing power. In this paper, we propose a novel Stochastic Window Transformer (SWFormer) framework to resolve these issues. First, the effective spatial and spectral feature projection networks are built independently based on hybrid-modal heterogeneous data composition using parallel feature extraction, which is conducive to excavating the perceptual features more representative along different dimensions. Furthermore, to construct local-global nonlinear feature maps more flexibly, we implement multi-scale strip convolution coupled with a transformer strategy. Moreover, in an innovative random window transformer structure, features are randomly masked to achieve sparse window pruning, alleviating the problem of information density redundancy, and reducing the parameters required for intensive attention. Finally, we designed a plug-and-play feature aggregation module that adapts domain offset between modal features adaptively to minimize semantic gaps between them and enhance the representational ability of the fusion feature. Three fiducial datasets demonstrate the effectiveness of the SWFormer in determining classification results.
Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Image Process.4
2024 SCFormer: Spectral Coordinate Transformer for Cross-Domain Few-Shot Hyperspectral Image Classification
abstract
Cross-domain (CD) hyperspectral image classification (HSIC) has been significantly boosted by methods employing Few-Shot Learning (FSL) based on CNNs or GCNs. Nevertheless, the majority of current approaches disregard the prior information of spectral coordinates with limited interpretability, leading to inadequate robustness and knowledge transfer. In this paper, we propose an asymmetric encoder-decoder architecture, Spectral Coordinate Transformer (SCFormer), for the CDFSL HSIC task. Several dense Spectral Coordinate blocks (SC blocks) are embedded in the backbone of the encoder to establish feature representation with better generalization, which integrates spectral coordinates via Rotary Position Embedding (RoPE) to minimize spectral position disturbance caused by the convolution operation. Due to a large amount of hyperspectral image data and the high demand for model generalization ability in cross-domain scenarios, we design two mask patterns (Random Mask and Sequential Mask) built on unexploited spectral coordinates within the SC blocks, which are unified with the asymmetric structure to learn high-capacity models efficiently and effectively with satisfactory generalization. Besides, from the perspective of the loss function, we devise an intra-domain loss function founded on the Orthogonal Complement Space Projection (OCSP) theory to facilitate the aggregation of samples in the metric space, which promotes intra-domain consistency and increases interpretability. Finally, the strengthened class expression capacity of the intra-domain loss function contributes to the inter-domain loss function constructed by Wasserstein Distance (WD) for realizing domain alignment. Experimental results on four benchmark data sets demonstrate the superiority of the SCFormer.
Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Image Process.3
2024 HPRN: Holistic Prior-Embedded Relation Network for Spectral Super-Resolution
abstract
Spectral super-resolution (SSR) refers to the hyperspectral image (HSI) recovery from an RGB counterpart. Due to the one-to-many nature of the SSR problem, a single RGB image can be reprojected to many HSIs. The key to tackle this ill-posed problem is to plug into multisource prior information such as the natural spatial context prior of RGB images, deep feature prior, or inherent statistical prior of HSIs so as to effectively alleviate the degree of ill-posedness. However, most current approaches only consider the general and limited priors in their customized convolutional neural networks (CNNs), which leads to the inability to guarantee the confidence and fidelity of reconstructed spectra. In this article, we propose a novel holistic prior-embedded relation network (HPRN) to integrate comprehensive priors to regularize and optimize the solution space of SSR. Basically, the core framework is delicately assembled by several multiresidual relation blocks (MRBs) that fully facilitate the transmission and utilization of the low-frequency content prior of RGBs. Innovatively, the semantic prior of RGB inputs is introduced to mark category attributes, and a semantic-driven spatial relation module (SSRM) is invented to perform the feature aggregation of clustered similar ranges for refining recovered characteristics. In addition, we develop a transformer-based channel relation module (TCRM), which breaks the habit of employing scalars as the descriptors of channelwise relations in the previous deep feature prior and replaces them with certain vectors to make the mapping function more robust and smoother. In order to maintain the mathematical correlation and spectral consistency between hyperspectral bands, the second-order prior constraints (SOPCs) are incorporated into the loss function to guide the HSI reconstruction. Finally, extensive experimental results on four benchmarks demonstrate that our HPRN can reach the state-of-the-art performance for SSR quantitatively and qualitatively. Furthermore, the effectiveness and usefulness of the reconstructed spectra are verified by the classification results on the remote sensing dataset. Codes are available at https://github.com/Deep-imagelab/HPRN.
Chaoxiong Wu, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Shape-Constraint Recurrent Flow for 6D Object Pose Estimation
abstract
Most recent 6D object pose methods use 2D optical flow to refine their results. However, the general optical flow methods typically do not consider the target's 3D shape information during matching, making them less effective in 6D object pose estimation. In this work, we propose a shape-constraint recurrent matching framework for 6D object pose estimation. We first compute a pose-induced flow based on the displacement of 2D reprojection between the initial pose and the currently estimated pose, which embeds the target's 3D shape implicitly. Then we use this pose-induced flow to construct the correlation map for the following matching iterations, which reduces the matching space significantly and is much easier to learn. Further-more, we use networks to learn the object pose based on the current estimated flow, which facilitates the computation of the pose-induced flow for the next iteration and yields an end-to-end system for object pose. Finally, we optimize the optical flow and object pose simultaneously in a recurrent manner. We evaluate our method on three challenging 6D object pose datasets and show that it outperforms the state of the art significantly in both accuracy and efficiency.
Yang Hai, Rui Song 0003, Jiaojiao Li 0001, Yinlin Hu
CVPR2
2023 Rigidity-Aware Detection for 6D Object Pose Estimation
abstract
Most recent 6D object pose estimation methods first use object detection to obtain 2D bounding boxes before actually regressing the pose. However, the general object detection methods they use are ill-suited to handle cluttered scenes, thus producing poor initialization to the subsequent pose network. To address this, we propose a rigidity-aware detection method exploiting the fact that, in 6D pose estimation, the target objects are rigid. This lets us introduce an approach to sampling positive object regions from the entire visible object area during training, instead of naively drawing samples from the bounding box center where the object might be occluded. As such, every visible object part can contribute to the final bounding box prediction, yielding better detection robustness. Key to the success of our approach is a visibility map, which we propose to build using a minimum barrier distance between every pixel in the bounding box and the box boundary. Our results on seven challenging 6D pose estimation datasets evidence that our method outperforms general detection frameworks by a large margin. Furthermore, combined with a pose regression network, we obtain state-of-the-art pose estimation results on the challenging BOP benchmark.
Yang Hai, Rui Song 0003, Jiaojiao Li 0001, Mathieu Salzmann, Yinlin Hu
CVPR2
2023 Pseudo Flow Consistency for Self-Supervised 6D Object Pose Estimation
abstract
Most self-supervised 6D object pose estimation methods can only work with additional depth information or rely on the accurate annotation of 2D segmentation masks, limiting their application range. In this paper, we propose a 6D object pose estimation method that can be trained with pure RGB images without any auxiliary information. We first obtain a rough pose initialization from networks trained on synthetic images rendered from the target’s 3D mesh. Then, we introduce a refinement strategy leveraging the geometry constraint in synthetic-to-real image pairs from multiple different views. We formulate this geometry constraint as pixel-level flow consistency between the training images with dynamically generated pseudo labels. We evaluate our method on three challenging datasets and demonstrate that it outperforms state-of-the-art self-supervised methods significantly, with neither 2D annotations nor additional depth images.
Yang Hai, Rui Song 0003, Jiaojiao Li 0001, David Ferstl, Yinlin Hu
ICCV2
2023 Class-Specific Autoaugment Architecture Based on Schmidt Mathematical Theory for Imbalanced Hyperspectral Classification
abstract
Hyperspectral image classification (HSIC) often suffers from severe imbalanced category distribution in real applications, which causes bias toward the dominated categories. As an effective method, the deep generative model (DGM) can be used to augment the features of imbalanced data through a learnable method to achieve superior classification performance. However, the features extracted by DGM are preset as a standard Gaussian distribution which results in low interclass difference. Besides, the generated features are too consistent with the original ones, which cannot play a positive role in the discriminability of minority categories (MCs). To conquer these drawbacks, we propose a class-specific autoaugment architecture based on Schmidt mathematical theory (CACS) for the challenging of imbalanced data which consists of two stages: one is training a superior features extractor, and the other one is augmenting features. The class-specific features of the whole HSI are extracted in stage one that supports the following feature augmented module. Specifically, we weighted the classifier in the first phase according to cost-sensitive learning, to prevent the classifier from overfitting. To expand the dispersion between categories, we construct feature prototypes obeying different Gaussian distributions for each class, respectively, and generate class-specific features. Then, the features are augmented in the second phase based on Schmidt’s mathematical theory, which enhances the discriminability of minority class features, thus further improving the classification accuracy with interpretability. Extensive experimental results on three benchmarking datasets demonstrate that CACS is outstanding in comparison algorithms, especially in MCs.
Jiaojiao Li 0001, Yan Diao, Rui Song 0003, Bobo Xi, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Sal²RN: A Spatial-Spectral Salient Reinforcement Network for Hyperspectral and LiDAR Data Fusion Classification
abstract
Hyperspectral image (HSI) and light detection and ranging (LiDAR) data fusion have been widely employed in HSI classification to promote interpreting performance. In the existing deep learning methods based on spatial–spectral features, the features extracted from different layers are treated fairly in the learning process. In reality, features extracted from the continuous layers contribute differentially to the final classification, such as large tracts of woodland and agriculture typically count on shallow contour features, whereas deep semantic spectral features have meaningful constraints for small entities like vehicles. Furthermore, the majority of existing classification algorithms employ a patch input scheme, which has a high probability to introduce pixels of different categories at the boundary. To acquire more accurate classification results, we propose a spatial–spectral saliency reinforcement network (Sal2RN) in this article. In spatial dimension, a novel cross-layer interaction module (CIM) is presented to adaptively alter the significance of features between various layers and integrate these diversified features. Moreover, a customized center spectrum correction module (CSCM) integrates neighborhood information and adaptively modifies the center spectrum to reduce intraclass variance and further improve the classification accuracy of the network. Finally, a statistically based feature weighted combination module is constructed to effectively fuse spatial, spectral, and LiDAR features. Compared with traditional and advanced classification methods, the Sal2RN achieves the state-of-the-art classification performance on three open benchmark datasets.
Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Kailiang Han, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 MFormer: Taming Masked Transformer for Unsupervised Spectral Reconstruction
abstract
Spectral reconstruction (SR) aims to recover the hyperspectral images (HSIs) from the corresponding RGB images directly. Most SR studies based on supervised learning require massive data annotations to achieve superior reconstruction performance, which are limited by complicated imaging techniques and laborious annotation calibration in practice. Thus, unsupervised strategies attract attention of the community, however, existing unsupervised SR works still face a fatal bottleneck from low accuracy. Besides, traditional CNN-based models are good at capturing local features but experience difficulty in global features. To ameliorate these drawbacks, we propose an unsupervised SR architecture with strong constraints, especially constructing a novel Masked Transformer (MFormer) to excavate latent hyperspectral characteristics to restore realistic HSIs further. Concretely, a Dual Spectral-wise Multi-head Self-attention (DSSA) mechanism embedded in transformer is proposed to firmly associate multi-head and channel dimensions and then capture the spectral representation in the implicit solution spaces. Furthermore, a plug-and-play Mask-guided Band Augment (MBA) module is presented to extract and further enhance the band-wise correlation and continuity to boost the robustness of the model. Innovatively, a customized loss based on the intrinsic mapping from HSIs to RGB images and the inherent spectral structural similarity is designed to restrain spectral distortion. Extensive experimental results on three benchmarks verify that our MFormer achieves superior performance over other state-of-the-art supervised and unsupervised methods under a no-label training process equally.
Jiaojiao Li 0001, Yihong Leng, Rui Song 0003, Wei Li 0032, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 RepCPSI: Coordinate-Preserving Proximity Spectral Interaction Network With Reparameterization for Lightweight Spectral Super-Resolution
abstract
Existing remarkable models for spectral super-resolution (SSR) achieve higher precision at the expense of computations with larger parameters. These algorithms require the heavy memory footprint and sufficient computing power, limiting their practical deployments and applications on portable devices. In this paper, we propose an efficient re-parameterizing coordinate-preserving proximity spectral interaction (RepCPSI) network for lightweight SSR. Specifically, the basic architecture is constituted of several polymorphic residual context restructuring (PRCR) modules to fully explore spatial and spectral contextual information with a multi-branch topology during the training stage. Using a structural re-parameterization scheme, the training-completed network is converted equivalently to a high-efficiency inference-time model, when it runs in the testing phase. To significantly improve the accuracy of SSR with an extra negligible computational overhead, a lightweight coordinate-preserving proximity spectral-aware attention (CPSA) block is developed. Such CPSA block can adaptively emphasize informative signatures and suppress useless ones among intermediate spatial-spectral features, which effectively enables the model to quickly locate features that are beneficial to the network learning and representation. Furthermore, considering the continuity of spectral variation for capturing real-world HSIs, a spectral physical consistency loss (SPCL) is added to the end-to-end network to constrain the changing trend of the spectral curve to be consistent with the ground-truth objects. Finally, our RepCPSI can accomplish a favorable balance between the reconstructed quality and model complexity. Extensive experimental results on six benchmarks demonstrate that our method obtains excellent performance with fewer parameters in terms of quantitative and qualitative measurements over the current advanced SSR approaches.
Chaoxiong Wu, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Structure-Aware Graph Convolution Network for Point Cloud Parsing
abstract
Point clouds are becoming a popular medium to describe 3D scenes, benefitting from their accuracy and completeness in expressing the spatial and geometrical information of objects. However, due to the disorder and uneven distribution nature, merely selecting neighbors for point clouds in Euclidean space is inefficient and position-ignoring. To fill this gap, we propose a structure-aware graph convolution network (SA-GCN), which consists of an adaptive dilated KNN module (ADKNN), a learnable graph filter (LGF), and a structure-aware feature transformation module (SFT). Specially, the ADKNN module can dynamically adjust the range of grouping neighbor points, while being universal to improve the performance of arbitrary KNN-based methods. Moreover, with the localized auxiliary information provided by LGF, our SFT module disentangles the spatial details as a sort of coding guidance for better deep feature representations. Extensive experimental results on point cloud classification and segmentation tasks demonstrate the superiority of our proposed network.
Fengda Hao, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Kailang Cao
IEEE Trans. Multim.3
2023 Deep Hybrid 2-D-3-D CNN Based on Dual Second-Order Attention With Camera Spectral Sensitivity Prior for Spectral Super-Resolution
abstract
A largely ignored fact in spectral super-resolution (SSR) is that the subsistent mapping methods neglect the auxiliary prior of camera spectral sensitivity (CSS) and only pay attention to wider or deeper network framework design while ignoring to excavate the spatial and spectral dependencies among intermediate layers, hence constraining representational capability of convolutional neural networks (CNNs). To conquer these drawbacks, we propose a novel deep hybrid 2-D-3-D CNN based on dual second-order attention with CSS prior (HSACS), which can excavate sufficient spatial-spectral context information. Specifically, dual second-order attention embedded in the residual block for more powerful spatial-spectral feature representation and relation learning is composed of a brand new trainable 2-D second-order channel attention (SCA) or 3-D second-order band attention (SBA) and a structure tensor attention (STA). Concretely, the band and channel attention modules are developed to adaptively recalibrate the band-wise and interchannel features via employing second-order band or channel feature statistics for more discriminative representations. Besides, the STA is promoted to rebuild the significant high-frequency spatial details for enough spatial feature extraction. Moreover, the CSS is first employed as a superior prior to avoid its effect of SSR quality, on the strength of which the resolved RGB can be calculated naturally through the super-reconstructed hyperspectral image (HSI); then, the final loss consists of the discrepancies of RGB and the HSI as a finer constraint. Experimental results demonstrate the superiority and progressiveness of the presented approach in terms of quantitative metrics and visual effect over SOTA SSR methods.
Jiaojiao Li 0001, Chaoxiong Wu, Rui Song 0003, Yunsong Li 0001, Weiying Xie, Lihuo He, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Semisupervised Cross-Scale Graph Prototypical Network for Hyperspectral Image Classification
abstract
In practice, the acquirement of labeled samples for hyperspectral image (HSI) is time-consuming and labor-intensive. It frequently induces the trouble of model overfitting and performance degradation for the supervised methodologies in HSI classification (HSIC). Fortunately, semisupervised learning can alleviate this deficiency, and graph convolutional network (GCN) is one of the most effective semisupervised approaches, which propagates the node information from each other in a transductive manner. In this study, we propose a cross-scale graph prototypical network (X-GPN) to achieve semisupervised high-quality HSIC. Specifically, considering the multiscale appearance of the land covers in the same remotely captured scene, we involve the neighborhoods of different scales to construct the adjacency matrices and simultaneously design a multibranch framework to investigate the abundant spectral-spatial features through graph convolutions. Furthermore, to exploit the complementary information between different scales, we simply employ the standard 1-D convolution to excavate the dependence of the intranode and concatenate the output with the features generated from other scales. Intuitively, different branches for various samples should have different importance to predict their categories. Thus, we develop a self-branch attentional addition (SBAA) module to adaptively highlight the most critical features produced by multiple branches. In addition, different from previous GCN for HSIC, we devise an innovative prototypical layer comprising a distance-based cross-entropy (DCE) loss function and a novel temporal entropy-based regularizer (TER), which can enhance the discrimination and representativeness of the node features and prototypes actively. Extensive experiments demonstrate that the proposed X-GPN is superior to the classic and state-of-the-art (SOTA) methods in terms of the classification performance.
Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Yuchao Xiao, Qian Du 0001, Jocelyn Chanussot
IEEE Trans. Neural Networks Learn. Syst.4
2022 Cascaded geometric feature modulation network for point cloud processing
Fengda Hao, Rui Song 0003, Jiaojiao Li 0001, Kailang Cao, Yunsong Li 0001
Neurocomputing2
2022 HASIC-Net: Hybrid Attentional Convolutional Neural Network With Structure Information Consistency for Spectral Super-Resolution of RGB Images
abstract
Spectral super-resolution (SSR), referring to the recovery of a reasonable hyperspectral image (HSI) from a single RGB image, has achieved satisfactory performance as part of the continued development of a convolutional neural network (CNN) in remote sensing image processing. However, the majority of existing algorithms focus on the pursuit of networks with deeper or broader architecture. Such algorithms have a poor channel or band feature extraction and fusing performance, and fail to fully leverage the input RGB images. To overcome these issues, we present a novel hybrid attentional CNN with structure information consistency (HASIC-net) that uses a two-pathway architecture. Specifically, both sides are stacked with several 2-D residual groups (2-DRGs) and residual groups (1-DRGs) equipped with channel or band attention (BA) modules, which mainly focuses on extracting channel statistics and bandwise features, respectively, by a parallel pooling architecture. We introduce several transversal connections from 2-DRG to 1-DRG to realize the interaction of information flow between both sides. In addition, we take the structure information of both RGB images and HSI into consideration and devise a structure information consistency (SIC) module to merge the structure tensor prior to the RGB images with the input of each 2-DRG. We then combine spectral gradient constraint loss with mean relative absolute error as a novel loss function to further restrain the spectral distortion and smooth the reconstructed spectral response curves. Experimental results on four benchmark datasets (i.e., NTIRE 2020, NTIRE 2018, CAVE, and Harvard) demonstrate that our proposed HASIC-net achieves state-of-the-art performance.
Jiaojiao Li 0001, Songcheng Du, Rui Song 0003, Chaoxiong Wu, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 A Triplet Semisupervised Deep Network for Fusion Classification of Hyperspectral and LiDAR Data
abstract
Data fusion of hyperspectral and light detection and ranging (LiDAR) is conducive to obtain more comprehensive surface information and thereby achieve better classification result in Earth Monitoring Systems. However, lack of labeled samples usually limits the performance of supervised classifiers, and the heterogeneity of multi-source data also brings great challenges to data fusion. Aiming to address these issues, we propose a triplet semi-supervised deep convolutional neural network (TSDN) for fusion classification of hyperspectral and LiDAR. Specifically, we utilize three basic pathways to extract deep learning features: 1D-CNN for spectral features in hyperspectral, 2D-CNN for spatial features in hyperspectral and Cascade Net for elevation features in LiDAR data. Furthermore, a novel label calibration module (LCM) is proposed to generate effective pseudo labels with high confidence based on the superpixel segmentation by comparing the multi-view classification results for assisting semi-supervised model training. In addition, we design a novel 3D-Cross Attention Block to enhance the complementary spatial features of multi-source data. Experiments on three public HSI-LiDAR benchmarks: Houston, Trento, and MUUFL Gulfport have demonstrated the effectiveness and superiority of our proposed method.
Jiaojiao Li 0001, Yinle Ma, Rui Song 0003, Bobo Xi, Danfeng Hong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Hyperspectral Pansharpening With Adaptive Feature Modulation-Based Detail Injection Network
abstract
Recently, deep learning-based methodologies have attained unprecedented performance in hyperspectral (HS) pansharpening, which aims to improve the spatial quality of HS images (HSIs) by making use of details extracted from the high-resolution panchromatic (HR-PAN) image. However, it remains challenging to incorporate the details into the pansharpened image effectively, while alleviating the spectral distortion simultaneously. To tackle this problem, in this article, we propose an adaptive feature modulation-based detail injection network (AFM-DIN) for HS pansharpening, which mainly consists of four phases: high-frequency details generation of the HR-PAN image, multiscale feature extraction of the upsampled HSI, AFM-based detail injection and reconstruction of the HR-HSI. First, a novel octave convolution unit is employed to decompose the HR-PAN image into high and low frequencies, and then merge the high-frequency features together to generate the comprehensive PAN-details. Second, the spatial and spectral separable 3D convolution units with multiple kernel sizes are designed to extract multiscale features of the upsampled HSI in a computationally efficient manner. Subsequently, by taking the critical PAN-details as prior, the proposed AFM module is able to not only incorporate the detail information effectively, but also adjust the injected details adaptively to ensure the spectral fidelity. Finally, the anticipated HR-HSI is obtained through adding the upsampled HSI to the predicted HSI-details reconstructed from informative modulated features. Extensive comparison experiments with several state-of-the-arts conducted on simulated and real HS data sets demonstrate that our proposed AFM-DIN can achieve superior pansharpening accuracy in both spatial and spectral aspects.
Yunsong Li 0001, Yuxuan Zheng, Jiaojiao Li 0001, Rui Song 0003, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2022 A Stepwise Domain Adaptive Segmentation Network With Covariate Shift Alleviation for Remote Sensing Imagery
abstract
Semantic segmentation for remote sensing images (RSI) is critical for the Earth monitoring system. However, the covariate shift between RSI datasets under different capture conditions cannot be alleviated by directly using the unsupervised domain adaptation (UDA) method, which negatively affects the segmentation accuracy in RSI. We propose a stepwise domain adaptive segmentation network with covariate shift alleviation (Cov-DA) for RSI parsing to solve this issue. Specifically, to alleviate domain shift generated by different sensors, both the source and target domains are projected into a colorspace with normalized distribution through an elaborate colorspace mapping unified module (CMUM). The color distributions of these two domains tend to be more uniform. Furthermore, in the target domain, the multistatistics joint evaluation module (MJEM) is proposed to capture different statistical characteristics of subscenarios for selecting plain scenarios regarded as high-confidence segmentation results to assist the further improvement of segmentation performance. In addition, a pyramid perceptual attention module (PPAM) containing omnidirectional features without computational burdens is added to our network for effectively enhancing the multiscale feature capture ability. In the cross-city DA experiments based on the International Society for Photogrammetry and Remote Sensing (ISPRS) and aerial benchmarks, the superiority of our algorithm is significantly demonstrated. Furthermore, we release a large-scale Martian terrain dataset noted as “Mars-Seg” containing 5 K images with pixel-level accurate annotations regarding issues, such as the lack of semantic segmentation datasets for unknown scenes.
Jiaojiao Li 0001, Shunyao Zi, Rui Song 0003, Yunsong Li 0001, Yinlin Hu, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Structure-Guided Feature Transform Hybrid Residual Network for Remote Sensing Object Detection
abstract
Object detection in remote sensing imagery (RSI) is a fundamental task for Earth monitoring. Objects captured from the bird’s eye view perspective in RSI can appear as multiscale in arbitrary orientations, most of which are small and dense. In specific, vehicles or ships only occupy a dozen pixels in the image, but are surrounded by roads and seas, which occupy thousands of pixels and comprise overwhelmingly dominant of all pixels. Although a large number of common object detection methods have been proposed, most of them cannot detect small and dense objects accurately because none of them has paid enough attention to the unique characteristic of RSI. In this work, we propose a novel structure-guided feature transform hybrid residual (SGFTHR) network, which can conquer the low performance of detection of objects at different scales, especially for small and dense objects, in an anchor-free manner. The structure-guided feature transform (SGFT) module is promoted to extract discriminative structural information and guide this information into high-level contextual feature maps, preventing the important low-level spatial and structural information from being lost when the network goes deeper. Furthermore, the hybrid residual (HR) module is embedded in the backbone to acquire multiscale features in a novel hybrid hierarchical residual-like manner. Extensive experiments are performed on the HRRSD and NWPU VHR-10 datasets to evaluate the performance of the SGFTHR network, which demonstrates that our SGFTHR network achieves state-of-the-art detection accuracy with high efficiency and robustness. Specifically, 4.12% improvements in mean average precision (mAP) on the HRRSD dataset compared with baseline powerfully demonstrate the effectiveness and superiority of the SGFTHR network.
Jiaojiao Li 0001, Huanqing Zhang, Rui Song 0003, Weiying Xie, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Corrections to "Multiscale Context-Aware Ensemble Deep KELM for Efficient Hyperspectral Image Classification"
abstract
In the above article[1],Fig. 19was incorrectly placed. The correct image and caption are provided here:
Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Multi-Direction Networks With Attentional Spectral Prior for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) have achieved prominent progress in recent years and demonstrated remarkable properties in spectral–spatial hyperspectral image (HSI) classification. However, conventional spatial-context-based CNNs commonly adopt the single patchwise scheme to represent the to-be-classified samples, which often fails to completely investigate the wealthy spectral–spatial information in complicated situations. For instance, it has great probability to cause misclassifications on the irregular or inhomogeneous areas, especially for the borders across different classes. To counteract this deficiency, we propose a unified multi-direction network (MDN) for HSI Classification (HSIC), which can exhaustively explore the abundant spectral and detailed spatial-context information through multi-direction samples. Additionally, considering the image-spectrum merged structure of the HSI, 3-D Squeeze-and-Excitation residual (3DSERes) blocks are devised in each stream of the framework to consecutively learn the spectral and spatial from low-level to high-level features. Specifically, 3DSERes can not only facilitate fluent gradient in backpropagation through skip connections, but also emphasize the significant spectral–spatial features and constrain the futile ones. This characteristic is beneficial to enhance the model’s generalization capability even with limited training samples. Furthermore, for properly aggregating the multi-direction deep features, we exploit the simple, yet effective attentional spectral prior (ASP) creatively through leveraging the original spectral correlations. Extensive experimental results on three benchmark data sets indicate that the proposed MDN-ASP can achieve promising classification performance compared to the state-of-the-art methods.
Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Yuchao Xiao, Yanzi Shi, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Few-Shot Learning With Class-Covariance Metric for Hyperspectral Image Classification
abstract
Recently, embedding and metric-based few-shot learning (FSL) has been introduced into hyperspectral image classification (HSIC) and achieved impressive progress. To further enhance the performance with few labeled samples, we in this paper propose a novel FSL framework for HSIC with a class-covariance metric (CMFSL). Overall, the CMFSL learns global class representations for each training episode by interactively using training samples from the base and novel classes, and a synthesis strategy is employed on the novel classes to avoid overfitting. During the meta-training and meta-testing, the class labels are determined directly using the Mahalanobis distance measurement rather than an extra classifier. Benefiting from the task-adapted class-covariance estimations, the CMFSL can construct more flexible decision boundaries than the commonly used Euclidean metric. Additionally, a lightweight cross-scale convolutional network (LXConvNet) consisting of 3D and 2D convolutions is designed to thoroughly exploit the spectral-spatial information in the high-frequency and low-frequency scales with low computational complexity. Furthermore, we devise a spectral-prior-based refinement module (SPRM) in the initial stage of feature extraction, which cannot only force the network to emphasize the most informative bands while suppressing the useless ones, but also alleviate the effects of the domain shift between the base and novel categories to learn a collaborative embedding mapping. Extensive experiment results on four benchmark data sets demonstrate that the proposed CMFSL can outperform the state-of-the-art methods with few-shot annotated samples.
Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Danfeng Hong, Jocelyn Chanussot
IEEE Trans. Image Process.4
2021 An Extreme Learning Machine Correction Network for High Precision Satellite Attitude Determination
abstract
The fusion framework of star sensor and gyro based on adaptive Kalman filter is widely used in satellite pose estimation. However, the discretization and linearization inevitably introduce system errors, which degrades of the filtering accuracy. To address this problem, we propose a high-precision satellite attitude determination algorithm based on extreme learning machine network correction. We design a dedicated network for error compensation and trained the parameters effectively. In attitude calculation procedure, the forward fusion filtering of star sensor and gyro data is performed firstly by using the adaptive Kalman filter. Then the filtering estimation results are compensated by the extreme learning machine network proposed in this paper. After that, backward smoothing is performed to solve the high-precision attitude. Simulation results show that armed with the compensation procedure of the proposed extreme learning machine network, the accuracy of estimated pose is significantly improved.
Kailang Cao, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Weijiao Jiang
IGARSS3
2021 Spectral Reconstruction Using Residual Channel Affinity Propagation Network with Structural Similarity Constraint
abstract
Recently, deep convolutional neural networks (CNNs) have been widely exploited for spectral reconstruction (SR) and achieved significant promotion. Nevertheless, most of the previous studies paid much attention to the design of the depth and width of the network, and neglected to explore the correlation between the intermediate feature maps, which hindered the representational ability of CNNs. To mitigate this problem, we propose a deep residual channel affinity propagation network (RCAPN) to learn the affinity matrix for more powerful feature expression. Specifically, the backbone consists of several dual residual blocks (DRB) with long and short skip connections to bypass plentiful low-frequency information. Furthermore, a novel channel affinity propagation module (CAPM) embedded in the DRB is investigated to learn the affinity among channels and adaptively integrate interdependencies to strengthen feature representations. Finally, a structural similarity (SSIM) constraint is employed to capture the structural information and recover more accurate edge positions. Experimental results demonstrate the superior performance of our proposed algorithm.
Chaoxiong Wu, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001
IGARSS3
2021 Multi-Scale Structure-Conditioned Feature Transform Network for Object Detection in Remote Sensing Imagery
abstract
In recent years, object detection in remote sensing imagery has attracted more and more attention. Accurate object detection in remote sensing imagery, especially for small objects, is still challenging. Most existing methods utilize the global information in the deep fully convolution layer and neglect the local information in the input image. However, the local information contains sufficient spatial information, which is beneficial for precise localization. Additionally, there still exists variable factors, such as the arbitrary aspect ratio and rotation, which interfere the object detection performance. To solve these problems mentioned above, we propose a novel multi-scale structure-conditioned feature transform network, which adopts FCOS as the baseline and ATSS as the method for training sample selection. On one hand, structural information is extracted to represent the spatial semantic information. On the other hand, multi-scale information is enhanced through a novel hierarchical residual-in-residual module. Experiments on the HRRSD data set have demonstrated the superiority of our method.
Huanqing Zhang, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001
IGARSS3
2021 ASFM-Net: Asymmetrical Siamese Feature Matching Network for Point Completion
abstract
We tackle the problem of object completion from point clouds and propose a novel point cloud completion network employing an Asymmetrical Siamese Feature Matching strategy, termed as ASFM-Net. Specifically, the Siamese auto-encoder neural network is adopted to map the partial and complete input point cloud into a shared latent space, which can capture detailed shape prior. Then we design an iterative refinement unit to generate complete shapes with fine-grained details by integrating prior information. Experiments are conducted on the PCN dataset and the Completion3D benchmark, demonstrating the state-of-the-art performance of the proposed ASFM-Net. Our method achieves the 1st place in the leaderboard of Completion3D and outperforms existing methods with a large margin, about 12%. The codes and trained models are released publicly at https://github.com/Yan-Xia/ASFM-Net.
Yaqi Xia, Yan Xia 0003, Wei Li 0111, Rui Song 0003, Kailang Cao, Uwe Stilla
ACM Multimedia4
2021 Hybrid 2-D-3-D Deep Residual Attentional Network With Structure Tensor Constraints for Spectral Super-Resolution of RGB Images
abstract
RGB image spectral super-resolution (SSR) is a challenging task due to its serious ill-posedness, which aims at recovering a hyperspectral image (HSI) from a corresponding RGB image. In this article, we propose a novel hybrid 2-D-3-D deep residual attentional network (HDRAN) with structure tensor constraints, which can take fully advantage of the spatial-spectral context information in the reconstruction progress. Previous works improve the SSR performance only through stacking more layers to catch local spatial correlation neglecting the differences and interdependences among features, especially band features; different from them, our novel method focuses on the context information utilization. First, the proposed HDRAN consists of a 2D-RAN following by a 3D-RAN, where the 2D-RAN mainly focuses on extracting abundant spatial features, whereas the 3D-RAN mainly simulates the interband correlations. Then, we introduce 2-D channel attention and 3-D band attention mechanisms into the 2D-RAN and 3D-RAN, respectively, to adaptively recalibrate channelwise and bandwise feature responses for enhancing context features. Besides, since structure tensor represents structure and spatial information, we apply structure tensor constraint to further reconstruct more accurate high-frequency details during the training process. Experimental results demonstrate that our proposed method achieves the state-of-the-art performance in terms of mean relative absolute error (MRAE) and root mean square error (RMSE) on both the “clean” and “real world” tracks in the NTIRE 2018 Spectral Reconstruction Challenge. As for competitive ranking metric MRAE, our method separately achieves a 16.06% and 2.90% relative reduction on two tracks over the first place. Furthermore, we investigate HDRAN on the other two HSI benchmarks noted as the CAVE and Harvard data sets, also demonstrating better results than state-of-the-art methods.
Jiaojiao Li 0001, Chaoxiong Wu, Rui Song 0003, Weiying Xie, Chiru Ge, Bo Li 0090, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Multiscale Context-Aware Ensemble Deep KELM for Efficient Hyperspectral Image Classification
abstract
Recently, multiscale spatial features have been widely utilized to improve the hyperspectral image (HSI) classification performance. However, fixed-size neighborhood involving the contextual information probably leads to misclassifications, especially for the boundary pixels. Additionally, it has been demonstrated that deep neural network (DNN) is practical to extract representative features for the classification tasks. Nevertheless, under the condition of high dimensionality versus small sample sizes, DNN tends to be over-fitting and it is generally time-consuming due to the deep-level feature learning process. To alleviate the aforementioned issues, we propose a multiscale context-aware ensemble deep kernel extreme learning machine (MSC-EDKELM) for efficient HSI classification. First, the scene of the HSI data set is over-segmented in multiscale via using the adaptive superpixel segmentation technique. Second, superpixel pattern (SP) and attentional neighboring superpixel pattern (ANSP) are generated by leveraging the superpixel maps, which can automatically comprise local and global contextual information, respectively. Afterward, an ensemble deep kernel extreme learning machine (EDKELM) is presented to investigate the deep-level characteristics in the SP and ANSP. Finally, the category of each pixel is accurately determined by the decision fusion and weighted output layer fusion strategy. Experimental results on four real-world HSI data sets demonstrate that the proposed frameworks outperform some classic and state-of-the-art methods with high computational efficiency, which can be employed to serve real-time applications.
Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2020 Dilated Residual Network Based on Dual Expectation Maximization Attention for Semantic Segmentation of Remote Sensing Images
abstract
Compared with common RGB images, remote sensing images (RSIs) have larger size and lower spatial resolution. RSIs are usually cropped into sub-images for training convolutional neural networks (CNNs), which loses amounts of context information, thus limiting the extraction of feature interdependencies and reducing the accuracy of semantic segmentation. In this paper, a novel dilated residual network based on dual expectation maximization attention (DE-MANet) is proposed for semantic segmentation of RSIs. In specific, we append a dual expectation maximization attention (DEMA) module on top of the dilated CNN. The spatial expectation maximization attention (SEMA) can model spatial feature interdependencies to acquire rich long-range contextual information. The channel expectation maximization attention (CEMA) enhances discriminant ability of channel-wise feature representations through extracting the channel dependencies. We evaluate the model on the dataset released in the Tianzhi Cup Artificial Intelligence Challenge and achieve 85.60% pixel accuracy and 69.00% mean intersection over union (mIoU).
Jiachao Liu, Xinyue Xiong, Jiaojiao Li 0001, Chaoxiong Wu, Rui Song 0003
IGARSS5
2020 Spectral Super-Resolution Using Hybrid 2D-3D Structure Tensor Attention Networks with Camera Spectral Sensitivity Prior
abstract
With the development of deep convolutional neural networks (CNNs), spectral super-resolution (SSR) has obtained a significant improvement, which aims to recover the hyperspectral image (HSI) from a single RGB. However, the existing mapping algorithms lack of utilization of the camera spectral sensitivity (CSS) and only focus on wider or deeper architecture design, neglecting to explore the feature correlations of intermediate layers, thus preventing the representational ability of CNNs. In our paper, a novel hybrid 2D-3D structure tensor attention networks (HSTAN) with CSS prior is proposed for SSR. In specific, a structure tensor attention (STA) embedded in the residual block is invented to extract the salient high-frequency spatial details for adequate spatial feature expression. Furthermore, the CSS is firstly exploited as a prior to avoid its influence of SSR quality, based on which the reconstructed RGB can be calculated naturally through the super-resolved HSI, then the final loss incorporates the discrepancies of RGB and the HSI as a finer constraint. Experimental results demonstrate the superiority of our proposed algorithm.
Chaoxiong Wu, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001
IGARSS3
2020 Hyperspectral Image Super-Resolution by Band Attention Through Adversarial Learning
abstract
Hyperspectral image (HSI) super-resolution (SR) is a challenging task due to the problems of texture blur and spectral distortion when the upscaling factor is large. To meet these two challenges, band attention through the adversarial learning method is proposed in this article. First, we put the SR process in a generative adversarial network (GAN) framework, so that the resulted high-resolution HSI can keep more texture details. Second, different from the other band-by-band SR method, the input of our method is of full bands. In order to explore the correlation of spectral bands and avoid the spectral distortion, a band attention mechanism is proposed in our generative network. A series of spatial-spectral constraints or loss functions is imposed to guide the training of our generative network so as to further alleviate spectral distortion and texture blur. The experiments on the Pavia and Cave data sets demonstrate that the proposed GAN-based SR method can yield very high-quality results, even under large upscaling factor (e.g., $8\times $ ). More importantly, it can outperform the other state-of-the-art methods by a margin which demonstrates its superiority and effectiveness.
Jiaojiao Li 0001, Ruxing Cui, Bo Li 0090, Rui Song 0003, Yunsong Li 0001, Yuchao Dai, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2019 Bilateral adaptive quantization in HEVC
Sanchun Li, Rui Song 0003
Multim. Tools Appl.2
2019 Local Spectral Similarity Preserving Regularized Robust Sparse Hyperspectral Unmixing
abstract
Spatial context has been demonstrated to be effective to constrain sparse unmixing (SU) of hyperspectral images. However, the existing algorithms employed simple spatial information without keeping spectral fidelity. By considering the fact that adjacent pixels own not only the endmembers with same variations but also approximated fractional abundances, in this paper, local spectral similarity preserving (LSSP) constraint is proposed to preserve spectral similarity in a local area during robust sparse unmixing (RSU). Specially, four LSSP constraints are constructed using different-norm-constrained pixel-level difference over abundance-level difference in a local area. Moreover, a convex optimization algorithm is proposed to solve the proposed LSSP-constrained RSU (LSSP-RSU). Experimental results on both synthetic and real hyperspectral data demonstrate that the developed algorithms yield better values of the signal-toreconstruction error (SRE). Especially, when using l2norm of pixel-level difference to weight the l1norm of abundance-level difference, the proposed LSSP-RSU algorithm can achieve superior unmixing performance.
Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Shaohui Mei, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2018 Minimum barrier superpixel segmentation
Yinlin Hu, Yunsong Li 0001, Rui Song 0003, Peng Rao, Yangli Wang
Image Vis. Comput.3
2018 Efficient, robust and divisible paired comparison for subjective quality assessment
Rui Song 0003, Yunsong Li 0001, Yuan Jia, Yangli Wang, Peng Rao
Multim. Tools Appl.1
2018 Coarse-to-Fine PatchMatch for Dense Correspondence
abstract
Although the matching technique has been studied in various areas of computer vision for decades, efficient dense correspondence remains an open problem. In this paper, we present a simple but powerful matching method that works in a coarse-to-fine scheme for optical flow and stereo matching. Inspired by the nearest neighbor field (NNF) algorithms, our approach, called coarse-to-fine PatchMatch, blends an efficient random search strategy with the coarse-to-fine scheme for efficient dense correspondence. Unlike existing NNF techniques, which are efficient but yield results that are often too noisy because of a lack of global regularization, we propose a propagation step involving a constrained random search radius between adjacent levels of a hierarchical architecture. The resulting correspondence has a built-in smoothing effect, making it more suited to dense correspondence than the NNF techniques. Furthermore, our approach can also capture tiny structures with large motions, which is a problem for traditional coarse-to-fine methods. Interpolated using an edge-preserving interpolation method, our method outperforms the state-of-the-art optical flow methods on the MPI-Sintel and KITTI data sets and is much faster than competing methods.
Yunsong Li 0001, Yinlin Hu, Rui Song 0003, Peng Rao, Yangli Wang
IEEE Trans. Circuits Syst. Video Technol.3
2017 Robust Interpolation of Correspondences for Large Displacement Optical Flow
abstract
The interpolation of correspondences (EpicFlow) was widely used for optical flow estimation in most-recent works. It has the advantage of edge-preserving and efficiency. However, it is vulnerable to input matching noise, which is inevitable in modern matching techniques. In this paper, we present a Robust Interpolation method of Correspondences (called RicFlow) to overcome the weakness. First, the scene is over-segmented into superpixels to revitalize an early idea of piecewise flow model. Then, each model is estimated robustly from its support neighbors based on a graph constructed on superpixels. We propose a propagation mechanism among the pieces in the estimation of models. The propagation of models is significantly more efficient than the independent estimation of each model, yet retains the accuracy. Extensive experiments on three public datasets demonstrate that RicFlow is more robust than EpicFlow, and it outperforms state-of-the-art methods.
Yinlin Hu, Yunsong Li 0001, Rui Song 0003
CVPR3
2017 A ParaBoost stereoscopic image quality assessment (PBSIQA) system
Hyunsuk Ko, Rui Song 0003, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.2
2016 Efficient Coarse-to-Fine Patch Match for Large Displacement Optical Flow
abstract
As a key component in many computer vision systems, optical flow estimation, especially with large displacements, remains an open problem. In this paper we present a simple but powerful matching method works in a coarse-to-fine scheme for optical flow estimation. Inspired by the nearest neighbor field (NNF) algorithms, our approach, called CPM (Coarse-to-fine PatchMatch), blends an efficient random search strategy with the coarse-to-fine scheme for optical flow problem. Unlike existing NNF techniques, which is efficient but the results is often too noisy for optical flow caused by the lack of global regularization, we propose a propagation step with constrained random search radius between adjacent levels on the hierarchical architecture. The resulting correspondences enjoys a built-in smoothing effect, which is more suited for optical flow estimation than NNF techniques. Furthermore, our approach can also capture the tiny structures with large motions which is a problem for traditional coarse-to-fine optical flow algorithms. Interpolated by an edge-preserving interpolation method (EpicFlow), our method outperforms the state of the art on MPI-Sintel and KITTI, and runs much faster than the competing methods.
Yinlin Hu, Rui Song 0003, Yunsong Li 0001
CVPR2
2016 Highly accurate optical flow estimation on superpixel tree
Yinlin Hu, Rui Song 0003, Yunsong Li 0001, Peng Rao, Yangli Wang
Image Vis. Comput.2
2015 MCL-V: A streaming video quality assessment database
Joe Yuchieh Lin, Rui Song 0003, Chihao Wu 0001, Tsung-Jung Liu, Haiqiang Wang, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.2
2015 Decoder side information generation techniques in Wyner-Ziv video coding: a review
Yuan Jia, Yangli Wang, Rui Song 0003, Jiandong Li 0001
Multim. Tools Appl.3
2014 High-efficiency pipeline design of binary arithmetic encoder
Rui Song 0003, Hongfei Cui, Yunsong Li 0001, Chengke Wu 0001
Sci. China Inf. Sci.1
2013 Statistically uniform intra-block refresh algorithm for very low delay video communication
abstract
This paper focuses on the mechanism underlying the overall delay of a real-time video communication system from the time of capture at the encoder to the time of display at the decoder. A detailed analysis is presented to illustrate the delay problem. We then describe a statistically uniform intra-block refresh scheme for very low delay video communication. By scattering intra-blocks uniformly into continuous frames, the overall delay is significantly decreased, and object changes in the scene could be presented to the end user instantly. For comparison, the overall delay and the peak signal-to-noise ratio (PSNR) performance are tested. The experiment results show that an average of approximately 0.1 dB PSNR gain on average is obtained relative to random intra-macroblock refresh algorithm in H.264 JM, and the end-to-end delay performance is significantly improved.
Rui Song 0003, Yangli Wang, Yunsong Li 0001
J. Zhejiang Univ. Sci. C1