EDBT 2026 Demo / reviewers in the wild / expert
Yi Xiao 0003
dblp:61/6921-3
· DBLP profile ↗
31ranked-venue papers
12as first author
31since 2021 · last 2026
0000-0001-9533-8917ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image DerainingabstractRain significantly degrades the performance of computer vision systems, particularly in applications like autonomous driving and video surveillance. While existing deraining methods have made considerable progress, they often struggle with fidelity of semantic and spatial details. To address these limitations, we propose the Multi-Prior Hierarchical Mamba (MPHM) network for image deraining. This novel architecture synergistically integrates macro-semantic textual priors (CLIP) for task-level semantic guidance and micro-structural visual priors (DINOv2) for scene-aware structural information. To alleviate potential conflicts between heterogeneous priors, we devise a progressive Priors Fusion Injection (PFI) that strategically injects complementary cues at different decoder levels. Meanwhile, we equip the backbone network with an elaborate Hierarchical Mamba Module (HMM) to facilitate robust feature representation, featuring a Fourier-enhanced dual-path design that concurrently addresses global context modeling and local detail recovery. Comprehensive experiments demonstrate MPHM's state-of-the-art performance, achieving a 0.57 dB PSNR gain on the Rain200H dataset while delivering superior generalization on real-world rainy scenarios. Zhaocheng Yu, Kui Jiang, Junjun Jiang, Xianming Liu 0005, Guanglu Sun, Yi Xiao 0003 |
AAAI | 6 |
| 2026 | Content-aware Information Compression and Selection for Whole Slide Image AnalysisabstractRecent advances in multi-instance learning (MIL) have demonstrated impressive performance in whole slide image (WSI) analysis. However, current methods search for cues and draw conclusions from all instances or regions, resulting in excessive redundant computation and suboptimal representation quality due to irrelevant and uninformative feature interference. To address these issues, we propose CICS, an efficient and general framework that performs compact information compression and selection for high-efficiency WSI analysis. In particular, CICS features two key components: (1) context-aware compression (CAC), which partitions the instance space into sub-regions and applies learnable compression to discard irrelevant components, reduce computational complexity while facilitating information selection, and (2) global-proximity selective attention (GPSA), which cherry-picks the most informative representation with a proximity-assisted global dynamic selection strategy. Building upon these innovations, CICS forms a plug-and-play module that reduces computational complexity through compact instance representations while improving feature quality by preserving the most informative cues. Extensive experiments on six WSI classification and survival prediction datasets show that CICS consistently improves the performance of multiple representative MIL methods. It achieves 2.5%, 7.7%, and 3.9% accuracy gain over the state-of-the-art Transformer-based TransMIL, Mamba-based MambaMIL, and graph-based WIKG methods on the ESCA dataset. Hongxun Yao, Sicheng Zhao, Yi Xiao 0003 |
AAAI | 4 |
| 2026 | Deep low-rank tensor embedded network for hyperspectral image super-resolution
Qiang Zhang 0011, Xianpeng Zhang, Yi Xiao 0003, Hongjie Xie |
Expert Syst. Appl. | 3 |
| 2025 | OODML: Whole Slide Image Classification Meets Online Pseudo-Supervision and Dynamic Mutual LearningabstractBag-label-based multi-instance learning (MIL) has demonstrated significant performance in whole slide image (WSI) analysis, particularly in pseudo-label-based learning schemes. However, due to inaccurate feature representation and interference, existing MIL methods often yield unreliable pseudo-labels, which spawn undesired predictions. To address these issues, we propose an Online Pseudo-Supervision and Dynamic Mutual Learning (OODML) framework that enhances pseudo-label generation and feature representation while exploring their mutual learning to improve bag-level prediction. Specifically, we design an Adaptive Memory Bank (AMB) to collect the most informative components of the current WSI. We also introduce a Self-Progressive Feature Fusion (SPFF) module that integrates label-related historical information from the AMB with current semantic variations, thereby enhancing the representation of pseudo-bag tokens. Furthermore, we propose a Decision Revision Pseudo-Label (DRPL) generation scheme to explore intrinsic connections between pseudo-bag representations and bag-label predictions, resulting in more reliable pseudo-label generation. To alleviate redundant and ambiguous representations, the class-wise prior of pseudo-label prediction is borrowed to facilitate label-related feature learning and to update the AMB, forming a mutual refinement between feature representation and pseudo-label generation. Additionally, a Dynamic Decision-Making (DDM) module is developed to harmonize explicit and implicit representations of bag information for more robust decision-making. Extensive experiments on four datasets demonstrate that our OODML surpasses the state-of-the-art by 3.3% and 6.9% on the CAMELYON16 and TCGA Lung datasets. Kui Jiang, Hongxun Yao, Yi Xiao 0003, Zhongyuan Wang 0001 |
AAAI | 4 |
| 2025 | M3amba: Memory Mamba is All You Need for Whole Slide Image ClassificationabstractMulti-instance learning (MIL) has demonstrated impressive performance in whole slide image (WSI) analysis. However, existing approaches struggle with undesirable results and unbearable computational overhead due to the quadratic complexity of Transformers. Recently, Mamba has offered a feasible solution for modeling long-range dependencies with linear complexity. However, vanilla Mamba inherently suffers from contextual forgetting issues, making it ill-suited for capturing global dependencies across instances in large-scale WSIs. To address this, we propose a memory-driven Mamba network, dubbed M3amba, to fully explore the global latent relations among instances. Specifically, M3amba retains and iteratively updates historical information with a dynamic memory bank (DMB), thus overcoming the catastrophic forgetting defects of Mamba for long-term context representation. For better feature representation, M3amba involves an intra-group bidirectional Mamba (BiMamba) block to refine local interactions within groups. Meanwhile, we additionally perform cross-attention fusion to incorporate relevant historical information across groups, facilitating richer inter-group connections. The joint learning of inter- and intra-group representations with memory merits enables M3amba with a more powerful capability for achieving accurate and comprehensive WSI representation. Extensive experiments on four datasets demonstrate that M3amba outperforms the state-of-the-art by 6.2% and 7.0% in accuracy on the TCGA BRCA and TCGA Lung datasets while maintaining low computational costs. Kui Jiang, Yi Xiao 0003, Sicheng Zhao, Hongxun Yao |
CVPR | 3 |
| 2025 | GMMamba: Group Masking Mamba for Whole Slide Image Classification
Hongxun Yao, Kui Jiang, Yi Xiao 0003, Sicheng Zhao |
ICCV | 4 |
| 2025 | Spiking Meets Attention: Efficient Remote Sensing Image Super-Resolution with Attention Spiking Neural NetworksabstractSpiking neural networks (SNNs) are emerging as a promising alternative to traditional artificial neural networks (ANNs), offering biological plausibility and energy efficiency. Despite these merits, SNNs are frequently hampered by limited capacity and insufficient representation power, yet remain underexplored in remote sensing image (RSI) super-resolution (SR) tasks. In this paper, we first observe that spiking signals exhibit drastic intensity variations across diverse textures, highlighting an active learning state of the neurons. This observation motivates us to apply SNNs for efficient SR of RSIs. Inspired by the success of attention mechanisms in representing salient information, we devise the spiking attention block (SAB), a concise yet effective component that optimizes membrane potentials through inferred attention weights, which, in turn, regulates spiking activity for superior feature representation. Our key contributions include: 1) we bridge the independent modulation between temporal and channel dimensions, facilitating joint feature correlation learning, and 2) we access the global self-similar patterns in large-scale remote sensing imagery to infer spatial attention weights, incorporating effective priors for realistic and faithful reconstruction. Building upon SAB, we proposed SpikeSR, which achieves state-of-the-art performance across various remote sensing benchmarks such as AID, DOTA, and DIOR, while maintaining high computational efficiency. Code of SpikeSR will be available at https://github.com/XY-boy/SpikeSR. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Wenke Huang 0003, Qiang Zhang 0011, Chia-Wen Lin, Liangpei Zhang 0001 |
NeurIPS | 1 |
| 2025 | Collaborative optimization for whole slide image classification
Hongxun Yao, Kui Jiang, Yi Xiao 0003 |
Knowl. Based Syst. | 4 |
| 2025 | GraphMamba: Whole slide image classification meets graph-driven selective state space model
Hongxun Yao, Sicheng Zhao, Kui Jiang, Yi Xiao 0003 |
Pattern Recognit. | 5 |
| 2025 | Errata to "Local-Global Temporal Difference Learning for Satellite Video Super-Resolution"abstractIn the above article, there exists a citation error related to the core technical foundation of the proposed method. Reference [1] was incorrectly cited. The correct citation is reference [2]. Yi Xiao 0003, Qiangqiang Yuan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Uncertainty-Aware With Adaptive Geometric Correction for Multimodal Land-Cover ClassificationabstractLand cover classification (LCC) is a fundamental task in remote sensing and geographic information science. Multi-modal fusion has shown great potential for enhancing LCC performance, for example, by combining optical and synthetic aperture radar (SAR) imagery to leverage their complementary strengths. However, two key challenges hinder effective fusion:1) local geometric mismatches caused by distinct imaging geometries, and2) inconsistent reliability (the ability of a modality to deliver accurate and stable information) in LCC arising from different modalities and their acquisition conditions. To address these issues, we propose Uncertainty-Aware Fusion with Adaptive Geometric Correction (UAG), which comprises three main components. First, the Adaptive Geometric Correction Module (AGCM) applies learnable pixel shifts to establish bidirectional local correlations between multiscale optical and SAR features, thereby mitigating spatial inconsistencies. Second, the Adaptive Uncertainty-Aware Dynamic Fusion Module (ADFM) employs evidential deep learning to model uncertainty, defined as the extent of reliability deficiency, for each modality using the Dirichlet distribution and subjective logic, enabling confidence-aware feature weighting. Third, a lightweight multiscale decoder integrates hierarchical features through a hybrid MLP-convolutional architecture, improving both segmentation efficiency and accuracy. We evaluate UAG on WHU-OPT-SAR and DFC23 datasets, where experimental results demonstrate substantial improvements over state-of-the-art methods. The code will be released at https://github.com/cccwbin/UAGNet. Xu Wang 0015, Yi Xiao 0003, Wenxin Huang, Bihan Wen, Xian Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | STAR: A Unified Spatiotemporal Fusion Framework for Satellite Video Object TrackingabstractSatellite video object tracking (SVOT) delivers comprehensive spatiotemporal insights for Earth surface observation, yet existing SVOT methods confront several critical challenges including data scarcity, modality restrictions, paradigm gaps, and underutilization of multidimensional features, sealing the performance ceiling. This study proposes STAR, a unified spatiotemporal fusion framework for satellite video object tracking, mitigating these issues. To optimize satellite video scenes, STAR first introduces a scene enhancement module for generating enhanced multi-modal representations. Then, the extraction-correlation-adaptation module is designed, incorporating a multi-modal hierarchical Transformer architecture with local and unified relation modeling, which jointly achieves feature extraction, relation learning, and domain adaptation. Additionally, the temporal decoding structure is introduced to integrate deep temporal features via attention propagation. Finally, the inertial navigation module models physical temporal features, including an awareness selector to assess the tracking confidence-uncertainty and an inertial navigation scheme to manage anomalous interferences and continuous trajectory. Inspired by the prompt learning pattern, STAR introduces a minimal number of tunable parameters yet achieves competitive performance across various SVOT benchmarks. Implementation details and evaluation results will be available at: https://github.com/YZCU/STAR. Yuzeng Chen, Qiangqiang Yuan, Yi Xiao 0003, Te Han |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Rep-Mamba: Re-Parameterization in Vision Mamba for Lightweight Remote Sensing Image Super-ResolutionabstractThe selective space model (Mamba) has recently demonstrated great potential in remote sensing image super-resolution (RSISR) tasks due to its capability for long-range dependency modeling with linear computational complexity. Despite these merits, existing Mamba architectures face two critical challenges in large-scale remote sensing scenarios: 1) neglecting the local semantic integrity due to the unfolding 1-D sequential representations and 2) facing the dilemma between effectiveness and efficiency. To address these issues, we propose Rep-Mamba, a lightweight progressive multiscale feature fusion architecture based on the state-space model (SSM) for RSISR. Specifically, we innovatively design a cross-scale state propagation (CSSP) mechanism and construct a lightweight progressive fusion module (LPFM) to dynamically capture hierarchical spatial dependencies in remote sensing scenes while maintaining high computational efficiency. Moreover, to achieve synergistic optimization between local semantic structure preservation and global context modeling, we introduce differentiable re-parameterization convolution (RepConv), which significantly enhances reconstruction accuracy and visual quality without compromising computational efficiency. Extensive experiments across multiple benchmarks demonstrate that Rep-Mamba achieves a superior tradeoff between accuracy and complexity, highlighting its effectiveness and scalability. The code is available athttps://github.com/meigeni0929/Rep-Mambahttps://github.com/meigeni0929/Rep-Mamba Kui Jiang, Mengru Yang, Yi Xiao 0003, Guangcheng Wang, Junjun Jiang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | MTCSAR: Fusing Multitemporal Interaction With Coherence Prior for SAR Image DenoisingabstractSynthetic aperture radar (SAR) images are inherently affected by speckle noise due to their imaging principles, significantly impacting downstream research. Denoising methods based on deep learning have garnered attention and are gradually maturing, yet they face certain challenges. Denoising single SAR images lacks temporal information, often resulting in structural fitting that introduces artifacts. The existing denoising techniques do not adequately account for the differences between SAR and optical images, despite their distinct structural characteristics. In addition, mainstream deep learning approaches heavily rely on data-driven methods, with limited consideration for statistical properties. In response to these challenges, we propose a novel network for denoising SAR images fusing multitemporal interaction with coherence prior, termed MTCSAR. The proposed network leverages multitemporal interaction (MTI) to gather information from different moments in various regions. The redundancy across these temporal dimensions helps mitigate artifacts and edge blurring. To address SAR image characteristics, we design a dual-branch network combining global and local information to handle large-scale regions and fine structures. In addition, we incorporate prior coherence information from SAR images into the network, utilizing statistical properties to enhance transparency during training. Experimental results on simulated and real datasets demonstrate that injecting MTIs and coherence improves our method’s qualitative and quantitative performance, surpassing current state-of-the-art algorithms. This validates the effectiveness of the proposed MTCSAR for multitemporal SAR image denoising. Xin Su 0003, Yi Xiao 0003, Jie Li 0022, Qiangqiang Yuan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Hyperspectral Video Tracking With Spectral-Spatial Fusion and Memory EnhancementabstractHyperspectral video (HSV) provides rich spectral-spatial-temporal information, enabling the capture of complex object dynamics beyond the limitations of conventional single- and multi-modal tracking. However, current HSV tracking methods face challenges such as data scarcity, band gaps, spectral fragmentation, temporal underutilization, and high computational load, which constrain performance. In this article, we present SpectralTrack, a novel HSV tracking framework with spectral-spatial fusion and memory enhancement. SpectralTrack incorporates an explicit visual prompting module to mitigate band gaps and spectral fragmentation. We further introduce an extraction-matching-interaction module, which leverages a template-bridging search adapter and a multi-layer perceptron adapter within a multi-modal Transformer architecture for efficient cross-modal feature extraction-matching-interaction. Additionally, a memory perception module enhances state reasoning by injecting temporal prompts to refine spectral and spatial cues. SpectralTrack follows parameter-efficient fine-tuning and feature-level fusion to alleviate data scarcity and reduce computational overhead. We instantiate two variants, SpectralTrack and SpectralTrack+, across nine HSV tracking datasets, demonstrating superior effectiveness over extensive trackers. Implementations and results will be available at https://github.com/YZCU/SpectralTrack. Yuzeng Chen, Qiangqiang Yuan, Hong Xie 0002, Yi Xiao 0003, Renxiang Guan, Xinwang Liu 0002, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Multi-Axis Feature Diversity Enhancement for Remote Sensing Video Super-ResolutionabstractHow to aggregate spatial-temporal information plays an essential role in video super-resolution (VSR) tasks. Despite the remarkable success, existing methods adopt static convolution to encode spatial-temporal information, which lacks flexibility in aggregating information in large-scale remote sensing scenes, as they often contain heterogeneous features (e.g., diverse textures). In this paper, we propose a spatial feature diversity enhancement module (SDE) and channel diversity enhancement module (CDE), which explore the diverse representation of different local patterns while aggregating the global response with compactly channel-wise embedding representation. Specifically, SDE introduces multiple learnable filters to extract representative spatial variants and encodes them to generate a dynamic kernel for enriched spatial representation. To explore the diversity in the channel dimension, CDE exploits the discrete cosine transform to transform the feature into the frequency domain. This enriches the channel representation while mitigating massive frequency loss caused by pooling operation. Based on SDE and CDE, we further devise a multi-axis feature diversity enhancement (MADE) module to harmonize the spatial, channel, and pixel-wise features for diverse feature fusion. These elaborate strategies form a novel network for satellite VSR, termed MADNet, which achieves favorable performance against state-of-the-art method BasicVSR++ in terms of average PSNR by 0.14 dB on various video satellites, including JiLin-1, Carbonite-2, SkySat-1, and UrtheCast. Code will be available at https://github.com/XY-boy/MADNet. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Yuzeng Chen, Shiqi Wang 0001, Chia-Wen Lin |
IEEE Trans. Image Process. | 1 |
| 2025 | Frequency-Assisted Mamba for Remote Sensing Image Super-ResolutionabstractRecent progress in remote sensing image (RSI) super-resolution (SR) has exhibited remarkable performance using deep neural networks, e.g., Convolutional Neural Networks and Transformers. However, existing SR methods often suffer from either a limited receptive field or quadratic computational overhead, resulting in sub-optimal global representation and unacceptable computational costs in large-scale RSI. To alleviate these issues, we develop the first attempt to integrate the Vision State Space Model (Mamba) for RSI-SR, which specializes in processing large-scale RSI by capturing long-range dependency with linear complexity. To achieve better SR reconstruction, building upon Mamba, we devise a Frequency-assisted Mamba framework, dubbed FMSR, to explore the spatial and frequent correlations. In particular, our FMSR features a multi-level fusion architecture equipped with the Frequency Selection Module (FSM), Vision State Space Module (VSSM), and Hybrid Gate Module (HGM) to grasp their merits for effective spatial-frequency fusion. Considering that global and local dependencies are complementary and both beneficial for SR, we further recalibrate these multi-level features for accurate feature fusion via learnable scaling adaptors. Extensive experiments on AID, DOTA, and DIOR benchmarks demonstrate that our FMSR outperforms state-of-the-art Transformer-based methods HAT-L in terms of PSNR by 0.11 dB on average, while consuming only 28.05% and 19.08% of its memory consumption and complexity, respectively. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Yuzeng Chen, Qiang Zhang 0011, Chia-Wen Lin |
IEEE Trans. Multim. | 1 |
| 2024 | Remote Sensing Image Super-Resolution with Top-K Token Selective TransformerabstractTransformer-based super-resolution (SR) method has recently demonstrated promising performance, due to its long-range and global aggregation capability. However, the existing Transformer brings two critical challenges for applying it in large-area earth observation scenes: (1) redundant token representation due to most irrelevant tokens; (2) single-scale representation which ignores scale correlation modeling of similar ground observation targets. To this end, this paper proposes to adaptively eliminate the interference of irreverent tokens for a more compact self-attention calculation. Specifically, we devise a Residual Token Selective Group (RTSG) to grasp the most crucial token by dynamically selecting the top-k keys in terms of score ranking for each query. For better feature aggregation, a Multi-scale Feed-forward Layer (MFL) is developed to generate an enriched representation of multi-scale feature mixtures during the feed-forward process. In particular, multiple cascaded RTSGs form our final Top-k Token Selective Transformer (TTST) to achieve progressive representation. Extensive experiments on three remote sensing benchmarks demonstrate our TTST performs favorably against state-of-the-art CNN-based and Transformer-based methods, both qualitatively and quantitatively. Yi Xiao 0003, Qiangqiang Yuan |
IGARSS | 1 |
| 2024 | Physics-Informed Multitemporal Ensemble Learning for Near Real-Time Precipitation Estimates From Himawari-8/-9abstractAs a significant element in the water cycle and a key parameter associated with atmospheric circulation, precipitation requires to be fast and accurately monitored. In this study, a new physics-informed multi-temporal model (PMDF) is proposed for the estimates of near real-time high-resolution (0.02°) precipitation, which adopts an ensemble learning framework. The PMDF model can introduce physical knowledge into itself and fast generate precipitation during the estimating phase with no auxiliary data. Validation results show that the PMDF model performs well, with the CC (CSI) of 0.44 (0.49) and 0.73 (0.63) at hourly and daily scales, respectively. The designed physics-informed and multi-temporal strategies both can effectively improve the model accuracy, which yields a significantly better performance than other widely used precipitation products. Furthermore, we can clearly observe the hourly variations of precipitation from estimated results at a high spatial resolution. Yuan Wang 0024, Yi Xiao 0003, Bincheng Wan, Yuanjian Yang, Qiangqiang Yuan, Guofang Wang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Local-Global Temporal Difference Learning for Satellite Video Super-ResolutionabstractOptical-flow-based and kernel-based approaches have been extensively explored for temporal compensation in satellite Video Super-Resolution (VSR). However, these techniques are less generalized in large-scale or complex scenarios, especially in satellite videos. In this paper, we propose to exploit the well-defined temporal difference for efficient and effective temporal compensation. To fully utilize the local and global temporal information within frames, we systematically modeled the short-term and long-term temporal discrepancies since we observe that these discrepancies offer distinct and mutually complementary properties. Specifically, we devise a Short-term Temporal Difference Module (S-TDM) to extract local motion representations from RGB difference maps between adjacent frames, which yields more clues for accurate texture representation. To explore the global dependency in the entire frame sequence, a Long-term Temporal Difference Module (L-TDM) is proposed, where the differences between forward and backward segments are incorporated and activated to guide the modulation of the temporal feature, leading to a holistic global compensation. Moreover, we further propose a Difference Compensation Unit (DCU) to enrich the interaction between the spatial distribution of the target frame and temporal compensated results, which helps maintain spatial consistency while refining the features to avoid misalignment. Rigorous objective and subjective evaluations conducted across five mainstream video satellites demonstrate that our method performs favorably against state-of-the-art approaches. Code will be available athttps://github.com/XY-boy/LGTD. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Xianyu Jin, Liangpei Zhang 0001, Chia-Wen Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | PHTrack: Prompting for Hyperspectral Video TrackingabstractHyperspectral (HS) video captures continuous spectral information of objects, enhancing material identification in tracking tasks. It is expected to overcome the inherent limitations of red-green–blue (RGB) and multimodal tracking, such as finite spectral cues and cumbersome modality alignment. However, HS tracking faces challenges such as data anxiety, bandgaps, and huge volumes. In this study, inspired by prompt learning in language models, we propose the prompting for hyperspectral video tracking (PHTrack) framework. PHTrack learns prompts to adapt foundation models, mitigating data anxiety and enhancing performance and efficiency. First, the modality prompter (MOP) is proposed to capture rich spectral cues and bridge bandgaps for improved model adaptation and knowledge enhancement. In addition, the distillation prompter (DIP) is developed to refine cross-modal features. PHTrack follows feature-level fusion, effectively managing huge volumes compared to traditional decision-level fusion fashions. Extensive experiments validate the proposed framework, offering valuable insights for future research. The code and data will be available athttps://github.com/YZCU/PHTrack Yuzeng Chen, Xin Su 0003, Jie Li 0022, Yi Xiao 0003, Qiangqiang Yuan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | SPIRIT: Spectral Awareness Interaction Network With Dynamic Template for Hyperspectral Object TrackingabstractHyperspectral (HS) video is able to capture abundant spectral, spatial, and temporal information about objects, which overcomes the limitations of common red-green-blue (RGB) video in complex scenarios such as similar appearances and background clutters (BCs). However, most trackers apply hand-crafted features extracted from manually selected bands instead of deep features for object representations due to limited HS data and the band gap problem. Each HS image consists of many bands, and it is challenging to fully interact with the band information while maintaining tracking speed. To this end, this article proposes a novel end-to-end spectral awareness interaction network with a dynamic template (SPIRIT) for HS video object tracking. First, a spectral awareness module (SAM) is proposed to learn band contributions with consideration of nonlinear and global interactions between HS bands. It can also cooperate with the feature extraction module pretrained with RGB data to attenuate the band gap and data-hungry. Second, an interaction module (IM) is proposed to achieve inter and intraband feature interactions to enhance tracking performance while improving efficiency. Furthermore, the proposed method contains a novel update module (UM) that evaluates the tracking confidence of the current state to adapt to object changes and attenuate tracking drifts. Extensive experiments demonstrate the superiority of our approach compared to state-of-the-arts (SOTAs) while meeting real-time demands. Yuzeng Chen, Qiangqiang Yuan, Yi Xiao 0003, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A CNN-Transformer Embedded Unfolding Network for Hyperspectral Image Super-ResolutionabstractHyperspectral images (HSIs) with rich spectral information have been widely used in surface classification, object detection, and other real application problems. However, due to the hardware limitations, the low spatial resolution HSIs hinder the exploration of their application potential. Deep learning-based methods are currently the most common solutions for single HSI super-resolution (HSI SR) tasks. However, such methods often overlook the degradation principle from high-resolution HSI to low-resolution HSI. In this article, we propose a CNN-transformer embedded unfolding network (CTUNet), in which an unfolding framework with an effective spatial-spectral prior network is designed for HSI SR by incorporating the degradation principle of HSIs. Specifically, a maximum posterior-based energy model is employed, enabling alternate optimization to seek the optimal solution in an iterative mechanism. To effectively utilize the structure prior of HSI, multiscale self-calibrated convolution (MSSC) and edge-guided transformer module are combined to learn latent spatial-spectral priors. Additionally, hidden feature connections between adjacent iterations enhance the representation of the image features. Extensive experiments conducted on three available HSI datasets demonstrate that our method outperforms several state-of-the-art HSI SR methods. The code will be available athttps://github.com/YoeTon/CTUNet. Jie Li 0022, Linwei Yue, Xinxin Liu 0002, Yi Xiao 0003, Qiangqiang Yuan |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | EDiffSR: An Efficient Diffusion Probabilistic Model for Remote Sensing Image Super-ResolutionabstractRecently, convolutional networks have achieved remarkable development in remote sensing image (RSI) super-resolution (SR) by minimizing the regression objectives, e.g., MSE loss. However, despite achieving impressive performance, these methods often suffer from poor visual quality with oversmooth issues. Generative adversarial networks (GANs) have the potential to infer intricate details, but they are easy to collapse, resulting in undesirable artifacts. To mitigate these issues, in this article, we first introduce diffusion probabilistic model (DPM) for efficient RSI SR, dubbed efficient diffusion model for RSI SR (EDiffSR). EDiffSR is easy to train and maintains the merits of DPM in generating perceptual-pleasant images. Specifically, different from previous works using heavy UNet for noise prediction, we develop an efficient activation network (EANet) to achieve favorable noise prediction performance by simplified channel attention and simple gate operation, which dramatically reduces the computational budget. Moreover, to introduce more valuable prior knowledge into the proposed EDiffSR, a practical conditional prior enhancement module (CPEM) is developed to help extract an enriched condition. Unlike most DPM-based SR models that directly generate conditions by amplifying LR images, the proposed CPEM helps to retain more informative cues for accurate SR. Extensive experiments on four remote sensing datasets demonstrate that EDiffSR can restore visual-pleasant images on simulated and real-world RSIs, both quantitatively and qualitatively. The code of EDiffSR will be available athttps://github.com/XY-boy/EDiffSR. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Xianyu Jin, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | TTST: A Top-k Token Selective Transformer for Remote Sensing Image Super-ResolutionabstractTransformer-based method has demonstrated promising performance in image super-resolution tasks, due to its long-range and global aggregation capability. However, the existing Transformer brings two critical challenges for applying it in large-area earth observation scenes: (1) redundant token representation due to most irrelevant tokens; (2) single-scale representation which ignores scale correlation modeling of similar ground observation targets. To this end, this paper proposes to adaptively eliminate the interference of irreverent tokens for a more compact self-attention calculation. Specifically, we devise a Residual Token Selective Group (RTSG) to grasp the most crucial token by dynamically selecting the top- k keys in terms of score ranking for each query. For better feature aggregation, a Multi-scale Feed-forward Layer (MFL) is developed to generate an enriched representation of multi-scale feature mixtures during feed-forward process. Moreover, we also proposed a Global Context Attention (GCA) to fully explore the most informative components, thus introducing more inductive bias to the RTSG for an accurate reconstruction. In particular, multiple cascaded RTSGs form our final Top- k Token Selective Transformer (TTST) to achieve progressive representation. Extensive experiments on simulated and real-world remote sensing datasets demonstrate our TTST could perform favorably against state-of-the-art CNN-based and Transformer-based methods, both qualitatively and quantitatively. In brief, TTST outperforms the state-of-the-art approach (HAT-L) in terms of PSNR by 0.14 dB on average, but only accounts for 47.26% and 46.97% of its computational cost and parameters. The code and pre-trained TTST will be available at https://github.com/XY-boy/TTST for validation. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Chia-Wen Lin, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Hyperspectral Image Denoising: From Model-Driven, Data-Driven, to Model-Data-DrivenabstractMixed noise pollution in HSI severely disturbs subsequent interpretations and applications. In this technical review, we first give the noise analysis in different noisy HSIs and conclude crucial points for programming HSI denoising algorithms. Then, a general HSI restoration model is formulated for optimization. Later, we comprehensively review existing HSI denoising methods, from model-driven strategy (nonlocal mean, total variation, sparse representation, low-rank matrix approximation, and low-rank tensor factorization), data-driven strategy [2-D convolutional neural network (CNN), 3-D CNN, hybrid, and unsupervised networks], to model-data-driven strategy. The advantages and disadvantages of each strategy for HSI denoising are summarized and contrasted. Behind this, we present an evaluation of the HSI denoising methods for various noisy HSIs in simulated and real experiments. The classification results of denoised HSIs and execution efficiency are depicted through these HSI denoising methods. Finally, prospects of future HSI denoising methods are listed in this technical review to guide the ongoing road for HSI denoising. The HSI denoising dataset could be found at https://qzhang95.github.io. Qiang Zhang 0011, Yaming Zheng, Qiangqiang Yuan, Meiping Song, Haoyang Yu 0001, Yi Xiao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Learning a Local-Global Alignment Network for Satellite Video Super-ResolutionabstractSatellite video is a novel data source for earth observation, which can be applied in multiple fields for dynamic monitoring. It is always equipped with high temporal resolution at the cost of low spatial resolution of tiny moving objects. Video super-resolution (VSR) is utilized to improve the spatial resolution of satellite video and obtain high spatial-temporal resolution data. However, most existing VSR methods mainly focus on the local interframe information during feature alignment, which lack the ability to model long-distance correspondence. In this letter, a novel two-branch alignment network with an efficient fusion module is proposed for satellite VSR. Both deformable convolution (DCN) and transformer-like attention are employed to fully explore the local and global information between frames. Furthermore, a fusion module is proposed to model the residuals between fusion features and compensate them for better fusion. Experiments on Jilin-1 satellite videos demonstrate that the proposed network can achieve comparable results to current state-of-the-art (SOTA) VSR methods with tiny parameters. Xianyu Jin, Yi Xiao 0003, Qiangqiang Yuan |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Deep Blind Super-Resolution for Satellite VideoabstractRecent efforts have witnessed remarkable progress in Satellite Video Super-Resolution (SVSR). However, most SVSR methods usually assume the degradation is fixed and known,e.g., bicubicdownsampling, which makes them vulnerable in real-world scenes with multiple and unknown degradations. To alleviate this issue, blind SR has thus become a research hotspot. Nevertheless, existing approaches are mainly engaged in blur kernel estimation while losing sight of another critical aspect for VSR tasks: temporal compensation, especially compensating for blurry and smooth pixels with vital sharpness from severely degraded satellite videos. Therefore, this paper proposes a practical Blind SVSR algorithm (BSVSR) to explore more sharp cues by considering the pixel-wise blur levels in a coarse-to-fine manner. Specifically, we employed multi-scale deformable convolution to coarsely aggregate the temporal redundancy into adjacent frames by window-slid progressive fusion. Then the adjacent features are finely merged into mid-feature using deformable attention, which measures the blur levels of pixels and assigns more weights to the informative pixels, thus inspiring the representation of sharpness. Moreover, we devise a pyramid spatial transformation module to adjust the solution space of sharp mid-feature, resulting in flexible feature adaptation in multi-level domains. Quantitative and qualitative evaluations on both simulated and real-world satellite videos demonstrate that our BSVSR performs favorably against state-of-the-art non-blind and blind SR models. Code will be available at https://github.com/XY-boy/Blind-Satellite-VSR. Yi Xiao 0003, Qiangqiang Yuan, Qiang Zhang 0011, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Learning an Intrinsic Graph Neural Network for Sartellite Video Super-ResolutionabstractExisting video super-resolution (VSR) methods usually merge the redundant temporal information along frames to achieve information enhancement, which naturally discards the spatial redundancy information. This paper proposes an intrinsic Graph Neural Network (GNN) framework for satellite VSR to fully explore the internal spatial prior while considering the temporal information in the video frame sequence. Firstly, a Multi-Scale Deformable convolution (MSD) is adopted to accurately model the spatial-temporal relationship between frames. Then, we search for k-nearest neighbors to construct the spatial graph and profoundly excavate the prior spatial information brought by patch recurrence. Finally, the spatial-temporal redundant information is integrated and complementary. Experiments on Jilin-1 satellite video demonstrate the effectiveness of our framework. Yi Xiao 0003, Xin Su 0003, Qiangqiang Yuan |
IGARSS | 1 |
| 2022 | Satellite Video Super-Resolution via Multiscale Deformable Convolution Alignment and Temporal Grouping ProjectionabstractAs a new earth observation tool, satellite video has been widely used in remote-sensing field for dynamic analysis. Video super-resolution (VSR) technique has thus attracted increasing attention due to its improvement to spatial resolution of satellite video. However, the difficulty of remote-sensing image alignment and the low efficiency of spatial–temporal information fusion make poor generalization of the conventional VSR methods applied to satellite videos. In this article, a novel fusion strategy of temporal grouping projection and an accurate alignment module are proposed for satellite VSR. First, we propose a deformable convolution alignment module with a multiscale residual block to alleviate the alignment difficulties caused by scarce motion and various scales of moving objects in remote-sensing images. Second, a temporal grouping projection fusion strategy is proposed, which can reduce the complexity of projection and make the spatial features of reference frames play a continuous guiding role in spatial–temporal information fusion. Finally, a temporal attention module is designed to adaptively learn the different contributions of temporal information extracted from each group. Extensive experiments on Jilin-1 satellite video demonstrate that our method is superior to current state-of-the-art VSR methods. Yi Xiao 0003, Xin Su 0003, Qiangqiang Yuan, Denghong Liu, Huanfeng Shen, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | A Recurrent Refinement Network for Satellite Video Super-ResolutionabstractDeep learning-based methods have shown superior performance in VSR tasks. However, satellite video frames are characterized by large width, low resolution, and lack of features. Consequently, the conventional VSR method is not suitable for satellite video. In this paper, a recurrent refinement network is proposed. Considering that the vast majority of remote sensing images belong to the static background, a single-image SR (SISR) method is first used to obtain high-resolution features for a specific target frame. To further complement the missing details, the network learns the complementary information enhanced by an Encoder-Decoder structure from adjacent frames to refine the results of SISR. To measure the contribution of different adjacent frames to the recovery of the target frame, a temporal attention mechanism is introduced in the final fusion stage. The experiment on the video data of Jilin-1 demonstrates the effectiveness of our method. Yi Xiao 0003, Xin Su 0003, Qiangqiang Yuan |
IGARSS | 1 |