Chenyu Li 0002

dblp:51/2854-2 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-1687-9676ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Any-Optical-Model: A Universal Foundation Model for Optical Remote Sensing
abstract
Optical satellites, with their diverse band layouts and ground sampling distances, supply indispensable evidence for tasks ranging from ecosystem surveillance to emergency response. However, significant discrepancies in band composition and spatial resolution across different optical sensors present major challenges for existing Remote Sensing Foundation Models (RSFMs). These models are typically pretrained on fixed band configurations and resolutions, making them vulnerable to real world scenarios involving missing bands, cross sensor fusion, and unseen spatial scales, thereby limiting their generalization and practical deployment. To address these limitations, we propose Any Optical Model (AOM), a universal RSFM explicitly designed to accommodate arbitrary band compositions, sensor types, and resolution scales. To preserve distinctive spectral characteristics even when bands are missing or newly introduced, AOM introduces a spectrum-independent tokenizer that assigns each channel a dedicated band embedding, enabling explicit encoding of spectral identity. To effectively capture texture and contextual patterns from sub-meter to hundred-meter imagery, we design a multi-scale adaptive patch embedding mechanism that dynamically modulates the receptive field. Furthermore, to maintain global semantic consistency across varying resolutions, AOM incorporates a multi-scale semantic alignment mechanism alongside a channel-wise self-supervised masking and reconstruction pretraining strategy that jointly models spectral-spatial relationships. Extensive experiments on over 10 public datasets, including those from Sentinel-2, Landsat, and HLS, demonstrate that AOM consistently achieves state-of-the-art (SOTA) performance under challenging conditions such as band missing, cross sensor, and cross resolution settings. These results highlight AOM as a crucial step toward building truly general-purpose RSFMs.
Chenyu Li 0002, Danfeng Hong
AAAI2
2025 A comprehensive survey for Hyperspectral Image Classification: The evolution from conventional to transformers and Mamba models
Muhammad Ahmad 0002, Salvatore Distefano, Adil Khan 0001, Manuel Mazzara, Chenyu Li 0002, Hao Li 0019, Jagannath Aryal, Yao Ding 0010, Gemine Vivone, Danfeng Hong
Neurocomputing5
2025 Hyperspectral Image Classification With Mamba
abstract
Local and global spectral and spatial information is crucial for hyperspectral image (HSI) classification. However, modeling the global context has been challenging due to the limitations of receptive fields and quadratic complexity. Mamba’s ability to leverage long-range dependencies with linear computational complexity offers an effective approach to alleviate this issue; however, it does lead to the loss of local detail information. To address this challenge, we propose a novel local-to-global Mamba for HSI classification, termed MambaLG. MambaLG consists of a dual-branch strategy, comprising two core modules: a local and global spatial modeling module (SpaM) and a short- and long-range spectral dynamic perception module (SpeM). In the SpaM, the local and global spatial information is sequentially extracted and integrated, aiming to capture global spatial semantics while preserving the integrity of local 2-D spatial structures. In the SpeM, we utilize local spectral extraction, spectral grouping, and spectral dynamic correlation clustering (SDCC) modules, leveraging Mamba’s strengths in exploring long-range dependencies for more precise short- and long-range spectral feature modeling. Additionally, we introduce a gate attention unit into MambaLG and design a more efficient and interpretable manner for merging spatial and spectral features. Experimental results across multiple datasets (encompassing urban and agricultural scenes) indicate that MambaLG surpasses state-of-the-art algorithms regarding classification accuracy (CA) and inference speed. Comprehensive ablation studies substantiate the advantages of MambaLG in modeling local and global spatial context, enhancing short- and long-range spectral perception, and fusing spatial and spectral information. The codes will be openly available athttps://github.com/danfenghong/IEEE_TGRS_MambaLGto facilitate the reproduction of experimental results.
Zhaojie Pan, Chenyu Li 0002, Antonio Plaza, Jocelyn Chanussot, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.2
2025 Learning Disentangled Priors for Hyperspectral Anomaly Detection: A Coupling Model-Driven and Data-Driven Paradigm
abstract
Accurately distinguishing between background and anomalous objects within hyperspectral images poses a significant challenge. The primary obstacle lies in the inadequate modeling of prior knowledge, leading to a performance bottleneck in hyperspectral anomaly detection (HAD). In response to this challenge, we put forth a groundbreaking coupling paradigm that combines model-driven low-rank representation (LRR) methods with data-driven deep learning techniques by learning disentangled priors (LDP). LDP seeks to capture complete priors for effectively modeling the background, thereby extracting anomalies from hyperspectral images more accurately. LDP follows a model-driven deep unfolding architecture, where the prior knowledge is separated into the explicit low-rank prior formulated by expert knowledge and implicit learnable priors by means of deep networks. The internal relationships between explicit and implicit priors within LDP are elegantly modeled through a skip residual connection. Furthermore, we provide a mathematical proof of the convergence of our proposed model. Our experiments, conducted on multiple widely recognized datasets, demonstrate that LDP surpasses most of the current advanced HAD techniques, exceling in both detection performance and generalization capability.
Chenyu Li 0002, Bing Zhang 0001, Danfeng Hong, Xiuping Jia, Antonio Plaza, Jocelyn Chanussot
IEEE Trans. Neural Networks Learn. Syst.1
2024 SpectralGPT: Spectral Remote Sensing Foundation Model
abstract
The foundation model has recently garnered significant attention due to its potential to revolutionize the field of visual representation learning in a self-supervised manner. While most foundation models are tailored to effectively process RGB images for various visual tasks, there is a noticeable gap in research focused on spectral data, which offers valuable information for scene understanding, especially in remote sensing (RS) applications. To fill this gap, we created for the first time a universal RS foundation model, named SpectralGPT, which is purpose-built to handle spectral RS images using a novel 3D generative pretrained transformer (GPT). Compared to existing foundation models, SpectralGPT 1) accommodates input images with varying sizes, resolutions, time series, and regions in a progressive training fashion, enabling full utilization of extensive RS Big Data; 2) leverages 3D token generation for spatial-spectral coupling; 3) captures spectrally sequential patterns via multi-target reconstruction; and 4) trains on one million spectral RS images, yielding models with over 600 million parameters. Our evaluation highlights significant performance improvements with pretrained SpectralGPT models, signifying substantial potential in advancing spectral RS Big Data applications within the field of geoscience across four downstream tasks: single/multi-label scene classification, semantic segmentation, and change detection.
Danfeng Hong, Bing Zhang 0001, Chenyu Li 0002, Jing Yao 0002, Naoto Yokoya, Hao Li 0019, Pedram Ghamisi, Xiuping Jia, Antonio Plaza, Paolo Gamba, Jón Atli Benediktsson, Jocelyn Chanussot
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Interpretable Networks for Hyperspectral Anomaly Detection: A Deep Unfolding Solution
abstract
Current hyperspectral anomaly detection (HAD) benchmark datasets suffer from low resolution, simple background, and small size of the anomalies. These factors also limit the performance of the well-known low-rank representation (LRR) models in terms of robustness on the separation of background and target features and the reliance on manual parameter selection. To this end, we build a new HAD benchmark dataset for improving the robustness in complex scenarios, AIR-HAD for short, and propose an interpretable network with deep unfolding a binary subspace learning, named LRR-Net+, which is capable of spectrally decoupling the background structure and object properties in a more generalized fashion and eliminating the bias introduced by vital interference targets simultaneously. In addition, LRR-Net+ integrates the solution process of the alternating direction method of multipliers (ADMM) optimizer with the deep network, guiding its search process and imparting a level of interpretability to parameter optimization. Additionally, the integration of physical models with DL techniques eliminates the need for manual parameter tuning. The manually tuned parameters are seamlessly transformed into trainable parameters for deep neural networks, facilitating a more efficient and automated optimization process. Extensive experiments conducted on the AIR-HAD dataset show the superiority of our LRR-Net+ in terms of detection performance and generalization ability, compared to top-performing competitors. Furthermore, our AIR-HAD benchmark datasets will be made available freely and openly athttps://github.com/danfenghong/IEEE_TGRS_LRR-Net.
Chenyu Li 0002, Bing Zhang 0001, Danfeng Hong, Jing Yao 0002, Xiuping Jia, Antonio Plaza, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2024 HyperSINet: A Synergetic Interaction Network Combined With Convolution and Transformer for Hyperspectral Image Classification
abstract
In hyperspectral images (HSIs), both local and non-local features play crucial roles in classification tasks. Vision Transformer (VIT) can extract non-local features through attention mechanisms, while Convolutional Neural Networks (CNN) excel at handling local components. However, in traditional dual-branch models based on VIT and CNN, there is a lack of interaction during feature processing, leading to potential compatibility issues when merging the two types of features. In this article, we propose HyperSINet, a Synergetic Interaction Network that combines VIT and CNN to establish interaction between the two branches, enabling mutual compensation between local and non-local features during the training process and ultimately enhancing the performance of classification tasks. Specifically, we devise a pair of interactors, namely Conv2Trans and Trans2Conv, which serve as intermediaries between the two branches, enabling the VIT branch to refine its local details, while allowing the CNN branch to process larger receptive field non-local features. Typical feature Maps are implemented to visualize the function of the interactors. Furthermore, within the VIT branch, a VIT Encoder with the local mask is developed to strike a balance between emphasizing non-local features and preserving local details, while a lightweight CNN block is designed to process spectral and spatial features in the CNN branch. Extensive experiments conducted on four real-world datasets demonstrate that, under a reasonable count of parameters, HyperSINet surpasses several current state-of-the-art methods.
Qixing Yu, Weibo Wei, Dantong Li, Zhenkuan Pan 0001, Chenyu Li 0002, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.5
2023 Decoupled-and-Coupled Networks: Self-Supervised Hyperspectral Image Super-Resolution With Subpixel Fusion
abstract
Enormous efforts have been recently made to super-resolve hyperspectral (HS) images with the aid of high spatial resolution multispectral (MS) images. Most prior works usually perform the fusion task by means of multifarious pixel-level priors. Yet the intrinsic effects of a large distribution gap between HS-MS data due to differences in the spatial and spectral resolution are less investigated. The gap might be caused by unknown sensor-specific properties or highly-mixed spectral information within one pixel (due to low spatial resolution). To this end, we propose a subpixel-level HS super-resolution framework by devising a novel decoupled-and-coupled network, called DC-Net, to progressively fuse HS-MS information from the pixel- to subpixel-level, from the image- to feature-level. As the name suggests, DC-Net first decouples the input into common (or cross-sensor) and sensor-specific components to eliminate the gap between HS-MS images before further fusion, and then thoroughly blends them by a model-guided coupled spectral unmixing (CSU) net. More significantly, we append a self-supervised learning module behind the CSU net by guaranteeing material consistency to enhance the detailed appearance of the restored HS product. Extensive experimental results show the superiority of our method both visually and quantitatively and achieve a significant improvement in comparison with the state-of-the-art.
Danfeng Hong, Jing Yao 0002, Chenyu Li 0002, Deyu Meng, Naoto Yokoya, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.3
2023 LRR-Net: An Interpretable Deep Unfolding Network for Hyperspectral Anomaly Detection
abstract
Considerable endeavors have been expended towards enhancing the representation performance for Hyperspectral Anomaly Detection (HAD) through physical model-based methods and recent deep learning-based approaches. Of these methods, the Low-Rank Representation (LRR) model is widely adopted for its formidable separation capabilities for background and target features, however, its practical applications are limited due to the reliance on manual parameter selection and subpar generalization performance. To this end, this paper presents a new HAD baseline network, referred to as LRR-Net, which synergizes the LRR model with deep learning techniques. LRR-Net leverages the alternating direction method of multipliers (ADMM) optimizer to solve the LRR model efficiently and incorporates the solution as prior knowledge into the deep network to guide the optimization of parameters. Moreover, LRR-Net transforms the regularized parameters into trainable parameters of the deep neural network, thus alleviating the need for manual parameter tuning. Additionally, this paper proposes a sparse neural network embedding to demonstrate the scalability of the LRR-Net framework. Empirical evaluations on eight distinct datasets illustrate the efficacy and superiority of the proposed approach compared to state-of-the-art methods.
Chenyu Li 0002, Bing Zhang 0001, Danfeng Hong, Jing Yao 0002, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2023 Extended Vision Transformer (ExViT) for Land Use and Land Cover Classification: A Multimodal Deep Learning Framework
abstract
The recent success of attention mechanism-driven deep models, like Vision Transformer (ViT) as one of the most representative, has intrigued a wave of advanced research to explore their adaptation to broader domains. However, current Transformer-based approaches in the remote sensing (RS) community pay more attention to single-modality data, which might lose expandability in making full use of the ever-growing multimodal Earth observation data. To this end, we propose a novel multimodal deep learning framework by extending conventional ViT with minimal modifications, abbreviated as ExViT, aiming at the task of land use and land cover classification. Unlike common stems that adopt either linear patch projection or deep regional embedder, our approach processes multimodal RS image patches with parallel branches of position-shared ViTs extended with separable convolution modules, which offers an economical solution to leverage both spatial and modality-specific channel information. Furthermore, to promote information exchange across heterogeneous modalities, their tokenized embeddings are then fused through a cross-modality attention module by exploiting pixel-level spatial correlation in RS scenes. Both of these modifications significantly improve the discriminative ability of classification tokens in each modality and thus further performance increase can be finally attained by a full tokens-based decision-level fusion module. We conduct extensive experiments on two multimodal RS benchmark datasets, i.e., the Houston2013 dataset containing hyperspectral and light detection and ranging (LiDAR) data, and Berlin dataset with hyperspectral and synthetic aperture radar (SAR) data, to demonstrate that our ExViT outperforms concurrent competitors based on Transformer or convolutional neural network (CNN) backbones, in addition to several competitive machine learning-based models. The source codes and investigated datasets of this work will be made publicly available at https://github.com/jingyao16/ExViT.
Jing Yao 0002, Bing Zhang 0001, Chenyu Li 0002, Danfeng Hong, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.3