Yan Li 0083

dblp:87/660-83 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-5564-5074ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LaViSE: Language-aware Vision Scale Enhancement for Referring Remote Sensing Image Segmentation
abstract
Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in aerial imagery based on natural language expressions. Although recent multi-scale feature aggregation methods have improved cross-modal alignment, and existing approaches still struggle with accurate localization and segmentation across scales because interactions between visual scales and language are not sufficiently modeled. To address these challenges, we propose a SAM-based framework termed LaViSE, which incorporates two key modules : the Language-Guided Hierarchical Fusion (LGHF) module integrates cross-modal features at multiple scales for precise object localization by injecting spatial coordinates into visual representations and combining visual features with aligned features; the Language-Attentive Scale-Unified Aggregation (LASA) module globally merges multi-level features while maintaining spatial consistency and boundary fidelity, and further preserves fine-grained structural details through and language-conditioned feature recalibration across scales, which is crucial for segmenting densely distributed and scale-varying targets in complex remote sensing imagery. Experiments on the widely used RefSegRS and RRSIS-D benchmarks demonstrate that LaViSE consistently outperforms state-of-the-art methods, particularly in challenging scenarios with small and densely distributed objects. Our code and pre-trained models will be released upon publication.
Yan Li 0083, Zhouchao Fu, Shengjie Yang, Jianwei Zheng 0001
ICMR1
2026 Subspace-frequency regularization for hyperspectral image super-resolution
Chuangjie Fang, Yan Li 0083, Hong Qiu, Honghui Xu 0002, Jianwei Zheng 0001
Knowl. Based Syst.2
2025 Pipeline-Centered Neighboring Network for Deep Unfolding Pansharpening
abstract
Pansharpening technique is dedicated to enriching the spatial details of low-resolution multispectral images (LRMS) under the guidance of a panchromatic (PAN) image. With the guarantee of promising results, Transformer-based methods have enjoyed a high reputation in this field. However, to reduce computational cost, existing solutions typically divide images into smaller, independent windows, which often weakens inter-window and channel-wise interactions as well as leads to unsmooth edges. To address these issues, we first formulate the pansharpening task as a variational optimization problem, and subsequently solve its data and prior subproblems alternately through an unrolling algorithm. In the prior extractor, we propose a Pipeline-Centered Neighboring Attention (PCNA), which holistically allows all pixels to share the same attention span while fully leveraging channel dependencies, thereby significantly improving the capability to process multispectral images. Moreover, a Multi-Scale Channel-Aware (MSCA) module is designed to capture the edges and structural details. Finally, by sequentially integrating the data and prior modules at each iteration stage, we unroll the iterations into a stage-wise unfolding network. Extensive experiments on three satellite datasets demonstrate the effectiveness and efficiency of our proposal compared to cutting-edge methods.
Yan Li 0083, Qiuju Chen, Chuangjie Fang, Ni Xu, Honghui Xu 0002, Jianwei Zheng 0001
ICASSP1
2025 Spatial-Spectral Fusion Neural Operator
abstract
With the rapid development of deep learning, spatial-spectral fusion (SSF) has emerged as an ideal alternative to traditional, costly hyperspectral image (HSI) acquisition methods. However, current solutions necessitate training and storing multiple models for different scaling factors. Besides, a meticulously designed network architecture to meet desirable performance often suffers from a severe computational burden. To counteract the dilemma, we propose SFNO, a lightweight spatial-spectral fusion neural operator for arbitrary-scale SSF. SFNO leverages approximation theory by embedding features from two degraded functions into a high-dimensional latent space, enabling efficient learning of basis functions. Kernel integration mechanisms are then used to approximate certain priors, followed by dimensionality reduction to generate high-resolution HSIs. Moreover, with the aid of discrete invariance property, we propose a new mechanism of progressive resampling (PR), which allows for the shrinkage of necessary spatial domain without any performance degradation. Extensive experiments on CAVE and Harvard datasets show that SFNO and its variant significantly improve performance, especially in out-of-domain fusion, requiring only 0.098M parameters and 0.966G FLOPs.
Wei Li 0034, Jiawei Jiang 0002, Ni Xu, Yan Li 0083, Jianwei Zheng 0001
ICME5
2025 Multi-Scale Core-Peripheral Attention Network for Camouflaged Object Detection
abstract
In recent years, camouflage object detection has remained a significant challenge due to the high similarities between objects and backgrounds. Relying solely on convolutions with limited receptive fields or attentions with fixed ranges is in trouble with handling the size variability of cared objects. Moreover, camouflaged targets are frequently covered by their surroundings, with existing methods prone to erroneously identifying the occluded portions. To break the dilemma, we propose a multi-scale core-peripheral attention network (CPANet), mainly including two elaborations: core-peripheral mask attention (CPMA) and multi-scale weighted fusion (MSWF). CPMA boosts camouflaged features by employing core- and peripheral-based attention mechanisms, mitigating the influence of surrounding obstacles and enabling precise localization of concealed targets. Additionally, MSWF captures multi-scale low-level features to refine local details and manifest complete object representations. Extensive evaluations demonstrate that CPANet outperforms state-of-the-art methods across four widely used benchmarks.
Yueqian Quan, Tiancheng Pan, Chuangjie Fang, Yan Li 0083, Jianwei Zheng 0001
ICME4
2025 Collaborative Cross-Complementary Unfolding Network for Pan-sharpening Remote Sensing Image
abstract
Due to the acquisition limitations of physical devices, pansharpening serves as a computational alternative, enhancing spatial details in low-resolution hyperspectral images with the guidance of corresponding panchromatic images. By leveraging the benefits of nonlinear network architectures and interpretable optimization schemes, deep unfolding networks (DUNs) have shed new light on pansharpening. However, current DUNs lack a dedicated design for both estimating the degradation matrices and extracting intricate information from the proximal operator. To address these challenges, we propose a novel Collaborative Cross-Complementary Unfolding Network (C3U), which is organized into two main steps: customized multi-scale convolution estimation (MSCE) and a data-driven prior extractor. In the MSCE step, the spatial and spectral degradation matrices are individually adapted through multiscale treatment and point convolution operations. Specifically, the overall estimation undergoes an end-to-end iterative block, allowing for adaptive modeling of complex spatial and spectral structures. Within the prior extractor, a cross-complementary attention mechanism is proposed to enable iterative information interaction between global and local Transformers, capturing holistic features and enhancing inductive capacity. Additionally, a collaborative scale-aware-channel mechanism is designed to enlarge the receptive field and capture multiscale channel features in a lightweight manner. More importantly, the principle of collaborative cross-complementary (CCC) permeates all the sub-assemblies, ensuring a desirable information flow. Experimental results on multiple remote sensing datasets demonstrate the superiority of the proposed method over previous state-of-the-art (SOTA) techniques, achieving a 0.8 dB PSNR gain on the GF-2 dataset.
Honghui Xu 0002, Yan Li 0083, Yutao Jia, Chuangjie Fang, Jianwei Zheng 0001
ICMR2
2025 Nonlinear Learnable Triple-Domain Transform Tensor Nuclear Norm for Hyperspectral Image Super-Resolution
abstract
Tensor Nuclear Norm (TNN) has been widely employed as a regularization term for hyperspectral image super-resolution (HSISR). However, conventional TNN constraints based on Discrete Fourier Transform (DFT) often suffer from rank estimation biases and an inability to effectively capture complex spectral-spatial correlations, limiting their efficacy in HSISR. To address these challenges, we propose a Nonlinear Learnable Triple-domain (NLT) transform framework that integrates nonlinear transform, DFT, and self-learning adaptation. This multi-stage process promotes singular value concentration, improving low-rank approximation and rank estimation accuracy. Building upon this framework, we develop an NL-transform-oriented tensor product, a truncated singular value decomposition (TSVD) operation, and a novel tensor nuclear norm (NLTN) tailored for HSISR. By incorporating spectral subspace estimation and clustering-based patch grouping, our approach effectively leverages spatial-spectral correlations and non-local self-similarities, leading to enhanced reconstruction quality. To further mitigate singular value over-penalization, we introduce a logarithmic-based generalized NLTNN (GNLTN) and formulate an optimization strategy based on the alternating direction method of multipliers (ADMM). Extensive experiments demonstrate that our method significantly outperforms existing approaches in terms of fusion accuracy and visual fidelity, setting new benchmarks for hyperspectral image super-resolution. The code is available at https://github.com/xuhonghui96/GNLTN.
Honghui Xu 0002, Yueqian Quan, Chuangjie Fang, Yan Li 0083, Jianwei Zheng 0001
IEEE Trans. Geosci. Remote. Sens.6
2021 A Lightweight Depth Estimation Network for Wide-Baseline Light Fields
abstract
Existing traditional and ConvNet-based methods for light field depth estimation mainly work on the narrow-baseline scenario. This paper explores the feasibility and capability of ConvNets to estimate depth in another promising scenario: wide-baseline light fields. Due to the deficiency of training samples, a large-scale and diverse synthetic wide-baseline dataset with labelled data is introduced for depth prediction tasks. Considering the practical goal for real-world applications, we design an end-to-end trained lightweight convolutional network to infer depths from light fields, called LLF-Net. The proposed LLF-Net is built by incorporating a cost volume which allows variable angular light field inputs and an attention module that enables to recover details at occlusion areas. Evaluations are made on the synthetic and real-world wide-baseline light fields, and experimental results show that the proposed network achieves the best performance when compared to recent state-of-the-art methods. We also evaluate our LLF-Net on narrow-baseline datasets, and it consequently improves the performance of previous methods.
Yan Li 0083, Lu Zhang 0037, Gauthier Lafruit
IEEE Trans. Image Process.1
2020 Manet: Multi-Scale Aggregated Network For Light Field Depth Estimation
abstract
We present a novel end-to-end network, MANet, for light field depth estimation. MANet is a parameter-effective and effi-cient multi-scale aggregated network, which is about 3 times smaller and 3 times faster than the current top-performing method Epinet. The MANet architecture is performed for estimating depth from light field plenoptic cameras, and experimental results show that the proposed MANet outperforms state-of-the-art methods on HCI, CVIA-HCI and EPFL Lytro light field datasets.
Yan Li 0083, Lu Zhang 0037, Gauthier Lafruit
ICASSP1
2020 Overview of deep-learning based methods for salient object detection in videos
Lu Zhang 0037, Yan Li 0083, Kidiyo Kpalma
Pattern Recognit.3