Ke Li 0024

dblp:75/6627-24 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0003-2873-7795ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images
abstract
Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing images based on free-form natural language expressions. Existing approaches are typically constrained to closed-set vocabularies, limiting their applicability in open-world scenarios. While recent attempts to leverage generic foundation models for open-vocabulary RSVG, they overly rely on expensive high-quality datasets and time-consuming fine-tuning. To address these limitations, we propose RSVG-ZeroOV, a training-free framework that aims to explore the potential of frozen generic foundation models for zero-shot open-vocabulary RSVG. Specifically, RSVG-ZeroOV comprises three key stages: (i) Overview: We utilize a vision-language model (VLM) to obtain cross-attention maps that capture semantic correlations between text queries and visual regions. (ii) Focus: By leveraging the fine-grained modeling priors of a diffusion model (DM), we fill in gaps in structural and shape information of objects, which are often overlooked by VLM. (iii) Evolve: A simple yet effective attention evolution module is introduced to suppress irrelevant activations, yielding purified segmentation masks over the referred objects. Without cumbersome task-specific training, RSVG-ZeroOV offers an efficient and scalable solution. Extensive experiments demonstrate that the proposed framework consistently outperforms existing weakly-supervised and zero-shot methods.
Ke Li 0024, Di Wang 0011, Fuyu Dong, Quan Wang 0006
AAAI1
2025 FD2-Net: Frequency-Driven Feature Decomposition Network for Infrared-Visible Object Detection
abstract
Infrared-visible object detection (IVOD) seeks to harness the complementary information in infrared and visible images, thereby enhancing the performance of detectors in complex environments. However, existing methods often neglect the frequency characteristics of complementary information, such as the abundant high-frequency details in visible images and the valuable low-frequency thermal information in infrared images, thus constraining detection performance. To solve this problem, we introduce a novel Frequency-Driven Feature Decomposition Network for IVOD, called FD2-Net, which effectively captures the unique frequency representations of complementary information across multimodal visual spaces. Specifically, we propose a feature decomposition encoder, wherein the high-frequency unit (HFU) utilizes discrete cosine transform to capture representative high-frequency features, while the low-frequency unit (LFU) employs dynamic receptive fields to model the multi-scale context of diverse objects. Next, we adopt a parameter-free complementary strengths strategy to enhance multimodal features through seamless inter-frequency recoupling. Furthermore, we innovatively design a multimodal reconstruction mechanism that recovers image details lost during feature extraction, further leveraging the complementary information from infrared and visible images to enhance overall representational capacity. Extensive experiments demonstrate that FD2-Net outperforms state-of-the-art (SoTA) models across various IVOD benchmarks, i.e. LLVIP (96.2% mAP), FLIR (82.9% mAP), and M3FD (83.5% mAP).
Ke Li 0024, Di Wang 0011, Zhangyuan Hu, Weiping Ni, Lin Zhao 0003, Quan Wang 0006
AAAI1
2025 A Reinforcement Learning Framework for Efficient Task Allocation Among AGVs in Smart Warehouse
abstract
In smart warehouses that use automated guided vehicles (AGVs) for goods transportation, task allocation has a great impact on operational efficiency. Currently, warehouse task allocation is typically modeled as a pickup and delivery problem (PDP), which requires vehicles to start and return from the same depot to construct several closed-loop routes. This approach increases the vehicle travel distance without load in high-throughput warehouses and results in resource wastage. Thus, we remodel the task allocation problem as an open-loop routing problem with heterogeneous starting points and name it capacitied multiagent open PDP (CMOPDP), which has more complex solution space and constraints than PDP. The solving speed of existing heuristic methods cannot meet the real-time processing demands of large-scale warehouses. And deep reinforcement learning (DRL)-based methods typically satisfy constraints through the output mask of decoders, which leads to unsatisfactory quality of solutions under complex constraints. To address these limitations, we design an DRL-based model with encoder-decoder architecture to solve the CMOPDP. Specifically, first, an encoder with heterogeneous attention is designed to fully explore constraint relationships between nodes. Second, we utilize dual decoders and information sharing to maximize vehicle-customer nodes matching. Finally, entropy rewards are introduced to enhance exploration during reinforcement learning, preventing the model from getting stuck in local optima. Extensive experiments on random datasets and various warehouse maps demonstrate that our method improves solution quality by at least 1.76% over baselines, while maintaining competitive solving time and exhibiting good generalization performance.
Zejian Zhao, Di Wang 0011, Ke Li 0024, Gang Liu 0006, Quan Wang 0006
IEEE Internet Things J.4
2025 Visual grounding of remote sensing images with multi-dimensional semantic-guidance
Yueli Ding, Di Wang 0011, Ke Li 0024, Xiaohong Zhao, Yifeng Wang 0004
Pattern Recognit. Lett.3
2025 Physical Adversarial Patch Attack for Optical Fine-Grained Aircraft Recognition
abstract
Deep neural networks (DNNs) have been widely used in remote sensing but demonstrated to be sensitive with adversarial examples. By introducing carefully designed perturbations to clean images, DNNs can be led to incorrect predictions. Adversarial patch is commonly used to conduct adversarial attack, where traditional methods optimize its content and position separately, neglecting the coupling relation of two factors. In this paper, we propose a black-box attack framework targeting fine-grained aircraft recognition, named PatchGen, simultaneously optimizing both content and position of physical adversarial patches. For the requirements of physical attack, we further constrain the patch in object region and utilize elaborate criteria to evaluate its naturalness to alleviate the distortion when applying the patch in real world. We comprehensively validate our method in fine-grained aircraft classification, extending to object detection subsequently. Extensive experiments demonstrate that the proposed method achieves superior attack performance efficiently for classification and detection tasks in digital domain. Moreover, we validate the effectiveness of the adversarial patch under diverse circumstances in the physical world and prove that our method can be applied to different models as well as various domains.
Ke Li 0024, Di Wang 0011, Wenxuan Zhu, Quan Wang 0006, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.1
2024 Unleashing Channel Potential: Space-Frequency Selection Convolution for SAR Object Detection
abstract
Deep Convolutional Neural Networks (DCNNs) have achieved remarkable performance in synthetic aperture radar (SAR) object detection, but this comes at the cost of tremendous computational resources, partly due to extracting redundant features within a single convolutional layer. Recent works either delve into model compression methods or focus on the carefully-designed lightweight models, both of which result in performance degradation. In this paper, we propose an efficient convolution module for SAR object detection, called SFS-Conv, which increases feature diversity within each convolutional layer through a shunt-perceive-select strategy. Specifically, we shunt input feature maps into space and frequency aspects. The former perceives the context of various objects by dynamically adjusting receptive field, while the latter captures abundant frequency variations and textural features via fractional Gabor transformer. To adaptively fuse features from space and frequency aspects, a parameter-free feature selection module is proposed to ensure that the most representative and distinctive information are preserved. With SFS-Conv, we build a lightweight SAR object detection network, called SFS-CNet. Experimental results show that SFS-CNet outperforms state-of-the-art (SoTA) models on a series of SAR object detection benchmarks, while simultaneously reducing both the model size and computational cost.
Ke Li 0024, Di Wang 0011, Zhangyuan Hu, Wenxuan Zhu, Quan Wang 0006
CVPR1
2024 Alignment and Multimodal Reasoning for Remote Sensing Visual Question Answering
abstract
Recently, visual question answering for remote sensing data (RSVQA) has emerged as a prominent research area in the field of remote sensing. Transformer-based approaches have demonstrated impressive results, attributed to their superior performance in jointly modeling visual and textual modalities. However, existing Remote Sensing Visual Question Answering (RSVQA) methods often overlook the modality biases present in visual-language interactions, leading to in-accuracies in answers. To address this issue, we propose a novel Transformer-based approach aimed at mitigating modality biases in RSVQA. Specifically, we introduce a contrastive learning loss to align image and text representations before cross-modal fusion, facilitating foundational learning of visual and language representations. Subsequently, we design a cross-modal decoder to comprehensively understand the correlations between images and text. Notably, in addition to predicting answers to questions, we incorporate an extra head for regression prediction of question types. Experimental results demonstrate that our approach achieves higher accuracy in answer prediction compared to state-of-the-art (SoTA) methods, establishing a new record.
Yumin Tian, Di Wang 0011, Ke Li 0024, Lin Zhao 0003
IGARSS4
2024 Kernel-Adaptive Change Detection Network in Remote Sensing Imagery
abstract
Effective representation of features at multiple scales is crucial for remote sensing change detection (RSCD). The latest advancements in Convolutional Neural Networks (CNNs) consistently demonstrate enhanced multiscale representation capabilities, leading to improved performance in RSCD. However, existing multiscale feature extraction methods often require additional module designs, resulting in higher model parameters and computation costs. In this paper, we propose a CNN building block called Kernel-Adaptive (KA) convolution, which utilizes spatial attention generated from different scales to seamlessly integrate effective receptive fields of various sizes within a single network layer. By stacking multiple KA blocks, we construct a lightweight deep network named Kernel-Adaptive Change Detection Network (KANet). Our experiments on widely used datasets, such as LEVIR-CD and CDD, demonstrate that KANet outperforms existing state-of-the-art (SoTA) methods with fewer parameters and FLOPs. Further ablation studies validate the superior multiscale perception capability of KANet compared to existing RSCD methods.
Di Wang 0011, Fuyu Dong, Ke Li 0024
IGARSS3
2024 Transferable Physical Adversarial Patch Attack for Remote Sensing Object Detection
abstract
Deep neural networks (DNNs) have been widely used in remote sensing but demonstrated to be vulnerable with adversarial examples. By adding elaborately designed perturbations on the clean images, DNNs may output wrong prediction. Research on adversarial attack contributes to the study of model robustness. However, previous methods mainly focus on white-box scenario or digital domain for classification tasks, while the vulnerability of remote sensing detectors has not been fully explored. Aiming at attacking black-box remote sensing detectors in physical domain, we propose to generate a transferable physical adversarial patch (TPAP) as the perturbations. Specifically, the initial patch is optimized by a U-Net and modified by the plane mask and position mask before applied to the clean image. By attacking a surrogate model, TPAP can be transferred to the target model. Abundant experimental results validate the attack ability of TPAP and evaluate the robustness of current one-stage detectors.
Di Wang 0011, Wenxuan Zhu, Ke Li 0024, Pengfei Yang 0001
IGARSS3
2024 Multi-object behavior recognition based on object detection for dense crowds
Min Dang, Gang Liu 0006, Qijie Xu, Ke Li 0024, Di Wang 0011, Lihuo He
Expert Syst. Appl.4
2024 Visual Selection and Multistage Reasoning for RSVG
abstract
Visual grounding of remote sensing (RSVG) is a task to locate targets indicated by referring expressions in remote sensing (RS) images. Previous approaches directly concatenate visual and language features, and stack a series of transformer encoders for cross-modal fusion. However, this fusion strategy fails to fully leverage attributes and contextual information of the targets in referring expressions, limiting the performance of existing methods. To address this issue, we propose a novel visual grounding framework for RSVG, named VSMR, which achieves accurate localization by adaptively selecting target-relevant features and performing multi-stage cross-modal reasoning. Specifically, we propose an Adaptive Feature Selection (AFS) module, which automatically selects visual features relevant to queries while suppressing background noises. A Multi-Stage Decoder (MSD) is designed to iteratively infer correlations between images and queries by leveraging abundant object attributes and contextual information in the referring expressions, thereby achieving accurate target localization. Experiments demonstrate our method is superior to other state-of-the-art (SoTA) methods, achieving accuracy of 78.24%.
Yueli Ding, Di Wang 0011, Ke Li 0024, Yumin Tian
IEEE Geosci. Remote. Sens. Lett.4
2024 DiagSWin: A multi-scale vision transformer with diagonal-shaped windows for object detection and segmentation
Ke Li 0024, Di Wang 0011, Gang Liu 0006, Wenxuan Zhu, Haodi Zhong, Quan Wang 0006
Neural Networks1
2024 GR-GAN: A unified adversarial framework for single image glare removal and denoising
Cong Niu, Ke Li 0024, Di Wang 0011, Wenxuan Zhu, Jinhui Dong
Pattern Recognit.2
2024 Language-Guided Progressive Attention for Visual Grounding in Remote Sensing Images
abstract
Visual grounding in remote sensing (RSVG) images aims to detect specific objects associated with referring expressions in remote sensing images. Existing methods typically combine outputs of pretrained visual and linguistic backbones to locate referred objects. However, due to the lack of interaction with the language modality during the visual feature extraction process, the visual backbone may suffer from attention drift, limiting RSVG’s performance. To avoid this, we propose a novel RSVG framework, namely, language-guided progressive visual attention (LPVA), which achieves precise attention on referred objects by adjusting visual features with a progressive attention (PA) module and a multilevel feature enhancement (MFE) decoder. Specifically, the former can dynamically generate multiscale weights and biases, enabling the visual backbone to gradually focus on expression-related features at spatial and channel levels. The latter is designed to aggregate visual contextual information of the referred objects to enhance features’ distinctiveness while simultaneously suppressing information of irrelevant regions. To thoroughly examine the localization capability of RSVG models, we construct a new large-scale benchmark dataset, namely, OPT-RSVG, which poses challenges in comprehensive understanding among complex scenarios. Experimental results show that the proposed method pushes the accuracy score to 82.27% (6.29% absolute improvement) on the DIOR-RSVG dataset and 78.03% on the OPT-RSVG dataset, thus setting new records. The source codes of the proposed method and OPT-RSVG dataset are available athttps://github.com/like413/OPT-RSVG.
Ke Li 0024, Di Wang 0011, Haodi Zhong, Cong Wang 0033
IEEE Trans. Geosci. Remote. Sens.1
2023 Mixing Self-Attention and Convolution: A Unified Framework for Multisource Remote Sensing Data Classification
abstract
Convolution and self-attention are two powerful techniques for multi-source remote sensing (RS) data fusion that have been widely adopted in Earth observation tasks. However, Convolutional Neural Networks (CNNs) are inadequate for fully mining contextual information and representing the sequence attributes of spectral signatures. Additionally, the specific self-attention mechanism often comes with high computation costs, which hinders its application in the field of RS. To overcome the above limitations, this paper proposes a unified framework called “Mixing Self-Attention and Convolution Network" for comprehensive feature extraction and efficient feature fusion. First, the proposed MACN utilizes two adaptive CNN encoders (ACEs) to extract shallow convolutional features from multi-source RS data. Secondly, taking the complexity and varying scales of RS data into account, the proposed mixing self-attention and convolution transformer (MACT) layer achieves local and global multiscale perception through an elegant integration of self-attention and convolution. MACT can extract abundant spatial and high-dimensional information (e.g., spectral and elevation information) while maintaining minimal computational overhead compared to pure convolution or self-attention counterparts. Finally, a multi-source cross-guided fusion (MCGF) module is designed to achieve deep fusion of multi-source RS data features. MCGF utilizes a carefully designed cross-modal attention mechanism to capture the interaction between multi-source data and aggregate contextual information. Extensive tests on six public RS datasets have shown that our method outperforms other multi-source fusion models, delivering state-of-the-art results on multiple RS data fusion tasks without specific tuning. The source code of the proposed method will be available publicly at https://github.com/like413/MACN.
Ke Li 0024, Di Wang 0011, Xu Wang 0057, Gang Liu 0006, Zili Wu, Quan Wang 0006
IEEE Trans. Geosci. Remote. Sens.1