EDBT 2026 Demo / reviewers in the wild / expert
Shuo Li 0010
dblp:49/595-10
· DBLP profile ↗
48ranked-venue papers
14as first author
48since 2021 · last 2026
0000-0003-2002-3894ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 9 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HTTrack: Learning to Perceive Targets via Historical Trajectories in Satellite Video TrackingabstractIn recent years, the rapid progress of deep learning has driven notable advancements in satellite video tracking, a critical task for applications such as environmental monitoring, disaster management, and defense. Despite these strides, existing approaches remain constrained by their inability to handle dynamic challenges, such as target appearance variations, complex motion patterns, and occlusions. Traditional methods often suffer from static template matching or overly complex update mechanisms, compromising their robustness and practicality in real-world scenarios. To address these limitations, we propose a paradigm shift in satellite video tracking by integrating historical trajectory knowledge with visual features. This fusion enhances the tracker's perceptual understanding of targets over time, enabling more adaptive and resilient tracking. By aligning spatial, temporal, and cross-modal information, our approach effectively bridges the gap between fragmented observations and coherent tracking performance, even under challenging conditions like small target detection and cluttered backgrounds. Extensive experiments conducted on multiple satellite video tracking benchmarks demonstrate the superiority of our method, with HTTrack achieving success rates of 51.5% on SV248S, 52.9% on SatSOT, and 32.6% on VISO, significantly outperforming state-of-the-art trackers and marking a step forward in achieving robust, accurate, and scalable satellite video tracking. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
AAAI | 5 |
| 2026 | Semantic Feature Purification for Adversarially-Aware RGB-T TrackingabstractRGB-T tracking is increasingly deployed in safety-critical applications such as autonomous driving, surveillance, and rescue robotics, where tracking reliability is essential under adverse conditions. Although the fusion of RGB and thermal infrared (TIR) modalities offers improved robustness in low-light and occluded scenes, recent findings show that RGB-T trackers remain highly susceptible to subtle input perturbations, human-imperceptible modifications that exploit cross-modal inconsistencies to mislead tracking outputs. In real-world scenarios, such perturbations can arise from sensor spoofing, infrared camouflage, or physical-world attacks, posing serious risks to operational safety. To address this, we propose SFPT, a Semantic Feature Purification framework that enhances RGB-T tracking at the representation level. Rather than filtering corrupted inputs at the pixel level, SFPT introduces task-specific semantic anchors into the feature space to reinforce perturbation-invariant cues. These anchors are derived from descriptive language, interact with visual features to purify representations. To further suppress modality-specific interference, we design an Adaptive Perturbation-Guided Cross-Modal Fusion (APG-CMF) module, which leverages language and visual signals to estimate reliability and dynamically reweight cross-modal features, ensuring robust fusion under perturbation conditions. Extensive experiments under diverse perturbation conditions validate the effectiveness of our approach. Notably, SFPT maintains performance comparable to clean settings even when subjected to perturbations of strength 1/255 and 4/255, demonstrating strong resilience to real-world interference. Jiahao Wang 0002, Fang Liu 0001, Hao Wang 0211, Shuo Li 0010, Puhua Chen |
AAAI | 4 |
| 2026 | Enhancing few-shot segmentation via mask combination learning
Shuo Li 0010, Fang Liu 0034, Licheng Jiao, Xuejian Gou, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
Neurocomputing | 1 |
| 2026 | Text augmentation for vision: Modality-preference aware few-shot learning
Zehua Hao, Fang Liu 0001, Shuo Li 0010, Yaoyang Du, Jiahao Wang 0002, Hao Wang 0211, Licheng Jiao |
Knowl. Based Syst. | 3 |
| 2026 | Concept-Aware Learning for Weakly Supervised Video Anomaly Detection
Shuo Li 0010, Fang Liu 0034, Licheng Jiao, Jiahao Wang 0002, Xu Liu 0006, Lingling Li 0002, Puhua Chen |
Pattern Recognit. | 1 |
| 2026 | PC2F: Language-guided Progressive Calibration and Cascade Filtering for Remote Sensing Visual Grounding
Licheng Jiao, Xu Liu 0006, Xiaoqiang Lu, Shuo Li 0010, Xiaolin Tian 0002 |
Pattern Recognit. | 8 |
| 2026 | VCGPrompt: Visual Concept Graph-Aware Prompt Learning for Vision-Language Models
Mengjia Wang, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
Pattern Recognit. | 4 |
| 2026 | Vision-by-prompt: Context-aware dual prompts for composed video retrieval
Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 5 |
| 2026 | TFBTrack: Target-Aware Foreground-Background Modeling for vision-language tracking
Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 5 |
| 2026 | Language-guided modulation-update for semi-supervised semantic segmentation
Libo Yan, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Jiahao Wang 0002, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Xuejian Gou |
Pattern Recognit. | 4 |
| 2026 | ERFC: Energy-Aware Reinforcement Feedback Calibration for Zero-Shot CaptioningabstractZero-shot captioning aims to generate descriptive captions for unseen image and video data by leveraging the potential of visual language models (VLMs) and language models (LMs) without requiring task-specific training. It has emerged as a critical task, but its performance is often hindered by the inherent gap between the training distribution and unseen test data. The fundamental challenge lies in the model’s strong dependence on the marginal distribution of the training data, which leads to biased predictions when handling test samples. To address this issue, we propose an Energy-aware Reinforcement Feedback Calibration (ERFC) framework to calibrate the distribution and predictions of caption models from a novel energy perspective. The calibration process of ERFC is divided into two key components: 1) We first construct an Energy Stabilizer (ES) based on the caption model, where energy is considered a measure of the affinity between the input sample and the model’s learned distribution. ES iteratively adjusts the embedding features of the input sample using Langevin Dynamics, reducing its energy to implicitly align the model’s distribution with the unseen target domain. 2) We deploy a Reinforcement Calibrator (RC) to refine and calibrate the generated captions through a reward-feedback mechanism. RC leverages the expert CLIP model as a reward signal to assess the quality of the generated captions and employs the policy gradient algorithm to reward or penalize the model, thereby improving its performance. By iteratively combining energy-based optimization and reward-driven calibration, ERFC achieves superior zero-shot generalization capabilities, as demonstrated on image benchmarks such as MSCOCO, Flickr30K, and NoCaps, as well as video benchmarks such as MSR-VTT and MSVD. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Learning to Prompt With Refining Text Knowledge for Zero-Shot Video Action RecognitionabstractFoundational vision-language models (VLMs) like CLIP are redefining the vision domain with their exceptional generalization capabilities. Prompt-based learning methods adapt pre-trained VLMs to video action recognition tasks using task-specific learnable text tokens. However, these tokens often struggle to generalize to unseen categories, as they tend to forget general textual knowledge. To address this, we construct knowledge prompts composed of handcrafted and descriptive prompts and introduce a novel knowledge-guided context mapping to enhance the generalization of learnable prompts to unseen categories. This approach mitigates the forgetting of fundamental knowledge by reducing the discrepancy between learnable prompts and knowledge prompts while simultaneously allowing the prompts to extract rich contextual knowledge from LLM data. Then, incorporating the knowledge-guided context mapping into the contrastive loss enables zero-shot transfer of prompts to new categories and data, providing discriminative prompts for both seen and unseen tasks. In addition, we propose an advanced temporal aggregation method that refines uniform mean pooling by incorporating frame-level textual relevance scoring. Extensive evaluations on multiple benchmarks demonstrate that learning to prompt with refining text knowledge is an effective quick-tuning method, achieving superior sample generalization performance without increasing training parameters. Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
IEEE Trans. Multim. | 5 |
| 2026 | Adaptive Multi-Modal Visual Tracking With Dynamic Semantic PromptsabstractRGB-based object tracking is a fundamental task in computer vision, aiming to identify, locate, and continuously track objects of interest across sequential video frames. Despite the significant advancements in the performance of traditional RGB trackers, they still face challenges in maintaining accuracy and robustness in the presence of complex backgrounds, occlusions, and rapid movements. To tackle these challenges, combining visual auxiliary modalities has gained significant attention. Beyond this, integrating natural language information offers additional advantages by providing high-level semantic context, enhancing robustness, and clarifying target priorities, further elevating tracker performance. This work proposes theAdaptiveMulti-modalVisual Tracking with Dynamic Semantic Prompts (AMVTrack) tracker, which efficiently incorporates image descriptions and avoids text dependency during tracking to improve flexibility and adaptability. AMVTrack significantly reduces computational resource consumption by freezing the parameters of the image encoder, text encoder, and Box Head and only optimizing a few learnable prompt parameters. Additionally, we introduce the Adaptive Dynamic Semantic Prompt Generator (ADSPG), which dynamically generates semantic prompts based on visual features, and theVisual-LanguageFusionAdaptation (V-L FA) method, which integrates multi-modal features to ensure consistency and complementarity of information. Additionally, we partition the Image Encoder to conduct an in-depth investigation into the relationship between the importance of features across different depth and width regions. Experimental results demonstrate that AMVTrack achieves significant performance improvements on multiple benchmark datasets, proving its effectiveness and robustness in complex scenarios. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | Adaptive Visual Prompting for Effective Satellite Video TrackingabstractSatellite video tracking presents significant challenges due to unpredictable target variations, environmental disturbances, and occlusions. Existing approaches either rely on auxiliary modalities or require full fine-tuning of foundation models, resulting in excessive parameter sensitivity and poor generalization. Meanwhile, conventional prompt-based tuning only updates parameters at a single location, limiting its ability to adapt to complex appearance changes. To address these limitations, we propose Adaptive Visual Prompting for Effective Satellite Video Tracking (AVPTrack). Unlike conventional prompts, introduced Super Prompts dynamically refine the original template at multiple distinct positions. This multi-location adaptation allows for fine-grained representation learning, enabling the tracker to better capture target variations and resist environmental disturbances. Additionally, Dynamic Templates are introduced to mitigate tracking failures in highly challenging scenarios, such as occlusions and background clutter, ensuring robust target localization. Furthermore, the Template Selection Adapter (TSA) selects the most relevant templates in real-time, enhancing tracking efficiency. These components are optimized during training while keeping other parameters frozen, ensuring parameter efficiency. We also investigate the relationship between fine-tuning proportions and learning rates to optimize model performance. Extensive evaluations on the SV248S, SatSOT, and VISO datasets demonstrate the superior adaptability and robustness of AVPTrack compared to existing methods. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Yanbiao Ma, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Mengjia Wang |
IEEE Trans. Multim. | 5 |
| 2026 | PromptVAD: Abnormal Prompt via Vision-Language ModelabstractWeakly supervised video anomaly detection (WSVAD) aims at predicting frame-level anomaly scores by modeling training videos with video-level annotations. The category names of abnormal events contain high-level knowledge abstracted by humans about abnormalities, which is of great help in identifying abnormal events. To utilize the knowledge implicit in category names, based on the visual-language pretraining model, we introduce a learnable abnormal prompt from three aspects: learnable domain prompt, learnable category prompt, and nonlearnable category definition prompt. Based on the learnable abnormal prompt, we propose a novel fine-grained WSVAD method: PromptVAD, which exploits a learnable abnormal prompt to reduce the semantic gap between visual images and anomaly categories. Through a similarity measure and our proposed coarse-grained two-class prompt module, our PromptVAD jointly learns coarse-grained and fine-grained VAD. Extensive experimental results on the ShanghaiTech, University of Central Florida (UCF)-Crime, and XD-Violence datasets show that our method achieves state-of-the-art performance. Specifically, our method achieves an area under the curve (AUC) of 88.62% on the UCF-Crime dataset. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Zehua Hao, Jiahao Wang 0002, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Logits DeConfusion with CLIP for Few-Shot LearningabstractWith its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP’s logits suffer from serious inter-class confusion problems in down-stream tasks, and the ambiguity between categories seriously affects the accuracy. To address this challenge, we propose a novel method called Logits DeConfusion, which effectively learns and eliminates inter-class confusion in logits by combining our Multi-level Adapter Fusion (MAF) module with our Inter-Class Deconfusion (ICD) module. Our MAF extracts features from different levels and fuses them uniformly to enhance feature representation. Our ICD learnably eliminates inter-class confusion in logits with a residual structure. Experimental results show that our method can significantly improve the classification performance and alleviate the inter-class confusion problem. The code is available at https://github.com/LiShuo1001/LDC. Shuo Li 0010, Fang Liu 0001, Zehua Hao, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
CVPR | 1 |
| 2025 | Knowledge-Guided Part Segmentation
Xuejian Gou, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Hao Wang 0211, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
ICCV | 4 |
| 2025 | Hierarchical Variational Test-Time Prompt Generation for Zero-Shot Generalization
Zhaoyang Wu, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, LiXu Liu, Puhua Chen, Wenping Ma 0001 |
ICCV | 4 |
| 2025 | FA3T: Feature-Aware Adversarial Attacks for Multi-modal TrackingabstractMulti-modal visual tracking leverages complementary sensor information to enhance robustness under challenging conditions. However, the security of multi-modal tracking systems remains largely unexplored. Existing attacks primarily target single-modal trackers or independently disrupt each modality, failing to exploit the inherent feature interactions and fusion mechanisms that define multi-modal tracking. As a result, these methods exhibit limited attack effectiveness and fail to assess multi-modal tracking systems' vulnerabilities accurately. Understanding these security risks is crucial, as adversarial threats could lead to severe failures in safety-critical applications. To address these challenges, a feature-aware adversarial attack, termed FA3T is proposed. It is designed to explicitly disrupt feature extraction and cross-modal alignment, thereby weakening the fusion process that multi-modal trackers rely on. To achieve this, a Frequency-Spatial Feature Separation (FSFS) module is constructed to perturb feature representations at multiple levels, weakening the modality-complementary advantages of multi-modal tracking. Furthermore, a Target Confusion Attack (TCA) module is devised to manipulate the target-background-template relationships, making it increasingly difficult for the tracker to distinguish the true target, significantly impairing tracking performance. Extensive experiments on five benchmark datasets (i.e., LasHeR, RGBT234, DepthTrack, VOT-RGBD2022, VisEvent) across three different modalities (RGB-T, RGB-D, and RGB-E) demonstrate that our attack substantially degrades state-of-the-art multi-modal trackers, exposing their susceptibility to adversarial threats. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
ACM Multimedia | 5 |
| 2025 | Imagining Vision From Language for Few-Shot Class-Incremental Learning
Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Yanbiao Ma, Puhua Chen, Lingling Li 0002, Xu Liu 0006, Xuejian Gou |
ACM Multimedia | 1 |
| 2025 | Preserving text space integrity for robust compositional zero-shot learning via mixture of pretrained experts
Zehua Hao, Fang Liu 0001, Licheng Jiao, Yaoyang Du, Shuo Li 0010, Hao Wang 0211, Pengfang Li, Xu Liu 0006, Puhua Chen |
Neurocomputing | 5 |
| 2025 | LLM Knowledge-Driven Target Prototype Learning for Few-Shot Segmentation
Pengfang Li, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Xu Liu 0006, Puhua Chen, Lingling Li 0002, Zehua Hao |
Knowl. Based Syst. | 4 |
| 2025 | Visual-Language Scene-Relation-Aware Zero-Shot CaptionerabstractZero-shot image captioning can harness the knowledge of pre-trained visual language models (VLMs) and language models (LMs) to generate captions for target domain images without paired sample training. Existing methods attempt to establish high-quality connections between visual and textual modalities in text-only pre-training tasks. These methods can be divided into two perspectives: sentence-level and entity-level. Although they achieve effective performance on some metrics, they suffer from hallucinations due to biased associations during training. In this paper, we propose a scene-relation-level pre-training task by considering relations as more valuable modal connection bridges. Based on this, we construct a novel Visual-Language Scene Relation Aware Captioner (SRACap), which expands the ability to predict scene relations while generating captions for images. In addition, SRACap possesses excellent cross-domain zero-shot generalization capability, which is driven by a well-designed scene reinforcement switching pipeline. We introduce a scene policy network to dynamically crop salient regions from images and feed them into a language model to generate captions. We integrate multiple expert CLIP models to form a mixture-of-rewards module (MoR) as a reward source, and deeply optimized SRACap through the policy gradient algorithm in the zero-shot inference stage. With the iteration of scene reinforcement switching, SRACap can gradually refine the generated caption details while maintaining high semantic consistency across visual-linguistic modalities. We conduct extensive experiments on multiple standard image captioning benchmarks, showing that SRACap can accurately understand scene structures and generate high-quality text, significantly outperforming other zero-shot inference methods. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Text generation and multi-modal knowledge transfer for few-shot object detection
Yaoyang Du, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Zehua Hao, Pengfang Li, Jiahao Wang 0002, Hao Wang 0211, Xu Liu 0006 |
Pattern Recognit. | 4 |
| 2025 | Knowledge-Driven Compositional Action Recognition
Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006 |
Pattern Recognit. | 5 |
| 2025 | VLPA-CLIP: Video Language Prompting and Adapting CLIP for efficient video action recognition
Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
Pattern Recognit. | 5 |
| 2025 | Prompt-Based Concept Learning for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) faces a huge stability-plasticity challenge due to continuously learning knowledge from new classes with a small number of training samples without forgetting the knowledge of previously seen old classes. To alleviate this challenge, we propose a novel method called Prompt-based Concept Learning (PCL) for FSCIL, which generalizes conceptual knowledge learned from old classes to new classes by simulating human learning capabilities. In our PCL, in the base session, we simultaneously learn common basic concepts from the training data and the class-concept weight of each class in a prompt learning manner, and in each incremental session, class-concept weights between new classes and previously learned basic concepts are learned to achieve incremental learning. Furthermore, in order to avoid catastrophic forgetting, we propose a distribution estimation module to retain feature distributions of previously seen classes and a data replay module to randomly sample features of previously seen classes in incremental sessions. We verify the effectiveness of our PCL on widely used benchmarks, such as miniImageNet, CIFAR-100, and CUB-200. Experimental results show that our PCL achieves competitive results compared with other state-of-the-art methods, especially we achieve an average accuracy of 94.02% across all sessions on the miniImageNet benchmark. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Fine-Grained Visual-Language Alignment for Remote Sensing Image-Text RetrievalabstractRemote sensing image-text retrieval (RSITR) is critical for applications, including environmental monitoring and disaster management. The main challenge in this field is that the multi-scale feature of remote sensing images and the semantic differences of professional texts make it difficult to achieve accurate alignment. Existing coarse-grained methods struggle to address the inherent difference between images and text. In light of this, we propose the Fine-Grained Visual-Language Alignment (FGVLA) method. Our FGVLA employs a hybrid loss function that combines coarse-grained contrastive and triplet loss with novel fine-grained loss. Fine-grained loss includes spatial mask loss and fine-grained contrastive loss to enhance semantic alignment. The method also introduces an inference process that works cooperatively with fine-grained loss to explicitly align image patches with textual nouns. Extensive experiments on RSICD, RSITMD, and UCM-Caption datasets demonstrate that FGVLA outperforms existing methods, achieving superior retrieval performance. The code of our FGVLA has been released at https://github.com/Ji-Haoyang/FGVLA. Shuo Li 0010, Haoyang Ji, Fang Liu 0001, Licheng Jiao, Xutong Min, Jiahao Wang 0002, Lingling Li 0002, Xu Liu 0006 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Change Knowledge-Guided Vision-Language Remote Sensing Change DetectionabstractRemote sensing image change detection plays a critical role in applications like video surveillance and geographic information systems. However, existing binary and semantic change detection methods often rely solely on visual information, neglecting language information, which limits interpretability and the ability to provide specific change details. This work proposes the Change Knowledge-Guided Vision-Language Remote Sensing Change Detection (CKCD) method to address these limitations. By introducing change knowledge as language information, CKCD enhances semantic understanding and change detail representation. A Cross-Modal Affinity (CMA) module is designed to effectively fuse visual and textual features, improving information complementarity and fusion coherence. CKCD further enhances data utilization efficiency by merging change area detection and change category information into a single output through endto- end learning. This design reduces redundant data representations and simplifies the detection process, leading to a more compact and efficient use of the input data without requiring additional branches or multiple output heads. Experimental results demonstrate consistent performance improvements over traditional methods across multiple change detection datasets. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Local-Global Spectral Feature-Aware Learning for Hyperspectral Imagery ClassificationabstractEffective modeling of the relationship between local spectral details and global contextual information remains a core challenge in hyperspectral image (HSI) classification. In this paper, a local-global spectral feature-aware network (LGSFA-Net) is proposed, which achieves local and global spectral feature learning through synergistic integration of local convolutional inductive biases and global state-space models (SSMs). The architecture of LGSFA-Net comprises three sequential components, including an embedding stage using convolutions for fundamental feature extraction, an encoding stage that cascades standard Mamba blocks with specialized interactive Mamba (IMamba) blocks and an enhanced spatial-spectral feature fusion (ESSFF) module. The proposed IMamba blocks employ separable convolutions and feature interaction learning for explicitly modeling the cross-channel spectral correlations learning, which can be effective in awareness of the spectral feature. And then, the ESSFF module utilizes self-attention mechanisms to dynamically balance local and global spatial-spectral feature contributions. The final prediction stage incorporates a lightweight classification head for efficient inference. Experimental results validate the effectiveness of the proposed methods for HSI classification on four benchmark datasets, including the PaviaU, Houston, Honghu, and Hanchuan datasets. The proposed LGSFA-Net achieves approximately 1.48%-2.51% increased overall accuracy (OA), 1.34%-2.06% increased average accuracy(AA), and 1.37%-3.75% increased Kappa on the aforementioned four datasets, respectively, outperforming the contrasting methods. The code implementation will be available at https://github.com/yutinyang/LGSFA-Net. Yuting Yang 0008, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0001, Shuo Li 0010, Wenping Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Anomaly-Led Prompting Learning Caption Generating Model and BenchmarkabstractVideo anomaly detection (VAD) is an important intelligent system application, but most current research views it as a coarse binary classification task that lacks a fine-grained understanding of abnormal video sequences. We explore a new task for video anomaly analysis called Comprehensive Video Anomaly Caption (CVAC), which aims to generate comprehensive textual captions (containing scene information such as time, location, anomalous subject, anomalous behavior, etc.) for surveillance videos. CVAC is more consistent with human understanding than VAD, but it has not been well explored. We constructed a large-scale benchmark CVACBench to lead this research. For each video clip, we provide 6 fine-grained annotations, including scene information and abnormal keywords. A new evaluation metric Abnormal-F1 (A-F1) is also proposed to more accurately evaluate the caption generation performance of the model. We also designed a method called Anomaly-Led Generating Prompting Transformer (AGPFormer) as a baseline. In AGPFormer, we introduce an anomaly-led language modeling mechanism (Anomaly-Led MLM, AMLM) to focus on anomalous events in videos. To achieve more efficient cross-modal semantic understanding, we design the Interactive Generating Prompting (IGP) module and Scene Alignment Prompting (SAP) module to explore the divide between video and text modalities from multiple perspectives, and to improve the model's performance in understanding and reasoning about the complex semantics of videos. We conducted experiments on CVACBench by using traditional caption metrics and the proposed metrics, and the experimental results demonstrate the effectiveness of AGPFormer in the field of anomaly caption. Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Baoliang Chen |
IEEE Trans. Multim. | 5 |
| 2025 | Chain-of-Situation Aware Progressive Inference LearningabstractThe grounded situation recognition (GSR) task aims to recognize the structured semantics of an image to achieve "human-like" event understanding. Most previous studies primarily focus on the visual features of the situation, overlooking the step-by-step cognitive reasoning process that humans employ in complex task settings. Recently, the emergence of multimodal large language models (MLLMs) has provided novel directions for addressing complex problems. However, directly deploying MLLMs on the GSR task is suboptimal due to their tendency to exhibit "hallucination" issues. Additionally, fine-tuning MLLMs for the GSR task incurs high training costs. To address these challenges, inspired by human cognitive theory and the chain-of-thought (CoT) strategy, we propose the chain-of-situation progressive inference learning (CoS-PIL) framework, a lightweight approach that progressively completes verb prediction, noun prediction, and role grounding. The prediction of each step depends on the historical information of the previous step. Specifically, we first design situation prompts tailored to the GSR task and utilize MLLMs to analyze the input image and language prompts, generating heuristic response text for the current situation in the image. Instead of fine-tuning the MLLM, we activate the reasoning capabilities of the frozen MLLM and adapt its generated responses into three lightweight modules: CoS-Verb, CoS-Noun, and CoS-Ground. Considering that MLLMs may generate redundant content, we carefully design the chain-of-interest predictor (CoI-Predictor) to extract key information from the extensive response text and inject it into the model as prompts to enhance the performance. Extensive experiments on the challenging SWiG benchmark demonstrate that CoS-PIL outperforms other state-of-the-art methods. The code is publically available at https://github.com/XDLiuyyy/CoS-PIL. Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationabstractPre-trained vision-language(V-L) models such as CLIP have demonstrated impressive Zero-Shot performance in many downstream tasks. Since adopting contrastive video-text pairs methods like CLIP to video tasks is limited by its high cost and scale, recent approaches focus on efficiently transferring the image-based CLIP to the video domain. A major finding is that fine-tuning the pre-trained model to achieve strong fully supervised performance leads to low zero shot, few shot, and base to novel generalization. Instead, freezing the backbone network to maintain generalization ability weakens fully supervised performance. Otherwise, no single prompt tuning branch consistently performs optimally. In this work, we proposed a multimodal prompt learning scheme that balances supervised and generalized performance. Our prompting approach contains three sections: 1) Independent prompt on both the vision and text branches to learn the language and visual contexts. 2) Inter-modal prompt mapping to ensure mutual synergy. 3) Reducing the discrepancy between the hand-crafted prompt (a video of a person doing [CLS]) and the learnable prompt, to alleviate the forgetting about essential video scenarios. Extensive validation of fully supervised, zero-shot, few-shot, base-to-novel generalization settings for video recognition indicates that the proposed approach achieves competitive performance with less commute cost. Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Zehua Hao, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
AAAI | 6 |
| 2024 | Self-restrained contrastive enhanced network for graph structure learning
Xiaolin Tian 0002, Shuo Li 0010, Licheng Jiao |
Expert Syst. Appl. | 4 |
| 2024 | Satellite Video Object Tracking Based on Location PromptsabstractObject Tracking in satellite videos is a challenging task due to the small target size, low spatial resolution, limited appearance and texture information, and the potential for background confusion. While current state-of-the-art tracking methods perform well on natural images, they often produce unsatisfactory results when applied to satellite videos. In this paper, we address these challenges by leveraging location prompts and refining the feature extractor and bounding box refinement module. Furthermore, we integrate motion features to effectively handle illumination variations that frequently arise in satellite videos, thereby enhancing the overall robustness of the tracker. Our proposed approach, abbreviated as SVLPNet, has been thoroughly evaluated through extensive experiments conducted on two authentic satellite video datasets. The obtained results unequivocally showcase the promising potential of SVLPNet in facilitating object tracking on satellite videos. The source code and raw results will be released at https://github.com/Wprofessor/SVLPNet. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Yingjia Gao, Hao Wang 0211, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Shuo Li 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2024 | Multi-Grained Gradual Inference Model for Multimedia Event ExtractionabstractWith the development of multimedia technology, events are usually presented in multimedia forms, thus multimedia event extraction (MEE) has become more and more important. Existing MEE works usually use simple strategies to align two modalities, making it difficult to precisely extract events and arguments in complex multimedia documents. To address this problem, we propose a novel Multi-grained Gradual Inference Model (MGIM) that focuses on inferring and interpreting events in complex multimedia structures in a coarse-to-fine manner. To efficiently integrate textual and visual modalities, we design a Coarse-grained Alignment (CA) module, which represents the two modalities in a graph structure and performs coarse-grained alignment. Based on the CA module, we further propose a Fine-grained Inference module (FI) that fine-grained aligns text and image by performing multiple rounds of gradual inference. MGIM provides a comprehensive interpretation of multimedia events at two information granularities (coarse and fine). Extensive experiments on the M2E2 dataset demonstrate the effectiveness of MGIM. Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Visual and Language Collaborative Learning for RGBT Object TrackingabstractDespite the extensive research on RGBT object tracking, there are still several challenges and issues in practical applications, such as modality differences, lighting variations and disappearance of the target, and changes in viewpoint. Existing methods mostly address these issues by fusing image features, while neglecting a significant amount of target label information. To address these challenges, this paper introduces text to drive the alignment of visible and infrared image features, transforming features from different modalities into the same feature space and fully using complementary features between different modalities. Furthermore, inspired by the success of prompt learning in various tasks, we utilize prior boxes and language as prompts to further guide the model in tracking the target. Extensive experiments demonstrate that the proposed VLCTrack tracker has excellent potential in RGBT object tracking. Compared to previous methods developed for this purpose, our approach achieves state-of-the-art performance on three benchmark datasets. Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Yingjia Gao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Mask-Guided Correlation Learning for Few-Shot Segmentation in Remote Sensing ImageryabstractFew-shot segmentation aims to segment specific objects in a query image based on a few densely annotated images and has been extensively studied in recent years. In remote sensing, image segmentation faces challenges such as less training data, large intraclass diversity, and low foreground-background contrast. In this work, we propose a novel few-shot segmentation method in remote sensing imagery based on mask-guided correlation learning (MGCL) to alleviate the above challenges. In our MGCL, a novel mask-guided feature enhancement (MGFE) module is proposed, which makes features have intramask consistency by leveraging oversegmented masks. In order to enhance the contrast between foreground and background, a novel foreground-background correlation (FBC) module is proposed, which enhances background correlation representation by learning foreground correlation and background correlation separately. Furthermore, a novel mask-guided correlation decoder (MGCD) module is proposed to guide the decoder to focus on the consistency within the mask, thereby learning how to segment complete objects and improving segmentation accuracy. Sufficient experiments on the iSAID-$5^{i}$and DLRSD-$5^{i}$datasets show that our MGCL outperforms all comparative methods. In particular, in the one-shot setting of the iSAID-$5^{i}$dataset, we achieve an mIoU of 39.92 based on ResNet50, which is an improvement of 4.25 over the state-of-the-art (SOAT) method. The visualization of features before and after the MGFE module further concretely demonstrates the motivation and advantages of our MGCL. The code is available athttps://github.com/LiShuo1001/MGCL. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Xu Liu 0006, Puhua Chen, Lingling Li 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Self-Supervised Self-Organizing Clustering Network: A Novel Unsupervised Representation Learning MethodabstractDeep learning-based clustering methods usually regard feature extraction and feature clustering as two independent steps. In this way, the features of all images need to be extracted before feature clustering, which consumes a lot of calculation. Inspired by the self-organizing map network, a self-supervised self-organizing clustering network ( [Formula: see text]OCNet) is proposed to jointly learn feature extraction and feature clustering, thus realizing a single-stage clustering method. In order to achieve joint learning, we propose a self-organizing clustering header (SOCH), which takes the weight of the self-organizing layer as the cluster centers, and the output of the self-organizing layer as the similarities between the feature and the cluster centers. In order to optimize our network, we first convert the similarities into probabilities which represents a soft cluster assignment, and then we obtain a target for self-supervised learning by transforming the soft cluster assignment into a hard cluster assignment, and finally we jointly optimize backbone and SOCH. By setting different feature dimensions, a Multilayer SOCHs strategy is further proposed by cascading SOCHs. This strategy achieves clustering features in multiple clustering spaces. [Formula: see text]OCNet is evaluated on widely used image classification benchmarks such as Canadian Institute For Advanced Research (CIFAR)-10, CIFAR-100, Self-Taught Learning (STL)-10, and Tiny ImageNet. Experimental results show that our method significant improvement over other related methods. The visualization of features and images shows that our method can achieve good clustering results. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Puhua Chen, Lingling Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Task context transformer and GCN for few-shot learning of cross-domain
Pengfang Li, Fang Liu 0001, Licheng Jiao, Lingling Li 0002, Puhua Chen, Shuo Li 0010 |
Neurocomputing | 6 |
| 2023 | MinEnt: Minimum entropy for self-supervised representation learning
Shuo Li 0010, Fang Liu 0001, Zehua Hao, Licheng Jiao, Xu Liu 0006, Yuwei Guo 0001 |
Pattern Recognit. | 1 |
| 2023 | Knowledge transduction for cross-domain few-shot learning
Pengfang Li, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006 |
Pattern Recognit. | 4 |
| 2023 | Knowledge transfer evolutionary search for lightweight neural architecture with dynamic inference
Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Shuo Li 0010, Puhua Chen, Xu Liu 0006 |
Pattern Recognit. | 6 |
| 2023 | Learning Salient Feature for Salient Object Detection Without LabelsabstractSupervised salient object detection (SOD) methods achieve state-of-the-art performance by relying on human-annotated saliency maps, while unsupervised methods attempt to achieve SOD by not using any annotations. In unsupervised SOD, how to obtain saliency in a completely unsupervised manner is a huge challenge. Existing unsupervised methods usually gain saliency by introducing other handcrafted feature-based saliency methods. In general, the location information of salient objects is included in the feature maps. If the features belonging to salient objects are called salient features and the features that do not belong to salient objects, such as background, are called nonsalient features, by dividing the feature maps into salient features and nonsalient features in an unsupervised way, then the object at the location of the salient feature is the salient object. Based on the above motivation, a novel method called learning salient feature (LSF) is proposed, which achieves unsupervised SOD by LSF from the data itself. This method takes enhancing salient feature and suppressing nonsalient features as the objective. Furthermore, a salient object localization method is proposed to roughly locate objects where the salient feature is located, so as to obtain the salient activation map. Usually, the object in the salient activation map is incomplete and contains a lot of noise. To address this issue, a saliency map update strategy is introduced to gradually remove noise and strengthen boundaries. The visualization of images and their salient activation maps show that our method can effectively learn salient visual objects. Experiments show that we achieve superior unsupervised performance on a series of datasets. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Xu Liu 0006, Puhua Chen |
IEEE Trans. Cybern. | 1 |
| 2022 | Self-Training Multi-Sequence Learning with Transformer for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised Video Anomaly Detection (VAD) using Multi-Instance Learning (MIL) is usually based on the fact that the anomaly score of an abnormal snippet is higher than that of a normal snippet. In the beginning of training, due to the limited accuracy of the model, it is easy to select the wrong abnormal snippet. In order to reduce the probability of selection errors, we first propose a Multi-Sequence Learning (MSL) method and a hinge-based MSL ranking loss that uses a sequence composed of multiple snippets as an optimization unit. We then design a Transformer-based MSL network to learn both video-level anomaly probability and snippet-level anomaly scores. In the inference stage, we propose to use the video-level anomaly probability to suppress the fluctuation of snippet-level anomaly scores. Finally, since VAD needs to predict the snippet-level anomaly scores, by gradually reducing the length of selected sequence, we propose a self-training strategy to gradually refine the anomaly scores. Experimental results show that our method achieves significant improvements on ShanghaiTech, UCF-Crime, and XD-Violence. Shuo Li 0010, Fang Liu 0001, Licheng Jiao |
AAAI | 1 |
| 2022 | Unsupervised Few-Shot Image Classification by Learning Features into Clustering Space
Shuo Li 0010, Fang Liu 0001, Zehua Hao, Kaibo Zhao 0001, Licheng Jiao |
ECCV (31) | 1 |
| 2022 | Augmentative contrastive learning for one-shot object detection
Yaoyang Du, Fang Liu 0001, Licheng Jiao, Zehua Hao, Shuo Li 0010, Xu Liu 0006, Jing Liu 0006 |
Neurocomputing | 5 |
| 2022 | MFNet: A Novel GNN-Based Multi-Level Feature Network With Superpixel PriorsabstractSince the superpixel segmentation method aggregates pixels based on similarity, the boundaries of some superpixels indicate the outline of the object and the superpixels provide prerequisites for learning structural-aware features. It is worthwhile to research how to utilize these superpixel priors effectively. In this work, by constructing the graph within superpixel and the graph among superpixels, we propose a novel Multi-level Feature Network (MFNet) based on graph neural network with the above superpixel priors. In our MFNet, we learn three-level features in a hierarchical way: from pixel-level feature to superpixel-level feature, and then to image-level feature. To solve the problem that the existing methods cannot represent superpixels well, we propose a superpixel representation method based on graph neural network, which takes the graph constructed by a single superpixel as input to extract the feature of the superpixel. To reflect the versatility of our MFNet, we apply it to an image-level prediction task and a pixel-level prediction task by designing different prediction modules. An attention linear classifier prediction module is proposed for image-level prediction tasks, such as image classification. An FC-based superpixel prediction module and a Decoder-based pixel prediction module are proposed for pixel-level prediction tasks, such as salient object detection. Our MFNet achieves competitive results on a number of datasets when compared with related methods. The visualization shows that the object boundaries and outline of the saliency maps predicted by our proposed MFNet are more refined and pay more attention to details. Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Puhua Chen, Xu Liu 0006, Lingling Li 0002 |
IEEE Trans. Image Process. | 1 |