Lunbo Li

dblp:188/6062 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0003-0834-2075ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FAMNet: Frequency-aware Matching Network for Cross-domain Few-shot Medical Image Segmentation
abstract
Existing few-shot medical image segmentation (FSMIS) models fail to address a practical issue in medical imaging: the domain shift caused by different imaging techniques, which limits the applicability to current FSMIS tasks. To overcome this limitation, we focus on the cross-domain few-shot medical image segmentation (CD-FSMIS) task, aiming to develop a generalized model capable of adapting to a broader range of medical image segmentation scenarios with limited labeled data from the novel target domain. Inspired by the characteristics of frequency domain similarity across different domains, we propose a Frequency-aware Matching Network (FAMNet), which includes two key components: a Frequency-aware Matching (FAM) module and a Multi-Spectral Fusion (MSF) module. The FAM module tackles two problems during the meta-learning phase: 1) intra-domain variance caused by the inherent support-query bias, due to the different appearances of organs and lesions, and 2) inter-domain variance caused by different medical imaging techniques. Additionally, we design an MSF module to integrate the different frequency features decoupled by the FAM module, and further mitigate the impact of inter-domain variance on the model's segmentation performance. Combining these two modules, our FAMNet surpasses existing FSMIS models and Cross-domain Few-shot Semantic Segmentation models on three cross-domain datasets, achieving state-of-the-art performance in the CD-FSMIS task.
Yuntian Bo, Yazhou Zhu 0001, Lunbo Li, Haofeng Zhang 0001
AAAI3
2025 Test-time Filtering Boosts Training-free Zero-shot Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) is an image retrieval task where users provide a reference image along with modification text to retrieve a target image. Zero-shot CIR (ZS-CIR) attracts significant research interest owing to its strong generalization capability and independence from labeled training data. Most ZS-CIR methods employ late fusion and textual inversion, but these approaches fail to precisely convert the reference image and modification text into a target-aligned query. Consequently, the final query may retain redundant information—such as elements present in the reference image but absent in the target image. To address these limitations, a training-free ZS-CIR method called Test-time Filtering (TTF) is proposed in this paper, leveraging Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Specifically, an MLLM generates captions for reference images, which, along with the modification text, are fed into an LLM to produce target image descriptions. This converts the final query into plain text. Using these results, an MLLM filters the selected candidate images to select positive and negative samples. Finally, the similarity scores of candidate images matching positive samples are amplified, while those of negative samples are suppressed. The proposed method is evaluated on three ZS-CIR datasets—CIRR, FashionIQ, and CIRCO—with experimental results demonstrating superior performance over prior ZS-CIR approaches. The source code is available at https://github.com/After-lifes/TTF.
Haoyue Chong, Lunbo Li, Haofeng Zhang 0001
MMAsia2
2025 Leveraging Pseudo-triplet and Flexible Prompt for Zero-shot Composed Image Retrieval
abstract
Compared to supervised composed image retrieval (CIR), which requires a large number of manually labeled triplets (reference image, modification text, target image) for training, Zero-shot Composed Image Retrieval (ZS-CIR) only needs easily available image-caption pairs for training. Prior works for ZS-CIR have primarily used textual inversion to convert images from image-caption pairs into pseudo-words using a pre-trained Vision-language Model (VLM). These pseudo-words are then concatenated to a template, such as “a photo of $”. We find that this fixed template representation hinders the generalization ability of ZS-CIR, and that converting an image into a single pseudo-token fails to accurately represent the information contained in the image. Additionally, given the high performance demonstrated by traditional supervised CIR after triplet training, we aim to achieve a similar structure in ZS-CIR. In this paper, we propose a novel ZS-CIR paradigm based on pseudo-triplets and flexible prompts, rather than fixed templates. Specifically, we fine-tune the vision encoder and split the caption to generate pseudo-triplets using random patch masking. The obtained masked image tokens and corresponding split captions are fed into a Q-former, which is guided by the learned queries and generates queries containing both image and caption information. Finally, the learned queries are concatenated with the split caption and fed into the text encoder to obtain a composed query. Extensive experiments on commonly used ZS-CIR benchmarks such as CIRR, FashionIQ, and CIRCO demonstrate that our method outperforms previous ZS-CIR approaches. The source code is available at https://github.com/After-lifes/PFP.
Haoyue Chong, Lunbo Li, Haofeng Zhang 0001
MMAsia2
2025 CREAM: Few-shot Object Counting with Cross REfinement and Adaptive density Map
Yuanwu Xu, Minxian Li, Qiaolin Ye, Lunbo Li, Haofeng Zhang 0001
Image Vis. Comput.5
2024 MTPNet: Learning Multiple Twin-support Prototypes for Few-shot Medical Image Segmentation
abstract
Given the high annotation costs and ethical considerations associated with medical images, leveraging a limited number of annotated samples for Few-Shot Medical Image Segmentation (FSMIS) has become increasingly prevalent. However, existing models tend to focus on visible foreground support information, often overlooking extreme foreground-background imbalances. In addition, query images sometimes have slight different appearance compared to support images of the same category due to the differences in size as well as slicing angle, thus employing only support images to generate prototypes inevitably leads to matching bias. To address these challenges, we present an innovative approach through learning a Multiple Twin-support Prototypes Network (MTPNet). Our approach includes the design of the Scale Consistent Sampling (SCS) module, which adaptively adjusts the foreground and background points within the support set, thereby balancing the influence of various structural elements in the image. Additionally, the Twin-support Prototypes Extraction (TPE) module facilitates the critical interaction between query and support features to extract twin-support prototypes. This module incorporates a Backtrace Interaction Filter (BIF) to eliminate erroneous interaction prototypes. Extensive experimental validation on three widely used medical image datasets demonstrates that our method surpasses current State-of-the-arts, showcasing its potential to address key limitations in FSMIS. The code is available at https://github.com/FeifanSong/MTPNet.
Feifan Song 0004, Ziming Cheng, Lunbo Li, Haofeng Zhang 0001
BIBM3
2023 Graph Attention Hashing via Contrastive Learning for Unsupervised Cross-Modal Retrieval
Shuyan Ding, Lunbo Li, Jianhui Guo
ICONIP (14)3
2023 CSTrack: A Comprehensive and Concise Vision Transformer Tracker
Shuyan Ding, Jianhui Guo, Lunbo Li
PRCV (12)5
2023 Exploiting spatial relationships for visual tracking
Lunbo Li, Jianhui Guo, Haofeng Zhang 0001
Pattern Recognit. Lett.2
2022 Self-Attentive CLIP Hashing for Unsupervised Cross-Modal Retrieval
abstract
With the explosive growth of multi-modal data such as video, images, and text on the Internet, cross-modal retrieval has received extensive attention, especially the deep hashing method. Compared with the real-value method, deep hashing has shown promising prospects due to its low memory consumption and high searching efficiency. However, most existing studies have difficulties in effectively utilizing the raw image-text pairs to generate discriminative feature representations. Moreover, these methods ignore the latent relationship between different modalities and fail to construct a robust similarity matrix, resulting in suboptimal retrieval performance. In this paper, we focus on the unsupervised cross-modal hashing tasks and propose a Self Attentive CLIP Hashing (SACH) model. Specifically, we construct the feature extraction network by employing the pre-trained CLIP model, which has shown excellent performance in zero-shot tasks. Besides, to fully exploit the semantic relationships, an attention module is introduced to reduce the disturbance of redundant information and focus on important information. On this basis, we construct a semantic fusion similarity matrix that capable of preserving the original semantic relationships from different modalities. Extensive experiments show the superiority of SACH compared with recent state-of-the-art unsupervised hashing methods.
Shuyan Ding, Lunbo Li, Jiexin Wu
MMAsia3
2022 Semi-supervised cross-modal hashing with multi-view graph representation
Haofeng Zhang 0001, Lunbo Li, Wankou Yang, Li Liu 0004
Inf. Sci.3
2021 Attention-Guided Semantic Hashing for Unsupervised Cross-Modal Retrieval
abstract
Recently, due to the low storage consumption and high search efficiency of hashing methods and the powerful feature extraction capability of deep neural networks, deep cross-modal hashing has received extensive attention in the field of multi-media retrieval. However, existing methods tend to ignore the latent relationships between heterogeneous data when learning a common semantic subspace, and cannot retain more important semantic information when mining deep correlations. In this paper, an attention mechanism which focuses on the characteristics of the associated features is employed to propose an attention-aware semantic fusion matrix that integrates important information from different modalities. We introduce a novel network that can pass the extracted features through the attention module to efficiently encode rich and relevant features, and can also generate hash codes under the self-supervision of the proposed attention-aware semantic fusion matrix. Our experimental results and detailed analysis prove that our method can achieve better retrieval performance on the three popular datasets, compared with the recent unsupervised cross-modal hashing methods.
Haofeng Zhang 0001, Lunbo Li, Li Liu 0004
ICME3
2021 Clustering-driven Deep Adversarial Hashing for scalable unsupervised cross-modal retrieval
Haofeng Zhang 0001, Lunbo Li, Zheng Zhang 0006, Debao Chen, Li Liu 0004
Neurocomputing3