Ming Lu 0008

dblp:15/5997-8 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-0577-1969ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 SPCL: Semantic Polymorphism and Commonality Learning for Text-Based Person Retrieval
abstract
Text-Based Person Retrieval (TBPR) refers to identifying a specific target pedestrian image based on natural language descriptions. Most previous methods rely on one-to-one alignment between paired text-image data, ignoring the polymorphic nature of visual and linguistic information. Moreover, constrained by ID, earlier methods have shown limited exploration of intra-individual and inter-individual relations. This limitation confines them to exploring characteristics within individuals, making it challenging to uncover commonalities and invariants that extend across IDs (e.g., attributes). Recently, due to the lack of accurate annotations, exploring attribute-based cross-modal interactions and alignments has become a significant challenge in TBPR. To address these issues, we propose a Semantic Polymorphism and Commonality Learning (SPCL) framework. First, we present Relation-Sensitive Semantic Polymorphism Alignment (RSSPA) and ID-Based Semantic Polymorphism Alignment (IBSPA) to explore ID-limited Feature Redistribution. Second, we transcend the constraints of ID, leveraging ID-Free Attribute Alignment (IFAA) from a macro perspective to explore commonalities and invariants based on attribute features. Finally, from a micro perspective, we design Attribute Prior Fusion Reconstruction (APFR) to optimize the attention of our model, exploring the positive impact of attribute priors on cross-modal interaction. Experiments on CUHK-PEDES, ICFG-PEDES and RSTPReid show that our method achieves state-of-the-art performance on Rank-1, mAP and mINP.
Jiayi Li 0003, Jun Kong 0001, Yunde Zhang, Ming Lu 0008, Min Jiang 0008
IEEE Trans. Circuits Syst. Video Technol.4
2026 Mitigating Inherent Bias of Answer Heuristic Based Frameworks in Knowledge-Based Visual Question Answering
abstract
Knowledge-based Visual Question Answering (KBVQA) aims to utilize external knowledge to answer image-related questions. Current KBVQA frameworks prompt large language models (LLMs) with answer heuristics, narrowing the focus to more relevant answers. However, these answer heuristic based frameworks exhibit the bias towards overemphasizing highest-scoring answers, neglecting potential answers with lower scores. This bias arises from the deficiencies of in-context example constructions and underperformances of multi-query ensemble strategy. In this paper, we propose MinBias, an approach designed tomitigateinherentbiasof answer heuristic based VQA framework. Firstly, to help LLMs learn dialectically, we propose Dialectical Learning Space Construction strategy (DLSC). This module mitigates bias via guaranteeing diversity in the selected examples. Secondly, to strengthen the connection between image captions and questions, we propose Large Language Model guided Visual Masking (LVM) algorithm. This module mitigates bias via enhancing visual clues in example content. Finally, we propose Visual Entailment based Answer Rerank module (VEAR). This module mitigates the bias arising from multi-query strategy via obtaining unbiased evaluations. Extensive experiments demonstrate that our method mitigates the bias of the answer heuristic based VQA framework and enhances the model performance.
Min Jiang 0008, Jun Kong 0001, Danfeng Zhuang, Ming Lu 0008
IEEE Trans. Multim.5
2025 Learning Distinct Semantic Information in Modality-Specific and Modality-Shared Features for Visible-Infrared Person Re-identification
Min Jiang 0008, Huang Luo, Jun Kong 0001, Ming Lu 0008
PRCV (17)4
2025 Vari-FocalFuse: Sparse Attention With Adaptive Attention Spans for Image Fusion
abstract
Existing image fusion methods introduce convolution or attention to perceive multi-scale targets in different source images. However, fixed operation spans, including convolution spans and attention spans, result in insufficient or redundant perception. In this work, we propose Vari-FocalFuse, which performs sparse attention with adaptive attention spans. Firstly, to perform local perception and capture intricate details, we propose Focal Attention (FAtten). It localizes attention spans to the nearest neighbor tokens like convolutions, modeling spatial local correlation. Secondly, we adapt the attention spans to multi-scale targets, and propose Memorization Content Awareness (MCAware). The attention spans are expanded until they can cover the corresponding targets. Thirdly, to perform global perception and minimize redundant perception, we propose Vari-Focal Attention (VFAtten) by combining FAtten and MCAware. The adaptive attention spans can capture multi-scale targets without taking irrelevant information into account. Finally, to mitigate noise caused by redundant perception, we propose Sandwich GLU (SwGLU). Spatial perception refined gating process is utilized to mitigate redundant information. Extensive experiments demonstrate that our Vari-FocalFuse achieves state-of-the-art (SOTA) performance on various image fusion tasks.
Shengchen Zhu, Min Jiang 0008, Jun Kong 0001, Ming Lu 0008, Jiayi Li 0003
IEEE Signal Process. Lett.4
2025 AU-Net: Adaptive Unified Network for Joint Multi-Modal Image Registration and Fusion
abstract
Joint multi-modal image registration and fusion (JMIRF) typically follows a register-first, fuse-later paradigm. It has a registration module to align parallax images and a fusion module to fuse registered images. Existing research typically focuses on the mutual enhancement between the two modules, but this is essentially a straightforward combination rather than an efficient, unified network. Moreover, executing the two modules separately may cause inefficiency, as the total runtime is merely the sum of both steps without investigating potential shared structures. In this paper, we propose an Adaptive Unified Network (AU-Net) following a novel end-to-end paradigm called Feature-Level Joint Training (FLJT). Firstly, AU-Net learns registration and fusion within a unified network through shared structure and hierarchical semantic interaction. A multi-level dynamic fusion module is designed to adaptively fuse input features from different scales and modalities. Secondly, the image-to-image translation based on Denoising Diffusion Probabilistic Models (DDPMs) is introduced to train AU-Net using simple and reliable single-modal metrics. Unlike previous unidirectional translation, we explore bidirectional translation to provide additional implicit branch supervision. Furthermore, a cache-like scheme is proposed to elegantly circumvent the additional computational overhead caused by the iterative denoising of DDPMs. Finally, our method was validated on two publicly available datasets, demonstrating advantages over state-of-the-art methods in terms of qualitative evaluation, quantitative evaluation, and computational complexity analysis. The code will be publically available at https://github.com/luming1314/AU-Net.
Ming Lu 0008, Min Jiang 0008, Xuefeng Tao, Jun Kong 0001
IEEE Trans. Image Process.1
2024 MIAFusion: Infrared and Visible Image Fusion via Multi-scale Spatial and Channel-Aware Interaction Attention
Teng Lin, Ming Lu 0008, Min Jiang 0008, Jun Kong 0001
PRCV (8)2
2024 Unsupervised Learning of Intrinsic Semantics With Diffusion Model for Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) aims to learn semantic representations for person retrieval without using identity labels. Most existing methods generate fine-grained patch features to reduce noise in global feature clustering. However, these methods often compromise the discriminative semantic structure and overlook the semantic consistency between the patch and global features. To address these problems, we propose a Person Intrinsic Semantic Learning (PISL) framework with diffusion model for unsupervised person Re-ID. First, we design the Spatial Diffusion Model (SDM), which performs a denoising diffusion process from noisy spatial transformer parameters to semantic parameters, enabling the sampling of patches with intrinsic semantic structure. Second, we propose the Semantic Controlled Diffusion (SCD) loss to guide the denoising direction of the diffusion model, facilitating the generation of semantic patches. Third, we propose the Patch Semantic Consistency (PSC) loss to capture semantic consistency between the patch and global features, refining the pseudo-labels of global features. Comprehensive experiments on three challenging datasets show that our method surpasses current unsupervised Re-ID methods. The source code will be publicly available at https://github.com/taoxuefong/Diffusion-reid.
Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Ming Lu 0008, Ajmal Mian
IEEE Trans. Image Process.4