Yan Zhang 0135

dblp:04/3348-135 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-2819-388XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 iEBAKER: Improved remote sensing image-text retrieval framework via eliminate before align and keyword explicit reasoning
Yan Zhang 0135, Zhong Ji, Changxu Meng, Yanwei Pang
Expert Syst. Appl.1
2026 A causality based multi-task framework for enhancing image and text interactions
Zhaomeng Cheng, Zhong Ji, Yan Zhang 0135
Knowl. Based Syst.3
2026 Underlying Semantic Diffusion for Effective and Efficient In-Context Learning
abstract
Diffusion models have emerged as a powerful framework for tasks like image controllable generation and dense prediction. However, existing models often struggle to capture underlying semantics (e.g., edges, textures, shapes) and effectively utilize in-context learning, limiting their contextual understanding and image generation quality. Furthermore, high computational costs and slow inference speeds hinder their real-time applications. To address these challenges, we propose Underlying Semantic Diffusion (US-Diffusion), an enhanced diffusion model that improves underlying semantics learning, computational efficiency, and in-context learning capabilities on multi-task scenarios. We introduce Separate & Gather Adapter (SGA), which decouples input conditions for different tasks while sharing the architecture, enabling better in-context learning and generalization across diverse visual domains. We also present a Feedback-Aided Learning (FAL) framework, which leverages feedback signals to guide the model in capturing semantic details and dynamically adapting to task-specific contextual cues. Furthermore, we propose a plug-and-play Efficient Sampling Strategy (ESS) for dense sampling at time steps with high-noise levels, which aims at optimizing training and inference efficiency while maintaining strong in-context learning performance. Experimental results demonstrate that US-Diffusion outperforms the state-of-the-art method, achieving an average reduction of 7.47 in FID on Map2Image tasks and an average reduction of 0.026 in RMSE on Image2Map tasks, while achieving approximately $9.45\times $ faster inference speed. Our method also demonstrates superior training efficiency and in-context learning capabilities, excelling in new datasets and tasks, highlighting its robustness and adaptability across diverse visual domains. The source code will be released at https://github.com/dragon-cao/US-Diffusion.
Zhong Ji, Weilong Cao, Yan Zhang 0135, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.3
2025 Video Wire Inpainting via Hierarchical Feature Mixture
Zhong Ji, Yimu Su, Yan Zhang 0135, Shuangming Yang, Yanwei Pang
Image Vis. Comput.3
2025 Hierarchical and complementary experts transformer with momentum invariance for image-text retrieval
Yan Zhang 0135, Zhong Ji, Yanwei Pang, Jungong Han
Knowl. Based Syst.1
2025 Raformer: Redundancy-Aware Transformer for Video Wire Inpainting
abstract
Video Wire Inpainting (VWI) is a prominent application in video inpainting, aimed at flawlessly removing wires in films or TV series, offering significant time and labor savings compared to manual frame-by-frame removal. However, wire removal poses greater challenges due to the wires being longer and slimmer than objects typically targeted in general video inpainting tasks, and often intersecting with people and background objects irregularly, which adds complexity to the inpainting process. Recognizing the limitations posed by existing video wire datasets, which are characterized by their small size, poor quality, and limited variety of scenes, we introduce a new VWI dataset with a novel mask generation strategy, namely Wire Removal Video Dataset 2 (WRV2) and Pseudo Wire-Shaped (PWS) Masks. WRV2 dataset comprises over 4,000 videos with an average length of 80 frames, designed to facilitate the development and efficacy of inpainting models. Building upon this, our research proposes the Redundancy-Aware Transformer (Raformer) method that addresses the unique challenges of wire removal in video inpainting. Unlike conventional approaches that indiscriminately process all frame patches, Raformer employs a novel strategy to selectively bypass redundant parts, such as static background segments devoid of valuable information for inpainting. At the core of Raformer is the Redundancy-Aware Attention (RAA) module, which isolates and accentuates essential content through a coarse-grained, window-based attention mechanism. This is complemented by a Soft Feature Alignment (SFA) module, which refines these features and achieves end-to-end feature alignment. Extensive experiments on both the traditional video inpainting datasets and our proposed WRV2 dataset demonstrate that Raformer outperforms other state-of-the-art methods. Our codes and the WRV2 dataset will be made available at: https://github.com/Suyimu/WRV2.
Zhong Ji, Yimu Su, Yan Zhang 0135, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.3
2025 Visual Semantic Contextualization Network for Multi-Query Image Retrieval
abstract
Multi-Query Image Retrieval (MQIR) aims to establish connections between vision and language by exploring fine-grained region-query alignments. It is still a challenging task owing to its intrinsical ambiguity, where a query matches with multiple semantically similar regions and introduces misleading noises. Although researchers have made great efforts to alleviate the ambiguity in many retrieval-related tasks, there are few attempts considering this bottleneck in MQIR, which greatly limits present performance. To this end, we propose a novel Visual Semantic Contextualization Network (VSCN) to mitigate ambiguity by capturing the contextual knowledge within each image-text pair. Specifically, we first develop a Context Semantic Perception (CSP) module to capture the dual-level context, where a visual context transformer explores the intra-context within regions, and a cross-modal context transformer mines the inter-context among concatenated visual-linguistic embeddings. Then, to yield superior contextual understanding, we strengthen the connotations in context via a Context Semantic Interaction (CSI) module. Particularly, knowledge distillation is first employed to transfer the CLIP-guided semantic into the regional intra-context to complement the potential background information. Then, the intra-context & inter-context interaction is conducted via the self-attention mechanism to link the dual-level context and obtain the interacted contextual knowledge. Our method is evaluated on the Visual Genome dataset and substantially outperforms the state-of-the-art methods (30.3% improvements on Recall@1 in the first round). Our source codes will be released athttps://github.com/zhli-cs/VSCN.
Zhong Ji, Zhihao Li 0006, Yan Zhang 0135, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Multim.3
2024 Eliminate Before Align: A Remote Sensing Image-Text Retrieval Framework with Keyword Explicit Reasoning
abstract
Mountains of researches center around the Remote Sensing Image-Text Retrieval (RSITR), aiming at retrieving the corresponding targets based on the given query. Among them, the transfer of Foundation Models (FMs), such as CLIP, to remote sensing domain shows promising results. However, existing FM-based approaches neglect the negative impact of weakly correlated sample pairs and the key distinctions among remote sensing texts, leading to biased and superficial exploration of sample pairs. To address these challenges, we propose a novel Eliminate Before Align strategy with Keyword Explicit Reasoning framework (EBAKER) for RSITR. Specifically, we devise an innovative Eliminate Before Align (EBA) strategy to filter out the weakly correlated sample pairs to mitigate their deviations from optimal embedding space during alignment. Moreover, we introduce a Keyword Explicit Reasoning (KER) module to facilitate the positive role of subtle key concept differences. Without bells and whistles, our method achieves a one-step transformation from FM to RSITR task, obviating the necessity for extra pretraining on remote sensing data. Extensive experiments on three popular benchmark datasets validate that our proposed EBAKER method outperform the state-of-the-art methods with fewer training data. Our source code will be released soon.
Zhong Ji, Changxu Meng, Yan Zhang 0135, Haoran Wang 0004, Yanwei Pang, Jungong Han
ACM Multimedia3
2024 Modality-experts coordinated adaptation for large multimodal models
Yan Zhang 0135, Zhong Ji, Yanwei Pang, Jungong Han, Xuelong Li 0001
Sci. China Inf. Sci.1
2024 Hierarchical matching and reasoning for multi-query image retrieval
Zhong Ji, Zhihao Li 0006, Yan Zhang 0135, Haoran Wang 0004, Yanwei Pang, Xuelong Li 0001
Neural Networks3
2024 USER: Unified Semantic Enhancement With Momentum Contrast for Image-Text Retrieval
abstract
As a fundamental and challenging task in bridging language and vision domains, Image-Text Retrieval (ITR) aims at searching for the target instances that are semantically relevant to the given query from the other modality, and its key challenge is to measure the semantic similarity across different modalities. Although significant progress has been achieved, existing approaches typically suffer from two major limitations: (1) It hurts the accuracy of the representation by directly exploiting the bottom-up attention based region-level features where each region is equally treated. (2) It limits the scale of negative sample pairs by employing the mini-batch based end-to-end training mechanism. To address these limitations, we propose a Unified Semantic Enhancement Momentum Contrastive Learning (USER) method for ITR. Specifically, we delicately design two simple but effective Global representation based Semantic Enhancement (GSE) modules. One learns the global representation via the self-attention algorithm, noted as Self-Guided Enhancement (SGE) module. The other module benefits from the pre-trained CLIP module, which provides a novel scheme to exploit and transfer the knowledge from an off-the-shelf model, noted as CLIP-Guided Enhancement (CGE) module. Moreover, we incorporate the training mechanism of MoCo into ITR, in which two dynamic queues are employed to enrich and enlarge the scale of negative sample pairs. Meanwhile, a Unified Training Objective (UTO) is developed to learn from mini-batch based and dynamic queue based samples. Extensive experiments on the benchmark MSCOCO and Flickr30K datasets demonstrate the superiority of both retrieval accuracy and inference efficiency. For instance, compared with the existing best method NAAF, the metric R@1 of our USER on the MSCOCO 5K Testing set is improved by 5% and 2.4% on caption retrieval and image retrieval without any external knowledge or pre-trained model while enjoying over 60 times faster inference speed. Our source code will be released at https://github.com/zhangy0822/USER.
Yan Zhang 0135, Zhong Ji, Di Wang 0026, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Image Process.1
2023 Consensus Knowledge Exploitation for Partial Query Based Image Retrieval
abstract
Partial Query based Image Retrieval (PQIR) enables a search engine to perform an interactive retrieval given by only an initial query and actively provide alternative feedbacks for a user to refine a set of retrieval results. It alleviates the deficiency in practice interactive image retrieval that requires the user to laboriously provide detailed feedbacks, and enables the retrieval on-the-fly with the incomplete initial query. Although significant progress has been made, existing works remain have challenge in actively providing more discriminative feedbacks. To address this challenge, we propose a novel Attributes&Objects-based Consensus Extraction and Representation (AoCer) framework. Specifically, we formulate a simple but effective Attribute&Object Feedback (AOF) paradigm, which employs both attributes and objects as intermediate feedbacks to carry out multiple rounds of interaction. To mine the intrinsic associations among concepts and enhance their feature representations, we further propose an Interventional Consensus Representation Learning (ICRL) module, which mainly constructs an interventional concept graph to yield the Interventional Consensus Representation (ICR). In addition, a Dual-Head Feedback Sampler (DHFS) is developed to sample objects and attributes for conducting the next round retrieval. Extensive experiments demonstrate the superiority of the proposed framework. Our source code will be released athttps://github.com/zhangy0822/AoCer.
Yan Zhang 0135, Zhong Ji, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Knowledge-Aided Momentum Contrastive Learning for Remote-Sensing Image Text Retrieval
abstract
Remote sensing image-text retrieval (RSITR) has attracted widespread attention due to its great potential for rapid information mining ability on remote sensing images. Although significant progress has been achieved, existing methods typically overlook the challenge posed by the extremely analogous descriptions, where the subtle differences remain largely unexploited or, in some cases, are entirely disregarded. To address the limitation, we propose a Knowledge Aided Momentum Contrastive Learning (KAMCL) method for RSITR. Specifically, we propose a novel Knowledge Aided Learning framework, including knowledge initialization, construction, filtration, and alignment operations, which aims at providing valuable concepts and learning discriminative representations. On this basis, we integrate Momentum Contrastive Learning to promote the capture of key concepts within the representation via expanding the scale of negative sample pairs. Moreover, we design a hierarchical aggregator module to better capture the multi-level information from remote sensing images. Finally, we introduce an innovative two-step training strategy designed to effectively harness the synergy among concepts and leverage their respective functionalities. Extensive experiments conducted on the three public datasets showcase the remarkable performance of our approach in terms of retrieval accuracy and computational efficiency. For instance, compared with the existing state-of-the-art method, our method exhibits notable performance improvements of 2.65% on the RSICD dataset, simultaneously achieving improvements in inference efficiency by 48%. Our source code will be released at https://github.com/mcx-mcx/KAMCL.
Zhong Ji, Changxu Meng, Yan Zhang 0135, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.3