EDBT 2026 Demo / reviewers in the wild / expert
Xingjian He
dblp:204/0216
· DBLP profile ↗
23ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0001-5396-6253ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human TrajectoriesabstractNavigating complex urban environments using natural language instructions poses significant challenges for embodied agents, including noisy language instructions, ambiguous spatial references, diverse landmarks, and dynamic street scenes. Current visual navigation methods are typically limited to simulated or off-street environments, and often rely on precise goal formats, such as specific coordinates or images. This limits their effectiveness for autonomous agents like last-mile delivery robots navigating unfamiliar cities. To address these limitations, we introduce UrbanNav, a scalable framework that trains embodied agents to follow free-form language instructions in diverse urban settings. Leveraging web-scale city walking videos, we develop an scalable annotation pipeline that aligns human navigation trajectories with language instructions grounded in real-world landmarks. UrbanNav encompasses over 1,500 hours of navigation data and 3 million instruction-trajectory-landmark triplets, capturing a wide range of urban scenarios. Our model learns robust navigation policies to tackle complex urban scenarios, demonstrating superior spatial reasoning, robustness to noisy instructions, and generalization to unseen urban settings. Experimental results show that UrbanNav significantly outperforms existing methods, highlighting the potential of large-scale web video data to enable language-guided, real-world urban navigation for embodied agents. Yanghong Mei, Yirong Yang, Longteng Guo, Qunbo Wang, Ming-Ming Yu, Xingjian He, Wenjun Wu 0001, Jing Liu 0001 |
AAAI | 6 |
| 2026 | Generalized referring expression segmentation driven by instance-oriented queries
Jiangyun Li, Zhaokun Wen, Yisi Zhang, Wenxuan Wang 0002, Yuanxiu Cai, Xingjian He, Jing Liu 0001 |
Pattern Recognit. | 7 |
| 2025 | ViPE: Visual Perception in Parameter Space for Efficient Video-Language UnderstandingabstractExisting video-language models (Video-LLMs) typically rely on concatenating visual tokens with textual inputs for joint modeling.However, this token-level alignment leads to significant inefficiency, especially when scaling to long videos with dense visual inputs.In this work, we propose a video-to-parameter efficiency paradigm named ViPE that eliminates redundant visual tokens by transforming video content into visual perceptual weights, which are directly injected into the LLM's parameters.ViPE consists of a visual injection module that compresses video features into a small set of perceptual queries using a hierarchical merge strategy, and a visual perception module that integrates the resulting representations into the LLM through a lightweight LoRA-like mechanism.ViPE achieves performance comparable to token-based baselines such as LLaVA, while reducing FLOPs by 85% and inference time by up to 65%, demonstrating a highly efficient and scalable solution for video understanding. Shichen Lu, Tongtian Yue, Longteng Guo, Handong Li, Xingjian He, Si Liu 0001, Jing Liu 0001 |
EMNLP | 5 |
| 2025 | MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge DistillationabstractIn recent years, Embodied Artificial Intelligence (Embodied AI) has advanced rapidly, yet the increasing size of models conflicts with the limited computational capabilities of Embodied AI platforms. To address this challenge, we aim to achieve both high model performance and practical deployability. Specifically, we focus on Vision-and-Language Navigation (VLN), a core task in Embodied AI. This paper introduces a two-stage knowledge distillation framework, producing a student model, MiniVLN, and showcasing the significant potential of distillation techniques in developing lightweight models. The proposed method aims to capture fine-grained knowledge during the pretraining phase and navigation-specific knowledge during the fine-tuning phase. Our findings indicate that the two-stage distillation approach is more effective in narrowing the performance gap between the teacher model and the student model compared to single-stage distillation. On the public R2R and REVERIE benchmarks, MiniVLN achieves performance on par with the teacher model while having only about 12 % of the teacher model's parameter count. Junyou Zhu, Yanyuan Qiao, Xingjian He, Qi Wu 0001, Jing Liu 0001 |
ICRA | 4 |
| 2025 | GroundingMate: Aiding Object Grounding for Goal-Oriented Vision-and-Language NavigationabstractGoal-Oriented Vision-and-Language Navigation (VLN) aims to enable agents to navigate to specified locations and identify designated target objects following natural language instruction. This approach has gained popularity due to its close alignment with real-world scenarios. However, existing studies have predominantly focused on enhancing navigation performance, neglecting the ability to locate objects at the navigation endpoint. This oversight has resulted in a significant discrepancy between the success rates of navigation and object grounding. The challenge is compounded by the complex reasoning required by the instructions and the necessity to synthesize multiperspective images of objects, which overwhelms traditional object grounding methods. We leverage the Multi-Modal Large Language Model (MLLM) to bridge this gap, allowing agents to seek assistance from these models when struggling to locate the target object. The agent conducts a multi-stage evaluation to discern the cause of its confusion and promptly extracts and updates the most relevant information for MLLM to assess. Our method is plug-and-play and model-agnostic, facilitating integration with numerous existing VLN strategies without the need for retraining. Implementing our approach across four distinct methods has improved performance on the REVERIE and SOON datasets, demonstrating the effectiveness and generalizability of our technique. Qianyi Liu, Yanyuan Qiao, Junyou Zhu, Longteng Guo, Qunbo Wang, Xingjian He, Qi Wu 0001, Jing Liu 0001 |
WACV | 8 |
| 2025 | VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and DatasetabstractIn this paper, we propose the Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multimodal understanding and generation. Unlike widely-studied vision-language pretraining models, VALOR jointly models the relationships among vision, audio, and language in an end-to-end manner. It consists of three separate encoders for single modality representations and a decoder for multimodal conditional text generation. We design two pretext tasks to pretrain the VALOR model: Multimodal Grouping Alignment (MGA) and Multimodal Grouping Captioning (MGC). MGA projects vision, language, and audio into the same common space, simultaneously building vision-language, audio-language, and audiovisual-language alignment. MGC learns to generate text tokens under conditions of vision, audio, or both. To promote vision-audio-language pretraining research, we construct a large-scale, high-quality tri-modality dataset named VALOR-1M, containing 1 million audible videos with human-annotated audiovisual captions. Extensive experiments show that VALOR can learn strong multimodal correlations and generalize to various downstream tasks (e.g., retrieval, captioning, and question answering) with different input modalities (e.g., vision-language, audio-language, and audiovisual-language). VALOR achieves new state-of-the-art performance on a series of public cross-modality benchmarks. Jing Liu 0001, Xingjian He, Longteng Guo, Weining Wang 0001, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | CGViT: Cross-image GroupViT for zero-shot semantic segmentation
Jie Jiang 0016, Xingjian He, Weining Wang 0001, Jing Liu 0001 |
Pattern Recognit. | 2 |
| 2025 | VLAB: Enhancing Video Language Pretraining by Feature Adapting and BlendingabstractLarge-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited research on learning video-text representations for general video multimodal tasks based on these powerful features. Towards this goal, we propose a novel video-text pre-training method dubbed VLAB:VideoLanguage pre-training by featureAdapting andBlending, which transfers CLIP representations to video pre-training tasks and develops unified video multimodal models for a wide range of video-text tasks. Specifically, VLAB is founded on two key strategies: feature adapting and feature blending. In the former, we introduce a new video adapter module to address CLIP's deficiency in modeling temporal information and extend the model's capability to encompass both contrastive and generative tasks. In the latter, we propose an end-to-end training method that further enhances the model's performance by exploiting the complementarity of image and video features. We validate the effectiveness and versatility of VLAB through extensive experiments on highly competitive video multimodal tasks, including video text retrieval, video captioning, and video question answering. Remarkably, VLAB outperforms competing methods significantly and sets new records in video question answering on MSRVTT, MSVD, and TGIF datasets. It achieves an accuracy of 49.6, 60.9, and 79.0, respectively. Xingjian He, Fan Ma, Zhicheng Huang 0002, Xiaojie Jin 0004, Dongmei Fu, Yi Yang 0001, Jing Liu 0001, Jiashi Feng |
IEEE Trans. Multim. | 1 |
| 2025 | Hierarchical Contrastive Learning for Semantic SegmentationabstractRecently, pixel-to-pixel contrastive learning in single-scale feature space has been widely studied in semantic segmentation to learn a unified feature expression for pixels of the same category. However, the unified representation is too extreme, and the receptive field of each single-scale pixel is limited, which is insufficient to reflect the representative features of the category. To address these problems, this article extends the single-scale feature space to that of multiscale and proposes a hierarchical contrastive learning (Hi-CL) method to explore pixel-to-component semantic relationships. First, we generate multiscale candidate samples by applying several pooling windows with different sizes on a feature map, where different windows may represent different parts of the objects in the image. Then, we prune the sample set through threshold-based criteria to select appropriate samples for feature representation learning. Finally, Hi-CL is performed to learn the pixel-to-component consistency with the pruned samples. Our method is easy to be applied on existing semantic segmentation models and obtains consistent improvement. Furthermore, we achieve state-of-the-art results on three popular benchmarks, including Cityscapes, ADE20K, and COCO Stuff datasets. Jie Jiang 0016, Xingjian He, Weining Wang 0001, Hanqing Lu, Jing Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression SegmentationabstractReferring expression segmentation (RES) aims at segmenting the foreground masks of the entities that match the descriptive natural language expression. Previous datasets and methods for classic RES task heavily rely on the prior assumption that one expression must refer to object-level targets. In this paper, we take a step further to finer-grained part-level RES task. To promote the object-level RES task towards finer-grained vision-language understanding, we put forward a new multi-granularity referring expression segmentation (MRES) task and construct an evaluation benchmark called RefCOCOm by manual annotations. By employing our automatic model-assisted data engine, we build the largest visual grounding dataset namely MRES-32M, which comprises over 32.2M high-quality masks and captions on the provided 1M images. Besides, a simple yet strong model named UniRES is designed to accomplish the unified object-level and part-level grounding task. Extensive experiments on our RefCOCOm for MRES and three datasets (i.e., RefCOCO(+/g)) for classic RES task demonstrate the superiority of our method over previous state-of-the-art methods. To foster future research into fine-grained visual grounding, our benchmark RefCOCOm, the MRES-32M dataset and model UniRES will be publicly available at https://github.com/Rubics-Xuan/MRES. Wenxuan Wang 0002, Tongtian Yue, Yisi Zhang, Longteng Guo, Xingjian He, Jing Liu 0001 |
CVPR | 5 |
| 2024 | SC- Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language ModelsabstractRecent trends in Large Vision Language Models (LVLMs) research have been increasingly focusing on ad-vancing beyond general image understanding towards more nuanced, object-level referential comprehension. In this paper, we present and delve into the self-consistency ca-pability of LVLMs, a crucial aspect that reflects the mod-els' ability to both generate informative captions for spe-cific objects and subsequently utilize these captions to ac-curately re-identify the objects in a closed-loop process. This capability significantly mirrors the precision and reli-ability of fine- grained visual-language understanding. Our findings reveal that the self-consistency level of existing LVLMs falls short of expectations, posing limitations on their practical applicability and potential. To address this gap, we introduce a novel fine-tuning paradigm named Self-Consistency Tuning (SC-Tune). It features the syn-ergistic learning of a cyclic describer-locator system. This paradigm is not only data-efficient but also exhibits gener-alizability across multiple LVLMs. Through extensive ex-periments, we demonstrate that SC- Tune significantly ele-vates performance across a spectrum of object-level vision-language benchmarks and maintains competitive or im-proved performance on image-level vision-language bench-marks. Both our model and code will be publicly available at https://github.com/ivattyue/SC-Tune. Tongtian Yue, Jie Cheng 0009, Longteng Guo, Xingyuan Dai, Zijia Zhao, Xingjian He, Gang Xiong 0001, Jing Liu 0001 |
CVPR | 6 |
| 2024 | Fuse and Calibrate: A Bi-directional Vision-Language Guided Framework for Referring Image Segmentation
Xingjian He, Shichen Lu, Jing Liu 0001 |
ICIC (11) | 2 |
| 2024 | COSA: Concatenated Sample Pretrained Vision-Language Foundation ModelabstractDue to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modeling visually semantic representations while disregarding temporal semantic representations and correlations. To address this issue, we propose COSA, a COncatenated SAmple pretrained vision-language foundation model. COSA can jointly model visual contents and event-level temporal cues using only image-text corpora. We achieve this by sequentially concatenating multiple image-text pairs as inputs for pretraining. This transformation effectively converts existing image-text corpora into a pseudo video-paragraph corpus, enabling richer scene transformations and explicit event-description correspondence. Extensive experiments demonstrate that COSA consistently improves performance across a broad range of semantic vision-language downstream tasks, including paragraph-to-video retrieval, text-to-video/image retrieval, video/image captioning and video QA. Notably, COSA achieves state-of-the-art results on various competitive benchmarks. Code and model are released at https://github.com/TXH-mercury/COSA. Xingjian He, Handong Li, Xiaojie Jin 0004, Jiashi Feng, Jing Liu 0001 |
ICLR | 2 |
| 2024 | Calibration & Reconstruction: Deeply Integrated Language for Referring Image SegmentationabstractReferring image segmentation aims to segment an object referred to by natural language expression from an image. The primary challenge lies in the efficient propagation of fine-grained semantic information from textual features to visual features. Many recent works utilize a Transformer to address this challenge. However, conventional transformer decoders can distort linguistic information with deeper layers, leading to suboptimal results. In this paper, we introduce CRFormer, a model that iteratively calibrates multi-modal features in the transformer decoder. We start by generating language queries using vision features, emphasizing different aspects of the input language. Then, we propose a novel Calibration Decoder (CDec) wherein the multi-modal features can iteratively calibrated by the input language features. In the Calibration Decoder, we use the output of each decoder layer and the original language features to generate new queries for continuous calibration, which gradually updates the language features. Based on CDec, we introduce a Language Reconstruction Module and a reconstruction loss. This module leverages queries from the final layer of the decoder to reconstruct the input language and compute the reconstruction loss. This can further prevent the language information from being lost or distorted. Our experiments consistently show the superior performance of our approach across RefCOCO, RefCOCO+, and G-Ref datasets compared to state-of-the-art methods. Xingjian He, Jing Liu 0001 |
ICMR | 2 |
| 2024 | CM-MaskSD: Cross-Modality Masked Self-Distillation for Referring Image SegmentationabstractReferring image segmentation (RIS) is a fundamental vision-language task that intends to segment a desired object from an image based on a given natural language expression. Due to the essentially distinct data properties between image and text, most of existing methods either introduce complex designs towards fine-grained vision-language alignment or lack required dense alignment, resulting in scalability issues or mis-segmentation problems such as over- or under-segmentation. To achieve effective and efficient fine-grained feature alignment in the RIS task, we explore the potential of masked multimodal modeling coupled with self-distillation and propose a novel cross-modality masked self-distillation framework named CM-MaskSD, in which our method inherits the transferred knowledge of image-text semantic alignment from CLIP model to realize fine-grained patch-word feature alignment for better segmentation accuracy. Moreover, our CM-MaskSD framework can considerably boost model performance in a nearly parameter-free manner, since it shares weights between the main segmentation branch and the introduced masked self-distillation branches, and solely introduces negligible parameters for coordinating the multimodal features. Comprehensive experiments on three benchmark datasets (ie RefCOCO, RefCOCO+, G-Ref) for the RIS task convincingly demonstrate the superiority of our proposed framework over previous state-of-the-art methods. Wenxuan Wang 0002, Xingjian He, Yisi Zhang, Longteng Guo, Jiangyun Li, Jing Liu 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | WL-MSR: Watch and Listen for Multimodal Subtitle RecognitionabstractVideo subtitles could be defined as the combination of visualized subtitles in frames and textual content recognized from speech, which play a significant role in video understanding for both humans and machines. In this paper, we propose a novel Watch and Listen for Multimodal Subtitle Recognition (WL-MSR) framework to obtain comprehensive video subtitles, by fusing the information provided by Optical Character Recognition (OCR) and Automatic Speech Recognition (ASR) models. Specifically, we build a Transformer model with mask and crop strategies and multi-level identity embeddings to aggregate both the textual results and features of the two modalities. To pre-filter out the noise items in OCR results before fusion, we adopt an OCR filter based on ASR results and confidence scores of OCR. By combining these techniques, our solution wins the 2nd place in Multimodal Subtitle Recognition Challenge on ICPR2022. Jiawei Liu 0001, Weining Wang 0001, Xingjian He, Jing Liu 0001 |
ICASSP | 4 |
| 2023 | CSDNet: Contrastive Similarity Distillation Network for Multi-lingual Image-Text Retrieval
Shichen Lu, Longteng Guo, Xingjian He, Jing Liu 0001, Si Liu 0001 |
ICIG (3) | 3 |
| 2023 | Enhancing Vision-Language Pre-Training with Jointly Learned Questioner and Dense CaptionerabstractLarge pre-trained multimodal models have demonstrated significant success in a range of downstream tasks, including image captioning, image-text retrieval, visual question answering (VQA), etc. However, many of these methods rely on image-text pairs collected from the web as pre-training data and unfortunately overlook the need for fine-grained feature alignment between vision and language modalities, which requires detailed understanding of images and language expressions. While integrating VQA and dense captioning (DC) into pre-training can address this issue, acquiring image-question-answer as well as image-location-caption triplets is challenging and time-consuming. Additionally, publicly available datasets for VQA and dense captioning are typically limited in scale due to manual data collection and labeling efforts. In this paper, we propose a novel method called Joint QA and DC GEneration (JADE), which utilizes a pre-trained multimodal model and easily-crawled image-text pairs to automatically generate and filter large-scale VQA and dense captioning datasets. We apply this method to the Conceptual Caption (CC3M) dataset to generate a new dataset called CC3M-QA-DC. Experiments show that when used for pre-training in a multi-task manner, CC3M-QA-DC can improve the performance with various backbones on various downstream tasks. Furthermore, our generated CC3M-QA-DC can be combined with larger image-text datasets (e.g., CC15M) and achieve competitive results compared with models using much more data. Code and dataset are available at https://github.com/johncaged/OPT_Questioner. Longteng Guo, Handong Li, Xingjian He, Jing Liu 0001 |
ACM Multimedia | 5 |
| 2023 | MAMO: Fine-Grained Vision-Language Representations Learning with Masked Multimodal ModelingabstractMultimodal representation learning has shown promising improvements on various vision-language tasks (e.g., image-text retrieval, visual question answering, etc) and has significantly advanced the development of multimedia information systems. Most existing methods excel at building global-level alignment between vision and language while lacking effective fine-grained image-text interaction. In this paper, we propose a jointly masked multimodal modeling method to learn fine-grained multimodal representations. Our method performs joint masking on image-text input and integrates both implicit and explicit targets for the masked signals to recover. The implicit target provides a unified and debiased objective for vision and language, where the model predicts latent multimodal representations of the unmasked input. The explicit target further enriches the multimodal representations by recovering high-level and semantically meaningful information: momentum visual features of image patches and concepts of word tokens. Through such a masked modeling process, our model not only learns fine-grained multimodal interaction, but also avoids the semantic gap between high-level representations and low-or mid-level prediction targets (e.g., image pixels, discrete vision tokens), thus producing semantically rich multimodal representations that perform well on both zero-shot and fine-tuned settings. Our pre-trained model (named MAMO) achieves state-of-the-art performance on various downstream vision-language tasks, including image-text retrieval, visual question answering, visual reasoning, and weakly-supervised visual grounding. Zijia Zhao, Longteng Guo, Xingjian He, Shuai Shao 0005, Zehuan Yuan, Jing Liu 0001 |
SIGIR | 3 |
| 2022 | An Efficient Sampling-Based Attention Network for Semantic SegmentationabstractSelf-attention is widely explored to model long-range dependencies in semantic segmentation. However, this operation computes pair-wise relationships between the query point and all other points, leading to prohibitive complexity. In this paper, we propose an efficient Sampling-based Attention Network which combines a novel sample method with an attention mechanism for semantic segmentation. Specifically, we design a Stochastic Sampling-based Attention Module (SSAM) to capture the relationships between the query point and a stochastic sampled representative subset from a global perspective, where the sampled subset is selected by a Stochastic Sampling Module. Compared to self-attention, our SSAM achieves comparable segmentation performance while significantly reducing computational redundancy. In addition, with the observation that not all pixels are interested in the contextual information, we design a Deterministic Sampling-based Attention Module (DSAM) to sample features from a local region for obtaining the detailed information. Extensive experiments demonstrate that our proposed method can compete or perform favorably against the state-of-the-art methods on the Cityscapes, ADE20K, COCO Stuff, and PASCAL Context datasets. Xingjian He, Jing Liu 0001, Weining Wang 0001, Hanqing Lu |
IEEE Trans. Image Process. | 1 |
| 2021 | Consistent-Separable Feature Representation for Semantic SegmentationabstractCross-entropy loss combined with softmax is one of the most commonly used supervision components in most existing segmentation methods. The softmax loss is typically good at optimizing the inter-class difference, but not good at reducing the intra-class variation, which can be suboptimal for semantic segmentation task. In this paper, we propose a Consistent-Separable Feature Representation Network to model the Consistent-Separable (C-S) features, which are intra-class consistent and inter-class separable, improving the discriminative power of the deep features. Specifically, we develop a Consistent-Separable Feature Learning Module to obtain C-S features through a new loss, called Class-Aware Consistency loss. This loss function is proposed to force the deep features to be consistent among the same class and apart between different classes. Moreover, we design an Adaptive feature Aggregation Module to fuse the C-S features and original features from backbone for the better semantic prediction. We show that compared with various baselines, the proposed method brings consistent performance improvement. Our proposed approach achieves state-of-the-art performance on Cityscapes (82.6% mIoU in test set), ADE20K (46.65% mIoU in validation set), COCO Stuff (41.3% mIoU in validation set) and PASCAL Context (55.9% mIoU in test set). Xingjian He, Jing Liu 0001, Jun Fu 0005, Jinqiao Wang, Hanqing Lu |
AAAI | 1 |
| 2020 | Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent LearningabstractMost image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently, non-autoregressive decoding has been proposed in machine translation to speed up the inference time by generating all words in parallel. Typically, these models use the word-level cross-entropy loss to optimize each word independently. However, such a learning process fails to consider the sentence-level consistency, thus resulting in inferior generation quality of these non-autoregressive models. In this paper, we propose a Non-Autoregressive Image Captioning (NAIC) model with a novel training paradigm: Counterfactuals-critical Multi-Agent Learning (CMAL). CMAL formulates NAIC as a multi-agent reinforcement learning system where positions in the target sequence are viewed as agents that learn to cooperatively maximize a sentence-level reward. Besides, we propose to utilize massive unlabeled images to boost captioning performance. Extensive experiments on MSCOCO image captioning benchmark show that our NAIC model achieves a performance comparable to state-of-the-art autoregressive models, while brings 13.9x decoding speedup. Longteng Guo, Jing Liu 0001, Xingjian He, Jie Jiang 0016, Hanqing Lu |
IJCAI | 4 |
| 2019 | Image fusion method based on simultaneous sparse representation with non-subsampled contourlet transformabstractThe image fusion method based on sparse representation in the single‐scale image domain has produced better fusion results than the classic methods based on multi‐scale analysis nowadays. However, due to the limited number of dictionary atoms, it is difficult to provide an accurate description for image details in the sparse‐representation‐based image fusion methods, and it requires a lot of time. A novel dictionary is constructed with non‐subsampled contourlet transform and sparse representation by using the proposed simultaneous strategy. Then the novel dictionary could combine the sparsity attribute of the learning dictionary with a multi‐scale feature of non‐subsampled contourlet transform. Moreover, the simultaneous strategy is combined with this novel dictionary so that sparse coefficients can be represented with the same dictionary atoms and thus they can be compared in a reasonable and accurate way. Finally, the image fusion method along with this novel dictionary is proposed and named non‐subsampled contourlet transform (NSCT)–simultaneous sparse representation (SSR). Experimental results show that the proposed fusion method NSCT–SSR, with its more excellent fusion effect and better anti‐noise capability, outperforms the existing fusion methods, which are based on both multi‐scale domain and sparse representation in the single‐scale image domain. Guiqing He, Xingjian He, Jun Wang 0078, Jianping Fan 0001 |
IET Comput. Vis. | 3 |