Kunpeng Wang 0005

dblp:86/1646-5 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-2788-7583ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Taming Cascaded Mixture-of-Experts for Modality-missing Multi-modal Salient Object Detection
abstract
Multi-modal Salient Object Detection (SOD) shows an improvement over its uni-modal counterpart by exploiting the complementary benefits between modalities. However, this improvement relies on complete multi-modal information, which is difficult to be guaranteed in practice due to sensor failures and transmission errors. To address this issue, we propose a robust multi-modal SOD framework that enhances the adaptability to modality-missing conditions, while maintaining comparable performance in the modality-complete condition. Nevertheless, flexibly handling modality-missing and modality-complete cases and integrating their corresponding multi-modal features in a unified framework is non-trivial. To this end, we achieve this framework by designing a Cascaded Mixture-of-Experts (CMoE) network that sequentially incorporates missing-aware and multi-modal MoE. Specifically, the missing-aware MoE employs three modality-reconstruction experts with a soft router to adaptively reconstruct feature representations for both missing and available modalities, assisted by an expert modulation loss that guides the router to assign expert weights according to missing conditions. The multi-modal MoE adopts two homogeneous uni-modal experts with learned modality-specific knowledge tailored for integrating modality features, which are dynamically combined via the soft router. The cascaded architecture fully empowers CMoE with the flexibility across varying input cases. Extensive experiments on modality-missing and modality-complete conditions demonstrate the effectiveness of the proposed method.
Kunpeng Wang 0005, Feifan Sun, Keke Chen
AAAI1
2026 Pixel-Level RGBT Fusion Tracking via Heterogeneous Multi-Expert Distillation and Decoupled Representation Learning
abstract
Pixel-level fusion is widely considered a lightweight yet limited strategy in RGB-Thermal (RGBT) tracking due to its shallow representational capacity. However, its actual limitations and potential remain largely unexplored. We systematically analyze fusion location, modality alignment, and tracking performance, revealing that despite lower modality gaps than feature-level fusion, pixel-level fusion lacks task-relevant discrimination, restricting its effectiveness. In this paper, we propose the Task-driven Pixel-level Fusion tracker (TPF), which preserves the efficiency of early fusion while enhancing discriminative capacity. Central to TPF is a lightweight pixel fusion adapter that ensures real-time image fusion with only 14.3KB extra parameters over the baseline at inference. To enhance its limited representational capacity, we propose a task-driven progressive learning framework consisting of two key stages. First, a heterogeneous multi-expert distillation scheme adaptively transfers image fusion knowledge from diverse models under tracking-guided evaluation, mitigating the generalization limitations of single-teacher distillation across varied tracking scenarios. Second, to overcome limited task discrimination caused by sparse, target-focused tracking supervision, we propose a decoupled representation learning strategy that offers dense, complementary guidance to improve target-background separation and fusion quality. A nearest-neighbor dynamic template update further enhances robustness to appearance changes. Extensive experiments on four RGBT tracking benchmarks show that TPF achieves competitive accuracy and speed, outperforming both feature-level and existing pixel-level fusion methods, offering new insights into efficient RGBT tracking.
Andong Lu, Yuanzhi Guo, Kunpeng Wang 0005, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Image Process.3
2025 Alignment-Free RGB-T Salient Object Detection: A Large-Scale Dataset and Progressive Correlation Network
abstract
Alignment-free RGB-Thermal (RGB-T) salient object detection (SOD) aims to achieve robust performance in complex scenes by directly leveraging the complementary information from unaligned visible-thermal image pairs, without requiring manual alignment. However, the labor-intensive process of collecting and annotating image pairs limits the scale of existing benchmarks, hindering the advancement of alignment-free RGB-T SOD. In this paper, we construct a large-scale and high-diversity unaligned RGB-T SOD dataset named UVT20K, comprising 20,000 image pairs, 407 scenes, and 1256 object categories. All samples are collected from real-world scenarios with various challenges, such as low illumination, image clutter, complex salient objects, and so on. To support the exploration for further research, each sample in UVT20K is annotated with a comprehensive set of ground truths, including saliency masks, scribbles, boundaries, and challenge attributes. In addition, we propose a Progressive Correlation Network (PCNet), which models inter- and intra-modal correlations on the basis of explicit alignment to achieve accurate predictions in unaligned image pairs. Extensive experiments conducted on two unaligned three weakly aligned three aligned datasets demonstrate the effectiveness of our method.
Kunpeng Wang 0005, Keke Chen, Chenglong Li 0002, Zhengzheng Tu, Bin Luo 0001
AAAI1
2025 Hierarchical semantics guided multi-scale correlation network for alignment-free red-green-blue and thermal salient object detection
Chengmei Han, Lei Liu 0049, Kunpeng Wang 0005
Eng. Appl. Artif. Intell.3
2025 Erasure-based interaction network for red-green-blue and thermal object detection and a unified benchmark
Qishun Wang, Zhengzheng Tu, Chenglong Li 0002, Hongshun Wang, Kunpeng Wang 0005
Eng. Appl. Artif. Intell.5
2025 Unified-Modal Salient Object Detection via Adaptive Prompt Learning
abstract
Existing single-modal and multi-modal salient object detection (SOD) methods focus on designing specific architectures tailored for their respective tasks. However, developing completely different models for different tasks leads to labor and time consumption, as well as high computational and practical deployment costs. In this paper, we attempt to address both single-modal and multi-modal SOD in a unified framework called UniSOD, which fully exploits the overlapping prior knowledge between different tasks. Nevertheless, assigning appropriate strategies to modality variable inputs is challenging. To this end, UniSOD learns modality-aware prompts with task-specific hints through adaptive prompt learning, which are seamlessly plugged into the proposed pre-trained baseline SOD model to handle corresponding tasks, while only requiring few learnable parameters compared to training the entire model from scratch. In particular, each modality-aware prompt is solely generated from a homogeneous switchable prompt generation (SPG) block, which adaptively performs structural switching based on single-modal and multi-modal inputs without manual intervention, ensuring that the framework can effectively handle diverse input cases (e.g., RGB-only, RGB-D, RGB-T) with a unified approach. Through end-to-end joint training, UniSOD achieves ovrall competitive performance on 14 benchmark datasets, demonstrating its ability to efficiently unify single-modal and multi-modal SOD tasks. Code has been available athttps://github.com/Angknpng/UniSOD
Kunpeng Wang 0005, Zhengzheng Tu, Chenglong Li 0002, Zhengyi Liu, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Learning Adaptive Fusion Bank for Multi-Modal Salient Object Detection
abstract
Multi-modal salient object detection (MSOD) aims to boost saliency detection performance by integrating visible sources with depth or thermal infrared ones. Existing methods generally design different fusion schemes to handle certain issues or challenges. Although these fusion schemes are effective at addressing specific issues or challenges, they may struggle to handle multiple complex challenges simultaneously. To solve this problem, we propose a novel adaptive fusion bank that makes full use of the complementary benefits from a set of basic fusion schemes to handle different challenges simultaneously for robust MSOD. We focus on handling five major challenges in MSOD, namely center bias, scale variation, image clutter, low illumination, and thermal crossover or depth ambiguity. The fusion bank proposed consists of five representative fusion schemes, which are specifically designed based on the characteristics of each challenge, respectively. The bank is scalable, and more fusion schemes could be incorporated into the bank for more challenges. To adaptively select the appropriate fusion scheme for multi-modal input, we introduce an adaptive ensemble module that forms the adaptive fusion bank, which is embedded into hierarchical layers for sufficient fusion of different source data. Moreover, we design an indirect interactive guidance module to accurately detect salient hollow objects via the skip integration of high-level semantic information and low-level spatial details. Extensive experiments on three RGBT datasets and seven RGBD datasets demonstrate that the proposed method achieves the outstanding performance compared to the state-of-the-art methods.
Kunpeng Wang 0005, Zhengzheng Tu, Chenglong Li 0002, Cheng Zhang 0010, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Alignment-Free RGBT Salient Object Detection: Semantics-Guided Asymmetric Correlation Network and a Unified Benchmark
abstract
RGB and Thermal (RGBT) Salient Object Detection (SOD) aims to achieve high-quality saliency prediction by exploiting the complementary information of visible and thermal image pairs, which are initially captured in an unaligned manner. However, existing methods are tailored for manually aligned image pairs, which are labor-intensive, and directly applying these methods to original unaligned image pairs could significantly degrade their performance. In this paper, we make the first attempt to address RGBT SOD for initially captured RGB and thermal image pairs without manual alignment. Specifically, we propose a Semantics-guided Asymmetric Correlation Network (SACNet) that consists of two novel components: 1) an asymmetric correlation module utilizing semantics-guided attention to model cross-modal correlations specific to unaligned salient regions; 2) an associated feature sampling module to sample relevant thermal features according to the corresponding RGB features for multi-modal feature integration. In addition, we construct a unified benchmark dataset called UVT2000, containing 2000 RGB and thermal image pairs directly captured from various real-world scenes without any alignment, to facilitate research on alignment-free RGBT SOD. Extensive experiments on both aligned and unaligned datasets demonstrate the effectiveness and superior performance of our method.
Kunpeng Wang 0005, Danying Lin, Chenglong Li 0002, Zhengzheng Tu, Bin Luo 0001
IEEE Trans. Multim.1
2023 Multimodal salient object detection via adversarial learning with collaborative generator
Zhengzheng Tu, Wenfang Yang, Kunpeng Wang 0005, Amir Hussain 0001, Bin Luo 0001, Chenglong Li 0002
Eng. Appl. Artif. Intell.3