VLDB 2026 Research / reviewers in the wild / expert
Yiming Sun 0003
dblp:36/2423-3
· DBLP profile ↗
9ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-4961-5239ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image FusionabstractInfrared and visible image fusion generates all-weather perception-capable images by combining complementary modalities, enhancing environmental awareness for intelligent unmanned systems. Existing methods either focus on pixel-level fusion while overlooking downstream task adaptability or implicitly learn rigid semantics through cascaded detection/segmentation models, unable to interactively address diverse semantic target perception needs. We propose CtrlFuse, a controllable image fusion framework that enables interactive dynamic fusion guided by mask prompts. The model integrates a multi-modal feature extractor, a reference prompt encoder (RPE), and a prompt-semantic fusion module (PSFM). The RPE dynamically encodes task-specific semantic prompts by fine-tuning pre-trained segmentation models with input mask guidance, while the PSFM explicitly injects these semantics into fusion features. Through synergistic optimization of parallel segmentation and fusion branches, our method achieves mutual enhancement between task performance and fusion quality. Experiments demonstrate state-of-the-art results in both fusion controllability and segmentation accuracy, with the adapted task branch even outperforming the original segmentation model. Yiming Sun 0003, Yuan Ruan, Qinghua Hu, Pengfei Zhu 0001 |
AAAI | 1 |
| 2026 | KSCNet: Exploring KAN and state space model collaboration network for small object detection from UAV imagery
Yiming Sun 0003, Pengfei Zhu 0001, Xinzhong Zhu |
Expert Syst. Appl. | 3 |
| 2026 | TEDFuse: Task-Driven Equivariant Consistency Decomposition Network for Multi-Modal Image FusionabstractMultimodal image fusion integrates infrared and visible images by leveraging their complementary strengths. However, most existing fusion techniques primarily focus on pixel level integration, often neglecting the preservation of semantic consistency between the source and fused images. To address this limitation, we propose TEDFuse, a Task-Driven Equivariant Consistency Decomposition Network that ensures semantic con sistency within the image space and across high-level semantic tasks. TEDFuse incorporates two key components: first, a robust decomposition framework with equivariant consistency, ensuring that the fused image retains consistent transformation properties under shifts, rotations, and reflections, thereby enhancing local detail preservation and global semantic alignment; In addition, a task-driven fusion framework that integrates a segmentation module, reinforcing semantic feature preservation through a semantic loss function and ensuring consistency in downstream tasks such as segmentation and detection. The proposed method not only preserves the semantic coherence of the fused image but also improves performance in high-level tasks, demonstrating superior capability in multimodal fusion for complex visual applications. Extensive experiments validate the effectiveness of TEDFuse by analyzing feature evolution, examining the relationship between fusion quality and task performance, and discussing calibration strategies for infrared-visible image fusion. The code is available at https://github.com/Claire-cxy/TEDFuse. Yiming Sun 0003, Zhen Wang 0033, Hao Cheng 0010, Yongfeng Dong, Pengfei Zhu 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Task-Gated Multi-Expert Collaboration Network for Degraded Multi-Modal Image FusionabstractMulti-modal image fusion aims to integrate complementary information from different modalities to enhance perceptual capabilities in applications such as rescue and security. However, real-world imaging often suffers from degradation issues, such as noise, blur, and haze in visible imaging, as well as stripe noise in infrared imaging, which significantly degrades model performance. To address these challenges, we propose a task-gated multi-expert collaboration network (TG-ECNet) for degraded multi-modal image fusion. The core of our model lies in the task-aware gating and multi-expert collaborative framework, where the task-aware gating operates in two stages: degradation-aware gating dynamically allocates expert groups for restoration based on degradation types, and fusion-aware gating guides feature integration across modalities to balance information retention between fusion and restoration tasks. To achieve this, we design a two-stage training strategy that unifies the learning of restoration and fusion tasks. This strategy resolves the inherent conflict in information processing between the two tasks, enabling all-in-one multi-modal image restoration and fusion. Experimental results demonstrate that TG-ECNet significantly enhances fusion performance under diverse complex degradation conditions and improves robustness in downstream applications. The code is available at https://github.com/LeeX54946/TG-ECNet. Yiming Sun 0003, Pengfei Zhu 0001, Qinghua Hu, Dongwei Ren, Xinzhong Zhu |
ICML | 1 |
| 2024 | Dynamic Brightness Adaptation for Robust Multi-modal Image Fusion
Yiming Sun 0003, Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu |
IJCAI | 1 |
| 2023 | Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionabstractInfrared and visible image fusion aims to integrate comprehensive information from multiple sources to achieve superior performances on various practical tasks, such as detection, over that of a single modality. However, most existing methods directly combined the texture details and object contrast of different modalities, ignoring the dynamic changes in reality, which diminishes the visible texture in good lighting conditions and the infrared contrast in low lighting conditions. To fill this gap, we propose a dynamic image fusion framework with a multi-modal gated mixture of local-to-global experts, termed MoE-Fusion, to dynamically extract effective and comprehensive information from the respective modalities. Our model consists of a Mixture of Local Experts (MoLE) and a Mixture of Global Experts (MoGE) guided by a multi-modal gate. The MoLE performs specialized learning of multi-modal local features, prompting the fused images to retain the local information in a sample-adaptive manner, while the MoGE focuses on the global information that complements the fused image with overall texture detail and contrast. Extensive experiments show that our MoE-Fusion outperforms state-of-the-art methods in preserving multi-modal image texture and contrast through the local-to-global dynamic learning paradigm, and also achieves superior performance on detection tasks. Our code is available: https://github.com/SunYM2020/MoE-Fusion. Bing Cao 0002, Yiming Sun 0003, Pengfei Zhu 0001, Qinghua Hu |
ICCV | 2 |
| 2022 | DetFusion: A Detection-driven Infrared and Visible Image Fusion NetworkabstractInfrared and visible image fusion aims to utilize the complementary information between the two modalities to synthesize a new image containing richer information. Most existing works have focused on how to better fuse the pixel-level details from both modalities in terms of contrast and texture, yet ignoring the fact that the significance of image fusion is to better serve downstream tasks. For object detection tasks, object-related information in images is often more valuable than focusing on the pixel-level details of images alone. To fill this gap, we propose a detection-driven infrared and visible image fusion network, termed DetFusion, which utilizes object-related information learned in the object detection networks to guide multimodal image fusion. We cascade the image fusion network with the detection networks of both modalities and use the detection loss of the fused images to provide guidance on task-related information for the optimization of the image fusion network. Considering that the object locations provide a priori information for image fusion, we propose an object-aware content loss that motivates the fusion model to better learn the pixel-level information in infrared and visible images. Moreover, we design a shared attention module to motivate the fusion network to learn object-specific information from the object detection networks. Extensive experiments show that our DetFusion outperforms state-of-the-art methods in maintaining pixel intensity distribution and preserving texture details. More notably, the performance comparison with state-of-the-art image fusion methods in task-driven evaluation also demonstrates the superiority of the proposed method. Our code will be available: https://github.com/SunYM2020/DetFusion. Yiming Sun 0003, Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu |
ACM Multimedia | 1 |
| 2022 | Drone-Based RGB-Infrared Cross-Modality Vehicle Detection Via Uncertainty-Aware LearningabstractDrone-based vehicle detection aims at detecting vehicle locations and categories in aerial images. It empowers smart city traffic management and disaster relief. Researchers have made a great deal of effort in this area and achieved considerable progress. However, because of the paucity of data under extreme conditions, drone-based vehicle detection remains a challenge when objects are difficult to distinguish, particularly in low-light conditions. To fill this gap, we constructed a large-scale drone-based RGB-infrared vehicle detection dataset called DroneVehicle, which contains 28, 439 RGB-infrared image pairs covering urban roads, residential areas, parking lots, and other scenarios from day to night. Cross-modal images provide complementary information for vehicle detection, but also introduce redundant information. To handle this dilemma, we further propose an uncertainty-aware cross-modality vehicle detection (UA-CMDet) framework to improve detection performance in complex environments. Specifically, we design an uncertainty-aware module using cross-modal intersection over union and illumination estimation to quantify the uncertainty of each object. Our method takes uncertainty as a weight to boost model learning more effectively while reducing bias caused by high-uncertainty objects. For more robust cross-modal integration, we further perform illumination-aware non-maximum suppression during inference. Extensive experiments on our DroneVehicle and two challenging RGB-infrared object detection datasets demonstrated the advanced flexibility and superior performance of UA-CMDet over competing methods. Our code and DroneVehicle will be available:https://github.com/VisDrone/DroneVehicle. Yiming Sun 0003, Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Multi-Drone-Based Single Object Tracking With Agent Sharing NetworkabstractDrones equipped with cameras (UAVs) can dynamically track the target in the air from a broader view compared with static cameras or moving sensors over the ground. However, it is still challenging to accurately track the target using a single drone due to several factors such as appearance variations and severe occlusions. To this end, we collect a newMulti-Drone singleObjectTracking (MDOT) dataset that consists of 92 groups of video clips with 113, 918 high resolution frames taken by two drones and 63 groups of video clips with 145, 875 high resolution frames taken by three drones. Besides, two evaluation metrics are specially designed for multi-drone single object tracking,i.e., automatic fusion score (AFS) and ideal fusion score (IFS). Moreover, the agent sharing network (ASNet) is proposed by integrating self-supervised template sharing, target re-detection, and view-aware fusion of the target from multiple drones into a unified framework, which can improve the tracking accuracy significantly compared with single drone tracking. Extensive experiments on MDOT show that our ASNet significantly outperforms recent state-of-the-art trackers. The dataset can be found inhttps://github.com/VisDrone/MultiDrone. Pengfei Zhu 0001, Jiayu Zheng, Dawei Du, Longyin Wen, Yiming Sun 0003, Qinghua Hu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |