Chunyang Cheng

dblp:244/1912 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 EvaNet: Toward More Efficient and Consistent Infrared and Visible Image Fusion Assessment
abstract
Evaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to capture the true quality of the fusion results but also are computationally demanding. To address these issues, we propose a unified evaluation framework specifically tailored for image fusion. At its core is a lightweight network designed efficiently to approximate widely used metrics, following a divide-and-conquer strategy. Unlike conventional approaches that directly assess similarity between fused and source images, we first decompose the fusion result into infrared and visible components. The evaluation model is then used to measure the degree of information preservation in these separated components, effectively disentangling the fusion evaluation process. During training, we incorporate a contrastive learning strategy and inform our evaluation model by perceptual scene assessment provided by a large language model. Last, we propose the first consistency evaluation framework, which measures the alignment between image fusion metrics and human visual perception, using both independent no-reference scores and downstream tasks performance as objective references. Extensive experiments show that our learning-based evaluation paradigm delivers both superior efficiency (up to 1,000 times faster) and greater consistency across a range of standard image fusion benchmarks.
Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Tao Zhou 0002, Hui Li 0037, Zhangyong Tang, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion
abstract
Advanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction through pixel-level supervision. This new paradigm provides strong guidance for unsupervised multimodal fusion without relying on abstract semantics, enhancing task-shared feature learning for broader applicability. Owning to the hybrid image features and enhanced universal representations, the proposed GIFNet supports diverse fusion tasks, achieving high performance across both seen and unseen scenarios with a single model. Uniquely, experimental results reveal that our framework also supports single-modality enhancement, offering superior flexibility for practical applications. Our code will be available at https://github.com/AWCXV/GIFNet.
Chunyang Cheng, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Zhangyong Tang, Hui Li 0037, Zeyang Zhang 0002, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler
CVPR1
2025 Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and Benchmarking
abstract
Unifying multiple multi-modal visual object tracking (MMVOT) tasks draws increasing attention due to the complementary nature of different modalities in building robust tracking systems. Existing practices mix all data sensor types in a single training procedure, structuring a parallel paradigm from the data-centric perspective and aiming for a global optimum on the joint distribution of the involved tasks. However, the absence of a unified benchmark where all types of data coexist forces evaluations on separated benchmarks, causing inconsistency between training and testing, thus leading to performance degradation. To address these issues, this work advances in two aspects: A unified benchmark, coined as UniBench300, is introduced to bridge the inconsistency by incorporating multiple task data, reducing inference passes from three to one and cutting time consumption by 27%. The unification process is reformulated in a serial format, progressively integrating new tasks. In this way, the performance degradation can be specified as knowledge forgetting of previous tasks, which naturally aligns with the philosophy of continual learning (CL), motivating further exploration of injecting CL into the unification process. Extensive experiments conducted on two baselines and four benchmarks demonstrate the significance of UniBench300 and the superiority of CL in supporting a stable unification process. Moreover, while conducting dedicated analyses, the performance degradation is found to be negatively correlated with network capacity. Additionally, modality discrepancies contribute to varying degradation levels across tasks (RGBT > RGBD > RGBE in MMVOT), offering valuable insights for future multi-modal vision research. Source codes and the proposed benchmark is available at https://github.com/Zhangyong-Tang/UniBench300.
Zhangyong Tang, Tianyang Xu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Tao Zhou 0002, Xiaojun Wu 0001, Josef Kittler
ACM Multimedia4
2025 MambaGuard: A CLIP-Mamba Approach for OOD Generated Image Detection
Xinchang Wang, Yuechen Zhang, Wenyao Qiu, Chunyang Cheng
PRCV (4)5
2025 FusionBooster: A Unified Image Fusion Boosting Paradigm
Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Hui Li 0037, Xi Li 0001, Josef Kittler
Int. J. Comput. Vis.1
2025 SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion
Huan Kang, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Chunyang Cheng, Josef Kittler
Int. J. Comput. Vis.6
2025 S4Fusion: Saliency-Aware Selective State Space Model for Infrared and Visible Image Fusion
abstract
The preservation and the enhancement of complementary features between modalities are crucial for multi-modal image fusion and downstream vision tasks. However, existing methods are limited to local receptive fields (CNNs) or lack comprehensive utilization of spatial information from both modalities during interaction (transformers), which results in the inability to effectively retain useful information from both modalities in a comparative manner. Consequently, the fused images may exhibit a bias towards one modality, failing to adaptively preserve salient targets from all sources. Thus, a novel fusion framework (S4Fusion) based on the Saliency-aware Selective State Space is proposed. S4Fusion introduces the Cross-Modal Spatial Awareness Module (CMSA), which is designed to simultaneously capture global spatial information from all input modalities and promote effective cross-modal interaction. This enables a more comprehensive representation of complementary features. Furthermore, to guide the model in adaptively preserving salient objects, we propose a novel perception-enhanced loss function. This loss aims to enhance the retention of salient features by minimizing ambiguity or uncertainty, as measured at a pre-trained model's decision layer, within the fused images. The code is available at https://github.com/zipper112/S4Fusion.
Haolong Ma, Hui Li 0037, Chunyang Cheng, Gaoang Wang, Xiaoning Song, Xiaojun Wu 0001
IEEE Trans. Image Process.3
2025 Revisiting RGBT Tracking Benchmarks From the Perspective of Modality Validity: A New Benchmark, Problem, and Solution
abstract
RGBT tracking draws increasing attention because of its robustness in multi-modal warranting (MMW) scenarios, such as nighttime and adverse weather conditions, where relying on a single sensing modality fails to ensure stable tracking results. However, existing benchmarks predominantly contain videos collected in common scenarios where both RGB and thermal infrared (TIR) information are of sufficient quality. This weakens the representativeness of existing benchmarks in severe imaging conditions, leading to tracking failures in MMW scenarios. To bridge this gap, we present a new benchmark considering the modality validity, MV-RGBT, captured specifically from MMW scenarios where either RGB (extreme illumination) or TIR (thermal truncation) modality is invalid. Hence, it is further divided into two subsets according to the valid modality, offering a new compositional perspective for evaluation and providing valuable insights for future designs. Moreover, MV-RGBT is the most diverse benchmark of its kind, featuring 36 different object categories captured across 19 distinct scenes. Furthermore, considering severe imaging conditions in MMW scenarios, a new problem is posed in RGBT tracking, named 'when to fuse', to stimulate the development of fusion strategies for such scenarios. To facilitate its discussion, we propose a new solution with a mixture of experts, named MoETrack, where each expert generates independent tracking results along with a confidence score. Extensive results demonstrate the significant potential of MV-RGBT in advancing RGBT tracking and elicit the conclusion that fusion is not always beneficial, especially in MMW scenarios. Besides, MoETrack achieves state-of-the-art results on several benchmarks, including MV-RGBT, GTOT, and LasHeR. Source codes and benchmarks are available at https://github.com/Zhangyong-Tang/MVRGBT.
Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Zhenhua Feng 0001, Josef Kittler
IEEE Trans. Image Process.5
2024 SMFuse: Two-Stage Structural Map Aware Network for Multi-focus Image Fusion
Tianyu Shen, Hui Li 0037, Chunyang Cheng, Xiaoning Song
ICPR (22)3
2024 MMDRFuse: Distilled Mini-Model with Dynamic Refresh for Multi-Modality Image Fusion
abstract
In recent years, Multi-Modality Image Fusion (MMIF) has been applied to many fields, which has attracted many scholars to endeavour to improve the fusion performance. However, the prevailing focus has predominantly been on the architecture design, rather than the training strategies. As a low-level vision task, image fusion is supposed to quickly deliver output images for observation and supporting downstream tasks. Thus, superfluous computational and storage overheads should be avoided. In this work, a lightweight Distilled Mini-Model with a Dynamic Refresh strategy (MMDRFuse) is proposed to achieve this objective. To pursue model parsimony, an extremely small convolutional network with a total of 113 trainable parameters (0.44 KB) is obtained by three carefully designed supervisions. First, digestible distillation is constructed by emphasising external spatial feature consistency, delivering soft supervision with balanced details and saliency for the target network. Second, we develop a comprehensive loss to balance the pixel, gradient, and perception clues from the source images. Third, an innovative dynamic refresh training strategy is used to collaborate history parameters and current supervision during training, together with an adaptive adjust function to optimise the fusion network. Extensive experiments on several public datasets demonstrate that our method exhibits promising advantages in terms of model efficiency and complexity, with superior performance in multiple image fusion tasks and downstream pedestrian detection application. The code of this work is publicly available at https://github.com/yanglinDeng/MMDRFuse.
Yanglin Deng, Tianyang Xu 0001, Chunyang Cheng, Xiaojun Wu 0001, Josef Kittler
ACM Multimedia3
2024 Learning Feature Restoration Transformer for Robust Dehazing Visual Object Tracking
Tianyang Xu 0001, Yifan Pan, Zhenhua Feng 0001, Xuefeng Zhu 0003, Chunyang Cheng, Xiaojun Wu 0001, Josef Kittler
Int. J. Comput. Vis.5
2023 LE2Fusion: A Novel Local Edge Enhancement Module for Infrared and Visible Image Fusion
Yongbiao Xiao, Hui Li 0037, Chunyang Cheng, Xiaoning Song
ICIG (1)3
2019 Optimizing Hospital Emergency Department Layout via Multiobjective Tabu Search
abstract
Hospital department layout problems (HDLPs) are significant in enhancing service quality and reducing patients' travel distance and time. Their studies are scarce in comparison with those for facility layout problems in manufacturing systems. Existing approaches to HDLPs usually adopt simplified models and thus gain very limited applications in a real world. HDLPs typically involve multiple objectives that may conflict with each other. There have been no studies on their multiobjective heuristic approaches to our best knowledge. In this paper, we propose multiobjective tabu search (MTS) for a real-world HDLP. Beside the frequently used objective of flow cost, ensuring the closeness among certain departments is introduced as another one. A solution coding scheme is designed to represent a solution. A penalty function is devised to handle infeasible solutions. Local search is integrated into tabu search to optimize the assignment of departments. Experiment results show that MTS is able to produce Pareto solutions that outperform those of the comparative method. Compared to the actually implemented layout, solutions produced by MTS can save about 5%-15% patients' travel time (distance).
Xingquan Zuo, Xuewen Huang, MengChu Zhou, Chunyang Cheng, Xinchao Zhao, Zhishuo Liu
IEEE Trans Autom. Sci. Eng.5