EDBT 2026 Demo / reviewers in the wild / expert
Hui Li 0037
dblp:66/3387-37
· DBLP profile ↗
35ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0003-4550-7879ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 11 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCAFNet: Multimodal stroke medical image synthesis and fusion network based on self attention and cross attention
Liqiang Song, Junli Zhao, Guodong Wang 0001, Hui Li 0037, Yi Li 0031 |
Comput. Vis. Image Underst. | 5 |
| 2026 | A Color Information Driven Collaborative Training of Dual Task Parallel Network for Visible and Thermal Infrared Image Fusion and Saliency Object Detection
Zeyang Zhang 0002, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Muhammad Awais 0001, Josef Kittler |
Int. J. Comput. Vis. | 2 |
| 2026 | EvaNet: Toward More Efficient and Consistent Infrared and Visible Image Fusion AssessmentabstractEvaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to capture the true quality of the fusion results but also are computationally demanding. To address these issues, we propose a unified evaluation framework specifically tailored for image fusion. At its core is a lightweight network designed efficiently to approximate widely used metrics, following a divide-and-conquer strategy. Unlike conventional approaches that directly assess similarity between fused and source images, we first decompose the fusion result into infrared and visible components. The evaluation model is then used to measure the degree of information preservation in these separated components, effectively disentangling the fusion evaluation process. During training, we incorporate a contrastive learning strategy and inform our evaluation model by perceptual scene assessment provided by a large language model. Last, we propose the first consistency evaluation framework, which measures the alignment between image fusion metrics and human visual perception, using both independent no-reference scores and downstream tasks performance as objective references. Extensive experiments show that our learning-based evaluation paradigm delivers both superior efficiency (up to 1,000 times faster) and greater consistency across a range of standard image fusion benchmarks. Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Tao Zhou 0002, Hui Li 0037, Zhangyong Tang, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Full-combination contrastive learning for multi-view clustering
Zhe Chen 0018, Heng Liu 0003, Hui Li 0037, Tianyang Xu 0001 |
Pattern Recognit. | 4 |
| 2026 | Visual complexity guided diffusion defender for video object tracking and recognition
Shao-Chuan Zhao, Tianyang Xu 0001, Hui Li 0037, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 3 |
| 2026 | Text-guided medical image fusion using unbalanced optimal transport: Semantic alignment and cross-modal interaction
Liqiang Song, Guodong Wang 0001, Junli Zhao, Hui Li 0037, Yi Li 0031 |
Signal Process. | 5 |
| 2026 | BusReF: Infrared-Visible Images Registration and Fusion Focus on Reconstructible Area Using One Set of FeaturesabstractIn multi-modal imaging scenarios, the misalignment of images presents a persistent challenge. Conventional image fusion algorithms, aiming to enhance the performance of downstream vision tasks, presuppose strictly registered inputs to achieve satisfactory results. To relax this assumption, a common approach is to register the images first; however, existing multi-modal registration methods are often hindered by complex architectures and a heavy reliance on semantic information. This article proposes BusRef, a unified framework that jointly addresses image registration and fusion, with a specific focus on the Infrared-Visible Image Registration and Fusion (IVRF) task. Within this framework, unaligned image pairs are processed through three sequential stages: coarse registration, fine registration, and fusion. We demonstrate that this integrated approach enables more robust and accurate IVRF. Key to our framework is a novel training and evaluation strategy that employs masks to mitigate the influence of non-reconstructible regions on the loss function, thereby significantly improving the model’s accuracy and robustness. Furthermore, we introduce a gradient-aware fusion network designed to effectively preserve complementary information from both modalities. Comprehensive experiments demonstrate that BusRef achieves superior performance when compared against various state-of-the-art registration and fusion algorithms. Our code is available at https://github.com/Yukarizz/BusReF . Zeyang Zhang 0002, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Congcong Bian, Josef Kittler |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent SystemabstractHaoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, Nanqing Dong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haoyang Su 0001, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng 0001, Jinzhe Li, Biqing Qi, Hui Li 0037, Wanli Ouyang, Philip Torr 0001, Bowen Zhou 0002, Nanqing Dong |
ACL (1) | 9 |
| 2025 | One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image FusionabstractAdvanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction through pixel-level supervision. This new paradigm provides strong guidance for unsupervised multimodal fusion without relying on abstract semantics, enhancing task-shared feature learning for broader applicability. Owning to the hybrid image features and enhanced universal representations, the proposed GIFNet supports diverse fusion tasks, achieving high performance across both seen and unseen scenarios with a single model. Uniquely, experimental results reveal that our framework also supports single-modality enhancement, offering superior flexibility for practical applications. Our code will be available at https://github.com/AWCXV/GIFNet. Chunyang Cheng, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Zhangyong Tang, Hui Li 0037, Zeyang Zhang 0002, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
CVPR | 6 |
| 2025 | Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive LawsabstractExisting infrared and visible image fusion methods often face the dilemma of balancing modal information. Generative fusion methods reconstruct fused images by learning from data distributions, but their generative capabilities remain limited. Moreover, the lack of interpretability in modal information selection further affects the reliability and consistency of fusion results in complex scenarios. This manuscript revisits the essence of generative image fusion under the inspiration of human cognitive laws and proposes a novel infrared and visible image fusion method, termed HCLFuse. First, HCLFuse investigates the quantification theory of information mapping in unsupervised fusion networks, which leads to the design of a multi-scale mask-regulated variational bottleneck encoder. This encoder applies posterior probability modeling and information decomposition to extract accurate and concise low-level modal information, thereby supporting the generation of high-fidelity structural details. Furthermore, the probabilistic generative capability of the diffusion model is integrated with physical laws, forming a time-varying physical guidance mechanism that adaptively regulates the generation process at different stages, thereby enhancing the ability of the model to perceive the intrinsic structure of data and reducing dependence on data quality. Experimental results show that the proposed method achieves state-of-the-art fusion performance in qualitative and quantitative evaluations across multiple datasets and significantly improves semantic segmentation metrics. This fully demonstrates the advantages of this generative image fusion method, drawing inspiration from human cognition, in enhancing structural consistency and detail quality. Xiaoqing Luo, Zhancheng Zhang, Hui Li 0037, Rui Wang 0050, Zhenhua Feng 0001, Xiaoning Song |
NeurIPS | 5 |
| 2025 | FusionBooster: A Unified Image Fusion Boosting Paradigm
Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Hui Li 0037, Xi Li 0001, Josef Kittler |
Int. J. Comput. Vis. | 4 |
| 2025 | SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion
Huan Kang, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Chunyang Cheng, Josef Kittler |
Int. J. Comput. Vis. | 2 |
| 2025 | OCCO: LVM-Guided Infrared and Visible Image Fusion Framework Based on Object-Aware and Contextual Contrastive Learning
Hui Li 0037, Congcong Bian, Zeyang Zhang 0002, Xiaoning Song, Xi Li 0001, Xiaojun Wu 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Anchor Graph Learning with Double Noise Removal for Multi-View Clustering
Zhe Chen 0018, Mingzhi Zhu, Hui Li 0037, Tianyang Xu 0001 |
Neural Networks | 3 |
| 2025 | MMAE: A universal image fusion method via mask attention mechanism
Lixing Fang, Junli Zhao, Zhenkuan Pan 0001, Hui Li 0037, Yi Li 0031 |
Pattern Recognit. | 5 |
| 2025 | Deep Discriminative Multi-View ClusteringabstractMulti-view clustering based on deep auto-encoder networks has garnered increasing attention and made significant progress in recent years. However, we argue that most existing methods inadequately explore the discriminability while learning clustering assignments, resulting in models struggling to accurately cluster data, particularly those with ambiguous semantics. To address this problem, we propose a novel framework termed deep discriminative multi-view clustering (DDMvC). This framework is designed to further increase the inter-cluster distances by learning a discriminative projection dictionary with global prior information. To begin with, we enhance the reliability of the dictionary atoms by initializing them with class-specific prototypes derived from concatenated global features across multiple views. Subsequently, we iteratively refine the atoms to guarantee their independence from any specific cluster. Simultaneously, we incorporate contrastive learning for the cluster assignments projected by these atoms, striving for inter-view consistent clustering results. Experimental results on benchmark multi-view datasets demonstrate that our framework achieves the state-of-the-art clustering performance. Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Hui Li 0037, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | S4Fusion: Saliency-Aware Selective State Space Model for Infrared and Visible Image FusionabstractThe preservation and the enhancement of complementary features between modalities are crucial for multi-modal image fusion and downstream vision tasks. However, existing methods are limited to local receptive fields (CNNs) or lack comprehensive utilization of spatial information from both modalities during interaction (transformers), which results in the inability to effectively retain useful information from both modalities in a comparative manner. Consequently, the fused images may exhibit a bias towards one modality, failing to adaptively preserve salient targets from all sources. Thus, a novel fusion framework (S4Fusion) based on the Saliency-aware Selective State Space is proposed. S4Fusion introduces the Cross-Modal Spatial Awareness Module (CMSA), which is designed to simultaneously capture global spatial information from all input modalities and promote effective cross-modal interaction. This enables a more comprehensive representation of complementary features. Furthermore, to guide the model in adaptively preserving salient objects, we propose a novel perception-enhanced loss function. This loss aims to enhance the retention of salient features by minimizing ambiguity or uncertainty, as measured at a pre-trained model's decision layer, within the fused images. The code is available at https://github.com/zipper112/S4Fusion. Haolong Ma, Hui Li 0037, Chunyang Cheng, Gaoang Wang, Xiaoning Song, Xiaojun Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | I Know How You Move: Explicit Motion Estimation for Human Action RecognitionabstractEnabled by hierarchical convolutions and nonlinear mappings, recent action recognition studies have continuously boosted performance with spatiotemporal modelling. In general, motion clues are essential in video-oriented tasks, while existing approaches aggregate the spatial and temporal signatures via specially designed modules in the middle or output stages. To highlight the privilege provided by temporal motions, in this paper, we propose a simple but effectiveMOTion Estimator(MOTE) to generate the motion patterns from every single frame, avoiding complex dense-frame input. In particular, MOTE follows an encoder-decoder structure, which takes the short-term motion features generated by the pretrained dense-frame network as the learning target. The spatial information of a single frame is utilized to estimate the instantaneous motion appearance. It can support the expression of vulnerable regions, such as the ‘hand’ in ‘waving hands’, which would otherwise be suppressed in the feature maps as the ‘hand’ suffers from motion blur. The training process of MOTE is independent of the action recognition system. Therefore, the trained MOTE can be transplanted to the input-end of existing action recognition methods to provide instantaneous motion estimation as feature enhancement according to practical requirements. Our experiments performed on Something-Something V1, V2, Kinetics-400, and Diving48 verify the effectiveness of the proposed method. Xiaojun Wu 0001, Hui Li 0037, Tianyang Xu 0001, Cong Wu 0006 |
IEEE Trans. Multim. | 3 |
| 2024 | SMFuse: Two-Stage Structural Map Aware Network for Multi-focus Image Fusion
Tianyu Shen, Hui Li 0037, Chunyang Cheng, Xiaoning Song |
ICPR (22) | 2 |
| 2024 | IFFusion: Illumination-Free Fusion Network for Infrared and Visible Images
Hui Li 0037, Tianyang Xu 0001, Zeyang Zhang 0002, Xiaojun Wu 0001 |
ICPR (5) | 2 |
| 2024 | CoMoFusion: Fast and High-Quality Fusion of Infrared and Visible Image with Consistency Model
Zhiming Meng, Hui Li 0037, Zeyang Zhang 0002, Yunlong Yu 0001, Xiaoning Song, Xiaojun Wu 0001 |
PRCV (8) | 2 |
| 2024 | An unsupervised multi-focus image fusion method via dual-channel convolutional network and discriminator
Lixing Fang, Junli Zhao, Zhenkuan Pan 0001, Hui Li 0037, Yi Li 0031 |
Comput. Vis. Image Underst. | 5 |
| 2024 | UUD-Fusion: An unsupervised universal image fusion approach via generative diffusion model
Lixing Fang, Junli Zhao, Zhenkuan Pan 0001, Hui Li 0037, Yi Li 0031 |
Comput. Vis. Image Underst. | 5 |
| 2024 | A method of test case set generation in the commutativity test of reduce functions
Xiangyu Mu, Lei Liu 0049, Hui Li 0037 |
Sci. Comput. Program. | 5 |
| 2024 | APMG: 3D Molecule Generation Driven by Atomic Chemical PropertiesabstractRecently, mask-fill-based 3D Molecular Generation (MG) methods have become very popular in virtual drug design. However, the existing MG methods ignore the chemical properties of atoms and contain inappropriate atomic position training data, which limits their generation capability. To mitigate the above issues, this paper presents a novel mask-fill-based 3D molecule generation model driven by atomic chemical properties (APMG). Specifically, we construct a new attention-MPNN-based encoder and introduce the electronic information into atom representations to enrich chemical properties. Also, a multi-functional classifier is designed to predict the electronic information of each generated atom, guiding the type prediction of elements and bonds. By design, the proposed method uses the chemical properties of atoms and their correlations for high-quality molecule generation. Second, to optimize the atomic position training data, we propose a novel atomic training position generation approach using the Chi-Square distribution. We evaluate our APMG method on the CrossDocked dataset and visualize the docking states of the pockets and generated molecules. The obtained results demonstrate the superiority and merits of APMG over the state-of-the-art approaches. Yang Hua 0002, Zhenhua Feng 0001, Xiaoning Song, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Dongjun Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | LE2Fusion: A Novel Local Edge Enhancement Module for Infrared and Visible Image Fusion
Yongbiao Xiao, Hui Li 0037, Chunyang Cheng, Xiaoning Song |
ICIG (1) | 2 |
| 2023 | LRRNet: A Novel Representation Learning Guided Fusion Network for Infrared and Visible ImagesabstractDeep learning based fusion methods have been achieving promising performance in image fusion tasks. This is attributed to the network architecture that plays a very important role in the fusion process. However, in general, it is hard to specify a good fusion architecture, and consequently, the design of fusion networks is still a black art, rather than science. To address this problem, we formulate the fusion task mathematically, and establish a connection between its optimal solution and the network architecture that can implement it. This approach leads to a novel method proposed in the paper of constructing a lightweight fusion network. It avoids the time-consuming empirical network design by a trial-and-test strategy. In particular we adopt a learnable representation approach to the fusion task, in which the construction of the fusion network architecture is guided by the optimisation algorithm producing the learnable model. The low-rank representation (LRR) objective is the foundation of our learnable model. The matrix multiplications, which are at the heart of the solution are transformed into convolutional operations, and the iterative process of optimisation is replaced by a special feed-forward network. Based on this novel network architecture, an end-to-end lightweight fusion network is constructed to fuse infrared and visible light images. Its successful training is facilitated by a detail-to-semantic information loss function proposed to preserve the image details and to enhance the salient features of the source images. Our experiments show that the proposed fusion network exhibits better fusion performance than the state-of-the-art fusion methods on public datasets. Interestingly, our network requires a fewer training parameters than other existing methods. Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Jiwen Lu, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Generalized n-Dimensional Rigid Registration: Theory and ApplicationsabstractThe generalized rigid registration problem in high-dimensional Euclidean spaces is studied. The loss function is minimized with an equivalent error formulation by the Cayley formula. The closed-form linear least-square solution to such a problem is derived which generates the registration covariances, i.e., uncertainty information of rotation and translation, providing quite accurate probabilistic descriptions. Simulation results indicate the correctness of the proposed method and also present its efficiency on computation-time consumption, compared with previous algorithms using singular value decomposition (SVD) and linear matrix inequality (LMI). The proposed scheme is then applied to an interpolation problem on the special Euclidean group SE(n) with covariance-preserving functionality. Finally, experiments on covariance-aided Lidar mapping show practical superiority in robotic navigation. Jin Wu 0002, Miaomiao Wang 0001, Hassen Fourati, Hui Li 0037, Yilong Zhu, Chengxi Zhang, Yi Jiang 0007, Xiangcheng Hu, Ming Liu 0001 |
IEEE Trans. Cybern. | 4 |
| 2021 | MSC-Fuse: An Unsupervised Multi-scale Convolutional Fusion Framework for Infrared and Visible Image
Guo-Yang Chen, Xiaojun Wu 0001, Hui Li 0037, Tianyang Xu 0001 |
ICIG (1) | 3 |
| 2020 | Subspace Clustering via Joint Unsupervised Feature SelectionabstractAny high-dimensional data arising from practical applications usually contains irrelevant features that may impact on the performance of existing subspace clustering methods. This paper proposes a novel subspace clustering method which reconstructs the feature matrix by the means of unsupervised feature selection (UFS) to achieve a better dictionary for subspace clustering (SC). Different from most existing clustering methods, the proposed approach uses the reconstructed feature matrix as the dictionary rather than the original data matrix. As the feature matrix reconstructed by representative features is more discriminative and closer to the ground-truth, it results in improved performance. The corresponding non-convex optimization problem is effectively solved using the half-quadratic and augmented Lagrange multiplier methods. Extensive experiments on four real datasets demonstrate the effectiveness of the proposed method. Wenhua Dong, Xiaojun Wu 0001, Hui Li 0037, Zhenhua Feng 0001, Josef Kittler |
ICPR | 3 |
| 2020 | MDLatLRR: A Novel Decomposition Method for Infrared and Visible Image FusionabstractImage decomposition is crucial for many image processing tasks, as it allows to extract salient features from source images. A good image decomposition method could lead to a better performance, especially in image fusion tasks. We propose a multi-level image decomposition method based on latent low-rank representation(LatLRR), which is called MDLatLRR. This decomposition method is applicable to many image processing fields. In this paper, we focus on the image fusion task. We build a novel image fusion framework based on MDLatLRR which is used to decompose source images into detail parts(salient features) and base parts. A nuclear-norm based fusion strategy is used to fuse the detail parts and the base parts are fused by an averaging strategy. Compared with other state-of-the-art fusion methods, the proposed algorithm exhibits better fusion performance in both subjective and objective evaluation. Hui Li 0037, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Image Process. | 1 |
| 2019 | MSDNet for Medical Image Fusion
Xu Song, Xiaojun Wu 0001, Hui Li 0037 |
ICIG (2) | 3 |
| 2019 | DenseFuse: A Fusion Approach to Infrared and Visible ImagesabstractIn this paper, we present a novel deep learning architecture for infrared and visible images fusion problem. In contrast to conventional convolutional networks, our encoding network is combined by convolutional layers, fusion layer and dense block in which the output of each layer is connected to every other layer. We attempt to use this architecture to get more useful features from source images in encoding process. And two fusion layers(fusion strategies) are designed to fuse these features. Finally, the fused image is reconstructed by decoder. Compared with existing fusion methods, the proposed fusion method achieves state-of-the-art performance in objective and subjective assessment. Hui Li 0037, Xiaojun Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Infrared and Visible Image Fusion using a Deep Learning FrameworkabstractIn recent years, deep learning has become a very active research tool which is used in many image processing fields. In this paper, we propose an effective image fusion method using a deep learning framework to generate a single image which contains all the features from infrared and visible images. First, the source images are decomposed into base parts and detail content. Then the base parts are fused by weighted-averaging. For the detail content, we use a deep learning network to extract multi-layer features. Using these features, we use$l_{1}$-norm and weighted-average strategy to generate several candidates of the fused detail content. Once we get these candidates, the max selection strategy is used to get the final fused detail content. Finally, the fused image will be reconstructed by combining the fused base part and the detail content. The experimental results demonstrate that our proposed method achieves state-of-the-art performance in both objective assessment and visual quality. The Code of our fusion method is available at https://github.com/exceptionLi/imagefusion_deeplearning. Hui Li 0037, Xiaojun Wu 0001, Josef Kittler |
ICPR | 1 |
| 2017 | Multi-focus Image Fusion Using Dictionary Learning and Low-Rank Representation
Hui Li 0037, Xiaojun Wu 0001 |
ICIG (1) | 1 |