EDBT 2026 Demo / reviewers in the wild / expert
Guanyao Wu
dblp:317/5208
· DBLP profile ↗
15ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-7956-1753ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discovering a nimble network for misaligned multi-exposure image fusion
Xiaoshan Liu, Guanyao Wu, Yichuan Peng, Weiqiang Kong, Jinyuan Liu 0001 |
Neurocomputing | 2 |
| 2025 | Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and BeyondabstractMulti-modality image fusion, particularly infrared and visible, plays a crucial role in integrating diverse modalities to enhance scene understanding. Although early research prioritized visual quality, preserving fine details and adapting to downstream tasks remains challenging. Recent approaches attempt task-specific design but rarely achieve "The Best of Both Worlds" due to inconsistent optimization goals. To address these issues, we propose a novel method that leverages the semantic knowledge from the Segment Anything Model (SAM) to Grow the quality of fusion results and Enable downstream task adaptability, namely SAGE. Specifically, we design a Semantic Persistent Attention (SPA) Module that efficiently maintains source information via the persistent repository while extracting high-level semantic priors from SAM. More importantly, to eliminate the impractical dependence on SAM during inference, we introduce a bi-level optimization-driven distillation mechanism with triplet losses, which allow the student network to effectively extract knowledge. Extensive experiments show that our method achieves a balance between high-quality visual results and downstream task adaptability while maintaining practical deployment efficiency. The code is available at https://github.com/RollingPlain/SAGE_IVIF. Guanyao Wu, Hongming Fu, Yichuan Peng, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu |
CVPR | 1 |
| 2025 | TextMEF: Text-guided Prompt Learning for Multi-exposure Image FusionabstractMulti-exposure image fusion~(MEF) aims to integrate a set of low dynamic range images, producing a single image with a higher dynamic range than either one. Despite significant advancements, current MEF approaches still struggle to handle extremely over- or under-exposed conditions, resulting in unsatisfactory visual effects such as hallucinated details and distorted color tones. With this regard, we propose TextMEF, a prompt-driven fusion method enhanced by prompt learning, for multi-exposure image fusion. Specifically, we learn a set of prompts based on text-image similarity among negative and positive samples (over-exposed, under-exposed images, and well-exposed ones). These learned prompts are seamlessly integrated into the loss function, providing high-level guidance for constraining non-uniform exposure regions. Furthermore, we develop a attention Mamba module effectively translates over-/under- exposed regional features into exposure invariant space and ensure them to build efficient long-range dependency to high dynamic range image. Extensive experimental results on three publicly available benchmarks demonstrate that our TextMEF significantly outperforms state-of-the-art approaches in both visual inspection and objective analysis. Jinyuan Liu 0001, Qianjun Huang, Guanyao Wu, Di Wang 0018, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001 |
IJCAI | 3 |
| 2025 | Infrared and Visible Image Fusion: From Data Compatibility to Task AdaptionabstractInfrared-visible image fusion (IVIF) is a fundamental and critical task in the field of computer vision. Its aim is to integrate the unique characteristics of both infrared and visible spectra into a holistic representation. Since 2018, growing amount and diversity IVIF approaches step into a deep-learning era, encompassing introduced a broad spectrum of networks or loss functions for improving visual enhancement. As research deepens and practical demands grow, several intricate issues like data compatibility, perception accuracy, and efficiency cannot be ignored. Regrettably, there is a lack of recent surveys that comprehensively introduce and organize this expanding domain of knowledge. Given the current rapid development, this paper aims to fill the existing gap by providing a comprehensive survey that covers a wide array of aspects. Initially, we introduce a multi-dimensional framework to elucidate the prevalent learning-based IVIF methodologies, spanning topics from basic visual enhancement strategies to data compatibility, task adaptability, and further extensions. Subsequently, we delve into a profound analysis of these new approaches, offering a detailed lookup table to clarify their core ideas. Last but not the least, We also summarize performance comparisons quantitatively and qualitatively, covering registration, fusion and follow-up high-level tasks. Beyond delving into the technical nuances of these learning-based fusion approaches, we also explore potential future directions and open issues that warrant further exploration by the community. Jinyuan Liu 0001, Guanyao Wu, Zhu Liu 0004, Di Wang 0018, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Risheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Hybrid-Supervised Dual-Search: Leveraging Automatic Learning for Loss-Free Multi-Exposure Image FusionabstractMulti-exposure image fusion (MEF) has emerged as a prominent solution to address the limitations of digital imaging in representing varied exposure levels. Despite its advancements, the field grapples with challenges, notably the reliance on manual designs for network structures and loss functions, and the constraints of utilizing simulated reference images as ground truths. Consequently, current methodologies often suffer from color distortions and exposure artifacts, further complicating the quest for authentic image representation. In addressing these challenges, this paper presents a Hybrid-Supervised Dual-Search approach for MEF, dubbed HSDS-MEF, which introduces a bi-level optimization search scheme for automatic design of both network structures and loss functions. More specifically, we harness a unique dual research mechanism rooted in a novel weighted structure refinement architecture search. Besides, a hybrid supervised contrast constraint seamlessly guides and integrates with searching process, facilitating a more adaptive and comprehensive search for optimal loss functions. We realize the state-of-the-art performance in comparison to various competitive schemes, yielding a 10.61% and 4.38% improvement in Visual Information Fidelity (VIF) for general and no-reference scenarios, respectively, while providing results with high contrast, rich details and colors. The code is available at https://github.com/RollingPlain/HSDS_MEF. Guanyao Wu, Hongming Fu, Jinyuan Liu 0001, Long Ma 0002, Xin Fan 0001, Risheng Liu |
AAAI | 1 |
| 2024 | Segmentation-Driven Infrared and Visible Image Fusion Via Transformer-Enhanced Architecture SearchingabstractA series of infrared and visible image fusion (IVIF) methods have emerged to improve the performance of segmentation task. However, existing perception-focused IVIF methods take visual effects and semantic information as a unified goal for training, ignoring the task conflicts. Moreover, these methods often involve manually designed modules, which are laborious and suboptimal. To solve the problems, we propose a collaborative feature learning framework based on neural architecture search (NAS). Specifically, we extract shared features of fusion and segmentation tasks into a unified space and separately process task objectives through a dual decoder. In light of the essential role that semantic information plays in the segmentation task, we construct a hybrid search space with transformers incorporated to enhance context dependence handling. Our method undergoes extensive experiments, showcasing exceptional visual effects and significant enhancements in segmentation tasks compared to other state-of-the-art methods. Hongming Fu, Guanyao Wu, Zhu Liu 0004, Tiantian Yan, Jinyuan Liu 0001 |
ICASSP | 2 |
| 2024 | Where Elegance Meets Precision: Towards a Compact, Automatic, and Flexible Framework for Multi-modality Image Fusion and Applications
Jinyuan Liu 0001, Guanyao Wu, Zhu Liu 0004, Long Ma 0002, Risheng Liu, Xin Fan 0001 |
IJCAI | 2 |
| 2024 | CoCoNet: Coupled Contrastive Learning Network with Multi-level Feature Ensemble for Multi-modality Image Fusion
Jinyuan Liu 0001, Runjia Lin, Guanyao Wu, Risheng Liu, Zhongxuan Luo, Xin Fan 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | Searching a Compact Architecture for Robust Multi-Exposure Image FusionabstractIn recent years, learning-based methods have achieved significant advancements in multi-exposure image fusion. However, two major stumbling blocks hinder the development, including pixel misalignment and inefficient inference. Reliance on aligned image pairs in existing methods causes susceptibility to artifacts due to device motion. Additionally, existing techniques often rely on handcrafted architectures with huge network engineering, resulting in redundant parameters, adversely impacting inference efficiency and flexibility. To mitigate these limitations, this study introduces an architecture search-based paradigm incorporating self-alignment and detail repletion modules for robust multi-exposure image fusion. Specifically, targeting the extreme discrepancy of exposure, we propose the self-alignment module, leveraging scene relighting to constrain the illumination degree for following alignment and feature extraction. Detail repletion is proposed to enhance the texture details of scenes. Additionally, incorporating a hardware-sensitive constraint, we present the fusion-oriented architecture search to explore compact and efficient networks for fusion. The proposed method outperforms various competitive schemes, achieving a noteworthy 3.19% improvement in PSNR for general scenarios and an impressive 23.5% enhancement in misaligned scenarios. Moreover, it significantly reduces inference time by 69.1%. The code will be available at https://github.com/LiuZhu-CV/CRMEF. Zhu Liu 0004, Jinyuan Liu 0001, Guanyao Wu, Xin Fan 0001, Risheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Learning to Search a Lightweight Generalized Network for Medical Image FusionabstractImage fusion is indispensable in a comprehensive medical imaging pipeline. By embracing deep learning technology, medical image fusion has achieved tremendous progress over the past few years. However, existing approaches make efforts on the specific type of medical image fusion task and may face difficulties in generalizing well. Moreover, most of them strain every nerve to design various architectures with an increase of the width of depth, placing an obstacle in running efficiency. To address the above problems, we propose an Auto-searching Light-weighted Multi-source Fusion network, namely ALMFnet, aiming at incorporating both software and hardware knowledge in a network architecture searching manner for medical image fusion. Specifically, the ALMFnet, consisting of two different feature-extracting modules and one fusion module, is developed to extract and refine multi-source features in a generalized model. Besides, motivated by the collaborative principle, we introduce hardware constraints for sufficient searching the each particular component, further reducing the complexity of the obtained model. Furthermore, to preserve important details in pathological image areas, we introduce a segmentation mask into the developed method. Experimental results demonstrate that our generalized model outperforms previous methods not only in terms of quantitative scores but also in model complexity. Source code will be available at https://github.com/RollingPlain/ALMFnet. Pan Mu, Guanyao Wu, Jinyuan Liu 0001, Yuduo Zhang, Xin Fan 0001, Risheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and SegmentationabstractMulti-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, e.g., fusion or segmentation, making it hard to reach ‘Best of Both Worlds’. To overcome this issue, in this paper, we propose a Multi-interactive Feature learning architecture for image fusion and Segmentation, namely SegMiF, and exploit dual-task correlation to promote the performance of both tasks. The SegMiF is of a cascade structure, containing a fusion sub-network and a commonly used segmentation sub-network. By slickly bridging intermediate features between two components, the knowledge learned from the segmentation task can effectively assist the fusion task. Also, the benefited fusion network supports the segmentation one to perform more pretentiously. Besides, a hierarchical interactive attention block is established to ensure fine-grained mapping of all the vital information between two tasks, so that the modality/semantic features can be fully mutual-interactive. In addition, a dynamic weight factor is introduced to automatically adjust the corresponding weights of each task, which can balance the interactive feature correspondence and break through the limitation of laborious tuning. Furthermore, we construct a smart multi-wave binocular imaging system and collect a full-time multi-modality benchmark with 15 annotated pixel-level categories for image fusion and segmentation. Extensive experiments on several public datasets and our benchmark demonstrate that the proposed method outputs visually appealing fused images and perform averagely 7.66% higher segmentation mIoU in the real-world scene than the state-of-the-art approaches. The source code and benchmark are available at https://github.com/JinyuanLiu-CV/SegMiF. Jinyuan Liu 0001, Zhu Liu 0004, Guanyao Wu, Long Ma 0002, Risheng Liu, Zhongxuan Luo, Xin Fan 0001 |
ICCV | 3 |
| 2023 | Bi-level Dynamic Learning for Jointly Multi-modality Image Fusion and BeyondabstractRecently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and neglecting others, seldom investigating their underlying connections for joint promotion. To overcome these limitations, we establish the hierarchical dual tasks-driven deep model to bridge these tasks. Concretely, we firstly construct an image fusion module to fuse complementary characteristics and cascade dual task-related modules, including a discriminator for visual effects and a semantic network for feature measurement. We provide a bi-level perspective to formulate image fusion and follow-up downstream tasks. To incorporate distinct task-related responses for image fusion, we consider image fusion as a primary goal and dual modules as learnable constraints. Furthermore, we develop an efficient first-order approximation to compute corresponding gradients and present dynamic weighted aggregation to balance the gradients for fusion learning. Extensive experiments demonstrate the superiority of our method, which not only produces visually pleasant fused results but also realizes significant promotion for detection and segmentation than the state-of-the-art approaches. Zhu Liu 0004, Jinyuan Liu 0001, Guanyao Wu, Long Ma 0002, Xin Fan 0001, Risheng Liu |
IJCAI | 3 |
| 2023 | A unified image fusion framework with flexible bilevel paradigm integration
Jinyuan Liu 0001, Zhiying Jiang, Guanyao Wu, Risheng Liu, Xin Fan 0001 |
Vis. Comput. | 3 |
| 2022 | Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionabstractThis study addresses the issue of fusing infrared and visible images that appear differently for object detection. Aiming at generating an image of high visual quality, previous approaches discover commons underlying the two modalities and fuse upon the common space either by iterative optimization or deep networks. These approaches neglect that modality differences implying the complementary information are extremely important for both fusion and subsequent detection task. This paper proposes a bilevel optimization formulation for the joint problem of fusion and detection, and then unrolls to a target-aware Dual Adversarial Learning (TarDAL) network for fusion and a commonly used detection network. The fusion network with one generator and dual discriminators seeks commons while learning from differences, which preserves structural information of targets from the infrared and textural details from the visible. Furthermore, we build a synchronized imaging system with calibrated infrared and optical sensors, and collect currently the most comprehensive benchmark covering a wide range of scenarios. Extensive experiments on several public datasets and our benchmark demonstrate that our method outputs not only visually appealing fusion but also higher detection mAP than the state-of-the-art approaches. The source code and benchmark are available at https://github.com/dlut-dimt/TarDAL. Jinyuan Liu 0001, Xin Fan 0001, Zhanbo Huang, Guanyao Wu, Risheng Liu, Zhongxuan Luo |
CVPR | 4 |
| 2022 | Learn to Search a Lightweight Architecture for Target-Aware Infrared and Visible Image FusionabstractDeep learning technology has recently achieved remarkable progress in infrared and visible image fusion. Nevertheless, existing methods have encountered blurred targets or unfaithful textural details on their fused results. They suffer from high computational expenses, thus unable to directly serve the subsequent high-level vision tasks. In this letter, to alleviate this issue, we proposed leveraging a lightweight architecture based on Neural Architecture Search (NAS) to realize the infrared and visible image fusion in an end-to-end manner, significantly reducing the computational expenses and runtime. Concretely, we construct a search-based architecture to explore the feature representation across different modalities automatically. Then a saliency-based loss function is designed to retain both the distinct target and texture details. Motivated by the cooperative principle, we also formulate a flexible hardware-sensitive regularization constraint in our loss function for discovering efficient operations. As a result, we can generate a target-distinct fused result with high efficiency. Extensive qualitative and quantitative experiments reveal that our method has superior performance against the state-of-the-art methods, especially highlighting the target, retaining realistic details, and achieving fast running speed. Specifically, our method increases by 150% in time, reduces the FLOPS by 21.3% and reduces the model parameters by 25%. Jinyuan Liu 0001, Guanyao Wu, Risheng Liu, Xin Fan 0001 |
IEEE Signal Process. Lett. | 3 |