Jinyuan Liu 0001

dblp:169/7173-1 · DBLP profile ↗
← Back
83ranked-venue papers
15as first author
82since 2021 · last 2026
0000-0003-2085-2676ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 66 · 11 first-author · 65 since 2021Artificial intelligence and machine learning · 41 · 9 first-author · 41 since 2021
YearPublicationVenuePosition
2026 Domain Adaptation Guided Infrared and Visible Image Fusion
abstract
Infrared and Visible Image Fusion (IVIF) integrates complementary information from distinct modalities to enhance image quality. However, the effectiveness declines under unseen conditions such as novel weather or scenes, due to domain shifts primarily from variations of data distribution in the visible modality, while the infrared modality remains relatively stable. To overcome domain shifts caused by the imbalance between modalities during image fusion, we propose a Domain Adaptation Guided Infrared and Visible Image Fusion method, termed DAFusion, leveraging a dual-rank domain adapter to enable fast adaptation to diverse adverse conditions during image fusion. Specifically, trainable low-rank and high-rank embedding spaces are respectively used to capture knowledge common across domains (domain-shared) and those unique to target domains (domain-specific). To leverage the dual-rank adapter more effectively, we develop a homeostatic knowledge allotment strategy to integrate the distinct types of knowledge dynamically based on the uncertainty value of target domains. Since domain adaptation typically optimizes for feature alignment across domains and emphasizes invariance rather than preserving specific cues critical for image fusion, while the fusion objective requires retaining discriminative and complementary features, a conflict between the two modules appears. To reconcile this, we further adopt a bi-level optimization framework that structurally decouples the two objectives, enabling the fusion module to steer the adaptation process while benefiting in return from domain-aligned representations. Experimental results on three benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, achieving both an enhancement in fusion quality and an improvement on subsequent high-level tasks.
Tianwei Guan, Haozhen Wei, Zecheng Xu, Zhiying Jiang, Jinyuan Liu 0001, Xingyuan Li 0005
AAAI7
2026 RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited Labels
abstract
Object detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their inability to provide precise annotation data for sonar images. Therefore, designing effective object detection methods for sonar images with extremely limited labels is particularly important. To address this, we propose a teacher-student framework called RSOD, which aims to fully learn the characteristics of sonar images and develop a pseudo-label strategy suitable for these images to mitigate the impact of limited labels. First, RSOD calculates a reliability score by assessing the consistency of the teacher's predictions across different views. To leverage this score, we introduce an object mixed pseudo-label method to tackle the shortage of labeled data in sonar images. Finally, we optimize the performance of the student by implementing a reliability-guided adaptive constraint. By taking full advantage of unlabeled data, the student can perform well even in situations with extremely limited labels. Notably, on the UATD dataset, our method, using only 5% of labeled data, achieves results that can compete against those of our baseline algorithm trained on 100% labeled data. We also collected a new dataset to provide more valuable data for research in the field of sonar.
Chengzhou Li, Guanchen Meng, Qi Jia 0001, Jinyuan Liu 0001, Zhu Liu 0004, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001
AAAI5
2026 A Hybrid Space Model for Misaligned Multi-modality Image Fusion
abstract
Infrared and visible image fusion aims to integrate complementary information, such as thermal saliency from infrared imagery and fine-grained texture details from visible imagery. However, real-world multi-modal misalignment and geometric deformation often introduce severe artifacts. Most existing methods focus on feature extraction within Euclidean space, thereby neglecting the inherent hierarchical structures embedded in multimodal representations. While Euclidean space excels at preserving local structural details and supporting efficient computation, hyperbolic space is naturally suited for modeling hierarchical relationships due to its geometric properties. Building upon these observations, this paper proposes a unified framework that jointly optimizes image registration and fusion through a dual-space architecture. This architecture synergistically combines the local fidelity of Euclidean geometry with the hierarchical modeling capability of hyperbolic geometry to enhance multimodal representation learning. Specifically, this paper introduces Hyperbolic Coupled Contrastive Learning Optimization (HCCLO), which aligns and optimizes the hierarchical structures of infrared and visible embeddings in hyperbolic space. Moreover, this paper designs a task-adaptive dual-space features fusion mechanism, which dynamically balances and fuses Euclidean local features with hyperbolic hierarchical representations, thereby improving adaptability for downstream tasks. Extensive experiments on misaligned multimodal datasets demonstrate that our method achieves state-of-the-art performance, while effectively capturing both spatial dependencies and hierarchical semantics.
Jia Wang 0036, Zhu Liu 0004, Di Wang 0018, Jinyuan Liu 0001, Risheng Liu
AAAI5
2026 HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-Resolution
abstract
Infrared video has been of great interest in visual tasks under challenging environments, but often suffers from severe atmospheric turbulence and compression degradation. Existing video super-resolution (VSR) methods either neglect the inherent modality gap between infrared and visible images or fail to restore turbulence-induced distortions. Directly cascading turbulence mitigation (TM) algorithms with VSR methods leads to error propagation and accumulation due to the decoupled modeling of degradation between turbulence and resolution. We introduce HATIR, a Heat-Aware Diffusion for Turbulent InfraRed Video Super-Resolution, which injects heat-aware deformation priors into the diffusion sampling path to jointly model the inverse process of turbulent degradation and structural detail loss. Specifically, HATIR constructs a Phasor-Guided Flow Estimator, rooted in the physical principle that thermally active regions exhibit consistent phasor responses over time, enabling reliable turbulence-aware flow to guide the reverse diffusion process. To ensure the fidelity of structural recovery under nonuniform distortions, a Turbulence-Aware Decoder is proposed to selectively suppress unstable temporal cues and enhance edge-aware feature aggregation via turbulence gating and structure-aware attention. We built FLIR-IVSR, the first dataset for turbulent infrared VSR, comprising paired LR-HR sequences from a FLIR T1050sc camera (1024 X 768) spanning 640 diverse scenes with varying camera and object motion conditions. This encourages future research in infrared VSR.
Yang Zou 0004, Xingyue Zhu, Kaiqi Han, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001
AAAI7
2026 Versatile Luminosity Tuning: Relighting Illumination via Dual-Prompt Exposure Correction
Jinyuan Liu 0001, Gehui Li, Zhiying Jiang, Long Ma 0002, Miao Zhang 0004, Xin Fan 0001, Risheng Liu
Int. J. Comput. Vis.1
2026 Contourlet Refinement Gate Framework for Thermal Spectrum Distribution Regularized Infrared Image Super-Resolution
Yang Zou 0004, Zhixin Chen, Xingyuan Li 0005, Long Ma 0002, Jinyuan Liu 0001
Int. J. Comput. Vis.6
2026 Discovering a nimble network for misaligned multi-exposure image fusion
Xiaoshan Liu, Guanyao Wu, Yichuan Peng, Weiqiang Kong, Jinyuan Liu 0001
Neurocomputing5
2026 Cross-domain time-frequency Mamba: A more effective model for long-term time series forecasting
Yuhang Duan, Lin Lin 0008, Jinyuan Liu 0001, Xin Fan 0001
Knowl. Based Syst.3
2026 Degradation-aware feature collaboration for underwater image stitching
Xiaoke Shang, Zengxi Zhang, Jinyuan Liu 0001
Pattern Recognit.4
2026 Structured Prompt-Guided Segmentation of Enhanced CT Images for Predicting Lateral Cervical Lymph Node Metastasis in Papillary Thyroid Carcinoma
abstract
Accurate segmentation of metastatic lymph nodes in head-and-neck CT is crucial for treatment planning in papillary thyroid carcinoma. Although deep networks have advanced medical image segmentation, most existing methods depend solely on visual features and neglect clinical semantic cues. In this work, we propose a text-guided segmentation framework that introduces structured clinical baseline feature into a YOLOv8-based segmentation network. To effectively fuse semantic and spatial features, we design a cross-modal fusion mechanism that aligns textual embeddings with image features at multiple scales. This allows the model to focus on diagnostically relevant regions and improves performance in challenging scenarios involving low contrast and ambiguous boundaries. Extensive experiments on a multi-center, pathology-confirmed CT dataset demonstrate that our method achieves superior results in AP@50, mAP@[.5:.95], and AR@[.5:.95], outperforming existing state-of-the-art instance segmentation models.
Xingyuan Li 0005, Jinyuan Liu 0001, Yanhui Peng
IEEE Signal Process. Lett.4
2026 Illumination Refinement via Textual Cues: A Prompt-Driven Approach for Low-Light NeRF Enhancement
abstract
In the realm of 3D scene modeling and rendering, the emergence of Neural Radiance Fields (NeRF) represents a significant leap forward. However, NeRF’s rendering performance suffers significantly when rendering images under low-light conditions. Existing approaches are optimized by enhancing low-light input images and combining NeRF models, but still fail to address the issues of multiview consistency and image quality. To address these challenges, our research introduces a textual constraint-prompted enhancement method that facilitates low-light image brightening and new view synthesis in an unsupervised manner. Specifically, we devise a semantic calibration strategy that employs positive and negative prompts to motivate and penalize the network towards attributes associated with high-quality images and exploits the capability of visual language models in semantic parsing to align the generated images with textual descriptors to improve image generation quality. In addition, to address the multiview consistency problem, we propose a two-layer optimization strategy, where the semantic cue optimization in the upper layer and the new view generation in the lower layer interact with each other to achieve a balance between luminance consistency and structural integrity by combining these improved images with text-driven semantic features. Comprehensive tests on two datasets with different resolutions, LOM and LLFF, show that our approach outperforms existing methods by significantly improving the brightness and clarity of low-light images to state-of-the-art while preserving the natural appearance and details.
Xinrui Ju, Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.7
2026 SeFENet: Robust Deep Homography Estimation via Semantic-Driven Feature Enhancement
abstract
Images captured in harsh environments often exhibit blurred details, reduced contrast, and color distortion, which hinder feature detection and matching, thereby affecting the accuracy and robustness of homography estimation. While visual enhancement can improve contrast and clarity, it may introduce visual-tolerant artifacts that obscure the structural integrity of images. Considering the resilience of semantic information against environmental interference, we propose a semantic-driven feature enhancement network for robust homography estimation, dubbed SeFENet. Concretely, in our homography estimation network —— Target Aware Homography Estimation Module(TAHEM), we first introduce an innovative hierarchical scale-aware module to expand the receptive field by aggregating multi-scale information, thereby effectively extracting image’s structural features under diverse harsh conditions. Subsequently, we employ a Semantic Extraction Module to extract multi-scale semantic features from the input images. Combined with a high-level perceptual framework, this enables degradation-tolerant semantic feature extraction. Building upon this, the Semantic-Guide Meta Constraints module leverages a meta-learning training strategy to effectively fuse the semantic features with structural features. By internal-external alternating optimization, the proposed network achieves implicit semantic-wise feature enhancement, thereby improving the robustness of homography estimation in adverse environments by strengthening the local feature comprehension and context information extraction. Experimental results under both normal and harsh conditions demonstrate that SeFENet significantly outperforms SOTA methods, reducing point match error by at least 41% on the large-scale datasets.
Zeru Shi, Zengxi Zhang, Kemeng Cui, Ruizhe An, Jinyuan Liu 0001, Zhiying Jiang
IEEE Trans. Circuits Syst. Video Technol.5
2026 Multi-Scale Spatial Channel Joint Representation for General Multi-Modality Image Fusion With Self-Supervision
abstract
The rapid advancement of multi-modality image fusion technology enables researchers to simultaneously acquire information from different modalities within a single fused image. In existing methods, some general approaches can implement both infrared and visible image fusion (IVIF) and medical image fusion (MIF) in the same framework. Nevertheless, these methods often ignore the learning of specific features in different modalities, resulting in unsatisfactory performance in fused results. To overcome this issue, we propose a multi-scale joint framework with self-supervision for general multi-modality image fusion, abbreviated as SCSFusion. It enables more targeted and robust implementation of IVIF and MIF. Specifically, in the fusion network, a joint attention module is employed to parallelly capture self-attention features in spatial and channel domains, which can keep fused results accurate in visual representation. Meanwhile, we utilize source images of different modalities to generate visual-focused maps as pseudo labels for self-supervised training of the fusion results. It effectively preserves the salient details in each fused image from being disrupted by other extracted information. Moreover, a medical dataset with segmentation labels, termed M2DF, is reorganized for fusion and down-stream tasks in MIF. With the help of M2DF, a pre-trained segmentation model can be cascaded with the fusion network, aiming to obtain high-level semantic features from inputs and enhance the data generalization in our general framework. We have conducted extensive experiments and analyses on SCSFusion in M$\rm ^{3}$FD, FMB, and M2DF datasets, respectively. The results indicate that the fused images generated by SCSFusion can not only achieve visually appealing results and superior performance metrics in MIF and IVIF, but also exhibit satisfactory performance in down-stream tasks.
Jiawei Li 0016, Jiansheng Chen 0001, Jinyuan Liu 0001, Xinlong Ding, Huimin Ma 0001
IEEE Trans. Multim.3
2025 A²RNet: Adversarial Attack Resilient Network for Robust Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion (IVIF) is a crucial technique for enhancing visual performance by integrating unique information from different modalities into one fused image. Exiting methods pay more attention to conducting fusion with undisturbed data, while overlooking the impact of deliberate interference on the effectiveness of fusion results. To investigate the robustness of fusion models, in this paper, we propose a novel adversarial attack resilient network, called A2RNet. Specifically, we develop an adversarial paradigm with an anti-attack loss function to implement adversarial attacks and training. It is constructed based on the intrinsic nature of IVIF and provide a robust foundation for future research advancements. We adopt a Unet as the pipeline with a transformer-based defensive refinement module (DRM) under this paradigm, which guarantees fused image quality in a robust coarse-to-fine manner. Compared to previous works, our method mitigates the adverse effects of adversarial perturbations, consistently maintaining high-fidelity fusion results. Furthermore, the performance of downstream tasks can also be well maintained under adversarial attacks.
Jiawei Li 0016, Jiansheng Chen 0002, Xinlong Ding, Jinyuan Liu 0001, Bochao Zou, Huimin Ma 0001
AAAI6
2025 CoA: Towards Real Image Dehazing via Compression-and-Adaptation
abstract
Learning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constraints and the diversity of real-world scenes. Therefore, there is an urgent need for an algorithm that excels in both efficiency and adaptability to address real image dehazing effectively. This work proposes a Compression-and-Adaptation (CoA) computational flow to tackle these challenges from a divide-and-conquer perspective. First, model compression is performed in the synthetic domain to develop a compact dehazing parameter space, satisfying efficiency demands. Then, a bilevel adaptation in the real domain is introduced to be fearless in unknown real environments by aggregating the synthetic dehazing capabilities during the learning process. Leveraging a succinct design free from additional constraints, our CoA exhibits domain-irrelevant stability and model-agnostic flexibility, effectively bridging the model chasm between synthetic and real domains to further improve its practical utility. Extensive evaluations and analyses underscore the approach's superiority and effectiveness. The code is publicly available at https://github.com/fyxnl/COA.
Long Ma 0002, Yan Zhang 0002, Jinyuan Liu 0001, Weimin Wang 0007, Guang-Yong Chen, Chengpei Xu, Zhuo Su 0001
CVPR4
2025 Rethinking Reconstruction and Denoising in the Dark: New Perspective, General Architecture and Beyond
abstract
Recently, enhancing image quality in the original RAW domain has garnered significant attention, with denoising and reconstruction emerging as fundamental tasks. Although some works attempt to couple these tasks, they primarily focus on cascade learning while neglecting task associativity within a broader parameter space, leading to suboptimal performance. This work introduces a novel approach by rethinking denoising and reconstruction from a "backbone-head" perspective, leveraging the stronger shared parameter space offered by the backbone, compared to the encoder used in existing works. We derive task-specific heads with fewer parameters to mitigate learning pressure. By incorporating chromaticity-and-noise perception module into the backbone and introducing task-specific supervision during training, we enable simultaneous high-quality results for reconstruction and denoising. Additionally, we design a dual-head interaction module to capture the latent correspondence between the two tasks, significantly enhancing multi-task accuracy. Extensive experiments validate the superiority of the proposed method. Code is available at: https://github.com/csmty/CANS.
Tengyu Ma 0004, Long Ma 0002, Ziye Li, Yuetong Wang, Jinyuan Liu 0001, Chengpei Xu, Risheng Liu
CVPR5
2025 DEAL: Data-Efficient Adversarial Learning for High-Quality Infrared Imaging
abstract
Thermal imaging is often compromised by dynamic, complex degradations caused by hardware limitations and unpredictable environmental factors. The scarcity of high-quality infrared data, coupled with the challenges of dynamic, intricate degradations, makes it difficult to recover details using existing methods. In this paper, we introduce thermal degradation simulation integrated into the training process via a mini-max optimization, by modeling these degraded factors as adversarial attacks on thermal images. The simulation is dynamic to maximize objective functions, thus capturing a broad spectrum of degraded data distributions. This approach enables training with limited data, thereby improving model performance. Additionally, we introduce a dual-interaction network that combines the benefits of spiking neural networks with scale transformation to capture degraded features with sharp spike signal intensities. This architecture ensures compact model parameters while preserving efficient feature representation. Extensive experiments demonstrate that our method not only achieves superior visual quality under diverse single and composited degradation, but also delivers a significant reduction in processing when trained on only fifty clear images, outperforming existing techniques in efficiency and accuracy. The source code will be available at https://github.com/LiuZhu-CV/DEAL.
Zhu Liu 0004, Jinyuan Liu 0001, Fanqi Meng, Long Ma 0002, Risheng Liu
CVPR3
2025 DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution
abstract
Infrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low spatial resolution and complex degradations, consistently challenge imaging quality and subsequent visual tasks. Hence, infrared image super-resolution (IISR) has been developed to address this challenge. While recent developments in diffusion models have greatly advanced this field, current methods to solve it either ignore the unique modal characteristics of infrared imaging or overlook the machine perception requirements. To bridge these gaps, we propose DifIISR, an infrared image super-resolution diffusion model optimized for visual quality and perceptual performance. Our approach achieves task-based guidance for diffusion by injecting gradients derived from visual and perceptual priors into the noise during the reverse process. Specifically, we introduce an infrared thermal spectrum distribution regulation to preserve visual fidelity, ensuring that the reconstructed infrared images closely align with high-resolution images by matching their frequency components. Subsequently, we incorporate various visual foundational models as the perceptual guidance for downstream visual tasks, infusing generalizable perceptual features beneficial for detection and segmentation. As a result, our approach gains superior visual results while attaining State-Of-The-Art downstream task performance. Code is available at https://github.com/zirui0625/DifIISR
Xingyuan Li 0005, Yang Zou 0004, Zhixin Chen, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001
CVPR8
2025 DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each modality. Existing approaches typically treat image fusion and subsequent high-level tasks as separate processes, resulting in fused images that offer only marginal gains in task performance and fail to provide constructive feedback for optimizing the fusion process. To overcome these limitations, we propose a Discriminative Cross-Dimension Evolutionary Learning Framework, termed DCEvo, which simultaneously enhances visual quality and perception accuracy. Leveraging the robust search capabilities of Evolutionary Learning, our approach formulates the optimization of dual tasks as a multi-objective problem by employing an Evolutionary Algorithm (EA) to dynamically balance loss function parameters. Inspired by visual neuroscience, we integrate a Discriminative Enhancer (DE) within both the encoder and decoder, enabling the effective learning of complementary features from different modalities. Additionally, our Cross-Dimensional Embedding (CDE) block facilitates mutual enhancement between high-dimensional task features and low-dimensional fusion features, ensuring a cohesive and efficient feature integration process. Experimental results on three benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, achieving an average improvement of 9.32% in visual quality while also enhancing subsequent high-level tasks. The code is available at https://github.com/Beate-Suy-Zhang/DCEvo.
Jinyuan Liu 0001, Qingyun Mei, Xingyuan Li 0005, Yang Zou 0004, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001
CVPR1
2025 Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond
abstract
Multi-modality image fusion, particularly infrared and visible, plays a crucial role in integrating diverse modalities to enhance scene understanding. Although early research prioritized visual quality, preserving fine details and adapting to downstream tasks remains challenging. Recent approaches attempt task-specific design but rarely achieve "The Best of Both Worlds" due to inconsistent optimization goals. To address these issues, we propose a novel method that leverages the semantic knowledge from the Segment Anything Model (SAM) to Grow the quality of fusion results and Enable downstream task adaptability, namely SAGE. Specifically, we design a Semantic Persistent Attention (SPA) Module that efficiently maintains source information via the persistent repository while extracting high-level semantic priors from SAM. More importantly, to eliminate the impractical dependence on SAM during inference, we introduce a bi-level optimization-driven distillation mechanism with triplet losses, which allow the student network to effectively extract knowledge. Extensive experiments show that our method achieves a balance between high-quality visual results and downstream task adaptability while maintaining practical deployment efficiency. The code is available at https://github.com/RollingPlain/SAGE_IVIF.
Guanyao Wu, Hongming Fu, Yichuan Peng, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
CVPR5
2025 TextMEF: Text-guided Prompt Learning for Multi-exposure Image Fusion
abstract
Multi-exposure image fusion~(MEF) aims to integrate a set of low dynamic range images, producing a single image with a higher dynamic range than either one. Despite significant advancements, current MEF approaches still struggle to handle extremely over- or under-exposed conditions, resulting in unsatisfactory visual effects such as hallucinated details and distorted color tones. With this regard, we propose TextMEF, a prompt-driven fusion method enhanced by prompt learning, for multi-exposure image fusion. Specifically, we learn a set of prompts based on text-image similarity among negative and positive samples (over-exposed, under-exposed images, and well-exposed ones). These learned prompts are seamlessly integrated into the loss function, providing high-level guidance for constraining non-uniform exposure regions. Furthermore, we develop a attention Mamba module effectively translates over-/under- exposed regional features into exposure invariant space and ensure them to build efficient long-range dependency to high dynamic range image. Extensive experimental results on three publicly available benchmarks demonstrate that our TextMEF significantly outperforms state-of-the-art approaches in both visual inspection and objective analysis.
Jinyuan Liu 0001, Qianjun Huang, Guanyao Wu, Di Wang 0018, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001
IJCAI1
2025 Physics-Guided Sonar Image Fine-grained Recognition under Scarce Annotations
abstract
Sonar image recognition is a key technology in underwater exploration systems. Compared with natural images, sonar images have fewer texture details and are easily affected by heavy noise, making it more challenging for specialists to distinguish the subtle differences among classes. In view of this, studying fine-grained classification methods for sonar images with scarce annotations is of significant importance. To address this issue, we propose a Physics-Guided Teacher-Student (PGTS) framework to explore the unique physical information of sonar images while simultaneously mitigating the effects of limited annotations. First, PGTS reconstructs sonar signals through physical simulation and a specially designed physics-guided feature generation module, which allows it to bypass the time-consuming physical simulation during inference. Then, we design a multi-modal teacher model combines the reconstructed sonar signals and sonar images to extract discriminative features to generate robust pseudo labels for fine-grained target categories. Finally, the knowledge is transferred to a single-modal student model through consistency loss. Under the joint constraints of the teacher model and the reconstructed sonar physical signals, the student model continuously improves its performance in annotation-scarce scenarios. Notably, when merely 1% of the data is labeled, our method outperforms other state-of-the-art approaches by 12.46% in terms of accuracy.
Chengzhou Li, Qi Jia 0001, Jinyuan Liu 0001, Zhiying Jiang, Longhan Feng, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001
ACM Multimedia4
2025 Toward a Training-Free Plug-and-Play Refinement Framework for Infrared and Visible Image Registration and Fusion
abstract
Infrared and Visible Image Fusion (IVIF) under unregistered conditions has been of great interest in various visual tasks under challenging environments. While existing approaches often demonstrate promising results on specific benchmarks, they tend to exhibit performance drops in unseen scenarios and incur high computational overhead when retrained on new datasets. To address these challenges, we propose TRACE, a Training-free Reinforcement-based Alignment method for Cross-modality Enhancement, which incorporates Evaluator, a rewarding network, into an evaluation-driven Reinforcement Learning (RL) framework, enabling efficient and plug-and-play refinement of any existing registration approach. Specifically, TRACE constructs the Evaluator network to assess the alignment quality of the given registration model, generating confidence scores and adjustment masks via spatial and channel attention. Leveraging these cues as RL rewards, TRACE iteratively refines the registration network to mitigate misalignments until the accumulated improvement is satisfied. Due to its training-free and plug-and-play nature, TRACE notably enhances fusion results across diverse and unseen scenarios. TRACE achieves impressive improvements in different methods across diverse datasets with minimal computational cost. The project page is available at https://github.com/pubyLu/TRACE.
Yang Zou 0004, Xingyuan Li 0005, Xingyue Zhu, Kaiqi Han, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001
ACM Multimedia8
2025 Depth-Supervised Fusion Network for Seamless-Free Image Stitching
abstract
Image stitching synthesizes images captured from multiple perspectives into a single image with a broader field of view. The significant variations in object depth often lead to large parallax, resulting in ghosting and misalignment in the stitched results. To address this, we propose a depth-consistency-constrained seamless-free image stitching method. First, to tackle the multi-view alignment difficulties caused by parallax, a multi-stage mechanism combined with global depth regularization constraints is developed to enhance the alignment accuracy of the same apparent target across different depth ranges. Second, during the multi-view image fusion process, an optimal stitching seam is determined through graph-based low-cost computation, and a soft-seam region is diffused to precisely locate transition areas, thereby effectively mitigating alignment errors induced by parallax and achieving natural and seamless stitching results. Furthermore, considering the computational overhead in the shift regression process, a reparameterization strategy is incorporated to optimize the structural design, significantly improving algorithm efficiency while maintaining optimal performance. Extensive experiments demonstrate the superior performance of the proposed method against the existing methods. Code is available at https://github.com/DLUT-YRH/DSFN.
Zhiying Jiang, Ruhao Yan, Zengxi Zhang, Jinyuan Liu 0001
NeurIPS5
2025 Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark
abstract
We engage in the relatively underexplored task named thermal infrared image enhancement. Existing infrared image enhancement methods primarily focus on tackling individual degradations, such as noise, contrast, and blurring, making it difficult to handle coupled degradations. Meanwhile, all-in-one enhancement methods, commonly applied to RGB sensors, often demonstrate limited effectiveness due to the significant differences in imaging models. In sight of this, we first revisit the imaging mechanism and introduce a Recurrent Prompt Fusion Network (RPFN). Specifically, the RPFN initially establishes prompt pairs based on the thermal imaging process. For each type of degradation, we fuse the corresponding prompt pairs to modulate the model's features, providing adaptive guidance that enables the model to better address specific degradations under single or multiple conditions.In addition, a selective recurrent training mechanism is introduced to gradually refine the model's handling of composite cases to align the enhancement process, which not only allows the model to remove camera noise and retain key structural details, but also enhancing the overall contrast of the thermal image. Furthermore, we introduce the most comprehensive high-quality infrared benchmark covering a wide range of scenarios. Extensive experiments substantiate that our approach not only delivers promising visual results under specific degradation but also significantly improves performance on complex degradation scenes, achieving a notable 8.76% improvement.
Jinyuan Liu 0001, Zhu Liu 0004, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Risheng Liu
NeurIPS1
2025 Efficient Rectified Flow for Image Fusion
abstract
Image fusion is a fundamental and important task in computer vision, aiming to combine complementary information from different modalities to fuse images. In recent years, diffusion models have made significant developments in the field of image fusion. However, diffusion models often require complex computations and redundant inference time, which reduces the applicability of these methods. To address this issue, we propose RFfusion, an efficient one-step diffusion model for image fusion based on Rectified Flow. We incorporate Rectified Flow into the image fusion task to straighten the sampling path in the diffusion model, achieving one-step sampling without the need for additional training, while still maintaining high-quality fusion results. Furthermore, we propose a task-specific variational autoencoder (VAE) architecture tailored for image fusion, where the fusion operation is embedded within the latent space to further reduce computational complexity. To address the inherent discrepancy between conventional reconstruction-oriented VAE objectives and the requirements of image fusion, we introduce a two-stage training strategy. This approach facilitates the effective learning and integration of complementary information from multi-modal source images, thereby enabling the model to retain fine-grained structural details while significantly enhancing inference efficiency. Extensive experiments demonstrate that our method outperforms other state-of-the-art methods in terms of both inference speed and fusion quality.
Tianwei Guan, Xingyuan Li 0005, Minjing Dong, Jinyuan Liu 0001
NeurIPS7
2025 Image Stitching in Adverse Condition: A Bidirectional-Consistency Learning Framework and Benchmark
abstract
Deep learning-based image stitching methods have achieved promising performance on conventional stitching datasets. However, real-world scenarios may introduce challenges such as complex weather conditions, illumination variations, and dynamic scene motion, which severely degrade image quality and lead to significant misalignment in stitching results. To solve this problem, we propose an adverse condition-tolerant image stitching network, dubbed ACDIS. We first introduce a bidirectional consistency learning framework, which ensures reliable alignment through an iterative optimization paradigm that integrates differentiable image restoration and Gaussian-distribute encoded homography estimation. Subsequently, we incorporate motion constraints into the seamless composition network to produce robust stitching results without interference from moving scenes. We further propose the first adverse scene image stitching dataset, which covers diverse parallax and scenes under low-light, haze, and underwater environments. Extensive experiments show that the proposed method can generate visually pleasing stitched images under adverse conditions, outperforming state-of-the-art methods.
Zengxi Zhang, Junchen Ge, Zhiying Jiang, Miao Zhang 0004, Jinyuan Liu 0001
NeurIPS5
2025 Bilevel Fast Scene Adaptation for Low-Light Image Enhancement
Long Ma 0002, Dian Jin 0003, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
Int. J. Comput. Vis.4
2025 HUPE: Heuristic Underwater Perceptual Enhancement with Semantic Collaborative Learning
Zengxi Zhang, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
Int. J. Comput. Vis.4
2025 Infrared and Visible Image Fusion: From Data Compatibility to Task Adaption
abstract
Infrared-visible image fusion (IVIF) is a fundamental and critical task in the field of computer vision. Its aim is to integrate the unique characteristics of both infrared and visible spectra into a holistic representation. Since 2018, growing amount and diversity IVIF approaches step into a deep-learning era, encompassing introduced a broad spectrum of networks or loss functions for improving visual enhancement. As research deepens and practical demands grow, several intricate issues like data compatibility, perception accuracy, and efficiency cannot be ignored. Regrettably, there is a lack of recent surveys that comprehensively introduce and organize this expanding domain of knowledge. Given the current rapid development, this paper aims to fill the existing gap by providing a comprehensive survey that covers a wide array of aspects. Initially, we introduce a multi-dimensional framework to elucidate the prevalent learning-based IVIF methodologies, spanning topics from basic visual enhancement strategies to data compatibility, task adaptability, and further extensions. Subsequently, we delve into a profound analysis of these new approaches, offering a detailed lookup table to clarify their core ideas. Last but not the least, We also summarize performance comparisons quantitatively and qualitatively, covering registration, fusion and follow-up high-level tasks. Beyond delving into the technical nuances of these learning-based fusion approaches, we also explore potential future directions and open issues that warrant further exploration by the community.
Jinyuan Liu 0001, Guanyao Wu, Zhu Liu 0004, Di Wang 0018, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Risheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Learning With Self-Calibrator for Fast and Robust Low-Light Image Enhancement
abstract
Convolutional Neural Networks (CNNs) have shown significant success in the low-light image enhancement task. However, most of existing works encounter challenges in balancing quality and efficiency simultaneously. This limitation hinders practical applicability in real-world scenarios and downstream vision tasks. To overcome these obstacles, we propose a Self-Calibrated Illumination (SCI) learning scheme, introducing a new perspective to boost the model's capability. Based on a weight-sharing illumination estimation process, we construct an embedded self-calibrator to accelerate stage-level convergence, yielding gains that utilize only a single basic block for inference, which drastically diminishes computation cost. Additionally, by introducing the additivity condition on the basic block, we acquire a reinforced version dubbed SCI++, which disentangles the relationship between the self-calibrator and illumination estimator, providing a more interpretable and effective learning paradigm with faster convergence and better stability. We assess the proposed enhancers on standard benchmarks and in-the-wild datasets, confirming that they can restore clean images from diverse scenes with higher quality and efficiency. The verification on different levels of low-light vision tasks shows our applicability against other methods.
Long Ma 0002, Tengyu Ma 0004, Chengpei Xu, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 LarTap: A Luminance-Aware Framework With Text-Correlation Priors for Multi-Exposure Image Fusion
abstract
Conventional imaging devices often struggle to produce high-dynamic-range (HDR) images that accurately represent natural scenes. To overcome this limitation, multi-exposure image fusion (MEF) techniques have been introduced as a viable solution. Existing MEF approaches aim to enhance performance by optimizing or searching architectures. However, they face challenges in precise feature extraction and scene reconstruction, leading to distortion in the fused images. Additionally, most methods do not adequately address luminance variations across different image regions, which may result in the loss of essential details. To address these challenges, we present a novel luminance-aware MEF framework that integrates text-correlation priors (LarTap). By embedding textual information into fusion process, the proposed framework enhances content extraction and comprehension. Specifically, it consist of two key components: the text-image correlation network (N1) and the multi-exposure fusion network (N2). First, N1 performs correlation training to achieve a holistic alignment between text and image pairs. Its iterative vision encoders (VEs) generate text-correlated prior knowledge to facilitate the fusion process in N2. Second, N2 leverages these priors for scene reconstruction and dynamically adjusts luminance based on comparative perception. Extensive experiments on multiple datasets demonstrate that LarTap outperforms state-of-the-art methods. The source code is available at https://github.com/EnLong-wang/LarTap.
Enlong Wang, Jiawei Li 0016, Tiantian Yan, Jia Lei 0001, Shihua Zhou, Bin Wang 0005, Jinyuan Liu 0001, Nikola K. Kasabov
IEEE Trans. Circuits Syst. Video Technol.7
2025 Harmonized Domain Enabled Alternate Search for Infrared and Visible Image Alignment
abstract
Infrared and visible image alignment is essential and critical to the fusion and multi-modal perception applications. It addresses discrepancies in position and scale caused by spectral properties and environmental variations, ensuring precise pixel correspondence and spatial consistency. Existing manual calibration requires regular maintenance and exhibits poor portability, challenging the adaptability of multi-modal application in dynamic environments. In this paper, we propose a harmonized representation based infrared and visible image alignment, achieving both high accuracy and scene adaptability. Specifically, with regard to the disparity between multi-modal images, we develop an invertible translation process to establish a harmonized representation domain that effectively encapsulates the feature intensity and distribution of both infrared and visible modalities. Building on this, we design a hierarchical framework to correct deformations inferred from the harmonized domain in a coarse-to-fine manner. Our framework leverages advanced perception capabilities alongside residual estimation to enable accurate regression of sparse offsets, while an alternate correlation search mechanism ensures precise correspondence matching. Furthermore, we propose the first ground truth available misaligned infrared and visible image benchmark for evaluation. Extensive experiments validate the effectiveness of the proposed method against the state-of-the-arts, advancing the subsequent applications further. Code and dataset are available at https://github.com/Jzy2017/HR4IR.
Zhiying Jiang, Zengxi Zhang, Jinyuan Liu 0001
IEEE Trans. Image Process.3
2025 MLFuse: Multi-Scenario Feature Joint Learning for Multi-Modality Image Fusion
abstract
Multi-modality image fusion (MMIF) entails synthesizing images with detailed textures and prominent objects. Existing methods tend to use general feature extraction to handle different fusion tasks. However, these methods have difficulty breaking fusion barriers across various modalities owing to the lack of targeted learning routes. In this work, we propose a multi-scenario feature joint learning architecture, MLFuse, that employs the commonalities of multi-modality images to deconstruct the fusion progress. Specifically, we construct a cross-modal knowledge reinforcing network that adopts a multipath calibration strategy to promote information communication between different images. In addition, two professional networks are developed to maintain the salient and textural information of fusion results. The spatial-spectral domain optimizing network can learn the vital relationship of the source image context with the help of spatial attention and spectral attention. The edge-guided learning network utilizes the convolution operations of various receptive fields to capture image texture information. The desired fusion results are obtained by aggregating the outputs from the three networks. Extensive experiments demonstrate the superiority of MLFuse for infrared-visible image fusion and medical image fusion. The excellent results of downstream tasks (i.e., object detection and semantic segmentation) further verify the high-quality fusion performance of our method.
Jia Lei 0001, Jiawei Li 0016, Jinyuan Liu 0001, Bin Wang 0005, Shihua Zhou, Qiang Zhang 0008, Xiaopeng Wei, Nikola K. Kasabov
IEEE Trans. Multim.3
2025 Robust One-Stop Multi-Modality Image Registration-Fusion-Segmentation Framework Against Misalignments and Adversarial Attacks
abstract
In complex open scenes, multi-modality image fusion and segmentation encounter two challenges: i) Imaging misalignments, manifested as pixel shifts and structural distortions, are perceptible. ii) Human-crafted adversarial attacks, reflected in pixel distribution variations, are imperceptible. They not only degrade the visual quality of fused images, e.g., noticeable edge ghosts but more critically undermine semantic perception. However, none of the existing works considered the coupled effect of these degradations. This paper proposes a One-Stop framework incorporating sequential task flows of “Registration-Fusion-Segmentation”, termed OS-RFS. Registration aims to mitigate the chained impact of misalignment on fusion and segmentation. We follow a coarse-to-fine registration paradigm and develop a Global-Local Incremental Registration (GLoIR) model, where the global shift registration (GSR) is performed initially for long-range pixel shifts, followed by incremental local deformation registration (LDR) for subtle local deformations. To improve segmentation robustness, we innovatively introduce auxiliary positive attacks and build a Cancellation Defense Strategy (CDS) in the fusion model. The CDS constrains the fusion model to fit fused images to the distribution of positive attacks, endowing fused images with a robust defense ability against adversarial attacks. This significantly mitigates the impact of adversarial attacks on semantic segmentation. Extensive experimental results reveal that our OS-RFS performs remarkable robustness on multi-modality image fusion and semantic segmentation against imaging misalignments and adversarial attacks.
Di Wang 0018, Xianghao Jiao, Jinyuan Liu 0001, Xin Fan 0001
IEEE Trans. Multim.3
2024 Trash to Treasure: Low-Light Object Detection via Decomposition-and-Aggregation
abstract
Object detection in low-light scenarios has attracted much attention in the past few years. A mainstream and representative scheme introduces enhancers as the pre-processing for regular detectors. However, because of the disparity in task objectives between the enhancer and detector, this paradigm cannot shine at its best ability. In this work, we try to arouse the potential of enhancer + detector. Different from existing works, we extend the illumination-based enhancers (our newly designed or existing) as a scene decomposition module, whose removed illumination is exploited as the auxiliary in the detector for extracting detection-friendly features. A semantic aggregation module is further established for integrating multi-scale scene-related semantic information in the context space. Actually, our built scheme successfully transforms the "trash" (i.e., the ignored illumination in the detector) into the "treasure" for the detector. Plenty of experiments are conducted to reveal our superiority against other state-of-the-art methods. The code will be public if it is accepted.
Xiaohan Cui, Long Ma 0002, Tengyu Ma 0004, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
AAAI4
2024 Towards Robust Image Stitching: An Adaptive Resistance Learning against Compatible Attacks
abstract
Image stitching seamlessly integrates images captured from varying perspectives into a single wide field-of-view image. Such integration not only broadens the captured scene but also augments holistic perception in computer vision applications. Given a pair of captured images, subtle perturbations and distortions which go unnoticed by the human visual system tend to attack the correspondence matching, impairing the performance of image stitching algorithms. In light of this challenge, this paper presents the first attempt to improve the robustness of image stitching against adversarial attacks. Specifically, we introduce a stitching-oriented attack (SoA), tailored to amplify the alignment loss within overlapping regions, thereby targeting the feature matching procedure. To establish an attack resistant model, we delve into the robustness of stitching architecture and develop an adaptive adversarial training (AAT) to balance attack resistance with stitching precision. In this way, we relieve the gap between the routine adversarial training and benign models, ensuring resilience without quality compromise. Comprehensive evaluation across real-world and synthetic datasets validate the deterioration of SoA on stitching performance. Furthermore, AAT emerges as a more robust solution against adversarial perturbations, delivering superior stitching results. Code is available at: https://github.com/Jzy2017/TRIS.
Zhiying Jiang, Xingyuan Li 0005, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
AAAI3
2024 Improving Cross-Modal Alignment with Synthetic Pairs for Text-Only Image Captioning
abstract
Although image captioning models have made significant advancements in recent years, the majority of them heavily depend on high-quality datasets containing paired images and texts which are costly to acquire. Previous works leverage the CLIP's cross-modal association ability for image captioning, relying solely on textual information under unsupervised settings. However, not only does a modality gap exist between CLIP text and image features, but a discrepancy also arises between training and inference due to the unavailability of real-world images, which hinders the cross-modal alignment in text-only captioning. This paper proposes a novel method to address these issues by incorporating synthetic image-text pairs. A pre-trained text-to-image model is deployed to obtain images that correspond to textual data, and the pseudo features of generated images are optimized toward the real ones in the CLIP embedding space. Furthermore, textual information is gathered to represent image features, resulting in the image features with various semantics and the bridged modality gap. To unify training and inference, synthetic image features would serve as the training prefix for the language decoder, while real images are used for inference. Additionally, salient objects in images are detected as assistance to enhance the learning of modality alignment. Experimental results demonstrate that our method obtains the state-of-the-art performance on benchmark datasets.
Zhiyue Liu, Jinyuan Liu 0001, Fanrong Ma
AAAI2
2024 Hybrid-Supervised Dual-Search: Leveraging Automatic Learning for Loss-Free Multi-Exposure Image Fusion
abstract
Multi-exposure image fusion (MEF) has emerged as a prominent solution to address the limitations of digital imaging in representing varied exposure levels. Despite its advancements, the field grapples with challenges, notably the reliance on manual designs for network structures and loss functions, and the constraints of utilizing simulated reference images as ground truths. Consequently, current methodologies often suffer from color distortions and exposure artifacts, further complicating the quest for authentic image representation. In addressing these challenges, this paper presents a Hybrid-Supervised Dual-Search approach for MEF, dubbed HSDS-MEF, which introduces a bi-level optimization search scheme for automatic design of both network structures and loss functions. More specifically, we harness a unique dual research mechanism rooted in a novel weighted structure refinement architecture search. Besides, a hybrid supervised contrast constraint seamlessly guides and integrates with searching process, facilitating a more adaptive and comprehensive search for optimal loss functions. We realize the state-of-the-art performance in comparison to various competitive schemes, yielding a 10.61% and 4.38% improvement in Visual Information Fidelity (VIF) for general and no-reference scenarios, respectively, while providing results with high contrast, rich details and colors. The code is available at https://github.com/RollingPlain/HSDS_MEF.
Guanyao Wu, Hongming Fu, Jinyuan Liu 0001, Long Ma 0002, Xin Fan 0001, Risheng Liu
AAAI3
2024 Enhancing Neural Radiance Fields with Adaptive Multi-Exposure Fusion: A Bilevel Optimization Approach for Novel View Synthesis
abstract
Neural Radiance Fields (NeRF) have made significant strides in the modeling and rendering of 3D scenes. However, due to the complexity of luminance information, existing NeRF methods often struggle to produce satisfactory renderings when dealing with high and low exposure images. To address this issue, we propose an innovative approach capable of effectively modeling and rendering images under multiple exposure conditions. Our method adaptively learns the characteristics of images under different exposure conditions through an unsupervised evaluator-simulator structure for HDR (High Dynamic Range) fusion. This approach enhances NeRF's comprehension and handling of light variations, leading to the generation of images with appropriate brightness. Simultaneously, we present a bilevel optimization method tailored for novel view synthesis, aiming to harmonize the luminance information of input images while preserving their structural and content consistency. This approach facilitates the concurrent optimization of multi-exposure correction and novel view synthesis, in an unsupervised manner. Through comprehensive experiments conducted on the LOM and LOL datasets, our approach surpasses existing methods, markedly enhancing the task of novel view synthesis for multi-exposure environments and attaining state-of-the-art results. The source code can be found at https://github.com/Archer-204/AME-NeRF.
Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001
AAAI4
2024 Contourlet Residual for Prompt Learning Enhanced Infrared Image Super-Resolution
Xingyuan Li 0005, Jinyuan Liu 0001, Zhixin Chen, Yang Zou 0004, Long Ma 0002, Xin Fan 0001, Risheng Liu
ECCV (3)2
2024 Local Contrast Prior-Guided Cross Aggregation Model for Effective Infrared Small Target Detection
abstract
Infrared small target detection, referring to discovering the precise shapes of dim targets from complex clutter background, has gradually become a hot spot. In recent years, learning-based methods have become the mainstream schemes with high efficiency. However, these methods seldom consider the nature characteristics of infrared targets and are with high demands on computational complexity. By investigating the local contrast property, we propose the deformable attention module to separate salient and representative features and reduce the complexity of calculation. Furthermore, in order to aggregate the global and local correlations, we present the cross aggregation combining the transformer and convolution modules with the supervision of targets edges. Comprehensive experiments on the IRSTD-1k dataset demonstrate the superiority of our method, improve 3.32% of IOU and accelerate 50.59% of inference time compared with the most advanced method.
Zhu Liu 0004, Jinyuan Liu 0001
ICASSP3
2024 Segmentation-Driven Infrared and Visible Image Fusion Via Transformer-Enhanced Architecture Searching
abstract
A series of infrared and visible image fusion (IVIF) methods have emerged to improve the performance of segmentation task. However, existing perception-focused IVIF methods take visual effects and semantic information as a unified goal for training, ignoring the task conflicts. Moreover, these methods often involve manually designed modules, which are laborious and suboptimal. To solve the problems, we propose a collaborative feature learning framework based on neural architecture search (NAS). Specifically, we extract shared features of fusion and segmentation tasks into a unified space and separately process task objectives through a dual decoder. In light of the essential role that semantic information plays in the segmentation task, we construct a hybrid search space with transformers incorporated to enhance context dependence handling. Our method undergoes extensive experiments, showcasing exceptional visual effects and significant enhancements in segmentation tasks compared to other state-of-the-art methods.
Hongming Fu, Guanyao Wu, Zhu Liu 0004, Tiantian Yan, Jinyuan Liu 0001
ICASSP5
2024 AEAM3D: Adverse Environment-Adaptive Monocular 3D Object Detection via Feature Extraction Regularization
abstract
3D object detection plays a crucial role in intelligent vision systems. Detection in the open world inevitably encounters various adverse scenes while most of existing methods fail in these scenes. To address this issue, this paper proposes a monocular 3D detection model, termed AEAM3D, which effectively mitigates the degradation of detection performance in various harsh environments. Additionally, we assemble a new adverse 3D object detection dataset encompassing some challenging scenes, including rainy, foggy, and low light weather conditions. Experimental results demonstrate that our proposed method outperforms current state-of-the-art approaches by an average of 3.12% in terms of APR40for car category across adverse environments.
Yixin Lei, Xingyuan Li 0005, Zhiying Jiang, Xinrui Ju, Jinyuan Liu 0001
ICASSP5
2024 Adaptive Multi-Exposure Fusion for Enhanced Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) have revolutionized 3D scene modeling and rendering. However, their performance dips when handling images with diverse exposure levels, mainly due to the intricate luminance dynamics. Addressing this, we present an innovative method that proficiently models and renders images across a spectrum of exposure conditions. Our approach utilizes an unsupervised classifier-generator structure for HDR fusion, significantly enhancing NeRF’s ability to comprehend and adjust to light variations, leading to the generation of images with appropriate brightness. Extensive evaluations on the LOM[1] and LOL[2] datasets underscore our method’s edge. Our approach significantly improves the task of novel view synthesis for multi-exposure images, attaining state-of-the-art results.
Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Tiantian Yan, Jinyuan Liu 0001
ICASSP5
2024 Joint edge detection learning for recurrent homography estimation
abstract
Homography estimation plays a pivotal role in aligning image pairs across multiple viewpoints. Existing methods focus mainly on texture alignment, whereas overlooking the influence of geometric structures, thereby resulting in inaccurate homography estimation. In this paper, we propose a novel recurrent homography estimation framework with joint edge detection learning. We find that edge detection explores extra anchors for homography estimation, and meanwhile homography provides complementary information of cross views for edge detection refinement. Unlike traditional edge detection applied to individual images, our approach establishes structural consistency constraints to reinforce mutual edges while suppressing unreliable structures. Specifically, the detected edges guide and enhance the texture features through a specifically designed edge-aware fusion module. Ultimately, we recurrently compute the correlation of fusion features from small to large scales for homography regression. Our experimental results demonstrate that the proposed method reduces the matching error by 41.7% than state-of-the-art methods. Furthermore, our network excels in detecting edges with extensive details even under dramatic perspective changes. Code is available at https://github.com/edmandzhao/edge-detection-for-RHE.
Qi Jia 0001, Zikun Zhao, Xiaomei Feng, Jinyuan Liu 0001, Yu Liu 0012, Xinwei Xue
ICME4
2024 Where Elegance Meets Precision: Towards a Compact, Automatic, and Flexible Framework for Multi-modality Image Fusion and Applications
Jinyuan Liu 0001, Guanyao Wu, Zhu Liu 0004, Long Ma 0002, Risheng Liu, Xin Fan 0001
IJCAI1
2024 SDFuse: Semantic-injected dual-flow learning for infrared and visible image fusion
Enlong Wang, Jiawei Li 0016, Jia Lei 0001, Jinyuan Liu 0001, Shihua Zhou, Bin Wang 0005, Nikola K. Kasabov
Expert Syst. Appl.4
2024 CoCoNet: Coupled Contrastive Learning Network with Multi-level Feature Ensemble for Multi-modality Image Fusion
Jinyuan Liu 0001, Runjia Lin, Guanyao Wu, Risheng Liu, Zhongxuan Luo, Xin Fan 0001
Int. J. Comput. Vis.1
2024 A Task-Guided, Implicitly-Searched and Meta-Initialized Deep Model for Image Fusion
abstract
Image fusion plays a key role in a variety of multi-sensor-based vision systems, especially for enhancing visual quality and/or extracting aggregated features for perception. However, most existing methods just consider image fusion as an individual task, thus ignoring its underlying relationship with these downstream vision problems. Furthermore, designing proper fusion architectures often requires huge engineering labor. It also lacks mechanisms to improve the flexibility and generalization ability of current fusion approaches. To mitigate these issues, we establish a Task-guided, Implicit-searched and Meta-initialized (TIM) deep model to address the image fusion problem in a challenging real-world scenario. Specifically, we first propose a constrained strategy to incorporate information from downstream tasks to guide the unsupervised learning process of image fusion. Within this framework, we then design an implicit search scheme to automatically discover compact architectures for our fusion model with high efficiency. In addition, a pretext meta initialization technique is introduced to leverage divergence fusion data to support fast adaptation for different kinds of image fusion tasks. Qualitative and quantitative experimental results on different categories of image fusion problems and related downstream tasks (e.g., visual enhancement and semantic understanding) substantiate the flexibility and effectiveness of our TIM.
Risheng Liu, Zhu Liu 0004, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Searching a Compact Architecture for Robust Multi-Exposure Image Fusion
abstract
In recent years, learning-based methods have achieved significant advancements in multi-exposure image fusion. However, two major stumbling blocks hinder the development, including pixel misalignment and inefficient inference. Reliance on aligned image pairs in existing methods causes susceptibility to artifacts due to device motion. Additionally, existing techniques often rely on handcrafted architectures with huge network engineering, resulting in redundant parameters, adversely impacting inference efficiency and flexibility. To mitigate these limitations, this study introduces an architecture search-based paradigm incorporating self-alignment and detail repletion modules for robust multi-exposure image fusion. Specifically, targeting the extreme discrepancy of exposure, we propose the self-alignment module, leveraging scene relighting to constrain the illumination degree for following alignment and feature extraction. Detail repletion is proposed to enhance the texture details of scenes. Additionally, incorporating a hardware-sensitive constraint, we present the fusion-oriented architecture search to explore compact and efficient networks for fusion. The proposed method outperforms various competitive schemes, achieving a noteworthy 3.19% improvement in PSNR for general scenarios and an impressive 23.5% enhancement in misaligned scenarios. Moreover, it significantly reduces inference time by 69.1%. The code will be available at https://github.com/LiuZhu-CV/CRMEF.
Zhu Liu 0004, Jinyuan Liu 0001, Guanyao Wu, Xin Fan 0001, Risheng Liu
IEEE Trans. Circuits Syst. Video Technol.2
2024 Learning to Search a Lightweight Generalized Network for Medical Image Fusion
abstract
Image fusion is indispensable in a comprehensive medical imaging pipeline. By embracing deep learning technology, medical image fusion has achieved tremendous progress over the past few years. However, existing approaches make efforts on the specific type of medical image fusion task and may face difficulties in generalizing well. Moreover, most of them strain every nerve to design various architectures with an increase of the width of depth, placing an obstacle in running efficiency. To address the above problems, we propose an Auto-searching Light-weighted Multi-source Fusion network, namely ALMFnet, aiming at incorporating both software and hardware knowledge in a network architecture searching manner for medical image fusion. Specifically, the ALMFnet, consisting of two different feature-extracting modules and one fusion module, is developed to extract and refine multi-source features in a generalized model. Besides, motivated by the collaborative principle, we introduce hardware constraints for sufficient searching the each particular component, further reducing the complexity of the obtained model. Furthermore, to preserve important details in pathological image areas, we introduce a segmentation mask into the developed method. Experimental results demonstrate that our generalized model outperforms previous methods not only in terms of quantitative scores but also in model complexity. Source code will be available at https://github.com/RollingPlain/ALMFnet.
Pan Mu, Guanyao Wu, Jinyuan Liu 0001, Yuduo Zhang, Xin Fan 0001, Risheng Liu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Improving Misaligned Multi-Modality Image Fusion With One-Stage Progressive Dense Registration
abstract
Misalignments between multi-modality images pose challenges in image fusion, manifesting as structural distortions and edge ghosts. Existing efforts commonly resort to registering first and fusing later, typically employing two separate stages for registration, i.e., coarse registration and fine registration. Both stages directly estimate the respective target deformation fields. This paper contends that the separate two-stage registration lacks compactness, and the direct estimation of their target deformation fields falls short in accuracy. To tackle these challenges, we introduce IMF, a framework for improving misaligned multi-modality image fusion. Central to IMF is a One-stage Progressive Dense Registration (OPDR) scheme, which accomplishes the coarse-to-fine registration through only a one-stage optimization. Specifically, two pivotal components are involved in OPDR, a dense Deformation Field Fusion (DFF) module and a Progressive Feature Fine (PFF) module. The DFF aggregates the predicted multi-scale deformation sub-fields at the current scale, while the PFF progressively refines the remaining misaligned features. Together, they effectively and accurately estimate the final deformation fields. In addition, we develop a Transformer-Conv-based Fusion (TCF) subnetwork that considers local and long-range feature dependencies, allowing us to capture more informative features from the registered infrared and visible images for the generation of high-quality fused images. Extensive experimental analysis demonstrates the superiority of the proposed method in the fusion of misaligned cross-modality images. The code will be available athttps://github.com/wdhudiekou/IMF.
Di Wang 0018, Jinyuan Liu 0001, Long Ma 0002, Risheng Liu, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Multispectral Image Stitching via Global-Aware Quadrature Pyramid Regression
abstract
Image stitching is a critical task in panorama perception that involves combining images captured from different viewing positions to reconstruct a wider field-of-view (FOV) image. Existing visible image stitching methods suffer from performance drops under severe conditions since environmental factors can easily impair visible images. In contrast, infrared images possess greater penetrating ability and are less affected by environmental factors. Therefore, we propose an infrared and visible image-based multispectral image stitching method to achieve all-weather, broad FOV scene perception. Specifically, based on two pairs of infrared and visible images, we employ the salient structural information from the infrared images and the textual details from the visible images to infer the correspondences within different modality-specific features. For this purpose, a multiscale progressive mechanism coupled with quadrature correlation is exploited to improve regression in different modalities. Exploiting the complementary properties, accurate and credible homography can be obtained by integrating the deformation parameters of the two modalities to compensate for the missing modality-specific information. A global-aware guided reconstruction module is established to generate an informative and broad scene, wherein the attentive features of different viewpoints are introduced to fuse the source images with a more seamless and comprehensive appearance. We construct a high-quality infrared and visible stitching dataset for evaluation, including real-world and synthetic sets. The qualitative and quantitative results demonstrate that the proposed method outperforms the intuitive cascaded fusion-stitching procedure, achieving more robust and credible panorama generation. Code and dataset are available at https://github.com/Jzy2017/MSGA.
Zhiying Jiang, Zengxi Zhang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
IEEE Trans. Image Process.3
2024 GeSeNet: A General Semantic-Guided Network With Couple Mask Ensemble for Medical Image Fusion
abstract
At present, multimodal medical image fusion technology has become an essential means for researchers and doctors to predict diseases and study pathology. Nevertheless, how to reserve more unique features from different modal source images on the premise of ensuring time efficiency is a tricky problem. To handle this issue, we propose a flexible semantic-guided architecture with a mask-optimized framework in an end-to-end manner, termed as GeSeNet. Specifically, a region mask module is devised to deepen the learning of important information while pruning redundant computation for reducing the runtime. An edge enhancement module and a global refinement module are presented to modify the extracted features for boosting the edge textures and adjusting overall visual performance. In addition, we introduce a semantic module that is cascaded with the proposed fusion network to deliver semantic information into our generated results. Sufficient qualitative and quantitative comparative experiments (i.e., MRI-CT, MRI-PET, and MRI-SPECT) are deployed between our proposed method and ten state-of-the-art methods, which shows our generated images lead the way. Moreover, we also conduct operational efficiency comparisons and ablation experiments to prove that our proposed method can perform excellently in the field of multimodal medical image fusion. The code is available at https://github.com/lok-18/GeSeNet.
Jiawei Li 0016, Jinyuan Liu 0001, Shihua Zhou, Qiang Zhang 0008, Nikola K. Kasabov
IEEE Trans. Neural Networks Learn. Syst.2
2024 Coarse-to-fine multi-scale attention-guided network for multi-exposure image fusion
Jingrun Zheng, Xiaoke Shang, Jinyuan Liu 0001
Vis. Comput.5
2023 Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and Segmentation
abstract
Multi-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, e.g., fusion or segmentation, making it hard to reach ‘Best of Both Worlds’. To overcome this issue, in this paper, we propose a Multi-interactive Feature learning architecture for image fusion and Segmentation, namely SegMiF, and exploit dual-task correlation to promote the performance of both tasks. The SegMiF is of a cascade structure, containing a fusion sub-network and a commonly used segmentation sub-network. By slickly bridging intermediate features between two components, the knowledge learned from the segmentation task can effectively assist the fusion task. Also, the benefited fusion network supports the segmentation one to perform more pretentiously. Besides, a hierarchical interactive attention block is established to ensure fine-grained mapping of all the vital information between two tasks, so that the modality/semantic features can be fully mutual-interactive. In addition, a dynamic weight factor is introduced to automatically adjust the corresponding weights of each task, which can balance the interactive feature correspondence and break through the limitation of laborious tuning. Furthermore, we construct a smart multi-wave binocular imaging system and collect a full-time multi-modality benchmark with 15 annotated pixel-level categories for image fusion and segmentation. Extensive experiments on several public datasets and our benchmark demonstrate that the proposed method outputs visually appealing fused images and perform averagely 7.66% higher segmentation mIoU in the real-world scene than the state-of-the-art approaches. The source code and benchmark are available at https://github.com/JinyuanLiu-CV/SegMiF.
Jinyuan Liu 0001, Zhu Liu 0004, Guanyao Wu, Long Ma 0002, Risheng Liu, Zhongxuan Luo, Xin Fan 0001
ICCV1
2023 Bi-level Dynamic Learning for Jointly Multi-modality Image Fusion and Beyond
abstract
Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and neglecting others, seldom investigating their underlying connections for joint promotion. To overcome these limitations, we establish the hierarchical dual tasks-driven deep model to bridge these tasks. Concretely, we firstly construct an image fusion module to fuse complementary characteristics and cascade dual task-related modules, including a discriminator for visual effects and a semantic network for feature measurement. We provide a bi-level perspective to formulate image fusion and follow-up downstream tasks. To incorporate distinct task-related responses for image fusion, we consider image fusion as a primary goal and dual modules as learnable constraints. Furthermore, we develop an efficient first-order approximation to compute corresponding gradients and present dynamic weighted aggregation to balance the gradients for fusion learning. Extensive experiments demonstrate the superiority of our method, which not only produces visually pleasant fused results but also realizes significant promotion for detection and segmentation than the state-of-the-art approaches.
Zhu Liu 0004, Jinyuan Liu 0001, Guanyao Wu, Long Ma 0002, Xin Fan 0001, Risheng Liu
IJCAI2
2023 Multi-Spectral Image Stitching via Spatial Graph Reasoning
abstract
Multi-spectral image stitching leverages the complementarity between infrared and visible images to generate a robust and reliable wide field-of-view~(FOV) scene. The primary challenge of this task is to explore the relations between multi-spectral images for aligning and integrating multi-view scenes. Capitalizing on the strengths of Graph Convolutional Networks (GCNs) in modeling feature relationships, we propose a spatial graph reasoning based multi-spectral image stitching method that effectively distills the deformation and integration of multi-spectral images across different viewpoints. To accomplish this, we embed multi-scale complementary features from the same view position into a set of nodes. The correspondence across different views is learned through powerful dense feature embeddings, where both inter- and intra-correlations are developed to exploit cross-view matching and enhance inner feature disparity. By introducing long-range coherence along spatial and channel dimensions, the complementarity of pixel relations and channel interdependencies aids in the reconstruction of aligned multi-view features, generating informative and reliable wide FOV scenes. Moreover, we release a challenging dataset named ChaMS, comprising both real-world and synthetic sets with significant parallax, providing a new option for comprehensive evaluation. Extensive experiments demonstrate that our method surpasses the state-of-the-arts.
Zhiying Jiang, Zengxi Zhang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
ACM Multimedia3
2023 Learning a Graph Neural Network with Cross Modality Interaction for Image Fusion
abstract
Infrared and visible image fusion has gradually proved to be a vital fork in the field of multi-modality imaging technologies. In recent developments, researchers not only focus on the quality of fused images but also evaluate their performance in downstream tasks. Nevertheless, the majority of methods seldom put their eyes on mutual learning from different modalities, resulting in fused images lacking significant details and textures. To overcome this issue, we propose an interactive graph neural network (GNN)-based architecture between cross modality for fusion, called IGNet. Specifically, we first apply a multi-scale extractor to achieve shallow features, which are employed as the necessary input to build graph structures. Then, the graph interaction module can construct the extracted intermediate features of the infrared/visible branch into graph structures. Meanwhile, the graph structures of two branches interact for cross-modality and semantic learning, so that fused images can maintain the important feature expressions and enhance the performance of downstream tasks. Besides, the proposed leader nodes can improve information propagation in the same modality. Finally, we merge all graph features to get the fusion result. Extensive experiments on different datasets (i.e. TNO, MFNet, and M3FD) demonstrate that our IGNet can generate visually appealing fused images while scoring averagely 2.59% [email protected] and 7.77% mIoU higher in detection and segmentation than the compared state-of-the-art methods. The source code of the proposed IGNet can be available at https://github.com/lok-18/IGNet.
Jiawei Li 0016, Jiansheng Chen 0002, Jinyuan Liu 0001, Huimin Ma 0001
ACM Multimedia3
2023 Fearless Luminance Adaptation: A Macro-Micro-Hierarchical Transformer for Exposure Correction
abstract
Photographs taken with less-than-ideal exposure settings often display poor visual quality. Since the correction procedures vary significantly, it is difficult for a single neural network to handle all exposure problems. Moreover, the inherent limitations of convolutions, hinder the models ability to restore faithful color or details on extremely over-/under-exposed regions. To overcome these limitations, we propose a Macro-Micro-Hierarchical transformer, which consists of a macro attention to capture long-range dependencies, a micro attention to extract local features, and a hierarchical structure for coarse-to-fine correction. In specific, the complementary macro-micro attention designs enhance locality while allowing global interactions. The hierarchical structure enables the network to correct exposure errors of different scales layer by layer. Furthermore, we propose a contrast constraint and couple it seamlessly in the loss function, where the corrected image is pulled towards the positive sample and pushed away from the dynamically generated negative samples. Thus the remaining color distortion and loss of detail can be removed. We also extend our method as an image enhancer for low-light face recognition and low-light semantic segmentation. Experiments demonstrate that our approach obtains more attractive results than state-of-the-art methods quantitatively and qualitatively.
Gehui Li, Jinyuan Liu 0001, Long Ma 0002, Zhiying Jiang, Xin Fan 0001, Risheng Liu
ACM Multimedia2
2023 Bilevel Generative Learning for Low-Light Vision
abstract
Recently, there has been a growing interest in constructing deep learning schemes for Low-Light Vision (LLV). Existing techniques primarily focus on designing task-specific and data-dependent vision models on the standard RGB domain, which inherently contain latent data associations. In this study, we propose a generic low-light vision solution by introducing a generative block to convert data from the RAW to the RGB domain. This novel approach connects diverse vision problems by explicitly depicting data generation, which is the first in the field. To precisely characterize the latent correspondence between the generative procedure and the vision task, we establish a bilevel model with the parameters of the generative block defined as the upper level and the parameters of the vision task defined as the lower level. We further develop two types of learning strategies targeting different goals, namely low cost and high accuracy, to acquire a new bilevel generative learning paradigm. The generative blocks embrace a strong generalization ability in other low-light vision tasks through the bilevel optimization on enhancement tasks. Extensive experimental evaluations on three representative low-light vision tasks, namely enhancement, detection, and segmentation, fully demonstrate the superiority of our proposed approach. The code will be available at https://github.com/Yingchi1998/BGL.
Yingchi Liu, Zhu Liu 0004, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
ACM Multimedia4
2023 PAIF: Perception-Aware Infrared-Visible Image Fusion for Attack-Tolerant Semantic Segmentation
abstract
Infrared and visible image fusion is a powerful technique that combines complementary information from different modalities for downstream semantic perception tasks. Existing learning-based methods show remarkable performance, but are suffering from the inherent vulnerability of adversarial attacks, causing a significant decrease in accuracy. In this work, a perception-aware fusion framework is proposed to promote segmentation robustness in adversarial scenes. We first conduct systematic analyses about the components of image fusion, investigating the correlation with segmentation robustness under adversarial perturbations. Based on these analyses, we propose a harmonized architecture search with a decomposition-based structure to balance standard accuracy and robustness. We also propose an adaptive learning strategy to improve the parameter robustness of image fusion, which can learn effective feature extraction under diverse adversarial perturbations. Thus, the goals of image fusion (i.e., extracting complementary features from source modalities and defending attack) can be realized from the perspectives of architectural and learning strategies. Extensive experimental results demonstrate that our scheme substantially enhances the robustness, with gains of 15.3% mIOU of segmentation in the adversarial scene, compared with advanced competitors. The source codes are available at https://github.com/LiuZhu-CV/PAIF.
Zhu Liu 0004, Jinyuan Liu 0001, Benzhuang Zhang, Long Ma 0002, Xin Fan 0001, Risheng Liu
ACM Multimedia2
2023 Little Strokes Fell Great Oaks: Boosting the Hierarchical Features for Multi-exposure Image Fusion
abstract
In recent years, deep learning networks have made remarkable strides in the domain of multi-exposure image fusion. Nonetheless, prevailing approaches often involve directly feeding over-exposed and under-exposed images into the network, which leads to the under-utilization of inherent information present in the source images. Additionally, unsupervised techniques predominantly employ rudimentary weighted summation for color channel processing, culminating in an overall desaturated final image tone. To partially mitigate these issues, this study proposes a gamma correction module specifically designed to fully leverage latent information embedded within source images. Furthermore, a modified transformer block, embracing self-attention mechanisms, is introduced to optimize the fusion process. Ultimately, a novel color enhancement algorithm is presented to augment image saturation while preserving intricate details. The source code is available at https://github.com/ZhiyingDu/BHFMEF.
Pan Mu, Zhiying Du, Jinyuan Liu 0001, Cong Bai
ACM Multimedia3
2023 WaterFlow: Heuristic Normalizing Flow for Underwater Image Enhancement and Beyond
abstract
Underwater images suffer from light refraction and absorption, which impairs visibility and interferes the subsequent applications. Existing underwater image enhancement methods mainly focus on image quality improvement, ignoring the effect on practice. To balance the visual quality and application, we propose a heuristic normalizing flow for detection-driven underwater image enhancement, dubbed WaterFlow. Specifically, we first develop an invertible mapping to achieve the translation between the degraded image and its clear counterpart. Considering the differentiability and interpretability, we incorporate the heuristic prior into the data-driven mapping procedure, where the ambient light and medium transmission coefficient benefit credible generation. Furthermore, we introduce a detection perception module to transmit the implicit semantic guidance into the enhancement procedure, where the enhanced images hold more detection-favorable features and are able to promote the detection performance. Extensive experiments prove the superiority of our WaterFlow, against state-of-the-art methods quantitatively and qualitatively.
Zengxi Zhang, Zhiying Jiang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
ACM Multimedia3
2023 Exploring a Distillation with Embedded Prompts for Object Detection in Adverse Environments
Hao Fu 0004, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
PRCV (10)3
2023 Breaking Free From Fusion Rule: A Fully Semantic-Driven Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion plays a vital role in the field of computer vision. Previous approaches make efforts to design various fusion rules in the loss functions. However, these experimental designed fusion rules make the methods more and more complex. Besides, most of them only focus on boosting the visual effects, thus showing unsatisfactory performance for the follow-up high-level vision tasks. To address these challenges, in this letter, we develop a semantic-level fusion network to sufficiently utilize the semantic guidance, emancipating the experimental designed fusion rules. In addition, to achieve a better semantic understanding of the feature fusion process, a fusion block based on the transformer is presented in a multi-scale manner. Moreover, we devise a regularization loss function, together with a training strategy, to fully use semantic guidance from the high-level vision tasks. Compared with state-of-the-art methods, our method does not depend on the hand-crafted fusion loss function. Still, it achieves superior performance on visual quality along with the follow-up high-level vision tasks.
Zhu Liu 0004, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
IEEE Signal Process. Lett.3
2023 Learning a Coordinated Network for Detail-Refinement Multiexposure Image Fusion
abstract
Nowadays, deep learning has made rapid progress in the field of multi-exposure image fusion. However, it is still challenging to extract available features while retaining texture details and color. To address this difficult issue, in this paper, we propose a coordinated learning network for detail-refinement in an end-to-end manner. Firstly, we obtain shallow feature maps from extreme over/under-exposed source images by a collaborative extraction module. Secondly, smooth attention weight maps are generated under the guidance of a self-attention module, which can draw a global connection to correlate patches in different locations. With the cooperation of the two aforementioned used modules, our proposed network can obtain a coarse fused image. Moreover, by assisting with an edge revision module, edge details of fused results are refined and noise is suppressed effectively. We conduct subjective qualitative and objective quantitative comparisons between the proposed method and twelve state-of-the-art methods on two available public datasets, respectively. The results show that our fused images significantly outperform others in visual effects and evaluation metrics. In addition, we also perform ablation experiments to verify the function and effectiveness of each module in our proposed method. The source code can be achieved athttps://github.com/lok-18/LCNDR.
Jiawei Li 0016, Jinyuan Liu 0001, Shihua Zhou, Qiang Zhang 0008, Nikola K. Kasabov
IEEE Trans. Circuits Syst. Video Technol.2
2023 A unified image fusion framework with flexible bilevel paradigm integration
Jinyuan Liu 0001, Zhiying Jiang, Guanyao Wu, Risheng Liu, Xin Fan 0001
Vis. Comput.1
2022 Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object Detection
abstract
This study addresses the issue of fusing infrared and visible images that appear differently for object detection. Aiming at generating an image of high visual quality, previous approaches discover commons underlying the two modalities and fuse upon the common space either by iterative optimization or deep networks. These approaches neglect that modality differences implying the complementary information are extremely important for both fusion and subsequent detection task. This paper proposes a bilevel optimization formulation for the joint problem of fusion and detection, and then unrolls to a target-aware Dual Adversarial Learning (TarDAL) network for fusion and a commonly used detection network. The fusion network with one generator and dual discriminators seeks commons while learning from differences, which preserves structural information of targets from the infrared and textural details from the visible. Furthermore, we build a synchronized imaging system with calibrated infrared and optical sensors, and collect currently the most comprehensive benchmark covering a wide range of scenarios. Extensive experiments on several public datasets and our benchmark demonstrate that our method outputs not only visually appealing fusion but also higher detection mAP than the state-of-the-art approaches. The source code and benchmark are available at https://github.com/dlut-dimt/TarDAL.
Jinyuan Liu 0001, Xin Fan 0001, Zhanbo Huang, Guanyao Wu, Risheng Liu, Zhongxuan Luo
CVPR1
2022 ReCoNet: Recurrent Correction Network for Fast and Efficient Multi-modality Image Fusion
Zhanbo Huang, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu, Zhongxuan Luo
ECCV (18)2
2022 Unsupervised Misaligned Infrared and Visible Image Fusion via Cross-Modality Image Generation and Registration
abstract
Recent learning-based image fusion methods have marked numerous progress in pre-registered multi-modality data, but suffered serious ghosts dealing with misaligned multi-modality data, due to the spatial deformation and the difficulty narrowing cross-modality discrepancy. To overcome the obstacles, in this paper, we present a robust cross-modality generation-registration paradigm for unsupervised misaligned infrared and visible image fusion (IVIF). Specifically, we propose a Cross-modality Perceptual Style Transfer Network (CPSTN) to generate a pseudo infrared image taking a visible image as input. Benefiting from the favorable geometry preservation ability of the CPSTN, the generated pseudo infrared image embraces a sharp structure, which is more conducive to transforming cross-modality image alignment into mono-modality registration coupled with the structure-sensitive of the infrared image. In this case, we introduce a Multi-level Refinement Registration Network (MRRN) to predict the displacement vector field between distorted and pseudo infrared images and reconstruct registered infrared image under the mono-modality setting. Moreover, to better fuse the registered infrared images and visible images, we present a feature Interaction Fusion Module (IFM) to adaptively select more meaningful features for fusion in the Dual-path Interaction Fusion Network (DIFN). Extensive experimental results suggest that the proposed method performs superior capability on misaligned cross-modality image fusion.
Di Wang 0018, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
IJCAI2
2022 Learn to Search a Lightweight Architecture for Target-Aware Infrared and Visible Image Fusion
abstract
Deep learning technology has recently achieved remarkable progress in infrared and visible image fusion. Nevertheless, existing methods have encountered blurred targets or unfaithful textural details on their fused results. They suffer from high computational expenses, thus unable to directly serve the subsequent high-level vision tasks. In this letter, to alleviate this issue, we proposed leveraging a lightweight architecture based on Neural Architecture Search (NAS) to realize the infrared and visible image fusion in an end-to-end manner, significantly reducing the computational expenses and runtime. Concretely, we construct a search-based architecture to explore the feature representation across different modalities automatically. Then a saliency-based loss function is designed to retain both the distinct target and texture details. Motivated by the cooperative principle, we also formulate a flexible hardware-sensitive regularization constraint in our loss function for discovering efficient operations. As a result, we can generate a target-distinct fused result with high efficiency. Extensive qualitative and quantitative experiments reveal that our method has superior performance against the state-of-the-art methods, especially highlighting the target, retaining realistic details, and achieving fast running speed. Specifically, our method increases by 150% in time, reduces the FLOPS by 21.3% and reduces the model parameters by 25%.
Jinyuan Liu 0001, Guanyao Wu, Risheng Liu, Xin Fan 0001
IEEE Signal Process. Lett.1
2022 Learning a Deep Multi-Scale Feature Ensemble and an Edge-Attention Guidance for Image Fusion
abstract
Image fusion integrates a series of images acquired from different sensors,e.g., infrared and visible, outputting an image with richer information than either one. Traditional and recent deep-based methods have difficulties in preserving prominent structures and recovering vital textural details for practical applications. In this article, we propose a deep network for infrared and visible image fusion cascading a feature learning module with a fusion learning mechanism. Firstly, we apply a coarse-to-fine deep architecture to learn multi-scale features for multi-modal images, which enables discovering prominent common structures for later fusion operations. The proposed feature learning module requires no well-aligned image pairs for training. Compared with the existing learning-based methods, the proposed feature learning module can ensemble numerous examples from respective modals for training, increasing the ability of feature representation. Secondly, we design an edge-guided attention mechanism upon the multi-scale features to guide the fusion focusing on common structures, thus recovering details while attenuating noise. Moreover, we provide a new aligned infrared and visible image fusion dataset, RealStreet, collected in various practical scenarios for comprehensive evaluation. Extensive experiments on two benchmarks, TNO and RealStreet, demonstrate the superiority of the proposed method over the state-of-the-art in terms of both visual inspection and objective analysis on six evaluation metrics. We also conduct the experiments on the FLIR and NIR datasets, containing foggy weather and poor light conditions, to verify the generalization and robustness of the proposed method.
Jinyuan Liu 0001, Xin Fan 0001, Ji Jiang, Risheng Liu, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.1
2022 Attention-Guided Global-Local Adversarial Learning for Detail-Preserving Multi-Exposure Image Fusion
abstract
Deep learning networks have recently demonstrated yielded impressive progress for multi-exposure image fusion. However, how to restore realistic texture details while correcting color distortion is still a challenging problem to be solved. To alleviate the aforementioned issues, in this paper, we propose an attention-guided global-local adversarial learning network for fusing extreme exposure images in a coarse-to-fine manner. Firstly, the coarse fusion result is generated under the guidance of attention weight maps, which acquires the essential region of interest from both sides. Secondly, we formulate an edge loss function, along with a spatial feature transform layer, for refining the fusion process. So that it can take full use of the edge information to deal with blurry edges. Moreover, by incorporating global-local learning, our method can balance pixel intensity distribution and correct the color distortion on spatially varying source images from both image/patch perspectives. Such a global-local discriminator ensures all the local patches of the fused images align with realistic normal-exposure ones. Extensive experimental results on two publicly available datasets show that our method drastically outperforms state-of-the-art methods in visual inspection and objective analysis. Furthermore, sufficient ablation experiments prove that our method has significant advantages in generating high-quality fused results with appealing details, clear targets, and faithful color. Source code will be available athttps://github.com/JinyuanLiu-CV/AGAL.
Jinyuan Liu 0001, Jingjie Shang, Risheng Liu, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Global structure-guided learning framework for underwater image enhancement
Runjia Lin, Jinyuan Liu 0001, Risheng Liu, Xin Fan 0001
Vis. Comput.2
2021 Multiple Task-Oriented Encoders for Unified Image Fusion
abstract
Image fusion methods have achieved incredible progress, but they are vulnerable to handling a certain type of fusion task rather than considering deeper relations between cross-realm task correlations. To achieve this, we integrate different image fusion tasks into a unified network. Our method is accomplished through multiple task-oriented encoders and a generic decoder, in addition to a self-adapting loss function. The taskoriented encoders are trained to learn task-specific features, while the generic decoder reconstructs the fused features to generate a comprehensive image. Subsequently, by introducing the self-adapting loss in our method, it can automatically adjust itself to source data characteristics on different tasks. Besides, we formulate a training strategy based on bilevel optimization to update the multi-encoder and generic decoder in an alternative manner. Extensive experimental results demonstrate the superior performance of our method over the stateof-the-art methods.
Zhuoxiao Li, Jinyuan Liu 0001, Risheng Liu, Xin Fan 0001, Zhongxuan Luo, Wen Gao 0001
ICME2
2021 Halder: Hierarchical Attention-Guided Learning with Detail-Refinement for Multi-Exposure Image Fusion
abstract
Deep learning techniques have yielded impressive progress in the field of computational imaging. Existing approaches ignore designing specific constrain on illumination or edges, making them limited in handling asymmetric halos and more likely to generate a fusion result with color discrepancy or blurred edges. To alleviate these issues, we propose a hierarchical attention-guided learning with detail-refinement, termed as HALDeR, to tackle the multi-exposure fusion (MEF) task in a coarse-to-fine manner. Firstly, a hierarchical attention network is designed to produce a fusion result by calculating well-exposed areas under different illumination. Secondly, we develop a collaborative-refine module for preventing the missing details and correcting distorted color simultaneously. Moreover, adversarial learning is employed at end of our network, which can effectively alleviate other remaining artifacts (e.g., ringing effect and noises). Extensive quantitative and qualitative results on two publicly available datasets demonstrate that our HALDeR performs favorably against the state-of-the-art methods in generating vivid color and faithful detail. Source code will be available at https://github.com/JinyuanLiu-CV/HALDeR.
Jinyuan Liu 0001, Jingjie Shang, Risheng Liu, Xin Fan 0001
ICME1
2021 Searching a Hierarchically Aggregated Fusion Architecture for Fast Multi-Modality Image Fusion
abstract
Multi-modality image fusion refers to generating a complementary image that integrates typical characteristics from source images. In recent years, we have witnessed the remarkable progress of deep learning models for multi-modality fusion. Existing CNN-based approaches strain every nerve to design various architectures for realizing these tasks in an end-to-end manner. However, these handcrafted designs are unable to cope with the high demanding fusion tasks, resulting in blurred targets and lost textural details. To alleviate these issues, in this paper, we propose a novel approach, aiming at searching effective architectures according to various modality principles and fusion mechanisms.
Risheng Liu, Zhu Liu 0004, Jinyuan Liu 0001, Xin Fan 0001
ACM Multimedia3
2021 SMoA: Searching a Modality-Oriented Architecture for Infrared and Visible Image Fusion
abstract
Nowadays, driven by the high demand for autonomous driving and surveillance, infrared and visible image fusion (IVIF) has attracted significant attention from both the industry and research community. Existing learning-based IVIF methods tried to design various architectures to extract features. Still, these hand-crafted designed architectures cannot adequately represent the typical features of different modalities, resulting in undesirable artifacts on the fused results. To alleviate this issue, we propose a Neural Architecture Search (NAS)-based deep learning network to realize the IVIF task, which can automatically discover the modality-oriented feature representation. Our network is accomplished through two modality-oriented encoders and a unified decoder, in addition to a self-visual saliency weight module (SvSW). The two modality-oriented encoders target to learn different intrinsic feature representations automatically from infrared-/visible- modality images. Subsequently, these intermediate features are merged via the SvSW module. Finally, the fused image is recovered by a unified decoder. Extensive experiments demonstrate that our method outperforms the state-of-the-art approaches by a large margin, especially in generating distinct targets and abundant details.
Jinyuan Liu 0001, Zhanbo Huang, Risheng Liu, Xin Fan 0001
IEEE Signal Process. Lett.1
2021 Learning Hadamard-Product-Propagation for Image Dehazing and Beyond
abstract
Image dehazing has evolved into an attractive research field in the computer vision community in the past few decades. Previous traditional approaches attempt to design energy-based objective functions. However, they cannot accurately express the intrinsic characteristics of the images, posing weak adaptation ability for real-world complex scenarios. More recently, deep learning techniques for image dehazing have matured and become more reliable, showing outstanding performance. Nevertheless, these methods heavily depend on training data, restricting their application ranges. More importantly, both traditional and deep learning approaches all ignore a common issue, noises/artifacts always appear in the recovery process. To this end, a new Hadamard-Product (HP) model is proposed, which consists of a series of data-driven priors. Based on this model, we derive a Learnable Hadamard-Product-Propagation (LHPP) by cascading a series of principle-inspired guidance and recovery modules. In which, the principle-inspired guidance related to transmission is endowed the smoothness property, the other recovery module satisfies the distribution of natural images. The Hadamard-product-based propagations is generated in our developed learnable framework for the task of image dehazing. In this way, we can eliminate noises/artifacts in the recovery procedure to obtain the ideal outputs. Subsequently, since the generality of our HP model, we successfully extend our LHPP to settle low-light image enhancement and underwater image enhancement problems. A series of analytical experiments are performed to verify our effectiveness. Plenty of performance evaluations on three complex tasks fully reveal our superiority against multiple state-of-the-art methods.
Risheng Liu, Jinyuan Liu 0001, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.3
2021 A Bilevel Integrated Model With Data-Driven Layer Ensemble for Multi-Modality Image Fusion
abstract
Image fusion plays a critical role in a variety of vision and learning applications. Current fusion approaches are designed to characterize source images, focusing on a certain type of fusion task while limited in a wide scenario. Moreover, other fusion strategies (i.e., weighted averaging, choose-max) cannot undertake the challenging fusion tasks, which furthermore leads to undesirable artifacts facilely emerged in their fused results. In this paper, we propose a generic image fusion method with a bilevel optimization paradigm, targeting on multi-modality image fusion tasks. Corresponding alternation optimization is conducted on certain components decoupled from source images. Via adaptive integration weight maps, we are able to get the flexible fusion strategy across multi-modality images. We successfully applied it to three types of image fusion tasks, including infrared and visible, computed tomography and magnetic resonance imaging, and magnetic resonance imaging and single-photon emission computed tomography image fusion. Results highlight the performance and versatility of our approach from both quantitative and qualitative aspects.
Risheng Liu, Jinyuan Liu 0001, Zhiying Jiang, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.2
2019 Compounded Layer-Prior Unrolling: A Unified Transmission-Based Image Enhancement Framework
abstract
Improving the quality of images degraded by various transmission media has important practical significance. Such enhancement tasks involve resolving both transmission degradation and residual contamination including imaging noise, color distortion, and occlusions. Existing methods typically develop the priors on natural scenes to resolve ill-posed problems separately. However, the solutions derived from hand-crafted priors may fail on specific regions where a priori assumptions break, and recent data-driven methods highly depend on training data owing to the absence of effective priors. Based on a unified formulation for transmission-based image enhancement tasks, we develop a compounded unrolling framework to generate hybrid image layer propagations. Specifically, as multiple deeply-trained priors are integrated into the iterative propagation scheme, the deep model can recognize specific task properties and data distributions for different applications. Both quantitative and qualitative experiments demonstrate the superior performance of the proposed framework on various transmission-based tasks (haze removal, underwater image enhancement and rain removal).
Risheng Liu, Minjun Hou, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo
ICME3