Long Ma 0002

dblp:93/5262-2 · DBLP profile ↗
← Back
83ranked-venue papers
10as first author
72since 2021 · last 2026
0000-0001-5125-0198ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 62 · 6 first-author · 52 since 2021Artificial intelligence and machine learning · 38 · 6 first-author · 36 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning 3D Occupancy from Beam Overlap in 2D Rotating mmWave Radar
abstract
Robust 3D perception under adverse weather is critical for autonomous systems. While mmWave Radars are inherently weather-resistant, conventional 2D rotating Radar sensors lack direct elevation resolution, limiting their 3D perception ability. Although 4D imaging radars can provide elevation information, they typically suffer from limited coverage and range. In this work, we exploit a key observation about mechanically rotating 2D mmWave Radars: in each sweep, an overlap exists between adjacent azimuth beam coverage due to the width of the main lobe, which makes the reflected intensity difference imply object materials and geometric shapes, including elevation. With this observation, we propose a method that learns 3D occupancy by disentangling bird’s-eye view (BEV) layout and elevation estimation from one frame Radar scan. Specifically, we partition one sweep into two interleaved subsets, corresponding to overlapping beam directions, and utilize them to infer coarse geometric structure through spatial differences and intensity patterns. Extensive quantitative and qualitative evaluations on two real-world datasets demonstrate that our proposed method outperforms existing baselines. The codes will be publicly available.
Ruifeng Nie, Long Ma 0002, Chengpei Xu, Yu Liu 0012, Weimin Wang 0007
AAAI3
2026 SNRD-Net: SNR-aware dual enhancement network for low-light images
Muhammad Zain Ul Abideen, Long Ma 0002, Risheng Liu
Comput. Vis. Image Underst.3
2026 Versatile Luminosity Tuning: Relighting Illumination via Dual-Prompt Exposure Correction
Jinyuan Liu 0001, Gehui Li, Zhiying Jiang, Long Ma 0002, Miao Zhang 0004, Xin Fan 0001, Risheng Liu
Int. J. Comput. Vis.4
2026 Contourlet Refinement Gate Framework for Thermal Spectrum Distribution Regularized Infrared Image Super-Resolution
Yang Zou 0004, Zhixin Chen, Xingyuan Li 0005, Long Ma 0002, Jinyuan Liu 0001
Int. J. Comput. Vis.5
2026 HiEn: Hierarchical ensemble learning for semi-supervised medical image segmentation
Long Ma 0002, Xinwei Xue, Chengpei Xu, Weimin Wang 0007, Yi Wang 0037
Neurocomputing1
2026 SWG-Fusion: Soft weather-guided multimodal fusion with VLM-assistance for BEV object detection under harsh weather
Weimin Wang 0007, Ruifeng Nie, Yingchi Liu, Long Ma 0002, Chengpei Xu, Qi Jia 0001, Yu Liu 0012, Na Lei
Pattern Recognit.4
2026 DSOS-UIE: Binarized Decoupled Synergistic Optimization Strategy for Underwater Image Enhancement
abstract
Underwater images typically suffer from two main types of degradation: reduced visibility caused by scattering and color distortion due to color cast. Most existing deep learning-based enhancement methods adopt end-to-end architectures to address both issues simultaneously. However, this design not only limits the model’s generalization capability but also hinders practical deployment due to excessive computational overhead. To this end, this paper proposes a binarized Decoupled Synergistic Optimization Strategy (DSOS), which explicitly decouples scattering and color cast degradations and performs collaborative optimization through specialized subtask modules. Each subtask learns purer features under the guidance of independent supervised signals, while a cascaded architecture ensures effective global restoration. Furthermore, cross-module collaborative optimization effectively mitigates the performance degradation caused by binarization, achieving a favorable balance between efficiency and high accuracy. Experimental results on multiple publicly available underwater image datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches in both restoration quality and computational efficiency.
Ruosheng Lu, Donghui Yang, Yan Zhang 0002, Boying Wang, Long Ma 0002
IEEE Trans. Circuits Syst. Video Technol.7
2026 Enhancing Underwater Images via Resonant Fusion
abstract
Recent advances in learning-based underwater image enhancement have achieved remarkable progress. However, the inherent diversity and complexity of underwater scenes still limit the ability of existing approaches to simultaneously restore fine structural details and global image layouts. To address this challenge, we propose a Resonant Fusion (ReFu) framework that explicitly leverages complementary information in both spatial and frequency domains. Specifically, we design a frequency decomposer and a spatial decomposer to capture high- and low-frequency cues from different perspectives. A resonant fuser is then introduced to adaptively integrate high-frequency resonances for detail refinement and low-frequency resonances for structural consistency. This fine-grained cross-domain fusion significantly improves structural preservation and detail enhancement, thereby generating visually more natural and perceptually friendly underwater images. Extensive quantitative and qualitative evaluations across diverse underwater benchmarks show that ReFu consistently surpasses state-of-the-art methods by a clear margin. Comprehensive ablation studies further validate the effectiveness of each module and prove the necessity of the proposed ReFu mechanism. Our code is available at https://github.com/CircleQa/ReFu-main.
Xinwei Xue, Zimeng Xu, Jincheng Yuan, Jingchun Zhou, Chengpei Xu, Xiaoke Shang, Long Ma 0002, Weimin Wang 0007
IEEE Trans. Image Process.8
2025 CoA: Towards Real Image Dehazing via Compression-and-Adaptation
abstract
Learning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constraints and the diversity of real-world scenes. Therefore, there is an urgent need for an algorithm that excels in both efficiency and adaptability to address real image dehazing effectively. This work proposes a Compression-and-Adaptation (CoA) computational flow to tackle these challenges from a divide-and-conquer perspective. First, model compression is performed in the synthetic domain to develop a compact dehazing parameter space, satisfying efficiency demands. Then, a bilevel adaptation in the real domain is introduced to be fearless in unknown real environments by aggregating the synthetic dehazing capabilities during the learning process. Leveraging a succinct design free from additional constraints, our CoA exhibits domain-irrelevant stability and model-agnostic flexibility, effectively bridging the model chasm between synthetic and real domains to further improve its practical utility. Extensive evaluations and analyses underscore the approach's superiority and effectiveness. The code is publicly available at https://github.com/fyxnl/COA.
Long Ma 0002, Yan Zhang 0002, Jinyuan Liu 0001, Weimin Wang 0007, Guang-Yong Chen, Chengpei Xu, Zhuo Su 0001
CVPR1
2025 Rethinking Reconstruction and Denoising in the Dark: New Perspective, General Architecture and Beyond
abstract
Recently, enhancing image quality in the original RAW domain has garnered significant attention, with denoising and reconstruction emerging as fundamental tasks. Although some works attempt to couple these tasks, they primarily focus on cascade learning while neglecting task associativity within a broader parameter space, leading to suboptimal performance. This work introduces a novel approach by rethinking denoising and reconstruction from a "backbone-head" perspective, leveraging the stronger shared parameter space offered by the backbone, compared to the encoder used in existing works. We derive task-specific heads with fewer parameters to mitigate learning pressure. By incorporating chromaticity-and-noise perception module into the backbone and introducing task-specific supervision during training, we enable simultaneous high-quality results for reconstruction and denoising. Additionally, we design a dual-head interaction module to capture the latent correspondence between the two tasks, significantly enhancing multi-task accuracy. Extensive experiments validate the superiority of the proposed method. Code is available at: https://github.com/csmty/CANS.
Tengyu Ma 0004, Long Ma 0002, Ziye Li, Yuetong Wang, Jinyuan Liu 0001, Chengpei Xu, Risheng Liu
CVPR2
2025 DEAL: Data-Efficient Adversarial Learning for High-Quality Infrared Imaging
abstract
Thermal imaging is often compromised by dynamic, complex degradations caused by hardware limitations and unpredictable environmental factors. The scarcity of high-quality infrared data, coupled with the challenges of dynamic, intricate degradations, makes it difficult to recover details using existing methods. In this paper, we introduce thermal degradation simulation integrated into the training process via a mini-max optimization, by modeling these degraded factors as adversarial attacks on thermal images. The simulation is dynamic to maximize objective functions, thus capturing a broad spectrum of degraded data distributions. This approach enables training with limited data, thereby improving model performance. Additionally, we introduce a dual-interaction network that combines the benefits of spiking neural networks with scale transformation to capture degraded features with sharp spike signal intensities. This architecture ensures compact model parameters while preserving efficient feature representation. Extensive experiments demonstrate that our method not only achieves superior visual quality under diverse single and composited degradation, but also delivers a significant reduction in processing when trained on only fifty clear images, outperforming existing techniques in efficiency and accuracy. The source code will be available at https://github.com/LiuZhu-CV/DEAL.
Zhu Liu 0004, Jinyuan Liu 0001, Fanqi Meng, Long Ma 0002, Risheng Liu
CVPR5
2025 DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution
abstract
Infrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low spatial resolution and complex degradations, consistently challenge imaging quality and subsequent visual tasks. Hence, infrared image super-resolution (IISR) has been developed to address this challenge. While recent developments in diffusion models have greatly advanced this field, current methods to solve it either ignore the unique modal characteristics of infrared imaging or overlook the machine perception requirements. To bridge these gaps, we propose DifIISR, an infrared image super-resolution diffusion model optimized for visual quality and perceptual performance. Our approach achieves task-based guidance for diffusion by injecting gradients derived from visual and perceptual priors into the noise during the reverse process. Specifically, we introduce an infrared thermal spectrum distribution regulation to preserve visual fidelity, ensuring that the reconstructed infrared images closely align with high-resolution images by matching their frequency components. Subsequently, we incorporate various visual foundational models as the perceptual guidance for downstream visual tasks, infusing generalizable perceptual features beneficial for detection and segmentation. As a result, our approach gains superior visual results while attaining State-Of-The-Art downstream task performance. Code is available at https://github.com/zirui0625/DifIISR
Xingyuan Li 0005, Yang Zou 0004, Zhixin Chen, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001
CVPR7
2025 DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each modality. Existing approaches typically treat image fusion and subsequent high-level tasks as separate processes, resulting in fused images that offer only marginal gains in task performance and fail to provide constructive feedback for optimizing the fusion process. To overcome these limitations, we propose a Discriminative Cross-Dimension Evolutionary Learning Framework, termed DCEvo, which simultaneously enhances visual quality and perception accuracy. Leveraging the robust search capabilities of Evolutionary Learning, our approach formulates the optimization of dual tasks as a multi-objective problem by employing an Evolutionary Algorithm (EA) to dynamically balance loss function parameters. Inspired by visual neuroscience, we integrate a Discriminative Enhancer (DE) within both the encoder and decoder, enabling the effective learning of complementary features from different modalities. Additionally, our Cross-Dimensional Embedding (CDE) block facilitates mutual enhancement between high-dimensional task features and low-dimensional fusion features, ensuring a cohesive and efficient feature integration process. Experimental results on three benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, achieving an average improvement of 9.32% in visual quality while also enhancing subsequent high-level tasks. The code is available at https://github.com/Beate-Suy-Zhang/DCEvo.
Jinyuan Liu 0001, Qingyun Mei, Xingyuan Li 0005, Yang Zou 0004, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001
CVPR7
2025 TextMEF: Text-guided Prompt Learning for Multi-exposure Image Fusion
abstract
Multi-exposure image fusion~(MEF) aims to integrate a set of low dynamic range images, producing a single image with a higher dynamic range than either one. Despite significant advancements, current MEF approaches still struggle to handle extremely over- or under-exposed conditions, resulting in unsatisfactory visual effects such as hallucinated details and distorted color tones. With this regard, we propose TextMEF, a prompt-driven fusion method enhanced by prompt learning, for multi-exposure image fusion. Specifically, we learn a set of prompts based on text-image similarity among negative and positive samples (over-exposed, under-exposed images, and well-exposed ones). These learned prompts are seamlessly integrated into the loss function, providing high-level guidance for constraining non-uniform exposure regions. Furthermore, we develop a attention Mamba module effectively translates over-/under- exposed regional features into exposure invariant space and ensure them to build efficient long-range dependency to high dynamic range image. Extensive experimental results on three publicly available benchmarks demonstrate that our TextMEF significantly outperforms state-of-the-art approaches in both visual inspection and objective analysis.
Jinyuan Liu 0001, Qianjun Huang, Guanyao Wu, Di Wang 0018, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001
IJCAI6
2025 Degradation-Aware One-Step Diffusion Model for Content-Sensitive Super-Resolution in the Dark
abstract
Diffusion-based super-resolution methods have achieved impressive results under normal lighting conditions. However, their performance in low-light scenarios faces fundamental limitations due to two inherent challenges. First, the characteristic noise patterns and complex degradation features in severely underexposed images create significant obstacles for diffusion models to establish reliable noise prediction mechanisms. Second, these methods often fail to establish effective coupling between the degradation priors of low-light observations and the reconstruction process, resulting in compromised detail recovery and unrealistic texture synthesis.To address these limitations, we propose Degradation-aware Adaptation with Representation Embedding (DARE) method, a novel one-step diffusion framework specifically designed for super-resolution in dark environments. DARE employs a degradation-aware low-rank adaptation strategy that dynamically adjusts model parameters conditioned on degradation-specific features, effectively addressing compound degradations such as low-light, blur, and noise. Furthermore, we introduce a content-sensitive representation embedding mechanism, integrating complementary spatial and frequency domain priors through a bilinear cross-attention module. This module explicitly captures second-order statistical correlations, enriching semantic understanding and detail recovery during the denoising process. Extensive experiments across diverse low-light scenarios demonstrate that DARE outperforms state-of-the-art methods in terms of both visual quality and perceptual accuracy. The code is available at https://github.com/csmty/DARE.
Tengyu Ma 0004, Jiafa Ruan, Yuetong Wang, Guangchao Han, Zhu Liu 0004, Long Ma 0002, Risheng Liu
ACM Multimedia6
2025 Inter-Task Weaving in Image Enhancement: From a New Unified Architecture to a Better Meta-Representation Learning
abstract
Image enhancement is a classical and enduring challenge in computer vision, seeking to produce high-quality images from corrupted observations. Unlike existing methods that target specific tasks, this work focuses on endowing the model with generic inductive capabilities, enabling fast adaptation to previously unseen enhancement tasks. Specifically, we investigate the inter-task weaving from both structural and parametric perspectives. Structurally, we establish inter-task weaving under a Hadamard view by designing a unified architecture called Degradation Unraveling Network (DUNet) tailored for diverse enhancement tasks, which incorporates a progressive degradation unraveling mechanism for fine-grained enhancement. Parametrically, we reveal the task-agnostic nature of degradation estimation parameters and treat them as meta-representations. A Bilevel Purify Modeling (BPM) framework is then proposed to reinforce their latent unified representation, where only the degradation-related parameters are optimized as meta-representations. Based on this design, a task-aware adaptation solution is further introduced, only the remaining parameters are allowed to be fine-tuned efficiently and enabling fast task adaptation. Extensive performance evaluations on three representative image enhancement tasks demonstrate the effectiveness and superiority of our method. The adaptability of our method is further verified by a series of algorithm analyses.
Siqi Xu, Long Ma 0002, Zhu Liu 0004, Guangchao Han, Tengyu Ma 0004, Risheng Liu
ACM Multimedia3
2025 Toward a Training-Free Plug-and-Play Refinement Framework for Infrared and Visible Image Registration and Fusion
abstract
Infrared and Visible Image Fusion (IVIF) under unregistered conditions has been of great interest in various visual tasks under challenging environments. While existing approaches often demonstrate promising results on specific benchmarks, they tend to exhibit performance drops in unseen scenarios and incur high computational overhead when retrained on new datasets. To address these challenges, we propose TRACE, a Training-free Reinforcement-based Alignment method for Cross-modality Enhancement, which incorporates Evaluator, a rewarding network, into an evaluation-driven Reinforcement Learning (RL) framework, enabling efficient and plug-and-play refinement of any existing registration approach. Specifically, TRACE constructs the Evaluator network to assess the alignment quality of the given registration model, generating confidence scores and adjustment masks via spatial and channel attention. Leveraging these cues as RL rewards, TRACE iteratively refines the registration network to mitigate misalignments until the accumulated improvement is satisfied. Due to its training-free and plug-and-play nature, TRACE notably enhances fusion results across diverse and unseen scenarios. TRACE achieves impressive improvements in different methods across diverse datasets with minimal computational cost. The project page is available at https://github.com/pubyLu/TRACE.
Yang Zou 0004, Xingyuan Li 0005, Xingyue Zhu, Kaiqi Han, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001
ACM Multimedia7
2025 Bright to Dark: Stage-wise Bilevel Knowledge Transfer for Seeing Text in the Dark
abstract
Localizing text under low-light conditions has gained attention, with typical approaches relying on two stage cascading modules that combine low-light enhancement and text localization. However, these often require additional enhancement modules and cause inefficiency in joint optimization. In this work, we address the challenge by adopting a novel approach: tailoring the detector for low light conditions through knowledge distillation from normal light conditions, without relying on any enhancement module. First, we design a Graph Topological Aggregation (GTA) model that utilizes the message passing mechanism of graph neural networks to structurally represent text topology and facilitate structured feature expression in knowledge transfer. We then introduce two specially designed knowledge transfer constraints aimed at enhancing the learning of text's multi-scale features and topological knowledge. Finally,we propose a Stage-wise Bilevel Knowledge Transfer learning strategy that designates the low-light learning process as the upper-level task, while treating normal light learning as the lower-level task, effectively addressing the coupling issues and sequential dependencies prevalent during the distillation process. Extensive experiments underscore the approach's superiority.
Chengpei Xu, Long Ma 0002, Weimin Wang 0007, Feng Xia 0001, Binghao Li, Wenjie Zhang 0001
ACM Multimedia3
2025 Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark
abstract
We engage in the relatively underexplored task named thermal infrared image enhancement. Existing infrared image enhancement methods primarily focus on tackling individual degradations, such as noise, contrast, and blurring, making it difficult to handle coupled degradations. Meanwhile, all-in-one enhancement methods, commonly applied to RGB sensors, often demonstrate limited effectiveness due to the significant differences in imaging models. In sight of this, we first revisit the imaging mechanism and introduce a Recurrent Prompt Fusion Network (RPFN). Specifically, the RPFN initially establishes prompt pairs based on the thermal imaging process. For each type of degradation, we fuse the corresponding prompt pairs to modulate the model's features, providing adaptive guidance that enables the model to better address specific degradations under single or multiple conditions.In addition, a selective recurrent training mechanism is introduced to gradually refine the model's handling of composite cases to align the enhancement process, which not only allows the model to remove camera noise and retain key structural details, but also enhancing the overall contrast of the thermal image. Furthermore, we introduce the most comprehensive high-quality infrared benchmark covering a wide range of scenarios. Extensive experiments substantiate that our approach not only delivers promising visual results under specific degradation but also significantly improves performance on complex degradation scenes, achieving a notable 8.76% improvement.
Jinyuan Liu 0001, Zhu Liu 0004, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Risheng Liu
NeurIPS5
2025 Bilevel Fast Scene Adaptation for Low-Light Image Enhancement
Long Ma 0002, Dian Jin 0003, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
Int. J. Comput. Vis.1
2025 HUPE: Heuristic Underwater Perceptual Enhancement with Semantic Collaborative Learning
Zengxi Zhang, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
Int. J. Comput. Vis.3
2025 Crossing the Chasm: A practical architecture augmentation for low-quality object detection
Xinwei Xue, Haoze Zheng, Yuechao Gao, Tengyu Ma 0004, Long Ma 0002, Qi Jia 0001
Neurocomputing5
2025 Infrared and Visible Image Fusion: From Data Compatibility to Task Adaption
abstract
Infrared-visible image fusion (IVIF) is a fundamental and critical task in the field of computer vision. Its aim is to integrate the unique characteristics of both infrared and visible spectra into a holistic representation. Since 2018, growing amount and diversity IVIF approaches step into a deep-learning era, encompassing introduced a broad spectrum of networks or loss functions for improving visual enhancement. As research deepens and practical demands grow, several intricate issues like data compatibility, perception accuracy, and efficiency cannot be ignored. Regrettably, there is a lack of recent surveys that comprehensively introduce and organize this expanding domain of knowledge. Given the current rapid development, this paper aims to fill the existing gap by providing a comprehensive survey that covers a wide array of aspects. Initially, we introduce a multi-dimensional framework to elucidate the prevalent learning-based IVIF methodologies, spanning topics from basic visual enhancement strategies to data compatibility, task adaptability, and further extensions. Subsequently, we delve into a profound analysis of these new approaches, offering a detailed lookup table to clarify their core ideas. Last but not the least, We also summarize performance comparisons quantitatively and qualitatively, covering registration, fusion and follow-up high-level tasks. Beyond delving into the technical nuances of these learning-based fusion approaches, we also explore potential future directions and open issues that warrant further exploration by the community.
Jinyuan Liu 0001, Guanyao Wu, Zhu Liu 0004, Di Wang 0018, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Risheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Learning With Self-Calibrator for Fast and Robust Low-Light Image Enhancement
abstract
Convolutional Neural Networks (CNNs) have shown significant success in the low-light image enhancement task. However, most of existing works encounter challenges in balancing quality and efficiency simultaneously. This limitation hinders practical applicability in real-world scenarios and downstream vision tasks. To overcome these obstacles, we propose a Self-Calibrated Illumination (SCI) learning scheme, introducing a new perspective to boost the model's capability. Based on a weight-sharing illumination estimation process, we construct an embedded self-calibrator to accelerate stage-level convergence, yielding gains that utilize only a single basic block for inference, which drastically diminishes computation cost. Additionally, by introducing the additivity condition on the basic block, we acquire a reinforced version dubbed SCI++, which disentangles the relationship between the self-calibrator and illumination estimator, providing a more interpretable and effective learning paradigm with faster convergence and better stability. We assess the proposed enhancers on standard benchmarks and in-the-wild datasets, confirming that they can restore clean images from diverse scenes with higher quality and efficiency. The verification on different levels of low-light vision tasks shows our applicability against other methods.
Long Ma 0002, Tengyu Ma 0004, Chengpei Xu, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 ASF-Net: Robust video deraining via temporal alignment and online adaptive learning
Xinwei Xue, Long Ma 0002, Risheng Liu
Pattern Recognit.3
2025 Striving for Faster and Better: A One-Layer Architecture With Auto Re-Parameterization for Low-Light Image Enhancement
abstract
Deep learning-based low-light image enhancers have made significant progress in recent years, with a trend towards achieving satisfactory visual quality while gradually reducing the number of parameters and improving computational efficiency. In this work, we aim to delving into the limits of image enhancers both from visual quality and computational efficiency, while striving for both better performance and faster processing. To be concrete, by rethinking the task demands, we build an explicit connection, i.e., visual quality and computational efficiency are corresponding to model learning and structure design, respectively. Around this connection, we enlarge parameter space by introducing the re-parameterization for ample model learning of a pre-defined minimalist network (e.g., just one layer), to avoid falling into a local solution. To strengthen the structural representation, we define a hierarchical search scheme for discovering a task-oriented re-parameterized structure, which also provides powerful support for efficiency. Ultimately, this achieves efficient low-light image enhancement using only a single convolutional layer, while maintaining excellent visual quality. Experimental results show our sensible superiority both in quality and efficiency against recently-proposed methods. Especially, our running time on various platforms (e.g., CPU, GPU, NPU, DSP) consistently moves beyond the existing fastest scheme. The source code will be released athttps://github.com/vis-opt-group/AR-LLIE.
Long Ma 0002, Guangchao Han, Xin Fan 0001, Risheng Liu
IEEE Trans. Circuits Syst. Video Technol.2
2024 Trash to Treasure: Low-Light Object Detection via Decomposition-and-Aggregation
abstract
Object detection in low-light scenarios has attracted much attention in the past few years. A mainstream and representative scheme introduces enhancers as the pre-processing for regular detectors. However, because of the disparity in task objectives between the enhancer and detector, this paradigm cannot shine at its best ability. In this work, we try to arouse the potential of enhancer + detector. Different from existing works, we extend the illumination-based enhancers (our newly designed or existing) as a scene decomposition module, whose removed illumination is exploited as the auxiliary in the detector for extracting detection-friendly features. A semantic aggregation module is further established for integrating multi-scale scene-related semantic information in the context space. Actually, our built scheme successfully transforms the "trash" (i.e., the ignored illumination in the detector) into the "treasure" for the detector. Plenty of experiments are conducted to reveal our superiority against other state-of-the-art methods. The code will be public if it is accepted.
Xiaohan Cui, Long Ma 0002, Tengyu Ma 0004, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
AAAI2
2024 Hybrid-Supervised Dual-Search: Leveraging Automatic Learning for Loss-Free Multi-Exposure Image Fusion
abstract
Multi-exposure image fusion (MEF) has emerged as a prominent solution to address the limitations of digital imaging in representing varied exposure levels. Despite its advancements, the field grapples with challenges, notably the reliance on manual designs for network structures and loss functions, and the constraints of utilizing simulated reference images as ground truths. Consequently, current methodologies often suffer from color distortions and exposure artifacts, further complicating the quest for authentic image representation. In addressing these challenges, this paper presents a Hybrid-Supervised Dual-Search approach for MEF, dubbed HSDS-MEF, which introduces a bi-level optimization search scheme for automatic design of both network structures and loss functions. More specifically, we harness a unique dual research mechanism rooted in a novel weighted structure refinement architecture search. Besides, a hybrid supervised contrast constraint seamlessly guides and integrates with searching process, facilitating a more adaptive and comprehensive search for optimal loss functions. We realize the state-of-the-art performance in comparison to various competitive schemes, yielding a 10.61% and 4.38% improvement in Visual Information Fidelity (VIF) for general and no-reference scenarios, respectively, while providing results with high contrast, rich details and colors. The code is available at https://github.com/RollingPlain/HSDS_MEF.
Guanyao Wu, Hongming Fu, Jinyuan Liu 0001, Long Ma 0002, Xin Fan 0001, Risheng Liu
AAAI4
2024 Contourlet Residual for Prompt Learning Enhanced Infrared Image Super-Resolution
Xingyuan Li 0005, Jinyuan Liu 0001, Zhixin Chen, Yang Zou 0004, Long Ma 0002, Xin Fan 0001, Risheng Liu
ECCV (3)5
2024 Where Elegance Meets Precision: Towards a Compact, Automatic, and Flexible Framework for Multi-modality Image Fusion and Applications
Jinyuan Liu 0001, Guanyao Wu, Zhu Liu 0004, Long Ma 0002, Risheng Liu, Xin Fan 0001
IJCAI4
2024 Seeing Text in the Dark: Algorithm and Benchmark
abstract
Localizing text in low-light environments is challenging due to visual degradations. Although a straightforward solution involves a two-stage pipeline with low-light image enhancement (LLE) as the initial step followed by detection, LLE is primarily designed for human vision rather than machine vision and can accumulate errors. In this work, we propose an efficient and effective single-stage approach for localizing text in the dark that circumvents the need for LLE. We introduce a constrained learning module as an auxiliary mechanism during the training stage of the text detector. This module is designed to guide the text detector in preserving textual spatial features amidst feature map resizing, thus minimizing the loss of spatial information in texts under low-light visual degradations. Specifically, we incorporate spatial reconstruction and spatial semantic constraints within this module to ensure the text detector acquires essential positional and contextual range knowledge. Our approach enhances the original text detector's ability to identify text's local topological features using a dynamic snake feature pyramid network and adopts a bottom-up contour shaping strategy with a novel rectangular accumulation technique for accurate delineation of streamlined text features. In addition, we present a comprehensive low-light dataset for arbitrary-shaped text, encompassing diverse scenes and languages. Notably, our method achieves state-of-the-art results on this low-light dataset and exhibits comparable performance on standard normal light datasets. The code and dataset will be released.
Chengpei Xu, Hao Fu 0004, Long Ma 0002, Wenjing Jia, Chengqi Zhang, Feng Xia 0001, Xiaoyu Ai, Binghao Li, Wenjie Zhang 0001
ACM Multimedia3
2024 Advancing Real-World Image Dehazing: Perspective, Modules, and Training
abstract
Restoring high-quality images from degraded hazy observations is a fundamental and essential task in the field of computer vision. While deep models have achieved significant success with synthetic data, their effectiveness in real-world scenarios remains uncertain. To improve adaptability in real-world environments, we construct an entirely new computational framework by making efforts from three key aspects: imaging perspective, structural modules, and training strategies. To simulate the often-overlooked multiple degradation attributes found in real-world hazy images, we develop a new hazy imaging model that encapsulates multiple degraded factors, assisting in bridging the domain gap between synthetic and real-world image spaces. In contrast to existing approaches that primarily address the inverse imaging process, we design a new dehazing network following the "localization-and-removal" pipeline. The degradation localization module aims to assist in network capture discriminative haze-related feature information, and the degradation removal module focuses on eliminating dependencies between features by learning a weighting matrix of training samples, thereby avoiding spurious correlations of extracted features in existing deep methods. We also define a new Gaussian perceptual contrastive loss to further constrain the network to update in the direction of the natural dehazing. Regarding multiple full/no-reference image quality indicators and subjective visual effects on challenging RTTS, URHI, and Fattal real hazy datasets, the proposed method has superior performance and is better than the current state-of-the-art methods.
Long Ma 0002, Xiaozhe Meng, Fan Zhou 0001, Risheng Liu, Zhuo Su 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Learning Deep Scene Curve for Fast and Robust Underwater Image Enhancement
abstract
Learning-based approaches inspired by the scattering model for enhancing underwater imagery have gained prominence. Nevertheless, these methods often suffer from time-consuming attributable to their sizable model dimensions. Moreover, they face challenges in adapting unknown scenes, primarily because the scattering model's original design was intended for atmospheric rather than marine condition. To address these obstacles, we begin by investigating the inherent differences in imaging characteristics between atmospheric and marine conditions based on statistical distributions. Building on these observations, we introduce an efficient and effective algorithm called Deep Scene Curve, abbreviated as DSC. This method comprises two essential steps: scene-irrelevant zero-mean adjustment and scene-oriented hyperparameter estimation. The first step transforms scene features into a unified zero-mean space, thereby reducing interference from scene-specific attributes. In the second step, we employ a lightweight neural network to estimate scene-oriented hyperparameters for a defined pixel-level curve based on underwater observations. This approach enables us to generate a deep curve that excels in both adaptability and efficiency, as substantiated by extensive experiments. Notably, our method achieves a significant 56% improvement in average inference time while reducing FLOPs by 92% compared to existing techniques. Furthermore, our extensive experiments in low-light image enhancement tasks highlight the potential advantages of DSC.
Xinwei Xue, Yidong Han, Long Ma 0002, Risheng Liu
IEEE Signal Process. Lett.4
2024 Bridging the Gap Between Haze Scenarios: A Unified Image Dehazing Model
abstract
In real-world scenarios, the haze presents diversity and complexity. However, current dehazing researches usually focus solely on specific categories or the removal of common white haze, frequently lacking the ability to adapt across various unknown haze types. In this study, our emphasis is on constructing a model that shows excellent adaptability across diverse haze conditions. Unlike approaches that solely rely on network structure design to enhance model adaptability, we comprehensively improve dehazing model adaptability from three key aspects: constructing the multitype haze dataset from designed haze degradation models, designing the network architecture, and formulating training strategies suitable for cross-scene generalization. Firstly, to meet the diverse haze training data requirements, we design a multitype haze degradation model to generate more realistic pairs of hazy images. Secondly, to ensure thorough haze removal and natural restoration of texture details in the recovered images, we construct a dual-branch ensemble network framework by leveraging pre-trained clear image prior features and the characteristics of 2D discrete wavelet priors. Finally, to further enhance the adaptability for removing various types of haze, we employ a sample reweighting decorrelation strategy during the network training phase to eliminate dependencies between haze and haze-free background features. Through extensive experiments, our approach shows remarkable performance across diverse haze scenarios. Our method not only outperforms state-of-the-art scene-specific dehazing methods in typical scenarios like daytime and nighttime, but it also excels in handling challenging scenarios such as dusty conditions, and color haze. See more resultshttps://github.com/fyxnl/Image-dehazing-CGID.
Zhuo Su 0001, Long Ma 0002, Xin Li 0175, Risheng Liu, Fan Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Improving Misaligned Multi-Modality Image Fusion With One-Stage Progressive Dense Registration
abstract
Misalignments between multi-modality images pose challenges in image fusion, manifesting as structural distortions and edge ghosts. Existing efforts commonly resort to registering first and fusing later, typically employing two separate stages for registration, i.e., coarse registration and fine registration. Both stages directly estimate the respective target deformation fields. This paper contends that the separate two-stage registration lacks compactness, and the direct estimation of their target deformation fields falls short in accuracy. To tackle these challenges, we introduce IMF, a framework for improving misaligned multi-modality image fusion. Central to IMF is a One-stage Progressive Dense Registration (OPDR) scheme, which accomplishes the coarse-to-fine registration through only a one-stage optimization. Specifically, two pivotal components are involved in OPDR, a dense Deformation Field Fusion (DFF) module and a Progressive Feature Fine (PFF) module. The DFF aggregates the predicted multi-scale deformation sub-fields at the current scale, while the PFF progressively refines the remaining misaligned features. Together, they effectively and accurately estimate the final deformation fields. In addition, we develop a Transformer-Conv-based Fusion (TCF) subnetwork that considers local and long-range feature dependencies, allowing us to capture more informative features from the registered infrared and visible images for the generation of high-quality fused images. Extensive experimental analysis demonstrates the superiority of the proposed method in the fusion of misaligned cross-modality images. The code will be available athttps://github.com/wdhudiekou/IMF.
Di Wang 0018, Jinyuan Liu 0001, Long Ma 0002, Risheng Liu, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and Segmentation
abstract
Multi-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, e.g., fusion or segmentation, making it hard to reach ‘Best of Both Worlds’. To overcome this issue, in this paper, we propose a Multi-interactive Feature learning architecture for image fusion and Segmentation, namely SegMiF, and exploit dual-task correlation to promote the performance of both tasks. The SegMiF is of a cascade structure, containing a fusion sub-network and a commonly used segmentation sub-network. By slickly bridging intermediate features between two components, the knowledge learned from the segmentation task can effectively assist the fusion task. Also, the benefited fusion network supports the segmentation one to perform more pretentiously. Besides, a hierarchical interactive attention block is established to ensure fine-grained mapping of all the vital information between two tasks, so that the modality/semantic features can be fully mutual-interactive. In addition, a dynamic weight factor is introduced to automatically adjust the corresponding weights of each task, which can balance the interactive feature correspondence and break through the limitation of laborious tuning. Furthermore, we construct a smart multi-wave binocular imaging system and collect a full-time multi-modality benchmark with 15 annotated pixel-level categories for image fusion and segmentation. Extensive experiments on several public datasets and our benchmark demonstrate that the proposed method outputs visually appealing fused images and perform averagely 7.66% higher segmentation mIoU in the real-world scene than the state-of-the-art approaches. The source code and benchmark are available at https://github.com/JinyuanLiu-CV/SegMiF.
Jinyuan Liu 0001, Zhu Liu 0004, Guanyao Wu, Long Ma 0002, Risheng Liu, Zhongxuan Luo, Xin Fan 0001
ICCV4
2023 Bi-level Dynamic Learning for Jointly Multi-modality Image Fusion and Beyond
abstract
Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and neglecting others, seldom investigating their underlying connections for joint promotion. To overcome these limitations, we establish the hierarchical dual tasks-driven deep model to bridge these tasks. Concretely, we firstly construct an image fusion module to fuse complementary characteristics and cascade dual task-related modules, including a discriminator for visual effects and a semantic network for feature measurement. We provide a bi-level perspective to formulate image fusion and follow-up downstream tasks. To incorporate distinct task-related responses for image fusion, we consider image fusion as a primary goal and dual modules as learnable constraints. Furthermore, we develop an efficient first-order approximation to compute corresponding gradients and present dynamic weighted aggregation to balance the gradients for fusion learning. Extensive experiments demonstrate the superiority of our method, which not only produces visually pleasant fused results but also realizes significant promotion for detection and segmentation than the state-of-the-art approaches.
Zhu Liu 0004, Jinyuan Liu 0001, Guanyao Wu, Long Ma 0002, Xin Fan 0001, Risheng Liu
IJCAI4
2023 Fearless Luminance Adaptation: A Macro-Micro-Hierarchical Transformer for Exposure Correction
abstract
Photographs taken with less-than-ideal exposure settings often display poor visual quality. Since the correction procedures vary significantly, it is difficult for a single neural network to handle all exposure problems. Moreover, the inherent limitations of convolutions, hinder the models ability to restore faithful color or details on extremely over-/under-exposed regions. To overcome these limitations, we propose a Macro-Micro-Hierarchical transformer, which consists of a macro attention to capture long-range dependencies, a micro attention to extract local features, and a hierarchical structure for coarse-to-fine correction. In specific, the complementary macro-micro attention designs enhance locality while allowing global interactions. The hierarchical structure enables the network to correct exposure errors of different scales layer by layer. Furthermore, we propose a contrast constraint and couple it seamlessly in the loss function, where the corrected image is pulled towards the positive sample and pushed away from the dynamically generated negative samples. Thus the remaining color distortion and loss of detail can be removed. We also extend our method as an image enhancer for low-light face recognition and low-light semantic segmentation. Experiments demonstrate that our approach obtains more attractive results than state-of-the-art methods quantitatively and qualitatively.
Gehui Li, Jinyuan Liu 0001, Long Ma 0002, Zhiying Jiang, Xin Fan 0001, Risheng Liu
ACM Multimedia3
2023 Bilevel Generative Learning for Low-Light Vision
abstract
Recently, there has been a growing interest in constructing deep learning schemes for Low-Light Vision (LLV). Existing techniques primarily focus on designing task-specific and data-dependent vision models on the standard RGB domain, which inherently contain latent data associations. In this study, we propose a generic low-light vision solution by introducing a generative block to convert data from the RAW to the RGB domain. This novel approach connects diverse vision problems by explicitly depicting data generation, which is the first in the field. To precisely characterize the latent correspondence between the generative procedure and the vision task, we establish a bilevel model with the parameters of the generative block defined as the upper level and the parameters of the vision task defined as the lower level. We further develop two types of learning strategies targeting different goals, namely low cost and high accuracy, to acquire a new bilevel generative learning paradigm. The generative blocks embrace a strong generalization ability in other low-light vision tasks through the bilevel optimization on enhancement tasks. Extensive experimental evaluations on three representative low-light vision tasks, namely enhancement, detection, and segmentation, fully demonstrate the superiority of our proposed approach. The code will be available at https://github.com/Yingchi1998/BGL.
Yingchi Liu, Zhu Liu 0004, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
ACM Multimedia3
2023 PAIF: Perception-Aware Infrared-Visible Image Fusion for Attack-Tolerant Semantic Segmentation
abstract
Infrared and visible image fusion is a powerful technique that combines complementary information from different modalities for downstream semantic perception tasks. Existing learning-based methods show remarkable performance, but are suffering from the inherent vulnerability of adversarial attacks, causing a significant decrease in accuracy. In this work, a perception-aware fusion framework is proposed to promote segmentation robustness in adversarial scenes. We first conduct systematic analyses about the components of image fusion, investigating the correlation with segmentation robustness under adversarial perturbations. Based on these analyses, we propose a harmonized architecture search with a decomposition-based structure to balance standard accuracy and robustness. We also propose an adaptive learning strategy to improve the parameter robustness of image fusion, which can learn effective feature extraction under diverse adversarial perturbations. Thus, the goals of image fusion (i.e., extracting complementary features from source modalities and defending attack) can be realized from the perspectives of architectural and learning strategies. Extensive experimental results demonstrate that our scheme substantially enhances the robustness, with gains of 15.3% mIOU of segmentation in the adversarial scene, compared with advanced competitors. The source codes are available at https://github.com/LiuZhu-CV/PAIF.
Zhu Liu 0004, Jinyuan Liu 0001, Benzhuang Zhang, Long Ma 0002, Xin Fan 0001, Risheng Liu
ACM Multimedia4
2023 Exploring a Distillation with Embedded Prompts for Object Detection in Adverse Environments
Hao Fu 0004, Long Ma 0002, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu
PRCV (10)2
2023 Learning With Nested Scene Modeling and Cooperative Architecture Search for Low-Light Vision
abstract
Images captured from low-light scenes often suffer from severe degradations, including low visibility, color casts, intensive noises, etc. These factors not only degrade image qualities, but also affect the performance of downstream Low-Light Vision (LLV) applications. A variety of deep networks have been proposed to enhance the visual quality of low-light images. However, they mostly rely on significant architecture engineering and often suffer from the high computational burden. More importantly, it still lacks an efficient paradigm to uniformly handle various tasks in the LLV scenarios. To partially address the above issues, we establish Retinex-inspired Unrolling with Architecture Search (RUAS), a general learning framework, that can address low-light enhancement task, and has the flexibility to handle other challenging downstream vision tasks. Specifically, we first establish a nested optimization formulation, together with an unrolling strategy, to explore underlying principles of a series of LLV tasks. Furthermore, we design a differentiable strategy to cooperatively search specific scene and task architectures for RUAS. Last but not least, we demonstrate how to apply RUAS for both low- and high-level LLV applications (e.g., enhancement, detection and segmentation). Extensive experiments verify the flexibility, effectiveness, and efficiency of RUAS.
Risheng Liu, Long Ma 0002, Tengyu Ma 0004, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Investigating intrinsic degradation factors by multi-branch aggregation for real-world underwater image enhancement
Xinwei Xue, Long Ma 0002, Qi Jia 0001, Risheng Liu, Xin Fan 0001
Pattern Recognit.3
2023 Low-Light Image Enhancement via Self-Reinforced Retinex Projection Model
abstract
Low-light image enhancement aims to improve the quality of images captured under low-lightening conditions, which is a fundamental problem in computer vision and multimedia areas. Although many efforts have been invested over the years, existing illumination-based models tend to generate unnatural-looking results (e.g., over-exposure). It is because that the widely-adopted illumination adjustment (e.g., Gamma Correction) breaks down the favorable smoothness property of the original illumination derived from the well-designed illumination estimation model. To settle this issue, a great-efficiency and high-quality Self-Reinforced Retinex Projection (SRRP) model is developed in this paper, which contains optimization modules of both illumination and reflectance layers. Specifically, we construct a new fidelity term with the self-reinforced function for the illumination optimization to eliminate the dependence of the illumination adjustment to obtain a desired illumination with the excellent smoothing property. By introducing a flexible feasible constraint, we obtain a reflectance optimization module with projection. Owing to its flexibility, we can extend our model to an enhanced version by integrating a data-driven denoising mechanism as the projection, which is able to effectively handle the generated noises/artifacts in the enhanced procedure. In the experimental part, on one side, we make ample comparative assessments on multiple benchmarks with considerable state-of-the-art methods. These evaluations fully verify the outstanding performance of our method, in terms of the qualitative and quantitative analyses and execution efficiency. On the other side, we also conduct extensive analytical experiments to indicate the effectiveness and advantages of our proposed model.
Long Ma 0002, Risheng Liu, Yiyang Wang 0001, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Multim.1
2022 Toward Fast, Flexible, and Robust Low-Light Image Enhancement
abstract
Existing low-light image enhancement techniques are mostly not only difficult to deal with both visual quality and computational efficiency but also commonly invalid in unknown complex scenarios. In this paper, we develop a new Self-Calibrated Illumination (SCI) learning framework for fast, flexible, and robust brightening images in real-world low-light scenarios. To be specific, we establish a cascaded illumination learning process with weight sharing to handle this task. Considering the computational burden of the cascaded pattern, we construct the self-calibrated module which realizes the convergence between results of each stage, producing the gains that only use the single basic block for inference (yet has not been exploited in previous works), which drastically diminishes computation cost. We then define the unsupervised training loss to elevate the model capability that can adapt general scenes. Further, we make comprehensive explorations to excavate SCI's inherent properties (lacking in existing works) including operation-insensitive adaptability (acquiring stable performance under the settings of different simple operations) and model-irrelevant generality (can be applied to illumination-based existing works to improve performance). Finally, plenty of experiments and ablation studies fully indicate our superiority in both quality and efficiency. Applications on low-light face detection and nighttime semantic segmentation fully reveal the latent practical values for SCI. The source code is available at https://github.com/vis-opt-group/SCI.
Long Ma 0002, Tengyu Ma 0004, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
CVPR1
2022 Learning to Fuse Heterogeneous Features for Low-Light Image Enhancement
abstract
To see clearly in low-light scenarios, a series of learning-based techniques have been developed to improve visual quality. However, due to the absence of semantic-level features, the existing methods are perhaps less effective on semantic-oriented visual analysis tasks (e.g., saliency detection). To break down the limitation, we propose a new classification-driven enhancement method with heterogeneous feature fusion. Specifically, we construct a new low-light image enhancement network by integrating features acquired from the pre-trained classification network. Then, to better exploit the semantic-level information, we establish a Heterogeneous Feature Fusion (HF2) operation with channel-and-spatial attention to strength the effects of cross-domain features. HF2 acts on not only the fusion between classification and encoded features but also the fusion between encoded and decoded features. Extensive experiments are conducted to indicate our superiority against other state-of-the-art methods. The application on saliency detection further reveals our effectiveness in settling the semantic-oriented visual tasks.
Zhenyu Tang 0004, Long Ma 0002, Xiaoke Shang, Xin Fan 0001
ICASSP2
2022 Semantic-aware Texture-Structure Feature Collaboration for Underwater Image Enhancement
abstract
Underwater image enhancement has become an attractive topic as a significant technology in marine engi-neering and aquatic robotics. However, the limited number of datasets and imperfect hand-crafted ground truth weaken its robustness to unseen scenarios, and hamper the application to high-level vision tasks. To address the above limitations, we develop an efficient and compact enhancement network in collaboration with a high-level semantic-aware pretrained model, aiming to exploit its hierarchical feature representation as an auxiliary for the low-level underwater image enhance-ment. Specifically, we tend to characterize the shallow layer features as textures while the deep layer features as structures in the semantic-aware model, and propose a multi-path Contextual Feature Refinement Module (CFRM) to refine features in multiple scales and model the correlation between different features. In addition, a feature dominative network is devised to perform channel-wise modulation on the aggregated texture and structure features for the adaptation to different feature patterns of the enhancement network. Extensive experiments on benchmarks demonstrate that the proposed algorithm achieves more appealing results and outperforms state-of-the-art meth-ods by large margins. We also apply the proposed algorithm to the underwater salient object detection task to reveal the favorable semantic-aware ability for high-level vision tasks.
Di Wang 0018, Long Ma 0002, Risheng Liu, Xin Fan 0001
ICRA2
2022 Hierarchical Bilevel Learning with Architecture and Loss Search for Hadamard-based Image Restoration
abstract
In the past few decades, Hadamard-based image restoration problems (e.g., low-light image enhancement) attract wide concerns in multiple areas related to artificial intelligence. However, existing works mostly focus on heuristically defining architecture and loss by the engineering experiences that came from extensive practices. This way brings about expensive verification costs for seeking out the optimal solution. To this end, we develop a novel hierarchical bilevel learning scheme to discover the architecture and loss simultaneously for different Hadamard-based image restoration tasks. More concretely, we first establish a new Hadamard-inspired neural unit to aggregate domain knowledge into the network design. Then we model a triple-level optimization that consists of the architecture, loss and parameters optimizations to deliver a macro perspective for network learning. Then we introduce a new hierarchical bilevel learning scheme for solving the built triple-level model to progressively generate the desired architecture and loss. We also define an architecture search space consisting of a series of simple operations and an image quality-oriented loss search space. Extensive experiments on three Hadamard-based image restoration tasks (including low-light image enhancement, single image haze removal and underwater image enhancement) fully verify our superiority against state-of-the-art methods.
Guijing Zhu, Long Ma 0002, Xin Fan 0001, Risheng Liu
IJCAI2
2022 PIA: Parallel Architecture with Illumination Allocator for Joint Enhancement and Detection in Low-Light
abstract
Visual perception in low-light conditions (e.g., nighttime) plays an important role in various multimedia-related applications (e.g., autonomous driving). The enhancement (provides a visual-friendly appearance) and detection (detects the instances of objects) in low-light are two fundamental and crucial visual perception tasks. In this paper, we make efforts on how to simultaneously realize low-light enhancement and detection from two aspects. First, we define a parallel architecture to satisfy the task demand for both two tasks. In which, a decomposition-type warm-start acting on the entrance of parallel architecture is developed to narrow down the adverse effects brought by low-light scenes to some extent. Second, a novel illumination allocator is designed by encoding the key illumination component (the inherent difference between normal-light and low-light) to extract hierarchical features for assisting in enhancement and detection. Further, we make a substantive discussion for our proposed method. That is, we solve enhancement in a coarse-to-fine manner and handle detection in a decomposed-to-integrated fashion. Finally, multidimensional analytical and evaluated experiments are performed to indicate our effectiveness and superiority. The code is available at \urlhttps://github.com/tengyu1998/PIA
Tengyu Ma 0004, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
ACM Multimedia2
2022 Best of Both Worlds: See and Understand Clearly in the Dark
abstract
Recently, with the development of intelligent technology, the perception of low-light scenes has been gaining widespread attention. However, existing techniques usually focus on only one task (e.g., enhancement) and lose sight of the others (e.g., detection), making it difficult to perform all of them well at the same time. To overcome this limitation, we propose a new method that can handle visual quality enhancement and semantic-related tasks (e.g., detection, segmentation) simultaneously in a unified framework. Specifically, we build a cascaded architecture to meet the task requirements. To better enhance the entanglement in both tasks and achieve mutual guidance, we develop a new contrastive-alternative learning strategy for learning the model parameters, to largely improve the representational capacity of the cascaded architecture. Notably, the contrastive learning mechanism establishes the communication between two objective tasks in essence, which actually extends the capability of contrastive learning to some extent. Finally, extensive experiments are performed to fully validate the advantages of our method over other state-of-the-art works in enhancement, detection, and segmentation. A series of analytical evaluations are also conducted to reveal our effectiveness. The code is available at https://github.com/k914/contrastive-alternative-learning.
Xinwei Xue, Long Ma 0002, Yi Wang 0037, Xin Fan 0001, Risheng Liu
ACM Multimedia3
2022 Video Deraining via Temporal Discrepancy Learning
Yirui Fan, Long Ma 0002, Risheng Liu
PRCV (4)2
2022 Task-Oriented Convex Bilevel Optimization With Latent Feasibility
abstract
This paper firstly proposes a convex bilevel optimization paradigm to formulate and optimize popular learning and vision problems in real-world scenarios. Different from conventional approaches, which directly design their iteration schemes based on given problem formulation, we introduce a task-oriented energy as our latent constraint which integrates richer task information. By explicitly re- characterizing the feasibility, we establish an efficient and flexible algorithmic framework to tackle convex models with both shrunken solution space and powerful auxiliary (based on domain knowledge and data distribution of the task). In theory, we present the convergence analysis of our latent feasibility re- characterization based numerical strategy. We also analyze the stability of the theoretical convergence under computational error perturbation. Extensive numerical experiments are conducted to verify our theoretical findings and evaluate the practical performance of our method on different applications.
Risheng Liu, Long Ma 0002, Xiaoming Yuan 0001, Shangzhi Zeng, Jin Zhang 0002
IEEE Trans. Image Process.2
2022 Underexposed Image Correction via Hybrid Priors Navigated Deep Propagation
abstract
Enhancing visual quality for underexposed images is an extensively concerning task that plays an important role in various areas of multimedia and computer vision. Most existing methods often fail to generate high-quality results with appropriate luminance and abundant details. To address these issues, we develop a novel framework, integrating both knowledge from physical principles and implicit distributions from data to address underexposed image correction. More concretely, we propose a new perspective to formulate this task as an energy-inspired model with advanced hybrid priors. A propagation procedure navigated by the hybrid priors is well designed for simultaneously propagating the reflectance and illumination toward desired results. We conduct extensive experiments to verify the necessity of integrating both underlying principles (i.e., with knowledge) and distributions (i.e., from data) as navigated deep propagation. Plenty of experimental results of underexposed image correction demonstrate that our proposed method performs favorably against the state-of-the-art methods on both subjective and objective assessments. In addition, we execute the task of face detection to further verify the naturalness and practical value of underexposed image correction. What is more, we apply our method to solve single-image haze removal whose experimental results further demonstrate our superiorities.
Risheng Liu, Long Ma 0002, Yuxi Zhang 0001, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.2
2022 Learning Deep Context-Sensitive Decomposition for Low-Light Image Enhancement
abstract
Enhancing the quality of low-light (LOL) images plays a very important role in many image processing and multimedia applications. In recent years, a variety of deep learning techniques have been developed to address this challenging task. A typical framework is to simultaneously estimate the illumination and reflectance, but they disregard the scene-level contextual information encapsulated in feature spaces, causing many unfavorable outcomes, e.g., details loss, color unsaturation, and artifacts. To address these issues, we develop a new context-sensitive decomposition network (CSDNet) architecture to exploit the scene-level contextual dependencies on spatial scales. More concretely, we build a two-stream estimation mechanism including reflectance and illumination estimation network. We design a novel context-sensitive decomposition connection to bridge the two-stream mechanism by incorporating the physical principle. The spatially varying illumination guidance is further constructed for achieving the edge-aware smoothness property of the illumination component. According to different training patterns, we construct CSDNet (paired supervision) and context-sensitive decomposition generative adversarial network (CSDGAN) (unpaired supervision) to fully evaluate our designed architecture. We test our method on seven testing benchmarks [including massachusetts institute of technology (MIT)-Adobe FiveK, LOL, ExDark, and naturalness preserved enhancement (NPE)] to conduct plenty of analytical and evaluated experiments. Thanks to our designed context-sensitive decomposition connection, we successfully realized excellent enhanced results (with sufficient details, vivid colors, and few noises), which fully indicates our superiority against existing state-of-the-art approaches. Finally, considering the practical needs for high efficiency, we develop a lightweight CSDNet (named LiteCSDNet) by reducing the number of channels. Furthermore, by sharing an encoder for these two components, we obtain a more lightweight version (SLiteCSDNet for short). SLiteCSDNet just contains 0.0301M parameters but achieves the almost same performance as CSDNet. Code is available at https://github.com/KarelZhang/CSDNet-CSDGAN.
Long Ma 0002, Risheng Liu, Jiaao Zhang, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Neural Networks Learn. Syst.1
2021 Physics-inspired Learning for Structure-Aware Texture-Sensitive Underwater Image Enhancement
abstract
Recently, improving the visual quality of underwater images using deep learning-based methods has drawn considerable attention. Unfortunately, diverse environmental factors (e.g., blue/green color distortion) severely limit their performance in real-world environments. Therefore, strengthening the superiority of the underwater image enhancement method is critical. In this paper, we devote ourselves to develop a new architecture with strong superiority and adaptability. Inspired by the underwater imaging principle, we establish a novel physics-inspired learning model that is easy to realize. A Structure-Aware Texture-Sensitive Network (SATS-Net) is further developed to portray the model. The structure-aware module is responsible for structural information, and the texture-sensitive module is responsible for textural information. Thus, SATS-Net successfully incorporates robust characterization absorbed from the physical principle to achieve strong robustness and adaptability. We conduct extensive experiments to demonstrate that SATS-Net outperforms existing advanced techniques in various real-world underwater environments.
Xinwei Xue, Long Ma 0002, Risheng Liu, Xin Fan 0001
ACML3
2021 Retinex-Inspired Unrolling With Cooperative Prior Architecture Search for Low-Light Image Enhancement
abstract
Low-light image enhancement plays very important roles in low-level vision areas. Recent works have built a great deal of deep learning models to address this task. However, these approaches mostly rely on significant architecture engineering and suffer from high computational burden. In this paper, we propose a new method, named Retinex-inspired Unrolling with Architecture Search (RUAS), to construct lightweight yet effective enhancement network for low-light images in real-world scenario. Specifically, building upon Retinex rule, RUAS first establishes models to characterize the intrinsic underexposed structure of low-light images and unroll their optimization processes to construct our holistic propagation structure. Then by designing a cooperative reference-free learning strategy to discover low-light prior architectures from a compact search space, RUAS is able to obtain a top-performing image enhancement network, which is with fast speed and requires few computational resources. Extensive experiments verify the superiority of our RUAS framework against recently proposed state-of-the-art methods. The project page is available at http://dutmedia.org/RUAS/.
Risheng Liu, Long Ma 0002, Jiaao Zhang, Xin Fan 0001, Zhongxuan Luo
CVPR2
2021 NASA: A Noise-Adaptive and Structure-Aware Learning Framework for Image Deblurring
abstract
Image deblurring is a classical low-level visual processing task, which aims to recover a potentially noise-free sharp image from the blurred image. Existing prior-based and learning-based methods usually need to manually set some vital auxiliary components (e.g., noise level). It brings about extremely weak adaptability and flexibility. To settle this issue, we develop a Noise-Adaptive Structure-Aware learning framework (NASA) to achieve fully intelligent manufacturing. Concretely, by introducing a new task-assisted module, we define a novel robust image deblurring model derived from a MAP-based energy function. Consequently, we establish the NASA which consists of three basic modules including the task-assisted, fidelity-term, and regularization-term modules, to solve our designed model. The task-assisted module generates the noise-adaptive and structure-aware maps, which are fed to the other two modules. By end-to-end training our NASA, we successfully avoid the cumbersome manually parameters-adjustment process. Quantitative and qualitative experiments demonstrate our superiority compared to the state-of-the-art methods, both in visual effect and numerical scores. A series of ablation study also verify the effectiveness and necessity of our designed mechanism.
Xiaokun Liu, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICASSP2
2021 Temporal Rain Decomposition with Spatial Structure Guidance for Video Deraining
abstract
Recently, removing rain streaks from videos has drawn wide concerns in vision and multimedia communities. But existing works ignore the depicts of image inherent structure and rain location to cause details loss, and their adopted manners of exploiting temporal information are still insufficient. In this work, we propose a multi-frame deraining network with temporal rain decomposition and spatial structure guidance to more effectively accomplish video deraining. A learnable decomposition method is defined to learn the distribution of rain, where the location map acts on a single-frame deraining block. We construct a multi-frame fusion module with a detailed guidance map to integrate temporal and spatial information. Many evaluated experiments demonstrate that our algorithm performs favorably on video deraining tasks compared with other methods. The elaborate ablation study in terms of network architecture fully indicates the effectiveness of our network.
Xinwei Xue, Ying Ding 0006, Long Ma 0002, Yi Wang 0037, Risheng Liu, Xin Fan 0001
ICASSP3
2021 GTA-Net: Gradual Temporal Aggregation Network for Fast Video Deraining
abstract
Recently, the development of intelligent technology arouses the requirements of high-quality videos. Rain streak is a frequent and inevitable factor to degrade the video. Many researchers have put their energies into eliminating the adverse effects of rainy video. Unfortunately, how to fully utilize the temporal information from rainy video is still in suspense. In this work, to effectively exploit temporal information, we develop a simple but effective network, Gradual Temporal Aggregation Network (GTA-Net for short). To be specific, according to the temporal distance between rainy frames and the reference frame, we divide the rainy frames into different groups. A multi-stream coarse temporal aggregation module is first performed to aggregate different temporal information with equal status and importance. Then we design a single-stream fine temporal aggregation module to further fuse the integrated frames that maintain the different distances with the target frame. In this way of coarse-to-fine, we not only achieve superior performance, but also gain the surprising execution speed owing to abandon the time-consuming alignment operation. Plenty of experimental results demonstrate that our GTA-Net performs favorably compared to other state-of-the-art approaches. The meticulous ablation study further indicates the effectiveness of our designed GTA-Net.
Xinwei Xue, Long Ma 0002, Risheng Liu, Xin Fan 0001
ICASSP3
2021 Video Deraining Via Temporal Aggregation-and-Guidance
abstract
Learning-based video deraining methods generally integrate temporal correlation within the network. But their non-transparency (i.e., difficult to comprehend how to exploit temporal correlation) seriously limits the development for video deraining. To conquer it, this paper proposes a novel Temporal Aggregation-and-Guidance Network (TAG-Net). Concretely, we define a new temporal ensemble model with set representation by modeling correspondence between rain regions of the current frame and rain-free regions of adjacent frames. Further, we build a TAG-Net that contains: 1) temporal aggregation network derived from the ensemble model, which is with newly-designed self-directed attention acting on video sequences, it automatically learns temporal correlation from multiple adjacent frames to optimize the current frame. 2) temporal guidance network, which aims at eliminating rain streaks in intersected rain regions between the current and adjacent frames to enhance the previously-recovered frame. Extensive evaluations verify that TAG-Net yields the best performance against other advanced methods.
Long Ma 0002, Risheng Liu, Xin Fan 0001
ICME1
2021 Spatial-Temporal Integration Network with Self-Guidance for Robust Video Deraining
abstract
Recently, video deraining has become a research focus. Network-based approaches are continuously showing extrusive performance. However, they lack precise control over the motion consistency in temporal information and characterize spatial distribution, so that their results are unsatisfying, especially in some real-world scenarios. To settle them, we develop a spatial-temporal integration network with self-guidance. It contains flow-induced alignment, self-guidance generation, and spatial-temporal integration modules. The alignment module not only preliminarily removes rain to provide more effective temporal correlation but also accurately keeps motion consistency between frames. The self-guidance map characterizes the pixel-level spatial distribution for the target to avoid injuring the background. Finally, we concatenate adjacent aligned frames, self-guidance map, and original current rain frame into the integration module to progressively fuse them in a coarse-to-fine way. Extensive evaluations demonstrate our superiority against other state-of-the-art methods qualitatively and quantitatively.
Xiaokun Liu, Risheng Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICME3
2021 Searching Frame-Recurrent Attentive Deformable Network for Real-Time Video Deraining
abstract
Video deraining has become an issue of great interest since rain streaks inevitably affect video quality. Most of the existing works focus on heuristically designing the network architecture to integrate available information derived from the temporal dimension. However, their inferences take a long time, so that the practicability is somewhat ignored. To solve this problem, we develop a real-time video deraining network in a frame-recurrent manner. It includes a fast attentive deformable alignment module and an automatically-discovered spatial-temporal reconstruction module. In which, the alignment is composed of a single newly-built deformable convolution under the channel attention mechanism to keep the accurate motion consistency and reduce time-consuming by a wide margin. The reconstruction part for the first time introduces the architecture search technique for video deraining to automatically discover a high-effective architecture by designing an effective and compact search space. Experimental results demonstrate remarkable superiority both in computational efficiency and actual performance compared to other state-of-the-art approaches.
Xinwei Xue, Long Ma 0002, Yi Wang 0037, Risheng Liu, Xin Fan 0001
ICME3
2021 Star-Net: Spatial-Temporal Attention Residual Network for Video Deraining
abstract
Learning-based video deraining has recently drawn increasing attention. They tend to directly package aligned frames to input a fully end-to-end network. However, the network is generally object-driven and cannot recognize how to utilize temporal information so that the results are unsatisfied. In this work, we design a novel Spatial-Temporal Attention Network (STAR-Net) to explicitly utilize the temporal information. Concretely, we define the self-spatial attention to characterizing the rain region of the target frame, and the temporal-spatial attention to learn the profitable information for remedying the rain region of the target frame from the adjacent frame. We also introduce a simple residual network to further strengthen the relationship between the target and the adjacent frame. These addressed frames are fused by a three-layers convolutional module to further improve the capability. Extensive evaluations indicate our superiority against state-of-the-art methods.
Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME3
2021 Collaborative Reflectance-And-Illumination Learning For High-Efficient Low-Light Image Enhancement
abstract
In this paper, we settle the low-light image enhancement problem by developing a collaborative learning framework, which not only improves lightness and suppresses noises simultaneously but also with fast speed and requires few computational resources. The approach is inspired by the fact that reflectance and illumination are highly correlated to satisfy the well-known Retinex decomposition principle. With this in mind, we establish a Reflectance-and-Illumination Collaborative (RIC) block to depict the compact physical relationship between reflectance and illumination. By cascading multiple RIC blocks, we obtain an end-to-end RICNet to interactively optimize these two components in a collaborative manner. Benefiting from the RIC block that integrates powerful task cues, RICNet just needs few parameters to simultaneously improve brightness and remove noises. Extensive experiments demonstrate our superiority against existing state-of-the-art methods. We also make meticulous analysis for the RIC block. The results reveal the rationality and effectiveness of our built mechanism.
Guijing Zhu, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME2
2021 Bridging the Gap between Low-Light Scenes: Bilevel Learning for Fast Adaptation
abstract
Brightening low-light images of diverse scenes is a challenging but widely concerned task in the multimedia community. Convolutional Neural Networks (CNNs) based approaches mostly acquire the enhanced model by learning the data distribution from the specific scenes. However, these works present poor adaptability (even fail) when meeting real-world scenarios that never encountered before. To conquer it, we develop a novel bilevel learning scheme for fast adaptation to bridge the gap between low-light scenes. Concretely, we construct a Retinex-induced encoder-decoder with an adaptive denoising mechanism, aiming at covering more practical cases. Different from existing works that directly learn model parameters by using the massive data, we provide a new hyperparameter optimization perspective to formulate a bilevel learning scheme towards general low-light scenarios. This scheme depicts the latent correspondence (i.e., scene-irrelevant encoder) and the respective characteristic (i.e., scene-specific decoder) among different data distributions. Due to the expensive inner optimization, estimating the hyper-parameter gradient exactly can be prohibitive, we develop an approximate hyper-parameter gradient method by introducing the one-step forward approximation and finite difference approximation to ensure the high-efficient inference. Extensive experiments are conducted to reveal our superiority against other state-of-the-art methods. A series of analytical experiments are also executed to verify our effectiveness.
Dian Jin 0003, Long Ma 0002, Risheng Liu, Xin Fan 0001
ACM Multimedia2
2021 Latency-Constrained Spatial-Temporal Aggregated Architecture Search for Video Deraining
Zhu Liu 0004, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo, Yuduo Zhang
PRCV (3)2
2021 Semantic-Driven Context Aggregation Network for Underwater Image Enhancement
Dongxiang Shi, Long Ma 0002, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
PRCV (3)2
2021 Deep Multi-Illumination Fusion for Low-Light Image Enhancement
Long Ma 0002, Risheng Liu, Xin Fan 0001
PRCV (3)3
2021 A convergent framework with learnable feasibility for Hadamard-based image recovery
Yiyang Wang 0001, Long Ma 0002, Risheng Liu
Comput. Vis. Image Underst.2
2021 Learning to Discover a Unified Architecture for Low-Level Vision
abstract
Neural Architecture Search (NAS) has pioneered various constructive principles to push forward the development of deep learning and achieved dramatic performances for diverse tasks recently. Existing NAS methods mainly focus on a single specific task to discover the architecture automatically. But actually, these methods lack ample exploitation and exploration for the latent ability of architecture search mechanism, e.g., from diverse cross-task distributions to discover a unified architecture automatically. In this work, we propose a Cross-task Differentiable ARchiTecture Search (Cross-DARTS for short) framework to discover a unified architecture for different low-level vision tasks automatically, to further widen the capacity of NAS. Specifically, we establish a new model to bridge different low-level vision tasks under the architecture search perspective. By performing a new data construction that integrates multi-task distributions, Cross-DARTS is obtained based on the differentiable search scheme. A multi-scale fusion cell with powerful contextual representation capacity is designed as the basic component of search space towards the low-level vision. Consistent achievements of promising results on three vision tasks, including noise, rain, joint rain and haze removal fully show our superiority.
Zhu Liu 0004, Long Ma 0002, Risheng Liu, Xin Fan 0001
IEEE Signal Process. Lett.2
2021 Joint Luminance and Chrominance Learning for Underwater Image Enhancement
abstract
Recently, learning-based works have been widely-investigated to enhance underwater images. However, interactions between various degradation factors (e.g., color distortion and haze effects) inevitably cause negative interference during the inference phase. Thus, these works cannot fully remove degraded factors. To address this problem, we propose a novel Joint Luminance and Chrominance Learning Network (JLCL-Net). Concretely, we reformulate the task as luminance reconstruction (for haze removal), and chrominance correction (for color correction) sub-tasks by separating the luminance and chrominance (i.e., color appearance) of the underwater images. In this way, we successfully realize the disentanglement in degraded factors to avoid introducing interference. We specify the reconstruction by integrating the atmospheric scattering model, which endows the adaptive dehazing ability over different scenarios. The correction learns to compensate for color by a simple network to reverse the color attenuation process. To this end, we obtain our JLCL-Net. To better train it, we design a new multi-stage cross-space training strategy, which progressively updates the network parameters to enlarge the network potentiality. Extensive evaluations are presented to fully verify our superiority against other methods.
Xinwei Xue, Zhenhua Hao, Long Ma 0002, Yi Wang 0037, Risheng Liu
IEEE Signal Process. Lett.3
2021 Learning Hadamard-Product-Propagation for Image Dehazing and Beyond
abstract
Image dehazing has evolved into an attractive research field in the computer vision community in the past few decades. Previous traditional approaches attempt to design energy-based objective functions. However, they cannot accurately express the intrinsic characteristics of the images, posing weak adaptation ability for real-world complex scenarios. More recently, deep learning techniques for image dehazing have matured and become more reliable, showing outstanding performance. Nevertheless, these methods heavily depend on training data, restricting their application ranges. More importantly, both traditional and deep learning approaches all ignore a common issue, noises/artifacts always appear in the recovery process. To this end, a new Hadamard-Product (HP) model is proposed, which consists of a series of data-driven priors. Based on this model, we derive a Learnable Hadamard-Product-Propagation (LHPP) by cascading a series of principle-inspired guidance and recovery modules. In which, the principle-inspired guidance related to transmission is endowed the smoothness property, the other recovery module satisfies the distribution of natural images. The Hadamard-product-based propagations is generated in our developed learnable framework for the task of image dehazing. In this way, we can eliminate noises/artifacts in the recovery procedure to obtain the ideal outputs. Subsequently, since the generality of our HP model, we successfully extend our LHPP to settle low-light image enhancement and underwater image enhancement problems. A series of analytical experiments are performed to verify our effectiveness. Plenty of performance evaluations on three complex tasks fully reveal our superiority against multiple state-of-the-art methods.
Risheng Liu, Jinyuan Liu 0001, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.4
2020 Sequential Deep Unrolling With Flow Priors For Robust Video Deraining
abstract
Video deraining has attracted wide attention since the urgent demand of high-quality video in recent years. The indistinct details and nonideal deraining effects are the most common defects in existing techniques, whose cause lies in the insufficient usage of single-frame image and temporal information. To effectively settle video deraining, we establish a new deraining model with flow priors to simultaneously introduce spatial and temporal information for accurately depicting the enhancement model of the current frame. A sequential deep unrolling framework is substantially presented by solving this model based on optimization techniques. The ablation study indicates our effectiveness as far as the design of architecture. Plenty of subjective and objective evaluations fully demonstrate our superiority in detail recovery and deraining effects against other state-of-the-are video deraining approaches.
Xinwei Xue, Ying Ding 0006, Pan Mu, Long Ma 0002, Risheng Liu, Xin Fan 0001
ICASSP4
2020 Principle-Inspired Multi-Scale Aggregation Network for Extremely Low-Light Image Enhancement
abstract
The under-exposure and low-light environments are common to degrade the image-quality with invisible information. To ameliorate this case, a copious of low-light image enhancement methods are developed. However, these existing works are hard to handle extremely low-light conditions with noises, even well-known network-based methods. To address this issue, we develop a Principle-inspired Multi-scale Aggregation Network (PMA-Net) to simultaneously achieve the exposure enhancement and noises removal. Specifically, we establish a pioneering principle-inspired connection to present the physical principle in the inside of the network, to strengthen the structural depict. Subsequently, we propose a multi-scale aggregation strategy to eliminate the noises in the enhanced results. Sufficient ablation studies manifest the effectiveness of our PMA-Net. Extensive qualitative and quantitative comparisons with other state-of-the-art methods are conducted to fully indicates our outstanding performance.
Jiaao Zhang, Risheng Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICASSP3
2020 Learning Multi-scale Retinex with Residual Network for Low-Light Image Enhancement
Long Ma 0002, Jingjie Shang, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
PRCV (1)1
2020 Joint Over and Under Exposures Correction by Aggregated Retinex Propagation for Image Enhancement
abstract
Since the interference of ambient light and the limitation of physical devices, it is quite a common phenomenon that images taken in real-world scenarios turn out to be incorrectly exposed. Most existing techniques emphasize underexposed image correction. On one hand, these works ignore the correction of over-exposure regions in the original input. On the other hand, it is likely to generate over-exposure images. To mitigate these issues, we have developed a novel aggregated Retinex propagations to simultaneously correct over and under-exposure correction of a single image. Concretely, we first manifest the necessity of concurrently correcting under and over-exposure appearances. We establish a Retinex image propagation framework with shared weights to correct different levels of exposure. Then by introducing the fusion computational module, we achieve the accurate exposure correction for a single image. Plenty of quantitative and qualitative comparisons are conducted to fully indicate our superiority against other state-of-the-art algorithms. The elaborated algorithmic analyses show our effectiveness. Experiments on face detection further verify our practicability.
Long Ma 0002, Dian Jin 0003, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
IEEE Signal Process. Lett.1
2019 Task Embedded Coordinate Update: A Realizable Framework for Multivariate Non-Convex Optimization
abstract
We in this paper propose a realizable framework TECU, which embeds task-specific strategies into update schemes of coordinate descent, for optimizing multivariate non-convex problems with coupled objective functions. On one hand, TECU is capable of improving algorithm efficiencies through embedding productive numerical algorithms, for optimizing univariate sub-problems with nice properties. From the other side, it also augments probabilities to receive desired results, by embedding advanced techniques in optimizations of realistic tasks. Integrating both numerical algorithms and advanced techniques together, TECU is proposed in a unified framework for solving a class of non-convex problems. Although the task embedded strategies bring inaccuracies in sub-problem optimizations, we provide a realizable criterion to control the errors, meanwhile, to ensure robust performances with rigid theoretical analyses. By respectively embedding ADMM and a residual-type CNN in our algorithm framework, the experimental results verify both efficiency and effectiveness of embedding task-oriented strategies in coordinate descent for solving practical problems.
Yiyang Wang 0001, Risheng Liu, Long Ma 0002, Xiaoliang Song
AAAI3
2019 Enhanced Residual Dense Intrinsic Network for Intrinsic Image Decomposition
abstract
Intrinsic image decomposition is a challenging task, which aims at recovering intrinsic components from the observation. Hand-crafted priors have been widely used in traditional methods, yet with unsatisfactory performance of quality and runtime. Recently, network-based approaches have been greatly developed, but the physical imaging principle is ignored causing the multiplication of estimated components is hard to reconstruct the observation. To overcome these limitations, we develop an enhanced residual dense intrinsic network (ERDIN) for intrinsic decomposition. Specifically, we construct the basic module (i.e., enhanced residual dense block (ERDB)) to fully exploit the hierarchical features. The physical imaging principle is designed as the reconstruction loss to ensure the consistency between the observation and the multiplication of estimated components, which is of equal importance with the data loss. Extensive experimental results illustrate our excellent performance compared with other state-of-the-art methods.
Risheng Liu, Long Ma 0002, Miao Zhang 0004, Xin Fan 0001, Zhongxuan Luo
ICME3
2019 Deep Proximal Unrolling: Algorithmic Framework, Convergence Analysis and Applications
abstract
Deep learning models have gained great success in many real-world applications. However, most existing networks are typically designed in heuristic manners, thus these approaches lack of rigorous mathematical derivations and clear interpretations. Several recent studies try to build deep models by unrolling a particular optimization model that involves task information. Unfortunately, due to the dynamic nature of network parameters, their resultant deep propagations do not possess the nice convergence property as the original optimization scheme does. In this work, we develop a generic paradigm to unroll nonconvex optimization for deep model design. Different from most existing frameworks, which just replace the iterations by network architectures, we prove in theory that the propagation generated by our proximally unrolled deep model can globally converge to the critical-point of the original optimization model. Moreover, even if the task information is only partially available (e.g., no prior regularization), we can still train a convergent deep propagations. We also extend these theoretical investigations on the more general multi-block models and thus a lot of real-world applications can be successfully handled by the proposed framework. Finally, we conduct experiments on various low-level vision tasks (i.e., non-blind deconvolution, dehazing, and low-light image enhancement) and demonstrate the superiority of our proposed framework, compared with existing state-of-the-art approaches.
Risheng Liu, Shichao Cheng, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.3
2019 Learning Converged Propagations With Deep Prior Ensemble for Image Enhancement
abstract
Enhancing visual qualities of images plays very important roles in various vision and learning applications. In the past few years, both knowledge-driven maximum a posterior (MAP) with prior modelings and fully data-dependent convolutional neural network (CNN) techniques have been investigated to address specific enhancement tasks. In this paper, by exploiting the advantages of these two types of mechanisms within a complementary propagation perspective, we propose a unified framework, named deep prior ensemble (DPE), for solving various image enhancement tasks. Specifically, we first establish the basic propagation scheme based on the fundamental image modeling cues and then introduce residual CNNs to help predicting the propagation direction at each stage. By designing prior projections to perform feedback control, we theoretically prove that even with experience-inspired CNNs, DPE is definitely converged and the output will always satisfy our fundamental task constraints. The main advantage against conventional optimization-based MAP approaches is that our descent directions are learned from collected training data, thus are much more robust to unwanted local minimums. While, compared with existing CNN type networks, which are often designed in heuristic manners without theoretical guarantees, DPE is able to gain advantages from rich task cues investigated on the bases of domain knowledges. Therefore, DPE actually provides a generic ensemble methodology to integrate both knowledge and data-based cues for different image enhancement tasks. More importantly, our theoretical investigations verify that the feedforward propagations of DPE are properly controlled toward our desired solution. Experimental results demonstrate that the proposed DPE outperforms state-of-the-arts on a variety of image enhancement tasks in terms of both quantitative measure and visual perception quality.
Risheng Liu, Long Ma 0002, Yiyang Wang 0001, Lei Zhang 0006
IEEE Trans. Image Process.2
2018 Deep Layer Prior Optimization for Single Image Rain Streaks Removal
abstract
Visible distortions caused by rain streaks have significant negative effects on the performance of many vision and learning algorithms. Most of the existing deraining approaches propose to build complex prior models to formulate the appearance of rain streaks. Unfortunately, these human-designed priors tend to over-smooth the background and leave too many rain streaks since the distribution of rain streaks is complex and disordered. In this work, we exploit a deep layer prior under the maximum a posterior framework to recover the intrinsic rain structure. The optimization of the resulted variational energy can be understood as simultaneously performing rain and image propagations based on data-dependent residual networks and task cues (e.g., total variation regularization), respectively. Experimental results on both synthetic and real test images demonstrate the effectiveness of our approach against both designed priors and fully data-dependent convolutional neural networks.
Risheng Liu, Zhiying Jiang, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICASSP3
2018 Robust Haze Removal Via Joint Deep Transmission and Scene Propagation
abstract
Haze is one of the most important factors which reduce the outdoor image quality. Existing approaches often aim to design their models based on principles of hazes. However, even with exactly modeled haze distribution, it is still a challenging task due to factors in real scenario, such as noises, halos and artifacts. To address limitations of existing approaches for real-world hazy removal problem, this paper proposes a novel framework to incorporate deep residual architectures into a propagation scheme to jointly estimate transmission and clean scene. We evaluate the proposed framework on both widely used benchmarks and real-world low-quality hazy images. Extensive experimental results demonstrate that our method performs favorably against approaches designed only based on haze cues and achieves the state-of-the-art results, compared with both conventional shallow models and deep dehzaing networks.
Risheng Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
ICASSP3
2018 A Bridging Framework for Model Optimization and Deep Propagation
abstract
Optimizing task-related mathematical model is one of the most fundamental methodologies in statistic and learning areas. However, generally designed schematic iterations may hard to investigate complex data distributions in real-world applications. Recently, training deep propagations (i.e., networks) has gained promising performance in some particular tasks. Unfortunately, existing networks are often built in heuristic manners, thus lack of principled interpretations and solid theoretical supports. In this work, we provide a new paradigm, named Propagation and Optimization based Deep Model (PODM), to bridge the gaps between these different mechanisms (i.e., model optimization and deep propagation). On the one hand, we utilize PODM as a deeply trained solver for model optimization. Different from these existing network based iterations, which often lack theoretical investigations, we provide strict convergence analysis for PODM in the challenging nonconvex and nonsmooth scenarios. On the other hand, by relaxing the model constraints and performing end-to-end training, we also develop a PODM based strategy to integrate domain knowledge (formulated as models) and real data distributions (learned by networks), resulting in a generic ensemble framework for challenging real-world applications. Extensive experiments verify our theoretical results and demonstrate the superiority of PODM against these state-of-the-art approaches.
Risheng Liu, Shichao Cheng, Xiaokun Liu, Long Ma 0002, Xin Fan 0001, Zhongxuan Luo
NeurIPS4