Pan Mu

dblp:234/5608 · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-5330-2673ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 9 first-author · 18 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Enhancing the Interpretation of Skin Lesion Diagnosis: Concept Adaptive Fine-Tuning of Vision-Language Models
abstract
Significant progress has been made in applying deep learning for the automatic diagnosis of skin lesions. However, most models remain unexplainable, which severely hinders their application in clinical settings. Concept-based ante-hoc interpretable models have the potential to clarify the decision-making process of diagnosis by learning high-level, human-understandable concepts, while they can only provide numerical values of conceptual contributions. Pre-trained Vision-Language Models (VLMs) can learn rich vision-language correlations from large-scale image-text pairs. Fine-tuning pre-trained VLMs for specific downstream tasks is an effective way to reduce data requirements. Nevertheless, when there is a substantial disparity between the pre-trained model and the target task, existing tuning methods frequently struggle to generalize, necessitating substantial training data to fully adapt VLMs to specialized medical tasks. In this work, we propose a concept adaptive fine-tuning (CptAFT) method based on the pre-trained VLM, BiomedCLIP, to develop a concept-based multi-modal interpretable skin lesion diagnosis model. By incorporating medical texts, such as reports and conceptual terms, our model can recognize fine-grained features and provide robust, natural language-driven interpretability. Moreover, our concept-adaptive method that reconstructs images using concept logits and imposes a consistency loss with the original image, enabling the VLM to quickly adapt to the task with a small amount of training data. Extensive experimental results demonstrate that our approach outperforms state-of-the-art closed box and interpretable models in both classification performance and medically relevant interpretability. In particular, after fine-tuning with a small amount of data, our model outperforms MONET, a model trained on the large Skin Disease Image-Report dataset, by 8.28% in concept recognition ability, demonstrating the interpretability of our model.
Yating Zhu, Xiaoyan Wang 0007, Ming Xia 0005, Pan Mu, Haigen Hu, Xiaoqin Zhang 0002
IEEE J. Biomed. Health Informatics5
2025 TC-Diffuser: Bi-Condition Multi-Modal Diffusion for Tropical Cyclone Forecasting
abstract
Tropical cyclones (TCs) are complex weather systems with strong winds and heavy rainfall, causing substantial loss of life and property. Therefore, accurate TC forecasting is crucial for the effective prevention of disasters caused by TCs. TC forecasting can be regarded as a spatio-temporal prediction problem. It has been proven that using multi-modal data can effectively introduce atmospheric information to achieve better prediction results and higher interpretability. But it also introduces inevitably introduces noise into the prediction process. The diffusion model's unique noise modeling capability can reduce prediction noise when using multi-modal datasets. However, adapting it to TC forecasting has two main challenges: how to extract valuable information from multi-modal data, and how to utilize them to guide the generation process. For the first challenge, while recent methods can predict multiple TC attributes using multi-modal data, they often overlook the interdependence of multiple attributes and the semantic gap between modalities. Considering the interdependence of attributes, we propose two condition generators that capture the commonalities and characteristics of TC attributes, extracting spatio-temporal and environmental features and incorporating expert knowledge. To reduce the semantic gap between multi-modal data, we introduce the PGSA-LSTM module to map primary and auxiliary modalities. For the second challenge, we propose a novel Bi-condition diffusion model that sequentially processes conditions from the characteristics to commonalities of attributes, thereby expanding the guidance information that the diffusion model can accept. Our results surpass state-of-the-art deep learning models and outperform the numerical weather prediction model used by the China Central Meteorological Observatory. TC-Diffuser shows high generalizability across global ocean areas, strong robustness in handling missing data, and higher computational efficiency.
Shiqi Zhang 0009, Pan Mu, Cong Bai
AAAI2
2025 NeighborRetr: Balancing Hub Centrality in Cross-Modal Retrieval
abstract
Cross-modal retrieval aims to bridge the semantic gap between different modalities, such as visual and textual data, enabling accurate retrieval across them. Despite significant advancements with models like CLIP that align cross-modal representations, a persistent challenge remains: the hubness problem, where a small subset of samples (hubs) dominate as nearest neighbors, leading to biased representations and degraded retrieval accuracy. Existing methods often mitigate hubness through post-hoc normalization techniques, relying on prior data distributions that may not be practical in real-world scenarios. In this paper, we directly mitigate hubness during training and introduce NeighborRetr, a novel method that effectively balances the learning of hubs and adaptively adjusts the relations of various kinds of neighbors. Our approach not only mitigates the hubness problem but also enhances retrieval performance, achieving state-of-the-art results on multiple cross-modal retrieval benchmarks. Furthermore, Neighbor-Retr demonstrates robust generalization to new domains with substantial distribution shifts, highlighting its effectiveness in real-world applications. We make our code publicly available at: https://github.com/NeighborRetr.
Zengrong Lin, Zheng Wang 0059, Tianwen Qian, Pan Mu, Sixian Chan 0001, Cong Bai
CVPR4
2025 Prompt-UIE: A Unified Prompt-Driven Framework for Underwater Image Enhancement
abstract
The complex and diverse underwater environment causes various types of degradation in underwater images. However, most existing methods focus on single underwater datasets, where the similarities in degradation limit the model’s exploration of different degradation characteristics. To address this challenge, we developed a new unified model for underwater image enhancement based on prompt learning, called PromptUIE, focusing on the common attributes of underwater images. Prompt-UIE is designed to adapt the pre-trained specific model to various underwater conditions without the need of multiple datasets. It builds upon a specific model and integrates a visual prompt module along with a reverse transmission map (RTM) guided loss function. First, the visual prompt module based on background light guides the specific model to perform enhancements based on different water types through a carefully designed visual prompt strategy. Next, the RTM guided loss function improves the model’s ability to handle non-uniform degradation. Experiments on both full-reference and no-reference datasets demonstrate the effectiveness and robustness of our method. The code is available at https://github.com/Zjut-MultimediaPlus/Prompt-UIE.
Yanling Zhang, Linxuan Luo, Pan Mu, Cong Bai
ICASSP3
2025 TCP-Diffusion: A Multi-modal Diffusion Model for Global Tropical Cyclone Precipitation Forecasting with Change Awareness
abstract
Deep learning methods have made significant progress in regular rainfall forecasting, yet the more hazardous tropical cyclone (TC) rainfall has not received the same attention. While regular rainfall models can offer valuable insights for designing TC rainfall forecasting models, most existing methods suffer from cumulative errors and lack physical consistency. Additionally, these methods overlook the importance of meteorological factors in TC rainfall and their integration with the numerical weather prediction (NWP) model. To address these issues, we propose Tropical Cyclone Precipitation Diffusion (TCP-Diffusion), a multi-modal model for forecasting of TC precipitation given an existing TC in any location globally. It forecasts rainfall around the TC center for the next 12 hours at 3 hourly resolution based on past rainfall observations and multi-modal environmental variables. Adjacent residual prediction (ARP) changes the training target from the absolute rainfall value to the rainfall trend and gives our model the capability of rainfall change awareness, reducing cumulative errors and ensuring physical consistency. Considering the influence of TC-related meteorological factors and the useful information from NWP model forecasts, we propose a multi-model framework with specialized encoders to extract richer information from environmental variables and results provided by NWP models. The results of extensive experiments show that our method outperforms other DL methods and the NWP method from the European Centre for Medium-Range Weather Forecasts (ECMWF).
Pan Mu, Cong Bai, Peter AG Watson
ICML2
2025 Physics-Coupled Frequency Dynamic Adaptation Network for Domain Generalized Underwater Object Detection
Linxuan Luo, Pan Mu, Cong Bai
ACM Multimedia2
2025 IDOL: Meeting Diverse Distribution Shifts with Prior Physics for Tropical Cyclone Multi-Task Estimation
abstract
Tropical Cyclone (TC) estimation aims to accurately estimate various TC attributes in real time. However, distribution shifts arising from the complex and dynamic nature of TC environmental fields, such as varying geographical conditions and seasonal changes, present significant challenges to reliable estimation. Most existing methods rely on multi-modal fusion for feature extraction but overlook the intrinsic distribution of feature representations, leading to poor generalization under out-of-distribution (OOD) scenarios. To address this, we propose an effective Identity Distribution-Oriented Physical Invariant Learning framework (IDOL), which imposes identity-oriented constraints to regulate the feature space under the guidance of prior physical knowledge, thereby dealing distribution variability with physical invariance. Specifically, the proposed IDOL employs the wind field model and dark correlation knowledge of TC to model task-shared and task-specific identity tokens. These tokens capture task dependencies and intrinsic physical invariances of TC, enabling robust estimation of TC wind speed, pressure, inner-core, and outer-core size under distribution shifts. Extensive experiments conducted on multiple datasets and tasks demonstrate the outperformance of the proposed IDOL, verifying that imposing identity-oriented constraints based on prior physical knowledge can effectively mitigates diverse distribution shifts in TC estimation.
Hanting Yan, Pan Mu, Shiqi Zhang 0009, Yuchao Zhu, Cong Bai
NeurIPS2
2025 Lightweight Multi-Stage Aggregation Transformer for robust medical image segmentation
Xiaoyan Wang 0007, Yating Zhu, Dongyan Guo, Pan Mu, Ming Xia 0005, Cong Bai, Zhongzhao Teng, Shengyong Chen
Medical Image Anal.6
2025 Enhancing multiple-style image colorization through context-aware codebook and multi-stage learning
Zheyuan Liu 0009, Hanning Xu, Pan Mu
Vis. Comput.3
2024 Phy-CoCo: Physical Constraint-Based Correlation Learning for Tropical Cyclone Intensity and Size Estimation
abstract
Tropical Cyclone (TC) estimation aims to estimate various attributes of TC in real-time to alleviate and prevent disasters caused by violent TCs. As artificial intelligence technology advances, various deep learning-based multi-task estimation approaches have been proposed. However, most of them only focus on extracting common features of tasks, disregarding potential negative transfer and task interactions between different tasks. This paper is thus motivated to propose a Physical Constraint-based Correlation (Phy-CoCo) learning framework from the perspective of Multi-Task Learning (MTL). Specifically, for task-specific feature learning, we introduce Correlation Modeling (CoM) based on Centrally Expanded Pooling (CEP). Furthermore, for cross-task interaction, we propose a Multi-Domain Recurrent Convolution (MDRC) module to incorporate physical constraints into MTL. These physical constraints enable the transformation of different task features by simulating the physical relations among different attributes of TC. Lastly, in combination with a task-shared network that leverages the hybrid fusion of multi-modal data, our MTL framework accurately estimates various TC attributes. Extensive experiments conducted on our constructed dataset demonstrate that the proposed Phy-CoCo outperforms previous methods in TC estimation in terms of estimation error, verifying the potential of the physics-incorporated MTL model.
Hanting Yan, Pan Mu, Cong Bai
ECAI2
2024 BFMEF: Brightness-Free Multi-exposure Image Fusion via Adaptive Correction
abstract
In recent years, deep learning has revolutionized the field of multi-exposure image fusion (MEF), overcoming the limitations of traditional techniques and proving more effective in complex and diverse scenarios. However, existing MEF methods mainly focus on paired extreme exposure dual inputs and pay little attention to single inputs or other extreme conditions. To address this gap, this paper introduces a fusion architecture that can adaptively correct various input forms with Auto-Gamma Correction (AGC). By leveraging the powerful information interaction capability of the Transformer, the Light-Guided Module (LGM) effectively extracts the brightness information from the input images. Furthermore, a specially designed color enhancement algorithm obtains a high-saturation fused image. Experimental results show that compared with the most advanced methods, our method achieves the best effects in terms of visuals and performance.
Pan Mu, Binjia Zhou, Zhiying Du, Xiaoyan Wang 0007
ICME1
2024 Learning to Search a Lightweight Generalized Network for Medical Image Fusion
abstract
Image fusion is indispensable in a comprehensive medical imaging pipeline. By embracing deep learning technology, medical image fusion has achieved tremendous progress over the past few years. However, existing approaches make efforts on the specific type of medical image fusion task and may face difficulties in generalizing well. Moreover, most of them strain every nerve to design various architectures with an increase of the width of depth, placing an obstacle in running efficiency. To address the above problems, we propose an Auto-searching Light-weighted Multi-source Fusion network, namely ALMFnet, aiming at incorporating both software and hardware knowledge in a network architecture searching manner for medical image fusion. Specifically, the ALMFnet, consisting of two different feature-extracting modules and one fusion module, is developed to extract and refine multi-source features in a generalized model. Besides, motivated by the collaborative principle, we introduce hardware constraints for sufficient searching the each particular component, further reducing the complexity of the obtained model. Furthermore, to preserve important details in pathological image areas, we introduce a segmentation mask into the developed method. Experimental results demonstrate that our generalized model outperforms previous methods not only in terms of quantitative scores but also in model complexity. Source code will be available at https://github.com/RollingPlain/ALMFnet.
Pan Mu, Guanyao Wu, Jinyuan Liu 0001, Yuduo Zhang, Xin Fan 0001, Risheng Liu
IEEE Trans. Circuits Syst. Video Technol.1
2024 Underwater image enhancement via color conversion and white balance-based fusion
Hanning Xu, Pan Mu, Zheyuan Liu 0009, Shichao Cheng
Vis. Comput.2
2023 Towards General and Fast Video Derain via Knowledge Distillation
abstract
As a common natural weather condition, rain can obscure video frames and thus affect the performance of the visual system, so video derain receives a lot of attention. In natural environments, rain has a wide variety of streak types, which increases the difficulty of the rain removal task. In this paper, we propose a Rain Review-based General video derain Network via knowledge distillation (named RRGNet) that handles different rain streak types with one pre-training weight. Specifically, we design a frame grouping-based encoder-decoder network that makes full use of the temporal information of the video. Further, we use the old task model to guide the current model in learning new rain streak types while avoiding forgetting. To consolidate the network’s ability to derain, we design a rain review module to play back data from old tasks for the current model. The experimental results show that our developed general method achieves the best results in terms of running speed and derain effect.
Defang Cai, Pan Mu, Sixian Chan 0001, Zhanpeng Shao, Cong Bai
ICME2
2023 Histogram-guided Video Colorization Structure with Spatial-Temporal Connection
abstract
Video colorization, aiming at obtaining colorful and plausible results from grayish frames, has aroused a lot of interest recently. Nevertheless, how to maintain temporal consistency while keeping the quality of colorized results remains challenging. To tackle the above problems, we present a Histogram-guided Video Colorization with Spatial-Temporal connection structure (named ST-HVC). To fully exploit the chroma and motion information, the joint flow and histogram module is tailored to integrate the histogram and flow features. To manage the blurred and artifact, we design a combination scheme attending to temporal detail and flow feature combination. We further recombine the histogram, flow and sharpness features via a U-shape network. Extensive comparisons are conducted with several state-of-the-art image and video-based methods, demonstrating that the developed method achieves excellent performance both quantitatively and qualitatively in two video datasets.
Zheyuan Liu 0009, Pan Mu, Hanning Xu, Cong Bai
ICME2
2023 Transmission and Color-guided Network for Underwater Image Enhancement
abstract
In recent years, with the continuous development of the marine industry, underwater image enhancement has attracted plenty of attention. Unfortunately, the propagation of light in water will be absorbed by water bodies and scattered by suspended particles, resulting in color deviation and low contrast. To solve these two problems, we propose an Adaptive Transmission and Dynamic Color guided network (named ATDCnet) for underwater image enhancement. In particular, to exploit the knowledge of physics, we design an Adaptive Transmission-directed Module (ATM) to better guide the network. To deal with the color deviation problem, we design a Dynamic Color-guided Module (DCM) to post-process the enhanced image color. Further, we design an Encoder-Decoder-based (EDC) structure with attention and a multistage feature fusion mechanism to perform color restoration and contrast enhancement simultaneously. Extensive experiments demonstrate the state-of-the-art performance of the ATDCnet on multiple benchmark datasets.
Pan Mu, Haotian Qian, Cong Bai
ICME1
2023 Little Strokes Fell Great Oaks: Boosting the Hierarchical Features for Multi-exposure Image Fusion
abstract
In recent years, deep learning networks have made remarkable strides in the domain of multi-exposure image fusion. Nonetheless, prevailing approaches often involve directly feeding over-exposed and under-exposed images into the network, which leads to the under-utilization of inherent information present in the source images. Additionally, unsupervised techniques predominantly employ rudimentary weighted summation for color channel processing, culminating in an overall desaturated final image tone. To partially mitigate these issues, this study proposes a gamma correction module specifically designed to fully leverage latent information embedded within source images. Furthermore, a modified transformer block, embracing self-attention mechanisms, is introduced to optimize the fusion process. Ultimately, a novel color enhancement algorithm is presented to augment image saturation while preserving intricate details. The source code is available at https://github.com/ZhiyingDu/BHFMEF.
Pan Mu, Zhiying Du, Jinyuan Liu 0001, Cong Bai
ACM Multimedia1
2023 A Generalized Physical-knowledge-guided Dynamic Model for Underwater Image Enhancement
abstract
Underwater images often suffer from color distortion and low contrast resulting in various image types, due to the scattering and absorption of light by water. While it is difficult to obtain high-quality paired training samples with a generalized model. To tackle these challenges, we design a Generalized Underwater image enhancement method via a Physical-knowledge-guided Dynamic Model (short for GUPDM). In particular, to cover complex underwater scenes, this study changes the global atmosphere light and the transmission to simulate various underwater image types through the formation model. We then design an Atmosphere-based Dynamic Structure (ADS) and Transmission-guided Dynamic Structure (TDS) that use dynamic convolutions to adaptively extract prior information from underwater images and generate parameters for Prior-based Multi-scale Structure (PMS). These two modules enable the network to select appropriate parameters for various water types adaptively. Besides, the multi-scale feature extraction module in PMS uses convolution blocks with different kernel sizes and obtains weights for each feature map via channel attention block. The source code will be available at https://github.com/shiningZZ/GUPDM
Pan Mu, Hanning Xu, Zheyuan Liu 0009, Zheng Wang 0059, Sixian Chan 0001, Cong Bai
ACM Multimedia1
2023 A General Descent Aggregation Framework for Gradient-Based Bi-Level Optimization
abstract
In recent years, a variety of gradient-based methods have been developed to solve Bi-Level Optimization (BLO) problems in machine learning and computer vision areas. However, the theoretical correctness and practical effectiveness of these existing approaches always rely on some restrictive conditions (e.g., Lower-Level Singleton, LLS), which could hardly be satisfied in real-world applications. Moreover, previous literature only proves theoretical results based on their specific iteration strategies, thus lack a general recipe to uniformly analyze the convergence behaviors of different gradient-based BLOs. In this work, we formulate BLOs from an optimistic bi-level viewpoint and establish a new gradient-based algorithmic framework, named Bi-level Descent Aggregation (BDA), to partially address the above issues. Specifically, BDA provides a modularized structure to hierarchically aggregate both the upper- and lower-level subproblems to generate our bi-level iterative dynamics. Theoretically, we establish a general convergence analysis template and derive a new proof recipe to investigate the essential theoretical properties of gradient-based BLO methods. Furthermore, this work systematically explores the convergence behavior of BDA in different optimization scenarios, i.e., considering various solution qualities (i.e., global/local/stationary solution) returned from solving approximation subproblems. Extensive experiments justify our theoretical results and demonstrate the superiority of the proposed algorithm for hyper-parameter optimization and meta-learning tasks. Source code is available at https://github.com/vis-opt-group/BDA.
Risheng Liu, Pan Mu, Xiaoming Yuan 0001, Shangzhi Zeng, Jin Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Optimization-Inspired Learning With Architecture Augmentations and Control Mechanisms for Low-Level Vision
abstract
In recent years, there has been a growing interest in combining learnable modules with numerical optimization to solve low-level vision tasks. However, most existing approaches focus on designing specialized schemes to generate image/feature propagation. There is a lack of unified consideration to construct propagative modules, provide theoretical analysis tools, and design effective learning mechanisms. To mitigate the above issues, this paper proposes a unified optimization-inspired learning framework to aggregate Generative, Discriminative, and Corrective (GDC for short) principles with strong generalization for diverse optimization models. Specifically, by introducing a general energy minimization model and formulating its descent direction from different viewpoints (i.e., in a generative manner, based on the discriminative metric and with optimality-based correction), we construct three propagative modules to effectively solve the optimization models with flexible combinations. We design two control mechanisms that provide the non-trivial theoretical guarantees for both fully- and partially-defined optimization formulations. Under the support of theoretical guarantees, we can introduce diverse architecture augmentation strategies such as normalization and search to ensure stable propagation with convergence and seamlessly integrate the suitable modules into the propagation respectively. Extensive experiments across varied low-level vision tasks validate the efficacy and adaptability of GDC.
Risheng Liu, Zhu Liu 0004, Pan Mu, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.3
2022 Structure-Inferred Bi-level Model for Underwater Image Enhancement
abstract
Very recently, with the development of underwater robots, underwater image enhancement arising growing interests in the computer vision community. However, owing to light being scattered and absorbed while it traveling in water, underwater captured images often suffer from color cast and low visibility. Existing methods depend on specific prior knowledge and training data to enhance underwater images in the absence of structure information, which results in poor and unnatural performance. To this end, we propose a Structural-Inferred Bi-level Model (SIBM) that incorporates different modalities of knowledge (i.e., semantic domain, gradient-domain, and pixel domain) hierarchically enhancing underwater images. In particular, by introducing a semantic mask, we individually optimize the forehand branch that avoids unnecessary interference arising from the background region. We design a gradient-based high-frequency branch to exploit gradient-space guidance for preserving texture structures. Moreover, we construct a pixel-based branch by feeding semantic and gradient information to enhance underwater images. To exploit different modalities, we introduce a hyper-parameter optimization scheme to fuse the above domain information. Experimental results illustrate that the developed method not only outperforms the previous methods in quantitative scores but also generalizes well on real-world underwater datasets. Source code is available at \hrefhttps://github.com/IntegralCoCo/SIBM https://github.com/IntegralCoCo/SIBM.
Pan Mu, Haotian Qian, Cong Bai
ACM Multimedia1
2022 Real-world Underwater Image Enhancement via Degradation-aware Dynamic Network
Haotian Qian, Wentao Tong, Pan Mu, Zheyuan Liu 0009, Hanning Xu
PRICAI (3)3
2022 Triple-Level Model Inferred Collaborative Network Architecture for Video Deraining
abstract
Video deraining is an important issue for outdoor vision systems and has been investigated extensively. However, designing optimal architectures by the aggregating model formation and data distribution is a challenging task for video deraining. In this paper, we develop a model-guided triple-level optimization framework to deduce network architecture with cooperating optimization and auto-searching mechanism, named Triple-level Model Inferred Cooperating Searching (TMICS), for dealing with various video rain circumstances. In particular, to mitigate the problem that existing methods cannot cover various rain streaks distribution, we first design a hyper-parameter optimization model about task variable and hyper-parameter. Based on the proposed optimization model, we design a collaborative structure for video deraining. This structure includes Dominant Network Architecture (DNA) and Companionate Network Architecture (CNA) that is cooperated by introducing an Attention-based Averaging Scheme (AAS). To better explore inter-frame information from videos, we introduce a macroscopic structure searching scheme that searches from Optical Flow Module (OFM) and Temporal Grouping Module (TGM) to help restore latent frame. In addition, we apply the differentiable neural architecture searching from a compact candidate set of task-specific operations to discover desirable rain streaks removal architectures automatically. Extensive experiments on various datasets demonstrate that our model shows significant improvements in fidelity and temporal consistency over the state-of-the-art works. Source code is available at https://github.com/vis-opt-group/TMICS.
Pan Mu, Zhu Liu 0004, Risheng Liu, Xin Fan 0001
IEEE Trans. Image Process.1
2021 Investigating Customization Strategies and Convergence Behaviors of Task-Specific ADMM
abstract
Alternating Direction Method of Multiplier (ADMM) has been a popular algorithmic framework for separable optimization problems with linear constraints. For numerical ADMM fail to exploit the particular structure of the problem at hand nor the input data information, leveraging task-specific modules (e.g., neural networks and other data-driven architectures) to extend ADMM is a significant but challenging task. This work focuses on designing a flexible algorithmic framework to incorporate various task-specific modules (with no additional constraints) to improve the performance of ADMM in real-world applications. Specifically, we propose Guidance from Optimality (GO), a new customization strategy, to embed task-specific modules into ADMM (GO-ADMM). By introducing an optimality-based criterion to guide the propagation, GO-ADMM establishes an updating scheme agnostic to the choice of additional modules. The existing task-specific methods just plug their task-specific modules into the numerical iterations in a straightforward manner. Even with some restrictive constraints on the plug-in modules, they can only obtain some relatively weaker convergence properties for the resulted ADMM iterations. Fortunately, without any restrictions on the embedded modules, we prove the convergence of GO-ADMM regarding objective values and constraint violations, and derive the worst-case convergence rate measured by iteration complexity. Extensive experiments are conducted to verify the theoretical results and demonstrate the efficiency of GO-ADMM.
Risheng Liu, Pan Mu, Jin Zhang 0002
IEEE Trans. Image Process.2
2020 Image Restoration Via Data-Dependent Proximal Averaged Optimization
abstract
Maximum A Posterior (MAP) acts as one of the most popular modeling scheme in image restoration and is usually reduced to a separable optimization model. Unfortunately, it is challenging to establish exact regularization term and the model with complex priors is hard to optimize. In additionally, it is still hard to incorporate different domain knowledge and data-dependent information into MAP model without changing the property of the objective. To partially address the above issues, we develop a Data-dependent Proximal Averaged (DPA) paradigm through optimizing objective and data-dependent feasibility constraint for the challenging Image Restoration (IR) tasks. Both visual and quantitative comparison results demonstrate that our method outperforms the state of the art.
Pan Mu, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICASSP1
2020 Sequential Deep Unrolling With Flow Priors For Robust Video Deraining
abstract
Video deraining has attracted wide attention since the urgent demand of high-quality video in recent years. The indistinct details and nonideal deraining effects are the most common defects in existing techniques, whose cause lies in the insufficient usage of single-frame image and temporal information. To effectively settle video deraining, we establish a new deraining model with flow priors to simultaneously introduce spatial and temporal information for accurately depicting the enhancement model of the current frame. A sequential deep unrolling framework is substantially presented by solving this model based on optimization techniques. The ablation study indicates our effectiveness as far as the design of architecture. Plenty of subjective and objective evaluations fully demonstrate our superiority in detail recovery and deraining effects against other state-of-the-are video deraining approaches.
Xinwei Xue, Ying Ding 0006, Pan Mu, Long Ma 0002, Risheng Liu, Xin Fan 0001
ICASSP3
2020 Flexible Bilevel Image Layer Modeling For Robust Deraining
abstract
Visual quality degradation by rain streaks in images/videos is a significant factor that makes many computer vision systems fail to function properly. However, existing rain removal methods tend to remove a specific type of rain streaks while cannot deal with diverse real rainy images. In this paper, we formulate a novel rain model collectively with two contrasting rain streaks and a weighting map. To self-adaptively handle the rain removal problem in the presence of various types of rain streaks, we further propose a bilevel optimization learning framework. Then, we synthesize a new dataset to evaluate the ability of our method to deal with diverse rain streaks. Extensive experiments show that our method can make better performance on both synthesized and real rainy images.
Pan Mu, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
ICME2
2020 A Generic First-Order Algorithmic Framework for Bi-Level Programming Beyond Lower-Level Singleton
abstract
In recent years, a variety of gradient-based bi-level optimization methods have been developed for learning tasks. However, theoretical guarantees of these existing approaches often heavily rely on the simplification that for each fixed upper-level variable, the lower-level solution must be a singleton (a.k.a., Lower-Level Singleton, LLS). In this work, by formulating bi-level models from the optimistic viewpoint and aggregating hierarchical objective information, we establish Bi-level Descent Aggregation (BDA), a flexible and modularized algorithmic framework for bi-level programming. Theoretically, we derive a new methodology to prove the convergence of BDA without the LLS condition. Furthermore, we improve the convergence properties of conventional first-order bi-level schemes (under the LLS simplification) based on our proof recipe. Extensive experiments justify our theoretical results and demonstrate the superiority of the proposed BDA for different tasks, including hyper-parameter optimization and meta learning.
Risheng Liu, Pan Mu, Xiaoming Yuan 0001, Shangzhi Zeng, Jin Zhang 0002
ICML2
2020 Investigating Task-Driven Latent Feasibility for Nonconvex Image Modeling
abstract
Properly modeling latent image distributions plays an important role in a variety of image-related vision problems. Most exiting approaches aim to formulate this problem as optimization models (e.g., Maximum A Posterior, MAP) with handcrafted priors. In recent years, different CNN modules are also considered as deep priors to regularize the image modeling process. However, these explicit regularization techniques require deep understandings on the problem and elaborately mathematical skills. In this work, we provide a new perspective, named Task-driven Latent Feasibility (TLF), to incorporate specific task information to narrow down the solution space for the optimization-based image modeling problem. Thanks to the flexibility of TLF, both designed and trained constraints can be embedded into the optimization process. By introducing control mechanisms based on the monotonicity and boundedness conditions, we can also strictly prove the convergence of our proposed inference process. We demonstrate that different types of image modeling problems, such as image deblurring and rain streaks removals, can all be appropriately addressed within our TLF framework. Extensive experiments also verify the theoretical results and show the advantages of our method against existing state-of-the-art approaches.
Risheng Liu, Pan Mu, Xin Fan 0001, Zhongxuan Luo
IEEE Trans. Image Process.2
2019 Learning Bilevel Layer Priors for Single Image Rain Streaks Removal
abstract
Rain streaks removal is an important issue of the outdoor vision system and recently has been investigated extensively. In the past decades, maximum a posterior and network-based architecture have been attracting considerable attention for this problem. However, it is challenging to establish effective regularization priors and the cost function with complex prior is hard to optimize. On the other hand, it is still hard to incorporate data-dependent information into conventional numerical iterations. To partially address the above limits and inspired by the leader-follower gaming perspective, we introduce an unrolling strategy to incorporate data-dependent network architectures into the established iterations, i.e., a learning bilevel layer priors method to jointly investigate the learnable feasibility and optimality of rain streaks removal problem. Both visual and quantitative comparison results demonstrate that our method outperforms the state of the art.
Pan Mu, Risheng Liu, Xin Fan 0001, Zhongxuan Luo
IEEE Signal Process. Lett.1