EDBT 2026 Demo / reviewers in the wild / expert
Faming Fang
dblp:96/8174
· DBLP profile ↗
77ranked-venue papers
12as first author
50since 2021 · last 2026
0000-0003-4511-4813ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 5 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 5 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Geometric Perspective on Optimizing Vector Quantized Latent Diffusion Model for Image RestorationabstractIn this paper, we investigate the limitations of the Vector Quantized Latent Diffusion Model (VQ-LDM) in restoration tasks. We identify a performance gap between the Vector Quantization (VQ) and Diffusion Model components, manifested as a significant discrepancy between the reconstruction quality of ground truth images processed via VQ autoregression and degraded images restored by VQ-LDM. Through experiments, we attribute this gap primarily to the lack of robustness in the mapped points of VQ within the original VQ-LDM framework. To address this issue, we propose a geometric based optimization approach. First, we introduce a simple yet effective method, termed interpolation-based latent initial state optimization, which mitigates the performance gap by replacing the original mapped points with interpolated values, supported by theoretical analysis. Here, the latent initial state refers specifically to the input of the diffusion model. Building upon this, we further propose a Chebyshev center-based latent initial state optimization, an elegant theoretical solution from a geometric perspective, that further enhances restoration performance. Our improvements consistently achieve superior results across nine benchmark datasets. Chen Hang, Haoming Chen, Xuwei Fang, Weisheng Xie, Xiangxiang Gao, Faming Fang, Guixu Zhang |
AAAI | 6 |
| 2026 | Towards Privacy-Protected Generalized Gaze Estimation Using Diffusion Models and Domain Stability Adaptation FrameworkabstractModern gaze estimation models can accurately predict human gaze from facial images. However, due to privacy concerns and intricate data collection procedures, gaze estimation datasets are typically smaller and less diverse compared to those for other vision tasks, which directly leads to poor generalization in gaze estimation models. Common solutions, such as domain adaptation models, require additional domain-specific data, yet such data is often difficult to obtain due to privacy restrictions. Meanwhile, domain generalization models suffer from limited performance due to insufficient training data. To address these fundamental challenges---privacy and data diversity---we explore privacy-preserving gaze data generation schemes and propose a novel data-driven generalization solution. Specifically, we develop two diffusion-based generative models, DDPM-Gaze and LDM-Gaze, for synthesizing gaze data. We demonstrate that synthetic data can significantly improve generalization performance when simply used with fine-tuning-based methods. Furthermore, we introduce the Domain Stability Adaptation (DSA) framework, a simple yet effective domain generalization approach that enhances model robustness by increasing the domain uncertainty of input samples while reducing prediction uncertainty. Extensive experiments validate the effectiveness of our synthetic data and demonstrate the superiority of our data-driven generalization solution. Shengcheng Ye, Faming Fang |
AAAI | 3 |
| 2026 | Deep Algorithm Unrolling with Alignment Embedding for Guided Image Super-resolution
Faming Fang, Tingting Wang 0007, Junkang Zhang, Aimin Zhou, Riquan Zhang, Guixu Zhang |
Int. J. Comput. Vis. | 1 |
| 2026 | Eliciting CLIP's intrinsic attribute knowledge through a dual-cache guided mechanism for class-incremental learning
Shengcheng Ye, Yaomin Huang, Faming Fang, Guixu Zhang |
Knowl. Based Syst. | 4 |
| 2026 | Task-aware all-in-one guided image super-resolution
Tingting Wang 0007, Jun Wang 0024, Qiuhai Yan, Junkang Zhang, Faming Fang, Guixu Zhang |
Pattern Recognit. | 5 |
| 2026 | Deep Unfolding Segmentation Network for Under-Sampled Magnetic Resonance ImagesabstractMagnetic Resonance (MR) image segmentation is a critical task in assisting disease diagnosis. Most existing methods assume that the images being segmented are fully-sampled. However, they ignore the fact that MR images obtained in clinics are often reconstructed from under-sampled k-space data. There are artifacts or distorted details in the reconstruction, leading to unsatisfactory segmentation performance. In this paper, we propose an end-to-end deep unfolding framework to segment desired lesions or organs from the under-sampled k-space data. Specifically, we build a new model to combine the compressive sensing-based under-sampled image reconstruction and level-set-based segmentation. In this model, we introduce an L0 norm on the reconstruction images to enforce smoothing while preserving important edge and boundary, boosting downstream segmentation performance. We employ the Augmented Lagrangian Method to seek the solution and unfold the iterative algorithm into a deep neural network, called deep unfolding segmentation network (DUSNet). To further enhance segmentation performance, we introduce a boundary loss function, which encourages the model to effectively capture edge details of the regions of interest and imposes geometric constraints on the segmentation results. Through end-to-end training, DUSNet can efficiently segment target regions from under-sampled k-space data. Comprehensive experiments demonstrate that the proposed DUSNet outperforms existing state-of-the-art methods for under-sampled MR image segmentation, achieving superior segmentation accuracy. Le Hu, Pengcheng Lei, Faming Fang, Guixu Zhang |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Decoupling Scattering: Pseudo-Label Guided NeRF for Scenes with Scattering MediaabstractNeural Radiance Fields (NeRF) has been widely used in computer vision and graphics, achieving impressive results in novel view synthesis and multi-view 3D reconstruction. However, despite its excellent performance under ideal conditions, NeRF struggles in challenging environments such as hazy, foggy, and underwater scenes, primarily due to the difficulty in decoupling objects from the scattering medium. To mitigate this limitation, we proposed a novel approach for NeRF in scenes with scattering media. Specifically, we leverage pseudo-labels during the early stage of training to guide NeRF in decoupling the densities of objects and the scattering medium, guiding the model toward a more appropriate search space. Furthermore, we introduce a Cyclical Progressive Dimensional Optimization Strategy (CPDOS) that focuses on optimizing a single or a few variables during specific periods. Experimental results demonstrate that our method can effectively simulate hazy and underwater scenes, accurately decouple the scattering medium from objects, estimate atmospheric parameters, and outperform existing methods in novel view synthesis and image restoration tasks. Junkang Zhang, Faming Fang, Guixu Zhang |
AAAI | 3 |
| 2025 | Explicit Depth-Aware Blurry Video Frame Interpolation Guided by Differential CurvesabstractBlurry video frame interpolation (BVFI), which aims to generate high-frame-rate clear videos from low-frame-rate blurry inputs, is a challenging yet significant task in computer vision. Current state-of-the-art approaches typically rely on linear or quadratic models to estimate intermediate motion. However, these methods often overlook depth variations that occur during fast object motion, leading to changes in object size and hindering interpolation performance.This paper proposes the Differential Curves-guided Blurry Video Frame Interpolation (DC-BVFI) framework, which leverages the differential curves theory to analyze and mitigate the effects of depth variations caused by object motion. Specifically, DC-BVFI consists of UBNet and MPNet. Unlike prior approaches that rely on optical flow for frame interpolation, MPNet is designed to estimate the 3D scene flow, which facilitates a more precise awareness of depth and velocity variations. Since scene flow cannot be directly inferred in the 2D frame space, UBNet is introduced to transform them into 3D point maps. Extensive experiments demonstrate that the proposed DC-BVFI framework surpasses state-of-the-art performance in simulated and real-world datasets. Zaoming Yan, Pengcheng Lei, Tingting Wang 0007, Faming Fang, Junkang Zhang, Yaomin Huang |
CVPR | 4 |
| 2025 | First-order State Space Model for Lightweight Image Super-resolutionabstractState space models (SSMs), particularly Mamba, have shown promise in NLP tasks and are increasingly applied to vision tasks. However, most Mamba-based vision models focus on network architecture and scan paths, with little attention to the SSM module. In order to explore the potential of SSMs, we modified the calculation process of SSM without increasing the number of parameters to improve the performance on lightweight super-resolution tasks. In this paper, we introduce the First-order State Space Model (FSSM) to improve the original Mamba module, enhancing performance by incorporating token correlations. We apply a first-order hold condition in SSMs, derive the new discretized form, and analyzed cumulative error. Extensive experimental results demonstrate that FSSM improves the performance of MambaIR on five benchmark datasets without additionally increasing the number of parameters, and surpasses current lightweight SR methods, achieving state-of-the-art results. Yekai Lu, Guang Yang 0068, Faming Fang, Guixu Zhang |
ICASSP | 5 |
| 2025 | Surface-Aware Feed-Forward Quadratic Gaussian for Frame Interpolation with Large MotionabstractMotion in the real world takes place in 3D space.
Existing Frame Interpolation methods often estimate global receptive fields in 2D frame space.
Due to the limitations of 2D space, these global receptive fields are limited, which makes it difficult to match object correspondences between frames, resulting in sub-optimal performance when handling large-motion scenarios.
In this paper, we introduce a novel pipeline for exploring object correspondences based on differential surface theory.
The differential surface coordinate system provides a better representation of the real world, enabling effective exploration of object correspondences.
Specifically, the pipeline first transforms an input pair of video frames from the image coordinate system to the differential surface coordinate system.
Subsequently, within this coordinate system, object correspondences are explored based on surface geometric properties and the surface uniqueness theorem.
Experimental findings showcase that our method attains state-of-the-art performance across large motion benchmarks.
Our method demonstrates the state-of-the-art performance on these VFI subsets with large motion. Zaoming Yan, Yaomin Huang, Pengcheng Lei, Qizhou Chen, Guixu Zhang, Faming Fang |
NeurIPS | 6 |
| 2025 | HMSFU: A hierarchical multi-scale fusion unit for video prediction and beyondabstractAbstract Video prediction is the process of learning necessary information from historical frames to predict future video frames. Learning features from historical frames is a crucial step in this process. However, most current methods have a relatively single‐scale learning approach, even if they learn features at different scales, they cannot fully integrate and utilise them, resulting in unsatisfactory prediction results. To address this issue, a hierarchical multi‐scale fusion unit (HMSFU) is proposed. By using a hierarchical multi‐scale architecture, each layer predicts future frames at different granularities using different convolutional scales. The abstract features from different layers can be fused, enabling the model not only to capture rich contextual information but also to expand the model's receptive field, enhance its expressive power, and improve its applicability to complex prediction scenarios. To fully utilise the expanded receptive field, HMSFU incorporates three fusion modules. The first module is the single‐layer historical attention fusion module, which uses an attention mechanism to fuse the features from historical frames into the current frame at each layer. The second module is the single‐layer spatiotemporal fusion module, which fuses complementary temporal and spatial features at each layer. The third module is the multi‐layer spatiotemporal fusion module, which fuses spatiotemporal features from different layers. Additionally, the authors not only focus on the frame‐level error using mean squared error loss, but also introduce the novel use of Kullback–Leibler (KL) divergence to consider inter‐frame variations. Experimental results demonstrate that our proposed HMSFU model achieves the best performance on popular video prediction datasets, showcasing its remarkable competitiveness in the field. Hongchang Zhu, Faming Fang |
IET Comput. Vis. | 2 |
| 2025 | DUGCN: Deep Unfolding Gradient Consistency Network for fast MRI reconstruction
Le Hu, Faming Fang, Shaoxin Li 0005 |
Neurocomputing | 4 |
| 2025 | ISGM-Fus: Internal structure-guided model for multispectral and hyperspectral image fusion
Cong Liu 0011, Jinming Qian, Faming Fang |
Neurocomputing | 3 |
| 2025 | Deep maximum a posterior estimator for accelerated MRI reconstruction
Tingting Wang 0007, Shengcheng Ye, Faming Fang, Guixu Zhang, Yuanyi Zheng |
Knowl. Based Syst. | 3 |
| 2025 | Digging Deeper in Gradient for Unrolling-Based Accelerated MRI ReconstructionabstractThere are two main methods that can be used to accelerate MRI reconstruction: parallel imaging and compressed sensing. To further accelerate the sampling process, the combination of these two methods has been extensively studied in recent years. However, existing MRI reconstruction methods often overlook the exploration of high-frequency information of images, leading to sub-optimal recovery of fine details in the reconstructed results. To address this issue, we conduct an in-depth analysis of image gradients and propose a novel MRI reconstruction model based on Maximum a Posteriori (MAP) estimation. We first establish the Cumulative Deviation from Maximum Gradient magnitude (CDMG) prior for fully sampled MR images through theoretical analysis, then incorporate this explicit CDMG prior along with an implicit deep prior to form the prior probability term. This combination of priors strikes a balance between physically informed constraints and data-driven adaptability, aiding in the recovery of meaningful high-frequency information. Additionally, we introduce a multi-order gradient operator to enhance the observation model, thereby improving the accuracy of the likelihood term. Through MAP estimation, we develop a novel accelerated MRI reconstruction model, the optimization of which is achieved by unrolling it into a convolutional neural network structure, referred to as DDGU-Net. Extensive experimental results demonstrate the effectiveness of our approach in reconstructing high-quality MR images and achieving state-of-the-art (SOTA) results, particularly at higher acceleration factors. Faming Fang, Tingting Wang 0007, Guixu Zhang, Fang Li 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Homography Estimation With Adaptive Query Transformer and Gated Interaction ModuleabstractHomography estimation is essential for aligning images captured from different viewpoints by accurately modeling the geometric relationship between them. In homography estimation, global information plays a critical role. To establish global correspondences, cross-attention has been widely used in recent studies. However, vanilla cross-attention mechanisms treat queries in redundant and low-texture areas the same as those in richly textured areas, leading to the accumulation and propagation of erroneous information. We define this phenomenon, where the model excessively attends to queries in redundant and low-texture areas, as query over-focusing. To alleviate query over-focusing and achieve fine-grained homography estimation, we propose a novel homography estimation network, termed AGNet, which integrates an Adaptive Query Transformer (AQFormer) and a Gated Interaction Module (GIM). The AQFormer is designed to dynamically adjust attention by applying a mask to queries, allowing the model to adaptively emphasize feature-rich regions while suppressing redundant or weakly textured areas. Meanwhile, the GIM selectively captures local information by adjusting convolutional kernels based on input, enhancing the extraction of shared features between image pairs. Extensive experiments on various datasets demonstrate that AGNet significantly improves accuracy in homography estimation, particularly in challenging scenarios with low overlap and large viewpoint variations. Faming Fang, Tingting Wang 0007, Guixu Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Robust Deep Convolutional Dictionary Model With Alignment Assistance for Multi-Contrast MRI Super-ResolutionabstractMulti-contrast magnetic resonance imaging (MCMRI) super-resolution (SR) methods aims to leverage the complementary information present in multi-contrast images. However, existing methods encounter several limitations. Firstly, most current networks fail to appropriately model the correlations of multi-contrast images and lack certain interpretability. Secondly, they often overlook the negative impact of spatial misalignment between modalities in clinical practice. Thirdly, existing methods do not effectively constrain the complementary information learned between multi-contrast images, resulting in information redundancy and limiting their model performance. In this paper, we propose a robust alignment-assisted multi-contrast convolutional dictionary (A2-CDic) model to address these challenges. Specifically, we develop an observation model based on convolutional sparse coding to explicitly represent multi-contrast images as common (e.g., consistent textures) and unique (e.g., inconsistent structures and contrasts) components. Considering there are spatial misalignments in real-world multi-contrast images, we incorporate a spatial alignment module to compensate for the misaligned structures. This approach enables the proposed model to fully exploit the valuable information in the reference image while mitigating interference from inconsistent information. We employ the proximal gradient algorithm to optimize the model and unroll the iterative steps into a multi-scale convolutional dictionary network. Furthermore, we utilize mutual information losses to constrain the extracted common and unique components. This constraint reduces the redundancy between the decomposed components, allowing each sub-module to learn more representative features. We evaluate our model on four publicly available datasets comprising internal, external, spatially aligned, and misaligned MCMRI images. The experimental results demonstrate that our model surpasses existing state-of-the-art MCMRI SR methods in terms of both generalization ability and overall performance. Code is available at https://github.com/lpcccc-cv/A2-CDic. Pengcheng Lei, Miaomiao Zhang 0002, Faming Fang, Guixu Zhang |
IEEE Trans. Medical Imaging | 3 |
| 2025 | FrDiff: Framelet-Based Conditional Diffusion Model for Multispectral and Panchromatic Image FusionabstractThe process of fusing low-resolution multispectral (LRMS) and high-resolution panchromatic (PAN) imagery, commonly referred to as pansharpening, is intended to generate high-resolution multispectral (HRMS) imagery. Typically, most pre-existing pansharpening frameworks mainly emphasize the straightforward learning of the mapping relationship among PAN and LRMS images to HRMS images. However, a key limitation of these frameworks is their potential overemphasis on spatial information, particularly the enhancement of low-frequency components. As a result, such an oversight potentially hinders the model's ability to simultaneously restore both spectral and spatial details. To address this issue, we propose a novel pansharpening model based on the denoising diffusion probabilistic model (DDPM), dubbed FrDiff. Specifically, we build a framelet-based conditional diffusion model that leverages the generative power of diffusion models to produce more refine results. Different from conventional methods directly inferring HRMS images, our strategy is designed to project their framelet coefficients, utilizing the available PAN and LRMS images as resources. This approach enables the separation of high-frequency and low-frequency components through framelet transformation, which are subsequently recombined to create a novel set of conditional embeddings that feed into the diffusion process. At the same time, the powerful predictive power of the diffusion model is exploited to simultaneously recover the high-frequency and low-frequency components of the HRMS. Moreover, we introduce a framelet-oriented cross-attention module dedicated to honing spectral fidelity. This module is crucial for improving the spectral precision of the HRMS images, ensuring a balanced emphasis on both spatial and spectral enhancements. Quantitative and qualitative experiments on multiple benchmark datasets demonstrate that the proposed method achieves more robustness and high-quality results than other state-of-the-art pansharpening methods. Junkang Zhang, Faming Fang, Tingting Wang 0007, Guixu Zhang |
IEEE Trans. Multim. | 2 |
| 2024 | Harmonizing Knowledge Transfer in Neural Network with Unified Distillation
Yaomin Huang, Zaomin Yan, Chaomin Shen 0001, Faming Fang, Guixu Zhang |
ECCV (33) | 4 |
| 2024 | Three-Stage Temporal Deformable Network for Blurry Video Frame InterpolationabstractBlurry video frame interpolation (BVFI) aims to generate high-frame-rate clear videos from low-frame-rate blurry videos, is a challenging but important topic in the computer vision community. Blurry videos not only provide spatial and temporal information like clear videos, but also contain additional motion information hidden in each blurry frame. However, existing BVFI methods usually fail to fully leverage all valuable information, which ultimately hinders their performance. In this paper, we propose a simple three-stage temporal deformable network to fully explore useful information from blurry videos. The frame interpolation stage designs a deformable network to directly sample useful information from blurry inputs and synthesize an intermediate frame at an arbitrary time interval. The temporal feature fusion stage explores the long-term temporal information for each target frame through a bi-directional recurrent deformable alignment network. And the deblurring stage applies a transformer-empowered Taylor approximation network to recursively recover the high-frequency details. Quantitative and qualitative results indicate that our model outperforms existing SOTA methods. Pengcheng Lei, Zaoming Yan, Tingting Wang 0007, Faming Fang, Guixu Zhang |
ICME | 4 |
| 2024 | HFF-Net: A High-Frequency Fidelity Model for Accelerated Parallel MRI ReconstructionabstractMagnetic Resonance Imaging (MRI) plays a crucial role in diagnosing and treating various diseases. However, the long acquisition time of MRI scans often leads to patient discomfort and motion artifacts. Consequently, accelerating MRI speed is essential. Researchers have combined Deep Learning with Compressed Sensing and Parallel Imaging to advance MRI. However, many existing methods fail to effectively recover the fine details and structures in Magnetic Resonance images. To address these challenges, we propose a novel model for accelerated parallel MRI reconstruction. Our model incorporates a high-frequency fidelity method into the reconstruction process, explicitly emphasizing the recovery of high-frequency information. Additionally, we consider the joint priori distribution between the reconstructed images from each coil. Using the variable splitting approach, the proposed model is unrolled as an end-to-end network termed HFF-Net. Experimental results demonstrate that our method outperforms state-of-the-art techniques, yielding high-quality MR images with enhanced detail and fine structure recovery. Zhenggang Yang, Faming Fang, Qiaosi Yi, Guixu Zhang, Fang Li 0004 |
ICME | 2 |
| 2024 | Exploring Fixed Point in Image Editing: Theoretical Support and Convergence OptimizationabstractIn image editing, Denoising Diffusion Implicit Models (DDIM) inversion has become a widely adopted method and is extensively used in various image editing approaches. The core concept of DDIM inversion stems from the deterministic sampling technique of DDIM, which allows the DDIM process to be viewed as an Ordinary Differential Equation (ODE) process that is reversible. This enables the prediction of corresponding noise from a reference image, ensuring that the restored image from this noise remains consistent with the reference image. Image editing exploits this property by modifying the cross-attention between text and images to edit specific objects while preserving the remaining regions. However, in the DDIM inversion, using the $t-1$ time step to approximate the noise prediction at time step $t$ introduces errors between the restored image and the reference image. Recent approaches have modeled each step of the DDIM inversion process as finding a fixed-point problem of an implicit function. This approach significantly mitigates the error in the restored image but lacks theoretical support regarding the existence of such fixed points. Therefore, this paper focuses on the study of fixed points in DDIM inversion and provides theoretical support. Based on the obtained theoretical insights, we further optimize the loss function for the convergence of fixed points in the original DDIM inversion, improving the visual quality of the edited image. Finally, we extend the fixed-point based image editing to the application of unsupervised image dehazing, introducing a novel text-based approach for unsupervised dehazing. Chen Hang, Haoming Chen, Xuwei Fang, Vincent Xie, Faming Fang, Guixu Zhang |
NeurIPS | 6 |
| 2024 | Self-supervised medical slice interpolation network using controllable feature flow
Pengcheng Lei, Faming Fang, Tingting Wang 0007, Cong Liu 0011, Guixu Zhang |
Expert Syst. Appl. | 2 |
| 2024 | Deep Richardson-Lucy Deconvolution for Low-Light Image Deblurring
Liang Chen 0026, Jiawei Zhang 0002, Yunxuan Wei, Faming Fang, Jimmy S. J. Ren, Jinshan Pan |
Int. J. Comput. Vis. | 5 |
| 2024 | HFGN: High-Frequency residual Feature Guided Network for fast MRI reconstruction
Faming Fang, Le Hu, Qiaosi Yi, Tieyong Zeng, Guixu Zhang |
Pattern Recognit. | 1 |
| 2024 | UGNet: Uncertainty aware geometry enhanced networks for stereo matching
Zhengkai Qi, Junkang Zhang, Faming Fang, Tingting Wang 0007, Guixu Zhang |
Pattern Recognit. | 3 |
| 2024 | MSCSCformer: Multiscale Convolutional Sparse Coding-Based Transformer for PansharpeningabstractWith the increasing significance of high-quality, high-resolution multispectral images (HRMS) in various domains, pansharpening, which fuses low-resolution multispectral images (LRMS) with high-resolution panchromatic images (PAN), has gained considerable attention. However, current deep learning methods have limitations in capturing global long-range dependencies and incorporating spectral characteristics across different spectral bands of multispectral images (MS). Additionally, model-based approaches do not effectively utilize the multi-scale information between LRMS and HRMS data, limiting their further performance enhancement. To address these limitations, we propose a new observation model based on Multi-Scale Convolutional Sparse Coding (MS-CSC) and design a novel Multi-Scale Hybrid Spatial-spectral Transformer (MSHST) for the unfolding networks. The MS-CSC based observation model aims to fuse multi-scale information, while the MSHST incorporates spatial self-attention to capture global long-range dependencies and spectral self-attention to capture the inter-band correlation. Experimental results demonstrate the superiority of our method over other state-of-the-art approaches in both reduced-resolution and full-resolution evaluations. Ablation experiments further validate the effectiveness of the proposed multi-scale model and MSHST. Code is available at https://github.com/Eternityyx/MSCSCformer. Yongxu Ye, Tingting Wang 0007, Faming Fang, Guixu Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Joint Under-Sampling Pattern and Dual-Domain Reconstruction for Accelerating Multi-Contrast MRIabstractMulti-Contrast Magnetic Resonance Imaging (MCMRI) utilizes the short-time reference image to facilitate the reconstruction of the long-time target one, providing a new solution for fast MRI. Although various methods have been proposed, they still have certain limitations. 1) existing methods featuring the preset under-sampling patterns give rise to redundancy between multi-contrast images and limit their model performance; 2) most methods focus on the information in the image domain, prior knowledge in the k-space domain has not been fully explored; and 3) most networks are manually designed and lack certain physical interpretability. To address these issues, we propose a joint optimization of the under-sampling pattern and a deep-unfolding dual-domain network for accelerating MCMRI. Firstly, to reduce the redundant information and sample more contrast-specific information, we propose a new framework to learn the optimal under-sampling pattern for MCMRI. Secondly, a dual-domain model is established to reconstruct the target image in both the image domain and the k-space frequency domain. The model in the image domain introduces a spatial transformation to explicitly model the inconsistent and unaligned structures of MCMRI. The model in the k-space learns prior knowledge from the frequency domain, enabling the model to capture more global information from the input images. Thirdly, we employ the proximal gradient algorithm to optimize the proposed model and then unfold the iterative results into a deep-unfolding network, called MC-DuDoN. We evaluate the proposed MC-DuDoN on MCMRI super-resolution and reconstruction tasks. Experimental results give credence to the superiority of the current model. In particular, since our approach explicitly models the inconsistent structures, it shows robustness on spatially misaligned MCMRI. In the reconstruction task, compared with conventional masks, the learned mask restores more realistic images, even under an ultra-high acceleration ratio ( ×30 ). Code is available at https://github.com/lpcccc-cv/MC-DuDoNet. Pengcheng Lei, Le Hu, Faming Fang, Guixu Zhang |
IEEE Trans. Image Process. | 3 |
| 2024 | Flow Guidance Deformable Compensation Network for Video Frame InterpolationabstractFlow-based and deformable convolution (DConv)-based methods are two mainstream approaches for solving the video frame interpolation (VFI) problem, which have made remarkable progress with the development of deep convolutional networks over the past years. However, flow-based VFI methods often suffer from the inaccuracy of flow map estimation, especially in dealing with complex and irregular real-world motions. DConv-based VFI methods have advantages in handling complex motions, while the increased degree of freedom makes the training of the DConv model difficult. To address these problems, in this article, we propose a flow guidance deformable compensation network (FGDCN) for the VFI task. FGDCN decomposes the frame sampling process into two steps: a flow step and a deformation step. Specifically, the flow step utilizes a coarse-to-fine flow estimation network to directly estimate the intermediate flows and synthesizes an anchor frame simultaneously. To ensure the accuracy of the estimated flow, a distillation loss and a task-oriented loss are jointly employed in this step. Under the guidance of the flow priors learned in step one, the deformation step designs a new pyramid deformable compensation network to compensate for the missing details of the flow step. In addition, a pyramid loss is proposed to supervise the model in both the image and frequency domains. Experimental results show that the proposed algorithm achieves excellent performance on various datasets with fewer parameters. Pengcheng Lei, Faming Fang, Tieyong Zeng, Guixu Zhang |
IEEE Trans. Multim. | 2 |
| 2023 | Indoor Depth Recovery Based on Deep Unfolding with Non-Local PriorabstractIn recent years, depth recovery based on deep networks has achieved great success. However, the existing state-of-the-art network designs perform like black boxes in depth recovery tasks, lacking a clear mechanism. Utilizing the property that there is a large amount of non-local common characteristics in depth images, we propose a novel model-guided depth recovery method, namely the DC-NLAR model. A non-local auto-regressive regular term is also embedded into our model to capture more non-local depth information. To fully use the excellent performance of neural networks, we develop a deep image prior to better describe the characteristic of depth images. We also introduce an implicit data consistency term to tackle the degenerate operator with high heterogeneity. We then unfold the proposed model into networks by using the half-quadratic splitting algorithm. This proposed method is experimented on the NYU-Depth V2 and SUN RGB-D datasets, and the experimental results achieve comparable performance to that of deep learning methods. Yuhui Dai, Junkang Zhang, Faming Fang, Guixu Zhang |
ICCV | 3 |
| 2023 | Decomposition-Based Variational Network for Multi-Contrast MRI Super-Resolution and ReconstructionabstractMulti-contrast MRI super-resolution (SR) and reconstruction methods aim to explore complementary information from the reference image to help the reconstruction of the target image. Existing deep learning-based methods usually manually design fusion rules to aggregate the multi-contrast images, fail to model their correlations accurately and lack certain interpretations. Against these issues, we propose a multi-contrast variational network (MC-VarNet) to explicitly model the relationship of multi-contrast images. Our model is constructed based on an intuitive motivation that multi-contrast images have consistent (edges and structures) and inconsistent (contrast) information. We thus build a model to reconstruct the target image and decompose the reference image as a common component and a unique component. In the feature interaction phase, only the common component is transferred to the target image. We solve the variational model and unfold the iterative solutions into a deep network. Hence, the proposed method combines the good interpretability of model-based methods with the powerful representation ability of deep learning-based methods. Experimental results on the multi-contrast MRI reconstruction and SR demonstrate the effectiveness of the proposed model. Especially, since we explicitly model the multi-contrast images, our model is more robust to the reference images with noises and large inconsistent structures. The code is available at https://github.com/lpcccccv/MC-VarNet. Pengcheng Lei, Faming Fang, Guixu Zhang, Tieyong Zeng |
ICCV | 2 |
| 2023 | Deep Unfolding Convolutional Dictionary Model for Multi-Contrast MRI Super-resolution and ReconstructionabstractMagnetic resonance imaging (MRI) tasks often involve multiple contrasts. Recently, numerous deep learning-based multi-contrast MRI super-resolution (SR) and reconstruction methods have been proposed to explore the complementary information from the multi-contrast images. However, these methods either construct parameter-sharing networks or manually design fusion rules, failing to accurately model the correlations between multi-contrast images and lacking certain interpretations. In this paper, we propose a multi-contrast convolutional dictionary (MC-CDic) model under the guidance of the optimization algorithm with a well-designed data fidelity term. Specifically, we bulid an observation model for the multi-contrast MR images to explicitly model the multi-contrast images as common features and unique features. In this way, only the useful information in the reference image can be transferred to the target image, while the inconsistent information will be ignored. We employ the proximal gradient algorithm to optimize the model and unroll the iterative steps into a deep CDic model. Especially, the proximal operators are replaced by learnable ResNet. In addition, multi-scale dictionaries are introduced to further improve the model performance. We test our MC-CDic model on multi-contrast MRI SR and reconstruction tasks. Experimental results demonstrate the superior performance of the proposed MC-CDic model against existing SOTA methods. Code is available at https://github.com/lpcccc-cv/MC-CDic. Pengcheng Lei, Faming Fang, Guixu Zhang, Ming Xu 0010 |
IJCAI | 2 |
| 2023 | Deep Algorithm Unrolling with Registration Embedding for PansharpeningabstractPansharpening aims to sharpen low resolution (LR) multispectral (MS) images with the help of corresponding high resolution (HR) panchromatic (PAN) images to obtain HRMS images. Model-based pansharpening methods manually design objective functions via observation model and hand-crafted priors. However, inevitable performance degradation may occur in the case that the prior is invalid. Although many deep learning based end-to-end pansharpening methods have been proposed recently, they still need to be improved due to the insufficient study on HRMS related domain knowledge. Besides, existing pansharpening methods rarely consider the misalignments between MS and PAN images, leading to poor performance. To tackle these issues, this paper proposes to unrolling the observation model with registration embedding for pansharpening. Inspired by the optical flow estimation, we embed the registration operation into the observation model to reconstruct the pansharpening function with the help of a deep prior of HRMS images, and then unroll the iterative solution into a novel deep convolutional network.. Apart from the single HRMS supervision, we also introduce a consistency loss to supervise the two degradation processes. The use of consistency loss enables the degradation sub-networks to learn more realistic degradation. Experimental results at reduced-resolution and full-resolution are reported to demonstrate the superiority of the proposed method to other state-of-the-art pansharpening methods. In GaoFen-2 dataset evaluation, our method achieves 1.2dB higher PSNR than SOTA techniques. Tingting Wang 0007, Yongxu Ye, Faming Fang, Guixu Zhang, Ming Xu 0010 |
ACM Multimedia | 3 |
| 2023 | Single image noise level estimation by artificial noise
Fang Li 0004, Faming Fang, Zhi Li 0080, Tieyong Zeng |
Signal Process. | 2 |
| 2023 | FrMLNet: Framelet-Based Multilevel Network for PansharpeningabstractMost modern satellites can provide two types of images: 1) panchromatic (PAN) image and 2) multispectral (MS) image. The former has high spatial resolution and low spectral resolution, while the latter has high spectral resolution and low spatial resolution. To obtain images with both high spectral and spatial resolution, pansharpening has emerged to fuse the spatial information of the PAN image and the spectral information of the MS image. However, most pansharpening methods fail to preserve spatial and spectral information simultaneously. In this article, we propose a framelet-based convolutional neural network (CNN) for pansharpening which makes it possible to pursue both high spectral and high spatial resolution. Our network consists of three subnetworks: 1) feature embedding net; 2) feature fusion net; and 3) framelet prediction net. Different from conventional CNN methods directly inferring high-resolution MS images, our approach learns to predict their framelet coefficients from available PAN and MS images. The introduction of multilevel feature aggregation and hybrid residual connection makes full use of spatial information of PAN image and spectral information of MS image. Quantitative and qualitative experiments at reduced- and full-resolution demonstrate that the proposed method achieves more appealing results than other state-of-the-art pansharpening methods. The source code and trained models are available at https://github.com/TingMAC/FrMLNet. Tingting Wang 0007, Faming Fang, Guixu Zhang |
IEEE Trans. Cybern. | 2 |
| 2023 | DMCSC: Deep Multisource Convolutional Sparse Coding Model for PansharpeningabstractPansharpening aims to produce a high-resolution multispectral (HRMS) image by combining a low-resolution multispectral (LRMS) image with a high-resolution panchromatic (PAN) image through a fusion process. Deep learning (DL)-based pansharpening methods have demonstrated impressive results in generating high-quality HRMS images. However, they suffer from a lack of interpretability due to their black-box network architectures. Recently, model-based deep unrolling networks have been proposed to improve the interpretability of networks. Among these approaches, the multi-source convolutional sparse coding (MCSC)-based models stand out by effectively learning common and unique features from both LRMS and PAN images, showing promising results. As the LRMS image provides limited information in MCSC-based models, it can result in weak feature response and even lead to incorrect fusion outcomes. To address this issue, we propose a novel deep MCSC-based method that enhances the robustness and performance. Specifically, we build an optimization model that integrates MCSC with a degradation model and a deep prior, which can sufficiently capture the common information shared by the latent HRMS images and PAN images, thereby enabling the recovery of more accurate spectral information. To optimize the proposed model, we adopt an iterative optimization strategy that unfolds the iterative solution into networks. Moreover, we propose an enhanced version of our method that utilizes multi-scale dictionaries to capture common and unique features at different scales, thereby facilitating the extraction of more abundant spectral and spatial details. We evaluate the effectiveness of our proposed method on multiple benchmark datasets. Experiment results demonstrate its effectiveness in improving the robustness and performance of MCSC-based models. Junkang Zhang, Yongxu Ye, Faming Fang, Tingting Wang 0007, Guixu Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Frequency Learning via Multi-Scale Fourier Transformer for MRI ReconstructionabstractSince Magnetic Resonance Imaging (MRI) requires a long acquisition time, various methods were proposed to reduce the time, but they ignored the frequency information and non-local similarity, so that they failed to reconstruct images with a clear structure. In this article, we propose Frequency Learning via Multi-scale Fourier Transformer for MRI Reconstruction (FMTNet), which focuses on repairing the low-frequency and high-frequency information. Specifically, FMTNet is composed of a high-frequency learning branch (HFLB) and a low-frequency learning branch (LFLB). Meanwhile, we propose a Multi-scale Fourier Transformer (MFT) as the basic module to learn the non-local information. Unlike normal Transformers, MFT adopts Fourier convolution to replace self-attention to efficiently learn global information. Moreover, we further introduce a multi-scale learning and cross-scale linear fusion strategy in MFT to interact information between features of different scales and strengthen the representation of features. Compared with normal Transformers, the proposed MFT occupies fewer computing resources. Based on MFT, we design a Residual Multi-scale Fourier Transformer module as the main component of HFLB and LFLB. We conduct several experiments under different acceleration rates and different sampling patterns on different datasets, and the experiment results show that our method is superior to the previous state-of-the-art method. Qiaosi Yi, Faming Fang, Guixu Zhang, Tieyong Zeng |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Adjustable super-resolution network via deep supervised learning and progressive self-distillation
Juncheng Li 0003, Faming Fang, Tieyong Zeng, Guixu Zhang, Xizhao Wang |
Neurocomputing | 2 |
| 2022 | Patch-based weighted SCAD prior for compressive sensing
Yamin Ru, Fang Li 0004, Faming Fang, Guixu Zhang |
Inf. Sci. | 3 |
| 2022 | Multi-Scale Grid Network for Image Deblurring With High-Frequency GuidanceabstractIt has been demonstrated that the blurring process reduces the high-frequency information of the original sharp image, so the main challenge for image deblurring is to reconstruct high-frequency information from the blurry image. In this paper, we propose a novel image deblurring framework to focus on the reconstruction of high-frequency information, which consists of two main subnetworks: a high-frequency reconstruction subnetwork (HFRSN) and a multi-scale grid subnetwork (MSGSN). The HFRSN is built to reconstruct latent high-frequency information from multiple scale blurry images. The MSGSN performs deblurring processes with high-frequency guidance at different scales simultaneously. Besides, in order to better use high-frequency information to restore sharpening images, we designed a high-frequency information aggregation (HFAG) module and a high-frequency information attention (HFAT) module in MSGSN. The HFAG module is designed to fuse high-frequency features and image features at the feature extraction stage, and the HFAT module is built to enhance the feature reconstruction stage. Extensive experiments on different datasets show the effectiveness and efficiency of our method. Yang Liu 0289, Faming Fang, Tingting Wang 0007, Juncheng Li 0003, Yun Sheng, Guixu Zhang |
IEEE Trans. Multim. | 2 |
| 2022 | Efficient and Accurate Multi-Scale Topological Network for Single Image DehazingabstractSingle image dehazing is a challenging ill-posed problem that has drawn significant attention in the last few years. Recently, convolutional neural networks have achieved great success in image dehazing. However, it is still difficult for these increasingly complex models to recover accurate details from the hazy image. In this paper, we pay attention to the feature extraction and utilization of the input image itself. To achieve this, we propose a Multi-scale Topological Network (MSTN) to fully explore the features at different scales. Meanwhile, we design a Multi-scale Feature Fusion Module (MFFM) and an Adaptive Feature Selection Module (AFSM) to achieve the selection and fusion of features at different scales, so as to achieve progressive image dehazing. This topological network provides a large number of search paths that enable the network to extract abundant image features as well as strong fault tolerance and robustness. In addition, ASFM and MFFM can adaptively select important features and ignore interference information when fusing different scale representations. Extensive experiments are conducted to demonstrate the superiority of our method compared with state-of-the-art methods. Qiaosi Yi, Juncheng Li 0003, Faming Fang, Aiwen Jiang, Guixu Zhang |
IEEE Trans. Multim. | 3 |
| 2021 | Blind Deblurring for Saturated ImagesabstractBlind deblurring has received considerable attention in recent years. However, state-of-the-art methods often fail to process saturated blurry images. The main reason is that pixels around saturated regions are not conforming to the commonly used linear blur model. Pioneer arts suggest excluding these pixels during the deblurring process, which sometimes simultaneously removes the informative edges around saturated regions and results in insufficient information for kernel estimation when large saturated regions exist. To address this problem, we introduce a new blur model to fit both saturated and unsaturated pixels, and all informative pixels can be considered during the deblurring process. Based on our model, we develop an effective maximum a posterior (MAP)-based optimization framework. Quantitative and qualitative evaluations on benchmark datasets and challenging real-world examples show that the proposed method performs favorably against existing methods. Liang Chen 0026, Jiawei Zhang 0002, Songnan Lin, Faming Fang, Jimmy S. J. Ren |
CVPR | 4 |
| 2021 | Learning a Non-Blind Deblurring Network for Night Blurry ImagesabstractDeblurring night blurry images is difficult, because the common-used blur model based on the linear convolution operation does not hold in this situation due to the influence of saturated pixels. In this paper, we propose a non-blind deblurring network (NBDN) to restore night blurry images. To mitigate the side effects brought by the pixels that violate the blur model, we develop a confidence estimation unit (CEU) to estimate a map which ensures smaller contributions of these pixels in the deconvolution steps which are optimized by the conjugate gradient (CG) method. Moreover, unlike the existing methods using manually tuned hyper-parameters in their frameworks, we propose a hyper-parameter estimation unit (HPEU) to adaptively estimate hyper-parameters for better image restoration. The experimental results demonstrate that the proposed network performs favorably against state-of-the-art algorithms both quantitatively and qualitatively. Liang Chen 0026, Jiawei Zhang 0002, Jinshan Pan, Songnan Lin, Faming Fang, Jimmy S. J. Ren |
CVPR | 5 |
| 2021 | Structure-Preserving Deraining with Residue Channel Prior GuidanceabstractSingle image deraining is important for many high-level computer vision tasks since the rain streaks can severely degrade the visibility of images, thereby affecting the recognition and analysis of the image. Recently, many CNN-based methods have been proposed for rain removal. Although these methods can remove part of the rain streaks, it is difficult for them to adapt to real-world scenarios and restore high-quality rain-free images with clear and accurate structures. To solve this problem, we propose a Structure-Preserving Deraining Network (SPDNet) with RCP guidance. SPDNet directly generates high-quality rain-free images with clear and accurate structures under the guidance of RCP but does not rely on any rain-generating assumptions. Specifically, we found that the RCP of images contains more accurate structural information than rainy images. Therefore, we introduced it to our deraining network to protect structure information of the rain-free image. Meanwhile, a Wavelet-based Multi-Level Module (WMLM) is proposed as the backbone for learning the background information of rainy images and an Interactive Fusion Module (IFM) is designed to make full use of RCP information. In addition, an iterative guidance strategy is proposed to gradually improve the accuracy of RCP, refining the result in a progressive path. Extensive experimental results on both synthetic and real-world datasets demonstrate that the proposed model achieves new state-of-the-art results. Code: https://github.com/Joyies/SPDNet Qiaosi Yi, Juncheng Li 0003, Qinyan Dai, Faming Fang, Guixu Zhang, Tieyong Zeng |
ICCV | 4 |
| 2021 | Feedback Network for Mutually Boosted Stereo Image Super-Resolution and Disparity EstimationabstractUnder stereo settings, the problem of image super-resolution (SR) and disparity estimation are interrelated that the result of each problem could help to solve the other. The effective exploitation of correspondence between different views facilitates the SR performance, while the high-resolution (HR) features with richer details benefit the correspondence estimation. According to this motivation, we propose a Stereo Super-Resolution and Disparity Estimation Feedback Network (SSRDE-FNet), which simultaneously handles the stereo image super-resolution and disparity estimation in a unified framework and interact them with each other to further improve their performance. Specifically, the SSRDE-FNet is composed of two dual recursive sub-networks for left and right views. Besides the cross-view information exploitation in the low-resolution (LR) space, HR representations produced by the SR process are utilized to perform HR disparity estimation with higher accuracy, through which the HR features can be aggregated to generate a finer SR result. Afterward, the proposed HR Disparity Information Feedback (HRDIF) mechanism delivers information carried by HR disparity back to previous layers to further refine the SR image reconstruction. Extensive experiments demonstrate the effectiveness and advancement of SSRDE-FNet. Qinyan Dai, Juncheng Li 0003, Qiaosi Yi, Faming Fang, Guixu Zhang |
ACM Multimedia | 4 |
| 2021 | Edge-guided Composition Network for Image Stitching
Qinyan Dai, Faming Fang, Juncheng Li 0003, Guixu Zhang, Aimin Zhou |
Pattern Recognit. | 2 |
| 2021 | MDCN: Multi-Scale Dense Cross Network for Image Super-ResolutionabstractConvolutional neural networks have been proven to be of great benefit for single-image super-resolution (SISR). However, previous works do not make full use of multi-scale features and ignore the inter-scale correlation between different upsampling factors, resulting in sub-optimal performance. Instead of blindly increasing the depth of the network, we are committed to mining image features and learning the inter-scale correlation between different upsampling factors. To achieve this, we propose a Multi-scale Dense Cross Network (MDCN), which achieves great performance with fewer parameters and less execution time. MDCN consists of multi-scale dense cross blocks (MDCBs), hierarchical feature distillation block (HFDB), and dynamic reconstruction block (DRB). Among them, MDCB aims to detect multi-scale features and maximize the use of image features flow at different scales, HFDB focuses on adaptively recalibrate channel-wise feature responses to achieve feature distillation, and DRB attempts to reconstruct SR images with different upsampling factors in a single model. It is worth noting that all these modules can run independently. It means that these modules can be selectively plugged into any CNN model to improve model performance. Extensive experiments show that MDCN achieves competitive results in SISR, especially in the reconstruction task with multiple upsampling factors. The code is provided athttps://github.com/MIVRC/MDCN-PyTorch. Juncheng Li 0003, Faming Fang, Kangfu Mei, Guixu Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Superpixel-Based Seamless Image Stitching for UAV ImagesabstractImage stitching aims to generate a natural seamless high-resolution panoramic image free of distortions or artifacts as fast as possible. In this article, we propose a new seam cutting strategy based on superpixels for unmanned aerial vehicle (UAV) image stitching. Explicitly, we decompose the issue into three steps: image registration, seam cutting, and image blending. First, we employ adaptive as-natural-as-possible (AANAP) warps for registration, obtaining two aligned images in the same coordinate system. Then, we propose a novel superpixel-based energy function that integrates color difference, gradient difference, and texture complexity information to search a perceptually optimal seam located in continuous areas with high similarity. We apply the graph cut algorithm to solve the problem and thereby conceal artifacts in the overlapping area. Finally, we utilize a superpixel-based color blending approach to eliminate visible seams and achieve natural color transitions. Experimental results demonstrate that our method can effectively and efficiently realize seamless stitching, and is superior to several state-of-the-art methods in UAV image stitching. Yiting Yuan, Faming Fang, Guixu Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Luminance-Aware Pyramid Network for Low-Light Image EnhancementabstractLow-light image enhancement based on deep convolutional neural networks (CNNs) has revealed prominent performance in recent years. However, it is still a challenging task since the underexposed regions and details are always imperceptible. Moreover, deep learning models are always accompanied by complex structures and enormous computational burden, which hinders their deployment on mobile devices. To remedy these issues, in this paper, we present a lightweight and efficient Luminance-aware Pyramid Network (LPNet) to reconstruct normal-light images in a coarse-to-fine strategy. The architecture is comprised of two coarse feature extraction branches and a luminance-aware refinement branch with an auxiliary subnet learning the luminance map of the input and target images. Besides, we propose a multi-scale contrast feature block (MSCFB) that involves channel split, channel shuffle strategies, and contrast attention mechanism. MSCFB is the essential component of our network, which achieves an excellent balance between image quality and model size. In this way, our method can not only brighten up low-light images with rich details and high contrast but also significantly ameliorate the execution speed. Extensive experiments demonstrate that our LPNet outperforms state-of-the-art methods both qualitatively and quantitatively. Juncheng Li 0003, Faming Fang, Fang Li 0004, Guixu Zhang |
IEEE Trans. Multim. | 3 |
| 2021 | Multilevel Edge Features Guided Network for Image DenoisingabstractImage denoising is a challenging inverse problem due to complex scenes and information loss. Recently, various methods have been considered to solve this problem by building a well-designed convolutional neural network (CNN) or introducing some hand-designed image priors. Different from previous works, we investigate a new framework for image denoising, which integrates edge detection, edge guidance, and image denoising into an end-to-end CNN model. To achieve this goal, we propose a multilevel edge features guided network (MLEFGN). First, we build an edge reconstruction network (Edge-Net) to directly predict clear edges from the noisy image. Then, the Edge-Net is embedded as part of the model to provide edge priors, and a dual-path network is applied to extract the image and edge features, respectively. Finally, we introduce a multilevel edge features guidance mechanism for image denoising. To the best of our knowledge, the Edge-Net is the first CNN model specially designed to reconstruct image edges from the noisy image, which shows good accuracy and robustness on natural images. Extensive experiments clearly illustrate that our MLEFGN achieves favorable performance against other methods and plenty of ablation studies demonstrate the effectiveness of our proposed Edge-Net and MLEFGN. The code is available at https://github.com/MIVRC/MLEFGN-PyTorch. Faming Fang, Juncheng Li 0003, Yiting Yuan, Tieyong Zeng, Guixu Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Stylization-Based Architecture for Fast Deep Exemplar ColorizationabstractExemplar-based colorization aims to add colors to a grayscale image guided by a content related reference image. Existing methods are either sensitive to the selection of reference images (content, position) or extremely time and resource consuming, which limits their practical application. To tackle these problems, we propose a deep exemplar colorization architecture inspired by the characteristics of stylization in feature extracting and blending. Our coarse- to-fine architecture consists of two parts: a fast transfer sub-net and a robust colorization sub-net. The transfer sub- net obtains a coarse chrominance map via matching basic feature statistics of the input pairs in a progressive way. The colorization sub-net refines the map to generate the final results. The proposed end-to-end network can jointly learn faithful colorization with a related reference and plausible color prediction with unrelated reference. Extensive experimental validation demonstrates that our approach outperforms the state-of-the-art methods in less time whether in exemplar-based colorization or image stylization tasks. Zhongyou Xu, Tingting Wang 0007, Faming Fang, Yun Sheng, Guixu Zhang |
CVPR | 3 |
| 2020 | Enhanced Sparse Model for Blind Deblurring
Faming Fang, Fang Li 0004, Guixu Zhang |
ECCV (25) | 2 |
| 2020 | OID: Outlier Identifying and Discarding in Blind Image Deblurring
Faming Fang, Guixu Zhang |
ECCV (25) | 2 |
| 2020 | Removing moiré patterns from single images
Faming Fang, Tingting Wang 0007, Shuyan Wu, Guixu Zhang |
Inf. Sci. | 1 |
| 2020 | Soft-Edge Assisted Network for Single Image Super-ResolutionabstractThe task of single image super-resolution (SISR) is a highly ill-posed inverse problem since reconstructing the highfrequency details from a low-resolution image is challenging. Most previous CNN-based super-resolution (SR) methods tend to directly learn the mapping from the low-resolution image to the high-resolution image through some complex convolutional neural networks. However, the method of blindly increasing the depth of the network is not the best choice because the performance improvement of such methods is marginal but the computational cost is huge. A more efficient method is to integrate the image prior knowledge into the model to assist the image reconstruction. Indeed, the soft-edge has been widely applied in many computer vision tasks as the role of an important image feature. In this paper, we propose a Soft-edge assisted Network (SeaNet) to reconstruct the high-quality SR image with the help of image soft-edge. The proposed SeaNet consists of three sub-nets: a rough image reconstruction network (RIRN), a soft-edge reconstruction network (Edge-Net), and an image refinement network (IRN). The complete reconstruction process consists of two stages. In Stage-I, the rough SR feature maps and the SR soft-edge are reconstructed by the RIRN and Edge-Net, respectively. In Stage-II, the outputs of the previous stages are fused and then feed to the IRN for high-quality SR image reconstruction. Extensive experiments show that our SeaNet converges rapidly and achieves excellent performance under the assistance of image soft-edge. The code is available at https://gitlab.com/junchenglee/seanet-pytorch. Faming Fang, Juncheng Li 0003, Tieyong Zeng |
IEEE Trans. Image Process. | 1 |
| 2020 | A Novel Retinex-Based Fractional-Order Variational Model for Images With Severely Low LightabstractIn this paper, we propose a novel Retinex-based fractional-order variational model for severely low-light images. The proposed method is more flexible in controlling the regularization extent than the existing integer-order regularization methods. Specifically, we decompose directly in the image domain and perform the fractional-order gradient total variation regularization on both the reflectance component and the illumination component to get more appropriate estimated results. The merits of the proposed method are as follows: 1) small-magnitude details are maintained in the estimated reflectance. 2) illumination components are effectively removed from the estimated reflectance. 3) the estimated illumination is more likely piecewise smooth. We compare the proposed method with other closely related Retinex-based methods. Experimental results demonstrate the effectiveness of the proposed method. Fang Li 0004, Faming Fang, Guixu Zhang |
IEEE Trans. Image Process. | 3 |
| 2020 | Variational Single Image Dehazing for Enhanced VisualizationabstractIn this paper, we investigate the challenging task of removing haze from a single natural image. The analysis on the haze formation model shows that the atmospheric veil has much less relevance to chrominance than luminance, which motivates us to neglect the haze in the chrominance channel and concentrate on the luminance channel in the dehazing process. Besides, the experimental study illustrates that the YUV color space is most suitable for image dehazing. Accordingly, a variational model is proposed in the Y channel of the YUV color space by combining the reformulation of the haze model and the two effective priors. As we mainly focus on the Y channel, most of the chrominance information of the image is preserved after dehazing. The numerical procedure based on the alternating direction method of multipliers (ADMM) scheme is presented to obtain the optimal solution. Extensive experimental results on real-world hazy images and synthetic dataset demonstrate clearly that our method can unveil the details and recover vivid color information, which is competitive among many existing dehazing algorithms. Further experiments show that our model also can be applied for image enhancement. Faming Fang, Tingting Wang 0007, Yang Wang 0020, Tieyong Zeng, Guixu Zhang |
IEEE Trans. Multim. | 1 |
| 2020 | A Superpixel-Based Variational Model for Image ColorizationabstractImage colorization refers to a computer-assisted process that adds colors to grayscale images. It is a challenging task since there is usually no one-to-one correspondence between color and local texture. In this paper, we tackle this issue by exploiting weighted nonlocal self-similarity and local consistency constraints at the resolution of superpixels. Given a grayscale target image, we first select a color source image containing similar segments to target image and extract multi-level features of each superpixel in both images after superpixel segmentation. Then a set of color candidates for each target superpixel is selected by adopting a top-down feature matching scheme with confidence assignment. Finally, we propose a variational approach to determine the most appropriate color for each target superpixel from color candidates. Experiments demonstrate the effectiveness of the proposed method and show its superiority to other state-of-the-art methods. Furthermore, our method can be easily extended to color transfer between two color images. Faming Fang, Tingting Wang 0007, Tieyong Zeng, Guixu Zhang |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Blind Image Deblurring With Local Maximum Gradient PriorabstractBlind image deblurring aims to recover sharp image from a blurred one while the blur kernel is unknown. To solve this ill-posed problem, a great amount of image priors have been explored and employed in this area. In this paper, we present a blind deblurring method based on Local Maximum Gradient (LMG) prior. Our work is inspired by the simple and intuitive observation that the maximum value of a local patch gradient will diminish after the blur process, which is proved to be true both mathematically and empirically. This inherent property of blur process helps us to establish a new energy function. By introducing an liner operator to compute the Local Maximum Gradient, together with an effective optimization scheme, our method can handle various specific scenarios. Extensive experimental results illustrate that our method is able to achieve favorable performance against state-of-the-art algorithms on both synthetic and real-world images. Faming Fang, Tingting Wang 0007, Guixu Zhang |
CVPR | 2 |
| 2019 | Bregman-Tanimoto Based Method for Contrast Preserving DecolorizationabstractThis paper presents a novel Bregman-Tanimoto based model for faithfully contrast-preserving decolorization. With regard to the defects of the traditional Euclidean metric approaches, the proposed method addresses the problems by introducing a Bregman-Tanimoto metric based maximum function, thus more details, features and visual distinctiveness can be preserved. Moreover, the model extends the solution space of the first-order linear parameter model used for color-to-gray mapping. An extended discrete searching algorithm is developed to solve the proposed model. Experimental results show that the proposed approach outperforms other state-of-the-art methods both quantitatively and qualitatively. Faming Fang |
ICME | 2 |
| 2019 | Cascaded Dilated Dense Network with Two-step Data Consistency for MRI ReconstructionabstractCompressed Sensing MRI (CS-MRI) aims at reconstrcuting de-aliased images from sub-Nyquist sampling k-space data to accelerate MR Imaging. Inspired by recent deep learning methods, we propose a Cascaded Dilated Dense Network (CDDN) for MRI reconstruction. Dense blocks with residual connection are used to restore clear images step by step and dilated convolution is introduced for expanding receptive field without taking more network parameters. After each sub-network, we use a novel two-step Data Consistency (DC) operation in k-space. We convert the complex result from first DC operation to real-valued images and applied another sampled \emph{k}-space data replacement. Extensive experiments demonstrate that the proposed CDDN with two-step DC achieves state-of-art result. Faming Fang, Guixu Zhang |
NeurIPS | 2 |
| 2019 | Fast Color Blending for Seamless Image StitchingabstractIn this letter, we propose a fast and robust method for stitching overlapped images captured by the unmanned aerial vehicle. First, we apply the shape-preserving half-projective method to precisely and stably align a pair of partially overlapped input images. Then, an optimal stitching line is searched to remove ghosts caused by the moving objects in the overlapped area. We subsequently propose a color blending method to eliminate all the color inconsistencies in the prealigned image. In accordance with the color differences of the pixels on the optimal stitching seam, we utilize weighted value coordinate interpolation algorithms to compute accurate color changes for all the pixels in the target image. The calculated color changes are then added to the target image to remove the color inconsistency. Furthermore, we introduce the superpixel segmentation to divide the target image into a reduced number of superpixels, and we assign each superpixel the same color change value. Such a superpixel level operation can greatly reduce the computational complexity. Experiments show that our method is promising to achieve effective and efficient stitching results. Faming Fang, Tingting Wang 0007, Yingying Fang, Guixu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | High-Quality Bayesian PansharpeningabstractPansharpening is a process of acquiring a multi-spectral image with high spatial resolution by fusing a low resolution multi-spectral image with a corresponding high resolution panchromatic image. In this paper, a new pansharpening method based on the Bayesian theory is proposed. The algorithm is mainly based on three assumptions: 1) the geometric information contained in the pan-sharpened image is coincident with that contained in the panchromatic image; 2) the pan-sharpened image and the original multi-spectral image should share the same spectral information; and 3) in each pan-sharpened image channel, the neighboring pixels not around the edges are similar. We build our posterior probability model according to above-mentioned assumptions and solve it by the alternating direction method of multipliers. The experiments at reduced and full resolution show that the proposed method outperforms the other state-of-the-art pansharpening methods. Besides, we verify that the new algorithm is effective in preserving spectral and spatial information with high reliability. Further experiments also show that the proposed method can be successfully extended to hyper-spectral image fusion. Tingting Wang 0007, Faming Fang, Fang Li 0004, Guixu Zhang |
IEEE Trans. Image Process. | 2 |
| 2018 | Topic Detection with Danmaku: A Time-Sync Joint NMF Approach
Qingchun Bai, Qinmin Hu, Faming Fang, Liang He 0001 |
DEXA (2) | 3 |
| 2018 | Multi-scale Residual Network for Image Super-Resolution
Juncheng Li 0003, Faming Fang, Kangfu Mei, Guixu Zhang |
ECCV (8) | 2 |
| 2016 | Single Image Dehazing Using Hölder Coefficient
Dehao Shang, Tingting Wang 0007, Faming Fang |
KSEM | 3 |
| 2016 | Single Image Super-Resolution Based on Nonlocal Sparse and Low-Rank Regularization
Chunhong Liu, Faming Fang, Chaomin Shen 0001 |
PRICAI | 2 |
| 2016 | Framelet-Based Sparse Unmixing of Hyperspectral ImagesabstractSpectral unmixing aims at estimating the proportions (abundances) of pure spectrums (endmembers) in each mixed pixel of hyperspectral data. Recently, a semi-supervised approach, which takes the spectral library as prior knowledge, has been attracting much attention in unmixing. In this paper, we propose a new semi-supervised unmixing model, termed framelet-based sparse unmixing (FSU), which promotes the abundance sparsity in framelet domain and discriminates the approximation and detail components of hyperspectral data after framelet decomposition. Due to the advantages of the framelet representations, e.g., images have good sparse approximations in framelet domain, and most of the additive noises are included in the detail coefficients, the FSU model has a better antinoise capability, and accordingly leads to more desirable unmixing performance. The existence and uniqueness of the minimizer of the FSU model are then discussed, and the split Bregman algorithm and its convergence property are presented to obtain the minimal solution. Experimental results on both simulated data and real data demonstrate that the FSU model generally performs better than the compared methods. Guixu Zhang, Faming Fang |
IEEE Trans. Image Process. | 3 |
| 2015 | Variational approach for multi-source image fusionabstractIn this study, the authors propose a variational model for image fusion using a gradient field to describe the features of all input images. The authors’ model is based on energy minimisation and the fused image corresponds to the minimiser of the energy functional. The authors first construct the gradient of fused image by using a weighted sum of the input gradients. Next, to increase the contrast in the fused image, the authors subtract the norm of gradient in the fused image from the functional. Finally, for the purpose of visual uniformity, the authors integrate the inputs using a ‘gray world’ assumption. The authors implement the algorithm using the augmented Lagrangian method. Three sets of images are used to verify the proposed method. Comparisons with other state‐of‐the‐art algorithms show that the proposed algorithm obtains remarkable results. Sizhang Tang, Faming Fang, Guixu Zhang |
IET Image Process. | 2 |
| 2015 | An Antinoise Method for Hyperspectral UnmixingabstractIn this letter, we propose an antinoise method for hyperspectral unmixing. In the antinoise method, all noises are addressed. The following techniques are applied: 1) an endmember dictionary is constructed first to initialize the solution; 2) an approximated L0norm constraint is employed to prune the dictionary and fulfill the sparse coding; and 3) the Itakura-Saito divergence, instead of the Square of Euclidean Distance divergence, is utilized to construct a novel optimization function. The experimental results on both synthetic and real hyperspectral data sets demonstrate the efficacy of the proposed method. Aimin Zhou, Guixu Zhang, Faming Fang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2015 | Similarity-Guided and ℓp-Regularized Sparse Unmixing of Hyperspectral DataabstractIn this letter, we propose a novel sparse unmixing model combined with two effective regularization terms: one is a similarity-weighting constraint, and the other is the ℓp(0p-norm, it has numerical advantages over the convex ℓ1-norm and better approximates the ℓ0-norm theoretically. Moreover, the ℓp-norm regularizer can simultaneously promote sparsity and enforce the abundance sum-to-one constraint. Therefore, this term yields more desirable results in practice. Experimental results on both simulated and real data demonstrate the effectiveness of the proposed model. Faming Fang, Guixu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | Framelet based pan-sharpening via a variational method
Faming Fang, Guixu Zhang, Fang Li 0004, Chaomin Shen 0001 |
Neurocomputing | 1 |
| 2014 | A Novel Blind Spectral Unmixing Method Based on Error Analysis of Linear Mixture ModelabstractIt is well known that the linear mixture model (LMM) is attracting much attention due to its simplicity. However, some theoretical analysis reveals that the traditional LMM also impedes the improvement of blind spectral unmixing. For this reason, we propose a novel blind spectral unmixing method (NBSUM) in this letter. NBSUM utilizes the conjugate gradient to calculate end-member spectral and abundance, which can not only overcome some shortcomings of the traditional LMM but also provide more accurate results. NBSUM is compared with some state-of-the-art approaches on both synthetic and real hyperspectral data sets, and the experimental results demonstrate the efficacy of the proposed method. Faming Fang, Aimin Zhou, Guixu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | Single Image Dehazing and Denoising: A Fast Variational ApproachabstractIn this paper, we propose a new fast variational approach to dehaze and denoise simultaneously. The proposed method first estimates a transmission map using a windows adaptive method based on the celebrated dark channel prior. This transmission map can significantly reduce the edge artifact in the resulting image and enhance the estimation precision. The transmission map is then converted to a depth map, with which the new variational model can be built to seek the final haze- and noise-free image. The existence and uniqueness of a minimizer of the proposed variational model is further discussed. A numerical procedure based on the Chambolle--Pock algorithm is given, and the convergence of the algorithm is ensured. Extensive experimental results on real scenes demonstrate that our method can restore vivid and contrastive haze- and noise-free images effectively. Faming Fang, Fang Li 0004, Tieyong Zeng |
SIAM J. Imaging Sci. | 1 |
| 2013 | A Variational Approach for Pan-SharpeningabstractPan-sharpening is a process of acquiring a high resolution multispectral (MS) image by combining a low resolution MS image with a corresponding high resolution panchromatic (PAN) image. In this paper, we propose a new variational pan-sharpening method based on three basic assumptions: 1) the gradient of PAN image could be a linear combination of those of the pan-sharpened image bands; 2) the upsampled low resolution MS image could be a degraded form of the pan-sharpened image; and 3) the gradient in the spectrum direction of pan-sharpened image should be approximated to those of the upsampled low resolution MS image. An energy functional, whose minimizer is related to the best pan-sharpened result, is built based on these assumptions. We discuss the existence of minimizer of our energy and describe the numerical procedure based on the split Bregman algorithm. To verify the effectiveness of our method, we qualitatively and quantitatively compare it with some state-of-the-art schemes using QuickBird and IKONOS data. Particularly, we classify the existing quantitative measures into four categories and choose two representatives in each category for more reasonable quantitative evaluation. The results demonstrate the effectiveness and stability of our method in terms of the related evaluation benchmarks. Besides, the computation efficiency comparison with other variational methods also shows that our method is remarkable. Faming Fang, Fang Li 0004, Chaomin Shen 0001, Guixu Zhang |
IEEE Trans. Image Process. | 1 |
| 2012 | Qualitative analysis and application of locally coupled neural oscillator network
Yuanhua Qiao, Yong Meng, Lijuan Duan, Faming Fang |
Neural Comput. Appl. | 4 |
| 2010 | Visual Selection and Attention Shifting Based on FitzHugh-Nagumo Equations
Yuanhua Qiao, Lijuan Duan, Faming Fang, Bingpeng Ma |
ISNN (2) | 4 |