EDBT 2026 Demo / reviewers in the wild / expert
Jinshan Pan
dblp:06/10816 · also Jin-shan Pan
· DBLP profile ↗
151ranked-venue papers
26as first author
95since 2021 · last 2026
0000-0003-0304-9507ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 108 · 17 first-author · 61 since 2021Artificial intelligence and machine learning · 107 · 22 first-author · 69 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neural Discrimination-Prompted Transformers for Efficient UHD Image Restoration and Enhancement
Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Yang Yang 0009 |
Int. J. Comput. Vis. | 2 |
| 2026 | Scene Prior Filtering for Depth Super-Resolution
Zhengxue Wang, Zhiqiang Yan 0001, Ming-Hsuan Yang 0001, Jinshan Pan, Guangwei Gao, Ying Tai, Jian Yang 0003 |
Int. J. Comput. Vis. | 4 |
| 2026 | Correction: Scene Prior Filtering for Depth Super-Resolution
Zhengxue Wang, Zhiqiang Yan 0001, Ming-Hsuan Yang 0001, Jinshan Pan, Guangwei Gao, Ying Tai, Jian Yang 0003 |
Int. J. Comput. Vis. | 4 |
| 2026 | Mamba-Driven Topology Fusion for monocular 3D human pose estimation
Zenghao Zheng, Lianping Yang, Jinshan Pan, Hegui Zhu |
Image Vis. Comput. | 3 |
| 2026 | Collaborative Feedback Discriminative Propagation for Video Super-ResolutionabstractThe key success of existing video super-resolution (VSR) methods stems mainly from exploring spatial and temporal information that is usually achieved by a temporal propagation with alignment strategies. However, inaccurate alignment usually leads to significant artifacts that will be accumulated during propagation and thus affect video restoration. Moreover, only propagating the same timestep features forward or backward does not handle the videos with complex motion or occlusion. To address these issues, we propose a collaborative feedback discriminative (CFD) method to correct inaccurate aligned features and better model spatial and temporal information for VSR. Specifically, we first develop a discriminative alignment correction (DAC) method to reduce the influences of the artifacts caused by inaccurate alignment. Then, we propose a collaborative feedback propagation (CFP) module based on feedback and gating mechanisms to explore spatial and temporal information of different timestep features from forward and backward propagation simultaneously. Finally, we embed the proposed DAC and CFP into commonly used VSR networks to verify the effectiveness of our method. Experimental results demonstrate that our method improves the performance of existing VSR models while maintaining a lower model complexity. Hao Li 0058, Xiang Chen 0015, Jiangxin Dong, Jinhui Tang 0001, Jinshan Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Adaptive Sparse Self-Attention for Efficient Image Super-Resolution and BeyondabstractBenefiting from the effectiveness of the self-attention mechanisms in the Transformer framework for modeling non-local features of images, significant progress has been achieved in image super-resolution. We note that existing self-attention mechanisms usually explore all similarities of the tokens between the queries and keys for the feature aggregation. However, using all the similarities does not effectively facilitate the high-quality image reconstruction as not all the tokens from the queries are relevant to those in keys. We further note that self-attention mechanisms are less effective for local feature exploration, which are less effective for the structural detail restoration. To overcome these problems, we develop a simple yet effective adaptive sparse self-attention method to utilize the most useful information of tokens for image restoration. We first develop a local spatial-variant feature estimation method to build the query and key used in the self-attention so that local information can be better modeled. Then, we present a simple yet effective sparse self-attention to adaptively select the most useful similarity values from the self-attention matrix for better the feature aggregation. We analyze that the proposed method models both local and non-local features and thus facilitates better structural detail restoration. We further show that the proposed method can serve as an alternative to existing self-attention mechanisms for better image restoration. Experimental results show that the proposed method performs favorably against state-of-the-art ones on benchmark datasets in terms of accuracy and model complexity. Jinshan Pan, Lianhong Song, Jiangxin Dong, Jian Yang 0003, Maocheng Zhao, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | A unified framework to learn invariant representations of graph neural networks for ECG biometrics
Tianbang Ma, Chunying Liu, Yilong Yin, Gongping Yang 0001, Jinshan Pan |
Pattern Recognit. | 6 |
| 2026 | Distilling Features From Vision Foundation Models for Effective Memory-Based Video ColorizationabstractAlthough memory-based video colorization methods can effectively leverage information from previous frames to guide the colorization of the current frame, they still struggle to capture the semantic features of grayscale frames. To overcome this limitation, we introduce a foundation model (Dinov3 [1]) with strong semantic priors to distill semantic features from grayscale frames, thereby improving the accuracy of feature retrieval from memory and enhancing the effect of subsequent object matching. Moreover, existing memory-based approaches often suffer from redundant information, which can degrade colorization quality. To address this issue, we design a dynamic feature selection module (DFSM) that adaptively selects the most relevant features from memory. In addition, we propose a multi-scale fusion module (MSFM) that fuses mid- and high-scale features derived from visual–semantic interactions to further refine video colorization. To minimize the loss of useful information, we discard the feature-compression strategy adopted in previous methods, thereby preserving temporal features to the greatest extent and further improving colorization performance. Extensive experiments on benchmark datasets and real-world videos demonstrate that the proposed method outperforms existing state-of-the-art approaches, validating its effectiveness and performance advantages. Zhongzheng Peng, Jinshan Pan, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Ultra-High-Definition Image Restoration: New Benchmarks and a Dual Interaction Prior-Driven SolutionabstractUltra-High-Definition (UHD) image restoration has acquired remarkable attention due to its practical demand. In this paper, we construct UHD snow and rain benchmarks, named UHD-Snow and UHD-Rain, to remedy the deficiency in this field. The UHD-Snow/UHD-Rain is established by simulating the physics process of rain/snow into consideration and each benchmark contains 3200 degraded/clear image pairs of 4K resolution. Furthermore, we propose an effective UHD image restoration solution by considering gradient and normal priors in model design, thanks to these priors’ spatial and detail contributions. Specifically, our method contains two branches: (a) feature fusion and reconstruction branch in high-resolution space and (b) prior feature interaction branch in low-resolution space. The former learns high-resolution features and fuses prior-guided low-resolution features to reconstruct clear images, while the latter utilizes normal and gradient priors to mine useful spatial features and detail features to guide high-resolution recovery better. To better utilize these priors, we introduce single prior feature interaction and dual prior feature interaction, where the former respectively fuses normal and gradient priors with high-resolution features to enhance prior ones, while the latter calculates the similarity between enhanced prior ones and further exploits dual guided filtering to boost the feature interaction of dual priors. We conduct experiments on both new and existing public datasets and demonstrate the state-of-the-art performance of our method on UHD image low-light enhancement, dehazing, deblurring, desnowing, and deraining. The source codes and benchmarks are available at https://github.com/wlydlut/UHDDIP. Cong Wang 0018, Jinshan Pan, Xiaofeng Liu 0001, Weixiang Zhou, Xiaoran Sun, Wei Wang 0335, Zhixun Su |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | UDMMColor: A Unified Diffusion Model for Multi-Modal ColorizationabstractDiffusion model-based networks have been widely applied in the field of image generation and have gradually demonstrated a strong potential in image colorization tasks. However, despite the emergence of various colorization diffusion models, two major challenges remain: (1) the lack of effective control over the colorization process and (2) the prevalent issue of color bleeding. Integrating suitable conditional control can effectively alleviate these challenges. To this end, we propose a unified multi-modal diffusion model that harnesses diverse modality information to achieve flexible and high-quality colorization. Specifically, we introduce a Stroke-Adapter that extracts and integrates stroke prompt, enhancing user control over color distribution. Additionally, we design an Edge-Guided Attention mechanism to effectively inject edge information into the colorization process, significantly reducing color bleeding artifacts. Extensive comparative experiments demonstrate that our method outperforms state-of-the-art image colorization approaches in both qualitative and quantitative evaluations, achieving superior colorization results with enhanced controllability. Yan Zhai, Zerui Han, Zhulin Tao, Xianglin Huang, Jinshan Pan, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Degradation-Aware Prompt Learning With Cross-Modal Compensation for Adverse Weather RemovalabstractAdverse weather causes diverse and complex image degradations, severely compromising the reliability of computer vision systems. Existing all-in-one restoration models attempt to address multiple degradation types within a unified framework, but often lack explicit spatial and semantic modeling of degradation characteristics, limiting their adaptability to diverse weather conditions. To address this limitation, we propose a Degradation-Aware Cross-Modal Prompt Compensation Network (DCMPC-Net) that leverages cross-modal degradation cues from a pre-trained vision-language model to condition restoration features within a unified backbone. Specifically, our DCMPC-Net mainly consists of the Cross-Modal Prompt Generator (CMPG), Prompt-Guided Attention Alignment Module (PGAAM), and Dual Feature Compensation Module (DFCM). The CMPG integrates textual embeddings with visual features to produce degradation-aware prompts that encode degradation-related semantic and contextual cues. These prompts are injected into the decoder via a PGAAM, which adaptively aligns semantic information with degraded regions to facilitate context-aware restoration. To further enhance structural fidelity, DFCM is introduced that disentangles degradation artifacts from scene structures, thereby improving the reconstruction of fine textures and detailed content. By integrating cross-modal semantic guidance with spatial alignment and structural enhancement, DCMPC-Net achieves robust and perceptually consistent restoration across diverse weather conditions. Extensive experiments show that DCMPC-Net outperforms state-of-the-art methods in both task-specific and unified settings, achieving superior accuracy and visual fidelity. The code is available at https://github.com/fanamber831/DCMPC-Net. Wanshu Fan, Yunzhe Zhang, Jing Qin 0007, Kin-Man Lam 0001, Cong Wang 0018, Jinshan Pan |
IEEE Trans. Image Process. | 8 |
| 2026 | Rethinking the Importance of High-Frequency Components in Transformers for Image RestorationabstractTransformer-based approaches have achieved superior performance in image restoration, since they can model long-term dependencies well. However, the limitation in capturing local information restricts their capacity to remove degradations. While existing approaches attempt to mitigate this issue by incorporating convolutional operations, the core component in Transformer, i.e., self-attention, which serves as a low-pass filter, could unintentionally dilute or even eliminate the acquired local patterns. In this paper, we propose HIT, a simple yet effective High-frequency Injected Transformer for image restoration. Specifically, we design a window-wise injection module (WIM), which incorporates abundant high-frequency details into the feature map, to provide reliable references for restoring high-quality images. We also develop a bidirectional interaction module (BIM) to aggregate features at different scales using a mutually reinforced paradigm, resulting in spatially and contextually improved representations. In addition, we introduce a spatial enhancement unit (SEU) to preserve essential spatial relationships that may be lost due to the computations carried out across channel dimensions in the BIM. Extensive experiments on 6 tasks (real noise, rain streak, blur, flare, underwater conditions, and low-light conditions) demonstrate that HIT with linear computational complexity performs favorably against the state-of-the-art methods. The source code is available at https://github.com/joshyZhou/HIT_. Shihao Zhou 0003, Jinshan Pan, Duosheng Chen, Yaopeng Dong, Jufeng Yang |
IEEE Trans. Image Process. | 2 |
| 2026 | Towards Ultra-High-Definition Image Deraining: A Benchmark and an Efficient MethodabstractDespite significant advancements in image deraining, most existing methods are carried out on low-resolution images, leaving their effectiveness on high-resolution images uncertain. This limitation becomes even more pronounced with the rise of ultra-high-definition (UHD) imaging. In this paper, we tackle the challenge of UHD image deraining and introduce 4K-Rain13 k, the first large-scale UHD image deraining dataset, featuring 13,000 paired images at 4 K resolution. Leveraging this dataset, we conduct a benchmark study on existing methods for processing UHD images. To better address this task, we propose UDR-Mixer, an efficient and effective architecture tailored for UHD image deraining. Our model comprises two key components: a spatial feature rearrangement layer, which captures long-range dependencies in UHD images, and a frequency feature modulation layer, which enhances high-fidelity image reconstruction. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods while maintaining lower model complexity. The source code and proposed dataset are available athttps://github.com/cschenxiang/UDR-Mixer. Hongming Chen 0004, Xiang Chen 0015, Chen Wu 0006, Zhuoran Zheng, Jinshan Pan, Xianping Fu |
IEEE Trans. Multim. | 5 |
| 2026 | Lightweight Multi-Dilated Transformer for Image DeblurringabstractWindow-based Transformers have achieved promising results in image deblurring. However, their limited ability to capture nonlocal information hinders further improvement in deblurring performance. In this article, we develop an effective multi-dilated Transformer, named MDFormer, to address this issue. Specifically, we first develop a multi-dilated feature aggregation (MDFA) module, which aims to extract and aggregate nonlocal information with reduced computational costs. As commonly used feed-forward networks are pixelwise operations, we propose a dilated feed-forward network (DiFFN) module to enhance the information interaction between pixels further. Moreover, to fully utilize the features of different scales, we introduce a multiscale feature fusion (MSFF) module to provide improved guidance for image reconstruction. Extensive experiments demonstrate that the proposed method generates comparable results against state-of-the-art approaches with reduced computational costs. Zhulin Tao, Jinshan Pan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Intra and Inter Parser-Prompted Transformers for Effective Image RestorationabstractWe propose Intra and Inter Parser-Prompted Transformers (PPTformer) that explore useful features from visual foundation models for image restoration. Specifically, PPTformer contains two parts: an Image Restoration Network (IRNet) for restoring images from degraded observations and a Parser-Prompted Feature Generation Network (PPFGNet) for providing IRNet with reliable parser information to boost restoration. To enhance the integration of the parser within IRNet, we propose Intra Parser-Prompted Attention (IntraPPA) and Inter Parser-Prompted Attention (InterPPA) to implicitly and explicitly learn useful parser features to facilitate restoration. The IntraPPA re-considers cross attention between parser and restoration features, enabling implicit perception of the parser from a long-range and intra-layer perspective. Conversely, the InterPPA initially fuses restoration features with those of the parser, followed by formulating these fused features within an attention mechanism to explicitly perceive parser information. Further, we propose a parser-prompted feed-forward network to guide restoration within pixel-wise gating modulation. Experimental results show that PPTformer achieves state-of-the-art performance on image deraining, defocus deblurring, desnowing, and low-light enhancement. Cong Wang 0018, Jinshan Pan, Wei Wang 0335 |
AAAI | 2 |
| 2025 | FaithDiff: Unleashing Diffusion Priors for Faithful Image Super-resolutionabstractFaithful image super-resolution (SR) not only needs to recover images that appear realistic, similar to image generation tasks, but also requires that the restored images maintain fidelity and structural consistency with the input. To this end, we propose a simple and effective method, named FaithDiff, to fully harness the impressive power of latent diffusion models (LDMs) for faithful image SR. In contrast to existing diffusion-based SR methods that freeze the diffusion model pre-trained on high-quality images, we propose to unleash the diffusion prior to identify useful information and recover faithful structures. As there exists a significant gap between the features of degraded inputs and the noisy latent from the diffusion model, we then develop an effective alignment module to explore useful features from degraded inputs to align well with the diffusion process. Considering the indispensable roles and interplay of the encoder and diffusion model in LDMs, we jointly fine-tune them in a unified optimization framework, facilitating the encoder to extract useful features that coincide with the diffusion process. Extensive experimental results demonstrate that FaithDiff outperforms state-of-the-art methods, providing high-quality and faithful SR results. Junyang Chen 0002, Jinshan Pan, Jiangxin Dong |
CVPR | 2 |
| 2025 | A Polarization-Aided Transformer for Image Deblurring via Motion Vector DecompositionabstractEffectively leveraging motion information is crucial for the image deblurring task. Existing methods typically build deep-learning models to restore a clean image by estimating blur patterns over the entire movement. This suggests that the blur caused by rotational motion components is processed together with the translational one. Exploring the movement without separation leads to limited performance for complex motion deblurring, especially rotational motion. In this paper, we propose Motion Decomposition Transformer (MDT), a transformer-based architecture augmented with polarized modules for deblurring via motion vector decomposition. MDT consists of a Motion Decomposition Module (MDM) for extracting hybrid rotation and translation features and a Radial Stripe Attention Solver (RSAS) for sharp image reconstruction with enhanced rotational information. Specifically, the MDM uses a deformable Cartesian convolutional branch to capture translational motion, complemented by a polar-system branch to capture rotational motion. The RSAS employs radial stripe windows and angular relative positional encoding in the polar system to enhance rotational information. This design preserves translational details while keeping computational costs lower than dual-coordinate design. Experimental results on 6 image deblurring datasets show that MDT outperforms state-of-the-art methods, particularly in handling blur caused by complex motions with significant rotational components. The code and pre-trained models are available at https://github.com/Calvin11311/MDT. Duosheng Chen, Shihao Zhou 0003, Jinshan Pan, Jinglei Shi, Lishen Qu, Jufeng Yang |
CVPR | 3 |
| 2025 | Efficient Visual State Space Model for Image DeblurringabstractConvolutional neural networks (CNNs) and Vision Transformers (ViTs) have achieved excellent performance in image restoration. While ViTs generally outperform CNNs by effectively capturing long-range dependencies and input-specific characteristics, their computational complexity increases quadratically with image resolution. This limitation hampers their practical application in high-resolution image restoration. In this paper, we propose a simple yet effective visual state space model (EVSSM) for image deblurring, leveraging the benefits of state space models (SSMs) for visual data. In contrast to existing methods that employ several fixed-direction scanning for feature extraction, which significantly increases the computational cost, we develop an efficient visual scan block that applies various geometric transformations before each SSM-based module, capturing useful non-local information and maintaining high efficiency. In addition, to more effectively capture and represent local information, we propose an efficient discriminative frequency domain-based feedforward network (EDFFN), which can effectively estimate useful frequency information for latent clear image restoration. Extensive experimental results show that the proposed EVSSM performs favorably against state-of-the-art methods on benchmark datasets and real-world images. Lingshun Kong, Jiangxin Dong, Jinhui Tang 0001, Ming-Hsuan Yang 0001, Jinshan Pan |
CVPR | 5 |
| 2025 | DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-ResolutionabstractRecent RGB-guided depth super-resolution methods have achieved impressive performance under the assumption of fixed and known degradation (e.g., bicubic downsampling). However, in real-world scenarios, captured depth data often suffer from unconventional and unknown degradation due to sensor limitations and complex imaging environments (e.g., low reflective surfaces, varying illumination). Consequently, the performance of these methods significantly declines when real-world degradation deviate from their assumptions. In this paper, we propose the Degradation Oriented and Regularized Network (DORNet), a novel framework designed to adaptively address unknown degradation in real-world scenes through implicit degradation representations. Our approach begins with the development of a self-supervised degradation learning strategy, which models the degradation representations of low-resolution depth data using routing selection-based degradation regularization. To facilitate effective RGB-D fusion, we further introduce a degradation-oriented feature transformation module that selectively propagates RGB content into the depth data based on the learned degradation priors. Extensive experimental results on both real and synthetic datasets demonstrate the superiority of our DORNet in handling unknown degradation, outperforming existing methods. Zhengxue Wang, Zhiqiang Yan 0001, Jinshan Pan, Guangwei Gao, Kai Zhang 0008, Jian Yang 0003 |
CVPR | 3 |
| 2025 | Efficient Video Super-Resolution for Real-time Rendering with Decoupled G-buffer GuidanceabstractLatency is a key driver for real-time rendering applications, making super-resolution techniques increasingly popular to accelerate rendering processes. In contrast to existing methods that directly concatenate low-resolution frames and G-buffers as input without discrimination, we develop an asymmetric UNet-based super-resolution network with decoupled G-buffer guidance, dubbed RDG, to facilitate the spatial and temporal feature exploration for minimizing performance overheads and latency. We first propose a dynamic feature modulator (DFM) to selectively encode the spatial information to capture precise structural information. We then incorporate auxiliary G-buffer information to guide the decoder to generate detail-rich, temporally stable results. Specifically, we adopt a high-frequency feature booster (HFB) to adaptively transfer the high-frequency information from the normal and bidirectional reflectance distribution function (BRDF) components of the G-buffer, enhancing the details of the generated results. To further enhance the temporal stability, we design a cross-frame temporal refiner (CTR) with depth and motion vector constraints to aggregate the previous and current frames. Extensive experimental results reveal that our proposed method is capable of generating high-quality and temporally stable results in real-time rendering. The proposed RDG-s produces 1080P rendering results on a RTX 3090 GPU with a speed of 126 FPS. Our source codes and pre-trained models are available at: https://github.com/sunny2109/RDG. Mingjun Zheng, Jiangxin Dong, Jinshan Pan |
CVPR | 4 |
| 2025 | Efficient Concertormer for Image Deblurring and Beyond
Pin-Hung Kuo, Jinshan Pan, Shao-Yi Chien, Ming-Hsuan Yang 0001 |
ICCV | 2 |
| 2025 | FoundIR: Unleashing Million-Scale Training Data to Advance Foundation Models for Image RestorationabstractDespite the significant progress made by all-in-one models in universal image restoration, existing methods suffer from a generalization bottleneck in real-world scenarios, as they are mostly trained on small-scale synthetic datasets with limited degradations. Therefore, large-scale high-quality real-world training data is urgently needed to facilitate the emergence of foundational models for image restoration. To advance this field, we spare no effort in contributing a million-scale dataset with two notable advantages over existing training data: real-world samples with larger-scale, and degradation types with higher diversity. By adjusting internal camera settings and external imaging conditions, we can capture aligned image pairs using our well-designed data acquisition system over multiple rounds and our data alignment criterion. Moreover, we propose a robust model, FoundIR, to better address a broader range of restoration tasks in real-world scenarios, taking a further step toward foundation models. Specifically, we first utilize a diffusion-based generalist model to remove degradations by learning the degradation-agnostic common representations from diverse inputs, where incremental learning strategy is adopted to better guide model training. To refine the model's restoration capability in complex scenarios, we introduce degradation-aware specialist models for achieving final high-quality results. Extensive experiments show the value of our dataset and the effectiveness of our method. Hao Li 0058, Xiang Chen 0015, Jiangxin Dong, Jinhui Tang 0001, Jinshan Pan |
ICCV | 5 |
| 2025 | PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-Resolution
Yong Liu 0031, Hang Dong 0001, Jinshan Pan, Qingji Dong, Kai Chen 0023, Rongxiang Zhang, Lean Fu, Fei Wang 0008 |
ICCV | 3 |
| 2025 | Frequency Domain-Based Diffusion Model for Unpaired Image DehazingabstractUnpaired image dehazing has attracted increasing attention due to its flexible data requirements during model training. Dominant methods based on contrastive learning not only introduce haze-unrelated content information, but also ignore haze-specific properties in the frequency domain (\ie,~haze-related degradation is mainly manifested in the amplitude spectrum). To address these issues, we propose a novel frequency domain-based diffusion model, named \ours, for fully exploiting the beneficial knowledge in unpaired clear data. In particular, inspired by the strong generative ability shown by Diffusion Models (DMs), we tackle the dehazing task from the perspective of frequency domain reconstruction and perform the DMs to yield the amplitude spectrum consistent with the distribution of clear images. To implement it, we propose an Amplitude Residual Encoder (ARE) to extract the amplitude residuals, which effectively compensates for the amplitude gap from the hazy to clear domains, as well as provide supervision for the DMs training. In addition, we propose a Phase Correction Module (PCM) to eliminate artifacts by further refining the phase spectrum during dehazing with a simple attention mechanism. Experimental results demonstrate that our \ours outperforms other state-of-the-art methods on both synthetic and real-world datasets. Chengxu Liu 0001, Lu Qi 0001, Jinshan Pan, Xueming Qian, Ming-Hsuan Yang 0001 |
ICCV | 3 |
| 2025 | Learning Deblurring Texture Prior From Unpaired Data with Diffusion Model
Chengxu Liu 0001, Lu Qi 0001, Jinshan Pan, Xueming Qian, Ming-Hsuan Yang 0001 |
ICCV | 3 |
| 2025 | CA2C: A Prior-Knowledge-Free Approach for Robust Label Noise Learning via Asymmetric Co-Learning and Co-Training
Mengmeng Sheng, Zeren Sun, Tianfei Zhou, Xiangbo Shu, Jinshan Pan, Yazhou Yao |
ICCV | 5 |
| 2025 | Devil is in the Uniformity: Exploring Diverse Learners Within Transformer for Image RestorationabstractTransformer-based approaches have gained significant attention in image restoration, where the core component, i.e, Multi-Head Attention (MHA), plays a crucial role in capturing diverse features and recovering high-quality results. In MHA, heads perform attention calculation independently from uniform split subspaces, and a redundancy issue is triggered to hinder the model from achieving satisfactory outputs. In this paper, we propose to improve MHA by exploring diverse learners and introducing various interactions between heads, which results in a Hierarchical multI-head atteNtion driven Transformer model, termed HINT, for image restoration. HINT contains two modules, i.e., the Hierarchical Multi-Head Attention (HMHA) and the Query-Key Cache Updating (QKCU) module, to address the redundancy problem that is rooted in vanilla MHA. Specifically, HMHA extracts diverse contextual features by employing heads to learn from subspaces of varying sizes and containing different information. Moreover, QKCU, comprising intra- and inter-layer schemes, further reduces the redundancy problem by facilitating enhanced interactions between attention heads within and across layers. Extensive experiments are conducted on 12 benchmarks across 5 image restoration tasks, including low-light enhancement, dehazing, desnowing, denoising, and deraining, to demonstrate the superiority of HINT. The source code is available in the supplementary materials. Shihao Zhou 0003, Dayu Li, Jinshan Pan, Juncheng Zhou, Jinglei Shi, Jufeng Yang |
ICCV | 3 |
| 2025 | UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion Space
Yong Liu 0031, Jinshan Pan, Yinchuan Li, Qingji Dong, Chao Zhu 0007, Yu Guo 0006, Fei Wang 0008 |
ACM Multimedia | 2 |
| 2025 | Rethinking Nighttime Image Deraining via Learnable Color Space TransformationabstractCompared to daytime image deraining, nighttime image deraining poses significant challenges due to inherent complexities of nighttime scenarios and the lack of high-quality datasets that accurately represent the coupling effect between rain and illumination. In this paper, we rethink the task of nighttime image deraining and contribute a new high-quality benchmark, HQ-NightRain, which offers higher harmony and realism compared to existing datasets. In addition, we develop an effective Color Space Transformation Network (CST-Net) for better removing complex rain from nighttime scenes. Specifically, we propose a learnable color space converter (CSC) to better facilitate rain removal in the Y channel, as nighttime rain is more pronounced in the Y channel compared to the RGB color space. To capture illumination information for guiding nighttime deraining, implicit illumination guidance is introduced enabling the learned features to improve the model's robustness in complex scenarios. Extensive experiments show the value of our dataset and the effectiveness of our method. The source code and datasets are available at https://github.com/guanqiyuan/CST-Net. Qiyuan Guan, Xiang Chen 0015, Guiyue Jin, Jiyu Jin, Shumin Fan, Tianyu Song 0003, Jinshan Pan |
NeurIPS | 7 |
| 2025 | DeblurDiff: Real-Word Image Deblurring with Generative Diffusion ModelsabstractDiffusion models have achieved significant progress in image generation and the pre-trained Stable Diffusion (SD) models are helpful for image deblurring by providing clear image priors. However, directly using a blurry image or a pre-deblurred one as a conditional control for SD will either hinder accurate structure extraction or make the results overly dependent on the deblurring network. In this work, we propose a Latent Kernel Prediction Network (LKPN) to achieve robust real-world image deblurring. Specifically, we co-train the LKPN in the latent space with conditional diffusion. The LKPN learns a spatially variant kernel to guide the restoration of sharp images in the latent space. By applying element-wise adaptive convolution (EAC), the learned kernel is utilized to adaptively process the blurry feature, effectively preserving the information of the blurry input. This process thereby more effectively guides the generative process of SD, enhancing both the deblurring efficacy and the quality of detail reconstruction. Moreover, the results at each diffusion step are utilized to iteratively estimate the kernels in LKPN to better restore the sharp latent by EAC in the subsequent step. This iterative refinement enhances the accuracy and robustness of the deblurring process. Extensive experimental results demonstrate that the proposed method outperforms state-of-the-art image deblurring methods on both benchmark and real-world images. Lingshun Kong, Jiawei Zhang 0002, Dongqing Zou, Fu Lee Wang, Jimmy S. J. Ren, Xiaohe Wu, Jiangxin Dong, Jinshan Pan |
NeurIPS | 8 |
| 2025 | FlareX: A Physics-Informed Dataset for Lens Flare Removal via 2D Synthesis and 3D RenderingabstractLens flare occurs when shooting towards strong light sources, significantly degrading the visual quality of images. Due to the difficulty in capturing flare-corrupted and flare-free image pairs in the real world, existing datasets are typically synthesized in 2D by overlaying artificial flare templates onto background images. However, the lack of flare diversity in templates and the neglect of physical principles in the synthesis process hinder models trained on these datasets from generalizing well to real-world scenarios. To address these challenges, we propose a new physics-informed method for flare data generation, which consists of three stages: parameterized template creation, the laws of illumination-aware 2D synthesis, and physical engine-based 3D rendering, which finally gives us a mixed flare dataset that incorporates both 2D and 3D perspectives, namely FlareX. This dataset offers 9,500 2D templates derived from 95 flare patterns and 3,000 flare image pairs rendered from 60 3D scenes. Furthermore, we design a masking approach to obtain real-world flare-free images from their corrupted counterparts to measure the performance of the model on real-world images. Extensive experiments demonstrate the effectiveness of our method and dataset. Lishen Qu, Jinshan Pan, Shihao Zhou 0003, Jinglei Shi, Duosheng Chen, Jufeng Yang |
NeurIPS | 3 |
| 2025 | Deep Unpaired Blind Image Super-Resolution Using Self-supervised Learning and Exemplar Distillation
Jiangxin Dong, Haoran Bai 0001, Jinhui Tang 0001, Jinshan Pan |
Int. J. Comput. Vis. | 4 |
| 2025 | Towards Unified Deep Image Deraining: A Survey and a New BenchmarkabstractRecent years have witnessed significant advances in image deraining due to the progress of effective image priors and deep learning models. As each deraining approach has individual settings (e.g., training and test datasets, evaluation criteria), how to fairly evaluate existing approaches comprehensively is not a trivial task. Although existing surveys aim to thoroughly review image deraining approaches, few of them focus on unifying evaluation settings to examine the deraining capability and practicality evaluation. In this paper, we provide a comprehensive review of existing image deraining methods and provide a unified evaluation setting to evaluate their performance. Furthermore, we construct a new high-quality benchmark named HQ-RAIN to conduct extensive evaluations, consisting of 5,000 paired high-resolution synthetic images with high harmony and realism. We also discuss existing challenges and highlight several future research opportunities worth exploring. To facilitate the reproduction and tracking of the latest deraining technologies for general users, we build an online platform to provide the off-the-shelf toolkit, involving the large-scale performance evaluation. Xiang Chen 0015, Jinshan Pan, Jiangxin Dong, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Learning Efficient Deep Discriminative Spatial and Temporal Networks for Video DeblurringabstractHow to effectively explore spatial and temporal information is important for video deblurring. In contrast to existing methods that directly align adjacent frames without discrimination, we develop a deep discriminative spatial and temporal network to facilitate the spatial and temporal feature exploration for better video deblurring. We first develop a channel-wise gated dynamic network to adaptively explore the spatial information. As adjacent frames usually contain different contents, directly stacking features of adjacent frames without discrimination may affect the latent clear frame restoration. Therefore, we develop a simple yet effective discriminative temporal feature fusion module to obtain useful temporal features for latent frame restoration. Moreover, to utilize the information from long-range frames, we develop a wavelet-based feature propagation method that takes the discriminative temporal feature fusion module as the basic unit to effectively propagate main structures from long-range frames for better video deblurring. Experimental results show that the proposed method performs favorably against state-of-the-art ones on benchmark datasets in terms of accuracy and model complexity. Jinshan Pan, Boming Xu, Jiangxin Dong, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Learning an Adaptive Sparse Transformer for Efficient Image RestorationabstractTransformer-based approaches have shown promising performance in image restoration tasks due to their ability to model long-range dependencies, which are essential for recovering clear images. Although various efficient attention mechanisms have been proposed to address the intensive computational loads of transformers, they often suffer from redundant information and noisy interactions from irrelevant regions, as they consider all available tokens. In this work, we propose an Adaptive Sparse Transformer (AST-v2) to mitigate these issues by reducing noisy interactions in irrelevant areas and removing feature redundancy along channel dimension. AST-v2 incorporates two core components: an Adaptive Sparse Self-Attention (ASSA) block and a Feature Refinement Feed-forward Network (FRFN). ASSA adopts a dual-branch design, where the sparse branch guides the modulation of standard dense attention weights. This paradigm reduces the negative impact of irrelevant token interactions while preserving the important ones. Meanwhile, FRFN utilizes an enhance-and-ease scheme to eliminate feature redundancy across channels, enhancing the restoration of clear images. Experimental results on commonly used benchmarks show the competitive performance of our method for 6 restoration tasks, including rain streak removal, haze removal, shadow removal, snow removal, blur removal, and low-light enhancement. The code is available in the supplementary materials. Shihao Zhou 0003, Jinshan Pan, Jufeng Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Efficient Non-Blind Image Deblurring With Discriminative Shrinkage Deep NetworksabstractMost existing non-blind deblurring methods formulate the problem into a maximum-a-posteriori framework and address it by manually designing a variety of regularization terms and data terms of the latent clear images. However, explicitly designing these two terms is quite challenging, which usually leads to complex optimization problems. In this paper, we propose a Discriminative Shrinkage Deep Network for fast and accurate deblurring. Most existing methods use deep convolutional neural networks (CNNs), or radial basis functions only to learn the regularization term. In contrast, we formulate both the data and regularization terms while splitting the deconvolution model into data-related and regularization-related sub-problems. We explore the properties of the Maxout function and develop a deep CNN model with Maxout layers to learn discriminative shrinkage functions, which directly approximate the solutions of these two sub-problems. Moreover, we develop a U-Net according to Krylov subspace method to restore the latent clear images effectively and efficiently, which plays a role but is better than the conventional fast-Fourier-transform-based or conjugate gradient method. Experimental results show that the proposed method performs favorably against the state-of-the-art methods regarding efficiency and accuracy. Pin-Hung Kuo, Jinshan Pan, Shao-Yi Chien, Ming-Hsuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Omni-Deblurring: Capturing Omni-Range Context for Image DeblurringabstractExisting CNN-based and Transformer-based methods have demonstrated remarkable performance in low-level visual tasks, including image deblurring. These methods generally capture spatial features only in a single way, such as by stacking blocks of CNNs and Transformers, resulting in inadequate utilization of spatial context. To address this issue, we propose a new feature aggregation scheme for image deblurring, named Omni-Deblurring. The core of our omni-deblurring is the omni-range context block, which enables explicitly aggregating the local-range, regional-range, and global-range features in a compact manner. With this design, it can bring a wider receptive field for modeling the contextual features. Extensive experiments on synthetic and real-world blurry datasets demonstrate the effectiveness of our proposed method in both quantitative and qualitative evaluations. Furthermore, the quality of our deblurring model is evaluated in the task of object detection, and the mean Average Precision (mAP) metric increases by 10% across all classes compared with other deblurring models. Code is available athttps://github.com/yaowli468/Omni-Deblurring. Hang An, Xiaoxuan Chen, Bo Jiang 0014, Jinshan Pan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Deep Frequency-Separable Temporal Network for Efficient Video Denoising
Zhulin Tao, Jinjuan Wang, Lifang Yang, Jinshan Pan, Jinhui Tang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Correlation Matching Transformation Transformers for UHD Image RestorationabstractThis paper proposes UHDformer, a general Transformer for Ultra-High-Definition (UHD) image restoration. UHDformer contains two learning spaces: (a) learning in high-resolution space and (b) learning in low-resolution space. The former learns multi-level high-resolution features and fuses low-high features and reconstructs the residual images, while the latter explores more representative features learning from the high-resolution ones to facilitate better restoration. To better improve feature representation in low-resolution space, we propose to build feature transformation from the high-resolution space to the low-resolution one. To that end, we propose two new modules: Dual-path Correlation Matching Transformation module (DualCMT) and Adaptive Channel Modulator (ACM). The DualCMT selects top C/r (r is greater or equal to 1 which controls the squeezing level) correlation channels from the max-pooling/mean-pooling high-resolution features to replace low-resolution ones in Transformers, which can effectively squeeze useless content to improve the feature representation in low-resolution space to facilitate better recovery. The ACM is exploited to adaptively modulate multi-level high-resolution features, enabling to provide more useful features to low-resolution space for better learning. Experimental results show that our UHDformer reduces about ninety-seven percent model sizes compared with most state-of-the-art methods while significantly improving performance under different training sets on 3 UHD image restoration tasks, including low-light image enhancement, image dehazing, and image deblurring. The source codes will be made available at https://github.com/supersupercong/UHDformer. Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Mengzhu Wang, Xiao-Ming Wu 0003, Jun Liu 0036 |
AAAI | 2 |
| 2024 | SelfPromer: Self-Prompt Dehazing Transformers with Depth-ConsistencyabstractThis work presents an effective depth-consistency Self-Prompt Transformer, terms as SelfPromer, for image dehazing. It is motivated by an observation that the estimated depths of an image with haze residuals and its clear counterpart vary. Enforcing the depth consistency of dehazed images with clear ones, therefore, is essential for dehazing. For this purpose, we develop a prompt based on the features of depth differences between the hazy input images and corresponding clear counterparts that can guide dehazing models for better restoration. Specifically, we first apply deep features extracted from the input images to the depth difference features for generating the prompt that contains the haze residual information in the input. Then we propose a prompt embedding module that is designed to perceive the haze residuals, by linearly adding the prompt to the deep features. Further, we develop an effective prompt attention module to pay more attention to haze residuals for better removal. By incorporating the prompt, prompt embedding, and prompt attention into an encoder-decoder network based on VQGAN, we can achieve better perception quality. As the depths of clear images are not available at inference, and the dehazed images with one-time feed-forward execution may still contain a portion of haze residuals, we propose a new continuous self-prompt inference that can iteratively correct the dehazing model towards better haze-free image generation. Extensive experiments show that our SelfPromer performs favorably against the state-of-the-art approaches on both synthetic and real-world datasets in terms of perception metrics including NIQE, PI, and PIQE. The source codes will be made available at https://github.com/supersupercong/SelfPromer. Cong Wang 0018, Jinshan Pan, Wanyu Lin, Jiangxin Dong, Wei Wang 0335, Xiao-Ming Wu 0003 |
AAAI | 2 |
| 2024 | Bidirectional Multi-Scale Implicit Neural Representations for Image DerainingabstractHow to effectively explore multi-scale representations of rain streaks is important for image deraining. In contrast to existing Transformer-based methods that depend mostly on single-scale rain appearance, we develop an end-to-end multi-scale Transformer that leverages the potentially useful features in various scales to facilitate high-quality image reconstruction. To better explore the common degradation representations from spatially-varying rain streaks, we incorporate intra-scale implicit neural representations based on pixel coordinates with the degraded inputs in a closed-loop design, enabling the learned features to facilitate rain removal and improve the robustness of the model in complex scenarios. To ensure richer collaborative representation from different scales, we embed a simple yet effective inter-scale bidirectional feedback operation into our multi-scale Transformer by performing coarse-to-fine and fine-to-coarse information communication. Extensive experiments demonstrate that our approach, named as NeRD-Rain, performs favorably against the state-of-the-art ones on both synthetic and real-world benchmark datasets. The source code and trained models are available at https://github.com/cschenxiang/NeRD-Rain. Xiang Chen 0015, Jinshan Pan, Jiangxin Dong |
CVPR | 2 |
| 2024 | Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationabstractTransformer-based approaches have achieved promising performance in image restoration tasks, given their ability to model long-range dependencies, which is crucial for recovering clear images. Though diverse efficient attention mechanism designs have addressed the intensive computations associated with using transformers, they often involve redundant information and noisy interactions from irrelevant regions by considering all available tokens. In this work, we propose an Adaptive Sparse Transformer (AST) to mitigate the noisy interactions of irrelevant areas and remove feature redundancy in both spatial and channel domains. AST comprises two core designs, i.e., an Adaptive Sparse Self-Attention (ASSA) block and a Feature Refinement Feed-forward Network (FRFN). Specifically, ASSA is adaptively computed using a two-branch paradigm, where the sparse branch is introduced to filter out the negative impacts of low query-key matching scores for aggregating features, while the dense one ensures sufficient information flow through the network for learning discriminative representations. Meanwhile, FRFN employs an enhance-and-ease scheme to eliminate feature redundancy in channels, enhancing the restoration of clear latent images. Experimental results on commonly used benchmarks have demonstrated the versatility and competitive performance of our method in several tasks, including rain streak removal, real haze removal, and raindrop removal. The code and pre-trained models are available at https://github.com/joshyZhou/AST. Shihao Zhou 0003, Duosheng Chen, Jinshan Pan, Jinglei Shi, Jufeng Yang |
CVPR | 3 |
| 2024 | ColorMNet: A Memory-Based Deep Spatial-Temporal Feature Propagation Network for Video Colorization
Yixin Yang 0005, Jiangxin Dong, Jinhui Tang 0001, Jinshan Pan |
ECCV (4) | 4 |
| 2024 | SMFANet: A Lightweight Self-Modulation Feature Aggregation Network for Efficient Image Super-Resolution
Mingjun Zheng, Jiangxin Dong, Jinshan Pan |
ECCV (50) | 4 |
| 2024 | Seeing the Unseen: A Frequency Prompt Guided Transformer for Image Restoration
Shihao Zhou 0003, Jinshan Pan, Jinglei Shi, Duosheng Chen, Lishen Qu, Jufeng Yang |
ECCV (16) | 2 |
| 2024 | Enhanced dual contrast representation learning with cell separation and merging for breast cancer diagnosis
Yang Liu 0119, Yiqi Zhu, Zhehao Gu, Jinshan Pan, Juncheng Li 0003, Ming Fan 0003, Lihua Li 0002, Tieyong Zeng |
Comput. Vis. Image Underst. | 4 |
| 2024 | Deep Richardson-Lucy Deconvolution for Low-Light Image Deblurring
Liang Chen 0026, Jiawei Zhang 0002, Yunxuan Wei, Faming Fang, Jimmy S. J. Ren, Jinshan Pan |
Int. J. Comput. Vis. | 7 |
| 2024 | Correction to: Deep Unpaired Blind Image Super-Resolution Using Self-supervised Learning and Exemplar Distillation
Jiangxin Dong, Haoran Bai 0001, Jinhui Tang 0001, Jinshan Pan |
Int. J. Comput. Vis. | 4 |
| 2024 | Video Colorization: A Survey
Zhongzheng Peng, Yixin Yang 0005, Jinhui Tang 0001, Jinshan Pan |
J. Comput. Sci. Technol. | 4 |
| 2024 | Deep self-supervised spatial-variant image deblurring
Bo Jiang 0014, Zhenghao Shi, Xiaoxuan Chen, Jinshan Pan |
Neural Networks | 5 |
| 2024 | Self-Supervised Deep Blind Video Super-ResolutionabstractExisting deep learning-based video super-resolution (SR) methods usually depend on the supervised learning approach, where the training data is usually generated by the blurring operation with known or predefined kernels (e.g., Bicubic kernel) followed by a decimation operation. However, this does not hold for real applications as the degradation process is complex and cannot be approximated by these idea cases well. Moreover, obtaining high-resolution (HR) videos and the corresponding low-resolution (LR) ones in real-world scenarios is difficult. To overcome these problems, we propose a self-supervised learning method to solve the blind video SR problem, which simultaneously estimates blur kernels and HR videos from the LR videos. As directly using LR videos as supervision usually leads to trivial solutions, we develop a simple and effective method to generate auxiliary paired data from original LR videos according to the image formation of video SR, so that the networks can be better constrained by the generated paired data for both blur kernel estimation and latent HR video restoration. In addition, we introduce an optical flow estimation module to exploit the information from adjacent frames for HR video restoration. Experiments show that our method performs favorably against state-of-the-art ones on benchmarks and real-world videos. Haoran Bai 0001, Jinshan Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | BiSTNet: Semantic Image Prior Guided Bidirectional Temporal Feature Fusion for Deep Exemplar-Based Video ColorizationabstractHow to effectively explore the colors of exemplars and propagate them to colorize each frame is vital for exemplar-based video colorization. In this article, we present a BiSTNet to explore colors of exemplars and utilize them to help video colorization by a bidirectional temporal feature fusion with the guidance of semantic image prior. We first establish the semantic correspondence between each frame and the exemplars in deep feature space to explore color information from exemplars. Then, we develop a simple yet effective bidirectional temporal feature fusion module to propagate the colors of exemplars into each frame and avoid inaccurate alignment. We note that there usually exist color-bleeding artifacts around the boundaries of important objects in videos. To overcome this problem, we develop a mixed expert block to extract semantic information for modeling the object boundaries of frames so that the semantic image prior can better guide the colorization process. In addition, we develop a multi-scale refinement block to progressively colorize frames in a coarse-to-fine manner. Extensive experimental results demonstrate that the proposed BiSTNet performs favorably against state-of-the-art methods on the benchmark datasets and real-world scenes. Moreover, the BiSTNet obtains one champion in NTIRE 2023 video colorization challenge (Kang et al. 2023). Yixin Yang 0005, Jinshan Pan, Zhongzheng Peng, Xiaoyu Du 0002, Zhulin Tao, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Deblurring Videos Using Spatial-Temporal Contextual Transformer With Feature PropagationabstractWe present a simple and effective approach to explore both local spatial-temporal contexts and non-local temporal information for video deblurring. First, we develop an effective spatial-temporal contextual transformer to explore local spatial-temporal contexts from videos. As the features extracted by the spatial-temporal contextual transformer does not model the non-local temporal information of video well, we then develop a feature propagation method to aggregate useful features from the long-range frames so that both local spatial-temporal contexts and non-local temporal information can be better utilized for video deblurring. Finally, we formulate the spatial-temporal contextual transformer with the feature propagation into a unified deep convolutional neural network (CNN) and train it in an end-to-end manner. We show that using the spatial-temporal contextual transformer with the feature propagation is able to generate useful features and makes the deep CNN model more compact and effective for video deblurring. Extensive experimental results show that the proposed method performs favorably against state-of-the-art ones on the benchmark datasets in terms of accuracy and model parameters. Liyan Zhang 0001, Boming Xu, Zhongbao Yang, Jinshan Pan |
IEEE Trans. Image Process. | 4 |
| 2023 | Hybrid CNN-Transformer Feature Fusion for Single Image DerainingabstractSince rain streaks exhibit diverse geometric appearances and irregular overlapped phenomena, these complex characteristics challenge the design of an effective single image deraining model. To this end, rich local-global information representations are increasingly indispensable for better satisfying rain removal. In this paper, we propose a lightweight Hybrid CNN-Transformer Feature Fusion Network (dubbed as HCT-FFN) in a stage-by-stage progressive manner, which can harmonize these two architectures to help image restoration by leveraging their individual learning strengths. Specifically, we stack a sequence of the degradation-aware mixture of experts (DaMoE) modules in the CNN-based stage, where appropriate local experts adaptively enable the model to emphasize spatially-varying rain distribution features. As for the Transformer-based stage, a background-aware vision Transformer (BaViT) module is employed to complement spatially-long feature dependencies of images, so as to achieve global texture recovery while preserving the required structure. Considering the indeterminate knowledge discrepancy among CNN features and Transformer features, we introduce an interactive fusion branch at adjacent stages to further facilitate the reconstruction of high-quality deraining results. Extensive evaluations show the effectiveness and extensibility of our developed HCT-FFN. The source code is available at https://github.com/cschenxiang/HCT-FFN. Xiang Chen 0015, Jinshan Pan, Jiyang Lu, Zhentao Fan, Hao Li 0058 |
AAAI | 2 |
| 2023 | FFHQ-UV: Normalized Facial UV-Texture Dataset for 3D Face ReconstructionabstractWe present a large-scale facial UV-texture dataset that contains over 50,000 high-quality texture UV-maps with even illuminations, neutral expressions, and cleaned facial regions, which are desired characteristics for rendering realistic 3D face models under different lighting conditions. The dataset is derived from a large-scale face image dataset namely FFHQ, with the help of our fully automatic and robust UV-texture production pipeline. Our pipeline utilizes the recent advances in StyleGAN-based facial image editing approaches to generate multi-view normalized face images from single-image inputs. An elaborated UV-texture extraction, correction, and completion procedure is then applied to produce high-quality UV-maps from the normalized face images. Compared with existing UV-texture datasets, our dataset has more diverse and higher-quality texture maps. We further train a GAN-based texture decoder as the nonlinear texture basis for parametric fitting based 3D face reconstruction. Experiments show that our method improves the reconstruction accuracy over state-of-the-art approaches, and more importantly, produces high-quality texture maps that are ready for realistic renderings. The dataset, code, and pre-trained texture decoder are publicly available at https://github.com/csbhr/FFHQ-UV. Haoran Bai 0001, Haoxian Zhang, Jinshan Pan, Linchao Bao |
CVPR | 4 |
| 2023 | Learning A Sparse Transformer Network for Effective Image DerainingabstractTransformers-based methods have achieved significant performance in image deraining as they can model the non-local information which is vital for high-quality image reconstruction. In this paper, we find that most existing Transformers usually use all similarities of the tokens from the query-key pairs for the feature aggregation. However, if the tokens from the query are different from those of the key, the self-attention values estimated from these tokens also involve in feature aggregation, which accordingly interferes with the clear image restoration. To overcome this problem, we propose an effective DeRaining network, Sparse Transformer (DRSformer) that can adaptively keep the most useful self-attention values for feature aggregation so that the aggregated features better facilitate high-quality image reconstruction. Specifically, we develop a learnable top-k selection operator to adaptively retain the most crucial attention scores from the keys for each query for better feature aggregation. Simultaneously, as the naive feed-forward network in Transformers does not model the multi-scale information that is important for latent clear image restoration, we develop an effective mixed-scale feed-forward network to generate better features for image deraining. To learn an enriched set of hybrid features, which combines local context from CNN operators, we equip our model with mixture of experts feature compensator to present a cooperation refinement deraining scheme. Extensive experimental results on the commonly used benchmarks demonstrate that the proposed method achieves favorable performance against state-of-the-art approaches. The source code and trained models are available at https://github.com/cschenxiang/DRSformer. Xiang Chen 0015, Hao Li 0058, Mingqiang Li, Jinshan Pan |
CVPR | 4 |
| 2023 | Efficient Frequency Domain-based Transformers for High-Quality Image DeblurringabstractWe present an effective and efficient method that explores the properties of Transformers in the frequency domain for high-quality image deblurring. Our method is motivated by the convolution theorem that the correlation or convolution of two signals in the spatial domain is equivalent to an element-wise product of them in the frequency domain. This inspires us to develop an efficient frequency domain-based self-attention solver (FSAS) to estimate the scaled dot-product attention by an element-wise product operation instead of the matrix multiplication in the spatial domain. In addition, we note that simply using the naive feed-forward network (FFN) in Transformers does not generate good deblurred results. To overcome this problem, we propose a simple yet effective discriminative frequency domain-based FFN (DFFN), where we introduce a gated mechanism in the FFN based on the Joint Photographic Experts Group (JPEG) compression algorithm to discriminatively determine which low- and high-frequency information of the features should be preserved for latent clear image restoration. We formulate the proposed FSAS and DFFN into an asymmetrical network based on an encoder and decoder architecture, where the FSAS is only used in the decoder module for better image deblurring. Experimental results show that the proposed method performs favorably against the state-of-the-art approaches. Lingshun Kong, Jiangxin Dong, Jianjun Ge, Mingqiang Li, Jinshan Pan |
CVPR | 5 |
| 2023 | Deep Discriminative Spatial and Temporal Network for Efficient Video DeblurringabstractHow to effectively explore spatial and temporal information is important for video deblurring. In contrast to existing methods that directly align adjacent frames without discrimination, we develop a deep discriminative spatial and temporal network to facilitate the spatial and temporal feature exploration for better video deblurring. We first develop a channel-wise gated dynamic network to adaptively explore the spatial information. As adjacent frames usually contain different contents, directly stacking features of adjacent frames without discrimination may affect the latent clear frame restoration. Therefore, we develop a simple yet effective discriminative temporal feature fusion module to obtain useful temporal features for latent frame restoration. Moreover, to utilize the information from long-range frames, we develop a wavelet-based feature propagation method that takes the discriminative temporal feature fusion module as the basic unit to effectively propagate main structures from long-range frames for better video deblurring. We show that the proposed method does not require additional alignment methods and performs favorably against state-of-the-art ones on benchmark datasets in terms of accuracy and model complexity. Jinshan Pan, Boming Xu, Jiangxin Dong, Jianjun Ge, Jinhui Tang 0001 |
CVPR | 1 |
| 2023 | Boosting Video Super Resolution with Patch-Based Temporal Redundancy Optimization
Hang Dong 0001, Jinshan Pan, Chao Zhu 0007, Boyang Liang, Yu Guo 0006, Ding Liu 0001, Lean Fu, Fei Wang 0008 |
ICANN (7) | 3 |
| 2023 | SVMV: Spatiotemporal Variance-Supervised Motion Volume for Video Frame InterpolationabstractHigh-performance video frame interpolation is challenging for complex scenes with diverse motion and occlusion characteristics. Existing methods, deploying off-the-shelf flow estimators to acquire initial characterizations refined by multiple subsequent models, often require heavy network architectures that are not practical for resource constrained systems. We investigate the unary potentials of the characterizations to improve efficiency. Specifically, we design a lightweight neural network to construct motion volumes via ensembles of offset approximations, and propose a spatiotemporal variance-aware loss to supervise the network learning. For network compactness, our spatiotemporal variance-supervised motion volume (SVMV) utilizes shared spatiotemporal representations via correlations among approximations, of which the diversifications are exploited to better leverage the network’s expressiveness through the spatiotemporal variances of motions and occlusions within the time interval to be interpolated. Experiments on publicly available datasets show that our method performs favorably against existing methods with a more compact network and less runtime. Yao Luo, Jinshan Pan, Jinhui Tang 0001 |
ICASSP | 2 |
| 2023 | Multi-scale Residual Low-Pass Filter Network for Image DeblurringabstractWe present a simple and effective Multi-scale Residual Low-Pass Filter Network (MRLPFNet) that jointly explores the image details and main structures for image deblurring. Our work is motivated by an observation that the difference between the blurry image and the clear one not only contains high-frequency contents1but also includes low-frequency information due to the influence of blur, while using the standard residual learning is less effective for modeling the main structure distorted by the blur. Considering that the low-frequency contents usually correspond to main global structures that are spatially variant, we first propose a learnable low-pass filter based on a self-attention mechanism to adaptively explore the global contexts for better modeling the low-frequency information. Then we embed it into a Residual Low-Pass Filter (RLPF) module, which involves an additional fully convolutional neural network with the standard residual learning to model the high-frequency information. We formulate the RLPF module into an end-to-end trainable network based on an encoder and decoder architecture and develop a wavelet-based feature fusion to fuse the multi-scale features. Experimental results show that our method performs favorably against state-of-the-art ones on commonly-used benchmarks. Jiangxin Dong, Jinshan Pan, Zhongbao Yang, Jinhui Tang 0001 |
ICCV | 2 |
| 2023 | DLGSANet: Lightweight Dynamic Local and Global Self-Attention Network for Image Super-ResolutionabstractWe propose an effective lightweight dynamic local and global self-attention network (DLGSANet) to solve image super-resolution. Our method explores the properties of Transformers while having low computational costs. Motivated by the network designs of Transformers, we develop a simple yet effective multi-head dynamic local self-attention (MHDLSA) module to extract local features efficiently. In addition, we note that existing Transformers usually explore all similarities of the tokens between the queries and keys for the feature aggregation. However, using all the similarities does not effectively facilitate the high-resolution image reconstruction as not all the tokens from the queries are relevant to those in keys. To overcome this problem, we develop a sparse global self-attention (SparseGSA) module to select the most useful similarity values so that the most useful global features can be better utilized for image reconstruction. We develop a hybrid dynamic-Transformer block (HDTB) that integrates the MHDLSA and SparseGSA for both local and global feature exploration. To ease the network training, we formulate the HDTBs into a residual hybrid dynamic-Transformer group (RHDTG). By embedding the RHDTGs into an end-to-end trainable network, we show that the proposed method has fewer network parameters and lower computational costs while achieving competitive performance against state-of-the-art ones in terms of accuracy. More information is available at https://neonleexiang.github.io/DLGSANet/. Xiang Li 0103, Jiangxin Dong, Jinhui Tang 0001, Jinshan Pan |
ICCV | 4 |
| 2023 | Spatially-Adaptive Feature Modulation for Efficient Image Super-ResolutionabstractAlthough deep learning-based solutions have achieved impressive reconstruction performance in image super-resolution (SR), these models are generally large, with complex architectures, making them incompatible with low-power devices with many computational and memory constraints. To overcome these challenges, we propose a spatially-adaptive feature modulation (SAFM) mechanism for efficient SR design. In detail, the SAFM layer uses independent computations to learn multi-scale feature representations and aggregates these features for dynamic spatial modulation. As the SAFM prioritizes exploiting non-local feature dependencies, we further introduce a convolutional channel mixer (CCM) to encode local contextual information and mix channels simultaneously. Extensive experimental results show that the proposed method is 3× smaller than state-of-the-art efficient SR methods, e.g., IMDN, and yields comparable performance with much less memory usage. Our source codes and pre-trained models are available at: https://github.com/sunny2109/SAFMN. Jiangxin Dong, Jinhui Tang 0001, Jinshan Pan |
ICCV | 4 |
| 2023 | Lightweight Deep Deblurring Model with Discriminative Multi-Scale Feature FusionabstractAlthough existing learning-based deblurring methods achieve significant progress, these approaches tend to require lots of network parameters and huge computational costs, which limits their practical applications. Instead of pursuing larger deep models for boosting deblurring performance, we propose a lightweight deep convolutional neural network with lower computational costs and comparable restoration performance, which is based on a multi-scale framework with an encoder and decoder network architecture. Specifically, we present an effective depth-wise separable convolution block (DSCB) as the fundamental building block of our method to reduce the model complexity. In addition, to better utilize the features from different scales, we develop a simple yet effective discriminative multi-scale feature fusion (DMFF) module for achieving high-quality results. Experimental results on the benchmarks show that our method is about 10× smaller than the state-of-the-art deblurring methods, e.g. MPRNet [1], in terms of model parameters and FLOPs while achieving competitive performance. The training code and models are available at https://github.com/cslvjt/LightweightDeblur. Jiangtao Lv, Jinshan Pan |
ICIP | 2 |
| 2023 | MBDFNet: Multi-scale Bidirectional Dynamic Feature Fusion Network for Efficient Image DeblurringabstractExisting deep image deblurring models achieve favorable results with growing model complexity. However, these models cannot be applied to those low-power devices with resource constraints (e.g., smart phones) as these models usually have lots of network parameters and require computational costs. To overcome this problem, we develop a multi-scale bidirectional dynamic feature fusion network (MBDFNet), a lightweight deep deblurring model, for efficient image deblurring. The proposed MBDFNet progressively restores multi-scale latent clear images from blurry input based on a multi-scale framework. To better utilize the features from coarse scales, we propose a bidirectional gated dynamic fusion module so that the most useful information of the features from coarse scales are kept to facilitate the estimations in the finer scales. We solve the proposed MBDFNet in an end-to-end manner and show that it has fewer network parameters and lower FLOPs values, where the FLOPs value of the proposed MBDFNet is at least 6× smaller than the state-of-the-art methods. Both quantitative and qualitative evaluations show that the proposed MBDFNet achieves favorable performance in terms of model complexity while having competitive performance in terms of accuracy against state-of-the-art methods. Zhongbao Yang, Jinshan Pan |
ICME | 2 |
| 2023 | PromptRestorer: A Prompting Image Restoration Method with Degradation PerceptionabstractWe show that raw degradation features can effectively guide deep restoration models, providing accurate degradation priors to facilitate better restoration. While networks that do not consider them for restoration forget gradually degradation during the learning process, model capacity is severely hindered. To address this, we propose a Prompting image Restorer, termed as PromptRestorer. Specifically, PromptRestorer contains two branches: a restoration branch and a prompting branch. The former is used to restore images, while the latter perceives degradation priors to prompt the restoration branch with reliable perceived content to guide the restoration process for better recovery. To better perceive the degradation which is extracted by a pre-trained model from given degradation observations, we propose a prompting degradation perception modulator, which adequately considers the characters of the self-attention mechanism and pixel-wise modulation, to better perceive the degradation priors from global and local perspectives. To control the propagation of the perceived content for the restoration branch, we propose gated degradation perception propagation, enabling the restoration branch to adaptively learn more useful features for better recovery. Extensive experimental results show that our PromptRestorer achieves state-of-the-art results on 4 image restoration tasks, including image deraining, deblurring, dehazing, and desnowing. Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Jiangxin Dong, Mengzhu Wang, Yakun Ju, Junyang Chen 0001 |
NeurIPS | 2 |
| 2023 | Memory-Augmented Deep Unfolding Network for Guided Image Super-resolution
Man Zhou 0003, Jinshan Pan, Wenqi Ren, Qi Xie 0002, Xiangyong Cao |
Int. J. Comput. Vis. | 3 |
| 2023 | Single Image Deraining Using Residual Channel Attention Networks
Di Wang 0018, Jinshan Pan, Jinhui Tang 0001 |
J. Comput. Sci. Technol. | 2 |
| 2023 | Cascaded Deep Video Deblurring Using Temporal Sharpness Prior and Non-Local Spatial-Temporal SimilarityabstractWe present compact and effective deep convolutional neural networks (CNNs) by exploring properties of videos for video deblurring. Motivated by the non-uniform blur property that not all the pixels of the frames are blurry, we develop a CNN to integrate a temporal sharpness prior (TSP) for removing blur in videos. The TSP exploits sharp pixels from adjacent frames to facilitate the CNN for better frame restoration. Observing that the motion field is related to latent frames instead of blurry ones in the image formation model, we develop an effective cascaded training approach to solve the proposed CNN in an end-to-end manner. As videos usually contain similar contents within and across frames, we propose a non-local similarity mining approach based on a self-attention method with the propagation of global features to constrain CNNs for frame restoration. We show that exploring the domain knowledge of videos can make CNNs more compact and efficient, where the CNN with the non-local spatial-temporal similarity is 3× smaller than the state-of-the-art methods in terms of model parameters while its performance gains are at least 1 dB higher in terms of PSNRs. Extensive experimental results show that our method performs favorably against state-of-the-art approaches on benchmarks and real-world videos. Jinshan Pan, Boming Xu, Haoran Bai 0001, Jinhui Tang 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Real-World Image Super-Resolution by Exclusionary Dual-LearningabstractReal-world image super-resolution is a practical image restoration problem that aims to obtain high-quality images from in-the-wild input, has recently received considerable attention with regard to its tremendous application potentials. Although deep learning-based methods have achieved promising restoration quality on real-world image super-resolution datasets, they ignore the relationship between L1- and perceptual- minimization and roughly adopt auxiliary large-scale datasets for pre-training. In this paper, we discuss the image types within a corrupted image and the property of perceptual- and Euclidean- based evaluation protocols. Then we propose a method, Real-World image Super-Resolution by Exclusionary Dual-Learning (RWSR-EDL) to address the feature diversity in perceptual- and L1- based cooperative learning. Moreover, a noise-guidance data collection strategy is developed to address the training time consumption in multiple datasets optimization. When an auxiliary dataset is incorporated, RWSR-EDL achieves promising results and repulses any training time increment by adopting the noise-guidance data collection strategy. Extensive experiments show that RWSR-EDL achieves competitive performance over state-of-the-art methods on four in-the-wild image super-resolution datasets. Hao Li 0058, Jinghui Qin, Zhijing Yang, Pengxu Wei, Jinshan Pan, Liang Lin 0004, Yukai Shi |
IEEE Trans. Multim. | 5 |
| 2022 | Online-Updated High-Order Collaborative Networks for Single Image DerainingabstractSingle image deraining is an important and challenging task for some downstream artificial intelligence applications such as video surveillance and self-driving systems. Most of the existing deep-learning-based methods constrain the network to generate derained images but few of them explore features from intermediate layers, different levels, and different modules which are beneficial for rain streaks removal. In this paper, we propose a high-order collaborative network with multi-scale compact constraints and a bidirectional scale-content similarity mining module to exploit features from deep networks externally and internally for rain streaks removal. Externally, we design a deraining framework with three sub-networks trained in a collaborative manner, where the bottom network transmits intermediate features to the middle network which also receives shallower rainy features from the top network and sends back features to the bottom network. Internally, we enforce multi-scale compact constraints on the intermediate layers of deep networks to learn useful features via a Laplacian pyramid. Further, we develop a bidirectional scale-content similarity mining module to explore features at different scales in a down-to-up and up-to-down manner. To improve the model performance on real-world images, we propose an online-update learning approach, which uses real-world rainy images to fine-tune the network and update the deraining results in a self-supervised manner. Extensive experiments demonstrate that our proposed method performs favorably against eleven state-of-the-art methods on five public synthetic datasets and one real-world dataset. Cong Wang 0018, Jinshan Pan, Xiao-Ming Wu 0003 |
AAAI | 2 |
| 2022 | Deep Recurrent Neural Network with Multi-Scale Bi-directional Propagation for Video DeblurringabstractThe success of the state-of-the-art video deblurring methods stems mainly from implicit or explicit estimation of alignment among the adjacent frames for latent video restoration. However, due to the influence of the blur effect, estimating the alignment information from the blurry adjacent frames is not a trivial task. Inaccurate estimations will interfere the following frame restoration. Instead of estimating alignment information, we propose a simple and effective deep Recurrent Neural Network with Multi-scale Bi-directional Propagation (RNN-MBP) to effectively propagate and gather the information from unaligned neighboring frames for better video deblurring. Specifically, we build a Multi-scale Bi-directional Propagation (MBP) module with two U-Net RNN cells which can directly exploit the inter-frame information from unaligned neighboring hidden states by integrating them in different scales. Moreover, to better evaluate the proposed algorithm and existing state-of-the-art methods on real-world blurry scenes, we also create a Real-World Blurry Video Dataset (RBVD) by a well-designed Digital Video Acquisition System (DVAS) and use it as the training and evaluation dataset. Extensive experimental results demonstrate that the proposed RBVD dataset effectively improve the performance of existing algorithms on real-world blurry videos, and the proposed algorithm performs favorably against the state-of-the-art methods on three typical benchmarks. The code is available at https://github.com/XJTU-CVLAB-LOWLEVEL/RNN-MBP. Chao Zhu 0007, Hang Dong 0001, Jinshan Pan, Boyang Liang, Lean Fu, Fei Wang 0008 |
AAAI | 3 |
| 2022 | Unpaired Deep Image Deraining Using Dual Contrastive LearningabstractLearning single image deraining (SID) networks from an unpaired set of clean and rainy images is practical and valuable as acquiring paired real-world data is almost infeasible. However, without the paired data as the supervision, learning a SID network is challenging. Moreover, simply using existing unpaired learning methods (e.g., unpaired adversarial learning and cycle-consistency constraints) in the SID task is insufficient to learn the underlying relationship from rainy inputs to clean outputs as there exists significant domain gap between the rainy and clean images. In this paper, we develop an effective unpaired SID adversarial framework which explores mutual properties of the unpaired exemplars by a dual contrastive learning manner in a deep feature space, named as DCD-GAN. The proposed method mainly consists of two cooperative branches: Bidirectional Translation Branch (BTB) and Contrastive Guidance Branch (CGB). Specifically, BTB exploits full advantage of the circulatory architecture of adversarial consistency to generate abundant exemplar pairs and excavates latent feature distributions between two domains by equipping it with bidirectional mapping. Simultaneously, CGB implicitly constrains the embeddings of different exemplars in the deep feature space by encouraging the similar feature distributions closer while pushing the dissimilar further away, in order to better facilitate rain removal and help image restoration. Extensive experiments demonstrate that our method performs favorably against existing unpaired deraining approaches on both synthetic and real-world datasets, and generates comparable results against several fully-supervised or semi-supervised models. Xiang Chen 0015, Jinshan Pan, Kui Jiang, Yufeng Li 0001, Caihua Kong, Longgang Dai, Zhentao Fan |
CVPR | 2 |
| 2022 | Learning Discriminative Shrinkage Deep Networks for Image Deconvolution
Pin-Hung Kuo, Jinshan Pan, Shao-Yi Chien, Ming-Hsuan Yang 0001 |
ECCV (19) | 2 |
| 2022 | ShuffleMixer: An Efficient ConvNet for Image Super-ResolutionabstractLightweight and efficiency are critical drivers for the practical application of image super-resolution (SR) algorithms. We propose a simple and effective approach, ShuffleMixer, for lightweight image super-resolution that explores large convolution and channel split-shuffle operation. In contrast to previous SR models that simply stack multiple small kernel convolutions or complex operators to learn representations, we explore a large kernel ConvNet for mobile-friendly SR design. Specifically, we develop a large depth-wise convolution and two projection layers based on channel splitting and shuffling as the basic component to mix features efficiently. Since the contexts of natural images are strongly locally correlated, using large depth-wise convolutions only is insufficient to reconstruct fine details. To overcome this problem while maintaining the efficiency of the proposed module, we introduce Fused-MBConvs into the proposed network to model the local connectivity of different features. Experimental results demonstrate that the proposed ShuffleMixer is about $3 \times$ smaller than the state-of-the-art efficient SR methods, e.g. CARN, in terms of model parameters and FLOPs while achieving competitive performance. Jinshan Pan, Jinhui Tang 0001 |
NeurIPS | 2 |
| 2022 | Dual Convolutional Neural Networks for Low-Level Vision
Jinshan Pan, Deqing Sun, Jiawei Zhang 0002, Jinhui Tang 0001, Jian Yang 0003, Yu-Wing Tai, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 1 |
| 2022 | DnSwin: Toward real-world denoising via a continuous Wavelet Sliding Transformer
Hao Li 0058, Zhijing Yang, Ziying Zhao, Junyang Chen 0001, Yukai Shi, Jinshan Pan |
Knowl. Based Syst. | 7 |
| 2022 | Learning Spatially Variant Linear Representation Models for Joint FilteringabstractJoint filtering mainly uses an additional guidance image as a prior and transfers its structures to the target image in the filtering process. Different from existing approaches that rely on local linear models or hand-designed objective functions to extract the structural information from the guidance image, we propose a new joint filtering method based on a spatially variant linear representation model (SVLRM), where the target image is linearly represented by the guidance image. However, learning SVLRMs for vision tasks is a highly ill-posed problem. To estimate the spatially variant linear representation coefficients, we develop an effective approach based on a deep convolutional neural network (CNN). As such, the proposed deep CNN (constrained by the SVLRM) is able to model the structural information of both the guidance and input images. We show that the proposed approach can be effectively applied to a variety of applications, including depth/RGB image upsampling and restoration, flash deblurring, natural image denoising, and scale-aware filtering. In addition, we show that the linear representation model can be extended to high-order representation models (e.g., quadratic and cubic polynomial representations). Extensive experimental results demonstrate that the proposed method performs favorably against the state-of-the-art methods that have been specifically designed for each task. Jiangxin Dong, Jinshan Pan, Jimmy S. J. Ren, Liang Lin 0004, Jinhui Tang 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Deblurring Dynamic Scenes via Spatially Varying Recurrent Neural NetworksabstractDeblurring images captured in dynamic scenes is challenging as the motion blurs are spatially varying caused by camera shakes and object movements. In this paper, we propose a spatially varying neural network to deblur dynamic scenes. The proposed model is composed of three deep convolutional neural networks (CNNs) and a recurrent neural network (RNN). The RNN is used as a deconvolution operator on feature maps extracted from the input image by one of the CNNs. Another CNN is used to learn the spatially varying weights for the RNN. As a result, the RNN is spatial-aware and can implicitly model the deblurring process with spatially varying kernels. To better exploit properties of the spatially varying RNN, we develop both one-dimensional and two-dimensional RNNs for deblurring. The third component, based on a CNN, reconstructs the final deblurred feature maps into a restored image. In addition, the whole network is end-to-end trainable. Quantitative and qualitative evaluations on benchmark datasets demonstrate that the proposed method performs favorably against the state-of-the-art deblurring algorithms. Wenqi Ren, Jiawei Zhang 0002, Jinshan Pan, Sifei Liu, Jimmy S. J. Ren, Junping Du 0001, Xiaochun Cao, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Learning an Occlusion-Aware Network for Video DeblurringabstractVideo deblurring is a challenging task since the blur is caused by camera shake, object motions, etc. The success of the state-of-the-art methods stems mainly from exploiting the temporal information of neighboring frames through alignment. When there exists occlusion among the sequence, these approaches become less effective for inaccurate alignment. In this paper, we propose an effective occlusion-aware network to handle the occlusion for video deblurring. The proposed module first generates a coarse pixel-wise alignment filter to explore the temporal information and then learns an adaptive affine transformation to deal with the occluded areas. In addition, a self-attention mechanism is developed to better model the occluded pixels. To further improve the performance, we progress a multi-scale strategy and train the network in an end-to-end manner. Both quantitative and qualitative experimental results show that the proposed method achieves favorable performance against state-of-the-art methods on the benchmark datasets. The code and trained models are available at:https://github.com/XQLuck/code.git Qian Xu 0014, Jinshan Pan, Yuntao Qian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Deep Dynamic Scene Deblurring From Optical FlowabstractDeblurring can not only provide visually more pleasant pictures and make photography more convenient, but also can improve the performance of objection detection as well as tracking. However, removing dynamic scene blur from images is a non-trivial task as it is difficult to model the non-uniform blur mathematically. Several methods first use single or multiple images to estimate optical flow (which is treated as an approximation of blur kernels) and then adopt non-blind deblurring algorithms to reconstruct the sharp images. However, these methods cannot be trained in an end-to-end manner and are usually computationally expensive. In this paper, we explore optical flow to remove dynamic scene blur by using the multi-scale spatially variant recurrent neural network (RNN). We utilize FlowNets to estimate optical flow from two consecutive images in different scales. The estimated optical flow provides the RNN weights in different scales so that the weights can better help RNNs to remove blur in the feature spaces. Finally, we develop a convolutional neural network (CNN) to restore the sharp images from the deblurred features. Both quantitatively and qualitatively evaluations on the benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art algorithms in terms of accuracy, speed, and model size. Jiawei Zhang 0002, Jinshan Pan, Daoye Wang, Shangchen Zhou, Xing Wei 0001, Furong Zhao, Jimmy S. J. Ren |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Self-Guided Image Dehazing Using Progressive Feature FusionabstractWe propose an effective image dehazing algorithm which explores useful information from the input hazy image itself as the guidance for the haze removal. The proposed algorithm first uses a deep pre-dehazer to generate an intermediate result, and takes it as the reference image due to the clear structures it contains. To better explore the guidance information in the generated reference image, it then develops a progressive feature fusion module to fuse the features of the hazy image and the reference image. Finally, the image restoration module takes the fused features as input to use the guidance information for better clear image restoration. All the proposed modules are trained in an end-to-end fashion, and we show that the proposed deep pre-dehazer with progressive feature fusion module is able to help haze removal. Extensive experimental results show that the proposed algorithm performs favorably against state-of-the-art methods on the widely-used dehazing benchmark datasets as well as real-world hazy images. Haoran Bai 0001, Jinshan Pan, Xinguang Xiang, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Deep Ranking Exemplar-Based Dynamic Scene DeblurringabstractDynamic scene deblurring is a challenging problem as it is difficult to be modeled mathematically. Benefiting from the deep convolutional neural networks, this problem has been significantly advanced by the end-to-end network architectures. However, the success of these methods is mainly due to simply stacking network layers. In addition, the methods based on the end-to-end network architectures usually estimate latent images in a regression way which does not preserve the structural details. In this paper, we propose an exemplar-based method to solve dynamic scene deblurring problem. To explore the properties of the exemplars, we propose a siamese encoder network and a shallow encoder network to respectively extract input features and exemplar features and then develop a rank module to explore useful features for better blur removing, where the rank modules are applied to the last three layers of encoder, respectively. The proposed method can be further extended to the way of multi-scale, which enables to recover more texture from the exemplar. Extensive experiments show that our method achieves significant improvements in both quantitative and qualitative evaluations. Jinshan Pan, Ye Luo 0004 |
IEEE Trans. Image Process. | 2 |
| 2022 | Bi-Directional Pseudo-Three-Dimensional Network for Video Frame InterpolationabstractRecent video frame interpolation methods have employed the curvilinear motion model to accommodate nonlinear motion among frames. The effectiveness of such model often hinges on motion estimation and occlusion detection, and therefore is greatly challenged when these methods are used to handle dynamic scenes that contain complex motions and occlusions. We address the challenges by proposing a bi-directional pseudo-three-dimensional network to exploit the correlation between motion estimation and depth-related occlusion estimation that considers the third dimension: depth. Specifically, the network exploits the correlation by learning shared multi-scale spatiotemporal representations, and by coupling the estimations, in both the past and future directions, to synthesize intermediate frames through a bi-directional pseudo-three-dimensional warping layer, where adaptive convolution kernels are estimated progressively from the coalescence of motion and depth-related occlusion estimations across multiple scales to acquire nonlocal and adaptive neighborhoods. The proposed network utilizes a novel multi-task collaborative learning strategy, which facilitates the supervised learning of video frame interpolation using complementary self-supervisory signals from motion and depth-related occlusion estimations. Across various benchmark datasets, the proposed method outperforms state-of-the-art methods in terms of accuracy, model size and runtime performance. Yao Luo, Jinshan Pan, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Learning Discriminative Cross-Modality Features for RGB-D Saliency DetectionabstractHow to explore useful information from depth is the key success of the RGB-D saliency detection methods. While the RGB and depth images are from different domains, a modality gap will lead to unsatisfactory results for simple feature concatenation. Towards better performance, most methods focus on bridging this gap and designing different cross-modal fusion modules for features, while ignoring explicitly extracting some useful consistent information from them. To overcome this problem, we develop a simple yet effective RGB-D saliency detection method by learning discriminative cross-modality features based on the deep neural network. The proposed method first learns modality-specific features for RGB and depth inputs. And then we separately calculate the correlations of every pixel-pair in a cross-modality consistent way, i.e., the distribution ranges are consistent for the correlations calculated based on features extracted from RGB (RGB correlation) or depth inputs (depth correlation). From different perspectives, color or spatial, the RGB and depth correlations end up at the same point to depict how tightly each pixel-pair is related. Secondly, to complemently gather RGB and depth information, we propose a novel correlation-fusion to fuse RGB and depth correlations, resulting in a cross-modality correlation. Finally, the features are refined with both long-range cross-modality correlations and local depth correlations to predict salient maps. In which, the long-range cross-modality correlation provides context information for accurate localization, and the local depth correlation keeps good subtle structures for fine segmentation. In addition, a lightweight DepthNet is designed for efficient depth feature extraction. We solve the proposed network in an end-to-end manner. Both quantitative and qualitative experimental results demonstrate the proposed algorithm achieves favorable performance against state-of-the-art methods. Fengyun Wang, Jinshan Pan, Shoukun Xu, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Learning a Non-Blind Deblurring Network for Night Blurry ImagesabstractDeblurring night blurry images is difficult, because the common-used blur model based on the linear convolution operation does not hold in this situation due to the influence of saturated pixels. In this paper, we propose a non-blind deblurring network (NBDN) to restore night blurry images. To mitigate the side effects brought by the pixels that violate the blur model, we develop a confidence estimation unit (CEU) to estimate a map which ensures smaller contributions of these pixels in the deconvolution steps which are optimized by the conjugate gradient (CG) method. Moreover, unlike the existing methods using manually tuned hyper-parameters in their frameworks, we propose a hyper-parameter estimation unit (HPEU) to adaptively estimate hyper-parameters for better image restoration. The experimental results demonstrate that the proposed network performs favorably against state-of-the-art algorithms both quantitatively and qualitatively. Liang Chen 0026, Jiawei Zhang 0002, Jinshan Pan, Songnan Lin, Faming Fang, Jimmy S. J. Ren |
CVPR | 3 |
| 2021 | Learning To Restore Hazy Video: A New Real-World Dataset and a New MethodabstractMost of the existing deep learning-based dehazing methods are trained and evaluated on the image dehazing datasets, where the dehazed images are generated by only exploiting the information from the corresponding hazy ones. On the other hand, video dehazing algorithms, which can acquire more satisfying dehazing results by exploiting the temporal redundancy from neighborhood hazy frames, receive less attention due to the absence of the video dehazing datasets. Therefore, we propose the first REal-world VIdeo DEhazing (REVIDE) dataset which can be used for the supervised learning of the video dehazing algorithms. By utilizing a well-designed video acquisition system, we can capture paired real-world hazy and haze-free videos that are perfectly aligned by recording the same scene (with or without haze) twice. Considering the challenge of exploiting temporal redundancy among the hazy frames, we also develop a Confidence Guided and Improved Deformable Network (CG-IDN) for video dehazing. The experiments demonstrate that the hazy scenes in the REVIDE dataset are more realistic than the synthetic datasets and the proposed algorithm also performs favorably against state-of-the-art dehazing methods. Xinyi Zhang 0005, Hang Dong 0001, Jinshan Pan, Chao Zhu 0007, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Fei Wang 0008 |
CVPR | 3 |
| 2021 | Image De-Raining via Continual LearningabstractWhile deep convolutional neural networks (CNNs) have achieved great success on image de-raining task, most existing methods can only learn fixed mapping rules between paired rainy/clean images on a single dataset. This limits their applications in practical situations with multiple and incremental datasets where the mapping rules may change for different types of rain streaks. However, the catastrophic forgetting of traditional deep CNN model challenges the design of generalized framework for multiple and incremental datasets. A strategy of sharing the network structure but in-dependently updating and storing the network parameters on each dataset has been developed as a potential solution. Nevertheless, this strategy is not applicable to compact systems as it dramatically increases the overall training time and parameter space. To alleviate such limitation, in this study, we propose a parameter importance guided weights modification approach, named PIGWM. Specifically, with new dataset (e.g. new rain dataset), the well-trained network weights are updated according to their importance evaluated on previous training dataset. With extensive experimental validation, we demonstrate that a single network with a single parameter set of our proposed method can process multiple rain datasets almost without performance degradation. The proposed model is capable of achieving superior performance on both inhomogeneous and incremental datasets, and is promising for highly compact systems to gradually learn myriad regularities of the different types of rain streaks. The results indicate that our proposed method has great potential for other computer vision tasks with dynamic learning environments. Man Zhou 0003, Jie Xiao 0002, Yifan Chang, Xueyang Fu, Aiping Liu, Jinshan Pan, Zhengjun Zha |
CVPR | 6 |
| 2021 | Unpaired Learning for Deep Image Deraining with Rain Direction RegularizerabstractWe present a simple yet effective unpaired learning based image rain removal method from an unpaired set of synthetic images and real rainy images by exploring the properties of rain maps. The proposed algorithm mainly consists of a semi-supervised learning part and a knowledge distillation part. The semi-supervised part estimates the rain map and reconstructs the derained image based on the well-established layer separation principle. To facilitate rain removal, we develop a rain direction regularizer to constrain the rain estimation network in the semi-supervised learning part. With the estimated rain maps from the semi-supervised learning part, we first synthesize a new paired set by adding to rain-free images based on the superimposition model. The real rainy images and the derained results constitute another paired set. Then we develop an effective knowledge distillation method to explore such two paired sets so that the deraining model in the semi-supervised learning part is distilled. We propose two new rainy datasets, named RainDirection and Real3000, to validate the effectiveness of the proposed method. Both quantitative and qualitative experimental results demonstrate that the proposed method achieves favorable results against state-of-the-art methods in benchmark datasets and real-world images. Yang Liu 0119, Ziyu Yue, Jinshan Pan, Zhixun Su |
ICCV | 3 |
| 2021 | Deep Blind Video Super-resolutionabstractExisting video super-resolution (SR) algorithms usually assume that the blur kernels in the degradation process are known and do not model the blur kernels in the restoration. However, this assumption does not hold for blind video SR and usually leads to over-smoothed super-resolved frames. In this paper, we propose an effective blind video SR algorithm based on deep convolutional neural networks (CNNs). Our algorithm first estimates blur kernels from low-resolution (LR) input videos. Then, with the estimated blur kernels, we develop an effective image deconvolution method based on the image formation model of blind video SR to generate intermediate latent frames so that sharp image contents can be restored well. To effectively explore the information from adjacent frames, we estimate the motion fields from LR input videos, extract features from LR videos by a feature extraction network, and warp the extracted features from LR inputs based on the motion fields. Moreover, we develop an effective sharp feature exploration method which first extracts sharp features from restored intermediate latent frames and then uses a transformation operation based on the extracted sharp features and warped features from LR inputs to generate better features for HR video restoration. We formulate the proposed algorithm into an end-to-end trainable framework and show that it performs favorably against state-of-the-art methods. Jinshan Pan, Haoran Bai 0001, Jiangxin Dong, Jiawei Zhang 0002, Jinhui Tang 0001 |
ICCV | 1 |
| 2021 | Learning a Tree-Structured Channel-Wise Refinement Network for Efficient Image DerainingabstractSignificant advances have been made in image deraining due to the use of kinds of deep neural networks. However, existing deep neural network-based methods usually contain significant abundant network parameters and thus leads to expensive computation cost, which limits the application of deraining technology in high-level vision tasks. In this paper, we propose a compact and flexible Tree-structured Channel-wise Refinement Block (TCRB) for efficient image deraining, which contains augmentation, refinement, and aggregation modules to better explore features. Specifically, the refinement module can progressively extract groups of more discriminative features from the channel augmented inputs, and then the aggregation module adaptively fuses features from the refinement module to preserve image details by leveraging the Enhanced Channel Attention (ECA) method. Moreover, we present a Tree-structured Channel-wise Refinement Network (TCRN) by stacking multiple TCRBs, which could achieve competitive performance as the complicated networks. We embed the TCRB into a Multi-scale Tree-structured Channel-wise Refinement Network (MTCRN) based on an encoder and decoder network architecture and show that it performs favorably against state-of-the-art deraining algorithms on both synthetic datasets and real-world rainy images, while reaching a better trade-off in terms of model parameters and inference time. Di Wang 0018, Hao Tang 0007, Jinshan Pan, Jinhui Tang 0001 |
ICME | 3 |
| 2021 | Editorial for CVIU_DL for image restoration
Jinshan Pan, Deqing Sun, Jian Yang 0003, Wangmeng Zuo, Paolo Favaro, Yasuyuki Matsushita, Ming-Hsuan Yang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2021 | Physics-Based Generative Adversarial Models for Image Restoration and BeyondabstractWe present an algorithm to directly solve numerous image restoration problems (e.g., image deblurring, image dehazing, and image deraining). These problems are ill-posed, and the common assumptions for existing methods are usually based on heuristic image priors. In this paper, we show that these problems can be solved by generative models with adversarial learning. However, a straightforward formulation based on a straightforward generative adversarial network (GAN) does not perform well in these tasks, and some structures of the estimated images are usually not preserved well. Motivated by an interesting observation that the estimated results should be consistent with the observed inputs under the physics models, we propose an algorithm that guides the estimation process of a specific task within the GAN framework. The proposed model is trained in an end-to-end fashion and can be applied to a variety of image restoration and low-level vision problems. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art algorithms. Jinshan Pan, Jiangxin Dong, Yang Liu 0119, Jiawei Zhang 0002, Jimmy S. J. Ren, Jinhui Tang 0001, Yu-Wing Tai, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Multi-Stage Degradation Homogenization for Super-Resolution of Face Images With Extreme DegradationsabstractFace Super-Resolution (FSR) aims to infer High-Resolution (HR) face images from the captured Low-Resolution (LR) face image with the assistance of external information. Existing FSR methods are less effective for the LR face images captured with serious low-quality since the huge imaging/degradation gap caused by the different imaging scenarios (i.e., the complex practical imaging scenario that generates test LR images, the simple manual imaging degradation that generates the training LR images) is not considered in these algorithms. In this paper, we propose an image homogenization strategy via re-expression to solve this problem. In contrast to existing methods, we propose a homogenization projection in LR space and HR space as compensation for the classical LR/HR projection to formulate the FSR in a multi-stage framework. We then develop a re-expression process to bridge the gap between the complex degradation and the simple degradation, which can remove the heterogeneous factors such as serious noise and blur. To further improve the accuracy of the homogenization, we extract the image patch set that is invariant to degradation changes as Robust Neighbor Resources (RNR), with which these two homogenization projections re-express the input LR images and the initial inferred HR images successively. Both quantitative and qualitative results on the public datasets demonstrate the effectiveness of the proposed algorithm against the state-of-the-art methods. Liang Chen 0026, Jinshan Pan, Junjun Jiang, Jiawei Zhang 0002, Zhen Han 0002, Linchao Bao |
IEEE Trans. Image Process. | 2 |
| 2021 | Deep Outlier Handling for Image DeblurringabstractOutlier handling has attracted considerable attention recently but remains challenging for image deblurring. Existing approaches mainly depend on iterative outlier detection steps to explicitly or implicitly reduce the influence of outliers on image deblurring. However, these outlier detection steps usually involve heuristic operations and iterative optimization processes, which are complex and time-consuming. In contrast, we propose to learn a deep convolutional neural network to directly estimate the confidence map, which can identify reliable inliers and outliers from the blurred image and thus facilitates the following deblurring process. We analyze that the proposed algorithm incorporated with the learned confidence map is effective in handling outliers and does not require ad-hoc outlier detection steps which are critical to existing outlier handling methods. Compared to existing approaches, the proposed algorithm is more efficient and can be applied to both non-blind and blind image deblurring. Extensive experimental results demonstrate that the proposed algorithm performs favorably against state-of-the-art methods in terms of accuracy and efficiency. Jiangxin Dong, Jinshan Pan |
IEEE Trans. Image Process. | 2 |
| 2020 | Learning to Deblur Face Images via Sketch SynthesisabstractThe success of existing face deblurring methods based on deep neural networks is mainly due to the large model capacity. Few algorithms have been specially designed according to the domain knowledge of face images and the physical properties of the deblurring process. In this paper, we propose an effective face deblurring algorithm based on deep convolutional neural networks (CNNs). Motivated by the conventional deblurring process which usually involves the motion blur estimation and the latent clear image restoration, the proposed algorithm first estimates motion blur by a deep CNN and then restores latent clear images with the estimated motion blur. However, estimating motion blur from blurry face images is difficult as the textures of the blurry face images are scarce. As most face images share some common global structures which can be modeled well by sketch information, we propose to learn face sketches by a deep CNN so that the sketches can help the motion blur estimation. With the estimated motion blur, we then develop an effective latent image restoration algorithm based on a deep CNN. Although involving the several components, the proposed algorithm is trained in an end-to-end fashion. We analyze the effectiveness of each component on face image deblurring and show that the proposed algorithm is able to deblur face images with favorable performance against state-of-the-art methods. Songnan Lin, Jiawei Zhang 0002, Jinshan Pan, Yicun Liu, Yongtian Wang, Jing S. J. Chen, Jimmy S. J. Ren |
AAAI | 3 |
| 2020 | Image Formation Model Guided Deep Image Super-ResolutionabstractWe present a simple and effective image super-resolution algorithm that imposes an image formation constraint on the deep neural networks via pixel substitution. The proposed algorithm first uses a deep neural network to estimate intermediate high-resolution images, blurs the intermediate images using known blur kernels, and then substitutes values of the pixels at the un-decimated positions with those of the corresponding pixels from the low-resolution images. The output of the pixel substitution process strictly satisfies the image formation model and is further refined by the same deep neural network in a cascaded manner. The proposed framework is trained in an end-to-end fashion and can work with existing feed-forward deep neural networks for super-resolution and converges fast in practice. Extensive experimental results show that the proposed algorithm performs favorably against state-of-the-art methods. Jinshan Pan, Yang Liu 0119, Deqing Sun, Jimmy S. J. Ren, Ming-Ming Cheng, Jian Yang 0003, Jinhui Tang 0001 |
AAAI | 1 |
| 2020 | Multi-Scale Boosted Dehazing Network With Dense Feature FusionabstractIn this paper, we propose a Multi-Scale Boosted Dehazing Network with Dense Feature Fusion based on the U-Net architecture. The proposed method is designed based on two principles, boosting and error feedback, and we show that they are suitable for the dehazing problem. By incorporating the Strengthen-Operate-Subtract boosting strategy in the decoder of the proposed model, we develop a simple yet effective boosted decoder to progressively restore the haze-free image. To address the issue of preserving spatial information in the U-Net architecture, we design a dense feature fusion module using the back-projection feedback scheme. We show that the dense feature fusion module can simultaneously remedy the missing spatial information from high-resolution features and exploit the non-adjacent features. Extensive evaluations demonstrate that the proposed model performs favorably against the state-of-the-art approaches on the benchmark datasets as well as real-world hazy images. Hang Dong 0001, Jinshan Pan, Xinyi Zhang 0005, Fei Wang 0008, Ming-Hsuan Yang 0001 |
CVPR | 2 |
| 2020 | Cascaded Deep Video Deblurring Using Temporal Sharpness PriorabstractWe present a simple and effective deep convolutional neural network (CNN) model for video deblurring. The proposed algorithm mainly consists of optical flow estimation from intermediate latent frames and latent frame restoration steps. It first develops a deep CNN model to estimate optical flow from intermediate latent frames and then restores the latent frames based on the estimated optical flow. To better explore the temporal information from videos, we develop a temporal sharpness prior to constrain the deep CNN model to help the latent frame restoration. We develop an effective cascaded training approach and jointly train the proposed CNN model in an end-to-end manner. We show that exploring the domain knowledge of video deblurring is able to make the deep CNN model more compact and efficient. Extensive experimental results show that the proposed algorithm performs favorably against state-of-the-art methods on the benchmark datasets as well as real-world videos. Jinshan Pan, Haoran Bai 0001, Jinhui Tang 0001 |
CVPR | 1 |
| 2020 | Physics-Based Feature Dehazing Networks
Jiangxin Dong, Jinshan Pan |
ECCV (30) | 2 |
| 2020 | Learning Event-Driven Video Deblurring and Interpolation
Songnan Lin, Jiawei Zhang 0002, Jinshan Pan, Dongqing Zou, Yongtian Wang, Jing Chen 0018, Jimmy S. J. Ren |
ECCV (8) | 3 |
| 2020 | Single Image Dehazing via Multi-scale Convolutional Neural Networks with Holistic Edges
Wenqi Ren, Jinshan Pan, Hua Zhang 0008, Xiaochun Cao, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | Modeling and Optimizing of the Multi-Layer Nearest Neighbor Network for Face Image Super-ResolutionabstractIn this paper, we propose a face super-resolution (FSR) method to handle the decreasing face recognition rate caused by low-quality images. To better model the input images, we build a nearest neighbor network (NNN) which consists of nodes and paths by introducing the second-layer nearest neighbors (SLNNs), where the paths of the network represent the distance between nodes. As the SLNN is trained in the high-resolution (HR) space and is exponentially supplementary to the traditional first-layer nearest neighbors (FLNNs), the neighbor inadequacy problem can be effectively solved by enriching the neighbor candidate set via NNN. Furthermore, we solve the NNN for the optimal weights of neighbors. Finally, we fuse the refined weights and neighbors for better reconstruction results. The effectiveness of this fusion strategy is validated by both quantitative and qualitative experimental results. The extensive experimental results on the public face datasets and real-world challenging low-resolution (LR) images demonstrate that the proposed method performs favorably against the state-of-the-art methods. Liang Chen 0026, Jinshan Pan, Ruimin Hu, Zhen Han 0002, Chao Liang 0001, Yi Wu 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Robust Face Super-Resolution via Position Relation Model Based on Global Face ContextabstractBecause Face Super-Resolution (FSR) tends to infer High-Resolution (HR) face image by breaking the given Low- Resolution (LR) image into individual patches and inferring the HR correspondence one patch by one separately, Super- Resolution (SR) of face images with serious degradation, especially with occlusion, is still a challenging problem of the computer vision field. To address this problem, we propose a patch-level face model for FSR, which we called the position relation model. This model consists of the mapping relationships in every face position to the rest of the face positions based on similarity. In other words, we build a constraint for each patch position via the relationship in this model from the global range of face. Once an individual input LR image patch is seriously deteriorated, the substitute patch in whole face range can be sought according to the relationship of the model at this position as the provider of the LR information. In this way, the lost facial structures can be compensated by knowledge located in remote pixels or structure information which leads to better high-resolution face images. The LR images with degradations, not only the serious low-quality degradation, e.g. noise, blur, but also the occlusions, can be effectively hallucinated into HR ones. Quantitative and qualitative evaluations on the public datasets demonstrate that the proposed algorithm performs favorably against state-of-theart methods. Liang Chen 0026, Jinshan Pan, Junjun Jiang, Jiawei Zhang 0002, Yi Wu 0010 |
IEEE Trans. Image Process. | 2 |
| 2020 | Semi-Supervised Image DehazingabstractWe present an effective semi-supervised learning algorithm for single image dehazing. The proposed algorithm applies a deep Convolutional Neural Network (CNN) containing a supervised learning branch and an unsupervised learning branch. In the supervised branch, the deep neural network is constrained by the supervised loss functions, which are mean squared, perceptual, and adversarial losses. In the unsupervised branch, we exploit the properties of clean images via sparsity of dark channel and gradient priors to constrain the network. We train the proposed network on both the synthetic data and real-world images in an end-to-end manner. Our analysis shows that the proposed semi-supervised learning algorithm is not limited to synthetic training datasets and can be generalized well to real-world images. Extensive experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art single image dehazing algorithms on both benchmark datasets and real-world images. Lerenhan Li, Yunlong Dong, Wenqi Ren, Jinshan Pan, Changxin Gao, Nong Sang, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Task-Oriented Network for Image DehazingabstractHaze interferes the transmission of scene radiation and significantly degrades color and details of outdoor images. Existing deep neural networks-based image dehazing algorithms usually use some common networks. The network design does not model the image formation of haze process well, which accordingly leads to dehazed images containing artifacts and haze residuals in some special scenes. In this paper, we propose a task-oriented network for image dehazing, where the network design is motivated by the image formation of haze process. The task-oriented network involves a hybrid network containing an encoder and decoder network and a spatially variant recurrent neural network which is derived from the hazy process. In addition, we develop a multi-stage dehazing algorithm to further improve the accuracy by filtering haze residuals in a step-bystep fashion. To constrain the proposed network, we develop a dual composition loss, content-based pixel-wise loss and total variation constraint. We train the proposed network in an end-to-end manner and analyze its effect on image dehazing. Experimental results demonstrate that the proposed algorithm achieves favorable performance against state-of-the-art dehazing methods. Runde Li, Jinshan Pan, Zechao Li, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Dynamic Scene Deblurring by Depth Guided ModelabstractDynamic scene blur is usually caused by object motion, depth variation as well as camera shake. Most existing methods usually solve this problem using image segmentation or fully end-to-end trainable deep convolutional neural networks by considering different object motions or camera shakes. However, these algorithms are less effective when there exist depth variations. In this work, we propose a deep neural convolutional network that exploits the depth map for dynamic scene deblurring. Given a blurred image, we first extract the depth map and adopt a depth refinement network to restore the edges and structure in the depth map. To effectively exploit the depth map, we adopt the spatial feature transform layer to extract depth features and fuse with the image features through scaling and shifting. Our image deblurring network thus learns to restore a clear image under the guidance of the depth map. With substantial experiments and analysis, we show that the depth information is crucial to the performance of the proposed model. Finally, extensive quantitative and qualitative evaluations demonstrate that the proposed model performs favorably against the state-of-the-art dynamic scene deblurring approaches as well as conventional depth-based deblurring algorithms. Lerenhan Li, Jinshan Pan, Wei-Sheng Lai, Changxin Gao, Nong Sang, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Deep Video Deblurring Using Sharpness Features From ExemplarsabstractVideo deblurring is a challenging problem as the blur in videos is usually caused by camera shake, object motion, depth variation, etc. Existing methods usually impose handcrafted image priors or use end-to-end trainable networks to solve this problem. However, using image priors usually leads to highly non-convex problems while directly using end-to-end trainable networks in a regression generates over-smoothes details in the restored images. In this paper, we explore the sharpness features from exemplars to help the blur removal and details restoration. We first estimate optical flow to explore the temporal information which can help to make full use of neighboring information. Then, we develop an encoder and decoder network and explore the sharpness features from exemplars to guide the network for better image restoration. We train the proposed algorithm in an end-to-end manner and show that using sharpness features from exemplars can help blur removal and details restoration. Both quantitative and qualitative evaluations demonstrate that our method performs favorably against state-of-the-art approaches on the benchmark video deblurring datasets and real-world images. Xinguang Xiang, Hao Wei 0005, Jinshan Pan |
IEEE Trans. Image Process. | 3 |
| 2020 | Robust dense correspondence using deep convolutional features
Yang Liu 0119, Jinshan Pan, Zhixun Su, Kewei Tang |
Vis. Comput. | 2 |
| 2019 | Spatially Variant Linear Representation Models for Joint FilteringabstractJoint filtering mainly uses an additional guidance image as a prior and transfers its structures to the target image in the filtering process. Different from existing algorithms that rely on locally linear models or hand-designed objective functions to extract the structural information from the guidance image, we propose a new joint filter based on a spatially variant linear representation model (SVLRM), where the target image is linearly represented by the guidance image. However, the SVLRM leads to a highly ill-posed problem. To estimate the linear representation coefficients, we develop an effective algorithm based on a deep convolutional neural network (CNN). The proposed deep CNN (constrained by the SVLRM) is able to estimate the spatially variant linear representation coefficients which are able to model the structural information of both the guidance and input images. We show that the proposed algorithm can be effectively applied to a variety of applications, including depth/RGB image upsampling and restoration, flash/no-flash image deblurring, natural image denoising, scale-aware filtering, etc. Extensive experimental results demonstrate that the proposed algorithm performs favorably against state-of-the-art methods that have been specially designed for each task. Jinshan Pan, Jiangxin Dong, Jimmy S. J. Ren, Liang Lin 0004, Jinhui Tang 0001, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2019 | DAVANet: Stereo Deblurring With View AggregationabstractNowadays stereo cameras are more commonly adopted in emerging devices such as dual-lens smartphones and unmanned aerial vehicles. However, they also suffer from blurry images in dynamic scenes which leads to visual discomfort and hampers further image processing. Previous works have succeeded in monocular deblurring, yet there are few studies on deblurring for stereoscopic images. By exploiting the two-view nature of stereo images, we propose a novel stereo image deblurring network with Depth Awareness and View Aggregation, named DAVANet. In our proposed network, 3D scene cues from the depth and varying information from two views are incorporated, which help to remove complex spatially-varying blur in dynamic scenes. Specifically, with our proposed fusion network, we integrate the bidirectional disparities estimation and deblurring into a unified framework. Moreover, we present a large-scale multi-scene dataset for stereo deblurring, containing 20,637 blurry-sharp stereo image pairs from 135 diverse sequences and their corresponding bidirectional disparities. The experimental results on our dataset demonstrate that DAVANet outperforms state-of-the-art methods in terms of accuracy, speed, and model size. Shangchen Zhou, Jiawei Zhang 0002, Wangmeng Zuo, Haozhe Xie, Jinshan Pan, Jimmy S. J. Ren |
CVPR | 5 |
| 2019 | Learning Deep Priors for Image DehazingabstractImage dehazing is a well-known ill-posed problem, which usually requires some image priors to make the problem well-posed. We propose an effective iteration algorithm with deep CNNs to learn haze-relevant priors for image dehazing. We formulate the image dehazing problem as the minimization of a variational model with favorable data fidelity terms and prior terms to regularize the model. We solve the variational model based on the classical gradient descent method with built-in deep CNNs so that iteration-wise image priors for the atmospheric light, transmission map and clear image can be well estimated. Our method combines the properties of both the physical formation of image dehazing as well as deep learning approaches. We show that it is able to generate clear images as well as accurate atmospheric light and transmission maps. Extensive experimental results demonstrate that the proposed algorithm performs favorably against state-of-the-art methods in both benchmark datasets and real-world images. Yang Liu 0119, Jinshan Pan, Jimmy S. J. Ren, Zhixun Su |
ICCV | 2 |
| 2019 | Spatio-Temporal Filter Adaptive Network for Video DeblurringabstractVideo deblurring is a challenging task due to the spatially variant blur caused by camera shake, object motions, and depth variations, etc. Existing methods usually estimate optical flow in the blurry video to align consecutive frames or approximate blur kernels. However, they tend to generate artifacts or cannot effectively remove blur when the estimated optical flow is not accurate. To overcome the limitation of separate optical flow estimation, we propose a Spatio-Temporal Filter Adaptive Network (STFAN) for the alignment and deblurring in a unified framework. The proposed STFAN takes both blurry and restored images of the previous frame as well as blurry image of the current frame as input, and dynamically generates the spatially adaptive filters for the alignment and deblurring. We then propose the new Filter Adaptive Convolutional (FAC) layer to align the deblurred features of the previous frame with the current frame and remove the spatially variant blur from the features of the current frame. Finally, we develop a reconstruction network which takes the fusion of two transformed features to restore the clear frames. Both quantitative and qualitative evaluation results on the benchmark datasets and real-world videos demonstrate that the proposed algorithm performs favorably against state-of-the-art methods in terms of accuracy, speed as well as model size. Shangchen Zhou, Jiawei Zhang 0002, Jinshan Pan, Wangmeng Zuo, Haozhe Xie, Jimmy S. J. Ren |
ICCV | 3 |
| 2019 | Blind Image Deblurring via Deep Discriminative Priors
Lerenhan Li, Jinshan Pan, Wei-Sheng Lai, Changxin Gao, Nong Sang, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 2 |
| 2019 | Joint Face Hallucination and Deblurring via Structure Generation and Detail Enhancement
Yibing Song, Jiawei Zhang 0002, Lijun Gong, Shengfeng He, Linchao Bao, Jinshan Pan, Qingxiong Yang, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 6 |
| 2019 | Learning to Deblur Images with ExemplarsabstractHuman faces are one interesting object class with numerous applications. While significant progress has been made in the generic deblurring problem, existing methods are less effective for blurry face images. The success of the state-of-the-art image deblurring algorithms stems mainly from implicit or explicit restoration of salient edges for kernel estimation. However, existing methods are less effective as only few edges can be restored from blurry face images for kernel estimation. In this paper, we address the problem of deblurring face images by exploiting facial structures. We propose a deblurring algorithm based on an exemplar dataset without using coarse-to-fine strategies or heuristic edge selections. In addition, we develop a convolutional neural network to restore sharp edges from blurry images for deblurring. Extensive experiments against the state-of-the-art methods demonstrate the effectiveness of the proposed algorithm for deblurring face images. In addition, we show that the proposed algorithms can be applied to image deblurring for other object classes. Jinshan Pan, Wenqi Ren, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Robust Face Image Super-Resolution via Joint Learning of Subdivided Contextual ModelabstractIn this paper, we focus on restoring high-resolution facial images under noisy low-resolution scenarios. This problem is a challenging problem as the most important structures and details of captured facial images are missing. To address this problem, we propose a novel local patch-based face super-resolution (FSR) method via the joint learning of the contextual model. The contextual model is based on the topology consisting of contextual sub-patches, which provide more useful structural information than the commonly used local contextual structures due to the finer patch size. In this way, the contextual models are able to recover the missing local structures in target patches. In order to further strengthen the structural compensation function of contextual topology, we introduce the recognition feature as additional regularity. Based on the contextual model, we formulate the super-resolved procedure as a contextual joint representation with respect to the target patch and its adjacent patches. The high-resolution image is obtained by weighting contextual estimations. Both quantitative and qualitative validations show that the proposed method performs favorably against state-of-the-art algorithms. Liang Chen 0026, Jinshan Pan, Qing Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Learning a Discriminative Prior for Blind Image DeblurringabstractWe present an effective blind image deblurring method based on a data-driven discriminative prior. Our work is motivated by the fact that a good image prior should favor clear images over blurred ones. In this work, we formulate the image prior as a binary classifier which can be achieved by a deep convolutional neural network (CNN). The learned prior is able to distinguish whether an input image is clear or not. Embedded into the maximum a posterior (MAP) framework, it helps blind deblurring in various scenarios, including natural, face, text, and low-illumination images. However, it is difficult to optimize the deblurring method with the learned image prior as it involves a non-linear CNN. Therefore, we develop an efficient numerical approach based on the half-quadratic splitting method and gradient decent algorithm to solve the proposed model. Furthermore, the proposed model can be easily extended to non-uniform deblurring. Both qualitative and quantitative experimental results show that our method performs favorably against state-of-the-art algorithms as well as domain-specific image deblurring approaches. Lerenhan Li, Jinshan Pan, Wei-Sheng Lai, Changxin Gao, Nong Sang, Ming-Hsuan Yang 0001 |
CVPR | 2 |
| 2018 | Single Image Dehazing via Conditional Generative Adversarial NetworkabstractIn this paper, we present an algorithm to directly restore a clear image from a hazy image. This problem is highly ill-posed and most existing algorithms often use hand-crafted features, e.g., dark channel, color disparity, maximum contrast, to estimate transmission maps and then atmospheric lights. In contrast, we solve this problem based on a conditional generative adversarial network (cGAN), where the clear image is estimated by an end-to-end trainable neural network. Different from the generative network in basic cGAN, we propose an encoder and decoder architecture so that it can generate better results. To generate realistic clear images, we further modify the basic cGAN formulation by introducing the VGG features and an L1-regularized gradient prior. We also synthesize a hazy dataset including indoor and outdoor scenes to train and evaluate the proposed algorithm. Extensive experimental results demonstrate that the proposed method performs favorably against the state-of-the-art methods on both synthetic dataset and real world hazy images. Runde Li, Jinshan Pan, Zechao Li, Jinhui Tang 0001 |
CVPR | 2 |
| 2018 | LSTM Pose MachinesabstractWe observed that recent state-of-the-art results on single image human pose estimation were achieved by multistage Convolution Neural Networks (CNN). Notwithstanding the superior performance on static images, the application of these models on videos is not only computationally intensive, it also suffers from performance degeneration and flicking. Such suboptimal results are mainly attributed to the inability of imposing sequential geometric consistency, handling severe image quality degradation (e.g. motion blur and occlusion) as well as the inability of capturing the temporal correlation among video frames. In this paper, we proposed a novel recurrent network to tackle these problems. We showed that if we were to impose the weight sharing scheme to the multi-stage CNN, it could be re-written as a Recurrent Neural Network (RNN). This property decouples the relationship among multiple network stages and results in significantly faster speed in invoking the network for videos. It also enables the adoption of Long Short-Term Memory (LSTM) units between video frames. We found such memory augmented RNN is very effective in imposing geometric consistency among frames. It also well handles input quality degradation in videos while successfully stabilizes the sequential outputs. The experiments showed that our approach significantly outperformed current state-of-the-art methods on two large-scale video pose estimation benchmarks. We also explored the memory cells inside the LSTM and provided insights on why such mechanism would benefit the prediction for video-based pose estimations.1 Jimmy S. J. Ren, Zhouxia Wang, Wenxiu Sun, Jinshan Pan, Jiahao Pang, Liang Lin 0004 |
CVPR | 5 |
| 2018 | Learning Dual Convolutional Neural Networks for Low-Level VisionabstractIn this paper, we propose a general dual convolutional neural network (DualCNN) for low-level vision problems, e.g., super-resolution, edge-preserving filtering, deraining and dehazing. These problems usually involve the estimation of two components of the target signals: structures and details. Motivated by this, our proposed DualCNN consists of two parallel branches, which respectively recovers the structures and details in an end-to-end manner. The recovered structures and details can generate the target signals according to the formation model for each particular application. The DualCNN is a flexible framework for low-level vision tasks and can be easily incorporated into existing CNNs. Experimental results show that the DualCNN can be effectively applied to numerous low-level vision tasks with favorable performance against the state-of-the-art methods. Jinshan Pan, Sifei Liu, Deqing Sun, Jiawei Zhang 0002, Yang Liu 0119, Jimmy S. J. Ren, Zechao Li, Jinhui Tang 0001, Huchuan Lu, Yu-Wing Tai, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2018 | Gated Fusion Network for Single Image DehazingabstractIn this paper, we propose an efficient algorithm to directly restore a clear image from a hazy input. The proposed algorithm hinges on an end-to-end trainable neural network that consists of an encoder and a decoder. The encoder is exploited to capture the context of the derived input images, while the decoder is employed to estimate the contribution of each input to the final dehazed result using the learned representations attributed to the encoder. The constructed network adopts a novel fusion-based strategy which derives three inputs from an original hazy image by applying White Balance (WB), Contrast Enhancing (CE), and Gamma Correction (GC). We compute pixel-wise confidence maps based on the appearance differences between these different inputs to blend the information of the derived inputs and preserve the regions with pleasant visibility. The final dehazed image is yielded by gating the important features of the derived inputs. To train the network, we introduce a multi-scale approach such that the halo artifacts can be avoided. Extensive experimental results on both synthetic and real-world images demonstrate that the proposed algorithm performs favorably against the state-of-the-art algorithms. Wenqi Ren, Lin Ma 0002, Jiawei Zhang 0002, Jinshan Pan, Xiaochun Cao, Wei Liu 0005, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2018 | Dynamic Scene Deblurring Using Spatially Variant Recurrent Neural NetworksabstractDue to the spatially variant blur caused by camera shake and object motions under different scene depths, deblurring images captured from dynamic scenes is challenging. Although recent works based on deep neural networks have shown great progress on this problem, their models are usually large and computationally expensive. In this paper, we propose a novel spatially variant neural network to address the problem. The proposed network is composed of three deep convolutional neural networks (CNNs) and a recurrent neural network (RNN). RNN is used as a deconvolution operator performed on feature maps extracted from the input image by one of the CNNs. Another CNN is used to learn the weights for the RNN at every location. As a result, the RNN is spatially variant and could implicitly model the deblurring process with spatially variant kernels. The third CNN is used to reconstruct the final deblurred feature maps into restored image. The whole network is end-to-end trainable. Our analysis shows that the proposed network has a large receptive field even with a small model size. Quantitative and qualitative evaluations on public datasets demonstrate that the proposed method performs favorably against state-of-the-art algorithms in terms of accuracy, speed, and model size. Jiawei Zhang 0002, Jinshan Pan, Jimmy S. J. Ren, Yibing Song, Linchao Bao, Rynson W. H. Lau, Ming-Hsuan Yang 0001 |
CVPR | 2 |
| 2018 | Learning Data Terms for Non-blind Deblurring
Jiangxin Dong, Jinshan Pan, Deqing Sun, Zhixun Su, Ming-Hsuan Yang 0001 |
ECCV (11) | 2 |
| 2018 | Deep Non-Blind Deconvolution via Generalized Low-Rank ApproximationabstractIn this paper, we present a deep convolutional neural network to capture the inherent properties of image degradation, which can handle different kernels and saturated pixels in a unified framework. The proposed neural network is motivated by the low-rank property of pseudo-inverse kernels. We first compute a generalized low-rank approximation for a large number of blur kernels, and then use separable filters to initialize the convolutional parameters in the network. Our analysis shows that the estimated decomposed matrices contain the most essential information of the input kernel, which ensures the proposed network to handle various blurs in a unified framework and generate high-quality deblurring results. Experimental results on benchmark datasets with noise and saturated pixels demonstrate that the proposed algorithm performs favorably against state-of-the-art methods. Wenqi Ren, Jiawei Zhang 0002, Lin Ma 0002, Jinshan Pan, Xiaochun Cao, Wangmeng Zuo, Wei Liu 0005, Ming-Hsuan Yang 0001 |
NeurIPS | 4 |
| 2018 | Blind image deblurring using elastic-net based rank prior
Jinshan Pan, Zhixun Su, Songxin Liang |
Comput. Vis. Image Underst. | 2 |
| 2018 | Deblurring Images via Dark Channel PriorabstractWe present an effective blind image deblurring algorithm based on the dark channel prior. The motivation of this work is an interesting observation that the dark channel of blurred images is less sparse. While most patches in a clean image contain some dark pixels, this is not the case when they are averaged with neighboring ones by motion blur. This change in sparsity of the dark channel pixels is an inherent property of the motion blur process, which we prove mathematically and validate using image data. Enforcing sparsity of the dark channel thus helps blind deblurring in various scenarios such as natural, face, text, and low-illumination images. However, imposing sparsity of the dark channel introduces a non-convex non-linear optimization problem. In this work, we introduce a linear approximation to address this issue. Extensive experiments demonstrate that the proposed deblurring algorithm achieves the state-of-the-art results on natural images and performs favorably against methods designed for specific scenarios. In addition, we show that the proposed method can be applied to image dehazing. Jinshan Pan, Deqing Sun, Hanspeter Pfister, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Non-uniform motion deblurring with Kernel grid regularization
Ziyi Shen, Tingfa Xu, Jinshan Pan, Jie Guo 0004 |
Signal Process. Image Commun. | 3 |
| 2018 | Motion Blur Kernel Estimation via Deep LearningabstractThe success of the state-of-the-art deblurring methods mainly depends on the restoration of sharp edges in a coarse-to-fine kernel estimation process. In this paper, we propose to learn a deep convolutional neural network for extracting sharp edges from blurred images. Motivated by the success of the existing filtering-based deblurring methods, the proposed model consists of two stages: suppressing extraneous details and enhancing sharp edges. We show that the two-stage model simplifies the learning process and effectively restores sharp edges. Facilitated by the learned sharp edges, the proposed deblurring algorithm does not require any coarse-to-fine strategy or edge selection, thereby significantly simplifying kernel estimation and reducing computation load. Extensive experimental results on challenging blurry images demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods on both synthetic and real-world images in terms of visual quality and run-time. Xiangyu Xu 0002, Jinshan Pan, Yu-Jin Zhang, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Learning Fully Convolutional Networks for Iterative Non-blind DeconvolutionabstractIn this paper, we propose a fully convolutional network for iterative non-blind deconvolution. We decompose the non-blind deconvolution problem into image denoising and image deconvolution. We train a FCNN to remove noise in the gradient domain and use the learned gradients to guide the image deconvolution step. In contrast to the existing deep neural network based methods, we iteratively deconvolve the blurred images in a multi-stage framework. The proposed method is able to learn an adaptive image prior, which keeps both local (details) and global (structures) information. Both quantitative and qualitative evaluations on the benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art algorithms in terms of quality and speed. Jiawei Zhang 0002, Jinshan Pan, Wei-Sheng Lai, Rynson W. H. Lau, Ming-Hsuan Yang 0001 |
CVPR | 2 |
| 2017 | Blind Image Deblurring with Outlier HandlingabstractDeblurring images with outliers has attracted considerable attention recently. However, existing algorithms usually involve complex operations which increase the difficulty of blur kernel estimation. In this paper, we propose a simple yet effective blind image deblurring algorithm to handle blurred images with outliers. The proposed method is motivated by the observation that outliers in the blurred images significantly affect the goodness-of-fit in function approximation. Therefore, we propose an algorithm to model the data fidelity term so that the outliers have little effect on kernel estimation. The proposed algorithm does not require any heuristic outlier detection step, which is critical to the state-of-the-art blind deblurring methods for images with outliers. We analyze the relationship between the proposed algorithm and other blind deblurring methods with outlier handling and show how to estimate intermediate latent images for blur kernel estimation principally. We show that the proposed method can be applied to generic image deblurring as well as non-uniform deblurring. Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art blind image deblurring methods on both synthetic and real-world images. Jiangxin Dong, Jinshan Pan, Zhixun Su, Ming-Hsuan Yang 0001 |
ICCV | 2 |
| 2017 | Learning Discriminative Data Fitting Functions for Blind Image DeblurringabstractSolving blind image deblurring usually requires defining a data fitting function and image priors. While existing algorithms mainly focus on developing image priors for blur kernel estimation and non-blind deconvolution, only a few methods consider the effect of data fitting functions. In contrast to the state-of-the-art methods that use a single or a fixed data fitting term, we propose a data-driven approach to learn effective data fitting functions from a large set of motion blurred images with the associated ground truth blur kernels. The learned data fitting function facilitates estimating accurate blur kernels for generic scenes and domain-specific problems with corresponding image priors. In addition, we extend the learning approach for data fitting function to latent image restoration and nonuniform deblurring. Extensive experiments on challenging motion blurred images demonstrate the proposed algorithm performs favorably against the state-of-the-art methods. Jinshan Pan, Jiangxin Dong, Yu-Wing Tai, Zhixun Su, Ming-Hsuan Yang 0001 |
ICCV | 1 |
| 2017 | Video Deblurring via Semantic Segmentation and Pixel-Wise Non-linear KernelabstractVideo deblurring is a challenging problem as the blur is complex and usually caused by the combination of camera shakes, object motions, and depth variations. Optical flow can be used for kernel estimation since it predicts motion trajectories. However, the estimates are often inaccurate in complex scenes at object boundaries, which are crucial in kernel estimation. In this paper, we exploit semantic segmentation in each blurry frame to understand the scene contents and use different motion models for image regions to guide optical flow estimation. While existing pixel-wise blur models assume that the blur kernel is the same as optical flow during the exposure time, this assumption does not hold when the motion blur trajectory at a pixel is different from the estimated linear optical flow. We analyze the relationship between motion blur trajectory and optical flow, and present a novel pixel-wise non-linear kernel model to account for motion blur. The proposed blur model is based on the non-linear optical flow, which describes complex motion blur more effectively. Extensive experiments on challenging blurry videos demonstrate the proposed algorithm performs favorably against the state-of-the-art methods. Wenqi Ren, Jinshan Pan, Xiaochun Cao, Ming-Hsuan Yang 0001 |
ICCV | 2 |
| 2017 | Learning to Super-Resolve Blurry Face and Text ImagesabstractWe present an algorithm to directly restore a clear highresolution image from a blurry low-resolution input. This problem is highly ill-posed and the basic assumptions for existing super-resolution methods (requiring clear input) and deblurring methods (requiring high-resolution input) no longer hold. We focus on face and text images and adopt a generative adversarial network (GAN) to learn a category-specific prior to solve this problem. However, the basic GAN formulation does not generate realistic high-resolution images. In this work, we introduce novel training losses that help recover fine details. We also present a multi-class GAN that can process multi-class image restoration tasks, i.e., face and text images, using a single generator network. Extensive experiments demonstrate that our method performs favorably against the state-of-the-art methods on both synthetic and real-world images at a lower computational cost. Xiangyu Xu 0002, Deqing Sun, Jinshan Pan, Hanspeter Pfister, Ming-Hsuan Yang 0001 |
ICCV | 3 |
| 2017 | Deep feature matching for dense correspondenceabstractImage matching is a challenging problem as different views often undergo significant appearance changes caused by deformation, abrupt motion, and occlusion. In this paper, we explore features extracted from convolutional neural networks to help the estimation of image matching so that dense pixel correspondence can be built. As the deep features are able to describe the image structures, the matching method based on these features is able to match across different scenes and/or object appearances. We analyze the deep features and compare them with other robust features, e.g., SIFT. Extensive experiments on 5 datasets demonstrate the proposed algorithm performs favorably against the state-of-the-art methods in terms of visually matching quality and accuracy. Yang Liu 0119, Jinshan Pan, Zhixun Su |
ICIP | 2 |
| 2017 | L0-Regularized Intensity and Gradient Prior for Deblurring Text Images and Beyondabstract-regularized prior based on intensity and gradient for text image deblurring. The proposed image prior is based on distinctive properties of text images, with which we develop an efficient optimization algorithm to generate reliable intermediate results for kernel estimation. The proposed algorithm does not require any heuristic edge selection methods, which are critical to the state-of-the-art edge-based deblurring methods. We discuss the relationship with other edge-based deblurring methods and present how to select salient edges more principally. For the final latent image restoration step, we present an effective method to remove artifacts for better deblurred results. We show the proposed algorithm can be extended to deblur natural images with complex scenes and low illumination, as well as non-uniform deblurring. Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art image deblurring methods. Jinshan Pan, Zhixun Su, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Blur kernel estimation via salient edges and low rank prior for blind image deblurring
Jiangxin Dong, Jinshan Pan, Zhixun Su |
Signal Process. Image Commun. | 2 |
| 2016 | Soft-Segmentation Guided Object Motion DeblurringabstractObject motion blur is a challenging problem as the foreground and the background in the scenes undergo different types of image degradation due to movements in various directions and speed. Most object motion deblurring methods address this problem by segmenting blurred images into regions where different kernels are estimated and applied for restoration. Segmentation on blurred images is difficult due to ambiguous pixels between regions, but it plays an important role for object motion deblurring. To address these problems, we propose a novel model for object motion deblurring. The proposed model is developed based on a maximum a posterior formulation in which soft-segmentation is incorporated for object layer estimation. We propose an efficient algorithm to jointly estimate object segmentation and camera motion where each layer can be deblurred well under the guidance of the soft-segmentation. Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art object motion deblurring methods on challenging scenarios. Jinshan Pan, Zhixun Su, Hsin-Ying Lee 0001, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2016 | Robust Kernel Estimation with Outliers Handling for Image DeblurringabstractEstimating blur kernels from real world images is a challenging problem as the linear image formation assumption does not hold when significant outliers, such as saturated pixels and non-Gaussian noise, are present. While some existing non-blind deblurring algorithms can deal with outliers to a certain extent, few blind deblurring methods are developed to well estimate the blur kernels from the blurred images with outliers. In this paper, we present an algorithm to address this problem by exploiting reliable edges and removing outliers in the intermediate latent images, thereby estimating blur kernels robustly. We analyze the effects of outliers on kernel estimation and show that most state-of-the-art blind deblurring methods may recover delta kernels when blurred images contain significant outliers. We propose a robust energy function which describes the properties of outliers for the final latent image restoration. Furthermore, we show that the proposed algorithm can be applied to improve existing methods to deblur images with outliers. Extensive experiments on different kinds of challenging blurry images with significant amount of outliers demonstrate the proposed algorithm performs favorably against the state-of-the-art methods. Jinshan Pan, Zhouchen Lin, Zhixun Su, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2016 | Blind Image Deblurring Using Dark Channel PriorabstractWe present a simple and effective blind image deblurring method based on the dark channel prior. Our work is inspired by the interesting observation that the dark channel of blurred images is less sparse. While most image patches in the clean image contain some dark pixels, these pixels are not dark when averaged with neighboring highintensity pixels during the blur process. This change in the sparsity of the dark channel is an inherent property of the blur process, which we both prove mathematically and validate using training data. Therefore, enforcing the sparsity of the dark channel helps blind deblurring on various scenarios, including natural, face, text, and low-illumination images. However, sparsity of the dark channel introduces a non-convex non-linear optimization problem. We introduce a linear approximation of the min operator to compute the dark channel. Our look-up-table-based method converges fast in practice and can be directly extended to non-uniform deblurring. Extensive experiments show that our method achieves state-of-the-art results on deblurring natural images and compares favorably methods that are well-engineered for specific scenarios. Jinshan Pan, Deqing Sun, Hanspeter Pfister, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2016 | Learning Recursive Filters for Low-Level Vision via a Hybrid Neural Network
Sifei Liu, Jinshan Pan, Ming-Hsuan Yang 0001 |
ECCV (4) | 2 |
| 2016 | Single Image Dehazing via Multi-scale Convolutional Neural Networks
Wenqi Ren, Si Liu 0001, Hua Zhang 0008, Jinshan Pan, Xiaochun Cao, Ming-Hsuan Yang 0001 |
ECCV (2) | 4 |
| 2016 | Image Deblurring via Enhanced Low-Rank PriorabstractLow-rank matrix approximation has been successfully applied to numerous vision problems in recent years. In this paper, we propose a novel low-rank prior for blind image deblurring. Our key observation is that directly applying a simple low-rank model to a blurry input image significantly reduces the blur even without using any kernel information, while preserving important edge information. The same model can be used to reduce blur in the gradient map of a blurry input. Based on these properties, we introduce an enhanced prior for image deblurring by combining the low rank prior of similar patches from both the blurry image and its gradient map. We employ a weighted nuclear norm minimization method to further enhance the effectiveness of low-rank prior for image deblurring, by retaining the dominant edges and eliminating fine texture and slight edges in intermediate images, allowing for better kernel estimation. In addition, we evaluate the proposed enhanced low-rank prior for both the uniform and the non-uniform deblurring. Quantitative and qualitative experimental evaluations demonstrate that the proposed algorithm performs favorably against the state-of-the-art deblurring methods. Wenqi Ren, Xiaochun Cao, Jinshan Pan, Xiaojie Guo 0001, Wangmeng Zuo, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | L0-Regularized Object Representation for Visual Tracking
Jinshan Pan, Jongwoo Lim, Zhixun Su, Ming-Hsuan Yang 0001 |
BMVC | 1 |
| 2014 | Deblurring Text Images via L0-Regularized Intensity and Gradient PriorabstractWe propose a simple yet effective L0-regularized prior based on intensity and gradient for text image deblurring. The proposed image prior is motivated by observing distinct properties of text images. Based on this prior, we develop an efficient optimization method to generate reliable intermediate results for kernel estimation. The proposed method does not require any complex filtering strategies to select salient edges which are critical to the state-of-the-art deblurring algorithms. We discuss the relationship with other deblurring algorithms based on edge selection and provide insight on how to select salient edges in a more principled way. In the final latent image restoration step, we develop a simple method to remove artifacts and render better deblurred images. Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art text image deblurring methods. In addition, we show that the proposed method can be effectively applied to deblur low-illumination images. Jinshan Pan, Zhixun Su, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2014 | Deblurring Face Images with Exemplars
Jinshan Pan, Zhixun Su, Ming-Hsuan Yang 0001 |
ECCV (7) | 1 |
| 2014 | Motion blur kernel estimation via salient edges and low rank priorabstractBlind image deblurring, i.e., estimating a blur kernel from a single input blurred image is a severely ill-posed problem. In this paper, we show how to effectively apply low rank prior to blind image deblurring and then propose a new algorithm which combines salient edges and low rank prior. Salient edges provide reliable edge information for kernel estimation, while low rank prior provides data-authentic priors for the latent image. When estimating the kernel, the salient edges are extracted from an intermediate latent image solved by combining the predicted edges and low rank prior, which help preserve more useful edges than previous deconvolution methods do. By solving the blind image deblurring problem in this fashion, high-quality blur kernels can be obtained. Extensive experiments testify to the superiority of the proposed method over state-of-the-art algorithms, both qualitatively and quantitatively. Jinshan Pan, Risheng Liu, Zhixun Su, Guili Liu |
ICME | 1 |
| 2013 | Saliency detection based on an edge-preserving filterabstractHow to detect visual salient regions is a challenging problem in computer vision. Recently, saliency detection methods that use boundaries or convex hulls under Bayesian framework have attracted lots of attention. Although these methods achieve state-of-the-art results, there still exist some limitations, e.g., the background will get highlighted when the initial convex hulls are not good enough. This paper presents a new algorithm that retains the advantages of such saliency maps while overcoming their shortcomings. First, the initial convex hull is improved by the image matting model which can be efficiently solved by an edge-preserving filter. Second, a more accurate prior map can be obtained by the improved convex hull. Third, the final convex hull is further refined by an edge-preserving filter to compute the observation likelihood. Finally, the Bayesian framework is employed to compute the saliency map. Extensive experiments compared with state-of-the-art saliency detection algorithms demonstrate the effectiveness of our method. Jinshan Pan, Zhixun Su, Maoran Bian, Risheng Liu |
ICIP | 1 |
| 2013 | Kernel estimation from salient structure for robust motion deblurring
Jinshan Pan, Risheng Liu, Zhixun Su, Xianfeng Gu |
Signal Process. Image Commun. | 1 |
| 2013 | Fast l0 -Regularized Kernel Estimation for Robust Motion DeblurringabstractBlind image deblurring is a challenging problem in computer vision and image processing. In this paper, we propose a newl0-regularized approach to estimate a blur kernel from a single blurred image by regularizing the sparsity property of natural images. Furthermore, by introducing an adaptive structure map in the deblurring process, our method is able to restore useful salient edges for kernel estimation. Finally, we propose an efficient algorithm which can solve the proposed model efficiently. Extensive experiments compared with state-of-the-art blind deblurring methods demonstrate the effectiveness of the proposed method. Jinshan Pan, Zhixun Su |
IEEE Signal Process. Lett. | 1 |
| 2011 | Robust optical flow estimation based on brightness correction fieldsabstractOptical flow estimation is still an important task in computer vision with many interesting applications. However, the results obtained by most of the optical flow techniques are affected by motion discontinuities or illumination changes. In this paper, we introduce a brightness correction field combined with a gradient constancy constraint to reduce the sensibility to brightness changes between images to be estimated. The advantage of this brightness correction field is its simplicity in terms of computational complexity and implementation. By analyzing the deficiencies of the traditional total variation regularization term in weakly textured areas, we also adopt a structure-adaptive regularization based on the robust Huber norm to preserve motion discontinuities. Finally, the proposed energy functional is minimized by solving its corresponding Euler-Lagrange equation in a more effective multi-resolution scheme, which integrates the twice downsampling strategy with a support-weight median filter. Numerous experiments show that our method is more effective and produces more accurate results for optical flow estimation. Wei Wang 0335, Zhixun Su, Jinshan Pan, Ye Wang 0023, Riming Sun |
J. Zhejiang Univ. Sci. C | 3 |