Jiawei Zhang 0002

dblp:10/239-2 · DBLP profile ↗
← Back
53ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0002-2292-4592ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 4 first-author · 20 since 2021Artificial intelligence and machine learning · 38 · 3 first-author · 17 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PSTAN: A JND-Aware Pairwise Spatio-Temporal Alignment Network for Compressed Videos Quality Enhancement
abstract
Compressed video quality enhancement (CVQE) is crucial for mitigating compression artifacts and improving perceptual visual quality, especially under diverse quantization parameters (QPs) and motion patterns. However, many existing approaches insufficiently exploit long-range temporal dependencies, and their reliance on QP-specific training often leads to limited robustness when compression conditions change. In this work, we propose a just noticeable difference (JND)-aware and perception-driven learning framework for CVQE, termed the Pairwise Spatio-Temporal Alignment Network (PSTAN). PSTAN incorporates perceptual priors primarily through a JND-guided training paradigm rather than relying solely on architectural modifications, where learning is driven by perceptuallypoorvideo segments identified in the VideoSet dataset. This strategy alleviates the reliance on QP-specific supervision and promotes more stable enhancement behavior across varying compression conditions. To effectively capture temporal dependencies, PSTAN employs a pairwise spatio-temporal interaction mechanism that models each reference-target frame pair independently, enabling adaptive utilization of both nearby and distant frames. In addition, a transformer-based alignment module combining temporal mutual attention with cascaded deformable convolution is introduced to handle complex and large motions. Extensive experiments on VideoSet, MFQE 2.0 and our constructed HEVC-comperssed dataset show that PSTAN achieves consistent improvements over state-of-the-art CVQE methods in both objective and perceptual quality metrics. The code of this work is available at https://github.com/leryong/PSTAN.git.
Yuan Yuan 0007, Eryong Li, Jiawei Zhang 0002, Jinchang Ren, Xu Lu 0002
IEEE Trans. Circuits Syst. Video Technol.3
2025 A Diffusion-Based Framework for Occluded Object Movement
abstract
Seamlessly moving objects within a scene is a common requirement for image editing, but it is still a challenge for existing editing methods. Especially for real-world images, the occlusion situation further increases the difficulty. The main difficulty is that the occluded portion needs to be completed before movement can proceed. To leverage the real-world knowledge embedded in the pre-trained diffusion models, we propose a Diffusion-based framework specifically designed for Occluded Object Movement, named DiffOOM. The proposed DiffOOM consists of two parallel branches that perform object de-occlusion and movement simultaneously. The de-occlusion branch utilizes a background color-fill strategy and a continuously updated object mask to focus the diffusion process on completing the obscured portion of the target object. Concurrently, the movement branch employs latent optimization to place the completed object in the target location and adopts local text-conditioned guidance to integrate the object into new surroundings appropriately. Extensive evaluations across various metrics demonstrate the superior performance of our method, which is further validated by a comprehensive user study.
Zheng-Peng Duan, Jiawei Zhang 0002, Zheng Lin 0005, Chunle Guo, Dongqing Zou, Jimmy S. J. Ren, Chongyi Li
AAAI2
2025 DiffRetouch: Using Diffusion to Retouch on the Shoulder of Experts
abstract
Image retouching aims to enhance the visual quality of photos. Considering the different aesthetic preferences of users, the target of retouching is subjective. However, current retouching methods mostly adopt deterministic models, which not only neglects the style diversity in the expert-retouched results and tends to learn an average style during training, but also lacks sample diversity during inference. In this paper, we propose a diffusion-based method, named DiffRetouch. Thanks to the excellent distribution modeling ability of diffusion, our method can capture the complex fine-retouched distribution covering various visual-pleasing styles in the training data. Moreover, four image attributes are made adjustable to provide a user-friendly editing mechanism. By adjusting these attributes in specified ranges, users are allowed to customize preferred styles within the learned fine-retouched distribution. Additionally, the affine bilateral grid and contrastive learning scheme are introduced to handle the problem of texture distortion and control insensitivity respectively. Extensive experiments have demonstrated the superior performance of our method on visually appealing and sample diversity.
Zheng-Peng Duan, Jiawei Zhang 0002, Zheng Lin 0005, Xin Jin 0005, Xundong Wang, Dongqing Zou, Chunle Guo, Chongyi Li
AAAI2
2025 DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution
abstract
Large-scale pre-trained diffusion models are becoming increasingly popular in solving the Real-World Image Super-Resolution (Real-ISR) problem because of their rich generative priors. The recent development of diffusion transformer (DiT) has witnessed overwhelming performance over the traditional UNet-based architecture in image generation, which also raises the question: Can we adopt the advanced DiT-based diffusion model for Real-ISR? To this end, we propose our DiT4SR, one of the pioneering works to tame the large-scale DiT model for Real-ISR. Instead of directly injecting embeddings extracted from low-resolution (LR) images like ControlNet, we integrate the LR embeddings into the original attention mechanism of DiT, allowing for the bidirectional flow of information between the LR latent and the generated latent. The sufficient interaction of these two streams allows the LR stream to evolve with the diffusion process, producing progressively refined guidance that better aligns with the generated latent at each diffusion step. Additionally, the LR guidance is injected into the generated latent via a cross-stream convolution layer, compensating for DiT's limited ability to capture local information. These simple but effective designs endow the DiT model with superior performance in Real-ISR, which is demonstrated by extensive experiments. Project Page: https://adam-duan.github.io/projects/dit4sr/.
Zheng-Peng Duan, Jiawei Zhang 0002, Xin Jin 0005, Zheng Xiong, Dongqing Zou, Jimmy S. J. Ren, Chunle Guo, Chongyi Li
ICCV2
2025 Event-Guided HDR Reconstruction with Diffusion Priors
Yixin Yang 0008, Jiawei Zhang 0002, Yunxuan Wei, Dongqing Zou, Jimmy S. J. Ren, Boxin Shi
ICCV2
2025 Multimodal Re-Ranking for Heterogeneous Face Re-Identification
abstract
Heterogeneous face re-identification (Re-ID), aiming to match low-quality faces captured by disjoint visible light (VIS) and near-infrared (NIR) cameras, has become a critical application in video surveillance. However, the domain discrepancy between the NIR-VIS faces degrades the Re-ID performance. To solve this problem, this paper proposes a multimodal re-ranking method including two stages. Firstly, we utilize the VIS-NIR face bi-directional modality transformation based on the positive and negative samples separate training strategy to reduce domain discrepancy and generate the multimodal ranking lists of face Re-ID with complementarities. Secondly, we propose linear and nonlinear multimodal ranking lists fusion strategies based on single-modal and multi-modal k-reciprocal nearest neighbors (K-RNNs) to obtain a more accurate fused ranking list for face Re-ID. Extensive experiments on heterogeneous face datasets demonstrate the superior performance of our method over existing methods.
Wenqin Song, Jiawei Zhang 0002, Zhen Han 0002, Yunfeng Xue, Xihao Wang, Zhongyuan Wang 0001
ICIP3
2025 UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
abstract
Large Language Model (LLM) agents equipped with external tools have become increasingly powerful for complex tasks such as web shopping, automated email replies, and financial trading. However, these advancements amplify the risks of adversarial attacks, especially when agents can access sensitive external functionalities. Nevertheless, manipulating LLM agents into performing targeted malicious actions or invoking specific tools remains challenging, as these agents extensively reason or plan before executing final actions. In this work, we present UDora, a unified red teaming framework designed for LLM agents that dynamically hijacks the agent's reasoning processes to compel malicious behavior. Specifically, UDora first generates the model’s reasoning trace for the given task, then automatically identifies optimal points within this trace to insert targeted perturbations. The resulting perturbed reasoning is then used as a surrogate response for optimization. By iteratively applying this process, the LLM agent will then be induced to undertake designated malicious actions or to invoke specific malicious tools. Our approach demonstrates superior effectiveness compared to existing methods across three LLM agent datasets. The code is available at https://github.com/AI-secure/UDora.
Jiawei Zhang 0002, Bo Li 0026
ICML1
2025 DeblurDiff: Real-Word Image Deblurring with Generative Diffusion Models
abstract
Diffusion models have achieved significant progress in image generation and the pre-trained Stable Diffusion (SD) models are helpful for image deblurring by providing clear image priors. However, directly using a blurry image or a pre-deblurred one as a conditional control for SD will either hinder accurate structure extraction or make the results overly dependent on the deblurring network. In this work, we propose a Latent Kernel Prediction Network (LKPN) to achieve robust real-world image deblurring. Specifically, we co-train the LKPN in the latent space with conditional diffusion. The LKPN learns a spatially variant kernel to guide the restoration of sharp images in the latent space. By applying element-wise adaptive convolution (EAC), the learned kernel is utilized to adaptively process the blurry feature, effectively preserving the information of the blurry input. This process thereby more effectively guides the generative process of SD, enhancing both the deblurring efficacy and the quality of detail reconstruction. Moreover, the results at each diffusion step are utilized to iteratively estimate the kernels in LKPN to better restore the sharp latent by EAC in the subsequent step. This iterative refinement enhances the accuracy and robustness of the deblurring process. Extensive experimental results demonstrate that the proposed method outperforms state-of-the-art image deblurring methods on both benchmark and real-world images.
Lingshun Kong, Jiawei Zhang 0002, Dongqing Zou, Fu Lee Wang, Jimmy S. J. Ren, Xiaohe Wu, Jiangxin Dong, Jinshan Pan
NeurIPS2
2024 An Efficient Transformer For Demosaicing Via Compressed Multi-Branch Attention Mechanism
abstract
Recent demosaicing approaches are not effective and efficient enough as they do not make full use of these two factors: (1) Capturing long-range spatial dependencies effiently. (2) Reducing the computational costs when utilizing channel attention. To take them into consideration, we propose an Efficient Compressed Multibranch Transformer (ECMT) to handle demosaicing problem on various color filter arrays (CFA). There are two key specifically designed components: First, a Local-Global Spatial-wise Multi-head Self-Attention (LGSMSA) captures both spatial long-range dependencies and local context details by an efficient manner. Second, an Inner-Compressed Channel-wise Multi-head Self-Attention (ICCMSA) clusters closely related channels into the same group and then calculates channel-wise MSA within each group to reduce the computational cost. Experimental results on both real and synthetic datasets show that our ECMT achieves promising reconstruction quality in Bayer and QuadBayer CFA while requiring cheaper computational, memory costs and less inference time compared to state-of-the-art methods.
Fanqing Meng, Jiawei Zhang 0002, Feng Zhang 0047
ICASSP4
2024 Deep Richardson-Lucy Deconvolution for Low-Light Image Deblurring
Liang Chen 0026, Jiawei Zhang 0002, Yunxuan Wei, Faming Fang, Jimmy S. J. Ren, Jinshan Pan
Int. J. Comput. Vis.2
2024 Learning Diverse Tone Styles for Image Retouching
abstract
Image retouching, aiming to regenerate the visually pleasing renditions of given images, is a subjective task where the users are with different aesthetic sensations. Most existing methods adopt a deterministic model to learn the retouching style from a specific expert, making it less flexible to meet diverse subjective preferences. Besides, the intrinsic diversity of an expert due to the targeted processing of different images is also deficiently described. To circumvent such issues, we propose to learn diverse image retouching with normalizing flow-based architectures. Unlike current flow-based methods which directly generate the output image, we argue that learning in a one-dimensional style space could 1) disentangle the retouching styles from the image content, 2) lead to a stable style presentation form, and 3) avoid the spatial disharmony effects. For obtaining meaningful image tone style representations, a joint-training pipeline is delicately designed, which is composed of a style encoder, a conditional RetouchNet, and the image tone style normalizing flow (TSFlow) module. In particular, the style encoder predicts the target style representation of an input image, which serves as the conditional information in the RetouchNet for retouching, while the TSFlow maps the style representation vector into a Gaussian distribution in the forward pass. After training, the TSFlow can generate diverse image tone style vectors by sampling from the Gaussian distribution. Extensive experiments on MIT-Adobe FiveK and PPR10K datasets show that our proposed method performs favorably against state-of-the-art methods and is effective in generating diverse results to satisfy different human aesthetic preferences. Source codeterministic and pre-trained models are publicly available at https://github.com/SSRHeart/TSFlow.
Haolin Wang 0004, Jiawei Zhang 0002, Ming Liu 0018, Xiaohe Wu, Wangmeng Zuo
IEEE Trans. Image Process.2
2024 Analysis and Benchmarking of Extending Blind Face Image Restoration to Videos
abstract
Recent progress in blind face restoration has resulted in producing high-quality restored results for static images. However, efforts to extend these advancements to video scenarios have been minimal, partly because of the absence of benchmarks that allow for a comprehensive and fair comparison. In this work, we first present a fair evaluation benchmark, in which we first introduce a Real-world Low-Quality Face Video benchmark (RFV-LQ), evaluate several leading image-based face restoration algorithms, and conduct a thorough systematical analysis of the benefits and challenges associated with extending blind face image restoration algorithms to degraded face videos. Our analysis identifies several key issues, primarily categorized into two aspects: significant jitters in facial components and noise-shape flickering between frames. To address these issues, we propose a Temporal Consistency Network (TCN) cooperated with alignment smoothing to reduce jitters and flickers in restored videos. TCN is a flexible component that can be seamlessly plugged into the most advanced face image restoration algorithms, ensuring the quality of image-based restoration is maintained as closely as possible. Extensive experiments have been conducted to evaluate the effectiveness and efficiency of our proposed TCN and alignment smoothing operation.
Zhouxia Wang, Jiawei Zhang 0002, Xintao Wang 0002, Tianshui Chen, Ying Shan, Wenping Wang 0001, Ping Luo 0002
IEEE Trans. Image Process.2
2023 Joint Demosaicing and Denoising with Gradient Guidance in Quad Bayer CFA
abstract
In this study, we introduce a challenging task called Joint Demosaicking and Denoising in Quad Bayer CFA (JDD-QBC). Inspired by the effectiveness of gradient prior in previous tasks, we design a gradient prior extraction algorithm to extract robust gradient profiles directly from Quad Bayer raw images. By leveraging contextual information, our algorithm overcomes the limitations of previous gradient extraction algorithms. Building on this prior, we propose an end-to-end JDD-QBC framework called Gradient Guidance Network (G2-Net), consisting of two specifically designed components: a Gradient Refinement Branch (GRB) to transform the extracted gradient prior, and a Multi-Scale Backbone Branch (MSB2) to recover all missing values under the guidance of transformed prior. Extensive experiments demonstrate our proposed G2-Net outperforms state-of-the-art methods across a wide range of noise levels while requiring 50% less memory cost.
Jiawei Zhang 0002, Feng Zhang 0047, Jimmy S. J. Ren
ICIP3
2023 Generic Attention-model Explainability by Weighted Relevance Accumulation
abstract
Attention-based Transformer models have achieved remarkable progress in multi-modal tasks, such as visual question answering. The explainability of attention-based methods has recently attracted wide interest as it can explain the inner changes of attention tokens by accumulating relevancy across attention layers. Current methods simply update relevancy by equally accumulating the token relevancy before and after the attention processes. However, the importance of token values is usually different during relevance accumulation.In this paper, we propose a weighted relevancy strategy, which takes the importance of token values into consideration, to reduce distortion when equally accumulating relevance. To evaluate our method, we propose a unified CLIP-based two-stage model, named CLIPmapper, to process Vision-and-Language tasks through CLIP encoder and a following mapper. CLIPmapper consists of self-attention, cross-attention, single-modality, and cross-modality attention, thus it is more suitable for evaluating our generic explainability method. Extensive perturbation tests on visual question answering and image captioning tasks validate that our explainability method outperforms existing methods.
Yiming Huang 0002, Aozhe Jia, Xiaodan Zhang 0003, Jiawei Zhang 0002
MMAsia4
2023 RestoreFormer++: Towards Real-World Blind Face Restoration From Undegraded Key-Value Pairs
abstract
Blind face restoration aims at recovering high-quality face images from those with unknown degradations. Current algorithms mainly introduce priors to complement high-quality details and achieve impressive progress. However, most of these algorithms ignore abundant contextual information in the face and its interplay with the priors, leading to sub-optimal performance. Moreover, they pay less attention to the gap between the synthetic and real-world scenarios, limiting the robustness and generalization to real-world applications. In this work, we propose RestoreFormer++, which on the one hand introduces fully-spatial attention mechanisms to model the contextual information and the interplay with the priors, and on the other hand, explores an extending degrading model to help generate more realistic degraded face images to alleviate the synthetic-to-real-world gap. Compared with current algorithms, RestoreFormer++ has several crucial benefits. First, instead of using a multi-head self-attention mechanism like the traditional visual transformer, we introduce multi-head cross-attention over multi-scale features to fully explore spatial interactions between corrupted information and high-quality priors. In this way, it can facilitate RestoreFormer++ to restore face images with higher realness and fidelity. Second, in contrast to the recognition-oriented dictionary, we learn a reconstruction-oriented dictionary as priors, which contains more diverse high-quality facial details and better accords with the restoration target. Third, we introduce an extending degrading model that contains more realistic degraded scenarios for training data synthesizing, and thus helps to enhance the robustness and generalization of our RestoreFormer++ model. Extensive experiments show that RestoreFormer++ outperforms state-of-the-art algorithms on both synthetic and real-world datasets.
Zhouxia Wang, Jiawei Zhang 0002, Tianshui Chen, Wenping Wang 0001, Ping Luo 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 RestoreFormer: High-Quality Blind Face Restoration from Undegraded Key-Value Pairs
abstract
Blind face restoration is to recover a high-quality face image from unknown degradations. As face image contains abundant contextual information, we propose a method, RestoreFormer, which explores fully-spatial attentions to model contextual information and surpasses existing works that use local operators. RestoreFormer has several benefits compared to prior arts. First, unlike the conventional multi-head self-attention in previous Vision Transformers (ViTs), RestoreFormer incorporates a multi-head cross-attention layer to learn fully-spatial interactions between corrupted queries and high-quality key-value pairs. Second, the key-value pairs in ResotreFormer are sampled from a reconstruction-oriented high-quality dictionary, whose elements are rich in high-quality facial features specifically aimed for face reconstruction, leading to superior restoration results. Third, RestoreFormer outperforms advanced state-of-the-art methods on one synthetic dataset and three real-world datasets, as well as produces images with better visual quality. Code is available at https://github.com/wzhouxiff/RestoreFormer.git.
Zhouxia Wang, Jiawei Zhang 0002, Runjian Chen, Wenping Wang 0001, Ping Luo 0002
CVPR2
2022 Self-Guided Video Super-Resolution Based on a Fast Deformable ConvGRU Model
abstract
Video super-resolution (VSR) aims at recovering a natural and realistic high-resolution (HR) video frame from the corre-sponding low-resolution (LR) counterpart and its consecutive neighboring frames. The challenge is how to make full use of spatio-temporal coherence among the input LR frames. In this work, we propose a self-guided deformable convolutional gated recurrent unit (GRU) framework for VSR. Specifically, convolutional GRU can efficiently extract temporal features of the input LR frames. Deformable convolution (DConv) is utilized to spatially align the hidden states of GRU cells with the input feature maps. Moreover, we argue that the reference LR frame itself is efficient to guide the aggregated features learning frame-specific feature maps, which are then used to generate rich and more realistic textures towards the corresponding HR frame. Extensive experimental results on benchmark datasets demonstrate that the proposed framework achieves better performance than state-of-the-art methods and has higher model efficiency.
Jingming Chen, Yuan Yuan 0007, Jiawei Zhang 0002, Jianping Luo
ICME3
2022 Dual Convolutional Neural Networks for Low-Level Vision
Jinshan Pan, Deqing Sun, Jiawei Zhang 0002, Jinhui Tang 0001, Jian Yang 0003, Yu-Wing Tai, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.3
2022 Deblurring Dynamic Scenes via Spatially Varying Recurrent Neural Networks
abstract
Deblurring images captured in dynamic scenes is challenging as the motion blurs are spatially varying caused by camera shakes and object movements. In this paper, we propose a spatially varying neural network to deblur dynamic scenes. The proposed model is composed of three deep convolutional neural networks (CNNs) and a recurrent neural network (RNN). The RNN is used as a deconvolution operator on feature maps extracted from the input image by one of the CNNs. Another CNN is used to learn the spatially varying weights for the RNN. As a result, the RNN is spatial-aware and can implicitly model the deblurring process with spatially varying kernels. To better exploit properties of the spatially varying RNN, we develop both one-dimensional and two-dimensional RNNs for deblurring. The third component, based on a CNN, reconstructs the final deblurred feature maps into a restored image. In addition, the whole network is end-to-end trainable. Quantitative and qualitative evaluations on benchmark datasets demonstrate that the proposed method performs favorably against the state-of-the-art deblurring algorithms.
Wenqi Ren, Jiawei Zhang 0002, Jinshan Pan, Sifei Liu, Jimmy S. J. Ren, Junping Du 0001, Xiaochun Cao, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Deep Dynamic Scene Deblurring From Optical Flow
abstract
Deblurring can not only provide visually more pleasant pictures and make photography more convenient, but also can improve the performance of objection detection as well as tracking. However, removing dynamic scene blur from images is a non-trivial task as it is difficult to model the non-uniform blur mathematically. Several methods first use single or multiple images to estimate optical flow (which is treated as an approximation of blur kernels) and then adopt non-blind deblurring algorithms to reconstruct the sharp images. However, these methods cannot be trained in an end-to-end manner and are usually computationally expensive. In this paper, we explore optical flow to remove dynamic scene blur by using the multi-scale spatially variant recurrent neural network (RNN). We utilize FlowNets to estimate optical flow from two consecutive images in different scales. The estimated optical flow provides the RNN weights in different scales so that the weights can better help RNNs to remove blur in the feature spaces. Finally, we develop a convolutional neural network (CNN) to restore the sharp images from the deblurred features. Both quantitatively and qualitatively evaluations on the benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art algorithms in terms of accuracy, speed, and model size.
Jiawei Zhang 0002, Jinshan Pan, Daoye Wang, Shangchen Zhou, Xing Wei 0001, Furong Zhao, Jimmy S. J. Ren
IEEE Trans. Circuits Syst. Video Technol.1
2021 Efficient Deep Image Denoising via Class Specific Convolution
abstract
Deep neural networks have been widely used in image denoising during the past few years. Even though they achieve great success on this problem, they are computationally inefficient which makes them inappropriate to be implemented in mobile devices. In this paper, we propose an efficient deep neural network for image denoising based on pixel-wise classification. Despite using a computationally efficient network cannot effectively remove the noises from any content, it is still capable to denoise from a specific type of pattern or texture. The proposed method follows such a divide and conquer scheme. We first use an efficient U-net to pixel-wisely classify pixels in the noisy image based on the local gradient statistics.Then we replace part of the convolution layers in existing denoising networks by the proposed Class Specific Convolution layers (CSConv) which use different weights for different classes of pixels. Quantitative and qualitative evaluations on public datasets demonstrate that the proposed method can reduce the computational costs without sacrificing the performance compared to state-of-the-art algorithms.
Jiawei Zhang 0002, Xuanye Cheng, Feng Zhang 0047, Xing Wei 0001, Jimmy S. J. Ren
AAAI2
2021 Blind Deblurring for Saturated Images
abstract
Blind deblurring has received considerable attention in recent years. However, state-of-the-art methods often fail to process saturated blurry images. The main reason is that pixels around saturated regions are not conforming to the commonly used linear blur model. Pioneer arts suggest excluding these pixels during the deblurring process, which sometimes simultaneously removes the informative edges around saturated regions and results in insufficient information for kernel estimation when large saturated regions exist. To address this problem, we introduce a new blur model to fit both saturated and unsaturated pixels, and all informative pixels can be considered during the deblurring process. Based on our model, we develop an effective maximum a posterior (MAP)-based optimization framework. Quantitative and qualitative evaluations on benchmark datasets and challenging real-world examples show that the proposed method performs favorably against existing methods.
Liang Chen 0026, Jiawei Zhang 0002, Songnan Lin, Faming Fang, Jimmy S. J. Ren
CVPR2
2021 Learning a Non-Blind Deblurring Network for Night Blurry Images
abstract
Deblurring night blurry images is difficult, because the common-used blur model based on the linear convolution operation does not hold in this situation due to the influence of saturated pixels. In this paper, we propose a non-blind deblurring network (NBDN) to restore night blurry images. To mitigate the side effects brought by the pixels that violate the blur model, we develop a confidence estimation unit (CEU) to estimate a map which ensures smaller contributions of these pixels in the deconvolution steps which are optimized by the conjugate gradient (CG) method. Moreover, unlike the existing methods using manually tuned hyper-parameters in their frameworks, we propose a hyper-parameter estimation unit (HPEU) to adaptively estimate hyper-parameters for better image restoration. The experimental results demonstrate that the proposed network performs favorably against state-of-the-art algorithms both quantitatively and qualitatively.
Liang Chen 0026, Jiawei Zhang 0002, Jinshan Pan, Songnan Lin, Faming Fang, Jimmy S. J. Ren
CVPR2
2021 Deep Blind Video Super-resolution
abstract
Existing video super-resolution (SR) algorithms usually assume that the blur kernels in the degradation process are known and do not model the blur kernels in the restoration. However, this assumption does not hold for blind video SR and usually leads to over-smoothed super-resolved frames. In this paper, we propose an effective blind video SR algorithm based on deep convolutional neural networks (CNNs). Our algorithm first estimates blur kernels from low-resolution (LR) input videos. Then, with the estimated blur kernels, we develop an effective image deconvolution method based on the image formation model of blind video SR to generate intermediate latent frames so that sharp image contents can be restored well. To effectively explore the information from adjacent frames, we estimate the motion fields from LR input videos, extract features from LR videos by a feature extraction network, and warp the extracted features from LR inputs based on the motion fields. Moreover, we develop an effective sharp feature exploration method which first extracts sharp features from restored intermediate latent frames and then uses a transformation operation based on the extracted sharp features and warped features from LR inputs to generate better features for HR video restoration. We formulate the proposed algorithm into an end-to-end trainable framework and show that it performs favorably against state-of-the-art methods.
Jinshan Pan, Haoran Bai 0001, Jiangxin Dong, Jiawei Zhang 0002, Jinhui Tang 0001
ICCV4
2021 Learning RAW-to-sRGB Mappings with Inaccurately Aligned Supervision
abstract
Learning RAW-to-sRGB mapping has drawn increasing attention in recent years, wherein an input raw image is trained to imitate the target sRGB image captured by another camera. However, the severe color inconsistency makes it very challenging to generate well-aligned training pairs of input raw and target sRGB images. While learning with inaccurately aligned supervision is prone to causing pixel shift and producing blurry results. In this paper, we circumvent such issue by presenting a joint learning model for image alignment and RAW-to-sRGB mapping. To diminish the effect of color inconsistency in image alignment, we introduce to use a global color mapping (GCM) module to generate an initial sRGB image given the input raw image, which can keep the spatial location of the pixels unchanged, and the target sRGB image is utilized to guide GCM for converting the color towards it. Then a pre-trained optical flow estimation network (e.g., PWC-Net) is deployed to warp the target sRGB image to align with the GCM output. To alleviate the effect of inaccurately aligned supervision, the warped target sRGB image is leveraged to learn RAW-to-sRGB mapping. When training is done, the GCM module and optical flow network can be detached, thereby bringing no extra computation cost for inference. Experiments show that our method performs favorably against state-of-the-arts on ZRR and SR-RAW datasets. With our joint learning model, a light-weight backbone can achieve better quantitative and qualitative performance on ZRR dataset. Codes are available at https://github.com/cszhilu1998/RAW-to-sRGB.
Zhilu Zhang 0001, Haolin Wang 0004, Ming Liu 0018, Ruohao Wang, Jiawei Zhang 0002, Wangmeng Zuo
ICCV5
2021 Physics-Based Generative Adversarial Models for Image Restoration and Beyond
abstract
We present an algorithm to directly solve numerous image restoration problems (e.g., image deblurring, image dehazing, and image deraining). These problems are ill-posed, and the common assumptions for existing methods are usually based on heuristic image priors. In this paper, we show that these problems can be solved by generative models with adversarial learning. However, a straightforward formulation based on a straightforward generative adversarial network (GAN) does not perform well in these tasks, and some structures of the estimated images are usually not preserved well. Motivated by an interesting observation that the estimated results should be consistent with the observed inputs under the physics models, we propose an algorithm that guides the estimation process of a specific task within the GAN framework. The proposed model is trained in an end-to-end fashion and can be applied to a variety of image restoration and low-level vision problems. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art algorithms.
Jinshan Pan, Jiangxin Dong, Yang Liu 0119, Jiawei Zhang 0002, Jimmy S. J. Ren, Jinhui Tang 0001, Yu-Wing Tai, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Multi-Stage Degradation Homogenization for Super-Resolution of Face Images With Extreme Degradations
abstract
Face Super-Resolution (FSR) aims to infer High-Resolution (HR) face images from the captured Low-Resolution (LR) face image with the assistance of external information. Existing FSR methods are less effective for the LR face images captured with serious low-quality since the huge imaging/degradation gap caused by the different imaging scenarios (i.e., the complex practical imaging scenario that generates test LR images, the simple manual imaging degradation that generates the training LR images) is not considered in these algorithms. In this paper, we propose an image homogenization strategy via re-expression to solve this problem. In contrast to existing methods, we propose a homogenization projection in LR space and HR space as compensation for the classical LR/HR projection to formulate the FSR in a multi-stage framework. We then develop a re-expression process to bridge the gap between the complex degradation and the simple degradation, which can remove the heterogeneous factors such as serious noise and blur. To further improve the accuracy of the homogenization, we extract the image patch set that is invariant to degradation changes as Robust Neighbor Resources (RNR), with which these two homogenization projections re-express the input LR images and the initial inferred HR images successively. Both quantitative and qualitative results on the public datasets demonstrate the effectiveness of the proposed algorithm against the state-of-the-art methods.
Liang Chen 0026, Jinshan Pan, Junjun Jiang, Jiawei Zhang 0002, Zhen Han 0002, Linchao Bao
IEEE Trans. Image Process.4
2020 Learning to Deblur Face Images via Sketch Synthesis
abstract
The success of existing face deblurring methods based on deep neural networks is mainly due to the large model capacity. Few algorithms have been specially designed according to the domain knowledge of face images and the physical properties of the deblurring process. In this paper, we propose an effective face deblurring algorithm based on deep convolutional neural networks (CNNs). Motivated by the conventional deblurring process which usually involves the motion blur estimation and the latent clear image restoration, the proposed algorithm first estimates motion blur by a deep CNN and then restores latent clear images with the estimated motion blur. However, estimating motion blur from blurry face images is difficult as the textures of the blurry face images are scarce. As most face images share some common global structures which can be modeled well by sketch information, we propose to learn face sketches by a deep CNN so that the sketches can help the motion blur estimation. With the estimated motion blur, we then develop an effective latent image restoration algorithm based on a deep CNN. Although involving the several components, the proposed algorithm is trained in an end-to-end fashion. We analyze the effectiveness of each component on face image deblurring and show that the proposed algorithm is able to deblur face images with favorable performance against state-of-the-art methods.
Songnan Lin, Jiawei Zhang 0002, Jinshan Pan, Yicun Liu, Yongtian Wang, Jing S. J. Chen, Jimmy S. J. Ren
AAAI2
2020 Visually Imbalanced Stereo Matching
abstract
Understanding of human vision system (HVS) has inspired many computer vision algorithms. Stereo matching, which borrows the idea from human stereopsis, has been extensively studied in the existing literature. However, scant attention has been drawn on a typical scenario where binocular inputs are qualitatively different (e.g., high-res master camera and low-res slave camera in a dual-lens module). Recent advances in human optometry reveal the capability of the human visual system to maintain coarse stereopsis under such visually imbalanced conditions. Bionically aroused, it is natural to question that: do stereo machines share the same capability? In this paper, we carry out a systematic comparison to investigate the effect of various imbalanced conditions on current popular stereo matching algorithms. We show that resembling the human visual system, those algorithms can handle limited degrees of monocular downgrading but also prone to collapses beyond a certain threshold. To avoid such collapse, we propose a solution to recover the stereopsis by a joint guided-view-restoration and stereo-reconstruction framework. We show the superiority of our framework on KITTI dataset and its extension on real-world applications.
Yicun Liu, Jimmy S. J. Ren, Jiawei Zhang 0002, Mude Lin
CVPR3
2020 Learning a Reinforced Agent for Flexible Exposure Bracketing Selection
abstract
Automatically selecting exposure bracketing (images exposed differently) is important to obtain a high dynamic range image by using multi-exposure fusion. Unlike previous methods that have many restrictions such as requiring camera response function, sensor noise model, and a stream of preview images with different exposures (not accessible in some scenarios e.g. mobile applications), we propose a novel deep neural network to automatically select exposure bracketing, named EBSNet, which is sufficiently flexible without having the above restrictions. EBSNet is formulated as a reinforced agent that is trained by maximizing rewards provided by a multi-exposure fusion network (MEFNet). By utilizing the illumination and semantic information extracted from just a single auto-exposure preview image, EBSNet enables to select an optimal exposure bracketing for multi-exposure fusion. EBSNet and MEFNet can be jointly trained to produce favorable results against recent state-of-the-art approaches. To facilitate future research, we provide a new benchmark dataset for multi-exposure selection and fusion.
Zhouxia Wang, Jiawei Zhang 0002, Mude Lin, Ping Luo 0002, Jimmy S. J. Ren
CVPR2
2020 Learning Event-Driven Video Deblurring and Interpolation
Songnan Lin, Jiawei Zhang 0002, Jinshan Pan, Dongqing Zou, Yongtian Wang, Jing Chen 0018, Jimmy S. J. Ren
ECCV (8)2
2020 EfficientFCN: Holistically-Guided Decoding for Semantic Segmentation
Junjun He, Jiawei Zhang 0002, Jimmy S. J. Ren, Hongsheng Li 0001
ECCV (26)3
2020 Cross-Scale Internal Graph Neural Network for Image Super-Resolution
abstract
Non-local self-similarity in natural images has been well studied as an effective prior in image restoration. However, for single image super-resolution (SISR), most existing deep non-local methods (e.g., non-local neural networks) only exploit similar patches within the same scale of the low-resolution (LR) input image. Consequently, the restoration is limited to using the same-scale information while neglecting potential high-resolution (HR) cues from other scales. In this paper, we explore the cross-scale patch recurrence property of a natural image, i.e., similar patches tend to recur many times across different scales. This is achieved using a novel cross-scale internal graph neural network (IGNN). Specifically, we dynamically construct a cross-scale graph by searching k-nearest neighboring patches in the downsampled LR image for each query patch in the LR image. We then obtain the corresponding k HR neighboring patches in the LR image and aggregate them adaptively in accordance to the edge label of the constructed graph. In this way, the HR information can be passed from k HR neighboring patches to the LR query patch to help it recover more detailed textures. Besides, these internal image-specific LR/HR exemplars are also significant complements to the external information learned from the training dataset. Extensive experiments demonstrate the effectiveness of IGNN against the state-of-the-art SISR methods including existing non-local networks on standard benchmarks.
Shangchen Zhou, Jiawei Zhang 0002, Wangmeng Zuo, Chen Change Loy
NeurIPS2
2020 Self-Guided Novel View Synthesis via Elastic Displacement Network
abstract
Synthesizing a novel view from different viewpoints has been an essential problem in 3D vision. Among a variety of view synthesis tasks, single image based view synthesis is particularly challenging. Recent works address this problem by a fixed number of image planes of discrete disparities, which tend to generate structurally inconsistent results on wide-baseline, scene-complicated datasets such as KITTI. In this paper, we propose the Self-Guided Elastic Displacement Network (SG-EDN), which explicitly models the geometric transformation by a novel non-discrete scene representation called layered displacement maps (LDM). To generate realistic views, we exploit the positional characteristics of the displacement maps and design a multi-scale structural pyramid for self-guided filtering on the displacement maps. To optimize efficiency and scene-adaptivity, we allow the effective range of each displacement map to be `elastic', with fully learnable parameters. Experimental results confirm that our framework outperforms existing methods in both quantitative and qualitative tests.
Yicun Liu, Jiawei Zhang 0002, Jimmy S. J. Ren
WACV2
2020 Cross-spectral stereo matching for facial disparity estimation in the dark
Songnan Lin, Jiawei Zhang 0002, Jing Chen 0018, Yongtian Wang, Yicun Liu, Jimmy S. J. Ren
Comput. Vis. Image Underst.2
2020 Robust Face Super-Resolution via Position Relation Model Based on Global Face Context
abstract
Because Face Super-Resolution (FSR) tends to infer High-Resolution (HR) face image by breaking the given Low- Resolution (LR) image into individual patches and inferring the HR correspondence one patch by one separately, Super- Resolution (SR) of face images with serious degradation, especially with occlusion, is still a challenging problem of the computer vision field. To address this problem, we propose a patch-level face model for FSR, which we called the position relation model. This model consists of the mapping relationships in every face position to the rest of the face positions based on similarity. In other words, we build a constraint for each patch position via the relationship in this model from the global range of face. Once an individual input LR image patch is seriously deteriorated, the substitute patch in whole face range can be sought according to the relationship of the model at this position as the provider of the LR information. In this way, the lost facial structures can be compensated by knowledge located in remote pixels or structure information which leads to better high-resolution face images. The LR images with degradations, not only the serious low-quality degradation, e.g. noise, blur, but also the occlusions, can be effectively hallucinated into HR ones. Quantitative and qualitative evaluations on the public datasets demonstrate that the proposed algorithm performs favorably against state-of-theart methods.
Liang Chen 0026, Jinshan Pan, Junjun Jiang, Jiawei Zhang 0002, Yi Wu 0010
IEEE Trans. Image Process.4
2019 DAVANet: Stereo Deblurring With View Aggregation
abstract
Nowadays stereo cameras are more commonly adopted in emerging devices such as dual-lens smartphones and unmanned aerial vehicles. However, they also suffer from blurry images in dynamic scenes which leads to visual discomfort and hampers further image processing. Previous works have succeeded in monocular deblurring, yet there are few studies on deblurring for stereoscopic images. By exploiting the two-view nature of stereo images, we propose a novel stereo image deblurring network with Depth Awareness and View Aggregation, named DAVANet. In our proposed network, 3D scene cues from the depth and varying information from two views are incorporated, which help to remove complex spatially-varying blur in dynamic scenes. Specifically, with our proposed fusion network, we integrate the bidirectional disparities estimation and deblurring into a unified framework. Moreover, we present a large-scale multi-scene dataset for stereo deblurring, containing 20,637 blurry-sharp stereo image pairs from 135 diverse sequences and their corresponding bidirectional disparities. The experimental results on our dataset demonstrate that DAVANet outperforms state-of-the-art methods in terms of accuracy, speed, and model size.
Shangchen Zhou, Jiawei Zhang 0002, Wangmeng Zuo, Haozhe Xie, Jinshan Pan, Jimmy S. J. Ren
CVPR2
2019 Spatio-Temporal Filter Adaptive Network for Video Deblurring
abstract
Video deblurring is a challenging task due to the spatially variant blur caused by camera shake, object motions, and depth variations, etc. Existing methods usually estimate optical flow in the blurry video to align consecutive frames or approximate blur kernels. However, they tend to generate artifacts or cannot effectively remove blur when the estimated optical flow is not accurate. To overcome the limitation of separate optical flow estimation, we propose a Spatio-Temporal Filter Adaptive Network (STFAN) for the alignment and deblurring in a unified framework. The proposed STFAN takes both blurry and restored images of the previous frame as well as blurry image of the current frame as input, and dynamically generates the spatially adaptive filters for the alignment and deblurring. We then propose the new Filter Adaptive Convolutional (FAC) layer to align the deblurred features of the previous frame with the current frame and remove the spatially variant blur from the features of the current frame. Finally, we develop a reconstruction network which takes the fusion of two transformed features to restore the clear frames. Both quantitative and qualitative evaluation results on the benchmark datasets and real-world videos demonstrate that the proposed algorithm performs favorably against state-of-the-art methods in terms of accuracy, speed as well as model size.
Shangchen Zhou, Jiawei Zhang 0002, Jinshan Pan, Wangmeng Zuo, Haozhe Xie, Jimmy S. J. Ren
ICCV2
2019 Recovering Extremely Degraded Faces by Joint Super-Resolution and Facial Composite
abstract
In the past a few years, we witnessed rapid advancement in face super-resolution from very low resolution(VLR) images. However, most of the previous studies focus on solving such problem without explicitly considering the impact of severe real-life image degradation (e.g. blur and noise). We can show that robustly recover details from VLR images is a task beyond the ability of current state-of-the-art method. In this paper, we borrow ideas from "facial composite" and propose an alternative approach to tackle this problem. We endow the degraded VLR images with additional cues by integrating existing face components from multiple reference images into a novel learning pipeline with both low level and high level semantic loss function as well as a specialized adversarial based training scheme. We show that our method is able to effectively and robustly restore relevant facial details from 16x16 images with extreme degradation. We also tested our approach against real-life images and our method performs favorably against previous methods.
Xiu Li 0001, Guichun Duan, Zhouxia Wang, Jimmy S. J. Ren, Yongbing Zhang 0002, Jiawei Zhang 0002, Kaixiang Song
ICTAI6
2019 Joint Face Hallucination and Deblurring via Structure Generation and Detail Enhancement
Yibing Song, Jiawei Zhang 0002, Lijun Gong, Shengfeng He, Linchao Bao, Jinshan Pan, Qingxiong Yang, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.2
2019 Interactive Hierarchical Object Proposals
abstract
Object proposal algorithms have been demonstrated to be very successful in accelerating object detection process. High object localization quality and detection recall can be obtained using thousands of proposals. However, the performance with a small number of proposals is still unsatisfactory. This paper demonstrates that the performance of a few proposals can be significantly improved with the minimal human interaction-a single touch point. To this end, we first generate hierarchical superpixels using an efficient tree-organized structure as our initial object proposals, and then select only a few proposals from them by learning an effective Convolutional neural network for objectness ranking. We explore and design an architecture to integrate human interaction with the global information of the whole image for objectness scoring, which is able to significantly improve the performance with a minimum number of object proposals. Extensive experiments show the proposed method outperforms all the state-of-the-art methods for locating the meaningful object with the touch point constraint. Furthermore, the proposed method is extended for video. By combining with the novel interactive motion segmentation cue for generating hierarchical superpixels, the performance on a single proposal is satisfactory and can be used in the interactive vision systems, such as selecting the input of a real-time tracking system.
Jiawei Zhang 0002, Shengfeng He, Qingxiong Yang, Qing Li 0001, Ming-Hsuan Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 Learning Dual Convolutional Neural Networks for Low-Level Vision
abstract
In this paper, we propose a general dual convolutional neural network (DualCNN) for low-level vision problems, e.g., super-resolution, edge-preserving filtering, deraining and dehazing. These problems usually involve the estimation of two components of the target signals: structures and details. Motivated by this, our proposed DualCNN consists of two parallel branches, which respectively recovers the structures and details in an end-to-end manner. The recovered structures and details can generate the target signals according to the formation model for each particular application. The DualCNN is a flexible framework for low-level vision tasks and can be easily incorporated into existing CNNs. Experimental results show that the DualCNN can be effectively applied to numerous low-level vision tasks with favorable performance against the state-of-the-art methods.
Jinshan Pan, Sifei Liu, Deqing Sun, Jiawei Zhang 0002, Yang Liu 0119, Jimmy S. J. Ren, Zechao Li, Jinhui Tang 0001, Huchuan Lu, Yu-Wing Tai, Ming-Hsuan Yang 0001
CVPR4
2018 Gated Fusion Network for Single Image Dehazing
abstract
In this paper, we propose an efficient algorithm to directly restore a clear image from a hazy input. The proposed algorithm hinges on an end-to-end trainable neural network that consists of an encoder and a decoder. The encoder is exploited to capture the context of the derived input images, while the decoder is employed to estimate the contribution of each input to the final dehazed result using the learned representations attributed to the encoder. The constructed network adopts a novel fusion-based strategy which derives three inputs from an original hazy image by applying White Balance (WB), Contrast Enhancing (CE), and Gamma Correction (GC). We compute pixel-wise confidence maps based on the appearance differences between these different inputs to blend the information of the derived inputs and preserve the regions with pleasant visibility. The final dehazed image is yielded by gating the important features of the derived inputs. To train the network, we introduce a multi-scale approach such that the halo artifacts can be avoided. Extensive experimental results on both synthetic and real-world images demonstrate that the proposed algorithm performs favorably against the state-of-the-art algorithms.
Wenqi Ren, Lin Ma 0002, Jiawei Zhang 0002, Jinshan Pan, Xiaochun Cao, Wei Liu 0005, Ming-Hsuan Yang 0001
CVPR3
2018 Dynamic Scene Deblurring Using Spatially Variant Recurrent Neural Networks
abstract
Due to the spatially variant blur caused by camera shake and object motions under different scene depths, deblurring images captured from dynamic scenes is challenging. Although recent works based on deep neural networks have shown great progress on this problem, their models are usually large and computationally expensive. In this paper, we propose a novel spatially variant neural network to address the problem. The proposed network is composed of three deep convolutional neural networks (CNNs) and a recurrent neural network (RNN). RNN is used as a deconvolution operator performed on feature maps extracted from the input image by one of the CNNs. Another CNN is used to learn the weights for the RNN at every location. As a result, the RNN is spatially variant and could implicitly model the deblurring process with spatially variant kernels. The third CNN is used to reconstruct the final deblurred feature maps into restored image. The whole network is end-to-end trainable. Our analysis shows that the proposed network has a large receptive field even with a small model size. Quantitative and qualitative evaluations on public datasets demonstrate that the proposed method performs favorably against state-of-the-art algorithms in terms of accuracy, speed, and model size.
Jiawei Zhang 0002, Jinshan Pan, Jimmy S. J. Ren, Yibing Song, Linchao Bao, Rynson W. H. Lau, Ming-Hsuan Yang 0001
CVPR1
2018 Grassmann Pooling as Compact Homogeneous Bilinear Pooling for Fine-Grained Visual Classification
Xing Wei 0001, Yihong Gong, Jiawei Zhang 0002, Nanning Zheng 0001
ECCV (3)4
2018 Deep Non-Blind Deconvolution via Generalized Low-Rank Approximation
abstract
In this paper, we present a deep convolutional neural network to capture the inherent properties of image degradation, which can handle different kernels and saturated pixels in a unified framework. The proposed neural network is motivated by the low-rank property of pseudo-inverse kernels. We first compute a generalized low-rank approximation for a large number of blur kernels, and then use separable filters to initialize the convolutional parameters in the network. Our analysis shows that the estimated decomposed matrices contain the most essential information of the input kernel, which ensures the proposed network to handle various blurs in a unified framework and generate high-quality deblurring results. Experimental results on benchmark datasets with noise and saturated pixels demonstrate that the proposed algorithm performs favorably against state-of-the-art methods.
Wenqi Ren, Jiawei Zhang 0002, Lin Ma 0002, Jinshan Pan, Xiaochun Cao, Wangmeng Zuo, Wei Liu 0005, Ming-Hsuan Yang 0001
NeurIPS2
2018 Specular highlight reduction with known surface geometry
Xing Wei 0001, Xiaobin Xu 0001, Jiawei Zhang 0002, Yihong Gong
Comput. Vis. Image Underst.3
2017 Learning Fully Convolutional Networks for Iterative Non-blind Deconvolution
abstract
In this paper, we propose a fully convolutional network for iterative non-blind deconvolution. We decompose the non-blind deconvolution problem into image denoising and image deconvolution. We train a FCNN to remove noise in the gradient domain and use the learned gradients to guide the image deconvolution step. In contrast to the existing deep neural network based methods, we iteratively deconvolve the blurred images in a multi-stage framework. The proposed method is able to learn an adaptive image prior, which keeps both local (details) and global (structures) information. Both quantitative and qualitative evaluations on the benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art algorithms in terms of quality and speed.
Jiawei Zhang 0002, Jinshan Pan, Wei-Sheng Lai, Rynson W. H. Lau, Ming-Hsuan Yang 0001
CVPR1
2017 CREST: Convolutional Residual Learning for Visual Tracking
abstract
Discriminative correlation filters (DCFs) have been shown to perform superiorly in visual tracking. They only need a small set of training samples from the initial frame to generate an appearance model. However, existing DCFs learn the filters separately from feature extraction, and update these filters using a moving average operation with an empirical weight. These DCF trackers hardly benefit from the end-to-end training. In this paper, we propose the CREST algorithm to reformulate DCFs as a one-layer convolutional neural network. Our method integrates feature extraction, response map generation as well as model update into the neural networks for an end-to-end training. To reduce model degradation during online update, we apply residual learning to take appearance changes into account. Extensive experiments on the benchmark datasets demonstrate that our CREST tracker performs favorably against state-of-the-art trackers.
Yibing Song, Chao Ma 0004, Lijun Gong, Jiawei Zhang 0002, Rynson W. H. Lau, Ming-Hsuan Yang 0001
ICCV4
2017 A hand pose tracking benchmark from stereo matching
abstract
In this paper we establish a long-term 3D hand pose tracking benchmark1. It contains 18,000 stereo image pairs as well as the ground-truth 3D positions of palm and finger joints from different scenarios. Meanwhile, to accurately segment hand from stereo images, we propose a novel stereo-based hand segmentation and depth estimation algorithm specially tailored for hand tracking here. The experiments indicate the effectiveness of the proposed algorithm by demonstrating that its tracking performance is comparable to the use of an active depth sensor under various of challenging scenarios.
Jiawei Zhang 0002, Jianbo Jiao, Liangqiong Qu, Xiaobin Xu 0001, Qingxiong Yang
ICIP1
2017 Fast Preprocessing for Robust Face Sketch Synthesis
abstract
Exemplar-based face sketch synthesis methods usually meet the challenging problem that input photos are captured in different lighting conditions from training photos. The critical step causing the failure is the search of similar patch candidates for an input photo patch. Conventional illumination invariant patch distances are adopted rather than directly relying on pixel intensity difference, but they will fail when local contrast within a patch changes. In this paper, we propose a fast preprocessing method named Bidirectional Luminance Remapping (BLR), which interactively adjust the lighting of training and input photos. Our method can be directly integrated into state-of-the-art exemplar-based methods to improve their robustness with ignorable computational cost
Yibing Song, Jiawei Zhang 0002, Linchao Bao, Qingxiong Yang
IJCAI2
2017 Learning to Hallucinate Face Images via Component Generation and Enhancement
abstract
We propose a two-stage method for face hallucination. First, we generate facial components of the input image using CNNs. These components represent the basic facial structures. Second, we synthesize fine-grained facial structures from high resolution training images. The details of these structures are transferred into facial components for enhancement. Therefore, we generate facial components to approximate ground truth global appearance in the first stage and enhance them through recovering details in the second stage. The experiments demonstrate that our method performs favorably against state-of-the-art methods.
Yibing Song, Jiawei Zhang 0002, Shengfeng He, Linchao Bao, Qingxiong Yang
IJCAI2
2017 RGBD Salient Object Detection via Deep Fusion
abstract
Numerous efforts have been made to design various low-level saliency cues for RGBD saliency detection, such as color and depth contrast features as well as background and color compactness priors. However, how these low-level saliency cues interact with each other and how they can be effectively incorporated to generate a master saliency map remain challenging problems. In this paper, we design a new convolutional neural network (CNN) to automatically learn the interaction mechanism for RGBD salient object detection. In contrast to existing works, in which raw image pixels are fed directly to the CNN, the proposed method takes advantage of the knowledge obtained in traditional saliency detection by adopting various flexible and interpretable saliency feature vectors as inputs. This guides the CNN to learn a combination of existing features to predict saliency more effectively, which presents a less complex problem than operating on the pixels directly. We then integrate a superpixel-based Laplacian propagation framework with the trained CNN to extract a spatially consistent saliency map by exploiting the intrinsic structure of the input image. Extensive quantitative and qualitative experimental evaluations on three data sets demonstrate that the proposed method consistently outperforms the state-of-the-art methods.
Liangqiong Qu, Shengfeng He, Jiawei Zhang 0002, Jiandong Tian, Yandong Tang, Qingxiong Yang
IEEE Trans. Image Process.3