Qing Zhang 0006

dblp:68/1429-6 · DBLP profile ↗
← Back
62ranked-venue papers
10as first author
43since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 54 · 8 first-author · 36 since 2021Artificial intelligence and machine learning · 28 · 2 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 SynerDetect: Hierarchical Synergistic Learning for Generalizable AI-Generated Image Detection
abstract
The rapid advancement of generative models, which produce increasingly realistic synthetic images, urgently demands robust and generalizable detection methods. Consequently, research has largely pivoted to leveraging large-scale Vision Foundation Models (VFMs) for enhanced generalization. However, existing VFM-based approaches primarily adhere to either perceptual or generative paradigms, each with limitations: perceptual models capture high-level semantics but often miss subtle artifacts, whereas generative models emphasize fine-grained flaws yet overlook semantic inconsistency. To resolve this inherent trade-off, we introduce SynerDetect, a novel hierarchical synergistic framework that fundamentally unifies the two paradigms. SynerDetect achieves deep integration of heterogeneous forensic representations through two levels of synergy: Cross-Model Interactive Distillation (CMID) distills generative forensic signals into perceptual encoders via prompt-guided reconstruction; and Optimal Transport-Guided Discriminative Contrastive Learning (OT-DCL) structurally aligns and integrates these heterogeneous representations, consolidating them into a robust, unified detection space. SynerDetect achieves superior performance on standard benchmarks (AIGCDetectBenchmark and GenImage) and attains a notable 5.20% accuracy gain on the challenging Chameleon benchmark, whose synthetic images consistently pass the Visual Turing Test. These results unequivocally validate the robust, real-world generalization of our unified cross-paradigm framework.
Shuaibo Li, Zhaohu Xing, Hongqiu Wang, Pengfei Hao, Zekai Liu, Qing Zhang 0006, Lei Zhu 0003
AAAI8
2026 DLIENet: A lightweight low-light image enhancement network via knowledge distillation
Ling Zhang 0017, Qing Zhang 0006, Zheng Liu 0004, Xiaolong Zhang 0002, Chunxia Xiao
Pattern Recognit.4
2025 CLIP-RestoreX: Restore Image Structure and Perception in Exposure Correction
abstract
Exposure correction aims to adjust the exposure of an under- and over-exposed image to enhance its overall visual quality. The core challenge of this task lies in that it requires to faithfully restore both the structure and perception information. In this work, we present a novel exposure correction method, referred to as CLIP-RestoreX, that leverages structural and perceptual priors from CLIP to tackle exposure correction. Specifically, we in CLIP-RestoreX propose to perform exposure correction by aligning CLIP-based structural and perceptual feature of the impaired image with its ground-truth image. To better restore the damaged structural information and perceptual information, we further design a frequency-domain based feature enhancement diffusion model, where we utilize the globality of Fourier transform to help reveal potential the relationship within the features. We conduct extensive experiments on several benchmark datasets. The results demonstrate that the proposed CLIP-RestoreX outperforms state-of-the-art exposure correction methods.
Qing Zhang 0006, Jianfang Hu, Wei-Shi Zheng 0001
AAAI2
2025 When Shadow Removal Meets Intrinsic Image Decomposition: A Joint Learning Framework Using Unpaired Data
abstract
We present a framework that achieves shadow removal by learning intrinsic image decomposition (IID) from unpaired shadow and shadow-free images. Although it is well-known that intrinsic images, \ie, illumination and reflectance, are highly beneficial to shadow removal, IID is rarely adopted by previous work due to its inherent ambiguity and the scarcity of training data. However, we find that by properly coupling shadow removal and IID into a joint learning framework, they can reinforce each other and enable promising results on both tasks, even with unpaired training data. Our framework is comprised of an IID network for separating the shadow input image into illumination and reflectance, and an illumination recovery network for predicting shadow-free illumination with which we are able to produce the shadow removal output by recombining with the estimated reflectance. We perform extensive experiments on various benchmark datasets to demonstrate the effectiveness of our method in shadow removal, and also showcase our advantage over previous IID methods in handling images with complex shadows.
Rongjia Zheng, Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001
AAAI2
2025 RoGSplat: Learning Robust Generalizable Human Gaussian Splatting from Sparse Multi-View Images
abstract
This paper presents RoGSplat, a novel approach for synthesizing high-fidelity novel views of unseen human from sparse multi-view images, while requiring no cumbersome per-subject optimization. Unlike previous methods that typically struggle with sparse views with few overlappings and are less effective in reconstructing complex human geometry, the proposed method enables robust reconstruction in such challenging conditions. Our key idea is to lift SMPL vertices to dense and reliable 3D prior points representing accurate human body geometry, and then regress human Gaussian parameters based on the points. To account for possible misalignment between SMPL model and images, we propose to predict image-aligned 3D prior points by leveraging both pixel-level features and voxel-level features, from which we regress the coarse Gaussians. To enhance the ability to capture high-frequency details, we further render depth maps from the coarse 3D Gaussians to help regress fine-grained pixel-wise Gaussians. Experiments on several benchmark datasets demonstrate that our method outperforms state-of-the-art methods in novel view synthesis and cross-dataset generalization. Our code is available at https://github.com/iSEE-Laboratory/RoGSplat.
Junjin Xiao, Qing Zhang 0006, Yonewei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001
CVPR2
2025 Diffusion-based Event Generation for High-Quality Image Deblurring
abstract
While event-based deblurring have demonstrated impressive results, they are impractical for consumer photos captured by cell phones and digital cameras that are not equipped with the event sensor. To address this problem, we in this paper propose a novel deblurring framework called Event Generation Deblurring (EGDeblurring), which allows to effectively deblur an image by generating event guidance describing the motion information using a diffusion model. Specifically, we design a motion prior generation diffusion model and a feature extractor to produce prior information beneficial for deblurring, rather than generating the raw event representation. In order to achieve effective fusion of motion prior information with blurry images and produce high-quality results, we develop a regression deblurring network embedded with a dual-attention channel fusion block. Experiments on multiple datasets demonstrate that our method outperforms state-of-the-art image deblurring methods. Our code is available at https://github.com/XinanXie/EGDeblurring.
Xinan Xie, Qing Zhang 0006, Wei-Shi Zheng 0001
CVPR2
2025 EntityErasure: Erasing Entity Cleanly via Amodal Entity Segmentation and Completion
abstract
This paper presents EntityErasure, a novel diffusion-based inpainting method that can effectively erase entities without inducing unwanted sundries. To this end, we propose to address this problem by dividing it into amodal entity segmentation and completion, such that the region to inpaint takes only entities in the non-inpainting area as reference, avoiding the possibility to generate unpredictable sundries. Moreover, we develop two entity segmentation based metrics for quantitatively assessing the performance of object erasure, which are shown be more effective than existing metrics. Experimental results demonstrate that our approach outperforms other state-of-the-art object erasure methods. Our code and data are available at https://zyxunh.github.io/EntityErasure-ProjectPage/.
Yixing Zhu, Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001
CVPR2
2025 Learning Rank Constrained Exposure Correction from Unpaired Data
abstract
This paper presents a novel unpaired learning based exposure correction network that can robustly deal with both under- and over-exposed images. Given an input image in any exposure conditions, we first obtain its intermediate under- and over-exposure corrected versions by predicting dual illuminations for Retinex-based enhancement, the two outputs together with the original input are then fed to a multi-exposure fusion module to adaptively locate the best-exposed regions in the three images and then seamlessly fuse them into a well-exposed output. To ensure that the generated result is visually natural and free of disturbing visual artifacts such as loss of details and contrast degradation, we introduce a novel rank loss. Experiments show that our method outperforms existing methods in terms of both quantitative and qualitative results.
Zhuoyue Gong, Qing Zhang 0006
ICASSP2
2025 Self-Supervised Image Harmonization via Holistic Feature Fusion
abstract
Image harmonization is crucial for image composition, aiming to adjust the appearance of the foreground objects to be visually consistent to the background image, so as to produce a realistic yet natural composite image. Due to the difficulty in collecting large-scale annotated datasets, investigating self-supervised image harmonization methods has become a trend, while existing self-supervised methods typically struggle with complex cases with large visual discrepancy between the foreground and background. To address their limitations, we in this paper present a novel self-supervised framework for image harmonization. To allow for semantic-aware localized style adjustment and also global lighting transfer, we propose to perform holistic feature fusion, where local attention feature and global style feature are fused to produce the harmonization result. Experiments demonstrate that our method outperforms existing self-supervised image harmonization methods.
Chenyang Tian, Qing Zhang 0006
ICASSP2
2025 Structure-Guided Diffusion Models for High-Fidelity Portrait Shadow Removal
abstract
We present a diffusion-based portrait shadow removal approach that can robustly produce high-fidelity results. Unlike previous methods, we cast shadow removal as diffusion-based inpainting. To this end, we first train a shadow-independent structure extraction network on a real-world portrait dataset with various synthetic lighting conditions, which allows to generate a shadow-independent structure map including facial details while excluding the unwanted shadow boundaries. The structure map is then used as condition to train a structure-guided inpainting diffusion model for removing shadows in a generative manner. Finally, to restore the fine-scale details (e.g., eyelashes, moles and spots) that may not be captured by the structure map, we take the gradients inside the shadow regions as guidance and train a detail restoration diffusion model to refine the shadow removal result. Extensive experiments on the benchmark datasets show that our method clearly outperforms existing methods, and is effective to avoid previously common issues such as facial identity tampering, shadow residual, color distortion, structure blurring, and loss of details. Our code is available at https://github.com/wanchang-yu/Structure-Guided-Diffusion-for-Portrait-Shadow-Removal.
Wanchang Yu, Qing Zhang 0006, Rongjia Zheng, Wei-Shi Zheng 0001
ICCV2
2025 DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse Rendering
abstract
Recent methods have shown that pre-trained diffusion models can be fine-tuned to enable generative inverse rendering by learning image-conditioned noise-to-intrinsic mapping. Despite their remarkable progress, they struggle to robustly produce high-quality results as the noise-to-intrinsic paradigm essentially utilizes noisy images with deteriorated structure and appearance for intrinsic prediction, while it is common knowledge that structure and appearance information in an image are crucial for inverse rendering. To address this issue, we present DNF-Intrinsic, a robust yet efficient inverse rendering approach fine-tuned from a pre-trained diffusion model, where we propose to take the source image rather than Gaussian noise as input to directly predict deterministic intrinsic properties via flow matching. Moreover, we design a generative renderer to constrain that the predicted intrinsic properties are physically faithful to the source image. Experiments on both synthetic and real-world datasets show that our method clearly outperforms existing state-of-the-art methods.
Rongjia Zheng, Qing Zhang 0006, Chengjiang Long, Wei-Shi Zheng 0001
ICCV2
2025 Occlusion-Preserved Surveillance Video Synopsis with Flexible Object Graph
Yongwei Nie, Siming Zeng, Qing Zhang 0006, Guiqing Li, Ping Li 0016, Hongmin Cai
Int. J. Comput. Vis.4
2025 Portrait Shadow Removal Using Context-Aware Illumination Restoration Network
abstract
Portrait shadow removal is a challenging task due to the complex surface of the face. Although existing work in this field makes substantial progress, these methods tend to overlook information in the background areas. However, this background information not only contains some important illumination cues but also plays a pivotal role in achieving lighting harmony between the face and the background after shadow elimination. In this paper, we propose a Context-aware Illumination Restoration Network (CIRNet) for portrait shadow removal. Our CIRNet consists of three stages. First, the Coarse Shadow Removal Network (CSRNet) mitigates the illumination discrepancies between shadow and non-shadow areas. Next, the Area-aware Shadow Restoration Network (ASRNet) predicts the illumination characteristics of shadowed areas by utilizing background context and non-shadow portrait context as references. Lastly, we introduce a Global Fusion Network to adaptively merge contextual information from different areas and generate the final shadow removal result. This approach leverages the illumination information from the background region while ensuring a more consistent overall illumination in the generated images. Our approach can also be extended to high-resolution portrait shadow removal and portrait specular highlight removal. Besides, we construct the first real facial shadow dataset for portrait shadow removal, consisting of 6200 pairs of facial images. Qualitative and quantitative comparisons demonstrate the advantages of our proposed dataset as well as our method.
Jiangjian Yu, Ling Zhang 0017, Qing Zhang 0006, Daiguo Zhou, Chao Liang 0001, Chunxia Xiao
IEEE Trans. Image Process.3
2025 Self-supervised Texture Filtering
abstract
Decomposing an image I into the combination of structure S and texture T components is an important problem in computational photography and image analysis. Traditional solutions are basically non-learning based, because it is difficult to construct datasets containing ground-truth decompositions or find effective structure/texture supervisions. In this article, we present a self-supervised framework for smoothing out textures while maintaining the image structures. At the core of our method is a texture-inversion observation — if structure S and texture T are well disentangled, then S-T will produce a texture-inverted image that is symmetric to the input image I=S+T and the two will be visually highly similar, while for other conditions that structure and texture are not effectively separated, the generated texture-inverted images will be less similar to the input. Based on the observation, we propose to learn texture filtering from unlabeled data by encouraging the texture inverted image generated from the filtering output to be visually more similar to the input via contrastive learning. Experiments show that our method can robustly produce high-quality texture smoothing results, and also enables various applications.
Hao Jiang 0057, Rongjia Zheng, Yongwei Nie, Chunxia Xiao, Wei-Shi Zheng 0001, Qing Zhang 0006
ACM Trans. Graph.6
2025 Single-Image SVBRDF Estimation Using Auxiliary Renderings as Intermediate Targets
abstract
Recently, single-image SVBRDF capture is formulated as a regression problem, which uses a network to infer four SVBRDF maps from a flash-lit image. However, the accuracy is still not satisfactory since previous approaches usually adopt end-to-end inference strategies. To mitigate the challenge, we propose "auxiliary renderings" as the intermediate regression targets, through which we divide the original end-to-end regression task into several easier sub-tasks, thus achieving better inference accuracy. Our contributions are threefold. First, we design three (or two pairs of) auxiliary renderings and summarize the motivations behind the designs. By our design, the auxiliary images are bumpiness-flattened or highlight-removed, containing disentangled visual cues about the final SVBRDF maps and can be easily transformed to the final maps. Second, to help estimate the auxiliary targets from the input image, we propose two mask images including a bumpiness mask and a highlight mask. Our method thus first infers mask images, then with the help of the mask images infers auxiliary renderings, and finally transforms the auxiliary images to SVBRDF maps. Third, we propose backbone UNets to infer mask images, and gated deformable UNets for estimating auxiliary targets. Thanks to the well-designed networks and intermediate images, our method outputs better SVBRDF maps than previous approaches, validated by the extensive comparisonal and ablation experiments.
Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li, Hongmin Cai
IEEE Trans. Vis. Comput. Graph.4
2025 Towards Photorealistic Portrait Style Transfer in Unconstrained Conditions
abstract
We present a photorealistic portrait style transfer approach that allows for producing high-quality results in previously challenging unconstrained conditions, e.g., large facial perspective difference between portraits, faces with complex illumination (e.g., shadow and highlight) and occlusion, and can test without portrait parsing masks. We achieve this by developing a framework to learn robust dense correspondence across portraits for semantically aligned style transfer, where a regional style contrastive learning strategy is devised to boost the effectiveness of semantic-aware style transfer while enhancing the robustness to complex illumination. Extensive experiments demonstrate the superiority of our method.
Xinbo Wang, Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001
IEEE Trans. Vis. Comput. Graph.2
2024 Face Expression Recognition via Product-Cross Dual Attention and Neutral-Aware Anchor Loss
Yongwei Nie, Qing Zhang 0006, Xuemiao Xu, Guiqing Li, Hongmin Cai
CVM (2)3
2024 NECA: Neural Customizable Human Avatar
abstract
Human avatar has become a novel type of 3D asset with various applications. Ideally, a human avatar should be fully customizable to accommodate different settings and environments. In this work, we introduce NECA, an approach capable of learning versatile human representation from monocular or sparse-view videos, enabling granular customization across aspects such as pose, shadow, shape, lighting and texture. The core of our approach is to represent humans in complementary dual spaces and predict disentangled neural fields of geometry, albedo, shadow, as well as an external lighting, from which we are able to derive realistic rendering with high-frequency details via volumetric rendering. Extensive experiments demonstrate the advantage of our method over the state-of-the-art methods in photorealistic rendering, as well as various editing tasks such as novel pose synthesis and relighting. Our code is available at https://github.com/iSEE-Laboratory/NECA.
Junjin Xiao, Qing Zhang 0006, Wei-Shi Zheng 0001
CVPR2
2024 Interleaving One-Class and Weakly-Supervised Models with Adaptive Thresholding for Unsupervised Video Anomaly Detection
Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Pradipta Maji, Hongmin Cai
ECCV (30)4
2024 Multi-RoI Human Mesh Recovery with Camera Consistency and Contrastive Losses
Yongwei Nie, Changzhen Liu, Chengjiang Long, Qing Zhang 0006, Guiqing Li, Hongmin Cai
ECCV (47)4
2024 Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh Recovery
abstract
Human Mesh Recovery (HMR) is the task of estimating a parameterized 3D human mesh from an image. There is a kind of methods first training a regression model for this problem, then further optimizing the pretrained regression model for any specific sample individually at test time. However, the pretrained model may not provide an ideal optimization starting point for the test-time optimization. Inspired by meta-learning, we incorporate the test-time optimization into training, performing a step of test-time optimization for each sample in the training batch before really conducting the training optimization over all the training samples. In this way, we obtain a meta-model, the meta-parameter of which is friendly to the test-time optimization. At test time, after several test-time optimization steps starting from the meta-parameter, we obtain much higher HMR accuracy than the test-time optimization starting from the simply pretrained regression model. Furthermore, we find test-time HMR objectives are different from training-time objectives, which reduces the effectiveness of the learning of the meta-model. To solve this problem, we propose a dual-network architecture that unifies the training-time and test-time objectives. Our method, armed with meta-learning and the dual networks, outperforms state-of-the-art regression-based and optimization-based HMR approaches, as validated by the extensive experiments. The codes are available at https://github.com/fmx789/Meta-HMR.
Yongwei Nie, Mingxian Fan, Chengjiang Long, Qing Zhang 0006, Jian Zhu 0001, Xuemiao Xu
NeurIPS4
2024 Make static person walk again via separating pose action from shape
abstract
This paper addresses the problem of animating a person in static images, the core task of which is to infer future poses for the person. Existing approaches predict future poses in the 2D space, suffering from entanglement of pose action and shape. We propose a method that generates actions in the 3D space and then transfers them to the 2D person. We first lift the 2D pose of the person to a 3D skeleton, then propose a 3D action synthesis network predicting future skeletons, and finally devise a self-supervised action transfer network that transfers the actions of 3D skeletons to the 2D person. Actions generated in the 3D space look plausible and vivid. More importantly, self-supervised action transfer allows our method to be trained only on a 3D MoCap dataset while being able to process images in different domains. Experiments on three image datasets validate the effectiveness of our method.
Yongwei Nie, Meihua Zhao, Qing Zhang 0006, Ping Li 0016, Jian Zhu 0001, Hongmin Cai
Graph. Model.3
2024 Towards High-Resolution Specular Highlight Detection
Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Qifeng Lin, Siyuan Fan, Chunxia Xiao
Int. J. Comput. Vis.2
2024 ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection
Lei Zhu 0003, Jiaxing Shen, Huazhu Fu, Qing Zhang 0006, Liansheng Wang 0002
Int. J. Comput. Vis.5
2024 Learning Motion-Guided Multi-Scale Memory Features for Video Shadow Detection
abstract
Natural images often contain multiple shadow regions, and existing video shadow detection methods tend to fail in fully identifying all shadow regions, since they mainly learned temporal features at single-scale and single memory. In this work, we develop a novel convolutional neural network (CNN) to learn motion-guided multi-scale memory features to obtain multi-scale temporal information based on multiple network memories for boosting video shadow detection. To do so, our network first constructs three memories (i.e., a global memory, a local memory, and a motion memory) to combine spatial context and object motion for detecting shadows. Based on these three memories, we then devise a multi-scale motion-guided long-short transformer (MMLT) module to learn multi-scale temporal and motion memory features for predicting a shadow detection map of the input video frame. Our MMLT module includes a dense-scale long transformer (DLT), a dense-scale short transformer (DST), and a dense-scale motion transformer (DMT) to read three memories for learning multi-scale transformer features. Our DLT, DST, and DMT consist of a set of memory-read pooling attention (MPA) blocks and densely connect these output features of multiple MPA blocks to learn multi-scale transformer features since the scales of these output features are varied. By doing so, we can more accurately identify multiple shadow regions with different sizes from the input video. Moreover, we devise a self-supervised pretext task to pre-training the feature encoder for enhancing the downstream video shadow detection. Experimental results on three benchmark datasets show that our video shadow detection network quantitatively and qualitatively outperforms 26 state-of-the-art methods.
Jiaxing Shen, Xin Yang 0011, Huazhu Fu, Qing Zhang 0006, Ping Li 0016, Bin Sheng 0001, Liansheng Wang 0002, Lei Zhu 0003
IEEE Trans. Circuits Syst. Video Technol.5
2024 Building Coarse to Fine Convex Hulls With Auxiliary Vertices for Palette-Based Image Recoloring
abstract
Constructing a convex hull for the pixel colors of an image by viewing them as 3D points can extract a set of palette colors for the image, then image recoloring can be achieved by modifying the palette colors. For better recoloring effect, the convex hull should contain more pixels (inclusive) and be more compact. Otherwise, reconstruction error would occur or the extracted palette color would be less representative, yielding wrong recoloring results or less effective edit. We observe that convex hulls constructed by prior methods can contain all the image pixels, but are far from compact. Efforts have been made to optimize the vertices of convex hull to increase the compactness but are still not perfect. In this paper, we propose a novel coarse to fine convex hull construction scheme with auxiliary vertices. We start by constructing a coarse convex hull whose vertices are directly image pixels which is thus the most compact but cannot contain all pixels. We then make a remedy by adding auxiliary vertices into the coarse convex hull to obtain a fine convex hull. More auxiliary vertices are added, more image pixels will be contained into the fine convex hull. The auxiliary vertices are image pixels too so that the compactness can still be maintained. During editing, the auxiliary vertices are not allowed to be edited for edit convenience, but deformed as-rigid-as-possible with the adjusting of other vertices. Our convex hull is both inclusive and compact. Extensive experiments validate the effectiveness of the proposed method.
Qiwei Sun, Yongwei Nie, Qing Zhang 0006, Guiqing Li
IEEE Trans. Vis. Comput. Graph.3
2023 Document Image Shadow Removal Guided by Color-Aware Background
abstract
Existing works on document image shadow removal mostly depend on learning and leveraging a constant background (the color of the paper) from the image. However, the constant background is less representative and frequently ignores other background colors, such as the printed colors, resulting in distorted results. In this paper, we present a color-aware background extraction network (CBENet) for extracting a spatially varying background image that accurately depicts the background colors of the document. Furthermore, we propose a background-guided document images shadow removal network (BGShadowNet) using the predicted spatially varying background as auxiliary information, which consists of two stages. At Stage I, a background-constrained decoder is designed to promote a coarse result. Then, the coarse result is refined with a background-based attention module (BAModule) to maintain a consistent appearance and a detail improvement module (DEModule) to enhance the texture details at Stage II. Experiments on two benchmark datasets qualitatively and quantitatively validate the superiority of the proposed approach over state-of-the-arts.
Ling Zhang 0017, Yinghao He, Qing Zhang 0006, Zheng Liu 0004, Xiaolong Zhang 0002, Chunxia Xiao
CVPR3
2023 Towards High-Quality Specular Highlight Removal by Leveraging Large-Scale Synthetic Data
abstract
This paper aims to remove specular highlights from a single object-level image. Although previous methods have made some progresses, their performance remains somewhat limited, particularly for real images with complex specular highlights. To this end, we propose a three-stage network to address them. Specifically, given an input image, we first decompose it into the albedo, shading, and specular residue components to estimate a coarse specular-free image. Then, we further refine the coarse result to alleviate its visual artifacts such as color distortion. Finally, we adjust the tone of the refined result to match the tone of the input as closely as possible. In addition, to facilitate network training and quantitative evaluation, we present a large-scale synthetic dataset of object-level images, covering diverse objects and illumination conditions. Extensive experiments illustrate that our network is able to generalize well to unseen real object-level images, and even produce good results for scene-level images with multiple background objects and complex lighting.
Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Chunxia Xiao, Ping Li 0016
ICCV2
2023 Learning to Remove Shadows from a Single Image
Hao Jiang 0057, Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001
Int. J. Comput. Vis.2
2023 S $^3$ Net: Self-Supervised Self-Ensembling Network for Semi-Supervised RGB-D Salient Object Detection
abstract
RGB-D salient object detection aims to detect visually distinctive objects or regions from a pair of the RGB image and the depth image. State-of-the-art RGB-D saliency detectors are mainly based on convolutional neural networks but almost suffer from an intrinsic limitation relying on the labeled data, thus degrading detection accuracy in complex cases. In this work, we present a self-supervised self-ensembling network (S$^3$Net) for semi-supervised RGB-D salient object detection by leveraging the unlabeled data and exploring a self-supervised learning mechanism. To be specific, we first build a self-guided convolutional neural network (SG-CNN) as a baseline model by developing a series of three-layer cross-model feature fusion (TCF) modules to leverage complementary information among depth and RGB modalities and formulating an auxiliary task that predicts a self-supervised image rotation angle. After that, to further explore the knowledge from unlabeled data, we assign SG-CNN to a student network and a teacher network, and encourage the saliency predictions and self-supervised rotation predictions from these two networks to be consistent on the unlabeled data. Experimental results on seven widely-used benchmark datasets demonstrate that our network quantitatively and qualitatively outperforms the state-of-the-art methods.
Lei Zhu 0003, Xiaoqiang Wang 0007, Ping Li 0016, Xin Yang 0011, Qing Zhang 0006, Weiming Wang 0002, Carola-Bibiane Schönlieb, C. L. Philip Chen
IEEE Trans. Multim.5
2023 Pyramid Texture Filtering
abstract
We present a simple but effective technique to smooth out textures while preserving the prominent structures. Our method is built upon a key observation---the coarsest level in a Gaussian pyramid often naturally eliminates textures and summarizes the main image structures. This inspires our central idea for texture filtering, which is to progressively upsample the very low-resolution coarsest Gaussian pyramid level to a full-resolution texture smoothing result with well-preserved structures, under the guidance of each fine-scale Gaussian pyramid level and its associated Laplacian pyramid level. We show that our approach is effective to separate structure from texture of different scales, local contrasts, and forms, without degrading structures or introducing visual artifacts. We also demonstrate the applicability of our method on various applications including detail enhancement, image abstraction, HDR tone mapping, inverse halftoning, and LDR image enhancement. Code is available at https://rewindl.github.io/pyramid_texture_filtering/.
Qing Zhang 0006, Hao Jiang 0057, Yongwei Nie, Wei-Shi Zheng 0001
ACM Trans. Graph.1
2022 Progressively Generating Better Initial Guesses Towards Next Stages for High-Quality Human Motion Prediction
abstract
This paper presents a high-quality human motion pre-diction method that accurately predicts future human poses given observed ones. Our method is based on the observation that a good “initial guess” of the future poses is very helpful in improving the forecasting accuracy. This mo-tivates us to propose a novel two-stage prediction frame-work, including an init-prediction network that just computes the good guess and then a formal-prediction network that predicts the target future poses based on the guess. More importantly, we extend this idea further and design a multi-stage prediction framework where each stage pre-dicts initial guess for the next stage, which brings more performance gain. To fulfill the prediction task at each stage, we propose a network comprising Spatial Dense Graph Convolutional Networks (S-DGCN) and Temporal Dense Graph Convolutional Networks (T-DGCN). Alternatively executing the two networks helps extract spatiotem-poral features over the global receptive field of the whole pose sequence. All the above design choices cooperating together make our method outperform previous approaches by large margins: 6%-7% on Human3.6M, 5%-10% on CMU-MoCap, and 13%-16% on 3DPW. Code is available at https://github.com/705062791/PGBIG.
Tiezheng Ma, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li
CVPR4
2022 Diverse Human Motion Prediction via Gumbel-Softmax Sampling from an Auxiliary Space
abstract
Diverse human motion prediction aims at predicting multiple possible future pose sequences from a sequence of observed poses. Previous approaches usually employ deep generative networks to model the conditional distribution of data, and then randomly sample outcomes from the distribution. While different results can be obtained, they are usually the most likely ones which are not diverse enough. Recent work explicitly learns multiple modes of the conditional distribution via a deterministic network, which however can only cover a fixed number of modes within a limited range. In this paper, we propose a novel sampling strategy for sampling very diverse results from an imbalanced multimodal distribution learned by a deep generative model. Our method works by generating an auxiliary space and smartly making randomly sampling from the auxiliary space equivalent to the diverse sampling from the target distribution. We propose a simple yet effective network architecture that implements this novel sampling strategy, which incorporates a Gumbel-Softmax coefficient matrix sampling method and an aggressive diversity promoting hinge loss function. Extensive experiments demonstrate that our method significantly improves both the diversity and accuracy of the samplings compared with previous state-of-the-art sampling approaches. Code and pre-trained models are available at https://github.com/Droliven/diverse_sampling.
Lingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li
ACM Multimedia4
2022 Learning Multi-Scale Deep Image Prior for High-Quality Unsupervised Image Denoising
abstract
Abstract Recent methods on image denoising have achieved remarkable progress, benefiting mostly from supervised learning on massive noisy/clean image pairs and unsupervised learning on external noisy images. However, due to the domain gap between the training and testing images, these methods typically have limited applicability on unseen images. Although several attempts have been made to avoid the domain gap issue by learning denoising from singe noisy image itself, they are less effective in handling real‐world noise because of assuming the noise corruptions are independent and zero mean. In this paper, we go step further beyond prior work by presenting a novel unsupervised image denoising framework trained from single noisy image without making any explicit assumptions on the noise statistics. Our approach is built upon the deep image prior (DIP), which enables diverse image restoration tasks. However, as is, the denoising performance of DIP will significantly deteriorate on nonzero‐mean noise and is sensitive to the number of iterations. To overcome this problem, we propose to utilize multi‐scale deep image prior by imposing DIP across different image scales under the constraint of a scale consistency. Experiments on synthetic and real datasets demonstrate that our method performs favorably against the state‐of‐the‐art methods for image denoising.
Hao Jiang 0057, Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001
Comput. Graph. Forum2
2022 Unsupervised Intrinsic Image Decomposition Using Internal Self-Similarity Cues
abstract
Recent learning-based intrinsic image decomposition methods have achieved remarkable progress. However, they usually require massive ground truth intrinsic images for supervised learning, which limits their applicability on real-world images since obtaining ground truth intrinsic decomposition for natural images is very challenging. In this paper, we present an unsupervised framework that is able to learn the decomposition effectively from a single natural image by training solely with the image itself. Our approach is built upon the observations that the reflectance of a natural image typically has high internal self-similarity of patches, and a convolutional generation network tends to boost the self-similarity of an image when trained for image reconstruction. Based on the observations, an unsupervised intrinsic decomposition network (UIDNet) consisting of two fully convolutional encoder-decoder sub-networks, i.e., reflectance prediction network (RPN) and shading prediction network (SPN), is devised to decompose an image into reflectance and shading by promoting the internal self-similarity of the reflectance component, in a way that jointly trains RPN and SPN to reproduce the given image. A novel loss function is also designed to make effective the training for intrinsic decomposition. Experimental results on three benchmark real-world datasets demonstrate the superiority of the proposed method.
Qing Zhang 0006, Lei Zhu 0003, Wei Sun 0007, Chunxia Xiao, Wei-Shi Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 A Blind Color Separation Model for Faithful Palette-Based Image Recoloring
abstract
Palette-based image recoloring provides a simple yet effective way for color adjustment, which allows users to interactively manipulate the color of an image by editing a compact color palette. While remarkable progress has been made by previous methods, they have the common limitations that may produce unfaithful image recoloring results i.e., the obtained result does not respond faithfully to the palette adjustment, and tend to induce visual artifacts such as color bleeding and distortion. To address these limitations, we in this paper present a novel color separation model for palette-based recoloring. Akin to previous methods, our color separation model is built upon the assumption that color of each pixel in an image can be formulated as a linear combination of a small set of same basis colors. However, different from previous palette-based recoloring methods which typically rely on heuristic rules to build the color separation model, we experimentally reveal the underlying relationship between the color separation and the palette-based recoloring, and summarize three specialized color separation priors that allow more faithful palette-based recoloring. Based on these priors, we devise a blind color separation model that not only does not require known palette as input as done in previous methods, but also enables more effective palette-based recoloring with much less visual artifacts. Experiments on two datasets demonstrate that our method outperforms the state-of-the-art palette-based recoloring methods. In addition, we show some applications enabled by the proposed color separation model, including automatic pattern coloring generation, green screen keying and region-controllable color transfer.
Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Chunxia Xiao, Wei-Shi Zheng 0001
IEEE Trans. Multim.1
2021 A Multi-Task Network for Joint Specular Highlight Detection and Removal
abstract
Specular highlight detection and removal are fundamental and challenging tasks. Although recent methods have achieved promising results on the two tasks by training on synthetic training data in a supervised manner, they are typically solely designed for highlight detection or removal, and their performance usually deteriorates significantly on real-world images. In this paper, we present a novel network that aims to detect and remove highlights from natural images. To remove the domain gap between synthetic training samples and real test images, and support the investigation of learning-based approaches, we first introduce a dataset with about 16K real images, each of which has the corresponding ground truths of highlight detection and removal. Using the presented dataset, we develop a multi-task network for joint highlight detection and removal, based on a new specular highlight image formation model. Experiments on the benchmark datasets and our new dataset show that our approach clearly outperforms state-of-the-art methods for both highlight detection and removal.
Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Ping Li 0016, Chunxia Xiao
CVPR2
2021 MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion Prediction
abstract
Human motion prediction is a challenging task due to the stochasticity and aperiodicity of future poses. Recently, graph convolutional network has been proven to be very effective to learn dynamic relations among pose joints, which is helpful for pose prediction. On the other hand, one can abstract a human pose recursively to obtain a set of poses at multiple scales. With the increase of the abstraction level, the motion of the pose becomes more stable, which benefits pose prediction too. In this paper, we propose a novel Multi-Scale Residual Graph Convolution Network (MSR-GCN) for human pose prediction task in the manner of end-to-end. The GCNs are used to extract features from fine to coarse scale and then from coarse to fine scale. The extracted features at each scale are then combined and decoded to obtain the residuals between the input and target poses. Intermediate supervisions are imposed on all the predicted poses, which enforces the network to learn more representative features. Our proposed approach is evaluated on two standard benchmark datasets, i.e., the Human3.6M dataset and the CMU Mocap dataset. Experimental results demonstrate that our method outperforms the state-of-the-art approaches. Code and pre-trained models are available at https://github.com/Droliven/MSRGCN.
Lingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li
ICCV4
2021 A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame Prediction
abstract
In this paper, we propose HF2-VAD, a Hybrid framework that integrates Flow reconstruction and Frame prediction seamlessly to handle Video Anomaly Detection. Firstly, we design the network of ML-MemAE-SC (Multi-Level Memory modules in an Autoencoder with Skip Connections) to memorize normal patterns for optical flow reconstruction so that abnormal events can be sensitively identified with larger flow reconstruction errors. More importantly, conditioned on the reconstructed flows, we then employ a Conditional Variational Autoencoder (CVAE), which captures the high correlation between video frame and optical flow, to predict the next frame given several previous frames. By CVAE, the quality of flow reconstruction essentially influences that of frame prediction. Therefore, poorly reconstructed optical flows of abnormal events further deteriorate the quality of the final predicted future frame, making the anomalies more detectable. Experimental results demonstrate the effectiveness of the proposed method. Code is available at https://github.com/LiUzHiAn/hf2vad.
Zhian Liu, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li
ICCV4
2021 From Synthetic to Real: Image Dehazing Collaborating with Unlabeled Real Data
abstract
Single image dehazing is a challenging task, for which the domain shift between synthetic training data and real-world testing images usually leads to degradation of existing methods. To address this issue, we propose a novel image dehazing framework collaborating with unlabeled real data. First, we develop a disentangled image dehazing network (DID-Net), which disentangles the feature representations into three component maps, i.e. the latent haze-free image, the transmission map, and the global atmospheric light estimate, respecting the physical model of a haze process. Our DID-Net predicts the three component maps by progressively integrating features across scales, and refines each map by passing an independent refinement network. Then a disentangled-consistency mean-teacher network (DMT-Net) is employed to collaborate unlabeled real data for boosting single image dehazing. Specifically, we encourage the coarse predictions and refinements of each disentangled component to be consistent between the student and teacher networks by using a consistency loss on unlabeled real data. We make comparison with 13 state-of-the-art dehazing methods on a new collected dataset (Haze4K) and two widely-used dehazing datasets (i.e., SOTS and HazeRD), as well as on real-world hazy images. Experimental results demonstrate that our method has obvious quantitative and qualitative improvements over the existing methods.
Lei Zhu 0003, Shunda Pei, Huazhu Fu, Harry Qin, Qing Zhang 0006, Wei Feng 0005
ACM Multimedia6
2021 Joint regression and learning from pairwise rankings for personalized image aesthetic assessment
abstract
Recent image aesthetic assessment methods have achieved remarkable progress due to the emergence of deep convolutional neural networks (CNNs). However, these methods focus primarily on predicting generally perceived preference of an image, making them usually have limited practicability, since each user may have completely different preferences for the same image. To address this problem, this paper presents a novel approach for predicting personalized image aesthetics that fit an individual user’s personal taste. We achieve this in a coarse to fine manner, by joint regression and learning from pairwise rankings. Specifically, we first collect a small subset of personal images from a user and invite him/her to rank the preference of some randomly sampled image pairs. We then search for the K -nearest neighbors of the personal images within a large-scale dataset labeled with average human aesthetic scores, and use these images as well as the associated scores to train a generic aesthetic assessment model by CNN-based regression. Next, we fine-tune the generic model to accommodate the personal preference by training over the rankings with a pairwise hinge loss. Experiments demonstrate that our method can effectively learn personalized image aesthetic preferences, clearly outperforming state-of-the-art methods. Moreover, we show that the learned personalized image aesthetic benefits a wide variety of applications.
Qing Zhang 0006, Jian-Hao Fan, Wei Sun 0007, Wei-Shi Zheng 0001
Comput. Vis. Media2
2021 Enhancing Underexposed Photos Using Perceptually Bidirectional Similarity
abstract
Although remarkable progress has been made, existing methods for enhancing underexposed photos tend to produce visually unpleasing results due to the existence of visual artifacts (e.g., color distortion, loss of details and uneven exposure). We observed that this is because they fail to ensure the perceptual consistency of visual information between the source underexposed image and its enhanced output. To obtain high-quality results free of these artifacts, we present a novel underexposed photo enhancement approach that is able to maintain the perceptual consistency. We achieve this by proposing an effective criterion, referred to as perceptually bidirectional similarity, which explicitly describes how to ensure the perceptual consistency. Particularly, we adopt the Retinex theory and cast the enhancement problem as a constrained illumination estimation optimization, where we formulate perceptually bidirectional similarity as constraints on illumination and solve for the illumination which can recover the desired artifact-free enhancement results. In addition, we describe a video enhancement framework that adopts the presented illumination estimation for handling underexposed videos. To this end, a probabilistic approach is introduced to propagate illuminations of sampled keyframes to the entire video by tackling a Bayesian Maximum A Posteriori problem. Extensive experiments demonstrate the superiority of our method over the state-of-the-art methods.
Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Chunxia Xiao, Wei-Shi Zheng 0001
IEEE Trans. Multim.1
2021 Monte Carlo denoising via auxiliary feature guided self-attention
abstract
While self-attention has been successfully applied in a variety of natural language processing and computer vision tasks, its application in Monte Carlo (MC) image denoising has not yet been well explored. This paper presents a self-attention based MC denoising deep learning network based on the fact that self-attention is essentially non-local means filtering in the embedding space which makes it inherently very suitable for the denoising task. Particularly, we modify the standard self-attention mechanism to an auxiliary feature guided self-attention that considers the by-products (e.g., auxiliary feature buffers) of the MC rendering process. As a critical prerequisite to fully exploit the performance of self-attention, we design a multi-scale feature extraction stage, which provides a rich set of raw features for the later self-attention module. As self-attention poses a high computational complexity, we describe several ways that accelerate it. Ablation experiments validate the necessity and effectiveness of the above design choices. Comparison experiments show that the proposed self-attention based MC denoising method outperforms the current state-of-the-art methods.
Yongwei Nie, Chengjiang Long, Wenjun Xu 0002, Qing Zhang 0006, Guiqing Li
ACM Trans. Graph.5
2020 Deep Camouflage Images
abstract
This paper addresses the problem of creating camouflage images. Such images typically contain one or more hidden objects embedded into a background image, so that viewers are required to consciously focus to discover them. Previous methods basically rely on hand-crafted features and texture synthesis to create camouflage images. However, due to lack of reliable understanding of what essentially makes an object recognizable, they typically result in either complete standout or complete invisible hidden objects. Moreover, they may fail to produce seamless and natural images because of the sensitivity to appearance differences. To overcome these limitations, we present a novel neural style transfer approach that adopts the visual perception mechanism to create camouflage images, which allows us to hide objects more effectively while producing natural-looking results. In particular, we design an attention-aware camouflage loss to adaptively mask out information that make the hidden objects visually standout, and also leave subtle yet enough feature clues for viewers to perceive the hidden objects. To remove the appearance discontinuities between the hidden objects and the background, we formulate a naturalness regularization to constrain the hidden objects to maintain the manifold structure of the covered background. Extensive experiments show the advantages of our approach over existing camouflage methods and state-of-the-art neural style transfer algorithms.
Qing Zhang 0006, Gelin Yin, Yongwei Nie, Wei-Shi Zheng 0001
AAAI1
2020 Learning to Detect Specular Highlights from Real-world Images
abstract
Specular highlight detection is a challenging problem, and has many applications such as shiny object detection and light source estimation. Although various highlight detection methods have been proposed, they fail to disambiguate bright material surfaces from highlights, and cannot handle non-white-balanced images. Moreover, at present, there is still no benchmark dataset for highlight detection. In this paper, we present a large-scale real-world highlight dataset containing a rich variety of material categories, with diverse highlight shapes and appearances, in which each image is with an annotated ground-truth mask. Based on the dataset, we develop a deep learning-based specular highlight detection network (SHDNet) leveraging multi-scale context contrasted features to accurately detect specular highlights of varying scales. In addition, we design a binary cross-entropy (BCE) loss and an intersection-over-union edge (IoUE) loss for our network. Compared with existing highlight detection methods, our method can accurately detect highlights of different sizes, while effectively excluding the non-highlight regions, such as bright materials, non-specular as well as colored lighting, and even light sources.
Gang Fu 0003, Qing Zhang 0006, Qifeng Lin, Lei Zhu 0003, Chunxia Xiao
ACM Multimedia2
2020 Interactive Contour Extraction via Sketch-Alike Dense-Validation Optimization
abstract
We propose an interactive contour extraction method inspired by a skill often adopted in sketching: an artist usually sketches an object by first drawing lots of short, directional, and redundant strokes, then following these small strokes to draw the final outline of the object. Our method simulates this process. To extract a contour, our method relies on user interaction, which provides us with a narrow band containing the target contour. Then, we densely sample sub-bands from the whole band, with each sub-band containing a local segment of the target contour. We design a curve-centered coordinate system in which a dynamic programming algorithm is proposed to extract the local segment in each sub-band. The local segment is guaranteed to be as evident and smooth as possible, to mimic the strokes sketched by the artist. Finally, we integrate all local segments of all sub-bands together to obtain the whole target contour based on the weighted principal component analysis. Our method can extract high-quality object contours due to the dense validations among local segments. That is, even if one segment deviates from the right location, several other segments in its local neighborhood can correct it in the integration stage. Both quantitative experiments and a user study demonstrate the effectiveness of the proposed method.
Yongwei Nie, Ping Li 0016, Qing Zhang 0006, Zhensong Zhang, Guiqing Li, Hanqiu Sun
IEEE Trans. Circuits Syst. Video Technol.4
2020 Collision-Free Video Synopsis Incorporating Object Speed and Size Changes
abstract
This paper presents a new surveillance video synopsis method which performs much better than previous approaches in terms of both compression ratio and artifact. Previously, a surveillance video was usually compressed by shifting the moving objects of that video forward along the time axis, which inevitably yielded serious collision and chronological disorder artifacts between the shifted objects. The main observation of this paper is that these artifacts can be alleviated by changing the speed or size of the objects, since with varied speed and size the objects can move more flexibly to avoid collision points or to keep chronological relationships. Based on this observation, we propose a video synopsis method that performs object shifting, speed changing, and size scaling simultaneously. We show how to integrate the three heterogeneous operations into a single optimization framework and achieve high-quality synopsis results. Unlike previous approaches that usually use alternative optimization strategies to solve synopsis optimizations, we develop a Metropolis sampling algorithm to find the solution for our three-variable optimization problem. A variety of experiments demonstrate the effectiveness of our method.
Yongwei Nie, Zhenkai Li, Zhensong Zhang, Qing Zhang 0006, Tiezheng Ma, Hanqiu Sun
IEEE Trans. Image Process.4
2020 Multi-View Video Synopsis via Simultaneous Object-Shifting and View-Switching Optimization
abstract
We present a method for synopsizing multiple videos captured by a set of surveillance cameras with some overlapped field-of-views. Currently, object-based approaches that directly shift objects along the time axis are already able to compute compact synopsis results for multiple surveillance videos. The challenge is how to present the multiple synopsis results in a more compact and understandable way. Previous approaches show them side by side on the screen, which however is difficult for user to comprehend. In this paper, we solve the problem by joint object-shifting and camera view-switching. Firstly, we synchronize the input videos, and group the same object in different videos together. Then we shift the groups of objects along the time axis to obtain multiple synopsis videos. Instead of showing them simultaneously, we just show one of them at each time, and allow to switch among the views of different synopsis videos. In this view switching way, we obtain just a single synopsis results consisting of content from all the input videos, which is much easier for user to follow and understand. To obtain the best synopsis result, we construct a simultaneous object-shifting and view-switching optimization framework instead of solving them separately. We also present an alternative optimization strategy composed of graph cuts and dynamic programming to solve the unified optimization. Experiments demonstrate that our single synopsis video generated from multiple input videos is compact, complete, and easy to understand.
Zhensong Zhang, Yongwei Nie, Hanqiu Sun, Qing Zhang 0006, Qiuxia Lai, Guiqing Li, Mingyu Xiao 0001
IEEE Trans. Image Process.4
2020 Effective Video Stabilization via Joint Trajectory Smoothing and Frame Warping
abstract
Video stabilization is usually composed of three stages: feature trajectory extraction, trajectory smoothing, and frame warping. Most previous approaches view them as three separate stages. This paper proposes a method combining the last two stages, namely the trajectory smoothing and frame warping stages, into a single optimization framework. The novelty exists in the way of how we combine them: the trajectory smoothing part plays a major role while the frame warping part plays an auxiliary role. With this kind of design, we can conveniently increase the strength of the trajectory smoothing part by a robust first-order derivative term, which makes it possible to produce very aggressive stabilization effects. On the other hand, we adopt adaptive weighting mechanisms in the frame warping part, to follow the smoothed trajectories as much as possible while regularizing other places as similar as possible. Our method is robust to utilize both foreground and background features, and very short trajectories. The utilization of all these information in turn increases the accuracy of the proposed method. We also provide a simplified implementation of our method, which is less accurate but more efficient. Experiments on various kinds of videos demonstrate the effectiveness of our method.
Tiezheng Ma, Yongwei Nie, Qing Zhang 0006, Zhensong Zhang, Hanqiu Sun, Guiqing Li
IEEE Trans. Vis. Comput. Graph.3
2019 Underexposed Photo Enhancement Using Deep Illumination Estimation
abstract
This paper presents a new neural network for enhancing underexposed photos. Instead of directly learning an image-to-image mapping as previous work, we introduce intermediate illumination in our network to associate the input with expected enhancement result, which augments the network's capability to learn complex photographic adjustment from expert-retouched input/output image pairs. Based on this model, we formulate a loss function that adopts constraints and priors on the illumination, prepare a new dataset of 3,000 underexposed image pairs, and train the network to effectively learn a rich variety of adjustment for diverse lighting conditions. By these means, our network is able to recover clear details, distinct contrast, and natural color in the enhancement results. We perform extensive experiments on the benchmark MIT-Adobe FiveK dataset and our new dataset, and show that our network is effective to deal with previously challenging images.
Ruixing Wang, Qing Zhang 0006, Chi-Wing Fu, Xiaoyong Shen, Wei-Shi Zheng 0001, Jiaya Jia
CVPR2
2019 Deep Multi-Model Fusion for Single-Image Dehazing
abstract
This paper presents a deep multi-model fusion network to attentively integrate multiple models to separate layers and boost the performance in single-image dehazing. To do so, we first formulate the attentional feature integration module to maximize the integration of the convolutional neural network (CNN) features at different CNN layers and generate the attentional multi-level integrated features (AMLIF). Then, from the AMLIF, we further predict a haze-free result for an atmospheric scattering model, as well as for four haze-layer separation models, and then fuse the results together to produce the final haze-free image. To evaluate the effectiveness of our method, we compare our network with several state-of-the-art methods on two widely-used dehazing benchmark datasets, as well as on two sets of real-world hazy images. Experimental results demonstrate clear quantitative and qualitative improvements of our method over the state-of-the-arts.
Zijun Deng, Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Qing Zhang 0006, Harry Qin, Pheng-Ann Heng
ICCV6
2019 Towards High-Quality Intrinsic Images in the Wild
abstract
We address the intrinsic image decomposition problem for separating an image into its intrinsic images, i.e, a reflectance layer and a shading layer. Although this problem has been studied for decades, it remains a significant challenge, particularly for real-world images. In this paper, we present a novel method for estimating high-quality intrinsic images for real-world images. Our method is built upon two observations on real-world images: (i) reflectance is generally sparse and there are limited number of reflectance values in an image; (ii) shading usually has locally smooth transition. Based on the two observations, we formulate the decomposition problem into an optimization framework, where we encourage the reflectance sparseness by globally confining the number of reflectance discontinuities among neighboring pixels using an L_0 norm, and utilize a total variation for maintaining locally smooth shading. We employ two benchmark datasets and perform various experiments to evaluate our method. Experimental results show that our method outperforms state-of-the-art methods, both qualitatively and quantitatively.
Gang Fu 0003, Qing Zhang 0006, Chunxia Xiao
ICME2
2019 DBDNet: Learning Bi-directional Dynamics for Early Action Prediction
abstract
Predicting future actions from observed partial videos is very challenging as the missing future is uncertain and sometimes has multiple possibilities. To obtain a reliable future estimation, a novel encoder-decoder architecture is proposed for integrating the tasks of synthesizing future motions from observed videos and reconstructing observed motions from synthesized future motions in an unified framework, which can capture the bi-directional dynamics depicted in partial videos along the temporal (past-to-future) direction and reverse chronological (future-back-to-past) direction. We then employ a bi-directional long short-term memory (Bi-LSTM) architecture to exploit the learned bi-directional dynamics for predicting early actions. Our experiments on two benchmark action datasets show that learning bi-directional dynamics benefits the early action prediction and our system clearly outperforms the state-of-the-art methods.
Guoliang Pang, Xionghui Wang, Jianfang Hu, Qing Zhang 0006, Wei-Shi Zheng 0001
IJCAI4
2019 Specular Highlight Removal for Real-world Images
abstract
Abstract Removing specular highlight in an image is a fundamental research problem in computer vision and computer graphics. While various methods have been proposed, they typically do not work well for real‐world images due to the presence of rich textures, complex materials, hard shadows, occlusions and color illumination, etc. In this paper, we present a novel specular highlight removal method for real‐world images. Our approach is based on two observations of the real‐world images: (i) the specular highlight is often small in size and sparse in distribution; (ii) the remaining diffuse image can be represented by linear combination of a small number of basis colors with the sparse encoding coefficients. Based on the two observations, we design an optimization framework for simultaneously estimating the diffuse and specular highlight images from a single image. Specifically, we recover the diffuse components of those regions with specular highlight by encouraging the encoding coefficients sparseness using L0 norm. Moreover, the encoding coefficients and specular highlight are also subject to the non‐negativity according to the additive color mixing theory and the illumination definition, respectively. Extensive experiments have been performed on a variety of images to validate the effectiveness of the proposed method and its superiority over the previous methods.
Gang Fu 0003, Qing Zhang 0006, Chengfang Song, Qifeng Lin, Chunxia Xiao
Comput. Graph. Forum2
2019 Dual Illumination Estimation for Robust Exposure Correction
abstract
Abstract Exposure correction is one of the fundamental tasks in image processing and computational photography. While various methods have been proposed, they either fail to produce visually pleasing results, or only work well for limited types of image (e.g., underexposed images). In this paper, we present a novel automatic exposure correction method, which is able to robustly produce high‐quality results for images of various exposure conditions (e.g., underexposed, overexposed, and partially under‐ and over‐exposed). At the core of our approach is the proposed dual illumination estimation, where we separately cast the under‐and over‐exposure correction as trivial illumination estimation of the input image and the inverted input image. By performing dual illumination estimation, we obtain two intermediate exposure correction results for the input image, with one fixes the underexposed regions and the other one restores the overexposed regions. A multi‐exposure image fusion technique is then employed to adaptively blend the visually best exposed parts in the two intermediate exposure correction images and the input image into a globally well‐exposed image. Experiments on a number of challenging images demonstrate the effectiveness of the proposed approach and its superiority over the state‐of‐the‐art methods and popular automatic exposure correction tools.
Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001
Comput. Graph. Forum1
2018 High-Quality Exposure Correction of Underexposed Photos
abstract
We address the problem of correcting the exposure of underexposed photos. Previous methods have tackled this problem from many different perspectives and achieved remarkable progress. However, they usually fail to produce natural-looking results due to the existence of visual artifacts such as color distortion, loss of detail, exposure inconsistency, etc. We find that the main reason why existing methods induce these artifacts is because they break a perceptually similarity between the input and output. Based on this observation, an effective criterion, termed as perceptually bidirectional similarity (PBS) is proposed. Based on this criterion and the Retinex theory, we cast the exposure correction problem as an illumination estimation optimization, where PBS is defined as three constraints for estimating illumination that can generate the desired result with even exposure, vivid color and clear textures. Qualitative and quantitative comparisons, and the user study demonstrate the superiority of our method over the state-of-the-art methods.
Qing Zhang 0006, Ganzhao Yuan, Chunxia Xiao, Lei Zhu 0003, Wei-Shi Zheng 0001
ACM Multimedia1
2017 Palette-Based Image Recoloring Using Color Decomposition Optimization
abstract
Previous works on palette-based color manipulation typically fail to produce visually pleasing results with vivid color and natural appearance. In this paper, we present an approach to edit colors of an image by adjusting a compact color palette. Different from existing methods that fail to preserve inherent color characteristics residing in the source image, we propose a color decomposition optimization for flexible recoloring while retaining these characteristics. For an input image, we first employ a variant of the k -means algorithm to create a palette consisting of a small set of most representative colors. Next, we propose a color decomposition optimization to decompose colors of the entire image into linear combinations of basis colors in the palette. The captured linear relationships then allow us to recolor the image by recombining the coding coefficients with a user-modified palette. Qualitative comparisons with existing methods show that our approach can more effectively recolor images. Further user study quantitatively demonstrates that our method is a good candidate for color manipulation tasks. In addition, we showcase some applications enabled by our method, including pattern colorings suggesting, color transfer, tissue staining analysis and color image segmentation.
Qing Zhang 0006, Chunxia Xiao, Hanqiu Sun
IEEE Trans. Image Process.1
2016 Video Background Completion Using Motion-Guided Pixel Assignment Optimization
abstract
Background completion for consumer videos captured by free-moving cameras is a challenging problem. In this paper, we present a new approach to complete the holes left by removing objects with motion-guided pixels assignment optimization. We first estimate the motion field in the holes by applying a two-step motion propagation method. Then, using estimated motion field as guidance, the missing parts of the video are completed by performing pixels assignment optimization based on the Markov random field, which optimally assigns available pixels from other neighboring video frames to the missing regions. Finally, we present an illumination-adjusting approach to eliminate the illumination inconsistency in the completed holes. We validate our method on a variety of videos captured by free-moving cameras. Compared with previous methods, our method works better to keep the completed background spatiotemporally coherent, to complete video background with much depth discontinuity and to make the illumination consistent in the completed region.
Qing Zhang 0006, Chunxia Xiao
IEEE Trans. Circuits Syst. Video Technol.2
2016 Underexposed Video Enhancement via Perception-Driven Progressive Fusion
abstract
Underexposed video enhancement aims at revealing hidden details that are barely noticeable in LDR video frames with noise. Previous work typically relies on a single heuristic tone mapping curve to expand the dynamic range, which inevitably leads to uneven exposure and visual artifacts. In this paper, we present a novel approach for underexposed video enhancement using an efficient perception-driven progressive fusion. For an input underexposed video, we first remap each video frame using a series of tentative tone mapping curves to generate an multi-exposure image sequence that contains different exposed versions of the original video frame. Guided by some visual perception quality measures encoding the desirable exposed appearance, we locate all the best exposed regions from multi-exposure image sequences and then integrate them into a well-exposed video in a temporally consistent manner. Finally, we further perform an effective texture-preserving spatio-temporal filtering on this well-exposed video to obtain a high-quality noise-free result. Experimental results have shown that the enhanced video exhibits uniform exposure, brings out noticeable details, preserves temporal coherence, and avoids visual artifacts. Besides, we demonstrate applications of our approach to a set of problems including video dehazing, video denoising and HDR video reconstruction.
Qing Zhang 0006, Yongwei Nie, Ling Zhang 0017, Chunxia Xiao
IEEE Trans. Vis. Comput. Graph.1
2015 Shadow Remover: Image Shadow Removal Based on Illumination Recovering Optimization
abstract
In this paper, we present a novel shadow removal system for single natural images as well as color aerial images using an illumination recovering optimization method. We first adaptively decompose the input image into overlapped patches according to the shadow distribution. Then, by building the correspondence between the shadow patch and the lit patch based on texture similarity, we construct an optimized illumination recovering operator, which effectively removes the shadows and recovers the texture detail under the shadow patches. Based on coherent optimization processing among the neighboring patches, we finally produce high-quality shadow-free results with consistent illumination. Our shadow removal system is simple and effective, and can process shadow images with rich texture types and nonuniform shadows. The illumination of shadow-free results is consistent with that of surrounding environment. We further present several shadow editing applications to illustrate the versatility of the proposed method.
Ling Zhang 0017, Qing Zhang 0006, Chunxia Xiao
IEEE Trans. Image Process.2
2014 Cloud Detection of RGB Color Aerial Photographs by Progressive Refinement Scheme
abstract
In this paper, we propose an automatic and effective cloud detection algorithm for color aerial photographs. Based on the properties derived from observations and statistical results on a large number of color aerial photographs with cloud layers, we present a novel progressive refinement scheme for detecting clouds in the color aerial photographs. We first construct a significance map which highlights the difference between cloud regions and noncloud regions. Based on the significance map and the proposed optimal threshold setting, we obtain a coarse cloud detection result which classifies the input aerial photograph into the candidate cloud regions and noncloud regions. In order to accurately detect the cloud regions from the candidate cloud regions, we then construct a robust detail map derived from a multiscale bilateral decomposition to guide us in removing noncloud regions from the candidate cloud regions. Finally, we further perform a guided feathering to achieve our final cloud detection result, which detects semitransparent cloud pixels around the boundaries of cloud regions. The proposed method is evaluated in terms of both visual and quantitative comparisons, and the evaluation results show that our proposed method works well for the cloud detection of color aerial photographs.
Qing Zhang 0006, Chunxia Xiao
IEEE Trans. Geosci. Remote. Sens.1
2013 Video retargeting combining warping and summarizing optimization
Yongwei Nie, Qing Zhang 0006, Renfang Wang, Chunxia Xiao
Vis. Comput.2