VLDB 2026 Research / reviewers in the wild / expert
Jue Wang 0001
dblp:69/393-1
· DBLP profile ↗
150ranked-venue papers
16as first author
60since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 125 · 13 first-author · 47 since 2021Artificial intelligence and machine learning · 88 · 9 first-author · 42 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency AdapterabstractIn this work, we present Patch-Adapter, an effective framework for high-resolution text-guided image inpainting. Unlike existing methods limited to lower resolutions, our approach achieves 4K+ resolution while maintaining precise content consistency and prompt alignment, two critical challenges in image inpainting that intensify with increasing resolution and texture complexity. Patch-Adapter leverages a two-stage adapter architecture to scale the diffusion model's resolution from 1K to 4K+ without requiring structural overhauls: (1) Dual Context Adapter learns coherence between masked and unmasked regions at reduced resolutions to establish global structural consistency; and (2) Reference Patch Adapter implements a patch-level attention mechanism for full-resolution inpainting, preserving local detail fidelity through adaptive feature fusion. This dual-stage architecture uniquely addresses the scalability gap in high-resolution inpainting by decoupling global semantics from localized refinement. Experiments demonstrate that Patch-Adapter not only resolves artifacts common in large-scale inpainting but also achieves state-of-the-art performance on the OpenImages and Photo-Concept-Bucket datasets, outperforming existing methods in both perceptual quality and text-prompt adherence. Qirui Sun, Wang Luyang, Chaoyu Feng, Jue Wang 0001, Shuaicheng Liu |
ICCV | 9 |
| 2025 | Understanding adversarial robustness against on-manifold adversarial examples
Jiancong Xiao, Liusha Yang, Yanbo Fan, Jue Wang 0001, Zhi-Quan Luo |
Pattern Recognit. | 4 |
| 2024 | Improving Fast Adversarial Training With Prior-Guided KnowledgeabstractFast adversarial training (FAT) is an efficient method to improve robustness in white-box attack scenarios. However, the original FAT suffers from catastrophic overfitting, which dramatically and suddenly reduces robustness after a few training epochs. Although various FAT variants have been proposed to prevent overfitting, they require high training time. In this paper, we investigate the relationship between adversarial example quality and catastrophic overfitting by comparing the training processes of standard adversarial training and FAT. We find that catastrophic overfitting occurs when the attack success rate of adversarial examples becomes worse. Based on this observation, we propose a positive prior-guided adversarial initialization to prevent overfitting by improving adversarial example quality without extra training time. This initialization is generated by using high-quality adversarial perturbations from the historical training process. We provide theoretical analysis for the proposed initialization and propose a prior-guided regularization method that boosts the smoothness of the loss function. Additionally, we design a prior-guided ensemble FAT method that averages the different model weights of historical models using different decay rates. Our proposed method, called FGSM-PGK, assembles the prior-guided knowledge, i.e., the prior-guided initialization and model weights, acquired during the historical training process. The proposed method can effectively improve the model's adversarial robustness in white-box attack scenarios. Evaluations of four datasets demonstrate the superiority of the proposed method. Xiaojun Jia, Yong Zhang 0034, Xingxing Wei 0001, Baoyuan Wu, Ke Ma 0001, Jue Wang 0001, Xiaochun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | CoordFill: Efficient High-Resolution Image Inpainting via Parameterized Coordinate QueryingabstractImage inpainting aims to fill the missing hole of the input. It is hard to solve this task efficiently when facing high-resolution images due to two reasons: (1) Large reception field needs to be handled for high-resolution image inpainting. (2) The general encoder and decoder network synthesizes many background pixels synchronously due to the form of the image matrix. In this paper, we try to break the above limitations for the first time thanks to the recent development of continuous implicit representation. In detail, we down-sample and encode the degraded image to produce the spatial-adaptive parameters for each spatial patch via an attentional Fast Fourier Convolution (FFC)-based parameter generation network. Then, we take these parameters as the weights and biases of a series of multi-layer perceptron (MLP), where the input is the encoded continuous coordinates and the output is the synthesized color value. Thanks to the proposed structure, we only encode the high-resolution image in a relatively low resolution for larger reception field capturing. Then, the continuous position encoding will be helpful to synthesize the photo-realistic high-frequency textures by re-sampling the coordinate in a higher resolution. Also, our framework enables us to query the coordinates of missing pixels only in parallel, yielding a more efficient solution than the previous methods. Experiments show that the proposed method achieves real-time performance on the 2048X2048 images using a single GTX 2080 Ti GPU and can handle 4096X4096 images, with much better performance than existing state-of-the-art methods visually and numerically. The code is available at: https://github.com/NiFangBaAGe/CoordFill. Weihuang Liu, Xiaodong Cun, Chi-Man Pun, Menghan Xia, Yong Zhang 0034, Jue Wang 0001 |
AAAI | 6 |
| 2023 | Truncate-Split-Contrast: A Framework for Learning from Mislabeled VideosabstractLearning with noisy label is a classic problem that has been extensively studied for image tasks, but much less for video in the literature. A straightforward migration from images to videos without considering temporal semantics and computational cost is not a sound choice. In this paper, we propose two new strategies for video analysis with noisy labels: 1) a lightweight channel selection method dubbed as Channel Truncation for feature-based label noise detection. This method selects the most discriminative channels to split clean and noisy instances in each category. 2) A novel contrastive strategy dubbed as Noise Contrastive Learning, which constructs the relationship between clean and noisy instances to regularize model training. Experiments on three well-known benchmark datasets for video classification show that our proposed truNcatE-split-contrAsT (NEAT) significantly outperforms the existing baselines. By reducing the dimension to 10% of it, our method achieves over 0.4 noise detection F1-score and 5% classification accuracy improvement on Mini-Kinetics dataset under severe noise (symmetric-80%). Thanks to Noise Contrastive Learning, the average classification accuracy improvement on Mini-Kinetics and Sth-Sth-V1 is over 1.6%. Zixiao Wang 0001, Junwu Weng, Chun Yuan 0003, Jue Wang 0001 |
AAAI | 4 |
| 2023 | UV Volumes for Real-time Rendering of Editable Free-view Human PerformanceabstractNeural volume rendering enables photo-realistic renderings of a human performer in free-view, a critical task in immersive VR/AR applications. But the practice is severely limited by high computational costs in the rendering process. To solve this problem, we propose the UV Volumes, a new approach that can render an editable free-view video of a human performer in real-time. It separates the high-frequency (i.e., non-smooth) human appearance from the 3D volume, and encodes them into 2D neural texture stacks (NTS). The smooth UV volumes allow much smaller and shallower neural networks to obtain densities and texture coordinates in 3D while capturing detailed appearance in 2D NTS. For editability, the mapping between the parameterized human model and the smooth texture coordinates allows us a better generalization on novel poses and shapes. Furthermore, the use of NTS enables interesting applications, e.g., retexturing. Extensive experiments on CMU Panoptic, ZJU Mocap, and H36M datasets show that our model can render$960\times 540$images in 30FPS on average with comparable photo-realism to state-of-the-art methods. The project and supplementary materials are available at https://fanegg.github.io/UV-Volumes. Xuan Wang 0009, Qi Zhang 0029, Xiaoyu Li 0002, Yu Guo 0006, Jue Wang 0001, Fei Wang 0008 |
CVPR | 7 |
| 2023 | Patch-Based 3D Natural Scene Generation from a Single ExampleabstractWe target a 3D generative model for general natural scenes that are typically unique and intricate. Lacking the necessary volumes of training data, along with the difficulties of having ad hoc designs in presence of varying scene characteristics, renders existing setups intractable. Inspired by classical patch-based image models, we advocate for synthesizing 3D scenes at the patch level, given a single example. At the core of this work lies important algorithmic designs w.r.t the scene representation and generative patch nearest-neighbor module, that address unique challenges arising from lifting classical 2D patch-based framework to 3D generation. These design choices, on a collective level, contribute to a robust, effective, and efficient model that can generate high-quality general natural scenes with both realistic geometric structure and visual appearance, in large quantities and varieties, as demonstrated upon a variety of exemplar scenes. Data and code can be found at http://wyysf-98.github.io/Sin3DGen. Xuelin Chen, Jue Wang 0001, Baoquan Chen |
CVPR | 3 |
| 2023 | Fine-Grained Face Swapping Via Regional GAN InversionabstractWe present a novel paradigm for high-fidelity face swapping that faithfully preserves the desired subtle geometry and texture details. We rethink face swapping from the perspective of fine-grained face editing, i.e., “editing for swapping” (E4S), and propose a framework that is based on the explicit disentanglement of the shape and texture of facial components. Following the E4S principle, our framework enables both global and local swapping of facial features, as well as controlling the amount of partial swapping specified by the user. Furthermore, the E4S paradigm is in-herently capable of handling facial occlusions by means of facial masks. At the core of our system lies a novel Regional GAN Inversion (RGI) method, which allows the explicit disentanglement of shape and texture. It also allows face swapping to be performed in the latent space of Style-GAN. Specifically, we design a multi-scale mask-guided encoder to project the texture of each facial component into regional style codes. We also design a mask-guided injection module to manipulate the feature maps with the style codes. Based on the disentanglement, face swapping is re-formulated as a simplified problem of style and mask swapping. Extensive experiments and comparisons with current state-of-the-art methods demonstrate the superiority of our approach in preserving texture and shape details, as well as working with high resolution images. The project page is https://e4s2022.github.io Zhian Liu, Maomao Li, Yong Zhang 0034, Cairong Wang, Qi Zhang 0029, Jue Wang 0001, Yongwei Nie |
CVPR | 6 |
| 2023 | CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorabstractSpeech-driven 3D facial animation has been widely studied, yet there is still a gap to achieving realism and vividness due to the highly ill-posed nature and scarcity of audio-visual data. Existing works typically formulate the cross-modal mapping into a regression task, which suffers from the regression-to-mean problem leading to over-smoothed facial motions. In this paper, we propose to cast speech-driven facial animation as a code query task in a finite proxy space of the learned codebook, which effectively promotes the vividness of the generated motions by reducing the cross-modal mapping uncertainty. The codebook is learned by self-reconstruction over real facial motions and thus embedded with realistic facial motion priors. Over the discrete motion space, a temporal autoregressive model is employed to sequentially synthesize facial motions from the input speech signal, which guarantees lip-sync as well as plausible facial expressions. We demonstrate that our approach outperforms current state-of-the-art methods both qualitatively and quantitatively. Also, a user study further justifies our superiority in perceptual quality. Code and video demo are available at https://doubiiu.github.io/projects/codetalker. Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang 0001, Tien-Tsin Wong |
CVPR | 5 |
| 2023 | ACR: Attention Collaboration-based Regressor for Arbitrary Two-Hand ReconstructionabstractReconstructing two hands from monocular RGB images is challenging due to frequent occlusion and mutual confusion. Existing methods mainly learn an entangled representation to encode two interacting hands, which are in-credibly fragile to impaired interaction, such as truncated hands, separate hands, or external occlusion. This paper presents ACR (Attention Collaboration-based Regres-sor), which makes the first attempt to reconstruct hands in arbitrary scenarios. To achieve this, ACR explicitly miti-gates interdependencies between hands and between parts by leveraging center and part-based attention for feature extraction. However, reducing interdependence helps re-lease the input constraint while weakening the mutual reasoning about reconstructing the interacting hands. Thus, based on center attention, ACR also learns cross-hand prior that handle the interacting hands better. We evaluate our method on various types of hand reconstruction datasets. Our method significantly outperforms the best interacting-hand approaches on the InterHand2.6M dataset while yielding comparable performance with the state-of-the-art single-hand methods on the FreiHand dataset. More qualitative results on in-the-wild and hand-object interaction datasets and web images/videos further demonstrate the effectiveness of our approach for arbitrary hand reconstruction. Our code is available at this link11https://github.com/ZhengdiYu/Arbitrary-Hands-3D-Reconstruction. Zhengdi Yu, Shaoli Huang, Toby P. Breckon, Jue Wang 0001 |
CVPR | 5 |
| 2023 | Skinned Motion Retargeting with Residual Perception of Motion Semantics & GeometryabstractA good motion retargeting cannot be reached without reasonable consideration of source-target differences on both the skeleton and shape geometry levels. In this work, we propose a novel Residual RETargeting network (R2ET) structure, which relies on two neural modification modules, to adjust the source motions to fit the target skeletons and shapes progressively. In particular, a skeleton-aware module is introduced to preserve the source motion semantics. A shape-aware module is designed to perceive the geometries of target characters to reduce interpenetration and contact-missing. Driven by our explored distance-based losses that explicitly model the motion semantics and geometry, these two modules can learn residual motion modifications on the source motion to generate plausible retargeted motion in a single inference without postprocessing. To balance these two modifications, we further present a balancing gate to conduct linear interpolation between them. Extensive experiments on the public dataset Mixamo demonstrate that our R2ET achieves the state-of-the-art performance, and provides a good balance between the preservation of motion semantics as well as the attenuation of interpenetration and contact-missing. Code is available at https://github.com/Kebii/R2ET. Junwu Weng, Fang Zhao 0006, Shaoli Huang, Xuefei Zhe, Linchao Bao, Ying Shan, Jue Wang 0001, Zhigang Tu 0001 |
CVPR | 9 |
| 2023 | Learning Anchor Transformations for 3D Garment AnimationabstractThis paper proposes an anchor-based deformation model, namely AnchorDEF, to predict 3D garment animation from a body motion sequence. It deforms a garment mesh template by a mixture of rigid transformations with extra nonlinear displacements. A set of anchors around the mesh surface is introduced to guide the learning of rigid transformation matrices. Once the anchor transformations are found, per-vertex nonlinear displacements of the garment template can be regressed in a canonical space, which reduces the complexity of deformation space learning. By explicitly constraining the transformed anchors to satisfy the consistencies of position, normal and direction, the physical meaning of learned anchor transformations in space is guaranteed for better generalization. Furthermore, an adaptive anchor updating is proposed to optimize the anchor position by being aware of local mesh topology for learning representative anchor transformations. Qualitative and quantitative experiments on different types of garments demonstrate that AnchorDEF achieves the state-of-the-art performance on 3D garment deformation prediction in motion, especially for loose-fitting garments. Fang Zhao 0006, Zekun Li 0002, Shaoli Huang, Junwu Weng, Tianfei Zhou, Guosen Xie, Jue Wang 0001, Ying Shan |
CVPR | 7 |
| 2023 | Content-Aware Unsupervised Deep Homography Estimation and its ExtensionsabstractHomography estimation is a basic image alignment method in many applications. It is usually done by extracting and matching sparse feature points, which are error-prone in low-light and low-texture images. On the other hand, previous deep homography approaches use either synthetic images for supervised learning or aerial images for unsupervised learning, both ignoring the importance of handling depth disparities and moving objects in real-world applications. To overcome these problems, in this work, we propose an unsupervised deep homography method with a new architecture design. In the spirit of the RANSAC procedure in traditional methods, we specifically learn an outlier mask to only select reliable regions for homography estimation. We calculate loss with respect to our learned deep features instead of directly comparing image content as did previously. To achieve the unsupervised training, we also formulate a novel triplet loss customized for our network. We verify our method by conducting comprehensive comparisons on a new dataset that covers a wide range of scenes with varying degrees of difficulties for the task. Experimental results reveal that our method outperforms the state-of-the-art, including deep solutions and feature-based solutions. Shuaicheng Liu, Nianjin Ye, Chuan Wang 0001, Jirong Zhang, Lanpeng Jia, Kunming Luo, Jue Wang 0001, Jian Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Robust Physical-World Attacks on Face Recognition
Xin Zheng 0008, Yanbo Fan, Baoyuan Wu, Yong Zhang 0034, Jue Wang 0001, Shirui Pan |
Pattern Recognit. | 5 |
| 2023 | VDTR: Video Deblurring With TransformerabstractVideo deblurring is still an unsolved problem due to the challenging spatio-temporal modeling process. While existing convolutional neural network (CNN)-based methods show a limited capacity of effective spatial and temporal modeling for video deblurring. This paper presents VDTR, an effective Transformer-based model that makes the first attempt to adapt pure Transformer for video deblurring. VDTR exploits the superior long-range and relation modeling capabilities of Transformer for both spatial and temporal modeling. However, it is challenging to design an appropriate Transformer-based model for video deblurring due to the complicated non-uniform blurs, misalignment across multiple frames and the high computational costs for high-resolution spatial modeling. To address these problems, VDTR advocates performing attention within non-overlapping windows and exploiting the hierarchical structure for long-range dependencies modeling. For frame-level spatial modeling, we propose an encoder-decoder Transformer that utilizes multi-scale features for deblurring. For multi-frame temporal modeling, we adapt Transformer to fuse multiple spatial features efficiently. Compared with CNN-based methods, the proposed method achieves highly competitive results on both synthetic and real-world video deblurring benchmarks, including DVD, GOPRO, REDS and BSD. We hope such a Transformer-based architecture can serve as a powerful alternative baseline for video deblurring and other video restoration tasks. The source code will be available athttps://github.com/ljzycmd/VDTR. Mingdeng Cao, Yanbo Fan, Yong Zhang 0034, Jue Wang 0001, Yujiu Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Fast Adversarial Training With Adaptive Step SizeabstractWhile adversarial training and its variants have shown to be the most effective algorithms to defend against adversarial attacks, their extremely slow training process makes it hard to scale to large datasets like ImageNet. The key idea of recent works to accelerate adversarial training is to substitute multi-step attacks (e.g., PGD) with single-step attacks (e.g., FGSM). However, these single-step methods suffer from catastrophic overfitting, where the accuracy against PGD attack suddenly drops to nearly 0% during training, and the network totally loses its robustness. In this work, we study the phenomenon from the perspective of training instances. We show that catastrophic overfitting is instance-dependent, and fitting instances with larger input gradient norm is more likely to cause catastrophic overfitting. Based on our findings, we propose a simple but effective method, Adversarial Training with Adaptive Step size (ATAS). ATAS learns an instance-wise adaptive step size that is inversely proportional to its gradient norm. Our theoretical analysis shows that ATAS converges faster than the commonly adopted non-adaptive counterparts. Empirically, ATAS consistently mitigates catastrophic overfitting and achieves higher robust accuracy on CIFAR10, CIFAR100, and ImageNet when evaluated on various adversarial budgets. Our code is released at https://github.com/HuangZhiChao95/ATAS. Zhichao Huang 0002, Yanbo Fan, Chen Liu 0027, Yong Zhang 0034, Mathieu Salzmann, Sabine Süsstrunk, Jue Wang 0001 |
IEEE Trans. Image Process. | 8 |
| 2023 | W-Net: Structure and Texture Interaction for Image InpaintingabstractRecent literature has developed two advanced tools for image inpainting: appearance propagation and attention matching. However, given the ineffective feature reorganization and vulnerable attention maps, existing works yield suboptimal results with distorted structures and inconsistent contents. Furthermore, we observe that deep sampling layers (DSL) and shallow skip connections (SSC) in U-Net separately promote image structure inference and texture synthesis. To address the above two issues, we devise a W-shaped network (W-Net), which consists of two key components: a texture spatial attention (TSA) module in SSC and a structure channel excitation (SCE) module in DSL. W-Net is a two-stage network, with coarse and refined structures derived at each stage. Meanwhile, the TSA module fills incomplete textures with reliable attention scores under the guidance of coarse structures, which effectively diminishes inconsistency from appearance to semantics. The SCE module rectifies structures according to the difference between coarse structures and refined structures enhanced by texture features. Then the module motivates them to produce more reasonable shapes. Complete textures and refined structures constitute desired inpainted images, as the output of W-Net. Experiments on multiple datasets demonstrate the superior performance of W-Net. Ruisong Zhang, Weize Quan, Yong Zhang 0034, Jue Wang 0001, Dong-Ming Yan 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Hallucinated Neural Radiance Fields in the WildabstractNeural Radiance Fields (NeRF) has recently gained popularity for its impressive novel view synthesis ability. This paper studies the problem of hallucinated NeRF: i.e., recovering a realistic NeRF at a different time of day from a group of tourism images. Existing solutions adopt NeRF with a controllable appearance embedding to render novel views under various conditions, but they cannot render view-consistent images with an unseen appearance. To solve this problem, we present an end-to-end framework for constructing a hallucinated NeRF, dubbed as Ha-NeRF. Specifically, we propose an appearance hallucination module to handle time-varying appearances and transfer them to novel views. Considering the complex occlusions of tourism images, we introduce an anti-occlusion module to decompose the static subjects for visibility accurately. Experimental results on synthetic data and real tourism photo collections demonstrate that our method can hallucinate the desired appearances and render occlusion-free images from different views. The project and supplementary materials are available at https://rover-xingyu.github.io/Ha-NeRF/. Qi Zhang 0029, Xiaoyu Li 0002, Xuan Wang 0009, Jue Wang 0001 |
CVPR | 7 |
| 2022 | Self-supervised Learning of Adversarial Example: Towards Good Generalizations for Deepfake DetectionabstractRecent studies in deepfake detection have yielded promising results when the training and testing face forgeries are from the same dataset. However, the problem remains challenging when one tries to generalize the detector to forgeries created by unseen methods in the training dataset. This work addresses the generalizable deepfake detection from a simple principle: a generalizable representation should be sensitive to diverse types of forgeries. Following this principle, we propose to enrich the “diversity” of forgeries by synthesizing augmented forgeries with a pool of forgery configurations and strengthen the “sensitivity” to the forgeries by enforcing the model to predict the forgery configurations. To effectively explore the large forgery augmentation space, we further propose to use the adversarial training strategy to dynamically synthesize the most challenging forgeries to the current model. Through extensive experiments, we show that the proposed strategies are surprisingly effective (see Figure 1), and they could achieve superior performance than the current state-of-the-art methods. Code is available at https://github.com/liangchen527/SLADD. Liang Chen 0030, Yong Zhang 0034, Yibing Song, Lingqiao Liu, Jue Wang 0001 |
CVPR | 5 |
| 2022 | Motion-aware Contrastive Video Representation Learning via Foreground-background MergingabstractIn light of the success of contrastive learning in the image domain, current self-supervised video representation learning methods usually employ contrastive loss to facilitate video representation learning. When naively pulling two augmented views of a video closer, the model however tends to learn the common static background as a shortcut but fails to capture the motion information, a phenomenon dubbed as background bias. Such bias makes the model suffer from weak generalization ability, leading to worse performance on downstream tasks such as action recognition. To alleviate such bias, we propose Foreground-background Merging (FAME) to deliberately compose the moving foreground region of the selected video onto the static background of others. Specifically, without any off-the-shelf detector, we extract the moving fore-ground out of background regions via the frame difference and color statistics, and shuffle the background regions among the videos. By leveraging the semantic consistency between the original clips and the fused ones, the model focuses more on the motion patterns and is debiased from the background shortcut. Extensive experiments demonstrate that FAME can effectively resist background cheating and thus achieve the state-of-the-art performance on downstream tasks across UCF101, HMDB51, and Diving48 datasets. The code and configurations are released at https://github.com/Mark12Ding/FAME. Shuangrui Ding, Maomao Li, Tianyu Yang 0003, Rui Qian 0001, Haohang Xu, Qingyi Chen, Jue Wang 0001, Hongkai Xiong |
CVPR | 7 |
| 2022 | LAS-AT: Adversarial Training with Learnable Attack StrategyabstractAdversarial training (AT) is always formulated as a minimax problem, of which the performance depends on the inner optimization that involves the generation of adver-sarial examples (AEs). Most previous methods adopt Projected Gradient Decent (PGD) with manually specifying attack parameters for AE generation. A combination of the attack parameters can be referred to as an attack strategy. Several works have revealed that using a fixed attack strategy to generate AEs during the whole training phase limits the model robustness and propose to exploit different attack strategies at different training stages to improve robustness. But those multi-stage handcrafted attack strategies need much domain expertise, and the robustness improvement is limited. In this paper, we propose a novel framework for adversarial training by introducing the concept of “learnable attack strategy”, dubbed LAS-AT, which learns to automatically produce attack strategies to improve the model robustness. Our framework is composed of a target network that uses AEs for training to improve robustness, and a strategy network that produces attack strategies to control the AE generation. Experimental evaluations on three benchmark databases demonstrate the superiority of the proposed method. The code is released at https://github.com/jiaxiaojunQAQ/LAS-AT. Xiaojun Jia, Yong Zhang 0034, Baoyuan Wu, Ke Ma 0001, Jue Wang 0001, Xiaochun Cao |
CVPR | 5 |
| 2022 | Exploring Denoised Cross-video Contrast for Weakly-supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization aims to localize actions in untrimmed videos with only video-level labels. Most existing methods address this problem with a “localization-by-classification” pipeline that localizes action regions based on snippet-wise classification sequences. Snippet-wise classifications are unfortunately error prone due to the sparsity of video-level labels. Inspired by recent success in unsupervised contrastive representation learning, we propose a novel denoised cross-video contrastive algorithm, aiming to enhance the feature discrimination ability of video snippets for accurate temporal action localization in the weakly-supervised setting. This is enabled by three key designs: 1) an effective pseudo-label denoising module to alleviate the side effects caused by noisy contrastive features, 2) an efficient region-level feature contrast strategy with a region-level memory bank to capture “global” contrast across the entire dataset, and 3) a diverse contrastive learning strategy to enable action-background separation as well as intra-class compactness & inter-class separability. Extensive experiments on THUMOS14 and ActivityNet v1.3 demonstrate the superior performance of our approach. Tianyu Yang 0003, Wei Ji 0011, Jue Wang 0001, Li Cheng 0001 |
CVPR | 4 |
| 2022 | Deblur-NeRF: Neural Radiance Fields from Blurry ImagesabstractNeural Radiance Field (NeRF) has gained considerable attention recently for 3D scene reconstruction and novel view synthesis due to its remarkable synthesis quality. However, image blurriness caused by defocus or motion, which often occurs when capturing scenes in the wild, significantly degrades its reconstruction quality. To address this problem, We propose Deblur-NeRF, the first method that can recover a sharp NeRF from blurry input. We adopt an analysis-by-synthesis approach that reconstructs blurry views by simulating the blurring process, thus making NeRF robust to blurry inputs. The core of this simulation is a novel Deformable Sparse Kernel (DSK) module that models spatially-varying blur kernels by deforming a canonical sparse kernel at each spatial location. The ray origin of each kernel point is Jointly optimized, inspired by the physical blurring process. This module is parameterized as an MLP that has the ability to be generalized to various blur types. Jointly optimizing the NeRF and the DSK module allows us to restore a sharp NeRF. We demonstrate that our method can be used on both camera motion blur and defocus blur: the two most common types of blur in real scenes. Evaluation results on both synthetic and real-world data show that our method outperforms several baselines. The synthetic and real datasets along with the source code is publicly available at https://limacv.github.io/deblurNeRF/. Xiaoyu Li 0002, Jing Liao 0001, Qi Zhang 0029, Xuan Wang 0009, Jue Wang 0001, Pedro V. Sander |
CVPR | 6 |
| 2022 | FENeRF: Face Editing in Neural Radiance FieldsabstractPrevious portrait image generation methods roughly fall into two categories: 2D GANs and 3D-aware GANs. 2D GANs can generate high fidelity portraits but with low view consistency. 3D-aware GAN methods can maintain view consistency but their generated images are not locally editable. To overcome these limitations, we propose FENeRF, a 3D-aware generator that can produce view-consistent and locally-editable portrait images. Our method uses two decoupled latent codes to generate corresponding facial semantics and texture in a spatial-aligned 3D volume with shared geometry. Benefiting from such underlying 3D representation, FENeRF can Jointly render the boundary-aligned image and semantic mask and use the semantic mask to edit the 3D volume via GAN inversion. We further show such 3D representation can be learned from widely available monocular image and semantic mask pairs. Moreover, we reveal that Joint learning semantics and texture helps to generate finer geometry. Our experiments demonstrate that FENeRF outperforms state-of-the-art methods in various face editing tasks. Code is available at https://github.com/MrTornado24/FENeRF. Jingxiang Sun, Xuan Wang 0009, Yong Zhang 0034, Xiaoyu Li 0002, Qi Zhang 0029, Yebin Liu, Jue Wang 0001 |
CVPR | 7 |
| 2022 | Long-Short Temporal Contrastive Learning of Video TransformersabstractVideo transformers have recently emerged as a competitive alternative to 3D CNNs for video understanding. However, due to their large number of parameters and reduced inductive biases, these models require supervised pretraining on large-scale image datasets to achieve top performance. In this paper, we empirically demonstrate that self-supervised pretraining of video transformers on video-only datasets can lead to action recognition results that are on par or better than those obtained with supervised pretraining on large-scale image datasets, even massive ones such as ImageNet-21K. Since transformer-based models are effective at capturing dependencies over extended temporal spans, we propose a simple learning procedure that forces the model to match a long-term view to a short-term view of the same video. Our approach, named Long-Short Temporal Contrastive Learning (LSTCL), enables video transformers to learn an effective clip-level representation by predicting temporal context captured from a longer temporal extent. To demonstrate the generality of our findings, we implement and validate our approach under three different self-supervised contrastive learning frameworks (MoCo v3, BYOL, SimSiam) using two distinct video-transformer architectures, including an improved variant of the Swin Transformer augmented with space-time attention. We conduct a thorough ablation study and show that LSTCL achieves competitive performance on multiple video benchmarks and represents a convincing alternative to supervised image-based pretraining. Jue Wang 0001, Gedas Bertasius, Du Tran, Lorenzo Torresani |
CVPR | 1 |
| 2022 | Deformable Video TransformerabstractVideo transformers have recently emerged as an effective alternative to convolutional networks for action classification. However, most prior video transformers adopt either global space-time attention or hand-defined strategies to compare patches within and across frames. These fixed attention schemes not only have high computational cost but, by comparing patches at predetermined locations, they neglect the motion dynamics in the video. In this paper, we introduce the Deformable Video Transformer (DVT), which dynamically predicts a small subset of video patches to attend for each query location based on motion information, thus allowing the model to decide where to look in the video based on correspondences across frames. Crucially, these motion-based correspondences are obtained at zero-cost from information stored in the compressed format of the video. Our deformable attention mechanism is optimized directly with respect to classification performance, thus eliminating the need for suboptimal hand-design of attention strategies. Experiments on four large-scale video benchmarks (Kinetics-400, Something-Something-V2, EPIC-KITCHENS and Diving-48) demonstrate that, compared to existing video transformers, our model achieves higher accuracy at the same or lower computational cost, and it attains state-of-the-art results on these four datasets. Jue Wang 0001, Lorenzo Torresani |
CVPR | 1 |
| 2022 | High-Fidelity GAN Inversion for Image Attribute EditingabstractWe present a novel highfidelity generative adversarial network (GAN) inversion framework that enables attribute editing with image-specific details well-preserved (e.g., background, appearance, and illumination). We first analyze the challenges of highfidelity GAN inversion from the perspective of lossy data compression. With a low bitrate latent code, previous works have difficulties in preserving highfidelity details in reconstructed and edited images. Increasing the size of a latent code can improve the accuracy of GAN inversion but at the cost of inferior editability. To improve image fidelity without compromising editability, we propose a distortion consultation approach that employs a distortion map as a reference for highfidelity reconstruction. In the distortion consultation inversion (DCI), the distortion map is first projected to a high-rate latent map, which then complements the basic low-rate latent code with more details via consultation fusion. To achieve high-fidelity editing, we propose an adaptive distortion alignment (ADA) module with a self-supervised training scheme, which bridges the gap between the edited and inversion images. Extensive experiments in the face and car domains show a clear improvement in both inversion and editing quality. The project page is https://tengfei-wang.github.io/HFGI/. Tengfei Wang 0002, Yong Zhang 0034, Yanbo Fan, Jue Wang 0001, Qifeng Chen 0001 |
CVPR | 4 |
| 2022 | Multi-Robot Active Mapping via Neural Bipartite Graph MatchingabstractWe study the problem of multi-robot active mapping, which aims for complete scene map construction in minimum time steps. The key to this problem lies in the goal position estimation to enable more efficient robot movements. Previous approaches either choose the frontier as the goal position via a myopic solution that hinders the time efficiency, or maximize the long-term value via reinforcement learning to directly regress the goal position, but does not guarantee the complete map construction. In this paper, we propose a novel algorithm, namely NeuralCoMapping, which takes advantage of both approaches. We reduce the problem to bipartite graph matching, which establishes the node correspondences between two graphs, denoting robots and frontiers. We introduce a multiplex graph neural network (mGNN) that learns the neural distance to fill the affinity matrix for more effective graph matching. We optimize the mGNN with a differentiable linear assignment layer by maximizing the long-term values that favor time efficiency and map completeness via reinforcement learning. We compare our algorithm with several state-of-the-art multi-robot active mapping approaches and adapted reinforcement-learning baselines. Experimental results demonstrate the superior performance and exceptional generalization ability of our algorithm on various indoor scenes and unseen number of robots, when only trained with 9 indoor scenes. Kai Ye 0007, Siyan Dong, Qingnan Fan, He Wang 0010, Li Yi 0001, Fei Xia 0002, Jue Wang 0001, Baoquan Chen |
CVPR | 7 |
| 2022 | Unsupervised Pre-training for Temporal Action Localization TasksabstractUnsupervised video representation learning has made remarkable achievements in recent years. However, most existing methods are designed and optimized for video classification. These pretrained models can be sub-optimal for temporal localization tasks due to the inherent discrepancy between video-level classification and clip-level localization. To bridge this gap, we make the first attempt to propose a self-supervised pretext task, coined as Pseudo Action Localization (PAL) to Unsupervisedly Pre-train feature encoders for Temporal Action Localization tasks (UP-TAL). Specifically, we first randomly select temporal regions, each of which contains multiple clips, from one video as pseudo actions and then paste them onto different temporal positions of the other two videos. The pretext task is to align the features of pasted pseudo action regions from two synthetic videos and maximize the agreement between them. Compared to the existing unsupervised video representation learning approaches, our PAL adapts better to downstream TAL tasks by introducing a temporal equivariant contrastive learning paradigm in a temporally dense and scale-aware manner. Extensive experiments show that PAL can utilize large-scale unlabeled video data to significantly boost the performance of existing TAL methods. Our codes and models will be made publicly available at https://github.com/zhang-can/UP-TAL. Can Zhang 0001, Tianyu Yang 0003, Junwu Weng, Meng Cao 0002, Jue Wang 0001, Yuexian Zou |
CVPR | 5 |
| 2022 | LocVTP: Video-Text Pre-training for Temporal Localization
Meng Cao 0002, Tianyu Yang 0003, Junwu Weng, Can Zhang 0001, Jue Wang 0001, Yuexian Zou |
ECCV (26) | 5 |
| 2022 | Towards Accurate Active Camera Localization
Qihang Fang, Yingda Yin, Qingnan Fan, Fei Xia 0002, Siyan Dong, Jue Wang 0001, Leonidas J. Guibas, Baoquan Chen |
ECCV (10) | 7 |
| 2022 | Prior-Guided Adversarial Initialization for Fast Adversarial Training
Xiaojun Jia, Yong Zhang 0034, Xingxing Wei 0001, Baoyuan Wu, Ke Ma 0001, Jue Wang 0001, Xiaochun Cao |
ECCV (4) | 6 |
| 2022 | Spatial-Separated Curve Rendering Network for Efficient and High-Resolution Image Harmonization
Jingtang Liang, Xiaodong Cun, Chi-Man Pun, Jue Wang 0001 |
ECCV (7) | 4 |
| 2022 | StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGAN
Yong Zhang 0034, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang 0009, Qingyan Bai, Baoyuan Wu, Jue Wang 0001, Yujiu Yang 0001 |
ECCV (17) | 9 |
| 2022 | EViT: Expediting Vision Transformers via Token Reorganizations
Youwei Liang, Chongjian Ge, Zhan Tong, Yibing Song, Jue Wang 0001, Pengtao Xie |
ICLR | 5 |
| 2022 | HyP2 Loss: Beyond Hypersphere Metric Space for Multi-label Image RetrievalabstractImage retrieval has become an increasingly appealing technique with broad multimedia application prospects, where deep hashing serves as the dominant branch towards low storage and efficient retrieval. In this paper, we carried out in-depth investigations on metric learning in deep hashing for establishing a powerful metric space in multi-label scenarios, where the pair loss suffers high computational overhead and converge difficulty, while the proxy loss is theoretically incapable of expressing the profound label dependencies and exhibits conflicts in the constructed hypersphere space. To address the problems, we propose a novel metric learning framework with Hybrid Proxy-Pair Loss (HyP$^2$ Loss) that constructs an expressive metric space with efficient training complexity w.r.t. the whole dataset. The proposed HyP$^2$ Loss focuses on optimizing the hypersphere space by learnable proxies and excavating data-to-data correlations of irrelevant pairs, which integrates sufficient data correspondence of pair-based methods and high-efficiency of proxy-based methods. Extensive experiments on four standard multi-label benchmarks justify the proposed method outperforms the state-of-the-art, is robust among different hash bits and achieves significant performance gains with a faster, more stable convergence speed. Our code is available at https://github.com/JerryXu0129/HyP2-Loss. Chengyin Xu, Zenghao Chai, Zhengzhuo Xu, Chun Yuan 0003, Yanbo Fan, Jue Wang 0001 |
ACM Multimedia | 6 |
| 2022 | AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionabstractPretraining Vision Transformers (ViTs) has achieved great success in visual recognition. A following scenario is to adapt a ViT to various image and video recognition tasks. The adaptation is challenging because of heavy computation and memory storage. Each model needs an independent and complete finetuning process to adapt to different tasks, which limits its transferability to different visual domains.To address this challenge, we propose an effective adaptation approach for Transformer, namely AdaptFormer, which can adapt the pre-trained ViTs into many different image and video tasks efficiently.It possesses several benefits more appealing than prior arts.Firstly, AdaptFormer introduces lightweight modules that only add less than 2% extra parameters to a ViT, while it is able to increase the ViT's transferability without updating its original pre-trained parameters, significantly outperforming the existing 100\% fully fine-tuned models on action recognition benchmarks.Secondly, it can be plug-and-play in different Transformers and scalable to many visual tasks.Thirdly, extensive experiments on five image and video datasets show that AdaptFormer largely improves ViTs in the target domains. For example, when updating just 1.5% extra parameters, it achieves about 10% and 19% relative improvement compared to the fully fine-tuned models on Something-Something~v2 and HMDB51, respectively. Code is available at https://github.com/ShoufaChen/AdaptFormer. Shoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang, Yibing Song, Jue Wang 0001, Ping Luo 0002 |
NeurIPS | 6 |
| 2022 | OST: Improving Generalization of DeepFake Detection via One-Shot Test-Time TrainingabstractState-of-the-art deepfake detectors perform well in identifying forgeries when they are evaluated on a test set similar to the training set, but struggle to maintain good performance when the test forgeries exhibit different characteristics from the training images e.g., forgeries are created by unseen deepfake methods. Such a weak generalization capability hinders the applicability of deepfake detectors. In this paper, we introduce a new learning paradigm specially designed for the generalizable deepfake detection task. Our key idea is to construct a test-sample-specific auxiliary task to update the model before applying it to the sample. Specifically, we synthesize pseudo-training samples from each test image and create a test-time training objective to update the model. Moreover, we proposed to leverage meta-learning to ensure that a fast single-step test-time gradient descent, dubbed one-shot test-time training (OST), can be sufficient for good deepfake detection performance. Extensive results across several benchmark datasets demonstrate that our approach performs favorably against existing arts in terms of generalization to unseen data and robustness to different post-processing steps. Liang Chen 0030, Yong Zhang 0034, Yibing Song, Jue Wang 0001, Lingqiao Liu |
NeurIPS | 4 |
| 2022 | Boosting the Transferability of Adversarial Attacks with Reverse Adversarial PerturbationabstractDeep neural networks (DNNs) have been shown to be vulnerable to adversarial examples, which can produce erroneous predictions by injecting imperceptible perturbations. In this work, we study the transferability of adversarial examples, which is significant due to its threat to real-world applications where model architecture or parameters are usually unknown. Many existing works reveal that the adversarial examples are likely to overfit the surrogate model that they are generated from, limiting its transfer attack performance against different target models. To mitigate the overfitting of the surrogate model, we propose a novel attack method, dubbed reverse adversarial perturbation (RAP). Specifically, instead of minimizing the loss of a single adversarial point, we advocate seeking adversarial example located at a region with unified low loss value, by injecting the worst-case perturbation (the reverse adversarial perturbation) for each step of the optimization procedure. The adversarial attack with RAP is formulated as a min-max bi-level optimization problem. By integrating RAP into the iterative process for attacks, our method can find more stable adversarial examples which are less sensitive to the changes of decision boundary, mitigating the overfitting of the surrogate model. Comprehensive experimental comparisons demonstrate that RAP can significantly boost adversarial transferability. Furthermore, RAP can be naturally combined with many existing black-box attack techniques, to further boost the transferability. When attacking a real-world image recognition system, Google Cloud Vision API, we obtain 22% performance improvement of targeted attacks over the compared method. Our codes are available at https://github.com/SCLBD/TransferattackRAP. Zeyu Qin, Yanbo Fan, Li Shen 0008, Yong Zhang 0034, Jue Wang 0001, Baoyuan Wu |
NeurIPS | 6 |
| 2022 | VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingabstractPre-training video transformers on extra large-scale datasets is generally required to achieve premier performance on relatively small datasets. In this paper, we show that video masked autoencoders (VideoMAE) are data-efficient learners for self-supervised video pre-training (SSVP). We are inspired by the recent ImageMAE and propose customized video tube masking with an extremely high ratio. This simple design makes video reconstruction a more challenging and meaningful self-supervision task, thus encouraging extracting more effective video representations during the pre-training process. We obtain three important findings with VideoMAE: (1) An extremely high proportion of masking ratio (i.e., 90% to 95%) still yields favorable performance for VideoMAE. The temporally redundant video content enables higher masking ratio than that of images. (2) VideoMAE achieves impressive results on very small datasets (i.e., around 3k-4k videos) without using any extra data. This is partially ascribed to the challenging task of video reconstruction to enforce high-level structure learning. (3) VideoMAE shows that data quality is more important than data quantity for SSVP. Domain shift between pre-training and target datasets is an important factor. Notably, our VideoMAE with the vanilla ViT backbone can achieve 87.4% on Kinects-400, 75.4% on Something-Something V2, 91.3% on UCF101, and 62.6% on HMDB51, without using any extra data. Code is available at https://github.com/MCG-NJU/VideoMAE. Zhan Tong, Yibing Song, Jue Wang 0001, Limin Wang 0002 |
NeurIPS | 3 |
| 2022 | Stability Analysis and Generalization Bounds of Adversarial TrainingabstractIn adversarial machine learning, deep neural networks can fit the adversarial examples on the training dataset but have poor generalization ability on the test set. This phenomenon is called robust overfitting, and it can be observed when adversarially training neural nets on common datasets, including SVHN, CIFAR-10, CIFAR-100, and ImageNet. In this paper, we study the robust overfitting issue of adversarial training by using tools from uniform stability. One major challenge is that the outer function (as a maximization of the inner function) is nonsmooth, so the standard technique (e.g., Hardt et al., 2016) cannot be applied. Our approach is to consider $\eta$-approximate smoothness: we show that the outer function satisfies this modified smoothness assumption with $\eta$ being a constant related to the adversarial perturbation $\epsilon$. Based on this, we derive stability-based generalization bounds for stochastic gradient descent (SGD) on the general class of $\eta$-approximate smooth functions, which covers the adversarial loss. Our results suggest that robust test accuracy decreases in $\epsilon$ when $T$ is large, with a speed between $\Omega(\epsilon\sqrt{T})$ and $\mathcal{O}(\epsilon T)$. This phenomenon is also observed in practice. Additionally, we show that a few popular techniques for adversarial training (\emph{e.g.,} early stopping, cyclic learning rate, and stochastic weight averaging) are stability-promoting in theory. Jiancong Xiao, Yanbo Fan, Ruoyu Sun 0001, Jue Wang 0001, Zhi-Quan Luo |
NeurIPS | 4 |
| 2022 | One Model to Edit Them All: Free-Form Text-Driven Image Manipulation with Semantic ModulationsabstractFree-form text prompts allow users to describe their intentions during image manipulation conveniently. Based on the visual latent space of StyleGAN[21] and text embedding space of CLIP[34], studies focus on how to map these two latent spaces for text-driven attribute manipulations. Currently, the latent mapping between these two spaces is empirically designed and confines that each manipulation model can only handle one fixed text prompt. In this paper, we propose a method named Free-Form CLIP (FFCLIP), aiming to establish an automatic latent mapping so that one manipulation model handles free-form text prompts. Our FFCLIP has a cross-modality semantic modulation module containing semantic alignment and injection. The semantic alignment performs the automatic latent mapping via linear transformations with a cross attention mechanism. After alignment, we inject semantics from text prompt embeddings to the StyleGAN latent space. For one type of image (e.g., human portrait'), one FFCLIP model can be learned to handle free-form text prompts. Meanwhile, we observe that although each training text prompt only contains a single semantic meaning, FFCLIP can leverage text prompts with multiple semantic meanings for image manipulation. In the experiments, we evaluate FFCLIP on three types of images (i.e.,human portraits', cars', andchurches'). Both visual and numerical results show that FFCLIP effectively produces semantically accurate and visually realistic images. Project page: https://github.com/KumapowerLIU/FFCLIP. Yibing Song, Ziyang Yuan, Xintong Han, Chun Yuan 0003, Qifeng Chen 0001, Jue Wang 0001 |
NeurIPS | 8 |
| 2022 | VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the WildabstractWe present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disentangles this objective into three sequential tasks: (1) face video generation with a canonical expression; (2) audio-driven lip-sync; and (3) face enhancement for improving photo-realism. Given a talking-head video, we first modify the expression of each frame according to the same expression template using the expression editing network, resulting in a video with the canonical expression. This video, together with the given audio, is then fed into the lip-sync network to generate a lip-syncing video. Finally, we improve the photo-realism of the synthesized faces through an identity-aware face enhancement network and post-processing. We use learning-based approaches for all three steps and all our modules can be tackled in a sequential pipeline without any user intervention. Furthermore, our system is a generic approach that does not need to be retrained to a specific person. Evaluations on two widely-used datasets and in-the-wild examples demonstrate the superiority of our framework over other state-of-the-art methods in terms of lip-sync accuracy and visual quality. Xiaodong Cun, Yong Zhang 0034, Menghan Xia, Mingrui Zhu, Xuan Wang 0009, Jue Wang 0001, Nannan Wang 0001 |
SIGGRAPH Asia | 8 |
| 2022 | UPHDR-GAN: Generative Adversarial Network for High Dynamic Range Imaging With Unpaired DataabstractThe paper proposes a method to effectively fuse multi-exposure inputs and generate high-quality high dynamic range (HDR) images with unpaired datasets. Deep learning-based HDR image generation methods rely heavily on paired datasets. The ground truth images play a leading role in generating reasonable HDR images. Datasets without ground truth are hard to be applied to train deep neural networks. Recently, Generative Adversarial Networks (GAN) have demonstrated their potentials of translating images from source domain$X$to target domain$Y$in the absence of paired examples. In this paper, we propose a GAN-based network for solving such problems while generating enjoyable HDR results, named UPHDR-GAN. The proposed method relaxes the constraint of the paired dataset and learns the mapping from the LDR domain to the HDR domain. Although the pair data are missing, UPHDR-GAN can properly handle the ghosting artifacts caused by moving objects or misalignments with the help of the modified GAN loss, the improved discriminator network and the useful initialization phase. The proposed method preserves the details of important regions and improves the total image perceptual quality. Qualitative and quantitative comparisons against the representative methods demonstrate the superiority of the proposed UPHDR-GAN. Ru Li 0002, Chuan Wang 0001, Jue Wang 0001, Guanghui Liu 0001, Heng-Yu Zhang, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | DeepOIS: Gyroscope-Guided Deep Optical Image Stabilizer CompensationabstractMobile captured images can be aligned using their gyroscope sensors. Optical image stabilizer (OIS) terminates this possibility by adjusting the images during the capturing. In this work, we propose a deep network that compensates for the motions caused by the OIS, such that the gyroscopes can be used for image alignment on the OIS cameras. To achieve this, we first record both videos and gyroscope readings with an OIS camera as training data. Then, we convert gyroscope readings into motion fields. Second, we propose an Essential Mixtures motion model for rolling shutter cameras, where an array of rotations within a frame are extracted as the ground-truth guidance. Third, we train a convolutional neural network with gyroscope motions as input to compensate for the OIS motion. Once finished, the compensation network can be applied for other scenes, where the image alignment is purely based on gyroscopes with no need for images contents, delivering strong robustness. Experiments show that our results are comparable with that of non-OIS cameras, and outperform image-based alignment results with a relatively large margin. Code and dataset is available at:https://github.com/lhaippp/DeepOIS. Shuaicheng Liu, Haipeng Li 0001, Zhengning Wang, Jue Wang 0001, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Boosting Fast Adversarial Training With Learnable Adversarial InitializationabstractAdversarial training (AT) has been demonstrated to be effective in improving model robustness by leveraging adversarial examples for training. However, most AT methods are in face of expensive time and computational cost for calculating gradients at multiple steps in generating adversarial examples. To boost training efficiency, fast gradient sign method (FGSM) is adopted in fast AT methods by calculating gradient only once. Unfortunately, the robustness is far from satisfactory. One reason may arise from the initialization fashion. Existing fast AT generally uses a random sample-agnostic initialization, which facilitates the efficiency yet hinders a further robustness improvement. Up to now, the initialization in fast AT is still not extensively explored. In this paper, focusing on image classification, we boost fast AT with a sample-dependent adversarial initialization, i.e., an output from a generative network conditioned on a benign image and its gradient information from the target network. As the generative network and the target network are optimized jointly in the training phase, the former can adaptively generate an effective initialization with respect to the latter, which motivates gradually improved robustness. Experimental evaluations on four benchmark databases demonstrate the superiority of our proposed method over state-of-the-art fast AT methods, as well as comparable robustness to advanced multi-step AT methods. The code is released at https://github.com//jiaxiaojunQAQ//FGSM-SDI. Xiaojun Jia, Yong Zhang 0034, Baoyuan Wu, Jue Wang 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 4 |
| 2022 | Image Inpainting With Local and Global RefinementabstractImage inpainting has made remarkable progress with recent advances in deep learning. Popular networks mainly follow an encoder-decoder architecture (sometimes with skip connections) and possess sufficiently large receptive field, i.e., larger than the image resolution. The receptive field refers to the set of input pixels that are path-connected to a neuron. For image inpainting task, however, the size of surrounding areas needed to repair different kinds of missing regions are different, and the very large receptive field is not always optimal, especially for the local structures and textures. In addition, a large receptive field tends to involve more undesired completion results, which will disturb the inpainting process. Based on these insights, we rethink the process of image inpainting from a different perspective of receptive field, and propose a novel three-stage inpainting framework with local and global refinement. Specifically, we first utilize an encoder-decoder network with skip connection to achieve coarse initial results. Then, we introduce a shallow deep model with small receptive field to conduct the local refinement, which can also weaken the influence of distant undesired completion results. Finally, we propose an attention-based encoder-decoder network with large receptive field to conduct the global refinement. Experimental results demonstrate that our method outperforms the state of the arts on three popular publicly available datasets for image inpainting. Our local and global refinement network can be directly inserted into the end of any existing networks to further improve their inpainting performance. Code is available at https://github.com/weizequan/LGNet.git. Weize Quan, Ruisong Zhang, Yong Zhang 0034, Zhifeng Li 0001, Jue Wang 0001, Dong-Ming Yan 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | High-Fidelity 3D Digital Human Head Creation from RGB-D SelfiesabstractWe present a fully automatic system that can produce high-fidelity, photo-realistic three-dimensional (3D) digital human heads with a consumer RGB-D selfie camera. The system only needs the user to take a short selfie RGB-D video while rotating his/her head and can produce a high-quality head reconstruction in less than 30 s. Our main contribution is a new facial geometry modeling and reflectance synthesis procedure that significantly improves the state of the art. Specifically, given the input video a two-stage frame selection procedure is first employed to select a few high-quality frames for reconstruction. Then a differentiable renderer-based 3D Morphable Model (3DMM) fitting algorithm is applied to recover facial geometries from multiview RGB-D data, which takes advantages of a powerful 3DMM basis constructed with extensive data generation and perturbation. Our 3DMM has much larger expressive capacities than conventional 3DMM, allowing us to recover more accurate facial geometry using merely linear basis. For reflectance synthesis, we present a hybrid approach that combines parametric fitting andConvolutional Neural Networks (CNNs)to synthesize high-resolution albedo/normal maps with realistic hair/pore/wrinkle details. Results show that our system can produce faithful 3D digital human faces with extremely realistic details. The main code and the newly constructed 3DMM basis is publicly available. Linchao Bao, Xiangkai Lin, Haoxian Zhang, Xuefei Zhe, Hao-Zhi Huang 0001, Xinwei Jiang, Jue Wang 0001, Dong Yu 0001, Zhengyou Zhang |
ACM Trans. Graph. | 10 |
| 2022 | Neural Parameterization for Dynamic Human Head EditingabstractImplicit radiance functions emerged as a powerful scene representation for reconstructing and rendering photo-realistic views of a 3D scene. These representations, however, suffer from poor editability. On the other hand, explicit representations such as polygonal meshes allow easy editing but are not as suitable for reconstructing accurate details in dynamic human heads, such as fine facial features, hair, teeth, and eyes. In this work, we present Neural Parameterization (NeP), a hybrid representation that provides the advantages of both implicit and explicit methods. NeP is capable of photo-realistic rendering while allowing fine-grained editing of the scene geometry and appearance. We first disentangle the geometry and appearance by parameterizing the 3D geometry into 2D texture space. We enable geometric editability by introducing an explicit linear deformation blending layer. The deformation is controlled by a set of sparse key points, which can be explicitly and intuitively displaced to edit the geometry. For appearance, we develop a hybrid 2D texture consisting of an explicit texture map for easy editing and implicit view and time-dependent residuals to model temporal and view variations. We compare our method to several reconstruction and editing baselines. The results show that the NeP achieves almost the same level of rendering accuracy while maintaining high editability. Xiaoyu Li 0002, Jing Liao 0001, Xuan Wang 0009, Qi Zhang 0029, Jue Wang 0001, Pedro V. Sander |
ACM Trans. Graph. | 6 |
| 2022 | IDE-3D: Interactive Disentangled Editing for High-Resolution 3D-Aware Portrait SynthesisabstractExisting 3D-aware facial generation methods face a dilemma in quality versus editability: they either generate editable results in low resolution, or high-quality ones with no editing flexibility. In this work, we propose a new approach that brings the best of both worlds together. Our system consists of three major components: (1) a 3D-semantics-aware generative model that produces view-consistent, disentangled face images and semantic masks; (2) a hybrid GAN inversion approach that initializes the latent codes from the semantic and texture encoder, and further optimizes them for faithful reconstruction; and (3) a canonical editor that enables efficient manipulation of semantic masks in canonical view and produces high-quality editing results. Our approach is competent for many applications, e.g. free-view face drawing, editing and style control. Both quantitative and qualitative results show that our method reaches the state-of-the-art in terms of photorealism, faithfulness and efficiency. Jingxiang Sun, Xuan Wang 0009, Yichun Shi, Lizhen Wang 0002, Jue Wang 0001, Yebin Liu |
ACM Trans. Graph. | 5 |
| 2022 | Disentangled Image Colorization via Global AnchorsabstractColorization is multimodal by nature and challenges existing frameworks to achieve colorful and structurally consistent results. Even the sophisticated autoregressive model struggles to maintain long-distance color consistency due to the fragility of sequential dependence. To overcome this challenge, we propose a novel colorization framework that disentangles color multimodality and structure consistency through global color anchors, so that both aspects could be learned effectively. Our key insight is that several carefully located anchors could approximately represent the color distribution of an image, and conditioned on the anchor colors, we can predict the image color in a deterministic manner by utilizing internal correlation. To this end, we construct a colorization model with dual branches, where the color modeler predicts the color distribution for anchor color representation, and the color generator predicts the pixel colors by referring the sampled anchor colors. Importantly, the anchors are located under two principles: color independence and global coverage, which is realized with clustering analysis on the deep color features. To simplify the computation, we creatively adopt soft superpixel segmentation to reduce the image primitives, which still nicely reserves the reversibility to pixel-wise representation. Extensive experiments show that our method achieves notable superiority over various mainstream frameworks in perceptual quality. Thanks to anchor-based color representation, our model has the flexibility to support diverse and controllable colorization as well. Menghan Xia, Wenbo Hu 0005, Tien-Tsin Wong, Jue Wang 0001 |
ACM Trans. Graph. | 4 |
| 2022 | Video Vectorization via Bipartite Diffusion Curves Propagation and OptimizationabstractWe propose a new video vectorization approach for converting videos in the raster format to vector representation with the benefits of resolution independence and compact storage. Through classifying extracted curves in each video frame into salient ones and non-salient ones, we introduce a novel bipartite diffusion curves (BDCs) representation in order to preserve both important image features such as sharp boundaries and regions with smooth color variation. This bipartite representation allows us to propagate non-salient curves across frames such that the propagation, in conjunction with geometry optimization and color optimization of salient curves, ensures the preservation of fine details within each frame and across different frames, and meanwhile, achieves good spatial-temporal coherence. Thorough experiments on a variety of videos show that our method is capable of converting videos to the vector representation with low reconstruction errors, low computational cost, and fine details, demonstrating our superior performance over the state of the art. We also show that, when used for video upsampling, our method produces results comparable to video super-resolution. Yuanqi Li, Chuan Wang 0001, Jie Guo 0001, Jue Wang 0001, Yanwen Guo 0001, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Vx2Text: End-to-End Learning of Video-Based Text Generation From Multimodal InputsabstractWe present VX2TEXT, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, which have been shown to be effective at modeling language, each modality is first converted into a set of language embeddings by a learnable tokenizer. This allows our approach to perform multimodal fusion in the language space, thus eliminating the need for ad-hoc cross-modal fusion modules. To address the non-differentiability of tokenization on continuous inputs (e.g., video or audio), we utilize a relaxation scheme that enables end-to-end training. Furthermore, unlike prior encoder-only models, our network includes an autoregressive decoder to generate open-ended text from the multimodal embeddings fused by the language encoder. This renders our approach fully generative and makes it directly applicable to different "video+x to text" problems without the need to design specialized network heads for each task. The proposed framework is not only conceptually simple but also remarkably effective: experiments demonstrate that our approach based on a single architecture outperforms the state-of-the-art on three videobased text-generation tasks—captioning, question answering and audio-visual scene-aware dialog. Xudong Lin 0003, Gedas Bertasius, Jue Wang 0001, Shih-Fu Chang, Devi Parikh, Lorenzo Torresani |
CVPR | 3 |
| 2021 | UPFlow: Upsampling Pyramid for Unsupervised Optical Flow LearningabstractWe present an unsupervised learning approach for optical flow estimation by improving the upsampling and learning of pyramid network. We design a self-guided upsample module to tackle the interpolation blur problem caused by bilinear upsampling between pyramid levels. Moreover, we propose a pyramid distillation loss to add supervision for intermediate levels via distilling the finest flow as pseudo labels. By integrating these two components together, our method achieves the best performance for unsupervised optical flow learning on multiple leading benchmarks, including MPI-SIntel, KITTI 2012 and KITTI 2015. In particular, we achieve EPE=1.4 on KITTI 2012 and F1=9.38% on KITTI 2015, which outperform the previous state-of-the-art methods by 22.2% and 15.7%, respectively. Kunming Luo, Chuan Wang 0001, Shuaicheng Liu, Haoqiang Fan, Jue Wang 0001, Jian Sun 0001 |
CVPR | 5 |
| 2021 | Person Image Synthesis in Arbitrary 3D Poses Based on Part Affinity FieldsabstractWe consider the person image synthesis problem, where an output image is generated from an arbitrary source image and an arbitrary target 3D pose. Prior person image synthesizing methods usually use 2D keypoint heatmaps to represent the target pose. However, this 2D representation can be ambiguous in self-occlusion scenarios due to the lack of depth information, resulting in generating inappropriate images. To solve this problem, we propose to synthesize person image from 3D poses. We introduce an improved part affinity field representation to describe the 3D configuration of the target pose and the 2D location of the person in the pixel space. Compared to using 2D poses, our synthesized images have visually better details, such as correct self-occlusion, brightness, face direction et al. Moreover, in contrast with prior person image generators, our method predicts the difference between the source image and target image instead of the output images. This strategy allows the network to generate better foreground and significantly reduces the noise in the image background. We evaluate our method on images of fifteen different actions on Human3.6M dataset. Extensive experiments demonstrate that our method can synthesize much better person images than 2D pose-based ones. If given a sequence of desired poses, our method can produce a sequence of temporally smooth and coherent images, even on another subject and actions, which means that our method has great potential to generate high-quality person videos. Jue Wang 0001, Shaoli Huang, Dacheng Tao |
IJCNN | 1 |
| 2021 | Revitalizing CNN Attention via Transformers in Self-Supervised Visual Representation LearningabstractStudies on self-supervised visual representation learning (SSL) improve encoder backbones to discriminate training samples without labels. While CNN encoders via SSL achieve comparable recognition performance to those via supervised learning, their network attention is under-explored for further improvement. Motivated by the transformers that explore visual attention effectively in recognition scenarios, we propose a CNN Attention REvitalization (CARE) framework to train attentive CNN encoders guided by transformers in SSL. The proposed CARE framework consists of a CNN stream (C-stream) and a transformer stream (T-stream), where each stream contains two branches. C-stream follows an existing SSL framework with two CNN encoders, two projectors, and a predictor. T-stream contains two transformers, two projectors, and a predictor. T-stream connects to CNN encoders and is in parallel to the remaining C-Stream. During training, we perform SSL in both streams simultaneously and use the T-stream output to supervise C-stream. The features from CNN encoders are modulated in T-stream for visual attention enhancement and become suitable for the SSL scenario. We use these modulated features to supervise C-stream for learning attentive CNN encoders. To this end, we revitalize CNN attention by using transformers as guidance. Experiments on several standard visual recognition benchmarks, including image classification, object detection, and semantic segmentation, show that the proposed CARE framework improves CNN encoder backbones to the state-of-the-art performance. Chongjian Ge, Youwei Liang, Yibing Song, Jianbo Jiao, Jue Wang 0001, Ping Luo 0002 |
NeurIPS | 5 |
| 2021 | SDP-GAN: Saliency Detail Preservation Generative Adversarial Networks for High Perceptual Quality Style TransferabstractThe paper proposes a solution to effectively handle salient regions for style transfer between unpaired datasets. Recently, Generative Adversarial Networks (GAN) have demonstrated their potentials of translating images from source domain X to target domain Y in the absence of paired examples. However, such a translation cannot guarantee to generate high perceptual quality results. Existing style transfer methods work well with relatively uniform content, they often fail to capture geometric or structural patterns that always belong to salient regions. Detail losses in structured regions and undesired artifacts in smooth regions are unavoidable even if each individual region is correctly transferred into the target style. In this paper, we propose SDP-GAN, a GAN-based network for solving such problems while generating enjoyable style transfer results. We introduce a saliency network, which is trained with the generator simultaneously. The saliency network has two functions: (1) providing constraints for content loss to increase punishment for salient regions, and (2) supplying saliency features to generator to produce coherent results. Moreover, two novel losses are proposed to optimize the generator and saliency networks. The proposed method preserves the details on important salient regions and improves the total image perceptual quality. Qualitative and quantitative comparisons against several leading prior methods demonstrates the superiority of our method. Ru Li 0002, Chihao Wu 0001, Shuaicheng Liu, Jue Wang 0001, Guangfu Wang, Guanghui Liu 0001, Bing Zeng 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | OIFlow: Occlusion-Inpainting Optical Flow Estimation by Unsupervised LearningabstractOcclusion is an inevitable and critical problem in unsupervised optical flow learning. Existing methods either treat occlusions equally as non-occluded regions or simply remove them to avoid incorrectness. However, the occlusion regions can provide effective information for optical flow learning. In this paper, we present OIFlow, an occlusion-inpainting framework to make full use of occlusion regions. Specifically, a new appearance-flow network is proposed to inpaint occluded flows based on the image content. Moreover, a boundary dilated warp is proposed to deal with occlusions caused by displacement beyond the image border. We conduct experiments on multiple leading flow benchmark datasets such as Flying Chairs, KITTI and MPI-Sintel, which demonstrate that the performance is significantly improved by our proposed occlusion handling framework. Shuaicheng Liu, Kunming Luo, Nianjin Ye, Chuan Wang 0001, Jue Wang 0001, Bing Zeng 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Semi-Supervised Pixel-Level Scene Text Segmentation by Mutually Guided NetworkabstractIn this paper we present a new data-driven method for pixel-level scene text segmentation from a single natural image. Although scene text detection, i.e. producing a text region mask, has been well studied in the past decade, pixel-level text segmentation is still an open problem due to the lack of massive pixel-level labeled data for supervised training. To tackle this issue, we incorporate text region mask as an auxiliary data into this task, considering acquiring large-scale of labeled text region mask is commonly less expensive and time-consuming. To be specific, we propose a mutually guided network which produces a polygon-level mask in one branch and a pixel-level text mask in the other. The two branches' outputs serve as guidance for each other and the whole network is trained via a semi-supervised learning strategy. Extensive experiments are conducted to demonstrate the effectiveness of our mutually guided network, and experimental results show our network outperforms the state-of-the-art in pixel-level scene text segmentation. We also demonstrate the mask produced by our network could improve the text recognition performance besides the trivial image editing application. Chuan Wang 0001, Shan Zhao 0010, Li Zhu 0003, Kunming Luo, Yanwen Guo 0001, Jue Wang 0001, Shuaicheng Liu |
IEEE Trans. Image Process. | 6 |
| 2021 | Aesthetic-guided outward image croppingabstractImage cropping is a commonly used post-processing operation for adjusting the scene composition of an input photography, therefore improving its aesthetics. Existing automatic image cropping methods are all bounded by the image border, thus have very limited freedom for aesthetics improvement if the original scene composition is far from ideal, e.g. the main object is too close to the image border. In this paper, we propose a novel, aesthetic-guided outward image cropping method. It can go beyond the image border to create a desirable composition that is unachievable using previous cropping methods. Our method first evaluates the input image to determine how much the content of the image should be extrapolated by a field of view (FOV) evaluation model. We then synthesize the image content in the extrapolated region, and seek an optimal aesthetic crop within the expanded FOV, by jointly considering the aesthetics of the cropped view, and the local image quality of the extrapolated image content. Experimental results show that our method can generate more visually pleasing image composition in cases that are difficult for previous image cropping tools due to the border constraint, and can also automatically degrade to an inward method when high quality image extrapolation is infeasible. Feng-Heng Li, Hao-Zhi Huang 0001, Yong Zhang 0034, Shao-Ping Lu, Jue Wang 0001 |
ACM Trans. Graph. | 6 |
| 2020 | When AWGN-Based Denoiser Meets Real NoisesabstractDiscriminative learning based image denoisers have achieved promising performance on synthetic noises such as Additive White Gaussian Noise (AWGN). The synthetic noises adopted in most previous work are pixel-independent, but real noises are mostly spatially/channel-correlated and spatially/channel-variant. This domain gap yields unsatisfied performance on images with real noises if the model is only trained with AWGN. In this paper, we propose a novel approach to boost the performance of a real image denoiser which is trained only with synthetic pixel-independent noise data dominated by AWGN. First, we train a deep model that consists of a noise estimator and a denoiser with mixed AWGN and Random Value Impulse Noise (RVIN). We then investigate Pixel-shuffle Down-sampling (PD) strategy to adapt the trained model to real noises. Extensive experiments demonstrate the effectiveness and generalization of the proposed approach. Notably, our method achieves state-of-the-art performance on real sRGB images in the DND benchmark among models trained with synthetic noises. Codes are available at https://github.com/yzhouas/PD-Denoising-pytorch. Yuqian Zhou, Jianbo Jiao, Yang Wang 0023, Jue Wang 0001, Humphrey Shi, Thomas S. Huang |
AAAI | 5 |
| 2020 | Practical Deep Raw Image Denoising on Mobile Devices
Yuzhi Wang, Yiqun Liu 0001, Jue Wang 0001 |
ECCV (6) | 6 |
| 2020 | Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards
Heming Zhang 0003, Yingru Liu, Chihao Wu 0001, Jianchao Tan, Dongliang Xie, Jue Wang 0001, Xin Wang 0001 |
ECCV (13) | 8 |
| 2020 | Content-Aware Unsupervised Deep Homography Estimation
Jirong Zhang, Chuan Wang 0001, Shuaicheng Liu, Lanpeng Jia, Nianjin Ye, Jue Wang 0001, Ji Zhou 0001, Jian Sun 0001 |
ECCV (1) | 6 |
| 2020 | Skin Textural Generation via Blue-noise Gabor Filtering based Generative Adversarial NetworkabstractFacial skin texture synthesis is a fundamental problem in high-quality facial image generation and enhancement. The key behind is how to effectively synthesize plausible textured noise for the faces. With the development of CNNs and GANs, most works cast the problem as an image to image translation problem. However, these methods lack an explicit mechanism to simulate the facial noise pattern, so that the generated images are of obvious artifacts. To this end, we propose a new facial noise generation method. Specifically, we utilize the property of blue noise and Gabor filter to implicitly guide the asymmetrical sampling for the face region as a guidance map, where non-uniform point sampling is conducted. Thus we propose a novel Blue-Noise Gabor Module to produce a spatial-variant noisy image. Our proposed two-branch framework combined facial identity enhancing with textures details generation to jointly produce a high-quality facial image. Experimental results demonstrate the superiority of our method compared with the state-of-the-art, which enables the generation of high-quality facial texture based on a 2D image only, without the involvement of any 3D models. Hui Zhang 0027, Chuan Wang 0001, Nenglun Chen, Jue Wang 0001, Wenping Wang 0001 |
ACM Multimedia | 4 |
| 2020 | Deep Portrait Image Completion and ExtrapolationabstractGeneral image completion and extrapolation methods often fail on portrait images where parts of the human body need to be recovered -a task that requires accurate human body structure and appearance synthesis. We present a twostage deep learning framework for tackling this problem. In the first stage, given a portrait image with an incomplete human body, we extract a complete, coherent human body structure through a human parsing network, which focuses on structure recovery inside the unknown region with the help of full-body pose estimation. In the second stage, we use an image completion network to fill the unknown region, guided by the structure map recovered in the first stage. For realistic synthesis the completion network is trained with both perceptual loss and conditional adversarial loss.We further propose a face refinement network to improve the fidelity of the synthesized face region. We evaluate our method on publicly-available portrait image datasets, and show that it outperforms other state-of-the-art general image completion methods. Our method enables new portrait image editing applications such as occlusion removal and portrait extrapolation. We further show that the proposed general learning framework can be applied to other types of images, e.g. animal images. Xian Wu 0004, Ruilong Li, Jian-Cheng Liu, Jue Wang 0001, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Video Inpainting by Jointly Learning Temporal Structure and Spatial DetailsabstractWe present a new data-driven video inpainting method for recovering missing regions of video frames. A novel deep learning architecture is proposed which contains two subnetworks: a temporal structure inference network and a spatial detail recovering network. The temporal structure inference network is built upon a 3D fully convolutional architecture: it only learns to complete a low-resolution video volume given the expensive computational cost of 3D convolution. The low resolution result provides temporal guidance to the spatial detail recovering network, which performs imagebased inpainting with a 2D fully convolutional network to produce recovered video frames in their original resolution. Such two-step network design ensures both the spatial quality of each frame and the temporal coherence across frames. Our method jointly trains both sub-networks in an end-to-end manner. We provide qualitative and quantitative evaluation on three datasets, demonstrating that our method outperforms previous learning-based video inpainting methods. Chuan Wang 0001, Xiaoguang Han 0001, Jue Wang 0001 |
AAAI | 4 |
| 2019 | Adaptation Strategies for Applying AWGN-Based Denoiser to Realistic NoiseabstractDiscriminative learning based denoising model trained with Additive White Gaussian Noise (AWGN) performs well on synthesized noise. However, realistic noise can be spatialvariant, signal-dependent and a mixture of complicated noises. In this paper, we explore multiple strategies for applying an AWGN-based denoiser to realistic noise. Specifically, we trained a deep network integrating noise estimating and denoiser with mixed Gaussian (AWGN) and Random Value Impulse Noise (RVIN). To adapt the model to realistic noises, we investigated multi-channel, multi-scale and super-resolution approaches. Our preliminary results demonstrated the effectiveness of the newly-proposed noise model and adaptation strategies. Yuqian Zhou, Jianbo Jiao, Jue Wang 0001, Thomas S. Huang |
AAAI | 4 |
| 2019 | GeoNet: Deep Geodesic Networks for Point Cloud AnalysisabstractSurface-based geodesic topology provides strong cues for object semantic analysis and geometric modeling. However, such connectivity information is lost in point clouds. Thus we introduce GeoNet, the first deep learning architecture trained to model the intrinsic structure of surfaces represented as point clouds. To demonstrate the applicability of learned geodesic-aware representations, we propose fusion schemes which use GeoNet in conjunction with other baseline or backbone networks, such as PU-Net and PointNet++, for down-stream point cloud analysis. Our method improves the state-of-the-art on multiple representative tasks that can benefit from understandings of the underlying surface topology, including point upsampling, normal estimation, mesh reconstruction and non-rigid shape classification. Tong He 0002, Li Yi 0001, Yuqian Zhou, Chihao Wu 0001, Jue Wang 0001, Stefano Soatto |
CVPR | 6 |
| 2019 | GIF2Video: Color Dequantization and Temporal Interpolation of GIF ImagesabstractGraphics Interchange Format (GIF) is a highly portable graphics format that is ubiquitous on the Internet. Despite their small sizes, GIF images often contain undesirable visual artifacts such as flat color regions, false contours, color shift, and dotted patterns. In this paper, we propose GIF2Video, the first learning-based method for enhancing the visual quality of GIFs in the wild. We focus on the challenging task of GIF restoration by recovering information lost in the three steps of GIF creation: frame sampling, color quantization, and color dithering. We first propose a novel CNN architecture for color dequantization. It is built upon a compositional architecture for multi-step color correction, with a comprehensive loss function designed to handle large quantization errors. We then adapt the SuperSlomo network for temporal interpolation of GIF frames. We introduce two large datasets, namely GIF-Faces and GIF-Moments, for both training and evaluation. Experimental results show that our method can significantly improve the visual quality of GIFs, and outperforms direct baseline and state-of-the-art approaches. Yang Wang 0097, Chuan Wang 0001, Tong He 0002, Jue Wang 0001, Minh Hoai |
CVPR | 5 |
| 2019 | Disentangled Image MattingabstractMost previous image matting methods require a roughly-specificed trimap as input, and estimate fractional alpha values for all pixels that are in the unknown region of the trimap. In this paper, we argue that directly estimating the alpha matte from a coarse trimap is a major limitation of previous methods, as this practice tries to address two difficult and inherently different problems at the same time: identifying true blending pixels inside the trimap region, and estimate accurate alpha values for them. We propose AdaMatting, a new end-to-end matting framework that disentangles this problem into two sub-tasks: trimap adaptation and alpha estimation. Trimap adaptation is a pixel-wise classification problem that infers the global structure of the input image by identifying definite foreground, background, and semi-transparent image regions. Alpha estimation is a regression problem that calculates the opacity value of each blended pixel. Our method separately handles these two sub-tasks within a single deep convolutional neural network (CNN). Extensive experiments show that AdaMatting has additional structure awareness and trimap fault-tolerance. Our method achieves the state-of-the-art performance on Adobe Composition-1k dataset both qualitatively and quantitatively. It is also the current best-performing method on the alphamatting.com online evaluation for all commonly-used metrics. Shaofan Cai, Xiaoshuai Zhang, Haoqiang Fan, Jiangyu Liu, Jiaying Liu 0001, Jue Wang 0001, Jian Sun 0001 |
ICCV | 8 |
| 2019 | Semi-Supervised Skin Detection by Network With Mutual GuidanceabstractWe present a new data-driven method for robust skin detection from a single human portrait image. Unlike previous methods, we incorporate human body as a weak semantic guidance into this task, considering acquiring large-scale of human labeled skin data is commonly expensive and time-consuming. To be specific, we propose a dual-task neural network for joint detection of skin and body via a semi-supervised learning strategy. The dual-task network contains a shared encoder but two decoders for skin and body separately. For each decoder, its output also serves as a guidance for its counterpart, making both decoders mutually guided. Extensive experiments were conducted to demonstrate the effectiveness of our network with mutual guidance, and experimental results show our network outperforms the state-of-the-art in skin detection. Jiayuan Shi, Chuan Wang 0001, Guanbin Li, Risheng Liu, Jue Wang 0001 |
ICCV | 8 |
| 2019 | Not All Parts Are Created Equal: 3D Pose Estimation by Modeling Bi-Directional Dependencies of Body PartsabstractNot all the human body parts have the same degree of freedom (DOF) due to the physiological structure. For example, the limbs may move more flexibly and freely than the torso does. Most of the existing 3D pose estimation methods, despite the very promising results achieved, treat the body joints equally and consequently often lead to larger reconstruction errors on the limbs. In this paper, we propose a progressive approach that explicitly accounts for the distinct DOFs among the body parts. We model parts with higher DOFs like the elbows, as dependent components of the corresponding parts with lower DOFs like the torso, of which the 3D locations can be more reliably estimated. Meanwhile, the high-DOF parts may, in turn, impose a constraint on where the low-DOF ones lie. As a result, parts with different DOFs supervise one another, yielding physically constrained and plausible pose-estimation results. To further facilitate the prediction of the high-DOF parts, we introduce a pose-attribution estimation, where the relative location of a limb joint with respect to the torso, which has the least DOF of a human body, is explicitly estimated and further fed to the joint-estimation module. The proposed approach achieves very promising results, outperforming the state of the art on several benchmarks. Jue Wang 0001, Shaoli Huang, Xinchao Wang, Dacheng Tao |
ICCV | 1 |
| 2019 | A Semi-Formal Requirement Modeling Pattern for Designing Industrial Cyber-Physical SystemsabstractRequirement engineering is a crucial part of the engineering process. The traditional methods of requirement engineering are time-consuming and human-centered. A well-established software requirement description model needs to ensure the accuracy and integrity of the transformation and is also hoped to be scalable, versatile, and efficient in transformation and transmission. This paper presents a method of requirement engineering, including constricted nature language requirement input pattern, and the formalized requirement description JSON model. This method provides convenience for requirement modification and validation that can satisfy the real-time constraints of industrial cyber-physical systems. Jue Wang 0001, Yineng Song, Xian Wu 0004, Wenbin William Dai |
IECON | 1 |
| 2019 | An Automatic Requirement Transformation Approach for Code Generation in Industrial Cyber-Physical SystemsabstractRequirement Engineering (RE) is a crucial step prior to building their control program for industries. Traditional requirement engineering is time-consuming and error-prone that an optimized way is necessary to build an accurate model within fractions of a second. In this paper, an automatic software requirement modeling method is proposed to achieve fast and accurate requirement modeling. A new requirement model is designed to prove the concept. The newly designed model is designed based on the flexible object-oriented structure with process flow and constraints. Furthermore, to be compatible with the UML model format, the transformation rules and algorithm are provided. The requirement input system uses a general defined model, which is also the same format as the newly designed model, to deal with information from any input form like drawings, semiformal text and uploaded UML model files. Requirement model is further transformed into control code as illustrated in this paper. Yineng Song, Xian Wu 0004, Jue Wang 0001, Wenbin William Dai |
INDIN | 4 |
| 2018 | Live Sketch: Video-driven Dynamic Deformation of Static DrawingsabstractCreating sketch animations using traditional tools requires special artistic skills, and is tedious even for trained professionals. To lower the barrier for creating sketch animations, we propose a new system, emphLive Sketch, which allows novice users to interactively bring static drawings to life by applying deformation-based animation effects that are extracted from video examples. Dynamic deformation is first extracted as a sparse set of moving control points from videos and then transferred to a static drawing. Our system addresses a few major technical challenges, such as motion extraction from video, video-to-sketch alignment, and many-to-one motion-driven sketch animation. While each of the sub-problems could be difficult to solve fully automatically, we present reliable solutions by combining new computational algorithms with intuitive user interactions. Our pilot study shows that our system allows both users with or without animation skills to easily add dynamic deformation to static drawings. Qingkun Su, Hongbo Fu 0001, Chiew-Lan Tai, Jue Wang 0001 |
CHI | 5 |
| 2018 | DocUNet: Document Image Unwarping via a Stacked U-NetabstractCapturing document images is a common way for digitizing and recording physical documents due to the ubiquitousness of mobile cameras. To make text recognition easier, it is often desirable to digitally flatten a document image when the physical document sheet is folded or curved. In this paper, we develop the first learning-based method to achieve this goal. We propose a stacked U-Net [25] with intermediate supervision to directly predict the forward mapping from a distorted image to its rectified version. Because large-scale real-world data with ground truth deformation is difficult to obtain, we create a synthetic dataset with approximately 100 thousand images by warping non-distorted document images. The network is trained on this dataset with various data augmentations to improve its generalization ability. We further create a comprehensive benchmark1 that covers various real-world conditions. We evaluate the proposed model quantitatively and qualitatively on the proposed benchmark, and compare it with previous non-learning-based methods. Ke Ma 0001, Zhixin Shu, Jue Wang 0001, Dimitris Samaras |
CVPR | 4 |
| 2018 | Scale-Recurrent Network for Deep Image DeblurringabstractIn single image deblurring, the "coarse-to-fine" scheme, i.e. gradually restoring the sharp image on different resolutions in a pyramid, is very successful in both traditional optimization-based methods and recent neural-network-based approaches. In this paper, we investigate this strategy and propose a Scale-recurrent Network (SRN-DeblurNet) for this deblurring task. Compared with the many recent learning-based approaches in [25], it has a simpler network structure, a smaller number of parameters and is easier to train. We evaluate our method on large-scale deblurring datasets with complex motion. Results show that our method can produce better quality results than state-of-the-arts, both quantitatively and qualitatively. Xin Tao 0001, Hongyun Gao 0001, Xiaoyong Shen, Jue Wang 0001, Jiaya Jia |
CVPR | 4 |
| 2018 | Deblurring Low-Light Images with Light StreaksabstractImages acquired in low-light conditions with handheld cameras are often blurry, so steady poses and long exposure time are required to alleviate this problem. Although significant advances have been made in image deblurring, state-of-the-art approaches often fail on low-light images, as a sufficient number of salient features cannot be extracted for blur kernel estimation. On the other hand, light streaks are common phenomena in low-light images that have not been extensively explored in existing approaches. In this work, we propose an algorithm that utilizes light streaks to facilitate deblurring low-light images. The light streaks, which commonly exist in the low-light blurry images, contain rich information regarding camera motion and blur kernels. A method is developed in this work to detect light streaks for kernel estimation. We introduce a non-linear blur model that explicitly takes light streaks and corresponding light sources into account, and pose them as constraints for estimating the blur kernel in an optimization framework. For practical applications, the proposed algorithm is extended to handle images undergoing non-uniform blur. Experimental results show that the proposed algorithm performs favorably against the state-of-the-art methods on deblurring real-world low-light images. Sunghyun Cho, Jue Wang 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Detecting and Removing Visual Distractors for Video Aesthetic EnhancementabstractPersonal videos often contain visual distractors, which are objects that are accidentally captured and can distract viewers from focusing on the main subjects. We propose a method to automatically detect and localize these distractors through learning from a manually labeled dataset. To achieve spatially and temporally coherent detection, we propose extracting features at the temporal-superpixel level using a traditional supporting vector machine based learning framework. We also experiment with end-to-end learning using convolutional neural networks, which achieves slightly higher performance than other methods. The classification result is further refined in a postprocessing step based on graph-cut optimization. Experimental results show that our method achieves an accuracy of 81% and a recall of 86%. We demonstrate several ways of removing the detected distractors to improve the video quality, including video hole filling, video frame replacement, and camera path replanning. The user study results show that our method can significantly improve the aesthetic quality of videos. Xian Wu 0004, Ruilong Li, Jue Wang 0001, Zhao-Heng Zheng, Shi-Min Hu 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | Deep Video Deblurring for Hand-Held CamerasabstractMotion blur from camera shake is a major problem in videos captured by hand-held devices. Unlike single-image deblurring, video-based approaches can take advantage of the abundant information that exists across neighboring frames. As a result the best performing methods rely on the alignment of nearby frames. However, aligning images is a computationally expensive and fragile procedure, and methods that aggregate information must therefore be able to identify which regions have been accurately aligned and which have not, a task that requires high level scene understanding. In this work, we introduce a deep learning solution to video deblurring, where a CNN is trained end-to-end to learn how to accumulate information across frames. To train this network, we collected a dataset of real videos recorded with a high frame rate camera, which we use to generate synthetic motion blur for supervision. We show that the features learned from this dataset extend to deblurring motion blur that arises due to camera shake in a wide range of videos, and compare the quality of results to a number of other baselines. Shuochen Su, Mauricio Delbracio, Jue Wang 0001, Guillermo Sapiro, Wolfgang Heidrich, Oliver Wang |
CVPR | 3 |
| 2017 | Detail-Revealing Deep Video Super-ResolutionabstractPrevious CNN-based video super-resolution approaches need to align multiple frames to the reference. In this paper, we show that proper frame alignment and motion compensation is crucial for achieving high quality results. We accordingly propose a “sub-pixel motion compensation” (SPMC) layer in a CNN framework. Analysis and experiments show the suitability of this layer in video SR. The final end-to-end, scalable CNN framework effectively incorporates the SPMC layer and fuses multiple frames to reveal image details. Our implementation can generate visually and quantitatively high-quality results, superior to current state-of-the-arts, without the need of parameter tuning. Xin Tao 0001, Hongyun Gao 0001, Renjie Liao 0001, Jue Wang 0001, Jiaya Jia |
ICCV | 4 |
| 2017 | Zero-Order Reverse FilteringabstractIn this paper, we study an unconventional but practically meaningful reversibility problem of commonly used image filters. We broadly define filters as operations to smooth images or to produce layers via global or local algorithms. And we raise the intriguingly problem if they are reservable to the status before filtering. To answer it, we present a novel strategy to understand general filter via contraction mappings on a metric space. A very simple yet effective zero-order algorithm is proposed. It is able to practically reverse most filters with low computational cost. We present quite a few experiments in the paper and supplementary file to thoroughly verify its performance. This method can also be generalized to solve other inverse problems and enables new applications. Xin Tao 0001, Chao Zhou 0001, Xiaoyong Shen, Jue Wang 0001, Jiaya Jia |
ICCV | 4 |
| 2017 | Time slice video synthesis by robust video alignmentabstractTime slice photography is a popular effect that visualizes the passing of time by aligning and stitching multiple images capturing the same scene at different times together into a single image. Extending this effect to video is a difficult problem, and one where existing solutions have only had limited success. In this paper, we propose an easy-to-use and robust system for creating time slice videos from a wide variety of consumer videos. The main technical challenge we address is how to align videos taken at different times with substantially different appearances, in the presence of moving objects and moving cameras with slightly different trajectories. To achieve a temporally stable alignment, we perform a mixed 2D-3D alignment, where a rough 3D reconstruction is used to generate sparse constraints that are integrated into a pixelwise 2D registration. We apply our method to a number of challenging scenarios, and show that we can achieve a higher quality registration than prior work. We propose a 3D user interface that allows the user to easily specify how multiple videos should be composited in space and time. Finally, we show that our alignment method can be applied in more general video editing and compositing tasks, such as object removal. Zhaopeng Cui, Oliver Wang, Ping Tan 0002, Jue Wang 0001 |
ACM Trans. Graph. | 4 |
| 2017 | PlenoPatch: Patch-Based Plenoptic Image ManipulationabstractPatch-based image synthesis methods have been successfully applied for various editing tasks on still images, videos and stereo pairs. In this work we extend patch-based synthesis to plenoptic images captured by consumer-level lenselet-based devices for interactive, efficient light field editing. In our method the light field is represented as a set of images captured from different viewpoints. We decompose the central view into different depth layers, and present it to the user for specifying the editing goals. Given an editing task, our method performs patch-based image synthesis on all affected layers of the central view, and then propagates the edits to all other views. Interaction is done through a conventional 2D image editing user interface that is familiar to novice users. Our method correctly handles object boundary occlusion with semi-transparency, thus can generate more realistic results than previous methods. We demonstrate compelling results on a wide range of applications such as hole-filling, object reshuffling and resizing, changing object depth, light field upscaling and parallax magnification. Jue Wang 0001, Eli Shechtman, Zi-Ye Zhou, Jiaxin Shi, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | Removing Shadows from Images of Documents
Steve Bako, Soheil Darabi, Eli Shechtman, Jue Wang 0001, Kalyan Sunkavalli, Pradeep Sen |
ACCV (3) | 4 |
| 2016 | Coherent Parametric Contours for Interactive Video Object SegmentationabstractInteractive video segmentation systems aim at producing sub-pixel-level object boundaries for visual effect applications. Recent approaches mainly focus on using sparse user input (i.e. scribbles) for efficient segmentation, however, the quality of the final object boundaries is not satisfactory for the following reasons: (1) the boundary on each frame is often not accurate, (2) boundaries across adjacent frames wiggle around inconsistently, causing temporal flickering, and (3) there is a lack of direct user control for fine tuning. We propose Coherent Parametric Contours, a novel video segmentation propagation framework that addresses all the above issues. Our approach directly models the object boundary using a set of parametric curves, providing direct user controls for manual adjustment. A spatiotemporal optimization algorithm is employed to produce object boundaries that are spatially accurate and temporally stable. We show that existing evaluation datasets are limited and demonstrate a new set to cover the common cases in professional rotoscoping. A new metric for evaluating temporal consistency is proposed. Results show that our approach generates higher quality, more coherent segmentation results than previous methods. Yao Lu 0028, Linda G. Shapiro, Jue Wang 0001 |
CVPR | 4 |
| 2016 | Automatic Fence Segmentation in Videos of Dynamic ScenesabstractWe present a fully automatic approach to detect and segment fence-like occluders from a video clip. Unlike previous approaches that usually assume either static scenes or cameras, our method is capable of handling both dynamic scenes and moving cameras. Under a bottom-up framework, it first clusters pixels into coherent groups using color and motion features. These pixel groups are then analyzed in a fully connected graph, and labeled as either fence or non-fence using graph-cut optimization. Finally, we solve a dense Conditional Random Filed (CRF) constructed from multiple frames to enhance both spatial accuracy and temporal coherence of the segmentation. Once segmented, one can use existing hole-filling methods to generate a fencefree output. Extensive evaluation suggests that our method outperforms previous automatic and interactive approaches on complex examples captured by mobile devices. Renjiao Yi, Jue Wang 0001, Ping Tan 0002 |
CVPR | 2 |
| 2016 | Robust Image and Video Dehazing with Visual Artifact Suppression via Gradient Residual Minimization
Chen Chen 0003, Minh N. Do, Jue Wang 0001 |
ECCV (2) | 3 |
| 2016 | Learning High-Order Filters for Efficient Blind Deconvolution of Document Photographs
Lei Xiao 0014, Jue Wang 0001, Wolfgang Heidrich, Michael Hirsch 0001 |
ECCV (3) | 2 |
| 2016 | PanoSwarm: Collaborative and Synchronized Multi-Device Panoramic PhotographyabstractTaking a picture has been traditionally a one-person task. In this paper we present a novel system that allows multiple mobile devices to work collaboratively in a synchronized fashion to capture a panorama of a highly dynamic scene, creating an entirely new photography experience that encourages social interactions and teamwork. Our system contains two components: a client app that runs on all participating devices, and a server program that monitors and communicates with each device. In a capturing session, the server collects in realtime the viewfinder images of all devices and stitches them on-the-fly to create a panorama preview, which is then streamed to all devices as visual guidance. The system also allows one camera to be the host and send direct visual instructions to others to guide camera adjustment. When ready, all devices take pictures at the same time for panorama stitching. Our preliminary study suggests that the proposed system can help users capture high quality panoramas with an enjoyable teamwork experience. Yan Wang 0059, Sunghyun Cho, Jue Wang 0001, Shih-Fu Chang |
IUI | 3 |
| 2016 | Re-Compositable Panoramic Selfie with Robust Multi-Frame Segmentation and StitchingabstractAbstract It is a challenging task for ordinary users to capture selfies with a good scene composition, given the limited freedom to position the camera. Creative hardware (e.g., selfie sticks) and software (e.g., panoramic selfie apps) solutions have been proposed to extend the background coverage of a selife, but to achieve a perfect composition on the spot when the selfie is captured remains to be difficult. In this paper, we propose a system that allows the user to shoot a selfie video by rotating the body first, then produce a final panoramic selfie image with user‐guided scene composition as postprocessing. Our key technical contribution is a fully Automatic, robust multi‐frame segmentation and stitching framework that is tailored towards the special characteristics of selfie images. We analyze the sparse feature points and employ a spatial‐temporal optimization for bilayer feature segmentation, which leads to more reliable background alignment than previous image stitching techniques. The sparse classification is then propagated to all pixels to create dense foreground masks for person‐background composition. Finally, based on a user‐selected foreground position, our system uses content‐preserving warping to produce a panoramic seflie with minimal distortion to the face region. Experimental results show that our approach can reliably generate high quality panoramic selfies, while a simple combination of previous image stitching and segmentation approaches often fails. Kai Li 0016, Jue Wang 0001, Yebin Liu, Qionghai Dai |
Comput. Graph. Forum | 2 |
| 2016 | Appearance Harmonization for Single Image Shadow RemovalabstractAbstract Shadow removal is a challenging problem and previous approaches often produce de‐shadowed regions that are visually inconsistent with the rest of the image. We propose an automaticshadow region harmonizationapproach that makes the appearance of a de‐shadowed region (produced using any previous technique) compatible with the rest of the image. We use a shadow‐guided patch‐based image synthesis approach that reconstructs the shadow region using patches sampled from non‐shadowed regions. This result is then refined based on the reconstruction confidence to handle unique textures. Qualitative comparisons over a wide range of images, and a quantitative evaluation on a benchmark dataset show that our technique significantly improves upon the state‐of‐the‐art. Li-Qian Ma, Jue Wang 0001, Eli Shechtman, Kalyan Sunkavalli, Shi-Min Hu 0001 |
Comput. Graph. Forum | 2 |
| 2016 | Blur-Kernel Bound Estimation From Pyramid StatisticsabstractThis letter presents an approach for automatically estimating the spatial bound of the blur kernel in a motion-blurred image based on the statistics of multilevel image gradients. We observe that blur has a significant impact on the histogram of oriented gradients (HOGs) at higher levels of an image pyramid, but has much less of an impact at coarser levels. Based on this fact, we estimate the spatial bound of the unknown blur kernel using a learning-based approach. We first learn a generic pyramid HOG model from natural sharp images, then given an HOG pyramid of a blurry image, we predict the corresponding model of its latent sharp image. Finally, we learn another model to predict the spatial kernel bound from the difference between the observed and the predicted HOG pyramids. Experimental results show that the proposed method can estimate accurate blur kernel sizes, enabling existing blind deconvolution methods to achieve best possible results. Shaoguo Liu, Jue Wang 0001, Chunhong Pan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Automatic triage for a photo seriesabstractPeople often take a series of nearly redundant pictures to capture a moment or scene. However, selecting photos to keep or share from a large collection is a painful chore. To address this problem, we seek a relative quality measure within a series of photos taken of the same scene, which can be used for automatic photo triage. Towards this end, we gather a large dataset comprised of photo series distilled from personal photo albums. The dataset contains 15, 545 unedited photos organized in 5,953 series. By augmenting this dataset with ground truth human preferences among photos within each series, we establish a benchmark for measuring the effectiveness of algorithmic models of how people select photos. We introduce several new approaches for modeling human preference based on machine learning. We also describe applications for the dataset and predictor, including a smart album viewer, automatic photo enhancement, and providing overviews of video clips. Huiwen Chang, Fisher Yu 0001, Jue Wang 0001, Douglas Ashley, Adam Finkelstein |
ACM Trans. Graph. | 3 |
| 2016 | Robust background identification for dynamic video editingabstractExtracting background features for estimating the camera path is a key step in many video editing and enhancement applications. Existing approaches often fail on highly dynamic videos that are shot by moving cameras and contain severe foreground occlusion. Based on existing theories, we present a new, practical method that can reliably identify background features in complex video, leading to accurate camera path estimation and background layering. Our approach contains a local motion analysis step and a global optimization step. We first divide the input video into overlapping temporal windows, and extract local motion clusters in each window. We form a directed graph from these local clusters, and identify background ones by finding a minimal path through the graph using optimization. We show that our method significantly outperforms other alternatives, and can be directly used to improve common video editing applications such as stabilization, compositing and background reconstruction. Xian Wu 0004, Hao-Tian Zhang, Jue Wang 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 4 |
| 2015 | Blind optical aberration correction by exploring geometric and visual priorsabstractOptical aberration widely exists in optical imaging systems, especially in consumer-level cameras. In contrast to previous solutions using hardware compensation or pre-calibration, we propose a computational approach for blind aberration removal from a single image, by exploring various geometric and visual priors. The global rotational symmetry allows us to transform the non-uniform degeneration into several uniform ones by the proposed radial splitting and warping technique. Locally, two types of symmetry constraints, i.e. central symmetry and reflection symmetry are defined as geometric priors in central and surrounding regions, respectively. Furthermore, by investigating the visual artifacts of aberration degenerated images captured by consumer-level cameras, the non-uniform distribution of sharpness across color channels and the image lattice is exploited as visual priors, resulting in a novel strategy to utilize the guidance from the sharpest channel and local image regions to improve the overall performance and robustness. Extensive evaluation on both real and synthetic data suggests that the proposed method outperforms the state-of-the-art techniques. Tao Yue 0003, Jin-Li Suo, Jue Wang 0001, Xun Cao, Qionghai Dai |
CVPR | 3 |
| 2015 | Structuring Lecture Videos by Automatic Projection Screen Localization and AnalysisabstractWe present a fully automatic system for extracting the semantic structure of a typical academic presentation video, which captures the whole presentation stage with abundant camera motions such as panning, tilting, and zooming. Our system automatically detects and tracks both the projection screen and the presenter whenever they are visible in the video. By analyzing the image content of the tracked screen region, our system is able to detect slide progressions and extract a high-quality, non-occluded, geometrically-compensated image for each slide, resulting in a list of representative images that reconstruct the main presentation structure. Afterwards, our system recognizes text content and extracts keywords from the slides, which can be used for keyword-based video retrieval and browsing. Experimental results show that our system is able to generate more stable and accurate screen localization results than commonly-used object tracking methods. Our system also extracts more accurate presentation structures than general video summarization methods, for this specific type of video. Kai Li 0016, Jue Wang 0001, Haoqian Wang, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Simultaneous Camera Path Optimization and Distraction Removal for Improving Amateur VideoabstractA major difference between amateur and professional video lies in the quality of camera paths. Previous work on video stabilization has considered how to improve amateur video by smoothing the camera path. In this paper, we show that additional changes to the camera path can further improve video aesthetics. Our new optimization method achieves multiple simultaneous goals: 1) stabilizing video content over short time scales; 2) ensuring simple and consistent camera paths over longer time scales; and 3) improving scene composition by automatically removing distractions, a common occurrence in amateur video. Our approach uses an L(1) camera path optimization framework, extended to handle multiple constraints. Two passes of optimization are used to address both low-level and high-level constraints on the camera path. The experimental and user study results show that our approach outputs video that is perceptually better than the input, or the results of using stabilization only. Jue Wang 0001, Ralph R. Martin, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Automatic blur-kernel-size estimation for motion deblurring
Shaoguo Liu, Jue Wang 0001, Sunghyun Cho, Chunhong Pan |
Vis. Comput. | 3 |
| 2014 | Deblurring Low-Light Images with Light StreaksabstractImages taken in low-light conditions with handheld cameras are often blurry due to the required long exposure time. Although significant progress has been made recently on image deblurring, state-of-the-art approaches often fail on low-light images, as these images do not contain a sufficient number of salient features that deblurring methods rely on. On the other hand, light streaks are common phenomena in low-light images that contain rich blur information, but have not been extensively explored in previous approaches. In this work, we propose a new method that utilizes light streaks to help deblur low-light images. We introduce a non-linear blur model that explicitly models light streaks and their underlying light sources, and poses them as constraints for estimating the blur kernel in an optimization framework. Our method also automatically detects useful light streaks in the input image. Experimental results show that our approach obtains good results on challenging real-world examples that no other methods could achieve before. Sunghyun Cho, Jue Wang 0001, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2014 | Investigating Haze-Relevant Features in a Learning Framework for Image DehazingabstractHaze is one of the major factors that degrade outdoor images. Removing haze from a single image is known to be severely ill-posed, and assumptions made in previous methods do not hold in many situations. In this paper, we systematically investigate different haze-relevant features in a learning framework to identify the best feature combination for image dehazing. We show that the dark-channel feature is the most informative one for this task, which confirms the observation of He et al. [8] from a learning perspective, while other haze-relevant features also contribute significantly in a complementary way. We also find that surprisingly, the synthetic hazy image patches we use for feature investigation serve well as training data for realworld images, which allows us to train specific models for specific applications. Experiment results demonstrate that the proposed algorithm outperforms state-of-the-art methods on both synthetic and real-world datasets. Ketan Tang, Jianchao Yang, Jue Wang 0001 |
CVPR | 3 |
| 2014 | Good Image Priors for Non-blind Deconvolution - Generic vs. Specific
Libin Sun, Sunghyun Cho, Jue Wang 0001, James Hays |
ECCV (4) | 3 |
| 2014 | Discriminative Indexing for Probabilistic Image Patch Priors
Yan Wang 0059, Sunghyun Cho, Jue Wang 0001, Shih-Fu Chang |
ECCV (4) | 3 |
| 2014 | Hybrid Image Deblurring by Fusing Edge and Power Spectrum Information
Tao Yue 0003, Sunghyun Cho, Jue Wang 0001, Qionghai Dai |
ECCV (7) | 3 |
| 2014 | Confidence-driven image co-matting
Linbo Wang 0001, Tianchen Xia, Yanwen Guo 0001, Ligang Liu 0001, Jue Wang 0001 |
Comput. Graph. | 5 |
| 2014 | Automatic Upright Adjustment of Photographs With Robust Camera CalibrationabstractMan-made structures often appear to be distorted in photos captured by casual photographers, as the scene layout often conflicts with how it is expected by human perception. In this paper, we propose an automatic approach for straightening up slanted man-made structures in an input image to improve its perceptual quality. We call this type of correction upright adjustment. We propose a set of criteria for upright adjustment based on human perception studies, and develop an optimization framework which yields an optimal homography for adjustment. We also develop a new optimization-based camera calibration method that performs favorably to previous methods and allows the proposed system to work reliably for a wide range of images. The effectiveness of our system is demonstrated by both quantitative comparisons and qualitative user study. Hyunjoon Lee, Eli Shechtman, Jue Wang 0001, Seungyong Lee 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | A Data-Driven Approach for Facial Expression Retargeting in VideoabstractThis paper presents a data-driven approach for facial expression retargeting in video, i.e., synthesizing a face video of a target subject that mimics the expressions of a source subject in the input video. Our approach takes advantage of a pre-existing facial expression database of the target subject to achieve realistic synthesis. First, for each frame of the input video, a new facial expression similarity metric is proposed for querying the expression database of the target person to select multiple candidate images that are most similar to the input. The similarity metric is developed using a metric learning approach to reliably handle appearance difference between different subjects. Secondly, we employ an optimization approach to choose the best candidate image for each frame, resulting in a retrieved sequence that is temporally coherent. Finally, a spatio-temporal expression mapping method is employed to further improve the synthesized sequence. Experimental results show that our system is capable of generating high quality facial expression videos that match well with the input sequences, even when the source and target subjects have big identity difference. In addition, extensive evaluations demonstrate the high accuracy of the learned expression similarity metric and the effectiveness of our retrieval strategy. Kai Li 0016, Qionghai Dai, Ruiping Wang 0001, Yebin Liu, Feng Xu 0005, Jue Wang 0001 |
IEEE Trans. Multim. | 6 |
| 2014 | TrackCam: 3D-aware tracking shots from consumer videoabstractPanning and tracking shots are popular photography techniques in which the camera tracks a moving object and keeps it at the same position, resulting in an image where the moving foreground is sharp but the background is blurred accordingly, creating an artistic illustration of the foreground motion. Such shots however are hard to capture even for professionals, especially when the foreground motion is complex (e.g., non-linear motion trajectories). In this work we propose a system to generate realistic, 3D-aware tracking shots from consumer videos. We show how computer vision techniques such as segmentation and structure-from-motion can be used to lower the barrier and help novice users create high quality tracking shots that are physically plausible. We also introduce a pseudo 3D approach for relative depth estimation to avoid expensive 3D reconstruction for improved robustness and a wider application range. We validate our system through extensive quantitative and qualitative evaluations. Shuaicheng Liu, Jue Wang 0001, Sunghyun Cho, Ping Tan 0002 |
ACM Trans. Graph. | 2 |
| 2014 | EZ-sketching: three-level optimization for error-tolerant image tracingabstractWe present a new image-guided drawing interface called EZ-Sketching , which uses a tracing paradigm and automatically corrects sketch lines roughly traced over an image by analyzing and utilizing the image features being traced. While previous edge snapping methods aim at optimizing individual strokes, we show that a co-analysis of multiple roughly placed nearby strokes better captures the user's intent. We formulate automatic sketch improvement as a three-level optimization problem and present an efficient solution to it. EZ-Sketching can tolerate errors from various sources such as indirect control and inherently inaccurate input, and works well for sketching on touch devices with small screens using fingers. Our user study confirms that the drawings our approach helped generate show closer resemblance to the traced images, and are often aesthetically more pleasing. Qingkun Su, Wing Ho Andy Li, Jue Wang 0001, Hongbo Fu 0001 |
ACM Trans. Graph. | 3 |
| 2013 | Supervised Semantic Gradient Extraction Using Linear-Time OptimizationabstractThis paper proposes a new supervised semantic edge and gradient extraction approach, which allows the user to roughly scribble over the desired region to extract semantically-dominant and coherent edges in it. Our approach first extracts low-level edge lets (small edge clusters) from the input image as primitives and build a graph upon them, by jointly considering both the geometric and appearance compatibility of edge lets. Given the characteristics of the graph, it cannot be effectively optimized by commonly-used energy minimization tools such as graph cuts. We thus propose an efficient linear algorithm for precise graph optimization, by taking advantage of the special structure of the graph. %Optimal parameter settings of the model are learnt from a dataset. Objective evaluations show that the proposed method significantly outperforms previous semantic edge detection algorithms. Finally, we demonstrate the effectiveness of the system in various image editing tasks. Shulin Yang, Jue Wang 0001, Linda G. Shapiro |
CVPR | 2 |
| 2013 | Handling Noise in Single Image Deblurring Using Directional FiltersabstractState-of-the-art single image deblurring techniques are sensitive to image noise. Even a small amount of noise, which is inevitable in low-light conditions, can degrade the quality of blur kernel estimation dramatically. The recent approach of Tai and Lin [17] tries to iteratively denoise and deblur a blurry and noisy image. However, as we show in this work, directly applying image denoising methods often partially damages the blur information that is extracted from the input image, leading to biased kernel estimation. We propose a new method for handling noise in blind image deconvolution based on new theoretical and practical insights. Our key observation is that applying a directional low-pass filter to the input image greatly reduces the noise level, while preserving the blur information in the orthogonal direction to the filter. Based on this observation, our method applies a series of directional filters at different orientations to the input image, and estimates an accurate Radon transform of the blur kernel from each filtered image. Finally, we reconstruct the blur kernel using inverse Radon transform. Experimental results on synthetic and real data show that our algorithm achieves higher quality results than previous approaches on blurry and noisy images. Lin Zhong 0002, Sunghyun Cho, Dimitris N. Metaxas, Sylvain Paris, Jue Wang 0001 |
CVPR | 5 |
| 2013 | Edge-based blur kernel estimation using patch priorsabstractBlind image deconvolution, i.e., estimating a blur kernel k and a latent image x from an input blurred image y, is a severely ill-posed problem. In this paper we introduce a new patch-based strategy for kernel estimation in blind deconvolution. Our approach estimates a “trusted” subset of x by imposing a patch prior specifically tailored towards modeling the appearance of image edge and corner primitives. To choose proper patch priors we examine both statistical priors learned from a natural image dataset and a simple patch prior from synthetic structures. Based on the patch priors, we iteratively recover the partial latent image x and the blur kernel k. A comprehensive evaluation shows that our approach achieves state-of-the-art results for uniformly blurred images. Libin Sun, Sunghyun Cho, Jue Wang 0001, James Hays |
ICCP | 3 |
| 2013 | Text Localization in Natural Images Using Stroke Feature Transform and Text Covariance DescriptorsabstractIn this paper, we present a new approach for text localization in natural images, by discriminating text and non-text regions at three levels: pixel, component and text line levels. Firstly, a powerful low-level filter called the Stroke Feature Transform (SFT) is proposed, which extends the widely-used Stroke Width Transform (SWT) by incorporating color cues of text pixels, leading to significantly enhanced performance on inter-component separation and intra-component connection. Secondly, based on the output of SFT, we apply two classifiers, a text component classifier and a text-line classifier, sequentially to extract text regions, eliminating the heuristic procedures that are commonly used in previous approaches. The two classifiers are built upon two novel Text Covariance Descriptors (TCDs) that encode both the heuristic properties and the statistical characteristics of text stokes. Finally, text regions are located by simply thresholding the text-line confident map. Our method was evaluated on two benchmark datasets: ICDAR 2005 and ICDAR 2011, and the corresponding F-measure values are 0.72 and 0.73, respectively, surpassing previous methods in accuracy by a large margin. Zhe Lin 0001, Jianchao Yang, Jue Wang 0001 |
ICCV | 4 |
| 2013 | Kinect depth restoration via energy minimization with TV21 regularizationabstractDepth maps generated by Kinect cameras often contain a significant amount of missing pixels and strong noise, limiting their usability in many computer vision applications. We present a new energy minimization method to fill the missing regions and remove noise in a depth map, by exploiting the strong correlation between color and depth values in local image neighborhoods. To preserve sharp edges and remove noise from the depth map, we propose to add a TV21regularization term into the energy function. Finally, we show how to effectively minimize the total energy using an alternating optimization approach. Experimental results show that the proposed method outperforms commonly-used depth inpainting approaches. Shaoguo Liu, Ying Wang 0008, Jue Wang 0001, Jixia Zhang, Chunhong Pan |
ICIP | 3 |
| 2013 | A Progressive Tri-level Segmentation Approach for Topology-Change-Aware Video MattingabstractAbstract Previous video matting approaches mostly adopt the “binary segmentation + matting” strategy, i.e., first segment each frame into foreground and background regions, then extract the fine details of the foreground boundary using matting techniques. This framework has several limitations due to the fact that binary segmentation is employed. In this paper, we propose a new supervised video matting approach. Instead of applying binary segmentation, we explicitly model segmentation uncertainty in a novel tri‐level segmentation procedure. The segmentation is done progressively, enabling us to handle difficult cases such as large topology changes, which are challenging to previous approaches. The tri‐level segmentation results can be naturally fed into matting techniques to generate the final alpha mattes. Experimental results show that our system can generate high quality results with less user inputs than the state‐of‐theart methods. Jinlong Ju, Jue Wang 0001, Yebin Liu, Haoqian Wang, Qionghai Dai |
Comput. Graph. Forum | 2 |
| 2013 | Inverse image editing: recovering a semantic editing history from a before-and-after image pairabstractWe study the problem of inverse image editing , which recovers a semantically-meaningful editing history from a source image and an edited copy. Our approach supports a wide range of commonly-used editing operations such as cropping, object insertion and removal, linear and non-linear color transformations, and spatially-varying adjustment brushes. Given an input image pair, we first apply a dense correspondence method between them to match edited image regions with their sources. For each edited region, we determine geometric and semantic appearance operations that have been applied. Finally, we compute an optimal editing path from the region-level editing operations, based on predefined semantic constraints. The recovered history can be used in various applications such as image re-editing, edit transfer, and image revision control. A user study suggests that the editing histories generated from our system are semantically comparable to the ones generated by artists. Shi-Min Hu 0001, Kun Xu 0003, Li-Qian Ma, Bi-Ye Jiang, Jue Wang 0001 |
ACM Trans. Graph. | 6 |
| 2013 | PatchNet: a patch-based image representation for interactive library-driven image editingabstractWe introduce PatchNets , a compact, hierarchical representation describing structural and appearance characteristics of image regions, for use in image editing. In a PatchNet, an image region with coherent appearance is summarized by a graph node, associated with a single representative patch, while geometric relationships between different regions are encoded by labelled graph edges giving contextual information. The hierarchical structure of a PatchNet allows a coarse-to-fine description of the image. We show how this PatchNet representation can be used as a basis for interactive, library-driven, image editing. The user draws rough sketches to quickly specify editing constraints for the target image. The system then automatically queries an image library to find semantically-compatible candidate regions to meet the editing goal. Contextual image matching is performed using the PatchNet representation, allowing suitable regions to be found and applied in a few seconds, even from a library containing thousands of images. Shi-Min Hu 0001, Miao Wang 0004, Ralph R. Martin, Jue Wang 0001 |
ACM Trans. Graph. | 5 |
| 2013 | A no-reference metric for evaluating the quality of motion deblurringabstractMethods to undo the effects of motion blur are the subject of intense research, but evaluating and tuning these algorithms has traditionally required either user input or the availability of ground-truth images. We instead develop a metric for automatically predicting the perceptual quality of images produced by state-of-the-art deblurring algorithms. The metric is learned based on a massive user study, incorporates features that capture common deblurring artifacts, and does not require access to the original images (i.e., is "noreference"). We show that it better matches user-supplied rankings than previous approaches to measuring quality, and that in most cases it outperforms conventional full-reference image-similarity measures. We demonstrate applications of this metric to automatic selection of optimal algorithms and parameters, and to generation of fused images that combine multiple deblurring results. Yiming Liu 0001, Jue Wang 0001, Sunghyun Cho, Adam Finkelstein, Szymon Rusinkiewicz |
ACM Trans. Graph. | 2 |
| 2012 | Automatic upright adjustment of photographsabstractMan-made structures often appear to be distorted in photos captured by casual photographers, as the scene layout often conflicts with how it is expected by human perception. In this paper we propose an automatic approach for straightening up slanted man-made structures in an input image to improve its perceptual quality. We call this type of correction upright adjustment. We propose a set of criteria for upright adjustment based on human perception studies, and develop an optimization framework which yields an optimal homography for adjustment. We also develop a new optimization-based camera calibration method that performs favorably to previous methods and allows the proposed system to work reliably for a wide variety of images. The effectiveness of our system is demonstrated by both quantitative comparisons and qualitative user studies. Hyunjoon Lee, Eli Shechtman, Jue Wang 0001, Seungyong Lee 0001 |
CVPR | 3 |
| 2012 | A data-driven approach for facial expression synthesis in videoabstractThis paper presents a method to synthesize a realistic facial animation of a target person, driven by a facial performance video of another person. Different from traditional facial animation approaches, our system takes advantage of an existing facial performance database of the target person, and generates the final video by retrieving frames from the database that have similar expressions to the input ones. To achieve this we develop an expression similarity metric for accurately measuring the expression difference between two video frames. To enforce temporal coherence, our system employs a shortest path algorithm to choose the optimal image for each frame from a set of candidate frames determined by the similarity metric. Finally, our system adopts an expression mapping method to further minimize the expression difference between the input and retrieved frames. Experimental results show that our system can generate high quality facial animation using the proposed data-driven approach. Kai Li 0016, Feng Xu 0005, Jue Wang 0001, Qionghai Dai, Yebin Liu |
CVPR | 3 |
| 2012 | Facial expression editing in video using a temporally-smooth factorizationabstractWe address the problem of editing facial expression in video, such as exaggerating, attenuating or replacing the expression with a different one in some parts of the video. To achieve this we develop a tensor-based 3D face geometry reconstruction method, which fits a 3D model for each video frame, with the constraint that all models have the same identity and requiring temporal continuity of pose and expression. With the identity constraint, the differences between the underlying 3D shapes capture only changes in expression and pose. We show that various expression editing tasks in video can be achieved by combining face reordering with face warping, where the warp is induced by projecting differences in 3D face shapes into the image plane. Analogously, we show how the identity can be manipulated while fixing expression and pose. Experimental results show that our method can effectively edit expressions and identity in video in a temporally-coherent way with high fidelity. Fei Yang 0001, Lubomir D. Bourdev, Eli Shechtman, Jue Wang 0001, Dimitris N. Metaxas |
CVPR | 4 |
| 2012 | Text Image Deblurring Using Text-Specific Properties
Hojin Cho, Jue Wang 0001, Seungyong Lee 0001 |
ECCV (5) | 2 |
| 2012 | Face morphing using 3D-aware appearance optimization
Fei Yang 0001, Eli Shechtman, Jue Wang 0001, Lubomir D. Bourdev, Dimitris N. Metaxas |
Graphics Interface | 3 |
| 2012 | Region-based active surface modelling and alpha matting for unsupervised tumour segmentation in PETabstractThis paper presents a combination of existing advanced methods to solve the partial volume segmentation problem. It uses region-based active surface modelling in a hierarchical scheme to eliminate segmentation errors, followed by an alpha matting step to further refine the segmentation. This method can have an interest in several applications in medical imaging. We have validated our method on real PET images of head-and-neck cancer patients as well as custom designed phantom PET images. Experiments show that our method can generate more accurate segmentation than existing approaches. Jue Wang 0001, Tony Shepherd, Reyer Zwiggelaar |
ICIP | 2 |
| 2012 | Unsupervised Template Learning for Fine-Grained Object RecognitionabstractFine-grained recognition refers to a subordinate level of recognition, such are recognizing different species of birds, animals or plants. It differs from recognition of basic categories, such as humans, tables, and computers, in that there are global similarities in shape or structure shared within a category, and the differences are in the details of the object parts. We suggest that the key to identifying the fine-grained differences lies in finding the right alignment of image regions that contain the same object parts. We propose a template model for the purpose, which captures common shape patterns of object parts, as well as the co-occurence relation of the shape patterns. Once the image regions are aligned, extracted features are used for classification. Learning of the template model is efficient, and the recognition results we achieve significantly outperform the state-of-the-art algorithms. Shulin Yang, Liefeng Bo, Jue Wang 0001, Linda G. Shapiro |
NIPS | 3 |
| 2012 | Video deblurring for hand-held cameras using patch-based synthesisabstractVideos captured by hand-held Cameras often contain significant camera shake, causing many frames to be blurry. Restoring shaky videos not only requires smoothing the camera motion and stabilizing the content, but also demands removing blur from video frames. However, video blur is hard to remove using existing single or multiple image deblurring techniques, as the blur kernel is both spatially and temporally varying. This paper presents a video deblurring method that can effectively restore sharp frames from blurry ones caused by camera shake. Our method is built upon the observation that due to the nature of camera shake, not all video frames are equally blurry. The same object may appear sharp on some frames while blurry on others. Our method detects sharp regions in the video, and uses them to restore blurry regions of the same content in nearby frames. Our method also ensures that the deblurred frames are both spatially and temporally coherent using patch-based synthesis. Experimental results show that our method can effectively remove complex video blur under the presence of moving objects and other outliers, which cannot be achieved using previous deconvolution-based approaches. Sunghyun Cho, Jue Wang 0001, Seungyong Lee 0001 |
ACM Trans. Graph. | 2 |
| 2011 | The video mesh: A data structure for image-based three-dimensional video editingabstractThis paper introduces the video mesh, a data structure for representing video as 2.5D “paper cutouts.” The video mesh allows interactive editing of moving objects and modeling of depth, which enables 3D effects and post-exposure camera control. The video mesh sparsely encodes optical flow as well as depth, and handles occlusion using local layering and alpha mattes. Motion is described by a sparse set of points tracked over time. Each point also stores a depth value. The video mesh is a triangulation over this point set and per-pixel information is obtained by interpolation. The user rotoscopes occluding contours and we introduce an algorithm to cut the video mesh along them. Object boundaries are refined with per-pixel alpha values. The video mesh is at its core a set of texture mapped triangles, we leverage graphics hardware to enable interactive editing and rendering of a variety of effects. We demonstrate the effectiveness of our representation with special effects such as 3D viewpoint changes, object insertion, depth-of-field manipulation, and 2D to 3D video conversion. Jiawen Chen 0001, Sylvain Paris, Jue Wang 0001, Wojciech Matusik, Michael F. Cohen, Frédo Durand |
ICCP | 3 |
| 2011 | Modeling and removing spatially-varying optical blurabstractPhoto deblurring has been a major research topic in the past few years. So far, existing methods have focused on removing the blur due to camera shake and object motion. In this paper, we show that the optical system of the camera also generates significant blur, even with professional lenses. We introduce a method to estimate the blur kernel densely over the image and across multiple aperture and zoom settings. Our measures show that the blur kernel can have a non-negligible spread, even with top-of-the-line equipment, and that it varies nontrivially over this domain. In particular, the spatial variations are not radially symmetric and not even left-right symmetric. We develop and compare two models of the optical blur, each of them having its own advantages. We show that our models predict accurate blur kernels that can be used to restore photos. We demonstrate that we can produce images that are more uniformly sharp unlike those produced with spatially-invariant deblurring techniques. Eric Kee, Sylvain Paris, Simon Chen, Jue Wang 0001 |
ICCP | 4 |
| 2011 | Handling outliers in non-blind image deconvolutionabstractNon-blind deconvolution is a key component in image deblurring systems. Previous deconvolution methods assume a linear blur model where the blurred image is generated by a linear convolution of the latent image and the blur kernel. This assumption often does not hold in practice due to various types of outliers in the imaging process. Without proper outlier handling, previous methods may generate results with severe ringing artifacts even when the kernel is estimated accurately. In this paper we analyze a few common types of outliers that cause previous methods to fail, such as pixel saturation and non-Gaussian noise. We propose a novel blur model that explicitly takes these outliers into account, and build a robust non-blind deconvolution method upon it, which can effectively reduce the visual artifacts caused by outliers. The effectiveness of our method is demonstrated by experimental results on both synthetic and real-world examples. Sunghyun Cho, Jue Wang 0001, Seungyong Lee 0001 |
ICCV | 2 |
| 2011 | Pause-and-play: automatically linking screencast video tutorials with applicationsabstractVideo tutorials provide a convenient means for novices to learn new software applications. Unfortunately, staying in sync with a video while trying to use the target application at the same time requires users to repeatedly switch from the application to the video to pause or scrub backwards to replay missed steps. We present Pause-and-Play, a system that helps users work along with existing video tutorials. Pause-and-Play detects important events in the video and links them with corresponding events in the target application as the user tries to replicate the depicted procedure. This linking allows our system to automatically pause and play the video to stay in sync with the user. Pause-and-Play also supports convenient video navigation controls that are accessible from within the target application and allow the user to easily replay portions of the video without switching focus out of the application. Finally, since our system uses computer vision to detect events in existing videos and leverages application scripting APIs to obtain real time usage traces, our approach is largely independent of the specific target application and does not require access or modifications to application source code. We have implemented Pause-and-Play for two target applications, Google SketchUp and Adobe Photoshop, and we report on a user study that shows our system improves the user experience of working with video tutorials. Suporn Pongnumkul, Mira Dontcheva, Wilmot Li, Jue Wang 0001, Lubomir D. Bourdev, Shai Avidan, Michael F. Cohen |
UIST | 4 |
| 2011 | Subspace video stabilizationabstractWe present a robust and efficient approach to video stabilization that achieves high-quality camera motion for a wide range of videos. In this article, we focus on the problem of transforming a set of input 2D motion trajectories so that they are both smooth and resemble visually plausible views of the imaged scene; our key insight is that we can achieve this goal by enforcing subspace constraints on feature trajectories while smoothing them. Our approach assembles tracked features in the video into a trajectory matrix, factors it into two low-rank matrices, and performs filtering or curve fitting in a low-dimensional linear space. In order to process long videos, we propose a moving factorization that is both efficient and streamable. Our experiments confirm that our approach can efficiently provide stabilization results comparable with prior 3D methods in cases where those methods succeed, but also provides smooth camera motions in cases where such approaches often fail, such as videos that lack parallax. The presented approach offers the first method that both achieves high-quality video stabilization and is practical enough for consumer applications. Feng Liu 0015, Michael Gleicher, Jue Wang 0001, Hailin Jin, Aseem Agarwala |
ACM Trans. Graph. | 3 |
| 2011 | Expression flow for 3D-aware face component transferabstractWe address the problem of correcting an undesirable expression on a face photo by transferring local facial components, such as a smiling mouth, from another face photo of the same person which has the desired expression. Direct copying and blending using existing compositing tools results in semantically unnatural composites, since expression is a global effect and the local component in one expression is often incompatible with the shape and other components of the face in another expression. To solve this problem we present Expression Flow, a 2D flow field which can warp the target face globally in a natural way, so that the warped face is compatible with the new facial component to be copied over. To do this, starting with the two input face photos, we jointly construct a pair of 3D face shapes with the same identity but different expressions. The expression flow is computed by projecting the difference between the two 3D shapes back to 2D. It describes how to warp the target face photo to match the expression of the reference photo. User studies suggest that our system is able to generate face composites with much higher fidelity than existing methods. Fei Yang 0001, Jue Wang 0001, Eli Shechtman, Lubomir D. Bourdev, Dimitris N. Metaxas |
ACM Trans. Graph. | 2 |
| 2010 | Dynamic Color Flow: A Motion-Adaptive Color Model for Object Segmentation in Video
Jue Wang 0001, Guillermo Sapiro |
ECCV (5) | 2 |
| 2010 | Content-aware dynamic timeline for video browsingabstractWhen browsing a long video using a traditional timeline slider control, its effectiveness and precision degrade as a video's length grows. When browsing videos with more frames than pixels in the slider, aside from some frames being inaccessible, scrolling actions cause sudden jumps in a video's continuity as well as video frames to flash by too fast for one to assess the content. We propose a content-aware dynamic timeline control that is designed to overcome these limitations. Our timeline control decouples video speed and playback speed, and leverages video content analysis to allow salient shots to be presented at an intelligible speed. Our control also takes advantage of previous work on elastic sliders, which allows us to produce an accurate navigation control. Suporn Pongnumkul, Jue Wang 0001, Gonzalo A. Ramos, Michael F. Cohen |
UIST | 2 |
| 2009 | A perceptually motivated online benchmark for image mattingabstractThe availability of quantitative online benchmarks for low-level vision tasks such as stereo and optical flow has led to significant progress in the respective fields. This paper introduces such a benchmark for image matting. There are three key factors for a successful benchmarking system: (a) a challenging, high-quality ground truth test set; (b) an online evaluation repository that is dynamically updated with new results; (c) perceptually motivated error functions. Our new benchmark strives to meet all three criteria. We evaluated several matting methods with our benchmark and show that their performance varies depending on the error function. Also, our challenging test set reveals problems of existing algorithms, not reflected in previously reported results. We hope that our effort will lead to considerable progress in the field of image matting, and welcome the reader to visit our benchmark at www.aIphamatting.com. Christoph Rhemann, Carsten Rother, Jue Wang 0001, Margrit Gelautz, Pushmeet Kohli, Pamela Rott |
CVPR | 3 |
| 2009 | Video SnapCut: robust video object cutout using localized classifiersabstractAlthough tremendous success has been achieved for interactive object cutout in still images, accurately extracting dynamic objects in video remains a very challenging problem. Previous video cutout systems present two major limitations: (1) reliance on global statistics, thus lacking the ability to deal with complex and diverse scenes; and (2) treating segmentation as a global optimization, thus lacking a practical workflow that can guarantee the convergence of the systems to the desired results. We present Video SnapCut , a robust video object cutout system that significantly advances the state-of-the-art. In our system segmentation is achieved by the collaboration of a set of local classifiers, each adaptively integrating multiple local image features. We show how this segmentation paradigm naturally supports local user editing and propagates them across time. The object cutout system is completed with a novel coherent video matting technique. A comprehensive evaluation and comparison is presented, demonstrating the effectiveness of the proposed system at achieving high quality results, as well as the robustness of the system against various types of inputs. Jue Wang 0001, David Simons, Guillermo Sapiro |
ACM Trans. Graph. | 2 |
| 2009 | Noise brush: interactive high quality image-noise separationabstractThis paper proposes aninteractiveapproach usingjoint image-noise filteringfor achieving high quality image-noise separation. The core of the system is our novel joint image-noise filter which operates in both image and noise domain, and can effectively separate noise from both high and low frequency image structures. A novel user interface is introduced, which allows the user to interact with both the image and the noise layer, and apply the filter adaptively and locally to achieve optimal results. A comprehensive and quantitative evaluation shows that our interactive system can significantly improve the initial image-noise separation results. Our system can also be deployed in various noise-consistent image editing tasks, where preserving the noise characteristics inherent in the input image is a desired feature. Jia Chen 0026, Chi-Keung Tang, Jue Wang 0001 |
ACM Trans. Graph. | 3 |
| 2008 | Creating map-based storyboards for browsing tour videosabstractWatching a long unedited video is usually a boring experience. In this paper we examine a particular subset of videos, tour videos, in which the video is captured by walking about with a running camera with the goal of conveying the essence of some place. We present a system that makes the process of sharing and watching a long tour video easier, less boring, and more informative. To achieve this, we augment the tour video with a map-based storyboard, where the tour path is reconstructed, and coherent shots at different locations are directly visualized on the map. This allows the viewer to navigate the video in the joint location-time space. To create such a storyboard we employ an automatic pre-processing component to parse the video into coherent shots, and an authoring tool to enable the user to tie the shots with landmarks on the map. The browser-based viewing tool allows users to navigate the video in a variety of creative modes with a rich set of controls, giving each viewer a unique, personal viewing experience. Informal evaluation shows that our approach works well for tour videos compared with conventional media players. Suporn Pongnumkul, Jue Wang 0001, Michael F. Cohen |
UIST | 2 |
| 2007 | Optimized Color Sampling for Robust MattingabstractImage matting is the problem of determining for each pixel in an image whether it is foreground, background, or the mixing parameter, "alpha", for those pixels that are a mixture of foreground and background. Matting is inherently an ill-posed problem. Previous matting approaches either use naive color sampling methods to estimate foreground and background colors for unknown pixels, or use propagation-based methods to avoid color sampling under weak assumptions about image statistics. We argue that neither method itself is enough to generate good results for complex natural images. We analyze the weaknesses of previous matting approaches, and propose a new robust matting algorithm. In our approach we also sample foreground and background colors for unknown pixels, but more importantly, analyze the confidence of these samples. Only high confidence samples are chosen to contribute to the matting energy function which is minimized by a Random Walk. The energy function we define also contains a neighborhood term to enforce the smoothness of the matte. To validate the approach, we present an extensive and quantitative comparison between our algorithm and a number of previous approaches in hopes of providing a benchmark for future matting research. Jue Wang 0001, Michael F. Cohen |
CVPR | 1 |
| 2007 | Simultaneous Matting and CompositingabstractRecent work in matting, hole filling, and compositing allows image elements to be mixed in a new composite image. Previous algorithms for matting foreground elements have assumed that the new background for compositing is unknown. We show that, if the new background is known, the matting algorithm has more freedom to create a successful matte by simultaneously optimizing the matting and compositing operations. We propose a new algorithm, that integrates matting and compositing into a single optimization process. The system is able to compose foreground elements onto a new background more efficiently and with less artifacts compared with previous approaches. In our examples, we show how one can enlarge the foreground while maintaining the wide angle view of the background. We also demonstrate composing a foreground element on top of similar backgrounds to help remove unwanted portions of the background or to re-scale or re-arrange the composite. We compare and contrast our method with a number of previous matting and compositing systems. Jue Wang 0001, Michael F. Cohen |
CVPR | 1 |
| 2007 | Discriminative Gaussian Mixtures for Interactive Image SegmentationabstractRecently graph-cut optimization has been extensively explored for interactive image segmentation. In this paper we propose discriminative Gaussian mixtures (DGMs) to boost the performance of graph-cut-based segmentation. Given the user specified pixels, our algorithm analyzes their distributions in color, texture and spatial spaces and produces optimized Gaussian mixtures to set the data cost in the image graph, under the criteria of maximizing the discriminant power. We also show how to assemble novel training data to train DGMs for the link cost in the graph. Experimental results demonstrate that DGMs can noticeably improve the performance of graph-cut segmentation on texture-rich images. Jue Wang 0001 |
ICASSP (1) | 1 |
| 2007 | Soft scissors: an interactive tool for realtime high quality mattingabstractWe present Soft Scissors , an interactive tool for extracting alpha mattes of foreground objects in realtime. We recently proposed a novel offline matting algorithm capable of extracting high-quality mattes for complex foreground objects such as furry animals [Wang and Cohen 2007]. In this paper we both improve the quality of our offline algorithm and give it the ability to incrementally update the matte in an online interactive setting. Our realtime system efficiently estimates foreground color thereby allowing both the matte and the final composite to be revealed instantly as the user roughly paints along the edge of the foreground object. In addition, our system can dynamically adjust the width and boundary conditions of the scissoring paint brush to approximately capture the boundary of the foreground object that lies ahead on the scissor's path. These advantages in both speed and accuracy create the first interactive tool for high quality image matting and compositing. Jue Wang 0001, Maneesh Agrawala, Michael F. Cohen |
ACM Trans. Graph. | 1 |
| 2006 | The cartoon animation filterabstractWe present the "Cartoon Animation Filter", a simple filter that takes an arbitrary input motion signal and modulates it in such a way that the output motion is more "alive" or "animated". The filter adds a smoothed, inverted, and (sometimes) time shifted version of the second derivative (the acceleration) of the signal back into the original signal. Almost all parameters of the filter are automated. The user only needs to set the desired strength of the filter. The beauty of the animation filter lies in its simplicity and generality. We apply the filter to motions ranging from hand drawn trajectories, to simple animations within PowerPoint presentations, to motion captured DOF curves, to video segmentation results. Experimental results show that the filtered motion exhibits anticipation, follow-through, exaggeration and squash-and-stretch effects which are not present in the original input motion data. Jue Wang 0001, Steven Mark Drucker, Maneesh Agrawala, Michael F. Cohen |
ACM Trans. Graph. | 1 |
| 2005 | Very Low Frame-Rate Video Streaming For Face-to-Face TeleconferenceabstractProviding the best possible face-to-face video over low bandwidth networks is a major challenge for current teleconferencing systems especially when implementing multi-party conversations. With the limitations of fixed webcams and low bandwidth, the transmitted faces may be significantly blurred, even unrecognizable. We describe a novel strategy for real-time coding of face video at a very low frame rate such as one frame per 2-3 seconds. By reducing the frame rate, we transmit high spatial fidelity face video at a very low bit-rate of 8 Kbit/s, with a reasonable representation of the original video, especially for conversants in "listening" mode. One challenge overcome is selecting which frames to transmit. We provide experimental results that show that our compression strategy provides very low bit-rate face-to-face teleconferencing. Jue Wang 0001, Michael F. Cohen |
DCC | 1 |
| 2005 | An Iterative Optimization Approach for Unified Image Segmentation and MattingabstractSeparating a foreground object from the background in a static image involves determining both full and partial pixel coverages, also known as extracting a matte. Previous approaches require the input image to be presegmented into three regions: foreground, background and unknown, which are called a trimap. Partial opacity values are then computed only for pixels inside the unknown region. This presegmentation based approach fails for images with large portions of semitransparent foreground where the trimap is difficult to create even manually. In this paper, we combine the segmentation and matting problem together and propose a unified optimization approach based on belief propagation. We iteratively estimate the opacity value for every pixel in the image, based on a small sample of foreground and background pixels marked by the user. Experimental results show that compared with previous approaches, our method is more efficient to extract high quality mattes for foregrounds with significant semitransparent regions Jue Wang 0001, Michael F. Cohen |
ICCV | 1 |
| 2005 | Combining shape and physical modelsfor online cursive handwriting synthesis
Jue Wang 0001, Ying-Qing Xu, Harry Shum |
Int. J. Document Anal. Recognit. | 1 |
| 2005 | Interactive video cutoutabstractWe present an interactive system for efficiently extracting foreground objects from a video. We extend previous min-cut based image segmentation techniques to the domain of video with four new contributions. We provide a novel painting-based user interface that allows users to easily indicate the foreground object across space and time. We introduce a hierarchical mean-shift preprocess in order to minimize the number of nodes that min-cut must operate on. Within the min-cut we also define new local cost functions to augment the global costs defined in earlier work. Finally, we extend 2D alpha matting methods designed for images to work with 3D video volumes. We demonstrate that our matting approach preserves smoothness across both space and time. Our interactive video cutout system allows users to quickly extract foreground objects from video sequences for use in a variety of applications including compositing onto new backgrounds and NPR cartoon style rendering. Jue Wang 0001, Pravin Bhat, Alex Colburn, Maneesh Agrawala, Michael F. Cohen |
ACM Trans. Graph. | 1 |
| 2004 | Image and Video Segmentation by Anisotropic Kernel Mean Shift
Jue Wang 0001, Bo Thiesson, Ying-Qing Xu, Michael F. Cohen |
ECCV (2) | 1 |
| 2004 | Video tooningabstractWe describe a system for transforming an input video into a highly abstracted, spatio-temporally coherent cartoon animation with a range of styles. To achieve this, we treat video as a space-time volume of image data. We have developed an anisotropic kernel mean shift technique to segment the video data into contiguous volumes. These provide a simple cartoon style in themselves, but more importantly provide the capability to semi-automatically rotoscope semantically meaningful regions.In our system, the user simply outlines objects on keyframes. A mean shift guided interpolation algorithm is then employed to create three dimensional semantic regions by interpolation between the keyframes, while maintaining smooth trajectories along the time dimension. These regions provide the basis for creating smooth two dimensional edge sheets and stroke sheets embedded within the spatio-temporal video volume. The regions, edge sheets, and stroke sheets are rendered by slicing them at particular times. A variety of styles of rendering are shown. The temporal coherence provided by the smoothed semantic regions and sheets results in a temporally consistent non-photorealistic appearance. Jue Wang 0001, Ying-Qing Xu, Harry Shum, Michael F. Cohen |
ACM Trans. Graph. | 1 |