EDBT 2026 Demo / reviewers in the wild / expert
Fang Wen 0001
dblp:66/1922-1
· DBLP profile ↗
65ranked-venue papers
1as first author
22since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 1 first-author · 19 since 2021Artificial intelligence and machine learning · 55 · 21 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | 3DFaceShop: Explicitly Controllable 3D-Aware Portrait GenerationabstractIn contrast to the traditional avatar creation pipeline which is a costly process, contemporary generative approaches directly learn the data distribution from photographs. While plenty of works extend unconditional generative models and achieve some levels of controllability, it is still challenging to ensure multi-view consistency, especially in large poses. In this work, we propose a network that generates 3D-aware portraits while being controllable according to semantic parameters regarding pose, identity, expression and illumination. Our network uses neural scene representation to model 3D-aware portraits, whose generation is guided by a parametric face model that supports explicit control. While the latent disentanglement can be further enhanced by contrasting images with partially different attributes, there still exists noticeable inconsistency in non-face areas when animating expressions. We solve this by proposing a volume blending strategy in which we form a composite output by blending dynamic and static areas, with two parts segmented from the jointly learned semantic field. Our method outperforms prior arts in extensive experiments, producing realistic portraits with vivid expression in natural lighting when viewed from free viewpoints. It also demonstrates generalization ability to real images as well as out-of-domain data, showing great promise in real applications. Junshu Tang, Bo Zhang 0025, Binxin Yang, Ting Zhang 0002, Dong Chen 0003, Lizhuang Ma, Fang Wen 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2023 | PeCo: Perceptual Codebook for BERT Pre-training of Vision TransformersabstractThis paper explores a better prediction target for BERT pre-training of vision transformers. We observe that current prediction targets disagree with human perception judgment. This contradiction motivates us to learn a perceptual prediction target. We argue that perceptually similar images should stay close to each other in the prediction target space. We surprisingly find one simple yet effective idea: enforcing perceptual similarity during the dVAE training. Moreover, we adopt a self-supervised transformer model for deep feature extraction and show that it works well for calculating perceptual similarity. We demonstrate that such learned visual tokens indeed exhibit better semantic meanings, and help pre-training achieve superior transfer performance in various downstream tasks. For example, we achieve 84.5% Top-1 accuracy on ImageNet-1K with ViT-B backbone, outperforming the competitive method BEiT by +1.3% under the same pre-training epochs. Our approach also gets significant improvement on object detection and segmentation on COCO and semantic segmentation on ADE20K. Equipped with a larger backbone ViT-H, we achieve the state-of-the-art ImageNet accuracy (88.3%) among methods using only ImageNet-1K data. Xiaoyi Dong, Jianmin Bao, Ting Zhang 0002, Dongdong Chen 0001, Weiming Zhang 0001, Lu Yuan 0001, Dong Chen 0003, Fang Wen 0001, Nenghai Yu, Baining Guo |
AAAI | 8 |
| 2023 | MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image PretrainingabstractThis paper presents a simple yet effective framework MaskCLIP, which incorporates a newly proposed masked self-distillation into contrastive language-image pretraining. The core idea of masked self-distillation is to distill representation from a full image to the representation predicted from a masked image. Such incorporation enjoys two vital benefits. First, masked self-distillation targets local patch representation learning, which is complementary to vision-language contrastive focusing on text-related representation. Second, masked self-distillation is also consistent with vision-language contrastive from the perspective of training objective as both utilize the visual encoder for feature aligning, and thus is able to learn local semantics getting indirect supervision from the language. We provide specially designed experiments with a comprehensive analysis to validate the two benefits. Symmetrically, we also introduce the local semantic supervision into the text branch, which further improves the pretraining performance. With extensive experiments, we show that MaskCLIP, when applied to various challenging downstream tasks, achieves superior results in linear probing, finetuning, and zeroshot performance with the guidance of the language encoder. Code will be release at https://github.com/LightDXY/MaskCLIP. Xiaoyi Dong, Jianmin Bao, Yinglin Zheng, Ting Zhang 0002, Dongdong Chen 0001, Hao Yang 0036, Ming Zeng 0008, Weiming Zhang 0001, Lu Yuan 0001, Dong Chen 0003, Fang Wen 0001, Nenghai Yu |
CVPR | 11 |
| 2023 | RODIN: A Generative Model for Sculpting 3D Digital Avatars Using DiffusionabstractThis paper presents a 3D diffusion model that automatically generates 3D digital avatars represented as neural radiance fields (NeRFs). A significant challenge for 3D diffusion is that the memory and processing costs are prohibitive for producing high-quality results with rich details. To tackle this problem, we propose the roll-out diffusion network (RODIN), which takes a 3D NeRF model represented as multiple 2D feature maps and rolls out them onto a single 2D feature plane within which we perform 3D-aware diffusion. The RODIN model brings much-needed computational efficiency while preserving the integrity of 3D diffusion by using 3D-aware convolution that attends to projected features in the 2D plane according to their original relationships in 3D. We also use latent conditioning to orchestrate the feature generation with global coherence, leading to high-fidelity avatars and enabling semantic editing based on text prompts. Finally, we use hierarchical synthesis to further enhance details. The 3D avatars generated by our model compare favorably with those produced by existing techniques. We can generate highly detailed avatars with realistic hairstyles and facial hair. We also demonstrate 3D avatar generation from image or text, as well as text-guided editability. Tengfei Wang 0002, Bo Zhang 0025, Ting Zhang 0002, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen 0003, Fang Wen 0001, Qifeng Chen 0001, Baining Guo |
CVPR | 9 |
| 2023 | Paint by Example: Exemplar-based Image Editing with Diffusion ModelsabstractLanguage-guided image editing has achieved great success recently. In this paper, we investigate exemplar-guided image editing for more precise control. We achieve this goal by leveraging self-supervised training to disentangle and re-organize the source image and the exemplar. However, the naive approach will cause obvious fusing artifacts. We carefully analyze it and propose a content bottleneck and strong augmentations to avoid the trivial solution of directly copying and pasting the exemplar image. Meanwhile, to ensure the controllability of the editing process, we design an arbitrary shape mask for the exemplar image and leverage the classifier-free guidance to increase the similarity to the exemplar image. The whole framework involves a single forward of the diffusion model without any iterative optimization. We demonstrate that our method achieves an impressive performance and enables controllable editing on in-the-wild images with high fidelity. The code and pretrained models are available at https://github.com/Fantasy-Studio/Paint-by-Example. Binxin Yang, Shuyang Gu, Bo Zhang 0025, Ting Zhang 0002, Xuejin Chen, Xiaoyan Sun 0001, Dong Chen 0003, Fang Wen 0001 |
CVPR | 8 |
| 2023 | MetaPortrait: Identity-Preserving Talking Head Generation with Fast Personalized AdaptationabstractIn this work, we propose an ID-preserving talking head generation framework, which advances previous methods in two aspects. First, as opposed to interpolating from sparse flow, we claim that dense landmarks are crucial to achieving accurate geometry-aware flow fields. Second, inspired by face-swapping methods, we adaptively fuse the source identity during synthesis, so that the network better preserves the key characteristics of the image portrait. Although the proposed model surpasses prior generation fidelity on established benchmarks, personalized fine-tuning is still needed to further make the talking head generation qualified for real usage. However, this process is rather computationally demanding that is unaffordable to standard users. To alleviate this, we propose a fast adaptation model using a metalearning approach. The learned model can be adapted to a high-quality personalized model as fast as 30 seconds. Last but not least, a spatial-temporal enhancement module is proposed to improve the fine details while ensuring temporal coherency. Extensive experiments prove the significant superiority of our approach over the state of the arts in both one-shot and personalized settings. Bowen Zhang 0010, Pan Zhang 0003, Bo Zhang 0025, HsiangTao Wu, Dong Chen 0003, Qifeng Chen 0001, Fang Wen 0001 |
CVPR | 9 |
| 2023 | X-Paste: Revisiting Scalable Copy-Paste for Instance Segmentation using CLIP and StableDiffusionabstractCopy-Paste is a simple and effective data augmentation strategy for instance segmentation. By randomly pasting object instances onto new background images, it creates new training data for free and significantly boosts the segmentation performance, especially for rare object categories. Although diverse, high-quality object instances used in Copy-Paste result in more performance gain, previous works utilize object instances either from human-annotated instance segmentation datasets or rendered from 3D object models, and both approaches are too expensive to scale up to obtain good diversity. In this paper, we revisit Copy-Paste at scale with the power of newly emerged zero-shot recognition models (e.g., CLIP) and text2image models (e.g., StableDiffusion). We demonstrate for the first time that using a text2image model to generate images or zero-shot recognition model to filter noisily crawled images for different object categories is a feasible way to make Copy-Paste truly scalable. To make such success happen, we design a data acquisition and processing framework, dubbed ``X-Paste", upon which a systematic study is conducted. On the LVIS dataset, X-Paste provides impressive improvements over the strong baseline CenterNet2 with Swin-L as the backbone. Specifically, it archives +2.6 box AP and +2.1 mask AP gains on all classes and even more significant gains with +6.8 box AP +6.5 mask AP on long-tail classes. Dianmo Sheng, Jianmin Bao, Dongdong Chen 0001, Dong Chen 0003, Fang Wen 0001, Lu Yuan 0001, Ce Liu 0001, Wenbo Zhou 0004, Qi Chu 0001, Weiming Zhang 0001, Nenghai Yu |
ICML | 6 |
| 2023 | Old Photo Restoration via Deep Latent Space TranslationabstractWe propose to restore old photos that suffer from severe degradation through a deep learning approach. Unlike conventional restoration tasks that can be solved through supervised learning, the degradation in real photos is complex and the domain gap between synthetic images and real old photos makes the network fail to generalize. Therefore, we propose a novel triplet domain translation network by leveraging real photos along with massive synthetic image pairs. Specifically, we train two variational autoencoders (VAEs) to respectively transform old photos and clean photos into two latent spaces. And the translation between these two latent spaces is learned with synthetic paired data. This translation generalizes well to real photos because the domain gap is closed in the compact latent space. Besides, to address multiple degradations mixed in one old photo, we design a global branch with a partial nonlocal block targeting the structured defects, such as scratches and dust spots, and a local branch targeting the unstructured defects, such as noises and blurriness. We also extend the global branch with a more memory-efficient scheme, named multi-scale patch-based attention to processing high-resolution photos. Two branches are fused in the latent space, leading to improved capability to restore old photos from multiple defects. Furthermore, we apply another face refinement network to recover fine details of faces in the old photos, thus ultimately generating photos with enhanced perceptual quality. With comprehensive experiments, the proposed pipeline demonstrates superior performance over state-of-the-art methods as well as existing commercial tools in terms of visual quality for old photos restoration. Both code and models could be found at https://github.com/microsoft/Bringing-Old-Photos-Back-to-Life. Ziyu Wan, Bo Zhang 0025, Dongdong Chen 0001, Pan Zhang 0003, Dong Chen 0003, Fang Wen 0001, Jing Liao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Protecting Celebrities from DeepFake with Identity Consistency TransformerabstractIn this work we propose Identity Consistency Transformer, a novel face forgery detection method that focuses on high-level semantics, specifically identity information, and detecting a suspect face by finding identity inconsistency in inner and outer face regions. The Identity Consistency Transformer incorporates a consistency loss for identity consistency determination. We show that Identity Consistency Transformer exhibits superior generalization ability not only across different datasets but also across various types of image degradation forms found in real-world applications including deepfake videos. The Identity Consistency Transformer can be easily enhanced with additional identity information when such information is available, and for this reason it is especially well-suited for detecting face forgeries involving celebrities.11Code will be released at https://github.com/LightDXY/ICT_DeepFake Xiaoyi Dong, Jianmin Bao, Dongdong Chen 0001, Ting Zhang 0002, Weiming Zhang 0001, Nenghai Yu, Dong Chen 0003, Fang Wen 0001, Baining Guo |
CVPR | 8 |
| 2022 | Large-Scale Pre-training for Person Re-identification with Noisy LabelsabstractThis paper aims to address the problem of pretraining for person re-identification (Re-ID) with noisy labels. To setup the pretraining task, we apply a simple online multi-object tracking system on raw videos of an existing un-labeled Re-ID dataset “LUPerson” and build the Noisy Labeled variant called “LUPerson-NL”. Since theses ID labels automatically derived from tracklets inevitably con-tain noises, we develop a large-scale Pre-training frame-work utilizing Noisy Labels (PNL), which consists of three learning modules: supervised Re-ID learning, prototype-based contrastive learning, and label-guided contrastive learning. In principle, joint learning of these three mod-ules not only clusters similar examples to one prototype, but also rectifies noisy labels based on the prototype as-signment. We demonstrate that learning directly from raw videos is a promising alternative for pre-training, which utilizes spatial and temporal correlations as weak super-vision. This simple pre-training task provides a scalable way to learn SOTA Re-ID representations from scratch on “LUPerson-NL” without bells and whistles. For example, by applying on the same supervised Re-ID method MGN, our pre-trained model improves the mAP over the unsu-pervised pre-training counterpart by 5.7%, 2.2%, 2.3% on CUHK03, DukeMTMC, and MSMT17 respectively. Under the small-scale or few-shot setting, the performance gain is even more significant, suggesting a better transferability of the learned representation. Code is available at https://github.com/DengpanFu/LUPerson-NL. Dengpan Fu, Dongdong Chen 0001, Hao Yang 0036, Jianmin Bao, Lu Yuan 0001, Lei Zhang 0001, Houqiang Li, Fang Wen 0001, Dong Chen 0003 |
CVPR | 8 |
| 2022 | Vector Quantized Diffusion Model for Text-to-Image SynthesisabstractWe present the vector quantized diffusion (VQ-Diffusion) model for text-to-image generation. This method is based on a vector quantized variational autoencoder (VQ-VAE) whose latent space is modeled by a conditional variant of the recently developed Denoising Diffusion Probabilistic Model (DDPM). We find that this latent-space method is well-suited for text-to-image generation tasks because it not only eliminates the unidirectional bias with existing methods but also allows us to incorporate a mask-and-replace diffusion strategy to avoid the accumulation of errors, which is a serious problem with existing methods. Our experiments show that the VQ-Diffusion produces significantly better text-to-image generation results when compared with conventional autoregressive (AR) models with similar numbers of parameters. Compared with previous GAN-based text-to-image methods, our VQ-Diffusion can handle more complex scenes and improve the synthesized image quality by a large margin. Finally, we show that the image generation computation in our method can be made highly efficient by reparameterization. With traditional AR methods, the text-to-image generation time increases linearly with the output image resolution and hence is quite time consuming even for normal size images. The VQ-Diffusion allows us to achieve a better trade-off between quality and speed. Our experiments indicate that the VQ-Diffusion model with the reparameterization is fifteen times faster than traditional AR methods while achieving a better image quality. The code and models are available at https://github.com/cientgu/VQ-Diffusion. Shuyang Gu, Dong Chen 0003, Jianmin Bao, Fang Wen 0001, Bo Zhang 0025, Dongdong Chen 0001, Lu Yuan 0001, Baining Guo |
CVPR | 4 |
| 2022 | StyleSwin: Transformer-based GAN for High-resolution Image GenerationabstractDespite the tantalizing success in a broad of vision tasks, transformers have not yet demonstrated on-par ability as ConvNets in high-resolution image generative modeling. In this paper, we seek to explore using pure transformers to build a generative adversarial network for high-resolution image synthesis. To this end, we believe that local attention is crucial to strike the balance between computational efficiency and modeling capacity. Hence, the proposed generator adopts Swin transformer in a style-based architecture. To achieve a larger receptive field, we propose double attention which simultaneously leverages the context of the local and the shifted windows, leading to improved generation quality. Moreover, we show that offering the knowledge of the absolute position that has been lost in window-based transformers greatly benefits the generation quality. The proposed StyleSwin is scalable to high resolutions, with both the coarse geometry and fine structures benefit from the strong expressivity of transformers. However, blocking artifacts occur during high-resolution synthesis because performing the local attention in a block-wise manner may break the spatial coherency. To solve this, we empirically investigate various solutions, among which we find that employing a wavelet discriminator to examine the spectral discrepancy effectively suppresses the artifacts. Extensive experiments show the superiority over prior transformer-based GANs, especially on high resolutions, e.g.,$1024 \times$1024. The StyleSwin, without complex training strategies, excels over StyleGAN on CelebA-HQ 1024, and achieves on-par performance on FFHQ-1024, proving the promise of using transformers for high-resolution image generation. The code and pretrained models are available at https://github.com/microsoft/StyleSwin. Bowen Zhang 0010, Shuyang Gu, Bo Zhang 0025, Jianmin Bao, Dong Chen 0003, Fang Wen 0001, Baining Guo |
CVPR | 6 |
| 2022 | General Facial Representation Learning in a Visual-Linguistic MannerabstractHow to learn a universal facial representation that boosts all face analysis tasks? This paper takes one step toward this goal. In this paper, we study the transfer performance of pre-trained models on face analysis tasks and introduce a framework, called FaRL, for general facial representation learning. On one hand, the framework involves a contrastive loss to learn high-level semantic meaning from image-text pairs. On the other hand, we propose exploring low-level information simultaneously to further enhance the face representation by adding a masked image modeling. We perform pre-training on LAION-FACE, a dataset containing a large amount of face image-text pairs, and evaluate the representation capability on multiple downstream tasks. We show that FaRL achieves better transfer performance compared with previous pre-trained models. We also verify its superiority in the low-data regime. More importantly, our model surpasses the state-of-the-art methods on face analysis tasks including face parsing and face alignment. Yinglin Zheng, Hao Yang 0036, Ting Zhang 0002, Jianmin Bao, Dongdong Chen 0001, Yangyu Huang, Lu Yuan 0001, Dong Chen 0003, Ming Zeng 0008, Fang Wen 0001 |
CVPR | 10 |
| 2022 | Bootstrapped Masked Autoencoders for Vision BERT Pretraining
Xiaoyi Dong, Jianmin Bao, Ting Zhang 0002, Dongdong Chen 0001, Weiming Zhang 0001, Lu Yuan 0001, Dong Chen 0003, Fang Wen 0001, Nenghai Yu |
ECCV (30) | 8 |
| 2022 | Real-Time Neural Character Rendering with Pose-Guided Multiplane Images
Hao Ouyang, Bo Zhang 0025, Pan Zhang 0003, Hao Yang 0036, Jiaolong Yang, Dong Chen 0003, Qifeng Chen 0001, Fang Wen 0001 |
ECCV (32) | 8 |
| 2022 | Group Sampling for Scale Invariant Face DetectionabstractDetectors based on deep learning tend to detect multi-scale objects on a single input image for efficiency. Recent works, such as FPN and SSD, generally use feature maps from multiple layers with different spatial resolutions to detect objects at different scales, e.g., high-resolution feature maps for small objects. However, we find that objects at all scales can also be well detected with features from a single layer of the network. In this paper, we carefully examine the factors affecting detection performance across a large range of scales, and conclude that the balance of training samples, including both positive and negative ones, at different scales is the key. We propose a group sampling method which divides the anchors into several groups according to the scale, and ensure that the number of samples for each group is the same during training. Our approach using only one single layer of FPN as features is able to advance the state-of-the-arts. Comprehensive analysis and extensive experiments have been conducted to show the effectiveness of the proposed method. Moreover, we show that our approach is favorably applicable to other tasks, such as object detection on COCO dataset, and to other detection pipelines, such as YOLOv3, SSD and R-FCN. Our approach, evaluated on face detection benchmarks including FDDB and WIDER FACE datasets, achieves state-of-the-art results without bells and whistles. Xiang Ming, Fangyun Wei, Ting Zhang 0002, Dong Chen 0003, Nanning Zheng 0001, Fang Wen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | High-Fidelity and Arbitrary Face EditingabstractCycle consistency is widely used for face editing. However, we observe that the generator tends to find a tricky way to hide information from the original image to satisfy the constraint of cycle consistency, making it impossible to maintain the rich details (e.g., wrinkles and moles) of non-editing areas. In this work, we propose a simple yet effective method named HifaFace to address the above-mentioned problem from two perspectives. First, we relieve the pressure of the generator to synthesize rich details by directly feeding the high-frequency information of the input image into the end of the generator. Second, we adopt an additional discriminator to encourage the generator to synthesize rich details. Specifically, we apply wavelet transformation to transform the image into multi-frequency domains, among which the high-frequency parts can be used to recover the rich details. We also notice that a fine-grained and wider-range control for the attribute is of great importance for face editing. To achieve this goal, we propose a novel attribute regression loss. Powered by the proposed framework, we achieve high-fidelity and arbitrary face editing, outperforming other state-of-the-art approaches. Yue Gao 0006, Fangyun Wei, Jianmin Bao, Shuyang Gu, Dong Chen 0003, Fang Wen 0001, Zhouhui Lian |
CVPR | 6 |
| 2021 | Style-Based Point Generator With Adversarial Rendering for Point Cloud CompletionabstractIn this paper, we proposed a novel Style-based Point Generator with Adversarial Rendering (SpareNet) for point cloud completion. Firstly, we present the channel-attentive EdgeConv to fully exploit the local structures as well as the global shape in point features. Secondly, we observe that the concatenation manner used by vanilla foldings limits its potential of generating a complex and faithful shape. Enlightened by the success of StyleGAN, we regard the shape feature as style code that modulates the normalization layers during the folding, which considerably enhances its capability. Thirdly, we realize that existing point supervisions, e.g., Chamfer Distance or Earth Mover’s Distance, cannot faithfully reflect the perceptual quality of the reconstructed points. To address this, we propose to project the completed points to depth maps with a differentiable renderer and apply adversarial training to advocate the perceptual realism under different viewpoints. Comprehensive experiments on ShapeNet and KITTI prove the effectiveness of our method, which achieves state-of-the-art quantitative performance while offering superior visual quality. Chulin Xie, Chuxin Wang, Bo Zhang 0025, Hao Yang 0036, Dong Chen 0003, Fang Wen 0001 |
CVPR | 6 |
| 2021 | Prototypical Pseudo Label Denoising and Target Structure Learning for Domain Adaptive Semantic SegmentationabstractSelf-training is a competitive approach in domain adaptive segmentation, which trains the network with the pseudo labels on the target domain. However inevitably, the pseudo labels are noisy and the target features are dispersed due to the discrepancy between source and target domains. In this paper, we rely on representative prototypes, the feature centroids of classes, to address the two issues for unsupervised domain adaptation. In particular, we take one step further and exploit the feature distances from prototypes that provide richer information than mere prototypes. Specifically, we use it to estimate the likelihood of pseudo labels to facilitate online correction in the course of training. Meanwhile, we align the prototypical assignments based on relative feature distances for two different views of the same target, producing a more compact target feature space. Moreover, we find that distilling the already learned knowledge to a self-supervised pretrained model further boosts the performance. Our method shows tremendous performance advantage over state-of-the-art methods. The code is available at https://github.com/microsoft/ProDA. Pan Zhang 0003, Bo Zhang 0025, Ting Zhang 0002, Dong Chen 0003, Fang Wen 0001 |
CVPR | 6 |
| 2021 | CoCosNet v2: Full-Resolution Correspondence Learning for Image TranslationabstractWe present the full-resolution correspondence learning for cross-domain images, which aids image translation. We adopt a hierarchical strategy that uses the correspondence from coarse level to guide the fine levels. At each hierarchy, the correspondence can be efficiently computed via PatchMatch that iteratively leverages the matchings from the neighborhood. Within each PatchMatch iteration, the ConvGRU module is employed to refine the current correspondence considering not only the matchings of larger context but also the historic estimates. The proposed Co-CosNet v2, a GRU-assisted PatchMatch approach, is fully differentiable and highly efficient. When jointly trained with image translation, full-resolution semantic correspondence can be established in an unsupervised manner, which in turn facilitates the exemplar-based image translation. Experiments on diverse translation tasks show that CoCosNet v2 performs considerably better than state-of-the-art literature on producing high-resolution images. Xingran Zhou, Bo Zhang 0025, Ting Zhang 0002, Pan Zhang 0003, Jianmin Bao, Dong Chen 0003, Zhongfei Zhang, Fang Wen 0001 |
CVPR | 8 |
| 2021 | Dual Path Learning for Domain Adaptation of Semantic SegmentationabstractDomain adaptation for semantic segmentation enables to alleviate the need for large-scale pixel-wise annotations. Recently, self-supervised learning (SSL) with a combination of image-to-image translation shows great effectiveness in adaptive segmentation. The most common practice is to perform SSL along with image translation to well align a single domain (the source or target). However, in this single-domain paradigm, unavoidable visual inconsistency raised by image translation may affect subsequent learning. In this paper, based on the observation that domain adaptation frameworks performed in the source and target domain are almost complementary in terms of image translation and SSL, we propose a novel dual path learning (DPL) framework to alleviate visual inconsistency. Concretely, DPL contains two complementary and interactive single-domain adaptation pipelines aligned in source and target domain respectively. The inference of DPL is extremely simple, only one segmentation model in the target domain is employed. Novel technologies such as dual path image translation and dual path adaptive segmentation are proposed to make two paths promote each other in an interactive manner. Experiments on GTA5→Cityscapes and SYNTHIA→Cityscapes scenarios demonstrate the superiority of our DPL model over the state-of-the-art methods. The code and models are available at: https://github.com/royee182/DPL. Yiting Cheng 0001, Fangyun Wei, Jianmin Bao, Dong Chen 0003, Fang Wen 0001 |
ICCV | 5 |
| 2021 | Exploring Temporal Coherence for More General Video Face Forgery DetectionabstractAlthough current face manipulation techniques achieve impressive performance regarding quality and controllability, they are struggling to generate temporal coherent face videos. In this work, we explore to take full advantage of the temporal coherence for video face forgery detection. To achieve this, we propose a novel end-to-end framework, which consists of two major stages. The first stage is a fully temporal convolution network (FTCN). The key insight of FTCN is to reduce the spatial convolution kernel size to 1, while maintaining the temporal convolution kernel size un-changed. We surprisingly find this special design can benefit the model for extracting the temporal features as well as improve the generalization capability. The second stage is a Temporal Transformer network, which aims to explore the long-term temporal coherence. The proposed frame-work is general and flexible, which can be directly trained from scratch without any pre-training models or external datasets. Extensive experiments show that our framework outperforms existing methods and remains effective when applied to detect new sorts of face forgery videos. Yinglin Zheng, Jianmin Bao, Dong Chen 0003, Ming Zeng 0008, Fang Wen 0001 |
ICCV | 5 |
| 2020 | Disentangled and Controllable Face Image Generation via 3D Imitative-Contrastive LearningabstractWe propose an approach for face image generation of virtual people with disentangled, precisely-controllable latent representations for identity of non-existing people, expression, pose, and illumination. We embed 3D priors into adversarial learning and train the network to imitate the image formation of an analytic 3D face deformation and rendering process. To deal with the generation freedom induced by the domain gap between real and rendered faces, we further introduce contrastive learning to promote disentanglement by comparing pairs of generated images. Experiments show that through our imitative-contrastive learning, the factor variations are very well disentangled and the properties of a generated face can be precisely controlled. We also analyze the learned latent space and present several meaningful properties supporting factor disentanglement. Our method can also be used to embed real images into the disentangled latent space. We hope our method could provide new understandings of the relationship between physical properties and deep image synthesis. Yu Deng 0006, Jiaolong Yang, Dong Chen 0003, Fang Wen 0001, Xin Tong 0001 |
CVPR | 4 |
| 2020 | Advancing High Fidelity Identity Swapping for Forgery DetectionabstractIn this work, we study various existing benchmarks for deepfake detection researches. In particular, we examine a novel two-stage face swapping algorithm, called FaceShifter, for high fidelity and occlusion aware face swapping. Unlike many existing face swapping works that leverage only limited information from the target image when synthesizing the swapped face, FaceShifter generates the swapped face with high-fidelity by exploiting and integrating the target attributes thoroughly and adaptively. FaceShifter can handle facial occlusions with a second synthesis stage consisting of a Heuristic Error Acknowledging Refinement Network (HEAR-Net), which is trained to recover anomaly regions in a self-supervised way without any manual annotations. Experiments show that existing deepfake detection algorithm performs poorly with FaceShifter, since it achieves advantageous quality over all existing benchmarks. However, our newly developed Face X-Ray method can reliably detect forged images created by FaceShifter. Lingzhi Li 0002, Jianmin Bao, Hao Yang 0036, Dong Chen 0003, Fang Wen 0001 |
CVPR | 5 |
| 2020 | Face X-Ray for More General Face Forgery DetectionabstractIn this paper we propose a novel image representation called face X-ray for detecting forgery in face images. The face X-ray of an input face image is a greyscale image that reveals whether the input image can be decomposed into the blending of two images from different sources. It does so by showing the blending boundary for a forged image and the absence of blending for a real image. We observe that most existing face manipulation methods share a common step: blending the altered face into an existing background image. For this reason, face X-ray provides an effective way for detecting forgery generated by most existing face manipulation algorithms. Face X-ray is general in the sense that it only assumes the existence of a blending step and does not rely on any knowledge of the artifacts associated with a specific face manipulation technique. Indeed, the algorithm for computing face X-ray can be trained without fake images generated by any of the state-of-the-art face manipulation methods. Extensive experiments show that face X-ray remains effective when applied to forgery generated by unseen face manipulation techniques, while most existing face forgery detection or deepfake detection algorithms experience a significant performance drop. Lingzhi Li 0002, Jianmin Bao, Ting Zhang 0002, Hao Yang 0036, Dong Chen 0003, Fang Wen 0001, Baining Guo |
CVPR | 6 |
| 2020 | Bringing Old Photos Back to LifeabstractWe propose to restore old photos that suffer from severe degradation through a deep learning approach. Unlike conventional restoration tasks that can be solved through supervised learning, the degradation in real photos is complex and the domain gap between synthetic images and real old photos makes the network fail to generalize. Therefore, we propose a novel triplet domain translation network by leveraging real photos along with massive synthetic image pairs. Specifically, we train two variational autoencoders (VAEs) to respectively transform old photos and clean photos into two latent spaces. And the translation between these two latent spaces is learned with synthetic paired data. This translation generalizes well to real photos because the domain gap is closed in the compact latent space. Besides, to address multiple degradations mixed in one old photo, we design a global branch with a partial nonlocal block targeting to the structured defects, such as scratches and dust spots, and a local branch targeting to the unstructured defects, such as noises and blurriness. Two branches are fused in the latent space, leading to improved capability to restore old photos from multiple defects. The proposed method outperforms state-of-the-art methods in terms of visual quality for old photos restoration. Ziyu Wan, Bo Zhang 0025, Dongdong Chen 0001, Pan Zhang 0003, Dong Chen 0003, Jing Liao 0001, Fang Wen 0001 |
CVPR | 7 |
| 2020 | Deep 3D Portrait From a Single ImageabstractIn this paper, we present a learning-based approach for recovering the 3D geometry of human head from a single portrait image. Our method is learned in an unsupervised manner without any ground-truth 3D data. We represent the head geometry with a parametric 3D face model together with a depth map for other head regions including hair and ear. A two-step geometry learning scheme is proposed to learn 3D head reconstruction from in-the-wild face images, where we first learn face shape on single images using self-reconstruction and then learn hair and ear geometry using pairs of images in a stereo-matching fashion. The second step is based on the output of the first to not only improve the accuracy but also ensure the consistency of overall head geometry. We evaluate the accuracy of our method both in 3D and with pose manipulation tasks on 2D images. We alter pose based on the recovered geometry and apply a refinement network trained with adversarial learning to ameliorate the reprojected images and translate them to the real image domain. Extensive evaluations and comparison with previous methods show that our new method can produce high-fidelity 3D head geometry and head pose manipulation results. Sicheng Xu, Jiaolong Yang, Dong Chen 0003, Fang Wen 0001, Yu Deng 0006, Yunde Jia, Xin Tong 0001 |
CVPR | 4 |
| 2020 | Cross-Domain Correspondence Learning for Exemplar-Based Image TranslationabstractWe present a general framework for exemplar-based image translation, which synthesizes a photo-realistic image from the input in a distinct domain (e.g., semantic segmentation mask, or edge map, or pose keypoints), given an exemplar image. The output has the style (e.g., color, texture) in consistency with the semantically corresponding objects in the exemplar. We propose to jointly learn the cross-domain correspondence and the image translation, where both tasks facilitate each other and thus can be learned with weak supervision. The images from distinct domains are first aligned to an intermediate domain where dense correspondence is established. Then, the network synthesizes images based on the appearance of semantically corresponding patches in the exemplar. We demonstrate the effectiveness of our approach in several image translation tasks. Our method is superior to state-of-the-art methods in terms of image quality significantly, with the image style faithful to the exemplar with semantic consistency. Moreover, we show the utility of our method for several applications. Pan Zhang 0003, Bo Zhang 0025, Dong Chen 0003, Lu Yuan 0001, Fang Wen 0001 |
CVPR | 5 |
| 2020 | GIQA: Generated Image Quality Assessment
Shuyang Gu, Jianmin Bao, Dong Chen 0003, Fang Wen 0001 |
ECCV (11) | 4 |
| 2019 | Mask-Guided Portrait Editing With Conditional GANsabstractPortrait editing is a popular subject in photo manipulation.The Generative Adversarial Network (GAN) advances the generating of realistic faces and allows more face editing. In this paper, we argue about three issues in existing techniques: diversity, quality, and controllability for portrait synthesis and editing. To address these issues, we propose a novel end-to-end learning framework that leverages conditional GANs guided by provided face masks for generating faces. The framework learns feature embeddings for every face component (e.g., mouth, hair, eye), separately, contributing to better correspondences for image translation, and local face editing. With the mask, our network is available to many applications, like face synthesis driven by mask, face Swap+ (including hair in swapping), and local manipulation. It can also boost the performance of face parsing a bit as an option of data augmentation. Shuyang Gu, Jianmin Bao, Hao Yang 0036, Dong Chen 0003, Fang Wen 0001, Lu Yuan 0001 |
CVPR | 5 |
| 2019 | Face Parsing With RoI Tanh-WarpingabstractFace parsing computes pixel-wise label maps for different semantic components (e.g., hair, mouth, eyes) from face images. Existing face parsing literature have illustrated significant advantages by focusing on individual regions of interest (RoIs) for faces and facial components. However,the traditional crop-and-resize focusing mechanism ignores all contextual area outside the RoIs, and thus is not suitable when the component area is unpredictable, e.g. hair. Inspired by the physiological vision system of human, we propose a novel RoI Tanh-warping operator that combines the central vision and the peripheral vision together. It addresses the dilemma between a limited sized RoI for focusing and an unpredictable area of surrounding context for peripheral information. To this end, we propose a novel hybrid convolutional neural network for face parsing. It uses hierarchical local based method for inner facial components and global methods for outer facial components. The whole framework is simple and principled, and can be trained end-to-end. To facilitate future research of face parsing, we also manually relabel the training data of the HELEN dataset and will make it public. Experiments on both HELEN and LFW-PL benchmarks demonstrate that our method surpasses state-of-the-art methods. Jinpeng Lin, Hao Yang 0036, Dong Chen 0003, Ming Zeng 0008, Fang Wen 0001, Lu Yuan 0001 |
CVPR | 5 |
| 2019 | Group Sampling for Scale Invariant Face DetectionabstractDetectors based on deep learning tend to detect multi-scale faces on a single input image for efficiency. Recent works, such as FPN and SSD, generally use feature maps from multiple layers with different spatial resolutions to detect objects at different scales, e.g., high-resolution feature maps for small objects. However, we find that such multi-layer prediction is not necessary. Faces at all scales can be well detected with features from a single layer of the network. In this paper, we carefully examine the factors affecting face detection across a large range of scales, and conclude that the balance of training samples, including both positive and negative ones, at different scales is the key. We propose a group sampling method which divides the anchors into several groups according to the scale, and ensure that the number of samples for each group is the same during training. Our approach using only the last layer of FPN as features is able to advance the state-of-the-arts. Comprehensive analysis and extensive experiments have been conducted to show the effectiveness of the proposed method. Our approach, evaluated on face detection benchmarks including FDDB and WIDER FACE datasets, achieves state-of-the-art results without bells and whistles. Xiang Ming, Fangyun Wei, Ting Zhang 0002, Dong Chen 0003, Fang Wen 0001 |
CVPR | 5 |
| 2018 | Towards Open-Set Identity Preserving Face SynthesisabstractWe propose a framework based on Generative Adversarial Networks to disentangle the identity and attributes of faces, such that we can conveniently recombine different identities and attributes for identity preserving face synthesis in open domains. Previous identity preserving face synthesis processes are largely confined to synthesizing faces with known identities that are already in the training dataset. To synthesize a face with identity outside the training dataset, our framework requires one input image of that subject to produce an identity vector, and any other input face image to extract an attribute vector capturing, e.g., pose, emotion, illumination, and even the background. We then recombine the identity vector and the attribute vector to synthesize a new face of the subject with the extracted attribute. Our proposed framework does not need to annotate the attributes of faces in any way. It is trained with an asymmetric loss function to better preserve the identity and stabilize the training process. It can also effectively leverage large amounts of unlabeled training face images to further improve the fidelity of the synthesized faces for subjects that are not presented in the labeled training face dataset. Our experiments demonstrate the efficacy of the proposed framework. We also present its usage in a much broader set of applications including face frontalization, face attribute morphing, and face adversarial example detection. Jianmin Bao, Dong Chen 0003, Fang Wen 0001, Houqiang Li, Gang Hua 0001 |
CVPR | 3 |
| 2017 | Neural Aggregation Network for Video Face RecognitionabstractThis paper presents a Neural Aggregation Network (NAN) for video face recognition. The network takes a face video or face image set of a person with a variable number of face images as its input, and produces a compact, fixed-dimension feature representation for recognition. The whole network is composed of two modules. The feature embedding module is a deep Convolutional Neural Network (CNN) which maps each face image to a feature vector. The aggregation module consists of two attention blocks which adaptively aggregate the feature vectors to form a single feature inside the convex hull spanned by them. Due to the attention mechanism, the aggregation is invariant to the image order. Our NAN is trained with a standard classification or verification loss without any extra supervision signal, and we found that it automatically learns to advocate high-quality face images while repelling low-quality ones such as blurred, occluded and improperly exposed faces. The experiments on IJB-A, YouTube Face, Celebrity-1000 video face recognition benchmarks show that it consistently outperforms naive aggregation methods and achieves the state-of-the-art accuracy. Jiaolong Yang, Peiran Ren, Dongqing Zhang, Dong Chen 0003, Fang Wen 0001, Hongdong Li, Gang Hua 0001 |
CVPR | 5 |
| 2017 | CVAE-GAN: Fine-Grained Image Generation through Asymmetric TrainingabstractWe present variational generative adversarial networks, a general learning framework that combines a variational auto-encoder with a generative adversarial network, for synthesizing images in fine-grained categories, such as faces of a specific person or objects in a category. Our approach models an image as a composition of label and latent attributes in a probabilistic model. By varying the fine-grained category label fed into the resulting generative model, we can generate images in a specific category with randomly drawn values on a latent attribute vector. Our approach has two novel aspects. First, we adopt a cross entropy loss for the discriminative and classifier network, but a mean discrepancy objective for the generative network. This kind of asymmetric loss function makes the GAN training more stable. Second, we adopt an encoder network to learn the relationship between the latent space and the real image space, and use pairwise feature matching to keep the structure of generated images. We experiment with natural images of faces, flowers, and birds, and demonstrate that the proposed models are capable of generating realistic and diverse samples with fine-grained category labels. We further show that our models can be applied to other tasks, such as image inpainting, super-resolution, and data augmentation for training better face recognition models. Jianmin Bao, Dong Chen 0003, Fang Wen 0001, Houqiang Li, Gang Hua 0001 |
ICCV | 3 |
| 2017 | An Efficient Joint Formulation for Bayesian Face VerificationabstractThis paper revisits the classical Bayesian face recognition algorithm from Baback Moghaddam et al. and proposes enhancements tailored to face verification, the problem of predicting whether or not a pair of facial images share the same identity. Like a variety of face verification algorithms, the original Bayesian face model only considers the appearance difference between two faces rather than the raw images themselves. However, we argue that such a fixed and blind projection may prematurely reduce the separability between classes. Consequently, we model two facial images jointly with an appropriate prior that considers intra- and extra-personal variations over the image pairs. This joint formulation is trained using a principled EM algorithm, while testing involves only efficient closed-formed computations that are suitable for real-time practical deployment. Supporting theoretical analyses investigate computational complexity, scale-invariance properties, and convergence issues. We also detail important relationships with existing algorithms, such as probabilistic linear discriminant analysis and metric learning. Finally, on extensive experimental evaluations, the proposed model is superior to the classical Bayesian face algorithm and many alternative state-of-the-art supervised approaches, achieving the best test accuracy on three challenging datasets, Labeled Face in Wild, Multi-PIE, and YouTube Faces, all with unparalleled computational efficiency. Dong Chen 0003, Xudong Cao, David P. Wipf, Fang Wen 0001, Jian Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Supervised Transformer Network for Efficient Face Detection
Dong Chen 0003, Gang Hua 0001, Fang Wen 0001, Jian Sun 0001 |
ECCV (5) | 3 |
| 2015 | Learning Discriminative Reconstructions for Unsupervised Outlier RemovalabstractWe study the problem of automatically removing outliers from noisy data, with application for removing outlier images from an image collection. We address this problem by utilizing the reconstruction errors of an autoencoder. We observe that when data are reconstructed from low-dimensional representations, the inliers and the outliers can be well separated according to their reconstruction errors. Based on this basic observation, we gradually inject discriminative information in the learning process of an autoencoder to make the inliers and the outliers more separable. Experiments on a variety of image datasets validate our approach. Xudong Cao, Fang Wen 0001, Gang Hua 0001, Jian Sun 0001 |
ICCV | 3 |
| 2014 | Well Begun Is Half Done: Generating High-Quality Seeds for Automatic Image Dataset Construction from Web
Xudong Cao, Fang Wen 0001, Jian Sun 0001 |
ECCV (4) | 3 |
| 2014 | Face Alignment by Explicit Shape Regression
Xudong Cao, Fang Wen 0001, Jian Sun 0001 |
Int. J. Comput. Vis. | 3 |
| 2013 | Blessing of Dimensionality: High-Dimensional Feature and Its Efficient Compression for Face VerificationabstractMaking a high-dimensional (e.g., 100K-dim) feature for face recognition seems not a good idea because it will bring difficulties on consequent training, computation, and storage. This prevents further exploration of the use of a high dimensional feature. In this paper, we study the performance of a high dimensional feature. We first empirically show that high dimensionality is critical to high performance. A 100K-dim feature, based on a single-type Local Binary Pattern (LBP) descriptor, can achieve significant improvements over both its low-dimensional version and the state-of-the-art. We also make the high-dimensional feature practical. With our proposed sparse projection method, named rotated sparse regression, both computation and model storage can be reduced by over 100 times without sacrificing accuracy quality. Dong Chen 0003, Xudong Cao, Fang Wen 0001, Jian Sun 0001 |
CVPR | 3 |
| 2013 | K-Means Hashing: An Affinity-Preserving Quantization Method for Learning Binary Compact CodesabstractIn computer vision there has been increasing interest in learning hashing codes whose Hamming distance approximates the data similarity. The hashing functions play roles in both quantizing the vector space and generating similarity-preserving codes. Most existing hashing methods use hyper-planes (or kernelized hyper-planes) to quantize and encode. In this paper, we present a hashing method adopting the k-means quantization. We propose a novel Affinity-Preserving K-means algorithm which simultaneously performs k-means clustering and learns the binary indices of the quantized cells. The distance between the cells is approximated by the Hamming distance of the cell indices. We further generalize our algorithm to a product space for learning longer codes. Experiments show our method, named as K-means Hashing (KMH), outperforms various state-of-the-art hashing encoding methods. Kaiming He, Fang Wen 0001, Jian Sun 0001 |
CVPR | 2 |
| 2013 | A Practical Transfer Learning Algorithm for Face VerificationabstractFace verification involves determining whether a pair of facial images belongs to the same or different subjects. This problem can prove to be quite challenging in many important applications where labeled training data is scarce, e.g., family album photo organization software. Herein we propose a principled transfer learning approach for merging plentiful source-domain data with limited samples from some target domain of interest to create a classifier that ideally performs nearly as well as if rich target-domain data were present. Based upon a surprisingly simple generative Bayesian model, our approach combines a KL-divergence based regularizer/prior with a robust likelihood function leading to a scalable implementation via the EM algorithm. As justification for our design choices, we later use principles from convex analysis to recast our algorithm as an equivalent structured rank minimization problem leading to a number of interesting insights related to solution structure and feature-transform invariance. These insights help to both explain the effectiveness of our algorithm as well as elucidate a wide variety of related Bayesian approaches. Experimental testing with challenging datasets validate the utility of the proposed algorithm. Xudong Cao, David P. Wipf, Fang Wen 0001, Genquan Duan, Jian Sun 0001 |
ICCV | 3 |
| 2013 | Joint Inverted IndexingabstractInverted indexing is a popular non-exhaustive solution to large scale search. An inverted file is built by a quantizer such as k-means or a tree structure. It has been found that multiple inverted files, obtained by multiple independent random quantizers, are able to achieve practically good recall and speed. Instead of computing the multiple quantizers independently, we present a method that creates them jointly. Our method jointly optimizes all code words in all quantizers. Then it assigns these code words to the quantizers. In experiments this method shows significant improvement over various existing methods that use multiple independent quantizers. On the one-billion set of SIFT vectors, our method is faster and more accurate than a recent state-of-the-art inverted indexing method. Kaiming He, Fang Wen 0001, Jian Sun 0001 |
ICCV | 3 |
| 2012 | Face alignment by Explicit Shape RegressionabstractWe present a very efficient, highly accurate, “Explicit Shape Regression” approach for face alignment. Unlike previous regression-based approaches, we directly learn a vectorial regression function to infer the whole facial shape (a set of facial landmarks) from the image and explicitly minimize the alignment errors over the training data. The inherent shape constraint is naturally encoded into the regressor in a cascaded learning framework and applied from coarse to fine during the test, without using a fixed parametric shape model as in most previous methods. To make the regression more effective and efficient, we design a two-level boosted regression, shape-indexed features and a correlation-based feature selection method. This combination enables us to learn accurate models from large training data in a short time (20 minutes for 2,000 training images), and run regression extremely fast in test (15 ms for a 87 landmarks shape). Experiments on challenging data show that our approach significantly outperforms the state-of-the-art in terms of both accuracy and efficiency. Xudong Cao, Fang Wen 0001, Jian Sun 0001 |
CVPR | 3 |
| 2012 | Bayesian Face Revisited: A Joint Formulation
Dong Chen 0003, Xudong Cao, Liwei Wang 0009, Fang Wen 0001, Jian Sun 0001 |
ECCV (3) | 4 |
| 2012 | Geodesic Saliency Using Background Priors
Fang Wen 0001, Wangjiang Zhu, Jian Sun 0001 |
ECCV (3) | 2 |
| 2012 | IntentSearch: Capturing User Intention for One-Click Internet Image SearchabstractWeb-scale image search engines (e.g., Google image search, Bing image search) mostly rely on surrounding text features. It is difficult for them to interpret users' search intention only by query keywords and this leads to ambiguous and noisy search results which are far from satisfactory. It is important to use visual information in order to solve the ambiguity in text-based image retrieval. In this paper, we propose a novel Internet image search approach. It only requires the user to click on one query image with minimum effort and images from a pool retrieved by text-based search are reranked based on both visual and textual content. Our key contribution is to capture the users' search intention from this one-click query image in four steps. 1) The query image is categorized into one of the predefined adaptive weight categories which reflect users' search intention at a coarse level. Inside each category, a specific weight schema is used to combine visual features adaptive to this kind of image to better rerank the text-based search result. 2) Based on the visual content of the query image selected by the user and through image clustering, query keywords are expanded to capture user intention. 3) Expanded keywords are used to enlarge the image pool to contain more relevant images. 4) Expanded keywords are also used to expand the query image to multiple positive visual examples from which new query specific visual and textual similarity metrics are learned to further improve content-based image reranking. All these steps are automatic, without extra effort from the user. This is critically important for any commercial web-based image search engine, where the user interface has to be extremely simple. Besides this key contribution, a set of visual features which are both effective and efficient in Internet image search are designed. Experimental evaluation shows that our approach significantly improves the precision of top-ranked images and also the user experience. Xiaoou Tang, Jingyu Cui, Fang Wen 0001, Xiaogang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | A rank-order distance based clustering algorithm for face taggingabstractWe present a novel clustering algorithm for tagging a face dataset (e. g., a personal photo album). The core of the algorithm is a new dissimilarity, called Rank-Order distance, which measures the dissimilarity between two faces using their neighboring information in the dataset. The Rank-Order distance is motivated by an observation that faces of the same person usually share their top neighbors. Specifically, for each face, we generate a ranking order list by sorting all other faces in the dataset by absolute distance (e. g., L1 or L2 distance between extracted face recognition features). Then, the Rank-Order distance of two faces is calculated using their ranking orders. Using the new distance, a Rank-Order distance based clustering algorithm is designed to iteratively group all faces into a small number of clusters for effective tagging. The proposed algorithm outperforms competitive clustering algorithms in term of both precision/recall and efficiency. Chunhui Zhu, Fang Wen 0001, Jian Sun 0001 |
CVPR | 2 |
| 2010 | User intention modeling for interactive image retrievalabstractWe propose three innovative interactive methods to let computer better understand user intention in content-based image retrieval: 1. Smart intention list induces user intention, thereby improves search results by intention-specific search schema; 2. Reference strokes interaction allows user to specify in detail about the intention by pointing out interested regions; 3. Natural user feedback easily collects data of user relevance feedbacks to boost the performance of the system. Systematic user study shows that the proposed interactive mechanism improves search efficiency, reduces user workload, and enhances user experience. Jingyu Cui, Fang Wen 0001, Xiaoou Tang |
ICME | 2 |
| 2008 | Transductive object cutoutabstractIn this paper, we address the issue of transducing the object cutout model from an example image to novel image instances. We observe that although object and background are very likely to contain similar colors in natural images, it is much less probable that they share similar color configurations. Motivated by this observation, we propose a local color pattern model to characterize the color configuration in a robust way. Additionally, we propose an edge profile model to modulate the contrast of the image, which enhances edges along object boundaries and attenuates edges inside object or background. The local color pattern model and edge model are integrated in a graph-cut framework. Higher accuracy and improved robustness of the proposed method are demonstrated through experimental comparison with state-of-the-art algorithms. Jingyu Cui, Qiong Yang, Fang Wen 0001, Qiying Wu, Changshui Zhang, Luc Van Gool, Xiaoou Tang |
CVPR | 3 |
| 2008 | Face Alignment Via Component-Based Discriminative Search
Rong Xiao 0003, Fang Wen 0001, Jian Sun 0001 |
ECCV (2) | 3 |
| 2008 | Scribble-a-Secret: Similarity-based password authentication using sketchesabstractThis paper presents a sketch-based password authentication system called Scribble-a-Secret as a graphical password scheme in which free-form drawings are used as a means to authenticate users. Unlike existing schemes, this approach requires no input of graphical passwords in particular sequences of strokes. Moreover, the system allows for a modicum of variation when users recreate their passwords. Our technique uses edge orientations extracted from sketch images to discern one user from another. Our experiments show that our recognition technique is robust for recognizing sketches while differentiating from others with both a false acceptance rate and false rejection rate of less than 1%. Mizuki Oka, Kazuhiko Kato, Ying-Qing Xu, Fang Wen 0001 |
ICPR | 5 |
| 2008 | Easytoon: an easy and quick tool to personalize a cartoon storyboard using family photo albumabstractA family photo album based cartoon personalization system, EasyToon, is proposed in this paper. Using state of the art computer vision and graphics technologies and effective UI design, the interactive tool can quickly generate a personalized cartoon storyboard, which naturally blends a real face chosen from the family photo album into a cartoon picture. The personalized cartoon image is easily and quickly obtained in two main steps. First, the best face candidate is selected from the album interactively. Then a personalized cartoon image is automatically synthesized by blending the selected face into the interesting cartoon image. Experiments show that most users express great interest in our system. Without any art background, they can make a personalized cartoon of high quality using the EasyToon within minutes. Shifeng Chen, Yuandong Tian, Fang Wen 0001, Ying-Qing Xu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2008 | Real time google and live image search re-rankingabstractNowadays, web-scale image search engines (e.g. Google, Live Image Search) rely almost purely on surrounding text features. This leads to ambiguous and noisy results. We propose to use adaptive visual similarity to re-rank the text-based search results. A query image is first categorized into one of several predefined intention categories, and a specific similarity measure is used inside each category to combine image features for re-ranking based on the query image. Extensive experiments demonstrate that using this algorithm to filter output of Google and Live Image Search is a practical and effective way to dramatically improve the user experience. A real-time image search engine is developed for on-line image search with re-ranking: http://mmlab.ie.cuhk.edu.hk/intentsearch Jingyu Cui, Fang Wen 0001, Xiaoou Tang |
ACM Multimedia | 2 |
| 2008 | IntentSearch: interactive on-line image search re-rankingabstractIn this demo, we present IntentSearch, an interactive system for realtime web based image retrieval. IntentSearch works directly on top of Microsoft Live Image Search, and re-ranks its results according to user specified query image(s) and the automatically inferred user intention. Besides searching in the interface of Microsoft Live Image Search, we also design a more flexible interface to let users browse and play with all the images in the current search session, which makes web image search more efficient and interesting. Please visit http://mmlab.ie.cuhk.edu.hk/intentsearch for the experience. Jingyu Cui, Fang Wen 0001, Xiaoou Tang |
ACM Multimedia | 2 |
| 2008 | EasyToon: cartoon personalization using face photosabstractIn this demo, we present a family photo album based cartoon personalization system, EasyToon. Using the family photo album as the candidate pool, a personalized cartoon image is obtained in two main steps. First, the best face candidate is selected from the album interactively. Then a personalized cartoon image is automatically synthesized by lending the selected face into the target cartoon image. By integrating state of the art computer vision and graphics technologies and effective UI design EasyToon can generate a personalized cartoon storyboard easily and quickly. Fang Wen 0001, Shifeng Chen, Xiaoou Tang |
ACM Multimedia | 1 |
| 2007 | EasyAlbum: an interactive photo annotation system based on face clustering and re-rankingabstractDigital photo management is becoming indispensable for the explosively growing family photo albums due to the rapid popularization of digital cameras and mobile phone cameras. In an effective photo management system photo annotation is the most challenging task. In this paper, we develop several innovative interaction techniques for semi-automatic photo annotation. Compared with traditional annotation systems, our approach provides the following new features: "cluster annotation" puts similar faces or photos with similar scene together, and enables user label them in one operation; "contextual re-ranking" boosts the labeling productivity by guessing the user intention; "ad hoc annotation" allows user label photos while they are browsing or searching, and improves system performance progressively through learning propagation. Our results show that these technologies provide a more user friendly interface for the annotation of person name, location, and event, and thus substantially improve the annotation performance especially for a large photo album. Jingyu Cui, Fang Wen 0001, Rong Xiao 0003, Yuandong Tian, Xiaoou Tang |
CHI | 2 |
| 2007 | A Face Annotation Framework with Partial Clustering and Interactive LabelingabstractFace annotation technology is important for a photo management system. In this paper, we propose a novel interactive face annotation framework combining unsupervised and interactive learning. There are two main contributions in our framework. In the unsupervised stage, a partial clustering algorithm is proposed to find the most evident clusters instead of grouping all instances into clusters, which leads to a good initial labeling for later user interaction. In the interactive stage, an efficient labeling procedure based on minimization of both global system uncertainty and estimated number of user operations is proposed to reduce user interaction as much as possible. Experimental results show that the proposed annotation framework can significantly reduce the face annotation workload and is superior to existing solutions in the literature. Yuandong Tian, Wei Liu 0026, Rong Xiao 0003, Fang Wen 0001, Xiaoou Tang |
CVPR | 4 |
| 2007 | Color Transfer BrushabstractIn this paper, we introduce an interactive tool for local color transfer. The new technique is based on the observation that color transfer operations are local in nature while at the same time should adhere global consistency. We introduce a brush by which the user specifies the source and destination image regions for color transfer. Color statistics in the source region are transferred to the destination region. A global optimization is then applied to eliminate vi- sual discontinuities that may result by the local operations. We demonstrate that our tool is easy to use yet effective in quickly generating diverse artistic effects. Qing Luan, Fang Wen 0001, Ying-Qing Xu |
PG | 2 |
| 2007 | Natural Image Colorization
Qing Luan, Fang Wen 0001, Daniel Cohen-Or, Ying-Qing Xu, Harry Shum |
Rendering Techniques | 2 |
| 2007 | Image vectorization using optimized gradient meshesabstractRecently, gradient meshes have been introduced as a powerful vector graphics representation to draw multicolored mesh objects with smooth transitions. Using tools from Abode Illustrator and Corel CorelDraw, a user can manually create gradient meshes even for photo-realistic vector arts, which can be further edited, stylized and animated. In this paper, we present an easy-to-use interactive tool, called optimized gradient mesh , to semi-automatically and quickly create gradient meshes from a raster image. We obtain the optimized gradient mesh by formulating an energy minimization problem. The user can also interactively specify a few vector lines to guide the mesh generation. The resulting optimized gradient mesh is an editable and scalable mesh that otherwise would have taken many hours for a user to manually create. Jian Sun 0001, Fang Wen 0001, Harry Shum |
ACM Trans. Graph. | 3 |
| 2006 | Accurate Face Alignment using Shape Constrained Markov NetworkabstractIn this paper, we present a shape constrained Markov network for accurate face alignment. The global face shape is defined as a set of weighted shape samples which are integrated into the Markov network optimization. These weighted samples provide structural constraints to make the Markov network more robust to local image noise. We propose a hierarchical Condensation algorithm to draw the shape samples efficiently. Specifically, a proposal density incorporating the local face shape is designed to generate more samples close to the image features for accurate alignment, based on a local Markov network search. A constrained regularization algorithm is also developed to weigh favorably those points that are already accurately aligned. Extensive experiments demonstrate the accuracy and effectiveness of our proposed approach. Fang Wen 0001, Ying-Qing Xu, Xiaoou Tang, Harry Shum |
CVPR (1) | 2 |
| 2006 | An Integrated Model for Accurate Shape Alignment
Fang Wen 0001, Xiaoou Tang, Ying-Qing Xu |
ECCV (4) | 2 |
| 2004 | Synthesizing Dynamic Texture with Closed-Loop Linear Dynamic System
Lu Yuan 0001, Fang Wen 0001, Ce Liu 0001, Harry Shum |
ECCV (2) | 2 |