EDBT 2026 Demo / reviewers in the wild / expert
Chuanxia Zheng
dblp:195/8988
· DBLP profile ↗
37ranked-venue papers
11as first author
32since 2021 · last 2026
0000-0002-3584-9640ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 10 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 20 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Compact 3D Gaussians via Feed-Forward Point FusionabstractWe present Splatt3RFusion, a feed-forward neural network that, given a set of unposed and uncalibrated images, directly reconstructs a compact and high-quality 3D Gaussian Splat representation of a scene. Unlike prior pixel-aligned feed-forward methods that typically predict one 3D Gaussian primitive per pixel in each image - producing severe redundancy, duplication, and ghosting on one physical surface - our approach efficiently fuses points in 3D space through a multi-scale octree structure, yielding a compact and coherent representation. Built upon VGGT, a foundation model for pose-free 3D geometry prediction, Splatt3RFusion introduces a Gaussian prediction branch that infers primitive parameters using only photometric supervision. We also introduce the ability to control the number of 3D Gaussians generated at test-time, allowing for a controllable tradeoff between PSNR and the number of 3D Gaussian primitives used. The model is efficient, reducing both memory usage and rendering cost, while achieving state-of-the-art results on RealEstate10k and ScanNet++. Brandon Smart, Chuanxia Zheng, Iro Laina, Victor Adrian Prisacariu |
3DV | 2 |
| 2025 | Flash3D: Feed-Forward Generalisable 3D Scene Reconstruction from a Single ImageabstractWe propose Flash3D, a method for scene reconstruction and novel view synthesis from a single image which is both very generalisable and efficient. For generalisability, we start from a 'foundation' model for monocular depth estimation and extend it to a full 3D shape and appearance reconstructor. For efficiency, we base this extension on feed-forward Gaussian Splatting. Specifically, we predict a first layer of 3D Gaussians at the predicted depth, and then add additional layers of Gaussians that are offset in space, allowing the model to complete the reconstruction behind occlusions and truncations. Flash3D is very efficient, trainable on a single GPU in a day, and thus accessible to most researchers. It achieves state-of-the-art results when trained and tested on RealEstate10k. When transferred to unseen datasets like NYU it outperforms competitors by a large margin. More impressively, when transferred to KITTI, Flash3D achieves better PSNR than methods trained specifically on that dataset. In some instances, it even outperforms recent methods that use multiple views as input. Code, models, demo, and more results are available at https://www.robots.ox.ac.uk/~vgg/research/flash3d/. Stanislaw Szymanowicz, Eldar Insafutdinov, Chuanxia Zheng, Dylan Campbell, João F. Henriques, Christian Rupprecht 0001, Andrea Vedaldi |
3DV | 3 |
| 2025 | Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
Zeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus, Andrea Vedaldi |
ICCV | 2 |
| 2025 | Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level DynamicsabstractWe introduce Puppet-Master, an interactive video generator that captures the internal, part-level motion of objects, serving as a proxy for modeling object dynamics universally. Given an image of an object and a set of "drags" specifying the trajectory of a few points on the object, the model synthesizes a video where the object's parts move accordingly. To build Puppet-Master, we extend a pre-trained image-to-video generator to encode the input drags. We also propose all-to-first attention, an alternative to conventional spatial attention that mitigates artifacts caused by fine-tuning a video generator on out-of-domain data. The model is fine-tuned on Objaverse-Animation-HQ, a new dataset of curated part-level motion clips obtained by rendering synthetic 3D animations. Unlike real videos, these synthetic clips avoid confounding part-level motion with overall object and camera motion. We extensively filter sub-optimal animations and augment the synthetic renderings with meaningful drags that emphasize the internal dynamics of objects. We demonstrate that Puppet-Master learns to generate part-level motions, unlike other motion-conditioned video generators that primarily move the object as a whole. Moreover, Puppet-Master generalizes well to out-of-domain real images, outperforming existing methods on real-world benchmarks in a zero-shot manner. Ruining Li, Chuanxia Zheng, Christian Rupprecht 0001, Andrea Vedaldi |
ICCV | 2 |
| 2025 | DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness
Ruining Li, Chuanxia Zheng, Christian Rupprecht 0001, Andrea Vedaldi |
ICCV | 2 |
| 2025 | Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
Tianhao Wu 0015, Chuanxia Zheng, Frank Guan, Andrea Vedaldi, Tat-Jen Cham |
ICCV | 2 |
| 2025 | Semantix: An Energy-guided Sampler for Semantic Style TransferabstractRecent advances in style and appearance transfer are impressive, but most methods isolate global style and local appearance transfer, neglecting semantic correspondence. Additionally, image and video tasks are typically handled in isolation, with little focus on integrating them for video transfer. To address these limitations, we introduce a novel task, *Semantic Style Transfer*, which involves transferring style and appearance features from a reference image to a target visual content based on semantic correspondence. We subsequently propose a training-free method, *Semantix*, an energy-guided sampler designed for Semantic Style Transfer that simultaneously guides both style and appearance transfer based on semantic understanding capacity of pre-trained diffusion models. Additionally, as a sampler, *Semantix* can be seamlessly applied to both image and video models, enabling semantic style transfer to be generic across various visual media. Specifically, once inverting both reference and context images or videos to noise space by SDEs, *Semantix* utilizes a meticulously crafted energy function to guide the sampling process, including three key components: *Style Feature Guidance*, *Spatial Feature Guidance* and *Semantic Distance* as a regularisation term. Experimental results demonstrate that *Semantix* not only effectively accomplishes the task of semantic style transfer across images and videos, but also surpasses existing state-of-the-art solutions in both fields. Huiang He, Minghui Hu 0001, Chuanxia Zheng, Tat-Jen Cham |
ICLR | 3 |
| 2025 | Explicit Correspondence Matching for Generalizable Neural Radiance FieldsabstractWe present a new generalizable NeRF method that is able to directly generalize to new unseen scenarios and perform novel view synthesis with as few as two source views. The key to our approach lies in the explicitly modeled correspondence matching information, so as to provide the geometry prior to the prediction of NeRF color and density for volume rendering. The explicit correspondence matching is quantified with the cosine similarity between image features sampled at the 2D projections of a 3D point on different views, which is able to provide reliable cues about the surface geometry. Unlike previous methods where image features are extracted independently for each view, we consider modeling the cross-view interactions via Transformer cross-attention, which greatly improves the feature matching quality. Our method achieves state-of-the-art results on different evaluation settings, with the experiments showing a strong correlation between our learned cosine feature similarity and volume density, demonstrating the effectiveness and superiority of our proposed method. Yuedong Chen, Haofei Xu, Qianyi Wu, Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | One-Shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural TexturingabstractHuman motion transfer aims at animating a static source image with a driving video. While recent advances in one-shot human motion transfer have led to significant improvement in results, it remains challenging for methods with 2D body landmarks, skeleton and semantic mask to accurately capture correspondences between source and driving poses due to the large variation in motion and articulation complexity. In addition, the accuracy and precision of DensePose degrade the image quality for neural-rendering-based methods. To address the limitations and by both considering the importance of appearance and geometry for motion transfer, in this work, we proposed a unified framework that combines multi-scale feature warping and neural texture mapping to recover better 2D appearance and 2.5D geometry, partly by exploiting the information from DensePose, yet adapting to its inherent limited accuracy. Our model takes advantage of multiple modalities by jointly training and fusing them, which allows it to robust neural texture features that cope with geometric errors as well as multi-scale dense motion flow that better preserves appearance. Experimental results with full and half-view body video datasets demonstrate that our model can generalize well and achieve competitive results, and that it is particularly effective in handling challenging cases such as those with substantial self-occlusions. Yuzhu Ji, Chuanxia Zheng, Tat-Jen Cham |
IEEE Trans. Multim. | 2 |
| 2024 | One More Step: A Versatile Plug-and-Play Module for Rectifying Diffusion Schedule Flaws and Enhancing Low-Frequency ControlsabstractIt is well known that many open-released foundational diffusion models have difficulty in generating images that substantially depart from average brightness, despite such images being present in the training data. This is due to an inconsistency: while denoising starts from pure Gaus-sian noise during inference, the training noise schedule retains residual data even in the final timestep distribution, due to difficulties in numerical conditioning in main-stream formulation, leading to unintended bias during in-ference. To mitigate this issue, certain ∊-prediction mod-els are combined with an ad-hoc offset-noise methodology. In parallel, some contemporary models have adopted zero-terminal SNR noise schedules together with v -prediction, which necessitate major alterations to pre-trained models. However, such changes risk destabilizing a large multitude of community-driven applications anchored on these pre-trained models. In light of this, our investigation revisits the fundamental causes, leading to our proposal of an inno-vative and principled remedy, called One More Step (OMS). By integrating a compact network and incorporating an ad-ditional simple yet effective step during inference, OMS ele-vates image fidelity and harmonizes the dichotomy between training and inference, while preserving original model pa-rameters. Once trained, various pre-trained diffusion mod-els with the same latent domain can share the same OMS module. Codes and models are released at here. Minghui Hu 0001, Chuanxia Zheng, Dacheng Tao, Tat-Jen Cham |
CVPR | 3 |
| 2024 | Amodal Ground Truth and Completion in the WildabstractThis paper studies amodal image segmentation: predicting entire object segmentation masks including both visible and invisible (occluded) parts. In previous work, the amodal segmentation ground truth on real images is usually pre-dicted by manual annotaton and thus is subjective. In contrast, we use 3D data to establish an automatic pipeline to determine authentic ground truth amodal masks for partially occluded objects in real images. This pipeline is used to construct an amodal completion evaluation benchmark, MP3D-Amodal, consisting of a variety of object categories and la-bels. To better handle the amodal completion task in the wild, we explore two architecture variants: a two-stage model that first infers the occluder, followed by amodal mask completion; and a one-stage model that exploits the representation power of Stable Diffusion for amodal segmentation across many categories. Without bells and whistles, our method achieves a new state-of-the-art performance on Amodal segmentation datasets that cover a large variety of objects, in-cluding COCOA and our new MP3D-Amodal dataset. The dataset, model, and code are available at https://www.robots.ox.ac.uk/~vgg/research/amodal/ Guanqi Zhan, Chuanxia Zheng, Weidi Xie, Andrew Zisserman |
CVPR | 2 |
| 2024 | Free3D: Consistent Novel View Synthesis Without 3D RepresentationabstractWe introduce Free3D, a simple accurate method for monocular open-set novel view synthesis (NVS). Similar to Zero-1-to-S, we start from a pre-trained 2D image generator for generalization, and fine-tune it for NVS. Compared to other works that took a similar approach, we obtain significant improvements without resorting to an explicit 3D representation, which is slow and memory-consuming, and without training an additional network for 3D reconstruction. Our key contribution is to improve the way the target camera pose is encoded in the network, which we do by introducing a new ray conditioning normalization (RCN) layer. The latter injects pose information in the underlying 2D image generator by telling each pixel its viewing direction. We further improve multi-view consistency by using light-weight multi-view attention layers and by sharing generation noise between the different views. We train Free3D on the Objaverse dataset and demonstrate excellent generalization to new categories in new datasets, including OmniObject3D and GSO. The project page is available at https://chuanxiar.com/free3d/. Chuanxia Zheng, Andrea Vedaldi |
CVPR | 1 |
| 2024 | MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-view Images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger 0001, Tat-Jen Cham, Jianfei Cai 0001 |
ECCV (21) | 3 |
| 2024 | DragAPart: Learning a Part-Level Motion Prior for Articulated Objects
Ruining Li, Chuanxia Zheng, Christian Rupprecht 0001, Andrea Vedaldi |
ECCV (2) | 2 |
| 2024 | ClusteringSDF: Self-Organized Neural Implicit Surfaces for 3D Decomposition
Tianhao Wu 0015, Chuanxia Zheng, Qianyi Wu, Tat-Jen Cham |
ECCV (57) | 2 |
| 2024 | PanoDiffusion: 360-degree Panorama Outpainting via DiffusionabstractGenerating complete 360\textdegree{} panoramas from narrow field of view images is ongoing research as omnidirectional RGB data is not readily available. Existing GAN-based approaches face some barriers to achieving higher quality output, and have poor generalization performance over different mask types. In this paper, we present our 360\textdegree{} indoor RGB panorama outpainting model using latent diffusion models (LDM), called PanoDiffusion. We introduce a new bi-modal latent diffusion structure that utilizes both RGB and depth panoramic data during training, which works surprisingly well to outpaint depth-free RGB images during inference. We further propose a novel technique of introducing progressive camera rotations during each diffusion denoising step, which leads to substantial improvement in achieving panorama wraparound consistency. Results show that our PanoDiffusion not only significantly outperforms state-of-the-art methods on RGB panorama outpainting by producing diverse well-structured results for different types of masks, but can also synthesize high-quality depth panoramas to provide realistic 3D indoor models. Tianhao Wu 0015, Chuanxia Zheng, Tat-Jen Cham |
ICLR | 2 |
| 2024 | MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse ViewsabstractWe introduce MVSplat360, a feed-forward approach for 360° novel view synthesis (NVS) of diverse real-world scenes, using only sparse observations. This setting is inherently ill-posed due to minimal overlap among input views and insufficient visual information provided, making it challenging for conventional methods to achieve high-quality results. Our MVSplat360 addresses this by effectively combining geometry-aware 3D reconstruction with temporally consistent video generation. Specifically, it refactors a feed-forward 3D Gaussian Splatting (3DGS) model to render features directly into the latent space of a pre-trained Stable Video Diffusion (SVD) model, where these features then act as pose and visual cues to guide the denoising process and produce photorealistic 3D-consistent views. Our model is end-to-end trainable and supports rendering arbitrary views with as few as 5 sparse input views. To evaluate MVSplat360's performance, we introduce a new benchmark using the challenging DL3DV-10K dataset, where MVSplat360 achieves superior visual quality compared to state-of-the-art methods on wide-sweeping or even 360° NVS tasks. Experiments on the existing benchmark RealEstate10K also confirm the effectiveness of our model. Readers are highly recommended to view the video results at [donydchen.github.io/mvsplat360](https://donydchen.github.io/mvsplat360). Yuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang, Andrea Vedaldi, Tat-Jen Cham, Jianfei Cai 0001 |
NeurIPS | 2 |
| 2024 | A General Protocol to Probe Large Vision Models for 3D Physical UnderstandingabstractOur objective in this paper is to probe large vision models to determine to what extent they ‘understand’ different physical properties of the 3D scene depicted in an image. To this end, we make the following contributions: (i) We introduce a general and lightweight protocol to evaluate whether features of an off-the-shelf large vision model encode a number of physical ‘properties’ of the 3D scene, by training discriminative classifiers on the features for these properties. The probes are applied on datasets of real images with annotations for the property. (ii) We apply this protocol to properties covering scene geometry, scene material, support relations, lighting, and view-dependent measures, and large vision models including CLIP, DINOv1, DINOv2, VQGAN, Stable Diffusion. (iii) We find that features from Stable Diffusion and DINOv2 are good for discriminative learning of a number of properties, including scene geometry, support relations, shadows and depth, but less performant for occlusion and material, while outperforming DINOv1, CLIP and VQGAN for all properties. (iv) It is observed that different time steps of Stable Diffusion features, as well as different transformer layers of DINO/CLIP/VQGAN, are good at different properties, unlocking potential applications of 3D physical understanding. Guanqi Zhan, Chuanxia Zheng, Weidi Xie, Andrew Zisserman |
NeurIPS | 2 |
| 2024 | Bridging Global Context Interactions for High-Fidelity Pluralistic Image CompletionabstractWe introduce PICFormer, a novel framework for Pluralistic Image Completion using a transFormer based architecture, that achieves both high quality and diversity at a much faster inference speed. Our key contribution is to introduce a code-shared codebook learning using a restrictive CNN on small and non-overlapping receptive fields (RFs) for the local visible token representation. This results in a compact yet expressive discrete representation, facilitating efficient modeling of global visible context relations by the transformer. Unlike the prevailing autoregressive approaches, we proposed to sample all tokens simultaneously, leading to more than 100× faster inference speed. To enhance appearance consistency between visible and generated regions, we further propose a novel attention-aware layer (AAL), designed to better exploit distantly related high-frequency features. Through extensive experiments, we demonstrate that the PICFormer efficiently learns semantically-rich discrete codes, resulting in significantly improved image quality. Moreover, our diverse image completion framework surpasses State-of-the-Art methods on multiple image completion datasets. Chuanxia Zheng, Guoxian Song, Tat-Jen Cham, Jianfei Cai 0001, Linjie Luo, Dinh Q. Phung |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Online Clustered CodebookabstractVector Quantisation (VQ) is experiencing a comeback in machine learning, where it is increasingly used in representation learning. However, optimizing the codevectors in existing VQ-VAE is not entirely trivial. A problem is codebook collapse, where only a small subset of codevectors receive gradients useful for their optimisation, whereas a majority of them simply "dies off" and is never updated or used. This limits the effectiveness of VQ for learning larger codebooks in complex computer vision tasks that require high-capacity representations. In this paper, we present a simple alternative method for online codebook learning, Clustering VQ-VAE (CVQ-VAE). Our approach selects encoded features as anchors to update the "dead" codevectors, while optimising the codebooks which are alive via the original loss. This strategy brings unused codevectors closer in distribution to the encoded features, increasing the likelihood of being chosen and optimized. We extensively validate the generalization capability of our quantiser on various datasets, tasks (e.g. reconstruction and generation), and architectures (e.g. VQ-VAE, VQGAN, LDM). CVQ-VAE can be easily integrated into the existing models with just a few lines of code. Chuanxia Zheng, Andrea Vedaldi |
ICCV | 1 |
| 2023 | Unified Discrete Diffusion for Simultaneous Vision-Language Generation
Minghui Hu 0001, Chuanxia Zheng, Zuopeng Yang, Tat-Jen Cham, Heliang Zheng, Dacheng Tao, Ponnuthurai N. Suganthan |
ICLR | 2 |
| 2023 | Vector Quantized Wasserstein Auto-EncoderabstractLearning deep discrete latent presentations offers a promise of better symbolic and summarized abstractions that are more useful to subsequent downstream tasks. Inspired by the seminal Vector Quantized Variational Auto-Encoder (VQ-VAE), most of work in learning deep discrete representations has mainly focused on improving the original VQ-VAE form and none of them has studied learning deep discrete representations from the generative viewpoint. In this work, we study learning deep discrete representations from the generative viewpoint. Specifically, we endow discrete distributions over sequences of codewords and learn a deterministic decoder that transports the distribution over the sequences of codewords to the data distribution via minimizing a WS distance between them. We develop further theories to connect it with the clustering viewpoint of WS distance, allowing us to have a better and more controllable clustering solution. Finally, we empirically evaluate our method on several well-known benchmarks, where it achieves better qualitative and quantitative performances than the other VQ-VAE variants in terms of the codebook utilization and image reconstruction/generation. Long Tung Vuong, Trung Le 0001, He Zhao 0001, Chuanxia Zheng, Mehrtash Harandi, Jianfei Cai 0001, Dinh Q. Phung |
ICML | 4 |
| 2023 | Cocktail: Mixing Multi-Modality Control for Text-Conditional Image GenerationabstractText-conditional diffusion models are able to generate high-fidelity images with diverse contents.
However, linguistic representations frequently exhibit ambiguous descriptions of the envisioned objective imagery, requiring the incorporation of additional control signals to bolster the efficacy of text-guided diffusion models.
In this work, we propose Cocktail, a pipeline to mix various modalities into one embedding, amalgamated with a generalized ControlNet (gControlNet), a controllable normalisation (ControlNorm), and a spatial guidance sampling method, to actualize multi-modal and spatially-refined control for text-conditional diffusion models.
Specifically, we introduce a hyper-network gControlNet, dedicated to the alignment and infusion of the control signals from disparate modalities into the pre-trained diffusion model.
gControlNet is capable of accepting flexible modality signals, encompassing the simultaneous reception of any combination of modality signals, or the supplementary fusion of multiple modality signals.
The control signals are then fused and injected into the backbone model according to our proposed ControlNorm.
Furthermore, our advanced spatial guidance sampling methodology proficiently incorporates the control signal into the designated region, thereby circumventing the manifestation of undesired objects within the generated image.
We demonstrate the results of our method in controlling various modalities, proving high-quality synthesis and fidelity to multiple external signals. Minghui Hu 0001, Daqing Liu, Chuanxia Zheng, Dacheng Tao, Tat-Jen Cham |
NeurIPS | 4 |
| 2022 | Bridging Global Context Interactions for High-Fidelity Image CompletionabstractBridging global context interactions correctly is important for high-fidelity image completion with large masks. Previous methods attempting this via deep or large receptive field (RF) convolutions cannot escape from the dominance of nearby interactions, which may be inferior. In this paper, we propose to treat image completion as a directionless sequence-to-sequence prediction task, and deploy a transformer to directly capture long-range depen-dence. Crucially, we employ a restrictive CNN with small and non-overlapping RF for weighted token representation, which allows the transformer to explicitly model the long-range visible context relations with equal importance in all layers, without implicitly confounding neighboring tokens when larger RFs are used. To improve appearance consistency between visible and generated regions, a novel attention-aware layer (AAL) is introduced to better exploit distantly related high-frequency features. Overall, extensive experiments demonstrate superior performance compared to state-of-the-art methods on several datasets. Code is available at https://github.com/lyndonzheng/TFill. Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001, Dinh Q. Phung |
CVPR | 1 |
| 2022 | Sem2NeRF: Converting Single-View Semantic Masks to Neural Radiance Fields
Yuedong Chen, Qianyi Wu, Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
ECCV (14) | 3 |
| 2022 | Object-Compositional Neural Implicit Surfaces
Qianyi Wu, Yuedong Chen, Kejie Li, Chuanxia Zheng, Jianfei Cai 0001, Jianmin Zheng |
ECCV (27) | 5 |
| 2022 | MoVQ: Modulating Quantized Vectors for High-Fidelity Image GenerationabstractAlthough two-stage Vector Quantized (VQ) generative models allow for synthesizing high-fidelity and high-resolution images, their quantization operator encodes similar patches within an image into the same index, resulting in a repeated artifact for similar adjacent regions using existing decoder architectures. To address this issue, we propose to incorporate the spatially conditional normalization to modulate the quantized vectors so as to insert spatially variant information to the embedded index maps, encouraging the decoder to generate more photorealistic images. Moreover, we use multichannel quantization to increase the recombination capability of the discrete codes without increasing the cost of model and codebook. Additionally, to generate discrete tokens at the second stage, we adopt a Masked Generative Image Transformer (MaskGIT) to learn an underlying prior distribution in the compressed latent space, which is much faster than the conventional autoregressive model. Experiments on two benchmark datasets demonstrate that our proposed modulated VQGAN is able to greatly improve the reconstructed image quality as well as provide high-fidelity image generation. Chuanxia Zheng, Tung Long Vuong, Jianfei Cai 0001, Dinh Q. Phung |
NeurIPS | 1 |
| 2021 | The Spatially-Correlative Loss for Various Image Translation TasksabstractWe propose a novel spatially-correlative loss that is simple, efficient and yet effective for preserving scene structure consistency while supporting large appearance changes during unpaired image-to-image (I2I) translation. Previous methods attempt this by using pixel-level cycle-consistency or feature-level matching losses, but the domain-specific nature of these losses hinder translation across large domain gaps. To address this, we exploit the spatial patterns of self-similarity as a means of defining scene structure. Our spatially-correlative loss is geared towards only capturing spatial relationships within an image rather than domain appearance. We also introduce a new self-supervised learning method to explicitly learn spatially-correlative maps for each specific translation task. We show distinct improvement over baseline models in all three modes of unpaired I2I translation: single-modal, multi-modal, and even single-image translation. This new loss can easily be integrated into existing network architectures and thus allows wide applicability. The code is available at https://github.com/lyndonzheng/F-LSeSim. Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
CVPR | 1 |
| 2021 | A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder∗abstractWe present a unified and flexible framework to address the generalized problem of 3D motion synthesis that covers the tasks of motion prediction, completion, interpolation, and spatial-temporal recovery. Since these tasks have different input constraints and various fidelity and diversity requirements, most existing approaches only cater to a specific task or use different architectures to address various tasks. Here we propose a unified framework based on Conditional Variational Auto-Encoder (CVAE), where we treat any arbitrary input as a masked motion series. Notably, by considering this problem as a conditional generation process, we estimate a parametric distribution of the missing regions based on the input conditions, from which to sample and synthesize the full motion series. To further allow the flexibility of manipulating the motion style of the generated series, we design an Action-Adaptive Modulation (AAM) to propagate the given semantic guidance through the whole sequence. We also introduce a cross-attention mechanism to exploit distant relations among decoder and encoder features for better realism and global consistency. We conducted extensive experiments on Human 3.6M and CMU-Mocap. The results show that our method produces coherent and realistic results for various motion synthesis tasks, with the synthesized motions distinctly adapted by the given action labels. Yujun Cai, Yiwei Wang 0001, Yiheng Zhu 0003, Tat-Jen Cham, Jianfei Cai 0001, Junsong Yuan 0001, Jun Liu 0036, Chuanxia Zheng, Sijie Yan, Henghui Ding, Xiaohui Shen, Ding Liu 0001, Nadia Magnenat-Thalmann |
ICCV | 8 |
| 2021 | Pluralistic Free-Form Image Completion
Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | Visiting the Invisible: Layer-by-Layer Completed Scene Decomposition
Chuanxia Zheng, Duy-Son Dao, Guoxian Song, Tat-Jen Cham, Jianfei Cai 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | AgileGAN: stylizing portraits by inversion-consistent transfer learningabstractPortraiture as an art form has evolved from realistic depiction into a plethora of creative styles. While substantial progress has been made in automated stylization, generating high quality stylistic portraits is still a challenge, and even the recent popular Toonify suffers from several artifacts when used on real input images. Such StyleGAN-based methods have focused on finding the best latent inversion mapping for reconstructing input images; however, our key insight is that this does not lead to good generalization to different portrait styles. Hence we propose AgileGAN, a framework that can generate high quality stylistic portraits via inversion-consistent transfer learning. We introduce a novel hierarchical variational autoencoder to ensure the inverse mapped distribution conforms to the original latent Gaussian distribution, while augmenting the original space to a multi-resolution latent space so as to better encode different levels of detail. To better capture attribute-dependent stylization of facial features, we also present an attribute-aware generator and adopt an early stopping strategy to avoid overfitting small training datasets. Our approach provides greater agility in creating high quality and high resolution (1024×1024) portrait stylization models, requiring only a limited number of style exemplars (~100) and short training time (~1 hour). We collected several style datasets for evaluation including 3D cartoons, comics, oil paintings and celebrities. We show that we can achieve superior portrait stylization quality to previous state-of-the-art methods, with comparisons done qualitatively, quantitatively and through a perceptual user study. We also demonstrate two applications of our method, image editing and motion retargeting. Guoxian Song, Linjie Luo, Wan-Chun Ma, Chun-Pong Lai, Chuanxia Zheng, Tat-Jen Cham |
ACM Trans. Graph. | 6 |
| 2019 | Pluralistic Image CompletionabstractMost image completion methods produce only one result for each masked input, although there may be many reasonable possibilities. In this paper, we present an approach for pluralistic image completion - the task of generating multiple and diverse plausible solutions for image completion. A major challenge faced by learning-based approaches is that usually only one ground truth training instance per label. As such, sampling from conditional VAEs still leads to minimal diversity. To overcome this, we propose a novel and probabilistically principled framework with two parallel paths. One is a reconstructive path that utilizes the only one given ground truth to get prior distribution of missing parts and rebuild the original image from this distribution. The other is a generative path for which the conditional prior is coupled to the distribution obtained in the reconstructive path. Both are supported by GANs. We also introduce a new short+long term attention layer that exploits distant relations among decoder and encoder features, improving appearance consistency. When tested on datasets with buildings (Paris), faces (CelebA-HQ), and natural images (ImageNet), our method not only generated higher-quality completion results, but also with multiple and diverse plausible outputs. Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
CVPR | 1 |
| 2019 | Task-in-all Domain Adaptation for Semantic SegmentationabstractIn this work we tackle the problem of unsupervised domain adaptation for semantic segmentation. One pipeline is to sequentially train image-translation model and the final task segmentation model. In such pipeline, image translation is aimed to generate the translated source-domain images which are visually similar to the target-domain images and then the final task model is trained using the translated images and its corresponding groundtruth. However, the visually optimal translated-images are not necessarily optimal for the final task of segmenting the target-domain images. Thus we propose a Task-in-all pipeline for unsupervised domain adaptation on semantic segmentation, which incorporates image translation and final segmentation task into an end-to-end training pipeline. Our aim is to generate the translated images which better assists the final task, instead of just being visually similar to the target domain images. We show that in the task of adapting from GTA5 to Cityscapes dataset, the segmentation performance of our Task-in-all pipeline outperforms the sequentially training pipeline, with simpler model structure and less training complexity. Tianyi Zhang 0004, Chuanxia Zheng, Guosheng Lin, Jianfei Cai 0001, Alex Chichung Kot |
VCIP | 3 |
| 2018 | T ^2 2 Net: Synthetic-to-Realistic Translation for Solving Single-Image Depth Estimation Tasks
Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai 0001 |
ECCV (7) | 1 |
| 2018 | Multi-class indoor semantic segmentation with deep structured model
Chuanxia Zheng, Weihai Chen, Xingming Wu |
Vis. Comput. | 1 |
| 2017 | Learning aggregated features and optimizing model for semantic labeling
Chuanxia Zheng, Weihai Chen, Xingming Wu |
Vis. Comput. | 2 |