Chun-Han Yao

dblp:184/9458 · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
14since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 10 since 2021
YearPublicationVenuePosition
2025 SViM3D: Stable Video Material Diffusion for Single Image 3D Generation
Andreas Engelhardt, Mark Boss, Vikram Voletti, Chun-Han Yao, Hendrik P. A. Lensch, Varun Jampani
ICCV4
2025 SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation
Chun-Han Yao, Vikram Voleti, Huaizu Jiang, Varun Jampani
ICCV1
2025 FaceCraft4D: Animated 3D Facial Avatar Generation from a Single Image
Mallikarjun B. R. 0001, Chun-Han Yao, Rafal Mantiuk, Varun Jampani
ICCV3
2025 Stable Virtual Camera: Generative View Synthesis with Diffusion Models
abstract
We present Stable Virtual Camera (Seva), a generalist diffusion model that creates novel views of a scene, given any number of input views and target cameras. Existing works struggle to generate either large viewpoint changes or temporally smooth samples, while relying on specific task configurations. Our approach overcomes these limitations through simple model design, optimized training recipe, and flexible sampling strategy that generalize across view synthesis tasks at test time. As a result, our samples maintain high consistency without requiring additional 3D representation-based distillation, thus streamlining view synthesis in the wild. Furthermore, we show that our method can generate high-quality videos lasting up to half a minute with seamless loop closure. Extensive benchmarking demonstrates that Seva outperforms existing methods across different datasets and settings. Project page with code and model: https://stable-virtual-camera.github.io/.
Jensen Zhou, Vikram Voleti, Aaryaman Vasishta, Chun-Han Yao, Mark Boss, Philip Torr 0001, Christian Rupprecht 0001, Varun Jampani
ICCV5
2025 SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
abstract
We present Stable Video 4D (SV4D) — a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and novel view synthesis, we design a unified diffusion model to generate novel view videos of dynamic 3D objects. Specifically, given a monocular reference video, SV4D generates novel views for each video frame that are temporally consistent. We then use the generated novel view videos to optimize an implicit 4D representation (dynamic NeRF) efficiently, without the need for cumbersome SDS-based optimization used in most prior works. To train our unified novel view video generation model, we curate a dynamic 3D object dataset from the existing Objaverse dataset. Extensive experimental results on multiple datasets and user studies demonstrate SV4D's state-of-the-art performance on novel-view video synthesis as well as 4D generation compared to prior works. Project page: https://sv4d.github.io.
Chun-Han Yao, Vikram Voleti, Huaizu Jiang, Varun Jampani
ICLR2
2025 Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
abstract
We present Stable Part Diffusion 4D (SP4D), a framework for generating paired RGB and kinematic part videos from monocular inputs. Unlike conventional part segmentation methods that rely on appearance-based semantic cues, SP4D learns to produce kinematic parts --- structural components aligned with object articulation and consistent across views and time. SP4D adopts a dual-branch diffusion model that jointly synthesizes RGB frames and corresponding part segmentation maps. To simplify architecture and flexibly enable different part counts, we introduce a spatial color encoding scheme that maps part masks to continuous RGB-like images. This encoding allows the segmentation branch to share the latents VAE from the RGB branch, while enabling part segmentation to be recovered via straightforward post-processing. A Bidirectional Diffusion Fusion (BiDiFuse) module enhances cross-branch consistency, supported by a contrastive part consistency loss to promote spatial and temporal alignment of part predictions. We demonstrate that the generated 2D part maps can be lifted to 3D to derive skeletal structures and harmonic skinning weights with few manual adjustments. To train and evaluate SP4D, we construct KinematicParts20K, a curated dataset of over 20K rigged objects selected and processed from Objaverse XL, each paired with multi-view RGB and part video sequences. Experiments show that SP4D generalizes strongly to diverse scenarios, including real-world videos, novel generated objects, and rare articulated poses, producing kinematic-aware outputs suitable for downstream animation and motion-related tasks.
Hao Zhang 0122, Chun-Han Yao, Simon Donné, Narendra Ahuja, Varun Jampani
NeurIPS2
2024 ANIM: Accurate Neural Implicit Model for Human Reconstruction from a Single RGB-D Image
abstract
Recent progress in human shape learning, shows that neural implicit models are effective in generating 3D hu-man surfaces from limited number of views, and even from a single RGB image. However, existing monocular approaches still struggle to recover fine geometric details such as face, hands or cloth wrinkles. They are also easily prone to depth ambiguities that result in distorted geome-tries along the camera optical axis. In this paper, we ex-plore the benefits of incorporating depth observations in the reconstruction process by introducing ANIM, a novel method that reconstructs arbitrary 3D human shapes from single-view RGB-D images with an unprecedented level of accuracy. Our model learns geometric details from both multi-resolution pixel-aligned and voxel-aligned features to leverage depth information and enable spatial relation-ships, mitigating depth ambiguities. We further enhance the quality of the reconstructed shape by introducing a depth-supervision strategy, which improves the accuracy of the signed distance field estimation of points that lie on the re-constructed surface. Experiments demonstrate that ANIM outperforms state-of-the-art works that use RGB, surface normals, point cloud or RGB-D data as input. In addition, we introduce ANIM-Real, a new multi-modal dataset comprising highquality scans paired with consumer-grade RGB-D camera, and our protocol to fine-tune ANIM, enabling highquality reconstruction from real-world human capture. https://marcopesavento.github.io/Anim/
Marco Pesavento, Yuanlu Xu, Nikolaos Sarafianos, Robert Maier 0001, Chun-Han Yao, Marco Volino, Edmond Boyer, Adrian Hilton 0001, Tony Tung
CVPR6
2024 SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image Using Latent Video Diffusion
Vikram Voleti, Chun-Han Yao, Mark Boss, Adam Letts, David Pankratz, Dmitry Tochilkin, Christian Laforte, Robin Rombach, Varun Jampani
ECCV (1)2
2023 Hi-LASSIE: High-Fidelity Articulated Shape and Skeleton Discovery from Sparse Image Ensemble
abstract
Automatically estimating 3D skeleton, shape, camera viewpoints, and part articulation from sparse in-the-wild image ensembles is a severely under-constrained and challenging problem. Most prior methods rely on large-scale image datasets, dense temporal correspondence, or human annotations like camera pose, 2D keypoints, and shape templates. We propose Hi-LASSIE, which performs 3D articulated reconstruction from only 20–30 online images in the wild without any user-defined shape or skeleton templates. We follow the recent work of LASSIE that tackles a similar problem setting and make two significant advances. First, instead of relying on a manually annotated 3D skeleton, we automatically estimate a class-specific skeleton from the selected reference image. Second, we improve the shape reconstructions with novel instance-specific optimization strategies that allow reconstructions to faithful fit on each instance while preserving the class-specific priors learned across all images. Experiments on in-the-wild image ensembles show that Hi-LASSIE obtains higher fidelity state-of-the-art 3D reconstructions despite requiring minimum user input. Project page: chhankyao.github.io/hi-lassie/
Chun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein, Ming-Hsuan Yang 0001, Varun Jampani
CVPR1
2023 ARTIC3D: Learning Robust Articulated 3D Shapes from Noisy Web Image Collections
abstract
Estimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture, lighting, etc. We propose ARTIC3D, a self-supervised framework to reconstruct per-instance 3D shapes from a sparse image collection in-the-wild. Specifically, ARTIC3D is built upon a skeleton-based surface representation and is further guided by 2D diffusion priors from Stable Diffusion. First, we enhance the input images with occlusions/truncation via 2D diffusion to obtain cleaner mask estimates and semantic features. Second, we perform diffusion-guided 3D optimization to estimate shape and texture that are of high-fidelity and faithful to input images. We also propose a novel technique to calculate more stable image-level gradients via diffusion models compared to existing alternatives. Finally, we produce realistic animations by fine-tuning the rendered shape and texture under rigid part transformations. Extensive evaluations on multiple existing datasets as well as newly introduced noisy web image collections with occlusions and truncation demonstrate that ARTIC3D outputs are more robust to noisy images, higher quality in terms of shape and texture details, and more realistic when animated.
Chun-Han Yao, Amit Raj, Wei-Chih Hung, Michael Rubinstein, Yuanzhen Li, Ming-Hsuan Yang 0001, Varun Jampani
NeurIPS1
2022 Learning Visibility for Robust Dense Human Body Estimation
Chun-Han Yao, Jimei Yang, Duygu Ceylan, Yi Zhou 0023, Yang Zhou 0009, Ming-Hsuan Yang 0001
ECCV (1)1
2022 LASSIE: Learning Articulated Shapes from Sparse Image Ensemble via 3D Part Discovery
abstract
Creating high-quality articulated 3D models of animals is challenging either via manual creation or using 3D scanning tools. Therefore, techniques to reconstruct articulated 3D objects from 2D images are crucial and highly useful. In this work, we propose a practical problem setting to estimate 3D pose and shape of animals given only a few (10-30) in-the-wild images of a particular animal species (say, horse). Contrary to existing works that rely on pre-defined template shapes, we do not assume any form of 2D or 3D ground-truth annotations, nor do we leverage any multi-view or temporal information. Moreover, each input image ensemble can contain animal instances with varying poses, backgrounds, illuminations, and textures. Our key insight is that 3D parts have much simpler shape compared to the overall animal and that they are robust w.r.t. animal pose articulations. Following these insights, we propose LASSIE, a novel optimization framework which discovers 3D parts in a self-supervised manner with minimal user intervention. A key driving force behind LASSIE is the enforcing of 2D-3D part consistency using self-supervisory deep features. Experiments on Pascal-Part and self-collected in-the-wild animal datasets demonstrate considerably better 3D reconstructions as well as both 2D and 3D part discovery compared to prior arts. Project page: https://chhankyao.github.io/lassie/
Chun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein, Ming-Hsuan Yang 0001, Varun Jampani
NeurIPS1
2022 Federated Multi-Target Domain Adaptation
abstract
Federated learning methods enable us to train machine learning models on distributed user data while preserving its privacy. However, it is not always feasible to obtain high-quality supervisory signals from users, especially for vision tasks. Unlike typical federated settings with labeled client data, we consider a more practical scenario where the distributed client data is unlabeled, and a centralized labeled dataset is available on the server. We further take the server-client and inter-client domain shifts into account and pose a domain adaptation problem with one source (centralized server data) and multiple targets (distributed client data). Within this new Federated Multi-Target Domain Adaptation (FMTDA) task, we analyze the model performance of existing domain adaptation methods and propose an effective DualAdapt method to address the new challenges. Extensive experimental results on image classification and semantic segmentation tasks demonstrate that our method achieves high accuracy, incurs minimal communication cost, and requires low computational resources on client devices.
Chun-Han Yao, Boqing Gong, Yin Cui, Yukun Zhu, Ming-Hsuan Yang 0001
WACV1
2021 Discovering 3D Parts from Image Collections
abstract
Reasoning 3D shapes from 2D images is an essential yet challenging task, especially when only single-view images are at our disposal. While an object can have a complicated shape, individual parts are usually close to geometric primitives and thus are easier to model. Furthermore, parts provide a mid-level representation that is robust to appearance variations across objects in a particular category. In this work, we tackle the problem of 3D part discovery from only 2D image collections. Instead of relying on manually annotated parts for supervision, we propose a self-supervised approach, latent part discovery (LPD). Our key insight is to learn a novel part shape prior that allows each part to fit an object shape faithfully while constrained to have simple geometry. Extensive experiments on the synthetic ShapeNet, PartNet, and real-world Pascal 3D+ datasets show that our method discovers consistent object parts and achieves favorable reconstruction accuracy compared to the existing methods with the same level of supervision. Our project page with code is at https://chhankyao.github.io/lpd/.
Chun-Han Yao, Wei-Chih Hung, Varun Jampani, Ming-Hsuan Yang 0001
ICCV1
2020 Video Object Detection via Object-Level Temporal Aggregation
Chun-Han Yao, Xiaohui Shen, Yangyue Wan, Ming-Hsuan Yang 0001
ECCV (14)1
2020 Progressive Domain Adaptation for Object Detection
abstract
Recent deep learning methods for object detection rely on a large amount of bounding box annotations. Collecting these annotations is laborious and costly, yet supervised models do not generalize well when testing on images from a different distribution. Domain adaptation provides a solution by adapting existing labels to the target testing data. However, a large gap between domains could make adaptation a challenging task, which leads to unstable training processes and sub-optimal results. In this paper, we propose to bridge the domain gap with an intermediate domain and progressively solve easier adaptation subtasks. This intermediate domain is constructed by translating the source images to mimic the ones in the target domain. To tackle the domain-shift problem, we adopt adversarial learning to align distributions at the feature level. In addition, a weighted task loss is applied to deal with unbalanced image quality in the intermediate domain. Experimental results show that our method performs favorably against the state-of-the-art method in terms of the performance on the target domain.
Han-Kai Hsu, Chun-Han Yao, Yi-Hsuan Tsai, Wei-Chih Hung, Hung-Yu Tseng, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
WACV2
2017 Occlusion-aware Video Temporal Consistency
abstract
Image color editing techniques such as color transfer, HDR tone mapping, dehazing, and white balance have been widely used and investigated in recent decades. However, naively employing them to videos frame-by-frame often leads to flickering or color inconsistency. To solve it generally, earlier methods rely on temporal filtering or warping from the previous frame, but they still fail in the cases of occlusion and produce blurry results. We introduce a new framework for these challenges: (1) We develop an online keyframe strategy to keep track of the dynamic objects, where more temporal information can be acquired than a single previous frame. (2) To preserve image details, local color affine model is employed. The main concept of this post-processing step is to capture the color transformation from editing algorithms and maintain the detail structures of the raw image simultaneously. Practically, our approach takes a raw video and its per-frame processed version, and generates a temporally consistent output. In addition, we propose a video quality metric to evaluate temporal coherence. Extensive experiments and subjective test are done to show the superiority of the proposed framework with respect to color fidelity, detail preservation, and temporal consistency.
Chun-Han Yao, Chia-Yang Chang, Shao-Yi Chien
ACM Multimedia1
2017 Outage reduction with joint scheduling and power allocation in 5G mmWave cellular networks
abstract
Millimeter-wave (mmWave) communications is a promising technology which supports high datarates (multi-Gbps) by utilizing high bandwidth and the directional antenna. While the directionality reduces interference significantly and compensates the high propagation loss, it brings about two major problems. Firstly, mmWave links are easily blocked by obstacles like human bodies and buildings. Secondly, user mobility can frequently cause misalignments between transmitter and receiver beams, which is known as the deafness problem. In this paper, these problems are addressed and a joint scheduling and power allocation framework is proposed to reduce the outage probability during user movement. Extensive simulations are done to demonstrate the pros and cons of the proposed algorithms and the improvement of system performance.
Chun-Han Yao, Yin-Yi Chen, B. P. S. Sahoo, Hung-Yu Wei 0001
PIMRC1
2017 Millimeter-Wave Multi-Hop Wireless Backhauling for 5G Cellular Networks
abstract
The millimeter-wave (mmWave) bands, roughly referred to 30-300 GHz, have been widely recognized as a promising candidate for dense deployment of small- cells backhaul network. The dense small-cell deployment produces a huge amount of backhaul traffic, and directional communication poses a significant challenge. In this paper, we propose a dynamic frame reconfiguration scheme which provides greater flexibility for dynamic traffic adaptation in multi-hop mmWave relay backhaul. This allows better exploitation of the traffic dynamics in smallcell. In addition, we present a traffic load and link-quality aware multi-hop relay backhaul scheduling algorithm to maximize the overall system performance. The extensive simulation results demonstrate the superiority of our proposed algorithm when compared with other schemes.
B. P. S. Sahoo, Chun-Han Yao, Hung-Yu Wei 0001
VTC Spring2
2016 Example-based video color transfer
abstract
Color transfer is an image processing technique commonly used to fix images with wrong colors, enhance the lighting conditions, or produce special styles to express specific emotions. With the aid of a reference image, the intended color characteristics can be properly transferred to the source images or videos. Since the applications start growing popular, many image color transfer methods have emerged; However, only few video color transfer algorithms have been proposed so far. In this paper, we propose an efficient example-based color transfer algorithm for both images and videos, which preserves the gradient details of the source by building Laplacian pyramids and improves the spatio-temporal consistency by employing patch matching method. Other than the existing quality metrics, we further propose a video quality metric regarding temporal consistency and compare our algorithm with other color transfer methods. The experimental results show that our algorithm generally produces outputs with high fidelity in terms of colors, scene details, and also spatiotemporal consistency.
Chun-Han Yao, Chia-Yang Chang, Shao-Yi Chien
ICME1