Yen-Chi Cheng

dblp:239/4170 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-0639-3154ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 63% 3D vision · 37%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
1.222023
SDFusion: Multimodal 3D Shape Completion, Reconstruction, and Generation · CVPR 2023
AutoSDF: Shape Priors for 3D Completion, Reconstruction and Generation · CVPR 2022
Machine learning › Generative modeling › diffusion model
3d shape generation
1.222023
SDFusion: Multimodal 3D Shape Completion, Reconstruction, and Generation · CVPR 2023
AutoSDF: Shape Priors for 3D Completion, Reconstruction and Generation · CVPR 2022
Machine learning › Generative modeling
diffusion model
0.712023
SDFusion: Multimodal 3D Shape Completion, Reconstruction, and Generation · CVPR 2023
Computer vision › 3D vision › 3d reconstruction
image-based 3d reconstruction
0.712023
SDFusion: Multimodal 3D Shape Completion, Reconstruction, and Generation · CVPR 2023
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.712023
SDFusion: Multimodal 3D Shape Completion, Reconstruction, and Generation · CVPR 2023
Machine learning › Generative modeling › generative adversarial network
GAN inversion
0.612022
InOut: Diverse Image Outpainting via GAN Inversion · CVPR 2022
Machine learning › Generative modeling
generative adversarial network
0.612022
InOut: Diverse Image Outpainting via GAN Inversion · CVPR 2022
Machine learning › Generative modeling
image generation
0.612022
InfinityGAN: Towards Infinite-Pixel Image Synthesis · ICLR 2022
Computer vision › 3D vision › 3d shape reconstruction
shape completion
0.612022
AutoSDF: Shape Priors for 3D Completion, Reconstruction and Generation · CVPR 2022
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
0.612022
AutoSDF: Shape Priors for 3D Completion, Reconstruction and Generation · CVPR 2022
Visual content generation and editing
image outpainting
0.612022
InOut: Diverse Image Outpainting via GAN Inversion · CVPR 2022
Machine learning › Generative modeling
variational autoencoder
0.412020
Controllable Image Synthesis via SegVAE · ECCV (7) 2020
Visual content generation and editing › image generation
controllable image generation
0.412020
Controllable Image Synthesis via SegVAE · ECCV (7) 2020
Machine learning › Generative modeling
video generation
0.412019
Point-to-Point Video Generation · ICCV 2019

Methods — techniques the papers use, named apart from their topics

patch-based generation · 1.1latent code optimization · 1.1semantic segmentation · 0.9VAE · 0.9text-to-image model · 0.7task-specific encoders · 0.7cross-attention · 0.7latent representation learning · 0.6generative adversarial network · 0.6autoregressive modeling · 0.6
YearPublicationVenuePosition
2025 DreaMo: Articulated 3D Reconstruction from a Single Casual Video
abstract
Articulated 3D reconstruction has valuable applications in various domains, yet it remains costly and demands intensive work from domain experts. Recent advancements in template-free learning methods show promising results with monocular videos. Nevertheless, these approaches necessitate a comprehensive coverage of all viewpoints of the subject in the input video, thus limiting their applicability to casually captured videos from online sources. In this work, we study articulated 3D shape reconstruction from a single and casually captured Internet video, where the subject's view coverage is incomplete. We propose DreaMo that jointly performs shape reconstruction while solving the challenging low-coverage regions with view-conditioned diffusion prior and several tailored regularizations. In addition, we introduce a skeleton generation strategy to create human-interpretable skeletons from the learned neural bones and skinning weights without any predefined skeleton structures. We conduct our study on a self-collected internet video collection characterized by incomplete view coverage. DreaMo shows promising quality in novel-view rendering, detailed articulated shape reconstruction, and skeleton generation. Extensive qualitative and quantitative studies validate the efficacy of each proposed component, and show existing methods are unable to solve correct geometry due to the incomplete view coverage.
Tao Tu 0002, Ming-Feng Li, Chieh Hubert Lin, Yen-Chi Cheng, Min Sun 0001, Ming-Hsuan Yang 0001
WACV4
2023 SDFusion: Multimodal 3D Shape Completion, Reconstruction, and Generation
abstract
In this work, we present a novel framework built to sim-plify 3D asset generation for amateur users. To enable interactive generation, our method supports a variety of input modalities that can be easily provided by a human, in-cluding images, text, partially observed shapes and combinations of these, further allowing to adjust the strength of each input. At the core of our approach is an encoder-decoder, compressing 3D shapes into a compact latent representation, upon which a diffusion model is learned. To enable a variety of multimodal inputs, we employ task-specific encoders with dropout followed by a cross-attention mechanism. Due to its flexibility, our model naturally supports a variety of tasks, outperforming prior works on shape completion, image-based 3D reconstruction, and text-to-3D. Most interestingly, our model can combine all these tasks into one swiss-army-knife tool, enabling the user to perform shape generation using incomplete shapes, images, and textual descriptions at the same time, providing the relative weights for each input and facilitating interactivity. Despite our approach being shape-only, we further show an efficient method to texture the generated shape using large-scale text-to-image models.
Yen-Chi Cheng, Hsin-Ying Lee 0001, Sergey Tulyakov, Alexander G. Schwing, Liangyan Gui
CVPR1
2022 InOut: Diverse Image Outpainting via GAN Inversion
abstract
Image outpainting seeks for a semantically consistent extension of the input image beyond its available content. Compared to inpainting - filling in missing pixels in a way coherent with the neighboring pixels - outpainting can be achieved in more diverse ways since the problem is less constrained by the surrounding pixels. Existing image outpainting methods pose the problem as a conditional image-to-image translation task, often generating repetitive structures and textures by replicating the content available in the input image. In this work, we formulate the problem from the perspective of inverting generative adversarial networks. Our generator renders micro-patches conditioned on their joint latent code as well as their individual positions in the image. To outpaint an image, we seek for multiple latent codes not only recovering available patches but also synthesizing diverse outpainting by patch-based generation. This leads to richer structure and content in the outpainted regions. Furthermore, our formulation allows for outpainting conditioned on the categorical input, thereby enabling flexible user controls. Extensive experimental results demonstrate the proposed method performs favorably against existing in- and outpainting methods, featuring higher visual quality and diversity.
Yen-Chi Cheng, Chieh Hubert Lin, Hsin-Ying Lee 0001, Jian Ren 0005, Sergey Tulyakov, Ming-Hsuan Yang 0001
CVPR1
2022 AutoSDF: Shape Priors for 3D Completion, Reconstruction and Generation
abstract
Powerful priors allow us to perform inference with in-sufficient information. In this paper, we propose an au-toregressive prior for 3D shapes to solve multimodal 3D tasks such as shape completion, reconstruction, and gener-ation. We model the distribution over 3D shapes as a non-sequential autoregressive distribution over a discretized, low-dimensional, symbolic grid-like latent representation of 3D shapes. This enables us to represent distributions over 3D shapes conditioned on information from an arbitrary set of spatially anchored query locations and thus perform shape completion in such arbitrary settings (e.g. generating a complete chair given only a view of the back leg). We also show that the learned autoregressive prior can be leveraged for conditional tasks such as single-view reconstruction and language-based generation. This is achieved by learning task-specific ‘naive’ conditionals which can be approxi-mated by light-weight models trained on minimal paired data. We validate the effectiveness of the proposed method using both quantitative and qualitative evaluation and show that the proposed method outperforms the specialized state-of-the-art methods trained for individual tasks. The project page with code and video visualizations can be found at https://yccyenchicheng.github.io/AutoSDF/.
Paritosh Mittal, Yen-Chi Cheng, Maneesh Kumar Singh 0001, Shubham Tulsiani
CVPR2
2022 InfinityGAN: Towards Infinite-Pixel Image Synthesis
Chieh Hubert Lin, Hsin-Ying Lee 0001, Yen-Chi Cheng, Sergey Tulyakov, Ming-Hsuan Yang 0001
ICLR3
2020 Controllable Image Synthesis via SegVAE
Yen-Chi Cheng, Hsin-Ying Lee 0001, Min Sun 0001, Ming-Hsuan Yang 0001
ECCV (7)1
2019 Point-to-Point Video Generation
abstract
While image synthesis achieves tremendous breakthroughs (e.g., generating realistic faces), video generation is less explored and harder to control, which limits its applications in the real world. For instance, video editing requires temporal coherence across multiple clips and thus poses both start and end constraints within a video sequence. We introduce point-to-point video generation that controls the generation process with two control points: the targeted start- and end-frames. The task is challenging since the model not only generates a smooth transition of frames but also plans ahead to ensure that the generated end-frame conforms to the targeted end-frame for videos of various lengths. We propose to maximize the modified variational lower bound of conditional data likelihood under a skip-frame training strategy. Our model can generate end-frame-consistent sequences without loss of quality and diversity. We evaluate our method through extensive experiments on Stochastic Moving MNIST, Weizmann Action, Human3.6M, and BAIR Robot Pushing under a series of scenarios. The qualitative results showcase the effectiveness and merits of point-to-point generation.
Tsun-Hsuan Wang, Yen-Chi Cheng, Chieh Hubert Lin, Hwann-Tzong Chen, Min Sun 0001
ICCV2