Chieh Hubert Lin

dblp:225/5489 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0003-2417-9992ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
3D vision · 46% Generative modeling · 33% Deep learning architectures and training · 12%
Computer graphics and multimedia
5 papers
Visual content generation and editing · 35% Rendering · 33% Computer animation and physical simulation · 16%

Topics — the 25 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
1.722025
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos · NeurIPS 2025
Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video · NeurIPS 2025
Machine learning › Generative modeling
generative adversarial network
1.332022
InOut: Diverse Image Outpainting via GAN Inversion · CVPR 2022
COCO-GAN: Generation by Parts via Conditional Coordinating · ICCV 2019
Escaping from Collapsing Modes in a Constrained Space · ECCV (7) 2018
Machine learning › Generative modeling
image generation
1.022022
InfinityGAN: Towards Infinite-Pixel Image Synthesis · ICLR 2022
COCO-GAN: Generation by Parts via Conditional Coordinating · ICCV 2019
Computer vision › 3D vision
3d scene reconstruction
0.912025
InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model · NeurIPS 2025
Computer vision › 3D vision › 3d scene reconstruction
deformable scene reconstruction
0.912025
Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video · NeurIPS 2025
Computer vision › 3D vision › 3d reconstruction › learning-based 3d reconstruction
feed-forward reconstruction
0.912025
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos · NeurIPS 2025
Computer vision › 3D vision › 3d reconstruction › learning-based 3d reconstruction
large reconstruction model
0.912025
InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model · NeurIPS 2025
Geometric modeling and processing
3d reconstruction
0.912025
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos · NeurIPS 2025
Rendering › gaussian splatting
deformable gaussian splatting
0.912025
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos · NeurIPS 2025
Computer animation and physical simulation
motion synthesis
0.912025
Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video · NeurIPS 2025
Rendering
neural radiance fields
0.812024
Taming Latent Diffusion Model for Neural Radiance Field Inpainting · ECCV (3) 2024
Machine learning › Deep learning architectures and training
convolutional neural network
0.712023
Unveiling The Mask of Position-Information Pattern Through the Mist of Image Features · ICML 2023
Machine learning › Deep learning architectures and training
positional encoding
0.712023
Unveiling The Mask of Position-Information Pattern Through the Mist of Image Features · ICML 2023
Visual content generation and editing › 3d content generation
3d city generation
0.712023
InfiniCity: Infinite-Scale City Synthesis · ICCV 2023
Visual content generation and editing
3d content creation
0.712023
InfiniCity: Infinite-Scale City Synthesis · ICCV 2023
Machine learning › Generative modeling › generative adversarial network
GAN inversion
0.612022
InOut: Diverse Image Outpainting via GAN Inversion · CVPR 2022
Visual content generation and editing
image outpainting
0.612022
InOut: Diverse Image Outpainting via GAN Inversion · CVPR 2022
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.412020
InstaNAS: Instance-Aware Neural Architecture Search · AAAI 2020
Machine learning › Generative modeling
video generation
0.412019
Point-to-Point Video Generation · ICCV 2019
Machine learning › Generative modeling › generative adversarial network › GAN training
mode collapse
0.312018
Escaping from Collapsing Modes in a Constrained Space · ECCV (7) 2018
Computer vision › Video understanding and tracking › object tracking
3d object tracking
0.312025
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.212024
Taming Latent Diffusion Model for Neural Radiance Field Inpainting · ECCV (3) 2024
Rendering
neural rendering
0.212023
InfiniCity: Infinite-Scale City Synthesis · ICCV 2023
Machine learning › Efficient and distributed learning › inference efficiency
latency-aware inference
0.112020
InstaNAS: Instance-Aware Neural Architecture Search · AAAI 2020
Machine learning › Efficient and distributed learning
model compression
0.112020
InstaNAS: Instance-Aware Neural Architecture Search · AAAI 2020

Methods — techniques the papers use, named apart from their topics

video-rewinding training · 1.7transformer · 1.7occlusion-aware ARAP regularization · 1.7large reconstruction model · 1.7disocclusion backtracing · 1.73d gaussian splatting · 1.7latent diffusion model · 1.5patch-based generation · 1.0masked finetuning · 0.9feed-forward reconstruction · 0.9octree-based voxel completion · 0.7neural rendering · 0.7infinite-pixel image synthesis · 0.7
YearPublicationVenuePosition
2025 Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video
abstract
Motion is one of the key components in deformable 3D scenes. Generative video models allow users to animate static scenes with text prompts for novel motion, but when it comes to 4D reconstruction, such reanimations often fall apart. The generated videos often suffer from geometric artifacts, implausible motion, and occlusions, which hinder physically consistent 4D reanimation. In this work, we introduce \textbf{Restage4D}, a geometry-preserving pipeline for deformable scene reconstruction from a single edited video. Our key insight is to leverage the unedited original video as an additional source of supervision, allowing the model to propagate accurate structure into occluded and disoccluded regions. To achieve this, we propose a video-rewinding training scheme that temporally bridges the edited and original sequences via a shared motion representation. We further introduce an occlusion-aware ARAP regularization to preserve local rigidity, and a disocclusion backtracing mechanism that supplements missing geometry in the canonical space. Together, these components enable robust reconstruction even when the edited input contains hallucinated content or inconsistent motion. We validate Restage4D on DAVIS and PointOdyssey, demonstrating improved geometry consistency, motion quality, and 3D tracking performance. Our method not only preserves deformable structure under novel motion, but also automatically corrects errors introduced by generative models, bridging the gap between flexible video synthesis and physically grounded 4D reconstruction.
Jixuan He, Chieh Hubert Lin, Lu Qi 0001, Ming-Hsuan Yang 0001
NeurIPS2
2025 DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
abstract
We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction has gained significant attention for its ability to rapidly create digital replicas of real-world environments. However, most existing models are limited to static scenes and fail to reconstruct the motion of moving objects. Developing a feed-forward model for dynamic scene reconstruction poses significant challenges, including the scarcity of training data and the need for appropriate 3D representations and training paradigms. To address these challenges, we introduce several key technical contributions: an enhanced large-scale synthetic dataset with ground-truth multi-view videos and dense 3D scene flow supervision; a per-pixel deformable 3D Gaussian representation that is easy to learn, supports high-quality dynamic view synthesis, and enables long-range 3D tracking; and a large transformer network that achieves real-time, generalizable dynamic scene reconstruction. Extensive qualitative and quantitative experiments demonstrate that DGS-LRM achieves dynamic scene reconstruction quality comparable to optimization-based methods, while significantly outperforming the state-of-the-art predictive dynamic reconstruction method on real-world examples. Its predicted physically grounded 3D deformation is accurate and can be readily adapted for long-range 3D tracking tasks, achieving performance on par with state-of-the-art monocular video 3D tracking methods.
Chieh Hubert Lin, Zhaoyang Lv, Songyin Wu, Thu Nguyen-Phuoc, Hung-Yu Tseng, Julian Straub, Numair Khan, Lei Xiao 0014, Ming-Hsuan Yang 0001, Yuheng Ren, Richard A. Newcombe, Zhao Dong 0001, Zhengqin Li
NeurIPS1
2025 InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model
abstract
Recent advances in 3D scene reconstruction enable real-time viewing in virtual and augmented reality. To support interactive operations for better immersiveness, such as moving or editing objects, 3D scene inpainting methods are proposed to repair or complete the altered geometry. To support users in interacting (such as moving or editing objects) with the scene for the next level of immersiveness, 3D scene inpainting methods are developed to repair the altered geometry. However, current approaches rely on lengthy and computationally intensive optimization, making them impractical for real-time or online applications. We propose InstaInpaint, a reference-based feed-forward framework that produces 3D-scene inpainting from a 2D inpainting proposal within 0.4 seconds. We develop a self-supervised masked-finetuning strategy to enable training of our custom large reconstruction model (LRM) on the large-scale dataset. Through extensive experiments, we analyze and identify several key designs that improve generalization, textural consistency, and geometric correctness. InstaInpaint achieves a 1000$\times$ speed-up from prior methods while maintaining a state-of-the-art performance across two standard benchmarks. Moreover, we show that InstaInpaint generalizes well to flexible downstream applications such as object insertion and multi-region inpainting.
Junqi You, Chieh Hubert Lin, Weijie Lyu, Zhengbo Zhang, Ming-Hsuan Yang 0001
NeurIPS2
2025 DreaMo: Articulated 3D Reconstruction from a Single Casual Video
abstract
Articulated 3D reconstruction has valuable applications in various domains, yet it remains costly and demands intensive work from domain experts. Recent advancements in template-free learning methods show promising results with monocular videos. Nevertheless, these approaches necessitate a comprehensive coverage of all viewpoints of the subject in the input video, thus limiting their applicability to casually captured videos from online sources. In this work, we study articulated 3D shape reconstruction from a single and casually captured Internet video, where the subject's view coverage is incomplete. We propose DreaMo that jointly performs shape reconstruction while solving the challenging low-coverage regions with view-conditioned diffusion prior and several tailored regularizations. In addition, we introduce a skeleton generation strategy to create human-interpretable skeletons from the learned neural bones and skinning weights without any predefined skeleton structures. We conduct our study on a self-collected internet video collection characterized by incomplete view coverage. DreaMo shows promising quality in novel-view rendering, detailed articulated shape reconstruction, and skeleton generation. Extensive qualitative and quantitative studies validate the efficacy of each proposed component, and show existing methods are unable to solve correct geometry due to the incomplete view coverage.
Tao Tu 0002, Ming-Feng Li, Chieh Hubert Lin, Yen-Chi Cheng, Min Sun 0001, Ming-Hsuan Yang 0001
WACV3
2024 Taming Latent Diffusion Model for Neural Radiance Field Inpainting
Chieh Hubert Lin, Changil Kim 0001, Jia-Bin Huang 0001, Qinbo Li, Chih-Yao Ma, Johannes Kopf 0001, Ming-Hsuan Yang 0001, Hung-Yu Tseng
ECCV (3)1
2023 InfiniCity: Infinite-Scale City Synthesis
abstract
Toward infinite-scale 3D city synthesis, we propose a novel framework, InfiniCity, which constructs and renders an unconstrainedly large and 3D-grounded environment from random noises. InfiniCity decomposes the seemingly impractical task into three feasible modules, taking advantage of both 2D and 3D data. First, an infinite-pixel image synthesis module generates arbitrary-scale 2D maps from the bird’s-eye view. Next, an octree-based voxel completion module lifts the generated 2D map to 3D octrees. Finally, a voxel-based neural rendering module texturizes the voxels and renders 2D images. InfiniCity can thus synthesize arbitrary-scale and traversable 3D city environments. We quantitatively and qualitatively demonstrate the efficacy of the proposed framework.
Chieh Hubert Lin, Hsin-Ying Lee 0001, Willi Menapace, Menglei Chai, Aliaksandr Siarohin, Ming-Hsuan Yang 0001, Sergey Tulyakov
ICCV1
2023 Unveiling The Mask of Position-Information Pattern Through the Mist of Image Features
abstract
Recent studies have shown that paddings in convolutional neural networks encode absolute position information which can negatively affect the model performance for certain tasks. However, existing metrics for quantifying the strength of positional information remain unreliable and frequently lead to erroneous results. To address this issue, we propose novel metrics for measuring and visualizing the encoded positional information. We formally define the encoded information as Position-information Pattern from Padding (PPP) and conduct a series of experiments to study its properties as well as its formation. The proposed metrics measure the presence of positional information more reliably than the existing metrics based on PosENet and tests in F-Conv. We also demonstrate that for any extant (and proposed) padding schemes, PPP is primarily a learning artifact and is less dependent on the characteristics of the underlying padding schemes.
Chieh Hubert Lin, Hung-Yu Tseng, Hsin-Ying Lee 0001, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
ICML1
2022 InOut: Diverse Image Outpainting via GAN Inversion
abstract
Image outpainting seeks for a semantically consistent extension of the input image beyond its available content. Compared to inpainting - filling in missing pixels in a way coherent with the neighboring pixels - outpainting can be achieved in more diverse ways since the problem is less constrained by the surrounding pixels. Existing image outpainting methods pose the problem as a conditional image-to-image translation task, often generating repetitive structures and textures by replicating the content available in the input image. In this work, we formulate the problem from the perspective of inverting generative adversarial networks. Our generator renders micro-patches conditioned on their joint latent code as well as their individual positions in the image. To outpaint an image, we seek for multiple latent codes not only recovering available patches but also synthesizing diverse outpainting by patch-based generation. This leads to richer structure and content in the outpainted regions. Furthermore, our formulation allows for outpainting conditioned on the categorical input, thereby enabling flexible user controls. Extensive experimental results demonstrate the proposed method performs favorably against existing in- and outpainting methods, featuring higher visual quality and diversity.
Yen-Chi Cheng, Chieh Hubert Lin, Hsin-Ying Lee 0001, Jian Ren 0005, Sergey Tulyakov, Ming-Hsuan Yang 0001
CVPR2
2022 InfinityGAN: Towards Infinite-Pixel Image Synthesis
Chieh Hubert Lin, Hsin-Ying Lee 0001, Yen-Chi Cheng, Sergey Tulyakov, Ming-Hsuan Yang 0001
ICLR1
2020 InstaNAS: Instance-Aware Neural Architecture Search
abstract
Conventional Neural Architecture Search (NAS) aims at finding a single architecture that achieves the best performance, which usually optimizes task related learning objectives such as accuracy. However, a single architecture may not be representative enough for the whole dataset with high diversity and variety. Intuitively, electing domain-expert architectures that are proficient in domain-specific features can further benefit architecture related objectives such as latency. In this paper, we propose InstaNAS—an instance-aware NAS framework—that employs a controller trained to search for a “distribution of architectures” instead of a single architecture; This allows the model to use sophisticated architectures for the difficult samples, which usually comes with large architecture related cost, and shallow architectures for those easy samples. During the inference phase, the controller assigns each of the unseen input samples with a domain expert architecture that can achieve high accuracy with customized inference costs. Experiments within a search space inspired by MobileNetV2 show InstaNAS can achieve up to 48.8% latency reduction without compromising accuracy on a series of datasets against MobileNetV2.
An-Chieh Cheng, Chieh Hubert Lin, Da-Cheng Juan, Wei Wei 0019, Min Sun 0001
AAAI2
2019 COCO-GAN: Generation by Parts via Conditional Coordinating
abstract
Humans can only interact with part of the surrounding environment due to biological restrictions. Therefore, we learn to reason the spatial relationships across a series of observations to piece together the surrounding environment. Inspired by such behavior and the fact that machines also have computational constraints, we propose COnditional COordinate GAN (COCO-GAN) of which the generator generates images by parts based on their spatial coordinates as the condition. On the other hand, the discriminator learns to justify realism across multiple assembled patches by global coherence, local appearance, and edge-crossing continuity. Despite the full images are never manipulated during training, we show that COCO-GAN can produce state-of-the-art-quality full images during inference. We further demonstrate a variety of novel applications enabled by our coordinate-aware framework. First, we perform extrapolation to the learned coordinate manifold and generate off-the-boundary patches. Combining with the originally generated full image, COCO-GAN can produce images that are larger than training samples, which we called "beyond-boundary generation". We then showcase panorama generation within a cylindrical coordinate system that inherently preserves horizontally cyclic topology. On the computation side, COCO-GAN has a built-in divide-and-conquer paradigm that reduces memory requisition during training and inference, provides high-parallelism, and can generate parts of images on-demand.
Chieh Hubert Lin, Chia-Che Chang, Yu-Sheng Chen, Da-Cheng Juan, Wei Wei 0019, Hwann-Tzong Chen
ICCV1
2019 Point-to-Point Video Generation
abstract
While image synthesis achieves tremendous breakthroughs (e.g., generating realistic faces), video generation is less explored and harder to control, which limits its applications in the real world. For instance, video editing requires temporal coherence across multiple clips and thus poses both start and end constraints within a video sequence. We introduce point-to-point video generation that controls the generation process with two control points: the targeted start- and end-frames. The task is challenging since the model not only generates a smooth transition of frames but also plans ahead to ensure that the generated end-frame conforms to the targeted end-frame for videos of various lengths. We propose to maximize the modified variational lower bound of conditional data likelihood under a skip-frame training strategy. Our model can generate end-frame-consistent sequences without loss of quality and diversity. We evaluate our method through extensive experiments on Stochastic Moving MNIST, Weizmann Action, Human3.6M, and BAIR Robot Pushing under a series of scenarios. The qualitative results showcase the effectiveness and merits of point-to-point generation.
Tsun-Hsuan Wang, Yen-Chi Cheng, Chieh Hubert Lin, Hwann-Tzong Chen, Min Sun 0001
ICCV3
2019 3D LiDAR and Stereo Fusion using Stereo Matching Network with Conditional Cost Volume Normalization
abstract
The complementary characteristics of active and passive depth sensing techniques motivate the fusion of the LiDAR sensor and stereo camera for improved depth perception. Instead of directly fusing estimated depths across LiDAR and stereo modalities, we take advantages of the stereo matching network with two enhanced techniques: Input Fusion and Conditional Cost Volume Normalization (CCVNorm) on the LiDAR information. The proposed framework is generic and closely integrated with the cost volume component that is commonly utilized in stereo matching neural networks. We experimentally verify the efficacy and robustness of our method on the KITTI Stereo and Depth Completion datasets, obtaining favorable performance against various fusion strategies. Moreover, we demonstrate that, with a hierarchical extension of CCVNorm, the proposed method brings only slight overhead to the stereo matching network in terms of computation time and model size.
Tsun-Hsuan Wang, Hou-Ning Hu, Chieh Hubert Lin, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Min Sun 0001
IROS3
2018 Escaping from Collapsing Modes in a Constrained Space
Chia-Che Chang, Chieh Hubert Lin, Che-Rung Lee, Da-Cheng Juan, Wei Wei 0019, Hwann-Tzong Chen
ECCV (7)2