Yao-Chih Lee

dblp:295/9855 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0006-8787-9397ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Video understanding and tracking · 43% 3D vision · 38% Generative modeling · 19%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 41% Rendering · 23% Computer animation and physical simulation · 20%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
Generative Omnimatte: Learning to Decompose Video into Layers · CVPR 2025
Computer vision › Video understanding and tracking › video reconstruction
video inpainting
0.912025
Generative Omnimatte: Learning to Decompose Video into Layers · CVPR 2025
Computer vision › Video understanding and tracking
video layer decomposition
0.912025
Generative Omnimatte: Learning to Decompose Video into Layers · CVPR 2025
Computer vision › 3D vision
novel view synthesis
0.812024
Fast View Synthesis of Casual Videos with Soup-of-Planes · ECCV (38) 2024
Rendering
novel view synthesis
0.812024
Fast View Synthesis of Casual Videos with Soup-of-Planes · ECCV (38) 2024
Visual content generation and editing › video editing
text-driven video editing
0.712023
Shape-Aware Text-Driven Layered Video Editing · CVPR 2023
Visual content generation and editing
video editing
0.712023
Shape-Aware Text-Driven Layered Video Editing · CVPR 2023
Computer vision › 3D vision
depth estimation
0.512021
3D Video Stabilization With Depth Estimation by CNN-Based Optimization · CVPR 2021
Computer vision › 3D vision › depth estimation › self-supervised depth estimation
self-supervised depth and ego-motion estimation
0.512021
3D Video Stabilization With Depth Estimation by CNN-Based Optimization · CVPR 2021
Image and video processing
video stabilization
0.512021
3D Video Stabilization With Depth Estimation by CNN-Based Optimization · CVPR 2021

Methods — techniques the papers use, named apart from their topics

soup-of-planes representation · 1.5self-supervised learning · 1.03d reconstruction · 1.0object mask conditioning · 0.9diffusion model fine-tuning · 0.9diffusion model · 0.7deformation field · 0.7UV mapping · 0.7
YearPublicationVenuePosition
2025 Generative Omnimatte: Learning to Decompose Video into Layers
abstract
Given a video and a set of input object masks, an omnimatte method aims to decompose the video into semantically meaningful layers containing individual objects along with their associated effects, such as shadows and reflections. Existing omnimatte methods assume a static background or accurate pose and depth estimation and produce poor decompositions when these assumptions are violated. Furthermore, due to the lack of generative prior on natural videos, existing methods cannot complete dynamic occluded regions. We present a novel generative layered video decomposition framework to address the omnimatte problem. Our method does not assume a stationary scene or require camera pose or depth information and produces clean, complete layers, including convincing completions of occluded dynamic regions. Our core idea is to train a video diffusion model to identify and remove scene effects caused by a specific object. We show that this model can be finetuned from an existing video inpainting model with a small, carefully curated dataset, and demonstrate high-quality decompositions and editing results for a wide range of casually captured videos containing soft shadows, glossy reflections, splashing water, and more.
Yao-Chih Lee, Erika Lu, Sarah Rumbley, Michal Geyer, Jia-Bin Huang 0001, Tali Dekel, Forrester Cole
CVPR1
2024 Fast View Synthesis of Casual Videos with Soup-of-Planes
Yao-Chih Lee, Zhoutong Zhang, Kevin Matzen, Simon Niklaus, Jianming Zhang 0001, Jia-Bin Huang 0001, Feng Liu 0015
ECCV (38)1
2023 Shape-Aware Text-Driven Layered Video Editing
abstract
Temporal consistency is essential for video editing applications. Existing work on layered representation of videos allows propagating edits consistently to each frame. These methods, however, can only edit object appearance rather than object shape changes due to the limitation of using a fixed UV mapping field for texture atlas. We present a shape-aware, text-driven video editing method to tackle this challenge. To handle shape changes in video editing, we first propagate the deformation field between the input and edited keyframe to all frames. We then leverage a pre-trained text-conditioned diffusion model as guidance for refining shape distortion and completing unseen regions. The experimental results demonstrate that our method can achieve shape-aware consistent video editing and compare favorably with the state-of-the-art.
Yao-Chih Lee, Ji-Ze Genevieve Jang, Elizabeth Qiu, Jia-Bin Huang 0001
CVPR1
2021 3D Video Stabilization With Depth Estimation by CNN-Based Optimization
abstract
Video stabilization is an essential component of visual quality enhancement. Early methods rely on feature tracking to recover either 2D or 3D frame motion, which suffer from the robustness of local feature extraction and tracking in shaky videos. Recently, learning-based methods seek to find frame transformations with high-level information via deep neural networks to overcome the robustness issue of feature tracking. Nevertheless, to our best knowledge, no learning-based methods leverage 3D cues for the transformation inference yet; hence they would lead to artifacts on complex scene-depth scenarios. In this paper, we propose Deep3D Stabilizer, a novel 3D depth-based learning method for video stabilization. We take advantage of the recent self-supervised framework on jointly learning depth and camera ego-motion estimation on raw videos. Our approach requires no data for pre-training but stabilizes the input video via 3D reconstruction directly. The rectification stage incorporates the 3D scene depth and camera motion to smooth the camera trajectory and synthesize the stabilized video. Unlike most one-size-fits-all learning-based methods, our smoothing algorithm allows users to manipulate the stability of a video efficiently. Experimental results on challenging benchmarks show that the proposed solution consistently outperforms the state-of-the-art methods on almost all motion categories.
Yao-Chih Lee, Kuan-Wei Tseng, Yu-Ta Chen, Chien-Cheng Chen, Chu-Song Chen, Yi-Ping Hung
CVPR1
2021 PixStabNet: Fast Multi-Scale Deep Online Video Stabilization with Pixel-Based Warping
abstract
Online video stabilizaton is increasingly needed for real-time applications such as live streaming, drone remote control, and video communication. We propose a multi-scale convolutional neural network (PixStabNet) which stabilizes video in real time without using future frames. Instead of calculating a global homography or multiple homographies, we estimate a pixel-based warping map to make the transformation of each pixel to achieve more precise modelling. In addition, we propose well-designed loss functions along with a two-stage training scheme to enhance network robustness. The quantitative result shows that our method outperforms other learning-based online methods in terms of stability with excellent geometric and temporal consistency. Moreover, to the best of our knowledge, the proposed algorithm is the most efficient approach for video stabilization. The models and results are available at: https://yu-ta-chen.github.io/PixStabNet.
Yu-Ta Chen, Kuan-Wei Tseng, Yao-Chih Lee, Yi-Ping Hung
ICIP3