Yizhen Lao

dblp:199/1270 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-6284-1724ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021
YearPublicationVenuePosition
2025 IOVS4NeRF: Incremental Optimal View Selection for Large-Scale NeRFs
abstract
Large-scale Neural Radiance Fields (NeRF) reconstructions are typically hindered by the requirement for extensive image datasets and substantial computational resources. This paper introduces IOVS4NeRF, a framework that employs an uncertainty-guided incremental optimal view selection strategy adaptable to various NeRF implementations. Specifically, by leveraging a hybrid uncertainty model that combines rendering and positional uncertainties, the proposed method calculates the most informative view from among the candidates, thereby enabling incremental optimization of scene reconstruction. Our detailed experiments demonstrate that IOVS4NeRF achieves high-fidelity NeRF reconstruction with minimal computational resources, making it suitable for large-scale scene applications.
Jingpeng Xie, Shiyu Tan, Yuanlei Wang, Tianle Du, Yifei Xue, Yizhen Lao
ICASSP6
2025 Faster and Better 3D Splatting via Group Training
abstract
3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, demonstrating remarkable capability in high-fidelity scene reconstruction through its Gaussian primitive representations. However, the computational overhead induced by the massive number of primitives poses a significant bottleneck to training efficiency. To overcome this challenge, we propose Group Training, a simple yet effective strategy that organizes Gaussian primitives into manageable groups, optimizing training efficiency and improving rendering quality. This approach shows universal compatibility with existing 3DGS frameworks, including vanilla 3DGS and Mip-Splatting, consistently achieving accelerated training while maintaining superior synthesis quality. Extensive experiments reveal that our straightforward Group Training strategy achieves up to 30\% faster convergence and improved rendering quality across diverse scenarios. Project Website: https://chengbo-wang.github.io/3DGS-with-Group-Training/
Guozheng Ma, Yifei Xue, Yizhen Lao
ICCV4
2025 ACM Multimedia 2025 Grand Challenge report for Image-to-Video Generation Model Acceleration
abstract
Recently, MGTV organized the Image-to-Video Model Acceleration Challenge, calling for participants to propose optimization solutions for the Wan 2.1-14B model. The challenge emphasizes techniques such as quantization and GPU acceleration to improve the model's inference efficiency. As AIGC technology advances rapidly, video generation large models exhibit great potential in content creation, yet they face critical challenges of high computing power consumption, long inference time, and excessive VRAM usage during inference, which severely hinder content production efficiency. This challenge aims to explore approaches for efficient video generation under limited computing resources, requiring participants to reduce the model's computing power and VRAM demands while improving inference speed, all without compromising generation quality. To support participants' development and evaluation, the challenge provides a baseline framework and test dataset. For further details, please refer to the official challenge website (https://challenge.ai.mgtv.com/#/track/53).
Jie Yang 0073, Shien Song, Haoyuan Xie, Yifei Xue, Yizhen Lao
ACM Multimedia7
2024 RSL-BA: Rolling Shutter Line Bundle Adjustment
Yongcong Zhang, Bangyan Liao, Yifei Xue, Peidong Liu 0001, Yizhen Lao
ECCV (59)6
2024 ACM Multimedia 2024 Grand Challenge Report for Artificial Intelligence Generated Image Detection
abstract
The AI Generated Image Detection Challenge, organized by MGTV, invites participants to develop advanced algorithms capable of accurately distinguishing between real and AI-generated images. These images may be created using various cutting-edge techniques, including but not limited to GAN and Stable Diffusion algorithms. Participants are encouraged to utilize open-source datasets or develop their own datasets to train their algorithms. This challenge presents a unique opportunity to enhance the field of AI-generated image detection, particularly in improving the algorithm's generalization capabilities to identify unknown and emerging samples. For more details and resources, please visit our official website (https://challenge.ai.mgtv.com/#/track/24).
Shien Song, Jie Yang 0073, Yifei Xue, Yizhen Lao
ACM Multimedia6
2023 Revisiting Rolling Shutter Bundle Adjustment: Toward Accurate and Fast Solution
abstract
We propose an accurate and fast bundle adjustment (BA) solution that estimates the 6-DoF pose with an independent RS model of the camera and the geometry of the environment based on measurements from a rolling shutter (RS) camera. This tackles the challenges in the existing works, namely, relying on high frame rate video as input, restrictive assumptions on camera motion and poor efficiency. To this end, we first verify the positive influence of the image point normalization to RSBA. Then we present a novel visual residual covariance model to standardize the reprojection error during RSBA, which consequently improves the overall accuracy. Besides, we demonstrate the combination of Normalization and covariance standardization Weighting in RSBA (NW-RSBA) can avoid common planar degeneracy without the need to constrain the filming manner. Finally, we propose an acceleration strategy for NW-RSBA based on the sparsity of its Jacobian matrix and Schur complement. The extensive synthetic and real data experiments verify the effectiveness and efficiency of the proposed solution over the state-of-the-art works.
Bangyan Liao, Delin Qu, Yifei Xue, Huiqing Zhang, Yizhen Lao
CVPR5
2023 Towards Nonlinear-Motion-Aware and Occlusion-Robust Rolling Shutter Correction
abstract
This paper addresses the problem of rolling shutter correction in complex nonlinear and dynamic scenes with extreme occlusion. Existing methods suffer from two main drawbacks. Firstly, they face challenges in estimating the accurate correction field due to the uniform velocity assumption, leading to significant image correction errors under complex motion. Secondly, the drastic occlusion in dynamic scenes prevents current solutions from achieving better image quality because of the inherent difficulties in aligning and aggregating multiple frames. To tackle these challenges, we model the curvilinear trajectory of pixels analytically and propose a geometry-based Quadratic Rolling Shutter (QRS) motion solver, which precisely estimates the high-order correction field of individual pixels. Besides, to reconstruct high-quality occlusion frames in dynamic scenes, we present a 3D video architecture that effectively Aligns and Aggregates multi-frame context, namely, RSA2-Net. We evaluate our method across a broad range of cameras and video sequences, demonstrating its significant superiority. Specifically, our method surpasses the state-of-the-art by +4.98, +0.77, and +4.33 of PSNR on CarlaRS, Fastec-RS, and BS-RSC datasets, respectively. Code is available at https://github.com/DelinQu/qrsc.
Delin Qu, Yizhen Lao, Zhigang Wang 0002, Dong Wang 0028, Bin Zhao 0001, Xuelong Li 0001
ICCV2
2023 ACM Multimedia 2023 Grand Challenge Report: Invisible Video Watermark
abstract
MGTV recently organized a pioneering Invisible Video Watermark Challenge, inviting participants to create a framework capable of embedding invisible watermarks into videos and extracting them from watermarked content.
Shien Song, Jie Yang 0073, Yifei Xue, Yizhen Lao
ACM Multimedia7
2023 Fast Rolling Shutter Correction in the Wild
abstract
This paper addresses the problem of rolling shutter correction (RSC) in uncalibrated videos. Existing works remove rolling shutter (RS) distortion by explicitly computing the camera motion and depth as intermediate products, followed by motion compensation. In contrast, we first show that each distorted pixel can be implicitly rectified back to the corresponding global shutter (GS) projection by rescaling its optical flow. Such a point-wise RSC is feasible with both perspective and non-perspective cases without the pre-knowledge of the camera used. Besides, it allows a pixel-wise varying direct RS correction (DRSC) framework that handles locally varying distortion caused by various sources, such as camera motion, moving objects, and even highly varying depth scenes. More importantly, our approach is an efficient CPU-based solution that enables undistorting RS videos in real-time (40fps for 480p). We evaluate our approach across a broad range of cameras and video sequences, including fast motion, dynamic scenes, and non-perspective lenses, demonstrating the superiority of our proposed approach over state-of-the-art methods in both effectiveness and efficiency. We also evaluated the ability of the RSC results to serve for downstream 3D analysis, such as visual odometry and structure-from-motion, which verifies preference for the output of our algorithm over other existing RSC methods.
Delin Qu, Bangyan Liao, Huiqing Zhang, Omar Ait-Aider, Yizhen Lao
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Augmenting TV Shows via Uncalibrated Camera Small Motion Tracking in Dynamic Scene
abstract
To augment the TV show in post-production, we propose a novel solution to uncalibrated camera small motion tracking in a dynamic scene that simultaneously reconstructs the sparse 3D scene and computes camera poses and focal lengths of each frame. The critical elements of our approach are a robust image feature tracking strategy in dynamic scenes followed by automatic local-window frames slicing, local and global bundle adjustment optimization initialized by a homography-based uncalibrated relative rotation solver. The proposed method allows us to add the virtual objects (elements) into the reconstructed 3D scene, then composite them back into the original shot while perfectly matched perspective and appear seamless.
Yizhen Lao, Shien Song
ACM Multimedia1
2021 Solving Rolling Shutter 3D Vision Problems using Analogies with Non-rigidity
Yizhen Lao, Omar Ait-Aider, Adrien Bartoli
Int. J. Comput. Vis.1
2021 Rolling Shutter Homography and its Applications
abstract
In this article we study the adaptation of the concept of homography to Rolling Shutter (RS) images. This extension has never been clearly adressed despite the many roles played by the homography matrix in multi-view geometry. We first show that a direct point-to-point relationship on a RS pair can be expressed as a set of 3 to 8 atomic 3x3 matrices depending on the kinematic model used for the instantaneous-motion during image acquisition. We call this group of matrices the RS Homography. We then propose linear solvers for the computation of these matrices using point correspondences. Finally, we derive linear and closed form solutions for two famous problems in computer vision in the case of RS images: image stitching and plane-based relative pose computation. Extensive experiments with both synthetic and real data from public benchmarks show that the proposed methods outperform state-of-art techniques.
Yizhen Lao, Omar Ait-Aider
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 A Robust Method for Strong Rolling Shutter Effects Correction Using Lines With Automatic Feature Selection
abstract
We present a robust method which compensates RS distortions in a single image using a set of image curves, basing on the knowledge that they correspond to 3D straight lines. Unlike in existing work, no a priori knowledge about the line directions (e.g. Manhattan World assumption) is required. We first formulate a parametric equation for the projection of a 3D straight line viewed by a moving rolling shutter camera under a uniform motion model. Then we propose a method which efficiently estimates ego angular velocity separately from pose parameters, using at least 4 image curves. Moreover, we propose for the first time a RANSAC-like strategy to select image curves which really correspond to 3D straight lines and reject those corresponding to actual curves in 3D world. A comparative experimental study with both synthetic and real data from famous benchmarks shows that the proposed method outperforms all the existing techniques from the state-of-the-art.
Yizhen Lao, Omar Ait-Aider
CVPR1
2018 Rolling Shutter Pose and Ego-Motion Estimation Using Shape-from-Template
Yizhen Lao, Omar Ait-Aider, Adrien Bartoli
ECCV (2)1
2018 Robustified Structure from Motion with rolling-shutter camera using straightness constraint
Yizhen Lao, Omar Ait-Aider, Helder Araújo
Pattern Recognit. Lett.1