VLDB 2026 Research / reviewers in the wild / expert
Zhaoyang Lv
dblp:183/6449
· DBLP profile ↗
18ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-7788-9982ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin DatasetabstractWe introduce Digital Twin Catalog (DTC), a new large-scale photorealistic 3D object digital twin dataset. A digital twin of a 3D object is a highly detailed, virtually indistinguishable representation of a physical object, accurately capturing its shape, appearance, physical properties, and other attributes. Recent advances in neural-based 3D reconstruction and inverse rendering have significantly improved the quality of 3D object reconstruction. Despite these advancements, there remains a lack of a large-scale, digital twin quality real-world dataset and benchmark that can quantitatively assess and compare the performance of different reconstruction methods, as well as improve reconstruction quality through training or fine-tuning. Moreover, to democratize 3D digital twin creation, it is essential to integrate creation techniques with next-generation egocentric computing platforms, such as AR glasses. Currently, there is no dataset available to evaluate 3D object reconstruction using egocentric captured images. To address these gaps, the DTC dataset features 2,000 scanned digital twin-quality 3D objects, along with image sequences captured under different lighting conditions using DSLR cameras and egocentric AR glasses. This dataset establishes the first comprehensive real-world evaluation benchmark for 3D digital twin creation tasks, offering a robust foundation for comparing and improving existing reconstruction methods. The DTC dataset is already released at https://www.projectaria.com/datasets/dtc/ and we will also make the baseline evaluations open-source. Zhao Dong 0001, Ka Chen, Zhaoyang Lv, Hong-Xing Yu, Yufeng Zhu, Stephen Tian, Zhengqin Li, Geordie Moffatt, Sean Christofferson, James Fort, Xiaqing Pan, Mingfei Yan, Jiajun Wu 0001, Carl Yuheng Ren, Richard A. Newcombe |
CVPR | 3 |
| 2025 | LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance FieldsabstractWe present Large Inverse Rendering Model (LIRM), a transformer architecture that jointly reconstructs high-quality shape, materials, and radiance fields with view-dependent effects in less than a second. Our model builds upon the recent Large Reconstruction Models (LRMs) that achieve state-of-the-art sparse-view reconstruction quality. However, existing LRMs struggle to reconstruct unseen parts accurately and cannot recover glossy appearance or generate relightable 3D contents that can be consumed by standard Graphics engines. To address these limitations, we make three key technical contributions to build a more practical multi-view 3D reconstruction framework. First, we introduce an update model that allows us to progressively add more input views to improve our reconstruction. Second, we propose a hexa-plane neural SDF representation to better recover detailed textures, geometry and material parameters. Third, we develop a novel neural directional-embedding mechanism to handle view-dependent effects. Trained on a large-scale shape and material dataset with a tailored coarse-to-fine training scheme, our model achieves compelling results. It compares favorably to optimization-based dense-view inverse rendering methods in terms of geometry and relighting accuracy, while requiring only a fraction of the inference time. Zhengqin Li, Dilin Wang, Ka Chen, Zhaoyang Lv, Thu Nguyen-Phuoc, Milim Lee, Jia-Bin Huang 0001, Lei Xiao 0014, Yufeng Zhu, Carl S. Marshall, Yuheng Ren, Richard A. Newcombe, Zhao Dong 0001 |
CVPR | 4 |
| 2025 | DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular VideosabstractWe introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction has gained significant attention for its ability to rapidly create digital replicas of real-world environments. However, most existing models are limited to static scenes and fail to reconstruct the motion of moving objects. Developing a feed-forward model for dynamic scene reconstruction poses significant challenges, including the scarcity of training data and the need for appropriate 3D representations and training paradigms. To address these challenges, we introduce several key technical contributions: an enhanced large-scale synthetic dataset with ground-truth multi-view videos and dense 3D scene flow supervision; a per-pixel deformable 3D Gaussian representation that is easy to learn, supports high-quality dynamic view synthesis, and enables long-range 3D tracking; and a large transformer network that achieves real-time, generalizable dynamic scene reconstruction. Extensive qualitative and quantitative experiments demonstrate that DGS-LRM achieves dynamic scene reconstruction quality comparable to optimization-based methods, while significantly outperforming the state-of-the-art predictive dynamic reconstruction method on real-world examples. Its predicted physically grounded 3D deformation is accurate and can be readily adapted for long-range 3D tracking tasks, achieving performance on par with state-of-the-art monocular video 3D tracking methods. Chieh Hubert Lin, Zhaoyang Lv, Songyin Wu, Thu Nguyen-Phuoc, Hung-Yu Tseng, Julian Straub, Numair Khan, Lei Xiao 0014, Ming-Hsuan Yang 0001, Yuheng Ren, Richard A. Newcombe, Zhao Dong 0001, Zhengqin Li |
NeurIPS | 2 |
| 2025 | 4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular VideosabstractWe propose 4DGT, a 4D Gaussian-based Transformer model for dynamic scene reconstruction, trained entirely on real-world monocular posed videos. Using 4D Gaussian as an inductive bias, 4DGT unifies static and dynamic components, enabling the modeling of complex, time-varying environments with varying object lifespans. We proposed a novel density control strategy in training, which enables our 4DGT to handle longer space-time input. Our model processes 64 consecutive posed frames in a rolling-window fashion, predicting consistent 4D Gaussians in the scene. Unlike optimization-based methods, 4DGT performs purely feed-forward inference, reducing reconstruction time from hours to seconds and scaling effectively to long video sequences. Trained only on large-scale monocular posed video datasets, 4DGT can outperform prior Gaussian-based networks significantly in real-world videos and achieve on-par accuracy with optimization-based methods on cross-domain videos. Zhengqin Li, Zhao Dong 0001, Richard A. Newcombe, Zhaoyang Lv |
NeurIPS | 6 |
| 2024 | VideoLLM-online: Online Video Large Language Model for Streaming VideoabstractRecent Large Language Models (LLMs) have been en-hanced with vision capabilities, enabling them to compre-hend images, videos, and interleaved vision-language con-tent. However, the learning methods of these large multi-modal models (LMMs) typically treat videos as predeter-mined clips, rendering them less effective and efficient at handling streaming video inputs. In this paper, we pro-pose a novel Learning-In- Video-Stream (LIVE) framework, which enables temporally aligned, long-context, and real-time dialogue within a continuous video stream. Our LIVE framework comprises comprehensive approaches to achieve video streaming dialogue, encompassing: (1) a training ob-jective designed to perform language modeling for contin-uous streaming inputs, (2) a data generation scheme that converts offline temporal annotations into a streaming di-alogue format, and (3) an optimized inference pipeline to speed up interactive chat in real-world video streams. With our LIVE framework, we develop a simplified model called VideoLLM-online and demonstrate its significant advan-tages in processing streaming videos. For instance, our VideoLLM-online-7B model can operate at over 10 FPS on an A100 GPU for a 5-minute video clip from Ego4D narration. Moreover, VideoLLM-online also showcases state-of-the-art performance on public offline video bench-marks, such as recognition, captioning, and forecasting. The code, model, data, and demo have been made available at showlab.github. iolvideollm-online. Joya Chen, Zhaoyang Lv, Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, Zheng Shou 0001 |
CVPR | 2 |
| 2024 | EgoLifter: Open-World 3D Segmentation for Egocentric Perception
Qiao Gu, Zhaoyang Lv, Duncan P. Frost, Simon Green, Julian Straub, Chris Sweeney |
ECCV (43) | 2 |
| 2024 | LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video EditingabstractVideo creation has become increasingly popular, yet the expertise and effort required for editing often pose barriers to beginners. In this paper, we explore the integration of large language models (LLMs) into the video editing workflow to reduce these barriers. Our design vision is embodied in LAVE, a novel system that provides LLM-powered agent assistance and language-augmented editing features. LAVE automatically generates language descriptions for the user’s footage, serving as the foundation for enabling the LLM to process videos and assist in editing tasks. When the user provides editing objectives, the agent plans and executes relevant actions to fulfill them. Moreover, LAVE allows users to edit videos through either the agent or direct UI manipulation, providing flexibility and enabling manual refinement of agent actions. Our user study, which included eight participants ranging from novices to proficient editors, demonstrated LAVE’s effectiveness. The results also shed light on user perceptions of the proposed LLM-assisted editing paradigm and its impact on users’ creativity and sense of co-creation. Based on these findings, we propose design implications to inform the future development of agent-assisted content editing. Bryan Wang, Yuliang Li 0001, Zhaoyang Lv, Haijun Xia, Raj Sodhi |
IUI | 3 |
| 2022 | Neural 3D Video Synthesis from Multi-view VideoabstractWe propose a novel approach for 3D video synthesis that is able to represent multi-view video recordings of a dynamic real-world scene in a compact, yet expressive representation that enables high-quality view synthesis and motion interpolation. Our approach takes the high quality and compactness of static neural radiance fields in a new direction: to a model-free, dynamic setting. At the core of our approach is a novel time-conditioned neural radiance field that represents scene dynamics using a set of compact latent codes. We are able to significantly boost the training speed and perceptual quality of the generated imagery by a novel hierarchical training scheme in combination with ray importance sampling. Our learned representation is highly compact and able to represent a 10 second 30 FPS multi-view video recording by 18 cameras with a model size of only 28MB. We demonstrate that our method can render high-fidelity wide-angle novel views at over 1K resolution, even for complex and dynamic scenes. We perform an extensive qualitative and quantitative evaluation that shows that our approach outperforms the state of the art. Project website: https://neural-3d-video.github.io/. Tianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green, Christoph Lassner, Changil Kim 0001, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard A. Newcombe, Zhaoyang Lv |
CVPR | 11 |
| 2022 | AdaNeRF: Adaptive Sampling for Real-Time Rendering of Neural Radiance Fields
Andreas Kurz, Thomas Neff, Zhaoyang Lv, Michael Zollhöfer, Markus Steinberger |
ECCV (17) | 3 |
| 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural RenderingabstractWe present STaR, a novel method that performs Self-supervised Tracking and Reconstruction of dynamic scenes with rigid motion from multi-view RGB videos without any manual annotation. Recent work has shown that neural networks are surprisingly effective at the task of compressing many views of a scene into a learned function which maps from a viewing ray to an observed radiance value via volume rendering. Unfortunately, these methods lose all their predictive power once any object in the scene has moved. In this work, we explicitly model rigid motion of objects in the context of neural representations of radiance fields. We show that without any additional human specified supervision, we can reconstruct a dynamic scene with a single rigid object in motion by simultaneously decomposing it into its two constituent parts and encoding each with its own neural representation. We achieve this by jointly optimizing the parameters of two neural radiance fields and a set of rigid poses which align the two fields at each frame. On both synthetic and real world datasets, we demonstrate that our method can render photorealistic novel views, where novelty is measured on both spatial and temporal axes. Our factored representation furthermore enables animation of unseen object motion. Zhaoyang Lv, Tanner Schmidt, Steven Lovegrove |
CVPR | 2 |
| 2021 | ASIC Design Principle Course with Combination of Online-MOOC and Offline-Inexpensive FPGA BoardabstractASIC Design Principle (ASICDP) is a compulsory course for undergraduate majors in microelectronics and integrated circuits, and the focus of this paper is the teaching methods of online theoretical teaching and offline experimental teaching of this course. As is well known, in order to prevent and control COVID-19, the use of online platforms to carry out online teaching has attracted worldwide attention. In this paper, the teaching strategy "Online-MOOC + Offline Inexpensive FPGA Board" in ASICDP in the Spring 2020 semester is demonstrated, in where MOOC means Massive Open Online Course. The theoretical teaching content of ASICDP is entirely replicated from Hardware Acceleration Design Methodology (HADM) released by the present authors on "China University MOOC," the largest MOOC platform in China. Meanwhile, with the support of the "Xilinx & Ministry of Education University-Industry Collaborative Education Program," an FPGA development board called the "Spartan Edge Accelerator Board" (SEA Board) designed by the authors was used in the experimental teaching of the ASICDP. This method can be used to establish the linkage between online courses and offline experiments, and cultivate students' practical VLSI design and FPGA prototype verification skills. It is believed that for educators that want to improve courses related to ASIC design and FPGA prototype verification, Online-MOOC + Offline-Inexpensive FPGA Board is an effective method with lower cost that is easily promotable and replicated. Zhixiong Di, Yongming Tang, Jiahua Lu, Zhaoyang Lv |
ACM Great Lakes Symposium on VLSI | 4 |
| 2019 | Taking a Deeper Look at the Inverse Compositional AlgorithmabstractIn this paper, we provide a modern synthesis of the classic inverse compositional algorithm for dense image alignment. We first discuss the assumptions made by this well-established technique, and subsequently propose to relax these assumptions by incorporating data-driven priors into this model. More specifically, we unroll a robust version of the inverse compositional algorithm and replace multiple components of this algorithm using more expressive models whose parameters we train in an end-to-end fashion from data. Our experiments on several challenging 3D rigid motion estimation tasks demonstrate the advantages of combining optimization with learning-based techniques, outperforming the classic inverse compositional algorithm as well as data-driven image-to-pose regression approaches. Zhaoyang Lv, Frank Dellaert, James M. Rehg, Andreas Geiger 0001 |
CVPR | 1 |
| 2019 | SENSE: A Shared Encoder Network for Scene-Flow EstimationabstractWe introduce a compact network for holistic scene flow estimation, called SENSE, which shares common encoder features among four closely-related tasks: optical flow estimation, disparity estimation from stereo, occlusion estimation, and semantic segmentation. Our key insight is that sharing features makes the network more compact, induces better feature representations, and can better exploit interactions among these tasks to handle partially labeled data. With a shared encoder, we can flexibly add decoders for different tasks during training. This modular design leads to a compact and efficient model at inference time. Exploiting the interactions among these tasks allows us to introduce distillation and self-supervised losses in addition to supervised losses, which can better handle partially labeled real-world data. SENSE achieves state-of-the-art results on several optical flow benchmarks and runs as fast as networks specifically designed for optical flow. It also compares favorably against the state of the art on stereo and scene flow, while consuming much less memory. Huaizu Jiang, Deqing Sun, Varun Jampani, Zhaoyang Lv, Erik G. Learned-Miller, Jan Kautz |
ICCV | 4 |
| 2019 | Multi-class classification without multi-class labels
Yen-Chang Hsu, Zhaoyang Lv, Joel Schlosser, Phillip Odom, Zsolt Kira |
ICLR (Poster) | 2 |
| 2018 | Learning Rigidity in Dynamic Scenes with a Moving Camera for 3D Motion Field Estimation
Zhaoyang Lv, Alejandro J. Troccoli, Deqing Sun, James M. Rehg, Jan Kautz |
ECCV (5) | 1 |
| 2018 | Learning to cluster in order to transfer across domains and tasks
Yen-Chang Hsu, Zhaoyang Lv, Zsolt Kira |
ICLR (Poster) | 2 |
| 2016 | KEIPD: Knowledge Extraction and Inference System for Personal Documents
Zhaoyang Lv, Xiaohui Yu 0001 |
APWeb (2) | 1 |
| 2016 | A Continuous Optimization Approach for Efficient and Accurate Scene Flow
Zhaoyang Lv, Chris Beall, Pablo Fernández Alcantarilla, Fuxin Li, Zsolt Kira, Frank Dellaert |
ECCV (8) | 1 |